ImageBench

ImageBench V1 —

192 evaluations across 6 categories

Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this ImageBench V1 run.

local/mage-flow-turbo

Local

Mage-Flow-Turbo is Microsoft Research Asia’s 4B MIT-licensed native-resolution text-to-image model — a 4-step distilled rectified-flow DiT paired with the lightweight high-fidelity Mage-VAE and a Qwen3-VL text encoder. It generates any aspect ratio from 512 to 2048 px and claims a GenEval lead over FLUX.2 while matching much larger models at a fraction of the size.

Maker
Microsoft
Family
Mage-Flow
Model Size
4B
high
Cost
local run; no API price
not_applicable
Run Target
gx10/mage-flow-turbo
Effective Request
response_format: b64_json · size: 1024x1024 · num_inference_steps: 4 · guidance_scale: 1 · seed: 7
47.3
Overall
49%
Capability
45.6
Est. Preference
94
Pass
98
Fail
4.6s
Avg Latency
4.0s
Min Latency
5.0s
Max Latency
Text Rendering67%Spatial Reasoning49%Human realism29%Truthfulness37%Professional Studio85%Graphical design46%Preference46%Latency55%

All 192 generations

Text Rendering67%

Typography Style67%

Writing accuracy67%

Spatial Reasoning49%

Attributes Binding56%

Compositionality67%

Counting22%

Negation33%

Relative Position67%

Scale & Proportions44%

Human realism29%

Faces & Expressions67%

Full Body17%

Hands0%

Multi-Subject33%

Truthfulness37%

Photorealism33%

Physics & Reflections58%

World Knowledge17%

Professional Studio85%

Camera & Lighting83%

Color Precision92%

Photorealism67%

Graphical design46%

Data Visualisation0%

Layout & Design11%

Style Diversity83%