ImageBench

ImageBench V1 —

192 evaluations across 6 categories

Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this ImageBench V1 run.

local/mage-flow-rl

Local

Mage-Flow (RL-aligned) — Microsoft Research Asia’s 4B MIT-licensed native-resolution text-to-image model (Mage-VAE + Qwen3-VL encoder). RL-aligned 20-step variant; RL alignment targets prompt adherence + aesthetics over the distilled turbo.

Maker
Microsoft
Family
Mage-Flow
Model Size
4B
high
Cost
local run; no API price
not_applicable
Run Target
gx10/mage-flow-rl
Effective Request
response_format: b64_json · size: 1024x1024 · num_inference_steps: 20 · guidance_scale: 5 · seed: 7
45.8
Overall
49%
Capability
42.7
Est. Preference
94
Pass
98
Fail
20.9s
Avg Latency
20.0s
Min Latency
22.0s
Max Latency
Text Rendering73%Spatial Reasoning44%Human realism33%Truthfulness26%Professional Studio85%Graphical design58%Preference43%Latency11%

All 192 generations

Text Rendering73%

Typography Style100%

Writing accuracy67%

Spatial Reasoning44%

Attributes Binding56%

Compositionality78%

Counting22%

Negation22%

Relative Position58%

Scale & Proportions22%

Human realism33%

Faces & Expressions58%

Full Body25%

Hands17%

Multi-Subject33%

Truthfulness26%

Photorealism33%

Physics & Reflections33%

World Knowledge17%

Professional Studio85%

Camera & Lighting75%

Color Precision92%

Photorealism100%

Graphical design58%

Data Visualisation0%

Layout & Design44%

Style Diversity83%