ImageBench

ImageBench V1 —

192 evaluations across 6 categories

Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this ImageBench V1 run.

local/mage-flow-base

Local

Mage-Flow (base) — Microsoft Research Asia’s 4B MIT-licensed native-resolution text-to-image model (Mage-VAE + Qwen3-VL encoder). Un-distilled, non-RL 30-step baseline — the quality ceiling of the 4B architecture.

Maker
Microsoft
Family
Mage-Flow
Model Size
4B
high
Cost
local run; no API price
not_applicable
Run Target
gx10/mage-flow-base
Effective Request
response_format: b64_json · size: 1024x1024 · num_inference_steps: 30 · guidance_scale: 5 · seed: 7
43.9
Overall
46%
Capability
42.0
Est. Preference
88
Pass
104
Fail
30.4s
Avg Latency
29.0s
Min Latency
32.0s
Max Latency
Text Rendering60%Spatial Reasoning42%Human realism24%Truthfulness37%Professional Studio89%Graphical design46%Preference42%Latency0%

All 192 generations

Text Rendering60%

Typography Style67%

Writing accuracy58%

Spatial Reasoning42%

Attributes Binding56%

Compositionality56%

Counting33%

Negation11%

Relative Position58%

Scale & Proportions33%

Human realism24%

Faces & Expressions58%

Full Body0%

Hands0%

Multi-Subject50%

Truthfulness37%

Photorealism67%

Physics & Reflections50%

World Knowledge17%

Professional Studio89%

Camera & Lighting83%

Color Precision92%

Photorealism100%

Graphical design46%

Data Visualisation0%

Layout & Design11%

Style Diversity83%