ImageBench

ImageBench V1 —

192 evaluations across 6 categories

Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this ImageBench V1 run.

local/mgflow-flux2-klein-4b

Local

MGFlow is an open-weight one-step text-to-image model post-trained from FLUX.2 [klein] 4B. It generates a 512x512 image in a single network evaluation with no classifier-free guidance, trained by matching real and generated features in frozen representation spaces with Gaussian-mixture statistics (MGFlow-KL). This entry is the joint image-text checkpoint, submitted by the authors for evaluation. Apache-2.0 checkpoints, MIT code.

Maker
Shi Haoyang et al. (MGFlow authors)
Family
MGFlow
Model Size
4B
high
Cost
local run; no API price
not_applicable
Run Target
gx10/mgflow-flux2-klein-4b
Effective Request
response_format: b64_json · size: 512x512 · num_inference_steps: 1 · guidance_scale: 0 · seed: 7
52.8
Overall
43%
Capability
62.4
Est. Preference
83
Pass
109
Fail
0.8s
Avg Latency
0.7s
Min Latency
0.9s
Max Latency
Text Rendering33%Spatial Reasoning47%Human realism29%Truthfulness15%Professional Studio85%Graphical design50%Preference62%Latency100%

All 192 generations

Text Rendering33%

Typography Style67%

Writing accuracy25%

Spatial Reasoning47%

Attributes Binding67%

Compositionality56%

Counting33%

Negation44%

Relative Position50%

Scale & Proportions33%

Human realism29%

Faces & Expressions42%

Full Body8%

Hands50%

Multi-Subject0%

Truthfulness15%

Photorealism0%

Physics & Reflections25%

World Knowledge8%

Professional Studio85%

Camera & Lighting75%

Color Precision100%

Photorealism67%

Graphical design50%

Data Visualisation0%

Layout & Design11%

Style Diversity92%