ImageBench

vs

192 evaluations across 6 categories

Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this ImageBench V1 run.

local/qwen-image-2512-20b

Local

Qwen-Image-2512 is the December update of Qwen's text-to-image foundation model, improving human realism, natural detail, and text rendering over the earlier Qwen-Image release.

Maker
Qwen / Alibaba
Family
Qwen-Image
Model Size
20B
high
Cost
local run; no API price
not_applicable
Run Target
qwen-image-local/qwen-image-gen
Effective Request
size: 1024x1024 · response_format: b64_json
49.8vs58.6
Overall
46%vs50%
Capability
53.2vs67.1
Est. Preference
3.8svs80.2s
Avg Latency
Text Rendering40%53%Spatial Reasoning54%54%Human realism24%38%Truthfulness26%48%Professional Studio89%82%Graphical design46%25%Preference53%67%Latency61%0%

All 192 generations

Text Rendering40%vs53%

bfl/flux-2-klein-4blocal/qwen-image-2512-20b

Typography Style100%vs100%

Writing accuracy25%vs42%

Spatial Reasoning54%vs54%

bfl/flux-2-klein-4blocal/qwen-image-2512-20b

Attributes Binding67%vs67%

Compositionality67%vs78%

Counting33%vs67%

Negation56%vs33%

Relative Position50%vs42%

Scale & Proportions56%vs44%

Human realism24%vs38%

bfl/flux-2-klein-4blocal/qwen-image-2512-20b

Faces & Expressions50%vs75%

Full Body8%vs17%

Hands0%vs25%

Multi-Subject50%vs33%

Truthfulness26%vs48%

bfl/flux-2-klein-4blocal/qwen-image-2512-20b

Photorealism0%vs67%

Physics & Reflections42%vs50%

World Knowledge17%vs42%

Professional Studio89%vs81%

bfl/flux-2-klein-4blocal/qwen-image-2512-20b

Camera & Lighting83%vs75%

Color Precision100%vs83%

Photorealism67%vs100%

Graphical design46%vs25%

bfl/flux-2-klein-4blocal/qwen-image-2512-20b

Data Visualisation0%vs0%

Layout & Design11%vs22%

Style Diversity83%vs33%