ImageBench

ImageBench V1 —

192 evaluations across 6 categories

Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this ImageBench V1 run.

local/qwen-image-2512-20b

Local

Qwen-Image-2512 is the December update of Qwen's text-to-image foundation model, improving human realism, natural detail, and text rendering over the earlier Qwen-Image release.

Maker
Qwen / Alibaba
Family
Qwen-Image
Model Size
20B
high
Cost
local run; no API price
not_applicable
Run Target
qwen-image-local/qwen-image-gen
Effective Request
size: 1024x1024 · response_format: b64_json
58.6
Overall
50%
Capability
67.1
Est. Preference
96
Pass
96
Fail
80.2s
Avg Latency
79.7s
Min Latency
82.0s
Max Latency
Text Rendering53%Spatial Reasoning54%Human realism38%Truthfulness48%Professional Studio82%Graphical design25%Preference67%Latency0%

All 192 generations

Text Rendering53%

Typography Style100%

Writing accuracy42%

Spatial Reasoning54%

Attributes Binding67%

Compositionality78%

Counting67%

Negation33%

Relative Position42%

Scale & Proportions44%

Human realism38%

Faces & Expressions75%

Full Body17%

Hands25%

Multi-Subject33%

Truthfulness48%

Photorealism67%

Physics & Reflections50%

World Knowledge42%

Professional Studio81%

Camera & Lighting75%

Color Precision83%

Photorealism100%

Graphical design25%

Data Visualisation0%

Layout & Design22%

Style Diversity33%