ImageBench

ImageBench V1 —

192 evaluations across 6 categories

Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this ImageBench V1 run.

local/flux-2-klein-4b

Local

The local FLUX.2 [klein] 4B run uses Black Forest Labs' compact FLUX.2 open-weight model, selected for maximum speed and deployment on modest local hardware.

Maker
Black Forest Labs
Family
FLUX.2 [klein]
Model Size
4B
high
Cost
local run; no API price
not_applicable
Run Target
gx10/flux2-klein-4b
Effective Request
Effective request fields unknown
51.5
Overall
48%
Capability
54.6
Est. Preference
93
Pass
99
Fail
4.5s
Avg Latency
4.2s
Min Latency
5.2s
Max Latency
Text Rendering40%Spatial Reasoning49%Human realism36%Truthfulness30%Professional Studio89%Graphical design50%Preference55%Latency56%

All 192 generations

Text Rendering40%

Typography Style100%

Writing accuracy25%

Spatial Reasoning49%

Attributes Binding78%

Compositionality67%

Counting33%

Negation11%

Relative Position50%

Scale & Proportions56%

Human realism36%

Faces & Expressions67%

Full Body25%

Hands8%

Multi-Subject50%

Truthfulness30%

Photorealism0%

Physics & Reflections50%

World Knowledge17%

Professional Studio89%

Camera & Lighting92%

Color Precision92%

Photorealism67%

Graphical design50%

Data Visualisation0%

Layout & Design22%

Style Diversity83%