ImageBench

ImageBench V1 —

192 evaluations across 6 categories

Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this ImageBench V1 run.

local/sana-1.5-1.6b

Local

Sana 1.5 1.6B is an NVIDIA/Sana linear Diffusion Transformer model for efficient text-to-image generation, focused on scaling training and inference with comparatively small parameter counts.

Maker
NVIDIA / Sana team
Family
Sana 1.5
Model Size
1.6B
high
Cost
local run; no API price
not_applicable
Run Target
sana-local/sana-1.5-1.6b
Effective Request
size: 1024x1024
33.0
Overall
32%
Capability
34.2
Est. Preference
61
Pass
131
Fail
11.1s
Avg Latency
10.6s
Min Latency
11.5s
Max Latency
Text Rendering13%Spatial Reasoning33%Human realism7%Truthfulness30%Professional Studio67%Graphical design46%Preference34%Latency29%

All 192 generations

Text Rendering13%

Typography Style33%

Writing accuracy8%

Spatial Reasoning33%

Attributes Binding33%

Compositionality67%

Counting22%

Negation22%

Relative Position25%

Scale & Proportions33%

Human realism7%

Faces & Expressions17%

Full Body0%

Hands0%

Multi-Subject17%

Truthfulness30%

Photorealism0%

Physics & Reflections42%

World Knowledge25%

Professional Studio67%

Camera & Lighting67%

Color Precision75%

Photorealism33%

Graphical design46%

Data Visualisation0%

Layout & Design0%

Style Diversity92%