ImageBench

ImageBench V1 —

192 evaluations across 6 categories

Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this ImageBench V1 run.

local/tinydit-256

Local

TinyDiT-256 is a 210M-parameter text-to-image DiT trained from scratch on a single GPU in 3.5 days (400k steps, 4.2M images at 256px). It uses rectified flow with a frozen FLUX.2 autoencoder and flan-t5-base text encoder, and renders only small trained shapes such as 256x256. Weights are CC BY-NC-4.0 (non-commercial).

Maker
Ivan Mikhnenkov (independent)
Family
TinyDiT
Model Size
210M
high
Cost
local run; no API price
not_applicable
Run Target
gx10/tinydit-256
Effective Request
response_format: b64_json · size: 256x256 · num_inference_steps: 20 · guidance_scale: 4 · seed: 7
21.5
Overall
22%
Capability
21.1
Est. Preference
42
Pass
150
Fail
0.5s
Avg Latency
0.5s
Min Latency
0.7s
Max Latency
Text Rendering0%Spatial Reasoning25%Human realism12%Truthfulness19%Professional Studio56%Graphical design13%Preference21%Latency100%

All 192 generations

Text Rendering0%

Typography Style0%

Writing accuracy0%

Spatial Reasoning25%

Attributes Binding22%

Compositionality44%

Counting33%

Negation22%

Relative Position17%

Scale & Proportions11%

Human realism12%

Faces & Expressions0%

Full Body0%

Hands42%

Multi-Subject0%

Truthfulness19%

Photorealism0%

Physics & Reflections42%

World Knowledge0%

Professional Studio56%

Camera & Lighting42%

Color Precision83%

Photorealism0%

Graphical design13%

Data Visualisation0%

Layout & Design0%

Style Diversity25%