ImageBench

ImageBench V1 —

192 evaluations across 6 categories

Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this ImageBench V1 run.

local/tld-101m

Local

Transformer Latent Diffusion (TLD) is a minimal 101M-parameter latent diffusion transformer by apapiu (MIT), trained for about 32 hours on one A100. It conditions on a pooled CLIP ViT-L/14 embedding, samples with DPM-Solver++(2M) and outputs 256x256 images; it is a small educational/research model.

Maker
apapiu (independent)
Family
Transformer Latent Diffusion
Model Size
101M
high
Cost
local run; no API price
not_applicable
Run Target
gx10/tld-101m
Effective Request
response_format: b64_json · size: 256x256 · num_inference_steps: 15 · guidance_scale: 6 · seed: 7
8.4
Overall
13%
Capability
3.8
Est. Preference
25
Pass
167
Fail
0.2s
Avg Latency
0.2s
Min Latency
1.1s
Max Latency
Text Rendering0%Spatial Reasoning14%Human realism0%Truthfulness4%Professional Studio37%Graphical design25%Preference4%Latency100%

All 192 generations

Text Rendering0%

Typography Style0%

Writing accuracy0%

Spatial Reasoning14%

Attributes Binding0%

Compositionality11%

Counting0%

Negation44%

Relative Position0%

Scale & Proportions33%

Human realism0%

Faces & Expressions0%

Full Body0%

Hands0%

Multi-Subject0%

Truthfulness4%

Photorealism0%

Physics & Reflections8%

World Knowledge0%

Professional Studio37%

Camera & Lighting50%

Color Precision33%

Photorealism0%

Graphical design25%

Data Visualisation0%

Layout & Design0%

Style Diversity50%