ImageBench

Blog

Analysis, methodology, and findings from the ImageBench project.

016 min
Does the score match your eye?
A blind test of our aesthetic metric: look at the images, pick your favourite, and see how often you disagree with the score — a critique of what a single preference number can tell you.
027 min
V1.1: Grading the Benchmark with Gemini 3.1 Pro
ImageBench V1.1 swaps the local open-VLM judges for Gemini 3.1 Pro and specialized per-prompt questions — a stricter, more consistent grader. Capability drops ~22 points; the ranking holds.
038 min
Estimated Preference Score
A second axis for ImageBench V1 — an aesthetic human-preference score from the HPSv3 reward model, combined with capability into a new Overall score.
047 min
Tuning Krea-2 Turbo for Better Prompt Following
A controlled strength sweep of one conditioning weight in Krea-2 Turbo — prompt adherence rises to a peak around strength 1, then degrades.
059 min
Why a Few Good Prompts Are Enough
The theory behind ImageBench — Item Response Theory and tinyBenchmarks explain why a small, discriminative prompt set can rank models almost as well as a full benchmark.
066 min
Can Local VLMs Recognize Celebrities?
A 90-image public-figure recognition study across global icons, field-famous specialists, and long-tail Wikipedia notables.
075 min
Which VLM Should Judge Style Diversity?
A VLM calibration study for ImageBench Style Diversity, selecting Qwen 3.5 122B as the route judge.
083 min
RealBench V1 Methodology
How RealBench measures photorealism — paired real/AI images and human votes from the ImageBench community.
094 min
Which VLM Best Detects Bad Hands?
A detector calibration study using Nano Banana Pro and Bonsai Image 4B hand outputs.
105 min
GPT Image 2 vs Flux 2 vs Nano Banana 2
Side-by-side benchmark results for leading AI image generation models.
115 min
Best AI Image Generator for Text Rendering
Which models handle spelling, posters, labels, typography, and small text best.
125 min
ImageBench V1 Methodology
How ImageBench V1 is designed, scored, and reported across 64 tests and 6 capability categories.
137 min
Quality Metrics Across 10 Models
MUSIQ, NIQE, NIMA, and TOPIQ scores for all 10 ImageBench V1 models — and why the results are still preliminary.
145 min
Glossary
A–Z reference of every term used across the guides.
1510 min
Safety & Bias
NSFW content, demographic bias, IP concerns, red teaming, and the EU AI Act.
167 min
Consistency & Reproducibility
Same prompt, different outputs — measuring variance and why it matters for production.
1711 min
Prompt Fidelity & Compositionality
Does the image match the text? Measuring attribute binding, spatial reasoning, and counting.
189 min
Comparing Image Models
Quality, speed, cost, consistency — how to compare fairly and find the Pareto frontier.
1910 min
Human Evaluation
ELO rankings, pairwise preference, rater calibration, and LLM-as-a-Judge approaches.
2012 min
Automated Metrics
FID, CLIP Score, LPIPS, VQAScore — what they measure, when to use them, and common pitfalls.
218 min
Introduction to Image Evaluation
Why evaluating image generation matters, and the two sides: automated metrics vs human judgment.