Blog
Analysis, methodology, and findings from the ImageBench project.
016 min
V1.2: A Better Judge for Hands
Gemini over-flags hands — it fails good ones. V1.2 routes just the Hands tests to Qwen3.5-122B, the judge that best matches human labels (F1 80% vs Gemini's 65%). Human-realism scores rise ~19 points; the ranking holds.
026 min
Does the score match your eye?
A blind test of our aesthetic metric: look at the images, pick your favourite, and see how often you disagree with the score — a critique of what a single preference number can tell you.
037 min
V1.1: Grading the Benchmark with Gemini 3.1 Pro
ImageBench V1.1 swaps the local open-VLM judges for Gemini 3.1 Pro and specialized per-prompt questions — a stricter, more consistent grader. Capability drops ~22 points; the ranking holds.
048 min
Estimated Preference Score
A second axis for ImageBench V1 — an aesthetic human-preference score from the HPSv3 reward model, combined with capability into a new Overall score.
057 min
Tuning Krea-2 Turbo for Better Prompt Following
A controlled strength sweep of one conditioning weight in Krea-2 Turbo — prompt adherence rises to a peak around strength 1, then degrades.
069 min
Why a Few Good Prompts Are Enough
The theory behind ImageBench — Item Response Theory and tinyBenchmarks explain why a small, discriminative prompt set can rank models almost as well as a full benchmark.
076 min
Can Local VLMs Recognize Celebrities?
A 90-image public-figure recognition study across global icons, field-famous specialists, and long-tail Wikipedia notables.
085 min
Which VLM Should Judge Style Diversity?
A VLM calibration study for ImageBench Style Diversity, selecting Qwen 3.5 122B as the route judge.
093 min
RealBench V1 Methodology
How RealBench measures photorealism — paired real/AI images and human votes from the ImageBench community.
104 min
Which VLM Best Detects Bad Hands?
A detector calibration study using Nano Banana Pro and Bonsai Image 4B hand outputs.
115 min
GPT Image 2 vs Flux 2 vs Nano Banana 2
Side-by-side benchmark results for leading AI image generation models.
125 min
Best AI Image Generator for Text Rendering
Which models handle spelling, posters, labels, typography, and small text best.
135 min
ImageBench V1 Methodology
How ImageBench V1 is designed, scored, and reported across 64 tests and 6 capability categories.
147 min
Quality Metrics Across 10 Models
MUSIQ, NIQE, NIMA, and TOPIQ scores for all 10 ImageBench V1 models — and why the results are still preliminary.
155 min
Glossary
A–Z reference of every term used across the guides.
1610 min
Safety & Bias
NSFW content, demographic bias, IP concerns, red teaming, and the EU AI Act.
177 min
Consistency & Reproducibility
Why image models produce different outputs from identical prompts. Seed control, deterministic generation, and measuring consistency at production scale.
1811 min
Prompt Fidelity & Compositionality
Does the image match the text? Measuring attribute binding, spatial reasoning, and counting.
199 min
Comparing Image Models
Quality, speed, cost, consistency — how to compare fairly and find the Pareto frontier.
2010 min
Human Evaluation
ELO rankings, pairwise preference, rater calibration, and LLM-as-a-Judge approaches.
2112 min
Automated Metrics
FID, CLIP Score, LPIPS, VQAScore — what they measure, when to use them, and common pitfalls.
228 min
Introduction to Image Evaluation
Text-to-image models lack a standard benchmark, so model choices get made on vibes. Automated metrics vs human evaluation — the tradeoffs, and when each applies.