The only generative image benchmark that shows the images
52 models, 192 prompts, 6 categories — every output published. Judge with your own eyes which model is best for your use case, your budget, your quality bar.

Prompt: The word 'CHAPTER ONE' typed on aged paper with a vintage typewriter font, complete with slightly uneven ink
Benchmark V1
Capability benchmark — prompt adherence, text, hands, truthfulness and more, graded pass/fail with every image published.
ExploreRealBench V1
Realism benchmark — can a model's output pass as a real photograph? Scored by real human votes.
ExploreEditBench
Image-editing arena — authenticated pairwise votes between OpenAI and Nano Banana on real edit prompts.
Try itGallery
Browse generated images across models and prompts, side by side.
BrowseThe results at a glance
A preview of the top 10 across both benchmarks, plus how quality trades off against model size and price. Follow any link for the full, filterable results.
Benchmark V1
Ranked by Overall — capability (graded pass/fail across 192 prompts) blended with aesthetic preference.
| # | ||||
|---|---|---|---|---|
| 1 | openai/gpt-image-2 | 80.0 | token-based | 45.3s |
| 2 | fal/google/nano-banana-2 | 74.9 | $0.08 / image at 1K | 28.1s |
| 3 | fal/reve/2.1 | 74.8 | $0.04 / image | 45.0s |
| 4 | xai/grok-imagine-2 | 73.6 | $0.04 / image | 69.0s |
| 5 | fal/google/nano-banana-pro | 68.8 | $0.15 / image | 23.4s |
| 6 | fal/bytedance/seedream-v5-pro | 68.0 | from $0.0675 / image | 147.4s |
| 7 | fal/microsoft/mai-image-2.5 | 68.0 | ~$0.05 / image | 25.4s |
| 8 | fal/qwen-image-2-pro | 66.9 | $0.075 / image | 12.8s |
| 9 | local/boogu-image-turbo | 64.3 | N/A | 9.9s |
| 10 | bfl/flux-2-pro | 62.2 | from $0.03 / image | 11.8s |
Quality vs. size
Overall score against model size for the models we run locally — deeper green is the sweet spot.
See it full-sizeQuality vs. price
Overall score against API price per image for hosted models — deeper green is the sweet spot.
See it full-sizeRealBench V1
Realism leaderboard — can a model's output pass as a real photograph? Scored by human votes.
| # | Model | Realism score | Rated real / votes | Images |
|---|---|---|---|---|
| 1 | fal/google/nano-banana-pro | 53% | 1,907 / 3,609 | 141 |
| 2 | fal/bytedance/seedream-v4 | 47% | 1,679 / 3,539 | 139 |
| 3 | openai/gpt-image-2 | 46% | 1,572 / 3,418 | 139 |
| 4 | local/z-image-turbo-6b | 46% | 1,646 / 3,603 | 140 |
| 5 | bfl/flux-2-max | 41% | 1,473 / 3,592 | 141 |
Frequently asked questions
ImageBench is an AI image model benchmark that publishes every generated image, not just aggregate scores. Two evaluations: Benchmark V1 (capabilities across 6 categories, 52 models) and RealBench (photorealism via human votes). Every image is visible in the gallery.
Look at real outputs, not just scores. The ImageBench gallery shows the same 192 prompts generated by every model on the benchmark, side by side — filter by category or difficulty, or pick just the models you care about. Scores tell you who wins; the gallery shows you why.
On ImageBench V1, GPT Image 2 (OpenAI) leads with an overall score of 80.0 and an 87% task pass rate, ahead of Nano Banana 2 (Google, 74.9) and Reve 2.1 (74.8). Full rankings and per-category scores live on the leaderboard, and you can inspect the top three side by side in the gallery.
Four models score 100% on the Text Rendering category of Benchmark V1: Nano Banana Pro, GPT Image 2, Reve 2.1, and Luma UNI-1 Max. See the ranking and outputs on our best AI image generator for text page, or compare their text samples directly in the gallery.
ImageBench benchmarks open-weight models under the local/ prefix — Flux 2 Klein, Sana 1.5, HiDream, Z-Image, Qwen-Image, Krea 2, Boogu, and more. Boogu Image Turbo currently leads local models with a 64.3 overall score. Browse them on the leaderboard or compare the top local models in the gallery.
Boogu Image Turbo is the strongest open-weight model on ImageBench V1, scoring 64.3 overall (62% pass rate) — ahead of Qwen Image 2512 (59.6) and Krea 2 Turbo (59.1). The best open models still trail the top APIs (GPT Image 2 scores 80.0), but the gap keeps shrinking. See how close they get in the gallery.
Yes — sampler settings matter more than people assume. We swept a single conditioning-strength weight in Krea-2 Turbo while holding the prompt, seed, and steps fixed: prompt adherence rises to a peak around strength 1, then degrades. Read the full sweep in Tuning Krea-2 Turbo for Better Prompt Following.
Two different questions. Benchmark V1 asks “how good is it?” — pass/fail on capability tasks graded by VLM judges. RealBench asks “how real does it look?” — how often humans mistake AI images for real photos.
192 prompts across 6 categories: Text Rendering, Spatial Reasoning, Human Realism, Truthfulness, Professional Studio, Graphical Design. VLM judges (Qwen 3.5 122B and per-category specialists) grade pass/fail, verified by human review. Every prompt, image, and verdict is published.
Hands are the hardest thing image models do — and the hardest thing judges grade. Gemini over-flags them, failing good hands. Since Benchmark V1.2, hand prompts are graded by Qwen3.5-122B, the VLM that best matches human labels on bad-hand detection (F1 80% vs Gemini's 65%). Read why in V1.2: A Better Judge for Hands.
RealBench uses votes from a quick in-site game where players guess whether each image is a real photo or AI. A model's realism score is the share of votes that judged its AI images as real. Play the game or read the methodology.
V1 doesn't use FID or CLIP Score. Verdicts come from VLM judges — Qwen 3.5 122B as primary judge with per-category specialists — plus human review. See the methodology and calibration blog posts.
Arena platforms measure crowd preference via blind votes — great for overall ranking, but they don't tell you why or where models fail. ImageBench runs structured category benchmarks (text, spatial, hands, truthfulness, studio, design) and publishes every image and verdict.
Yes. Open-weight models like Flux 2, Sana, HiDream, Z-Image, Qwen-Image, and Boogu are evaluated under the same conditions as API-based models. Local models are labeled `local/` in the leaderboard.
Whenever a major model is released or updated. The leaderboard reflects the latest available versions; historical results are preserved so you can track progress over time.
Yes. If your model has a public API or downloadable weights, reach out and we'll add it to upcoming benchmark runs. All models are evaluated under the same conditions.
Yes. Every benchmark publishes its prompt suite, scoring code, and evaluation criteria. Transparent methodology is what separates useful benchmarks from marketing.
ImageBench.ai is built by Damien Henry — former co-founder of Clipdrop (YC W21, acquired by Stability AI), Google Arts & Culture innovation lead, and current SVP Image Research at Jasper.
Email hello@imagebench.ai