AI Image Model Benchmark Leaderboard
Benchmark V1 · Ranked by pass rate on 192 prompts
35 models ranked by Overall — capability (graded pass/fail by VLM judges across 192 prompts in six categories) blended 50/50 with aesthetic Estimated Preference. Every generated image is published so you can judge with your own eyes.
V1 Leaderboard
Click any model to see every image it generated, or open two models to compare them side-by-side.
| # | ||||||
|---|---|---|---|---|---|---|
| 1 | openai/gpt-image-2 | 78.5 | 83.9% | 73.0 | token-based | 45.3s |
| 2 | fal/google/nano-banana-2 | 73.0 | 73.4% | 72.6 | $0.08 / image at 1K | 28.1s |
| 3 | fal/google/nano-banana-pro | 66.2 | 73.4% | 59.1 | $0.15 / image | 23.4s |
| 4 | fal/bytedance/seedream-v5-pro | 65.6 | 70.3% | 60.9 | from $0.0675 / image | 147.4s |
| 5 | local/boogu-image-turbo | 61.9 | 57.3% | 66.6 | N/A | 9.9s |
| 6 | bfl/flux-2-pro | 60.1 | 63.0% | 57.3 | from $0.03 / image | 11.8s |
| 7 | bfl/flux-2-max | 59.6 | 62.5% | 56.6 | from $0.07 / image | 26.7s |
| 8 | fal/bytedance/seedream-v4 | 59.4 | 62.5% | 56.2 | $0.03 / image | 14.1s |
| 9 | local/qwen-image-2512-20b | 58.6 | 50.0% | 67.1 | N/A | 80.2s |
| 10 | local/flux-2-klein-9b | 56.6 | 53.1% | 60.1 | N/A | 8.5s |
| 11 | local/krea-2-turbo | 56.5 | 57.8% | 55.2 | N/A | 73.2s |
| 12 | bfl/flux-2-klein-9b | 55.3 | 53.6% | 57.1 | from $0.015 / image | 4.1s |
| 13 | local/krea-2-turbo-no-filter | 53.8 | 53.6% | 54.0 | N/A | 71.5s |
| 14 | fal/bytedance/seedream-v5-lite | 52.7 | 72.4% | 33.1 | $0.035 / image | 35.5s |
| 15 | local/flux-2-klein-4b | 51.5 | 48.4% | 54.6 | N/A | 4.5s |
| 16 | bfl/flux-2-klein-4b | 49.8 | 46.4% | 53.2 | from $0.014 / image | 3.8s |
| 17 | local/sefi-image-5b-base | 49.7 | 58.9% | 40.4 | N/A | 130.6s |
| 18 | local/z-image-turbo-6b | 49.2 | 48.4% | 49.9 | N/A | 18.1s |
| 19 | fal/krea/v2-medium-turbo | 48.7 | 56.8% | 40.5 | not found | 15.5s |
| 20 | fal/ideogram/v3 | 48.5 | 43.8% | 53.3 | $0.06 / image at BALANCED | 12.9s |
| 21 | fal/bria/fast | 47.9 | 51.0% | 44.8 | $0.028 / generation | 12.4s |
| 22 | local/sefi-image-5b-rl | 47.7 | 56.3% | 39.1 | N/A | 130.7s |
| 23 | local/z-image-6b | 47.4 | 53.1% | 41.7 | N/A | 130.7s |
| 24 | fal/krea/v2-medium | 46.9 | 58.9% | 34.9 | $0.030 / image | 18.5s |
| 25 | local/bonsai-image-ternary-4b | 46.5 | 43.2% | 49.9 | N/A | 4.1s |
| 26 | local/prxpixel-t2i-7b | 46.3 | 42.7% | 50.0 | N/A | 64.7s |
| 27 | fal/ideogram/v4 | 46.1 | 58.3% | 34.0 | $0.015 / MP at BALANCED on fal | 16.6s |
| 28 | fal/krea/v2-large | 45.9 | 59.9% | 31.9 | $0.060 / image | 30.1s |
| 29 | fal/recraft/v3 | 45.5 | 36.5% | 54.6 | $0.04 / image | 10.1s |
| 30 | local/hidream-i1-full-17b | 42.9 | 35.9% | 49.9 | N/A | 91.3s |
| 31 | local/krea-2-raw | 42.1 | 57.3% | 26.9 | N/A | 912.2s |
| 32 | local/sefi-image-5b-turbo | 41.9 | 52.6% | 31.2 | N/A | 5.9s |
| 33 | local/nucleus-image-17b-a2b | 36.6 | 41.7% | 31.5 | N/A | 39.1s |
| 34 | local/sefi-image-2b-turbo | 33.6 | 41.1% | 26.1 | N/A | 3.6s |
| 35 | local/sana-1.5-1.6b | 33.0 | 31.8% | 34.2 | N/A | 11.1s |
Local model analysis: Quality vs. size
Overall score against model size (billions of parameters) for all 18 models we ran locally on the NVIDIA DGX Spark. The green gradient deepens toward the sweet spot — smaller models that still score highly.
API model analysis: Quality vs. price
Overall score against API price per image for the 14 hosted models with a published flat per-image price. The green gradient deepens toward the sweet spot — cheaper models that still score highly.