AI Image Model Benchmark Leaderboard
Benchmark V1.2 · Ranked by pass rate on 192 prompts
50 models ranked by Overall — capability (graded pass/fail by VLM judges across 192 prompts in six categories) blended 50/50 with aesthetic Estimated Preference. Every generated image is published so you can judge with your own eyes.
V1.2 Leaderboard
Click any model to see every image it generated, or open two models to compare them side-by-side.
Don't take this ranking as the final word. Overall blends an automated VLM judge with an aesthetic preference model — both can be wrong, and neither replaces your own eyes. Every image every model generated is published, so look at them and form your own opinion. Browse the gallery.
| # | ||||||
|---|---|---|---|---|---|---|
| 1 | openai/gpt-image-2 | 80.0 | 87.0% | 73.0 | token-based | 45.3s |
| 2 | fal/google/nano-banana-2 | 74.9 | 77.1% | 72.6 | $0.08 / image at 1K | 28.1s |
| 3 | fal/reve/2.1 | 74.8 | 79.7% | 69.9 | /bin/bash.04 / image | 45.0s |
| 4 | fal/google/nano-banana-pro | 68.8 | 78.6% | 59.1 | $0.15 / image | 23.4s |
| 5 | fal/bytedance/seedream-v5-pro | 68.0 | 75.0% | 60.9 | from $0.0675 / image | 147.4s |
| 6 | fal/microsoft/mai-image-2.5 | 68.0 | 69.3% | 66.8 | ~$0.05 / image | 25.4s |
| 7 | local/boogu-image-turbo | 64.3 | 62.0% | 66.6 | N/A | 9.9s |
| 8 | bfl/flux-2-pro | 62.2 | 67.2% | 57.3 | from $0.03 / image | 11.8s |
| 9 | fal/bytedance/seedream-v4 | 62.2 | 68.2% | 56.2 | $0.03 / image | 14.1s |
| 10 | replicate/google/imagen-4-ultra | 62.0 | 62.0% | 61.9 | $0.06 / image | 14.6s |
| 11 | fal/luma/uni-1-max | 62.0 | 78.6% | 45.5 | ~$0.102 / image | 88.2s |
| 12 | bfl/flux-2-max | 61.9 | 67.2% | 56.6 | from $0.07 / image | 26.7s |
| 13 | local/qwen-image-2512-20b | 59.6 | 52.1% | 67.1 | N/A | 80.2s |
| 14 | local/krea-2-turbo | 59.1 | 63.0% | 55.2 | N/A | 73.2s |
| 15 | local/flux-2-klein-9b | 59.0 | 57.8% | 60.1 | N/A | 8.5s |
| 16 | fal/hunyuan-image-3 | 57.7 | 52.1% | 63.2 | ~$0.09 / image | 24.4s |
| 17 | bfl/flux-2-klein-9b | 57.2 | 57.3% | 57.1 | from $0.015 / image | 4.1s |
| 18 | local/krea-2-turbo-no-filter | 56.2 | 58.3% | 54.0 | N/A | 71.5s |
| 19 | replicate/recraft-ai/recraft-v4 | 55.5 | 62.5% | 48.5 | $0.04 / image | 9.5s |
| 20 | fal/bytedance/seedream-v5-lite | 54.8 | 76.6% | 33.1 | $0.035 / image | 35.5s |
| 21 | local/flux-2-klein-4b | 53.6 | 52.6% | 54.6 | N/A | 4.5s |
| 22 | replicate/google/imagen-4-fast | 53.0 | 51.6% | 54.5 | $0.02 / image | 5.8s |
| 23 | bfl/flux-2-klein-4b | 51.8 | 50.5% | 53.2 | from $0.014 / image | 3.8s |
| 24 | replicate/minimax/image-01 | 51.8 | 50.0% | 53.6 | $0.01 / image | 22.4s |
| 25 | local/sefi-image-5b-base | 51.5 | 62.5% | 40.4 | N/A | 130.6s |
| 26 | local/z-image-turbo-6b | 51.3 | 52.6% | 49.9 | N/A | 18.1s |
| 27 | fal/ideogram/v3 | 50.8 | 48.4% | 53.3 | $0.06 / image at BALANCED | 12.9s |
| 28 | fal/krea/v2-medium-turbo | 50.7 | 60.9% | 40.5 | not found | 15.5s |
| 29 | fal/bria/fast | 50.5 | 56.2% | 44.8 | $0.028 / generation | 12.4s |
| 30 | local/sefi-image-5b-rl | 50.0 | 60.9% | 39.1 | N/A | 130.7s |
| 31 | local/mage-flow-turbo | 49.9 | 54.2% | 45.6 | N/A | 4.6s |
| 32 | local/z-image-6b | 49.3 | 56.8% | 41.7 | N/A | 130.7s |
| 33 | replicate/leonardoai/lucid-origin | 48.8 | 52.1% | 45.6 | $0.0167 / image | 8.2s |
| 34 | fal/krea/v2-medium | 48.7 | 62.5% | 34.9 | $0.030 / image | 18.5s |
| 35 | local/prxpixel-t2i-7b | 48.2 | 46.4% | 50.0 | N/A | 64.7s |
| 36 | fal/ideogram/v4 | 48.0 | 62.0% | 34.0 | $0.015 / MP at BALANCED on fal | 16.6s |
| 37 | fal/krea/v2-large | 48.0 | 64.1% | 31.9 | $0.060 / image | 30.1s |
| 38 | fal/recraft/v3 | 47.8 | 41.1% | 54.6 | $0.04 / image | 10.1s |
| 39 | replicate/google/imagen-4 | 47.8 | 46.4% | 49.2 | $0.04 / image | 10.3s |
| 40 | local/bonsai-image-ternary-4b | 47.6 | 45.3% | 49.9 | N/A | 4.1s |
| 41 | local/mage-flow-rl | 46.8 | 51.0% | 42.7 | N/A | 20.9s |
| 42 | local/mage-flow-base | 46.0 | 50.0% | 42.0 | N/A | 30.4s |
| 43 | local/hidream-i1-full-17b | 45.0 | 40.1% | 49.9 | N/A | 91.3s |
| 44 | local/krea-2-raw | 44.2 | 61.5% | 26.9 | N/A | 912.2s |
| 45 | local/sefi-image-5b-turbo | 44.0 | 56.8% | 31.2 | N/A | 5.9s |
| 46 | replicate/luma/photon | 42.6 | 51.6% | 33.7 | $0.03 / image | 11.1s |
| 47 | local/nucleus-image-17b-a2b | 37.9 | 44.3% | 31.5 | N/A | 39.1s |
| 48 | local/sefi-image-2b-turbo | 35.2 | 44.3% | 26.1 | N/A | 3.6s |
| 49 | local/sana-1.5-1.6b | 34.5 | 34.9% | 34.2 | N/A | 11.1s |
| 50 | replicate/luma/photon-flash | 32.9 | 45.8% | 20.0 | $0.01 / image | 6.4s |
Local model analysis: Quality vs. size
Overall score against model size (billions of parameters) for all 21 models we ran locally on the NVIDIA DGX Spark. The green gradient deepens toward the sweet spot — smaller models that still score highly.
API model analysis: Quality vs. price
Overall score against API price per image for the 26 hosted models with a published flat per-image price. The green gradient deepens toward the sweet spot — cheaper models that still score highly.