ImageBench

AI Image Model Benchmark Leaderboard

Benchmark V1.2 · Ranked by pass rate on 192 prompts

50 models ranked by Overall — capability (graded pass/fail by VLM judges across 192 prompts in six categories) blended 50/50 with aesthetic Estimated Preference. Every generated image is published so you can judge with your own eyes.

V1.2 Leaderboard

Click any model to see every image it generated, or open two models to compare them side-by-side.

Don't take this ranking as the final word. Overall blends an automated VLM judge with an aesthetic preference model — both can be wrong, and neither replaces your own eyes. Every image every model generated is published, so look at them and form your own opinion. Browse the gallery.

50/50
#
1openai/gpt-image-2
80.0
87.0%45.3s
2fal/google/nano-banana-2
74.9
77.1%28.1s
3fal/reve/2.1
74.8
79.7%45.0s
4fal/google/nano-banana-pro
68.8
78.6%23.4s
5fal/bytedance/seedream-v5-pro
68.0
75.0%147.4s
6fal/microsoft/mai-image-2.5
68.0
69.3%25.4s
7local/boogu-image-turbo
64.3
62.0%9.9s
8bfl/flux-2-pro
62.2
67.2%11.8s
9fal/bytedance/seedream-v4
62.2
68.2%14.1s
10replicate/google/imagen-4-ultra
62.0
62.0%14.6s
11fal/luma/uni-1-max
62.0
78.6%88.2s
12bfl/flux-2-max
61.9
67.2%26.7s
13local/qwen-image-2512-20b
59.6
52.1%80.2s
14local/krea-2-turbo
59.1
63.0%73.2s
15local/flux-2-klein-9b
59.0
57.8%8.5s
16fal/hunyuan-image-3
57.7
52.1%24.4s
17bfl/flux-2-klein-9b
57.2
57.3%4.1s
18local/krea-2-turbo-no-filter
56.2
58.3%71.5s
19replicate/recraft-ai/recraft-v4
55.5
62.5%9.5s
20fal/bytedance/seedream-v5-lite
54.8
76.6%35.5s
21local/flux-2-klein-4b
53.6
52.6%4.5s
22replicate/google/imagen-4-fast
53.0
51.6%5.8s
23bfl/flux-2-klein-4b
51.8
50.5%3.8s
24replicate/minimax/image-01
51.8
50.0%22.4s
25local/sefi-image-5b-base
51.5
62.5%130.6s
26local/z-image-turbo-6b
51.3
52.6%18.1s
27fal/ideogram/v3
50.8
48.4%12.9s
28fal/krea/v2-medium-turbo
50.7
60.9%15.5s
29fal/bria/fast
50.5
56.2%12.4s
30local/sefi-image-5b-rl
50.0
60.9%130.7s
31local/mage-flow-turbo
49.9
54.2%4.6s
32local/z-image-6b
49.3
56.8%130.7s
33replicate/leonardoai/lucid-origin
48.8
52.1%8.2s
34fal/krea/v2-medium
48.7
62.5%18.5s
35local/prxpixel-t2i-7b
48.2
46.4%64.7s
36fal/ideogram/v4
48.0
62.0%16.6s
37fal/krea/v2-large
48.0
64.1%30.1s
38fal/recraft/v3
47.8
41.1%10.1s
39replicate/google/imagen-4
47.8
46.4%10.3s
40local/bonsai-image-ternary-4b
47.6
45.3%4.1s
41local/mage-flow-rl
46.8
51.0%20.9s
42local/mage-flow-base
46.0
50.0%30.4s
43local/hidream-i1-full-17b
45.0
40.1%91.3s
44local/krea-2-raw
44.2
61.5%912.2s
45local/sefi-image-5b-turbo
44.0
56.8%5.9s
46replicate/luma/photon
42.6
51.6%11.1s
47local/nucleus-image-17b-a2b
37.9
44.3%39.1s
48local/sefi-image-2b-turbo
35.2
44.3%3.6s
49local/sana-1.5-1.6b
34.5
34.9%11.1s
50replicate/luma/photon-flash
32.9
45.8%6.4s

Local model analysis: Quality vs. size

Overall score against model size (billions of parameters) for all 21 models we ran locally on the NVIDIA DGX Spark. The green gradient deepens toward the sweet spot — smaller models that still score highly.

smaller & higher quality203040506070800B4B8B12B16B20BModel size (billions of parameters)Overall scoreboogu-image-turbo — 10B, 64.3boogu-image-turboqwen-image-2512-20b — 20B, 59.6qwen-image-2512-20bkrea-2-turbo — 12B, 59.1krea-2-turboflux-2-klein-9b — 9B, 59.0flux-2-klein-9bkrea-2-turbo-no-filter — 12B, 56.2krea-2-turbo-no-filterflux-2-klein-4b — 4B, 53.6flux-2-klein-4bsefi-image-5b-base — 5B, 51.5sefi-image-5b-basez-image-turbo-6b — 6B, 51.3z-image-turbo-6bsefi-image-5b-rl — 5B, 50.0sefi-image-5b-rlmage-flow-turbo — 4B, 49.9mage-flow-turboz-image-6b — 6B, 49.3z-image-6bprxpixel-t2i-7b — 7B, 48.2prxpixel-t2i-7bbonsai-image-ternary-4b — 4B, 47.6bonsai-image-ternary-4bmage-flow-rl — 4B, 46.8mage-flow-rlmage-flow-base — 4B, 46.0mage-flow-basehidream-i1-full-17b — 17B, 45.0hidream-i1-full-17bkrea-2-raw — 12B, 44.2krea-2-rawsefi-image-5b-turbo — 5B, 44.0sefi-image-5b-turbonucleus-image-17b-a2b — 17B, 37.9nucleus-image-17b-a2bsefi-image-2b-turbo — 2B, 35.2sefi-image-2b-turbosana-1.5-1.6b — 1.6B, 34.5sana-1.5-1.6b

API model analysis: Quality vs. price

Overall score against API price per image for the 26 hosted models with a published flat per-image price. The green gradient deepens toward the sweet spot — cheaper models that still score highly.

cheaper & higher quality20406080100$0.00$0.04$0.08$0.12$0.16API price per image (USD)Overall scoregoogle/nano-banana-2 — $0.080, 74.9google/nano-banana-2reve/2.1 — $0.040, 74.8reve/2.1google/nano-banana-pro — $0.150, 68.8google/nano-banana-probytedance/seedream-v5-pro — $0.068, 68.0bytedance/seedream-v5-promicrosoft/mai-image-2.5 — $0.050, 68.0microsoft/mai-image-2.5bfl/flux-2-pro — $0.030, 62.2bfl/flux-2-probytedance/seedream-v4 — $0.030, 62.2bytedance/seedream-v4replicate/google/imagen-4-ultra — $0.060, 62.0replicate/google/imagen-4-ultraluma/uni-1-max — $0.102, 62.0luma/uni-1-maxbfl/flux-2-max — $0.070, 61.9bfl/flux-2-maxhunyuan-image-3 — $0.090, 57.7hunyuan-image-3bfl/flux-2-klein-9b — $0.015, 57.2bfl/flux-2-klein-9breplicate/recraft-ai/recraft-v4 — $0.040, 55.5replicate/recraft-ai/recraft-v4bytedance/seedream-v5-lite — $0.035, 54.8bytedance/seedream-v5-litereplicate/google/imagen-4-fast — $0.020, 53.0replicate/google/imagen-4-fastbfl/flux-2-klein-4b — $0.014, 51.8bfl/flux-2-klein-4breplicate/minimax/image-01 — $0.010, 51.8replicate/minimax/image-01ideogram/v3 — $0.060, 50.8ideogram/v3bria/fast — $0.028, 50.5bria/fastreplicate/leonardoai/lucid-origin — $0.017, 48.8replicate/leonardoai/lucid-originkrea/v2-medium — $0.030, 48.7krea/v2-mediumkrea/v2-large — $0.060, 48.0krea/v2-largerecraft/v3 — $0.040, 47.8recraft/v3replicate/google/imagen-4 — $0.040, 47.8replicate/google/imagen-4replicate/luma/photon — $0.030, 42.6replicate/luma/photonreplicate/luma/photon-flash — $0.010, 32.9replicate/luma/photon-flash