ImageBench

AI Image Model Benchmark Leaderboard

Benchmark V1 · Ranked by pass rate on 192 prompts

35 models ranked by Overall — capability (graded pass/fail by VLM judges across 192 prompts in six categories) blended 50/50 with aesthetic Estimated Preference. Every generated image is published so you can judge with your own eyes.

V1 Leaderboard

Click any model to see every image it generated, or open two models to compare them side-by-side.

35/35
#
1openai/gpt-image-2
78.5
83.9%45.3s
2fal/google/nano-banana-2
73.0
73.4%28.1s
3fal/google/nano-banana-pro
66.2
73.4%23.4s
4fal/bytedance/seedream-v5-pro
65.6
70.3%147.4s
5local/boogu-image-turbo
61.9
57.3%9.9s
6bfl/flux-2-pro
60.1
63.0%11.8s
7bfl/flux-2-max
59.6
62.5%26.7s
8fal/bytedance/seedream-v4
59.4
62.5%14.1s
9local/qwen-image-2512-20b
58.6
50.0%80.2s
10local/flux-2-klein-9b
56.6
53.1%8.5s
11local/krea-2-turbo
56.5
57.8%73.2s
12bfl/flux-2-klein-9b
55.3
53.6%4.1s
13local/krea-2-turbo-no-filter
53.8
53.6%71.5s
14fal/bytedance/seedream-v5-lite
52.7
72.4%35.5s
15local/flux-2-klein-4b
51.5
48.4%4.5s
16bfl/flux-2-klein-4b
49.8
46.4%3.8s
17local/sefi-image-5b-base
49.7
58.9%130.6s
18local/z-image-turbo-6b
49.2
48.4%18.1s
19fal/krea/v2-medium-turbo
48.7
56.8%15.5s
20fal/ideogram/v3
48.5
43.8%12.9s
21fal/bria/fast
47.9
51.0%12.4s
22local/sefi-image-5b-rl
47.7
56.3%130.7s
23local/z-image-6b
47.4
53.1%130.7s
24fal/krea/v2-medium
46.9
58.9%18.5s
25local/bonsai-image-ternary-4b
46.5
43.2%4.1s
26local/prxpixel-t2i-7b
46.3
42.7%64.7s
27fal/ideogram/v4
46.1
58.3%16.6s
28fal/krea/v2-large
45.9
59.9%30.1s
29fal/recraft/v3
45.5
36.5%10.1s
30local/hidream-i1-full-17b
42.9
35.9%91.3s
31local/krea-2-raw
42.1
57.3%912.2s
32local/sefi-image-5b-turbo
41.9
52.6%5.9s
33local/nucleus-image-17b-a2b
36.6
41.7%39.1s
34local/sefi-image-2b-turbo
33.6
41.1%3.6s
35local/sana-1.5-1.6b
33.0
31.8%11.1s

Local model analysis: Quality vs. size

Overall score against model size (billions of parameters) for all 18 models we ran locally on the NVIDIA DGX Spark. The green gradient deepens toward the sweet spot — smaller models that still score highly.

smaller & higher quality203040506070800B4B8B12B16B20BModel size (billions of parameters)Overall scoreboogu-image-turbo — 10B, 61.9boogu-image-turboqwen-image-2512-20b — 20B, 58.6qwen-image-2512-20bflux-2-klein-9b — 9B, 56.6flux-2-klein-9bkrea-2-turbo — 12B, 56.5krea-2-turbokrea-2-turbo-no-filter — 12B, 53.8krea-2-turbo-no-filterflux-2-klein-4b — 4B, 51.5flux-2-klein-4bsefi-image-5b-base — 5B, 49.7sefi-image-5b-basez-image-turbo-6b — 6B, 49.2z-image-turbo-6bsefi-image-5b-rl — 5B, 47.7sefi-image-5b-rlz-image-6b — 6B, 47.4z-image-6bbonsai-image-ternary-4b — 4B, 46.5bonsai-image-ternary-4bprxpixel-t2i-7b — 7B, 46.3prxpixel-t2i-7bhidream-i1-full-17b — 17B, 42.9hidream-i1-full-17bkrea-2-raw — 12B, 42.1krea-2-rawsefi-image-5b-turbo — 5B, 41.9sefi-image-5b-turbonucleus-image-17b-a2b — 17B, 36.6nucleus-image-17b-a2bsefi-image-2b-turbo — 2B, 33.6sefi-image-2b-turbosana-1.5-1.6b — 1.6B, 33.0sana-1.5-1.6b

API model analysis: Quality vs. price

Overall score against API price per image for the 14 hosted models with a published flat per-image price. The green gradient deepens toward the sweet spot — cheaper models that still score highly.

cheaper & higher quality20406080100$0.00$0.04$0.08$0.12$0.16API price per image (USD)Overall scoregoogle/nano-banana-2 — $0.080, 73.0google/nano-banana-2google/nano-banana-pro — $0.150, 66.2google/nano-banana-probytedance/seedream-v5-pro — $0.068, 65.6bytedance/seedream-v5-probfl/flux-2-pro — $0.030, 60.1bfl/flux-2-probfl/flux-2-max — $0.070, 59.6bfl/flux-2-maxbytedance/seedream-v4 — $0.030, 59.4bytedance/seedream-v4bfl/flux-2-klein-9b — $0.015, 55.3bfl/flux-2-klein-9bbytedance/seedream-v5-lite — $0.035, 52.7bytedance/seedream-v5-litebfl/flux-2-klein-4b — $0.014, 49.8bfl/flux-2-klein-4bideogram/v3 — $0.060, 48.5ideogram/v3bria/fast — $0.028, 47.9bria/fastkrea/v2-medium — $0.030, 46.9krea/v2-mediumkrea/v2-large — $0.060, 45.9krea/v2-largerecraft/v3 — $0.040, 45.5recraft/v3