ImageBench

ImageBench V1 —

192 evaluations across 6 categories

Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this ImageBench V1 run.

local/sensenova-u1.5-8b-mot-8step

Local

SenseNova-U1.5-8B-MoT is SenseTime OpenSenseNova's Apache-2.0 native unified multimodal model: one 8B checkpoint covers text-to-image, image editing and visual question answering. Built on a Qwen3-8B backbone with Mixture-of-Transformers generation branches and a pixel-space generator, it targets legible Chinese and English text, infographics, native 4K and complex instruction following. This entry runs the official 8-step distillation LoRA.

Maker
SenseTime (OpenSenseNova)
Family
SenseNova-U1.5
Model Size
8B
high
Cost
local run; no API price
not_applicable
Run Target
gx10/sensenova-u1.5-8b-mot-8step
Effective Request
response_format: b64_json · size: 1024x1024 · num_inference_steps: 8 · guidance_scale: 1 · seed: 7
64.0
Overall
64%
Capability
64.4
Est. Preference
122
Pass
70
Fail
4.1s
Avg Latency
3.9s
Min Latency
4.7s
Max Latency
Text Rendering67%Spatial Reasoning68%Human realism45%Truthfulness33%Professional Studio89%Graphical design88%Preference64%Latency58%

All 192 generations

Text Rendering67%

Typography Style100%

Writing accuracy58%

Spatial Reasoning68%

Attributes Binding100%

Compositionality78%

Counting56%

Negation67%

Relative Position67%

Scale & Proportions44%

Human realism45%

Faces & Expressions67%

Full Body8%

Hands75%

Multi-Subject17%

Truthfulness33%

Photorealism33%

Physics & Reflections25%

World Knowledge42%

Professional Studio89%

Camera & Lighting92%

Color Precision83%

Photorealism100%

Graphical design88%

Data Visualisation33%

Layout & Design100%

Style Diversity92%