ImageBench

ImageBench V1 —

192 evaluations across 6 categories

Benchmark V1 verdicts are produced by VLM judges and can contain mistakes. Treat PASS/FAIL labels as machine-assisted assessments, and inspect the images yourself. Learn more about the methodology.

Generation Details

Source-backed model context, size, cost, and request settings for this ImageBench V1 run.

local/sensenova-u1.5-8b-mot-base

Local

SenseNova-U1.5-8B-MoT is SenseTime OpenSenseNova's Apache-2.0 native unified multimodal model: one 8B checkpoint covers text-to-image, image editing and visual question answering. Built on a Qwen3-8B backbone with Mixture-of-Transformers generation branches and a pixel-space generator, it targets legible Chinese and English text, infographics, native 4K and complex instruction following. This entry runs the base model without the 8-step LoRA.

Maker
SenseTime (OpenSenseNova)
Family
SenseNova-U1.5
Model Size
8B
high
Cost
local run; no API price
not_applicable
Run Target
gx10/sensenova-u1.5-8b-mot-base
Effective Request
response_format: b64_json · size: 1024x1024 · num_inference_steps: 50 · guidance_scale: 4 · seed: 7
64.8
Overall
70%
Capability
59.3
Est. Preference
135
Pass
57
Fail
44.1s
Avg Latency
43.8s
Min Latency
44.2s
Max Latency
Text Rendering80%Spatial Reasoning79%Human realism57%Truthfulness41%Professional Studio96%Graphical design71%Preference59%Latency0%

All 192 generations

Text Rendering80%

Typography Style100%

Writing accuracy75%

Spatial Reasoning79%

Attributes Binding100%

Compositionality78%

Counting78%

Negation89%

Relative Position75%

Scale & Proportions56%

Human realism57%

Faces & Expressions83%

Full Body17%

Hands75%

Multi-Subject50%

Truthfulness41%

Photorealism33%

Physics & Reflections50%

World Knowledge33%

Professional Studio96%

Camera & Lighting92%

Color Precision100%

Photorealism100%

Graphical design71%

Data Visualisation33%

Layout & Design56%

Style Diversity92%