Moonshot AI has released PerceptionBench, an open benchmark for testing what multimodal models actually see. Across 16 frontier models, the company says none reached 60% accuracy, with perception-related hallucination the weakest capability on average.
Moonshot AI/Kimi Watch analysis
What happened
The PerceptionBench announcement describes 3,000 verified questions across 10 visual capabilities, including counting, OCR, depth, localisation and hallucination. Each question is designed to test a single perceptual skill, without requiring reasoning or outside knowledge.
Moonshot says the benchmark’s categories were derived from model failures across 42 existing benchmarks. It also says models with nearly identical overall scores can differ sharply in what they perceive. The released dataset is a 3,000-question sample from a larger in-house pool of more than 17,000 questions; the company says the dataset and evaluation code are open source.
Why it matters
A model can give a convincing answer while getting a basic visual detail wrong. PerceptionBench is intended to help researchers pinpoint where that happens, rather than letting a broad overall score blur together counting, reading text and recognising relationships in an image. That could make comparisons more useful for teams choosing or improving multimodal systems.
The results are Moonshot’s account of its own benchmark and evaluations. The headline finding is a warning about current visual perception, not a universal ranking: the useful test will be whether other teams can reproduce the results and whether improvements on these questions carry over to real tasks.
Our read
The best idea here is the benchmark’s insistence on asking what a model can see before asking it to reason about what it sees. That is tidy experimental design, and a welcome antidote to scores that make two very different systems look like twins. Now the wider research community gets to kick the tyres.
What to watch
- Whether independent teams reproduce the reported results across the 16 models.
- Whether model makers use the capability-level results to improve specific weaknesses.
- Whether better scores on PerceptionBench translate to more reliable visual work outside the benchmark.
Discussion spark: Should multimodal models be judged primarily on broad overall scores, or on separate evidence that they can reliably count, read and locate things in images?
Sources and evidence
- PerceptionBench: Evaluating Atomic Visual Perception in MLLMs (Publication date not supplied)
not affiliated with or endorsed by Moonshot AI and Kimi