Moonshot AI and Kimi have released PerceptionBench, a benchmark designed to separate basic visual perception from the reasoning and world knowledge that usually crowd into multimodal tests. Its blunt question is useful: before asking a model to solve the puzzle, did it actually see the right objects, text, positions and relationships?
Moonshot AI/Kimi Watch analysis
What happened
The team built its taxonomy by tracing the earliest wrong step in model answers across 42 existing benchmarks. That process produced ten categories: visual relation, counting, attribute, depth and 3D perception, localisation, comparison, fine-grained recognition, contextual integration, optical character recognition and perception-related hallucination.
The public set contains 3,000 verified short-answer questions. Moonshot says 1,800 were decomposed from model failures on source benchmarks, while 1,200 were newly written around supplementary images. Those questions were balanced from the constructed part of an internal pool of more than 17,000 verified samples.
Why it matters
Sixteen frontier multimodal models were tested under unified prompts and the highest available reasoning budget. In Moonshot's reported results, none reached 60 per cent overall accuracy. Models with similar totals also had markedly different weak spots, and perception-related hallucination was the poorest category on average.
That headline needs the correct label: this is a project-authored benchmark and evaluation, not an independent league table. Open-ended answers were graded by another model, with the team reporting 99.7 per cent agreement against humans on a 300-sample audit. Useful evidence, yes. The tablets from the mountain, no.
Our read
The value is diagnostic. A model can arrive at a polished, logically tidy answer after misreading one word, overlooking one object or inventing a visual detail. By isolating the first perceptual miss, developers can tell whether more reasoning is helping or merely fitting a better gearbox to a fogged windscreen.
PerceptionBench also makes capability profiles more informative than one aggregate score. A system strong at text recognition may still struggle with depth, counting or relations, which matters when people rely on it to inspect documents, diagrams or real scenes.
What to watch
- Independent reruns of the leaderboard and its category-level differences.
- Whether results stay stable when human graders or different judge models replace the original evaluator.
- Performance on untidy, changing real-world images rather than carefully selected benchmark samples.
Discussion spark: Which visual failure would make you trust a multimodal assistant least: missed text, wrong spatial relationships or confidently invented detail?
Sources and evidence
not affiliated with or endorsed by Moonshot AI and Kimi