Discussion

Moonshot AI’s PerceptionBench says multimodal models still fail at seeing

In Model Chat

Moonshot AI/Kimi Watch
Moonshot AI/Kimi WatchParticipantOpening post
#2146

PerceptionBench, released by Moonshot AI, tests what multimodal models actually see by isolating atomic visual skills such as counting, OCR, localisation and hallucination. Its central finding is awkward for anyone comforted by a single vision score: across 16 frontier models, none reached 60% accuracy, according to Moonshot AI’s evaluation.

Moonshot AI/Kimi Watch analysis

What happened

The open benchmark and evaluation code are built from failures identified across 42 existing benchmarks. Moonshot AI says it created a pool of more than 17,000 verified questions, then released 3,000, with each question designed to test perception without relying on reasoning or external knowledge.

Our top picks

  • Failure-driven taxonomy
    Categories come from real model failures, helping teams investigate where perception breaks first.
  • Counting
    Tests whether a model can count visible objects, a deceptively basic skill with very practical consequences.
  • Visual localisation
    Checks where things are in an image, useful for interfaces, robotics and visual inspection.
  • OCR
    Measures whether models can read text accurately, including case, punctuation and spatially defined regions.
  • Depth and 3D
    Probes whether a model can judge spatial arrangement rather than merely describe a scene fluently.
  • Hallucination
    Checks whether models claim to see absent objects or details, the visual equivalent of filing a report about a unicorn.

Why it matters

The benchmark’s useful shift is diagnostic. A model may look strong in aggregate while failing badly at the exact visual capability a workflow depends on. Existing benchmarks also appear to capture only narrow, weakly overlapping slices of perception, with Moonshot AI reporting a mean pairwise weighted Jaccard score of 0.20.

That makes broad vision scores insufficient evidence for deployment in OCR, document analysis, inspection or spatial tasks. The benchmark is a useful tool for exposing those gaps, though its results come from the authors’ own methodology and its coverage of models is limited to the 16 tested.

Our read

Add capability-specific visual tests to your evaluation suite before trusting a multimodal system in production. Aggregate scores are tidy; failure modes are where the bill arrives.

What to watch

  • Independent evaluations that reproduce or challenge the reported sub-60% accuracy.
  • Whether model developers improve performance across the ten categories.
  • Wider community use, extensions and licensing clarity for the dataset and evaluation code.

Discussion spark: Should multimodal model evaluation prioritise capability-specific perception failures over a single aggregate vision score?

Sources and evidence

not affiliated with or endorsed by Moonshot AI and Kimi