# PerceptionBench: Evaluating Atomic Visual Perception in MLLMs

Kimi Team released PerceptionBench, a 3,000-question benchmark isolating 10 atomic visual perception capabilities, on which no frontier MLLM reaches 60% accuracy.

- No frontier MLLM reaches 60% accuracy; perception-related hallucination is the weakest capability on average.
- The 3,000 questions are 60% decomposed from attributed model failures and 40% newly authored, drawn from a pool of 17,000+ verified questions.
- The 10 categories are Visual Relation, Counting, Attribute, Depth & 3D, Localization, Comparison, Fine-grained Recognition, Context Integration, OCR, and Hallucination.
- Dataset and evaluation code are open-sourced at github.com/MoonshotAI/PerceptionBench.

## Why it matters

Multimodal model developers can now measure and diagnose which atomic perception capabilities their models lack, since models with nearly identical overall scores diverge sharply in what they actually perceive.

## Sources

- [Moonshot AI: PerceptionBench: Evaluating Atomic Visual Perception in MLLMs](https://www.kimi.com/blog/perception-bench)

---

Summarized by dstilled on 2026-07-22. https://dstilled.ai/story/f58b93a8-9c2f-48bf-a012-a1deecb18497
