An Exam for Active Observers
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%Jiarui Zhang →
“An Exam for Active Observers”
- LinkedLinked via arxiv author · 85%Muzi Tao →
“An Exam for Active Observers”
- LinkedLinked via arxiv author · 85%Shangshang Wang →
“An Exam for Active Observers”
- LinkedLinked via arxiv author · 85%Ollie Liu →
“An Exam for Active Observers”
- LinkedLinked via arxiv author · 85%Xuezhe Ma →
“An Exam for Active Observers”
