newsReddit r/MachineLearningTrust 52 · CommunityPublished 23d agoLive · 21d ago
GPT-5.5 Scores 10.6% on ActiveVision, Humans Hit 96.1% [R]
The interesting finding from a new [arXiv paper]( https://arxiv.org/abs/2607.16165 ) isn't that a frontier vision model failed a new benchmark, that happens weekly, but the specific shape of the failure and the fact that the models cannot patch it by writing their own code. The benchmark, called ActiveVision, contains 17 tasks across 3 categories designed, in the authors' words, to "force repeated visual p
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 59%An Exam for Active Observers →
- PossiblePossibly related (embedding) · 53%SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation →
- PossiblePossibly related (embedding) · 50%From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA →
- PossiblePossibly related (embedding) · 47%VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context →
- PossiblePossibly related (embedding) · 47%The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric →
Covers
paperAn Exam for Active ObserverspaperSHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report GenerationpaperFrom Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQApaperVisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual ContextpaperThe Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric
Related across the graph
paperThe Many Senses of Visual Similarity: A Text-Prompted Image Perceptual MetricpaperAn Exam for Active ObserverspaperSHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report GenerationpaperVisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual ContextpaperFrom Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA
