Modality-Driven Search with Holistic Trace Judging for ARC-AGI-2
Large language models can produce fluent, internally coherent reasoning traces for abstract reasoning tasks while still being confidently wrong - making selection among candidates, not just generation, the central challenge. I present a solver for ARC-AGI-2, a few-shot visual reasoning benchmark, built around two principles: (i) treating reasoning modalities as search operators, generating diverse candidates independently across text, image, and code channels, and (ii) context-preserving holistic judging, in which a judge model jointly compares all candidate reasoning traces within a single lo
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownNew benchmark exposes reasoning gaps in top models →
- LinkedLinked via unknownNorthwind AI →
- LinkedLinked via unknownRetrieval is underrated →
- LinkedLinked via unknownSenior SWE Bench: a new benchmark focussed on realistically underspecified feature tasks →
- PossiblePossibly related (embedding) · 53%wbopan/flashtrace →
- PossiblePossibly related (embedding) · 54%Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova →
