MIRROR: Learning from the Other View for Multi-Modal Reasoning
Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diagram, and combined diagram+text views. We show that these views often elicit different behaviors: a model may solve a problem from text but fail on the corresponding diagram, or succeed visually while failing textually. This inconsistency suggests that different views expose complementary reasoning paths and failure modes that standard multimodal post-training does not fully exploit. To study and e
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%aymericdamien/TopDeepLearning →
“Fuzzy title match (0.73): “MIRROR: Learning from the Other View for Multi-Modal Reasoni” ≈ “aymericdamien/TopDeepLearning””
- LinkedLinked via arxiv author · 85%Wen Ye →
“MIRROR: Learning from the Other View for Multi-Modal Reasoning”
- LinkedLinked via arxiv author · 85%Yuxiao Qu →
“MIRROR: Learning from the Other View for Multi-Modal Reasoning”
- LinkedLinked via arxiv author · 85%Aviral Kumar →
“MIRROR: Learning from the Other View for Multi-Modal Reasoning”
- LinkedLinked via arxiv author · 85%Xuezhe Ma →
“MIRROR: Learning from the Other View for Multi-Modal Reasoning”
- PossiblePossibly related (embedding) · 47%[BIG DATASET RELEASE] - SupraLabs/reasoning-corpus-4K-5M-v1 - Train your tiny SLMs to think! →
