Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

Fine-grained visual reasoning remains challenging for vision-language models, especially when small but critical visual cues are buried in high-resolution images. Existing approaches rely on repeated cropping or test-time visual search to introduce local evidence, but they typically do not explicitly distinguish perception from reasoning. In this paper, we propose Perceive-to-Reason (P2R), a unified framework that formulates fine-grained visual reasoning as a two-stage process: the model first localizes question-relevant evidence as a Perceiver, and then answers the question as a Reasoner base

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via unknownVioletVision-3B
  • LinkedLinked via arxiv author · 85%Hongxing Li

    Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

  • LinkedLinked via arxiv author · 85%Xiufeng Huang

    Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

  • LinkedLinked via arxiv author · 85%Dingming Li

    Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

  • LinkedLinked via arxiv author · 85%Wenjing Jiang

    Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

  • LinkedLinked via arxiv author · 85%Zixuan Wang

    Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

  • LinkedLinked via arxiv author · 85%Haolei Xu

    Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

  • LinkedLinked via arxiv author · 85%Hanrong Zhang

    Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

Has model

authored (incoming)

Implements (incoming)

Related across the graph

Topics