Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

Visually Grounded Self-Reflection for Vision-Language Models via Reinforcement Learning

Large vision-language models can reason over multimodal inputs by generating textual chains of thought (CoT). A key capability exhibited in CoT reasoning is self-reflection: revisiting earlier decisions and correcting previous errors. However, existing LVLMs often fail to properly attend to visual inputs during reflection, limiting their ability to translate feedback into grounded corrections, especially for out-of-distribution images. To address this issue, we propose a novel reinforcement learning training framework VRRL, with two components explicitly designed to elicit visually grounded se

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B

    Fuzzy title match (0.73): “Visually Grounded Self-Reflection for Vision-Language Models” ≈ “VioletVision-3B”

  • LinkedLinked via arxiv author · 85%Liyan Tang

    Visually Grounded Self-Reflection for Vision-Language Models via Reinforcement Learning

  • LinkedLinked via arxiv author · 85%Fangcong Yin

    Visually Grounded Self-Reflection for Vision-Language Models via Reinforcement Learning

  • LinkedLinked via arxiv author · 85%Greg Durrett

    Visually Grounded Self-Reflection for Vision-Language Models via Reinforcement Learning

  • PossiblePossibly related (embedding) · 48%om-ai-lab/VLM-R1
  • FuzzySimilar title/name (fuzzy) · 84%pytorch/vision

    Fuzzy title match (0.92): “Visually Grounded Self-Reflection for Vision-Language Models” ≈ “pytorch/vision”

  • FuzzyOverlapping authors or contributors · 62%rasbt/LLMs-from-scratch

    Shared author/contributor keys: yin

  • FuzzySimilar title/name (fuzzy) · 59%aymericdamien/TopDeepLearning

    Fuzzy title match (0.73): “Visually Grounded Self-Reflection for Vision-Language Models” ≈ “aymericdamien/TopDeepLearning”

Has model

authored (incoming)

Implements (incoming)

Covers (incoming)

Related across the graph

Topics