Read original ↗
paperarXivTrust 82 · PrimaryPublished 5d agoLive · 4d ago

LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning

Multi-view spatial reasoning requires vision-language models to compare visual evidence across images, align object correspondences, and infer spatial relations over long visual contexts, a setting where chain-of-thought reasoning tends to grow verbose without becoming more accurate. Reinforcement learning with verifiable rewards is a natural fit for this task, but standard GRPO reward relies on sparse outcome-level feedback and gives no signal about where a reasoning trajectory goes wrong, nor any control over its length. We propose LenGuard-GPC, a dense reward framework that addresses both p

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%ray-project/ray

    Shared author/contributor keys: wang

  • FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow

    Shared author/contributor keys: wang

  • FuzzySimilar title/name (fuzzy) · 59%aymericdamien/TopDeepLearning

    Fuzzy title match (0.73): “LenGuard-GPC: Length Guarding with Guided-Prompt Consistency” ≈ “aymericdamien/TopDeepLearning”

  • FuzzySimilar title/name (fuzzy) · 59%linshenkx/prompt-optimizer

    Fuzzy title match (0.73): “LenGuard-GPC: Length Guarding with Guided-Prompt Consistency” ≈ “linshenkx/prompt-optimizer”

  • FuzzySimilar title/name (fuzzy) · 59%NirDiamant/Prompt_Engineering

    Fuzzy title match (0.73): “LenGuard-GPC: Length Guarding with Guided-Prompt Consistency” ≈ “NirDiamant/Prompt_Engineering”

  • LinkedLinked via arxiv author · 85%Xingjian Tao

    LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning

  • LinkedLinked via arxiv author · 85%Yiwei Wang

    LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning

  • LinkedLinked via arxiv author · 85%Yujun Cai

    LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning

Implements (incoming)

authored (incoming)

Related across the graph

Topics