Read original ↗
paperarXivTrust 82 · PrimaryPublished 6d agoLive · 2d ago

Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models

Multimodal large language models (MLLMs) combine linguistic reasoning with visual perception, yet their ability to perform visual spatial planning under explicit or previously unseen rule constraints remains underexplored. This setting requires models to jointly understand spatial layouts, interpret natural-language rules, and plan valid actions accordingly. To address this gap, we introduce RuleMaze, a controllable benchmark in which MLLMs must navigate mazes while obeying natural-language rules of varying complexity. RuleMaze isolates rule-compliant spatial planning by requiring accurate per

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%modular/modular

    Shared author/contributor keys: liu

  • LinkedLinked via arxiv author · 85%Zhiyu Chen

    Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models

  • LinkedLinked via arxiv author · 85%Ting Lei

    Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models

  • LinkedLinked via arxiv author · 85%Yaoyi Li

    Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models

  • LinkedLinked via arxiv author · 85%Jia Cai

    Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models

  • LinkedLinked via arxiv author · 85%Zhecen Wu

    Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models

  • LinkedLinked via arxiv author · 85%Dongyang Liu

    Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models

Implements (incoming)

authored (incoming)

Related across the graph

Topics