Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO
This paper introduces Actor as Its Own Critic, a unified reinforcement learning framework, Cycle Group Relative Policy Optimization (CycleGRPO), that jointly optimizes region understanding and localization for Multimodal Large Language Models (MLLMs). Unlike existing separate pipelines, we leverage the inherent duality between the two tasks to construct a self-evaluating reinforcement learning paradigm: "region $\to$ text $\to$ region''. Specifically, a single MLLM first acts as the actor to generate region captions, then immediately transitions to a critic to ground its generated text back in
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%Atomic-man007/Awesome_Multimodel_LLM →
- LinkedLinked via arxiv author · 85%Qingxin Zhang →
“Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO”
- LinkedLinked via arxiv author · 85%Haochen Wang →
“Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO”
- LinkedLinked via arxiv author · 85%Yikang Zhou →
“Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO”
- LinkedLinked via arxiv author · 85%Jason Li →
“Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO”
- LinkedLinked via arxiv author · 85%Robby T. Tan →
“Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO”
- PossiblePossibly related (embedding) · 54%Seeking collaborators for scaling and independent evaluation of a new recurrent language model architecture (preprint + code) [R] →
