Read original ↗
paperarXivTrust 82 · PrimaryPublished 4d agoLive · 2d ago

MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models

Large language models achieve strong reasoning performance, but often at prohibitive training cost - a challenge that is especially acute for compact models ($\leq 4 \, \mathrm{B}$ parameters) trained under limited budgets. We introduce MADA-RL, a post-training framework that specializes compact models into generator and critic roles and trains them with a debate-aware learning signal, fine-tuning only a small subset of parameters via LoRA adapters. Our central contribution is a counterfactual critic advantage: a dynamic, role-conditioned baseline that redefines the critic's advantage as its r

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 59%AgentCore-8B

    Fuzzy title match (0.73): “MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for” ≈ “AgentCore-8B”

  • FuzzySimilar title/name (fuzzy) · 87%SWE-agent/SWE-agent

    Fuzzy title match (0.94): “MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for” ≈ “SWE-agent/SWE-agent”

  • FuzzySimilar title/name (fuzzy) · 87%zhayujie/CowAgent

    Fuzzy title match (0.94): “MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for” ≈ “zhayujie/CowAgent”

  • FuzzySimilar title/name (fuzzy) · 66%open-multi-agent/open-multi-agent

    Fuzzy title match (0.78): “MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for” ≈ “open-multi-agent/open-multi-agent”

  • FuzzySimilar title/name (fuzzy) · 59%NousResearch/hermes-agent

    Fuzzy title match (0.73): “MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for” ≈ “NousResearch/hermes-agent”

  • FuzzySimilar title/name (fuzzy) · 59%bojieli/ai-agent-book

    Fuzzy title match (0.73): “MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for” ≈ “bojieli/ai-agent-book”

  • LinkedLinked via arxiv author · 85%Martino M. L. Pulici

    MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models

  • LinkedLinked via arxiv author · 85%Cuong Xuan Chu

    MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models

Has model

Implements (incoming)

authored (incoming)

Related across the graph

Topics