DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind setting, where ground-truth success signals are not available. Multi-criteria rubrics are a popular way to supply such a reward; they are scored once per trajectory, but a single scalar is a poor signal across tens of steps. We propose DRACO: Distributing Rubric-based Advantage for Credit Optimization. It generates rubrics dynamically during training to track the policy's evolving capability, scores those rubrics once per
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%AgentCore-8B →
“Fuzzy title match (0.73): “DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics f” ≈ “AgentCore-8B””
- FuzzySimilar title/name (fuzzy) · 87%SWE-agent/SWE-agent →
“Fuzzy title match (0.94): “DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics f” ≈ “SWE-agent/SWE-agent””
- FuzzySimilar title/name (fuzzy) · 87%zhayujie/CowAgent →
“Fuzzy title match (0.94): “DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics f” ≈ “zhayujie/CowAgent””
- FuzzySimilar title/name (fuzzy) · 84%Thysrael/Horizon →
“Fuzzy title match (0.92): “DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics f” ≈ “Thysrael/Horizon””
- FuzzyOverlapping authors or contributors · 62%huggingface/transformers →
“Shared author/contributor keys: gandhi”
- FuzzySimilar title/name (fuzzy) · 59%NousResearch/hermes-agent →
“Fuzzy title match (0.73): “DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics f” ≈ “NousResearch/hermes-agent””
