ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a monitor to detect covert sabotage before deployment. We evaluate AI control for automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization. Because the deliverable in AI R&D is an artifact that
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 64%Securing the future of AI agents →
- PossiblePossibly related (embedding) · 58%Safety and alignment in an era of long-horizon models →
- FuzzySimilar title/name (fuzzy) · 59%google-research/google-research →
“Fuzzy title match (0.73): “ResearchArena: Evaluating Sabotage and Monitoring in Automat” ≈ “google-research/google-research””
- LinkedLinked via arxiv author · 85%Lena Libon →
“ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D”
- LinkedLinked via arxiv author · 85%Ben Rank →
“ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D”
- LinkedLinked via arxiv author · 85%Jehyeok Yeon →
“ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D”
- LinkedLinked via arxiv author · 85%David Schmotz →
“ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D”
- LinkedLinked via arxiv author · 85%Jeremy Qin →
“ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D”
