Read original ↗
paperarXivTrust 82 · PrimaryPublished 3d agoLive · yesterday

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a monitor to detect covert sabotage before deployment. We evaluate AI control for automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization. Because the deliverable in AI R&D is an artifact that

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 64%Securing the future of AI agents
  • PossiblePossibly related (embedding) · 58%Safety and alignment in an era of long-horizon models
  • FuzzySimilar title/name (fuzzy) · 59%google-research/google-research

    Fuzzy title match (0.73): “ResearchArena: Evaluating Sabotage and Monitoring in Automat” ≈ “google-research/google-research”

  • LinkedLinked via arxiv author · 85%Lena Libon

    ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

  • LinkedLinked via arxiv author · 85%Ben Rank

    ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

  • LinkedLinked via arxiv author · 85%Jehyeok Yeon

    ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

  • LinkedLinked via arxiv author · 85%David Schmotz

    ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

  • LinkedLinked via arxiv author · 85%Jeremy Qin

    ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

Covers

Implements (incoming)

authored (incoming)

Related across the graph

Topics