SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essential to bridge this gap, among which Sparse Autoencoders (SAEs) serve as a cornerstone by isolating interpretable features for model inspection and steering. In this paper, we introduce SAEScientist-Bench to evaluate whether AI agents can act as scientists utilizing SAE tools for autonomous mechanist
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 60%Can AI Improve Itself? RSI Might Be the Answer [R] →
- PossiblePossibly related (embedding) · 59%An Anthropic researcher just gave us a peek at self-improving AI →
- FuzzySimilar title/name (fuzzy) · 87%NirDiamant/GenAI_Agents →
“Fuzzy title match (0.94): “SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Int” ≈ “NirDiamant/GenAI_Agents””
- FuzzySimilar title/name (fuzzy) · 84%Unity-Technologies/ml-agents →
“Fuzzy title match (0.92): “SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Int” ≈ “Unity-Technologies/ml-agents””
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzySimilar title/name (fuzzy) · 59%datawhalechina/hello-agents →
“Fuzzy title match (0.73): “SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Int” ≈ “datawhalechina/hello-agents””
- FuzzySimilar title/name (fuzzy) · 59%jnMetaCode/agency-agents-zh →
“Fuzzy title match (0.73): “SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Int” ≈ “jnMetaCode/agency-agents-zh””
- LinkedLinked via arxiv author · 85%Yuqiao Tan →
“SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?”
