The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology
Model organisms (MOs) - language models trained to exhibit undesired or unnatural behaviours - are frequently used as testbeds for evaluating white-box interpretability techniques. Current MOs are typically constructed via post-hoc supervised fine-tuning (SFT) on behavioural transcripts or synthetic documents. Prior research has shown that interpretability methods can easily identify hidden behaviours in these MOs. However, recent work suggests that such post-hoc training methods may make interpretability unrealistically easy. We investigate this claim by constructing a suite of 54 $\verb|OLMo
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownIEEE Rolls Out Large Language Models Virtual Training Course →
- LinkedLinked via unknownKnowledge Distillation of Black-Box Large Language Models →
- LinkedLinked via unknownKnowledge Distillation of Black-Box Large Language Models (2024) →
- PossiblePossibly related (embedding) · 30%interpretml/interpret →
“Possibly related via embedding similarity 0.58 (not asserted). Timestamp check: artifact after paper (+12d).”
- PossiblePossibly related (embedding) · 53%Understanding Annotator Safety Policy with Interpretability - Apple Machine Learning Research →
- PossiblePossibly related (embedding) · 48%A unifying framework from neural superposition to sparse interpretable codes - Nature →
- PossiblePossibly related (embedding) · 49%Large language models for interpretation of health checkup results - Nature →
