Surrogate Fidelity: When Can Open LLMs Explain Closed Ones?
Mechanistic interpretability (MI) requires full access to model internals, yet the APIs for most widely deployed language models at best expose log-probabilities over output tokens. This creates a surrogate problem: when do measurements made on open models allow us to make claims about a closed model? We evaluate surrogate fidelity at the prediction, attribution, and representation levels. For binary classification tasks, log-odds provide an API-compatible scalar readout of the model's representation space, and leave-one-out attributions provide insight into model behavior. Across eleven model
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownIs it agentic enough? Benchmarking open models on your own tooling →
- LinkedLinked via unknownEvaluate a model properly →
- LinkedLinked via unknownAre there good closed vs open LLM rankings? Also, are 70B–350B models actually worth it? →
- LinkedLinked via unknownNew Server Hopes to Break Through AI’s “Memory Wall” →
- LinkedLinked via unknownIEEE Rolls Out Large Language Models Virtual Training Course →
- PossiblePossibly related (embedding) · 47%Understanding Annotator Safety Policy with Interpretability - Apple Machine Learning Research →
- PossiblePossibly related (embedding) · 53%interpretml/interpret →
