Mechanistic Interpretability
3 items across the graph — tagged with Mechanistic Interpretability.
From the graph · 3
repo
ninjahawk/Subtext
→repoTo know what models don't say out loud.
moudrkat/brainscope
→repoOpenAI-compatible server with a live view into any HF model's residual stream. pip install brainscope
moudrkat/steeropathy
→Agents that talk through model internals — activations & J-space — instead of text. No words pass between them. The lab on top of brainscope + hidden-directions…
