Read original ↗
paperarXivTrust 82 · PrimaryPublished 4d agoLive · yesterday

SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization

Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from collecting such behavioral evidence at scale. We introduce SAEVerbalizer, a framework that injects SAE decoder directions into an LLM's representations and fine-tunes the LLM's downstream layers to generate natural-language explanations of the injected features. Once trained, the r

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%mudler/LocalAI

    Shared author/contributor keys: guo

  • FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow

    Shared author/contributor keys: wang

  • FuzzyOverlapping authors or contributors · 62%modular/modular

    Shared author/contributor keys: liu

  • FuzzyOverlapping authors or contributors · 62%ray-project/ray

    Shared author/contributor keys: wang

  • LinkedLinked via arxiv author · 85%Weihan Meng

    SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization

  • LinkedLinked via arxiv author · 85%Hongzhu Guo

    SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization

  • LinkedLinked via arxiv author · 85%Yi Jing

    SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization

  • LinkedLinked via arxiv author · 85%Dewen Liu

    SAEVerbalizer: Generating Explanations for Sparse Autoencoder Features via Representation Verbalization

Implements (incoming)

authored (incoming)

Related across the graph

Topics