From Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASP
A theoretical understanding of Transformers is crucial to better understand the capacities and limitations of large language models (LLMs). There is much work analyzing the expressivity of attention-based models. By proposing handcrafted weights or using computational complexity arguments, a large amount of past theoretical works have sought to characterize which tasks are and which are not in the hypothesis class of Transformer models. However, little work investigates the learnability of such solutions. In this work, we make progress towards this goal. Inspired by recent loss landscape analy
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 64%Transformer →
- PossiblePossibly related (embedding) · 57%Furyton/awesome-language-model-analysis →
- PossiblePossibly related (embedding) · 57%engineering87/llm-atlas →
- PossiblePossibly related (embedding) · 56%Transformers in Deep Learning: How Self-Attention Changed Modern AI - Snowflake →
- PossiblePossibly related (embedding) · 56%chrisliu298/awesome-llm-unlearning →
- LinkedLinked via arxiv author · 85%Michael Rizvi-Martel →
“From Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASP”
- LinkedLinked via arxiv author · 85%Satwik Bhattamishra →
“From Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASP”
- LinkedLinked via arxiv author · 85%Guillaume Rabusseau →
“From Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASP”
