Prompt Compression via Activation Aggregation
Large language models process prompts by propagating activations through dozens of layers before generating a response. We ask whether the task-relevant information contained in an instruction prompt can be compressed into a single activation vector and re-injected into the model, replacing the original token sequence? We show this is achievable using a learned weighted sum of activations extracted at an intermediate layer and injected at an early layer of the target LLM. The compressed vector preserves task-relevant information, incurring an accuracy drop of under $2\%$ relative to full promp
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%AgustiPuigserver/opus-prompt-architect →
- PossiblePossibly related (embedding) · 54%chrisliu298/awesome-llm-unlearning →
- PossiblePossibly related (embedding) · 51%Transformer →
- PossiblePossibly related (embedding) · 50%Looking for feedback on a small test SLM I built completely from scratch [P] →
- LinkedLinked via arxiv author · 85%Thibaud Ardoin →
“Prompt Compression via Activation Aggregation”
- LinkedLinked via arxiv author · 85%Semira Einsele →
“Prompt Compression via Activation Aggregation”
- LinkedLinked via arxiv author · 85%Evis Bregu →
“Prompt Compression via Activation Aggregation”
- LinkedLinked via arxiv author · 85%Gerhard Wunder →
“Prompt Compression via Activation Aggregation”
