Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference
Small language models are usually built like large ones and then squeezed onto a CPU afterwards. We did the opposite: we fixed the target first, one user, one token at a time, 4-bit weights, ordinary CPU, and chose the architecture to suit it. The result keeps full attention in only 6 of its 18 blocks. The other 12 use short convolutions whose memory is two timesteps wide no matter how long the conversation gets, so two thirds of the network never re-reads a growing cache. Trained from scratch on 59.9B tokens, the model scores 47.31 on a five-task benchmark against a bar of 42.20 that was fi
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%New Server Hopes to Break Through AI’s “Memory Wall” →
- PossiblePossibly related (embedding) · 56%Breakthrough in long-context efficiency announced →
- PossiblePossibly related (embedding) · 53%Introducing DWARF-55M-Base →
- PossiblePossibly related (embedding) · 52%deepseek-v4-flash-0731 - surprisingly usable →
- FuzzySimilar title/name (fuzzy) · 84%xorbitsai/inference →
“Fuzzy title match (0.92): “Daedalus-150M: A Convolution-Attention Hybrid Designed for C” ≈ “xorbitsai/inference””
- LinkedLinked via arxiv author · 85%Christos Koutsiaris →
“Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference”
