newsReddit r/LocalLLaMATrust 52 · CommunityPublished 26d agoLive · 26d ago
Introducing DWARF-55M-Base
Finally after months of research, the very first model made from the DWARF architecture is available for folks to check out and experiment with! DWARF is a nearly all-sparse attention architecture that uses 9 Dynamic Sparse Query-Gather (DSQG) layers as the backbone for transportation and a single full causal attention layer at 25% layer depth. For example, if there are 32 layers in the model it requires only a single full causal attention layer at L7 as
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%Morphing into Hybrid Attention Models →
- PossiblePossibly related (embedding) · 51%Attention →
- PossiblePossibly related (embedding) · 50%FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference →
- PossiblePossibly related (embedding) · 47%Understanding Large Language Models →
- PossiblePossibly related (embedding) · 46%Long-Context Fine-Tuning with Limited VRAM →
