repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
artalis-io/bitnet.c
Minimal, zero-dependency LLM inference in pure C11. CPU-first with NEON/AVX2 SIMD. Flash MoE (pread + LRU expert cache). TurboQuant 3-bit KV compression (8.9x less memory per session). 20+ GGUF quant formats. Compiles to WASM.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 60%A barebones CPU-only inference engine for Qwen 3, written from scratch in pure C →
- PossiblePossibly related (embedding) · 51%OpenAI and Broadcom announce chip designed for LLM inference at scale →
- PossiblePossibly related (embedding) · 50%BiSCo-LLM: Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model Compression →
- PossiblePossibly related (embedding) · 51%Recent llama.cpp updates for SYCL/Intel →
- PossiblePossibly related (embedding) · 47%Show HN: misa77 - a codec that decodes 2x faster than LZ4 (at better ratios) →
- PossiblePossibly related (embedding) · 50%DeepSeek V4 Flash | IQ3_XXS-AS & IQ2_S Bench | mainline b10064 vs fairydreaming | 1xRTX 3090 + 128GB DDR4 | 250PP/11TG on 50K CTX →
Covers
Implements
Covers (incoming)
Related across the graph
paperBiSCo-LLM: Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model CompressionnewsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsRecent llama.cpp updates for SYCL/IntelnewsShow HN: misa77 - a codec that decodes 2x faster than LZ4 (at better ratios)newsA barebones CPU-only inference engine for Qwen 3, written from scratch in pure CnewsDeepSeek V4 Flash | IQ3_XXS-AS & IQ2_S Bench | mainline b10064 vs fairydreaming | 1xRTX 3090 + 128GB DDR4 | 250PP/11TG on 50K CTX
