repoGitHubTrust 82 · PrimaryPublished 21d agoLive · 21d ago
arizqi/cpubrrr
Frontier-class LLM inference on a laptop CPU — gpt-oss:20b at ~110 tok/s on Apple M4 Max, 7.5x llama.cpp, no GPU. From-scratch NEON/SME kernels in Rust.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 58%Ask HN: MacBook vs. Dedicated GPU for LLM →
- PossiblePossibly related (embedding) · 55%[Paper] Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices →
- PossiblePossibly related (embedding) · 54%OpenAI and Broadcom announce chip designed for LLM inference at scale →
- PossiblePossibly related (embedding) · 53%MOREH Showcases High-Performance LLM Inference on AMD GPUs at AMD Advancing AI 2026 - bastillepost.com →
- PossiblePossibly related (embedding) · 53%Gemma 4 12B - MLX Kernel →
Covers
newsAsk HN: MacBook vs. Dedicated GPU for LLMnews[Paper] Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer DevicesnewsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsMOREH Showcases High-Performance LLM Inference on AMD GPUs at AMD Advancing AI 2026 - bastillepost.comnewsGemma 4 12B - MLX Kernel
Related across the graph
newsOpenAI and Broadcom announce chip designed for LLM inference at scalenews[Paper] Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer DevicesnewsAsk HN: MacBook vs. Dedicated GPU for LLMnewsGemma 4 12B - MLX KernelnewsMOREH Showcases High-Performance LLM Inference on AMD GPUs at AMD Advancing AI 2026 - bastillepost.com
