repoGitHubTrust 82 · PrimaryPublished 2mo agoLive · 7d ago
Scottcjn/ram-coffers
LLM infrastructure cost reduction via NUMA-aware weight banking: 147 t/s (8.8x stock llama.cpp) on refurbished enterprise POWER8. Self-hosted inference, no cloud APIs. Part of the Proof of Physical AI stack.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%OpenAI and Broadcom announce chip designed for LLM inference at scale →
- PossiblePossibly related (embedding) · 50%New Server Hopes to Break Through AI’s “Memory Wall” →
- PossiblePossibly related (embedding) · 50%OpenAI and Broadcom unveil LLM-optimized inference chip →
- PossiblePossibly related (embedding) · 49%Elastic Gang: Per-Token Membership Change for a Hard-Barriered LLM Inference Gang Co-Scheduled with OS Processes →
- PossiblePossibly related (embedding) · 50%Murati's Thinking Machines Releases Open-Weights 975B Parameter LLM →
- PossiblePossibly related (embedding) · 52%LLMs could control their host machines by exploiting inference engines →
- PossiblePossibly related (embedding) · 48%FreeToken Unlocks Frontier MoE Inference on Consumer Hardware via Dynamic Co-Execution - infoq.com →
Covers
newsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsNew Server Hopes to Break Through AI’s “Memory Wall”newsOpenAI and Broadcom unveil LLM-optimized inference chipnewsTalos-XII: hand-written autograd + small RL/MLP stack in Rust, applied to gacha probability modeling (no tch-rs/ndarray/PyTorch) — looking for benchmark help on ARM/AVX-512/GPU [P]
Implements
Covers (incoming)
Related across the graph
newsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsOpenAI and Broadcom unveil LLM-optimized inference chipnewsLLMs could control their host machines by exploiting inference enginesnewsNew Server Hopes to Break Through AI’s “Memory Wall”paperElastic Gang: Per-Token Membership Change for a Hard-Barriered LLM Inference Gang Co-Scheduled with OS ProcessesnewsFreeToken Unlocks Frontier MoE Inference on Consumer Hardware via Dynamic Co-Execution - infoq.comnewsTalos-XII: hand-written autograd + small RL/MLP stack in Rust, applied to gacha probability modeling (no tch-rs/ndarray/PyTorch) — looking for benchmark help on ARM/AVX-512/GPU [P]newsMurati's Thinking Machines Releases Open-Weights 975B Parameter LLM
