repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
dphnAI/aphrodite-engine
Large-scale LLM inference engine
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%DSpark: Speculative decoding accelerates LLM inference [pdf] →
- PossiblePossibly related (embedding) · 55%DeepSeek open-sources inference optimizations with 60–85% faster generation [pdf] →
- PossiblePossibly related (embedding) · 54%OpenAI and Broadcom announce chip designed for LLM inference at scale →
- PossiblePossibly related (embedding) · 52%OpenAI and Broadcom unveil LLM-optimized inference chip →
- PossiblePossibly related (embedding) · 52%A barebones CPU-only inference engine for Qwen 3, written from scratch in pure C →
- PossiblePossibly related (embedding) · 48%I wrote a GGUF inferencer from scratch, AMA →
- PossiblePossibly related (embedding) · 52%Disaggregated prefill and decode for LLM inference on SageMaker HyperPod →
- PossiblePossibly related (embedding) · 52%Show HN: Goku – WASM (wllama)-powered LLM inference and model manager →
Covers
newsDSpark: Speculative decoding accelerates LLM inference [pdf]newsDeepSeek open-sources inference optimizations with 60–85% faster generation [pdf]newsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsOpenAI and Broadcom unveil LLM-optimized inference chipnewsA barebones CPU-only inference engine for Qwen 3, written from scratch in pure C
Covers (incoming)
Related across the graph
newsHetzner is working on LLM InferencenewsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsOpenAI and Broadcom unveil LLM-optimized inference chipnewsDSpark: Speculative decoding accelerates LLM inference [pdf]newsDisaggregated prefill and decode for LLM inference on SageMaker HyperPodnewsDeepSeek open-sources inference optimizations with 60–85% faster generation [pdf]newsShow HN: Goku – WASM (wllama)-powered LLM inference and model managernewsA barebones CPU-only inference engine for Qwen 3, written from scratch in pure CnewsI wrote a GGUF inferencer from scratch, AMA
