repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
jmaczan/tiny-vllm
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 61%OpenAI and Broadcom announce chip designed for LLM inference at scale →
- PossiblePossibly related (embedding) · 59%OpenAI and Broadcom unveil LLM-optimized inference chip →
- PossiblePossibly related (embedding) · 57%Hardware startup unveils inference accelerator →
- PossiblePossibly related (embedding) · 54%A barebones CPU-only inference engine for Qwen 3, written from scratch in pure C →
- PossiblePossibly related (embedding) · 53%I mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset) →
- PossiblePossibly related (embedding) · 49%Gemma 4 12B - MLX Kernel →
- PossiblePossibly related (embedding) · 56%We'll benchmark an Open weights LLM on any GPU you choose — drop your model + hardware and we'll run it. [D] →
Covers
newsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsOpenAI and Broadcom unveil LLM-optimized inference chipnewsHardware startup unveils inference acceleratornewsA barebones CPU-only inference engine for Qwen 3, written from scratch in pure CnewsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)
Covers (incoming)
Related across the graph
newsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsOpenAI and Broadcom unveil LLM-optimized inference chipnewsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsWe'll benchmark an Open weights LLM on any GPU you choose — drop your model + hardware and we'll run it. [D]newsHardware startup unveils inference acceleratornewsA barebones CPU-only inference engine for Qwen 3, written from scratch in pure CnewsGemma 4 12B - MLX Kernel
