repoGitLabTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
luziyao1995/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 65%OpenAI and Broadcom announce chip designed for LLM inference at scale →
- PossiblePossibly related (embedding) · 63%OpenAI and Broadcom unveil LLM-optimized inference chip →
- PossiblePossibly related (embedding) · 55%Hardware startup unveils inference accelerator →
- PossiblePossibly related (embedding) · 54%Evaluate a model properly →
- PossiblePossibly related (embedding) · 49%I mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset) →
- PossiblePossibly related (embedding) · 52%H64LM: A 249M-parameter Mixture-of-Experts Transformer built from scratch in PyTorch [P] →
- PossiblePossibly related (embedding) · 45%Qwen 3.6 27B - VLLM Performance Benchmark Results (BF16, FP8, NVFP4) →
- PossiblePossibly related (embedding) · 59%FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference →
Covers
Related to
Covers (incoming)
newsH64LM: A 249M-parameter Mixture-of-Experts Transformer built from scratch in PyTorch [P]newsQwen 3.6 27B - VLLM Performance Benchmark Results (BF16, FP8, NVFP4)newsDisaggregated prefill and decode for LLM inference on SageMaker HyperPodnewsCloud-vLLM Benchmark Differences [R]newsTried testing qwen 35b moe model on s26 ultra , without compromising on precision [R] ,[D]
Implements (incoming)
Related across the graph
newsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsCloud-vLLM Benchmark Differences [R]newsQwen 3.6 27B - VLLM Performance Benchmark Results (BF16, FP8, NVFP4)newsOpenAI and Broadcom unveil LLM-optimized inference chipnewsI mapped which local LLMs actually fit each RAM tier, 8 to 128GB (open dataset)newsTried testing qwen 35b moe model on s26 ultra , without compromising on precision [R] ,[D]paperFreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM InferencenewsH64LM: A 249M-parameter Mixture-of-Experts Transformer built from scratch in PyTorch [P]newsDisaggregated prefill and decode for LLM inference on SageMaker HyperPodtutorialEvaluate a model properlynewsHardware startup unveils inference accelerator
