repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 24d ago
alibaba/rtp-llm
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 63%OpenAI and Broadcom announce chip designed for LLM inference at scale →
- PossiblePossibly related (embedding) · 63%OpenAI and Broadcom unveil LLM-optimized inference chip →
- PossiblePossibly related (embedding) · 57%Hardware startup unveils inference accelerator →
- PossiblePossibly related (embedding) · 55%DSpark: Speculative decoding accelerates LLM inference [pdf] →
- PossiblePossibly related (embedding) · 54%DeepSeek open-sources inference optimizations with 60–85% faster generation [pdf] →
- PossiblePossibly related (embedding) · 52%H64LM: A 249M-parameter Mixture-of-Experts Transformer built from scratch in PyTorch [P] →
- PossiblePossibly related (embedding) · 45%WattGPU Predicts LLM Inference Power Without Profiling - Let's Data Science →
- PossiblePossibly related (embedding) · 50%I wrote a GGUF inferencer from scratch, AMA →
Covers
newsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsOpenAI and Broadcom unveil LLM-optimized inference chipnewsHardware startup unveils inference acceleratornewsDSpark: Speculative decoding accelerates LLM inference [pdf]newsDeepSeek open-sources inference optimizations with 60–85% faster generation [pdf]
Covers (incoming)
newsH64LM: A 249M-parameter Mixture-of-Experts Transformer built from scratch in PyTorch [P]newsWattGPU Predicts LLM Inference Power Without Profiling - Let's Data SciencenewsI wrote a GGUF inferencer from scratch, AMAnewsDisaggregated prefill and decode for LLM inference on SageMaker HyperPodnewsDisaggregated prefill and decode for LLM inference on SageMaker HyperPod | Artificial Intelligence - Amazon Web Services (AWS)newsShow HN: Goku – WASM (wllama)-powered LLM inference and model managernewsTried testing qwen 35b moe model on s26 ultra , without compromising on precision [R] ,[D]newsHetzner is working on LLM Inference
Related across the graph
newsHetzner is working on LLM InferencenewsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsDisaggregated prefill and decode for LLM inference on SageMaker HyperPod | Artificial Intelligence - Amazon Web Services (AWS)newsOpenAI and Broadcom unveil LLM-optimized inference chipnewsTried testing qwen 35b moe model on s26 ultra , without compromising on precision [R] ,[D]newsDSpark: Speculative decoding accelerates LLM inference [pdf]newsH64LM: A 249M-parameter Mixture-of-Experts Transformer built from scratch in PyTorch [P]newsDisaggregated prefill and decode for LLM inference on SageMaker HyperPodnewsDeepSeek open-sources inference optimizations with 60–85% faster generation [pdf]newsShow HN: Goku – WASM (wllama)-powered LLM inference and model managernewsHardware startup unveils inference acceleratornewsWattGPU Predicts LLM Inference Power Without Profiling - Let's Data SciencenewsI wrote a GGUF inferencer from scratch, AMA
