repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 3h ago
llm-d/llm-d
Achieve state of the art inference performance with modern accelerators on Kubernetes
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%Hardware startup unveils inference accelerator →
- PossiblePossibly related (embedding) · 55%OpenAI and Broadcom announce chip designed for LLM inference at scale →
- PossiblePossibly related (embedding) · 52%OpenAI and Broadcom unveil LLM-optimized inference chip →
- PossiblePossibly related (embedding) · 45%Understanding dynamic resource allocation in Kubernetes →
- PossiblePossibly related (embedding) · 52%12 Ways to Reduce LLM Latency and Inference Costs in Production - KDnuggets →
- PossiblePossibly related (embedding) · 50%Disaggregated prefill and decode for LLM inference on SageMaker HyperPod →
Covers
Covers (incoming)
Related across the graph
newsUnderstanding dynamic resource allocation in KubernetesnewsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsOpenAI and Broadcom unveil LLM-optimized inference chipnews12 Ways to Reduce LLM Latency and Inference Costs in Production - KDnuggetsnewsDisaggregated prefill and decode for LLM inference on SageMaker HyperPodnewsHardware startup unveils inference accelerator
