repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 23h ago
defilantech/LLMKube
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 61%Understanding dynamic resource allocation in Kubernetes →
- PossiblePossibly related (embedding) · 55%OpenAI and Broadcom announce chip designed for LLM inference at scale →
- PossiblePossibly related (embedding) · 54%OpenAI and Broadcom unveil LLM-optimized inference chip →
- PossiblePossibly related (embedding) · 52%How're you deploying LLMs in production now-a-days? What's the best and most affordable way? [D] →
- PossiblePossibly related (embedding) · 47%Run NVIDIA Nemotron and OpenAI GPT OSS models on Amazon Bedrock in AWS GovCloud (US) →
- PossiblePossibly related (embedding) · 45%Self-hosted GitHub Actions runners on Lambda MicroVMs →
- PossiblePossibly related (embedding) · 70%Running a self-hosted LLM in Kubernetes with vLLM →
Covers
newsUnderstanding dynamic resource allocation in KubernetesnewsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsOpenAI and Broadcom unveil LLM-optimized inference chipnewsHow're you deploying LLMs in production now-a-days? What's the best and most affordable way? [D]newsRun NVIDIA Nemotron and OpenAI GPT OSS models on Amazon Bedrock in AWS GovCloud (US)
Covers (incoming)
Related across the graph
newsUnderstanding dynamic resource allocation in KubernetesnewsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsRun NVIDIA Nemotron and OpenAI GPT OSS models on Amazon Bedrock in AWS GovCloud (US)newsRunning a self-hosted LLM in Kubernetes with vLLMnewsOpenAI and Broadcom unveil LLM-optimized inference chipnewsSelf-hosted GitHub Actions runners on Lambda MicroVMsnewsHow're you deploying LLMs in production now-a-days? What's the best and most affordable way? [D]
