repoGitHubTrust 82 · PrimaryPublished 2d agoLive · yesterday
pegainfer-project/pegainfer
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%543 tok/s single-request Qwen3.6-35B-A3B on one RTX 5090 over a 65K-token decode →
- PossiblePossibly related (embedding) · 54%High-Performance MoE Inference: Qwen3.6–35B-A3B on an AI PC with OpenVINO - Medium →
- PossiblePossibly related (embedding) · 53%Chinese startup Moonshot AI unveils Kimi model it says rivals OpenAI, Anthropic - CNBC →
- PossiblePossibly related (embedding) · 52%China's Moonshot AI claims Kimi K3 can rival OpenAI and Anthropic →
- PossiblePossibly related (embedding) · 52%A barebones CPU-only inference engine for Qwen 3, written from scratch in pure C →
Covers
news543 tok/s single-request Qwen3.6-35B-A3B on one RTX 5090 over a 65K-token decodenewsHigh-Performance MoE Inference: Qwen3.6–35B-A3B on an AI PC with OpenVINO - MediumnewsChinese startup Moonshot AI unveils Kimi model it says rivals OpenAI, Anthropic - CNBCnewsChina's Moonshot AI claims Kimi K3 can rival OpenAI and AnthropicnewsA barebones CPU-only inference engine for Qwen 3, written from scratch in pure C
Related across the graph
news543 tok/s single-request Qwen3.6-35B-A3B on one RTX 5090 over a 65K-token decodenewsChina's Moonshot AI claims Kimi K3 can rival OpenAI and AnthropicnewsHigh-Performance MoE Inference: Qwen3.6–35B-A3B on an AI PC with OpenVINO - MediumnewsChinese startup Moonshot AI unveils Kimi model it says rivals OpenAI, Anthropic - CNBCnewsA barebones CPU-only inference engine for Qwen 3, written from scratch in pure C
