repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
marzukia/qMLX
qMLX: Custom inference engine for Qwen 3.5 122B on Apple Silicon, extending MLX with hybrid attention support, SSD-backed KV cache, and RYS layer duplication for efficient large-model serving on consumer hardware.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 60%A barebones CPU-only inference engine for Qwen 3, written from scratch in pure C →
- PossiblePossibly related (embedding) · 54%Hardware startup unveils inference accelerator →
- PossiblePossibly related (embedding) · 53%OpenAI and Broadcom unveil LLM-optimized inference chip →
- PossiblePossibly related (embedding) · 51%Gemma 4 12B - MLX Kernel →
- PossiblePossibly related (embedding) · 51%OpenAI and Broadcom announce chip designed for LLM inference at scale →
- PossiblePossibly related (embedding) · 55%High-Performance MoE Inference: Qwen3.6–35B-A3B on an AI PC with OpenVINO - Medium →
- PossiblePossibly related (embedding) · 63%SOTA Apple Silicon Inference (August 15, 2026) →
- PossiblePossibly related (embedding) · 61%Conjure cash with old Macs by linking them to AI inference Borg →
Covers
Covers (incoming)
Related across the graph
newsOpenAI and Broadcom announce chip designed for LLM inference at scalenewsOpenAI and Broadcom unveil LLM-optimized inference chipnewsMac Studio M5 Max Cost AnalysisnewsSOTA Apple Silicon Inference (August 15, 2026)newsConjure cash with old Macs by linking them to AI inference BorgnewsHigh-Performance MoE Inference: Qwen3.6–35B-A3B on an AI PC with OpenVINO - MediumnewsHardware startup unveils inference acceleratornewsA barebones CPU-only inference engine for Qwen 3, written from scratch in pure CnewsGemma 4 12B - MLX Kernel
