Read original ↗
newsReddit r/LocalLLaMATrust 52 · CommunityPublished 28d agoLive · 28d ago

543 tok/s single-request Qwen3.6-35B-A3B on one RTX 5090 over a 65K-token decode

An example TL;DR I have open-sourced NInfer , a from-scratch C++/CUDA inference engine currently specialized for two exact Qwen3.6 checkpoints on a single RTX 5090. Both the engine and the converted model artifacts are publicly available: Github : https://github.

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Covers (incoming)

Related across the graph