newsReddit r/LocalLLaMATrust 58 · CommunityPublished 1mo agoLive · 1mo ago
A barebones CPU-only inference engine for Qwen 3, written from scratch in pure C
TL;DR: The (very messy) code and writeups can be found at https://github.com/jakint0sh/qwen3-engine Read the README for instructions on how to get started. And for those who just want a bulleted list: - Inference engine for Qwen 3 sizes 4B and below - Written from scratch in pure C - No dependencies except libc, libm, and cJSON (and OpenMP if compiled with parallelization) - Loads directly fro
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%jmaczan/tiny-vllm →
- PossiblePossibly related (embedding) · 52%openinfer-project/openinfer →
- PossiblePossibly related (embedding) · 46%quic/efficient-transformers →
- PossiblePossibly related (embedding) · 52%dphnAI/aphrodite-engine →
- PossiblePossibly related (embedding) · 50%mohitsoni48/TurboLLM →
- PossiblePossibly related (embedding) · 48%TilelliLab/atome-lm →
- PossiblePossibly related (embedding) · 49%QodeXcli/QodeX →
- PossiblePossibly related (embedding) · 60%marzukia/qMLX →
