newsReddit r/LocalLLaMATrust 52 · CommunityPublished yesterdayLive · 20h ago
ExLlamaV3 is underrated
I moght get shit on for posting this but, I feel like i don't see this being talked enough and it feels like such a waste of a good piece of software. Exl3 is incredible, albeit only if you have NVIDIA cards I think? Exl3 quants are higher quality for its size, much lower KLD metrics, faster, all compared to llama.cpp just from a few personal sets of tests I like to give my local models (these are not benchmarks). From what I have been reading CPU MoE of
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 49%AmusementClub/vs-mlrt →
- PossiblePossibly related (embedding) · 48%AMD-AGI/Magpie →
- PossiblePossibly related (embedding) · 47%open-edge-platform/geti →
- PossiblePossibly related (embedding) · 47%TristanBilot/mlx-benchmark →
- PossiblePossibly related (embedding) · 46%hogeheer499-commits/strix-halo-guide →
