Read original ↗
newsReddit r/LocalLLaMATrust 52 · CommunityPublished 1mo agoLive · 1mo ago

Hy3 (295B MoE) and NVIDIA Nemotron-Labs-Audex-30B-A3B (audio-capable 30B MoE) GGUF quants

Sharing two GGUF quant sets, both with the same treatment: imatrix quantization, KLD/PPL measured against BF16 reference logits, llama-bench throughput numbers, and all raw benchmark data included in the repos. No vibes-based "quality tested" claims — every number is reproducible from the files in the repo. 1. Hy3 — Tencent's 295B MoE (21B active) LordNeel/Hy3-GGUF

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Related across the graph