newsReddit r/LocalLLaMATrust 52 · CommunityPublished 1mo agoLive · 1mo ago
Hy3 (295B MoE) and NVIDIA Nemotron-Labs-Audex-30B-A3B (audio-capable 30B MoE) GGUF quants
Sharing two GGUF quant sets, both with the same treatment: imatrix quantization, KLD/PPL measured against BF16 reference logits, llama-bench throughput numbers, and all raw benchmark data included in the repos. No vibes-based "quality tested" claims — every number is reproducible from the files in the repo. 1. Hy3 — Tencent's 295B MoE (21B active) LordNeel/Hy3-GGUF
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 47%notwitcheer/llm-bench-rig →
- PossiblePossibly related (embedding) · 45%Erwan923/gpu-infra-lab →
- PossiblePossibly related (embedding) · 45%SemiAnalysisAI/InferenceX →
