newsReddit r/LocalLLaMATrust 52 · CommunityPublished 20d agoLive · 20d ago
GLM 5.2 and ik_llama.ccp
Running GLM-5.2 (the new glm-dsa arch), Unsloth UD-Q4_K_XL, on a 4-socket Xeon E7-8880 v4 box with 1TB RAM and a single RTX 3060 12GB. ik_llama.cpp, experts on CPU (--cpu-moe), 24 attention layers on the GPU. Works great at 8k context — rock solid, ~3.7 tok/s gen. Problem: the second I raise context (anywhere past ~32–64k), generation crashes on the very first token. Fatal error in llama-sampling.cpp, and the dumped probabilities.txt shows every logit is
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%FedericoTs/quantprobe →
