newsReddit r/LocalLLaMATrust 52 · CommunityPublished yesterdayLive · 18h ago
Unpopular opinion : Qwen 3.8 27b is not an overthinker
Yes it uses a ton more reasoning tokens than 3.6 did But test in on the same tasks with the other chinese models, glm 5.3, deepseek v4 flash and pro, etc it's really similar, and they are needed The reality is, we're just frustrated because our hardware do not allow most of us to have 1M context (I know that it's not supported yet) with 150 tps decode Furthermore, if you don't mind the quality drop, you can just add a reasoning budget, it wi
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%Test-Time Scaling for Small VLMs on Multilingual Visual MCQ →
- PossiblePossibly related (embedding) · 48%AarambhDevHub/aarambh-ai →
- PossiblePossibly related (embedding) · 47%DeaMoE: Efficient MoE Structure for Fast Small-Batch Decoding →
- PossiblePossibly related (embedding) · 47%pzqpzq/LSF_MDia →
