GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning
Reinforcement learning (RL), particularly RL with Verifiable Rewards (RLVR), has recently emerged as a central paradigm for enhancing large language models' (LLMs) reasoning abilities, demonstrating remarkable effectiveness across reasoning tasks. Recent studies suggest that high-entropy tokens play an exceptionally important role in model training, since training with only the highest 20% entropy tokens yields significant performance gains. However, why such high-entropy tokens are beneficial remains insufficiently understood. In this work, we find that although high-entropy tokens within one
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%Qwen 3.8 27B Overthinking, It has to be done, it has to be overthinking to punch Opus 4.6 →
- LinkedLinked via arxiv author · 85%Outongyi Lv →
“GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning”
- LinkedLinked via arxiv author · 85%Yuanwei Zhang →
“GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning”
- LinkedLinked via arxiv author · 85%Xiaoqun Zhang →
“GMTS: Gradient Magnitude-based Token Selection Improves RLVR Training for LLM Reasoning”
