Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index
Reinforcement learning (RL) has become a powerful tool for propelling Large Language Models (LLMs) beyond imitation-based training towards more robust reasoning capabilities. Among existing approaches, RL with Verifiable Rewards (RLVR) has emerged as a pivotal paradigm for advancing LLM reasoning. Despite its empirical success, recent studies have offered different insights. One line of inquiry advocates prioritizing high-entropy token positions during training, while another perspective cautions against allowing low-probability tokens to dominate gradient updates. Notably, although high-entro
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownRLHF →
- LinkedLinked via unknownRL without TD learning →
- PossiblePossibly related (embedding) · 52%rllm-org/rllm →
- PossiblePossibly related (embedding) · 65%hscspring/rl-llm-nlp →
- FuzzySimilar title/name (fuzzy) · 59%run-llama/llama_index →
“Fuzzy title match (0.73): “Which Tokens Matter? Adaptive Token Selection for RLVR with ” ≈ “run-llama/llama_index””
- FuzzySimilar title/name (fuzzy) · 59%VectifyAI/PageIndex →
“Fuzzy title match (0.73): “Which Tokens Matter? Adaptive Token Selection for RLVR with ” ≈ “VectifyAI/PageIndex””
- PossiblePossibly related (embedding) · 56%Paper claims RL for reasoning only changes 1-3% of tokens, and they replicate the gains without RL at ~1000x less compute →
