newsReddit r/LocalLLaMATrust 52 · CommunityPublished yesterdayLive · yesterday
Paper claims RL for reasoning only changes 1-3% of tokens, and they replicate the gains without RL at ~1000x less compute
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index →
- PossiblePossibly related (embedding) · 49%wbopan/flashtrace →
- PossiblePossibly related (embedding) · 48%Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning →
- PossiblePossibly related (embedding) · 47%The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation →
- PossiblePossibly related (embedding) · 46%Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations →
- PossiblePossibly related (embedding) · 46%SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning →
Covers
paperWhich Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Indexrepowbopan/flashtracepaperToken-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement LearningpaperThe Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine TranslationpaperTrain the Model, Not the Reader: Decodability Supervision for Verifiable Activation ExplanationspaperSimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning
Related across the graph
paperToken-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement LearningpaperWhich Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal IndexpaperTrain the Model, Not the Reader: Decodability Supervision for Verifiable Activation ExplanationspaperSimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context ReasoningpaperThe Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translationrepowbopan/flashtrace
