newsBAIR (Berkeley)Trust 88 · LabPublished 9mo agoLive · 1mo ago
RL without TD learning
In this post, I’ll introduce a reinforcement learning (RL) algorithm based on an “alternative” paradigm: divide and conquer . Unlike traditional methods, this algorithm is not based on temporal difference (TD) learning (which has scalability challenges ), and scales well to long-horizon tasks.
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownReinforcement Learning without Ground-Truth Solutions can Improve LLMs →
- LinkedLinked via unknownJoint Learning of Experiential Rules and Policies for Large Language Model Agents →
- LinkedLinked via unknownRegularized Reward-Punishment Reinforcement Learning →
- LinkedLinked via unknownTandem Reinforcement Learning with Verifiable Rewards →
- LinkedLinked via unknownjaimasih05-commits/swarm-foraging-qlearn →
- LinkedLinked via unknownWhich Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index →
- LinkedLinked via unknownZ-1: Efficient Reinforcement Learning for Vision-Language-Action Models →
Covers (incoming)
paperReinforcement Learning without Ground-Truth Solutions can Improve LLMspaperJoint Learning of Experiential Rules and Policies for Large Language Model AgentspaperLearning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)paperRegularized Reward-Punishment Reinforcement LearningpaperTandem Reinforcement Learning with Verifiable Rewardsrepojaimasih05-commits/swarm-foraging-qlearnpaperWhich Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal IndexpaperZ-1: Efficient Reinforcement Learning for Vision-Language-Action ModelspaperIs One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL TrainingpaperGeneralization in offline RL: The structure is more important than the amount of pessimismreporedai-infra/Relaxrepopytorch/rlrepoAgileRL/AgileRLpaperSILO: Simulation-in-the-Loop Sim-to-Real Transfer for Multi-Stage Cable Routingrepohscspring/rl-llm-nlppaperSCOPE-RL: Optimizing Reasoning Paths Before and After SuccesspaperActive Offline-to-Online Reinforcement LearningpaperWhen Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry AnalysispaperDADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement LearningrepoinclusionAI/AReno
Related across the graph
repopytorch/rlpaperGeneralization in offline RL: The structure is more important than the amount of pessimismpaperSCOPE-RL: Optimizing Reasoning Paths Before and After Successrepohscspring/rl-llm-nlprepoinclusionAI/ARenopaperWhich Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Indexrepojaimasih05-commits/swarm-foraging-qlearnrepoAgileRL/AgileRLpaperActive Offline-to-Online Reinforcement LearningpaperRegularized Reward-Punishment Reinforcement LearningpaperDADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement LearningpaperSILO: Simulation-in-the-Loop Sim-to-Real Transfer for Multi-Stage Cable RoutingpaperJoint Learning of Experiential Rules and Policies for Large Language Model AgentspaperTandem Reinforcement Learning with Verifiable RewardspaperWhen Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry AnalysispaperIs One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL TrainingpaperLearning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)paperReinforcement Learning without Ground-Truth Solutions can Improve LLMspaperZ-1: Efficient Reinforcement Learning for Vision-Language-Action Modelsreporedai-infra/Relax
