repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 4h ago
rllm-org/rllm
Democratizing Reinforcement Learning for LLMs
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 61%Reinforcement Learning without Ground-Truth Solutions can Improve LLMs →
- PossiblePossibly related (embedding) · 52%Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index →
- PossiblePossibly related (embedding) · 52%Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs →
- PossiblePossibly related (embedding) · 51%Tandem Reinforcement Learning with Verifiable Rewards →
- PossiblePossibly related (embedding) · 51%Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training →
- PossiblePossibly related (embedding) · 45%ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning →
- PossiblePossibly related (embedding) · 53%Generalization in offline RL: The structure is more important than the amount of pessimism →
- PossiblePossibly related (embedding) · 54%DecompRL: Solving Harder Problems by Learning Modular Code Generation →
Implements
paperReinforcement Learning without Ground-Truth Solutions can Improve LLMspaperWhich Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal IndexpaperTriadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMspaperTandem Reinforcement Learning with Verifiable RewardspaperIs One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training
Implements (incoming)
paperART for Diffusion Sampling: Continuous-Time Control and Actor-Critic LearningpaperGeneralization in offline RL: The structure is more important than the amount of pessimismpaperDecompRL: Solving Harder Problems by Learning Modular Code GenerationpaperA Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-TrainingpaperCompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon AgentspaperTREK: Distill to Explore, Reinforce to RefinepaperWeak-to-Strong Generalization via Direct On-Policy DistillationpaperImproving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function DesignpaperInformation Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM AgentspaperSCOPE-RL: Optimizing Reasoning Paths Before and After SuccesspaperRing-Zero: Scaling Zero RL to a Trillion Parameters for Emergent ReasoningpaperFrom Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation
Covers (incoming)
newsWho've told you that distributed training is impossible? Democratizing AI: The Psyche Network ArchitecturenewsThe Little Book of Reinforcement LearningnewsIt only took 200 update steps to flip Qwen2.5-7B-Instruct from denying sentience to developing a robust identity of being a "sentient machine" [P]
Related to (incoming)
Related across the graph
paperTriadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMsmodelzai-org/chatglm-6bpaperGeneralization in offline RL: The structure is more important than the amount of pessimismpaperSCOPE-RL: Optimizing Reasoning Paths Before and After Successmodelzai-org/GLM-5.2paperImproving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function DesignpaperWhich Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal IndexpaperInformation Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM AgentspaperFrom Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence EstimationpaperDecompRL: Solving Harder Problems by Learning Modular Code GenerationpaperTandem Reinforcement Learning with Verifiable RewardspaperWeak-to-Strong Generalization via Direct On-Policy DistillationnewsThe Little Book of Reinforcement LearningpaperCompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon AgentspaperRing-Zero: Scaling Zero RL to a Trillion Parameters for Emergent ReasoningpaperIs One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL TrainingpaperA Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-TrainingnewsIt only took 200 update steps to flip Qwen2.5-7B-Instruct from denying sentience to developing a robust identity of being a "sentient machine" [P]paperReinforcement Learning without Ground-Truth Solutions can Improve LLMspaperTREK: Distill to Explore, Reinforce to RefinepaperART for Diffusion Sampling: Continuous-Time Control and Actor-Critic LearningnewsWho've told you that distributed training is impossible? Democratizing AI: The Psyche Network Architecture
