glossary termAngestromTrust 60Published 2mo agoLive · 2mo ago
RLHF
Reinforcement learning from human feedback — tuning a model toward preferred answers.
Reinforcement learning from human feedback — tuning a model toward preferred answers. Reinforcement learning from human feedback — tuning a model toward preferred answers.
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownReinforcement Learning without Ground-Truth Solutions can Improve LLMs →
- LinkedLinked via unknownJoint Learning of Experiential Rules and Policies for Large Language Model Agents →
- LinkedLinked via unknownAutomating Potential-based Reward Shaping with Vision Language Model Guidance →
- LinkedLinked via unknownTandem Reinforcement Learning with Verifiable Rewards →
- LinkedLinked via unknownHierarchical Experimentalist Agents →
- LinkedLinked via unknownWhich Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index →
- LinkedLinked via unknownIs One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training →
Related to (incoming)
paperReinforcement Learning without Ground-Truth Solutions can Improve LLMspaperJoint Learning of Experiential Rules and Policies for Large Language Model AgentspaperAutomating Potential-based Reward Shaping with Vision Language Model GuidancepaperTandem Reinforcement Learning with Verifiable RewardspaperHierarchical Experimentalist AgentspaperPessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning ModelspaperWhich Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal IndexpaperIs One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL TrainingpaperGeneralization in offline RL: The structure is more important than the amount of pessimismpaperAttention Limited Reward LearningpaperMulti-Modal, Multi-Environment Machine Teaching for Robust Reward LearningpaperSCOPE-RL: Optimizing Reasoning Paths Before and After SuccesspaperFrom Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence EstimationpaperUnderstanding Reasoning from Pretraining to Post-TrainingpaperBeyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric RewardspaperISO: An RLVR-Native Optimization StackpaperThe Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually Works
Related across the graph
paperGeneralization in offline RL: The structure is more important than the amount of pessimismpaperSCOPE-RL: Optimizing Reasoning Paths Before and After SuccesspaperISO: An RLVR-Native Optimization StackpaperBeyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric RewardspaperWhich Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal IndexpaperFrom Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence EstimationpaperMulti-Modal, Multi-Environment Machine Teaching for Robust Reward LearningpaperAutomating Potential-based Reward Shaping with Vision Language Model GuidancepaperPessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning ModelspaperJoint Learning of Experiential Rules and Policies for Large Language Model AgentspaperTandem Reinforcement Learning with Verifiable RewardspaperAttention Limited Reward LearningpaperUnderstanding Reasoning from Pretraining to Post-TrainingpaperIs One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL TrainingpaperReinforcement Learning without Ground-Truth Solutions can Improve LLMspaperHierarchical Experimentalist AgentspaperThe Dark Room in the Reward Channel: Dense Prediction Rewards Collapse GRPO-Trained LLM Agents -- and What Actually Works
