Reinforcement Learning without Ground-Truth Solutions can Improve LLMs
Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically rely on ground-truth answers to assign rewards, limiting their applicability to tasks where the ground-truth solution is unknown. We introduce a \textbf{R}anking-\textbf{i}nduced \textbf{VER}ifiable framework (RiVER) that trains LLMs on score-based optimization tasks without ground-truth solutions, using deterministic execution feedback as continuous-valued supervision. When applying group-relative RL to such continuous rewards, we identify two key challenges: \emph{scale dominance}, where uncalibrated score magn
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownRL without TD learning →
- LinkedLinked via unknownRLHF →
- PossiblePossibly related (embedding) · 50%teilomillet/retrain →
- PossiblePossibly related (embedding) · 61%rllm-org/rllm →
- PossiblePossibly related (embedding) · 60%hscspring/rl-llm-nlp →
- PossiblePossibly related (embedding) · 50%Maxims for machines: operationalizing Kant’s universal law via inverse reinforcement learning - Springer Nature Link →
