The Rise of Verbal Reinforcement Learning
Natural language is emerging as a primary feedback channel for improving language agents, capable of conveying intent, preferences, and causal structure in forms interpretable by both humans and modern language models. We call this paradigm Verbal Reinforcement Learning (VRL) and offer the first unified account of it. We organize the field around a single axis, \textit{when} verbal feedback takes effect in an agent's lifecycle and \textit{what} it modifies, yielding three pillars: (1) \textbf{Language as Grounding Signal}, where language defines the task itself by specifying goals, states, and
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 54%RLHF →
- FuzzySimilar title/name (fuzzy) · 84%amitness/learning →
“Fuzzy title match (0.92): “The Rise of Verbal Reinforcement Learning” ≈ “amitness/learning””
- FuzzyOverlapping authors or contributors · 62%ultralytics/yolov5 →
“Shared author/contributor keys: sharma”
- FuzzyOverlapping authors or contributors · 62%deepfakes/faceswap →
“Shared author/contributor keys: sharma”
- FuzzySimilar title/name (fuzzy) · 59%MathFoundationRL/Book-Mathematical-Foundation-of-Reinforcement-Learning →
“Fuzzy title match (0.73): “The Rise of Verbal Reinforcement Learning” ≈ “MathFoundationRL/Book-Mathematical-Foundation-of-Reinforceme””
- LinkedLinked via arxiv author · 85%Kshitij Tayal →
“The Rise of Verbal Reinforcement Learning”
- LinkedLinked via arxiv author · 85%Arun Sharma →
“The Rise of Verbal Reinforcement Learning”
- LinkedLinked via arxiv author · 85%Genta Indra Winata →
“The Rise of Verbal Reinforcement Learning”
