LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks
Reinforcement learning (RL) on open-ended tasks compresses an LLM's rubric-based evaluation into a scalar reward, discarding rich textual feedback and conflating responses with distinct quality profiles. We propose Experiential Learning (EL), which repurposes the feedback model from an LLM-as-a-Judge into an LLM-as-a-Coach. The coach distills its assessment of each on-policy response into transferable experiential knowledge, which conditions a teacher model and is internalized by the policy through on-policy context distillation. Compared with scalar rewards, this higher-bandwidth feedback cha
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%Reinforcement Learning With Metacognitive Feedback Is Offered As A Next-Gen Way To Shape AI LLMs - Forbes →
- FuzzySimilar title/name (fuzzy) · 59%aymericdamien/TopDeepLearning →
“Fuzzy title match (0.73): “LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tas” ≈ “aymericdamien/TopDeepLearning””
- LinkedLinked via arxiv author · 85%Tianzhu Ye →
“LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks”
- LinkedLinked via arxiv author · 85%Li Dong →
“LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks”
- LinkedLinked via arxiv author · 85%Guanheng Chen →
“LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks”
- LinkedLinked via arxiv author · 85%He Zhu →
“LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks”
- LinkedLinked via arxiv author · 85%Xun Wu →
“LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks”
- LinkedLinked via arxiv author · 85%Shaohan Huang →
“LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks”
