Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning
As autonomous agents are increasingly deployed across diverse operational contexts, aligning their behavior with human intent demands reward functions that remain robust to such changes rather than overfitting to any single environment. Inverse reinforcement learning (IRL) provides a principled way to infer such objectives from human feedback. However, existing analyses of optimal teaching approaches for IRL focus on single-environment, demonstration-only settings, leaving underexplored how heterogeneous feedback modalities and environment dynamics jointly constrain reward functions that gener
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%RLHF →
- PossiblePossibly related (embedding) · 48%Best practices for multi-turn reinforcement learning in Amazon SageMaker AI →
- PossiblePossibly related (embedding) · 48%jaimasih05-commits/swarm-foraging-qlearn →
- PossiblePossibly related (embedding) · 47%hud-evals/hud-python →
- PossiblePossibly related (embedding) · 46%redai-infra/Relax →
- LinkedLinked via arxiv author · 85%Ali Larian →
“Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning”
- LinkedLinked via arxiv author · 85%Qian Lin →
“Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning”
- LinkedLinked via arxiv author · 85%Chang Zong Wu →
“Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning”
