Automating Potential-based Reward Shaping with Vision Language Model Guidance
Sparse rewards are inherently challenging for reinforcement learning agents as they lack intermediate feedback to guide exploration and to correctly attribute the sparse success rewards to relevant parts of the trajectory. Naive reward shaping can induce reward hacking, yielding policies that exploit auxiliary signals instead of solving the intended task. Potential-based reward shaping (PBRS) guarantees preservation of the optimal policy set, but requires the definition of a heuristic potential function over the state space. In this work, we introduce the VLM-guided PBRS framework VLM-PBRS tha
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownvlm-starter →
- LinkedLinked via unknownGradient-based Planning for World Models at Longer Horizons →
- LinkedLinked via unknownRLHF →
- FuzzySimilar title/name (fuzzy) · 59%VioletVision-3B →
“Fuzzy title match (0.73): “Automating Potential-based Reward Shaping with Vision Langua” ≈ “VioletVision-3B””
- PossiblePossibly related (embedding) · 48%airbus/scikit-decide →
- PossiblePossibly related (embedding) · 49%pytorch/rl →
- FuzzySimilar title/name (fuzzy) · 84%pytorch/vision →
“Fuzzy title match (0.92): “Automating Potential-based Reward Shaping with Vision Langua” ≈ “pytorch/vision””
