Read original ↗
paperarXivTrust 82 · PrimaryPublished 4d agoLive · 3d ago

Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents

Combining large language models with reinforcement learning is increasingly explored, yet the theoretical status of LLM-derived reward signals is often left implicit. We formalize the hybrid LLM-planner and RL-controller architecture as a Goal-Augmented Markov Decision Process and show that when the LLM per-state progress score is used as a bounded potential function, the resulting shaping term preserves the optimal policy set even when the LLM scores are inaccurate. This guarantee is stronger than what general LLM-as-reward approaches provide. We verify the result numerically on a small MDP u

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 87%NirDiamant/GenAI_Agents

    Fuzzy title match (0.94): “Policy-Invariant Reward Shaping from LLM Feedback: A Framewo” ≈ “NirDiamant/GenAI_Agents”

  • FuzzySimilar title/name (fuzzy) · 84%Unity-Technologies/ml-agents

    Fuzzy title match (0.92): “Policy-Invariant Reward Shaping from LLM Feedback: A Framewo” ≈ “Unity-Technologies/ml-agents”

  • FuzzySimilar title/name (fuzzy) · 59%datawhalechina/hello-agents

    Fuzzy title match (0.73): “Policy-Invariant Reward Shaping from LLM Feedback: A Framewo” ≈ “datawhalechina/hello-agents”

  • FuzzySimilar title/name (fuzzy) · 59%jnMetaCode/agency-agents-zh

    Fuzzy title match (0.73): “Policy-Invariant Reward Shaping from LLM Feedback: A Framewo” ≈ “jnMetaCode/agency-agents-zh”

  • FuzzySimilar title/name (fuzzy) · 59%TauricResearch/TradingAgents

    Fuzzy title match (0.73): “Policy-Invariant Reward Shaping from LLM Feedback: A Framewo” ≈ “TauricResearch/TradingAgents”

  • LinkedLinked via arxiv author · 85%Christophe D. Hounwanou

    Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents

  • LinkedLinked via arxiv author · 85%John Emeka Eze

    Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents

  • LinkedLinked via arxiv author · 85%Yaé U. Gaba

    Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents

Implements (incoming)

authored (incoming)

Related across the graph

Topics