Joint Learning of Experiential Rules and Policies for Large Language Model Agents
For LLM agents in multi-step interactive environments, a key challenge is to make effective use of accumulated interaction experience. Existing work has typically separated two uses of such experience: keeping it outside the model as natural-language rules for later prompting, or using trajectories and feedback to update the model parameters. The former is easy to interpret but can fall out of sync with the evolving policy; the latter improves the policy more broadly but provides only limited correction for local mistakes in sparse-reward settings. We present Joint Learning of Experiential Rul
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownRLHF →
- LinkedLinked via unknownAgentCore-8B →
- LinkedLinked via unknownagent-tools →
- LinkedLinked via unknownRL without TD learning →
- PossiblePossibly related (embedding) · 53%redai-infra/Relax →
- PossiblePossibly related (embedding) · 53%AgileRL/AgileRL →
- PossiblePossibly related (embedding) · 55%KennispuntTwente/tidyprompt →
- PossiblePossibly related (embedding) · 50%ModelOriented/DALEX →
