Read original ↗
paperarXivTrust 82 · PrimaryPublished 16d agoLive · 14d ago

Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning

Generative pretraining established reusable task representations; later work on language-based task conditioning and in-context learning showed that a fixed model could adapt its behavior from instructions and demonstrations. Policy Iteration with Human Feedback (PIHF) builds on this development and the recurrent evaluate-and-improve structure of generalized policy iteration. PIHF uses a pretrained language model as its execution substrate and moves persistent revision to a versioned natural-language policy and tool set. A language-model critic and clinical expert review complete-panel reasoni

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 53%RLHF
  • FuzzyOverlapping authors or contributors · 62%open-webui/open-webui

    Shared author/contributor keys: nguyen

  • FuzzySimilar title/name (fuzzy) · 59%aymericdamien/TopDeepLearning

    Fuzzy title match (0.73): “Policy Iteration with Human Feedback: Bringing Post-Training” ≈ “aymericdamien/TopDeepLearning”

  • LinkedLinked via arxiv author · 85%Minh-Ha Nguyen

    Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning

  • LinkedLinked via arxiv author · 85%Cathy Shyr

    Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning

Related to

Implements (incoming)

authored (incoming)

Related across the graph

Topics