Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning
Generative pretraining established reusable task representations; later work on language-based task conditioning and in-context learning showed that a fixed model could adapt its behavior from instructions and demonstrations. Policy Iteration with Human Feedback (PIHF) builds on this development and the recurrent evaluate-and-improve structure of generalized policy iteration. PIHF uses a pretrained language model as its execution substrate and moves persistent revision to a versioned natural-language policy and tool set. A language-model critic and clinical expert review complete-panel reasoni
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 53%RLHF →
- FuzzyOverlapping authors or contributors · 62%open-webui/open-webui →
“Shared author/contributor keys: nguyen”
- FuzzySimilar title/name (fuzzy) · 59%aymericdamien/TopDeepLearning →
“Fuzzy title match (0.73): “Policy Iteration with Human Feedback: Bringing Post-Training” ≈ “aymericdamien/TopDeepLearning””
- LinkedLinked via arxiv author · 85%Minh-Ha Nguyen →
“Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning”
- LinkedLinked via arxiv author · 85%Cathy Shyr →
“Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning”
