TTPO: Test-Time Policy Optimization
Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replacing ground truth with majority-vote pseudo-labels is a natural alternative, yet it is fragile: an incorrect vote corrupts the teacher and misleads every token. We observe that this failure mode is asymmetric: rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is corre
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%TauricResearch/TradingAgents →
“Shared author/contributor keys: xiao”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- LinkedLinked via arxiv author · 85%Haozhe Wang →
“TTPO: Test-Time Policy Optimization”
- LinkedLinked via arxiv author · 85%Zhengxi Lu →
“TTPO: Test-Time Policy Optimization”
- LinkedLinked via arxiv author · 85%Jianze Wang →
“TTPO: Test-Time Policy Optimization”
- LinkedLinked via arxiv author · 85%Shangke Lv →
“TTPO: Test-Time Policy Optimization”
