Read original ↗
paperarXivTrust 82 · PrimaryPublished 4d agoLive · yesterday

TTPO: Test-Time Policy Optimization

Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT). Replacing ground truth with majority-vote pseudo-labels is a natural alternative, yet it is fragile: an incorrect vote corrupts the teacher and misleads every token. We observe that this failure mode is asymmetric: rollouts that disagree with the pseudo-label are typically wrong regardless of whether the vote itself is corre

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%modular/modular

    Shared author/contributor keys: liu

  • FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow

    Shared author/contributor keys: wang

  • FuzzyOverlapping authors or contributors · 62%TauricResearch/TradingAgents

    Shared author/contributor keys: xiao

  • FuzzyOverlapping authors or contributors · 62%ray-project/ray

    Shared author/contributor keys: wang

  • LinkedLinked via arxiv author · 85%Haozhe Wang

    TTPO: Test-Time Policy Optimization

  • LinkedLinked via arxiv author · 85%Zhengxi Lu

    TTPO: Test-Time Policy Optimization

  • LinkedLinked via arxiv author · 85%Jianze Wang

    TTPO: Test-Time Policy Optimization

  • LinkedLinked via arxiv author · 85%Shangke Lv

    TTPO: Test-Time Policy Optimization

Implements (incoming)

authored (incoming)

Related across the graph

Topics