Read original ↗
paperarXivTrust 82 · PrimaryPublished 6d agoLive · 3d ago

More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It

Power Sampling sharpens a language model's distribution over complete generation trajectories, offering a verifier-free way to improve reasoning at inference time. It also has the potential to serve as a general-purpose front end for a broad range of downstream sampling methods. However, we uncover a striking paradox: Power Sampling can drive more probability mass toward correct trajectories while degrading the downstream inference it is intended to enhance. Using self-consistency as a representative case, we observe accuracy drops of up to 18.5 percentage points across models and reasoning be

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%google-research/google-research

    Shared author/contributor keys: sun

  • LinkedLinked via arxiv author · 85%Haohui Yang

    More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It

  • LinkedLinked via arxiv author · 85%Jiaxing Sun

    More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It

  • LinkedLinked via arxiv author · 85%Xiujun Ma

    More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It

Implements (incoming)

authored (incoming)

Related across the graph

Topics