Read original ↗
paperarXivTrust 82 · PrimaryPublished 22h agoLive · 54m ago

STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment

Multi-preference alignment is often framed as scalarization: combine reward dimensions, then optimize. This leaves a temporal decision underspecified: when should each preference dimension enter policy optimization? We propose \methodname, a stability-guided active-set controller for controlled objective admission. \methodname starts from a small active set, retains admitted objectives, and expands when reward-deviation gates indicate low recent deviation or a patience budget is exhausted. A probing phase estimates a hard-to-easy order, and adaptive weighting emphasizes underperforming active

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%ray-project/ray

    Shared author/contributor keys: wang

  • FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow

    Shared author/contributor keys: wang

  • FuzzyOverlapping authors or contributors · 62%Zeyi-Lin/HivisionIDPhotos

    Shared author/contributor keys: lin

  • FuzzyOverlapping authors or contributors · 62%hiyouga/LlamaFactory

    Shared author/contributor keys: lin

  • LinkedLinked via arxiv author · 85%Yongqi Tong

    STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment

  • LinkedLinked via arxiv author · 85%Zhenyu Zhang

    STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment

  • LinkedLinked via arxiv author · 85%Ruirui Wang

    STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment

  • LinkedLinked via arxiv author · 85%Kewei Fu

    STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment

Implements (incoming)

authored (incoming)

Related across the graph

Topics