Read original ↗
paperarXivTrust 82 · PrimaryPublished 25d agoLive · 23d ago

ISO: An RLVR-Native Optimization Stack

Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimization layer that converts reward feedback into weight-space updates remains poorly understood. Building on our prior analysis (Zhu et al., 2025), we study this missing layer through the singular structure of model weights and identify spectral inheritance: RLVR can reuse the base model's weight spectra while acquiring new behavior through changes in the associated input and output singular frames. We operationalize spectral inheritance as Isospectral Optimi

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 54%RLHF
  • FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow

    Shared author/contributor keys: wang

  • FuzzyOverlapping authors or contributors · 62%modular/modular

    Shared author/contributor keys: liu

  • FuzzyOverlapping authors or contributors · 62%ray-project/ray

    Shared author/contributor keys: wang

  • LinkedLinked via arxiv author · 85%Hanqing Zhu

    ISO: An RLVR-Native Optimization Stack

  • LinkedLinked via arxiv author · 85%Wenyan Cong

    ISO: An RLVR-Native Optimization Stack

  • LinkedLinked via arxiv author · 85%Zhizhou Sha

    ISO: An RLVR-Native Optimization Stack

  • LinkedLinked via arxiv author · 85%Sagnik Mukherjee

    ISO: An RLVR-Native Optimization Stack

Related to

Implements (incoming)

authored (incoming)

Related across the graph

Topics