Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training

For LLM agents, supervised fine-tuning is not only about teacher labels' quality, but also about which interaction contexts those labels condition on. Pure behavioral cloning uses full teacher demonstrations, creating a mismatch between teacher-induced contexts seen in training and student-induced contexts encountered at test time. Recent work addresses this mismatch by querying a teacher at contexts reached by the student, often with increasingly elaborate filtering of the teacher's continuations. We instead frame on-policy data construction as a budget-allocation problem: under matched super

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • PossiblePossibly related (embedding) · 47%agent-tools
  • PossiblePossibly related (embedding) · 45%rllm-org/rllm
  • FuzzySimilar title/name (fuzzy) · 59%AgentCore-8B

    Fuzzy title match (0.73): “A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy ” ≈ “AgentCore-8B”

  • LinkedLinked via arxiv author · 85%Junze Ye

    A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training

  • LinkedLinked via arxiv author · 85%Jiayi Cheng

    A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training

  • LinkedLinked via arxiv author · 85%Miao Lu

    A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training

  • LinkedLinked via arxiv author · 85%Michal Mankowski

    A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training

  • LinkedLinked via arxiv author · 85%Jose Blanchet

    A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training

Implements

Has model

authored (incoming)

Implements (incoming)

Related across the graph

Topics