Read original ↗
paperarXivTrust 82 · PrimaryPublished 5d agoLive · 4d ago

StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized actions. Existing guardrails often evaluate completed trajectories, leaving pre-execution monitoring of step-level actions underexplored. We propose StepGuard, a step-level guard model that can audit completed agent trajectories and check tool actions before they are executed. To train StepGuard, we introduce StepGen, an automatic data engine that generates safe and unsafe trajectories with the same

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • LinkedLinked via arxiv author · 85%Zhijie Zheng

    StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

  • LinkedLinked via arxiv author · 85%Zhenyu Liu

    StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

  • LinkedLinked via arxiv author · 85%Chen Qian

    StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

  • LinkedLinked via arxiv author · 85%Yuqian Fu

    StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

  • LinkedLinked via arxiv author · 85%Yanwei Fu

    StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

  • LinkedLinked via arxiv author · 85%Lu Sheng

    StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

  • LinkedLinked via arxiv author · 85%Jing Shao

    StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

  • LinkedLinked via arxiv author · 85%Dongrui Liu

    StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

authored (incoming)

Implements (incoming)

Related across the graph

Topics