StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized actions. Existing guardrails often evaluate completed trajectories, leaving pre-execution monitoring of step-level actions underexplored. We propose StepGuard, a step-level guard model that can audit completed agent trajectories and check tool actions before they are executed. To train StepGuard, we introduce StepGen, an automatic data engine that generates safe and unsafe trajectories with the same
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via arxiv author · 85%Zhijie Zheng →
“StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing”
- LinkedLinked via arxiv author · 85%Zhenyu Liu →
“StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing”
- LinkedLinked via arxiv author · 85%Chen Qian →
“StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing”
- LinkedLinked via arxiv author · 85%Yuqian Fu →
“StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing”
- LinkedLinked via arxiv author · 85%Yanwei Fu →
“StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing”
- LinkedLinked via arxiv author · 85%Lu Sheng →
“StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing”
- LinkedLinked via arxiv author · 85%Jing Shao →
“StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing”
- LinkedLinked via arxiv author · 85%Dongrui Liu →
“StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing”
