SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment
The performance of LLM-based agents is jointly shaped by the base model and the harness used when interacting with the environment. This exposes them to safety risks in both harmful final responses and multi-step execution trajectories. Existing safety alignment mechanisms often rely on either external harness updates or policy optimization, yet applying either paradigm in isolation fails to bridge runtime control with intrinsic safety. We propose SafeEvolve, an experience-driven self-evolving framework for agent safety alignment. SafeEvolve leverages safety experience from completed on-policy
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%AgentCore-8B →
“Fuzzy title match (0.73): “SafeEvolve: Harness-Policy Co-Evolution from Agent Experienc” ≈ “AgentCore-8B””
- PossiblePossibly related (embedding) · 65%Safety and alignment in an era of long-horizon models →
- PossiblePossibly related (embedding) · 54%EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses [R] →
- LinkedLinked via arxiv author · 85%Qinghua Mao →
“SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment”
- LinkedLinked via arxiv author · 85%Wanying Qu →
“SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment”
- LinkedLinked via arxiv author · 85%Dadi Guo →
“SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment”
- LinkedLinked via arxiv author · 85%Leitao Yuan →
“SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment”
- LinkedLinked via arxiv author · 85%Qingyu Liu →
“SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment”
