Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring
As Large Language Models (LLMs) and agentic systems become integrated into real-world applications, ensuring their safety and security is critical. Guardrail systems that detect and block malicious instructions sent to and from an LLM are an essential component of AI security. However, researchers conducting black-box adversarial emulation against production AI systems often struggle to determine whether a guardrail block or an LLM rejection has occurred. This distinction is important because the techniques used to bypass guardrails can differ substantially from those used to bypass LLM safety
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 62%Prompt injection is exploiting enterprise AI's biggest design flaws by targeting agents, RAG pipelines and model routers →
- PossiblePossibly related (embedding) · 61%Securing the future of AI agents →
- PossiblePossibly related (embedding) · 60%AI browsers can be lulled into a dream world where guardrails no longer apply →
- PossiblePossibly related (embedding) · 57%"Dangerous" AI models are coming no matter what →
- PossiblePossibly related (embedding) · 54%How to Secure AI Agents With Container Sandboxing - HackerNoon →
- PossiblePossibly related (embedding) · 35%guardrails-ai/guardrails →
“Possibly related via embedding similarity 0.69 (not asserted). Timestamp check: artifact after paper (+14d).”
- LinkedLinked via arxiv author · 85%William Hackett →
“Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring”
- LinkedLinked via arxiv author · 85%Peter Garraghan →
“Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring”
