Read original ↗
paperarXivTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago

Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring

As Large Language Models (LLMs) and agentic systems become integrated into real-world applications, ensuring their safety and security is critical. Guardrail systems that detect and block malicious instructions sent to and from an LLM are an essential component of AI security. However, researchers conducting black-box adversarial emulation against production AI systems often struggle to determine whether a guardrail block or an LLM rejection has occurred. This distinction is important because the techniques used to bypass guardrails can differ substantially from those used to bypass LLM safety

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Related to

authored (incoming)

Covers (incoming)

Implements (incoming)

Related across the graph

repoJoasASantos/NeuroSploitrepoggwhite/4xnewsWhat does "Safe AI" look like? [D]newsHackers Use Fake API Documentation to Trick AI Agents Into Sending Crypto Payments - gbhackers.comrepoguardrails-ai/guardrailsnewsSingGuard-NSFA: Open-source guardrails for agentic AI - Help Net SecurityrepoOrdo-Engine/OrdonewsNew Open-Source AI Security Guardrails - Open Source For YounewsArtificial Intelligence in Behavioral Threat Assessment: Early Identification and Threat Hunting - Campus Security TodaynewsHow AI guardrails are impeding the work of offensive cybersecurity researchersnewsTrapping Malicious AI Knowledge Into On/Off Switchable Modules Gets Underway - ForbesnewsPrompt Injection Attacks Are Thwarting AI Hacking AgentsnewsNew Research: AI models can give dangerous responses despite output guardrails - The AI JournalnewsWhy prompt-level guardrails aren't enough: The platform security layers production agents needrepokillertcell428/aigisrepoVishisht16/Humane-ProxynewsPrompt injection is exploiting enterprise AI's biggest design flaws by targeting agents, RAG pipelines and model routersrepoLuD1161/agentjailnewsA sociotechnical threat model for AI-driven smart home devicespersonWilliam HackettnewsSouth Korea making its own security-centric AI modelnewsChain-of-Thought Spoofing Targets Reasoning AI Models - Hackadayrepolinghungegeg/LinghunnewsTo defend your software, first teach AI to break it - Virginia Tech NewsnewsAI browsers can be lulled into a dream world where guardrails no longer applynewsManaging third-party model risk and AI dependencies - Security BoulevardnewsBehavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policies - Apple Machine Learning ResearchnewsSecuring the future of AI agentspersonPeter GarraghannewsAre LLMs Stifling Political Speech? An Assessment of How AI Models Protect Free Expression - The Oversight BoardrepoTalEliyahu/Awesome-AI-Securityrepoanmolksachan/AI-ML-Free-Resources-for-Security-and-Prompt-Injectionnews"Dangerous" AI models are coming no matter whatnewsHow to Secure AI Agents With Container Sandboxing - HackerNoon

Topics