Read original ↗
newsReddit r/MachineLearningTrust 72 · CommunityPublished 1mo agoLive · 1mo ago

A system-level approach to prompt injection: separating instruction and data channels in LLM agents [P]

Prompt injection has emerged as one of the most persistent failure modes in tool-using LLM systems, particularly in agentic workflows where models interact with external data sources. Most mitigation strategies focus on input filtering or model-side alignment, but these approaches struggle because the core issue is structural: Approach I explored a system-level mitigation strategy by introducing a middleware laye

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Covers (incoming)

paperAutomating Cause-Effect Specification with Knowledge Graphs and Large Language ModelspaperA Tutorial on Autonomous Fault-Tolerant Control Using Knowledge-Grounded LLM AgentspaperEntity Binding Failures in Tool-Augmented AgentspaperWords Speak Louder Than Code: Investigating Cognitive Heuristics in LLM-Based Code Vulnerability DetectionpaperMESA: Prioritizing Vulnerable Communication Channels for Securing Multi-Agent SystemspaperMCP Server Architecture Patterns for LLM-Integrated ApplicationspaperSWE-Doctor: Guiding Software Engineering Agents with Runtime Diagnosis from Multi-Faceted Bug Reproduction TestspaperDistill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillationrepoaallan/verarepopromptfoo/promptfoorepolanggenius/difyrepoPipelex/pipelexrepolotus-data/lotusrepoNirDiamant/Prompt_Engineeringrepoagentjido/req_llmpaperLLM-Driven CI-CD Workflow Intelligence for Cyber Systems Engineeringrepoatharva557/Prompt-ChainingpaperScoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution ShiftpaperWhen Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systemsrepomicrosoft/PromptKitrepopromptfoo/promptfoo-actionpaperTracing Agentic Failure from the Flow of Successrepoucbepic/docetlpaperPrompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Modelsrepocyqlelabs/palrepoEdgarOrtegaRamirez/agent-abtest-frameworkrepoSuppieRK/cmdshapepaperQuoteBench: How Matched Scores Can Hide Command-Path Failures

Related across the graph

paperMCP Server Architecture Patterns for LLM-Integrated ApplicationspaperWhen Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent SystemspaperTracing Agentic Failure from the Flow of SuccesspaperWords Speak Louder Than Code: Investigating Cognitive Heuristics in LLM-Based Code Vulnerability DetectionpaperA Tutorial on Autonomous Fault-Tolerant Control Using Knowledge-Grounded LLM Agentsrepolanggenius/difyrepopromptfoo/promptfoo-actionpaperScoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution ShiftpaperLinguistic Firewall: Geometry as Defense in Multi-Agent Systems RoutingrepoSuppieRK/cmdshaperepoatharva557/Prompt-Chainingrepomicrosoft/PromptKitrepoEdgarOrtegaRamirez/agent-abtest-frameworkrepocyqlelabs/palpaperPrompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language ModelsrepoPipelex/pipelexpaperPolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM Agentsrepoucbepic/docetlrepoagentjido/req_llmpaperMESA: Prioritizing Vulnerable Communication Channels for Securing Multi-Agent SystemspaperDistill to Detect: Exposing Stealth Biases in LLMs through Cartridge DistillationpaperEntity Binding Failures in Tool-Augmented Agentsrepolotus-data/lotuspaperLLM-Driven CI-CD Workflow Intelligence for Cyber Systems EngineeringrepoNirDiamant/Prompt_EngineeringpaperSWE-Doctor: Guiding Software Engineering Agents with Runtime Diagnosis from Multi-Faceted Bug Reproduction Testsrepopromptfoo/promptfoopaperAutomating Cause-Effect Specification with Knowledge Graphs and Large Language ModelspaperA Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open ProblemspaperWhen the Database Fails: Prompting LLM Dialogue Agents for Safe Recovery in Task-Oriented Dialoguerepoagent-toolsrepoaallan/verapaperQuoteBench: How Matched Scores Can Hide Command-Path Failures