newsBAIR (Berkeley)Trust 88 · LabPublished 1y agoLive · 1mo ago
Defending against Prompt Injection with Structured Queries (StruQ) and Preference Optimization (SecAlign)
Recent advances in Large Language Models (LLMs) enable exciting LLM-integrated applications. However, as LLMs have improved, so have the attacks against them. Prompt injection attack is listed as the #1 threat by OWASP to LLM-integrated applications, where an LLM input contains a trusted prompt (instruction) and an untr
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownDistill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation →
- PossiblePossibly related (embedding) · 47%beyefendi/awesome-llm-security →
- PossiblePossibly related (embedding) · 46%yegor256/prompt →
- PossiblePossibly related (embedding) · 47%Retroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations →
- PossiblePossibly related (embedding) · 48%When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems →
Covers (incoming)
paperPrompt Injection in Automated Résumé Screening with Large Language Models: Single and Multi-Injection SettingspaperWords Speak Louder Than Code: Investigating Cognitive Heuristics in LLM-Based Code Vulnerability DetectionpaperA Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open ProblemspaperDistill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillationrepobeyefendi/awesome-llm-securityrepoyegor256/promptpaperRetroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model GenerationspaperWhen Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systemsrepopromptfoo/promptfoo-actionpaperPretraining Data Can Be Poisoned through Computational Propaganda
Related across the graph
paperWhen Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent SystemspaperWords Speak Louder Than Code: Investigating Cognitive Heuristics in LLM-Based Code Vulnerability Detectionrepopromptfoo/promptfoo-actionrepobeyefendi/awesome-llm-securitypaperPretraining Data Can Be Poisoned through Computational PropagandapaperDistill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillationrepoyegor256/promptpaperPrompt Injection in Automated Résumé Screening with Large Language Models: Single and Multi-Injection SettingspaperRetroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model GenerationspaperA Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems
