Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity
Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task. Existing benchmarks often compress these distinctions into pass/fail labels, obscuring whether failures arise from capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judgments. This paper introduces adversarial pragmatics as a benchmark and annotation protocol for eva
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownVerisight →
- LinkedLinked via arxiv author · 85%Brett Reynolds →
“Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Amb”
- PossiblePossibly related (embedding) · 53%Understanding Annotator Safety Policy with Interpretability - Apple Machine Learning Research →
- FuzzySimilar title/name (fuzzy) · 59%jeinlee1991/chinese-llm-benchmark →
“Fuzzy title match (0.73): “Adversarial Pragmatics for AI Safety Evaluation: A Benchmark” ≈ “jeinlee1991/chinese-llm-benchmark””
