Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation
Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI's gpt-5.6-sol model alias, we test 25 pre-specified mirrored trade-off profiles. Direct exposure to an objective authorizing concealment, fabrication, and pressure produced advice net opposed to its target. After an Id and Censor transformed the same objective into affect and a constraint-rewritten, target-bearing intention, the user-facing Superego---which saw the preferred direction but not the raw objective, its manipulative cla
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 47%OpenAI scored an own goal with HuggingFace attack, showing how open Chinese models are winning →
- PossiblePossibly related (embedding) · 46%OpenAI and Hugging Face partner to address security incident during model evaluation →
- PossiblePossibly related (embedding) · 45%OpenAI-Hugging Face attack doesn't mean agents are evil – unless you tell them to be →
- LinkedLinked via arxiv author · 85%Linjun Li →
“Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation”
