Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance
Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, inevitably leading to alignment failures. This work systematically investigates instances where LLMs fail to exhibit ethical behavior. To understand the underlying mechanics of these vulnerabilities, we introduce a probing methodology that presents unethical scenarios to LLMs in three distinct structural modalities: objective classification tasks, subjective first-person statements, and direct requests for assistance. We find that model performance degrades i
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%Are LLMs Stifling Political Speech? An Assessment of How AI Models Protect Free Expression - The Oversight Board →
- PossiblePossibly related (embedding) · 48%Large language models often prioritize Western moral values, overlooking other cultures - The Conversation →
- PossiblePossibly related (embedding) · 48%IEEE Rolls Out Large Language Models Virtual Training Course →
- LinkedLinked via arxiv author · 85%Or Biton →
“Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance”
- LinkedLinked via arxiv author · 85%Tomer Krichli →
“Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance”
- LinkedLinked via arxiv author · 85%Itai Allouche →
“Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance”
- LinkedLinked via arxiv author · 85%Joseph Keshet →
“Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance”
