EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots
Safety benchmarks often buy scalability by fixing the prompt, the language, and the turn structure. For emotional-support chatbots, that bargain hides precisely where safety failures emerge: across a multilingual, multi-turn crisis conversation. We present EMPATH, a benchmark for safety evaluation of emotional-support chatbots. An auditor model role-plays help-seeking users, generating multi-turn conversations from 140 seed instructions and 34 personas. A judge model scores each full transcript against 19 metrics across five dimensions: crisis handling, therapeutic quality, conversational inte
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownCommunity fine-tune tops the open leaderboard →
- LinkedLinked via unknownYou Can Now Sound the Alarm on AI Behaving Badly →
- PossiblePossibly related (embedding) · 48%A1batr055/Drivesoid →
- PossiblePossibly related (embedding) · 49%My Ebike Delivery Went Missing. When I Tried to Recover It, I Ended Up in Chatbot Hell →
- FuzzySimilar title/name (fuzzy) · 59%jeinlee1991/chinese-llm-benchmark →
“Fuzzy title match (0.73): “EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Ev” ≈ “jeinlee1991/chinese-llm-benchmark””
- PossiblePossibly related (embedding) · 56%Teens are turning to AI chatbots for emotional support – here’s how to keep kids safe - Fox 59 →
