ECHO: A Matched-Contrast Benchmark for Context-Sensitive Turn-Taking in Full-Duplex Dialogue
Full-duplex spoken dialogue systems must distinguish interruptions that require yielding the floor from backchannels that permit continued speaking. Existing benchmarks typically evaluate events independently and may therefore reward fixed action preferences rather than context-sensitive decisions. We introduce ECHO, a paired diagnostic benchmark for Chinese full-duplex turn-taking. ECHO pairs examples with the same overlap transcript but contrasting preceding multi-turn dialogue contexts, with one requiring Yield and the other Keep. It additionally includes off-talk examples for diagnosing un
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%The bottleneck for meeting transcription tools isn't accurate anymore, it's speaker attribution →
- PossiblePossibly related (embedding) · 49%Building an AI voice agent from scratch: the parts that actually took our time →
- PossiblePossibly related (embedding) · 48%ChatGPT-Live vs Pi vs Lucy OS1 vs Gemini-Live: best AI assistant to talk with? →
- PossiblePossibly related (embedding) · 46%Launch HN: Speko (YC S26) – OpenRouter for Voice AI →
- PossiblePossibly related (embedding) · 46%Agent Evaluation Metric for multi-turn conversations →
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%sgl-project/sglang →
“Shared author/contributor keys: zhou”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
