QuoteBench: How Matched Scores Can Hide Command-Path Failures
LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 56 one-shot tasks from 14 incident-derived families, crossing the generation contract with the execution transport around one deliberately unescaped added parser. Escaping at the interpolation point reproduces each replayed reply's raw-path outcome, so any recovery under a disclosed boundary must com
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 48%A system-level approach to prompt injection: separating instruction and data channels in LLM agents [P] →
- LinkedLinked via arxiv author · 85%Shangao Li →
“QuoteBench: How Matched Scores Can Hide Command-Path Failures”
- LinkedLinked via arxiv author · 85%Yuyao Zhang →
“QuoteBench: How Matched Scores Can Hide Command-Path Failures”
- LinkedLinked via arxiv author · 85%Volker Tresp →
“QuoteBench: How Matched Scores Can Hide Command-Path Failures”
- LinkedLinked via arxiv author · 85%Yuanyuan Yang →
“QuoteBench: How Matched Scores Can Hide Command-Path Failures”
