Teach it to stop, not just to click
Agentic computer-use RL is reported in single runs, and those numbers mislead. Using verifier-guided repair of a 35B computer-use agent (CUA) across five oracle-graded environments, we show a repaired policy's success rate is dominated by upstream variance: a variance-components decomposition across three cells (crossed data-draw $\times$ seed grid, bootstrap CIs) finds evaluation variance negligible ($σ_{\mathrm{eval}} \approx 0$) and the training-seed effect small everywhere ($\leq 10\%$); instead it splits between the data draw and run-to-run nondeterminism, the data draw's share rising to
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 49%I benchmarked 13 models at 65K-128K context to find out what actually matters for agentic workloads →
- PossiblePossibly related (embedding) · 47%Pluggable by design: An agent mesh for software modernization that adopts the next model release →
- PossiblePossibly related (embedding) · 47%I made a superhuman Generals.io agent with self-play RL [P] →
- PossiblePossibly related (embedding) · 47%Production AI & The False Finish Line [D] →
- LinkedLinked via arxiv author · 85%Barada Sahu →
“Teach it to stop, not just to click”
- LinkedLinked via arxiv author · 85%Shivesh Pandey →
“Teach it to stop, not just to click”
