Show-Harness: Just a VLM Agent Can Play Robots
Foundation vision-language models (VLMs) exhibit broad intelligence about the world, yet translating this intelligence into robot control remains challenging. We present Show-Harness, an Embodied Harness that enables VLMs to "play" robots through a compact semantic interface linking intent to action. Show-Harness exposes discrete semantic action units that VLMs can naturally reason over, while embodiment-specific interpreters deterministically ground them into local robot actions, keeping the VLM directly responsible for fine-grained physical decisions. Through the same interface, Show-Harness
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- FuzzySimilar title/name (fuzzy) · 59%AgentCore-8B →
“Fuzzy title match (0.73): “Show-Harness: Just a VLM Agent Can Play Robots” ≈ “AgentCore-8B””
- LinkedLinked via arxiv author · 85%Yanzhe Chen →
“Show-Harness: Just a VLM Agent Can Play Robots”
- LinkedLinked via arxiv author · 85%Zechen Bai →
“Show-Harness: Just a VLM Agent Can Play Robots”
- LinkedLinked via arxiv author · 85%Zhijun Cao →
“Show-Harness: Just a VLM Agent Can Play Robots”
- LinkedLinked via arxiv author · 85%Wenzheng Zeng →
“Show-Harness: Just a VLM Agent Can Play Robots”
- LinkedLinked via arxiv author · 85%Kevin Qinghong Lin →
“Show-Harness: Just a VLM Agent Can Play Robots”
- LinkedLinked via arxiv author · 85%Yiqi Lin →
“Show-Harness: Just a VLM Agent Can Play Robots”
- LinkedLinked via arxiv author · 85%Guoqiang Liang →
“Show-Harness: Just a VLM Agent Can Play Robots”
