Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction
Agentic systems offer a promising way to automate embodied benchmark construction, but existing approaches typically cover isolated stages or remain specialized to predefined environments and task families. More importantly, multi-step construction produces dependent intermediate artifacts that are often passed downstream without artifact-specific verification, allowing local defects to propagate into the final benchmark. We present Embodied-BenchForge, an agentic framework that transforms user-specified evaluation intents into complete embodied benchmark artifacts. It formulates construction
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 56%Beyond benchmarks: The 5 pillars of AI evaluation systems →
- PossiblePossibly related (embedding) · 55%What would a fair benchmark for agent architecture look like? [D] →
- LinkedLinked via arxiv author · 85%Baoyang Jiang →
“Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction”
- LinkedLinked via arxiv author · 85%Fengchun Zhang →
“Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction”
- LinkedLinked via arxiv author · 85%Leyuan Wang →
“Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction”
- LinkedLinked via arxiv author · 85%Haotian Li →
“Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction”
- LinkedLinked via arxiv author · 85%Yida Wang →
“Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction”
- LinkedLinked via arxiv author · 85%Zhe Jiang →
“Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction”
