repoGitHubTrust 82 · PrimaryPublished 1mo agoLive · 24d ago
OpenDataBox/Workspace-Bench
Benchmark self-evolving Agent upon realistic large-scale file workspaces
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%Is it agentic enough? Benchmarking open models on your own tooling →
- PossiblePossibly related (embedding) · 54%Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use →
- PossiblePossibly related (embedding) · 54%TraceLab: Characterizing Coding Agent Workloads for LLM Serving →
- PossiblePossibly related (embedding) · 54%TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution →
- PossiblePossibly related (embedding) · 54%PACE: A Proxy for Agentic Capability Evaluation →
Covers
Implements
paperCan Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool UsepaperTraceLab: Characterizing Coding Agent Workloads for LLM ServingpaperTestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-EvolutionpaperPACE: A Proxy for Agentic Capability Evaluation
Related across the graph
paperTraceLab: Characterizing Coding Agent Workloads for LLM ServingpaperPACE: A Proxy for Agentic Capability EvaluationpaperTestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-EvolutionpaperCan Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool UsenewsIs it agentic enough? Benchmarking open models on your own tooling
