newsHugging FaceTrust 88 · LabPublished 1mo agoLive · 1mo ago
ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownSWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions →
- LinkedLinked via unknownMirrorCode: AI can rebuild entire programs from behavior alone →
- LinkedLinked via unknownGovern the Repository, Not the Agent: Measuring Ecosystem-Level Risk in AI-Native Software →
- LinkedLinked via unknownFabiojvv/ai-cortex-hub →
- LinkedLinked via unknownTraceLab: Characterizing Coding Agent Workloads for LLM Serving →
- LinkedLinked via unknownAxDafny: Agentic Verified Code Generation in Dafny →
- LinkedLinked via unknownAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents? →
Covers
paperSWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding SessionspaperMirrorCode: AI can rebuild entire programs from behavior alonepaperGovern the Repository, Not the Agent: Measuring Ecosystem-Level Risk in AI-Native SoftwarerepoFabiojvv/ai-cortex-hubpaperTraceLab: Characterizing Coding Agent Workloads for LLM Serving
Covers (incoming)
paperAxDafny: Agentic Verified Code Generation in DafnypaperCan Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool UsepaperAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents?repohammadhaqqani/awesome-devops-airepoopensandbox-group/OpenSandboxreposamuelhm/42Jobsreporocketride-org/rocketride-serverrepoDashAISoftware/dashAIrepolangchain-ai/langgraphjsrepoTimefoldAI/timefold-solverrepopotpie-ai/potpierepospring-projects/spring-airepospiceai/spiceairepovercel/airepobytechefhq/bytechefrepozenml-io/zenmlreporun-llama/ParseBenchrepoedwardcapriolo/deliverancerepoidan-rubin/browserclawrepoalibaba/spring-ai-alibabapaperPACE: A Proxy for Agentic Capability EvaluationpaperTestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolutionrepogoogle/adk-javarepotrymirai/uzurepoAdamBien/airailsrepoTIGER-AI-Lab/ClawBenchrepocqfn/aibolitrepozapier/AutomationBenchrepoeunomia-bpf/bpf-benchmarkrepodeepjavalibrary/djlpaperOmniaBench: Benchmarking General AI Agents Across Diverse Scenariosrepoxlang-ai/OSWorld-V2repoai-for-decision-making-tue/Job_Shop_Scheduling_Benchmark_Environments_and_InstancespaperPACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments
Related across the graph
paperAre Performance-Optimization Benchmarks Reliably Measuring Coding Agents?repoDashAISoftware/dashAIpaperTraceLab: Characterizing Coding Agent Workloads for LLM ServingpaperAxDafny: Agentic Verified Code Generation in Dafnyrepodeepjavalibrary/djlrepoTIGER-AI-Lab/ClawBenchpaperMirrorCode: AI can rebuild entire programs from behavior alonerepoTimefoldAI/timefold-solverrepocqfn/aibolitrepoai-for-decision-making-tue/Job_Shop_Scheduling_Benchmark_Environments_and_Instancesrepoidan-rubin/browserclawrepopotpie-ai/potpierepoeunomia-bpf/bpf-benchmarkrepoFabiojvv/ai-cortex-hubrepozenml-io/zenmlreporocketride-org/rocketride-serverrepogoogle/adk-javapaperPACE: A Proxy for Agentic Capability Evaluationrepoedwardcapriolo/deliverancerepotrymirai/uzurepospring-projects/spring-airepospiceai/spiceaipaperPACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic EnvironmentspaperGovern the Repository, Not the Agent: Measuring Ecosystem-Level Risk in AI-Native Softwarerepohammadhaqqani/awesome-devops-aireporun-llama/ParseBenchrepoxlang-ai/OSWorld-V2paperTestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolutionrepoopensandbox-group/OpenSandboxrepovercel/aipaperCan Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Userepoalibaba/spring-ai-alibabapaperSWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessionsrepobytechefhq/bytechefrepoAdamBien/airailsreposamuelhm/42Jobsrepolangchain-ai/langgraphjsrepozapier/AutomationBenchpaperOmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
