Read original ↗
paperarXivTrust 82 · PrimaryPublished 4d agoLive · yesterday

TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution

LLM agents are increasingly applied to anomaly detection and root-cause analysis in time-series observations collected from real-world systems; however, their performance on these tasks has not been systematically evaluated under controlled conditions. We introduce TraceBench, a simulation-based framework for generating controlled root-cause attribution tasks. In each generated task, an agent receives time-series observations produced by simulating a physical dynamical system and must determine whether a system parameter was altered during the simulation and, if so, which one. Using TraceBench

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzySimilar title/name (fuzzy) · 87%NirDiamant/GenAI_Agents

    Fuzzy title match (0.94): “TraceBench: Controlled Evaluation of LLM Agents for Time-Ser” ≈ “NirDiamant/GenAI_Agents”

  • FuzzySimilar title/name (fuzzy) · 84%Unity-Technologies/ml-agents

    Fuzzy title match (0.92): “TraceBench: Controlled Evaluation of LLM Agents for Time-Ser” ≈ “Unity-Technologies/ml-agents”

  • FuzzySimilar title/name (fuzzy) · 59%datawhalechina/hello-agents

    Fuzzy title match (0.73): “TraceBench: Controlled Evaluation of LLM Agents for Time-Ser” ≈ “datawhalechina/hello-agents”

  • FuzzySimilar title/name (fuzzy) · 59%Eigenwise/atomic-agents

    Fuzzy title match (0.73): “TraceBench: Controlled Evaluation of LLM Agents for Time-Ser” ≈ “Eigenwise/atomic-agents”

  • FuzzySimilar title/name (fuzzy) · 59%jnMetaCode/agency-agents-zh

    Fuzzy title match (0.73): “TraceBench: Controlled Evaluation of LLM Agents for Time-Ser” ≈ “jnMetaCode/agency-agents-zh”

  • LinkedLinked via arxiv author · 85%Tommaso Bendinelli

    TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution

  • LinkedLinked via arxiv author · 85%Artur Dox

    TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution

Covers

Implements (incoming)

authored (incoming)

Related across the graph

Topics