A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis
LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimodal observability data. However, existing benchmarks remain largely outcome-oriented: they score only the final answer and fail to assess the systematic reasoning process in failure diagnosis. We address this gap by introducing two large-scale datasets (AIOps2025 and RCA100) under a reasoning-process evaluation paradigm that assesses agentic diagnostic capability along three dimensions: Localization (where the fault occurs), Identification (what type of fau
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownDebugging production agents with Amazon Bedrock AgentCore Observability →
- LinkedLinked via unknownAgentTrace →
- LinkedLinked via unknownagent-tools →
- LinkedLinked via unknownVerisight →
- PossiblePossibly related (embedding) · 50%semantica-agi/semantica →
- PossiblePossibly related (embedding) · 52%OpenDCAI/One-Eval →
- PossiblePossibly related (embedding) · 52%agamm/awesome-ai-sre →
