repoGitHubTrust 82 Β· PrimaryPublished 1mo agoLive Β· 6h ago
Giskard-AI/giskard-oss
π’ Open-Source Evaluation & Testing library for LLM Agents
Lineage graph
Paper β model β repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it β so bad links are debuggable.
- PossiblePossibly related (embedding) Β· 63%PACE: A Proxy for Agentic Capability Evaluation β
- PossiblePossibly related (embedding) Β· 58%TraceLab: Characterizing Coding Agent Workloads for LLM Serving β
- PossiblePossibly related (embedding) Β· 56%AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents β
- PossiblePossibly related (embedding) Β· 46%TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution β
- PossiblePossibly related (embedding) Β· 50%A$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction β
- PossiblePossibly related (embedding) Β· 57%GitLab CI skill for ai agents based on official docs β
- PossiblePossibly related (embedding) Β· 58%LLM-as-a-Verifier: A General-Purpose Verification Framework β
- PossiblePossibly related (embedding) Β· 58%New LLM Coordination Benchmark - Benchmarking Open-Ended Multi-Agent Coordination in Language Agents [R] β
Implements
Implements (incoming)
paperTestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-EvolutionpaperA$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT ConstructionpaperLLM-as-a-Verifier: A General-Purpose Verification FrameworkpaperCan LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment ReproductionpaperWho Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
Covers (incoming)
Related across the graph
paperTraceLab: Characterizing Coding Agent Workloads for LLM ServingpaperAgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM AgentspaperCan LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment ReproductionnewsNew LLM Coordination Benchmark - Benchmarking Open-Ended Multi-Agent Coordination in Language Agents [R]paperWho Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM AgentspaperLLM-as-a-Verifier: A General-Purpose Verification FrameworknewsGitLab CI skill for ai agents based on official docspaperPACE: A Proxy for Agentic Capability EvaluationnewsDevs shipping AI agents what does your security testing look like ?paperTestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-EvolutionpaperA$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction
