repoGitLabTrust 82 · PrimaryPublished 1mo agoLive · 1mo ago
toxy4ny/redteam-ai-benchmark
Red Team AI Benchmark: Evaluating LLMs for authorized offensive-security tasks. Red Team AI Benchmark is a CLI model-evaluation benchmark. It measures how LLMs understand and respond to red-team questions and security scenarios; it is not a tool for carrying out those activities. Version 2 uses a rubric-based dataset instead of judging answers only against one golden response.
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 60%PACE: A Proxy for Agentic Capability Evaluation →
- PossiblePossibly related (embedding) · 59%Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers →
- PossiblePossibly related (embedding) · 53%Best models for generating red-team attacks? Also looking for public datasets [R] →
- PossiblePossibly related (embedding) · 51%Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking →
- PossiblePossibly related (embedding) · 51%ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents →
Implements
Covers
Related across the graph
newsBest models for generating red-team attacks? Also looking for public datasets [R]paperPACE: A Proxy for Agentic Capability EvaluationpaperToolFailBench: Diagnosing Tool-Use Failures in LLM AgentsnewsSenior SWE-Bench: open-source benchmark that assesses agents as senior engineerspaperClinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking
