BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance
As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis. We present BioSecBench-Surveillance, a verifiable benchmark of 100 evaluations testing whether AI agents can infer the right analysis pipeline from raw sequencing data and surveillance context. Each evaluation gives an agent only the data and context a human analyst would have, then grades its structured answer deterministically. The tasks span seven categories, from taxonomic classification to genetic-engineering detection, across diverse sample types and sequencing technologies. Across 3,962
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 65%Introducing GeneBench-Pro →
- PossiblePossibly related (embedding) · 53%Data-driven surrogates of rational design enable antimicrobial peptide optimization →
- PossiblePossibly related (embedding) · 52%Using AI to help physicians diagnose rare genetic diseases affecting children →
- FuzzySimilar title/name (fuzzy) · 87%NirDiamant/GenAI_Agents →
“Fuzzy title match (0.94): “BioSecBench-Surveillance: A Verifiable Benchmark for AI Agen” ≈ “NirDiamant/GenAI_Agents””
- FuzzySimilar title/name (fuzzy) · 84%Unity-Technologies/ml-agents →
“Fuzzy title match (0.92): “BioSecBench-Surveillance: A Verifiable Benchmark for AI Agen” ≈ “Unity-Technologies/ml-agents””
- FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzySimilar title/name (fuzzy) · 59%datawhalechina/hello-agents →
“Fuzzy title match (0.73): “BioSecBench-Surveillance: A Verifiable Benchmark for AI Agen” ≈ “datawhalechina/hello-agents””
