ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions
Benchmark saturation and data contamination increasingly obscure genuine scientific reasoning in frontier LLMs. We introduce \textsc{ScienceArena}, an olympiad-style benchmark from thirteen public science competitions in physics, chemistry, and biology, including IPhO and IChO 2025--2026, IBO 2023, USAPhO 2026, and USNCO 2025. Its open-ended, multi-step problems use process-credit rubrics, making faithful scoring difficult. We build ScienceArena through an expert-audited digitization pipeline that converts official exams, figures, solutions, and rubrics into structured items verified by olympi
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%Terminal-Bench-Science: Evaluating AI agents on scientific research workflows →
- PossiblePossibly related (embedding) · 49%We compared different LLMs on IMO 2026 [R] →
- PossiblePossibly related (embedding) · 48%USC leads national AI research project to accelerate scientific discovery - USC Today →
- FuzzyOverlapping authors or contributors · 62%ray-project/ray →
“Shared author/contributor keys: wang”
- FuzzyOverlapping authors or contributors · 62%google-research/google-research →
“Shared author/contributor keys: sun”
- FuzzyOverlapping authors or contributors · 62%modular/modular →
“Shared author/contributor keys: liu”
- FuzzyOverlapping authors or contributors · 62%TauricResearch/TradingAgents →
“Shared author/contributor keys: xiao”
- LinkedLinked via arxiv author · 85%Guangxiang Zhao →
“ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions”
