AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification
Large language models (LLMs) have achieved remarkable performance on high-school and olympiad-style mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provide limited disciplinary coverage and often rely on final-answer correctness or coarse judgments, leaving the validity of the reasoning process inadequately assessed. To bridge this gap, we introduce AdvancedMathBench, a benchmark suite designed to evaluate advanced mathematical reasoning capabilities. Its core proof-gen
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 57%kyegomez/Lets-Verify-Step-by-Step →
- PossiblePossibly related (embedding) · 56%frenzymath/Danus →
- PossiblePossibly related (embedding) · 54%New benchmark exposes reasoning gaps in top models →
- PossiblePossibly related (embedding) · 50%Mistral Open-Sources AI Model That Can Verify Code and Mathematical Proofs - ProPakistani →
- PossiblePossibly related (embedding) · 49%sileod/reasoning-core →
- PossiblePossibly related (embedding) · 51%We compared different LLMs on IMO 2026 [R] →
- PossiblePossibly related (embedding) · 50%1.7B model leading strict-7 formal reasoning above Qwen3-8B and Gemma-4-26B - specialists eating generalist territory? →
- PossiblePossibly related (embedding) · 55%AI Used to Verify Toughest Mathematics Proof Yet →
