Read original ↗
paperarXivTrust 82 · PrimaryPublished 5d agoLive · 2d ago

FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

Large language models (LLMs) have shown growing potential for automated theoretical computer science (TCS) research, yet existing benchmarks remain far from realistic research settings. We introduce \ourbenchmark, an expert-validated benchmark for evaluating LLMs on frontier, end-to-end TCS research. \ourbenchmark contains $175$ instances drawn from papers accepted to STOC, FOCS, SODA, and COLT in 2025-2026, preserving paper-specific definitions, assumptions, and proof dependencies, with expert-verified Lean formalizations and proofs. Evaluations of leading LLMs reveal that current models rema

Lineage graph

Paper → model → repo connections mined from source citations (Tier-1 exact match).

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

  • FuzzyOverlapping authors or contributors · 62%ray-project/ray

    Shared author/contributor keys: wang

  • FuzzyOverlapping authors or contributors · 62%bytedance/deer-flow

    Shared author/contributor keys: wang

  • FuzzySimilar title/name (fuzzy) · 59%google-research/google-research

    Fuzzy title match (0.73): “FormalTCS: Benchmarking End-to-End Frontier Formal Theoretic” ≈ “google-research/google-research”

  • LinkedLinked via arxiv author · 85%Dingzirui Wang

    FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

  • LinkedLinked via arxiv author · 85%Xuanliang Zhang

    FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

  • LinkedLinked via arxiv author · 85%Keyan Xu

    FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

  • LinkedLinked via arxiv author · 85%Qingfu Zhu

    FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

  • LinkedLinked via arxiv author · 85%Wanxiang Che

    FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models

Implements (incoming)

authored (incoming)

Related across the graph

Topics