newsHacker NewsTrust 52 · CommunityPublished 4d agoLive · 4d ago
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
91points27comments
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 66%InternScience/ResearchClawBench →
- PossiblePossibly related (embedding) · 64%Starlight143/crucible →
- PossiblePossibly related (embedding) · 63%drpedapati/sciclaw →
- PossiblePossibly related (embedding) · 63%synthetic-sciences/openscience →
- PossiblePossibly related (embedding) · 62%SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery →
- PossiblePossibly related (embedding) · 50%Future-House/aviary →
- PossiblePossibly related (embedding) · 50%ScienceArena: Benchmarking LLMs on Latest Scientific Olympiad Competitions →
- PossiblePossibly related (embedding) · 59%Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents →
Covers
Covers (incoming)
Related across the graph
repoFuture-House/aviarypaperSciForge: An AI-Native, Multimodal Workbench for Scientific DiscoveryrepoStarlight143/cruciblerepodrpedapati/sciclawreposynthetic-sciences/opensciencepaperLearning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research AgentspaperScienceArena: Benchmarking LLMs on Latest Scientific Olympiad CompetitionsrepoInternScience/ResearchClawBench
