Read original ↗
newsAI NewsTrust 60Published 2mo agoLive · 3mo ago

New benchmark exposes reasoning gaps in top models

A harder evaluation suite shows even leading models struggle on multi-hop tasks.

Why these links exist

Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.

Covers

Covers (incoming)

paperNuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language ModelscompanyNorthwind AItutorialEvaluate a model properlypaperCOCOLogic-V2: Identifying Logical Inconsistencies via Truly Hard-NegativespaperCan LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QApaperCognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty PredictionpaperThe Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth ScalingpaperThinkProbe: Beyond Accuracy -- Structural Profiling of Open-Ended LLM Reasoning Traces via Non-Generative Thought GraphspaperTravel-Oriented Reasoning Large Language Model via Domain-Specific Knowledge GraphspaperQuery-Aware Spreading Activation for Multi-Hop Retrieval over Knowledge GraphspaperAre We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math ReasoningpaperParametric SkillspaperDoes Verbose Chain-of-Thought Really Help? In-Distribution Evidence that Content, Not Length, MatterspaperBefore Thinking, Learn to Decide: Proactive Routing for Efficient Visual ReasoningpaperGrounding LLM Reasoning under Incomplete Graph EvidencepaperOn the Faithfulness of Post-Hoc Concept Bottleneck ModelspaperLearning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAspaperFork-Think with ConfidencepaperModality-Driven Search with Holistic Trace Judging for ARC-AGI-2paperThink in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using AgentspaperBridging the Gap Between Latent and Explicit Reasoning with Looped TransformerspaperMessage Passing Enables Efficient ReasoningpaperDiffusion-GR2: Diffusion Generative Reasoning Re-rankerrepoZefan-Cai/R-KVpaperTUDUM: A Turkish-Thinking Reasoning Pipeline for Qwen3.5-27BpaperCheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented ReasoningpaperA rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning taskspaperHNSW with Accuracy Guarantees Using Graph Spanners -- A Technical ReportpaperPluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource LanguagespaperEstimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMspaperTest-Time Scaling for Small VLMs on Multilingual Visual MCQpaperAdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verificationrepos-JoL/open-reasoningpaperPost-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape Pre-, Intra-, and Post-CoT CalibrationpaperAIMO Interpretability ChallengepaperIs Your Model Thinking or Just Stagnating? PUMA: Diagnosing Reasoning Pathology via Phase-Momentum AlignmentpaperTwo-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual CompletenesspaperPoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity

Related across the graph

paperCheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented ReasoningpaperDiffusion-GR2: Diffusion Generative Reasoning Re-rankerpaperCOCOLogic-V2: Identifying Logical Inconsistencies via Truly Hard-NegativespaperBefore Thinking, Learn to Decide: Proactive Routing for Efficient Visual ReasoningpaperDoes Verbose Chain-of-Thought Really Help? In-Distribution Evidence that Content, Not Length, MatterspaperEstimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMspaperThe Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth ScalingpaperPost-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape Pre-, Intra-, and Post-CoT CalibrationpaperHNSW with Accuracy Guarantees Using Graph Spanners -- A Technical ReportcompanyNorthwind AIpaperAre We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoningrepos-JoL/open-reasoningpaperMessage Passing Enables Efficient ReasoningpaperLearning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAspaperGrounding LLM Reasoning under Incomplete Graph EvidencepaperAIMO Interpretability ChallengerepoZefan-Cai/R-KVpaperPoTRE: Test-Time Reasoning inspired by Cognitive HeterogeneitypaperTravel-Oriented Reasoning Large Language Model via Domain-Specific Knowledge GraphspaperPluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource LanguagespaperA rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning taskspaperParametric SkillspaperFork-Think with ConfidencepaperThink in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using AgentspaperQuery-Aware Spreading Activation for Multi-Hop Retrieval over Knowledge GraphspaperAdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and VerificationpaperTwo-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual CompletenesspaperCan LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QApaperThinkProbe: Beyond Accuracy -- Structural Profiling of Open-Ended LLM Reasoning Traces via Non-Generative Thought GraphspaperModality-Driven Search with Holistic Trace Judging for ARC-AGI-2paperTUDUM: A Turkish-Thinking Reasoning Pipeline for Qwen3.5-27BtutorialEvaluate a model properlypaperBridging the Gap Between Latent and Explicit Reasoning with Looped TransformerspaperIs Your Model Thinking or Just Stagnating? PUMA: Diagnosing Reasoning Pathology via Phase-Momentum AlignmentarticleRetrieval is underratedpaperNuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language ModelspaperTest-Time Scaling for Small VLMs on Multilingual Visual MCQpaperCognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty PredictionpaperOn the Faithfulness of Post-Hoc Concept Bottleneck Models