LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos
Large language models (LLMs) excel at pattern recognition and text generation, but their capacity for abductive inference - inferring latent hypotheses that explain observed behavior - remains poorly understood. Here, we introduce Elenchos (named after the Socratic method of cross-examination), a generative evaluation framework that measures abductive reasoning as a structural inverse problem. Given a reference formal system, such as the lambda-calculus, and a potentially mutated counterpart, agents must determine whether a mutation has occurred and infer the rule modifications responsible for
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%benjaminzwhite/reasoning-models →
- PossiblePossibly related (embedding) · 52%amitshekhariitbhu/llm-internals →
- PossiblePossibly related (embedding) · 51%sileod/reasoning-core →
- PossiblePossibly related (embedding) · 51%KennispuntTwente/tidyprompt →
- PossiblePossibly related (embedding) · 50%Retrace-1.5B →
- LinkedLinked via arxiv author · 85%Julius Steiglechner →
“LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos”
- LinkedLinked via arxiv author · 85%Lucas Mahler →
“LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos”
- LinkedLinked via arxiv author · 85%Gabriele Lohmann →
“LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos”
