NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models
Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly technical domains remains a significant challenge. In nuclear engineering, problem solving often requires not only factual knowledge but also quantitative reasoning and conceptual understanding. To address the need for systematic evaluation in this domain, we introduce NuclearQAv2, a benchmark for assessing LLMs on nuclear engineering knowledge. The benchmark comprises approximately 1,240 question-answer pairs spanning three categories: boolean, numeric, and
Lineage graph
Paper → model → repo connections mined from source citations (Tier-1 exact match).
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- LinkedLinked via unknownNew benchmark exposes reasoning gaps in top models →
- LinkedLinked via unknownDeepSWE: new benchmark looking at how well today's frontier models can actually write code [R] →
- LinkedLinked via unknownNorthwind AI →
- LinkedLinked via unknownBook Review: Domain-Specific Small Language Models by Guglielmo Iozzia →
- LinkedLinked via unknownIEEE Rolls Out Large Language Models Virtual Training Course →
- LinkedLinked via unknownKnowledge Distillation of Black-Box Large Language Models →
- LinkedLinked via unknownKnowledge Distillation of Black-Box Large Language Models (2024) →
- PossiblePossibly related (embedding) · 51%VectorPeak/LLM-Wiki →
