newsReddit r/artificialTrust 52 · CommunityPublished 16d agoLive · 15d ago
1.7B model leading strict-7 formal reasoning above Qwen3-8B and Gemma-4-26B - specialists eating generalist territory?
Most of the reasoning gains coming out of the big labs are still tied to scale. More params, more compute, better reasoning. That's been the play for a while. Ran into TwIL-LM2 which flips the script for narrow tasks. PEFT LoRA adapter on SmolLM2-1.7B, specialized purely for formal logic translation. On strict-7 scoring (no partial credit, exact-format required) it hits 0.2386 - ahead of Qwen3-8B at 0.2093 and Gemma-4-26B at 0.2050. On the loose-match six
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 55%pzqpzq/LSF_MDia →
- PossiblePossibly related (embedding) · 52%Test-Time Scaling for Small VLMs on Multilingual Visual MCQ →
- PossiblePossibly related (embedding) · 51%LLM-as-a-Verifier: A General-Purpose Verification Framework →
- PossiblePossibly related (embedding) · 51%PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages →
- PossiblePossibly related (embedding) · 50%AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification →
Covers
repopzqpzq/LSF_MDiapaperTest-Time Scaling for Small VLMs on Multilingual Visual MCQpaperLLM-as-a-Verifier: A General-Purpose Verification FrameworkpaperPluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource LanguagespaperAdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification
Related across the graph
paperLLM-as-a-Verifier: A General-Purpose Verification FrameworkpaperPluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource LanguagespaperAdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verificationrepopzqpzq/LSF_MDiapaperTest-Time Scaling for Small VLMs on Multilingual Visual MCQ
