newsAWS Machine LearningTrust 88 · LabPublished 25d agoLive · 23d ago
Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova
In this post, we explore an idea for generating thinking tokens for datasets that lack reasoning traces in SFT customization. We first examine the reasoning suppression problem, then introduce Self-Distilled Reasoning (SDR), validate it across three benchmarks, and provide practical recommendations.
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 59%ThinkProbe: Beyond Accuracy -- Structural Profiling of Open-Ended LLM Reasoning Traces via Non-Generative Thought Graphs →
- PossiblePossibly related (embedding) · 58%Better Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed Reasoning →
- PossiblePossibly related (embedding) · 57%Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering →
- PossiblePossibly related (embedding) · 56%sileod/reasoning-core →
- PossiblePossibly related (embedding) · 54%Modality-Driven Search with Holistic Trace Judging for ARC-AGI-2 →
- PossiblePossibly related (embedding) · 45%RS-RIE-Bench: Benchmarking Reasoning-Guided Remote Sensing Image Editing →
- PossiblePossibly related (embedding) · 46%Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models →
Covers
paperThinkProbe: Beyond Accuracy -- Structural Profiling of Open-Ended LLM Reasoning Traces via Non-Generative Thought GraphspaperBetter Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed ReasoningpaperCan We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steeringreposileod/reasoning-corepaperModality-Driven Search with Holistic Trace Judging for ARC-AGI-2
Covers (incoming)
Related across the graph
paperBetter Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed ReasoningpaperRS-RIE-Bench: Benchmarking Reasoning-Guided Remote Sensing Image Editingreposileod/reasoning-corepaperToken Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought ModelspaperThinkProbe: Beyond Accuracy -- Structural Profiling of Open-Ended LLM Reasoning Traces via Non-Generative Thought GraphspaperModality-Driven Search with Holistic Trace Judging for ARC-AGI-2paperCan We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering
