newsReddit r/MachineLearningTrust 52 · CommunityPublished 1mo agoLive · 1mo ago
A trained fast-weight memory: a 3M-param transformer installs never-trained rules at inference, forward-only — where test-time training transfers nothing (single RTX 3090, fully reproducible) [R]
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 52%The State-Prediction Separation Hypothesis →
- PossiblePossibly related (embedding) · 49%From SRA to Self-Flow: Data Augmentation or Self-Supervision? →
- PossiblePossibly related (embedding) · 49%Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training →
- PossiblePossibly related (embedding) · 49%CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention →
- PossiblePossibly related (embedding) · 46%CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield →
- PossiblePossibly related (embedding) · 48%MemDefrag: Latent Memory Defragmentation for Large Language Models →
- PossiblePossibly related (embedding) · 48%Systematic Evaluation of Learning Rate Scheduling Strategies Across Heterogeneous Architectures →
- PossiblePossibly related (embedding) · 47%Super Weights in LLMs and the Failure of Selective Training →
Covers
paperThe State-Prediction Separation HypothesispaperFrom SRA to Self-Flow: Data Augmentation or Self-Supervision?paperIs One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL TrainingpaperCARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear AttentionpaperCHERRY: Compressed Hierarchical Experts with Recurrent Representational YieldpaperMemDefrag: Latent Memory Defragmentation for Large Language Models
Covers (incoming)
paperSystematic Evaluation of Learning Rate Scheduling Strategies Across Heterogeneous ArchitecturespaperSuper Weights in LLMs and the Failure of Selective TrainingpaperT^2MLR: Transformer with Temporal Middle-Layer RecurrencerepoinclusionAI/AwexpaperMixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language ModelpaperReduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM InferencepaperDoes the LM Head Create a Harmful Gradient Bottleneck? A Causal Test
Related across the graph
paperThe State-Prediction Separation HypothesisrepoinclusionAI/AwexpaperFrom SRA to Self-Flow: Data Augmentation or Self-Supervision?paperCHERRY: Compressed Hierarchical Experts with Recurrent Representational YieldpaperT^2MLR: Transformer with Temporal Middle-Layer RecurrencepaperReduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM InferencepaperCARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear AttentionpaperSuper Weights in LLMs and the Failure of Selective TrainingpaperMemDefrag: Latent Memory Defragmentation for Large Language ModelspaperIs One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL TrainingpaperSystematic Evaluation of Learning Rate Scheduling Strategies Across Heterogeneous ArchitecturespaperMixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language ModelpaperDoes the LM Head Create a Harmful Gradient Bottleneck? A Causal Test
