newsReddit r/MachineLearningTrust 52 · CommunityPublished 7d agoLive · 5d ago
Same GRPO recipe on three from-scratch LLMs (353M/316M/672M) gave three different outcomes, with no clean relationship to scale [P]
I trained three LLMs from scratch in raw PyTorch then post-trained each one with SFT and then GRPO. Same process every time: same synthetic arithmetic curriculum, same reward function, same hyperparameters, same KL coefficient. Pre-training went as expected, the val loss went down as the model got more modern techniques (V1 to V2) and bigger (V3 being the biggest). However, GRPO hurt both V2 and V3 and I'm not sure why. Setup
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 50%modelscope/ms-swift →
- PossiblePossibly related (embedding) · 49%An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning →
- PossiblePossibly related (embedding) · 48%Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners →
- PossiblePossibly related (embedding) · 46%Evaluate a model properly →
- PossiblePossibly related (embedding) · 46%Super Weights in LLMs and the Failure of Selective Training →
Covers
repomodelscope/ms-swiftpaperAn Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM UnlearningpaperProcess Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM ReasonerstutorialEvaluate a model properlypaperSuper Weights in LLMs and the Failure of Selective Training
Related across the graph
repomodelscope/ms-swiftpaperProcess Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM ReasonerspaperSuper Weights in LLMs and the Failure of Selective TrainingtutorialEvaluate a model properlypaperAn Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning
