Grpo
14 items across the graph — tagged with Grpo.
From the graph · 14
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3…
Solve Visual Understanding with Reinforced VLMs
Implement a reasoning LLM in PyTorch from scratch, step by step
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
A systematic AI Agent development tutorial covering LLM agents, RAG, tool use, memory systems, multi-agent systems, LangChain, LangGraph, MCP, and agentic RL.|从…
RL environments + evals for AI agents. Define once, train anything.
An easy-to-use, fast toolkit to scale up RL post-training on a single node.
从 MiniMind 源码读起,再延伸到现代大模型技术体系的中文学习笔记。主线逐行精读预训练 / SFT / DPO / PPO / GRPO 与训练机制;附录 17 篇进阶卷覆盖量化、投机解码、RLHF 全景、模型代际史等 MiniMind 没涉及、但进阶绕不开的主题。
🦀 Decoder-only LLM built from scratch in pure Rust using Candle — no Python, no PyTorch. Gated DeltaNet + sparse attention, fine-grained MoE, native video/docu…
Curated, opinionated index of post-R1 LLM × Reinforcement Learning. Many deep-dive blog posts cross-linked to many papers — GRPO, DAPO, DPO, PPO, RLHF, GSPO, CI…
Procedural data generators for verifiable reasoning, synthetic pretraining, post-training, evaluation, and RL.
🦀 Decoder-only LLM built from scratch in pure Rust using Candle — no Python, no PyTorch. Gated DeltaNet + sparse attention, fine-grained MoE, native video/docu…
