Topic

Grpo

14 items across the graph — tagged with Grpo.

From the graph · 14

repo
modelscope/ms-swift

Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3…

repo
om-ai-lab/VLM-R1

Solve Visual Understanding with Reinforced VLMs

repo
rasbt/reasoning-from-scratch

Implement a reasoning LLM in PyTorch from scratch, step by step

repo
walkinglabs/hands-on-modern-rl

🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.

repo
JudgmentLabs/judgeval

The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.

repo
redai-infra/Relax

An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

repo
Haozhe-Xing/agent_learning

A systematic AI Agent development tutorial covering LLM agents, RAG, tool use, memory systems, multi-agent systems, LangChain, LangGraph, MCP, and agentic RL.|从…

repo
hud-evals/hud-python

RL environments + evals for AI agents. Define once, train anything.

repo
inclusionAI/AReno

An easy-to-use, fast toolkit to scale up RL post-training on a single node.

repo
Enping-Hu/minimind-deep-dive

从 MiniMind 源码读起,再延伸到现代大模型技术体系的中文学习笔记。主线逐行精读预训练 / SFT / DPO / PPO / GRPO 与训练机制;附录 17 篇进阶卷覆盖量化、投机解码、RLHF 全景、模型代际史等 MiniMind 没涉及、但进阶绕不开的主题。

repo
AarambhDevHub/aarambh-studio

🦀 Decoder-only LLM built from scratch in pure Rust using Candle — no Python, no PyTorch. Gated DeltaNet + sparse attention, fine-grained MoE, native video/docu…

repo
hscspring/rl-llm-nlp

Curated, opinionated index of post-R1 LLM × Reinforcement Learning. Many deep-dive blog posts cross-linked to many papers — GRPO, DAPO, DPO, PPO, RLHF, GSPO, CI…

repo
sileod/reasoning-core

Procedural data generators for verifiable reasoning, synthetic pretraining, post-training, evaluation, and RL.

repo
AarambhDevHub/aarambh-ai

🦀 Decoder-only LLM built from scratch in pure Rust using Candle — no Python, no PyTorch. Gated DeltaNet + sparse attention, fine-grained MoE, native video/docu…

Related topics