Dpo
22 items across the graph — tagged with Dpo.
From the graph · 22
Hugging Face model with 14191 likes. Tags: diffusers, safetensors, text-to-image, image-generation, flux, en, license:other, endpoints_compatible, diffusers:Flu…
Hugging Face model with 11540 likes. Tags: transformers, safetensors, qwen3_5, image-text-to-text, conversational, license:apache-2.0, eval-results, endpoints_c…
Hugging Face model with 5577 likes. Tags: diffusers, safetensors, text-to-image, image-generation, flux, en, license:apache-2.0, endpoints_compatible, diffusers…
Hugging Face model with 5463 likes. Tags: transformers, safetensors, deepseek_v4, text-generation, conversational, arxiv:2606.19348, license:mit, eval-results,…
Hugging Face model with 5112 likes. Tags: transformers, safetensors, gpt_oss, text-generation, vllm, conversational, arxiv:2508.10925, license:apache-2.0, eval-…
Hugging Face model with 4929 likes. Tags: transformers, safetensors, gpt_oss, text-generation, vllm, conversational, arxiv:2508.10925, license:apache-2.0, eval-…
Hugging Face model with 4173 likes. Tags: transformers, safetensors, deepseek_v3, text-generation, conversational, custom_code, arxiv:2412.19437, eval-results,…
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
Hugging Face model with 3930 likes. Tags: diffusers, stable-diffusion, text-to-image, dataset:Nerfgun3/bad_prompt, license:creativeml-openrail-m, endpoints_comp…
Hugging Face model with 3648 likes. Tags: transformers, pytorch, multi_modality, muiltimodal, text-to-image, unified-model, any-to-any, arxiv:2501.17811, licens…
Hugging Face model with 3603 likes. Tags: transformers, safetensors, gemma4, image-text-to-text, conversational, arxiv:2607.02770, base_model:google/gemma-4-31B…
Hugging Face model with 3557 likes. Tags: transformers, safetensors, deepseek_v4, text-generation, conversational, arxiv:2606.19348, license:mit, eval-results,…
Hugging Face model with 3501 likes. Tags: transformers, safetensors, phi, text-generation, nlp, code, en, license:mit, text-generation-inference, endpoints_comp…
Hugging Face model with 3238 likes. Tags: diffusers, safetensors, stable-diffusion, text-to-image, en, license:creativeml-openrail-m, endpoints_compatible, diff…
Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
Slips, a free software behavioral Python intrusion prevention system (IDS/IPS) that uses machine learning to detect malicious behaviors in the network traffic.…
Pivotal Token Search
从 MiniMind 源码读起,再延伸到现代大模型技术体系的中文学习笔记。主线逐行精读预训练 / SFT / DPO / PPO / GRPO 与训练机制;附录 17 篇进阶卷覆盖量化、投机解码、RLHF 全景、模型代际史等 MiniMind 没涉及、但进阶绕不开的主题。
Curated, opinionated index of post-R1 LLM × Reinforcement Learning. Many deep-dive blog posts cross-linked to many papers — GRPO, DAPO, DPO, PPO, RLHF, GSPO, CI…
Powerful no-code LLM fine-tuner: upload data → train → deploy in minutes. Unsloth 2-5× acceleration · QLoRA/DPO/RLHF/PPO/ORPO · Reward Model training · GGUF exp…
Research platform for model training, evaluation, and experimentation across architectures, benchmarks, and recipes.
Enterprise-grade browser extension bringing multilingual voice interaction to AI chatbots (Pi, Claude, ChatGPT). Features real-time speech detection with Silero…
