Topic

Dpo

22 items across the graph — tagged with Dpo.

From the graph · 22

model
black-forest-labs/FLUX.1-dev

Hugging Face model with 14191 likes. Tags: diffusers, safetensors, text-to-image, image-generation, flux, en, license:other, endpoints_compatible, diffusers:Flu…

model
Qwen/Qwen3.8-27B

Hugging Face model with 11540 likes. Tags: transformers, safetensors, qwen3_5, image-text-to-text, conversational, license:apache-2.0, eval-results, endpoints_c…

model
black-forest-labs/FLUX.1-schnell

Hugging Face model with 5577 likes. Tags: diffusers, safetensors, text-to-image, image-generation, flux, en, license:apache-2.0, endpoints_compatible, diffusers…

model
deepseek-ai/DeepSeek-V4-Pro

Hugging Face model with 5463 likes. Tags: transformers, safetensors, deepseek_v4, text-generation, conversational, arxiv:2606.19348, license:mit, eval-results,…

model
openai/gpt-oss-120b

Hugging Face model with 5112 likes. Tags: transformers, safetensors, gpt_oss, text-generation, vllm, conversational, arxiv:2508.10925, license:apache-2.0, eval-…

model
openai/gpt-oss-20b

Hugging Face model with 4929 likes. Tags: transformers, safetensors, gpt_oss, text-generation, vllm, conversational, arxiv:2508.10925, license:apache-2.0, eval-…

model
deepseek-ai/DeepSeek-V3

Hugging Face model with 4173 likes. Tags: transformers, safetensors, deepseek_v3, text-generation, conversational, custom_code, arxiv:2412.19437, eval-results,…

repo
walkinglabs/hands-on-modern-rl

🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.

model
WarriorMama777/OrangeMixs

Hugging Face model with 3930 likes. Tags: diffusers, stable-diffusion, text-to-image, dataset:Nerfgun3/bad_prompt, license:creativeml-openrail-m, endpoints_comp…

model
deepseek-ai/Janus-Pro-7B

Hugging Face model with 3648 likes. Tags: transformers, pytorch, multi_modality, muiltimodal, text-to-image, unified-model, any-to-any, arxiv:2501.17811, licens…

model
google/gemma-4-31B-it

Hugging Face model with 3603 likes. Tags: transformers, safetensors, gemma4, image-text-to-text, conversational, arxiv:2607.02770, base_model:google/gemma-4-31B…

model
deepseek-ai/DeepSeek-V4-Flash-0731

Hugging Face model with 3557 likes. Tags: transformers, safetensors, deepseek_v4, text-generation, conversational, arxiv:2606.19348, license:mit, eval-results,…

model
microsoft/phi-2

Hugging Face model with 3501 likes. Tags: transformers, safetensors, phi, text-generation, nlp, code, en, license:mit, text-generation-inference, endpoints_comp…

model
prompthero/openjourney

Hugging Face model with 3238 likes. Tags: diffusers, safetensors, stable-diffusion, text-to-image, en, license:creativeml-openrail-m, endpoints_compatible, diff…

repo
MakazhanAlpamys/Soup

Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.

repo
stratosphereips/StratosphereLinuxIPS

Slips, a free software behavioral Python intrusion prevention system (IDS/IPS) that uses machine learning to detect malicious behaviors in the network traffic.…

repo
codelion/pts

Pivotal Token Search

repo
Enping-Hu/minimind-deep-dive

从 MiniMind 源码读起,再延伸到现代大模型技术体系的中文学习笔记。主线逐行精读预训练 / SFT / DPO / PPO / GRPO 与训练机制;附录 17 篇进阶卷覆盖量化、投机解码、RLHF 全景、模型代际史等 MiniMind 没涉及、但进阶绕不开的主题。

repo
hscspring/rl-llm-nlp

Curated, opinionated index of post-R1 LLM × Reinforcement Learning. Many deep-dive blog posts cross-linked to many papers — GRPO, DAPO, DPO, PPO, RLHF, GSPO, CI…

repo
Yog-Sotho/LLM-fine-tuner

Powerful no-code LLM fine-tuner: upload data → train → deploy in minutes. Unsloth 2-5× acceleration · QLoRA/DPO/RLHF/PPO/ORPO · Reward Model training · GGUF exp…

repo
ruimalheiro/gradient-garden

Research platform for model training, evaluation, and experimentation across architectures, benchmarks, and recipes.

repo
Pedal-Intelligence/saypi-userscript

Enterprise-grade browser extension bringing multilingual voice interaction to AI chatbots (Pi, Claude, ChatGPT). Features real-time speech detection with Silero…

Related topics