惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

H
Help Net Security
L
LINUX DO - 最新话题
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
罗磊的独立博客
宝玉的分享
宝玉的分享
博客园 - 聂微东
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Cyberwarzone
Cyberwarzone
S
Securelist
博客园_首页
Know Your Adversary
Know Your Adversary
S
Schneier on Security
雷峰网
雷峰网
L
LINUX DO - 热门话题
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Simon Willison's Weblog
Simon Willison's Weblog
Last Week in AI
Last Week in AI
P
Privacy & Cybersecurity Law Blog
Scott Helme
Scott Helme
Schneier on Security
Schneier on Security
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
P
Proofpoint News Feed
AI
AI
K
Kaspersky official blog
爱范儿
爱范儿
H
Heimdal Security Blog
S
Secure Thoughts
T
Threatpost
B
Blog RSS Feed
NISL@THU
NISL@THU
C
CERT Recently Published Vulnerability Notes
云风的 BLOG
云风的 BLOG
Spread Privacy
Spread Privacy
Microsoft Azure Blog
Microsoft Azure Blog
IT之家
IT之家
Security Latest
Security Latest
V
Vulnerabilities – Threatpost
V2EX - 技术
V2EX - 技术
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Google DeepMind News
Google DeepMind News
Vercel News
Vercel News
人人都是产品经理
人人都是产品经理
Recent Announcements
Recent Announcements
The Cloudflare Blog
T
Troy Hunt's Blog
Stack Overflow Blog
Stack Overflow Blog
MyScale Blog
MyScale Blog
D
Docker
C
Cyber Attacks, Cyber Crime and Cyber Security

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
RLHF in 2026: when to pick PPO, DPO, or verifier-based RL
saurabh naik · 2026-05-16 · via DEV Community

The famous InstructGPT result is still the cleanest argument for post-training: a 1.3B aligned model was preferred over the 175B GPT-3 base ~85% of the time on instruction-following. Alignment beat a 100x scale gap.

That number got a lot of people to implement RLHF. Most of them later ripped it out and switched to DPO. A smaller group skipped both and went to verifier-based RL.

This post is the decision tree I wish I'd had when I started: what each pipeline actually looks like in TRL, where it breaks, and which one you should reach for first in 2026. The code blocks are runnable end-to-end against open weights — pick one and you have a working stack by tomorrow.

The three-way choice

Before any code, the picture:

  • PPO RLHF — sample, score with a reward model, update with PPO under a KL leash. The original InstructGPT recipe. Powerful, fiddly, expensive.
  • DPO — collapse the reward model and the RL loop into a single supervised loss on preference pairs. Trains like SFT, no sampling loop.
  • RLVR — verifier-based RL. The reward is ground truth (unit tests pass, math answer is correct, JSON parses). No human preferences at all.

A rough rule that holds in most post-training shops I've talked to:

  • Style, tone, instruction-following → DPO by default, PPO only if you can afford on-policy sampling.
  • Math, code, structured output, tool-use → RLVR. Don't waste a reward model on something a checker can score.
  • Mixed product behavior → SFT first, then DPO, then a verifier-RL pass on the verifiable slices.

The rest of this post is the why behind that table, and the actual training code.

SFT first, always

Every pipeline below assumes you've done SFT. The SFT model is both the starting policy for the RL/DPO step and the frozen reference the KL term anchors against.

from trl import SFTTrainer, SFTConfig
from datasets import load_dataset
from transformers import AutoTokenizer

MODEL = "Qwen/Qwen2.5-0.5B"
tokenizer = AutoTokenizer.from_pretrained(MODEL)

ds = load_dataset("HuggingFaceH4/ultrachat_200k", split="train_sft").select(range(5000))

trainer = SFTTrainer(
    model=MODEL,
    train_dataset=ds,
    args=SFTConfig(
        output_dir="qwen-sft",
        per_device_train_batch_size=4,
        learning_rate=2e-5,
        num_train_epochs=1,
        bf16=True,
    ),
    tokenizer=tokenizer,
)
trainer.train()

Enter fullscreen mode Exit fullscreen mode

SFT teaches the model to imitate a fixed target. It runs out of road the moment "good" isn't a single sentence away — helpfulness, tone, "did you actually answer the question" are comparative judgments, not next-token predictions. That's the whole reason the other stages exist.

Path A: classical PPO RLHF

Step 1 — train the reward model

The reward model (RM) is a scalar head on top of a transformer: (prompt, response) → r. You train it on pairwise comparisons with the Bradley-Terry loss:

L = -log σ(r(x, y_chosen) - r(x, y_rejected))

Enter fullscreen mode Exit fullscreen mode

Translation: push the score of the chosen response above the rejected one, by enough margin that softmax probabilities match human preferences.

from trl import RewardTrainer, RewardConfig
from transformers import AutoModelForSequenceClassification, AutoTokenizer
from datasets import load_dataset

RM_BASE = "Qwen/Qwen2.5-0.5B"
tokenizer = AutoTokenizer.from_pretrained(RM_BASE)
model = AutoModelForSequenceClassification.from_pretrained(RM_BASE, num_labels=1)

ds = load_dataset("trl-lib/ultrafeedback_binarized", split="train").select(range(10000))

trainer = RewardTrainer(
    model=model,
    args=RewardConfig(
        output_dir="qwen-rm",
        per_device_train_batch_size=8,
        learning_rate=1e-5,
        num_train_epochs=1,
        bf16=True,
        max_length=1024,
    ),
    train_dataset=ds,
    tokenizer=tokenizer,
)
trainer.train()

Enter fullscreen mode Exit fullscreen mode

Warning: the RM overfits fast. Track validation pairwise accuracy, not training loss. If train accuracy keeps climbing while eval plateaus around 0.65–0.70, stop training. A slightly underfit RM is far better than a sharp one — sharp RMs are the easiest to exploit.

OpenAI used a 6B RM against a 175B policy. The RM doesn't need to be as big as the policy; it just needs to be a stable judge.

Step 2 — PPO with a KL penalty

PPO samples completions from the current policy, scores them with the RM, and updates the policy with clipped policy-gradient. The KL penalty is what keeps the run from imploding:

r_total = r_RM(x, y) - β · KL(π_θ(·|x) || π_ref(·|x))

Enter fullscreen mode Exit fullscreen mode

Drop the KL term and the policy walks off the manifold the RM was trained on, finds a strange region of token space that scores high, and produces nonsense. With KL, every step is leashed to the SFT reference.

from trl import PPOTrainer, PPOConfig, AutoModelForCausalLMWithValueHead
from transformers import AutoModelForSequenceClassification, AutoTokenizer

policy = AutoModelForCausalLMWithValueHead.from_pretrained("qwen-sft")
ref = AutoModelForCausalLMWithValueHead.from_pretrained("qwen-sft")  # frozen
rm = AutoModelForSequenceClassification.from_pretrained("qwen-rm")
tokenizer = AutoTokenizer.from_pretrained("qwen-sft")

config = PPOConfig(
    output_dir="qwen-ppo",
    learning_rate=1e-6,
    per_device_train_batch_size=4,
    mini_batch_size=2,
    num_ppo_epochs=4,
    kl_coef=0.05,        # β — start here
    cliprange=0.2,
    cliprange_value=0.2,
    bf16=True,
)

trainer = PPOTrainer(
    args=config,
    model=policy,
    ref_model=ref,
    reward_model=rm,
    train_dataset=prompt_dataset,
    tokenizer=tokenizer,
)
trainer.train()

Enter fullscreen mode Exit fullscreen mode

Three dashboards to keep open:

  • Mean reward — should rise, then plateau. If it keeps climbing past your RM's eval accuracy ceiling, the policy is hacking the RM.
  • KL to reference — should stay bounded. A spike means the policy is sprinting away from SFT. Raise kl_coef.
  • A separate judge on held-out prompts — never trust the RM as ground truth. Read samples, or score with a different model entirely.

kl_coef between 0.02 and 0.2 covers most cases. I start at 0.05 and only move it when the KL graph misbehaves.

Why PPO breaks

After a few runs the failure modes get predictable:

  • Reward hacking — the policy finds outputs the RM loves and humans don't. Karpathy's line that RLHF is "just barely RL" is exactly this — the RM is a vibe check trained on a few thousand comparisons, and the policy is a much stronger optimizer than the RM is a judge.
  • Sycophancy — if labelers preferred responses that agreed with them, the RM learns "agreement = good," and the policy agrees with factual errors. Fix the data, not the optimizer.
  • Mode collapse — the policy narrows onto a few high-reward templates. Entropy drops, and you'll see the same opener over and over at temperature 1.0.
  • Alignment tax — RLHF'd models often regress on raw capability benchmarks like MMLU. You're trading capability for instruction-following, which is the right call for chat products and the wrong one for a model used as a backbone.

Path B: DPO — skip the RL loop

Direct Preference Optimization (Rafailov et al., 2023) folds the RM and PPO into a single supervised loss directly on preference pairs:

L_DPO = -log σ(β · [log π_θ(y_w|x)/π_ref(y_w|x) - log π_θ(y_l|x)/π_ref(y_l|x)])

Enter fullscreen mode Exit fullscreen mode

No reward model. No sampling loop. No value head. Same (chosen, rejected) data as the RM stage above, plus your frozen reference policy.

from trl import DPOTrainer, DPOConfig
from transformers import AutoModelForCausalLM, AutoTokenizer
from datasets import load_dataset

policy = AutoModelForCausalLM.from_pretrained("qwen-sft")
ref = AutoModelForCausalLM.from_pretrained("qwen-sft")
tokenizer = AutoTokenizer.from_pretrained("qwen-sft")

ds = load_dataset("trl-lib/ultrafeedback_binarized", split="train").select(range(10000))

trainer = DPOTrainer(
    model=policy,
    ref_model=ref,
    args=DPOConfig(
        output_dir="qwen-dpo",
        per_device_train_batch_size=4,
        learning_rate=5e-7,
        beta=0.1,                # KL strength, same role as PPO's kl_coef
        num_train_epochs=1,
        bf16=True,
    ),
    train_dataset=ds,
    tokenizer=tokenizer,
)
trainer.train()

Enter fullscreen mode Exit fullscreen mode

DPO wins when:

  • You have static preference data and don't want to maintain an RM service.
  • You want a training run that looks like SFT operationally — same trainer pattern, same monitoring, same failure profile.
  • You don't need on-policy exploration. DPO learns from a fixed dataset; PPO can sample fresh comparisons.

PPO still wins when:

  • You can generate fresh comparisons mid-training (online RLHF).
  • The preference signal is non-stationary and DPO's frozen dataset goes stale.
  • You're running a frontier-scale RM whose inference cost is justified.

For most teams shipping post-training in 2026, DPO (or its variants — IPO, KTO, SimPO) is the default. PPO RLHF still earns its place at the top of the budget curve.

Path C: RLVR — when you have a real checker

For domains with a real verifier — math, code, structured output, tool-use success — RLVR sidesteps the reward-model problem entirely. The reward is ground truth, not a learned vibe check. DeepSeek-R1 and o1-style training are the canonical examples.

The pipeline shape:

  1. Sample completions from the policy.
  2. Run them through a checker — execute the code, check the math answer, validate the JSON.
  3. Reward = 1 if pass, 0 if fail (or a richer shaped reward if you have partial credit).
  4. PPO-update against that reward, KL-anchored to the SFT reference exactly like classical RLHF.

The big advantage is that reward hacking gets much harder. A unit test either passes or it doesn't — there's no spurious phrase the policy can latch onto. The reward signal scales with capability instead of fighting it, which is why the recent capability jumps on reasoning benchmarks came from this direction rather than from bigger RMs.

The catch: it only works where you can build a cheap, reliable checker. For "is this helpful and polite," you're still in RLHF/DPO territory.

Putting it together

The minimal mental model:

  1. SFT gives you a model that follows instructions in form.
  2. Reward modeling lets you express comparative preferences when no ground truth exists.
  3. PPO + KL tunes the policy against those preferences without letting it wander.
  4. DPO collapses 2 and 3 into one supervised step — usually the right call for offline preference data.
  5. RLVR replaces all of the above wherever you have a real checker.

If I were standing up alignment from scratch this quarter: SFT, then DPO on offline preference data for style and helpfulness, then a verifier-RL pass on the math/code/tool-use slices where checkers are cheap. PPO RLHF would only show up if I had budget for online sampling and a serious RM team to back it.

What does your alignment stack look like in 2026 — PPO, DPO, or have you moved on to verifier-based RL where you can? I'm curious which step everyone is keeping versus dropping.

If you want to go deeper: