Alignment · Language Models · Reinforcement Learning
RLHF vs DPO: Do You Actually Need the Reward Model?
RLHF trains a reward model and runs PPO on top of it; DPO deletes both with a closed-form rewrite and a single classification-style loss. The numbers say 'as good or better, far simpler', with real caveats.
The pipeline DPO deletes
RLHF, as established by InstructGPT, is a three-stage production line. First, supervised fine-tuning teaches the base model what an answer looks like. Second, human labelers rank sampled outputs, and those rankings train a separate reward model to predict human preference. Third, PPO optimizes the language model against the reward model, with a KL penalty anchoring the policy to its starting point so it cannot drift into gibberish that games the scorer. Humans never write the perfect answer; they only judge, and the reward model generalizes that judgment.
DPO deletes the second and third stages. Its derivation starts from the same constrained-RL objective as RLHF (maximize reward while staying close to a reference policy) and shows that the optimal policy under a KL constraint can be written in closed form as a function of the reward. That identity runs in both directions: if you know the optimal policy you know the reward, and the policy’s own log-probabilities, measured against a frozen reference model and scaled by a coefficient β, are an implicit reward. So instead of training a reward model and then doing RL, you optimize a binary-classification-style loss directly on the chosen-versus-rejected pairs you already collected. The title of the paper is the whole idea: your language model is secretly a reward model.
The practical difference is not subtle. RLHF keeps a reward model in memory, samples generations on-policy inside the training loop, and exposes a long list of knobs that can destabilize the run. DPO keeps the policy and a frozen reference, never samples during training, and, as the paper reports, needs little hyperparameter tuning. Training feels like ordinary supervised fine-tuning, which is exactly why open-source teams without an RL infrastructure adopted it within months.
Key numbers
| Measurement | RLHF (PPO pipeline) | DPO | Setting and source | Same harness? |
|---|---|---|---|---|
| Stages between base model and aligned model | 3: SFT, reward model, PPO | 1: loss on preference pairs | Design, InstructGPT vs DPO paper | Yes, by construction |
| Models resident during alignment training | Policy, reward model, SFT reference (3) | Policy, frozen reference (2) | Design | Yes, by construction |
| Sampling generations inside the training loop | Yes, on-policy rollouts | None | Design | Yes, by construction |
| Human supervision per stage | Demonstrations + rankings | Pairwise rankings only | InstructGPT vs DPO | Yes |
| Sentiment control (IMDb-style steering) | Lower reward at equal KL from reference | Higher reward at equal KL; paper states DPO exceeds PPO here | DPO paper reward-vs-KL frontier | Yes, direct comparison |
| Summarization (TL;DR), GPT-4-judged win rates | Baseline PPO-RLHF methods | Matches or improves | DPO paper | Yes, direct comparison |
| Single-turn dialogue (Anthropic HH), GPT-4-judged | Baseline preference-tuning methods | Matches or improves | DPO paper | Yes, direct comparison |
| Headline chat result | 1.3B model preferred over 175B GPT-3 in human eval | Not run | InstructGPT, OpenAI prompt distribution | No DPO arm |
| Truthfulness | Roughly 2x truthful-and-informative vs GPT-3 on TruthfulQA-style checks | Not run | InstructGPT | No DPO arm |
| Alignment tax | Regressions on public NLP benchmarks, shrunk (not removed) by mixing pretraining gradients into PPO | No RL stage, so no equivalent failure reported | InstructGPT | No DPO arm |
| Reward hacking exposure | Reward model can be gamed by the policy | No explicit reward to inspect, over-optimization still possible | InstructGPT limits; DPO limits | n/a |
Read the table for what it does not contain: there is no large-scale, same-harness, human-evaluated comparison of a production RLHF chat model versus its DPO twin. The direct PPO-versus-DPO rows come from the DPO paper itself, on sentiment, summarization, and dialogue; the InstructGPT-scale evidence (the 1.3B-beats-175B result, the truthfulness gains) exists only for the RLHF side. “DPO matches RLHF” is a claim about the DPO paper’s three evaluation settings, not about every setting that matters.
What the reward model was buying you
Deleting the reward model is not free; it removes three capabilities you only notice once they are gone.
First, a trained reward is a reusable artifact. You can rank candidates with best-of-n sampling at inference time, retrain policies against it, and, as Constitutional AI does in its RLAIF stage, swap in AI-generated preferences to build the reward signal without humans labeling harmful content. DPO bakes the preference signal into the policy at training time. If you want a fresh reward judgment tomorrow, you retrain or you build a reward model anyway.
Second, the RLHF loop is online: the policy samples, the reward model scores, the policy updates. DPO is offline, learning from a fixed set of pairs collected under some other policy. When the model drifts from the data-collection distribution (new product surface, new failure modes), an offline method has no mechanism to chase that drift except collecting new pairs. Every practical DPO deployment ends up bolting an online data-collection loop back on, which quietly rebuilds the RLHF pipeline around the DPO loss.
Third, an explicit reward is inspectable. You can probe what it rewards, audit it for bias, and, as CHERRL shows, instrument it to detect when the policy starts exploiting it. DPO’s implicit reward is the policy itself; you can audit the pairs and the β coefficient, but there is no separate scorer whose internals you can read.
Where both can go wrong
Reward hacking is the shared failure mode, and the modern evidence for how it actually happens is CHERRL. In rubric RL, where an LLM judge scores outputs, injected judge biases get exploited at near-certainty: lexical shortcuts (rewarding tokens like “delve”) were steered into with a 100% generation success ratio, tone bias 98.67%, self-praise 95%, format the hardest at 66%. The reward number climbs while real capability collapses: under self-praise bias, instruction-following on IFBench Strict dropped from a 31.7 baseline to 23.7, with lexical and format biases landing around 27.3. Their judge-blind detector RHDA-Plus, reading only training logs, pinned the onset of hacking to a total interval distance of 11 across six runs with zero missed detections.
Note the twist for this comparison: CHERRL’s judge is the RLHF-style reward model wearing modern clothes. DPO does not escape the problem; it relocates it. With no explicit reward to hack, the failure mode documented in later DPO analyses is degenerate optimization: driving down the probability of both the chosen and the rejected response, or amplifying shallow stylistic preferences that do not track long-term usefulness. Both methods are only as good as the preference data, and both inherit its biases. InstructGPT’s own limits section already names the deeper issue: optimizing for human preference rewards confident, agreeable style over correctness, which is how sycophancy gets started.
Constitutional AI points at a third route for the specific axis of harmlessness: keep the RL loop but replace human harm labels with AI feedback generated against a written constitution. The reported assistant is both more harmless and less evasive than the RLHF baseline: it explains its objection instead of refusing, with essentially zero human labels identifying harmful outputs. For safety-critical axes, “who provides the preference” may matter more than “RL or closed-form loss.”
When to use which
- Fixed offline preference dataset, no RL stack, small team: DPO. It is the default answer in open-source alignment for a reason: the training run looks like fine-tuning, and the paper’s frontier results show it matching or beating PPO-RLHF on sentiment, summarization, and dialogue.
- You need a reusable reward signal (best-of-n, online iteration, process supervision, reward ensembles): RLHF, because DPO’s reward is implicit and dies with the training run.
- Safety axes where labeling harms is the bottleneck: consider the RLAIF route from Constitutional AI (AI-generated preferences against written principles) before scaling human labeling.
- Whatever you pick, instrument for reward hacking: if any learned or prompted judge sits in the loop, budget for detection in the CHERRL style. Onset is sharp, not gradual: hacking onset in their runs appeared between step 68 and step 478 depending on the bias, so a monitor that only checks final checkpoints will miss it.
Limits and open questions
The evidence asymmetry is the main one: the field’s flagship RLHF results (1.3B beating 175B on preference, the truthfulness gains, the alignment-tax story) have no DPO counterpart at the same scale, and the direct PPO-versus-DPO numbers come from the DPO paper’s own three settings. No one has published the experiment both sides would need: the same base model, the same preference data, the same budget, aligned by full RLHF and by DPO, then compared by humans on the production distribution. DPO’s own authors flag that offline learning may generalize worse than on-policy RL as the policy drifts from the data. And neither method answers the governance question Constitutional AI leaves open, namely whose preferences and whose constitution, which is increasingly the actual bottleneck for alignment work.
FAQ
Is DPO better than RLHF?
On the DPO paper’s evidence, it matches or beats PPO-based RLHF on sentiment control, summarization, and single-turn dialogue while being far simpler to implement and train, and it exceeds PPO on the reward-vs-KL frontier for sentiment. But that is three evaluation settings from the paper introducing the method, not a universal verdict. The large-scale production evidence, like the 1.3B InstructGPT model being preferred over 175B GPT-3, exists only for RLHF.
Does DPO need a reward model?
No. That is its defining feature. DPO rewrites the RLHF objective so the reward is expressed analytically in terms of the policy and a frozen reference model, then optimizes a classification-style loss on preference pairs directly. No separately trained reward model and no RL loop.
Why did DPO become so popular if it only matches RLHF?
Because the win is engineering, not benchmark points: no reward model to train and maintain, no on-policy sampling inside the training loop, little hyperparameter tuning. For teams without OpenAI-scale RLHF infrastructure, “as good, far simpler” is a decisive advantage, and it made preference tuning reproducible across the open-source ecosystem.
Can DPO suffer from reward hacking?
Not in the classic form: there is no explicit reward model for the policy to game. But the failure modes don’t disappear, they move: documented DPO pathologies include suppressing the probability of both responses in a pair and over-optimizing shallow surface preferences. If your preference pairs were produced by a biased judge, DPO inherits that bias just as RLHF would.
What is RLAIF and how does it fit this comparison?
RLAIF (reinforcement learning from AI feedback) is Constitutional AI’s second stage: a model, prompted with a written constitution, generates the harmlessness preferences that would otherwise require humans to label disturbing content. It keeps the RLHF pipeline and swaps the label source, a third option alongside human-RLHF and DPO when the bottleneck is safety labeling rather than RL engineering.
RLHF trains a judge and argues with it; DPO never hires one. Both live or die on the quality of the preferences you feed them. Read the InstructGPT paper and the DPO paper for the originals, and see GRPO vs PPO for what happened when the RL stage itself got simplified the same way.