In a common RLHF pipeline, preference labels first train a reward model, then reinforcement learning uses that model’s scores to update the language model. The original DPO paper derives a loss that lets preference pairs update the language model directly, avoiding that separate reward-model training step.
Each training example contains a prompt, a preferred response and a rejected response. The original objective compares the trainable model’s response probabilities with a reference policy. Skipping a separate reward model does not mean training uses no reference model or that the preference labels establish factual correctness.
In episode 69 at 15:29, Joel Hron places DPO within Thomson Reuters’ training pipeline. He says human experts curate task datasets and evaluations for this stage, followed by agentic reinforcement learning with access to the company’s applications.
DPO describes a training objective, not a restriction to tone or writing style. What it can teach depends on the examples and preferences supplied. For builders, the practical question is whether preferred and rejected responses capture the behavior they want: a fluent answer with an unsupported claim should not win merely because it reads well. Hron’s sequence describes his team’s pipeline, not a universal requirement that these methods be used in that order.
Hear it from the guest
“The next stage is really around fine-tuning and direct preference optimization.”
Quotes lightly edited to remove filler words.
Sources
- Rafailov and colleagues: Direct Preference Optimization — Introduces an objective that learns from preference pairs without fitting a separate reward model; it still uses a reference policy.
- Episode 69: Joel Hron on Thomson Reuters' training pipeline — Hron describes his team's expert-curated datasets and evaluations, followed by agentic reinforcement learning; this is one team's sequence.