Skip to content

AI Glossary · Concepts

What is Direct Preference Optimization (DPO)?

Direct Preference Optimization (DPO)

Direct Preference Optimization (DPO) fine-tunes a model on preferred and rejected responses to the same prompt, without training a separate reward model.

Also known as: DPO, direct preference optimization

· Updated · Chain of Thought

AI EngineeringModel Training

  • 1 explainer
  • 2 episodes
  • 1 guest clip
  • 2 sources

In a common RLHF pipeline, preference labels first train a reward model, then reinforcement learning uses that model’s scores to update the language model. The original DPO paper derives a loss that lets preference pairs update the language model directly, avoiding that separate reward-model training step.

Each training example contains a prompt, a preferred response and a rejected response. The original objective compares the trainable model’s response probabilities with a reference policy. Skipping a separate reward model does not mean training uses no reference model or that the preference labels establish factual correctness.

In episode 69 at 15:29, Joel Hron places DPO within Thomson Reuters’ training pipeline. He says human experts curate task datasets and evaluations for this stage, followed by agentic reinforcement learning with access to the company’s applications.

DPO describes a training objective, not a restriction to tone or writing style. What it can teach depends on the examples and preferences supplied. For builders, the practical question is whether preferred and rejected responses capture the behavior they want: a fluent answer with an unsupported claim should not win merely because it reads well. Hron’s sequence describes his team’s pipeline, not a universal requirement that these methods be used in that order.

Hear it from the guest

“The next stage is really around fine-tuning and direct preference optimization.”

Quotes lightly edited to remove filler words.

Sources

Go deeper

From the conversation