Post-training
Post-training is everything done to a language model after pre-training to make it useful: fine-tuning on example conversations, preference training such as RLHF or DPO, reinforcement learning on tasks with checkable answers, and safety training.
Also known as: posttraining
A base model from pre-training continues text; it doesn’t reliably answer questions or follow instructions. Post-training changes that. It usually starts with supervised fine-tuning on examples of good conversations, then preference training, where the model learns from human or AI rankings of its answers (RLHF or DPO). Newer reasoning models add reinforcement learning on problems with checkable answers, like math and code.
Adapting an existing model this way is usually much cheaper than pre-training one, though large reinforcement learning runs can be expensive too. That is why companies now do their own post-training on top of open models. On Chain of Thought, Intercom described post-training a small open model to take over one high-volume task from a frontier model, and Thomson Reuters described post-training an open model on its legal expertise. How large language models actually work explains where post-training fits.
Go deeper
- How do large language models actually work? AI, decoded · How Large Language Models Actually Work
- Should you use prompting, RAG, or fine-tuning to customize an AI model? AI, decoded · Fine-Tuning vs. RAG vs. Prompting: How to Choose
- When should you use a small language model instead of a frontier model in production? AI, decoded · Small Language Model vs. Frontier Model
From the conversation
-
How Intercom Cut $250K/Month by Ditching GPT for Open Models | Fergal Reid -
Thomson 1: The New $40M Legal AI Model | Thomson Reuters Joel Hron -
Most of the Web Will Never Get APIs for AI Agents | Dhruv Batra, Yutori -
Beyond Transformers: How Liquid AI Is Rethinking LLM Architecture | Maxime Labonne -
The AI Framework Era Is Over: Why Context Is the Moat | Jerry Liu of LlamaIndex -
11 Cameras, Dozens of Mics: The AI That Reads the Room | Tormod Ree, Neat -
Hallucinations Are a Data Architecture Problem | Sudhir Hasbe, Neo4j