Model Training
Training, fine-tuning, and adapting models for real-world work.
Model training teaches an AI model from data, from large-scale pretraining through post-training and fine-tuning that adapt a general model to a specific domain, task, or product.
Training happens in stages. Pretraining on a large general corpus produces a base model, and how to spend that compute is itself a research result: the Chinchilla study trained over 400 models and found that model size and training tokens should grow together, and its 70B-parameter model, trained on four times more data than the 280B-parameter Gopher with the same compute, outperformed it across a wide range of tasks (Hoffmann et al., Training Compute-Optimal Large Language Models).
Adaptation comes next. Suchin Gururangan and colleagues showed that a second phase of pretraining on in-domain text led to gains across the four domains they tested (biomedical and computer science papers, news, and reviews), even for already broad models (Don't Stop Pretraining); the continued pre-training entry covers how teams use it. Post-training then shapes behavior, through supervised fine-tuning on demonstrations and preference methods such as RLHF.
Most teams never pretrain. Their training work is fine-tuning, and parameter-efficient methods made it affordable: LoRA freezes the base model and trains small low-rank matrices, which its authors report cuts trainable parameters by 10,000 times and GPU memory by three times compared with fully fine-tuning GPT-3 175B, with no added inference latency (Hu et al., LoRA); see LoRA. The risk to watch is catastrophic forgetting, where gains in a new domain cost general capability.
Start here
- Thomson 1: The New $40M Legal AI Model | Thomson Reuters Joel Hron
Trace the continued pre-training stage on legal content before later tuning.
- How Intercom Cut $250K/Month by Ditching GPT for Open Models | Fergal Reid
Compare LoRA experiments with the later distributed training setup.
Go deeper
3 episodes
- Thomson 1: The New $40M Legal AI Model | Thomson Reuters Joel Hron
- How Intercom Cut $250K/Month by Ditching GPT for Open Models | Fergal Reid
- Beyond Transformers: How Liquid AI Is Rethinking LLM Architecture | Maxime Labonne
Explainers on this topic
- How Reinforcement Learning Improves AI Coding Models
- How to Prepare Data for Fine-Tuning an LLM
- How to Use Synthetic Data for AI Evaluation
- What Is Model Distillation in AI?
- How Large Language Models Actually Work
- Why LLMs Hallucinate
Terms on this topic
- Algorithm
- Boosting
- Catastrophic Forgetting
- Continued Pre-training
- Dataset
- Deep Learning
- Direct Preference Optimization (DPO)
- Epoch
- Generalization
- Hyperparameter
- Knowledge Cutoff
- Machine Learning
- Model
- Overfitting
- Post-training
- Pre-training
- Regularization