Pre-training
Pre-training is the first phase of building a language model, and usually the most expensive, in which it learns general patterns of language and knowledge by predicting the next token across a huge body of text. The result is a base model.
Also known as: pretraining
During pre-training, a model reads text drawn from the web, books, code and other sources and learns to predict the next token at every position. Nobody labels the data; the text itself supplies the answers. Over trillions of tokens, the model picks up grammar, facts, styles and reasoning patterns. Pre-training a frontier model takes weeks to months on thousands of specialized chips, which is why relatively few organizations do it from scratch.
What comes out is a base model: good at continuing text, not yet good at following instructions. Post-training turns it into an assistant. Some companies take an open base model and run continuous pre-training on their own domain’s text before post-training it. Pre-training is also where some hallucinations start, since rare facts are hard to learn from patterns alone; see why LLMs hallucinate.
Go deeper
From the conversation
-
Thomson 1: The New $40M Legal AI Model | Thomson Reuters Joel Hron -
Most of the Web Will Never Get APIs for AI Agents | Dhruv Batra, Yutori -
The Critical Infrastructure Behind the AI Boom | Cisco CPO Jeetu Patel -
Beyond Transformers: How Liquid AI Is Rethinking LLM Architecture | Maxime Labonne -
First Code, Then AGI: Software’s Event Horizon with Poolside Founders Jason Warner & Eiso Kant