How do large language models actually work?
A large language model is a neural network trained to predict the next token, a word or a piece of a word, from the text that came before it. Pre-training on a huge amount of text teaches it the patterns of language and a lot about the world. Post-training on example conversations and feedback from people or other models then turns that text predictor into an assistant that follows instructions.
Level 1: How AI works · 1.1 What an LLM actually is
Model ArchitectureModel Training
It predicts the next piece of text
Strip away the chat window and a large language model (LLM) does one thing: given some text, it predicts what comes next. The text is broken into tokens, which are words or pieces of words. The model assigns a probability to every token it knows. The software picks one, usually by sampling from those probabilities, adds it to the text, and runs the model again on the longer text. An answer of a few hundred words is that loop run several hundred times, once per token.
That job is older than chatbots. Dan Klein, a computer science professor at UC Berkeley and co-founder and CTO of Scaled Cognition, traces it back to speech recognition. A recognizer could hear the sounds but needed something to decide which of several sound-alike transcriptions a person would actually say. In episode 54 he told Conor, “So language models were just there to score plausible from implausible. And that core aspect of being a plausibility box really has just scaled up and up and up.” Today the model captures much more than which words sit next to each other. Klein lists long-distance knowledge, topics, syntax and real-world plausibility, all in one model.
Pre-training: learning the patterns
The model is a neural network with billions of adjustable numbers, called parameters. Pre-training shows it a vast amount of text and nudges those numbers, over and over, so its next-token guesses get less wrong. The architecture almost every LLM uses is the transformer, introduced by Google researchers in the 2017 paper “Attention Is All You Need”. It is built around attention, which lets each token draw on the tokens before it in the text the model can see.
Pre-training is usually the most expensive stage. Meta’s paper on its open Llama 3 models describes the largest as a transformer with 405 billion parameters. What comes out is a base model. It is very good at continuing text, but it doesn’t reliably answer questions or follow instructions, because nothing has taught it to yet.
Post-training: from text predictor to assistant
Post-training teaches the base model to behave like an assistant. A common recipe starts by fine-tuning it on example conversations, then trains it further on rankings of its answers, from people or from other models. When the rankings come from people, the method is called reinforcement learning from human feedback (RLHF). Reasoning models add training on problems with checkable answers, like math and code. OpenAI’s InstructGPT paper showed how much this matters. On the prompts the paper tested, human evaluators preferred answers from a 1.3-billion-parameter model trained this way over answers from the 175-billion-parameter GPT-3, despite the trained model having 100 times fewer parameters.
Companies now do their own post-training on top of open models. Joel Hron, Thomson Reuters’ chief technology officer, describes starting from an open-source model, realigning it to the company’s standards, then continuing pre-training and post-training it on the company’s legal and tax expertise. The bet, he said in episode 69, was that open models would keep improving, “And as the tide of open source rose, our ability to post-train our unique knowledge on top of that would rise as well with that tide.” Intercom did the same on a smaller scale, post-training a 14-billion-parameter open model for one high-volume job (how to cut AI agent costs has the numbers).
What this means when you use one
Three things follow from how these models are built, and each has its own page in this level.
- The output is plausible, not checked. Nothing in next-token prediction tests an answer against the facts, which is one reason models hallucinate.
- The cost runs on tokens. On API pricing, every token read and written is billed, so long conversations and multi-step agents get expensive. See what a token is.
- Its knowledge has a cutoff date. A model’s training doesn’t cover events after its knowledge cutoff, or your private data, unless it is trained further on them or you put that information into its context. RAG is the common way to supply it in context.
One caution against taking “next-token predictor” too literally. Anthropic’s interpretability team traced Claude writing rhyming poetry and found it picks the rhyme for the end of a line before writing the words that lead there. The model is trained to produce one token at a time, but it can still plan ahead.
Hear it from the guest
“So language models were just there to score plausible from implausible. And that core aspect of being a plausibility box really has just scaled up and up and up.”
“And as the tide of open source rose, our ability to post-train our unique knowledge on top of that would rise as well with that tide.”
Quotes lightly edited to remove filler words.
Go deeper
- [1hr Talk] Intro to Large Language ModelsA general-audience talk from a founding member of OpenAI on what a model file is, how training works and where the field is heading.
- Deep Dive into LLMs like ChatGPTThe full training stack for a general audience, from pre-training data to reinforcement learning. Take it in chapters.
- The Illustrated TransformerFor readers who want to see inside the transformer architecture itself, step by step.
Common questions
- Is a large language model the same thing as generative AI?
- No. Generative AI is the broad category of models that produce new content, including images, audio and video. A large language model is the kind that produces text, and it is the engine behind chat assistants and most AI agents. Many current LLMs are also multimodal, so they can read images or audio as well as text.
- Does an LLM look things up when it answers?
- Not on its own. A bare model answers from patterns it learned in training, which mostly stop at its knowledge cutoff. Products that search the web or your company's documents do the lookup outside the model and paste the results into its context before it answers. That pattern is called retrieval-augmented generation, or RAG.
- Why does the same question get a different answer each time?
- Because the model doesn't output one fixed next token. It produces a probability for every possible next token, and the software usually samples from those probabilities rather than always taking the top one. A setting called temperature controls how much randomness that sampling allows.
From the conversation
This explainer is drawn from these episodes — each carries its full transcript.
Concepts in this explainer
Model ParametersNeural NetworkKnowledge CutoffLarge Language ModelPost-trainingPre-trainingGenerative AIMachine Learning