Language model
A language model learns statistical patterns in language to estimate which words or tokens fit a context. Some predict the next token in a sequence; others predict hidden tokens from surrounding text.
A small next-word model can estimate how likely “number” is after “tracking your” in a collection of support messages. An n-gram model uses counts from a short history of words. A neural model learns representations that can use a much wider context.
Next-token prediction is not the only language-modeling setup. The masked training objective of Bidirectional Encoder Representations from Transformers (BERT) predicts hidden tokens from text on both sides. An autoregressive generator instead predicts successive tokens using the preceding context, which lets it continue a passage one token at a time.
These predictions make language models useful for completion and language processing. They do not, by themselves, check whether a sentence is true. A chatbot adds conversation behavior around a model, while retrieval can supply evidence beyond its learned patterns. A large language model is one kind of language model, not the definition of the whole category.
Sources
- Jurafsky and Martin: N-gram Language Models — Explains sequence probabilities and next-word prediction using word histories.
- Devlin et al.: BERT — Describes masked-token pretraining using bidirectional context.
Go deeper
- How do large language models actually work? AI, decoded · How Large Language Models Actually Work
- Why do LLMs hallucinate? AI, decoded · Why LLMs Hallucinate
- Andrej Karpathy: Building makemore video
Build a small next-character model, from counts to a trainable neural version.