AI Glossary

Language model

A language model learns statistical patterns in language to estimate which words or tokens fit a context. Some predict the next token in a sequence; others predict hidden tokens from surrounding text.

· Updated · Chain of Thought

Model Architecture

A small next-word model can estimate how likely “number” is after “tracking your” in a collection of support messages. An n-gram model uses counts from a short history of words. A neural model learns representations that can use a much wider context.

Next-token prediction is not the only language-modeling setup. The masked training objective of Bidirectional Encoder Representations from Transformers (BERT) predicts hidden tokens from text on both sides. An autoregressive generator instead predicts successive tokens using the preceding context, which lets it continue a passage one token at a time.

These predictions make language models useful for completion and language processing. They do not, by themselves, check whether a sentence is true. A chatbot adds conversation behavior around a model, while retrieval can supply evidence beyond its learned patterns. A large language model is one kind of language model, not the definition of the whole category.

Sources

Go deeper

  • Andrej Karpathy: Building makemore video

    Build a small next-character model, from counts to a trainable neural version.