Token
A token is a unit of text a language model processes, such as a whole word, part of a word, or punctuation. Token counts help determine how much text fits in a request and how usage is charged.
Tokens are the pieces; tokenization is the process that splits text into them. A word can occupy one token or several. The research by Sennrich and colleagues explains one reason to use smaller pieces: a system can represent unfamiliar words by combining subword units.
Example: an operations team wants to summarize a long incident report. The report, instructions, and answer all need space within the model’s limits. Counting pages or words alone will not establish whether the request fits; use the selected model’s token counter and check its input and output limits.
For services priced by tokens, input and output can have different rates. Some reasoning models also generate internal reasoning tokens that count toward usage even though they are not part of the visible answer. OpenAI’s guide documents these distinctions for its models; check the corresponding rules for the service you use.
When comparing costs, measure complete tasks with representative documents. A short answer does not tell you how much input or internal reasoning the task required.
See AI vocabulary in plain English for how tokens fit alongside context, prompts, and the other terms.
Sources
- OpenAI: Understanding and counting tokens — Tokens can represent words, word parts, or punctuation; counts depend on the model and encoding. Explains context limits and input, output, and reasoning usage. Checked October 6, 2026.
- Sennrich, Haddow and Birch: Neural Machine Translation of Rare Words with Subword Units — Describes representing rare words as sequences of subword units. Supports the word-piece concept, not vendor pricing or context limits. Checked October 6, 2026.