AI, decoded

What is a token in AI, and why does it decide what AI costs?

A token is the unit a language model reads and writes: a word, part of a word, or a piece of punctuation. In English a token averages about four characters, or three-quarters of a word. Providers bill per token, with separate prices for what goes in and what comes out, so the cost of an AI feature depends on how many tokens each task uses, not just on the price per token.

· Chain of Thought

Level 1: How AI works · 1.2 Tokens, context windows and cost

Enterprise AIAI Infrastructure

What a token is

Language models don’t read words or letters. They read tokens, which are chunks of text from a fixed vocabulary: a common word is often one token, a rarer word is split into several, and punctuation and spaces count too. OpenAI’s rule of thumb for English is that one token is about four characters, or about three-quarters of a word, so 100 tokens is roughly 75 words.

Everything a model handles is measured in tokens. The context window, the most text a model can consider at once, is a token limit. In a chat, every new message is added to the conversation, and unless the application trims or summarizes it, the history goes back to the model as input on the next turn, so the token count climbs as you go. Anthropic’s documentation notes that bigger is not automatically better: as the token count grows, accuracy and recall degrade, which it calls context rot.

How tokens turn into a bill

Model APIs charge per token, and they price the tokens you send (input) separately from the tokens the model writes (output). Subscriptions and per-seat plans hide this from the user, but the provider’s costs still run on tokens. On Anthropic’s price list as of October 2026, output costs five times as much as input for each Claude model. Two other details change the math:

  • Reasoning models think in tokens. The hidden reasoning a model does before answering is billed as output, so a short visible answer can carry a large bill.
  • Repeated input can be cached. Text that stays the same across requests, like long instructions, can be cached. Anthropic charges a tenth of the normal input price, or less depending on the model, to read a cached prompt, though writing it to the cache the first time costs more than normal input.

The cost of one AI task is roughly the tokens in, times the input price, plus the tokens out, times the output price, for every model call the task makes. That last part is where agents surprise people. One request to an agent can turn into dozens of model calls as it plans, uses tools, checks its work and retries, and each call typically carries a growing context.

The price per token falls, and the bill can still rise

The price of a given level of capability has dropped fast. Epoch AI’s March 2025 analysis found that the price of reaching GPT-4’s score on a set of PhD-level science questions fell about 40-fold per year. Across the milestones it measured, the rate ranged from 9-fold to 900-fold per year, and Epoch cautions that the fastest drops may not last. That doesn’t mean bills fall. Jeetu Patel, Cisco’s president and chief product officer, told Conor in episode 67 that “price per token might eventually become a vanity metric because you might have the price per token go down, but the amount of tokens that an agent uses go up because they’re long running agents.” What matters, he argues, is the price per unit of outcome.

At scale the volumes are large. Fergal Reid, Intercom’s chief AI officer, said Intercom’s Fin agent had passed 10 trillion tokens in production, which is why moving even part of that traffic to a smaller, cheaper model shows up as real margin. For teams just starting out, the problem is visibility. Jiaona Zhang, in episode 64: “I think that today a lot of organizations are not thinking about their token spend regularly and they’re getting hit with their bills.” Her example of an inefficient use is asking AI to redo the font on a slide deck when one click would have done it. She ties it to token maxxing: using AI everywhere because a mandate says to, rather than where it pays off.

What to do with this

  • Measure cost per completed task, not per call. Count every call a task makes, retries and failed attempts included, and divide by the tasks that succeeded.
  • Watch the context. Long histories, pasted documents and verbose tool results are paid for on every turn. Trimming and caching them is often the biggest saving.
  • Budget tokens like headcount. Patel expects teams to be given a set number of tokens and asked to deliver a set output from them, the same constraint as a hiring budget.

How to cut AI agent costs goes deeper on tracing and cutting the biggest drivers, and how to measure whether AI is paying off covers the return side.

Hear it from the guest

“Price per token might eventually become a vanity metric because you might have the price per token go down, but the amount of tokens that an agent uses go up because they're long running agents.”
“I think that today a lot of organizations are not thinking about their token spend regularly and they're getting hit with their bills.”

Quotes lightly edited to remove filler words.

Go deeper

“Deep Dive into LLMs like ChatGPT” by Andrej Karpathy (2025). Embedded from YouTube; all rights remain with the creator. Why it’s here: The tokenization chapter shows real text being cut into tokens and explains why models see text this way. It starts at the right moment; the rest of the video covers the full training stack.

Common questions

How many words is 1,000 tokens?
About 750 English words, using OpenAI's rule of thumb that a token is roughly three-quarters of a word. Other languages, code and unusual spellings usually take more tokens per word, and each model family splits text a little differently.
Why do output tokens cost more than input tokens?
Providers set the prices, and on Anthropic's published list for its Claude models output tokens cost five times as much as input tokens. Part of the reason is mechanical: a model can read the whole input in one parallel pass, but it has to write its output one token at a time, with another pass through the model for each one.
Is the cheapest model per token the cheapest to run?
Not necessarily. Models split the same text into different numbers of tokens, write answers of different lengths, and some spend extra tokens reasoning before they answer. OpenAI's own guidance says a lower price per million tokens does not necessarily produce a lower total cost, and recommends testing representative tasks. Compare cost per completed task, not price per token.

From the conversation

This explainer is drawn from these episodes — each carries its full transcript.

Concepts in this explainer

Context WindowTokenizationAccuracyAI AgentInferencePrecision and RecallReasoning Models