Temperature
Temperature controls randomness when sampling a model's output. Lower temperature concentrates probability on more likely tokens; higher temperature spreads it more broadly. Lower temperature can reduce variation, but it does not guarantee repeatable output.
At each step a language model produces a score (a logit) for every token in its vocabulary, and a softmax turns those scores into probabilities. Temperature divides the logits by a number t before the softmax. Below 1, the distribution skews toward the most likely tokens and the tail shrinks; above 1, it flattens and unlikely tokens get picked more often (Holtzman et al., The Curious Case of Neural Text Degeneration). Near 0, sampling approaches always taking the top token.
The same paper documents the tradeoff. Lowering temperature improves quality at the cost of diversity, and always picking the most likely continuation tends to produce bland, repetitive text, while sampling from the full distribution drifts into incoherence. The authors proposed nucleus (top-p) sampling, which samples only from the smallest set of tokens covering a set share of the probability, as a way to cut the unreliable tail without fixing the candidate count.
Practical guidance follows the task. Anthropic’s API documentation recommends a temperature closer to 0 for analytical and multiple-choice work and closer to 1 for creative and generative tasks, and adds that even at 0 results will not be fully deterministic (Anthropic, Messages API reference). So low temperature reduces variation; it does not guarantee repeatable output, and it does not make a model more accurate.
The control is also changing. The same reference marks temperature as deprecated and states that models released after Claude Opus 4.6 accept only the default of 1.0. If your evaluation or extraction pipeline relies on low temperature for consistency, check what each model version supports, and get consistency from structured outputs, validation and evaluation instead. How to test an AI system covers testing outputs that vary from run to run.