Perplexity
Perplexity measures how surprised a language model is by a piece of text — lower means the model found it more predictable. It's a quick intrinsic gauge of how well a model fits a dataset, but it says little about whether the model is actually useful or correct.
Perplexity is the exponentiated average negative log-likelihood of a sequence: take the probability the model assigned to each actual next token, average the logs, flip the sign, and exponentiate (Hugging Face, Perplexity of fixed-length models). Jurafsky and Martin describe the same quantity as the inverse probability of the test set, normalized by the number of words or tokens, which is why a lower perplexity means the text surprised the model less (Speech and Language Processing, chapter 3). It applies to autoregressive models that predict the next token; Hugging Face notes it is not well defined for masked models like BERT.
It is useful during training and for comparing language models on the same held-out text. Two conditions keep that comparison honest. The test set has to be unseen, since a model that has trained on it gets an artificially low score, and the models have to share a vocabulary: Jurafsky and Martin state that perplexity is only comparable between models with identical vocabularies. In practice that means the same tokenizer. Long documents add a third detail, because a fixed context window forces you to score text in chunks, and Hugging Face recommends a sliding window, which gives the model more context at each prediction, over disjoint chunks, which report a higher (worse) perplexity than the model deserves.
The limit for builders is that perplexity measures prediction, not usefulness. The same textbook is explicit that an improvement in perplexity does not guarantee an improvement on the task you care about; only measuring the application end to end tells you that. A model can predict text well and still hallucinate, ignore instructions, or fail your use case. Treat it as a training and model-comparison signal, and judge production behavior with task evaluations. How to test an AI system covers what goes into those.