Inference Throughput
Inference throughput is the amount of work a model-serving system completes per unit of time, such as output tokens per second or requests per minute. The unit and workload must be specified for the number to be useful.
Imagine a server completes 20 requests in a second, each producing 50 output tokens. Across those completed requests, that is 1,000 output tokens per second. It does not mean any one user received 1,000 tokens in a second, or that every request finished within one second of arrival.
Processing more requests together can increase aggregate throughput while increasing an individual user’s wait. Report throughput alongside latency, concurrency, input and output lengths, and the hardware and serving configuration. Distinguish generated tokens from input tokens processed; they involve different work.
This matters when sizing a service. A batch summarizer may care about documents completed per hour, while chat users care about time to the first token and the rate of the streamed answer. Measure on representative requests and under the response-time limits your application needs. A capacity number alone says nothing about answer quality.
Sources
- vLLM: Metrics — Documents token counters, completed requests, queue time and per-request latency measurements.
Go deeper
- vLLM: Benchmark CLI docs
Benchmark serving capacity together with request rate, concurrency and latency constraints.