For a transformer using a KV cache, prefill computes attention state for the prompt. Decode reuses that state while generating subsequent tokens. Some systems split a large prefill into chunks so it can be scheduled alongside other work. Continuous batching concerns how a serving engine schedules multiple requests; prefill and decode describe phases within a request.
Example: compare a long-document request that returns one sentence with a short prompt that generates a long report. The first places more work into reading the input; the second keeps generation running longer. Requests with similar total token counts can still have different bottlenecks.
This distinction is useful when debugging a slow application. Measure time to first token separately from the rate of later output and from full task completion. Client-measured time to first token can include queueing and network delays, so it does not isolate prefill. Compare it with server phase timings. vLLM’s metrics documentation defines prefill and decode intervals separately from request arrival and queueing. Retrieval and tools can also add delays outside both model phases.
Sources
- Hugging Face: Continuous batching from first principles — Explains input prefill, incremental decoding, and their scheduling interaction.
- vLLM: Metrics — Defines request arrival, queueing, prefill and decode intervals separately; checked October 7, 2026.