Skip to content

AI Glossary · Concepts

What is Continuous Batching?

Continuous Batching

Continuous batching schedules inference requests so new work can join and finished work can leave while other requests are still generating. It avoids holding an entire batch open until its longest request completes.

· Updated · Chain of Thought

AI Infrastructure

  • 2 sources

In a static batch, a short answer can finish while a long answer keeps the batch occupied. A continuous scheduler can reuse the freed capacity for another request. The serving engine must coordinate prefill and decode with available cache memory. Those phases describe work within a request; continuous batching describes how work from multiple requests is scheduled.

Example: an overnight summarization job may favor processing many documents per minute, while a live voice interaction needs a tight response deadline. A schedule that works well for one can be unsuitable for the other.

Our suggested comparison is to load-test realistic input and output lengths, then look at both aggregate inference throughput and the slow tail of request latency. Continuous batching admits new work during generation. Other forms of dynamic batching may only group waiting requests before execution; a provider’s asynchronous batch API may accept a file of jobs and return results later. Check which scheduling behavior a service actually implements.

Sources