Time to First Token
Time to first token (TTFT) measures the delay from a defined request start until the first output token arrives. It describes initial responsiveness, rather than how long the complete answer or task takes.
Also known as: TTFT
Define where the clock starts: a client timer includes network and application delays, while a model-server timer may measure a narrower interval. Retrieval, queuing, and input processing can all contribute before output begins.
Example: a chatbot that says “Let me check” immediately but takes 20 seconds to produce the answer has a fast first visible response and slow useful completion. Keep both measurements rather than letting an early filler token conceal the delay.
In episode 59 at 2:20, Superhuman’s Loïc Houssier explains why adding model inference challenges a product built around very fast interactions. That discussion concerns the whole experience; its 100-millisecond product goal is not a claim that every model request finishes in that time.
Sources
- vLLM: Metrics — Defines serving-time measurements, including first-token timing.
Go deeper
- vLLM: Benchmark CLI docs
Measure first-token delay alongside throughput and other latency metrics on a stated workload.