AI Glossary

Time to First Token

Time to first token (TTFT) measures the delay from a defined request start until the first output token arrives. It describes initial responsiveness, rather than how long the complete answer or task takes.

Also known as: TTFT

· Updated · Chain of Thought

AI Infrastructure

Define where the clock starts: a client timer includes network and application delays, while a model-server timer may measure a narrower interval. Retrieval, queuing, and input processing can all contribute before output begins.

Example: a chatbot that says “Let me check” immediately but takes 20 seconds to produce the answer has a fast first visible response and slow useful completion. Keep both measurements rather than letting an early filler token conceal the delay.

In episode 59 at 2:20, Superhuman’s Loïc Houssier explains why adding model inference challenges a product built around very fast interactions. That discussion concerns the whole experience; its 100-millisecond product goal is not a claim that every model request finishes in that time.

Sources

  • vLLM: Metrics — Defines serving-time measurements, including first-token timing.

Go deeper

  • vLLM: Benchmark CLI docs

    Measure first-token delay alongside throughput and other latency metrics on a stated workload.

From the conversation