How do you make an AI feature respond faster?
Measure the full path from the user’s action to a useful result, then shorten the steps that block it. Prepare stable context ahead of time, reuse eligible cached work and run independent steps together, while checking that freshness, answer quality and total cost still meet the task’s requirements.
Level 4: How AI systems are built · 4.6 Cost, latency and safety basics
AI EngineeringAI Infrastructure
Measure the wait the user experiences
A model call is one part of an interaction. Before changing models, record when the user starts, when retrieval finishes, when generation begins and when the result becomes useful. For a voice feature, include transcription. For an agent, include tool calls and retries. Look at slow requests as well as the typical one.
In episode 59, Loïc Houssier describes the problem through a voice-to-email interaction. The feature needs the speaker’s intent, the email thread and the appropriate tone. His proposed approach is to prepare context before the dictation is ready, so every piece of work does not have to wait for the previous piece.
Move independent work off the waiting path
Illustrative example: a user opens a reply composer and starts speaking. The application could load the thread and the user’s approved writing preferences while recording, then combine them with the completed transcription. Those preparation steps do not need to wait for the words of the new message.
Draw the dependencies before making the change. Loading the existing thread can happen early. Deciding whether the dictated message promises a delivery date must wait for the dictation. Running that decision early would require guessing information you do not yet have.
Give each request a version of the context it used. If another message arrives while the user speaks, decide whether to refresh the draft context. A speed improvement is not useful if the reply ignores the latest message in the thread.
Cache the right work
Prompt caching and caching a finished answer are different techniques. Anthropic’s prompt-caching documentation describes reusing computation for matching prompt prefixes. A stable block of instructions can be eligible for reuse while a new request still gets its own generated answer. Changing the prefix affects whether a cache entry matches.
An application cache might instead hold a retrieved thread or a completed summary. For that, define the user, source version and expiration conditions. Do not reuse another user’s material just because two requests look similar.
Measure cache misses as well as hits. Compare the first interaction, a repeated interaction and one after the source changes. Houssier also identifies the tradeoff in preparing work early: the user may abandon the interaction after preparation has already incurred a cost. Include those abandoned attempts in the comparison.
Where it falls short
You cannot parallelize steps that depend on each other’s results. A shorter prompt may also omit something essential, and a faster model may fail the task. Keep a fixed set of quality checks while changing the timing.
Showing progress or streaming text can make the wait easier to understand, but measure completion separately from the first visible output. For a draft email, the useful milestone is a reviewable draft. For an action, it is the verified outcome. Pair these latency checks with cost per completed task so faster interactions do not conceal more wasted work.
Hear it from the guest
“Providing the context, I would say, for your voice and tone independently than your dictation”
“A lot of that can be pre-fetched, prepared for the LLM to be just waiting for the dictation.”
Quotes lightly edited to remove filler words.
Go deeper
- Prompt caching Explains matching prompt prefixes, cache boundaries and diagnostics for misses; useful for separating caching from answer reuse.
From the conversation
This explainer is drawn from these episodes — each carries its full transcript.