AI, decoded

What is context poisoning, and how do you stop it?

Context poisoning is what happens when an agent's context window fills with wrong, stale, or badly shaped data and the agent reasons off it. The fix sits upstream of the model: control what reaches the window, because one ordinary query against a badly shaped API can burn tens of thousands of extra tokens on data the agent did not need.

· Chain of Thought

Context ManagementAgent MemoryAI Agents

A context window overflowing with stacked pages of API results, beside the number ~30,000 and the words: extra tokens, one ordinary query. The figure is Airbyte’s Michel Tricot’s approximate count from a live demo in episode 52.

1. The failure is upstream of the model

Airbyte co-founder Michel Tricot’s thesis is that agents do not fail because the models are bad. They fail because the data feeding them is wrong. Swapping in a stronger model does nothing about that, which is why teams keep upgrading and keep getting the same class of failure. Context poisoning is the name for the input problem, and it is a data-engineering problem wearing an AI costume.

2. Raw API access is usually the cause

Tricot demonstrated it live. Gong’s API, as used in the demo, does not let you filter calls by user, so an agent asked for one person’s recent calls has no choice but to page through lists. A query as ordinary as “retrieve all calls since February 1” burned about 30,000 extra tokens, and run that way it normally takes about three minutes, most of it spent pulling pages the agent did not need. Nothing was malicious and nothing was broken. The API was simply not shaped for an agent, and the agent paid for that in context window.

3. The precision cost is worse than the token cost

The token cost is the visible half. The expensive half is that a window stuffed with near-duplicate, low-signal records leaves the model uncertain about what to prioritize. You get an agent that had the right answer available and did not weight it. That failure can look like a reasoning failure in a trace, which can send teams off to evaluate models instead of inputs.

4. Put something between the agent and the systems

Tricot’s answer is a context store: a layer that centralizes and pre-shapes data from source systems so the agent queries a store built for retrieval instead of paging an API built for a web app. The pattern he runs internally is hybrid: discovery and understanding through search over the store, and a direct API call when the task needs the freshest record. The decision to make deliberately is which reads have to be live; for those, fetch the one record by ID rather than paging through a list.

Why it matters

Most agent debugging starts at the model and works backward. Start at the input instead: log what actually entered the window on a failed run. If the answer is 30,000 tokens of paginated list, no model upgrade will fix it.

Context poisoning: the same question, two ways to fill the window The same ordinary query, retrieve one person's Gong calls since February 1, run two ways. Top, raw API access built for a web app: the API cannot filter calls by user, so the agent pages through list after list, which normally takes about three minutes and filled the context window with about 30,000 extra tokens. Bottom, a context store built for retrieval: data is centralized and pre-shaped, the agent searches it and gets the relevant calls, and the window stays much smaller. Underneath: the token cost is the visible half; the precision cost is a window stuffed with low-signal records. The hybrid pattern: search the store to discover and understand, and call the API directly only when the task needs the freshest record. Same question, two ways to fill the window “Retrieve all calls since February 1,” for one person, from Airbyte’s live Gong demo in episode 52. RAW API · SHAPED FOR A WEB APP agent API can’t filter by user page after page of calls · normally ~3 min context window at startup ≈30,000 extra tokens CONTEXT STORE · SHAPED FOR RETRIEVAL agent context store search, pre-shaped the calls it asked for, from data synced ahead of time context window at startup room left for the work Tokens are the visible cost. Precision is the expensive one. A window stuffed with low-signal records can leave the agent holding the right answer without weighting it. Tricot’s hybrid: search the store to discover and understand; call the API directly only when the task needs the freshest record. Decide deliberately which reads have to be live.
Control what reaches the window before blaming the model. The token figure is approximate, from Airbyte co-founder Michel Tricot’s live demo in episode 52; the three minutes is his estimate of the usual run time. Download the image

Hear it from the guest

“Because as everyone is saying, the problem is not in the model today, the problem is in the data that you provide.”
“If you get stuff from an API you might just get like a list of records that you might not need. And all of that is just going to pollute your context.”

Quotes lightly edited to remove filler words.

From the conversation

This explainer is drawn from these episodes — each carries its full transcript.

Concepts in this explainer

Context PoisoningContext WindowTokenizationPrecision and Recall