AI, decoded

Do you have to centralize your data before AI agents can use it?

No, and Starburst's Jitender Aswani argues that moving everything into one place stops scaling as sources multiply. A traditional enterprise runs 52 to 200 data sources and the count keeps climbing, so a plan that ends with everything in one place keeps chasing new sources. Federation puts a query layer over the data where it already lives, which he thinks is the model that keeps working as the number of sources grows.

· Chain of Thought

Level 3: Leading with AI

Enterprise AIAI InfrastructureRAG & Retrieval

Scattered database cylinders under one glowing query layer, beside the hook: 52 to 200 data sources. Query them where they live.

1. The source count is the thing that broke centralization

Jitender Aswani, SVP of Engineering and Security at Starburst, puts the number plainly: “A traditional enterprise will have anywhere between 52 to 200 data sources.” His account of how they got there is ordinary rather than dramatic. Every SaaS tool and cloud service a company buys generates and stores its own data, so the source count grows with every tool a company adds.

Centralization was, in his words, “a fine, fine strategy. when the data was growing not as fast as your ability to move it.” That condition stopped holding. His framing is that federation “is the only model that scales with entropy,” a claim he states flatly in the episode’s cold open and hedges to “I think” when he makes it in full.

2. Moving the data does not end the work, it changes it

The pitch for a lake was that once everything lands in one place the problem is solved. Aswani’s read is the opposite: “Your problem starts to compound when you have moved all your data. Now you have a governance problems. you have this continuous ETL challenges.”

The mechanism is that data is not static. A SaaS vendor adds a column, and the pipeline that imported that table has to be retooled; until it is, someone is debugging why it fails. He describes leaders “spending days debugging why my pipelines are failing” as the steady state, not the incident.

3. Agents made the bill arrive sooner

Agents query differently from people. Aswani describes them operating in a reason-and-act loop, “generating queries at insane speed,” including unbounded ones, because an agent that lacks an answer fires another query rather than stopping to ask whether it should. When Starburst shipped its own MCP server, “we saw the query volume go through roof,” and customers had to scale compute behind it.

The failure he offers is his own team’s. Their internal FinOps agent had access to AWS cost data but not GCS or Azure. Asked to break spend down by cloud provider, “it starts fighting queries. It’s just frantically fighting queries without stepping back and thinking, do you even have GCS and Azure data?” That agent did not report its partial access. It worked harder.

4. Federation is a compute layer, not another copy

The alternative Aswani describes is to accept the fragmentation and move the compute instead: “the best way to live with that challenge is to actually put a virtual compute layer on top through data federation.” The data stays where it is; the query engine reaches across.

He extends the same argument to context. Knowing where the data lives is not sufficient, because meaning is scattered too: whether “customer” means account or opportunity, whether the fiscal year starts in February or July. His term for the layer that holds this is a context graph, which he distinguishes from a knowledge graph: the knowledge graph carries entities and relationships, and the context graph adds the business rules on top of them.

Why it matters

If a data-consolidation program is still underway when the agent pilot starts, the pilot does not wait for it. If the honest answer to “which systems can this agent reach” is a subset, decide that deliberately and tell the agent. Aswani’s FinOps agent did not work it out on its own. It just kept querying.

Federation solves reach, not authority. The separate question is how identity, source-level policy, and action controls combine once an agent can query across systems. See agent identity and permissions for that ownership model.

Centralize or federate: move the compute, not the data Left, centralization: every data source is piped by ETL into one lake. When a SaaS vendor adds a column, the pipeline breaks and has to be retooled, and once the data is moved you still have governance problems and continuous ETL. Right, federation: the sources stay where they are, a virtual compute layer queries across them, and a context graph on top holds business meaning such as what customer means, so an agent can query through it. Below: federation solves reach, not authority. Jitender Aswani's example from episode 66 is a FinOps agent that only had AWS data; asked for spend by cloud provider, it kept firing queries instead of asking whether it had GCS and Azure data at all. Move the compute, not the data A traditional enterprise runs 52 to 200 data sources, and the count keeps climbing. CENTRALIZE · MOVE EVERY SOURCE ⋮ one lake A vendor adds a column:retool the ETL, debug for days Once it is all moved, you still have governance problems and continuous ETL. FEDERATE · QUERY WHERE IT LIVES Agent Context graph: what “customer” means Virtual compute layer: queries reach across The data stays where it is. The query engine reaches across the sources. Federation solves reach, not authority. Tell the agent what it can’t reach. Jitender Aswani, Starburst, ep 66: a FinOps agent with only AWS data, asked for spend by cloud provider, kept firing queries instead of stepping back to ask whether it had GCS and Azure data at all.
Move the compute, not the data. The 52-to-200 range and the FinOps agent that kept querying data it did not have are Jitender Aswani’s, from episode 66. Download the image

Hear it from the guest

“So, up until a point, centralization works fine. But as the data volume continues to explode, the ETL tech debt starts to really accumulate.”
“And sometimes I've seen that like after 30 minutes, they're still doing it. And I know you are wasting my tokens. Just stop. I sent you on a wild goose chase and you did go on that wild goose chase.”

Quotes lightly edited to remove filler words.

From the conversation

This explainer is drawn from these episodes — each carries its full transcript.

Concepts in this explainer

Context GraphKnowledge GraphData FederationModel Context Protocol (MCP)ReAct (Reason + Act)AI Agent