Cover art for Why Context Alone Isn't Enough for Enterprise AI Agents | WisdomAI CPO Kapil Chhabra

Episodes · S4 E75

Why Context Alone Isn't Enough for Enterprise AI Agents | WisdomAI CPO Kapil Chhabra

· Kapil Chhabra, WisdomAI · 1 hr 23 min

AI AgentsAgent MemoryContext ManagementEnterprise AI

Concepts in this episode

AI terms discussed here — each links to a plain-language definition.

Model Context Protocol (MCP)AccuracyModel DriftAI AgentTokenizationContext EngineeringExplainabilityContext WindowAI HallucinationFrontier Model

Chapters

  1. 0:00Do your agents have the right context?
  2. 1:56The four ingredients of trust
  3. 5:13The criticality and impact 2x2
  4. 8:12Data, context, harness: the hospital analogy
  5. 11:12What the context layer actually means
  6. 11:48Specialized harnesses: legal, support, analytics
  7. 13:17Why only 7% of data leaders have scaled AI
  8. 15:46What models can't guess: ARR, churn, fiscal years
  9. 16:45The data stack collapses into the context layer
  10. 20:55Memory vs. context
  11. 25:26Are agents the new users of software?
  12. 27:19Where humans should spend their time
  13. 28:14Commissioning an AI agent, and who verifies it
  14. 31:28Context drift and the learning loop
  15. 34:04Context is a multiplayer game
  16. 35:12Decompose, query, verify, repair
  17. 38:47Replacing a $5M analytics pipeline with federation
  18. 42:04The context development life cycle
  19. 43:29The AI context engineer
  20. 46:02Jobs are changing, not disappearing
  21. 46:59Product, people and process
  22. 51:20Who decides? Why FDEs can't own your context
  23. 52:22Data context vs. business context
  24. 54:11The benchmark: specialized harness vs. general agent
  25. 56:08Meeting users in ChatGPT, Claude and Slack
  26. 58:48Static vs. runtime context
  27. 1:00:08Harness engineering as models change
  28. 1:01:44Right-sizing AI and Live Apps
  29. 1:04:42The boring parts: governance, security, caching
  30. 1:06:181,000 dashboards, 50 human-years
  31. 1:08:08What "live" means
  32. 1:09:19Are dashboards going away?
  33. 1:12:08A pipeline app built on a weekend walk
  34. 1:15:53Who owns the apps?
  35. 1:17:27Data teams now provide context, not insights
  36. 1:18:56Building with the WisdomAI MCP
  37. 1:19:43Your context is your IP
  38. 1:20:48Closing thoughts: none of that work goes to waste

Show notes

An AI agent answering questions about your business needs to understand how that business works. How do you calculate churn? When does your fiscal year end? Who has the authority to change those definitions? Connecting a model to company data leaves those questions unresolved.

That’s the work of context engineering: giving agents the business definitions, instructions and examples they need to interpret your data. But that context changes as the business evolves, and someone has to keep it accurate.

In this episode of Chain of Thought, WisdomAI co-founder and CPO Kapil Chhabra joins Conor Bronsdon to explain how his team approaches that challenge. We explore how companies maintain shared context, why data teams are taking on the role of AI context engineers, and how a specialized harness plans queries, checks results and repairs errors before returning an answer.

Recorded at WisdomAI’s San Mateo office, this episode is sponsored by WisdomAI.

We cover:

  • Why the context layer is broader than a semantic layer or catalog, and how it differs from memory
  • How context drifts, and why a subject-matter expert has to approve changes to shared definitions
  • Why Kapil sees an AI context engineer role emerging, and why data teams are moving from providing insights to providing context
  • How WisdomAI's harness decomposes a question, federates queries across data sources and repairs errors, including a customer that replaced a $5M-a-year analytics pipeline
  • The four ingredients of a trustworthy AI answer: accuracy, consistency, governance and explainability
  • Live Apps: governed analytics apps built from a single prompt, and what keeps them live
  • Why your context is your IP and should stay portable

Chapters:
(0:00) Do your agents have the right context?
(1:56) The four ingredients of trust
(5:13) The criticality and impact 2x2
(8:12) Data, context, harness: the hospital analogy
(11:12) What the context layer actually means
(11:48) Specialized harnesses: legal, support, analytics
(13:17) Why only 7% of data leaders have scaled AI
(15:46) What models can't guess: ARR, churn, fiscal years
(16:45) The data stack collapses into the context layer
(20:55) Memory vs. context
(25:26) Are agents the new users of software?
(27:19) Where humans should spend their time
(28:14) Commissioning an AI agent, and who verifies it
(31:28) Context drift and the learning loop
(34:04) Context is a multiplayer game
(35:12) Decompose, query, verify, repair
(38:47) Replacing a $5M analytics pipeline with federation
(42:04) The context development life cycle
(43:29) The AI context engineer
(46:02) Jobs are changing, not disappearing
(46:59) Product, people and process
(51:20) Who decides? Why FDEs can't own your context
(52:22) Data context vs. business context
(54:11) The benchmark: specialized harness vs. general agent
(56:08) Meeting users in ChatGPT, Claude and Slack
(58:48) Static vs. runtime context
(1:00:08) Harness engineering as models change
(1:01:44) Right-sizing AI and Live Apps
(1:04:42) The boring parts: governance, security, caching
(1:06:18) 1,000 dashboards, 50 human-years
(1:08:08) What "live" means
(1:09:19) Are dashboards going away?
(1:12:08) A pipeline app built on a weekend walk
(1:15:53) Who owns the apps?
(1:17:27) Data teams now provide context, not insights
(1:18:56) Building with the WisdomAI MCP
(1:19:43) Your context is your IP
(1:20:48) Closing thoughts: none of that work goes to waste

Links from the episode:
Meet the Modern Data Team (WisdomAI CDO report)
AI Context Engineer (ACE) certification
Live Apps
WisdomAI in ChatGPT Work
Avoid AI Writing
ssot-check
I Paid an AI Agent $8 to Write About its 'Life'
Slack Wants to Be the Context Harness for Code | CPO Jaime DeLanghe
The AI Framework Era Is Over: Why Context Is the Moat | Jerry Liu

Connect with Kapil Chhabra:
LinkedIn
WisdomAI
WisdomAI on X

Connect with Chain of Thought host Conor Bronsdon:
Newsletter
Twitter/X
LinkedIn
YouTube

More episodes: https://chainofthought.show

Thanks to WisdomAI for sponsoring this episode. WisdomAI is the agentic analytics platform for trusted enterprise intelligence: governed context, an analytics harness that makes every answer consistent and verifiable, and Live Apps built from a single prompt. Try Live Apps: https://wisdom.ai/liveapps

Transcript

150 segments

Conor Bronsdon 0:52 As companies bring AI into more business decisions, we need to know, do our agents understand the context behind the data they're exploring? Do they have the verification they need to ensure answers are correct? And will that knowledge and accuracy change or continue as business context changes alongside it? Today, I'm at the offices of Wisdom AI here in San Mateo. Thank you to them for sponsoring this episode and having me down. And I'm here to talk with Kapil, who is the CPO and co-founder of Wisdom AI. We're going to discuss software, shared context, agent analytics, and what it takes to build trustworthy AI in the enterprise for both people and agents. Kapil, it's great to see you.

Kapil Chhabra 1:39 Great to meet you, Conor. Thank you.

Conor Bronsdon 1:41 And Kapil, I know you have a lot of knowledge in this area. What are most people missing about the context needs for enterprise AI?

Kapil Chhabra 1:50 [OVERLAP] Yeah, first of all, thank you for having me here and coming into our offices today.

Conor Bronsdon 1:54 [OVERLAP] I mean, thank you for having me here. I'm showing up.

Kapil Chhabra 1:56 Wonderful. And this topic is actually pretty close to my heart. We've been building in this space for the last three plus years now and very proud to kind of have the large enterprises as customers who use the product in production with their most critical data. including data for their finance, closing their quarters, building out investor relationship decks and so on, which actually requires not just precision, but as you mentioned, trust in that data. Somebody is actually signing off on that stuff like an underwriter. So how do you do that becomes a question. And trust comes in various different flavors. At a high level, if you think about what delivers trust in an answer that is generated by a probabilistic system like an LLM, it comprises of four different high-level things. The first one has to be accuracy. The answer has to be accurate. That somebody asks the question, let's say somebody asks 100 different questions, how many of those are correct? The second aspect is consistency. That if you and I together ask the same question a hundred different times, how many of those times is the answer coming back to the same result? That's consistency. The third thing is governance. You may have slightly more access into the data than I do, which is very true for enterprises. And in that case, the system needs to honor that, the governance aspect as well. So these things actually together, oh, and the fourth thing is explainability that providing visibility into the answers when it is generating and being able to trace back into those answers and see where exactly what thing was delivered in terms of like that audit becomes very important. So a combination of all of these four things delivers the trust and coming back to your question on what is the importance of context in all of these it delivers that accuracy without the context the models are just flying blind

Conor Bronsdon 4:13 And a couple, first of all, thank you so much for having me here at the office today. And thank you to Wisdom.ai for sponsoring this conversation. In particular, I'll say explainability is obviously very close to my heart. The podcast is called Chain of Thought for a reason. We deeply care about it. And as we've talked about on this show multiple times, there are varying needs around which of these dimensions is most important to a specific enterprise. And there are varying needs around what a startup may care about or what an individual may care about in here. And there's a reason that we have to architect our context and our harnesses and all the data that is pouring into these models. differently depending on the use case. So what changes when a company is, let's say, moving from, hey, I'm using AI for individual productivity, it's helping us write some emails, maybe it's helping me do drafts, to, okay, now AI is analyzing and ordering information that may actually cause the business to make crucial decisions.

Kapil Chhabra 5:13 [OVERLAP] Yes. Yeah. So think about it this way. If there's like a two by two on the y-axis, let's draw criticality of a use case.

Conor Bronsdon 5:24 [OVERLAP] So we should put a whiteboard back here.

Kapil Chhabra 5:26 And on the x-axis, let's draw the impact that it can have. So now with that two by two, the top right will have the most critical high impact kind of use cases. If you start to think about those use cases, they're not about summarize this document or help me write this email, et cetera. They're more about, I want to optimize my supply chain in the organization. How do I do that? I have a financial planning and analysis task to run and I want to do flux analysis on my data or I want to do anomaly detection on this data and then go into root cause analysis. Now, all of those are more business critical and more high impact type of use cases. If you start to think about those use cases, you need all of these four aspects of explainability, accuracy, consistency, and governance for those to actually bring up the trust. Now, individual users may not have those needs, right? We are all happy writing those emails and summarizing, maybe doing design work and things like that with general purpose AI, which is fantastic. We are pretty big users of Cloud, Codex, everything within WisdomAI as well. But for all the data analysis and high critical tasks, we all internally and our customers use WisdomAI for that reason.

Conor Bronsdon 6:55 And I think an important element of this conversation is also the fact that there may be a one-off need. Let's say I could use a cloud artifact for a one-time send to a colleague. But if this is something that is going to be consistently used throughout the business, I have to ensure that updates are made, that there's accuracy, that they hit all these explainability elements and actually build that trust you talked about. There's such a need for trustworthiness in enterprise AI, particularly depending on the use case. So I know if you're working with a bank, for example, there may be much higher trust needs than if you're working with a dev tool startup. And I love my dev tool startups. I know a lot of you work there. But a bank is actually moving customers' money around. They can't afford to have a non-deterministic and a deterministic system that can flit. They can't afford to have, you know, only 89 of those hundred questions you asked off of it earlier actually work, let's say, if you're trying to move money around. So, I mean, this illustrates what you're talking about here with the importance of trust and the need to have these different vectors aligned. What's the right way for folks who are on the enterprise side to be thinking through the architecture behind that trustworthy enterprise app?

Kapil Chhabra 8:12 That's such a good question because the whole architecture is evolving as we speak right now.

Conor Bronsdon 8:18 [OVERLAP] Yeah, by the time I pull this out, it may have changed again.

Kapil Chhabra 8:21 [OVERLAP] That's very likely. But what we have seen is that there is gravity towards a few different layers that have clearly shaped out of this architecture. At the bottom, there is the enterprise data. And we all know that we are all here to kind of use the LLMs and these large language models or small language models in some cases. So those things are constant. Now on top of that, it is also starting to get very well established that we need to bridge the gap between the LLM's understanding of the world and their understanding of the company's data with context. So the context layer kind of find its way into that architecture. But the other piece that is actually very little talked about these days, and I am 100% sure this will become a big deal and everybody will start talking about it, is a specialized harness on top of that. And that harness is specialized for a given task. In our case, we are building an analytics specialized harness on top of that context layer, which is what then powers the front end surfaces, be it WisdomAI's UI, or a cloud or a chat GPT or a customer's own internal chatbot. A good way to think about this is with an analogy of something that we all are aware of, like hospital system. Think about the surgeon as the LLM, the model. They're highly special, they're trained at their job. They're specialized for doing that kind of thing, surgery, which is pretty complicated. But they're not going to take a patient and start operating right away on any table.

Conor Bronsdon 10:03 [OVERLAP] I hope so.

Kapil Chhabra 10:04 Yes, exactly. So what's missing there? First of all, what is it that they need? They need information about that data. The patient history becomes very critical. That patient history is the context. Based on that, the surgeon knows how to operate on this patient. But that alone isn't sufficient either because you cannot take a surgeon with the patient history and start operating in this room. Right? That wouldn't work either. So you need a specialized harness around that, which is the hospital system that provides it. And the hospital system is giving things like, here is the chart, here is the checklist, here is all the instrument and the equipment. And the back office is taking care of all the insurance needs and bringing in any other specialty whenever needed. So all of that, think of that as the harness. Now it's together, this hospital system being the harness, the context, which is the patient history, and the surgeon, which is equivalent of the LLM, coming together to make sure that the patient who's walking into a hospital is going out healthy.

Conor Bronsdon 11:12 I really like this analogy. I think it helps illustrate distinctly what the different elements of this context layer and harness around it are. But I want to drill down a bit more into what it may look like for a specific enterprise. So this healthcare example is great, but when we say the context layer, when we say agent analytics, what does that actually mean here? Because I think we have a lot of conversations, particularly in software development and a lot of folks listening to the show, where we maybe are referring to something else. And let's be clear on definitions.

Kapil Chhabra 11:48 Okay, good point. Let's push on this analogy a little bit as well and then tie it back into this enterprise use cases. So let's take the example of a maternity ward and oncology ward. You would not take a pregnant woman who's about to deliver a baby into an oncology ward to do that. Now, this is the same hospital system, but it's a specialized harness. One is for maternity labor and delivery. That is what they specialize in. The other one is for oncology, which is treating cancer patients. So these are specialized harnesses. Now, let's use that same analogy and see what it means for the technology world. If you look at Departments like Legal, Harvey and Legora, these are the companies that are specialized harnesses for legal use. Sierra, Decagon, et cetera, are specialized harnesses for support use cases and customer success. They are trained, those harnesses, to go ahead and optimize for retention of the customer. They know when to hand off the case to a specialized human. Similarly, for whenever there is a need to access the data internally, the structured data internally, and there is a need to go ahead and apply trust and governance on those things, you need a specialized harness. That is the analytics harness that's provided by WisdomAI.

Conor Bronsdon 13:17 [OVERLAP] This is very interesting in the context of some of the data I was reading that Wisdom put out earlier this year. You have this survey called the Meet the Modern Data Team. And yes, if you're on YouTube, I've got a copy with me. But you surveyed over 200 senior AI analytics and data leaders. And you know, one of the largest sample sizes was from healthcare. Um, and you know, these are roles like VP of AI transformation, VP data science, chief data officer. Uh, folks who are making these crucial decisions are working with not just data engineers, but developers across the organization. And as you've mentioned, their needs vary widely depending upon the industry they're involved in. Uh, over 90% of them. were, and this is earlier this year, so I suspect it's higher now, were using AI to analyze data, to at least experiment with collecting this intelligence. How would those types of roles need to think about the data that comes around this harness and how to manage it and ensure it is kept up to date so that they have the right context, whether it's for the surgery decisions or for agents that may be operating within the business?

Kapil Chhabra 14:24 [OVERLAP] Yeah.

Kapil Chhabra 14:27 [OVERLAP] On the, on the survey, one thing that we found, which is very interesting is that all the 93%, as you rightly quoted, of the respondents said that they are doing this work. They're, they're accessing the, they're using LLMs and agents to access their internal data. Only 7% across the whole set were able to say that they have scaled the operations. And that's the missing gap.

Conor Bronsdon 14:53 [OVERLAP] Okay,

Kapil Chhabra 14:53 [OVERLAP] So.

Conor Bronsdon 14:53 [OVERLAP] so that points to the challenge here that you're directly addressing.

Kapil Chhabra 14:56 [OVERLAP] Exactly. So they've been able to do that. Yes. A general purpose hospital system will be able to take in a cancer patient, right? But they will, they will not be able to completely treat them. Maybe they are able to, uh, right, but they are not specialty for that. Now, same analogy. That's exactly the reason only 7% are successful. Why are the rest of these people, these CDOs and VPs of data not able to go ahead and scale these operations is because they're not able to deliver the trust in the data that is coming out of these chat systems. And the gap that needs to be filled is with this context. And I'll go into a few examples of what this context would look like and the harness on top of it. To take some examples there,

Kapil Chhabra 15:46 [OVERLAP] let's use an example of a B2B software company that is selling their software to other businesses. One of the things that they may want to track is ARR. And every company has a slightly different definition of ARR, as we all know. And there is no way that out of the box, LLM would know what the definition of the ARR is for that company, or the definition of churn for that company.

Conor Bronsdon 16:13 [OVERLAP] or they may not know when the fiscal year ends, because that varies. And that, of course, impacts those IAR measurements.

Kapil Chhabra 16:19 [OVERLAP] That's exactly right. So these right between us right now, in the last 10 seconds, we've come up with three examples, the ARR, the churn definition, the fiscal year definition. All of this, when provided to the LLM, to the model, the model will then know without guessing. So it reduces the chances of hallucination there. So that is the context that needs to be fed into the model for it to work properly.

Conor Bronsdon 16:45 And a lot of this comes down to a change in how the data stack works as well. We've seen, and I think I saw this data in your survey as well, that there has been a collapse around this legacy data stack that is now being more broadly fed into the context layer. We're seeing things, you know, the modeling and prep, the semantic layer, governance that were previously distinct buckets now, we need to make sure our agents have access to those so that they can make the right decisions as you point out. What are other changes that people should be aware of around how the data layer is changing that feeds into this whole contextual system and harness?

Kapil Chhabra 17:25 Yeah, I keep getting the same question in few different flavors. Like how is the context layer different from a semantic layer? Or how is, what's the role of a catalog in this new world? They're all kind of, these questions are pointing to the same direction on what is this new stack? What is this amorphous category of context that is getting formed and what's in that? So let me kind of try to tackle that from all of these different angles.

Kapil Chhabra 17:57 Traditionally, the way we as data teams, we started explaining the data for business use cases was, let's start with a semantic definition. This is where the definition of a churn or ARR in our example would go and sit in. when we provided that to the bi tools because bi tools were the primary consumers of the semantic layer and we provided that the bi tools which are deterministic in nature would actually go ahead and use those in the specific format so they were designed for a deterministic bi layer now there is a new different world in this world there is a chat agent user can ask any question without having pre-formatted or determined by the data team And now in that world, is the semantic layer sufficient? It is necessary, but it is not sufficient. There are a few different gaps that it creates. We definitely need those definitions of ARR in turn, but how would you encode the example that you said, which is that my fiscal year ends on January 31st rather than December 31st in a particular context. And when somebody is saying, asking questions about Q1, the finance team might mean it's a financial quarter versus the marketing team might mean that it is the calendar quarter.

Conor Bronsdon 19:22 This is a headache I have dealt with so much. So you're, you're giving me a little bit of PTSD here.

Kapil Chhabra 19:28 I know this is, so we are dealing with these non-deterministic systems who are trying to reason on the fly. The reasoning hasn't been baked into the BI layer already. Now in that case, The semantic layer requires a lot more to be fed in. And that is why we call it context. I'll give you examples like this, this example of putting in the calendar and the fiscal quarter difference, and depending on who's asking what, how should you treat it? That's just a natural language instruction that when fed into the LLM deterministically, every time somebody is asking a question about calendars and quarters, That is what makes that more accurate. The other thing that goes into this, to the context, is examples of what good looks like. There are so many nuances when somebody, some analyst is writing an SQL statement or a Python query. It's not just the generic query that will fetch the results. Sometimes there are nuances, as you would know, that feed into that. And how would you capture those nuances and tell that to the LLM is through examples. instead of doing a single short answer generation, you feed in those examples as well. And that also needs to sit in this context layer. So context layer is broader than a catalog and a semantic layer. And it needs to be built and needs to be kept updated as the business evolves.

Conor Bronsdon 20:57 How do you think about the differences between memory and context? Because I think this is another area where we're seeing a lot of conversations that may have overlap or unclear definitions that could

Kapil Chhabra 21:11 [OVERLAP] the terms are being used interchangeably. And rightly so. So let's peel the onion a little bit more. When you and I send an instruction, a prompt over to the LLM, what we are sending is some piece of text. And that piece of text goes in this thing called a context window. and we all know that context windows are various different sizes depending on the model and so on so that is the amount of text that can go into the prompt now if you put in a lot of text let me put the whole company history and all the documents in that context window and let's expand the context window we are actually not optimizing. What we are doing is we are letting the LLM again take a guess on what out of that is the right piece of context to use. And we are also burning the tokens. So it's not optimized. If we send less context in that window, even then we are leaving room for the llm to make guesses that is where the hallucination comes in so the trick is to actually put in the just the right amount of context given every prompt so now there is the prompt there is a context that has been injected into that and that is what will produce the right results now coming back to the question of the memory Now, who produces this context and how it gets used is where the distinction between the context layer and the memory layer comes in. This information that we were just talking about, ARR definition or the definition of churn, it does not vary. This is our company's truth. That's organizational context. Somebody who's in the position of authority, let's say an SME in this case, can go ahead and put that in. It's not something that the model has to learn on the fly. But yours and mine preferences as we start to use the system, I personally do not like pie charts. I'd much rather have

Conor Bronsdon 23:15 [OVERLAP] Okay,

Kapil Chhabra 23:15 [OVERLAP] a donut

Conor Bronsdon 23:15 [OVERLAP] interesting.

Kapil Chhabra 23:15 [OVERLAP] chart, right? Just my preference.

Conor Bronsdon 23:18 [OVERLAP] Fair enough.

Kapil Chhabra 23:19 Now, that does not impact how the company operates or when somebody else asks a question, they also may not like pie charts. Just that doesn't make sense. So that is my personalized memory. And that instruction also goes into that context. So it's the context engineering that needs to kind of massage this information and put all of this together. But that's the distinction. My user level usage and preferences go into that memory layer. And the context is more institutional.

Conor Bronsdon 23:50 This is where I think if we look back at your examples note, there are examples that belong in the broader context layer and there are examples that belong into the individual memory layer for an individual user. One example of this from my personal software development would be that I have this expanding skill that's got a few thousand stars on it that I built called Avoid AI Writing to help rewrite AI text for much of the slop that you may experience as you try to have AI write for you. And we've built specific examples into the context layer for that skill about what good writing looks like, how it should think about this, how it should do rewrites. And then we have it call out to individualized context that we can have, memory in this case, that would be my own writing voice. Here are the examples for me personally, where I want an LLM to write in this way. This is how I like to speak. This is how I write. And I think that this is just an example from my own personal life, but there are so many instances of this. And it feels like we're just scratching the surface of how individualized AI is going to get as we see things like Muse launch, for example, and all these different personalized assistants I would anticipate that context engineering becomes even more a topic of conversation.

Kapil Chhabra 25:09 [OVERLAP] 100% even memory has, it's not just a one single block. Uh, there is short term memory. There is long term memory. There is episode memory. There's all sorts of different memories as well that play a role into this. We're just scratching the surface on all of these things together.

Conor Bronsdon 25:26 [OVERLAP] And if you look at like OpenAI or Anthropic or any of the major players, they've changed how they approach memory over the last couple of years. There have been multiple iterations, just like how we're seeing the context and analytics harness evolve as well. I'm going to pull us off track here for a minute because I'm very curious to get your opinion on this. It's something that's come up on a couple of recent episodes, and it feels relevant to the example you brought up earlier around an individual user prompting an agent that can pull from the data and context of the organization. That agent is really becoming the user of many apps that we used to use. Like I, as an individual, may not really be a user of many tools in Salesforce, for example, anymore. My agent is probably going out and doing that for me. Do you view agents as the new users of many applications?

Kapil Chhabra 26:15 Oh, 100%. That's such an important point that you're bringing up here because the tasks that humans used to do. will be done by humans, but will be augmented by agents as well. So if we are looking at this kind of task, which is access data internally in the organization, there, so far it's been us, the humans who go in, ask the question, get the answer, either by looking at a chart or in natural language or any of that stuff.

Kapil Chhabra 26:47 In the near future, there will be autonomous agents that do the same task of asking the question, get access to the data. And guess what? They're not there just to get the data out. They're there to take some action on top of that data. which means that essentially we are slowly in that use case eliminating the human in the loop from interpreting the response using the judgment whether the response is right or wrong before taking an action because the agent is going to go take an action.

Conor Bronsdon 27:19 I feel like this also speaks to the problem of where agents and humans should be operating within the stack, as we increasingly see agents being the users of system of record, the ones who are gathering data and providing it to us, who are picking up the context that we are putting out as little crumbs for them. Where should humans be spending their time?

Kapil Chhabra 27:39 The humans need to be spending time in decision-making. We are not at a point where we should let the AI make the decisions, primarily because the trust in the responses for that agent is missing. We need to learn collectively on how to deal with this probabilistic system. So far in enterprises, we have always dealt with deterministic systems, and we know how to work with those. Now introducing a probabilistic system into the mix, we are all learning on how to go about doing that.

Conor Bronsdon 28:14 totally think you're right. And I think this deeply impacts how we need to think about building these context harnesses and the layers that enterprises need to go through to achieve the trust that they need for different applications and for different agent use cases. To use a recent example, not specific to analytics, but I actually commissioned an AI agent to write an article for me recently. I had a cold email that I got from Sarah Kelly from ILANDS. It lets you spin up autonomous AI agents and put them out there and they can go commission work. And she wrote an article for me that's on my Stumpetstack about her experience of the world and how they're metering by tokens. And it's, it's the first time I have had an independent agent that I have gone to commission work from and interact with in that way. But I think we're starting to see those come into businesses already. Like this is, this is an exterior agent, but we're, that's not the first time I'm going to assess an agentic business, let alone an agent that's already within my enterprise systems. So how do we do the verification layer that is so important to ensure that we get that trust?

Kapil Chhabra 29:27 [OVERLAP] Okay, question back to you first before

Conor Bronsdon 29:28 [OVERLAP] Please.

Kapil Chhabra 29:28 [OVERLAP] I

Conor Bronsdon 29:29 [OVERLAP] Yeah,

Kapil Chhabra 29:29 [OVERLAP] answer that.

Conor Bronsdon 29:29 [OVERLAP] yeah.

Kapil Chhabra 29:31 [OVERLAP] In this example, did you let Sarah go ahead and publish that article to the public? Did not, so there is the human in the loop.

Conor Bronsdon 29:38 [OVERLAP] I was. Yeah, it's a really good point. It was it was like a freelancer experience where Sarah, you know, sends me a draft. I got to send back a couple of suggestions. Sarah sends me a second draft and then I publish that draft. But I was the one who was handling the publishing, to your point.

Kapil Chhabra 29:54 [OVERLAP] Yep, exactly. So you are the one who is reviewing that content. making a judgment, whether this is good enough for you to publish or not. And your,

Conor Bronsdon 30:03 [OVERLAP] With my

Kapil Chhabra 30:03 [OVERLAP] it

Conor Bronsdon 30:03 [OVERLAP] context.

Kapil Chhabra 30:04 [OVERLAP] is your,

Conor Bronsdon 30:04 [OVERLAP] Yeah.

Kapil Chhabra 30:04 [OVERLAP] it has to have your context. It has to have

Conor Bronsdon 30:06 [OVERLAP] OK.

Kapil Chhabra 30:06 [OVERLAP] your tone of writing. It needs to know your recent, uh, short term memory as well, that you just had a conversation with me. Maybe there's one or two things that you learned through this conversation and how do you apply that into the new article, which Sarah doesn't have the context. So bringing all of those things together. If you were to get to a point where Sarah is actually going ahead and publishing that article without your involvement, you need to have trust in their writing.

Conor Bronsdon 30:35 Oh man. Uh, all right. You're, this is totally off track now, but, uh, you're, you're giving me all these ideas about commissioning a team of agents to act as, uh, you know, freelance writers for a media business. I mean, maybe we'll pursue that retain a thought. Well, when you come back on in a few months, I'll be like, guess what? Sarah and five others are doing this work for me. But I mean, this is me as an individual, like media company podcast or thinking about this idea. Businesses are much farther along here, right? They have agents who are taking actions. They have verification that's happening around the SQL queries that different agents are running. When you design that harness to ensure accuracy, to ensure that, you know, memories are being pulled in, that broader business context is being pulled in, what kind of loops are you creating within that harness to ensure that the data stays up to date instead of being right at one point and then stale a few weeks later?

Kapil Chhabra 31:28 Yeah. So there's, there's so much to unpack onto that one. You spoke about the loops. So let's start with the context and then come back into the harness as well. The context itself is not a one and done. The business is a living, breathing entity. It keeps changing. New definitions keep coming in. Churn may be good. 10% churn may be good right now. But in a few months, even 5% churn may look bad. So how do you keep up with that changing context? We call it context drift. When the definitions of something are changing and guess what? The best way to capture that drift is to tap into where the conversation is actually happening. The good news is that a lot of conversation is happening within WisdomAI platform as well. Even if somebody is using a chat GPT or a cloud, the conversation is happening via WisdomAI and we do get that visibility. What we have done is built a learning loop within our harness, which is tapping into seeing what all pieces of context have been used for a given answer. And if there's any missing piece of context or there's limited context within an answer, we highlight that upfront. Now that gives us the ability to point out what is the missing piece of context that needs to go in back to the organizational level and ask that SME saying that, hey, looks like the definition is changing here. Are you okay with refining this definition? Or you want to stick with the previous definition. And by the way, here's the evidence of why that we are, why we think the definition is changing. Maybe we saw a Slack thread on that. Maybe there's a new document that has been published, or maybe just because somebody is asking that question and saying that known right now, the definition of churn looks like a new definition.

Conor Bronsdon 33:25 Maybe there's a meeting that happened where two of your VPs decided we have to target this metric in particular. And it's so important to be able to pull in all these different sources and to your point, access humans to also weigh them. Because just because I may have said something in a Slack thread, I may not be the decision maker for this. Maybe there is an SME and a key decision maker who are actually making the decision on what churn targets we have. And if the agent starts to prioritize my individual voice where I argue against it and say, no, we should be able to have higher churn, we could have very wrong decisions that start to cascade if we are not careful.

Kapil Chhabra 34:04 [OVERLAP] Yes, and that's what makes context a multiplayer game rather than a single-player game, where there is an SME in the loop who has to be the voice of authority, give it a thumbs up, and say that, yes, this is the definition that we all agree with. If there are definitions that we don't agree with, which means that there are individuals who have their own definitions, then that's more like their individual personalized memory rather than the institution context.

Conor Bronsdon 34:33 [OVERLAP] Yeah, I recently talked to Slack's chief product officer, Jamie DeLange, about this, and I think it's one of the reasons that I see Slack as such a crucial part of the context harness, because most organizations, or at least most tech businesses, are using Slack or Teams or something like that. they have continual conversations going on and you need to be able to ingest that data. But to your point, it also needs to be prioritized against other data sources. And I know that Wisdom has also thought about loops to not only update data but elsewhere in their harness to ensure that the context is staying up to date and that this accuracy and trust is passed forward in the platform. Can you tell me a bit more about how you're designing this overall architecture for the system?

Kapil Chhabra 35:12 One thing that we see very commonly is that enterprises do not have all their data stored into a single database or a data warehouse or a single system.

Conor Bronsdon 35:21 Very rarely.

Kapil Chhabra 35:22 Everybody has so many different data stores. Now, for us to provide intelligence across these data sources, what we need to do is when a new question comes in, we need to understand the intent of the question. and we need to decompose the question into its substituent parts. So what happens is we take the question, we find out various different questions behind the question that now align with the data source that can answer that question. So first thing that happens in that loop is decomposition. and then we are running those queries we are generating the queries we are running executing those queries on different data sources some data source may speak in one dialect of sql another one may speak a python another one is a rag system so it needs to understand and build all of those loops then when it gets the data back, it needs to verify whether it is getting the right information back or not. If there is an error, it needs to self-rectify that error. And then it needs to recompose the response to present it back to the user who is asking the question.

Conor Bronsdon 36:30 Tell me a bit more about this error handling. How are you solving this so that users are getting the right updates?

Kapil Chhabra 36:38 [OVERLAP] There are so many different ways we need to do that. I'll give you a very simple example. Let's take one type of a database, Clickhouse, one of the more popular databases that our customers tend to use. The other one is a Snowflake. Now, doing the same question and asking an LLM to write an SQL statement in a Clickhouse dialect versus a Snowflake dialect, the places where the LLMs trip are different. we know where the LLM trips in writing the Clickhouse SQL versus in writing the Snowflake SQL.

Conor Bronsdon 37:16 [OVERLAP] This is really fascinating. So there's different tripwires in both of them that are typically causing agents to have problems. Is this an area where you start to use those examples we talked about earlier to help agents to get around these potential pitfalls? Or how are you approaching

Kapil Chhabra 37:31 [OVERLAP] Yes.

Conor Bronsdon 37:31 [OVERLAP] this?

Kapil Chhabra 37:32 So there's, there's the examples play a major role because it's not like single shot, but then the harness plays the other role, which says that, Hey, we know that when talking about a

Kapil Chhabra 37:45 time zone conversion in Clickhouse, LLMs, even the best LLMs right now are making a mistake three out of 10 times.

Conor Bronsdon 37:54 [OVERLAP] Now, okay, that's pretty significant.

Kapil Chhabra 37:56 [OVERLAP] It's a yes. And in those cases, what we need to do is have a specialized deterministic loop, making sure that whenever the question requires that time zone conversion, we can inject a little piece of code onto that ClickHouse code before executing onto the Data Warehouse and letting the query fail.

Conor Bronsdon 38:16 Um, one of the big challenges for enterprises is obviously how many data sources they have access to and the disparate nature of that data. And previously it has taken hours, hours of work from data analytics teams to bring that data together, to build these BI dashboards, to actually ensure that they're centralized, useful data. How is wisdom AI helping unblock enterprises that are struggling to not only bring the data together, but with the cost of doing so?

Kapil Chhabra 38:47 Let me answer that with an example of a real customer use case. This is what was the setup. This is the customer that has many web properties and millions of views on their website. They have been using Google Analytics to capture all of that event's data. and then they had been doing Google Analytics to BigQuery ETL and running a LookML and dashboarding on top of that BigQuery. That was the pipeline. So users come in onto their website, the website is sending an event into Google Analytics, Google Analytics sends all of that data periodically to BigQuery and Looker is running on top of that. This pipeline was costing the millions, especially this ETL portion tied to Looker costs about $5 million a year, given the scale that this company is operating at. And guess what? The way all of this data was getting used was not necessarily on the Looker dashboards. There are people who are going to Slack chat and ask the data team this question that, hey, how many subscribers actually visited the website? And that's actually a pretty difficult question to answer if you think about it, because website visits do not include the subscriber count. You do not know if the visitor is a subscriber or not. So in their case, what they had was a database of subscribers also sitting in BigQuery. And that was the primary reason where they had to bring in the Google Analytics data into BigQuery so they could join the data and answer these questions of how many of the visitors were actually subscribers. With Wisdom, they saved that $5 million worth of pipeline cost. What they did was instead, took Wisdom, connect directly to the Google Analytics MCP, and connect the BigQuery database of subscribers. And when a user asks a question directly in Slack, The wisdom agent is spinning up on the back and it's answering that question. Behind the scenes, what it is doing is, as we mentioned, decomposes the question first. takes the first part of the question, sends it on the Google Analytics MCP to get the response of how many visitors were there and what all visitors. The second part of the question on how many of these are subscribers, it needs to get the subscriber information and do a join between these on the fly and then get the data response back to the user. All of this is happening on the fly in this agentic harness without having to ETL this data into BigQuery. So what this gives us is like a way to federate the data access across all of these systems which are disparate in nature. Google Analytics is talking the language of an MCP. BigQuery is talking its own dialect of SQL. And there may be other information as well, some latest news, some manuals and other documentation. All of that can be federated and the harness does the heavy lifting of joining that information and providing the answer back.

Conor Bronsdon 42:04 I mean, this all is part of the emerging discipline that many folks will say sounds familiar of context development life cycle. You know, everyone listening here is probably familiar with the SDLC, the software development life cycle. You have a context development life cycle that is occurring. We've talked about a lot of the parts of it and some of the problems with context drift that we find throughout this, and I think anyone who has to manage a coding agent, or frankly any other type of agent, has probably experienced some context drift themselves. And that only gets worse when you have multiplayer AI, and you and I are both pretty convinced that multiplayer AI is the future, whether it's individuals running mass swarms of agents, or agent swarms running themselves as you've seen with the hugging face hack and other things, but that's another conversation. But for enterprises too, like I can't only be operating my own agents. I have to have my agents be able to interact with your agents, with other team members. They need to be able to work in Slack and understand the context of other business decisions. How do you define this life cycle of context development and how should a business or maybe a team within a business work through transforming from existing dashboards and disparate pieces of context to a more established layer where they can actually have a real life cycle to continue to grow.

Kapil Chhabra 43:29 As a technology industry, we are not at a point where building and managing this context is fully automated. This is actually a pretty big significant missing piece in the LLM's understanding of the enterprise and a pretty defensible mode as well. The companies that we see that are successful with this are the ones that have adopted a new role in the organization. It's called the AI context engineer. We call it ACE. We even have a training that is coming out. It's a free of charge training. Anybody can go take that. It'll be a certification course, just up-leveling the people in the organization.

Conor Bronsdon 44:08 [OVERLAP] We will link on the show notes, so folks go find that. That's

Kapil Chhabra 44:10 [OVERLAP] Yeah.

Conor Bronsdon 44:11 [OVERLAP] awesome.

Kapil Chhabra 44:11 [OVERLAP] Awesome. Yeah, thank you. It's pretty straightforward. Wisdom.ai slash ace is where you can find that certification course. And what it is talking about is that, let's say the data analyst who were in the line of questioning. What I mean by that is let's say there's a business user who has a question. They go ask the analyst. Analyst is the one who's interpreting this question, writing, understanding, bringing their context in. translating that into code executing the code onto a dashboard or onto a data warehouse and bringing the results back they were in that loop not anymore because this business leader is now asking their agent the same question so where they need to inject themselves is behind the scenes to power this the context they need to be able to inject the context in that loop to the agent That role is this AI context engineer role. They follow the full CBLC.

Conor Bronsdon 45:12 [OVERLAP] I know we're still doing it. SDLC, CDLC.

Kapil Chhabra 45:14 [OVERLAP] Exactly. Yes. But the context development life cycle is, starts with how do we gather all the context that has been documented in the organization, bring that together, make sure it is accessible to the agents and then keep monitoring for this context drift. And as it is getting, uh, getting drifted, uh, keep capturing that as well. Maybe they themselves are those voices of authority who can say that, yes, the definition of churn has changed. Let me go ahead and put that. Or they are the ones who are going and asking the, the person in authority that, Hey, hasn't this changed? Should I make it universally acceptable? Uh, and so on. So that's, that's a new role or a new responsibility for the existing members of the data team that is emerging fast.

Conor Bronsdon 46:02 I think this is interesting in the context of broader concerns for many folks around jobs are going away, AI is taking jobs. And I will just remind our audience, and we've talked about this on the show multiple times, that yes, jobs are changing. You have to be prepared for that change. But there are so many new opportunities that are being created by AI. We are just changing the level of abstraction that people work at and the tasks that people are working at. And personally, I'm very excited that we can potentially do more impactful work around data and analytics. But I know this continues to become more and more complicated as you work with an enterprise that has multiple teams that are committing data and context at the same time. Obviously, there are different priorities across those teams on how to keep context from drifting, what type of context should go in the harness. How should organizations and leaders who are driving these changes think about this multiplayer aspect that we've talked a little bit about in this conversation?

Kapil Chhabra 46:59 Like the first thing to accept is that it's not just a product change. This is a people and the process change as well. So three P's, product, people, and process. They have to accept that. And companies and leaders that have accepted this are the most successful. That's 7% that I was talking about that have been able to scale. One thing that is common across all of them is that they have accepted that this is not just a product change. Now, what does that people and process change mean? From the process perspective, it is similar to what you were talking about just now, which is, it's not just the humans that are interacting with the data, agents are also interacting with the data. So there's a lot more automation that is happening. One needs to accept that, that this is going to happen and trust in the data becomes even more important if somebody is autonomously going ahead and letting them publish an article on Substack, for example. So that trust becomes even more important. So that's the process change. We also need to accept the risk in terms of some of these answers, if not enough context is provided, may not be accurate. So need to build in the processes on when the agent is flagging that, hey, I did not have full context to answer this question. What needs to happen in that case? Maybe they should ask the agent to actually not go ahead and take the action. So that's a process change. The second thing is about the people change. And what we were talking about just now on this ace role, the AI context engineer, that is the other thing that needs to be accepted that the people who were in the position of this data analysis and data engineering, who really understand the business aspects and the data aspect and can really bridge the gap, put them in this position of a context engineer. Let them be the ones who are providing the context into this layer that is powering all the agents in the organization.

Conor Bronsdon 49:06 [OVERLAP] think to your point, Kabul, we have a change management challenge that leaders and individuals within an organization need to address and not just address. I think it's an opportunity. It's an opportunity to transform your business in a way that can be much more efficient and let you take on these different layers of abstraction. We're seeing teams that are effective, whether with AI coding, whether with AI implement of their go-to-market teams accelerate past those that aren't. But it is not a magic bullet. It may be a magic bullet for an individual if they are spending the time doing this context engineering, this work uh, themselves to build their own individual harnesses, or they're picking up ones that are right size for them and adding context. But when you get to organization scale, and that can be five people, but it's certainly true at a thousand plus, there are many challenges that come in here because of you have processes that are already in there. You have technology that's already built in. You have people that are already steeped in the ways of the organization. They may have a lot of tribal knowledge that you need to contribute. They need to figure out how to pull out of that. And maybe that's because, you know, what I think one of the personal areas where i've seen a huge gain is from ingesting all of my meeting notes and i think we're seeing this for go-to-market organizations in particular where sales people are like fantastic i do want to record all my sales calls i want to pull that data in it's really helpful for me to keep things up to date and avoid context drift to your point But this does require people to manage it. It's not something that systems can do automatically in all cases. We do need human oversight of these. And when these systems are learning from interactions, there is a need to have a structured approach to approvals and improvement. how would you recommend teams go through this approach? And I'll say, we're going to be talking to Cloudera's chief data analytics officer about their change management approach. And in fact, I've actually already recorded it. Don't tell anybody. And one of the things he told me was he was like, look, I don't feel like I did the best job of this. Like I've been working on this for two and a half, three years and, you know, trying to do this internally for my own team. And yet like, there's all these learnings I can take from organizations like yourselves.

Kapil Chhabra 51:12 [OVERLAP] Yes, 100%. It's a change management, as you rightly said.

Kapil Chhabra 51:20 Who is in the right position to accept a change in the context that is going to impact all the agents that are using it? is something that the organizations need to decide.

Kapil Chhabra 51:35 No amount of forward deployed engineers can do that because at the end of the day, they are external to the organization. If somebody is thinking about building in an FDE army, bringing them in and help that they are going to build this context, No, the best they can do is go through all of those call recordings and say that, Hey, these are the ones that these are the definitions that I see. There's multiple meanings for this yet. They're going to come back knocking the door of that chief data officer and say that, Hey, who's going to give me the right answer between these two, because I'm saying, I need a decision here. So somebody needs to be in that position of deciding. Otherwise the LLM is deciding for us, which may or may not be right.

Conor Bronsdon 52:22 And another thing that I think is really important is having regular checks or loops, as you mentioned, for context drafts. So as an individual developer, I've done this, I built an SSOT check tool where it just checks to make sure that there's a single source of truth and it says, okay, is there drift across my docs and read me, okay, let's fix this, right? And that can work for me as a GitHub action on my repos. But again, like that's only so useful as I start to roll it out to a larger team. I'm an individual, I'm running a very small media org. How do companies handle these challenges? Because, yes, you need to identify people who are going to be the decision makers in these cases, but you also want to automate as much of the underlying grunt work as you can.

Kapil Chhabra 53:06 Yes, 100%. So there's two different dimensions here. First dimension is the type of context that we're talking about. So let's split that into data context and functional context. Data context, think about that as bottoms up. This is how we had designed this data system. This is the modeling of the data. And then on top of that, here is what we have defined as the semantics for that data. The second type of context is this functional or business context. So sticking to our example of churn, the definition, the metric of churn is part of the data context. But is 20% churn good, bad, ugly? Is the functional or the business context on top of that. Now, person in the data team's capacity may or may not be able to define what good, bad, or ugly looks like for that churn, but they are the ones to define what the metric churn looks like when they have to calculate it in SQL.

Conor Bronsdon 54:11 And you've tested this actually with and without your harness, my understanding. What kind of data did you see coming out of those tests?

Kapil Chhabra 54:19 [OVERLAP] It's fascinating to see those results. So what we did was, just for context for everybody, what we did was...

Conor Bronsdon 54:27 [OVERLAP] I know.

Kapil Chhabra 54:29 It's everywhere. So what we did was we took the best frontier model at the time. So we took Opus model and we fed it exactly the same context that we fed Wisdom AI. We did a side-by-side comparison. And this has nothing to do with like Claude or Opus. This is essentially a comparison with state-of-the-art versus a specialized harness for a task. What we saw was that the difference was on multiple matrices. The first one was time to value. When you ask a question into Cloud, in this example, it is going through a brute force loop to go find the context across all of the Slack threads, across all of those meeting nodes. It's running multiple tool calls and searching through that to come up with its own best definition. In some cases, it asks back the question and then comes back with the answer. versus wisdom where the context has been provided, it is not going making guesses around that. It is taking a faster, shorter, more optimized path. So there is a 3x difference in the time to getting the first answer. We also saw a 3.4x difference in the cost.

Conor Bronsdon 55:53 Interesting, so that's related to tokens spent to actually find the answer. So not only is it faster, but it's spending significantly less tokens.

Kapil Chhabra 56:00 That's exactly right. Yes, it is spending a lot less input as well as output tokens to get the responses back.

Conor Bronsdon 56:08 Obviously, we're seeing work start to happen in different places than we're used to. People are using their cloud code or Glean agents to go ask questions, do coding. Customers are spending time in these different surfaces. How does WisdomAI play a role in that?

Kapil Chhabra 56:25 Yeah, so what wisdom does is it provides an MCP server out of the box. It's completely hosted server. Nobody has to do or raise a finger either. Just connect the client where the work is happening with the wisdoms MCP server and the server. is also powered by mcp apps which essentially means that it is not just sending a text response back but it can also send a response in the html react format that if the client is capable it can render it which chat gpt and claude are capable of and other clients are also becoming capable of doing that We recently announced our partnership with OpenAI as well, where the data product that OpenAI announced, which is part of the ChatGPT work suite, Wisdom is powering the data analysis behind that. So if the work is actually happening in ChatGPT, which in a lot of organization is happening like that, when they ask a data question and wisdom is plugged in behind the scenes, the widgets, the answers that they see are powered by wisdom, which essentially gives them the benefit of the cost saving on tokens that we spoke about. And it also gives faster responses with all the governance and the trust that the data team wants to impose on it.

Conor Bronsdon 57:43 This seems like a major differentiator when you compare it to traditional approaches to gathering data and context.

Kapil Chhabra 57:49 Yes, 100%. See, the way we need to think about it is in the traditional world, the interface for the end user was a dashboard layer. It was the data team that was building and doing all of the work so that these tools were built for the data team, not really for the end users to build these. End users were the consumers. In this world, end users are the consumers, but then they are doing work on top of that. So what we need to do is meet them where they are. Are they doing work in Slack, in a Glean agent, in a Cloud agent, in a chat GPT, or do they want a specialized interface like a Wisdom AI interface? So we bring all of those things together depending on how that agent wants to interact with it. We have a Slack and Teams app natively available in the platform. We also have an MCP server that plugs into all of these different clients.

Conor Bronsdon 58:49 [OVERLAP] I think this is a crucial piece too, is to not just let people access Wesdem AI where they're already working, but also give developers tools to build on top of and build with Wesdem AI. And another important part of all this architecture you're building is obviously when and where to provide context. You can have static context or you can inject context at runtime. How are you making those

Kapil Chhabra 59:11 [OVERLAP] decisions.

Conor Bronsdon 59:11 [OVERLAP] decisions?

Kapil Chhabra 59:12 Yeah. So the, this, this distinction between referenceable static context, which is sitting somewhere in a repository and actually being able to use the context. And we spoke about like the right sizing of the context itself. Too much is problematic and too little is problematic as well. So you just need to find the right balance of the right amount of context. And the decisioning for that is all part of this harness. We call it the context runtime as well. So on the runtime, who is responsible for picking the right piece of context then generating the query code based on that, executing that code, getting the results back, is all tied into the runtime of this context, not just the repository part of the context.

Conor Bronsdon 1:00:08 I think this speaks to the value of the idea of harness engineering, which is increasingly being talked about by the industry. We're all talking about how to engineer our harnesses. And it's very clear that specialization around these frontier models can enable you to achieve better results. You know, we talked to Lama Index's Jerry Liu about this earlier this year. about the need to provide incredible context. And from their perspective, they're doing it by saying, okay, we're going to be the best in the world at unlocking context from PDFs. And then you can feed it into a system like Wisdom.ai. And I wonder how much iteration you're having to do, I assume quite a bit, around actually continuing to redevelop your context harness as each new series of models comes out.

Kapil Chhabra 1:00:51 What we've done is we've taken this context building process outside of the loop of a single question when it is getting asked. So think of it as an offline process, part of the CDLC tooling, not in the line of query. What that gives us is the opportunity to actually run that at a set frequency that we determine, rather than when thousands of people are asking questions and every time we are trying to do that brute force. So that gives a major optimization. And whenever a new model comes in, we are using that model in the line of query. It is not impacting how we are building out this context from all the known documented sources right now. So that level of architecture change that we have implemented is giving us the opportunity to leverage the best of the models that come out while not impacting how we build and manage this context.

Conor Bronsdon 1:01:44 And I've seen your team talk about this idea of right-sizing AI and ensuring that AI systems are aligned to business use cases, to those trust needs. And part of that is that you're building on top of your harness in a way where you have these great models you're working with, you have this harness that is helping keep the data and context around them accurate, and then you're building things like your newly announced live apps on top of that. Can you tell us a bit about how you're thinking about this kind of next layer around or on top of the harness.

Kapil Chhabra 1:02:16 There are two ways to go about any kind of transformation effort. AI transformation is no different from that perspective. One way is to think about, let's just go ahead and transform the whole thing. We want to go from zero to 100 in 2.3 seconds. The other way to think about it is that here is the use case that I have, and let me go ahead and address this use case, and then to the next one, and then to the next one. And by the time I get to the third or the fourth use case, I'll know exactly what the playbook for transformation looks like, and I will be able to scale. We are big proponents of this latter approach. Within Wisdom.ai, we have a concept of a domain, which essentially maps to the use case. let's say somebody wants to do this sales analysis or pipeline analysis in the example that we've been talking about. In that case, you connect your data from Salesforce, you connect the data from the ERP system, you connect the data of the contracts, and all of those things are the ones that can provide a pretty fantastic analysis for the sales operations as well as the pipeline operations there. Now, that is a atomic unit that you can go ahead and address, and that is what we call right-sizing the AI rather than boiling the ocean. And layering in the new feature that we have launched of live apps, which we are very excited about, we're getting tremendous usage on that feature, is essentially, you just give a single prompt, And from that prompt comes out a beautiful application that you can interact with. It's not just read the data from there, but you can use that for right back. You can do use that for what if type analysis or triggering deeper analysis, like root cause analysis, et cetera. And all of that within a matter of minutes, which would have otherwise taken days or weeks for somebody to build.

Conor Bronsdon 1:04:15 And I appreciate the opportunity to play around with this new live apps feature a bit, because I think one of the really interesting things we've seen is that everyone wants to create a pretty dashboard for themselves. And, you know, we have cloud artifacts, so I will spin up a quick HTML one for you if you want. If you're using coding agents, you can spin up a quick dashboard, sure. You can send it to someone. But live apps seems really differentiated in that it actually is continuing to adapt over time. And it's much more durable, it seems like.

Kapil Chhabra 1:04:42 Yes, and the big difference I'd say is the governance, the boring ticks. We all do white coding, right? Just give a prompt and there comes a beautiful artifact and that's white coded artifact. There's no novelty in that anymore. But all the boring things around that, that when you are accessing versus I am accessing that, are the row-level security features applied? Is the PII information masked? Is

Kapil Chhabra 1:05:13 this hosted at an environment where only the right set of users can access this? Can I go see the SQL or the Python code behind it? And if there is a change that I want to implement, I can go ahead and do that. If my metric definition changes, is it impacting all of the live app that is powered? All of those boring things is what is packaged within this feature of live app.

Conor Bronsdon 1:05:37 [OVERLAP] Yeah, just because coding has gotten very cheap as far as a person's time by throwing tokens at it does not mean the architecture, the governance, the decision making around it has also gotten cheap. It in fact has gotten, I would argue, maybe more important because we have these incredible, I guess, spray and praise of code where we can say, yes, I can throw code at it and, you know, eventually I can iterate towards what I need to do. But if I'm able to help guide it from the start, I can be much more effective. So What does the context harness handle that enables more effective dashboard and more effective analytics that a team would otherwise need to spend time on for these kind of generated apps?

Kapil Chhabra 1:06:19 [OVERLAP] Yeah, we were just doing analysis with one of our customers who is migrating all of their 1000 Tableau dashboards over to WisdomAI. And we were doing the math behind that. and so the time it took to build out each of those dashboards in Tableau actually varied but we took an average let's say two and a half weeks of human time which is about 100 hours makes the math easy a thousand dashboards 100 hours each we're talking about 100,000 hours more mind-blowing number if you divide that by 2000 which is the number of hours each of us works in a year it comes to 50 human years wow was the time it took to build those 1000 dashboards.

Conor Bronsdon 1:07:08 [OVERLAP] I think this is so interesting in the context of coding too, which is again where I think people are seeing these initial gains because we used to spend so much time saying, oh, I need to port this from Linux to Windows or I want to rewrite a system in a new language. I need to re-architect or refactor. And now we can throw coding agents at it, or we can throw these contextual live apps at it. And suddenly we can have a transformation that occurs. And there's still a ton of work to be done on top of that. AI is a magical bullet that is very self-contained. It cannot solve all your problems. You need to do a lot of work around it. And I think, frankly, I think that's the fun part.

Kapil Chhabra 1:07:45 That exactly is the fun part. And here at Wisdom, we are taking care of all the boring aspects of it. The governance that I spoke about, the caching behind that, all of that stuff, nobody wants to build that. They just want to kind of send out this prompt and get the beautiful artifact and be able to share that within 10 minutes rather than the hundred hours that we spoke about.

Conor Bronsdon 1:08:08 I have to admit that's a compelling pitch here. So, okay, you know, we're using this phrase live apps. What does live mean in this context? What's refreshing and how does someone know when and if data is accurate?

Kapil Chhabra 1:08:22 [OVERLAP] yeah so it's not just a static html that we produce this the html and the react application that we produce is actually connected behind the scenes into the into the live database or the data warehouse or the context repository it is using the whole wisdom harness as well as the context layer that we just spoke about And all of the queries are connected live, which means that whenever somebody is refreshing that live app, they are getting the latest data. The other improvement that we have done within this is that, let's say there are 100 people in the organization who are going ahead and using the same application. we're not going to the data warehouse and sending those 100 queries individually we have built a caching layer in between so we're only sending that once and the rest of the 99 people are just getting served through the cache which is still the latest data but at one percent of the compute

Conor Bronsdon 1:09:19 [OVERLAP] Huge token savings there. And I'm curious, are you transitioning most of your internal dashboard into live apps, or are you still holding certain things in, say, a conventional dashboard or spreadsheet or other form of data?

Kapil Chhabra 1:09:35 The purpose of the dashboard, which had existed for so long, is not changing. It is not going away anywhere. We all want to monitor our business through KPIs, right? These are the metrics that matter to us. So that essence, that idea does not go away. The format in which that gets represented is changing beneath our feet right now. So many years we've been building out these dashboards and these legacy tools. Now we are in this new era. We're building out live apps with a single prompt. So at Wisdom, we do not use any other dashboarding product all our operations are running on VisDim, and most of our customers are also transitioning away from Tableau Power BI Looker into VisDim AI for their dashboarding needs.

Conor Bronsdon 1:10:24 It's fascinating to see how the abstraction layers we operate on with AI have rapidly changed. It used to be if I was coding, I would be reading the code. And frankly, I'm typically not reading most code these days. I often am not doing code review myself, or at least I am taking a level of abstraction where I'm not reading the specific line-by-line code. Instead, I'm having multiple models that are doing that review for me and then are reporting back to me. I have friends who have transitioned to, I guess, almost a dashboard situation with their code, where they're saying, hey, I actually want a constantly generated artifact that runs alongside my coding window that shows the architecture of this harness I'm developing and how things are moving. So I can see, oh, look, now this has been lost or is no longer necessary. But AI can code faster than we can. It can write faster than we can. We have to change the extraction layer we operate at. And data is obviously a huge example here. I mean, we've seen to mention the hugging face hack again we if anyone who read the meter investigation or have seen anything about it they had to use agents as their front line to understand this because there were so many conversations that happened and you as a human we just simply don't have enough human hours to handle all the data we now have available to us today so it's crucial that we start to have these loops that can build this data for us and have agents that operate across these dashboards So, I'm curious to understand from you, as you have transitioned your organization to all running within these wisdom analytics, within web apps, what kind of agents do you have operating across these dashboards to help with context management, to help with identifying trends and those sort of things?

Kapil Chhabra 1:12:08 One of the most recent live apps. So this is a recency bias on this one. I was just doing it over the weekend. One of the most recent ones that I've built is to monitor my pipeline within the company. And that pipeline, essentially the thing that I wanted to monitor was pretty unique. It wasn't that which deal is in what stage of the pipeline. What I wanted to see was that in the last week, which deals have transitioned from one stage to the other. And for each stage, I wanted to see how many deals have progressed, or if there are any that have regressed. And I wanted a specialized visualization just for that. And I wanted to be able to filter when I click that bar in the chart and all of those things. And this was one of the ideas that I got while taking the stroll. And I came back, gave it a prompt, and in in less than 10 minutes I had this working application which now I have subscribed to. So every Monday morning as it gave one this Monday morning before my pipeline review meeting with the team I am getting that in my inbox on what changes have happened. So the meeting that I'm having with the team now in that pipeline review has completely changed. I am not asking them for that information. Now what I'm doing is asking them where can I help in making these deals progress even faster.

Conor Bronsdon 1:13:34 I think that's a fantastic example of this changing abstraction layer and being able to just ingest and understand data faster. And it's interesting to bring it up, this idea of a pipeline dashboard, because obviously, you know, we've been using pipeline dashboards and go to market for a long time, but the effectiveness of them has been variable depending on how good our SEs and our sales folks are at actually putting information into our systems. And I love my sales people most of the time, but they maybe are not always the best about data entry. And I worked at a previous organization where just last year we tried to vibe code a similar version of the dashboard. And the problem we had was around data management of it and keeping it up to date. And it was extremely useful at first, but rapidly we started to have problems with it getting out to date and we had to do a lot of fine tuning to actually get it to work over time. What kind of data and definitional work do you feel like needs to be in place first to enable this kind of next layer of analytics that can unlock you to spend more time on crucial unblocking routine numbers?

Kapil Chhabra 1:14:39 Yeah, and the answer there is context. And that context on what I'm asking for, is there sufficient information for the LLM to be able to generate the right set of queries for that? That is what it boils down to. Now that's such a amorphous definition that there's no way to know whether this is sufficient or not. Have I put in enough context or not? So what's important to do is the other way around, which is when that app has been built. What Wisdom does is highlight each of those sections in the app and says that, hey, this one I had full confidence because everything here was powered by the context that you have provided. But this is where some of the context was missing. So this is the widget where I want you to go ahead and double click and make sure that whatever assumptions I made, which Wisdom highlight, are correct or not. If there are some assumptions that are not correct, then it asks for me to go ahead and provide those corrections. If I say that they're all correct assumptions, then that goes into the learning loop and back into the context layer.

Conor Bronsdon 1:15:53 How are you thinking about the ownership of this app or other apps over time? Who should be responsible for ensuring these updates as they become business critical live applications that are now being used in a weekly pipeline meeting or something else?

Kapil Chhabra 1:16:12 AI should be responsible for that. This needs to be completely automated. Otherwise, we will be swimming in the sea of these apps and dashboards that we've always found this trouble in. Like the 1,000 dashboard example that I was talking about. Should there really be 1,000 dashboards for the business to run? Absolutely not. only the right ones, like on average, and Forrester did this study, only 20% of the dashboards actually get utilized in the organization. And less than 20% of the people who have bought a BI product actually use it, they're active users of the BI product. And that's a big gap. Companies are just spending so much money on shelfware that is getting produced. Now, if there are humans or processes built to go ahead and keep these things updated, we will run into the same problem, unless there is an automation. And that is why the name Live Apps, because you do not need to go ahead and update these once it is built. It always stays up to date. It is completely automated.

Conor Bronsdon 1:17:27 How do you think about the role of context engineers or data engineers in this new life cycle of how data is being used and leveraged?

Kapil Chhabra 1:17:37 They are no longer in the business of providing insights. That's just the reality. They are now in the business of providing context, so the insights are accurate. And that is a shift in the responsibilities. Just last week, I had this webinar with my dear friend and a passionate customer from Rubrik. And what they did was, they were talking about how their processes have changed. And one of the things that fascinated me completely was that in the world prior to wisdom, what they were doing was every dashboard the team had to build, they had to go ahead and build a go to the data prep area, go ahead and build a model for that, and then tie the dashboard to that model. So which means that If you take the iceberg analogy, the dashboard is the tip of the iceberg and everything beneath that is also powering that single dashboard. So the biggest transition that has happened for their team is that they've only had to build that definition once in this context layer within Wisdom and all of the dashboards or live apps are now just a prompt away.

Conor Bronsdon 1:18:56 We have a lot of developers who listen to the show and many of them are looking to unblock, not just themselves, but their teams. Often they're partnering with data engineers and folks who are on the data side. How can developers use the Wisdom AI MCP to unblock themselves and create more value?

Kapil Chhabra 1:19:15 Yeah, so Wisdom AI provides the MCP server, not just for the consumption of the data, but also for building out these domains. And the context management aspect of it, if somebody is building their own context harness around a cloud code or a codex, for example, they can very easily plug in the Wisdom's domain MCP servers and start to build out those domains in Wisdom.

Conor Bronsdon 1:19:43 What would you say to people who are maybe doubtful about handing off their context management to another company? Maybe they've gotten used to a lot of the toil involved with this.

Kapil Chhabra 1:19:55 Context is the IP of the company and it needs to stay with the company. They should have the ability to take that context with them anytime, even if they want to rip apart the vendor that they have selected for this context management. So it's very important. And in case of Wisdom, what we do is our context is portable. the customer can click a button and download the context in our standard YAML format, or in the Google's OKF format, or in the Apache Aussie format. With a click of a button, they can download that, or they can even set up a Git repository for their context, where Wisdom will keep writing back into that Git for their context. So the context belongs to the customer, and with Wisdom, we make sure that they get to keep it, even if they have Wisdom or not.

Conor Bronsdon 1:20:48 Well, thank you so much for this conversation. I really appreciate you taking us through your perspective on the context development lifecycle, how wisdom is approaching analytics, live apps, and so much more. What closing thoughts do you have for us about what this new world of context engineering and changing agent analytics look like?

Kapil Chhabra 1:21:07 It all seems like there is so much that is changing fast underneath us. The one thing that I want to kind of assure everybody of is that any work that they have already put in in AI data readiness or data cleansing or building the semantic layer or their catalogs, none of that, even the BI dashboards, none of that goes to waste. The right system actually leverages and reuses all of that and builds upon that rather than having the company start from scratch. So that's the, that's the biggest learning for me that we do not want any of the human work that has happened to go waste. That's actually more crucial than ever at this time.

Conor Bronsdon 1:21:52 And for folks who have stuck with us throughout this conversation, A, thank you, and B, would love to hear from you in the comments around what you found interesting or what areas you want us to explore next time. Couple, what's the best place for folks to go to learn more about Wisdom.AI?

Kapil Chhabra 1:22:07 It's as simple as you just said it, wisdom.ai.

Conor Bronsdon 1:22:10 Make sure to check out all the data from this survey at wisdom.ai slash CDO dash report. If you are interested in taking the new course that Wisdom.ai is offering, you can learn more about that at wisdom.ai slash ace. And if you want to try out Wisdom.ai's live apps, you can go to wisdom.ai slash live apps.

Kapil Chhabra 1:22:36 [OVERLAP] Conor, thank you so much. Had so much fun doing this. Everybody who's listening, please rate and review this podcast. I'm sure it means a lot to Conor and also hit that subscribe button. Thank you.

Conor Bronsdon 1:22:47 [OVERLAP] Thank you so much for having me down to the Wisdom AI offices. It's been fantastic chatting with you and I'm super excited to continue to explore the world of context engineering and context harnesses with you in the future.

Kapil Chhabra 1:22:58 Fantastic. It was great to have you here and it was a lovely conversation. Thank you so much.