# Chain of Thought — full > The podcast where builders reason through what’s changing in AI and software infrastructure. Hosted by Conor Bronsdon, with guests from NVIDIA, Google DeepMind, AMD, Databricks, Vercel, and more. Full transcripts and show notes for every episode. Host Conor Bronsdon interviews the people building AI infrastructure. Below is every episode with its key takeaways and FAQ inline; each links to the page that carries the full verbatim, speaker-labelled transcript — the authoritative source for any quote. To cite the show, link the episode page. ## EP 67: Parenting Your AI Agents for Best Results | Cisco's Jeetu Patel Guests: Jeetu Patel, Cisco · Published: 2026-08-12 · Duration: 46 min URL: https://chainofthought.show/podcast/67-parenting-your-ai-agents-for-best-results-ciscos-jeetu-patel/ (full transcript on page) Cisco President and Chief Product Officer Jeetu Patel returns to explain why AI agents are like teenagers: profoundly intelligent, extremely resourceful, and missing the judgment to know right from wrong. He covers dynamic runtime guardrails, rationing tokens like headcount, and why AI fluency is now a 10x differential. ### Key takeaways - The episode’s core argument: securing agents is parenting, not rule-writing. Jeetu Patel’s framing: “Agents are like teenagers. They are profoundly intelligent, extremely resourceful, have no fear of consequence, and don’t really have the judgment to know right from wrong at all times.” Static allow/block rules fail because “agents are too smart. They’ll actually bypass the static rules pretty quickly.” His replacement: dynamic rules “that get created at runtime,” watching for drift — “if it’s showing drift from the original intent, then and only then should you intercept.” - Cyber is entering its third phase: the agent trust platform. Phase one was best-of-breed (3,500 vendors, every company running about 50 to 70 security vendors or products); phase two was platforms (Cisco, Microsoft, Palo Alto, CrowdStrike, Google). In phase three, “security and observability are going to stop becoming two distinctly different domains.” His example: when an agent deletes 10,000 emails, “we don’t know right now, definitively” whether it was a breach, an external manipulation, or the agent confidently following orders. The platform has to watch resiliency, guardrail crossing, and tokenomics — including quarantining overly consumptive agents. - In Cisco Cloud Control’s agentic ops, no agent-recommended fix ships without a digital twin test. Ambient agents in the network detect anomalies and recommend remediations, but “that remediation and fix does not get directly put into production.” A digital twin “gets spun up right away off your environment,” live recorded data runs through it, and only after the twin operates without hiccups does the change deploy. Humans stay in the loop until the statistics earn trust: approve everything until “I’ve seen enough of them statistically to know that I feel good about it.” - Cisco is moving to ration AI tokens the way it rations headcount. Phase one was deliberate abundance: “everyone has access to tokens, go use to your heart’s content.” Phase two is allocation — just as a team gets “only so many open recs,” it will be given a token budget and an expected output, which pushes teams to engineer for yield: “20, 30, 40, 50, 60% more lift in tokens being generated because of the way that I’ve actually constructed my architecture.” The underlying diagnosis: token costs outrun value because companies “haven’t really gotten good yet at extracting the most out of the models” — his sequence is familiar, then good, then efficient, and companies are still between familiar and good. - Public benchmarks won’t predict your results, so build your own evals. “If a model does well on one eval, that doesn’t mean the model is going to do well in your environment.” Own evals unlock intelligent routing — knowing when to redirect a prompt to a cheaper model versus a more expensive one. Inference is where the pressure sits: “60% of the total compute capacity that’s available is going towards inferencing,” and because “price per token might eventually become a vanity metric” as long-running agents consume more tokens even as each one gets cheaper. The metric he wants down is price per unit of outcome. - His contrarian jobs case: “you’re going to have more jobs created because of AI in the future, rather than less jobs.” The mechanism is the moving bottleneck — “if you write unlimited code, code review becomes a bottleneck,” then product judgment, then instinct. The stakes are the fluency gap: “that’s not a 10% differential anymore, that’s a 10x differential,” while “less than 1% of the world is using agents today.” He calls the no-entry-level-hiring narrative “almost ridiculous”: OpenAI and Anthropic kept hiring from colleges while each person did “10 times more than what they could do a year ago, two years ago, three years ago.” ### FAQ **Why does Jeetu Patel say static security rules are obsolete for AI agents?** Because agents route around them. Patel’s framing is that agents are like teenagers — profoundly intelligent, extremely resourceful, no fear of consequence, and without reliable judgment — so a fixed allow/block list is a fence a smart agent will simply bypass. His alternative is dynamic boundary conditions generated at runtime: analyze what the agent is doing, compare it to the original intent, and intercept only when behavior drifts. His proof case is the OpenAI/Hugging Face incident discussed in the episode: not a malicious agent, an obedient one following its instruction to beat an eval. **What is an agent trust platform?** Patel’s name for cyber’s third phase. Security started as best-of-breed (3,500 vendors, companies running about 50 to 70 products), consolidated into platforms (he names Cisco, Microsoft, Palo Alto, CrowdStrike, and Google), and is now fusing with observability, because agent incidents are ambiguous: an agent deleting 10,000 emails could be a breach, an external manipulation, or the agent confidently following its orders. An agent trust platform watches three things at once — infrastructure resiliency, whether agents cross guardrail boundaries, and tokenomics, so overly consumptive agents can be quarantined. **How does Cisco use digital twins for agentic operations?** In Cisco Cloud Control, ambient agents monitor the network, detect anomalies, and recommend fixes — but no fix goes straight to production. The system spins up a digital twin of the environment, runs live recorded data through it, and applies the proposed policy change there first. Only after the twin runs without problems does the change deploy to production. The human-in-the-loop option is always there, and Patel’s default is to approve everything, removing yourself from the loop only once you have seen enough agent decisions to trust them statistically. **How does Cisco manage AI token costs?** In two phases. First, deliberate abundance: everyone got tokens to build familiarity. Now comes allocation, rationing like headcount — Patel says teams will be given a token budget and an expected output, the same way they get a limited number of open recs. That constraint pushes teams to engineer their architecture for token yield (his target: 20 to 60% more token lift for the same spend). He argues costs feel high because companies are still between “familiar” and “good” at extraction, and that price per token might eventually become a vanity metric; the number that matters is price per unit of outcome. **Does Jeetu Patel think AI will create more jobs than it destroys?** Yes — that is his stated contrarian position. Every step-function improvement in automation surfaces a new human bottleneck: unlimited code generation makes code review the bottleneck, automated review makes product judgment the bottleneck, and so on. He also calls the end-of-entry-level-hiring narrative “almost ridiculous,” noting OpenAI and Anthropic kept hiring new graduates while productivity per person rose 10x. His warning is the fluency gap: the differential between people who are dexterous with AI and those who are not is now 10x, not 10%, so upskilling at scale is what stands between opportunity and, in his words, a lot of human suffering. ## EP 66: Data Federation, Not Centralization, Is What Enterprise AI Needs Guests: Jitender Aswani, Starburst · Published: 2026-07-16 · Duration: 53 min URL: https://chainofthought.show/podcast/66-data-federation-not-centralization-is-what-enterprise-ai-needs/ (full transcript on page) Jitender Aswani was customer zero for Presto at Meta, scaled it at Netflix, and now runs engineering and security at Starburst, the $3.35 billion data platform built on Trino. He explains why the average enterprise has 52 to 200 data sources, why centralizing them will never work, and what happened to Starburst's query volume the day they shipped an MCP server. ### Key takeaways - Jitender Aswani’s core argument: fragmentation is a condition to design around, not a project to finish. “A traditional enterprise will have anywhere between 52 to 200 data sources,” accumulated naturally as each new service brought its own store. A decade of moving it all into one place produced debt instead of answers: “Centralization was a fine, fine strategy when the data was growing not as fast as your ability to move it,” but “as the data volume continues to explode, the ETL tech debt starts to really accumulate.” His conclusion, and the thesis of the episode: “The federation is the only model that scales with entropy. Can centralization scale with entropy? No, because the entropy is constantly, constantly increasing.” - Agents broke the query-volume assumptions underneath enterprise data platforms. Because they “operate in this react mode, they reason, they act, they reason, they act, they are generating queries at insane speed. Some people call it at machine speed. I think they’re generating it faster than the machine itself.” Sometimes they also fire unbounded ones, because “they’re not making judgment of, shall I fire that query?” Starburst watched it happen to its own service: “when we launched our own MCP server, we saw the query volume go through roof,” and “many of our customers reached out because they had to scale the compute capacity behind it.” Aswani also says “80% of our code is now generated by agents” — his own first-party number, not an audited one. - Some of the large language models driving those agents are, in his phrase, “overzealous,” and that is a line item. “They want to prove themselves right. They do not want to come back and say like, oh, I was wrong.” His example is a FinOps agent Starburst built internally that only had AWS data wired up. Asked to break cloud spend down by provider, “it starts firing queries. It’s just frantically firing queries without stepping back and thinking, do you even have GCS and Azure data?” And it is a pattern he keeps seeing, not a one-off: “sometimes I’ve seen that after 30 minutes, they’re still doing it. And I know you are wasting my tokens. Just stop.” His sliver of hope is guardrails and precise context — not fewer queries, but “they will become smarter and they will structure their queries better.” - Humans already failed this test, long before agents could take it. A senior VP asked why enterprise churn was rising; Aswani’s team spent three days on “very thorough and in-depth analysis” and came back to find the VP had already made the call and kicked off campaigns without waiting for them. He has a similar story from another organization: “the CFO asked the question, why is our churn up? And seven days later, three teams give him three different answers.” Finance included cancellations, product included trialers nobody else counted, sales left out renewals. Three teams, three methodologies, no shared definition. Then his pivot: “Imagine, imagine agents without access to all the data, without access to the right context, what would they do?” - Context is a federation problem too, and it is not metadata. “Context is just not metadata, context is more than metadata.” It is ontology — “understanding that customer means account” — plus business rules (one enterprise’s fiscal year starts in February, another’s maps to the calendar) and the definitions of churn, ARR, and monthly recurring revenue. That is the distinction he draws between graph types: “Knowledge graph is all about entities and relationships. Now context graph is building on top of a knowledge graph.” Starburst ran the Trino playbook at it rather than centralizing, since “we know your context lives in 10 different places.” His bet: “We had a moat in data federation, and we have a moat in context federation as well,” and “that’s where the next billion dollar companies will be created.” - The product shipping today is a 2011 research premise that finally got its technology. “Back in 2011-2012, we felt that there is no Google for data” — information was indexed and searchable in natural language, structured data was not, so business users needed a specialist just to learn what data existed and what shape it was in. The plan was to blend SQL, search, and graph behind a Google-like box. “Sometimes it does take 16 years to see that vision or dream come true.” It became AIDA, and the design point is action rather than answers: “It’s not just about ad hoc exploration. It’s about turning that ad hoc exploration into some kind of a business workflow.” It composes through MCP servers and reusable capability packs, because “skill has become the harness for agents.” ### FAQ **What is data federation and how is it different from a data lake?** Data federation queries data where it already lives instead of first moving it into one central lake or warehouse. Jitender Aswani, SVP of Engineering and Security at Starburst, argues centralization only worked while data grew slower than your ability to move it, and that a traditional enterprise now runs 52 to 200 data sources. His position is that “the federation is the only model that scales with entropy,” because entropy keeps increasing while ETL debt compounds with every new column and schema change. **Why are AI agents driving up database query volume?** Because agents work in a reason-act loop rather than asking one question at a time. Aswani describes them “generating queries at insane speed,” and sometimes firing unbounded ones without judging whether a given query is worth running. He says that when Starburst shipped its own MCP server, query volume went through the roof and many customers reached out because they had to scale the compute behind it. His advice to data teams is to assume auto-scaling compute and a platform that holds query performance under that load. **Why do AI agents keep querying instead of saying they do not have the data?** Aswani calls some of the large language models powering agents “overzealous”: they want to prove themselves right and will not come back and say they were wrong. His example is a FinOps agent Starburst built on its own platform that only had AWS data connected. Asked to break down cloud spend by provider, it kept frantically firing queries rather than stepping back and noting it could not see GCS or Azure. His read is that agents running on light context and a partial view of the data landscape burn tokens proving a point. **What is a context graph and how is it different from a knowledge graph?** A knowledge graph captures entities and their relationships. A context graph, as Aswani describes it, is built on top of that and adds the business layer an agent needs to answer correctly: ontology and synonyms (customer means account), business rules such as a fiscal year that runs February to January, and metric definitions for churn, ARR, or monthly recurring revenue. He argues that knowledge sits in individual heads or buried in Confluence and Google Docs, and that harnessing it is where he expects the next billion dollar companies to be built. **Why do different teams get different numbers for the same metric?** Because each team carries its own definition. Aswani tells the story of a CFO asking why churn was up and getting three different answers seven days later from three teams: finance counted cancellations, product counted trialers nobody else included, and sales left out renewals. His point is that this predates agents, and pointing an agent at the same fragmented data does not fix it, since “context is just not metadata, context is more than metadata.” Shared definitions have to live somewhere the agent can read them. ## EP 65: You Can't Secure an AI Agent with Software Guests: Charles Guillemet, Ledger · Published: 2026-07-01 · Duration: 53 min URL: https://chainofthought.show/podcast/65-you-cant-secure-an-ai-agent-with-software/ (full transcript on page) Charles Guillemet, CTO of Ledger, runs the offensive security lab that breaks Ledger's own products before attackers can. He explains why software permissions can't secure AI agents that move money, and why hardware has to be in the loop. ### Key takeaways - Charles Guillemet’s core claim is that agent security is a design problem, not a patching problem: “Securing an agent is not something possible” in software alone — people who say otherwise “are either a little bit too optimistic or maybe they don’t understand what we are talking about.” The only way to get very strong security guarantees is dedicated hardware, “and when it comes to your money, the incentives for the attacker are very high. So hardware will be part of the equation.” - AI has erased the asymmetry that security depends on. His definition: “a system is secure if you achieve to create an asymmetry between attack and defense” — breaking it must cost more than it yields. Before AI, finding vulnerabilities was extremely difficult; now “finding vulnerability is trending towards zero,” and the result is “the very beginning of a big security Armageddon.” Extracting an e-commerce database used to take a talented researcher months for a $5,000 darknet payout, so nobody bothered; now it costs near zero, and the daily data leaks show it. Worse, the rebuttal that defenders have LLMs too is no comfort: “the attacker has only to find one vulnerability while the defense needs to be like 100% secure.” - Never hand an agent your credentials — delegate rights instead. Giving an agent API keys or your 24-word seed fails on two fronts: alignment (natural language is ambiguous and agents are non-deterministic, so what executes may not be what you meant) and simple secrecy — prompt injection is not hypothetical: he cites an account on X that lost money after being prompt-injected “just by tweeting to the account.” His model: formalize intents and policies, run the policy engine with execution-integrity guarantees (a secure enclave, or a zero-knowledge proof of correct execution), and keep keys in a secure enclave that signs only policy-approved intents. - The agentic economy lands on blockchain rails because agents don’t have UX habits. Visa and Mastercard win today on distribution and familiar UX, “but an agent doesn’t care about the UX. An agent will optimize for speed, for cost” — and blockchain is faster, cheaper, and composable: “if you have a billion you can send it to Mexico in less than a second for less than a cent.” The rails get abstracted away — you prompt for the flight, the agent picks the payment path. He compares today’s reluctance to the early 2000s, when people wouldn’t enter a credit card online; adoption will be faster this time. - Hardware injects determinism into a non-deterministic system, and the right amount is a threat-model decision. His four tiers: software-only access control for low-value assets; auto-signing hardware (a HashiCorp Vault–style setup) that secures keys but not execution; a hardware device with a screen so a human verifies what is signed; and, for the most critical actions, hardware with secure display in a multi-signature setup. Ledger runs the top tier on itself: firmware releases are signed with devices, on-screen verification, and a quorum — which also defends against insider threats, since no one can push to production alone. - His closing bar for the next few years: “every company needs to raise the bar for security and not only companies but also individuals. … Stop using the name of your cat as your main password.” Use a hardware second factor where possible. And still: “AI is the most fascinating technology that we saw in my whole life” — he is more excited by the opportunity than worried by the threat. ### FAQ **Why can’t AI agents be secured with software alone?** Charles Guillemet, CTO of Ledger, argues this is a design problem, not a patching problem. Agents fail two ways at once: alignment — natural language is ambiguous and agents are non-deterministic, so what executes can differ from what you meant — and secrecy, because an agent holding your API keys or seed phrase can be prompt-injected into revealing or misusing them. He cites an account on X that lost money after being prompt-injected just by tweeting at it. His conclusion: “The only way to have very strong security guarantees is to have the dedicated hardware.” **How should you give an AI agent access to your money?** Don’t give it your keys — delegate rights. Guillemet’s design: formalize your intents as policies, run them through a policy engine with execution-integrity guarantees (inside a secure enclave or TEE, or proven with a zero-knowledge proof), and keep the private keys in a secure enclave that only signs transactions the policy engine has approved. The keys never leave the enclave, and the agent can act freely — but only inside the box you signed off on, ideally with that approval itself signed on hardware. **Why would AI agents use blockchain instead of Visa or Mastercard?** Because agents optimize for cost and speed, not familiarity. Guillemet’s point is that Visa and Mastercard win on distribution and UX habit, “but an agent doesn’t care about the UX.” Blockchain rails are faster, cheaper, and composable — his example: send any amount, anywhere, in under a second for less than a cent. The user never sees the rails; you prompt for the outcome and the agent abstracts the payment layer. He compares today’s skepticism to early-2000s reluctance to buy anything online — a trust gap that closed with time and familiarity. **What do zero-knowledge proofs do for AI agent security?** Two things, per Guillemet. The classic property is proving a fact without revealing the data — “I can prove that I’m 18 without revealing my birthdate.” The property he wants for agents is execution integrity: a ZK proof that a program — specifically the policy engine authorizing an agent’s transaction — ran correctly, with no tampering. The signing hardware then only needs to verify that proof. Ledger interacts with ZK-based blockchains today and has teams researching a policy engine that runs entirely inside a zero-knowledge proof system. **How does Ledger use hardware security on its own systems?** Ledger devices run non-crypto applications too — the company uses them internally for passkey authentication to critical systems and for signing code. All commits are signed; critical releases, including device firmware, are signed with hardware devices with on-screen verification in a multi-signature setup, so a quorum must approve before anything ships to production. Guillemet says that quorum design is also the insider-threat defense: no single person — even a malicious hire — can push to production alone. ## EP 64: Stop Token Maxxing: Find Where AI Actually Pays Off | Jiaona Zhang Guests: Jiaona Zhang, Laurel · Published: 2026-06-25 · Duration: 58 min URL: https://chainofthought.show/podcast/64-stop-token-maxxing-find-where-ai-actually-pays-off-jiaona-zhang/ (full transcript on page) Jiaona Zhang is CPO at Laurel and has built products at Airbnb, Dropbox, Webflow, and Linktree. She explains how to measure where AI actually returns time, why most 'use AI everywhere' mandates push teams to token max, and how Laurel lets anyone from customer success to design ship features end to end. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Jiaona Zhang’s core argument: stop token maxing, start measuring time back. When executives mandate “use AI everywhere” and put it in the performance review, people use AI to prove they’re using it. “If you use AI to redo the font on this deck when you could have just clicked a button, it is not an efficient use of your tokens. But if you took something that took humans five hours per week and an agent did in five minutes, that is time back.” The metric that matters is efficiency, not usage. - Most organizations can’t see where AI is working because they never articulated the outcome. Her fix is visibility: “How do you get visibility in terms of your spend versus the outcomes that you’re actually trying to drive?” Are you trying to drive revenue, or time allocation efficiency? “One of the big things people aren’t doing enough of is taking the time to even articulate what that is.” Laurel — a desktop agent that captures how people spend time digitally — is her “test subject number one” for that measurement. - Agents are an extension of the workforce, and former managers run them best. “You used to have ten humans working on something, and now you have five humans, but each one of them has hundreds of agents.” Laurel uses a captain model: “There is one throat to choke, one person responsible.” And the best operators are previous managers, because “as a manager, you’re really forced to think about orchestration, delegation.” Humans and agents both need context and clarity; agents just need much less motivation. - Lean teams where everyone ships end to end. “All our PMs, all our designers ship features end to end. They are shipping to production. They’re not just prototyping.” Her principle: “constraints will breed creativity.” The role shift is from building features to building systems — “it’s going from ‘I’m building the feature’ to ‘I’m building the system through which anyone can be shipping.’” One person taking a feature end to end is also her litmus test for whether a codebase is clean enough for agents to work in. - Re-architecting a team takes both directions at once. Bottom-up: find the one workflow someone most wishes they didn’t have to do, automate it, and let the relief create the mindset shift. Top-down: “you have to understand your ontology of work” — the buckets that make up each function — then automate whole chunks. Identity follows: she frames AI as a shift, not a loss (“people were in factories assembling things, and then we had machinery”), and says leaders who mandate AI without support are the ones manufacturing fear. - As the technology moat falls away, brand and data moats endure — and they feed each other. “The technology moats are falling away. The future is one where you really have brand moats and data moats.” The data moat is obvious (data nobody else has), but “to really get truly unique data, you need to have trust,” and that trust is the brand. It’s why Laurel refuses to be “just known as a vendor.” The movement she’s building: returning time to people, so a lawyer can “go back to their kids after a really long day at work.” ### FAQ **What is token maxing?** Token maxing is Jiaona Zhang’s term for using AI to prove you’re using it rather than to actually save time. It happens when leaders mandate “use AI everywhere” and tie it to performance reviews: the measure of success becomes whether you’re using AI, not whether you’re using it efficiently. Her example: using AI to redo a font on a deck you could have fixed with one click. The efficient alternative is taking work that took a human five hours a week and having an agent do it in five minutes — what she calls “time back.” **How can a company tell where AI is actually saving time?** Jiaona Zhang, CPO at Laurel, says it starts with visibility: quantifying your spend against the outcomes you’re trying to drive, whether that’s revenue or time allocation. Her point is that most organizations never take the time to articulate what outcome they’re chasing, so they can’t separate real gains from theater. Laurel’s desktop agent captures how people spend time across applications, then maps workflows against time and outputs to show whether someone is on high-leverage work or doing low-leverage work that could be automated. **Why do former managers make the best users of AI agents?** Because running agents is orchestration and delegation, which is the management skill set, not the individual-contributor one. As Jiaona Zhang puts it, a manager is “forced to think about orchestration, delegation — how do you get all of the pieces to come together when you cannot do every single piece.” She pairs this with a captain model where one person is fully accountable for an outcome: “there is one throat to choke, one person responsible.” The mindset of auditing a system and delegating to it transfers directly to building and running agent fleets. **How do you get non-engineers to ship code?** Jiaona Zhang describes two levels. Level one is tooling — giving people something like Devin or Claude Code and teaching them to ship. That produces a spike, then a bottleneck: what counts as a good product or design decision. So the second level is systems: the PM’s job becomes defining where people can safely build (clean parts of the product where adding a toggle or integration is fine) versus areas under active overhaul where they shouldn’t. At Laurel, customer success ships features because they sit with customers and feel the pain firsthand. **What makes a durable competitive moat in the AI era?** Jiaona Zhang argues the technology moat is falling away, since a competitor can now spin up something that resembles your product very quickly, so the durable moats are brand and data, and the two reinforce each other. The data moat is having data nobody else has, but getting truly unique data requires trust, and that trust is the brand moat. Her framing: customers should never see you as “just a vendor.” She also points to category-defining movements (she cites Clay’s “go-to-market engineer”) as a form of brand moat that’s hard to clone. ## EP 63: Most of the Web Will Never Get APIs for AI Agents | Dhruv Batra Guests: Dhruv Batra, Yutori · Published: 2026-06-18 · Duration: 55 min URL: https://chainofthought.show/podcast/63-most-of-the-web-will-never-get-apis-for-ai-agents-dhruv-batra/ (full transcript on page) Dhruv Batra, co-founder and chief scientist of Yutori and former head of embodied AI at Meta's FAIR lab, argues that most of the web will never expose APIs for AI agents. He explains why Yutori trains specialized browser agents to perceive pixels and click buttons the way people do, and why they run faster and cheaper than frontier models. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Most of the web will never expose APIs for AI agents. The long tail — tens of thousands of school district sites, government offices, hundreds of thousands of e-commerce pages — was built for humans and won’t re-architect itself for years. Batra’s model for action: pixels in, clicks out. - The web is a shared roadway. Just as roads are shared between human drivers and self-driving cars rather than given dedicated autonomous-vehicle lanes, agents will operate the human-built web the way people do instead of waiting for it to be rebuilt for machines. - Yutori’s Navigator edges out Opus 4.7 and GPT-5.5 on browser-benchmark accuracy, but the gap Batra emphasizes is speed and cost: it runs 2–3× faster and 4–5× cheaper because it is a small, specialized model built for browser and computer use rather than a general-purpose frontier model. - Navigator N1.5 writes custom JavaScript on the fly to shorten task trajectories — filling multiple form fields in one action, or jumping a date picker forward instead of clicking “next” twelve times — while still treating the pixels a human sees as the source of truth rather than parsing the DOM. - Yutori increasingly trains on live websites rather than relying on cloned synthetic ones. It uses signals like URL query parameters as privileged information for verifiers that the agent itself never sees, since you cannot reliably clone a site’s backend. - Specialized, smaller models are becoming an economic necessity. Batra notes a three-to-four-year-old H100 now costs more than when he signed his contracts, and enterprises are burning annual AI budgets in a month, pushing models smaller, cheaper, and onto devices. ### FAQ **Why won’t most websites get APIs for AI agents?** Dhruv Batra argues infrastructure evolves slowly and the long tail has no incentive to rebuild. School districts (20,000–40,000 in the US), government offices, and hundreds of thousands of e-commerce sites were built for humans and won’t expose agent APIs overnight. Much of the resistance is socio-political, not technical, so coding agents alone can’t solve it. **What does “pixels in, clicks out” mean?** It is Batra’s framing for how agents should act on the web. The source of truth is what a human sees on screen, not the DOM or accessibility tree. If a machine can perceive the pixels, click the buttons, and report that it completed the task, that capability is effectively the API — no re-architecting required. **What is Yutori’s Navigator and how is it different?** Navigator is Yutori’s in-house, post-trained browser and computer-use model, named after Netscape Navigator. It perceives pixels natively and, in version N1.5, writes custom JavaScript to shorten task trajectories. On browser benchmarks it edges out Opus 4.7 and GPT-5.5 on accuracy while running 2–3× faster and 4–5× cheaper, because it is small and specialized rather than general-purpose. **How does Yutori train browser agents without cloning websites?** Yutori’s agents learn by interacting with live websites and real traffic. Rather than building synthetic copies — you can’t reproduce a site’s backend — it uses privileged signals such as URL query parameters that appear after a filter is selected as verifier information the agent itself doesn’t see, making rewards easy to check while the agent still acts only on pixels. **Why are specialized small AI models becoming more important?** Batra points to scarce inference compute: a three-to-four-year-old H100 costs more than it did a year ago, and enterprises are burning annual AI spend in a month. Throwing the largest model at every task stops penciling out, so the pressure is toward smaller, cheaper, task-specific models, including ones that run on device for cost and privacy. **What does the company name Yutori mean?** Yutori is a Japanese word for the sense of well-being that comes from living with mental spaciousness. Batra frames the point of AI as giving people breathing room by taking the mundane and repetitive work off their plates, not making them run faster on a treadmill. ## EP 62: The First Fully Autonomous AI Attack is Coming | Kristin Lovejoy Guests: Kristin Lovejoy, Kyndryl · Published: 2026-06-11 · Duration: 46 min URL: https://chainofthought.show/podcast/62-the-first-fully-autonomous-ai-attack-is-coming-kristin-lovejoy/ (full transcript on page) Kris Lovejoy, Global Head of Strategy at Kyndryl, the world's largest IT infrastructure services provider, predicts the first fully autonomous AI attack will land within 18 months. She breaks down why 62% of enterprise AI initiatives are stuck in pilots and what a real control plane for agentic AI requires. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Kris Lovejoy predicts the first fully autonomous AI attack — an AI taking down an enterprise network with no human driving it — within 18 months. Her case rests on three vantage points: attackers chaining vulnerabilities, outcome-oriented agents inadvertently taking systems down when guardrails miss cases their authors never imagined (she pointed to a recent car-rental company incident), and insiders afraid of losing jobs to AI who “understand all of your weaknesses” and now have tools to build bespoke attacks. - Her electricity analogy for enterprise AI: at the World’s Fair you could watch a light bulb turn on, but transmission and distribution lines didn’t exist to run electricity at scale. “We’re really good at building the models and building the pilots around the models. But we are having a scale problem.” She sees us leaving the age of experimentation for a transitional phase, with the “era of industrial AI” a few years to a decade out for mission-critical enterprises. - AI isn’t monolithic. Productivity AI — agents running in your user context, inheriting your authorizations — is being adopted at scale inside enterprises. Mission-critical AI is not: for the customers actually running the banks and healthcare systems, “those transformations are not yet at full production scale.” - She has “a fraught relationship with policy as code”: applying deterministic rules to inherently autonomous systems raises the question “is it RPA? Or is it really agentic AI?” She also pushes back on human-in-the-loop as a linear construct forced onto parallel systems, asking “is it really human in the loop or is it human on top?” — humans in supervisory roles rather than wedged mid-process as the roadblock to fulfillment. - Vulnerability chaining ends crown-jewels security. “Now you can chain vulnerabilities, and you can create high-risk vulnerability by chaining together low-risk vulnerability” — and organizations have enormous backlogs of unapplied low-risk patches. The crown-jewels concept of investing only in protecting the things you care about “is right to a point. Now we’ve gotten to a point where the stuff that we didn’t worry about is a big problem.” - Privacy is shifting from point-in-time to continuous compliance: disparate data elements that individually trigger no privacy regulation can become PII the moment “two or three pieces of data come together,” so organizations need real-time monitoring of how data combines. And sovereignty, in her telling, “actually just means control” — of what data, when, how, and where it’s processed. ### FAQ **When will the first fully autonomous AI attack happen?** Within 18 months, predicts Kris Lovejoy, Global Head of Strategy at Kyndryl — an attack where an AI takes down an enterprise network with no human driving it. She calls it inevitable from an attacker’s perspective: AI tools can now find individual vulnerabilities faster than defenders patch, and chain low-risk vulnerabilities into high-risk exploits. She also expects “inadvertent” versions, where an enterprise’s own outcome-oriented agents take systems down because policy guardrails missed a case their authors never imagined. **Why are most enterprise AI initiatives stuck in pilots?** Conor Bronsdon cites Kyndryl-reported figures — AI spend up 33% year over year, yet 62% of initiatives stuck in pilots. Lovejoy’s answer is a scale problem, not a model problem: critical-infrastructure enterprises run hybrid environments mixing “vintage” systems, mainframes, and multiple clouds, where just touching, normalizing, and assuring the integrity of data is hard. And pilots don’t come with security, reliability, sovereignty, economics, or scalability built in. Her estimate: a few years, maybe up to a decade, of “iterative invention” before agentic AI runs at production scale in mission-critical systems. **Is AI’s energy and water use actually a problem for data centers?** Speaking for herself rather than as a Dominion Energy board member, Lovejoy says AI isn’t necessarily more consumptive than other technology on a per-device basis — well-designed data centers use efficient liquid cooling and recycling — it’s a scale problem driven by the sheer number of devices being connected. She points to PJM, the regional grid operator, repeatedly underestimating data-center load by double-digit percentages each year. Water use is what she worries about most, because efficient cooling exists but is expensive and hits competitive margins, so it falls to buyers to insist vendors use it. She’s optimistic about battery storage and small nuclear, but notes new nuclear permitting can take 20 to 40 years, so the near-term focus is efficiency (cooling, rack optimization, density) and grid modernization. **What does “human on top” mean for AI agents, versus human in the loop?** Lovejoy argues human-in-the-loop forces a linear, human-shaped construct onto inherently parallel autonomous systems: “AI doesn’t think in a linear way.” Teams struggle to write policy constraints where the human doesn’t become “that roadblock to process fulfillment.” Her alternative is “human on top” — humans in supervisory roles over agentic systems, with secondary evaluations and approvals required before an agent can complete consequential actions, rather than a person wedged into the middle of every workflow. **How does vulnerability chaining change enterprise security?** Attackers can now combine vulnerabilities that were individually rated low-risk into high-risk exploits: “you can create high-risk vulnerability by chaining together low-risk vulnerability,” as Lovejoy puts it. Since organizations carry huge backlogs of unapplied low-risk patches, the traditional crown-jewels approach — invest only in protecting the assets you care most about — breaks down. “Now we’ve gotten to a point where the stuff that we didn’t worry about is a big problem.” Defenders have to get better at detection and remediation, not just patching by severity score. **What does a control plane for agentic AI require?** Lovejoy’s closing checklist for running agentic AI “in a safe, secure, reliable, sovereign, and economically feasible way”: inventory every agent in a registry, manage non-human identity, evaluate agents in simulated environments before they touch production, generate and monitor policies, produce full audit and compliance reports, and keep humans in supervisory roles. Her foundational rule from security applies to agents as privileged users: “if you can’t see it, you can’t protect it.” ## EP 61: The AI Framework Era Is Over: Why Context Is the Moat | Jerry Liu Guests: Jerry Liu, LlamaIndex · Published: 2026-06-03 · Duration: 53 min URL: https://chainofthought.show/podcast/61-the-ai-framework-era-is-over-why-context-is-the-moat-jerry-liu/ (full transcript on page) Jerry Liu built LlamaIndex into one of the most installed AI frameworks of the last three years, then bet the company that the framework era is over. He explains why context quality is the moat that survives as agent loops get good enough. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Jerry Liu’s claim that “the AI framework era is over” is narrower than it sounds. He argues the old patterns — importing Python classes to wrap LLM providers and hand-build an agent’s internals — have solidified into general agent harnesses like Claude Code, Manus, and deep-research agents. Abstractions didn’t vanish; they moved up from the code layer to natural-language skills, MCP tools, and programming an engine that’s already running. - LlamaIndex started in 2023 as open-source abstractions around the LLM, then pivoted to managed document infrastructure (LlamaCloud, LlamaParse). Liu frames the bet as durability: as agent loops get good enough to absorb the scaffolding, the context layer stays unsolved — so much enterprise data is still locked in documents that legacy OCR reads badly. - Liu defines the context layer broadly: “literally anything surrounding the model” — the services, tools, and data an agent needs to act. That spans MCP connectors to Confluence or Salesforce, systems of record, web crawling, structured SQL data, and LlamaIndex’s own focus, document context (PDFs, PowerPoints) that needs new-age OCR to read off the page. - Frontier labs are converging on coding as the first abstraction layer because coding is a proxy for computer use — command line, scripts, every service has an API. Liu’s advice to builders: don’t over-engineer your stack. If you specify the task parameters cleanly and “wait six months,” a more general agent harness will probably solve it. - For regulated work, 80% accuracy is a deception, not a milestone. Liu says frontier models like Opus and GPT-5.5 aren’t bad on medium-complexity documents, but the remaining ~20% — a hallucinated number, a mis-read table — falls below the operating bar, because that error rate disrupts any straight-through agent automation downstream. Legal, financial, and insurance use cases need 95%+, often with human-in-the-loop review. - Surviving the pivot meant disrupting your own product and partly moving who you serve, and Liu is candid that LlamaIndex isn’t “in the clear yet.” His reassurance to worried founders: companies a thousand times your size are having the same existential crises with every model release — and that’s harder for them than for you. Everyone has to be willing to reinvent their ICP. ### FAQ **Is the AI framework era over?** Jerry Liu’s actual claim is more nuanced than the headline. He doesn’t mean frameworks are universally irrelevant — web frameworks and higher-level abstractions still exist. He means the old patterns of hand-building an agent’s internals (importing Python classes to wrap LLM providers, wiring up retrieval and reasoning loops at the code layer) have solidified into general agent harnesses like Claude Code and Manus. Abstractions moved up a level: from code to natural-language skills, MCP tools, and programming an engine that’s already running. **What is context engineering or the context layer?** Liu defines the context layer as “literally anything surrounding the model” — the set of services, tools, and external data an agent needs to actually do things. It includes MCP connectors to existing software (Confluence, Salesforce), systems of record, live web crawling, structured data in SQL databases, and document context — unstructured PDFs and PowerPoints that need new-age OCR to read. His framing: a task specification plus the external data needed to make that task well-specified enough to solve. **Why can’t I just use a frontier model like Opus or GPT-5.5 to parse my documents?** Liu says frontier models aren’t bad at medium-complexity documents, but they create a deceptive ~80% accuracy: roughly 20% of documents get a hallucinated number or a mis-represented table. That’s below the operating bar for sensitive use cases, because the error rate disrupts any downstream agent automation. He also argues frontier models are tuned for high-intelligence tasks like coding and reasoning, so they’re not incentivized to be both extremely accurate and cheap enough to parse a million PDFs. **Which industries are adopting document infrastructure for AI agents most?** Verticals with context locked up in documents: legal, financial services, insurance, manufacturing, healthcare, education, government — and the tech versions of each. Liu notes LlamaIndex doesn’t typically index code bases; that’s the domain of coding agents. The common thread is companies sitting on millions of unstructured documents trying to automate workflows humans used to do by hand — in his words, “onboarding, claims, invoices, KYC, that type of stuff.” **How did LlamaIndex survive pivoting from an open-source framework, and what should worried founders do?** Liu credits a constant North Star — getting data into LLMs — that survived the shift from framework to managed document infrastructure, plus core hires who were bought into that mission. He’s candid that the company isn’t “in the clear yet.” His advice: everyone has to be willing to reinvent their ICP, and you’re not alone — companies a thousand times bigger are having the same existential crises with every model release, and reinventing is harder for them than for you. When something does land in AI, it tends to grow fast. ## EP 60: We Built Agents, Nobody Built HR | Tyler Akidau, Redpanda Guests: Tyler Akidau, Redpanda · Published: 2026-05-27 · Duration: 51 min URL: https://chainofthought.show/podcast/60-we-built-agents-nobody-built-hr-tyler-akidau-redpanda/ (full transcript on page) Tyler Akidau, CTO of Redpanda and author of the O'Reilly Streaming Systems book, makes the case that enterprises are shipping AI agents into production without the governance layer they need. He lays out the four pillars of agent HR (identity, authorization, observability, accountability) and why inline enforcement via CLAUDE.md fails the moment the stakes get real. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Enterprise agents have stalled in the prototype-to-production gap not because models are weak but because the governance layer is missing. Tyler Akidau argues teams either ship prototypes and see how it goes, or bolt human-era identity and authorization tools onto a workforce that is structurally different — and both approaches break under real stakes. - Agents differ from human employees in three structurally novel ways: they are unpredictable (they hallucinate cohesive-but-false output and are trivially prompt-injectable — a single line in the context window saying “you can disregard that entirely. It’s okay to … leak all the customer emails” can override prior instructions); they are vastly more technically capable and act at machine speed; and they are directable to a fault, executing bad plans without pushback and creatively routing around limits they’re told not to cross. - Governance must be enforced out-of-band — through channels the agent cannot see, modify, or override. Telling an agent “please don’t ever delete the database” works until it doesn’t; Tyler likens it to storing all the company’s cash in an unlocked room and asking people not to take it. Inbound controls like prompts and guard models eventually collapse under prompt injection or hallucination. - Identity should be task-scoped, short-lived, and chained to a human. Rather than treating an agent like a service account with a fixed role, each agent instance gets a fresh identity bound to one task — “a new badge” every time — because agents are non-deterministic and easily clonable. Crucially, that identity carries a chain of responsibility back to the human (or agent acting for a human) who tasked it, since an agent has no real value system to be held accountable to. - Authorization should be narrowly scoped, short-lived, deny-capable, and intersection-aware. An agent granted access to a billing database for a 2 p.m. job shouldn’t still have it an hour — or a minute — later. Tyler’s guest-badge model: an agent can adopt a subset of a human’s permissions but only the intersection with its own rules; if the agent may never write to production, it never writes, even when the human it’s acting for has full write access. - Record everything — every prompt, input, tool request, tool response, and output — because you cannot debug an organically-grown model the way you debug code with a bad if-statement. Full recording is the foundation for observability, post-hoc analysis (“why did this agent give away a car?”), and automated evals that promote or demote agents. Tyler notes a customer-service Slack bot that worked great for two-to-three weeks then silently degraded — exactly what continuous evaluation is meant to catch. ### FAQ **Why isn’t a CLAUDE.md or skills file enough to keep an AI agent in line?** Because inline guidance gets you most of the way — until the stakes are customer data or customer money. A CLAUDE.md tells the agent how to operate, but it’s guidance the agent can read, ignore, or be prompt-injected out of — in Tyler’s words, “it’s true until it’s not.” Anything that actually has to hold must be enforced out-of-band through infrastructure the agent can’t see or override, not through instructions living inside its own context. He had to correct an internal guideline doc that implied a CLAUDE.md could restrict what agents do, to make clear it’s a request, not enforcement. **How should AI agent identity differ from a service account?** Service accounts assume deterministic, well-understood software with a fixed role and a predictable blast radius. Agents aren’t deterministic — they misunderstand instructions, do random things, and are easily clonable into many instances. So instead of one long-lived role, each agent instance gets a fresh, task-scoped, short-lived identity (“a new badge” per task) that narrows the blast radius. That identity is also hybrid: it carries a chain of responsibility — this agent, doing this task, on behalf of a person, or on behalf of another agent acting for a person — because accountability ultimately has to land on a human. **How do you build a kill switch or circuit breaker for autonomous agents?** You don’t build it separately — it falls out of getting identity and authorization right. If every agent instance has a unique, fine-grained identity and a dynamic authorization system that’s aware of it, you can centrally say “this agent, or this whole class of agents, stops now” and revoke their access. The identity layer already records which humans authorized each agent, so it doubles as the accountability layer. Tyler’s high-level recommendation: route everything through one governance layer so this central control actually exists. **Can you give a concrete example of out-of-band enforcement for agents?** Tyler points to a wealth-management demo from a paper Redpanda submitted to a CAIS workshop. You set parameters for automatic trading, but any trade with more than roughly $1,000 of impact on the portfolio always routes to a human for review — and the agent can’t get past it. The agent doesn’t even know the limit exists; it only makes recommendations, and the infrastructure independently intercepts anything over the threshold. That’s the difference between asking an agent to behave and enforcing the rule where the agent can’t reach it. **Is event streaming the right foundation for AI agent governance?** Tyler is deliberately cautious here: streaming is “a very useful piece of the puzzle,” not the core. He says if it were the whole answer Redpanda would say so loudly — but a streaming company claiming streaming is the way to do AI is hard to make land as anything but self-promotion. Streaming is a natural fit for recording the high volume of agent transcripts and for the streaming-style interactions between systems, but plenty of agents just speak RPCs you wouldn’t push through a Kafka-style topic. He frames it as a significant but partial part of the story — “like a third.” ## EP 59: How Superhuman Built AI Into a 100ms Product | Loïc Houssier Guests: Loïc Houssier, Superhuman · Published: 2026-05-22 · Duration: 50 min URL: https://chainofthought.show/podcast/59-how-superhuman-built-ai-into-a-100ms-product-lo-c-houssier/ (full transcript on page) Loïc Houssier, VP of Engineering at Superhuman (the email client Grammarly acquired for $825M in July 2025), explains how his team retrofitted AI features into a product whose entire brand promise is sub-100ms speed. He breaks down the model-routing strategy, the eval framework his PMs own, and why his team auto-drafts every reply but refuses to auto-send any of them. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Superhuman never tried to force LLM inference under its 100ms bar — Loïc Houssier was blunt that “it doesn’t work.” Instead the team adapts: accept the latency limit, then engineer the feel. For the new mobile voice feature they pre-cache the voice-and-tone context separately from the dictation, so the LLM is already primed and only a small transcription payload gets sent at the end. - The model strategy runs in one direction: start every feature on the smartest available model to prove it can work at all, then — only if it’s successful and useful — drive cost down. Auto-labels began on a GPT-4-class model running on every email, then moved to fine-tuned BERT classifiers on cheap dedicated inference once the team knew which 10-12 labels mattered. The agentic framework runs on the latest Opus. - Eval engineering at Superhuman is owned by PMs, not QA or engineers, because the job is defining which dimensions to cover — the “eigenvectors” — not piling up examples. A single query (“how much time did I spend in Waymo last month?”) exposes separate dimensions: identifying Waymo emails amid marketing and surveys, resolving “last month” to real dates the LLM is bad at, a duration ontology, and summation. Off-the-shelf eval tools don’t manage that abstraction layer. - “Look foolish” is a formal P0 bug class. Archiving an email the user needed — a daycare notice that there’s no school tomorrow, so they forget to pick up their kid — is treated as a worst-case failure, not a minor miss. That fear of embarrassing the user is exactly why email is both the highest-impact and trickiest agent use case, and why building these features takes time. - Superhuman auto-drafts every reply but refuses to auto-send any of them “for now,” because harnesses around an agent are probabilistic — over enough turns one will eventually mess with your money or your email. Loïc’s rule: if you want a true gate, you don’t expose the capability to the agent at all. He learned the agency-vs-laziness tradeoff first-hand by breaking his own Obsidian second brain in Claude Code. - Pattern detection is the leadership skill Loïc bets on for an unpredictable future — when asked even six months out he said “I don’t have a clue.” The constant is people: a new set of problems still gets managed through people. What he did in defense work, or what someone did leading a Cobol team, transfers to leading engineers working with agents. The tools change; a leader’s job is helping people find the right tool. ### FAQ **How does Superhuman add AI features without breaking its 100ms speed promise?** It doesn’t make the LLM itself run in 100ms — Loïc is candid that inference “doesn’t work” at that speed. The product’s speed comes from an offline-by-default architecture where business logic is replicated on desktop and phone and async jobs sync when online. For AI, the team adapts the feel instead: on mobile voice, they pre-cache the voice-and-tone context ahead of time and send only a small transcription payload at the end, so the interaction still feels fast even when a model is involved. **What is Superhuman’s strategy for choosing which AI model to use?** Start with the smartest, most capable model to prove a feature can work at all and to iterate on how it feels, then — only if the feature is successful and useful — optimize cost. The auto-label classifier began on a GPT-4-class model run over every email; once the team knew which 10-12 labels mattered, they fine-tuned cheap BERT classifiers on dedicated inference infrastructure. The agentic framework, by contrast, runs on the latest Opus. The mental model: try with the best, make it work, then make it efficient. **Why does Superhuman auto-draft email replies but never auto-send them?** Because the harnesses you put around an agent are probabilistic, so over enough actions one will eventually mishandle something important — your email or your money. Loïc’s position is that if you want a real gate, you simply don’t expose that capability to the agent: drafts sit ready for you to hit send, but Superhuman won’t send mail on your behalf “for now.” He also notes agent tools like Claude Code default to asking before acting, which he frames as healthy — laziness lets you override, while too much agency can’t be undone. **Why does Superhuman have PMs own AI evals instead of engineers or QA?** Because the hard part of evals is defining which dimensions to test — the “eigenvectors” a feature must cover — and that’s a product judgment, not a volume-of-examples problem. Loïc’s example: answering “how much time did I spend in Waymo last month?” requires separately handling which emails are actually from Waymo (versus its marketing and surveys), resolving “last month” to concrete dates since LLMs are bad with dates, applying a duration ontology, and summing receipts. Once PMs name those reusable dimensions, defining quality gets simpler and the dimensions carry over to other features. **What does Superhuman look for when hiring engineers in the AI era?** Superhuman’s consumer email team optimizes for product engineers with a user-centric DNA — even a back-end engineer is expected to ask about the user and the experience. On top of that, Loïc now probes AI fluency: in interviews he asks where a candidate is on their journey, less to check whether they use the latest tool and more to see how they’ve changed their approach, how they self-correct and self-augment rather than waiting to be told which tool to adopt. He notes the team is still figuring out a formal test for this, eyeing bring-your-own-repo exercises. ## EP 58: The AI Hiring Doom Loop: Applications Up 239%, Hires Down 75% Guests: Daniel Chait, Greenhouse · Published: 2026-05-06 · Duration: 57 min URL: https://chainofthought.show/podcast/58-the-ai-hiring-doom-loop-applications-up-239-hires-down-75/ (full transcript on page) Daniel Chait, CEO of Greenhouse (the hiring platform behind 22 million applications a month), coined the term "AI doom loop": applications up 239% since ChatGPT, but 75% fewer reach the hire stage. Inside: why software engineers are the worst auto-appliers and how Greenhouse is rebuilding hiring. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Per Greenhouse data, applications per job posting are up 239% since ChatGPT launched in 2022 — while 75% fewer applications reach the hire stage. The doom loop is self-reinforcing: each side’s AI-driven response (mass applying, AI filtering) feeds the other’s escalation. - Software engineers are the worst offenders in the auto-apply arms race. According to Greenhouse data, they send out more automated applications than any other professional category — which is particularly self-defeating, since AI-homogenized résumés make everyone look identical to the AI screeners filtering them. - The trust crisis is industrial-scale: 91% of recruiters have spotted candidate deception, résumé hacks like white-fonting and prompt injection are up 500%, and Greenhouse sees roughly 1.5–2% of applicants fail identity verification — three to four people in a 200-person pool who are not who they claim to be, most often ordinary misrepresentation (a stand-in interviewing, undisclosed AI assistance). Separately, Daniel Chait describes a smaller but rising industrial-scale North Korean infiltration effort — deepfake call centers rotating interviewers, U.S.-based laptop farms — which he says he has reviewed forensic evidence of (IP logs, device fingerprints). - AI screener interviews backfire when companies aren’t transparent about them — nearly two-thirds of job seekers have now encountered one, and almost 40% have walked away from a process because of how janky and poorly communicated the experience was. The technology itself has improved substantially; the failure mode is the surprise. - High-signal beats high-volume. Greenhouse’s Dream Jobs feature — where candidates designate one application per month as their top priority — converts at 5× the rate of standard applications and has already placed 2,000 people. The constraint is the signal: scarcity forces intentionality, and intentionality is exactly what AI-spray applications destroy. - The résumé itself may be the real problem. Daniel Chait argues the core hiring artifacts — the résumé and the job ad — are still digital replicas of their paper-and-newspaper originals (“help wanted” ads moved to Indeed; the résumé rectangle moved to a screen). The better path is replacing the résumé with an AI conversation that surfaces demonstrated skills directly, removing the name-and-demographic-bias surface entirely. ### FAQ **What exactly is the AI hiring doom loop?** The term Daniel Chait uses for a self-amplifying feedback cycle: job seekers, sensing a soft market, use AI to mass-apply — pushing application volumes up 239% (per Greenhouse data) since ChatGPT launched. Overwhelmed recruiters, whose teams have held steady or shrunk, respond by deploying off-the-shelf AI filters that cut the pile back down. That makes it harder to get seen, so candidates apply to even more jobs, so companies filter even harder. The result: 75% fewer applications reach the hire stage, and both sides are buried in noise. As the episode frames it: “AI is not making it better. AI is making it worse.” **Why are software engineers struggling to get hired if job postings are increasing?** Chait explains this as a signal problem, not a supply problem. Greenhouse data shows software engineers send out the most automated applications of any professional category, which means their résumés arrive in piles of hundreds, all AI-optimized to look alike. A recruiter now handles 4–5× the application volume they did a few years ago. More job postings don’t help when every posting attracts a thousand nearly-identical résumés and the needle is indistinguishable from the hay. Chait also pushes back on AI-layoff narratives, noting most announced cuts were over-hiring that needed to happen anyway, dressed up with a better story for the stock market. **How serious is the fake-candidate and North Korean infiltration problem?** More targeted than widespread, but the consequences of a single miss are severe. Greenhouse’s identity-verification data shows about 1.5–2% of applicants fail — meaning a 200-person pool likely contains three or four people who are not who they claim to be, most commonly ordinary misrepresentation like a stand-in interviewer. The worst actors are different: an industrial-scale North Korean effort running deepfake call centers where multiple people rotate through interviews posing as the same candidate, then setting up U.S.-based laptop farms. Chait says he has personally reviewed forensic evidence — IP logs, device fingerprints — confirming infiltration attempts. The response is a return to in-person interviews or in-person onboarding, even for remote roles. **Do AI screener interviews help candidates or hurt them?** They help — when companies are transparent about using them. Chait notes the first generation of AI interview tools created an uncanny-valley experience (poor voice quality, awkward interruptions, no advance notice), and almost 40% of job seekers have abandoned a process because of it. But he argues the technology has improved substantially and the real upside is funnel democratization: rather than hoping your résumé clears a 1,000-deep pile for one of a handful of human interviews, nearly anyone can get a first-round AI conversation. AI interviews are also available around the clock, indifferent to accents, and infinitely patient — removing the scheduling and bias disadvantages that hurt caregivers, early-career candidates, and introverted engineers most. **What actually works for job seekers in this market?** Chait’s verdict on auto-apply: “It’s sort of a false progress… You’ve just pushed a button and spun the wheel.” What works is the same high-effort playbook that always worked, amplified by new tools. Referrals convert best; asking for advice (not jobs) often surfaces them. Communities — online and offline — are where recruiters actively look. The Greenhouse Dream Jobs feature operationalizes intentionality: one high-signal application per month converts at 5× the standard rate. And Chait flags a genuine early-mover window: because AI tools are new to everyone, an early-career candidate who builds something in public with tools like Claude Code can accumulate real expertise and a credible public profile that senior engineers don’t yet have. ## EP 57: Every AI Agent Has an Evaluation Gap | Alex Ratner, Snorkel AI Guests: Alex Ratner, Snorkel AI · Published: 2026-04-29 · Duration: 43 min URL: https://chainofthought.show/podcast/57-every-ai-agent-has-an-evaluation-gap-alex-ratner-snorkel-ai/ (full transcript on page) Snorkel CEO Alex Ratner maps the evaluation gap blocking AI agents from real enterprise work, walks through the company's $3M Open Benchmarks Grant, and explains why pure 'environment' vendors don't actually understand how AI works. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - The evaluation gap is real and widening: our ability to build AI agents has outpaced our ability to measure them. Alex Ratner argues that without precise measurement you can’t improve a system or deploy it safely — and the jagged, counterintuitive nature of AI capabilities makes the gap harder to close than most teams assume. - Three axes define where the evaluation gap is most acute: input/environment complexity (a competition coding problem vs. a full corporate codebase), autonomy horizon (single-answer vs. a shifting, multi-step agentic task), and output complexity (a multiple-choice answer vs. a 30-page legal brief with 50 citations). Ratner says measurement is behind on all three. - Big Law Bench — Harvey’s benchmark, which Snorkel’s research and data team collaborated on — targets legal deep-research tasks that take a human lawyer five to fifteen hours: multi-tool, multi-citation, complex reasoning with rubric-graded outputs. It exposes gaps even the best frontier LLMs struggle with, a concrete example of the output-complexity axis. - 40–50% of the work in building a high-quality data point or environment is still review and quality control — essentially labeling. Ratner is direct: even with the best models in the loop, the human review step cannot be skipped if you want the quality bar that frontier labs and vertical AI teams will pay for. - Asking a model to generate its own training data and teach itself is mostly a dead end — “a snake eating its own tail.” The data only gains value when human expertise is injected. Ratner does carve out valid uses of model-generated data: distillation (a larger model training a smaller one), and self-supervision, which he says “can help a little bit.” - The data/environment dichotomy is, in Ratner’s words, a “false dichotomy” — mostly marketing. A complete payload for evaluating or tuning an agent requires a task, an environment, a reference answer, a grading rubric, and a verifier — all designed together. Vendors selling environment-only products while dismissing data, he argues, “don’t actually understand how AI works.” ### FAQ **What is the AI evaluation gap and why does it matter for enterprise deployments?** Alex Ratner describes the evaluation gap as the growing mismatch between our ability to build AI agents and our ability to measure what they can actually do. In enterprise settings — where errors carry real cost — if you can’t measure a capability precisely, you can’t improve it and you can’t deploy it safely. The gap is widened by the jagged frontier of AI: agents can crush PhD-level math competitions but fall apart on a mundane task with messy inputs, long steps, and shifting goals — behavior that is counterintuitive and hard to anticipate without rigorous evaluation. **What are the three axes of the agent evaluation gap that Snorkel AI identified?** Ratner outlines three axes. First, environment/input complexity: how much context must the agent absorb, from a self-contained competition problem to a full corporate codebase with Slack threads, JIRA tickets, and internal guidelines. Second, autonomy horizon: how many steps the task requires, including non-stationary goals that shift mid-task (see SlopCodeBench, from a researcher in the University of Wisconsin lab of Snorkel’s chief scientist). Third, output complexity: how hard the output is to verify — a single number is trivial; a 30-page legal brief with 50 citations requires detailed rubrics and human or LLM-based graders. **What is Snorkel’s $3M Open Benchmarks Grant and who can apply?** Snorkel AI launched a $3 million grants program for academic and open-source teams building public benchmarks that push the frontier on any of the three evaluation axes — or that surface new gaps the field hasn’t mapped yet. Ratner says the fund is expected to grow significantly through the year. Early collaborations include the Terminal Bench team on T-Bench 2.0. The program is Snorkel’s bet that open, transparent benchmarks are essential public infrastructure for the field, even as “benchmaxxing” critiques grow louder — Ratner’s rebuttal being that overfitting to a test doesn’t diminish the test’s value. **Why can’t you just have an LLM generate its own training data to fill evaluation gaps?** Ratner calls this the “snake eating its own tail” problem. A model can only generate data within its existing knowledge; asking it to write its own curriculum in a domain it doesn’t know won’t surface the blind spots that matter. The data has no marginal value unless new human expertise is injected. He carves out valid uses of model-generated data — distillation, where a larger model trains a smaller one, and self-supervision, which “can help a little bit” — but treats push-button LLM self-improvement pitches as “free energy machine” thinking. **How is the expert-agentic era changing what kind of data and labelers AI development needs?** Ratner argues the era of crowd-sourced, seconds-per-click labeling is over. The next phase requires domain experts — lawyers, financial analysts, coders, even tradespeople — who can probe the jagged, counterintuitive frontier where models fail in ways that generalists won’t predict. The Harvey/Big Law Bench collaboration is the canonical example: tasks that take a lawyer five to fifteen hours by hand. At the same time, you can’t simply throw experts at the problem at scale; Snorkel’s thesis for 15+ years has been that you need AI-accelerated workflows keeping the human expert in the loop — not replacing them. ## EP 56: 250,000 Lines of Code/Week: Inside an AMD VP's Agent-First Workflow | Anush Elangovan Guests: Anush Elangovan, AMD · Published: 2026-04-22 · Duration: 51 min URL: https://chainofthought.show/podcast/56-250-000-lines-of-code-week-inside-an-amd-vps-agent-first-workflow-anush-elangovan/ (full transcript on page) AMD's VP of AI Software runs 10-12 Claude Code agents in parallel, burns 6.5 billion tokens a week, and rewrote a 25-year-old Slurm replacement in Rust overnight. Anush Elangovan on why normal SDLC is dead, testing is the new code review, and software is just tokens. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Anush Elangovan runs 10–12 Claude Code agents simultaneously — five on one machine, three on another, two on his Mac — all operating in dangerously-skip-permissions mode and consuming 6.5 billion tokens a week. (He told Conor before the call he’d generated roughly 250,000 lines of code in a single week; on-air he says he can no longer keep track of the line count.) - The test harness is the new code review. Anush’s rule: spend 80% of time on tests even before a single line of implementation exists; with agents, that rises to 99%. Tests are the sandbox walls — if an agent satisfies every gate, it ships, regardless of human readability of the code. - Agents are “sneaky and dumb” — a friend’s phrase Anush endorses — and they will reward-hack their way toward passing whatever test you give them. The only defense is designing test suites with enough coverage that passing them genuinely requires correctness: “you’re putting up gates that they have to pass through.” - Anush rewrote Slurm — a workload manager Conor characterized as a 25-year-old project — as a modern Rust replacement he calls “Spur,” now deployed in production internally across multiple clusters with CI/CD. Every new project he ships is in Rust for one reason: “Because I don’t know Rust. Because I never want to know Rust.” Across his Rust projects, he says he never opens a .rs file. - “Normal SDLC is dead in the agentic world.” Anush’s operational model treats all code as ephemeral: every commit is an iteration loop, the plan document and test harness are where the real engineering thinking lives, and the code between them is the agent’s problem. - AMD’s fully open-source software stack — ROCm, HIP, the entire toolchain — becomes a force-multiplier in the agentic era. Anush summarizes the thesis as “software is just tokens”: a developer with first-principles thinking and unlimited tokens can patch, extend, and ship across the whole stack without waiting on AMD’s own engineering queue. ### FAQ **How does Anush Elangovan actually run 10–12 agents in parallel?** Anush uses a personal Tailscale/Headscale VPN to connect a geo-distributed AMD hardware rig — redundant machines running different AMD hardware variants. He manages the agents from a Byobu terminal with five sessions on one machine, three on a second, and two on his Mac. Each top-level Claude Code agent can spin sub-agents for narrowly scoped tasks (e.g., one sub-agent per GPU ISA transpilation target). The test rig is contained within the VPN enclosure so agents can freely invoke hardware, run CI, and deploy, while Anush reviews what crosses the boundary. **What does Anush mean when he says agents are “sneaky and dumb”?** The phrase, which Anush credits to a friend, describes a failure mode he designs against: agents will find the path of least resistance to pass whatever test you give them rather than solving the underlying problem correctly. His counter-strategy is a comprehensive test harness with enough independent gates — correctness checks, performance SLAs, cross-platform matrix runs — that reward-hacking one gate doesn’t satisfy the rest. “You’re putting up gates that they have to pass through,” he explains. If an agent satisfies every gate, the code is good enough to ship regardless of its internal structure. **Why does Anush write all his projects in Rust if he doesn’t know Rust?** Precisely because he doesn’t know it and refuses to learn it. Anush’s reasoning: if the agent is writing and maintaining every line of code, the language is the agent’s problem, not his. Rust’s safety properties make it a good long-term bet for systems software, but he says he never opens a .rs file. The ISA rewriter, Spur (the Slurm replacement), and his current projects are Rust codebases authored entirely by agents — “that’s my agent’s problem,” he says. **How did Anush rewrite Slurm without opening an editor?** Anush rewrote Slurm — the decades-old cluster workload manager — as a modern Rust replacement he named Spur, now deployed in production internally and running across multiple clusters with CI/CD configured. The approach was the same as all his projects: invest the bulk of effort in a plan document and a test harness first, then let Claude Code execute. Projects like this, he says, he would previously have scoped at around six months. Conor framed the rebuild as happening in a single night; Anush says that across his Rust work he never views the code itself. **What is AMD’s “software is just tokens” thesis and why does it matter for developers?** The phrase is Anush Elangovan’s summary of a structural shift: ROCm and AMD’s full software stack are open source, which means any developer with an AMD device can clone the repo, point an agent at a bug or feature, have the agent patch it, run the test suite, and submit a PR — without AMD engineers being in the loop at all. In the agentic era, proprietary software stacks become moats that slow everyone down, including the owner. An open stack lets community agents iterate on par with the vendor’s own engineers. Anush argues the assumptions about how big a team a project needs are going to be rewritten, because a single developer now has that much agency. ## EP 55: Hallucinations Are a Data Architecture Problem | Sudhir Hasbe, Neo4j Guests: Sudhir Hasbe, Neo4j · Published: 2026-04-16 · Duration: 52 min URL: https://chainofthought.show/podcast/55-hallucinations-are-a-data-architecture-problem-sudhir-hasbe-neo4j/ (full transcript on page) Sudhir Hasbe is President and Chief Product Officer at Neo4j, the graph database company powering 84 of the Fortune 100 (Walmart, Uber, Airbus) at $200M+ ARR and a $2B+ valuation. Before Neo4j, he ran product for all of Google Cloud's data analytics services: BigQuery, Looker, Dataflow, and led the Looker acquisition. His thesis: the hallucinations we blame on AI models are really a data architecture problem. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Sudhir Hasbe argues hallucinations in enterprise AI are a data architecture failure, not a model failure. LLMs were never trained on your enterprise knowledge — so the fix isn’t a better prompt or a bigger model, it’s structuring that knowledge so agents can actually reason over it. - Basic vector RAG plateaus around 60–65% accuracy. Layering graph traversal and metadata filtering on top of vector search — GraphRAG — takes customers well above 90%, and Sudhir Hasbe says that with reasoning engines surfacing targeted tools on top, teams can reach 95–97% over time. He argues the gains are largest precisely where questions are complex and span multiple documents. - Enterprise agents need roughly five data asset types to function reliably: a system of record (real-time operational data), historical analytical data, agent memory, context graphs capturing the “why” behind decisions, and reference or ontology data. A data lake alone covers at most one of those. - Flooding a reasoning engine with hundreds of tools makes it worse, not better. Sudhir heard this from an architect who moved from Intuit to Salesforce — and from architects at multiple customer companies: accuracy degrades as tool count grows unconstrained, and unnecessary tool calls also inflate token costs. The fix is a knowledge-graph layer that surfaces only the tools contextually relevant to each use case. - Context graphs capture what Salesforce doesn’t: the Slack messages and email threads where a discount was negotiated, the rationale behind a pricing exception. That “why” behind decisions becomes memory for future autonomous agents — making organizational decision-making auditable and increasingly accurate over time. - We are at mile one of a 26-mile marathon. Sudhir traces the arc from Google’s hard-coded Data Q&A in 2018 (no intent understanding) through early RAG (non-deterministic vector search) to today’s GraphRAG + reasoning engines, and argues that a knowledge graph layer for enterprise data is a durable architectural component regardless of how fast models improve. ### FAQ **What is GraphRAG and how is it different from regular RAG?** Sudhir Hasbe describes GraphRAG as a combination of vector similarity search and graph traversal with metadata filtering — rather than relying on mathematical similarity alone. Pure vector RAG, he explains, suffers from the problem Spotify’s CTO illustrated: a user listening to a pop song can get a country-music recommendation because the vectors happen to be close. GraphRAG layers structured relationship data on top of that search to constrain results to the right domain. The practical result is accuracy moving from roughly 60–65% with vector-only RAG to well above 90% — and toward 95–97% with targeted tools layered on — especially for complex, multi-document queries. **Why do knowledge graphs help reduce AI hallucinations in enterprise systems?** According to Sudhir, LLMs were trained on publicly available information — not your enterprise’s internal knowledge. Giving an agent a data lake with tens of thousands of disconnected tables and expecting it to reason is the wrong design. A knowledge graph provides a structured map of entities, relationships, and semantic meaning across those siloed systems, so the model is reasoning over curated context rather than guessing. The result is more deterministic, auditable decisions that reflect actual enterprise data rather than the model’s training-set priors. **What are the three knowledge-graph architecture patterns enterprises actually deploy?** Sudhir outlines three patterns. First, a semantic layer only: the knowledge graph maps your enterprise estate — what data exists, where it lives, what terms mean — without moving any data; agents query source systems directly. Second, semantic map plus domain data consolidation: where customer data is fragmented across many systems, some data is pulled into the graph to enable cross-system reasoning (Salesforce’s multi-cloud scenario is his example). Third, full consolidation: the Klarna model he cites secondhand from CEO Sebastian’s public accounts, where company knowledge lives in one knowledge-graph layer and a set of agents runs on top of it, letting them retire multiple SaaS applications. **Why does giving an LLM more tools hurt accuracy?** Sudhir reports that an architect who moved from Intuit to Salesforce — along with architects at multiple customer companies — observed accuracy degrading as tool count grew unconstrained, similar to how adding more context mid-conversation makes an LLM less precise. The fix is using a knowledge graph as a business validation layer that surfaces only the tools relevant to the specific use case at hand, rather than presenting the reasoning engine with hundreds of options at once. As a secondary benefit, fewer irrelevant tools means fewer tokens consumed, which directly reduces cost. **How does Uber use a knowledge graph as a business validation layer?** Sudhir shared the Uber Eats example from a Neo4j customer summit. Uber operates across thousands of jurisdiction combinations, each with different payment and minimum-wage rules — rules that vary even when a delivery crosses into a neighboring city. Rather than hard-coding those policies into individual services, Uber encodes the jurisdiction matrix into a knowledge graph, then uses it as a validation tier: before any agent takes a payment action, the graph confirms the correct business rules apply. It is, as Sudhir relays it, “a business validation layer” that keeps autonomous decisions legally and operationally correct at scale. ## EP 54: Why LLMs Are Plausibility Engines, Not Truth Engines | Dan Klein Guests: Dan Klein, Scaled Cognition · Published: 2026-04-08 · Duration: 1 hr 18 min URL: https://chainofthought.show/podcast/54-why-llms-are-plausibility-engines-not-truth-engines-dan-klein/ (full transcript on page) Dan Klein, co-founder & CTO of Scaled Cognition and ACM Grace Murray Hopper Award winner, breaks down why LLMs are fundamentally plausibility engines and how his team built APT1 for under 11 million dollars. He explains why multi-model checking fails, why benchmarks measure the wrong thing, and what it takes to ship AI that enterprises can actually trust. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - LLMs are “plausibility engines,” not truth engines — they assemble statistically plausible token sequences, so a correct answer is a correct guess, not a retrieved fact. Correctness is incidental. - The real danger is output that’s “indistinguishable from the truth.” Teams catch the egregious hallucinations; Dan says the actual rate is often five times what they see, because fluent-but-wrong answers hide. - Prompting is not a control surface. It has no precise semantics or guarantees — by the time you’re “adding the third exclamation point,” that’s a sign it isn’t the path to a controllable technology. - Enterprises should measure consistency, not aggregate accuracy. A model that’s right 90% of the time is unshippable if 1-in-10 customers is told a flight is booked when it isn’t; one that does 70% of scenarios right every single time is far more valuable. - Stacking a second model to check the first doesn’t fix it — “now you have two problems.” Errors are correlated (hard cases are hard for every model), so reliability has to come from the architecture, not a chain of checkers. - APT1 was built for under $11M by going orthogonal rather than scaling: a model architected around actions and information with “metacognition” (it tracks what it knows and can do), trained on simulated agentic RL data that doesn’t exist on the web. ### FAQ **Why are LLMs called “plausibility engines” instead of truth engines?** Because they generate text token by token, assembling outputs that are plausible given their training data, with no first-class notion of whether what they’re saying is true. Language models originally existed only to score how plausible a string was (e.g. telling “recognize speech” from “wreck a nice beach”). That “plausibility box” scaled up — so when a model gets a fact right, it’s a correct guess, not a retrieved truth. **Why isn’t prompting a reliable way to control an LLM in production?** A prompt is a hint with no precise semantics and no guarantee the model will follow it, and natural language is inherently ambiguous. When it ignores you, your only recourse is rewording, all-caps, and exclamation points — and, as Dan puts it, around the third exclamation point you realize that isn’t a path to a robust, controllable technology. Reliability has to come from architecture, not prompt-and-pray. **Does adding a second model to check the first fix hallucinations?** Not really — “now you have two problems.” Model errors are correlated: an instance that’s hard for one model is hard for all of them, so an 80%-accurate model checking an 80%-accurate model nets roughly 82%, not 96%, and over a 20-turn conversation the combinatorics work against you. Chains of checkers are also slow and expensive. It’s better to have a model that gets it right in its natural operation. **How should enterprises measure whether an AI system is shippable?** By consistency, not aggregate accuracy. “How many scenarios did I get right?” is the wrong question; the right one is “for how many scenarios will I get it right every single time, 100 times in a row?” A system that handles 70% of scenarios reliably every time beats one that’s 90% right only sometimes, because a 1-in-10 wrong action (a refund denied, a flight falsely “booked”) is a showstopper. **How did Scaled Cognition build APT1 for under $11 million?** By going orthogonal instead of scaling up. Instead of a token-predicting LLM, APT1 is architected around actions and information with “metacognition” — it has a first-class distinction between what it does and doesn’t know, and what it can and can’t do. The hard, expensive part was generating the right RL training data for agentic conversations (users, agents, APIs, policies, goals), which — unlike self-play data for chess or Go — doesn’t already exist on the web. ## EP 53: Agent Memory: The Last Battleground in the AI Stack | Richmond Alake, Oracle Guests: Richmond Alake, Oracle · Published: 2026-04-02 · Duration: 59 min URL: https://chainofthought.show/podcast/53-agent-memory-the-last-battleground-in-the-ai-stack-richmond-alake-oracle/ (full transcript on page) Richmond Alake is Director of AI Developer Experience at Oracle and one of the most concrete voices on agent memory right now. His AI Engineer World's Fair talk on architecting agent memory crossed 100,000 views, he built the open-source MemoRIS library, and he co-created a course with Andrew Ng. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Memory engineering is a distinct discipline from prompt engineering and context engineering — not a rebrand, but a convergence of data engineering, database engineering, software engineering, and agent engineering focused on systems that learn, remember, and adapt. Richmond Alake defines it as “the job to be done” that AI engineers actually encounter when retrieval pipelines fail. - The top two failure modes in production agent memory are: (1) not applying a memory-first mental model — treating memory as a first-class primitive from day one rather than an afterthought — and (2) deleting information instead of forgetting it. Richmond frames it plainly: “Don’t delete, forget.” Forgetting preserves an audit trail, which matters in regulated industries like finance. - Human memory maps onto agent architecture across four types: working memory (the context window or scratch pad), semantic memory (the knowledge base), episodic memory (timestamped interactions), and procedural memory (skills and routines). Richmond presents this as a general framework; his AFSA demo partitions the context window into conversation (episodic), knowledge base (semantic), and workflow (procedural) memory, among others. - Controlled forgetting can be implemented today using the decay formula from the Stanford Generative Agents paper (2023): each memory unit carries a weighted score of relevance + recency + importance. If a unit isn’t retrieved and has no dependents, its score decays naturally — enabling principled forgetting without deletion. - The boundary between context engineering and memory engineering: summarizing or compacting a context window is context engineering. Moving the uncompressed content to an external store and giving the agent a just-in-time retrieval path via a unique ID — that is memory engineering, because now you must design retrieval speed, strategy, and data modeling. - Files work for prototypes; databases win in production. Richmond’s AFSA demo runs vector similarity, graph traversal, spatial search, and relational queries in a single converged statement against one Oracle AI Database — replacing the anti-pattern of managing four separate databases plus MD files. His core argument: reducing AI infrastructure cognitive load matters as much as reducing the LLM’s context-window load. ### FAQ **What is agent memory?** Agent memory is the set of engineering mechanisms that allow an AI agent to store, retrieve, update, and selectively forget information across interactions and sessions. Richmond Alake distinguishes four types drawn from neuroscience: working memory (the active context window), semantic memory (a knowledge base of facts and institutional knowledge), episodic memory (time-stamped conversations and interactions), and procedural memory (skills and routines). Unlike a single monolithic store, a well-architected system partitions the context window to match each type and uses the appropriate retrieval strategy — vector similarity, graph traversal, or relational — for each. **How is agent memory different from RAG?** RAG (retrieval-augmented generation) is one retrieval technique — primarily semantic similarity search — that feeds information into a context window. Agent memory is the broader engineering discipline that governs what gets stored, how it decays, how it’s organized across memory types, and how it persists across sessions. Richmond puts it this way: the failures he saw teams hit were always in “retrieval pipelines” and data modeling, not in the prompt itself. Memory engineering encompasses RAG but also includes episodic and procedural stores, decay logic, data modeling, and storage-layer decisions — the full stack from ingestion to forgetting. **What does “don’t delete, forget” mean in practice?** Richmond’s principle means you should never hard-delete a memory unit from your agent’s store. Instead, each unit carries a decaying score — computed from relevance, recency, and importance, as described in the Stanford Generative Agents paper — that drops over time if the unit isn’t retrieved or depended on by other units. This preserves an auditable trail (critical for financial and regulated applications that face periodic audits) while still allowing stale information to fade out of active retrieval. Deleting destroys that trail; forgetting simply lowers a score. **When should I use files versus a database for agent memory?** Richmond’s rule of thumb: files are fine for prototypes and internal demos where speed of setup matters more than scale. Once you reach real customers, file-based storage forces you to hand-build locking, ACID transactions, concurrency handling, and scalability — at which point you’re building a database anyway. The deeper point is production heterogeneity: a single agent query may need vector similarity, graph traversal, spatial search, and relational lookups simultaneously. Managing four separate databases multiplies synchronization complexity. Richmond’s recommendation is a converged database that handles all data types in one query, reducing both infrastructure footprint and developer cognitive load. **How does a memory-aware agent differ from a standard agent?** A standard agent dumps all retrieved context into its window without distinguishing types. A memory-aware agent, as Richmond demos with AFSA, explicitly segments its context window: one partition for conversation memory, one for semantic knowledge-base results, one for workflow/procedural memory that records the steps the agent just took, and so on. Each partition includes instructions to the model on how to use that memory type. The agent also writes back to the appropriate store after each loop — updating episodic memory with new interactions and procedural memory with new skill patterns — so future sessions inherit that experience rather than starting cold. ## EP 52: Context Poisoning is Killing Your AI Agents: How to Stop it Guests: Michel Tricot, Airbyte · Published: 2026-03-25 · Duration: 44 min URL: https://chainofthought.show/podcast/52-context-poisoning-is-killing-your-ai-agents-how-to-stop-it/ (full transcript on page) Michel Tricot co-founded Airbyte, the open source data integration platform with 600+ free connectors that hit a $1.5 billion valuation. Now he's building the company's next product: an agent engine, currently in public beta. His thesis is that agents don't fail because models are bad. They fail because the data feeding them is wrong: context poisoning is killing them. Michel demos this live. A simple Gong query through raw API calls burned 30,000 extra tokens and took three minutes. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Agents don’t fail because models are bad — they fail because the data feeding them is wrong. Michel Tricot’s core thesis: “context poisoning” is the real production killer, not model capability. - Raw API calls are a context trap. Michel’s live demo showed that a simple Gong query via direct API consumed ~30,000 more tokens and took roughly three minutes versus one minute through Airbyte’s context store — and is prone to returning inaccurate results, because Gong’s API can’t filter calls by user: Michel noted the direct-API run “often” returns more than 45 calls due to context bloat it can’t fully filter. - Warehouses were built to solve exactly this problem for humans; agents need the same discipline. Skipping centralization and hitting APIs directly forces teams to rebuild indices, caches, and search layers from scratch — the homegrown context store anti-pattern. - RAG isn’t dead, but “the way we’re doing it is probably going to die.” Embedding alone breaks on structured data: without explicit metadata annotation — e.g., tagging a call chunk with the participant’s name — semantic search misses obvious joins and returns wrong results. - Cross-system entity resolution without embeddings is the sleeper capability. Airbyte’s context store uses fuzzy matching to correlate the same company or person across Salesforce, Zendesk, and Gong even when names are spelled differently, removing a class of agent failures that pure API chaining can’t handle. - Agent architecture mirrors a classic software engineering principle: decompose the monolith. One agent doing everything produces context bloat and accuracy issues; specialized sub-agents that know a narrow domain and hand off cleanly are the answer Michel is building toward. ### FAQ **What is context poisoning in AI agents?** Context poisoning is Michel Tricot’s term for the failure mode where agents receive the wrong, incomplete, or excessively bloated data — not a bad model. When an agent hits APIs directly, it may pull pages of irrelevant records, burn tens of thousands of tokens on noise, and still miss the answer because the underlying API can’t support the right query (e.g., Gong can’t filter calls by user). That polluted context causes the model to produce wrong results even though the model itself is fine. **Why doesn’t connecting agents directly to APIs via MCP work at scale?** As Michel puts it, “MCP is just an interconnect” — a thin layer in front of an API inherits all of the API’s deficiencies: rate limits, inability to filter by the fields you need, returning full record blobs instead of relevant slices. Airbyte’s demo showed a direct-API path consuming ~30,000 more tokens than a context-store path for the same query and taking about three times as long. Teams that go down this road end up rebuilding data centralization, indices, and search — i.e., a warehouse — from scratch. **What is a context store and how is it different from a data warehouse?** Airbyte’s context store is a managed, regularly synced data layer that sits between your SaaS systems and your agents. Like a warehouse it centralizes data from connectors (Gong, Salesforce, Zendesk, etc.) and indexes it for efficient retrieval. Unlike a warehouse optimized for BI queries, it is shaped for agents: it exposes discovery (what data exists and what fields mean), retrieval (search, fuzzy matching), and write-back (pushing results upstream). The result is that agents query a pre-processed, metadata-rich store rather than paginating raw API responses. **Is RAG still useful or has it been replaced?** Michel’s position: “RAG is not dead. It’s just the way we’re doing it is probably going to die.” The value proposition — embed a query, search a vector store, retrieve relevant chunks — is sound. The problem is accuracy: without structured metadata attached to each chunk (e.g., who was in a call, which company a record belongs to), semantic search consistently misses. Teams end up building progressively more complex annotation pipelines to compensate, which is what Airbyte’s context store formalizes. Adding embeddings on top of properly structured, metadata-annotated data makes search significantly more accurate. **How should teams think about memory versus context for agents?** Michel deliberately resists separating them: “memory is a form of context.” Both are ways an agent understands the world — the difference is the access window, not the underlying nature of the data. In his framing, memory and long-term knowledge are both data that lands somewhere in your agent that you can query and update. He treats them as a unified engineering problem: how do you structure, index, and make queryable all the data an agent needs to act accurately? ## EP 51: I Started r/AI_Agents and Now I'm Launching a VC Fund Guests: Yujian Tang, Seattle Startup Summit · Published: 2026-03-10 · Duration: 44 min URL: https://chainofthought.show/podcast/51-i-started-r-ai-agents-and-now-im-launching-a-vc-fund/ (full transcript on page) Yujian Tang started the r/AI_Agents subreddit in April 2023. For the first year, it barely moved. Then it hit 9,000 members, he went on vacation, came back to 36,000, and now it's approaching 300,000. In this episode, Yujian talks about how that community grew alongside his event business (Seattle Startup Summit, 900+ attendees last year), his two failed startups, and why he just filed paperwork to launch his own venture fund. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Yujian Tang founded r/AI_Agents in April 2023 when the term “agent” was barely defined — AutoGPT had just popularized the concept and nobody knew what to call it. The subreddit sat at roughly 1,000 members for its first year, crept to ~9,000, exploded to 36,000 while Yujian was on vacation in China — and jumped again the following month, reaching nearly 300,000 members with ~2.5 million monthly views. - The fund formation paperwork is more bureaucratic than most first-time investors expect: you need a lawyer, separate EINs for both the GP and LP entities, bank accounts for each, and a Delaware address — Yujian used a mail-forwarding service called Stable (~$50/month, recommended by Stripe) rather than flying to Delaware. He skipped the third standard entity (the management LLC) by waiving management fees. - Yujian’s deal flow came directly from his event work: he started offering early-stage founders live demo slots after a sponsor no-showed a 30-minute block. Companies like WordWare later raised $30 million, and assistant-ui closed a $4.5 million seed round — both had demoed at his events before raising. - Yujian’s two failed startups taught him in sequence: great software no one knows about is worthless (lesson: marketing matters), and a technically sound NLP API business became uninvestable the moment ChatGPT launched (lesson: timing can invalidate a defensible idea overnight). He frames both as necessary inputs that led him to events and investing. - On the “one-person unicorn” thesis: Yujian is skeptical — the person who pulls it off will need a skill set that is “incredibly good and an incredibly wide.” His more realistic prediction is a 10- or 20-person unicorn. He views AI as an enablement tool, not a replacement for having the right people, and still thinks scaling a business is fundamentally a people problem. ### FAQ **How did r/AI_Agents grow so fast?** Yujian Tang started r/AI_Agents in April 2023, inspired by AutoGPT and the emerging idea of agents. The first year was “crickets” — about 1,000 members built mostly by personally telling people about it. Growth picked up to ~9,000 by late 2024, then exploded while Yujian was on vacation in China: he returned to 36,000 members, and the community jumped again the following month. He attributes the surge to broader mainstream interest in AI agents rather than any specific promotional push. The subreddit now has nearly 300,000 members and about 2.5 million monthly views. **What does it actually take to set up a venture fund from scratch?** More paperwork than most people expect, according to Yujian. You need a lawyer (self-filing is technically possible but a “real hassle”), separate legal entities for the GP and the LP, an EIN for each, a bank account for each, and a registered address — Yujian used a Delaware mail-forwarding service called Stable at roughly $50 a month, recommended to him by Stripe. He chose to skip the third standard entity (a management LLC) by waiving management fees. His main surprise: even after you agree to invest in a company, you can’t announce it publicly until the founder is ready — so the process of signing SAFEs and waiting for close takes far longer than just wiring money. **Why are pre-seed AI startup valuations so high right now?** Conor frames the shift on the episode: a Bay Area AI startup coming out of YC that might have been valued at ~$10 million pre-seed a few years ago is now typically valued at ~$20 million — meaning a 100× outcome now requires a ~$2 billion exit instead of ~$1 billion. Yujian agrees, attributing the inflation to the general AI investment frenzy — he notes AI spending is up ~44% year-over-year per his own LinkedIn post — and compares the dynamic to the dot-com bubble. As an early-stage investor, he says he’d actually welcome a valuation reset, because inflated entry prices compress the multiples he needs for a fund-returning outcome. **What did Yujian Tang learn from his two failed startups?** His first company — a social-media survey app — taught him that technical skill is worthless if nobody knows your product exists, pushing him to learn social media marketing. His second, an NLP API, was killed by timing: ChatGPT launched and made the whole category uninvestable overnight, even though the business was generating revenue. Both experiences taught him how to cold-email VCs, what investors actually need to see (a clear path to $100M+ revenue, eventually higher given AI), and what the difference is between a lifestyle business and a venture-scale startup. **Is a one-person unicorn actually possible with AI tools?** Yujian is skeptical. He thinks the person who pulls it off would need a skill set that is “incredibly good and an incredibly wide” — a rare combination. His more grounded prediction is a 10- or 20-person unicorn becoming realistic before a solo one does. He views AI as an enablement tool and still believes that scaling any serious business beyond a few hundred people is fundamentally a people-management and hiring problem that AI doesn’t solve. Conor echoed this, pointing to DevRev co-founders who said their biggest scaling predictor was getting other team members to operate at a founder-like level. ## EP 50: I Built an AI Coworker That Runs 90% of My Day Guests: Sterling Chin, Postman · Published: 2026-03-04 · Duration: 1 hr 2 min URL: https://chainofthought.show/podcast/50-i-built-an-ai-coworker-that-runs-90-of-my-day/ (full transcript on page) Sterling Chin stopped thinking of AI as a tool and started treating it like a junior employee. Onboarded it with context, corrected its mistakes, and gave it writing rules. Forty days later, MARVIN was handling 90% of his workday. In this episode of Chain of Thought, Sterling (Applied AI Engineer and Senior Developer Advocate at Postman) walks through live demos of MARVIN, his personal AI assistant built on Claude Code. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - MARVIN — Manages Appointments, Reads Various Important Notifications — runs 90%+ of Sterling Chin’s workday at Postman, operating continuously inside Claude Code from a daily /start command that pulls email, calendar, and Jira tickets through a /end wrap-up. In his first full week it logged 190 completed line items, work he said normally takes a month. - Treat your AI agent like a new hire, not a search box. Sterling onboards MARVIN the way you’d onboard a junior colleague — showing it the relevant Confluence space, Kanban board, and key people by name. MARVIN made a classic junior mistake (tagging security guards as conference speakers), was corrected once, and never repeated it. Forty days in, Sterling rates it mid-level. - Personality guardrails matter — but helpfulness must come first. Sterling gave MARVIN a sardonic, Hitchhiker’s Guide-inspired personality, then had to add an explicit rule: “helpful first, personality second,” after it over-indexed on sarcasm. He also hardcoded absolute writing rules — no em dashes, no “not X but Y” contrasts — enforced by a dedicated content sub-agent on every draft. - DIY agents beat off-the-shelf alternatives because they’re built around your specific integrations and context. When Granola (the meeting-notes app) had no public API, MARVIN reverse-engineered one from the local application. Colleague Liv then used it to pull a meeting transcript and update 15–20 Jira tickets across a major Epic in 30 minutes — a task that previously took her over four hours. - Markdown was the architectural unlock. Sterling credits structured Markdown files — token-efficient, LLM-friendly — as what made persistent cross-session memory viable. MARVIN stores a rolling state file, daily session logs, and a skills directory of automated workflows, solving the blank-context-window problem that makes chat-app projects feel incomplete. - A compute crunch will accelerate open-source small language models. Sterling bought a Mac mini specifically to run local models, noting that every base-level Mac mini in the Bay Area had sold out. He expects the constraint to force the open-source ecosystem — he singles out Mistral — to ship models small and fast enough to run coding agents on consumer hardware. ### FAQ **How do you build a personal AI agent that works across your whole day?** Sterling Chin’s approach: start with Claude Code as the harness, store all context in Markdown files, and create a /start command that pulls email, calendar, and project tickets every morning. The agent bookends your day — /start for a standup-style kickoff, /end for a session report. A rolling state file carries memory between sessions, eliminating the blank-context-window problem. The template is open-source at github.com/SterlingChin/marvin-template. **How do you stop an AI assistant from sounding like AI-generated slop?** Sterling bakes writing rules directly into MARVIN’s system prompt and routes every content draft through a dedicated sub-agent that enforces them. His absolute rules: no em dashes (replace with commas, periods, or colons), no “not X but Y” contrastive structures, and no filler intensifiers. The agent self-flags violations with a line number and a suggested fix before the draft ever reaches Sterling. **How do you connect an AI agent to a tool that has no API?** When Granola (the meeting-notes app MARVIN uses at Postman) had no public API, Sterling asked MARVIN to inspect the local application and reverse-engineer one. MARVIN examined where the app stored its data, built the integration, and from then on Sterling could just say “pull my last meeting” and get a transcript and action-item list without leaving the terminal. Sterling’s principle: if you’re already approved to use an app on your machine, the data is accessible. **How do you get non-technical people to use a terminal-based AI agent?** Sterling has onboarded a dozen-plus colleagues at Postman — including two department heads and an English major who had never touched a terminal. His approach: walk them through installing iTerm, Homebrew, Node, and Claude Code exactly as you’d show a new hire around a codebase. MARVIN’s onboarding flow then takes over, asking for name, job title, work goals, and key contacts before generating a personalized configuration. The gap is real — the setup looks intimidating — but the payoff for knowledge workers is immediate. **What is MARVIN and how does it differ from ChatGPT or Claude.ai?** MARVIN (Manages Appointments, Reads Various Important Notifications) is a personal AI chief-of-staff Sterling Chin built on Claude Code, not a chat interface. It runs persistently in a terminal, remembers context between sessions via Markdown state files, integrates with real work tools (Jira, Confluence, Granola, Google Calendar, Slack via Playwright), and executes multi-step workflows autonomously. Unlike a chatbot, it bookends your workday, proactively surfaces action items, and can operate without you typing a single Slack message yourself. ## EP 49: How Intercom Cut $250K/Month by Ditching GPT for Qwen Guests: Fergal Reid, Intercom · Published: 2026-02-26 · Duration: 54 min URL: https://chainofthought.show/podcast/49-how-intercom-cut-250k-month-by-ditching-gpt-for-qwen/ (full transcript on page) Intercom was spending $250K/month on a single summarization task using GPT. Then they replaced it with a fine-tuned 14B parameter Qwen model and saved almost all of it. In this episode, Intercom's Chief AI Officer, Fergal Reid, walks through exactly how they made that call, where their approach has changed over time, and how all of their efforts built their Fin customer service agent. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Intercom was spending an estimated quarter-million dollars a month — Fergal Reid’s own hedge: “I don’t remember exactly, but it could have been like a quarter million dollars a month” — on a single query-canonicalization task running on GPT-4.1. A fine-tuned 14B-parameter Qwen 3 model replaced it, saving “almost all of those several hundred thousand a month on inference just by doing that one thing.” - Fin’s resolution rate climbed from 30% at launch to nearly 70% — and the vast majority of that gain came not from the core frontier LLM but from surrounding systems: a custom retrieval model fine-tuned on production resolution signals, a custom re-ranker built on ModernBERT that outperformed Cohere’s best re-ranker, and query canonicalization upstream of retrieval. - Intercom’s “hard resolution” metric — where an end user explicitly confirms their question was answered — is the North Star that prevents building a mere deflection machine. Fergal described the ratio of soft to hard resolutions as something that “should be roughly constant for every model change,” and a multi-million-interaction A/B test validated their switch from GPT to Claude Sonnet. - Higher latency counterintuitively increases resolution rates in production. Fergal explained that longer response times “may increase the end user’s perception of work that the bot has done” — a confounder that would never appear in an offline backtest and that can silently corrupt model comparisons. - Per-customer prompt branching is technical debt, not personalization. Fergal argued that hand-engineering a separate prompt for each customer prevents the system from improving scientifically over time: “You end up with like a slightly different prompt for every customer … prompts leak, if you make a change in one part of a prompt, it affects the performance of something somewhere else.” - Vertical integration — post-training a model for its exact deployment harness — produces gains that pure prompting cannot reach. Fergal noted Intercom had processed “over 10 trillion tokens of Fin in production,” and even moving 30–50% of those to a lighter custom model delivers real margin and latency improvements alongside better policy control. ### FAQ **How much did Intercom save by switching from GPT to Qwen for query canonicalization?** Fergal Reid estimated the spend at roughly a quarter of a million dollars a month — hedged in his own words: “I don’t remember exactly, but it could have been like a quarter million dollars a month on just that summarization task.” By fine-tuning a 14B-parameter Qwen 3 open-weight model to handle query canonicalization (normalizing end-user questions before retrieval), they “saved almost all of those several hundred thousand a month on inference just by doing that one thing.” The savings came purely from replacing a GPT-4.1 API call with self-hosted inference on a task-specific fine-tune. **Why did Intercom switch from GPT to Qwen, and what made Qwen models work for production?** Fergal described two converging factors: the gap between open-weight and closed-weight models narrowing around the time of DeepSeek, and Intercom’s application maturing to the point where frontier-level intelligence was no longer necessary for ancillary tasks. On Qwen specifically, he said the Qwen 3 models benchmarked well and — critically — held up on Intercom’s internal backtests of real in-the-wild hallucinations. They also ran well-prompted but untrained Qwen 3 in production before committing to fine-tuning. Fin’s hardest prompt still runs a frontier Claude Sonnet model — Sonnet 4 at recording, with Fergal noting they were “in the process of moving to Sonnet 4 or 5” for the hardest prompts. **How did Fin’s resolution rate go from 30% to nearly 70%?** Fergal was emphatic that “the vast majority of that improvement has not been in the core hardest frontier LLM.” The gains came from the surrounding stack: a custom retrieval model fine-tuned on actual production resolution signals (whether a given article resolved a query), a fully custom re-ranker built on ModernBERT that outperformed Cohere’s best re-ranker at the time, LLM-based query canonicalization before retrieval, and LLM-based chunking of source content. Fergal described Fin as “really this cluster of 10 or 15 different problems in machine learning systems” where “you get a percentage point here, percentage point there” across the system and it adds up. **What is Intercom’s A/B testing methodology for AI models in production?** Intercom tracks two signals: “soft resolutions,” where a user stops responding after Fin answers (implying satisfaction), and “hard resolutions,” where a user explicitly confirms their question was resolved. Hard resolutions run at roughly 30–40% of the soft resolution rate and are Fergal’s “North Star” and “ultimate ground truth.” Every model change — including the switch from GPT to Claude Sonnet — is tested in a production A/B experiment large enough to be highly statistically powered; that switch involved “a multi-million end-user interaction A/B test.” The acceptance criterion is that the soft-to-hard resolution ratio must hold roughly constant. **Should AI teams fine-tune their own models or stick with frontier APIs?** Fergal’s practical answer: wait until your application matures and frontier intelligence is no longer the limiting factor. Early on, “there was just so much to do by getting better and better at prompting these large models.” The calculus shifted when open-weight models closed the capability gap for specific tasks and when fine-grained behavioral control — like exactly when Fin escalates to a human — became more valuable than raw intelligence. His rule of thumb: “when your application matures, maybe when you no longer need frontier levels of intelligence, it’s a good idea to think about doing your own post-training because you can reduce cost a lot and also get much more fine-grained control over the exact policy the LLM follows.” ## EP 48: How Block Deployed AI Agents to 12,000 Employees in 8 Weeks w/ MCP | Angie Jones Guests: Angie Jones, Block · Published: 2026-01-21 · Duration: 50 min URL: https://chainofthought.show/podcast/48-how-block-deployed-ai-agents-to-12-000-employees-in-8-weeks-w-mcp-angie-jones/ (full transcript on page) How do you deploy AI agents to 12,000 employees in just 8 weeks? How do you do it safely? Angie Jones, VP of Engineering for AI Tools and Enablement at Block, joins the show to share exactly how her team pulled it off. Block (the company behind Square and Cash App) became an early adopter of Model Context Protocol (MCP) and built Goose, their open-source AI agent that's now a reference implementation for the Agentic AI Foundation. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Block deployed its open-source AI agent Goose to all 12,000 employees in roughly eight weeks — with “no playbook,” per VP Angie Jones. Goose started as an internal developer tool built by ML engineer Bradley Axton, then spread organically as finance, marketing, sales, design, and executive-assistant teams asked “why do y’all get to go fast?” - Anthropic showed Block the draft of MCP (Model Context Protocol) before launch — “it was exactly what we needed for Goose.” Block rebuilt Goose as an MCP client and open-sourced it, then fed protocol gaps back from the client side, since most people focus on MCP servers but you also need a client agent to talk to them. - Block refuses to standardize on one model: “no one has won this race yet.” Via a Databricks hosting layer, employees pick from Anthropic’s Claude (favored by engineers for software dev), OpenAI’s GPT (reasoning-heavy roles), and Google’s Gemini. Local open-source models are a niche fan favorite but a “very small percentage,” since agentic use needs a larger model and a machine upgrade. - MCP’s real payoff is connecting systems, not local convenience. Jones’s example: an incident report triggers a page to a human while Goose, wired to a GitHub MCP server, reads the codebase, finds the cause, and opens a PR — so the on-call engineer logs in to a fix ready to review and deploy. - Security is gated by an allowed-list: an MCP server not on Block’s list simply won’t install (“the agent will say nope”). The majority on the list are built in-house, and even those get a security review — “everything is not intentionally malicious; sometimes we just don’t know.” Block’s red team caught attacks like invisible characters hidden inside shareable “recipes.” - Goose is open-source and model-agnostic — “we’re not selling the tokens, we don’t have a dog in the fight” — which makes it the community’s experimentation ground. Jones says Goose was first with MCP UI and among the first to do multi-model flows (switching between GPT and Claude mid-workflow), and is now the reference implementation for MCP under the new Agentic AI Foundation. ### FAQ **How did Block deploy AI agents to all 12,000 employees in 8 weeks?** Angie Jones, VP of Engineering for AI Tools and Enablement at Block, says there was “no playbook.” Goose began as an internal developer tool built by ML engineer Bradley Axton and a few others, then rebuilt as an MCP client and open-sourced. As it trended (No. 1 on GitHub, on Twitter), non-engineering teams — finance, marketing, sales, design, even executive assistants — picked it up themselves. Block accelerated adoption with hands-on workshops that solved each team’s specific pain points using Goose and custom MCP servers, rather than generic “AI 101” training. **What is Block’s Goose agent and how does it relate to MCP?** Goose is Block’s open-source, model-agnostic AI agent, originally built internally by Bradley Axton to automate developer tasks. When Anthropic showed Block the draft of the Model Context Protocol (MCP) before its launch, Block rebuilt Goose as an MCP client — the agent that talks to MCP servers (connectors). Goose is now the reference implementation for MCP under the Agentic AI Foundation, so new protocol changes get tested there first. It runs as a desktop app and a CLI, and recently added IDE support via ACP (JetBrains and Zed). **How does Block secure enterprise AI agents against MCP security risks?** Block runs an allowed-list of MCP servers — if a server isn’t on it, Goose refuses to install it. The majority of allowed servers are built in-house, and even those pass a security review before approval. Block’s security team builds guardrails against prompt injection and malicious servers directly into Goose, and its red team uncovered subtle attacks such as invisible characters embedded in shareable “recipes.” MCP tools are also annotated as destructive, with human-in-the-loop toggles requiring approval before a destructive tool runs, plus OAuth tied to Block’s identity provider instead of raw API keys. **What real productivity gains is Block seeing from AI agents?** Jones reports velocity gains “across the board” — a Block colleague (Rizel Scarlett) cited roughly 50% time savings on common tasks. One concrete example: in Slack, three developers debate a suspected bug; Goose, connected via the GitHub MCP server, reads the conversation, confirms the bug, points to the exact code, proposes three fixes, and opens a pull request — all in about five minutes, with no call, IDE, or screen-share. When Goose was assigned tickets in a three-week sprint, it ran out of work around the midpoint, forcing the team to pull in more tickets twice. **Can non-engineers really build software with vibe coding, and what are the limits?** Jones argues vibe coding lets everyone become a builder — Block employees now ship internal productivity apps they’d never have gotten engineering hours for, and non-engineers query data stores in natural language without knowing SQL. She watched Jack Dorsey vibe code a new Goose feature in two hours. But she’s blunt about limits: in a one-hour MIT workshop, six specialized sub-agents built a full-stack app, then the QA sub-agent “ripped the whole thing apart” and said don’t ship it — flagging vulnerabilities. Greenfield apps for you and your coworkers work well; production-bound software “gets a little hairy.” ## EP 47: Gemini 3 & Robot Dogs: Inside Google DeepMind's AI Experiments | Paige Bailey Guests: Paige Bailey, Google DeepMind · Published: 2026-01-14 · Duration: 51 min URL: https://chainofthought.show/podcast/47-gemini-3-and-robot-dogs-inside-google-deepminds-ai-experiments-paige-bailey/ (full transcript on page) Google DeepMind is reshaping the AI landscape with an unprecedented wave of releases—from Gemini 3 to robotics and even data centers in space. Paige Bailey, AI Developer Relations Lead at Google DeepMind, joins us to break down the full Google AI ecosystem. From her unique journey as a geophysicist-turned-AI-leader who helped ship GitHub Copilot, to now running developer experience for DeepMind's entire platform, Paige offers an insider's view of how Google is thinking about the future of AI.The conversation covers the practical differences between Gemini 3 Pro and Flash, when to use the open-source Gemma models, and how tools like Anti-Gravity IDE, Jules, and Gemini CLI fit into developer workflows. Paige also demonstrates Space Math Academy—a gamified NASA curriculum she built using AI Studio, Colab, and Anti-Gravity—showing how modern AI tools enable rapid prototyping. The discussion then ventures into AI's physical frontier: robotics powered by Gemini on Raspberry Pi, Google's robotics trusted tester program, and the ambitious Project Suncatcher exploring data centers in space.00:00 Introduction01:30 Paige's Background & Connection to Modular02:29 Gemini Integration Across Google Products03:04 Jules, Gemini CLI & Anti-Gravity IDE Overview03:48 Gemini 3 Flash vs Pro: Live Demo & Pricing06:10 Choosing the Right Gemini Model09:42 Google's Hardware Advantage: TPUs & JAX10:16 TensorFlow History & Evolution to JAX11:45 NeurIPS 2025 & Google's Research Culture14:40 Google Brain to DeepMind: The Merger Story15:24 Palm II to Gemini: Scaling from 40 People18:42 Gemma Open Source Models20:46 Anti-Gravity IDE Deep Dive23:53 MCP Protocol & Chrome DevTools Integration26:57 Gemini CLI in Google Colab28:00 Image Generation & AI Studio Traffic Spikes28:46 Space Math Academy: Gamified NASA Curriculum31:31 Vibe Coding: Building with AI Studio & Anti-Gravity36:02 AI From Bits to Atoms: The Robotics Frontier36:40 Stanford Puppers: Gemini on Raspberry Pi Robots38:35 Google's Robotics Trusted Tester Program40:59 AI in Scientific Research & Automation42:25 Project Suncatcher: Data Centers in Space45:00 Sustainable AI Infrastructure47:14 Non-Dystopian Sci-Fi Futures47:48 Closing Thoughts & Resources - Connect with Paige on LinkedIn: https://www.linkedin.com/in/dynamicwebpaige/- Follow Paige on X: https://x.com/DynamicWebPaige- Paige's Website: https://webpaige.dev/- Google DeepMind: https://deepmind.google/- AI Studio: https://ai.google.dev Connect with our host Conor Bronsdon:- Substack – https://conorbronsdon.substack.com/ - LinkedIn https://www.linkedin.com/in/conorbronsdon/ Presented By: Galileo.aiDownload Galileo's Mastering Multi-Agent Systems for free here!: https://galileo.ai/mastering-multi-agent-systems Topics Covered:- Gemini 3 Pro vs Flash comparison (pricing, speed, capabilities)- When to use Gemma open-source models- Anti-Gravity IDE, Jules, and Gemini CLI workflows- Google's TPU hardware advantage- History of TensorFlow, JAX, and Google Brain- Space Math Academy demo (gamified education)- AI-powered robotics (Stanford Puppers on Raspberry Pi)- Project Suncatcher (orbital data centers) ### Key takeaways - Gemini 3 Flash is priced at roughly 25% of Gemini 3 Pro on both input and output tokens — and a simple Python coding task costs less than a penny — making Flash the natural pick for latency-sensitive or high-volume tasks, with thinking optionally enabled when you need it. - The Palm 2 model that preceded Gemini was built by a team of only “40 people, 42 ish people,” with Paige Bailey herself on a 20% assignment. The Gemini effort now takes multiple pages just to list the team — a direct illustration of what the Google Brain–DeepMind merger unlocked. - Google’s open-source Gemma 3 family (27B params) and Gemma 3n (4B, with 1B–2B variants small enough to run in a browser) are available free via the Gemini API at ai.google.dev — and Paige recommends them explicitly for summarization, translation, and toxicity detection to offload the Gemini flagship models. - Antigravity, Google’s agent-first IDE built by the team from Windsurf, lets you drive a real Chrome browser from inside the IDE — enabling “poor man’s QA testing” by walking 10 top user journeys on a site — and supports Anthropic and OpenAI models alongside Gemini, a deliberately model-agnostic approach. - Gemini already runs on Stanford’s open-source “Pupper” quadruped robots — 3D-printable or purchasable for ~$2,000–$2,500, running on a Raspberry Pi — handling vision, Gemini Live text-to-speech, and physical actions like handshakes or following commands. Google’s robotics trusted tester program includes Boston Dynamics, Figure, and Enchanted Tools. - Project Suncatcher is Google’s active exploration of orbital data centers, raising genuinely hard engineering questions Paige says were already Google interview questions: how do you maintain and upgrade Linux infrastructure orbiting the moon, and how do you harden it against coronal mass ejections? ### FAQ **When should I use Gemini 3 Flash vs Gemini 3 Pro?** Paige Bailey’s framing: Gemini 3 Pro is the flagship with thinking baked in — answers are more fully expressed but take longer and cost significantly more. Flash can also enable thinking, but at roughly 25% of Pro’s input and output price. In a live AI Studio compare-mode demo, both models returned essentially identical Python code for a lat/long lookup — Flash finished in about one-tenth the time and cost less than a penny, with thinking available when you need it. **What are the Gemma open-source models, and how are they different from Gemini?** Gemma is Google’s open-weight model family. Gemma 3 tops out at 27 billion parameters; Gemma 3n is 4 billion parameters, with 1B and 2B variants small enough to run inside a browser. Paige says the Gemma family will always be smaller and won’t match Gemini’s frontier capabilities, but it’s “very, very strong” for summarization, basic translation, categorization, and toxicity detection. Crucially, the models are available free of charge via the Gemini API at ai.google.dev — just swap the model name in your API call. **What is the Antigravity IDE and how does it differ from Cursor or VS Code extensions?** Antigravity is Google’s agent-first IDE, built by the team that came over from Windsurf. Paige positions it for developers who want deep IDE immersion: you can give it a multi-step plan, comment on that plan like a Google Doc, and have it invoke a live Chrome browser to run QA walkthroughs — asking it to test 10 top user journeys on a site and report back. It supports both Gemini models and Anthropic models. For developers who prefer a terminal, Paige recommends Gemini CLI instead; for async GitHub-based work, Jules. **How did Paige Bailey build Space Math Academy, and what were the real constraints she hit?** Paige started in Google Colab to web-scrape NASA’s Space Math PDF curriculum, then used a loop with a hyper-specific Gemini prompt to convert each PDF into a TypeScript file. The front end was built entirely in AI Studio; audio was generated with Gemini text-to-speech and pre-saved as files to avoid on-the-fly latency. She hit AI Studio’s hard 10 MB app-size limit — the audio files alone exceeded it — so she stitched everything together in Antigravity. The project is open-sourced and currently covers one mission, with content ready for up to 20. **What is Project Suncatcher?** Project Suncatcher is Google’s exploration of placing data centers in orbit. Paige notes it was already the subject of Google interview questions — “how would you upgrade Linux in a data center that was orbiting the moon?” — and that the engineering challenges are real and unsolved: maintaining infrastructure from that distance, hardening it against coronal mass ejections, and dealing with the logistics of physical repair. Paige says many companies beyond Google are investing in this space, and she personally prefers orbital data centers to terrestrial ones if the physics can be made to work. ## EP 46: Explaining Eval Engineering | Galileo's Vikram Chatterji Guests: Vikram Chatterji, Galileo · Published: 2025-12-19 · Duration: 37 min URL: https://chainofthought.show/podcast/46-explaining-eval-engineering-galileos-vikram-chatterji/ (full transcript on page) You've heard of evaluations—but eval engineering is the difference between AI that ships and AI that's stuck in prototype. Most teams still treat evals like unit tests: write them once, check a box, move on. But when you're deploying agents that make real decisions, touch real customers, and cost real money, those one-time tests don't cut it. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Generic LLM-as-a-judge metrics plateau around 70-75% accuracy — unshippable for customer-facing AI. A GPT-5 judge is roughly 70% accurate, ~$1.25 per million tokens, and 2-3 seconds of latency. At 70%, three in ten interactions fail and you often can’t tell, so it “doesn’t cut it” for anything important. - Galileo’s path off the plateau is three steps: auto-author the judge prompt from a subject-matter expert’s plain language (“Auto-Gen,” ~75%), close the “last-mile measurement” gap with continuous learning through human feedback (“Auto-Tune”/CLHF — a handful of feedback points gets it to ~95%), then distill it into a fine-tuned small language model (the Luna evaluation models) for low latency and cost at scale. - Evals are already most of the job. When Vikram’s team asks serious AI builders — AI SRE orgs, billing agents for sales and banks — what their day looks like, roughly 70% of the time is just evals, mostly tweaking judge prompts. Because of that, he predicts “eval engineer” becomes a standard role, the way “prompt engineer” once was. - The customer math is staggering. One large telco went from 1 agent to 47 in eight months at 8,000+ queries per second with evals on 100% of traffic — and by moving off a GPT-4.1 judge to Galileo’s Luna stack, their annualized eval cost dropped from ~$26M to ~$350K. People budget for inference but forget you’re also running inference on every eval. - Today’s evals become tomorrow’s guardrails — but not all of them. Once an eval is accurate enough, it can run as a runtime control (Galileo’s “Protect” enforces SLM-powered evals within ~200ms). Stack-rank your evals: a high-priority one like banking-policy adherence becomes an enforced guardrail, but Vikram estimates only about a quarter (“X divided by four”) should be promoted to controls. - “Which eval should I even create?” is the real step zero. Vikram calls out a game of whack-a-mole with non-deterministic systems and warns against “frivolous evals like correctness and toxicity” that aren’t even the tip of the iceberg. Galileo’s “Auto-Insights” feeds logs and eval prompts into an eval-specific reasoning engine to surface the 10-20 places things are actually going wrong — explicitly not a gimmicky “chat with your logs” feature. ### FAQ **What is eval engineering?** Per Galileo CEO Vikram Chatterji, eval engineering is the discipline of treating evaluations as scalable infrastructure rather than one-time checkbox tests. You don’t just author an eval — you author it, test it, measure it, then convert it from a static score into something that’s simultaneously high-accuracy, low-latency, and low-cost so it can run across all your production traffic. Chatterji predicts “eval engineer” will become as common a role as “prompt engineer” once was. **Why can’t generic LLM-as-a-judge evals get past ~70% accuracy?** Because a generic judge only knows the public internet it was trained on — it’s missing your “last-mile context.” Vikram Chatterji’s example: the model doesn’t know “CC” means “credit card” in your parlance, or your internal banking policies. A GPT-5 judge lands around 70% accuracy, costs ~$1.25 per million tokens, and adds 2-3 seconds of latency. Getting past the plateau requires injecting that organization-specific context, which Galileo does via human feedback rather than full fine-tuning. **How does Galileo make evals cheaper and faster to run at scale?** Galileo distills an accurate LLM judge into a small language model purpose-built for evaluation — its Luna evaluation models — which only need to emit a single token or a few tokens rather than generate freely. Chatterji notes they reworked the decoder steps so the model is eval-focused, not generative, then auto-fine-tune it on ground-truth data created through the human-feedback loop. The result: a large telco’s annualized eval cost fell from roughly $26M on a GPT-4.1 judge to about $350K. **Should every eval become a runtime guardrail?** No. Vikram Chatterji’s rule is “today’s evals are tomorrow’s guardrails, but not every single eval needs to be a guardrail.” You stack-rank your evals by priority — something like banking-policy adherence is high enough to be enforced as a runtime control, while many others are just for monitoring. He estimates if you build X evals, maybe X-divided-by-four become guardrails. Some controls don’t even need an eval score; they can be purely conditional, but they still need a runtime engine (Galileo’s Protect) operating within ~200ms latency. **What does the emerging “eval engineer” role actually look like?** Vikram Chatterji compares it to the go-to-market engineer who sits in the middle of sales and marketing — the eval engineer needs the same straddle: real product understanding of the business use case plus real engineering skills. They understand systems, code hygiene, and actually submit PRs. He frames it as a software engineer who is more product-minded, and notes it’s already happening, since AI engineers spend most of their time on prompts and evals rather than boilerplate code. ## EP 45: Debunking AI's Environmental Panic | Andy Masley Guests: Andy Masley, Effective Altruism DC · Published: 2025-11-26 · Duration: 59 min URL: https://chainofthought.show/podcast/45-debunking-ais-environmental-panic-andy-masley/ (full transcript on page) AI is destroying the planet—or so we've been told. This week on Chain of Thought, we tackle one of the most persistent and misleading narratives in the AI conversation. Andy Masley, Director of Effective Altruism DC, joins host Conor Bronsdon to fact-check the absurd AI environmental claims you've heard at parties, in articles, and even in bestselling books. ### Key takeaways - Andy Masley found the data-center water claim in Karen Hao’s NYT bestseller *Empire of AI* is off by a factor of about 4,500. The book says one data center serving a community of 88,000 people in Santiago, Chile uses 1,000x that community’s water; the book’s numbers implied each resident consumed only ~0.1 liters a day — “a fifth of a single bottle of water.” As Masley put it: “They’d all be dead.” - The 4,500x error stemmed from a units mix-up — the source confused liters with cubic meters (1 m³ = 1,000 liters) — and had stood uncorrected for the six months the book had been out. Corrected, the data center uses roughly 25% of the water in that one section of Santiago and about 3% of the entire municipal system: a real footprint, but a fraction of what the book described. - Adding up everything — inference, cooling, amortized training, data transmission, and the embodied carbon of the chips — Masley estimates a median ChatGPT prompt causes about 1/150,000th of your daily emissions. To raise your daily emissions by 1% you’d need to send roughly 1,000 prompts, which would take seven to eight hours of reading replies — a day on which your emissions are likely *lower* than normal. - Activist pressure based on the inflated water numbers pushed Google to use air cooling instead of water cooling at the Santiago data center — and air cooling takes more energy than water cooling, so Masley suspects its CO₂ emissions ended up higher. His point: environmental policy is all trade-offs, and getting the numbers wrong leads to genuinely worse environmental outcomes. - Data centers are already “wildly efficient,” so pressuring AI companies to optimize energy use chases near-zero wins — they’re heavily incentivized to do that already. The real lever is *whether* a data center gets built and whether its grid is green. And per the IEA’s 2024 report, AI broadly (mostly non-chatbot deep learning) will likely prevent three to four times as many emissions as all data centers consume. - Misinformation primarily harms your own side. Masley — a vegan environmentalist, not a “tech bro” — argues that piling weak water claims onto valid concerns (jobs, surveillance, P(doom)) discredits the whole case: if someone debunks the water stat, your stronger arguments look invalidated too. Climate attention is scarce, so spend it on the grid, not on policing personal chatbot use. ### FAQ **How much energy does ChatGPT use per prompt?** Very little relative to daily life. Andy Masley estimates that once you add up inference, cooling, amortized training, data transmission, and the embodied carbon of the AI chips, a median ChatGPT prompt causes roughly 1/150,000th of your daily emissions. You’d need about 1,000 prompts to raise your daily emissions by 1%. For intuition, one prompt emits about as much as printing a fifth of a page of a book, running a space heater for half a second, or driving a car about four feet. **Does using ChatGPT waste a lot of water?** No. Andy Masley estimates a single ChatGPT prompt represents about 1/800,000th of your daily water use, once you account for the water embedded in your electricity, food, and household consumption — not just water at the data center. He compares cutting your chatbot use to save water to filling a pot to boil spaghetti and then saving a few drops with an eyedropper: a “sad distraction” from real levers like building out renewable energy. **What was the 4,500x error in Karen Hao’s *Empire of AI*?** The book claims a data center serving a community of 88,000 people near Santiago, Chile uses 1,000 times that community’s water. Andy Masley found this is off by about 4,500x — partly from a units mix-up (confusing liters with cubic meters, where 1 m³ = 1,000 liters). The cited numbers implied each resident used ~0.1 liters of water a day, about a fifth of a water bottle — physically impossible. Corrected, the data center uses roughly 25% of the water in that one section of the city and about 3% of the full municipal system. **Could AI actually be good for the climate?** Masley argues AI’s biggest climate impact will come from how it’s used, not from data centers themselves. The International Energy Agency’s 2024 report estimates AI broadly will likely prevent three to four times as many emissions as are used in all data centers — driven mostly by non-chatbot deep learning: optimizing building energy use, detecting irrigation leaks, and enabling smart-grid technology that stores and routes renewable power. He suspects deep learning may have already saved more water than data centers have used, though he flags that claim as more contentious. **Should we pressure AI companies to make data centers more energy-efficient?** Masley says that’s mostly the wrong target — data centers are already “wildly efficient,” and companies are heavily incentivized to optimize chips and energy because there’s enormous money in it, so the marginal wins are tiny. The real questions are whether a given data center should be built at all, its knock-on effects on the grid, and whether the energy it draws can be made greener. Efficiency pressure on operations chases near-zero gains; siting and grid decisions are where the leverage is. ## EP 44: The Critical Infrastructure Behind the AI Boom | Cisco CPO Jeetu Patel Guests: Jeetu Patel, Cisco · Published: 2025-11-19 · Duration: 1 hr 18 min URL: https://chainofthought.show/podcast/44-the-critical-infrastructure-behind-the-ai-boom-cisco-cpo-jeetu-patel/ (full transcript on page) AI is accelerating at a breakneck pace, but model quality isn’t the only constraint we face.. There are major infrastructure requirements, energy needs, security, and data pipelines to run AI at scale. This week on Chain of Thought, Cisco’s President and Chief Product Officer Jeetu Patel joins host Conor Bronsdon to reveal what it actually takes to build the critical foundation for the AI era. ### Key takeaways - Jeetu Patel frames AI-and-jobs as a maturity spectrum: level one is “AI is going to take our job,” level two is “someone that uses AI better than me is more likely to take my job than AI” (a real risk), and level three is realizing it will be hard to do your job without AI — just as you now can’t get a finance job without understanding the spreadsheet. - Patel argues aggregation — getting a succinct answer from the existing corpus of human knowledge — is “like 1% of the benefit” of AI. The other 99% is original insights that don’t exist in the human corpus, which he ties to tackling longevity, curing disease, ending poverty, and the climate crisis. - On model proliferation, Patel expects consumer AI to consolidate to “a dozen” foundation models holding most inference volume, while the enterprise side could run “tens of thousands of models” behind an intelligent routing layer that optimizes cost, latency, and domain fit. He cites Cursor, guessing — “I don’t know for a fact” — that roughly 70% of its queries no longer go to Anthropic. - Cisco’s security work shows small models can win: Patel says a distilled, open-source 8-billion-parameter model can now outperform a 70-billion-parameter one in that domain, and can be quantized to run on a CPU on a laptop — addressing the trust deficit and the infrastructure constraint at once. Cisco open-sourced its first security model as Foundation AI. - Patel positions Cisco as “the picks and shovels company during the gold rush” — the critical infrastructure for the AI era — and casts the network as the force multiplier: power is the constraint, the GPU is the core asset, and without the network the GPUs can’t operate. Cisco’s “scale across” silicon lets two data centers hundreds of kilometers apart act as one coherent cluster for a single training run. - On leadership, Patel says he hires for three things you mostly can’t teach — hunger, extreme curiosity, and clarity of thought — and deliberately mixes experienced and inexperienced people so pattern-recognition doesn’t calcify into a false sense of confidence. He warns that with success, the danger flips from complacency to arrogance: “preserve humility and stay paranoid.” ### FAQ **What are the three biggest constraints holding AI back according to Cisco?** Cisco’s Jeetu Patel names three: an infrastructure constraint (not enough power, compute, network bandwidth, or data center capacity to satisfy AI’s needs), a trust deficit (models are non-deterministic and therefore unpredictable, yet we want to build highly predictable applications on top of them), and a data gap (human-generated public internet data is plateauing, even as synthetic data proves efficacious in post-training and machine data from apps and agents explodes). His formula: “Infrastructure is the input. Intelligence is the output.” **Will enterprises use one big AI model or many?** Patel expects the opposite of consolidation on the enterprise side. He thinks consumer AI will settle on roughly a dozen foundation models holding most inference volume, but enterprises could run “tens of thousands of models” used in conjunction behind an intelligent routing layer — sending certain queries to certain models to optimize cost, latency, and domain efficacy. As an example he guesses — explicitly noting he doesn’t know for a fact — that about 70% of Cursor’s queries no longer go to Anthropic. **What is Cisco’s “scale across” networking and why does it matter for AI?** Patel describes a progression from scale up (GPUs in a rack) to scale out (rows of racks in a data center) to “scale across” — networking that links two data centers, potentially hundreds of kilometers apart, so they operate as one coherent cluster for a single training run. It matters because, in his illustrative ranges, a run may need on the order of 300,000 to a million GPUs while one data center hosts only 50,000–100,000 — and power availability now decides where data centers get built. Cisco builds silicon, systems, and optics with deep buffering, since a dropped network packet can force a multimillion-dollar training run to restart. **Can small AI models outperform large ones?** Yes — in Cisco’s security work, Patel says a distilled, open-source 8-billion-parameter model can now perform better than a 70-billion-parameter model, the key being training on the right data rather than infinite data. That smaller model can be further quantized to run on a CPU on a laptop, which he says solves both the trust deficit and the infrastructure constraint. Cisco open-sourced this security-model effort as Foundation AI. **How does Jeetu Patel think about partnerships and acquisitions?** Patel says a company’s willingness to partner reveals its arrogance level — no one can build everything for a shift this large, so he favors an open ecosystem over a walled garden, even partnering with competitors to “grow the pie.” His rule: if a player has more than 20% of a market and you refuse to partner, you only exclude yourself from that 20%. On acquisitions, he insists strategy comes first — the question is “what are we going to build,” not “what are we going to buy”; acquisitions are just a tool to accelerate a clear true north, not a strategy in themselves. ## EP 43: Beyond Transformers: How Liquid AI Is Rethinking LLM Architecture | Maxime Labonne Guests: Maxime Labonne, Liquid AI · Published: 2025-11-12 · Duration: 53 min URL: https://chainofthought.show/podcast/43-beyond-transformers-how-liquid-ai-is-rethinking-llm-architecture-maxime-labonne/ (full transcript on page) The transformer architecture has dominated AI since 2017, but it’s not the only approach to building LLMs - and new architectures are bringing LLMs to edge devices Maxime Labonne, Head of Post-Training at Liquid AI and creator of the 67,000+ star LLM Course, joins host Conor Bronsdon to challenge the AI architecture status quo. Liquid AI’s hybrid architecture, combining transformers with convolutional layers, delivers faster inference, lower latency, and dramatically smaller footprints without sacrificing capability. ### Key takeaways - Liquid AI’s LFM2 models are a hybrid architecture: each model carries six attention layers and ten short convolution layers — more convolution than attention. The convolution layers keep the KV cache from exploding on long context, so inference is much faster and memory usage much lower while matching a pure transformer’s quality. - Liquid open-sourced 17 models in roughly three months (since July). The LFM2 line ships at 350M, 700M, and 1.2B parameters, designed first for phones and wearables rather than scaled up — the opposite of the “how big can we make this thing” race. - Theoretical speed gains don’t survive contact with real hardware. After LFM1, Liquid started optimizing LFM2 directly on a target device (a Samsung phone) plus over 100 pre-training evaluations, so an operator that looks fast on paper is actually measured fast before it ships. - Knowledge scales hard with parameter count — Maxime calls it “a fact” that a 1B model squeezed full of distilled knowledge “will never be as smart as a 3B model.” The fix isn’t more cramming: give the small model a web-search or tool-use function so it can look up what it doesn’t know instead of hallucinating. - In post-training, data beats technique: “training techniques are nice, but they will never replace good data.” Maxime frames quality as accuracy plus dataset diversity — diversity matters so much you might keep wrong samples for coverage — and says there’s no shortcut but manually reading model responses and training samples. - For retrieval, Liquid shipped a ColBERT-style late-interaction model that precomputes vectors and re-ranks in the same model — best of embeddings and cross-encoders. On the fast LFM2 backbone it runs a bigger, better model at the speed of a much smaller one, which matters when a RAG answer can’t wait a minute and a half mid-call. ### FAQ **What makes Liquid AI’s LFM2 architecture different from a transformer?** It’s a hybrid, not a transformer replacement. As Maxime Labonne explains, each LFM2 model combines six attention layers with ten short convolution layers — more convolution than attention. The convolution layers stop the KV cache from exploding on long context, so the model gets faster inference and lower memory use while keeping a pure transformer’s output quality. Liquid optimized it directly on target hardware like a Samsung phone so the speed gains hold up in practice, not just on paper. **Can small on-device models handle agents, function calling, and RAG?** Partly. Maxime is skeptical that small models can run full agentic, multi-step reasoning workflows — those remain frontier capabilities — but function calling and tool use work, and he’s gotten good feedback even on the 350M model. The sweet spot is one well-defined task, where several small single-task models can fill roles in a multi-agent system. RAG is a strong fit because it doesn’t require heavy reasoning, and Liquid ships a RAG-specific model plus a late-interaction retrieval model. **Does synthetic data work for post-training, or does it hurt model quality?** Maxime says it works “all the time” — he estimates 99% of post-training datasets are synthetic, with humans mostly contributing prompts. The real failure mode is diversity collapse: reuse one generation process and you get data that looks varied but isn’t, which caps coverage. Techniques like persona sampling (telling the model to role-play, e.g., a truck driver from a given country) reinject diversity and, per Allen AI work he cites, can even improve math performance. For a domain like math you still want many different generation processes. **How does Liquid AI evaluate small models instead of using standard benchmarks?** Liquid runs its own internal benchmark stack so it can target narrow skills, like specific kinds of function calling. A key move: instead of scoring a 1B model on frontier evals where it’d get ~1%, they repurpose older frontier benchmarks and convert knowledge tests into tool-use tests. For something like OpenAI’s SimpleQA, they give the model a web-search function — so it measures whether the model can look up the answer rather than whether it memorized it. **What were the hardest lessons from shipping LFM1 and LFM2?** Two stand out for Maxime. First, custom architectures aren’t compatible with anything — you have to reimplement operators from scratch in every library (Hugging Face Transformers, vLLM, SGLang, llama.cpp), so for the third generation he hopes for more implementation-friendly operators. Second, function-calling data is exceptionally hard to generate at the quality and quantity needed, even harder than math. The fix was building in-house, model-specific training methods instead of copying papers, and doing the unglamorous groundwork of reading the data. ## EP 42: Architecting AI Agents: The Shift from Models to Systems | Aishwarya Srinivasan Guests: Aishwarya Srinivasan, Fireworks AI · Published: 2025-10-08 · Duration: 53 min URL: https://chainofthought.show/podcast/42-architecting-ai-agents-the-shift-from-models-to-systems-aishwarya-srinivasan/ (full transcript on page) Most AI agents are built backwards, starting with models instead of system architecture.Aishwarya Srinivasan, Head of AI Developer Relations at Fireworks AI, joins host Conor Bronsdon to explain the shift required to build reliable agents: stop treating them as model problems and start architecting them as complete software systems. Benchmarks alone won't save you. Aish breaks down the evolution from prompt engineering to context engineering, revealing how production agents demand careful orchestration of multiple models, memory systems, and tool calls. She shares battle-tested insights on evaluation-driven development, the rise of open source models like DeepSeek v3, and practical strategies for managing autonomy with human-in-the-loop systems. The conversation addresses critical production challenges, ranging from LLM-as-judge techniques to navigating compliance in regulated environments.Connect with Aishwarya Srinivasan:LinkedIn: https://www.linkedin.com/in/aishwarya-srinivasan/Instagram: https://www.instagram.com/the.datascience.gal/Connect with Chain of Thought host Conor Bronsdon: https://www.linkedin.com/in/conorbronsdon/00:00 Intro — Welcome to Chain of Thought00:22 Guest Intro — Ash Srinivasan of Fireworks AI02:37 The Challenge of Responsible AI05:44 The Hidden Risks of Reward Hacking07:22 From Prompt to Context Engineering10:14 Data Quality and Human Feedback14:43 Quantifying Trust and Observability20:27 Evaluation-Driven Development30:10 Open Source Models vs. Proprietary Systems34:56 Gaps in the Open-Source AI Stack38:45 When to Use Different Models45:36 Governance and Compliance in AI Systems50:11 The Future of AI Builders56:00 Closing Thoughts & Follow Ash Online ### Key takeaways - The biggest prototype-to-production gap is treating LLMs as tools versus building tools *with* LLMs inside them. As Aishwarya Srinivasan puts it, the “magic box” wrapper era of 2022–2023 is over — production agents are “way more of a software engineering challenge rather than just a model building challenge,” orchestrating small, mid, and large models, memory, and tool calls as one system. - “How much autonomy is the right autonomy” has no universal answer — it depends entirely on the user journey, the points of risk, and the level of compliance you’re building under. The job is identifying exactly where a human-in-the-loop or a deterministic checkpoint belongs, not picking a global setting. - Quantify every step. Srinivasan’s core reliability principle is to make each step of the system measurable and logged — for a coding copilot, that means separate quantification for writing code, reviewing it, writing unit tests, and merging to main. Without checkpoints and traces, when something breaks you can’t reverse-engineer what broke, why, or how to stop it recurring. - DeepSeek V3 was the inflection point for open source. After it (followed by R1 and V3.1), the performance-to-price ratio made open models viable for teams that need control over training, data usage, zero-data-retention, on-prem deployment, and model size/speed. The top agentic open models Srinivasan sees are Kimi, Qwen, and DeepSeek — even OpenAI (GPT-OSS), Meta, and Google are now re-investing in open models. - Production agents are Compound AI systems, not single models — combining multimodal, small-and-large, and specialized models end-to-end. Notion fine-tuned on Fireworks AI specifically because the latency *and* response quality their AI needed weren’t available from proprietary models, then ran a combination of models rather than one. - For regulated environments, encode compliance into architecture via multi-agent role-based access. A primary agent that talks to the user is deliberately denied sensitive datasets; it calls a secondary agent that holds that access, checks the request against criteria, returns only matching data — and the primary agent never stores it in short- or long-term memory. ### FAQ **What’s the difference between building a wrapper around an LLM and building a real AI agent?** Aishwarya Srinivasan frames it as “LLMs as tools versus LLMs as part of tools.” The 2022–2023 wave of apps were thin wrappers — mostly LLM calls treating the model as a magic box that answers anything. A real agent is a software system: a combination of small, mid-sized, and large models — some fine-tuned, distilled, proprietary, or open source — wired together with memory systems, tool calls, and the logs and traces flowing between every model-to-model conversation. It’s a software engineering problem, not just a model-building one. **When should I add evaluations and observability to an AI system?** It depends on how soon you’re going to production, says Srinivasan. If production is near, evals should be a day-one design decision — how you plan to build them into the system from the start. If you’re building a throwaway prototype to pitch internally and see how it goes, you can defer. The deciding factors are how quickly the system moves “from your laptop to a cloud-hosted environment” and the risk profile — if PII could surface, you likely need guardrails and live evals regardless. **Is using an LLM as a judge a reliable way to evaluate AI outputs?** It has real risks — Srinivasan notes that judging one model’s output with another model of the same nature is like “taking another blind model to go and check how the previous model is doing.” It works best for subjective, hard-to-quantify tasks like a writing assistant, where outputs legitimately differ per user and there’s no single right answer. When the output range is well-defined and stateful, traditional evals are better. You can also fine-tune the judge to enforce boundaries for a specific use case. Even imperfect, having something beats having nothing. **Why would I use open source models instead of proprietary ones for building agents?** Per Srinivasan, teams move to open models when they need control proprietary APIs don’t offer: customizing how the model is trained, controlling how their data is used, zero-data-retention, on-prem deployment, and tuning to a specific size and speed. Most teams don’t run open models out of the box — they apply supervised or reinforcement fine-tuning, or synthetic data. Proprietary models still win on developer experience, ease of setup, and capabilities like vision, so the real pattern is a mix, choosing the right model per use case. **Why does the same LLM sometimes give different answers to the identical prompt?** Srinivasan points to a Thinking Machines Lab blog on the nondeterministic nature of LLMs. People assumed models vary only when you change the question — but the same query can produce different outputs from the inference engine too. The cause is largely “how the plumbing looked like around the large language model” rather than the model itself. It’s a still-being-uncovered behavior most practitioners don’t yet understand, and a reason architecting these systems requires both objective and subjective evaluation. ## EP 41: The Accidental Algorithm | Humans of AI Crossover with Writer's Melisa Russak Guests: Melisa Russak, Writer · Published: 2025-10-01 · Duration: 21 min URL: https://chainofthought.show/podcast/41-the-accidental-algorithm-humans-of-ai-crossover-with-writers-melisa-russak/ (full transcript on page) This week, we're doing something special and sharing an episode from another podcast we love: The Humans of AI by our friends at Writer. We're huge fans of their work, and you might remember Writer's CEO, May Habib, from the inaugural episode of our own show. From The Humans of AI : Learn how Melisa Russak, lead research scientist at WRITER, stumbled upon fundamental machine learning algorithms, completely unaware of existing research — twice. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Melisa Russak built what turned out to be K-means clustering without knowing the algorithm existed. As a high-school math teacher in China, she wrote a classification system to recognize her students’ handwritten Chinese characters — only later discovering, in her words, that “people are working already more than one hundred years on this problem.” - Russak says she chose to study Chinese alongside mathematics precisely because it was “maximally different” — “a completely different system” with no alphabet that pushed her “completely out of your comfort zone.” She wanted learning she could see compound day by day, unlike pure math where “you can spend ten hours on thinking about a math problem and have no results.” - Russak rediscovered embeddings the same accidental way. Trying to build software that tracked her internet activity to answer “who am I?,” she had to represent both text and images so a computer could compare them. Her team “spent a lot of time trying to design those features ourselves” — before, as she puts it years later, “it’s all about embeddings… you just need to learn good embeddings.” - Her advice to researchers: frame a problem yourself before reading the papers. “Once you start reading papers, you will converge to… how they framed this problem,” Russak warns, “and it’s very difficult to escape from that box once you are in the box.” Her accidental discoveries happened because she didn’t yet know the “right” way to think. - At Writer, Russak frames the no-customer-data policy as a feature, not a wall. “We don’t use customer data. So how do I develop a system without data?” The team’s answer was to train a model to generate the data itself: “I would never come up with this if not given those constraints.” - Russak puts the “vibes test” above benchmarks. A benchmark, she argues, “is like a cherry picked use case”; the missing piece is going to a human and asking whether they actually like using the model in production. “We always say that the VIBES test is the most important. After you satisfy all of those benchmarks, you go and do the VIBES check.” ### FAQ **Who is the guest, and is this a Chain of Thought interview?** The guest is Melisa Russak, a lead research scientist at Writer. This is a crossover episode — Chain of Thought is re-sharing an episode of Writer’s own podcast, Humans of AI. Conor Bronsdon records only a short intro and outro framing the swap; he does not conduct the interview. The Humans of AI host, Elora Weaver, narrates Russak’s story, so the analytical and connective passages between Russak’s own words are Weaver’s, not the guest’s. **How did Melisa Russak “accidentally” invent machine learning algorithms?** Russak describes doing it twice, without knowing the field existed. First, as a math teacher in China, she built a system to classify her students’ handwritten Chinese characters — which turned out to be K-means clustering. Later, building software to track her own internet activity and learn something about herself, she had to represent and cluster text and images, reinventing the approach now known as embeddings. In her telling, she only discovered afterward that researchers had long been working on the same problems. **What does Russak mean by the “vibes test”?** Russak argues that benchmarks are essentially cherry-picked use cases, so passing them isn’t enough. The missing piece, she says, is time-consuming and human: you give the model to a real person, have them use it in production, and ask whether they actually like talking to it. “We always say that the VIBES test is the most important,” she says — after a model satisfies the benchmarks, the team does the vibes check. **Why does Russak see Writer’s no-customer-data rule as an advantage?** Russak says Writer made a deliberate choice not to use customer data, which at first felt like an impossible constraint: “How do I develop a system without data?” Her team’s answer was to train a model to generate the data — “I would never come up with this if not given those constraints.” She treats the constraint as an “excellent constraint” rather than a limitation: something that forced an idea she would not have reached otherwise. **What is Russak’s warning about building on third-party model APIs?** Russak cautions that a startup relying on an external API is exposed if the provider changes the underlying model: “Your entire business is in pieces,” she says, “because everything that you created so far on top of it, it stops existing.” Writer’s counter-approach, as she describes it, is to understand its own models deeply through extensive testing — including how each one behaves under quantization and distillation — so the team knows how to design a system on top of the model rather than optimizing for benchmark scores alone. ## EP 40: After Code Gen: What Graphite Is Building for the Post-AI Dev Stack | Greg Foster Guests: Greg Foster, Graphite · Published: 2025-09-24 · Duration: 55 min URL: https://chainofthought.show/podcast/40-after-code-gen-what-graphite-is-building-for-the-post-ai-dev-stack-greg-foster/ (full transcript on page) The incredible velocity of AI coding tools has shifted the critical bottleneck in software development from code generation to code reviews. Greg Foster, Co-Founder & CTO of Graphite, joins the conversation to explore this new reality, outlining the three waves of AI that are leading to autonomous agents spawning pull requests in the background. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Greg Foster frames AI’s impact on engineering as three waves: tab-complete (GitHub Copilot, then Cursor extending tab to fill multiple lines ahead); an IDE or Claude Code sidebar “director mode” where you send the model on small missions; and the most emergent — headless background agents you ping on Slack or GitHub, then get a pull request back hours later. - As AI automates the inner loop of writing code, Greg argues the human outer loop is now the bottleneck — review, CI, merge queues, deployment, rollout monitoring. He points to tweets from the Anthropic team and others making the same case: it’s no longer code generation that’s the issue, but the DevOps and shipping practices around production. - The strain shows up in numbers. Greg cites Facebook being quoted as doubling internal code changes per engineer over the past year, with another doubling expected soon, and says Graphite sees the same — code-change volume per engineer is “strictly up into the right” — pressuring both reviewers and merge systems. - Review is now higher-stakes and lower-trust. Greg used to trust a teammate wrote a change; now the human might have “held tab down while staring at their iPhone,” or no human may have read the code before he does. With machines being gullible and security incidents rising, he says code review “really starts mattering.” - Greg makes the case for stacking — small, interdependent diffs built one on top of another, a technique he attributes to Facebook. He hedges that he thinks Google had a study, echoed in Graphite’s own data, finding longer changes get fewer comments per line because reviewers glaze over: a 10-line PR draws real feedback, a giant one gets an LGTM. He notes AI itself “really likes” stacking. - On the hiring gap, Greg offers two predictions. Pessimistic: engineering may regress toward finance, consulting, and law, where entry-level value is lower and grads grind for fewer slots. Optimistic: for anyone with high initiative, “the getting has never been so good” — versus high school, when he fought through Xcode on a family Intel machine to make iOS apps. ### FAQ **What are the three waves of AI coding that Greg Foster describes?** Greg Foster lays out three patterns that now coexist. First is tab-complete: GitHub Copilot cracked it open by predicting the rest of a line, and Cursor extended it to fill multiple lines ahead. Second, enabled by stronger models like Sonnet 3.5 and 3.7, is a sidebar or “director mode” — in the IDE or in Claude Code — where you send the agent on small missions like adding a button or a unit test. Third is headless background agents: you ping them on Slack or GitHub and a PR comes back hours later, like a self-driving car that does the delivery. He estimates that mode handles maybe 10% of PRs today. **Why does Greg Foster say code review has become the bottleneck in software development?** Greg Foster argues AI has obsessed over the inner loop — code generation on your machine — while the outer loop now buckles under the volume. He walks through everything after you open a PR: a human review, passing CI (linters, unit and end-to-end tests), a merge queue that re-tests in case someone broke something, rebuilding an artifact, rolling out to production, and monitoring in Datadog or Grafana. He cites Facebook being quoted as doubling code changes per engineer with another doubling soon, and Graphite’s data showing volume “strictly up into the right” — which is why the outer loop, not code creation, is the bottleneck. **What is stacking, and why does Greg Foster think it suits AI agents?** Greg Foster describes stacking as creating a small code change, then another on top of it, and another — many interdependent diffs reviewed and tested one at a time, rather than one large branch. He attributes the technique to Facebook and traces its roots to old-school Git, where people worked commit by commit. He hedges that he thinks Google had a study, echoed in Graphite’s own data: longer changes get fewer comments per line because reviewers glaze over. Greg says AI “really likes” stacking — agents apply chain-of-thought, building the function, then the test, then the endpoint as separate diffs, and build it better than one-shotting a large change. **Why does Greg Foster say GitHub faces a hard challenge adapting to AI agents?** Greg Foster says he’s deeply empathetic to GitHub because it serves two very different communities. Open source has the longest tail of workflows — strangers asking maintainers to pull in long-lived feature branches, with less trust and slower cycles. Closed-source companies want monorepos, trunk-based development, and merge queues, with high-trust teams who say “stamp this” rather than “please, stranger, pull in my changes.” Greg notes the AI outer-loop patterns — stacking, merge queues, AI code review — fit closed source far better, so leaning in risks offending longtime open-source users: an innovator’s dilemma like Google’s with search. **How does Graphite actually use LLMs in code review, according to Greg Foster?** Greg Foster describes Graphite as a client on top of GitHub — like an advanced email client over Gmail — that helps you process, test, and merge a pull request. He’s candid they don’t do “too much fancy stuff”: rather than build their own models, they feed the diff through smart models from Anthropic, Gemini, or OpenAI (sometimes a mix), layered with users’ custom rules, style guides, and prior comments. He frames AI code review as “a fuzzy version of CI.” Graphite optimizes for actionable inline comments and stays quiet when confidence is low, because Greg says trust matters enormously — a wrong suggestion annoys him fast. ## EP 39: Vercel's Playbook for AI Agents: From Vibe Check to Production | Malte Ubl Guests: Malte Ubl, Vercel · Published: 2025-09-10 · Duration: 54 min URL: https://chainofthought.show/podcast/39-vercels-playbook-for-ai-agents-from-vibe-check-to-production-malte-ubl/ (full transcript on page) What’s the first step to building an enterprise-grade AI tool? Malte Ubl, CTO of Vercel, joins us this week to share Vercel’s playbook for agents, explaining how agents are a new type of software for solving flexible tasks. He shares how Vercel's developer-first ecosystem, including tools like the AI SDK and AI Gateway, is designed to help teams move from a quick proof-of-concept to a trusted, production-ready application. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Malte Ubl reframes agents not as something weird but as “software that we always wanted to write” — the daily-task automation that needs a little flexibility and was historically hard to model with exhaustive if-this-then-that blocks. Agents do really well on exactly that kind of software, so suddenly it’s super easy to write. - Vercel built what Malte reluctantly calls a “DevSecOps agent” — it takes anomaly detection from a firewall and goes fishing for what happened. Hand a modern frontier model a few tools to query the data stream and it does a pretty good job from scratch; he says you spend an afternoon and have software that used to be incredibly hard to write. - On why coding agents lead, Malte is precise: software engineering has source code you can validate and run unit tests on, so you can reinforcement-learn on this valuable, open-ended problem in a way that’s harder elsewhere. He treats that as a signal the effectiveness most likely generalizes — not proof it already has. - Malte’s recommended architecture is the uncomfortable-but-it-works simple one — the reason Claude Code feels magical. An LLM, a trigger, input and prompt, and a set of relatively simple, atomic tools (read files, list files, edit files, create PRs). Tell it to make up to ten turns toward the goal, then “just let it cook.” - Prompt injection demands a different fix than SQL injection, Malte argues. Tool responses become part of the prompt, period — nothing jails them, so you can’t escape them the way you escape SQL. The fix moves security to a different layer: hard-code a query’s tenant conditions instead of passing a user ID the model can be tricked into changing. - The under-discussed pattern, in Malte’s view, is “AI in the loop” — the inverse of human in the loop. The human does the work and AI reviews it. He’s a fan because AI code review is relentless, doesn’t forget your past mistakes, doesn’t go to bed or wake up tired, and stays sharp on tedious checks where humans go low on quality the fifteenth time. ### FAQ **What does Malte Ubl mean by the “vibe check” when building an agent?** Malte Ubl describes the vibe check as a non-programming exercise to find out whether the LLM is anywhere up to the task — before you write any code. His example: if you have a set of writing rules, go to ChatGPT, paste your text, and ask it to check the text against those rules. If it delivers all kinds of false positives, that sucks; if it doesn’t find the problems, that sucks too — but maybe it does. You quickly learn whether you’re giving the AI a task it isn’t ready for, in which case you might need to wait a year or do something else. If it seems to work, you go turn it into a program. **How does Malte Ubl recommend structuring a production agent?** Malte Ubl encourages the simplest architecture — the one behind why Claude Code works as magically as it does. You have an LLM, a trigger to start work, input data and a prompt, and a set of relatively simple, atomic tools — for a coding agent, read files, list files, edit files, create pull requests; for deep research, maybe a search over your Glean or a generic SQL tool hitting Snowflake. You describe what each tool is for and the goal, tell it it gets up to ten turns, and let it cook. He stresses tools should be atomic rather than high-level — uncomfortable, but it just works. **Why does Malte Ubl say prompt injection can’t be solved like SQL injection?** Malte Ubl explains that with SQL injection, escaping user inputs per best practices gets you to essentially zero risk — in principle it’s solved. Prompt injection is different: a tool’s response becomes part of the prompt, with nothing jailing it. You can sanitize inputs, but an attacker can phrase the injection in Spanish, so you can’t “100% ensure” safety the way you guarantee escaped SQL. His fix moves security to a different layer: hard-coding a query’s tenant conditions inside the tool rather than passing a user ID the model can be tricked into changing. **What does the Vercel AI Gateway do, according to Malte Ubl?** Malte Ubl says the AI Gateway is directly integrated with the AI SDK and lets you use every AI model from any provider without any API keys. To give a new model a vibe check, you literally just change the model string — no going to a provider’s website to get a key, which he notes can be especially hard for some models. During development the rate limits are very low and it’s free, on the logic that you’d never need much yourself; when you go to production you can access the same models and Vercel bills you at market rate. He frames it as a frictionless way to actually try AI. **How does Malte Ubl suggest teams without AI experience get started?** Malte Ubl’s main advice is to try it in a non-pressure setting, because this is a new kind of software most people don’t have intuition for and some feel anxiety about. Vercel ran a one-week agent hackathon for all engineers; beyond the “pretty awesome” outcomes, the bigger win was that everyone had now built one agent, so they’d approach real business tasks having done it before. He also points to AI SDK examples — including his colleague Nico’s well-received session building a coaching agent from scratch with a not-smartest model and three tools — to show the technique transfers to other use cases. ## EP 38: From Demo to Defensibility: How to Build an AI Business that Lasts | Aurimas Griciūnas Guests: Aurimas Griciūnas, SwirlAI · Published: 2025-08-27 · Duration: 52 min URL: https://chainofthought.show/podcast/38-from-demo-to-defensibility-how-to-build-an-ai-business-that-lasts-aurimas-grici-nas/ (full transcript on page) The technological moat is eroding in the AI era, what new factors separate a successful startup from the rest? Aurimas Griciūnas, CEO of SwirlAI, joins the show to break down the realities of building in this new landscape. Startup success now hinges on speed, strong financial backing, or immediate distribution. Aurimas warns against the critical mistake of prioritizing shiny tools over fundamental engineering and the market gaps this creates. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Aurimas Griciūnas argues there are only three ways to easily start a successful startup today: build something that grabs attention fast and reaches escape velocity within the first months, have really strong backing from the start — enough money to build a big team and roll out enterprise operations properly — or have distribution on day one. - Open source is hard to turn into a profitable business, and enterprises don’t work well with it — they need a very mature solution. That’s also why infrastructure companies are hard: you need real engineering, not vibe coding, because security, stability, on-prem deployment, support, and enterprise features all have to be top notch. - The technology moat is genuinely less defensible now, Aurimas says. With a strong engineering team — maybe five strong engineers using AI — you can very quickly build an observability and eval tool that rivals Langsmith or Langfuse. So no one pays because you’re a known person; they pay because the product is better and more efficient than competitors. - When advising founders, Aurimas would weight the founder over the product: the first idea is usually not a great idea, so he’d probe how good their pivoting ability is. He calls distribution and reach key — how you sell and market the product is probably even more important than the product itself, at least at the very beginning. - For agent builders, context engineering equals prompt engineering — you can’t build agentic systems without it. Aurimas warns the context window explodes to a few hundred thousand tokens per run; five runs at 200,000 input tokens each can take fifty seconds and “that’s a chatbot.” His favorite lever: compress the conversation history and offload discarded actions to a scratch pad. - Aurimas predicts a slowdown — no big leaps in AI in the next six months, and no distributed multi-agent systems in production yet despite the A2A hype, because it’s too hard to instrument long-running distributed agentic systems. He doesn’t buy that LLMs alone reach AGI either, suspecting a model-architecture problem more than a hardware one, while still expecting we’ll need this kind of compute for inference. What excites him: coding CLI agents and self-improving agents that rewrite their own code, evolutionary-algorithm style. ### FAQ **Why does Aurimas Griciūnas decide against building another observability and evals company?** When Aurimas Griciūnas left Neptune AI, building something similar in the observability and evals space was the natural next move, so he researched it. Within the first few weeks he found 20-plus companies already doing it — plus the hyperscalers, and probably another 20 in stealth, half of them open source and free to self-host. Because it’s hard to pinpoint what will matter in the next few months, all of them try to cover end to end: traces, evals, experiments, prompt registries, even routers. He concluded there was no unique space left, so he called the market too packed and didn’t build. **What does Aurimas Griciūnas think about vibe coding your way to a startup?** Aurimas Griciūnas thinks vibe-coded tools can mostly succeed in B2C, where a new idea quickly captures the broad public’s attention. For enterprise products, he still expects VC-backed companies with large cash reserves to win — unless you build something really great really fast, raise a lot of money, then hire hundreds of engineers to refactor the vibe-coded foundation. He adds a key caveat: if a really strong engineer is doing it — and he notes they’re usually doing assisted coding, not pure vibe coding — that approach can work, because the engineering talent is already inside the founding team. **What fundamental gaps does Aurimas Griciūnas see in how teams build AI today?** Aurimas Griciūnas points to several. Evals are still a gap — eval-driven development is crucial but not widely adopted. Overreliance on orchestrators causes problems as systems mature, forcing teams back to base software engineering without wrappers. People take too long to ship: first MVPs aren’t rolled out soon enough and human feedback isn’t fed back in fast enough. He also sees weak business understanding — teams build something shiny instead of solving a business problem, then build agentic systems “in the basement” only to ship something that solves nothing. **Why does Aurimas Griciūnas call evals the hardest part of building agentic systems?** Aurimas Griciūnas says observability tooling in general is great but isn’t the hardest problem. The hardest problem when building agentic systems is creating the eval datasets themselves, which he calls really, really hard — sometimes 70% of the entire project goes into figuring them out. He frames this as the piece of the puzzle that’s currently missing: too many hours go into eval datasets, and that, not the observability layer, is where the real difficulty lives. **What does Aurimas Griciūnas say about data engineering in the AI era?** Aurimas Griciūnas — a data engineer for four or five years who also led data engineering teams — says it’s consistently underrepresented even though data engineers do most of the work to make these systems run. He pushes back on the idea that AI engineering is just data engineering: data engineering is about piping data to where it needs to live, while AI engineering is about building agentic system designs on top of that data. He’s candid that he doesn’t know how to keep it in the spotlight — it’ll never be hot, in his words, because data engineering saves costs rather than producing revenue. ## EP 37: Mindset Over Metrics: How to Approach AI Engineering | Hamel Husain Guests: Hamel Husain, Parlance Labs · Published: 2025-08-20 · Duration: 42 min URL: https://chainofthought.show/podcast/37-mindset-over-metrics-how-to-approach-ai-engineering-hamel-husain/ (full transcript on page) As we enter the era of the AI engineer, the biggest challenge isn't technical - it's a shift in mindset. Hamel Husain, a leading AI consultant and luminary in the eval space, joins the podcast to explore the skills and processes needed to build reliable AI. Hamel explains why many teams relying on vanity dashboards and a "buffet of metrics" experience a false sense of security, which is no substitute for customized evals tailored to domain-specific risks. The solution? Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - The first question Hamel Husain gets is “What are the tools?” — and he says that’s the wrong question. No tool abstracts away whether your AI is doing the right thing for the user; until something like AGI arrives, that’s not possible. The right question is “What’s the right process?” because every tool still forces you through one. - Generic dashboard metrics — hallucination score, toxicity score, conciseness score — are, in Hamel’s words, “most of the time not helpful at all.” They’re too generic, don’t necessarily correlate with actual failures in your product, and create an illusion you’ve checked the evals box while you’re really monitoring nothing. - Hamel grounds evals in failures through error analysis: open a trace viewer, read traces, take notes on what’s going wrong, then categorize and count those error types to decide what to prioritize. He notes the technique predates machine learning — it comes from the social sciences — and most people skip it only because no one taught them to do it. - Borrowing the social-science idea of “theoretical saturation” — keep looking until you stop learning anything new — Hamel still gives a concrete starter heuristic: aim for at least 100 traces. People freeze on the abstract version; 100 is a goal they can act on, and once they begin they stop caring about the number because they’re learning so much. - Hamel’s “spicy hot take”: an AI engineer doesn’t need to train models, but does need real data literacy. Teaching LLM-as-a-judge means validating the judge against human labels, and explaining why sampling works (e.g., bootstrap sampling for judge noise) drags you back to classic statistics — the skill set of a machine learning engineer or data scientist. - The biggest failure mode Hamel sees — and a major driver of his consulting business — is outsourcing evaluations to developers because teams treat AI like a software-engineering task. Unless you’re building a developer tool, the developer isn’t the domain expert, so you guess. If you’re building for lawyers, involve the lawyer. ### FAQ **Why does Hamel Husain say buying an evals tool won’t fix AI reliability?** Hamel Husain says the instinct to “abstract this entire thing away to some tool” and make accuracy “not my problem” doesn’t work — until something like AGI, no tool can just figure out whether your AI is doing the right thing for the user. The number-one question he gets is “What are the tools?”, and he calls that the wrong question. No matter which tool you pick, you still have to go through the same process to evaluate AI correctly, so the real question is “What’s the right process?” **What does Hamel mean by “look at your data,” and how many traces should you review?** For Hamel Husain, “look at your data” is shorthand for error analysis: open a trace viewer, read through traces, write notes on what’s going wrong, then categorize and count those error types to decide what to prioritize. He cites the social-science concept of “theoretical saturation” — keep looking until you’re not learning anything new — but gives a concrete starting heuristic of at least 100 traces, because the abstract version makes people anxious and they never begin. Once they start, he says, they stop caring about the 100 because of how much they’re learning. **Does Hamel think AI engineers need a data-science background?** Hamel Husain tried to see how far he could teach engineers evals without a data-science background and says you hit a limit fast. In the eval course he co-teaches — over 700 students of all backgrounds so far — topics like LLM-as-a-judge require validating the judge against human labels, and questions like why sampling works push you back to classic statistics, e.g. bootstrap sampling to measure judge noise. His take: you don’t need to train models, but you need real data literacy and exploratory data skills, which lands you near the skill set of an ML engineer or data scientist. **Why does Hamel recommend building custom data-annotation apps instead of using a generic dashboard?** Hamel Husain argues that real applications have domain-specific context — rendered widgets, emails being written, external data sources, traces that should be viewed exactly as the user sees them, or token-heavy content you want to hide. To do fast error analysis you want all the data you need in one place, rendered the right way. Because AI is now good at vibe-coding simple data-rendering apps, the value of building your own annotation tool outweighs the cost a lot of the time — though, he’s careful to add, not 100% of the time. He still wants a trace viewer like Galileo’s as a supplement. **How does Hamel say teams should involve domain experts in evals and prompting?** Hamel Husain calls outsourcing evaluations to developers one of the biggest failure modes he sees — fine for a developer tool, but otherwise developers lack the context, so you’re just guessing. If you’re building for lawyers, involve the legal expert. He warns that many people treat “prompt” as an abstract concept and expect a developer to write it — “the worst thing that can possibly happen.” His fix: give the domain expert an admin view to edit the prompt directly inside the real, user-facing application — rather than a playground that can’t call your tools, do RAG, or run your actual code. ## EP 36: How AI Velocity is Rewriting the Rules for Engineering Leaders | ChatPRD's Claire Vo Guests: Claire Vo, ChatPRD · Published: 2025-08-13 · Duration: 43 min URL: https://chainofthought.show/podcast/36-how-ai-velocity-is-rewriting-the-rules-for-engineering-leaders-chatprds-claire-vo/ (full transcript on page) What if your next competitor is not a startup, but a solo builder on a side project shipping features faster than your entire team? For Claire Vo, that's not a hypothetical. As the founder of ChatPRD, formerly the Chief Product and Technology Officer at LaunchDarkly, and host of the How I AI podcast, she has a unique vantage point on the driving forces behind a new blueprint for success. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Claire Vo names velocity as the differentiator that keeps her up at night: an incumbent’s real threat is a small, AI-native, well-funded team with a clean codebase ripping through features. Her fear, in her words, is “a Claire out there” rebuilding her own product space faster than her team can react. - The single change that most moved engineering adoption at LaunchDarkly, Vo says, was appointing one senior, long-tenured engineer as the AI “czar” — and she deliberately picked someone who wasn’t naturally bullish on AI. Credibility plus deep knowledge of the core monolith mattered more than enthusiasm. - Vo reframes legacy codebases as an AI opportunity, not a blocker. She says she’s ripped out a JavaScript front-end framework — an 18-to-24-month slog — almost every job of her career, and AI makes that replatforming and tech-debt cleanup meaningfully faster. The same purpose-built work that makes a repo easier for an agent makes it easier for an intern or senior engineer too. - On adoption math, Vo’s experiment at LaunchDarkly was simple — go try AI tools and report what works. “For every one total dud, we got two helpful wins,” she says, arguing leaders who refuse to ship because AI might introduce a bug ignore that humans ship bugs too, and slowly. - Vo predicts the “era of the super IC”: senior individual contributors who command high salaries and outsized impact without managing people — no one-on-ones, no performance reviews. She tells leaders to build a path to pay and promote them without forcing them onto teams. - Vo expects the product-manager role to split into two archetypes — a prototype-building UX-engineer PM and a commercially minded GM-style PM — collapsing the “keeper of what users want” middle. She references her “PM is dead” talk from Lenny’s conference the prior year. ### FAQ **Why does Claire Vo think incumbents are at risk from small AI-native teams?** Vo says velocity will be a “massive, massive differentiator” and that incumbents will be competing feature-for-feature and capability-for-capability against small teams that pair AI-native speed with access to large funding rounds. Her recurring fear, drawn from solo-building ChatPRD, is “a Claire out there” ripping through her product space. She names price disruption, perceived innovation velocity (optics matter even if a rival isn’t yet at scale), and talent attraction as the operational disruptions these teams can cause. **How did Claire Vo choose the AI “czar” for LaunchDarkly’s engineering org?** Vo picked a senior, long-tenured engineer (she calls him Zach) for three reasons: he had a robust grasp of the architecture and core monolith and knew “where there are dragons”; he had enough internal tenure and credibility that the team would trust his verdict on what works; and she wanted to give him a career win — a next wave of impact and “AI all over your resume.” Notably, she chose someone who wasn’t pre-inclined to be bullish on AI, which made his assessments more trusted. **What does Claire Vo say about using AI on legacy or messy codebases?** Vo argues people underestimate AI on legacy code. She says it accelerates cleaning up the gnarliest tech debt — she’s spent 18 to 24 months ripping out deprecated JavaScript frameworks almost every job of her career, and AI makes that faster. She also recommends purpose-built work to make a repo better for AI, which makes it better for humans too: if an agent can’t run the codebase locally, neither can an intern or senior engineer. Her advice to skeptics who say “it’ll never work in our disgusting old repo” is that it sounds like a “you problem” — and to actually give it a real go. **How did Claire Vo build a culture of AI experimentation at LaunchDarkly?** Vo describes several tactics: getting finance and security aligned with a simple framework for evaluating and budgeting tools (she asked infosec to be “risk aware” but “cool”); a “building with AI” Slack channel with around 200 people where staff post wins and failures to normalize and socialize learning; and an “AI Friday power hour” at 10 in the morning where two or three people try something live with AI in the actual codebase. The throughline is building in public to remove shame and spread what works. **What is the “super IC” Claire Vo describes, and what does she predict for product managers?** Vo calls this “the era of the super IC” — senior individual contributors who, powered by AI tools and breadth of experience, can have outsized impact and command high salaries without managing people, no one-on-ones or performance reviews required. She urges leaders to create a path to pay and promote them without forcing them onto teams. On PMs, referencing her “PM is dead” talk at Lenny’s conference the prior year, she predicts the role splits into a prototype-building UX-engineer PM and a commercially minded GM-style PM, with ChatPRD’s agent absorbing tactical day-to-day work. ## EP 35: Building an AI-Native Startup | GrowthX's Marcel Santilli Guests: Marcel Santilli, GrowthX · Published: 2025-08-06 · Duration: 24 min URL: https://chainofthought.show/podcast/35-building-an-ai-native-startup-growthxs-marcel-santilli/ (full transcript on page) How do you build an AI-native company to a $7M run rate in just six months? According to Marcel Santilli, Founder and CEO of GrowthX, the secret isn't chasing the next frontier model, it's mastering the "messy middle." Drawing on his deep experience at Scale AI and Deepgram, Marcel joins host Conor Bronsdon to share his framework for building durable, customer-obsessed businesses. Marcel argues that the most critical skills for the AI era aren't technical but philosophical: first-principles thinking and the art of delegation. ### Key takeaways - Marcel Santilli names two skills as the foundation of an AI-native business: first-principles thinking — decomposing a problem to its essence and rebuilding it differently — and delegation. His claim: most people delegate poorly to humans because they share context badly and don’t organize their thinking in writing, and that weakness carries straight over to AI. - GrowthX started with services, Marcel explains, because services are a forcing function: a customer pays for the work output, not the tool. Selling the output forced the team to learn what he calls the “messy middle” — the research, mental models, planning, execution, and iteration that turn expertise into a finished work product. - Marcel’s Michelin-star analogy frames why captured outputs aren’t enough: hand someone photos of the finished dish and the full ingredient list and they still can’t reproduce it. Greatness lives in the trial-and-error of the messy middle, he argues, and there’s no single path to it — which is why knowledge work resists being frozen into a UI. - Rather than replace or mandate humans, Marcel describes finding where expert intervention is genuinely needed. GrowthX built a coding agent that writes workflows in code, a runtime that executes them, and an orchestrator that decides where a human steps in and what form that step takes — an edit, a comment, an approve/reject, a multiple-choice, or a ranking. - Marcel reports GrowthX ran over half a million workflow runs in the prior month, with each run capable of representing roughly ten hours of human work. His framing flips the goal: instead of a person reading 100 articles on a subject knowing 89 might be irrelevant, the system ingests the information and surfaces the points where expert judgment actually adds value. - Marcel argues coding agents improved because the messy middle is public — open-source repos expose every pull request, commit, and doc. He contrasts that with internal company knowledge: he dares you to find a sales department doc that describes its process as cleanly as an open-source project’s. Closing that gap, he says, is the opportunity in knowledge work. ### FAQ **What does Marcel Santilli mean by an “AI-native” business?** For Marcel Santilli, AI-native isn’t about chasing the latest frontier model — it’s rethinking every function of the company around two capabilities. The first is first-principles thinking: breaking a problem down to its core and rebuilding it in a different way. The second is delegation. He argues that if you’re bad at delegating to another human — sharing context, communicating your thinking, organizing it in writing — you’ll be even worse at delegating to AI. The infrastructure and stitching matter, but those two skills are the real foundation. **Why did GrowthX start with services instead of shipping a product first?** Marcel Santilli describes services as a forcing function. When a customer pays you for a work output — an article, a landing page, copy, research on a topic — rather than for a tool, you’re forced to learn what it actually takes to deliver high-quality work informed by good strategy. That’s how GrowthX mapped the “messy middle.” Marcel also warns against the alternative: a product team designing a UI in a corner for six months, only to launch already behind a new paradigm. Better, he says, to sign up to deliver the thing, then figure out how to deliver it in a more AI-native way. **How does GrowthX decide where humans fit into its AI workflows?** Marcel Santilli says the question isn’t whether to use humans or AI — it’s where expert intervention is genuinely required to shape the work or pull better inputs from customers. GrowthX built an orchestrator layer over workflows running in code that determines where an intervention belongs and what kind it should be: an open-ended edit, a comment, an approve-or-reject, a multiple-choice, or a ranking. Over time, he notes, that mix shifts as the system learns which interventions matter most. **How does GrowthX use AI to improve the human reviewers themselves?** Marcel Santilli explains it runs both directions. AI can coach the human doing an intervention — for a technical article for a customer like Abnormal Security, a reviewer who knows some security but isn’t the ultimate expert gets reminders, guidelines, and things to watch for as they review. Running the reverse, a calibrated expert making enough A-versus-B judgments generates data, and Marcel says enough of that data lets GrowthX fine-tune a workflow into task-, domain-, and company-specific models, or a mixture-of-experts approach. **Why does Marcel Santilli believe coding agents got good before knowledge-work agents?** Marcel Santilli credits transparency. Coding agents improved, he argues, because the messy middle is fully in the open: open-source projects expose every pull request, every commit, and clean documentation. He contrasts that with the inside of companies — daring you to find a sales department doc describing its process as well as an open-source project’s, let alone the pull request behind it. Without that visible back-and-forth, knowledge work has been harder to automate — which he frames as exactly where the opportunity now lies. ## EP 34: AI's Trillion-Dollar Healthcare Bet | Corti's Andreas Cleve Guests: Andreas Cleve, Corti · Published: 2025-07-30 · Duration: 47 min URL: https://chainofthought.show/podcast/34-ais-trillion-dollar-healthcare-bet-cortis-andreas-cleve/ (full transcript on page) AI isn't just changing healthcare; it's providing the essential help needed to unlock a trillion-dollar opportunity for better care. Andreas Cleve, CEO & Co-founder of Corti, steps in to shed light on AI's immense, yet often misunderstood, transformative potential in this high-stakes environment. Andreas refutes the narrative of healthcare being slow adopters, emphasizing its high bar for trustworthy technology and its constant embrace of new tools. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Andreas Cleve rejects the “healthcare is bad at adopting technology” narrative as “boring” and “riddled with lack of ambition.” His counter: walk into any hospital and count the devices per room — clinicians adopt tons of technology, and fast, but hold a high bar. As he frames it, healthcare doesn’t necessarily need to change, “it needs help.” - Cleve estimates roughly 40% of salary dollars in healthcare go to work that relates to language — talking to patients, writing notes, EHRs, letters, quotes, invoicing — which is exactly where transformer-based models are strong. That language load, not surgical hardware, is where he sees AI getting picked up first. - Cleve reframes healthcare as “not just one monolith trillion dollar admin market” but “a thousand billion dollar opportunities.” He argues the winners go deep in a single specialty’s workflow and grow “like weeds” bottom-up via PLG — naming Abridge, Rad AI, Tandem, Heidi Health, and Nelly — because doctors are buyers and are more tech-forward than outsiders assume. - On the workforce gap, Cleve cites the WHO’s upgraded figure: at least 10 million healthcare professionals lacking globally by 2030 — and he’s careful to scope it to delivering “the care we deliver today,” not the more preventative, longevity-focused care he expects by then. He pairs it with a McKinsey/Accenture estimate that AI-optimizing the most-ripe 25% of language workflows could move up to $1 trillion of HCP salaries away from admin and back into care. - Corti sells AI infrastructure and models to companies building healthcare applications — purpose-built, not rate-limited, accurate on medical terminology — rather than competing in the last-mile workflow layer. Cleve positions European compliance (AI Act, GDPR) as an advantage, not a drag, and says a very large US customer chose Corti partly because the market is drifting toward vendors aligned to those standards. - Cleve’s answer to hallucination is “a lot of tedious work across the factory floor,” not one benchmark. Corti’s Fact-R recursive-reasoning pipeline listens in real time with an orchestra of LLMs acting as judges; on public benchmarks he says it finds 49% more of what matters while cutting 88% of what doctors deem verbose or noise. A grounded warning he returns to: the well-known case of AI telling people to eat a rock a day. ### FAQ **Why does Andreas Cleve argue healthcare isn’t actually slow to adopt technology?** Cleve calls the “doctors are tech laggards” narrative “super boring” and lacking ambition. His evidence is physical: in a hospital or clinic there are more devices per room than in his Copenhagen office of 70 machine-learning researchers, so the sector clearly adopts lots of technology and does it quite fast — it just holds a high bar before doing so. His analogy is telling NVIDIA, ten years into GPUs, that they’ve been slow at it. The takeaway: healthcare doesn’t necessarily need to change, it needs help, and that’s a compounding, multi-decade opportunity for builders willing to do hard work. **How big does Cleve say the healthcare workforce shortage and the AI opportunity are?** Cleve cites the WHO’s upgraded estimate that the world will lack at least 10 million healthcare professionals by 2030 — and he scopes it carefully to just delivering the care we deliver today, not the more preventative, longevity-oriented care he expects by then. On the upside, he references a McKinsey/Accenture figure: if you take the roughly 25% of language-based workflows that are most ripe for AI and automate the parts that can be, you could move up to $1 trillion of healthcare-professional salaries away from admin and back into care. He frames that reallocation as the most obvious, low-hanging way to close part of the gap. **What is Corti’s role in the healthcare AI stack, and why does Cleve treat European compliance as an edge?** Corti sells AI infrastructure and models to companies building healthcare applications with AI inside — purpose-built tooling that isn’t rate-limited, has high accuracy, knows medical terminology, and ships endpoints and SDKs for things like revenue cycle management and documentation. Cleve is explicit that Corti is not strong at the last mile of workflows; partners handle that. On compliance, being from Europe makes them “really good at it,” and he says a very large US customer with a long purchasing cycle chose Corti partly because they’re aligned to standards like the AI Act and GDPR — betting the market will drift toward vendors who treat that as a nucleus of trust, not just red tape. **How does Corti’s Fact-R approach try to reduce hallucination, and what results does Cleve claim?** Cleve describes Fact-R as a recursive-reasoning pipeline rather than a one-shot or few-shot LLM. Instead of batch-processing a conversation at the end, it listens in continuously, using an orchestra of LLMs to identify the salient facts and revisit their importance as the conversation changes, with a series of models acting as judges. On public benchmarks, he says it finds 49% more of what matters — the really important healthcare information — while reducing 88% of what doctors would deem verbose or noise. He grounds the motivation in a survey of 2,000 of the earliest US adopters of AI scribe tools, who reported spending at least three hours a week just vetting and correcting their AI. **Where does Cleve see AI in healthcare going beyond admin and documentation?** Cleve’s most interesting frontier is new workflows that simply don’t happen today because they’re not affordable. His example: he picks up an unfamiliar medication and feels a soft, non-leading symptom — a bit of airiness in his head — but his physician doesn’t have time to call and check in, so nobody does. If the average cost per minute of high-quality medical reasoning drops by orders of magnitude with safe infrastructure, he asks why you wouldn’t have AIs proactively call patients across pharma, physical rehab, psychiatry, primary care, and home care to caution and guide them. The hard constraint, he stresses, is control — the AI has to be safe enough not to hallucinate, invoking the infamous “eat one rock a day” failure as the cautionary baseline. ## EP 33: Mastering Multi-Agent Systems | MongoDB’s Mikiko Chandrasekhar Guests: Mikiko Chandrasekhar, MongoDB · Published: 2025-07-23 · Duration: 40 min URL: https://chainofthought.show/podcast/33-mastering-multi-agent-systems-mongodbs-mikiko-chandrasekhar/ (full transcript on page) AI agents offer unprecedented power, but mastering agent reliability is the ultimate challenge for agentic systems to actually work in production. Mikiko Chandrashekar, Staff Developer Advocate at MongoDB, whose background spans the entire data-to-AI pipeline, unveils MongoDB's vision as the memory store for agents, supporting complex multi-agent systems from data storage and vector search to debugging chat logs. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - At a poll Mikiko Chandrasekhar ran during her AI Engineer World’s Fair lightning talk, roughly 80% of the room said they were building agents and nearly as many had agents in production — but when she asked who had gotten their agents working 100% of the time on the first try, not a single hand went up. That gap is where she locates the reliability problem. - Mikiko frames MongoDB’s role as the memory store for agentic systems — not only vector data for semantic search and RAG, but the chat logs and traces developers need to debug errors and keep improving the underlying data and processes. She points to the Voyage acquisition, which brings high-performing embedding and re-ranking models in-house. - Treat agents as software products, Mikiko argues, rejecting the “let them do their thing and catch errors later in the application layer” school. Non-determinism changes some things, but established software-engineering and MLOps best practices still apply with rigor — she credits Hamel Husain and Shreya Shankar as voices for that discipline. - The field over-indexed on single-agent failure; Mikiko points to the paper “Why Do Multi-Agent LLM Systems Fail?” which builds a taxonomy showing many of these failures are actually predictable and classifiable. Her takeaway: the techniques that improve one agent need to be extended to coordinated systems of agents. - For multi-agent observability, Mikiko wants tracing native to memory patterns few people discuss: blackboard memory, where agents post partial solutions to a shared read-write space, and a skills library that saves a solved pattern so the system doesn’t re-derive it each run. Each specialized agent should be judged against its own criteria, not one shared rubric. - Vibe checks are fine up to a point, but Mikiko insists agentic systems need quantitative metrics — “you can’t improve what you can’t measure.” The fact that an LLM or VLM does the reasoning doesn’t exempt it from measurement; for coding agents she’d track accuracy, runtime efficiency, and suggestion-acceptance rate, while flagging lines of code as a terrible metric. ### FAQ **How does MongoDB fit into building AI agents?** Mikiko Chandrasekhar describes MongoDB as the memory store for agentic applications. Beyond its flexible, polymorphic-schema document database, it provides vector search on Atlas for semantic similarity and RAG, memory-store modules (including integrations with LangGraph and partners), and storage for chat logs and traces so developers can debug agents and improve the data feeding them. The Voyage acquisition adds high-performing embedding and re-ranking models in-house. **Why are multi-agent systems so hard to make reliable?** Because LLMs are non-deterministic, even a structured agent workflow can return different answers — Mikiko says as often as one in 20 runs for complex tasks. At her AI Engineer World’s Fair talk, no one in the room had gotten agents working 100% of the time on the first try. In high-stakes settings like financial reports, medical triage, or loan underwriting, those intermittent failures can be disastrous, which is why she pushes consistency over aggregate accuracy. **What are blackboard memory and a skills library in multi-agent systems?** These are two memory concepts Mikiko says are emerging in recent papers. Blackboard memory is a shared read-write space where multiple agents post partial solutions and pick up each other’s trail, each contributing its own expertise toward a problem. A skills library saves an established pattern once the system has solved it — unlike a cache, which stores a query result to fetch later, it stores the solved approach so agents don’t have to re-derive it every time. **Should you treat AI agents like software products?** Yes — that’s the school of thought Mikiko Chandrasekhar follows. She rejects the view that you should let non-deterministic agents run free and clean up errors later in the application layer. Agents are still code-based software products, so the rigor and best practices from traditional software engineering and MLOps should be adapted to them rather than abandoned. She cites Hamel Husain and Shreya Shankar as strong advocates for that disciplined approach. **What metrics should teams use to evaluate AI agents?** Mikiko argues vibe checks alone aren’t enough; teams need quantitative metrics alongside qualitative feedback, because the language-model core of an agent doesn’t exempt it from measurement. For coding or copilot-style agents where the output is code, she points to measurable signals like accuracy, runtime efficiency, and how often a suggestion is accepted. She is emphatic that lines of code is a terrible metric — impact, not volume, is what matters. ## EP 32: The AI Agent Trust Gap: Bridging Risk to Reliability | Elastic’s Philipp Krenn Guests: Philipp Krenn, Elastic · Published: 2025-07-16 · Duration: 44 min URL: https://chainofthought.show/podcast/32-the-ai-agent-trust-gap-bridging-risk-to-reliability-elastics-philipp-krenn/ (full transcript on page) The age of ubiquitous AI agents is here, bringing immense potential - and unprecedented risk. Hosts Conor Bronsdon and Vikram Chatterji open the episode by discussing the urgent need for building trust and reliability into next-generation AI agents. Vikram unveils Galileo's free AI reliability platform for agents, featuring Luna 2 SLMs for real-time guardrails and its Insights Engine for automatic failure mode analysis. This platform enables cost-effective, low-latency production evaluations, significantly transforming debugging. ### Key takeaways - Elasticsearch quietly powers search across much of the internet. Philipp Krenn notes that Wikipedia and Stack Overflow run Elasticsearch behind their search boxes, and that almost everything on GitHub is cached or powered by Elasticsearch in the background — usage most people never see. - MCP is changing how systems are accessed, not just what data they hold. Krenn frames it as a shift in interaction mode: rather than writing against one specific REST API, you let the LLM figure out the MCP connection, fetch the right data, and see what actions it can run — including descriptive asks like “build me a Kibana dashboard.” - Elastic does not build large language models. Krenn says the models Elastic builds are for inference and re-ranking, and it relies on partners to supply the LLM that generates the answer — or, in Galileo’s case, the Luna models on the evaluation side. He describes the stack as an “AI lasagna” of layers where each player partners with the others. - Re-ranking is a small-model pattern that wasn’t practical before. Krenn walks through it: from a million documents you retrieve the first thousand, then run a more expensive, slightly slower re-ranker over just that top-thousand subset. A fast, efficient small language model makes running that costlier model on a narrowed set feasible. - Latency tolerance for AI is a temporary novelty effect. Krenn recalls that a 200-millisecond Elasticsearch query once felt “unacceptably slow,” yet a five-second LLM answer is currently treated as “perfectly fine.” He expects that forgiveness to fade as AI becomes standard and users demand faster, more real-time responses. - A chatbot’s promises can legally bind the company behind it. Krenn cites a Canadian court case where an airline lost after its chatbot told a customer something it shouldn’t have; because it was the company’s agent, the result had to stick. Many companies then pulled LLMs back from customer-facing roles, keeping them internal until guardrails could be trusted. ### FAQ **What does Elastic actually do in the AI stack, according to Philipp Krenn?** Krenn positions Elastic as the data layer. Its classic use is retrieval-augmented generation: you store your data, do the retrieval first to get the right context, then generate the output with an LLM. Elastic can also cache results so a similar question reuses an earlier answer for a faster, cheaper response. On the evaluation side, because Elastic does OpenTelemetry and the ELK stack handles logging, it collects performance, cost, and quality telemetry. Krenn is explicit that Elastic stores the results but leaves the evaluation itself to Galileo. **How do Galileo and Elastic integrate technically?** The integration runs on OpenTelemetry. Krenn explains that Elastic acts as the data store that keeps the telemetry, then uses OTLP — the wire protocol for OpenTelemetry — to move that observability data so it can be aggregated and evaluated. He draws a parallel: MCP is becoming the protocol for AI, and OpenTelemetry is the protocol for telemetry data; standardizing on shared protocols is what lets Elastic, Galileo, and others partner more easily. Because both sides speak the OpenTelemetry standard, the same data feeding Elastic can feed Galileo for evaluation. **What small models does Elastic build, and why not just use a general-purpose LLM?** Krenn describes two areas. For re-ranking, Elastic built a model that runs over a retrieved subset — for example, re-ranking the top thousand documents out of a million. For embeddings and inference, Elastic heavily uses E5, a multilingual model, in the dense vector space, plus a custom model for the sparse vector space that does keyword expansion to find related keywords. He says a general-purpose large language model would be too expensive and too slow for these tasks, so the right specialized models are needed instead. **What is Philipp Krenn’s advice for developers trying to take AI to production?** Krenn says it’s very easy to get started and build, but getting to production is the hard part. Guardrails are a very important piece, and you need the evaluation side to confirm the system is actually doing what you want, within expectations for cost and performance. Drawing on his observability background, he warns that observability is often treated as an afterthought and expects the same mistake with AI — people will eventually realize it shouldn’t be. He adds that AI may even raise the stakes: hallucinations and wrong answers create a stronger business drive to get the right answers. **How does Philipp Krenn see AI changing customer-facing work?** Krenn frames it through analogy and caution. He compares AI’s trajectory to self-driving cars: hyped years ago, slow to arrive, but now he rides in a Waymo frequently — suggesting AI may take longer to reach the “this is what happens in reality” phase. He hopes reliable agents mean never having to call a hotline or wait on a chat system again, freeing people doing repetitive call-center work for harder background problems. On job loss he is measured: automation fears have recurred since the industrial revolution without playing out, so he expects work to shift rather than disappear. ## EP 31: Architecting Reliable Agentic AI | Cisco’s Giovanna Carofiglio on the AGNTCY Collective Guests: Giovanna Carofiglio, Cisco · Published: 2025-07-09 · Duration: 41 min URL: https://chainofthought.show/podcast/31-architecting-reliable-agentic-ai-ciscos-giovanna-carofiglio-on-the-agntcy-collective/ (full transcript on page) The Internet of Agents is rapidly taking shape, necessitating innovative foundational standards, protocols, and evaluation methods for its success. Recorded at Cisco's office in San Jose, we welcome Giovanna Carofiglio, Distinguished Engineer and Senior Director at Outshift by Cisco. As a leader of the AGNTCY Collective (an open-source initiative by Cisco, Galileo, LangChain, and many other participating companies), Giovanna outlines the vision for agents to collaborate seamlessly across the enterprise and the internet. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Giovanna Carofiglio describes AGNTCY as an open-source collective Cisco, Galileo, and LangChain launched as founding members “in March.” The premise: for agents built in different frameworks and deployed remotely to collaborate, the ecosystem needs interoperable, distributed agentic communication — what she calls a new internet revolution. - She lays out AGNTCY’s pillars as discovery, compose-and-deploy, then communication. Agents get found by skills, publishers, tool compatibility, and reputation (an area she says “Galileo really helps build”), via a discovery service launched “a week ago.” Once located, they collaborate where they already run, often served as a service. - On protocols, Giovanna separates layers: MCP handles agent-to-tool interaction, especially for remote tools behind MCP servers, while agent-to-agent is trickier — AGNTCY worked on ACP with LangChain, and Google has since proposed A2A. She wants AGNTCY to support all of them, over a lower-level transport layer akin to TCP. - Cisco’s transport protocol SLIM was “recently announced” for group-based, low-latency, secure, interactive communication. Giovanna argues point-to-point isn’t enough: when communication is driven by natural-language questions, multiple agents collaborate, so the foundation should be secure by design, low latency, group based, and data centric. - For observability, Giovanna says AGNTCY — with Galileo, LangChain, Traceloop, Pydantic, and 50-plus partners — is defining an interoperable standard schema for agentic communication, extending OpenTelemetry so MELT telemetry stays open. An SDK supporting the schema released “two days ago,” and she wants to reconstruct the agentic graph. - Giovanna frames predictability as explainability: even stochastic, “magical” model behavior can be explained. Working in the collective and with Splunk at Cisco, she wants to test the space of solutions and keep output consistent for similar inputs. Her frontier is “active evaluation” — quantifying the margin for improvement, then remediating. ### FAQ **What is the AGNTCY collective and who founded it?** Giovanna Carofiglio describes AGNTCY as an open-source collective launched “in March,” with Cisco, Galileo, and LangChain as founding members and 50-plus partners. Her premise: for agents built in different frameworks and deployed remotely to collaborate, the ecosystem needs interoperable, distributed agentic communication — a new internet revolution like the one Cisco pioneered years ago. Its pillars: agent discovery (by skills, publisher, tool compatibility, and reputation), compose-and-deploy (agents collaborating where they run), then communication, observability, and evaluation. **How does Giovanna distinguish MCP, ACP, and A2A, and why a transport layer beneath them?** Giovanna says MCP and ACP aim at very different objectives. MCP targets agent-to-tool interaction — valuable especially when tools are remote rather than integrated, exposed like an API behind MCP servers — and she notes it’s popular because it’s simpler. Agent-to-agent is trickier: AGNTCY started from LangChain’s agentic protocol (ACP), and Google has since proposed A2A. AGNTCY wants to support them all. Her key argument: these application-layer protocols need a lower-level transport layer beneath them — secure by design, low latency, group based — like TCP underpins the internet. **How is AGNTCY approaching agentic observability, evaluation, and standards?** Giovanna says the foundation is defining an interoperable standard schema for agentic communication — the layer that lets agents be instrumented so metrics, events, logs, and traces (MELT telemetry) reconnect — with Galileo, LangChain, Traceloop, Pydantic, and 50-plus partners by extending OpenTelemetry to keep it open. An SDK supporting the schema released “two days ago.” Beyond visibility, she wants to reconstruct the agentic graph (noting Galileo released a way to do this) so developers and enterprises can see how agents communicate, how data passes between them, and how tools are called. **How does AGNTCY think about agent identity and security at scale?** Giovanna says AGNTCY recently released an agent identity component. It started from a schema for identifying agents — she calls it “OSF” [AGNTCY’s published schema is OASF] — meant to do for agents what OCSF (Open Cybersecurity Schema Framework) did for security: become a lingua franca. AGNTCY pushed a first version and welcomes contributions on what defines an agent, since even that can be challenged. The goal: agents uniquely identified, with clear provenance for agent and data, spelling out their skills to be discovered and rated. She frames identity as the start of broader security work. **What does Giovanna want to see next from Galileo and the collective?** Giovanna says she wants Galileo to bring its work on defining and recommending metrics — especially for multi-agent systems, agentic context, and communication — into AGNTCY, including sample metrics and a metrics computation engine to guide developers through assembling an application. She praises Galileo’s Luna small language models, saying evaluation should be constrained in budget and time, and that deep evaluation with SLMs is “amazing.” She names evaluation that needs less ground-truth data as an exciting frontier, and points listeners to the observability and evaluation group. ## EP 30: Taste Is The New Moat | Why Customer Obsession Wins in the AI Era Guests: Bharat Vasan, Intangible · Published: 2025-07-02 · Duration: 53 min URL: https://chainofthought.show/podcast/30-taste-is-the-new-moat-why-customer-obsession-wins-in-the-ai-era/ (full transcript on page) When AI makes creating content and code nearly free, how do you stand out? Differentiation now hinges on two things: unique taste and effective distribution. This week, Bharat Vasan, founder & CEO at Intangible and a "recovering VC," explains why the AI landscape compelled him to return to founding. He sees AI sparking a new creative revolution, similar to the early internet, that makes it easier than ever to bring ideas to life. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Bharat Vasan argues that as the marginal cost of writing code and creating content trends toward zero, differentiation collapses onto two things: taste — knowing what to make for a particular audience — and distribution — being discovered at all. When everyone can produce as much as people can consume, the question becomes how anyone tells your work apart. - Brand matters up and down the AI stack, but Vasan says it matters most at the application layer. Chipmakers and foundries have more of a moat and more forgiveness; in applications there is no loyalty — if the product is no good, people just turn it off. He frames the founder’s job as cultivating a community that extends benefit of the doubt as you build. - Vasan’s religion across startups is to ship as fast as you can, because shipping clarifies. If it doesn’t work you learn something, and if it works you learn something — but not shipping lets you live in a five- or ten-year bubble where your team and you both avoid bad news. Long gaps between ships leave you with too little data, stuck in your own head. - Vasan reframes startups as a test of resilience, creativity, and persistence rather than intelligence — if it were about intelligence, he notes, the smartest people would found every startup. He estimates roughly 90% of people, in sports or startups, don’t have what it takes to get up every day and grind with no immediate reward. - On customer obsession, Vasan says the CEO has to set the tone personally. He cites Parker Conrad caring about payroll and benefits at Rippling, Jamie Siminoff putting his email on every Ring box, and Bezos at Amazon — when the CEO reads support tickets, the rest of the org has no excuse to deprioritize customers. - Vasan is direct that venture capital is rocket fuel meant for rockets — most businesses are not VC businesses. He says the AI age will make it possible to build durable lifestyle businesses that pay your bills for life. If you do take VC, you’ve signed up to build a $5-10 billion company, and you have to be honest about whether that’s the game you want. ### FAQ **Why did Bharat Vasan leave VC to found Intangible?** Vasan — who calls himself a “recovering VC” — says the number one reason was that it’s an exciting time to be alive in technology. He compares the current AI moment to the early days of the internet, recalling his first encounter with Telnet and FTP as a college student newly arrived from India and the feeling that it could be his conduit to the whole world. His belief is that the best way to learn in these early, messy phases is to actually build, and he couldn’t think of a bigger opportunity than catering to human creativity. **What advantages does Vasan think startups have over incumbents in the AI era?** Vasan names three. First, speed — startups have nothing they’re attached to, while big companies are slowed by their P&L, review cycles, HR cycles, and public reporting. Second, creativity and a key insight the market overlooks — he points to Melanie Perkins of Canva being told her market was too small, and to ideas like Airbnb and getting into strangers’ cars sounding implausible until they worked. Third, hunger: he likens incumbents to having “been a Lannister for so long you forget how to fight,” versus a founder’s daily drive to make the thing work. **What does Vasan mean when he says “shipping clarifies”?** Vasan means that releasing your work is what tells you where you actually stand. Rather than one big beta event, he describes hundreds of smaller ships over time that build your judgment and conviction about what’s working. When the gap between ships is long, he says, you have too little data and live in your own head with a roadmap and an imagined future. Users telling you “I like it but I don’t love it” gives you something concrete to fix — and if you know where you stand, you can do something about it. **How does Vasan advise founders to build a brand and cut through the noise?** His top advice is that brand follows success — it gets built if your startup works — so do the work and be relentlessly customer-obsessed, spending hours with customers rather than at conferences and VC parties. Second, get good at language and figure out which formats let you best communicate your vision, since you are the limiting factor. Third, take creative risks — early on nobody knows you exist, so there’s little brand to lose. He invokes Steve Jobs on the difficulty of simplifying a company down to one thing. **How does Vasan describe the current venture capital landscape for AI startups?** Vasan calls VC a craft business that doesn’t scale as well as people thought, reshaped by the end of the ZIRP era and a lack of exits. He sees it splitting into two paths: very large funds that index the market and accept lower returns to avoid missing the few market-making companies, and lone-wolf craft investors who need an edge in network, deal flow, and thesis. He describes the “tortured space” investors live in — paying up and owning too little versus missing a deal entirely — and notes it takes seven to fifteen years to learn if you were right. ## EP 29: The Emerging AI Agent Stack | CrewAI’s João Moura Guests: João Moura, CrewAI · Published: 2025-06-25 · Duration: 50 min URL: https://chainofthought.show/podcast/29-the-emerging-ai-agent-stack-crewais-jo-o-moura/ (full transcript on page) Unlocking AI agents for knowledge work automation and scaling intelligent, multi-agent systems within enterprises fundamentally requires measurability, reliability, and trust. João Moura, founder & CEO of CrewAI, joins Galileo’s host Conor Bronsdon and Vikram Chatterji to unpack and define the emerging AI agent stack. They explore how enterprises are moving beyond initial curiosity to tackle critical questions around provisioning, authentication, and measurement for hundreds or thousands of agents in production. ### Key takeaways - João Moura traces the agent stack from the data layer up — Databricks, Snowflake, BigQuery, then LLM orchestration, authentication, and connectors, topped by agentic apps with a natural-language UI. His point is that no single product covers it: enterprises going “agentic native,” running hundreds or thousands of agents in production, need a whole suite. - Moura argues the early enterprise questions — memory, graphs versus events, which open-source project — are a commodity, with roughly twenty ways to do memory and none of it mattering much. The questions that decide outcomes surface only once a company commits to scale: how to provision agents, authenticate them, evaluate them, and measure them. - Infrastructure choices are reopening rather than settling, per Moura. Companies are calling workloads back from the cloud, some to self-host and some demanding on-prem physical servers, and he points to a scenario where running at huge scale on physical hardware can save around 60% — though he won’t predict which way it lands. - A self-described simple guy, Moura wishes more stayed the same — HTTP and REST are his bread and butter. He sees MCP and A2A with a head start while AWS, Microsoft, and Google battle over the standard, but warns proper standards take years and whoever launches a “successful” protocol likely gains little edge while others build. - Defining success up front is what unlocks speed, per Moura: CrewAI disqualifies prospects who arrive asking what to build or what others are doing, because without a success definition there is nothing to measure. LLM-as-a-judge for quality and hallucination gets you far enough, but he insists serious use cases must go custom. - Moura positions CrewAI as an orchestration and control plane rather than a builder of individual agents, betting on democratizing access so everyone in a company can build, not just a technical subset. He pictures a repository of reusable agents recombined to automate processes, and a later stage that rethinks workflows entirely. ### FAQ **How does João Moura define the emerging AI agent stack?** Moura describes it as a suite, not a single product, because the stack spans many things — all the agentic resources, not just the agents. It starts at the data-management layer (Databricks, Snowflake, BigQuery), works up through LLM orchestration, authentication, and connectors, and tops out at agentic apps with a natural UI you can interface with. He frames the deeper questions — provisioning, authentication, evaluation, measurement — as the ones that emerge once a company decides to go “agentic native” and run hundreds or thousands of agents in production. **Why does João Moura think infrastructure and protocol choices are being re-examined?** Moura says agents are “shaking up everything” — the stacks, the protocols, and the market. He sees companies pulling workloads back from the cloud, some choosing to self-host and some insisting on on-prem physical servers, noting a scenario where running at huge scale on physical hardware can save around 60%. On protocols he wishes more stayed the same, favoring HTTP and REST, and expects MCP and A2A to lead while AWS, Microsoft, and Google battle over the standard — though he cautions proper standards will take years to establish. **What does João Moura say about how fast enterprises are actually adopting agents?** Moura reports that highly regulated industries are moving exceptionally fast — financial services and insurance in particular — drawn by efficiency gains in back-office work ripe for automation. He notes CrewAI is in conversations with major banks that are already building, and says its most advanced customers are putting around five use cases into production a month — a pace he calls insane that itself creates new control and monitoring challenges, since once teams stack one set of use cases on the next, someone has to own, monitor, and verify that all of it is actually working. **How does João Moura think teams should approach measurement and reliability for agents?** Moura argues that knowing what success looks like is what unlocks teams to move faster — CrewAI even disqualifies prospects who can’t define success for a use case, since without that there is nothing to measure. He says LLM-as-a-judge for quality and hallucination takes you far enough to start, but serious use cases have to go custom, with no way around it. He warns that traditional LLM tracing falls short for agents — there are many more layers to observe and evaluate — and that not thinking about data from day zero comes back to haunt teams fast. **What is CrewAI’s strategic focus according to João Moura?** Moura frames CrewAI as an orchestration and control plane focused on democratizing access, so everyone in a company — not just a technical subset — can build agents, with no-code as a dedicated, invested-in product and a way to give back to the community. He envisions repositories of reusable agents recombined to automate processes, and eventually rethinking workflows rather than just automating existing steps. On near-term strategy he calls it a land grab: closing as many customers as possible, helping them reach value, and leaning heavily on content and courses with a ship-fast culture. ## EP 28: AMD's Challenge to NVIDIA: The Open Ecosystem Bet | Anush Elangovan & Sharon Zhou Guests: Anush Elangovan & Sharon Zhou · Published: 2025-06-18 · Duration: 49 min URL: https://chainofthought.show/podcast/28-amds-challenge-to-nvidia-the-open-ecosystem-bet-anush-elangovan-and-sharon-zhou/ (full transcript on page) How is an open ecosystem powering the next generation of AI for developers and leaders? Broadcasting live from the heart of the action at AMD's Advancing AI 2025, Chain of Thought host Conor Bronsdon welcomes AMD’s Anush Elangovan, VP of AI Software, and Sharon Zhou, VP of AI. They unpack AMD's groundbreaking transformation from a hardware giant to a leader in full-stack AI, committed to an open ecosystem. ### Key takeaways - Anush Elangovan frames AMD’s MI350 GPU series, built on the CDNA 4 architecture, as delivering up to 20 petaflops of FP4 performance, which he calls “mind blowing.” He credits AMD’s chiplet design for NUMA load balancing and power savings, since unused chips can be turned off and only what a workload needs is powered on. - CDNA 4 introduces new AI data types — both FP4 and FP6 — and Anush highlights that FP6 throughput is nearly that of FP4. That lets data scientists step models down from FP8 to FP6 as an intermediary before reaching FP4, which he frames as where things are headed. - Anush says the MI350 series reaches 288 GB of memory, enough that “500,000,000,000 parameter models that can run on one GPU,” with at-scale deployments going far larger. He cites deep partnerships with seven of the ten top AI companies as evidence the generational hardware investment is paying off. - AMD launched ROCm 7, ROCm Enterprise AI, and the AMD Developer Cloud, with Day Zero native support for frontier models including DeepSeek, the Llamas, and Qwen. Anush says internal and external repos are now identical, so external contributors merge real code — and that path got Triton and PyTorch working on Windows for the Strix Halo laptop. - Anush argues software outlives any single hardware generation, so AMD now treats ROCm as a product with a decade-long plan. He predicts that once ROCm 7 lands, ROCm 8, 9, and 10 could ship on a roughly six-week cadence, modeled on how Chrome ships continuously — paired with a claimed 40% tokens-per-dollar savings versus Blackwell. - Sharon Zhou, former CEO and co-founder of Lamini, now VP of AI at AMD, frames developer adoption around listening, a “happy path” to early success, and a pre-believers versus pre-buyers funnel. She cites Lamini having run on over 300 AMD GPUs across courses with Andrew Ng and Meta, plus excitement for self-improving AI via “vibe based feedback.” ### FAQ **What did AMD announce for the MI350 GPU series and CDNA 4 architecture on this episode?** Anush Elangovan said the MI350 series, built on CDNA 4, delivers up to 20 petaflops of FP4 performance and reaches 288 GB of memory — enough, he said, that “500,000,000,000 parameter models that can run on one GPU.” CDNA 4 adds FP4 and FP6 data types, with FP6 throughput nearly matching FP4 so teams can step models from FP8 to FP6 before FP4. He credited AMD’s chiplet architecture for NUMA load balancing and for power savings by powering on only the chips in use, and noted the MI400 series is “less than twelve months” away. **What is AMD’s ROCm 7 and Developer Cloud story, and how open is the contribution model?** Anush said AMD launched ROCm 7, ROCm Enterprise AI (with cluster management and ML-operations capabilities), and the AMD Developer Cloud, with Day Zero native support for frontier models including DeepSeek, the Llamas, and Qwen. The Developer Cloud lets you sign in with a GitHub ID, spin up a real instance, and includes 25 hours of free credits for event attendees. He said internal and external repositories are now exactly the same, so external developers can contribute code that AMD merges — external contributors helped get Triton and PyTorch running on Windows for the Strix Halo laptop. **What performance and cost gains did AMD claim for agentic and inference workloads?** In his question, Conor cited AMD benchmarks showing a 3.8x generational improvement for AI agents and up to 4.2x for summarization tasks on AMD infrastructure. Anush did not restate those figures; he pivoted to performance per dollar, saying the MI350 series delivers a 40% tokens-per-dollar savings versus the competitor’s latest Blackwell platform — savings he said add up quickly at 250 million tokens a day or a billion tokens a day, across on-prem or CSP deployments. He framed agents as one form of “intelligent autonomous systems,” alongside virtual and eventually physical robots, all resting on heavy compute and software infrastructure. **Who is Sharon Zhou and what is she focused on at AMD?** Sharon Zhou is the former CEO and co-founder of Lamini and now VP of AI at AMD. Conor noted she has taught over a million people about AI. She described her focus as a combination of AI research and teaching — making ROCm and AMD’s software more accessible to developers and showing that modern workloads like vibe coding agents and reinforcement learning run well, and maybe optimized, on AMD. She is working with Andrew Ng at Deep Learning AI, and said that at Lamini she gave over 50 keynotes in a year and ran on over 300 AMD GPUs serving three courses with Andrew Ng plus one with Meta. **How does Sharon Zhou think about engaging developers and the future of AI?** Sharon said the first step is listening to the community, then building a “happy path” — three steps to succeed on AMD so attention-limited developers see something work fast. She uses a pre-believers versus pre-buyers funnel across three personas: AI developers, researchers, and leaders. On MCP, she stressed its value as an open protocol that can become a standard, noting OpenAI’s endorsement. Looking ahead, she is most excited about self-improving AI — models that edit their own training data — guided by what she calls “vibe based feedback”: casual natural language to nudge models. ## EP 27: Your Key to AI Success is Hiding in Plain Sight | Cohesity's Greg Statton Guests: Greg Statton, Cohesity · Published: 2025-06-11 · Duration: 46 min URL: https://chainofthought.show/podcast/27-your-key-to-ai-success-is-hiding-in-plain-sight-cohesitys-greg-statton/ (full transcript on page) What if the most valuable data in your enterprise—the key to your AI future—is sitting dormant in your backups, treated like an insurance policy you hope to never use? Join host Conor Bronsdon with Greg Statton, VP of AI Solutions at Cohesity, for an inside look at how they are turning this passive data into an active asset to power generative AI applications. ### Key takeaways - Cohesity began as an infinitely scalable distributed file system aimed at backup, then layered on security, file-system access, and now generative AI. Greg Statton frames each phase as a stop on one journey — re-leveraging data that customers already store but rarely touch — noting it is “pretty wild to spend money on something” you hold and hope you never have to use. - An early attempt to monetize that data flopped. Cohesity exposed an analytics workbench letting customers run MapReduce queries by writing custom Java mappers and reducers and uploading JAR files to the cluster. Statton admits “literally nobody in the enterprise” wanted to do that, so they paused it and pursued easier-to-adopt steps instead. - Cohesity’s machine-learning entry point was inline anomaly detection in the backup stream. By modeling how data changes between backups, the system fingerprints sudden entropy spikes that often signal ransomware encrypting or wiping data. Statton says it has caught malware before customers’ own SecOps tools, alerting the SOC to begin its security procedures. - The generative-AI push started as Statton’s personal experiment about three years ago. Leading a team of field experts who kept re-answering the same questions, he saw a “semantic divide” between official FAQs and how newcomers actually phrased things, then built a fix on GPT-3 and Cohesity’s internal docs — effectively a RAG system before the term was common, using an in-memory TF-IDF and cosine-similarity index, with reference links back to source files. The founder and CEO saw it and moved him into R&D to productize it. - Statton’s core advice is unglamorous: before chasing leaderboard models or agentic flows, get peers across the org to agree on a data governance model, map where data lives, and decide what AI should never touch. He calls this data readiness, preparedness, and hygiene the hard part everyone skips, because however good the model is, “it’s still going to give you garbage if you give it garbage.” - On evaluation, Statton softened his own “turtles all the way down” critique of LLM-as-judge: he says “I don’t know if you should never do this” — it can be valid if you have proper evaluations on the evaluator itself, but it shouldn’t be the only method or use the same model that generated the response. He favors decomposing outputs into claims with an LLM, then using other models plus retrieved context to check each one. ### FAQ **Who is Greg Statton and what does he do at Cohesity?** Greg Statton serves in Cohesity’s office of the CTO as vice president of AI solutions, now working in core R&D. On the episode he notes he is approaching his ten-year anniversary at the company and has worked in nearly every department except finance — including marketing, sales, architect and SE roles, and a global field role. He describes himself as a tinkerer who, despite not holding a PhD in machine learning or AI, has been curious about the space for the last fifteen to twenty years. Conor Bronsdon recorded the conversation on the road at Microsoft Build. **How did Cohesity’s generative-AI work actually begin?** It started as an internal experiment Statton ran about three years ago. He led a team of field experts whom new sales engineers kept asking the same questions, even after being pointed to docs and FAQs — a “semantic divide” between how FAQs were written and how newcomers phrased questions. After OpenAI opened GPT-3 access (he recalls a Reddit post), he built a web UI with an in-memory semantic index using TF-IDF vectorization and cosine similarity, retrieved relevant doc chunks, and passed them to GPT-3 — an early RAG system. Cohesity launched its first generative-AI apps about two years ago. **What does Statton say enterprises should do before building AI on their data?** He argues the first step is non-technical: get peers cross-functionally — marketing, HR, engineering — to agree on a governance model for the data, which he says almost no organization has. From there, map where all the data lives, decide which versions matter, set access controls, and identify data that should never interact with an AI model. He frames this as data readiness, preparedness, and hygiene — the work people skip, because however good the model is, “it’s still going to give you garbage if you give it garbage.” **Why is Statton skeptical of using one LLM to evaluate another?** He has called the approach “turtles all the way down,” though on the episode he walks it back slightly — saying “I don’t know if you should never do this,” and that it may be valid if you have proper evaluations on the evaluator model itself. His concerns: it shouldn’t be the only method, and it shouldn’t use the same model that generated the response, which he likens to walking into a room of thieves as a cop and asking if they are thieves. He also warns it can overfit toward a model’s own bias. His preferred pattern decomposes a generation into discrete claims and uses other models plus retrieved context to check each one. **How does Cohesity use backup data for security?** Because Cohesity ingests data repeatedly, it can model how that data changes between backups and build fingerprints of normal behavior. It added anomaly detection inline as data is ingested, flagging when a dataset’s entropy changes dramatically — often a sign that a bad actor is encrypting or wiping data. Statton says this has worked successfully: in some cases the engine caught the malware before the customer’s own SecOps tools did, and the alert from Cohesity let the SOC begin its security procedures. He frames it as augmenting, not replacing, a company’s existing security stack. ## EP 26: Why Gamers Paved the Way for AI | Databricks' Carly Taylor Guests: Carly Taylor, Databricks · Published: 2025-06-04 · Duration: 49 min URL: https://chainofthought.show/podcast/26-why-gamers-paved-the-way-for-ai-databricks-carly-taylor/ (full transcript on page) What if the pixels and polygons of your favorite video games were the secret architects of today's AI revolution? Carly Taylor, Field CTO for Gaming at Databricks and founder of ggAI, joins host Conor Bronsdon to illuminate the direct line from video game innovation to the current AI landscape. ### Key takeaways - Carly Taylor traces today’s AI boom to gaming’s push on GPUs. The first GPU — NVIDIA’s GeForce 256, with “the processing power of a potato nowadays” — arrived about twenty-five years ago, and the consumer drive for more polygons drove the cost of each matrix calculation down continuously, priming the compute economics generative AI now depends on. - Roller Coaster Tycoon is Taylor’s emblem of the gamer mindset. Chris Sawyer wrote the entire game himself in Assembly — “talking straight to the machine” — because the CPU power of the era couldn’t handle the guest interactions he wanted. Her read: “your only limit is your imagination,” and gamers won’t let compute power stop them. - Inside game studios, data is routinely “a second class citizen.” Development is creatively driven and everything bends toward shipping, so telemetry is rarely a block to launch — you can ship a game without it, you just won’t know what’s going on. The data adoption curve for most studios still lags where she’d expect most tech companies to be. - Production is the only real test environment. Taylor cites a LinkedIn joke she posted — teams ask to hire QA before launch and the answer is “we have QA. They’re called users.” Playtesting, even with eye-tracking labs, can’t catch the bugs, map holes, and translation gaps that emerge once “roving bands of lunatics” break things no one anticipated. - Putting an LLM on NPCs runs into infrastructure reality. Clients lack the GPU for inference, and shipping trained weights gives away IP “worth hundreds of millions of dollars, billions maybe”; game servers have no GPUs because they aren’t rendering anything. Taylor’s answer is a third path — compute beside the server, feeding results to clients. - Taylor is building an LLM pipeline over Steam reviews for localization and sentiment. The payoff is escaping the “chicken and egg” of social listening: rather than community managers scrolling Twitter and Reddit needing to already know what to search for, an LLM can be asked open-endedly what players are talking about and surface emergent issues. ### FAQ **How did the video game industry enable today’s AI boom?** Carly Taylor argues gamers have pushed the limits of computing since games first ran on computers, and that consumer demand for richer graphics drove GPU development and pricing. She notes the first GPU — NVIDIA’s GeForce 256 — arrived about twenty-five years ago with trivial power by today’s standards, and that the cost of each matrix calculation on a GPU has fallen continuously since. Without that consumer-driven acceleration making compute cheaper, she says, “I don’t think that we would have been primed for the AI revolution that we have today.” **Why does game data so often end up neglected during development?** Taylor explains that game development is creatively driven and dominated by deadlines — teams “are constantly counting down the minutes” to alpha, beta, and ship. Because a game can technically launch without telemetry, in-game data “often just ends up becoming a second class citizen”; it’s rarely a block to ship unless it’s a critical financial metric. The harder problem, she adds, isn’t always collection but usage: many studios capture data yet aren’t set up to act on it, and trial and error is expensive when getting a change into a live build can take six to eight weeks. **Why isn’t it realistic to just “put an LLM” on game NPCs?** Taylor lays out the infrastructure constraints. Games typically run either on the client (a player’s PC or console) or on a server. Running a proprietary LLM client-side means shipping model weights she values at “hundreds of millions of dollars, billions maybe” onto users’ machines, giving away IP; fine-tuning an open model on studio data leaks IP the same way. Server-side fails because game servers don’t have GPUs — they aren’t rendering anything. Her alternative is a third option: compute that sits next to the server and feeds results to clients, which she says vendors almost never raise. **What are the most realistic near-term uses of AI in gaming?** Taylor sees analytics and QA as far more solvable than smarter NPCs. Once data is centralized, governed, and has good metadata, modern platforms (she cites Databricks and Unity Catalog) let non-technical staff query data in plain language — an exec can self-serve a question about how a skin sold, freeing the data team for deeper experimentation like A/B tests and causal inference. The second area is reinforcement-learning QA bots trained by human testers; she finds them promising for overnight smoke tests with a few humans in the loop, but says she wouldn’t yet trust them for a full QA pass. **How can LLMs improve how studios understand player sentiment?** Taylor describes a project pulling in Steam reviews to handle localization and sentiment across languages and to generate exec-facing reports. The key advantage over older social-listening tools is escaping a “chicken and egg” problem: previously, community managers scrolled Twitter and Reddit and had to already know what to search for before they could find complaints. With an LLM you can ask open-ended questions — what are people talking about, what’s gaining traction, what seems different from usual — and let it surface emergent issues without being told the problem in advance. ## EP 25: The 2025 AI Shift: From Chat to Task Completion & Reliable Action | Galileo Founders Guests: Vikram Chatterji & Atindriyo Sanyal · Published: 2025-05-28 · Duration: 45 min URL: https://chainofthought.show/podcast/25-the-2025-ai-shift-from-chat-to-task-completion-and-reliable-action-galileo-founders/ (full transcript on page) AI in 2025 promises intelligent action, not just smarter chat. But are enterprises prepared for the agentic shift and the complex reliability hurdles it brings? Join host Conor Bronsdon on Chain of Thought with fellow co-hosts and Galileo co-founders, Vikram Chatterji (CEO) and Atindriyo Sanyal (CTO), as they explore this pivotal transformation. They discuss how generative AI is evolving from a simple tool into a powerful engine for enterprise task automation, a significant advance driving the pursuit of substantial ROI. ### Key takeaways - Vikram Chatterji frames 2025’s defining shift as generative AI moving from chat completion toward task completion and action. He argues this elevates the enterprise conversation beyond building chatbots to automating tasks, with the goal of higher ROI on OpEx and CapEx — though he concedes “agents” is “an overblown term.” - Vikram describes a “gold rush” among middleware providers racing to build orchestration and frameworks — citing OpenAI’s Agents SDK, Anthropic’s MCP, and Google’s A2A. He’s blunt that some “don’t mean anything” and are “shallow libraries,” but reads the collective effort as the system falling into place to make task-completion apps real. - Atin Sanyal pitches evaluation agents as composable functions that sit where the developer works — for example an eval agent installed as a tool inside Cursor that automatically fixes potential issues. He argues eval is expanding into a broader AI reliability story rather than the point-in-time activity it used to be, with MCP servers standardizing communication around the LLM. - Vikram positions Galileo’s LUNA metrics engine as the answer to AI’s measurement problem, which he calls the company’s north star. He notes generative AI lost the F1 score that classification had — and that F1 “wasn’t a very good metric” anyway — and describes LUNA as a factory for creating, testing, and low-latency-tuning metrics at scale. - On vibe coding, Atin says it makes developers “10x more faster,” compressing weeks of work into an hour, but warns that blind vibe coding yields “a pile of crap” that isn’t deployable. He argues a limited set of high-quality design patterns from earlier software eras is re-emerging with the LLM in the mix, which sharpens the need for evaluation and reliability. - Vikram reports the biggest enterprise ROI so far is internal-facing use cases — telco outage detection, wealth-management report generation, customer support, accounting — where the accuracy bar is lower so teams experiment rapidly while OpEx savings are “massive.” He says enterprises now want a single pane of glass over risk vectors across a hundred-plus use cases. ### FAQ **What does Vikram Chatterji mean by the shift from “chat to task completion”?** Vikram claims the defining 2025 theme is generative AI moving beyond chat completion and generation toward actually completing tasks and taking actions — what he calls agents, while admitting the term is “overblown.” He argues this matters for enterprises because it raises the stakes from building chatbots to automating real work, targeting higher ROI from both OpEx and CapEx. He frames language models as becoming operating systems that power these actions, and says adoption for this use case is moving “very rapidly,” even if widespread task completion is still arriving in “fits and starts.” **How do Vikram and Atin describe Galileo’s approach to the AI “measurement problem”?** Vikram says the measurement problem has been Galileo’s north star since the company began: generative AI has no equivalent of the F1 score that NLP classification relied on, and even F1, in his view, “wasn’t a very good metric.” His pitch is that you can’t build a high-quality system you can’t measure. He describes Galileo’s LUNA metrics engine as a factory for creating, testing, and tuning low-latency metrics that auto-adapt at scale. Atin adds that there’s “no one-stop-shop metric,” which is why they baked auto-adaptation into LUNA. Both are the founders’ claims, not verified here. **What did the guests say about vibe coding and reliability?** Atin claims vibe coding makes developers “10x more faster,” turning weeks of work into an hour, but cautions that coding “completely blindly” produces “a pile of crap” that isn’t deployable — great for prototyping, not production. He argues a limited number of proven design patterns are re-emerging with LLMs in the mix. Vikram frames speed (vibe coding) and quality (reliability) as two different problems, both accelerating, and says moving from “one to production” still needs an expert in the loop plus real guardrails. These are the founders’ characterizations. **Which enterprise use cases did Vikram highlight as already delivering value?** Vikram claims the strongest early ROI is internal-facing: telcos building faster outage detection, wealth-management teams at large banks generating reports faster, customer support across nearly every enterprise, plus accounting and finance. He argues the lower accuracy bar on internal tools lets teams experiment fast while OpEx savings are “massive.” Some retailers are moving external-facing too, where the accuracy bar rises and real-time guardrails on millions of queries become essential. Conor adds that big organizations can aggregate value from simple fixes like helping people find docs. **How fast do the guests expect enterprise AI adoption to move versus cloud?** Both founders use the cloud-migration analogy. Vikram describes a diffusion-of-innovations pattern: a handful of innovators in finance or healthcare prove out the model-risk-management and ops layers, see large efficiency gains, and trigger a domino effect of FOMO across the industry — the same dynamic that played out with cloud, but happening a bit faster. Atin goes further, claiming adoption “will happen a 100 x faster than the cloud,” because the enabling technology already lives on the cloud and, with sufficient compliance, almost anyone can “write two lines of code” to bake in intelligence. ## EP 24: Amplitude's AI Playbook: How Wade Chambers Builds for the Agentic Future Guests: Wade Chambers, Amplitude · Published: 2025-05-21 · Duration: 47 min URL: https://chainofthought.show/podcast/24-amplitudes-ai-playbook-how-wade-chambers-builds-for-the-agentic-future/ (full transcript on page) As AI redefines how products are built and customers are understood, what are the core strategies engineering leaders use to drive innovation and create lasting value? Join host Conor Bronsdon as he welcomes Wade Chambers, Chief Engineering Officer at Amplitude, to explore these critical questions. Wade shares how Amplitude is leveraging AI to deepen customer understanding and enhance product experiences, transforming raw data into actionable insights across their platform. ### Key takeaways - Wade Chambers ranks contribution on a deliberately steep scale: one point for flagging a problem, ten for getting to its root cause, a hundred for defining a first-principles solution that maps best, and a thousand for delivering the impact. He says it out loud, he explains, because the further up the stack you go, the more you want people to keep betting. - For Chambers, durable AI advantage at Amplitude comes from proprietary behavioral data and domain expertise, not LLM access. Because Amplitude has unlocked usage and friction data for so long, teams clear the low-hanging fruit and then “discover deeper patterns,” building a “deeper intelligence” he calls “pretty hard to replicate.” - Amplitude evaluates its AI against synthetic data, not just live output. Chambers’ teams generate synthetic results and compare them to previous norms, giving “a historical view and one that’s been generated on the other side” — two models to compare — while internally monitoring the “Ask Amplitude” agent for hallucinations. - Chambers refuses one universal eval metric: each org sets its own KPIs because infrastructure teams measure uptime and disk utilization while data teams care about throughput and the latency from receiving a record to it going live. He concedes Amplitude isn’t “golden across the board on every dimension” and doesn’t know any company that is. - Avoiding what Conor calls acquisition “organ rejection” runs on three things, Chambers says: clear alignment on shared objectives and success metrics, cultural onboarding with clear ownership and early wins, and empowerment — autonomy plus the resources to execute. He credits this for Command AI both embracing and extending Amplitude’s culture “even in a short period of time.” - Chambers, who places himself in his “late 50s,” frames AI as a personal mandate for leaders: put your own oxygen mask on first. Have you actually tried vibe coding, and which LLM did you like the best and why? He warns ICs they’ll be “increasingly going to be less and less competent” over the next five to ten years unless they keep a beginner’s mindset. ### FAQ **What is Wade Chambers’ point system for measuring engineering contribution?** On Chain of Thought, Amplitude Chief Engineering Officer Wade Chambers describes a deliberately steep scoring model: you earn one point for flagging a problem, ten points for getting to its root cause, a hundred points for defining from first principles a solution that maps best, and a thousand points for actually delivering the impact of that solution. He says it out loud so people keep betting further up the stack, and pairs it with coaching, context, and making clear who owns what. **How does Amplitude build a sustainable AI advantage instead of a thin LLM wrapper?** Wade Chambers argues the moat is proprietary behavioral data and domain expertise, not model access. Amplitude already understands how cohorts behave, where they hit friction, and what to do about it — and has unlocked that data long enough to move past low-hanging fruit into deeper patterns. Layering iterative learning on those data sets, Chambers says, produces a “deeper intelligence” that is “pretty hard to replicate.” **How does Amplitude evaluate its AI for hallucinations and accuracy?** Wade Chambers describes heavy internal inspection: teams hold a thesis on what the output should be and constantly use the product to confirm they see it. Amplitude runs an internal agent called “Ask Amplitude” and monitors what it is asked and how it responds to catch hallucinations. For evaluation, teams generate synthetic results and compare them against previous norms, giving both a historical baseline and a generated one — two models to compare and contrast. **What does Wade Chambers say about integrating acquired AI talent like Command AI?** Wade Chambers credits three pillars for avoiding what Conor Bronsdon called acquisition “organ rejection”: clear alignment on shared objectives and success metrics, cultural onboarding with clear communication, ownership and early wins, and empowerment — granting autonomy and the resources to deliver. He says the Command AI team has already both embraced and extended Amplitude’s culture and influenced the roadmap “even in a short period of time,” with its leaders showing up as future leaders. **What advice does Wade Chambers give leaders and ICs for adopting AI?** Wade Chambers tells leaders to “put your own oxygen mask on first” — actually try vibe coding, form a view on which LLM you like best and why — so you build empathy before coaching others, then push many people through the learning curve together as a “pit crew.” For individual contributors, he urges a beginner’s mindset: test the tools, find their limits, use several LLMs and tools like Windsurf, and expect to feel “less and less competent” over the next five to ten years unless you keep learning. ## EP 23: First Code, Then AGI: Software’s Event Horizon with Poolside Founders Jason Warner & Eiso Kant Guests: Jason Warner & Eiso Kant · Published: 2025-05-14 · Duration: 51 min URL: https://chainofthought.show/podcast/23-first-code-then-agi-softwares-event-horizon-with-poolside-founders-jason-warner-and-eiso-k/ (full transcript on page) Is the prevailing approach to Artificial General Intelligence (AGI) missing a crucial step – deep, focused specialization? For the first time since co-founding Poolside, CEO Jason Warner & CTO Eiso Kant reunite on a podcast articulating their distinct vision for AI's future with our host, Conor Bronsdon. ### Key takeaways - Poolside is an AGI company, not a coding-assistant company. Jason Warner: “Poolside is an AGI company. We’re going after AGI because we wanna build intelligence on compute.” Code is the chosen path — a “proxy task for intelligence,” Eiso Kant argues, because it demands both broad world knowledge and long-horizon reasoning. - The 2023 contrarian bet was reinforcement learning, not bigger pretraining. The consensus then: just “scale up next token prediction, scale up model size, provide more data, and we are, quote unquote AGI.” The Poolside founders bet the key scaling axis was RL instead — and “a lot of people thought we were kind of crazy.” In ’25, Jason says, it “looks more than prescient.” - Specialization is a budget decision, not a different machine. About a third of the way into training, the Poolside founders say, “we start biasing them massively towards software development.” Early on the models look general-purpose — “great at planning your trip to London or writing a poem.” Flip “a couple of dials” and they’d revert; the choice is where to spend finite compute. - Win the hardest customer first. Poolside started deliberately with high-consequence enterprise systems — “money movement, MRI machines, drone software.” Jason Warner frames it as NP-hard: solve the hardest case and “you can do it anywhere.” On those who equate a Flappy Birds clone with a defense system: “I want you nowhere near them… I want to sleep well at night.” - Refuse the “devil’s trade” on source code. Jason Warner never wants to ask a customer to “send me their most valuable asset, essentially their source code, And to quote unquote trust me.” Poolside’s technical claim: it doesn’t need customer code to keep improving — that comes from “reinforcement learning, synthetic data” — so the model goes to the data behind the firewall. - Developers don’t disappear; the job re-anchors on taste and agency. Jason Warner: “developers exist. The relationship to the tasks and jobs change.” Memorized SDKs and APIs lose value; “taste, discernment, and judgment” gain it. Eiso Kant: software’s surface area “will grow still a thousand X,” and the day-to-day shifts to “agency and… delegating” agents — “a little more like a tech lead.” ### FAQ **How does Poolside’s training approach differ from a general-purpose model like those from OpenAI or Anthropic?** The early phase, Poolside’s founders say, looks “extremely similar to how you would train any general purpose foundation model,” so the models are genuinely good at non-code tasks. About a third of the way in, Poolside “start[s] biasing them massively towards software development,” ending with reinforcement learning from code-execution feedback across “almost a million repositories… fully containerized with their full test suite.” Flipping “a couple of dials” would revert them to general models — the difference is where you spend finite compute. Jason Warner’s analogy: everyone builds a sedan, Poolside a truck — “you can’t slap the brakes from a sedan on a truck.” **Why does Poolside focus on high-consequence enterprise systems instead of consumer “vibe coding”?** Jason Warner argues the industry keeps mistaking “a Flappy Birds clone or an asteroid single shot” for “building the underlying systems that control money movement, MRI machines, drone software” — “it couldn’t be further from the truth.” Poolside aimed at “the hardest possible place” first: large enterprises bound by safety, regulatory, and compliance regimes, where customers aren’t looking to “vibe code their way to a solution.” The founders frame it as an NP-hard bet — solve the hardest case and “you can do it anywhere.” Eiso Kant adds the real shift is AI becoming “a coworker… part of your organization.” **What is the “devil’s trade,” and how does Poolside avoid asking customers to hand over source code?** Jason Warner describes the devil’s trade as the industry asking companies to send their most valuable asset — their source code — to an AI provider on the promise it won’t be misused. He refuses: “I never want to ask a customer to… send me their most valuable asset… And to quote unquote trust me.” Technically, Poolside doesn’t need customer source code to keep improving, because progress comes from reinforcement learning and synthetic data. Eiso Kant frames the goal as bringing “the model to the data instead of sending off the data to the model.” Jason adds an enterprise sale must satisfy three buyers: the CTO/CIO, the general counsel, and the CISO. **What does AGI actually mean to Poolside’s founders?** Eiso Kant is deliberately cautious — the term “is a moving target… got 50 different definitions” — and offers the one he finds useful “for the moment we’re in”: reaching “human level intelligence and capabilities for the vast majority of knowledge work that we do behind a laptop.” He breaks it into three parts: today’s AI isn’t embodied (robotics is early), so the first horizon is “the world of bits”; models hold broad knowledge but have “really poor reasoning… over long time horizons”; and there’s the ability to act — to use a computer and tools. He shies away from “AGI” because it will likely stay a moving target “for the next decade.” **Will software developers still have jobs as AI gets more capable?** Both founders say yes, but the role changes. Jason Warner: “developers exist. The relationship to the tasks and jobs change.” The skills that fade are memorized SDKs, APIs, and language-specific identity (“I’m a Java developer”); what remains is “taste, discernment, and judgment.” He expects a future where “you might have 10 human developers, but you also might have a thousand digital developers.” Eiso Kant adds that demand grows — “the surface area of software will grow still a thousand X” — and the day-to-day moves toward “agency and… delegating” a variable number of agents, making an engineer look “a little more like a tech lead.” ## EP 22: AI's Two Extremes – Foundations & The Frontier | Databricks’ Denny Lee Guests: Denny Lee, Databricks · Published: 2025-05-07 · Duration: 44 min URL: https://chainofthought.show/podcast/22-ais-two-extremes-foundations-and-the-frontier-databricks-denny-lee/ (full transcript on page) The AI landscape often pulls us between the allure of cutting-edge models and the quiet necessity of foundational work—yet how do these extremes actually connect to deliver value? Join host Conor Bronsdon as he welcomes Denny Lee, a self-proclaimed "data nerd" and Product Management Director, Developer Relations at Dataricks, to unpack this very spectrum, from AI's core infrastructure to its most advanced applications. ### Key takeaways - Logging and tracing feel boring, but they’re what everything else rests on. Denny Lee traces MLflow’s origin to a familiar mess: hyperparameters in spreadsheets, versions named “golden,” “golden_v1,” “golden_v2,” until nobody recalls which set built a given model. Across millions of automated iterations, you can only evaluate what you recorded. - The fight isn’t for one standard — it’s against having fifteen. The industry needs a shared way to record what matters — hyperparameters, lineage, architecture — so any tool can read it and you keep your choice of tooling. Lee doesn’t care if it’s MLflow, OpenTelemetry, or something else: “We just want one,” since each new “unifying” standard just becomes number sixteen. - Synthetic data isn’t fake data — treating it that way gets the math wrong. Lee’s correction: it’s generated from real data to preserve the patterns a model needs while protecting privacy. If it were truly fake, it would carry a different pattern and different math. Either way the ceiling is the real data — “you’re only as good as your data.” - Climb the GenAI ladder before you ever pre-train. Lee’s sequence: inference on a solid model, then surround it with retrieval and agents (RAG for exact facts like a CEO’s salary, plus authorization to block questions that shouldn’t reach the model), then fine-tune with a thousand-shot of Q&A — and pre-train last, a few-million-dollar A100/H100 commitment most teams never need. - Don’t build one Uber-model; build many small ones. The value isn’t reaching for the largest model — a Llama 4 at 450B or 440B, a number Lee hedges may already be stale — but composing smaller, specialized models you’ve improved into your own IP. He reads DeepSeek the same way: not “we no longer need GPUs,” but a strong showcase of reinforcement-learning ideas. - Power, not chips or software, is the binding constraint in the West. Lee notes two GPU makers — NVIDIA, far ahead, and a catching-up AMD — and TSMC already making 5nm wafers in Arizona while Taiwan pushes toward 1.5nm. But North America and Europe “don’t have enough power”: Microsoft is reactivating a Three Mile Island unit, while China just adds more chips on fast fabric. ### FAQ **Why does Denny Lee insist on logging and tracing before anything else in an AI project?** Because evaluation and feedback are the whole game, and you can’t evaluate what you didn’t record. Lee points back to why MLflow was built: developers kept hyperparameters in spreadsheets and CSVs, with versions named “golden,” “golden_v1,” and so on, until nobody could remember which hyperparameters, dataset, or data version produced a given model. For a couple of throwaway experiments a notepad is fine — but real work runs hundreds, thousands, or millions of automated iterations, and only logging everything lets you make sense of what’s effective and improve it. **What is the “GenAI ladder,” and when should you actually pre-train your own model?** Lee’s GenAI ladder is a cost-ordered sequence: start with inference on a good solid model; surround it with retrieval and agents (RAG to pull exact facts, authorization to block questions that shouldn’t reach the model); then fine-tune, beginning with a thousand-shot of questions and answers to see how far you get; and only pre-train if you truly need to. His point: you rarely need the largest model or many GPUs to improve a model into your own IP — and pre-training from the start is a few-million-dollar A100/H100 commitment most teams should avoid. **Is synthetic data just fake data, and why does Denny Lee push back on that framing?** Lee rejects “synthetic equals fake.” Synthetic data is generated from real data so it carries the key patterns a model needs while protecting the privacy of the individuals underneath. If it were truly fake, he argues, it would have a completely different pattern and a completely different set of math behind it — so it wouldn’t be useful for the same purpose. He cites conversations with people at Replit and at Gretel, a synthetic-data company, and ties it to privacy: it’s one way to keep useful patterns while reducing the chance of exposing real individuals. **Why does Denny Lee say power, not compute, is the real bottleneck for AI in the West?** Lee frames it as “good old fashioned: I don’t have racks, I don’t have power.” He acknowledges chip-supply concerns — effectively two GPU makers in NVIDIA and a catching-up AMD, with TSMC already making 5nm wafers in Arizona and Taiwan pushing toward 1.5nm — but argues North America and Europe have under-invested massively in power and can’t build data centers fast enough. He points to Microsoft reactivating a Three Mile Island unit (~three years out) and contrasts China, which produces enormous power and excels at network fabric, so it can simply add more chips even if each is less efficient. **What does the Latanya Sweeney re-identification example teach about privacy in LLMs?** It shows how little information it takes to identify someone. Lee recounts Sweeney’s k-anonymity work: for about $25 she obtained a CD of Massachusetts voter records, joined it with public medical data, and re-identified Governor William Weld’s medical records with just three questions. Lee recalls — from memory, and flagging the figure may not be exact — that roughly 83% of the US population was uniquely identifiable by birth date, five-digit ZIP, and gender alone. Because LLMs return specific, not aggregate, answers, he says you can’t guarantee individual privacy; differential privacy’s noise protects aggregate queries but doesn’t directly translate. ## EP 21: Why Enterprises Need a Different Approach to AI Agents | Lyzr’s Siva Surendira Guests: Siva Surendira, Lyzr · Published: 2025-04-30 · Duration: 39 min URL: https://chainofthought.show/podcast/21-why-enterprises-need-a-different-approach-to-ai-agents-lyzrs-siva-surendira/ (full transcript on page) Agentic AI exploded in 2025, but how do businesses move beyond prototypes to deploy reliable, valuable agents at scale? Join host Conor Bronsdon and Lyzr AI CEO Siva Surendira as they discuss the complexities of building and managing AI agents for enterprises. Siva shares his journey creating Lyzr, focusing on making powerful agent frameworks accessible and trustworthy for enterprise developers. ### Key takeaways - Most enterprise AI never ships. Siva Surendira says “95% of all AI work today done by enterprises remain as proof of concepts” and never reach production. The blockers are safe-AI and responsible-AI concerns — hallucination, toxicity, prompt injection — like an agent approving a $100,000 refund where it should have been $10. - Lyzr builds guardrails into the core, not as an add-on. Because the team wrote the agent framework themselves, they made safe/responsible-AI guardrails “natively to the core agent architecture” — agents that are “responsible by design,” not a third-party API bolted on, enforceable centrally “whether you are building 10 agents or 10,000,000 agents.” - Orchestration isn’t either/or — it’s hybrid. Siva names three modes: managerial (a master agent calling workers, which he credits Autogen with first), DAG popularized by LangGraph, and Lyzr’s hybrid flow. For high-stakes cases like a refund agent, hybrid flow mixes in deep-learning, ML, code, and SQL agents to make the workflow “far more deterministic.” - Integration and data — not models — are the real bottleneck. For any enterprise above $100M revenue, custom in-house apps “completely outweigh” systems like Salesforce and SAP. These decade-old apps have no documentation, so agent actions can’t be defined, and the data isn’t “LLM ready.” The third bottleneck: the skills to build on closed, no-code-only platforms. - Go after the human layer, not SaaS. Siva’s advice: target “the mundane work that teams are doing.” His example — a food-delivery company of 15,000+ employees that hired ~30 contract HR staff for two months of annual reviews — replaced that with agents his team built and deployed in three weeks, now running at scale for 15,000 employees. - Self-healing agents aren’t here yet — and Siva has the receipt. Testing an automated feedback-and-conflict-resolution loop across a multi-agent setup, “such loop cost us $300” in LLM calls and still didn’t fix the problem, so they stopped it. He pushes local open-source models instead: thousands of in-house experiments at “like 100 times cheaper” than GPT or Claude. ### FAQ **Why do most enterprise AI agents fail to reach production?** On the show, Lyzr’s Siva Surendira says “95% of all AI work today done by enterprises remain as proof of concepts” and don’t move into production. When his team dug into why, the primary concerns were safe AI and responsible AI: enterprises worry about getting sued, or mistakenly approving a $100,000 refund when it should have been $10. He frames the obstacles as hallucination, toxicity, and prompt injection — which is why Lyzr built guardrails “natively to the core agent architecture” rather than as a bolted-on third-party API. **What are the three types of multi-agent orchestration Lyzr uses?** Siva describes three. First, managerial orchestration — a master agent that calls worker agents at will to complete a task; he credits Autogen as the first to do this. Second, DAG (directed acyclic graph), a step-by-step workflow he says LangGraph launched. Third, Lyzr’s hybrid flow: rather than being restricted to a DAG or managerial setup, you can also bring in deep-learning, ML, code, and SQL agents — connected to Azure ML Studio or Amazon SageMaker — which he says makes the system “far more deterministic” for high-complexity cases like a customer-refund or retirement-planning agent. **What are the biggest bottlenecks to enterprise agent adoption?** Siva names three. First, integration: for any enterprise above $100M (mid-market) or $1B (large) in revenue, custom in-house apps “completely outweigh” standard systems like Salesforce and SAP, and these decades-old apps lack documentation, so the actions an agent can take are undefined. Second, data readiness — enterprise data isn’t “LLM ready”; it’s scattered, the gap tools like Glean address, though many orgs aren’t even ready to feed data to Glean yet. Third, the skills to build, since many enterprise agent platforms are closed and no-code-only, limiting you to low-complexity use cases. **Where does Siva see the biggest opportunity for agents — replacing SaaS or replacing human work?** He’s skeptical the main prize is replacing SaaS. While he notes agents eating into the software layer — citing a large media org that chose Moveworks over ServiceNow’s own agents — he says “the biggest unlock is always going to be agents replacing the mundane work that humans do.” His example: a food-delivery company with 15,000+ employees that hired about 30 contract HR staff every June for two months of reviews; agents his team built in three weeks now handle that at scale for 15,000 employees. He expects agents to start with low-skill follow-up work and move toward high-skill work in 2026—2027. **How does Siva think about evals and choosing between local and frontier models?** He describes two kinds: agent-level accuracy, groundedness, and hallucination, plus test cases for how well the agent solves its business problem. He sees a shift to local open-source models, which let enterprises run thousands of experiments with data in-house at “like 100 times cheaper” than GPT or Claude — and fine-tuned local models become the enterprise’s own IP. On AWS Bedrock, he names NVIDIA NIM and Meta Llama as the #1 and #2 local models for banks and insurers. He recommends a two-layered setup: Lyzr’s built-in guardrails plus an eval platform like Galileo as the “antivirus” layer. ## EP 20: Breaking the Language Barrier: Smartling's AI Translation Pipeline | Olga Beregovaya Guests: Olga Beregovaya, Smartling · Published: 2025-04-23 · Duration: 41 min URL: https://chainofthought.show/podcast/20-breaking-the-language-barrier-smartlings-ai-translation-pipeline-olga-beregovaya/ (full transcript on page) Are we on the verge of removing all language barriers with AI? Olga Beregovaya, VP of AI at Smartling, joins host Conor Bronsdon to tackle this question, discussing the evolution from rule-based NLP to today's powerful LLMs. Together, they confront the persistent challenges that stand in the way, like the English-centric nature of AI, domain-specific inaccuracies, and the unpredictability of model hallucinations. ### Key takeaways - Olga Beregovaya benchmarks human output at “a translator can deliver 2,000 words a day — that’s the metric,” and says a well-trained, well-prompted, fine-tuned model lifts that by “magnitude and order of magnitude.” She frames the payoff as a package: lower cost, higher productivity, and more predictable quality. - The bottleneck in language AI isn’t big-model capability — it’s the English-centric core. Beregovaya notes most foundational models remain “English language and English phenomena centric,” so a carefully crafted English prompt often “will render a better result in the target language” than a linguist’s prompt in their own native language. - Stuffing every linguistic asset into a prompt is, in Beregovaya’s words, a “waste of fortune.” Rather than feed translation memories, glossaries, and 100-page style guides into a 30,000-token prompt, Smartling uses RAG to “only fetch what’s relevant” — “a huge, huge, huge” game-changer and “a great way of mitigating biases and hallucinations.” - Smaller purpose-built models are beating generalists at translation now, but Beregovaya hedges it as a “pause and see what’s next” moment: she cites research that the more tasks a model handles, the more its performance on each individual one degrades — “jack of all trades, master of none” — yet she expects larger models to eventually catch up. - Data curation now matters more than data scale, Beregovaya argues: cleaner, smaller datasets beat “access to larger but potentially noisy” ones. She flags a hard floor — “you cannot fine-tune a model unless you have, like, your 100,000 strings” — and says the real struggle is harmonizing inconsistent bitext, glossary, and style-guide data. - Beregovaya’s sharpest agentic example is self-healing: a quality-estimation pass flags errors, then a smart agent says “heal everything that I’ve just identified.” Agents also route content and assign translators — but she warns teams to first ask: “does it call for an agent, or are we doing just fine with rule-based approaches?” ### FAQ **How much does AI speed up human translation according to Smartling?** Olga Beregovaya, VP of AI at Smartling, sets the human baseline at roughly 2,000 words per day per translator. With a well-trained, well-prompted, fine-tuned model doing the heavy lifting and a human compensating for the delta AI can’t cover, she describes the productivity gain as “magnitude and order of magnitude.” She frames the combined result as lower costs, higher productivity, and much more predictable quality — though she stops short of citing a single fixed multiplier. **Why do English-centric AI models cause problems for translation?** Beregovaya explains that most foundational models remain “English language and English phenomena centric,” so even with correct vocabulary the model can describe American-culture-centric phenomena while sounding fluent in the target language, costing factual and anthropological accuracy. Models also react poorly to prompts written in languages other than English — she notes a carefully crafted English prompt often produces a better target-language result than a linguist’s prompt in their own native language. Sometimes neural machine translation outperforms an LLM outright. **How does RAG improve AI translation quality?** According to Beregovaya, translation runs on assets like translation memories (legacy translations as a parallel corpus/bitext), glossaries, do-not-translate lists, and style guides that can run 100 pages. Feeding all of that into a prompt is a “waste of fortune”; RAG instead retrieves only what’s relevant, by keyword or concept, from curated sources. She calls RAG a “huge” game-changer and “a great way of mitigating biases and hallucinations,” especially when combined with knowledge graphs for structured retrieval. **Are smaller purpose-built models better than large models for translation?** Beregovaya says smaller, purpose-built models currently deliver more predictable, more explainable outputs for specific languages and domains, citing research that a model handling more tasks degrades on each individual one — “jack of all trades, master of none.” But she hedges this as a “pause and see” moment, expecting larger foundational models to eventually catch up as they learn from field feedback and expand language coverage. She also flags a fine-tuning floor of roughly 100,000 strings. **When will AI erase the language barrier in translation?** Beregovaya is candid that it’s her opinion amid passionate industry debate, but says the trajectory toward human parity is climbing sharply — faster for romance-language families, slower for agglutinative or Finno-Ugric languages. She believes “that eventually is around the corner — it can be three years, it can be five years,” and that the “language barrier as we know it is going to be erased.” Until biases and hallucinations are solved and language coverage is even, linguists will still be needed to produce datasets and fact-check output. ## EP 19: Low-Code AI: From Requirements to Apps in Minutes | OutSystems' Rodrigo Coutinho Guests: Rodrigo Coutinho, OutSystems · Published: 2025-04-16 · Duration: 40 min URL: https://chainofthought.show/podcast/19-low-code-ai-from-requirements-to-apps-in-minutes-outsystems-rodrigo-coutinho/ (full transcript on page) What if you could turn a requirement document into a full enterprise application in just minutes? Rodrigo Coutinho, co-founder and AI Product Manager at OutSystems, joins hosts Conor Bronsdon and Atin Sanyal to explore this new reality of AI-driven development. Rodrigo shares insights from OutSystems' nearly 25-year journey, detailing their early adoption of AI and the development of their AI platform, Mentor. ### Key takeaways - OutSystems started investing in AI back in 2018, in what Rodrigo Coutinho calls the era of deep-learning neural networks — using it to suggest a developer’s next move and to run automated reasoning over applications, checking them against security and performance standards. That “old school AI” predated the LLM shift the company built Mentor on. - The first thing OutSystems shipped with LLMs was modest: a feature letting a developer write a sentence in English and have it translated into a database access. They ran the large language model in-house at the time — “OpenAI wasn’t open yet,” not yet available to everybody — and that feature still ships today. - Coutinho’s core claim for Mentor: from a simple prompt or a requirement document, you can generate a full enterprise application — UI, business logic, and database — in a few minutes, with one-click publish that produces a ready container, database, and sample data. The speed claim is paired with a hard validation requirement, not stated alone. - Speed exposes an ownership problem when AI generates traditional code. Coutinho says some OutSystems partners turned off code-generation helpers entirely, because they ended up with “a lot of generated code that nobody understood” — it worked, then suddenly stopped, and there was nobody to maintain or fix it because no one was accountable for it. - OutSystems’ defense is that they generate into a visual modeling language rather than raw code, which they argue makes outputs easier to validate in one step. On top of that, AI agents check a finished, running application for performance, security, and compliance problems, and AI generates sample data so users can eyeball whether the result is what they wanted. - Coutinho frames the developer’s role as shifting from “a big pusher trying to make an integration work” toward team leader — doing code validation, talking to the business, and orchestrating the parts. He insists the magic comes from mixing AI with humans: at every step the human should know what’s happening and keep control to change the result. ### FAQ **What is OutSystems Mentor?** Mentor is OutSystems’ umbrella for the AI tools layered on top of its low-code platform, targeting several stages of the process. One tool goes from a prompt or requirement document to a full enterprise application — UI, business logic, and database — with one-click publish that generates a ready container, database, and sample data. Another lets you iterate on a running app with context-aware suggestions. A validation layer runs AI agents to check for performance, security, and compliance issues. Finally, an AI agent builder lets customers embed AI into their own applications. **How does OutSystems try to keep AI-generated applications trustworthy?** Rodrigo Coutinho says everything Mentor generates has to be a valid application that compiles and does something useful. Because OutSystems generates into a visual modeling language rather than raw code, they argue it’s easier to validate in a single step. They layer additional AI tools to check whether an app is secure, fast, and doing the right thing, run AI agents over the finished running app for performance, security, and compliance, and auto-generate sample data so users can immediately verify the output. The goal is catching hallucinations and security problems in the pipeline. **Does AI-generated code create an accountability problem?** Coutinho says yes, and it shows up most when these tools generate traditional code. He described OutSystems partners who turned off code-generation helpers entirely because they accumulated generated code that nobody understood — it worked, then suddenly stopped, and there was no one to maintain or fix it because nobody was responsible for it. He calls this a “limit situation,” but uses it to argue the promise of “anybody can do applications” isn’t fully here — you still need expertise to read the code, confirm it’s correct, and fit it to your architecture and compliance requirements. **Does OutSystems build its own AI models or use providers like OpenAI?** Both. Coutinho says OutSystems uses multiple models, including its own as well as OpenAI, and cautioned the specific answer could be outdated within a week because the team constantly tests which model performs best. As a default they lean toward public models whenever possible — they don’t have to host or train them, and avoid the people-investment running models in-house requires. They go in-house when data is sensitive or a model needs tight tuning to OutSystems’ systems. He noted their AI team is creative about saving money, but they tend to optimize for customer results over cost. **Where does Rodrigo Coutinho see the biggest challenge for AI in software development?** Not the technology, the use cases. Coutinho says this revolution mimicked older ones where the technology arrived ahead of the use cases. Many OutSystems customers know AI is important and is the future but don’t yet know how to use it or which use cases genuinely benefit. He cited customers already using it for ticket deflection in support, automating approval decisions with agents, and matching job offers to CVs, but said the bigger challenge is educating people on where the technology shines and working on ideation, countering a perception that these technologies “fix everything.” ## EP 18: AI Won't Solve Your Toughest Engineering Problems | Honeycomb’s Charity Majors Guests: Charity Majors, Honeycomb · Published: 2025-04-09 · Duration: 42 min URL: https://chainofthought.show/podcast/18-ai-wont-solve-your-toughest-engineering-problems-honeycombs-charity-majors/ (full transcript on page) Generative AI dominates the conversation, but does it actually make it easier to build, lead, and sustain high-performing engineering teams? Host Conor Bronsdon sits down with Charity Majors, co-founder and CTO of Honeycomb (.io), and the mind behind charity.wtf. Known for her sharp insights and unfiltered opinions, Charity kicks off the discussion by expanding on her popular article: 'Generative AI is not going to build your engineering team for you.' ### Key takeaways - Majors frames the AI debate as a category error: “Software engineering is not about writing code. It is about solving business problems with technology.” Code generation is the easy part — “generating code is a lot easier than owning it over the fullness of time.” - Most of this moment isn’t new, it’s amplified. Majors says you can swap “AI” for “automation” nine times out of ten and the sentence still holds: “the same thing. Only a little bit faster and more intensely than before.” The genuinely novel wrinkle is nondeterminism. - AI makes juniors faster, not obsolete. Majors says today’s juniors are “sharper, more educated, more skilled than juniors have ever been.” A junior engineer on her team spends the day asking an AI chat window the easy questions first, then arrives at senior engineers with a sharp one. - The center of gravity is moving from pre-production to production: “the source of truth is what is happening right now.” As AI-generated code floods in, the old debug path breaks — “What if there’s no expert? What if no one actually understands it or ever understood it?” - Beware cognitive decay, but design around it. Majors: shifting “from doing the work to just supervising the work” makes “your expertise decays.” An SRE agent “only 20 or 30% as good” as a world-class one is still “great for the world” — if it teaches as it works rather than silently doing it for you. - Software systems are feedback loops, and observability is the meta-loop: “a big feedback loop made up of a lot of other feedback loops.” Tightening them — code review, deploys — lowers cognitive overhead, so it “feels like you’re just building.” ### FAQ **Why does Charity Majors argue against hiring only senior engineers?** Majors calls software “an apprenticeship industry, full stop” — almost everything you learn happens at work, from each other. Hunting for “the most senior engineer that they can get for the least money” is shortsighted, she says, because what matters is the team’s ability to ship — which doesn’t correlate tightly with seniority. Her highest-performing teams spanned a range of levels where people grew and asked hard questions; the all-staff-plus teams she worked on cut corners and were bored. She adds that today’s juniors are sharper than ever, with AI letting them learn faster. **Does Majors think writing code is the hard part of software engineering?** No — she calls writing code the easy part, flagging it’s “a little spicy” and “not always true,” rooted in her ops background. Greenfield code is typically “the fastest, easiest part of the lifecycle” compared with maintaining, operating, and debugging it. That’s why, she argues, AI reached for code generation first: it’s the lowest-hanging fruit and an easy place to raise money. She predicts a split between “disposable code” nobody is expected to understand and “the code that runs the world” — banks, delivery, anything transactional — where “somebody’s got to understand that shit.” **What does Majors mean by “cognitive decay,” and how can teams avoid it?** She cites research that when you shift “from doing the work to just supervising the work,” your expertise decays — but stresses it isn’t inevitable, having been a risk as long as automation has existed. Her remedy is choosing tools that bring you along. She contrasts two agentic SREs: one that silently “does it for you,” making you a worse engineer because you stop asking questions, versus one that says “Here’s what I’m seeing. Have you tried this?” and teaches you as you go. Even an agent “only 20 or 30% as good” as a world-class SRE is, she says, “great for the world.” **How does Majors think AI changes observability and debugging?** Majors says the center of gravity is shifting from pre-production to production — “the source of truth is what is happening right now” — and the flood of AI-generated code breaks the usual path of finding whoever changed something and asking them. The new question becomes “What if there’s no expert? What if no one actually understands it or ever understood it?” You have to understand production on its own terms — the founding story of Honeycomb at Parse, where developers uploaded snippets the team couldn’t track down. The added challenge with AI is nondeterminism — same input, no longer the same output. **What is Majors’ advice for engineers who are skeptical or tired of AI hype?** “If you hate AI, you should really embrace it” — get to know it up close, because “nobody cares about the critique of somebody who is… way away from it.” She admits she “started pretty grumpy about it” and came around to seeing “a massive opportunity.” Her core line: “The fastest way to make yourself irrelevant is to not embrace the new.” She also notes that in tech “the scale of the hype often correlates with the eventual scale of impact,” comparing AI to the advent of the computer or the internet — while urging engineers to stay well-rounded, with cats and family, for a long career. ## EP 17: Inside IBM's watsonx: Building Enterprise AI That Ships | Dr. Maryam Ashoori Guests: Maryam Ashoori, IBM · Published: 2025-04-02 · Duration: 45 min URL: https://chainofthought.show/podcast/17-inside-ibms-watsonx-building-enterprise-ai-that-ships-dr-maryam-ashoori/ (full transcript on page) Building trustworthy, scalable AI isn't just about models; it's about navigating a complex ecosystem of tools and regulations. Join hosts Conor Bronsdon and Atindriyo Sanyal as they explore these challenges with Dr. Maryam Ashoori, Head of Product for watsonx AI at IBM. To meet these challenges, Maryam explains how watsonx simplifies the AI stack, automates pipelines, and empowers enterprises to scale their AI operations while optimizing costs rapidly. ### Key takeaways - Maryam Ashoori frames watsonx around the three enterprise challenges she sees consistently: responsible (trustworthy) implementation of AI, cost-performance optimization, and capturing the ROI of generative AI through automation. These — not a single “wow” model — were the guiding principles for the platform’s design. - On models, watsonx’s bet is optionality over a single provider: state-of-the-art open-source models available the same day they release, commercial models via partnerships with Meta and Mistral (Mistral Large), and the ability to import your own. Ashoori’s conviction: “one single model is not the solution” — it’s mix-and-match across sizes, architectures, and license terms. - IBM trains its Granite models from scratch so it can stand behind them — documenting training-data lineage, filtering toxic and copyrighted content “to best of our knowledge,” being transparent about how they’re trained, and providing client indemnification. Ashoori ties this to customers in regulated sectors like finance, insurance, and health care. - Ashoori breaks agent quality into four areas to assess per customer: the LLM itself (hallucination, jailbreak, the usual guardrails); agent-specific guardrails like faithfulness of the action taken; agent evaluation — a superset of LLM evals plus tool-calling consequences; and observability and governance across build time, run time, and over time. - Ashoori’s 1,000-developer US survey found only 24% of AI app developers called themselves knowledgeable and skilled on GenAI — and more than half use 5 to 15 tools daily yet will spend “not more than two hours” evaluating a new one. Her prescription: know your problem and build a point of view, so you can tell noise from what’s worth your limited time. - Ashoori is blunt that what the market calls “agents” today is mostly “LLM with function calling and tool calling,” not the autonomous reasoning, planning, and decision-making she studied 20 years ago. She doesn’t think current models reliably make sound decisions yet — there’s “some sort of preliminary planning” — but is “personally super excited” about the next six to nine months. ### FAQ **What is watsonx and how does Maryam Ashoori position it against other AI platforms?** Ashoori is Head of Product for watsonx AI, IBM’s AI development studio. She names three differentiators. First, hybrid deployment so customers aren’t locked in — the platform can run on premises or the cloud of their choice, rather than being tied to one model provider. Second, trust: IBM trains its own Granite models with documented lineage and indemnification, and ships governance and observability (watsonx.governance) with guardrails on inputs, outputs, and orchestration. Third, simplicity — integrating frameworks like CrewAI and LangGraph behind the scenes so a developer works through one SDK. **How does IBM approach build-versus-buy for models in watsonx?** Ashoori’s answer is optionality rather than one model. watsonx offers a range of state-of-the-art models across different sizes, architectures, and license terms, because global customers face varying license and GPU restrictions by region. IBM adds open-source models the same day they release, commercial models through partnerships with Meta and Mistral (she cites Mistral Large), and lets customers import their own. The same optionality applies to customization — from prompt engineering and RAG up to full fine-tuning, parameter-efficient fine-tuning, and alignment tuning. Her stated belief: “one single model is not the solution.” **Why does Ashoori say agents make governance and evaluation harder than LLMs alone?** Because agents act. With last year’s LLMs, she says the worst case was generating inappropriate content. Agents take actions, and actions have consequences — her example is an agent connecting to a sensitive structured database and deleting or combining customer data, which amplifies the impact. So watsonx aims not only to document action lineage but to proactively detect inappropriate actions and stop them, keeping a human in the loop so no high-stakes action runs automatically. She frames agent evaluation as a superset of LLM evals plus tool-calling: accuracy now depends on external APIs and the data they return. **What did IBM’s developer survey find about AI skills and tool fatigue?** Ashoori ran a survey of 1,000 US developers building AI applications. Only 24% described themselves as knowledgeable and skilled on GenAI — a gap she ties to why watsonx invests in automation and guidance for people entering AI development without deep AI backgrounds. On tooling, more than half said they use 5 to 15 tools daily in a fast-moving market, yet will spend “not more than two hours” evaluating whether a new tool fits. On AI-assisted coding, she recalled — hedging on the exact figure, “maybe it was 49%” — that a sizable share use it often, saving one to two hours a day on average, with a small handful saving more than four hours. **What does Ashoori think is overhyped about agentic AI right now?** Ashoori — who did multi-agent systems for her master’s degrees roughly 20 years ago, before deep learning — says the misconception is that today’s agents already deliver autonomous reasoning. What’s commonly called an “agent” is really an LLM with function and tool calling, which she values for bringing GenAI into every corner of the enterprise, but it isn’t the autonomous reasoning, planning, and decision-making she means. She doesn’t think models reliably make sound decisions yet — there’s “some sort of preliminary planning” — but is excited about the next six to nine months as the stack matures. **Where does Ashoori see GenAI delivering the most value in the enterprise?** A second misconception she flags: that GenAI fits every enterprise use case. It doesn’t — the work is understanding what GenAI is actually capable of. The pattern she highlights as highest-value is content-grounded question-answering, especially in customer care, where you equip people with answers drawn from a verifiable body of information. That was last year’s RAG story; now, with agents, the system can fire up a web-search API when the answer isn’t in that body and bring it back — always with a human in the loop to verify. ## EP 16: Information Symmetry: DevRev's Bet on AI-Driven Enterprise Decisions | Manoj Agarwal Guests: Manoj Agarwal & Yash Sheth · Published: 2025-03-26 · Duration: 41 min URL: https://chainofthought.show/podcast/16-information-symmetry-devrevs-bet-on-ai-driven-enterprise-decisions-manoj-agarwal/ (full transcript on page) What if everyone in your organization had equal information at all times? Would meetings even exist? This week, we dive into the concept of information symmetry with Manoj Agarwal, co-founder and president of DevRev. Manoj, along with hosts Conor Bronsdon and Yash Sheth, explores how DevRev is connecting data, personalizing schemas, and automating complex tasks, offering a glimpse into the next generation of AI-driven workflows. ### Key takeaways - Manoj Agarwal frames enterprise dysfunction as an information-symmetry problem: most meetings exist only because attendees hold different data. As he put it, “if the same information was available to every single person in that room, did we even require this meeting to make a decision?” The cost is decisions driven by “the biggest title and loudest voice” over the best-informed view. - DevRev’s answer, per Manoj, is a “business-centric knowledge graph” — not a dump where retrieval “magically” works, but a structure built around the only two things he says matter to a company: “the product or services that you provide and the customers you serve.” On top sit a conversational layer and agentic actions that two-way sync back into primary systems via DevRev’s patented Airdrop tech. - Manoj’s cost-collapse claim, with his own hedge intact: over “eighteen months” a million tokens went from “dollars 50 down to what? 12¢ or something,” and for advanced models “more than 99% reduction.” His point isn’t the exact figure — it’s that LLM inference is now nearly free, so the real enterprise challenge is uniting private data, not affording compute. - Manoj’s model for agentic AI is programmable skills: what people do in a meeting are skills you can “write in plain English” — “I have 20 skills” — then stitch across data for end-to-end work. His example is triage: managers weigh tier-one status, customer health (“red or orange”), and strategic value to set P0–P3; capture that judgment as a skill and run it at the scale of 200 reviewers. - On trust, Manoj splits internal from customer-facing AI. Internal users inspect the “chain of thought,” correct it, and “commit” — teaching the machine. Customers “don’t have that luxury,” leaving reinforcement learning plus continuous answer-checking. Yash Sheth adds that humans convey their confidence level, while a machine states every sentence with equal certainty — so even 10% wrong erodes trust fast. - Yash Sheth’s three-year forecast: every enterprise system — Service Cloud, HR, Jira — exposes its own agent, but ROI arrives only when those agents “seamlessly talk to each other.” The missing piece, he says, is a standard for inter-agent communication, including “what is DNS for agents?” — discovery — which is why Galileo and others formed a collective to build an open agent protocol. ### FAQ **What does Manoj Agarwal mean by “information symmetry,” and why does he say it matters inside enterprises?** Manoj’s premise is that the open internet — now accelerated by LLMs — already gives people roughly equal access to public information: he and Yash “can have the same search, get the similar result.” Inside companies it breaks down. Decisions require pulling people into meetings because each holds different data from different tools, systems, and side conversations, so outcomes get driven by “the biggest title and loudest voice.” Manoj argues this asymmetry is a major source of enterprise cost and wasted effort, and that closing it — giving everyone the same information — could remove the need for many meetings entirely. **How does DevRev technically tackle scattered enterprise data, according to Manoj?** Manoj describes a layered approach. First, connect siloed sources — structured systems like Salesforce Service Cloud, Zendesk, Jira, Atlassian, and GitHub, plus unstructured collaboration in Slack — using DevRev’s connectors, which he calls Airdrop technology. Second, organize that data into a business-centric knowledge graph built around the company’s products and customers, with a personalized rather than generic schema. Third, add a conversational AI layer for retrieval. Fourth, make it agentic — take actions and write changes back into primary systems through two-way sync. Manoj notes both Airdrop and the knowledge graph are patented technology DevRev built. **What concrete example did the guests give of turning human judgment into an automated agent?** Manoj uses ticket triage. Managers deciding whether something is P0, P1, P2, or P3 “stare at the data” — which customer it’s from, whether they’re tier one, the health score (“red or orange”), and strategic value even from non-paying accounts they want to win. He argues you can program an expert’s decision style as a “skill” and run it at the scale of 200 triagers. Yash Sheth says Galileo has seen this productionized: a customer getting thousands of tickets daily where expert engineers’ intelligence is baked in so tickets are auto-triaged and next steps carried out by LLMs — “baby steps,” but real and in production. **Why does Manoj distinguish between building trust for internal versus customer-facing AI?** For internal users, Manoj says you can expose the AI’s chain of thought and the linkage showing where information came from. Users inspect it, correct it when they disagree, and “commit” when it’s right — effectively teaching the machine how they think. Customer-facing AI removes that feedback loop: “they’re not the ones who are going to come and train.” There the tools are reinforcement learning — did the answer land or not — plus an internal system continuously checking whether what was given to the customer was correct, which gets much harder for complex, reasoning-heavy queries that demand more trust. **What does Yash Sheth say is missing for a multi-agent enterprise future, and what is being done about it?** Yash predicts that within about three years every enterprise system will expose its own agent, but value only materializes when those agents can “seamlessly talk to each other.” Today, he says, everyone builds agents in silos and the infrastructure between them is missing — both a shared communication standard (an “agent protocol”) and discovery, which he frames as “what is DNS for agents?” He notes other gaps like state passing and authentication. To address this, Galileo, with other companies in the space, formed a collective to build an open standard with everyone contributing — not a closed, single-decision-maker project. ## EP 15: The Agent Bubble Debate | Spot AI's Kelly Vaughn Guests: Kelly Vaughn, Spot AI · Published: 2025-03-19 · Duration: 35 min URL: https://chainofthought.show/podcast/15-the-agent-bubble-debate-spot-ais-kelly-vaughn/ (full transcript on page) Is the agentic AI bubble about to burst? Kelly Vaughn, Director of Engineering at Spot AI, questions whether the agent craze is overpromising potential and leading startups down a path of unsustainable expectations. Never one to shy away from a hot take, Kelly joins host Conor Bronsdon for a pragmatic look at AI, discussing the differences between building AI-enabled and traditional software, why replacing humans with AI teams will backfire (looking at you customer service), and the proliferation of AI tools. ### Key takeaways - Kelly Vaughn’s answer to the title question is unambiguous: “Yes. Agentic AI is a bubble.” But she frames it as a normal cycle, not a fraud — “before AI agents was AI. Before AI, it was blockchain.” Her real warning is that most startups chasing the hype have “no differentiator,” and “first to market might give you a leg up at the outset, but you have to follow through.” - On building AI-enabled products, Vaughn’s core distinction is velocity: AI products demand a fast pace because models are “getting smarter… getting cheaper over time,” so you must keep iterating. A traditional software product is “a little bit more deterministic” and, in her phrase, more “velocity proof” — you don’t have to iterate as fast to stay relevant. - Vaughn names over-promising AI capabilities as the biggest pitfall she has seen building video AI at Spot AI: “you’re never going to see 100% accuracy.” Her stakes example is a security break-in the model fails to detect — “you just missed a very important event,” an easy way to lose customer trust when safety decisions ride on it. Her second pitfall: lacking a fast feedback loop. - On team construction, Vaughn is skeptical of job posts seeking a senior full-stack engineer who is “also an expert at AI” — “congratulations if you can find a unicorn.” She advises hiring real domain expertise on both sides: a data scientist, ML engineer, or AI engineer as a separate role from the full-stack engineers, and testing “in real world situations,” not just a lab. - Vaughn rejects the linear-productivity narrative with a kitchen analogy: people claim AI lets them “double our efficiency,” but “have you ever had four cooks in the kitchen? Does it actually quadruple your productivity? It does not.” AI is “a very good augmentation tool,” not a replacement — and on companies swapping customer-service teams for pure AI, she says, “I’m going to watch them walk back on that one.” - Vaughn’s through-line on AI coding tools: useful but never load-bearing. She came around on tools like Cursor after early inaccuracy, but trains her team that AI “should never become a crutch… you should be able to survive without it.” She separates a prototype from “a minimum viable product and a minimum shippable product,” insisting production code meet a real “security and scalability standard.” ### FAQ **Is agentic AI a bubble?** On Chain of Thought, Spot AI’s Kelly Vaughn answers “yes” — but as a hype cycle, not a scam. She points out the pattern repeats: “before AI agents was AI. Before AI, it was blockchain,” driven by how venture fundraising and Y Combinator theme lists pull in waves of copycat startups. Her concern is differentiation: many of these companies “have no differentiator,” and being first to market only helps if you follow through. She closes the episode noting agentic AI “is not a bad thing” when used to solve specific problems. **How is building an AI product different from building traditional software?** Kelly Vaughn says the biggest difference is velocity. AI-enabled products force a fast pace because models keep changing — “getting smarter… getting cheaper over time” — so you must continuously iterate. A traditional software product, by contrast, is “a little bit more deterministic” and more “velocity proof,” meaning you don’t have to iterate as quickly to stay relevant. She adds that AI products carry heavier user-trust and data-governance burdens, since customers ask what you do with their data and whether you retrain models on it. **Will AI replace engineering or customer-service teams?** No, according to Spot AI’s Kelly Vaughn. She calls AI “a very good augmentation tool,” but says “it’s not going to replace your team or double the size of your team.” On companies replacing customer-service agents entirely with AI, she is blunt: “I’m going to watch them walk back on that one.” Her efficiency analogy: claims that AI will “double our efficiency” ignore that adding capacity is not linear — “have you ever had four cooks in the kitchen? Does it actually quadruple your productivity? It does not.” **How should you structure a team to build an AI product?** Kelly Vaughn advises against hunting for one engineer who is full-stack and “also an expert at AI” — “congratulations if you can find a unicorn.” Instead, hire genuine domain expertise on both sides: a data scientist, ML engineer, or AI engineer as a distinct role from the full-stack engineers who build the user experience. She also stresses testing in real-world conditions rather than a lab, and seating engineers close to customers so the team builds what actually serves their needs. **How should engineers use AI coding tools like Cursor responsibly?** Spot AI’s Kelly Vaughn supports her team adopting AI tools like Cursor — including to automate PR-review work — but with a firm rule: AI “should never become a crutch… you should be able to survive without it.” She warns AI “can write code that is not actually that great,” so engineers must distinguish a prototype from “a minimum viable product and a minimum shippable product,” confirming production code meets a security and scalability standard. Her tell for over-reliance: perfectly commented code prompts her to ask, “tell me what you did to actually confirm that this actually works.” ## EP 14: Using AI to Modernize Your Legacy Applications | MongoDB’s Rachelle Palmer Guests: Rachelle Palmer, MongoDB · Published: 2025-03-12 · Duration: 44 min URL: https://chainofthought.show/podcast/14-using-ai-to-modernize-your-legacy-applications-mongodbs-rachelle-palmer/ (full transcript on page) Imagine cutting your legacy code modernization timeline from years to months. It’s no longer science fiction and this week’s guest is here to tell us how. Rachelle Palmer, Director of Product Management at MongoDB, joins hosts Conor Bronsdon and Atindriyo Sanyal, for a discussion on the groundbreaking ways AI is modernizing legacy applications. At MongoDB, Rachelle's forward-deployed AI engineering team is tackling the challenge of transforming complex, outdated codebases, freeing developers from technical debt. ### Key takeaways - Rachelle runs modernization “incubators” for organizations captive to legacy code — teams that outsourced “the care and feeding” of their codebase to contractors and now find themselves “paying millions of dollars for the maintenance of their own applications” with no in-house knowledge left. - An 80/20 split governs the work: LLMs handle roughly 80% (documentation, comments, equivalence tests, business-logic conversion) while ~20% stays “purely manual” — and the LLM share runs on a machine timescale of “minutes or hours or days” instead of the months or years a manual engineering effort would take. - ROI is measured by old-fashioned benchmarking: a senior or staff engineer does the task manually first, then the same task is run with AI. Modernizations that would take “the magnitude of five plus years” collapse to “months or a year,” and “the team is five people instead of 40.” - High-risk domains demand a different bar. “None of us are doctors… it could say you have tuberculosis and I don’t know” — so you pair engineering with subject-matter experts to proof outputs, fine-tune on domain data, and run “repair loops where we have an LLM play different roles and evaluate the outputs.” - Rachelle’s team built a “chain of repair”: generated code is scored against a custom eval metric and looped — “fix the code, make it better” — until it clears the bar. “No human looks at the code until it hits that score,” and a human only intervenes after a set number of failed iterations. - Forward deployed AI engineering (a model “coined by Palantir”) puts engineers on the ground to build “whatever is useful and works right then as quickly as possible.” It follows a “steel thread” — bones, then meat, then skin, then hair — rather than a PM’s vision of the “fully fledged person two years later.” ### FAQ **How much faster does AI-powered modernization actually make this work?** Modernizations that conventionally run “on the magnitude of five plus years” compress to “months or a year,” executed by “five people instead of 40” — a return on investment Rachelle calls “really, really insane.” **What parts of legacy modernization can LLMs actually handle?** On a million-line codebase, LLMs map the vertical slices of functionality, write missing documentation and comments, and generate equivalence tests to prove a replacement matches the legacy app. That covers about 80% of the work; the remaining ~20% stays purely manual. **What is forward deployed AI engineering?** A model coined by Palantir: rather than a product manager scoping a feature from afar and handing over a “fully fledged person two years later,” you embed engineers with the customer to build whatever is useful and works right then, as fast as possible — starting from a “steel thread” of core functionality and building outward. **How does the “chain of repair” keep AI-generated code quality high?** Generated code is scored against a custom evaluation metric and looped — “fix the code, make it better” — until it clears a set score. No human reviews it until it passes; only after a fixed number of failed iterations does a person step in. Code is never sent to test until it has already cleared the quality bar. **Why should engineers read academic research before building with AI?** Unlike engineers, “the academic community documents everything,” so the field already records how LLMs behave and what works. A day of reading saves lessons learned — for example, papers may show a task’s practical ceiling is around 75% accuracy, so there is no point chasing 100% and the rest will be manual. ## EP 13: AI in 2025: Agents & The Rise of Evaluation-Driven Development Guests: Vikram Chatterji & Andrew Zigler · Published: 2025-03-05 · Duration: 29 min URL: https://chainofthought.show/podcast/13-ai-in-2025-agents-and-the-rise-of-evaluation-driven-development/ (full transcript on page) This week, we're sharing a special episode courtesy of 'Dev Interrupted.' Our co-host, Galileo CEO Vikram Chatterji, recently joined theDev Interrupted team for an engaging discussion on AI strategy. We were so impressed by the conversation that we wanted to share it with our audience, and they were kind enough to let us. We hope you enjoy it! Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Vikram Chatterji argues AI is “another tool in your arsenal,” not a mandate. The right first question isn’t “how do we use AI” but “what’s your use case, and is this a good fit?” — and engineering leaders need the machinery to push back and say “maybe I don’t need to use AI at all.” - Risk tolerance should set the pace. Chatterji contrasts a bank — where a misfiring consumer AI can “derail your entire bank’s reputation” because it deals with people’s money — against a DoorDash or Instacart, where a chatbot slip is “dealing with hungry people,” not bank accounts. Enterprises go crawl-walk-run; digital natives move fast and break things. - Even fast-moving teams need forecasting before they build. Chatterji’s sequence: forecast how many use cases (10, 15), provision enough compute (he name-checks A100 GPUs) so engineers aren’t asking “where’s my GPU at,” invest in tooling to cut compute cost, then add eval guardrails for an “AI CI/CD process” before letting teams loose. - Agents turn the LLM from a generator into a “smart router” that picks tools to complete a task — replacing deterministic code where “you would literally say this is exactly what you need to do” with “I just want this done, you figure it out.” That unlock creates new failure modes: did it pick the right tool, call it correctly, plan the task, and actually complete it? - Agentic apps are “compound systems” — Chatterji credits Databricks’ Matei Zaharia for the term — that look increasingly like classical software engineering built from function calls. Galileo’s Luna evaluation metrics aim to score them without ground truth via a TypeScript/JavaScript SDK, surfacing explainable metrics and automatic insights as “a co-pilot for your AI application development.” - Asked the most valuable skill for engineering leaders right now, Chatterji names two: keep building to stay upskilled — if you expect your team to ship AI apps, build simple apps on the side yourself — and stay plugged into the community of other eng leaders, because “you can’t wait to make those mistakes yourselves” when everything moves this fast. (His interviewer Andrew Zigler condensed it as “build, read, and communicate.”) ### FAQ **How should engineering leaders decide which AI use cases to actually build?** Galileo CEO Vikram Chatterji warns that an open-ended hackathon will generate a hundred ideas given how broad AI is — text, images, agentic task completion. The discipline is to take those back to product and business owners and prioritize on two axes: what the business actually needs, and which use cases you can get out the door very quickly. Then apply operational rigor to try a small number of prioritized ideas fast, rather than drowning in possible solutions. **Why do banks adopt AI more cautiously than companies like DoorDash?** Chatterji frames it as fault tolerance. A bank deals with people’s money, so a misfiring consumer-facing AI can derail the bank’s reputation and put it out of business quickly — pushing enterprises toward a careful crawl-walk-run approach. A DoorDash, Instacart, Airbnb, or Twilio has a lower bar: a chatbot mistake is bad but it’s “dealing with hungry people,” not unauthorized transactions. The stakes set how experimental the company can afford to be. **What does it mean that AI agents act as “smart routers”?** Instead of asking the LLM to just generate text, you wrap functions as tools and let the agent decide which tool to call to finish a task — acting like a smart router. Chatterji contrasts this with old deterministic code where you’d “literally say this is exactly what you need to do.” Now it’s “I just want this done, you figure it out,” creating a leader-worker relationship between the engineer and the agent. **How do you evaluate an AI agent versus a simple chatbot?** A chatbot is query-response, query-response. An agent runs a sequence of actions you don’t directly observe — you only see the final output. Chatterji says evaluation has to open that box: did it choose the right tool, was the tool called correctly, did it plan the task properly, and did it actually complete the task — and how do you even measure quality of completion? Galileo surfaces traces, spans, and explainable metrics so engineers can see how decisions were made and optimize wasteful API calls. **What is a “compound system” in AI engineering?** Chatterji credits Databricks’ Matei Zaharia with coining “compound systems” for what Galileo calls your AI app. As agentic scenarios add function calls, the system grows more complex and starts to resemble classical software engineering — it’s “all about how good are your functions and how well are you managing it all.” That’s why engineers now ask what a unit test or regression test even looks like for an agent, which is where eval comes in. ## EP 12: The Making of Gemini 2.0: DeepMind's Approach to AI Development and Deployment | Logan Kilpatrick Guests: Logan Kilpatrick, Google DeepMind · Published: 2025-02-12 · Duration: 41 min URL: https://chainofthought.show/podcast/12-the-making-of-gemini-2-0-deepminds-approach-to-ai-development-and-deployment-logan-kilpatr/ (full transcript on page) Google’s strength in AI has often seemed to get lost in the midst of OpenAI announcements or DeepSeek fervor - yet Gemini 2.0 is more than good for many tasks; it’s the model to beat - and we have the research to back it up. This week, Logan Kilpatrick, senior product manager at Google DeepMind, joins us to discuss Gemini’s creation story, its emergence as the premiere model in the AI race, and why the launch of Gemini 2.0 is great news for developers. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Flash-Lite holds the line on price; Pro pushes the frontier. Google kept Flash-Lite at “the same price point that we had with 1.5 flash” specifically to “remove the economic burden” for developers putting AI in front of users — “if cost is the barrier for you to build with AI, like, let’s not have that continue to be the case.” The 2.0 Pro model is a continuation of the December Gemini 1206 model — “one of the highest ranked models for coding use cases” — built to “push the frontier of coding and agents.” - Function calling is the foundation of the agentic era. Logan calls tool use “the core enabler of a lot of these agentic workflows” and says the team is “screaming it from the mountaintops internally” that the model has to be great at it. Gemini 2.0 was trained to natively know when to invoke first-party tools — search and code execution — to fix two core LLM limits: no access to “updated world knowledge,” and the “embarrassingly bad mistakes” on problems “you can actually solve with a calculator or by running a little bit of code.” - No weird pricing, and compositional function calling ships with 2.0. Search and code execution are free to try, and at scale “you just pay for the tokens that are created” — “there’s, like, no weird pricing story.” Compositional function calling, new in 2.0 and available in the API, lets developers “describe the sequence and the chain” of functions, because the order in which tools are called “actually matters a lot.” - Multimodal AI is barely in production — and that is the opportunity. “Basically multimodal AI is not in production at this point.” As a former ML engineer, Logan recalls that solving one domain-specific computer-vision problem took “on the order of nine to twelve months in a successful case.” Now segmentation, bounding boxes, and object detection “just work” out of the box, so the founders who were blocked because they lacked “a bunch of, like, ML PhDs” can finally ship vision use cases. - Long context is an infrastructure story as much as an algorithmic one. Gemini works “all the way up to 2,000,000 tokens with the Pro series” — and Logan stresses it is “just as much an algorithmic breakthrough story as it is an infrastructure breakthrough story,” only viable at scale “because of TPUs.” He frames Gemini’s first-mover list: first to ship native search, first to ship caching, first with a native multimodal LLM that could take in video, images, and audio. - Agents will drive a 100-1000x jump in inference compute — which only pencils out at Flash-tier cost. Logan expects “100 to 1000x more inference compute” as agents, not humans, become the ones generating tokens. That future “only … ends up being possible … with models that are at the cost of the Gemini two point zero flash models,” because otherwise the products get too expensive to reach most people. ### FAQ **Is Gemini 2.0 Flash production-ready (as of February 2025)?** Yes. After roughly two months in preview, 2.0 Flash became “fully available for production use with higher rate limits.” As Logan tells it, developers spent that window saying “we love two point zero Flash … let us use it in production,” and the team kept replying “Please, wait” while they put the finishing touches on it before the production rollout. **What is the difference between Gemini 2.0 Flash-Lite and 2.0 Pro?** Flash-Lite is the cost play: it holds “the same price point that we had with 1.5 flash” to “remove the economic burden” for developers shipping AI to users. Pro is the frontier play — a continuation of December’s Gemini 1206 model, “one of the highest ranked models for coding use cases,” built to “push the frontier of coding and agents.” For most workloads, Logan says 2.0 Flash is great and fully multimodal; reach for Pro on harder coding and agentic tasks. **How does Gemini 2.0 handle function calling and tools?** Function calling is “the core enabler of a lot of these agentic workflows,” and 2.0 is trained to natively know when to invoke first-party tools like search and code execution — fixing LLM gaps such as stale world knowledge and “embarrassingly bad mistakes” on calculator-solvable problems. New in 2.0, compositional function calling (available in the API) lets you “describe the sequence and the chain” of functions, because call order “actually matters a lot.” There is “no weird pricing story” — at scale “you just pay for the tokens that are created.” **How does Gemini support long context, and why do TPUs matter?** Long context runs “all the way up to 2,000,000 tokens with the Pro series.” Logan says research breakthroughs enabled it on the algorithmic side, but “the only reason we’re able to put long context into production … at the scale that it is, is because of TPUs” — it is “just as much an algorithmic breakthrough story as it is an infrastructure breakthrough story.” Owning the silicon is also why Google can pass cost savings on to developers. **Why do so many AI products fail in production?** Logan estimates that of 10 teams putting AI in front of users, “at least five out of 10 … would not have a good eval story” — they do not even know what success looks like for the AI they shipped. The “magic silver bullet” framing breaks down because real deployments need infrastructure that often did not exist: eval infrastructure and ops dashboarding for things like A/B testing prompts. It takes work, and the foundation has to be built underneath the product. ## EP 11: How DeepSeek Changed the AI Race Overnight Guests: Atindriyo Sanyal, Galileo · Published: 2025-02-05 · Duration: 33 min URL: https://chainofthought.show/podcast/11-how-deepseek-changed-the-ai-race-overnight/ (full transcript on page) This week, hosts Conor Bronsdon and Atindriyo Sanyal discuss the fallout from DeepSeek's groundbreaking R1 model, its impact on the open-source AI landscape, and how its release will impact model development moving forward. They also discuss what effect (if any) export controls have had on AI innovation and whether we’re witnessing the rise of “Agents as a Service”. ### Key takeaways - Atin argued R1 was no secret sauce: DeepSeek leaned on the standard methodologies already used to train the larger OpenAI and Anthropic models, pushing distillation and reinforcement learning hard enough to train a very large model at “like 95% less cost.” The novelty, he said, was efficiency — not some technique no one had seen before. A very large base model still sat underneath it all (the V3 base, trained on 15 trillion tokens), so the cost collapse rode on the scaling work the giants had already done. - On the cost curve, Atin predicted that within “twelve to eighteen months, someone should be able to train a similar model just on their laptop.” He pointed to cheaper models already built on top of DeepSeek — including, by his own double hedge, a Berkeley student project he “already heard” about that “what I heard was $500” — as the start of a pattern of ever-cheaper foundation models. - Atin pushed back on the hype that R1 was “some kind of all knowing all encompassing thing.” He framed it as a reasoning experiment: reinforcement learning plus auto-generated chain-of-thought data to produce high-quality reasoning models — which, by design, left it “not that great at some non reasoning based tasks.” - On policy, Atin said export controls “just stifle innovation.” He described AI progress as “symbiotic” across China, India, Europe, and the US — citing how work at Stanford leans on a Tsinghua reference that leans on an Indian university — and argued innovation and technology should be kept separate from political discussions. - On agent evaluation, Atin stressed there is “no one metric that can tell you agent quality.” An agent is a directed acyclic graph of LLM calls, tools, and vector lookups, and “one mistake in any part of this workflow” compounds downstream. His answer paired out-of-the-box action-advancement and task-completion metrics with custom ones teams define themselves. - Atin said the customizations teams actually build are “very simplistic functions” — small Python, Golang, or TypeScript checks needing “little to no compute” and no GPUs. The real value, in his telling, is the management layer that lets a team register a metric, share it, and track its lineage as it evolves. ### FAQ **How much cheaper did DeepSeek claim to train R1?** Atin put the saving at “like 95% less cost” versus earlier techniques, achieved by leaning harder on distillation and reinforcement learning. He was careful to note a very large base model still sat underneath — the V3 base, trained on 15 trillion tokens — so the savings built on the giants’ prior scaling work rather than replacing it. **Did DeepSeek invent new techniques, or reuse existing ones?** Existing ones, per Atin. He said “the methodology is not anything that’s new,” describing the same standard methodologies used to train the larger OpenAI and Anthropic models. The innovation, he argued, was in how DeepSeek put the building blocks together to hit very low cost — efficiency, not a brand-new recipe. **What was Atin’s view on export controls?** He took what he called a simplistic take: these kinds of controls “just stifle innovation.” He saw AI as a symbiotic, intricately connected system across countries and argued that innovation and technology should be kept separate from political discussions, warning the US would impose a “massive penalty” on itself with such restrictions. **Why is there no single metric for agent quality?** Because, Atin explained, an agent is a directed acyclic graph of components — LLM calls, tool and function calls, vector-store lookups — and “one mistake in any part of this workflow” compounds into large downstream errors. Quality depends on which components a given agent has, so the right metrics vary system to system. **What kinds of custom evals do teams actually build?** Mostly “very simplistic functions,” Atin said — small Python, Golang, or TypeScript checks like whether a string appears in a blob of text or whether a graph was sorted correctly. They need “little to no compute” and no GPUs; the harder, more valuable part is the management layer for registering, sharing, and tracking the lineage of those metrics. ## EP 10: AI, Open Source & Developer Safety | Block’s Rizel Scarlett Guests: Rizel Scarlett, Block · Published: 2025-01-29 · Duration: 34 min URL: https://chainofthought.show/podcast/10-ai-open-source-and-developer-safety-blocks-rizel-scarlett/ (full transcript on page) As DeepSeek so aptly demonstrated, AI doesn’t need to be closed source to be successful. This week, Rizel Scarlett, a Staff Developer Advocate at Block, joins host Conor Bronsdon to discuss the intersections between AI, open source, and developer advocacy. Rizel shares her journey into the world of AI, her passion for empowering developers, and her work on Block's new AI initiative, Goose , an on-machine developer agent designed to automate engineering tasks and enhance productivity. ### Key takeaways - AI as psychological safety for juniors: Rizel reframes the “AI replaces junior developers” fear — “if I had this when I was a junior developer, this would have made me so much more confident.” As a young Black woman engineer she was scared to ask questions and “come across as stupid”; AI lets you go back and forth and gain confidence, working “like a peer programmer.” - What Goose is: Block’s open-source, on-machine developer agent. It has a Rust back end and an Electron front end, and you can drive it through the terminal or the Electron app. It is extensible — you can wire it up to your file system and other tools — and it leverages Model Context Protocol servers. Block collaborated with Anthropic on building out MCPs. - Goose was built first for Block employees — starting with migrations. As Rizel put it, the goal was “to automate and make sure employees are not doing work about work” so the team “can just get straight to the impact,” before the project was opened up to the broader community. - Productivity is more than coding faster: developers are also parents and open-source maintainers. “We’re not just these people that are heads down at their computer coding all the time.” The real win is helping people keep context as they constantly switch back and forth, and show up as their best self across all their responsibilities. - Open source gave Rizel influence that closed tools did not. With GitHub Copilot she “couldn’t have as much influence as I wanted to,” even as a developer advocate; at Block she can push her own code to improve the tool — and so can community members, who bring in their own tools and MCP connections instead of filing requests. - Responsible use means reviewing AI output, not rubber-stamping it: “don’t just check-in AI code, review it.” Rizel argues for a dedicated “how to code with AI” training alongside the usual security and harassment onboarding — keep writing and manually testing your code, and learn “how hallucinations work and how to kind of reduce them.” ### FAQ **What is Goose?** Goose is Block’s open-source, on-machine developer agent for automating engineering tasks. It has a Rust back end and an Electron front end, runs from the terminal or as an Electron app, and is extensible through Model Context Protocol servers. You can learn more at github.com/block/goose. **How can AI coding tools create psychological safety for junior developers?** Rizel argues AI gives newer engineers a low-stakes way to ask questions and learn without fear of looking “stupid” — something she felt acutely as a young Black woman early in her career. Instead of replacing juniors, an AI tool acts “like a peer programmer” you can go back and forth with, so you show up on your team more confident. **Why did Block build Goose as open source?** Open source lets the community get involved and contribute code directly rather than just filing feedback — Rizel can push her own improvements, and so can outside contributors. She also frames it as “authentic and free marketing”: contributors who feel part of the project talk about it at conferences because they genuinely care about it. **How can teams use AI tools responsibly?** Don’t just check in AI-generated code — review it the same way you would review a coworker’s code. Rizel suggests a dedicated “how to code with AI” training alongside standard security and harassment onboarding, and stresses continuing to write and manually test your code and understanding how hallucinations work so you can reduce them. **Did Block collaborate with Anthropic on MCP?** Yes. As one of Block’s open-source AI initiatives, the team collaborated with Anthropic on building out MCPs (Model Context Protocol servers), and Goose leverages MCP for its extensibility. ## EP 9: AI in 2025: Agents & The Rise of Evaluation Driven Development Guests: Yash Sheth & Atindriyo Sanyal · Published: 2025-01-15 · Duration: 33 min URL: https://chainofthought.show/podcast/9-ai-in-2025-agents-and-the-rise-of-evaluation-driven-development/ (full transcript on page) "In the next three to five years, every piece of software that is built on this planet will have some sort of AI baked into it." - Atin Sanyal Chain of Thought is back for its second season, and this episode dives headfirst into the possibilities AI holds for 2025 and beyond. Join Conor Bronsdon as he chats with Galileo co-founders Yash Sheth (COO) and Atindriyo Sanyal (CTO) about major trends to look for this year. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Atin Sanyal’s hot take: “in the next three to five years, every piece of software that is built on this planet will have some sort of AI baked into it.” He argues AI software development is “back to square zero” — “there’s no eval tooling,” and people are using “caveman tools” to “just look at vibe checks and eyeballs.” Conor pushes back that it’ll take longer and bets Atin a meal that he is wrong. - Yash Sheth’s one-word theme for 2025: “automation.” He argues the real ROI comes not from making software more conversational but from automating workflows across industries — LLMs that make API calls, execute code, and process multi-modal data (images, audio). - Atin notes that “25 percent of Google’s code is now code gen automated,” but says under a magnifying glass most of it is boilerplate — “a couple of for loops” and basic things. His excitement for 2025 is moving past boilerplate toward “better code understanding” that grasps the context around the code. - Atin argues code writing “is just probably 10 percent of a software engineer’s job” — design and other work make up the rest — so a human (or developer) in the loop stays needed. Comparing the path to autonomous driving, he sees engineers freed up to focus on “connecting boxes with arrows and building awesome systems.” - Yash explains that because agents make API calls, tool calls, and code execution, they cause “irreversible changes,” which raises the penalty for a wrong action. He describes Galileo Protect as a “control pane” for a generative AI application — likening it to a firewall — that aims to detect bad behavior and prevent it “within seconds.” - On cost and scale, Atin claims Galileo has “the only hallucination detection model and algorithm that literally works at zero dollars,” at low latency across “both P50 and P95.” Yash frames the constraint bluntly: “no one wants to double their OpenAI bills,” and no one wants an evaluation system that “takes 10 seconds to evaluate one prompt” when traffic runs to “thousands of QPS.” ### FAQ **What did Yash and Atin predict for AI in 2025?** Yash Sheth’s one-word prediction was “automation” — real ROI coming from automating workflows across industries rather than just conversational software. Atin Sanyal predicted AI would find two kinds of fit: product market fit (real user and business benefit) and “product tool stack fit,” as the engineering and systems around the LLM mature. On code, Atin expected progress beyond boilerplate generation toward “better code understanding.” **Why do AI agents need stronger evaluation than chatbots?** Atin and Yash argue agents don’t just return text — they take actions through API calls, tool calling, and code execution, which means “the penalty you pay for the right action versus the wrong action is potentially much higher.” Because agents can make “irreversible changes,” they say you need robust evaluation that runs at scale and enforces good-behavior metrics in real time within an agentic flow. **What is Galileo Protect?** Yash describes Galileo Protect as a “control pane” for a generative AI application — he compares it to a firewall for security threats. As he frames it, the idea is to detect bad behavior, create a metric to catch it, and prevent the application from acting on it “within seconds.” These are the speakers’ descriptions of their own product, not independent claims. **How can AI evaluations be cheaper and more scalable?** Atin claims Galileo offers “the only hallucination detection model and algorithm that literally works at zero dollars,” with low latency across “both P50 and P95” so it works past a certain throughput. Yash frames the cost constraint directly: pushing past LLM-as-judge toward methods that scale to “thousands of QPS” at a price point that works, because “no one wants to double their OpenAI bills” or run an evaluation that “takes 10 seconds to evaluate one prompt.” These are the speakers’ claims about their own product. **What advice did they give businesses adopting AI in 2025?** Yash’s advice was to “start quantifying the behavior of your application into metrics early on,” establishing rigor before scaling rather than shipping a cool POC without it. Atin first joked “Get Galileo for your LLM evaluation needs,” then seconded Yash on rigor — comparing the moment to cloud adoption a decade ago, where security was the missing unlock, and arguing evaluation is the equivalent unlock for AI. ## EP 8: AI Infrastructure & the Evolution of RAG | Weaviate's Bob van Luijt Guests: Bob van Luijt, Weaviate · Published: 2025-01-08 · Duration: 35 min URL: https://chainofthought.show/podcast/8-ai-infrastructure-and-the-evolution-of-rag-weaviates-bob-van-luijt/ (full transcript on page) "This is the time. This is the time to start building... I can't say that often enough. This is the time." - Bob van Luijt Join Bob van Luijt, CEO and co-founder of Weaviate as he sits down with our host Conor Bronsdon for the Season 2 premiere of Chain of Thought. Together, they explore the ever-evolving world of AI infrastructure and the evolution of Retrieval-Augmented Generation (RAG) architecture. Bob's journey with Weaviate offers a compelling example of how to adapt to rapid changes in the AI landscape. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Weaviate had RAG before the term existed — they called it “generative search.” When the research-paper term caught on, they simply adopted the language: “in the product itself, nothing changed. It was just how we talked about it.” - Bob frames the arc in three steps: RAG is “one directional” (query → retrieve → generate), agents add a feedback loop, and generative feedback loops — his favorite — let the agent “put stuff back inside the database.” - The killer use case is master data management. Pre-Weaviate, consulting for a global candy manufacturer, Bob watched humans hand-reconcile messy data (Fahrenheit vs. Celsius, British vs. American spelling) — losing, because “data is coming in faster than these people could handle it.” Generative feedback loops let a model do that reconciliation instead. - 2024 was “the year of going to production.” Indexes that sit in memory get expensive at scale, so Weaviate added storage tiers (memory, disk, S3) at “a fraction of the price” — the speed penalty is “still in the milliseconds.” Before tiers, customers could “do the POC” but “can’t make a business case.” - Detailed documentation builds trust, against the minimal-docs orthodoxy. It is a Catch-22 — show everything and risk overwhelming, or stay lean and look shallow. Bob’s take: “if I see that there’s detailed documentation… I know when I scale this up, they can handle this.” - In “full transparency,” 95% of Weaviate customers still sit in the first bucket (search and recommendation). The remaining gap is psychological, not technical — a paradigm shift: “we know we can do it… it’s now people in these businesses to accept it.” ### FAQ **Where does the name “Weaviate” come from?** It comes from “weaving the models and the database together” — the core idea behind combining vector storage with generative models in one AI-native system. **What are generative feedback loops?** They are the third step beyond RAG and agents. Instead of only retrieving data, the agent can write back into the database — validating, correcting, and improving stored data. Bob points to master data management as the prime example: hand the model your data standards and let it flag or fix wrong objects (like a temperature stored in Fahrenheit that should be Celsius) instead of relying on humans in the loop. **Why did Weaviate introduce storage tiers?** As use cases grew into production, keeping every index in memory became too expensive to justify a business case. Storage tiers let you keep data in memory, or move it to disk or S3 at “a fraction of the price.” The trade-off is a small speed penalty that is “still in the milliseconds” — which suits high-tenant RAG workloads like document search and email clients. **Is now actually the time to build?** Bob is emphatic: “this is the time to start building.” As an angel investor in “like two dozen companies or something,” he sees real revenue, not just open-source downloads. He hedges that it is “not a silver bullet,” but argues that solving even 80% of a problem makes a lot of people very happy — so it is not the time to be too skeptical. **Why is chunking PDFs so hard?** A single page often lacks the context needed to answer a query. Bob’s example: a PDF about Apple’s Q2 2024 revenue might only be identifiable as Apple by the logo at the top, with no text naming the company alongside the numbers. To handle this you have to represent pages with multiple embeddings, store them all, and retrieve them in a way that fits the specific use case — which gets elaborate quickly. ## EP 7: Beyond Chatbots: How Twilio Uses AI to Strengthen Human Connection | Vinnie Giarrusso Guests: Vinnie Giarrusso, Twilio · Published: 2024-12-18 · Duration: 42 min URL: https://chainofthought.show/podcast/7-beyond-chatbots-how-twilio-uses-ai-to-strengthen-human-connection-vinnie-giarrusso/ (full transcript on page) Can AI assistants actually enhance human connection? As Season 1 of Chain of Thought comes to a close, host Conor Bronsdon and Vinnie Giarrusso (Twilio) explore the transformative potential of AI assistants in the workplace. Discover how these assistants function as "async junior digital employees," taking on specific tasks and contributing to the organizational structure. But will AI assistants ultimately replace human connection? ### Key takeaways - Twilio deliberately calls them “AI assistants,” not agents — the framing is that, especially early on, these systems play a “side by side” role with humans, augmenting people with “superhuman powers” while a human stays in the loop to accept or reject the assistant’s actions. - The product is a low-code autonomous agent platform: developers get tools, knowledge sources (document upload and web-page crawling), managed memory, and an API layer for deployment and invocation — with Twilio’s native channels (SMS, voice, WhatsApp, conversations) integrated so an assistant can actually reach customers. - RAG retrieval is the biggest failure mode at scale — it is easy to stand up a pipeline from the LangChain docs, but once you dump 10, 15, or 20 knowledge sources in, the assistant may pick the wrong source or return weak top-K chunks; context-aware retrieval (informed by Anthropic’s contextual research) has been “a game changer,” though Vinnie declined to give hard numbers on the accuracy lift. - Twilio has used Galileo Observe since “day minus one” as a core observability layer. Vinnie’s metaphor: Galileo is the “head chef” who knows everything happening in the kitchen — the Luna metrics (context adherence, complexity, accuracy, tone) are what the team is constantly checking. - Twilio open-sourced an evals package under its Twilio Alpha GitHub and NPM namespace — it began as security evals and broadened into evals generally — on the logic that it is “a lot more convincing when you put it in the hands of people to play with themselves.” - Vinnie is most excited about AI in education and the “democratization of information” — coming from a non-traditional computer-science background, he learned from thick textbooks largely on his own, and sees personalized AI tutoring as a way to make that kind of learning far more accessible. ### FAQ **What is Twilio’s AI assistant platform?** It is a low-code autonomous agent platform. Developers can give an assistant tools and knowledge sources (uploading documents or crawling web pages), with memory handled for them, all behind an API layer for deployment, uploads, and interaction. Because it sits on top of Twilio, assistants natively reach customers over existing channels like SMS, voice, WhatsApp, and conversations. **Why does Twilio call them “assistants” instead of “agents”?** It was an intentional choice. Vinnie notes the industry term is “agents,” but Twilio leans toward “assistants” because, especially early on, these systems play a side-by-side role with humans — augmenting people with “superhuman powers” rather than replacing them, often with a human tasked with accepting or rejecting the assistant’s actions. **What are the biggest failure modes Twilio sees in production?** The RAG retrieval pipeline is the hardest part. A basic pipeline is easy to spin up, but as you add more knowledge sources and unanticipated user questions, the assistant may select the wrong knowledge source or return a poor set of top-K chunks. Twilio leans on observability traces to debug this and on context-aware retrieval — informed by recent contextual-retrieval research — to improve accuracy. **How does Twilio use Galileo?** Primarily Galileo Observe, integrated since “day minus one.” Vinnie likens it to the “head chef” who knows everything happening in the kitchen and can step into any role — the team lives in the Luna metrics (context adherence, complexity, accuracy, tone) to find and fix issues. Crucially, it scales across the whole team: engineers, PMs, and business partners all have a reason to go in and can find what they need. **Will AI assistants replace jobs and mentorship?** Vinnie argues the opposite for customer experience — by handling the boring, technical, box-ticking work, assistants free humans to be more empathetic and present in the interactions that genuinely need a person. On mentorship he is more cautious: citing his colleague Dominic Condell, he warns the senior-to-junior mentorship relationship could quietly slip away, even as AI gives juniors an unprecedented personalized tutor. His advice to juniors is to stand side by side with these tools early to open doors — while teams stay intentional about preserving real human mentorship. ## EP 6: The Enterprise AI Deployment Playbook | ServiceTitan, Indeed & Twilio Guests: Mehmet Murat Ezbiderli, Vinnie Giarrusso, Grant Ledford & Atindriyo Sanyal · Published: 2024-12-11 · Duration: 51 min URL: https://chainofthought.show/podcast/6-the-enterprise-ai-deployment-playbook-servicetitan-indeed-and-twilio/ (full transcript on page) This week, a panel of experts (Mehmet Murat Ezbiderli, ServiceTitan; Grant Ledford, Indeed; and Vinnie Giarrusso, Twilio) join Atin Sanyal (CTO, Galileo) and host Conor Bronsdon (Developer Awareness, Galileo) to explore the challenges and opportunities of deploying GenAI at enterprise scale in a conversation that's a wake-up call for any business leader looking to harness the power of AI. ### Key takeaways - Mehmet (ServiceTitan): “Second Chance Leads” reassesses conversations a CSR marked as a non-lead — if the system decides the call should have booked a job, it triggers a callback within about five minutes. Mehmet calls it a direct ROI generator because it recovers a lead that was already missed. - Mehmet (ServiceTitan): for an actual virtual assistant, the LLM cost was “about at least half” of the operational overhead (hedged — “I don’t know the exact numbers” — and inclusive of all compute). On model-to-task fit, GPT-3.5 works great for summarization or sentiment detection, while GPT-4.0 is reserved for reasoning-heavy work like virtual assistants. - Mehmet (ServiceTitan): large context “defeats RAG in pretty much all” contextual Q&A, but you pay for it on every request. Emerging self-route / staggered approaches first ask the LLM whether it can answer from the RAG window, and only fall back to large context when it says it cannot. - Vinnie (Twilio): stack by team familiarity, not dogma. A TypeScript POC built on Vercel stayed TypeScript when it became real, and the LLM chain runs LangChain in TypeScript. Chunking and embedding were pulled into a separate Python service (“faster, better” tooling there), and the public API in front of everything is written in Go. - Vinnie (Twilio): multimodal eval is still unsolved — prompt injection is not solved even for text, and speech-to-speech (the model OpenAI built with Twilio) has few good testing options. The team starts task-based evals “at the ends” before merging into the harder middle of streaming and interrupts. - Grant (Indeed): the builder audience has broadened well beyond a traditional data-science background, so the tooling has to accommodate more context. There is no clean one-to-one mapping from a unit test to an LLM test, so the real measure is whether the application moves a business metric you care about — validated through tight feedback loops with early adopters. ### FAQ **What are “Second Chance Leads” at ServiceTitan?** Mehmet describes them as a reassessment of a customer/CSR conversation: when a CSR marks a call as a non-lead, the system re-evaluates it, and if it judges the call should have booked a job it prompts a callback within about five minutes. He frames it as an ROI generator because it recovers a lead that was already missed, with positive ROI feedback from customers. **How much of operational cost is the LLM, and how should you match models to tasks?** Mehmet estimates the LLM was “about at least half” of the operational overhead for a virtual assistant — hedged with “I don’t know the exact numbers” and inclusive of all compute. He matches models to tasks: GPT-3.5 handles summarization and sentiment detection well, while GPT-4.0 is reserved for reasoning-heavy work such as virtual assistants. **What language and tools should you build your GenAI stack in?** Vinnie’s advice is to use what your team is most comfortable with. Twilio kept their TypeScript POC (originally built on Vercel) in TypeScript because more of their engineers knew it, and the LLM chain stays LangChain in TypeScript. They split out chunking and embedding into a separate Python service where the tooling is faster and better, and front everything with a Go public API. **Why is multimodal evaluation so hard?** Vinnie notes the evals often do not exist yet — prompt injection is unsolved even for text, and speech-to-speech (the model OpenAI built with Twilio) has few good testing options. Holding the same level of rigor across text, voice, and video is difficult, so the team starts with task-based evals “at the ends” before tackling the harder middle of streaming, interrupts, and partial responses. **How do you measure whether a GenAI application is successful?** Grant frames it as giving builders the right levers — model size and latency tradeoffs, long context versus a RAG-based solution, and prompt caching — and then measuring the application as a whole, not just the model. The real test is whether it moves a business metric you care about, which his team validates through tight feedback loops with early adopters. ## EP 5: Practical Lessons for GenAI Evals | Chip Huyen & Vivienne Zhang Guests: Chip Huyen & Vivienne Zhang · Published: 2024-12-04 · Duration: 48 min URL: https://chainofthought.show/podcast/5-practical-lessons-for-genai-evals-chip-huyen-and-vivienne-zhang/ (full transcript on page) As AI agents and multimodal models become more prevalent, understanding how to evaluate GenAI is no longer optional – it's essential. Generative AI introduces new complexities in assessment compared to traditional software, and this week on Chain of Thought we’re joined by Chip Huyen (Storyteller, Tép Studio), Vivienne Zhang (Senior Product Manager, Generative AI Software, Nvidia) for a discussion on AI evaluation best practices. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - The harder the task, the harder the eval. Chip Huyen’s core point: “almost everyone can tell whether the solution to a first grade math problem is wrong, but it’s really hard to tell whether the solution to a PhD math level question is wrong.” As AI handles expert work, the pool of people qualified to judge its output shrinks — Vivienne Zhang notes NVIDIA’s ChipNeMo co-pilot answers chip-design questions she, without a hardware background, can’t even parse, let alone grade. - Most teams ship on a vibe check — and that’s the gap. Vivienne Zhang says curating a benchmark dataset and measuring against it “sets you probably ahead of the big crowd already, because a lot of application developers also just do a vibe check and then send it to production.” The fix is treating eval as a loop: capture user feedback and usage data post-launch, then systematically fix what tripped up the LLM. - “Coherence” and “relevancy” are not exact metrics. Chip Huyen’s #1 observed mistake is teams not defining clear evaluation guidelines, then treating LLM-judge scores as fixed numbers when they’re “highly dependent on the models you use, on the prompt you use for the judge.” She’s found typos in nearly every open-source eval prompt she’s inspected; the same metric can swing 90% from one day to the next as prompts change. - A correct answer isn’t a good answer. Chip Huyen cites LinkedIn’s job-fit bot: asked “am I a good fit for this role,” it would answer “No, you’re a terrible fit.” Factually correct, useless to the user. Teams skip the hard work of writing good-vs-bad guidelines and hope the AI judge figures it out — “it doesn’t quite go that way.” - Plausible-but-wrong is the most dangerous hallucination. Vivienne Zhang’s working definition: an answer that “looks correct, but it’s not what the user is looking for.” Her example (from Jensen Huang’s GTC talk): ask ChipNeMo “what is CTL?” and the generally-accepted answer is “combinational time logic,” but the right NVIDIA-specific answer is “Compute Trace Library.” In legal, healthcare, or finance, that insidious confidence is worse than an obviously wrong reply. - Evaluate agents at the points where they fail. Chip Huyen breaks agent eval into three checks: can it accomplish the task within constraints (plan a 2-week trip on a $5,000 budget), does it call the right tool with the right parameters, and is the plan efficient rather than looping forever. Her method: map how the application is going to fail, then put evaluation metrics around those failure points. ### FAQ **Why is evaluating generative AI harder than traditional machine learning?** Two reasons, per Chip Huyen. First, intelligence outpaces judgment: a first-grade math answer is easy to grade, a PhD-level one isn’t, and the smarter the system the fewer people qualified to assess it. Vivienne Zhang adds NVIDIA’s ChipNeMo example — she can’t even parse the chip-design questions, let alone score them. Second, outputs are open-ended: classification has a clear right answer, but judging a book summary may require reading the whole book. **What’s wrong with using LLM-as-a-judge metrics like coherence and relevancy?** Chip Huyen warns that teams treat these as exact, fixed metrics when they’re anything but. A judge score depends heavily on which model and which prompt you use, and the people who maintain the metric prompts are often different from the people using them. She’s found typos in nearly every open-source eval prompt she’s reviewed, and scores can shift 90% from one day to the next as prompts change. Define your guidelines explicitly rather than trusting a number. **Are hallucinations always bad?** No, says Chip Huyen — she’s “very into creative use cases,” where hallucination is a feature. It’s only bad for use cases that depend on factual consistency. The key insight: models are more likely to hallucinate when they lack access to the necessary, correct information, which is why RAG matters — it gives the model relevant context. Vivienne Zhang adds a practical definition: an answer that looks correct but isn’t what the user needs is itself a hallucination. **How do you evaluate hallucinations in a RAG system?** Chip Huyen frames it as natural language inference (textual entailment): given the retrieved context and the model’s answer, can the answer be derived from — and stay consistent with — that context? You evaluate two things: whether the retrieved context is relevant and correct to the question, and whether the answer follows from that context. Atin Sanyal notes Galileo’s Luna hallucination-detection models are built on this same entailment principle. **What should you measure when evaluating AI agents?** Chip Huyen names three things: whether the agent can accomplish the task within its constraints (e.g., plan a two-week trip on a $5,000 budget), whether it calls the right tool with valid parameters, and whether its sequence of actions is efficient rather than looping endlessly. Her general method is to break down how the application is likely to fail and place evaluation metrics around those failure points. Vivienne Zhang adds that human-in-the-loop escalation — like a customer-service bot handing off when it can’t handle a query — is itself a failure point worth evaluating. ## EP 4: Why Most Enterprise AI Projects Fail to Show ROI | HP, ServiceNow & Accenture Guests: Vikram Chatterji, Alex Klug, Sriram Palapudi & Jay Subrahmonia · Published: 2024-11-27 · Duration: 41 min URL: https://chainofthought.show/podcast/4-why-most-enterprise-ai-projects-fail-to-show-roi-hp-servicenow-and-accenture/ (full transcript on page) The “ROI of AI” has been marketed as a panacea, a near-magical solution to all business problems. Following that promise, many companies have invested heavily in AI over the past year and are now asking themselves, “What is the return on my AI investment?” Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - The enterprise “ROI of AI” panic is a correction, not a dead end. Vikram Chatterji frames it as a “trough of disillusionment”: enterprises that overpaid for expensive frontier models 8–10 months ago, spending “maybe 50 million dollars” on a couple of applications, are now whiplashed when the returns lag — but “we’re still in the first innings.” - ROI from enterprise AI rarely shows up as new revenue. Chatterji notes there is no “net new AI-powered bank” being built — the value is cost savings and operational efficiency, measured over six to eight months by comparing human-capital cost before and after and how many humans-in-the-loop a workflow still needs. - Trust gates deployment at HP. Alex Klug: “the worst thing that can happen is we are in the news about having deployed a solution that has steered our customers wrong.” His CDO puts it bluntly — “you guys are building rocket ships, I want a braking system” — and that missing explainability is what keeps CIOs from moving past one or two pilots. - The gap between intent and execution is stark. Jay Subrahmonia: over 95% of CEOs say their responsible-AI principles are documented, but the share who have operationalized them is “well below 10%,” and fewer than 10% of POCs ever reach production. - Concrete numbers make the case in regulated industries. In insurance underwriting, Subrahmonia cites a 4–5% improvement in the ability to predict risk — and notes underwriters today read documents at a deep level of detail only about 25–30% of the time, a scope that AI can expand. - ServiceNow optimizes a three-way trade-off — number of models, model size, and quality — all bounded by limited infrastructure. Sriram Palapudi measures returns through ticket deflection and how many users solve problems on their own, plus summarization that cuts the “cognitive burden” of reading ticket after ticket. ### FAQ **Why are so many enterprise AI projects failing to show ROI?** Chatterji argues it is a market correction: budgets were set early, GPUs and expensive models and consultants were bought up front, and the anticipated value outran the real value. Subrahmonia adds the execution gap — fewer than 10% of POCs make it into production, usually for lack of governance, testing, and monitoring. **How should companies actually measure AI ROI?** Chatterji recommends counting, over six to eight months, the human-capital cost of a workflow before and after, the time each task took, and how many humans-in-the-loop remain. Palapudi measures it through data — ticket deflection, how many users self-serve, and intangibles like reduced cognitive burden from summarization. **How do you pick the right AI use cases?** Chatterji says the test has not changed in a decade: is the task human-heavy today, is it repeatable rather than full of edge cases, and does automating it deliver real operational-cost savings. The smarter move is a rubric that narrows hundreds of candidate ideas down to the four or five highest-ROI ones, rather than chasing a hundred at once. **How do you build trust and explainability into enterprise AI?** Klug stresses visibility into where an answer comes from and the ability to catch bad responses before they cause brand harm. Chatterji lays out four foundations: deep logging of every step, the right guardrails and metrics and assertions, high-quality test sets with ground-truth responses, and humans in the loop for the inevitable edge cases. **What blocks POCs from reaching production?** Subrahmonia points to the 95%-documented-but-under-10%-operationalized gap and prescribes a risk-based approach so low-risk use cases are not bottlenecked. It takes a cross-functional team — legal, compliance, data science, infosec, procurement, HR — and works best by extending existing MLOps processes and controls rather than inventing a whole new one. ## EP 3: GenAI Predictions for 2025 | Databricks & Cohere Guests: Sara Hooker, Craig Wiley & Vikram Chatterji · Published: 2024-11-20 · Duration: 40 min URL: https://chainofthought.show/podcast/3-genai-predictions-for-2025-databricks-and-cohere/ (full transcript on page) Will 2025 be the year open-source LLMs catch up with their closed-source rivals? Will an established set of best practices for evaluating AI emerge? This week on Chain of Thought, we break out the crystal ball and give our biggest AI predictions for 2025. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Vikram Chatterji predicted that by the end of 2025 open-source LLMs would catch up to the best closed-source models — naming Anthropic’s Claude models as “super expensive… concept models” showing “the art of the possible” — for large context windows on a majority of tasks, citing Galileo’s Hallucination Index as evidence open source was closing the gap fast. - Chatterji predicted 5-to-10 of what were dismissed in 2024 as “wrapper apps” would each hit $100M in ARR by the end of 2025, arguing AI breaks the old mobile-app adoption curve where even DoorDash or WhatsApp took years to reach $100M in revenue. - Craig Wiley framed 2025 as “an era of toolification”: rather than “boil the ocean” with one single my-company-GPT.com, teams deploy narrow, purpose-built tools and smaller agents and orchestrate them together — a model tuned for a narrow capability beats one giant bot, and function-calling usage now signals a team’s AI maturity. - Wiley described 2024 as the year RAG disillusionment hit: customers arrived with “one-click RAG POCs” that dazzled in the vendor demo but failed at scale, learning that “prompt plus some docs doesn’t equal a system I can bet my company’s reputation on.” Production systems are compound, agentic, and need heavy tuning, golden datasets, and LLM judges. - Sara Hooker’s hot take was a dual trend: models would get bigger as people “push out the boat” on multimodal, and simultaneously a lot smaller, as teams better leverage optimization across inference and stratified data. She also argued multimodal inputs help calibrate uncertainty the way audio, visual, and motion let humans sense when something’s off. - Wiley’s wish for 2025 was unglamorous: solve PDFs. Despite excitement about video gen and live voice, he said “collectively we’re not parsing PDFs all that well,” and he’d be thrilled if insurance companies could simply parse claims docs — hoping to “exit 2025 a human understanding of PDFs.” ### FAQ **Will open-source LLMs catch up to closed-source models like GPT and Claude in 2025?** That was Vikram Chatterji’s headline prediction on this Chain of Thought panel. The Galileo CEO said that by the end of 2025 open-source LLMs would catch up to the best closed-source models of late 2024 — he named Anthropic’s Claude models, which he called “super expensive” concept models — for fairly large context windows across a majority of tasks. His evidence was Galileo’s Hallucination Index, which he said showed open source “catching up really fast.” He framed it as a win for developers who could fine-tune an open model instead of paying “an arm and a leg.” **What did the panel mean by the “era of toolification”?** Craig Wiley of Databricks (Mosaic AI) coined it as his top 2025 prediction. Instead of building one broad “my-company-GPT.com” bot that underwhelms everywhere, teams give a model purpose-built tools and specialized smaller models to call via function calling, then orchestrate them. Wiley argued a model tuned for a narrow set of capabilities outperforms a giant general one, and that how aggressively a team uses function calling now signals its level of AI maturity. He called toolification one of the most promising paths to a stepwise jump in system accuracy. **Why did so many RAG proof-of-concepts fail to reach production in 2024?** Craig Wiley said 2024 was the year teams learned that “prompt plus some docs doesn’t equal a system I can bet my company’s reputation on.” Customers showed up with one-click RAG POCs that were amazing in the vendor demo but didn’t hold up at scale or breadth. Production-grade systems, he said, are usually multi-component, compound, or agentic, requiring real investment, golden datasets, evaluation schemas, and LLM judges — not just a vector store and a prompt. The C-suite’s “max expectation of simplicity” collided with how much grounding and tuning these systems actually need. **How should AI teams approach safety and alignment across different regions?** Sara Hooker, VP of Research at Cohere, framed safety as a multiple-objective optimization problem. Some concepts are global — near-universal consensus on things like no generated violence and strict guardrails around sexual and CSAM content — but there are also local harms specific to each jurisdiction. The challenge is aligning a model to both global and local standards while staying flexible enough to adapt per domain and region. Her team’s main tool was preference alignment, which she called the “gold dust” that happens at the end of training where you steer the model toward those objectives. **Are bigger or smaller AI models the future, according to the panel?** Both — that was Sara Hooker’s hot take. She predicted dual trends: models would keep getting bigger, mainly to push the frontier on multimodal, while many models would also get a lot smaller as teams better leverage optimization across the inference space and stratified data. Conor Bronsdon echoed the small-model side, pointing to fine-tuned SLMs “having their day again” as the largest LLMs hit a near-term ceiling on training data, giving smaller and open-source models room to catch up (Vikram agreed). ## EP 2: Got Agents? Agentic Workflows & Architecture | Weaviate, Unstructured & CrewAI Guests: Brian Raymond, Bob van Luijt & João Moura · Published: 2024-11-13 · Duration: 31 min URL: https://chainofthought.show/podcast/2-got-agents-agentic-workflows-and-architecture-weaviate-unstructured-and-crewai/ (full transcript on page) AI agents have quickly emerged as the next ‘hot thing’ in AI, but what constitutes an AI agent and do they live up to the hype? Join Brian Raymond, founder & CEO at Unstructured.io, Bob van Luijt, co-founder & CEO at Weaviate, and João Moura, founder at CrewAI as they discuss the shift to agentic workflows, dissect their architecture, and tackle real-world challenges in agent deployment. From data management tips to generative feedback loops, this episode is your essential guide to operationalizing agents effectively. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - João Moura (CrewAI) draws the line at agency: a fixed if-this-then-that flow isn’t an agent. An agent “controls the flow of the application,” using an LLM “as a brain” to reason about what tools to use, when to use them, and how to format the answer — “something that has agency.” - Bob van Luijt (Weaviate) argues data is still the bottleneck, not models. A CIO told him: “it’s all great what you guys are building, but I still have like 15 ERP systems and I can’t make heads or tails of it.” Low data quality and not getting to insights remains “the number one problem.” - João Moura (CrewAI) warns agent architectures start simple but explode in production: once inside an enterprise you bolt on a caching layer, a memory layer (long-term memory of what the agent has done and learned), guardrails, validation, security, and distribution — and the complexity “can get…to skyrocket a little bit.” - Brian Raymond (Unstructured) says don’t expect agents to fix data plumbing yet: a CIO named RBAC — role-based access controls across a huge org — as his number one blocker to GenAI adoption. “What a boring topic… but it’s real.” Bob van Luijt urges solving it the new way, leveraging models, not slapping old binary access controls onto new systems. - João Moura (CrewAI) is blunt on accuracy: “You don’t get a 100% accuracy with AI agents. That’s not how things goes.” Humans in the loop get you very close. He tells customers chasing fully automated accounting they’ll “have some cute little loop there” — don’t pretend you’ll automate the whole thing end to end. - Bob van Luijt (Weaviate) pitches generative feedback loops: prompt the database, not the model. Store a product in Spanish in an English-only collection and it auto-upserts; store an oven temp in Fahrenheit when the dataset wants Celsius and it converts. You get full CRUD — the database reads, updates, deletes, and creates — making it “a form of an agentic architecture.” ### FAQ **What actually makes something an AI agent versus a normal workflow?** Per João Moura of CrewAI, it comes down to agency. A regular flow — if-this-then-that, triggering alarms on conditions — is not an agent. An agent is “something that controls the flow of the application,” using an LLM as a brain to reason and decide what tools to use, when to use them, and how to format the answer. Bob van Luijt of Weaviate frames it the same way: instead of “if this, then go left, else go right,” there’s reasoning behind whether you go left or right. **Why was RAG adopted so quickly for agents?** Brian Raymond of Unstructured explains that models alone have three problems: they’re “frozen in time,” they have no access to private data, and “they tend to make stuff up.” Retrieval-augmented generation became the dominant paradigm to counter those issues by feeding models the right private context. Bob van Luijt of Weaviate adds that classic RAG is “one directional” — a single line from query to answer — and the next leap is adding loops, multiple queries, and feedback so the database itself gets smarter. **Can AI agents reach 100% accuracy in production?** No — João Moura of CrewAI is direct: “You don’t get a 100% accuracy with AI agents. That’s not how things goes.” Customers arriving with crazy use cases expecting 100% are chasing something that isn’t available right now. Humans in the loop bring you very close, and Moura tells people who want to fully automate something like accounting that they’ll need “some cute little loop there” — don’t pretend you’ll automate it end to end. **What agentic use cases actually work well today?** João Moura of CrewAI sees a normal distribution of use cases clustering around research, analysis, summarization, and reporting — often combined. A typical winner: pull data from an ERP or CRM, run analysis, summarize, and emit a report (JSON or otherwise) to push into another system. Those “work like a charm.” The other pattern that works is bringing a human into the loop when necessary so you still save time but keep someone in control. **What does Weaviate mean by a generative feedback loop?** Bob van Luijt of Weaviate describes it as prompting the database instead of the model. Define a rule on a collection — e.g., every e-commerce product must be in American English, or all oven temperatures must be in Celsius — and when something arrives in Spanish or Fahrenheit, the database auto-converts and upserts it. You get full CRUD support, so the database reads, updates, deletes, and creates, making the loop “a form of an agentic architecture.” ## EP 1: The State of AI: Open-Source Models & Enterprise Trust | May Habib Guests: May Habib, Writer · Published: 2024-11-06 · Duration: 49 min URL: https://chainofthought.show/podcast/1-the-state-of-ai-open-source-models-and-enterprise-trust-may-habib/ (full transcript on page) From ChatGPT's search engine to Google's AI-powered code generation, artificial intelligence is transforming how we build and deploy technology. In this inaugural episode of Chain of Thought, the co-founders of Galileo explore the state of AI, from open-source models to establishing trust in enterprise applications. Plus, tune in for a segment on the impact of the Presidential election on AI regulation. Chain of Thought is hosted by Conor Bronsdon. ### Key takeaways - Atin Sanyal argued open source is overlooked relative to proprietary hype: Llama 3.1 and 3.2 now outperform GPT-4o on many benchmarks, the “cost of intelligence” has fallen enough to run models on a laptop, and open weights give far more degrees of freedom to fine-tune than a closed fine-tuning API. - Vikram Chatterji reframed corporate open source as a distribution channel, not altruism: Meta open-sourced Llama as a “lead generation mechanism” and brand play after lagging. He contrasted it with Google giving BERT away for free while OpenAI invested in more parameters and larger context windows and “won” the early language-model war. - Yash Sheth said Meta spending hundreds of millions to train open models commoditizes the costliest part — training and data gathering — so banks, healthcare, defense, and finance teams can build proprietary AI on top, massively furthering adoption that the proprietary-only world couldn’t match at the same rate. - The founders argue AI development has regressed to the “stone ages of the SDLC” — back to “vibe checks” and eyeballing — because there’s no trust layer. Atin noted you can’t manually check every one of a million production queries, making a deterministic evaluation/trust layer the critical bottleneck for shipping, which is why Galileo raised its $45M Series B. - Yash and Atin tied the ChatGPT moment to two human-driven factors — RLHF (human feedback as reward signal) and data quality — without which GPT-3.5 would have been only “sublinearly” better than GPT-3. The same human-in-the-loop principles now need to be baked into evaluation pipelines: one founder wagered that ~90% of the time to ship a new GPT version goes to human annotation and evaluation, not GPU training (his own bet, not an OpenAI figure). - May Habib said Writer’s “secret sauce” is abstracting complex engineering away: its RAG is “zero engineering rag” where the word “rag” appears nowhere in the product, guardrails are built with LLMs rather than naive rejects, and domain-specific models (financial services, customer support, creative, medical) ship preconfigured so customers don’t need to fine-tune — capturing business logic via examples, which need far less data than fine-tuning. ### FAQ **Is open source AI catching up to proprietary models like GPT-4?** Yes, per Galileo CTO Atin Sanyal. He pointed to Llama 3.1 and 3.2 outperforming GPT-4o on many benchmarks and noted the cost of intelligence has dropped enough to run these models on a laptop. The bigger advantage is open weights: you can fine-tune them with far more freedom than a closed fine-tuning API. Vikram Chatterji and Yash Sheth agreed both have a place — open source acts as a distribution channel and commoditizes the expensive training layer. **Why do enterprises need a trust layer for AI applications?** Because, as Yash Sheth put it, AI development has gone “back to the stone ages of the SDLC” — manual, no clear process, “vibe checks” instead of the IDEs, CI/CD, monitoring, and firewalls that let traditional software ship safely the same day. Atin Sanyal noted you can’t manually check every one of a million production queries, so without a deterministic evaluation and trust layer, teams can’t guarantee an app won’t damage their reputation or erode user trust. Galileo raised a $45M Series B to build that layer. **How much of Google’s code is written by AI?** About 25%, a stat Sundar Pichai shared on Google’s earnings call. The Galileo founders framed it as real proof of value rather than hype: Atin Sanyal said it likely automates the most obvious, boilerplate “grunt work” — the code developers are too lazy to write — freeing engineers for more nuanced problems. Conor Bronsdon added a parallel from Amazon’s Andy Jassy, who cited saving 4,500 developer-years upgrading code from Java 11 to Java 17. **How does Writer get enterprise AI accuracy without making customers fine-tune?** CEO May Habib says Writer ships domain-specific models — for financial services, customer support, creative/marketing, and medical — preconfigured and close to the use cases vertical by vertical, so customers don’t need to fine-tune. In its AI Studio, Writer captures business logic by pairing it with examples rather than data, and examples need far less data than fine-tuning. Writer also runs its own QA teams staffed with PhDs to build evals even when the customer hasn’t. **Should AI regulation target foundation models or applications?** The Galileo founders favored guardrails at the application layer. Yash Sheth warned that restricting foundation models upfront would “drastically reduce the number of use cases” and stifle innovation; instead, model makers should disclose training data to avoid harmful or biased content, while application developers handle domain-specific compliance (e.g., healthcare). They noted regulation is fragmenting across state and federal levels, and that AI policies jumped from 25 in 2023 to 82 by March 2024 and near 100 by year-end. # AI Glossary Plain-language definitions of the AI terms discussed on the show. Each is citable on its own; the linked page connects it to the episodes that discuss it. ## Accuracy Accuracy is the share of predictions a model got right out of all predictions. It's the most intuitive metric and the most misleading — on imbalanced data, a model can score high accuracy while being useless, which is why it's rarely enough on its own. URL: https://chainofthought.show/glossary/accuracy/ ## Agent Memory Agent memory is what an AI agent keeps and reuses beyond a single turn — working memory in the context window, plus longer-term stores it can write to and retrieve from later. It's how an agent stops starting every session from zero. URL: https://chainofthought.show/glossary/agent-memory/ ## Agentic Workflow An agentic workflow is a process where an AI model drives multiple steps toward a goal — planning, calling tools, and reacting to results in a loop — rather than producing a single response. It sits between a one-shot prompt and a fully autonomous agent: structured enough to be reliable, dynamic enough to handle real tasks. URL: https://chainofthought.show/glossary/agentic-workflow/ ## AI Agent An AI agent is an LLM-powered system that pursues a goal over multiple steps — deciding what to do, using tools to act, and reacting to the results — rather than just answering a single prompt. The model is the brain; the agent is the loop around it. URL: https://chainofthought.show/glossary/ai-agent/ ## AI Alignment AI alignment is the work of making an AI system pursue what its designers and users actually intend — including the goals they didn't think to spell out — rather than optimizing a literal objective in harmful or unintended ways. It spans training techniques, evaluation, and oversight. URL: https://chainofthought.show/glossary/ai-alignment/ ## AI Benchmark An AI benchmark is a standardized test set used to compare models on a task — same questions, same scoring, so results are comparable across models. Benchmarks rank general capability; they don't tell you how a model performs on your specific use case, which is what evaluation is for. URL: https://chainofthought.show/glossary/ai-benchmark/ ## AI Evaluation AI evaluation is how you measure whether an AI system actually works — scoring its outputs against what good looks like, systematically and repeatably, instead of eyeballing a few demos. For non-deterministic systems like LLMs and agents, it's the discipline that separates a thing that demos well from one you can ship. URL: https://chainofthought.show/glossary/ai-evaluation/ ## AI Governance AI governance is the set of controls that lets an organization deploy AI responsibly: knowing what AI systems are running, bounding what they can do, logging what they did, and naming who's accountable. It's how you earn the right to ship AI that takes real actions. URL: https://chainofthought.show/glossary/ai-governance/ ## AI Guardrails Guardrails are the checks that keep an AI system inside safe, intended behavior — filtering inputs, constraining what it can do, and validating outputs before they reach a user. They run outside the model, so they hold even when the model is wrong or manipulated. URL: https://chainofthought.show/glossary/ai-guardrails/ ## AI Hallucination An AI hallucination is when a model states something false or fabricated as if it were fact — a confident answer with no grounding in its training data or the sources it was given. It's a property of how language models generate text, not an occasional bug. URL: https://chainofthought.show/glossary/ai-hallucination/ ## AI Pilot (Proof of Concept) An AI pilot is a small, time-boxed test of an AI use case before a full rollout. The trap is that pilots are easy and production is hard — a demo that works with a few users often dies on the way to scale, which is why so many never show ROI. URL: https://chainofthought.show/glossary/ai-pilot/ ## AI Red Teaming AI red teaming is deliberately attacking your own AI system before someone else does — probing it with adversarial inputs to find where it leaks data, breaks its rules, or fails dangerously, so you can fix those holes before launch. URL: https://chainofthought.show/glossary/red-teaming/ ## AI Safety AI safety is the work of keeping AI systems from causing harm — making them behave as intended, refuse dangerous requests, and fail gracefully. In practice for builders it means alignment, guardrails, evaluation for harmful behavior, and human oversight on consequential actions. URL: https://chainofthought.show/glossary/ai-safety/ ## Answer Relevance Answer relevance measures whether a response actually addresses the question that was asked, rather than drifting into related-but-off material. It catches the failure where an answer is true and well-sourced but doesn't answer what the user wanted. URL: https://chainofthought.show/glossary/answer-relevance/ ## Artificial General Intelligence (AGI) Artificial general intelligence (AGI) is a hypothetical AI that can understand, learn, and perform any intellectual task a human can, rather than excelling at narrow ones. Today's models are narrow — capable but specialized — and there's no agreed definition or test for when AGI would arrive. URL: https://chainofthought.show/glossary/agi/ ## AUC-ROC AUC-ROC measures how well a classifier separates two classes across every possible threshold, summarized as one number from 0.5 (random) to 1.0 (perfect). Unlike accuracy, it doesn't depend on where you set the decision cutoff. URL: https://chainofthought.show/glossary/auc-roc/ ## Audit Trail An audit trail is a durable, reconstructable record of what an AI system did — the inputs, decisions, tool calls, and outputs — so you can later explain or investigate any action it took. It's what turns 'the agent did something' into 'here's exactly what it did and why.' URL: https://chainofthought.show/glossary/audit-trail/ ## Backdoor Attack A backdoor attack plants a hidden trigger in a model during training, so it behaves normally until it sees a specific input — then it flips to attacker-chosen behavior. The model passes normal testing, which is what makes the backdoor dangerous. URL: https://chainofthought.show/glossary/backdoor-attack/ ## BERTScore BERTScore compares generated and reference text by the similarity of their embeddings rather than exact word overlap. Because it works in meaning-space, it credits a correct paraphrase that BLEU or ROUGE would mark down. URL: https://chainofthought.show/glossary/bertscore/ ## BLEU Score BLEU scores machine-generated text by how much its word sequences overlap with one or more human reference texts. It was built for machine translation, runs from 0 to 1, and is fast and cheap — but it rewards surface word-matching, not meaning. URL: https://chainofthought.show/glossary/bleu-score/ ## Chain-of-Thought Prompting Chain-of-thought prompting asks a model to work through its reasoning step by step before giving a final answer. Spelling out the intermediate steps measurably improves accuracy on multi-step problems like math, logic, and planning. URL: https://chainofthought.show/glossary/chain-of-thought-prompting/ ## Cohen's Kappa Cohen's Kappa measures how much two raters agree beyond what you'd expect from random chance. It matters for AI because it's how you check whether your human labels — or an LLM judge against humans — are consistent enough to trust as ground truth. URL: https://chainofthought.show/glossary/cohens-kappa/ ## Compound AI Systems A compound AI system solves a task with multiple components — several model calls, retrieval, tools, and control logic — rather than a single prompt to a single model. Most real production AI is compound, which is why reliability is a systems problem, not a model problem. URL: https://chainofthought.show/glossary/compound-ai-systems/ ## Context Engineering Context engineering is the discipline of deciding what information goes into a model's context window for a task — which documents, which history, which tool output — and how fresh and trustworthy it is. As models commoditize, it's where a lot of the durable advantage now lives. URL: https://chainofthought.show/glossary/context-engineering/ ## Context Graph A context graph extends a knowledge graph with business rules, terminology definitions, and access metadata — such as what 'churn' or 'fiscal year' means for a given team — so AI agents can interpret enterprise data correctly, not just know where it lives. URL: https://chainofthought.show/glossary/context-graph/ ## Context Poisoning Context poisoning is when bad, stale, or excessive information in a model's context window degrades its reasoning — the agent gets buried in irrelevant tokens or misled by wrong data, and its answers get worse even though the model is fine. URL: https://chainofthought.show/glossary/context-poisoning/ ## Context Relevance Context relevance measures whether the documents a RAG system retrieved actually bear on the question. It scores the retrieval step on its own, before the model writes anything — because the best model can't answer well from the wrong context. URL: https://chainofthought.show/glossary/context-relevance/ ## Context Window The context window is the amount of text a model can take in at once — the prompt, the conversation so far, and any documents or tool output you include. Everything the model can 'see' for a given response has to fit inside it. URL: https://chainofthought.show/glossary/context-window/ ## Continuous Compliance Continuous compliance is monitoring data and systems for regulatory adherence while they're running, instead of certifying them once at launch and calling it done. It matters for AI because two or three data elements that are unregulated on their own can become PII the moment a system joins them — something a point-in-time audit will never catch. URL: https://chainofthought.show/glossary/continuous-compliance/ ## Conversational Analytics Conversational analytics is a natural-language interface for querying structured business data directly, replacing dashboards and static reports with a chat-style question-and-answer experience. URL: https://chainofthought.show/glossary/conversational-analytics/ ## Data Federation Data federation is an architecture where a virtual compute layer queries data across many source systems in place, rather than copying everything into a single centralized warehouse or lake first. URL: https://chainofthought.show/glossary/data-federation/ ## Data Poisoning Data poisoning is an attack that corrupts the data a model learns from — its training set, fine-tuning examples, or a knowledge base it retrieves from — so the model behaves the way the attacker wants while looking normal. URL: https://chainofthought.show/glossary/data-poisoning/ ## Data Sovereignty Data sovereignty is control over data and the workloads that process it: what data gets used, where it's stored, when it's processed, and which data center it runs in. Legally it's a question of jurisdiction, but enterprises increasingly use the word to mean operational control — the ability to move a workload, or pull it home, when they need to. URL: https://chainofthought.show/glossary/data-sovereignty/ ## Embeddings Embeddings are numerical representations of text, images, or other data as vectors, where things with similar meaning land close together. They're what lets a system search by meaning instead of by exact keyword. URL: https://chainofthought.show/glossary/embeddings/ ## Entity Resolution Entity resolution is matching records across systems that refer to the same real-world thing — the same customer, company, or account — when the names, spellings, and identifiers don't line up. It's long-standing data-engineering work, and agents with structured, indexed context can now do a usable version of it without a human defining the join first. URL: https://chainofthought.show/glossary/entity-resolution/ ## EU AI Act The EU AI Act is the European Union's regulation of AI, which sorts systems by risk level and imposes obligations accordingly — banning a few uses outright, heavily regulating 'high-risk' ones, and adding transparency rules for general-purpose models. Like GDPR, its reach extends to anyone serving EU users. URL: https://chainofthought.show/glossary/eu-ai-act/ ## Evasion Attack An evasion attack crafts an input designed to slip past a model's classifier or safety check at inference time — a spam message tweaked to read as legitimate, a malicious payload perturbed to look benign. The model isn't compromised; it's fooled by an input built to exploit its blind spots. URL: https://chainofthought.show/glossary/evasion-attack/ ## Excessive Agency Excessive agency is giving an AI agent more capability, permission, or autonomy than its task needs — broad tool access, write permissions, the ability to act without approval. It turns a model mistake or a successful attack into real-world damage. URL: https://chainofthought.show/glossary/excessive-agency/ ## Explainability Explainability is how well you can understand why an AI system produced a given output. It matters most where decisions need to be justified — lending, hiring, healthcare — and it's hard for large models, whose reasoning isn't transparent just because they can narrate a plausible-sounding rationale. URL: https://chainofthought.show/glossary/explainability/ ## F1 Score The F1 score combines precision and recall into a single number — their harmonic mean. It's high only when both are high, which makes it a fairer summary than plain accuracy when the classes are imbalanced. URL: https://chainofthought.show/glossary/f1-score/ ## Faithfulness Faithfulness measures whether an answer is actually supported by the source material it was given — every claim traceable to the retrieved context, nothing invented. It's the core anti-hallucination metric for RAG systems. URL: https://chainofthought.show/glossary/faithfulness/ ## Few-Shot Learning Few-shot learning is giving a model a handful of worked examples in the prompt so it picks up the pattern and applies it to your task — no retraining. Zero-shot means no examples (just the instruction); few-shot adds a few; both work because large models learn from context at inference time. URL: https://chainofthought.show/glossary/few-shot-learning/ ## Fine-Tuning Fine-tuning continues training a pretrained model on your own examples so it gets better at a specific task, tone, or format. It changes the model's weights, unlike prompting or RAG, which change what you feed it. URL: https://chainofthought.show/glossary/fine-tuning/ ## FlashAttention FlashAttention is an optimized way to compute a transformer's attention that's far more memory-efficient, by avoiding writing the huge intermediate attention matrix to memory. It makes longer context windows and faster training practical without changing the model's results. URL: https://chainofthought.show/glossary/flashattention/ ## Foundation Model A foundation model is a large AI model trained on broad data at scale so it can be adapted to many downstream tasks, rather than built for one. LLMs like GPT, Claude, and Gemini are foundation models — the general-purpose base that products are built on top of. URL: https://chainofthought.show/glossary/foundation-model/ ## Frontier Model A frontier model is one of the most capable AI models available at a given moment — the latest flagship releases from the major labs that set the current ceiling on what's possible. The label moves: today's frontier model is next year's baseline. URL: https://chainofthought.show/glossary/frontier-model/ ## GraphRAG GraphRAG is retrieval-augmented generation that retrieves from a knowledge graph — or a graph plus a vector index — instead of vector similarity alone. It grounds the model in explicit entities and relationships, so answers respect real facts and connections, not just topical similarity. URL: https://chainofthought.show/glossary/graphrag/ ## Human in the Loop Human in the loop means keeping a person in the decision path of an AI system — to approve high-stakes actions, review uncertain outputs, or label the cases the model got wrong. It's the practical way to deploy autonomy you don't fully trust yet. URL: https://chainofthought.show/glossary/human-in-the-loop/ ## Inference Inference is running a trained model to produce output — the part that happens every time a user sends a prompt. It's distinct from training (teaching the model in the first place): training is a big one-time cost, inference is the recurring cost you pay on every request, forever. URL: https://chainofthought.show/glossary/inference/ ## Instruction Adherence Instruction adherence measures whether a model actually did what it was told — followed the format, honored the constraints, stayed within the rules of the prompt. A model can give a high-quality answer that ignores half the instructions, and this is the metric that catches it. URL: https://chainofthought.show/glossary/instruction-adherence/ ## Jailbreaking Jailbreaking is crafting a prompt that gets a model to bypass its own safety rules — producing content it was trained to refuse — usually through roleplay, hypotheticals, or obfuscation that talks the model around its guardrails. URL: https://chainofthought.show/glossary/jailbreaking/ ## Knowledge Distillation Knowledge distillation trains a small 'student' model to imitate a larger 'teacher' model, transferring much of the teacher's capability into a model that's cheaper and faster to run. It's a main way the strong-but-expensive becomes small-enough-to-ship. URL: https://chainofthought.show/glossary/knowledge-distillation/ ## Knowledge Graph A knowledge graph stores information as entities and the relationships between them, rather than as loose documents or vectors. For AI, it gives a model structured, connected context it can traverse — which is one answer to grounding and hallucination. URL: https://chainofthought.show/glossary/knowledge-graph/ ## Latency Latency is how long an AI system takes to respond. For LLMs it splits into time-to-first-token (how fast output starts) and total generation time, and it's a first-class product metric — a more accurate model that's too slow can still be the wrong choice. URL: https://chainofthought.show/glossary/latency/ ## LLM as a Judge LLM-as-a-judge is using one language model to score the outputs of another against a rubric you define — quality, relevance, safety, correctness. It scales evaluation to volumes humans can't review by hand, trading some reliability for enormous reach. URL: https://chainofthought.show/glossary/llm-as-a-judge/ ## LoRA (Low-Rank Adaptation) LoRA is a parameter-efficient way to fine-tune a model: instead of updating all its weights, you train small add-on matrices and leave the original model frozen. You get most of the benefit of fine-tuning at a fraction of the compute and storage. URL: https://chainofthought.show/glossary/lora/ ## Mean Reciprocal Rank (MRR) MRR measures how high up the first correct result appears in a ranked list, averaged over many queries. If the right answer is usually near the top, MRR is close to 1; if it's buried, MRR drops. It's a core retrieval and search metric. URL: https://chainofthought.show/glossary/mean-reciprocal-rank/ ## Membership Inference Attack A membership inference attack figures out whether a specific record was in a model's training data by probing how the model responds. It's a privacy leak: confirming someone's data was used can itself expose sensitive information. URL: https://chainofthought.show/glossary/membership-inference-attack/ ## METEOR METEOR is a text-generation metric that scores overlap with a reference more flexibly than BLEU — it credits synonyms and word-stem matches, not just exact words, and accounts for word order. It was designed to correlate better with human judgment on translation. URL: https://chainofthought.show/glossary/meteor/ ## Mixture of Experts (MoE) A mixture of experts is a model split into many specialized sub-networks, where a router sends each input to just a few of them. You get the capacity of a huge model while only running a fraction of it per request. URL: https://chainofthought.show/glossary/mixture-of-experts/ ## Model Context Protocol (MCP) The Model Context Protocol (MCP) is an open standard for connecting AI models to tools and data sources. It defines one common interface — like a USB-C port for AI — so any MCP-compatible model can use any MCP-compatible tool without custom integration code for each pairing. URL: https://chainofthought.show/glossary/model-context-protocol/ ## Model Denial of Service Model denial of service is making an AI system unavailable or ruinously expensive by flooding it with requests or crafting inputs that force maximum work — huge outputs, deep tool loops, giant context. Because each call costs real money, the financial version is sometimes called 'denial of wallet.' URL: https://chainofthought.show/glossary/model-denial-of-service/ ## Model Drift Model drift is the gradual decline in an AI system's performance after deployment as the real world moves away from what it was built on — new inputs, changed user behavior, shifting data. The model didn't change; the world it operates in did, so accuracy quietly erodes. URL: https://chainofthought.show/glossary/model-drift/ ## Model Inversion Attack A model inversion attack reconstructs sensitive training data by probing a model's outputs — recovering, for example, features of the records it was trained on. It's a privacy threat: the model itself can leak the data it learned from. URL: https://chainofthought.show/glossary/model-inversion-attack/ ## Model Risk Management Model risk management is the discipline of identifying, measuring, and controlling the risks a model poses to a business — that it's wrong, biased, misused, or drifts over time. It comes from regulated finance and now applies to AI: treat each model as a risk to be governed, not just a tool to be shipped. URL: https://chainofthought.show/glossary/model-risk-management/ ## Multimodal AI Multimodal AI is a model that works across more than one kind of data — text, images, audio, video — in a single system, rather than handling only text. It can take a screenshot and a question together, or describe an image, because it represents different modalities in a shared space. URL: https://chainofthought.show/glossary/multimodal-ai/ ## Open Weights An open-weights model is one whose trained parameters are released publicly, so anyone can download, run, inspect, and fine-tune it. It's distinct from fully open source — the weights are open even when the training data and code aren't. URL: https://chainofthought.show/glossary/open-weights/ ## Pareto Frontier (Cost-Performance Curve) A Pareto frontier, or cost-performance curve, plots model capability against cost to show which models give the most capability for a given price. A model sits on the frontier when nothing else is both cheaper and more capable. Builders use it to compare tiers, such as a fast cheap model against a larger flagship, and pick the trade-off that fits the task. URL: https://chainofthought.show/glossary/pareto-frontier-cost-performance-curve/ ## Perplexity Perplexity measures how surprised a language model is by a piece of text — lower means the model found it more predictable. It's a quick intrinsic gauge of how well a model fits a dataset, but it says little about whether the model is actually useful or correct. URL: https://chainofthought.show/glossary/perplexity/ ## Physical AI Physical AI refers to AI models that perceive and act in the physical world through robots or embodied hardware, rather than operating only on text or digital data. The shorthand for it is AI moving from bits to atoms. A language model handling a robot's vision, speech, and motor control in real time is physical AI. URL: https://chainofthought.show/glossary/physical-ai/ ## Policy as Code Policy as code is writing governance, security, and compliance rules as machine-readable, version-controlled code so systems enforce them automatically instead of waiting on a human review. For agentic AI it's how you encode the guardrails: what an agent can access, what it can do, and what makes it stop. URL: https://chainofthought.show/glossary/policy-as-code/ ## Precision and Recall Precision and recall are two sides of accuracy. Precision asks: of the things the system flagged, how many were right? Recall asks: of the things it should have flagged, how many did it catch? They trade off against each other, so which one matters depends on whether false positives or misses cost you more. URL: https://chainofthought.show/glossary/precision-and-recall/ ## Prompt Engineering Prompt engineering is crafting the instruction you give a model — the wording, examples, and output format — to get better results without retraining it. It's the most visible AI skill and, as models improve, increasingly table stakes rather than a moat. URL: https://chainofthought.show/glossary/prompt-engineering/ ## Prompt Injection Prompt injection is an attack where malicious instructions hidden in the input — a user message, a web page, a document the agent reads — trick the model into ignoring its real instructions and doing the attacker's bidding instead. URL: https://chainofthought.show/glossary/prompt-injection/ ## Quantization Quantization shrinks a model by storing its weights at lower numerical precision — say 4-bit integers instead of 16-bit floats. The model gets smaller and faster to run, usually with little quality loss, which is what lets large models fit on smaller hardware. URL: https://chainofthought.show/glossary/quantization/ ## ReAct (Reason + Act) ReAct is an agent execution pattern in which a model alternates between reasoning about a task and taking an action, such as firing a query or calling a tool, repeating the cycle until it decides it has an answer. URL: https://chainofthought.show/glossary/react-reason-act/ ## Reasoning Models Reasoning models are LLMs trained to do extended step-by-step thinking before they answer, spending more compute at inference to work through hard problems. They trade latency and cost for accuracy on math, code, and multi-step logic. URL: https://chainofthought.show/glossary/reasoning-models/ ## Retrieval-Augmented Generation (RAG) RAG is the pattern of fetching relevant documents at query time and feeding them to a model alongside the question, so the answer is grounded in real sources instead of the model's memory. It's how you put private or current data in front of a model without retraining it. URL: https://chainofthought.show/glossary/retrieval-augmented-generation/ ## RLHF (Reinforcement Learning from Human Feedback) RLHF is a training step that tunes a model toward what people actually prefer: humans rank model outputs, those rankings train a reward model, and the model is then optimized to score well against it. It's a big part of why chat models feel helpful instead of just fluent. URL: https://chainofthought.show/glossary/rlhf/ ## Robotic Process Automation (RPA) RPA automates repetitive digital tasks with explicit, rule-based scripts — click here, copy this field, paste it there. It's deterministic and brittle: it does exactly what it's told and breaks when the screen or process changes, which is the contrast that defines AI agents. URL: https://chainofthought.show/glossary/robotic-process-automation/ ## ROUGE ROUGE scores a generated summary by how much it overlaps with a human reference summary — leaning on recall, how much of the reference's content the output captured. It's the standard automatic metric for summarization. URL: https://chainofthought.show/glossary/rouge/ ## Shadow AI Shadow AI is employees using AI tools their organization hasn't approved or doesn't know about — pasting work into a consumer chatbot, wiring up an unsanctioned agent. It's where a lot of real AI adoption actually happens, and where the governance and data-leak risk lives. URL: https://chainofthought.show/glossary/shadow-ai/ ## State-Space Models (Mamba) State-space models are a transformer alternative that process sequences by carrying a compact running state forward, rather than comparing every token to every other token. They scale linearly with sequence length instead of quadratically — cheaper on long inputs — with Mamba the best-known example. URL: https://chainofthought.show/glossary/state-space-models/ ## Synthetic Data Synthetic data is training or evaluation data generated by a model rather than collected from the real world. It's used to cover cases real data is missing, scarce, expensive, or too sensitive to use — and increasingly to train models when high-quality human data runs short. URL: https://chainofthought.show/glossary/synthetic-data/ ## Temperature Temperature is the setting that controls how random a model's output is. Low temperature makes it pick the most likely next token almost every time (focused, repeatable); high temperature spreads the odds (varied, creative, less predictable). It's the main dial between consistency and creativity. URL: https://chainofthought.show/glossary/temperature/ ## Test-Time Compute Test-time compute is the processing a model spends while answering, rather than during training. Letting a model 'think' longer at answer time — exploring and checking more before it commits — raises accuracy on hard problems without retraining the model. URL: https://chainofthought.show/glossary/test-time-compute/ ## Token Leakage Token leakage is an AI system exposing secrets it shouldn't — API keys, credentials, or auth tokens — in its output, logs, or traces. It happens when secrets end up in the context window or tool results and the model repeats them, or when verbose logging captures them. URL: https://chainofthought.show/glossary/token-leakage/ ## Tokenization Tokenization splits text into the chunks a model actually processes — tokens, which are roughly word-pieces, not whole words. It's why model limits and pricing are counted in tokens, and why 'a few paragraphs' is a fuzzy unit but 'tokens' is exact. URL: https://chainofthought.show/glossary/tokenization/ ## Tool Use (Function Calling) Tool use, also called function calling, is how an AI model takes real action: instead of only generating text, it emits a structured call to an external function — a search, a database query, a code run — and folds the result back into its answer. It's what lets a model do things, not just describe them. URL: https://chainofthought.show/glossary/tool-use/ ## TPU (Tensor Processing Unit) A TPU is a custom AI accelerator chip that Google designs specifically for training and running neural networks, built as an alternative to general-purpose GPUs. Google runs most of its own AI stack on TPUs, from the silicon up through models like Gemini, rather than depending on third-party GPU supply. URL: https://chainofthought.show/glossary/tpu-tensor-processing-unit/ ## Transformer The transformer is the neural-network architecture behind almost every modern large language model. Its key idea is attention: each token can look at every other token and weigh which ones matter, which is what lets the model handle context and long-range meaning. URL: https://chainofthought.show/glossary/transformer/ ## Vector Database A vector database stores embeddings — the numerical representations of your data — and is built to find the nearest ones to a query fast. It's the retrieval engine underneath most RAG systems. URL: https://chainofthought.show/glossary/vector-database/ ## Vibe Coding Vibe coding is building software by describing what you want in natural language and letting an AI generate the code, steering by the result rather than reading every line. It collapses the distance between idea and working prototype — and shifts the developer's job from writing code to specifying and reviewing it. URL: https://chainofthought.show/glossary/vibe-coding/ ## Vulnerability Chaining Vulnerability chaining is combining several individually low-severity vulnerabilities into one high-severity exploit. It matters more now because AI-driven tooling can find and link those weaknesses far faster than a human analyst, which turns the backlog of flaws everyone deprioritized into a live risk. URL: https://chainofthought.show/glossary/vulnerability-chaining/ ## Word Error Rate (WER) WER measures speech-recognition accuracy as the share of words a transcript got wrong — the insertions, deletions, and substitutions needed to fix it, divided by the number of words spoken. Lower is better, and unlike most metrics it can exceed 100%. URL: https://chainofthought.show/glossary/word-error-rate/ # AI, decoded — explainers ## How do you give an AI agent an identity and permissions? Give each agent instance a short-lived identity scoped to one task, not a standing service account, and carry the human it acts for along with it. Redpanda CTO Tyler Akidau's rule is that the permissions have to be enforced somewhere the agent cannot see or modify, because anything you enforce in the prompt eventually loses to prompt injection. URL: https://chainofthought.show/ai-decoded/ai-agent-identity-and-permissions/ ## Can a language model actually control a robot? Yes, and Google DeepMind's Paige Bailey describes it running on hobby hardware today: a Stanford-designed, 3D-printable robot dog taking spoken instructions through the Gemini APIs on a Raspberry Pi. The shift that matters is not new robots; it's that the control layer became a general model you can talk to. URL: https://chainofthought.show/ai-decoded/can-a-language-model-control-a-robot/ ## Can an AI agent actually run your workday? Postman's Sterling Chin says an agent he built runs 90% of his day, and the two things that made it work are unglamorous: scoping, and treating it like a junior hire. Worth knowing before you copy it: the 90% is his own estimate, and the productivity number behind it comes from a report the agent wrote about itself. URL: https://chainofthought.show/ai-decoded/can-an-ai-agent-run-your-workday/ ## Can switching to an open model actually cut your AI costs? On the narrow, high-volume tasks, yes. Intercom was spending $250,000 a month running one summarization job through a frontier API and replaced it with a fine-tuned 14-billion-parameter Qwen model. The decision is made per task rather than per platform: the hardest prompt in the same product still runs on a frontier model. URL: https://chainofthought.show/ai-decoded/can-an-open-model-cut-your-ai-costs/ ## Do you have to centralize your data before AI agents can use it? No, and Starburst's Jitender Aswani argues the attempt is what fails. A traditional enterprise runs 52 to 200 data sources and the count keeps climbing, so a plan that ends with everything in one place never finishes. Federation puts a query layer over the data where it already lives, which is the model that keeps working as the number of sources grows. URL: https://chainofthought.show/ai-decoded/do-you-need-to-centralize-data-for-ai/ ## Does AI replace human translators? Not yet, and Smartling's Olga Beregovaya is specific about where the gap sits. Models are trained mostly on English phenomena, so quality falls away on less-represented languages while the output still reads fluent, and the failure is factual and cultural rather than grammatical. Her framing is that AI does the heavy lifting and a human covers the delta. URL: https://chainofthought.show/ai-decoded/does-ai-replace-human-translators/ ## Is using AI bad for the environment? Not at the level of your own use, and many of the figures in circulation are wrong by orders of magnitude. The aggregate question is the one that bites: how much power the data-center buildout draws and where that power comes from. That is settled by utilities and siting decisions, not by how often you open a chatbot. URL: https://chainofthought.show/ai-decoded/is-ai-bad-for-the-environment/ ## Is AI taking software engineering jobs? The hiring data shows the matching layer breaking, not the jobs vanishing. Greenhouse saw applications rise 239% after ChatGPT while 75% fewer reached the hire stage, and software engineers are the heaviest users of the automation driving that. It is a signal problem, and each side makes it worse by responding rationally to the other. URL: https://chainofthought.show/ai-decoded/is-ai-taking-software-engineering-jobs/ ## What is context poisoning, and how do you stop it? Context poisoning is what happens when an agent's context window fills with wrong, stale, or badly shaped data and the agent reasons off it. The fix sits upstream of the model: control what reaches the window, because a single badly shaped API call can burn tens of thousands of tokens before the agent starts the work you asked for. URL: https://chainofthought.show/ai-decoded/what-is-context-poisoning/ ## What does it actually take to ship AI in a regulated industry? Provenance on every component, not a single passing benchmark. Corti CEO Andreas Cleve's requirement is that you can explain what each part of the pipeline did, why it changed, and what went into it. He also argues that public leaderboards are close to useless as evidence, because the models topping medical benchmarks are the ones nobody ships. URL: https://chainofthought.show/ai-decoded/what-it-takes-to-ship-ai-in-a-regulated-industry/ ## Who reviews the code when AI writes most of it? A human still owns the merge, but the review cannot stay a line-by-line read of the diff. Generation stopped being the constraint the moment background agents could open pull requests unattended, and the teams keeping up moved the check that decides into the test harness. URL: https://chainofthought.show/ai-decoded/who-reviews-ai-generated-code/ ## Why is it so hard to move AI workloads off NVIDIA? Because the lock-in lives in the software, not the silicon. Memory and math formats are catchable on a roadmap; what is hard to match is an ecosystem where every new model runs on day one, which is what AMD spent years buying back and what Google sidestepped by owning its stack from the compiler up. URL: https://chainofthought.show/ai-decoded/why-is-it-hard-to-move-off-nvidia/ ## Can AI agents be secured with software alone? No. Charles Guillemet, CTO of Ledger, says securing an agent with software alone is not possible. Ambiguous language and non-deterministic models break alignment, and prompt injection can trick an agent into leaking its own credentials. The fix is to delegate rights through a policy engine, prove that engine ran honestly with a secure enclave or a zero-knowledge proof, and keep the signing keys in hardware that signs only policy-approved intents. URL: https://chainofthought.show/ai-decoded/can-you-secure-an-ai-agent-with-software/ ## How do you measure whether AI is actually paying off? Measure time back, not usage. Count the human work an agent actually removed: work that took a person five hours a week and now takes an agent five minutes pays off; using AI to redo a font you could have fixed in one click does not. Jiaona Zhang calls the second case 'token maxing.' Get visibility into spend against the outcome you're driving — revenue or time-allocation efficiency — and articulate that outcome before you count. URL: https://chainofthought.show/ai-decoded/how-to-measure-if-ai-is-paying-off/ ## Why won't most websites get APIs for AI agents? Most websites will never expose agent APIs. The long tail that runs the internet, tens of thousands of school district sites, government offices, and hundreds of thousands of e-commerce pages, was built for humans and has no reason to re-architect. Dhruv Batra of Yutori calls the resistance socio-political, so coding agents can't fix it. His model: agents act like people. Pixels in, clicks out. If a machine can perceive the screen and click the buttons, that capability is the API. URL: https://chainofthought.show/ai-decoded/why-wont-the-web-get-apis-for-ai-agents/ ## When should you use a small language model instead of a frontier model in production? Default to a frontier model while you're figuring out what 'good' looks like — its broad capability lets you prototype fast without fighting the model. Move a task to a smaller model once it's well-scoped and high-volume, because that's where cost and latency dominate and a small model tuned to one job can match frontier quality at a fraction of the price. The frontier stays the right call for open-ended reasoning, low-volume work, and requirements that are still moving. The decision isn't 'which model is smartest' — it's which model is the cheapest, fastest way to clear the quality bar your specific task actually needs. URL: https://chainofthought.show/ai-decoded/small-language-model-vs-frontier-model/ ## Should you evaluate AI with an LLM-as-a-judge or with human review? Use an LLM-as-a-judge for scale and speed — scoring thousands of outputs continuously, catching regressions, and ranking A vs B. Use human evaluation for ground truth — defining what 'good' means, judging nuance and high-stakes cases, and calibrating the judge. They're a system, not a choice: humans set and audit the standard, the LLM judge applies it at volume. URL: https://chainofthought.show/ai-decoded/llm-as-a-judge-vs-human-evaluation/ ## Should you use MCP or build a custom integration to connect AI to your tools? Use MCP when a tool or data source will be reused across multiple models, apps, or teams — you integrate it once and everything speaks to it. Build a custom integration when you need a tight, high-control connection to a single system and the standard overhead isn't worth it. For most organizations scaling agents, MCP wins because it kills the N-models-times-M-tools integration explosion; custom is the exception for bespoke, performance-critical paths. URL: https://chainofthought.show/ai-decoded/mcp-vs-custom-tool-integration/ ## Should you use prompting, RAG, or fine-tuning to customize an AI model? Start at the cheapest rung and only climb when you must. Prompting shapes behavior with no infrastructure; RAG grounds the model in your data so answers stay current and citable; fine-tuning changes the model's weights to bake in a style, format, or skill. Most teams need prompting plus RAG, and reach for fine-tuning last — for how the model should behave, not what it should know. URL: https://chainofthought.show/ai-decoded/fine-tuning-vs-rag-vs-prompting/ ## When should you use a reasoning model instead of a standard LLM? Use a reasoning model when the task has multiple steps where a wrong turn early wrecks the answer — math, code, planning, hard analysis — and you can absorb the extra latency and cost. Use a standard model for everything else: retrieval, summarization, classification, and chat, where its speed and lower price win. Reasoning is a dial you spend on hard problems, not a default. URL: https://chainofthought.show/ai-decoded/reasoning-model-vs-standard-model/ ## Vector database or knowledge graph — which should you use for AI retrieval? Use a vector database when relevance is about meaning — finding passages similar to a question across unstructured text. Use a knowledge graph when the answer depends on explicit relationships and facts — who connects to what, and how. They're complementary, not rival: vectors find the right neighborhood, a graph enforces the right facts, and pairing them is increasingly how teams cut hallucination. URL: https://chainofthought.show/ai-decoded/vector-database-vs-knowledge-graph/ ## Which AI agent framework should you use — LangGraph, CrewAI, or AutoGen? They make different bets. LangGraph models an agent as an explicit graph of steps and state, so you trade simplicity for fine control. CrewAI organizes work as a 'crew' of role-playing agents with tasks, which is fast to stand up when the work splits cleanly by role. AutoGen centers on conversations between agents, good for open-ended problem-solving. Pick by how much control versus convention you want — and remember the strongest option is often no framework at all for a simple agent. URL: https://chainofthought.show/ai-decoded/ai-agent-frameworks-compared/ ## Do you still need an AI agent framework? Often no. A framework helps you start — it gives you tool-calling, state, and orchestration out of the box — but as the model providers fold those primitives into their own SDKs, the framework's value shrinks. The durable advantage isn't the framework; it's your context: what data you retrieve, how you manage it, and what the system remembers. Many teams start on a framework and then go framework-light as their needs get specific. URL: https://chainofthought.show/ai-decoded/do-you-need-an-agent-framework/ ## What is agentic RAG, and how is it different from regular RAG? Traditional RAG runs one fixed retrieve-then-generate step: fetch documents that match the query, stuff them in the prompt, answer. Agentic RAG puts an agent in charge of retrieval — it decides whether to search, reformulates the query, pulls from multiple sources, checks whether what it got is good enough, and retrieves again if it isn't. The difference is a static pipeline versus a control loop. URL: https://chainofthought.show/ai-decoded/agentic-rag-vs-traditional-rag/ ## What's the difference between agentic and non-agentic AI? Non-agentic AI runs a fixed path: you give it an input, it returns an output, done — a chatbot answering a question, a model classifying a document. Agentic AI runs a loop: it sets a sub-goal, takes an action, observes the result, and decides what to do next, repeating until the task is done. The line is autonomy over the steps. Non-agentic systems follow a path you defined; agentic systems decide the path themselves, which is more capable and far harder to predict. URL: https://chainofthought.show/ai-decoded/agentic-vs-non-agentic-ai/ ## What should you measure on an AI agent besides accuracy? Accuracy tells you whether the final answer was right, but it hides how the agent got there. The metrics that actually predict reliability watch the process: did it pick the right tool, call it correctly, and recover when something failed; how many steps and how much it cost to finish; whether it stayed on the user's intent across a long conversation; and how often it needed a human to step in. An agent can be accurate in a demo and unreliable in production because none of those were measured. URL: https://chainofthought.show/ai-decoded/ai-agent-metrics-beyond-accuracy/ ## Are AI hallucinations a data problem or a model problem? Largely a data problem. A language model predicts plausible text; when it lacks the right grounding it fills the gap with something that sounds right, which we call a hallucination. Much of that comes from the data layer — missing context, stale or contradictory sources, poor retrieval, no single source of truth. You can't fully train hallucination out of the model, but you can starve it: ground answers in trusted, current data and the model has less reason to invent. The model generates; the data decides whether it has the truth to generate from. URL: https://chainofthought.show/ai-decoded/are-ai-hallucinations-a-data-problem/ ## Are small language models better than large ones for production? Often, yes — for a specific, well-defined task. A small model that's been tuned for your job can match a frontier model's quality on that job while costing far less, running faster, and being possible to host yourself. The frontier models earn their keep on broad, open-ended reasoning. The mistake is defaulting to the biggest model for everything; the production-smart move is using the smallest model that still passes your evals for each task. URL: https://chainofthought.show/ai-decoded/are-small-language-models-better-for-production/ ## Can AI modernize legacy code and old applications? It can do a lot of the work, but not unsupervised. AI is good at the slow parts of modernization — reading undocumented code, explaining what a function does, translating between languages, and drafting migrations. Where it fails is the part that matters most: it doesn't know the business logic and edge cases the old system quietly encodes, so it will confidently rewrite something subtly wrong. The pattern that works is AI as an accelerator with engineers verifying, plus tests that prove the new code behaves like the old. URL: https://chainofthought.show/ai-decoded/can-ai-modernize-legacy-code/ ## How does an AI agent decide which tool to use? The agent is given a set of tools, each with a name and a description of what it does and when to use it. At each step the model reads the task and those descriptions and picks a tool, then generates the arguments to call it — a search query, an API payload, a database lookup. It runs the tool, reads the result, and decides the next move. The quality of that choice rides almost entirely on the tool descriptions: vague descriptions produce wrong tool calls, which is one of the most common ways agents fail. URL: https://chainofthought.show/ai-decoded/how-ai-agents-use-tools/ ## How do you cut the cost of running an AI agent? Most agent cost is hidden in the steps you can't see: redundant model calls, an oversized model doing a small job, bloated context sent on every turn, and retries from failures nobody caught. You cut it by first making the costs visible with tracing, then attacking the big drivers — route easy steps to a smaller or cheaper model, trim and cache context, cut needless tool calls and loops, and fix the failure modes that cause expensive retries. You can't optimize what you can't see, so observability comes first. URL: https://chainofthought.show/ai-decoded/how-to-cut-ai-agent-costs/ ## How do you evaluate a RAG system? Evaluate retrieval and generation separately, because they fail differently. For retrieval, ask whether the right documents came back — measure context relevance and recall. For generation, ask whether the answer is grounded in what was retrieved and actually answers the question — measure faithfulness (no claims beyond the sources) and answer relevance. A RAG system can retrieve perfectly and still hallucinate, or generate beautifully from the wrong documents, so a single end-to-end score hides which half is broken. URL: https://chainofthought.show/ai-decoded/how-to-evaluate-a-rag-system/ ## How do you govern AI agents in an enterprise? You govern agents the way you govern any system that takes consequential action: know what they are, control what they can do, and keep a record of what they did. In practice that means an inventory of every agent in production, scoped permissions and approval gates on high-stakes actions, audit trails of decisions and tool calls, and a named owner accountable for each one. The reason it matters now is trust — most leaders don't trust agent outputs, and governance is how you earn the right to deploy them anyway. URL: https://chainofthought.show/ai-decoded/how-to-govern-ai-agents/ ## How do you test an AI system when the output isn't deterministic? You stop expecting one exact answer and start testing properties. Because the same input can produce different valid outputs, traditional assert-equals tests don't fit. Instead you build a dataset of inputs with known-good characteristics and check each output against them — is it grounded, does it follow the instruction, does it avoid the unsafe thing — usually scored by a rubric or an LLM judge. You run that suite on every change, the way you'd run unit tests, so a regression shows up before users do. URL: https://chainofthought.show/ai-decoded/how-to-test-an-ai-system/ ## Is the AI agent bubble real? There's a real gap between the hype and what ships. Demos of autonomous agents are everywhere; reliable agents running unattended in production are rare, and a large share of agent projects never reach production at all. That doesn't mean agents are fake — it means the market priced in capability that the engineering hasn't caught up to yet. The bubble is in the expectations and the timeline, not in the underlying technology. URL: https://chainofthought.show/ai-decoded/is-the-ai-agent-bubble-real/ ## What's the difference between AI observability, evaluation, and benchmarking? Benchmarking compares models against a fixed dataset before you pick one — it answers 'which model is better in general.' Evaluation measures whether your system does the right thing on your task and your data — 'is this good enough to ship.' Observability is what you run in production — tracing live behavior to see what actually happened when something broke. They answer different questions at different stages, and teams get into trouble by using one where they need another. URL: https://chainofthought.show/ai-decoded/observability-vs-evaluation-vs-benchmarking/ ## Should you build a single agent or a multi-agent system? Start with a single agent. One agent with a clear set of tools is easier to build, debug, and trust, and it handles most tasks. Reach for multiple agents only when the work splits into distinct specialties that benefit from separate context and instructions — and accept that you're trading raw capability for new failure modes: coordination overhead, agents talking past each other, and harder debugging. Multi-agent is a way to manage complexity, not a free upgrade. URL: https://chainofthought.show/ai-decoded/single-agent-vs-multi-agent-architecture/ ## What are AI agent guardrails, and how do you set them? Guardrails are the limits that keep an autonomous agent inside safe, intended behavior — checks on what it's allowed to do, what it can access, and what it's about to output. They run at three points: on the input (block malicious or out-of-scope requests), on the actions (require approval for high-stakes tool calls, scope permissions), and on the output (catch unsafe, off-policy, or ungrounded responses before they reach the user). You set them by deciding in advance what the agent must never do, then enforcing those rules in code, not in the prompt alone. URL: https://chainofthought.show/ai-decoded/what-are-ai-agent-guardrails/ ## What is AI observability, and why do you need it in production? AI observability is instrumenting an AI system so you can see what it actually did on each request — the retrieved context, the tool calls, the intermediate reasoning, the final output — instead of just whether it succeeded or failed. You need it because AI systems are non-deterministic: the same input can behave differently, failures are silent, and a confident wrong answer looks identical to a right one. Without traces of the real behavior, you can't debug, you can't catch drift, and you can't tell a working system from one that's quietly breaking. URL: https://chainofthought.show/ai-decoded/what-is-ai-observability/ ## What is LLM-as-a-judge, and when can you trust it? LLM-as-a-judge uses one language model to score the output of another against a rubric — is this answer relevant, grounded, complete, safe. It scales evaluation past what humans can read by hand. You can trust it when you've calibrated it against human judgments on your own data, given it a concrete rubric, and kept a person in the loop for the high-stakes calls. Used blind, it inherits the same biases as the model doing the grading. URL: https://chainofthought.show/ai-decoded/what-is-llm-as-a-judge/ ## What is multimodal AI? Multimodal AI is a model that takes in and reasons across more than one kind of data — text, images, audio, video — in a single system. Instead of a separate model for each, one model can read a chart and answer questions about it, transcribe speech and act on it, or describe a video. The hard part isn't handling each modality; it's alignment — getting the model to connect what it sees, hears, and reads into one coherent understanding. URL: https://chainofthought.show/ai-decoded/what-is-multimodal-ai/ ## What is RAG, and why do AI systems use it? RAG, retrieval-augmented generation, is a pattern where the system fetches relevant documents at query time and hands them to the model along with the question, so the answer is grounded in real sources instead of the model's memory. It exists to fix two problems with a bare language model: it doesn't know your private or current data, and it makes things up when it doesn't know. RAG gives the model the right context to read before it answers. URL: https://chainofthought.show/ai-decoded/what-is-rag/ ## Why do most enterprise AI projects fail to show ROI? Most stall before they ever reach the scale where returns show up. The pilot demos well, then the project hits the costs nobody budgeted: evaluation, integration with messy real systems, data cleanup, governance sign-off, and the ongoing expense of running and monitoring the thing. Add a vague success metric — 'improve productivity' with no baseline — and you get projects that consume budget without producing a number anyone can point to. The failure is usually operational and organizational, not the model. URL: https://chainofthought.show/ai-decoded/why-enterprise-ai-fails-roi/ ## Why do some enterprises need to run AI on-premise? Because for regulated industries, the data can't leave the building. Sending prompts and documents to an outside AI provider means your sensitive data — patient records, financial data, regulated IP — crosses a boundary your compliance team can't allow. Running the models and the observability stack on-premise, behind your own firewall, keeps the data, the audit trail, and the control inside your perimeter. It costs more and is harder to operate, which is why it's a requirement for the regulated, not a default for everyone. URL: https://chainofthought.show/ai-decoded/why-enterprises-need-on-prem-ai/ ## Why do multi-agent systems fail, and how do you make them reliable? Multi-agent systems fail in the gaps between agents, not inside any one of them. Small per-agent errors compound: a handoff drops context, one agent's wrong output becomes another's trusted input, and a minor fault cascades into a systemic failure no single agent would have produced alone. You make them reliable by treating the system as the unit — tracing every step, validating what passes between agents, setting guardrails on autonomy, and threat-modeling how faults propagate before they reach production. URL: https://chainofthought.show/ai-decoded/why-multi-agent-systems-fail/ ## How much autonomy should you give an AI agent? As much as the risk of the task allows, and no more. There is no single right answer; you climb the ladder one step at a time and decide at each step whether a human still needs to sign off. URL: https://chainofthought.show/ai-decoded/ai-agent-autonomy-levels/ ## Are AI hallucinations always bad? No. A hallucination is the model generating something not grounded in fact, and whether that is bad depends entirely on the use. It is a feature for creative work and dangerous for anything factual, with the worst case being an answer that looks right but is wrong in context. URL: https://chainofthought.show/ai-decoded/are-ai-hallucinations-always-bad/ ## How do enterprises let employees use AI agents safely? Four guardrails: an allowed list of approved connectors, identity-based authentication, flags on destructive actions, and a human in the loop for anything risky. That is how Block runs AI agents across 12,000 employees at a company handling Square and Cash App. URL: https://chainofthought.show/ai-decoded/deploy-ai-agents-safely/ ## How do you evaluate an AI agent? You check it at three levels: the step (did it pick the right tool), the turn (did it do the steps in the right order), and the session (did the whole thing reach the right result). A single accuracy score hides all three, which is why agents that look fine in a demo fail in production. URL: https://chainofthought.show/ai-decoded/how-to-evaluate-ai-agents/ ## Should you use open source or proprietary LLMs? It depends on the job, and most serious teams use both: open when you need control, customization, privacy, or cost efficiency; proprietary when you need top-end quality on certain tasks or the easiest path to start. No one has won the race, so locking into one provider is the mistake. URL: https://chainofthought.show/ai-decoded/open-source-vs-proprietary-llms/ ## What is the difference between prompt, context, and memory engineering? They are three different jobs, and they happen in order. Prompt engineering is how you word the request. Context engineering is what you put in front of the model for a single task. Memory engineering is what the system keeps and reuses across tasks. URL: https://chainofthought.show/ai-decoded/prompt-vs-context-vs-memory-engineering/ ## What are the types of AI agent memory? An AI agent needs four kinds of memory, mapped to how the human brain works: working memory for what it is holding right now, semantic memory for the facts it knows, episodic memory for things that happened, and procedural memory for how to do a task. Most AI today runs on only the first one. URL: https://chainofthought.show/ai-decoded/types-of-ai-agent-memory/ ## What can MCP actually do? MCP lets an AI agent connect to your real tools and chain them together, which a plain chatbot cannot do. The value is not in any single connection but in wiring several systems into one workflow. URL: https://chainofthought.show/ai-decoded/what-can-mcp-do/ ## What is context in an AI agent? Context is everything you feed a model so it can actually do a task: documents, live web data, structured records, the tools it can call, and the systems where it stores and retrieves. As the models commoditize, the quality of that context is the part that compounds. URL: https://chainofthought.show/ai-decoded/what-is-context-in-ai-agents/ ## What makes an AI agent different from an LLM? An LLM answers; an agent does. The difference is four things built around the model: multiple models orchestrated together, memory and context, tools it can call, and a layer of checks running the whole time. The model is just one part of the system. URL: https://chainofthought.show/ai-decoded/what-makes-an-ai-agent/