Best Chain of Thought Episodes on AI Evaluation
Measuring whether AI actually works — evals, reliability, trust.
Evaluation is the part of the stack everyone agrees matters and then postpones. It’s also what decides whether anything else you built can go in front of a customer. The order runs backwards here, from the evaluation gap every agent eventually hits to the practical starting point. None of these let you settle for a vibe check.
- 1 Every AI Agent Has an Evaluation Gap | Alex Ratner, Snorkel AI Alex Ratner, Snorkel AI Alex Ratner on the evaluation gap every team hits the moment an agent leaves the demo — and how to close it.
- 2 Explaining Eval Engineering | Galileo's Vikram Chatterji Vikram Chatterji, Galileo Vikram Chatterji on treating evaluation as an engineering discipline, not a one-time check.
- 3 The AI Agent Trust Gap: Bridging Risk to Reliability | Elastic’s Philipp Krenn Philipp Krenn, Elastic Philipp Krenn on bridging the gap from risky to reliable — why trust is the real bar for shipping agents.
- 4 AI in 2025: Agents & The Rise of Evaluation-Driven Development (ft. Dev Interrupted) Vikram Chatterji & Andrew Zigler The case for evaluation-driven development — building the eval loop into how you ship, not bolting it on after.
- 5 Practical Lessons for GenAI Evals | Chip Huyen & Vivienne Zhang Chip Huyen & Vivienne Zhang Chip Huyen and Vivienne Zhang on practical GenAI evaluation — the most-cited starting point for evals on the show.
- Best Chain of Thought Episodes for AI Founders Building an AI company — defensibility, GTM, and the market reality.
- Best Chain of Thought Episodes for Engineers Shipping with AI — the craft, the workflow, and the reality checks.
- Best Chain of Thought Episodes on AI Agents How agents actually work — architecture, memory, context, frameworks.
- Best Chain of Thought Episodes on Agent Memory How agents remember — memory stores, context engineering, and what breaks without them.
- Best Chain of Thought Episodes on Multi-Agent Systems Many agents, one system — orchestration patterns, protocols, and the overhead they add.
- Best Chain of Thought Episodes on Open Source AI Open weights, open ecosystems, and what they actually cost in production.
- Best Chain of Thought Episodes on AI Coding Agents What changes downstream when agents write most of the code.
- Best Chain of Thought Episodes for Enterprise AI Leaders Getting AI past the pilot — ROI, rollout, and the data underneath it.
- Best Chain of Thought Episodes on AI Security Attacking agents, defending them, and the part software cannot cover.
- Best Chain of Thought Episodes on RAG & Retrieval Getting the right context to the model: chunking, vectors, graphs, and what breaks at scale.