AI, decoded

How do you check whether an AI answer is supported?

Break the answer into claims and compare each material claim with an authoritative source. Check dates, scope and qualifications, then verify calculations or recorded actions independently. A citation proves that a source was named; it does not prove that the source supports the answer.

· Chain of Thought

Level 2: Using AI well · 2.3 Check its work

AI Evaluation & Reliability

Check the claim, not the confidence

An answer can be fluent, relevant and unsupported. Start with the claim that changes a decision: an amount, deadline, eligibility condition or reported action. Find the actual passage or record and compare what it says with the answer. Do this before asking whether the answer sounds professional.

Worked example, with invented facts: policy P-17 allows refunds within 30 days. After that, an approved exception is required. The assistant replies, “You are eligible for an automatic refund through day 60,” and cites P-17.

Answer claimEvidence to inspectVerdict in this example
The deadline is 60 daysThe current policy’s deadline clauseUnsupported: the clause says 30
Exceptions are automaticThe exception-approval clauseUnsupported: approval is required
The refund has been issuedThe refund ledgerUnknown until the record is checked

The citation is real. The interpretation is wrong. Keeping those failures separate tells you what to fix.

In episode 5, Chip Huyen separates whether the retrieved context is relevant and correct from whether the answer based on it is relevant and correct. The RAG evaluation explainer applies that distinction: receiving the right policy and using it correctly are separate checks.

Check scope and dates

A policy for one region may not apply to another. An old page can accurately describe a policy that has since changed. A transcript records what a guest said at the time; it does not establish that a vendor still uses the same model today. Ask which source has authority for this question, when it took effect, and whether its qualifications survive in the answer.

For a calculation, recompute the result from the supplied inputs. For an action, inspect the system of record. “I sent the message” needs a sent-message record; a draft or an accepted tool request is a different state. These are suggested verification steps, not claims that one source can settle every decision.

Use another model as a helper

A second model can identify claims to check or compare a passage with an answer. It can also repeat the first model’s mistake. Hamel Husain’s discussion in episode 37 connects useful evaluation to reading actual failures and comparing judge decisions with human labels. The judge explainer turns that into a calibration process.

For an individual answer, ask the helper to identify the supporting passage and any missing evidence. Read those passages yourself when the decision matters. For a repeated workflow, keep reviewed cases and track false approvals; a high agreement score can hide the very mistakes you need to catch.

Check yourself

An answer quotes the policy accurately, then applies it to the wrong customer category. Is it grounded? The quotation is supported, but the conclusion is not established for that customer. Check applicability as well as wording. Continue with RAG evaluation to diagnose whether the wrong context was retrieved or the right context was misused.

Go deeper

From the conversation

This explainer is drawn from these episodes — each carries its full transcript.

Concepts in this explainer

Retrieval-Augmented Generation (RAG)LLM as a Judge