How do you check whether an AI answer is supported?
Break the answer into claims and compare each material claim with an authoritative source. Check dates, scope and qualifications, then verify calculations or recorded actions independently. A citation proves that a source was named; it does not prove that the source supports the answer.
Level 2: Using AI well · 2.3 Check its work
Check the claim, not the confidence
An answer can be fluent, relevant and unsupported. Start with the claim that changes a decision: an amount, deadline, eligibility condition or reported action. Find the actual passage or record and compare what it says with the answer. Do this before asking whether the answer sounds professional.
Worked example, with invented facts: policy P-17 allows refunds within 30 days. After that, an approved exception is required. The assistant replies, “You are eligible for an automatic refund through day 60,” and cites P-17.
| Answer claim | Evidence to inspect | Verdict in this example |
|---|---|---|
| The deadline is 60 days | The current policy’s deadline clause | Unsupported: the clause says 30 |
| Exceptions are automatic | The exception-approval clause | Unsupported: approval is required |
| The refund has been issued | The refund ledger | Unknown until the record is checked |
The citation is real. The interpretation is wrong. Keeping those failures separate tells you what to fix.
In episode 5, Chip Huyen separates whether the retrieved context is relevant and correct from whether the answer based on it is relevant and correct. The RAG evaluation explainer applies that distinction: receiving the right policy and using it correctly are separate checks.
Check scope and dates
A policy for one region may not apply to another. An old page can accurately describe a policy that has since changed. A transcript records what a guest said at the time; it does not establish that a vendor still uses the same model today. Ask which source has authority for this question, when it took effect, and whether its qualifications survive in the answer.
For a calculation, recompute the result from the supplied inputs. For an action, inspect the system of record. “I sent the message” needs a sent-message record; a draft or an accepted tool request is a different state. These are suggested verification steps, not claims that one source can settle every decision.
Use another model as a helper
A second model can identify claims to check or compare a passage with an answer. It can also repeat the first model’s mistake. Hamel Husain’s discussion in episode 37 connects useful evaluation to reading actual failures and comparing judge decisions with human labels. The judge explainer turns that into a calibration process.
For an individual answer, ask the helper to identify the supporting passage and any missing evidence. Read those passages yourself when the decision matters. For a repeated workflow, keep reviewed cases and track false approvals; a high agreement score can hide the very mistakes you need to catch.
Check yourself
An answer quotes the policy accurately, then applies it to the wrong customer category. Is it grounded? The quotation is supported, but the conclusion is not established for that customer. Check applicability as well as wording. Continue with RAG evaluation to diagnose whether the wrong context was retrieved or the right context was misused.
Go deeper
- Using LLM-as-a-Judge for Evaluation Shows how human judgments and observed failures shape a useful rubric rather than relying on a generic model score.
From the conversation
This explainer is drawn from these episodes — each carries its full transcript.