How do you use synthetic data to test an AI system?
Use synthetic data to draft test cases from trusted source material and known failure patterns, then have qualified reviewers check the questions, expected answers and coverage. Keep real-world cases and an independent holdout alongside generated tests so the generator does not become its own answer key.
Level 4: How AI systems are built · 4.4 Evaluation and observability in practice
AI Evaluation & ReliabilityModel Training
Give the generator something to work from
Synthetic test data consists of constructed cases rather than cases collected directly from users. It can help you explore a behavior before you have many examples of it. The hard part is deciding whether the generated case represents a real task and whether its expected answer is correct.
In episode 57, Alex Ratner argues for expert involvement in generating, curating and reviewing data. Asking a model to invent questions and certify its own answers leaves you without an independent check. Start with approved documents, reviewed examples or a failure you have already observed.
Generate situations, not just paraphrases
Illustrative example: a workplace equipment policy allows replacement after three years, with a separate approval process for accidental damage. Give the generator that policy and ask for cases that exercise different decisions: an ordinary replacement, an early request, a damaged device and a request with no purchase date.
For each case, ask it to propose a question, the relevant policy passage, an expected response and a reason the case belongs in the test set. A reviewer should verify all four. When the purchase date is missing, the expected response may be a clarification question; the generator should not invent a date to make the answer easy.
Ragas’s test-generation guide distinguishes questions answerable from one piece of information from questions that combine several. That is a useful starting structure for a retrieval test set. It does not establish that a generated answer is correct for your policy.
Review diversity and keep a separate check
In episode 43, Maxime Labonne points out that many generated variations can still come from the same process and have limited diversity. Inspect the decisions your cases exercise. Twenty rewrites of an ordinary replacement request do not cover the damaged-device exception.
Make a small coverage table before generating more rows. Include the intended decision, available evidence, missing information and expected next action. Add real failures as they become available, with appropriate data handling. Compare whether generated cases reflect the wording, ambiguity and conversation length that actual users bring.
Keep tests used for final evaluation apart from the cases used to tune prompts or train models. If a case becomes a development example, treat it as development material and preserve an untouched check elsewhere. Group near-duplicates so a reworded answer key does not cross that boundary.
Where it falls short
Synthetic data can help test a specified scenario, but the number of generated failures does not tell you how often that scenario occurs in the world. A deliberately difficult test set and a representative sample answer different questions. Label them accordingly.
Also inspect the scoring process. A correct question with an incorrect expected answer can punish the system for being right. When reviewers disagree, resolve the source interpretation or mark the case unresolved before using it as a gate. Testing nondeterministic AI covers what to do once you have trustworthy cases.
Hear it from the guest
“But, actually, it's the same data generation process. And because it's the same data generation process, it will limit the diversity.”
“You have to have sophisticated ways of having kind of human or expert in the loop, but AI-accelerated processes for all of the core operations of generating data, generating environments, curating those, reviewing those, et cetera.”
Quotes lightly edited to remove filler words.
Go deeper
- Testset Generation for RAG Shows how document relationships and different query types can guide test generation.
From the conversation
This explainer is drawn from these episodes — each carries its full transcript.