What is data lineage, and why does an AI system need it?
Data lineage records where data came from, how it changed and which downstream datasets or systems used it. For AI, connect that history to model, dataset and pipeline versions so you can investigate a bad result, identify affected outputs and decide what needs to be rebuilt or retested.
Level 4: How AI systems are built · 4.2 Context and memory engineering
Context ManagementAI Infrastructure
Keep the history behind the result
In episode 22, Denny Lee describes a basic investigation problem: remembering which data, assumptions, model settings and architecture produced a result. Without that history, you may know the answer looks wrong but have little basis for deciding what changed.
Data lineage supplies the connections between sources, transformations and downstream uses. OpenLineage’s object model organizes that information around datasets, jobs and individual runs. A job describes work; a run records an execution of it. Inputs and outputs link the execution to the data it read and produced.
For an AI workflow, use those connections alongside application records. A retrieval trace tells you which passages reached a particular request. Lineage helps explain how those passages entered the collection and which preparation steps produced them. The records are useful together.
Follow one incorrect answer backward
Illustrative example: a support assistant cites a delivery guide with an incorrect regional deadline. Start with the answer record and identify the retrieved passage. Follow its document version to the parsing job, then to the source file that job read.
You might discover that the correct file was uploaded but an older extracted version remained in the index. Or the extraction could have attached a deadline to the wrong table row. Those are different repairs. Preserve enough history to distinguish them before replacing the answering model.
Then follow the same source forward. Which indexes, summaries or evaluation cases were built from it? That gives you a review list after correcting the document. Do not assume that replacing one source file automatically refreshes every derived artifact.
Record versions at the boundaries
For each ingestion or training run, record stable input and output identifiers, the relevant versions, the transformation configuration and the execution result. For an answer, retain references to the model version, prompt or application version and retrieved source versions where your system makes them available.
Maryam Ashoori gives the model-training version of this problem in episode 17: knowing which version of a model was trained on which version of data. She also calls out distinguishing synthetic and real data. Keep provenance explicit when you combine generated examples with reviewed human examples.
Choose a small workflow and test whether a colleague can reconstruct a result from the records. If they must guess which file a generic name refers to, improve the identifiers. If an external provider does not expose some history, record that limit rather than filling it with an assumption.
Where it falls short
Lineage records a chain of custody and transformation; it does not certify that the original data is true. A completely traceable answer can still rely on an incorrect document or misinterpret a correct one.
Nor does a dependency map prove that every downstream artifact is affected in the same way. Use it to scope inspection, then test the actual outputs. Avoid collecting sensitive prompt contents merely to make the graph more detailed when references or restricted records would serve the investigation. Continue with AI observability for the runtime view of what happened during a request.
Hear it from the guest
“… Where it came from … what it's going into.”
“I need to know exactly what version of what model was trained on what version of data.”
Quotes lightly edited to remove filler words.
Go deeper
- OpenLineage Object Model Explains jobs, runs, datasets and metadata facets as building blocks for recording data movement.
From the conversation
This explainer is drawn from these episodes — each carries its full transcript.