←
CAP Certification
Strategic · M20 · lesson 20 of 60 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Data Lineage and Provenance Tracking

15 min

Gabriel Santos is a data engineering lead at a global insurance company. When his company's AI underwriting model produced an unexpected spike in claim rejection rates for one policy type, the investigation team asked a single question first: where did the training data come from, and what happened to it between the source system and the model? Gabriel could not answer. The model had been built twelve months earlier by a team that had since restructured. The preparation scripts were version controlled, but the original data sources were not documented, the transformations had been applied in at least four stages, and nobody had recorded what the data looked like after each one. Eight weeks later the team found the cause, a data cleaning step that had introduced a systematic error in one feature, but those eight weeks were spent rebuilding a picture that should have existed from the first day of the project.

What Data Lineage Actually Means

Data lineage is the documented history of a piece of data: where it came from, what happened to it, and where it went. Provenance, a term borrowed from archival and scientific contexts, refers more narrowly to the origin and ownership history of data, establishing its authenticity and its chain of custody. The two are often used interchangeably in vendor material, but the distinction is useful in practice. Lineage tells you the route the data travelled inside your systems; provenance tells you whether what entered those systems was what it claimed to be, and whether you had the right to use it. Together they answer the only question that matters when something goes wrong: can we trust this data for this purpose, and can we prove it?

For AI systems that question is not abstract. The performance, fairness, and regulatory compliance of a model are only as reliable as the data it was trained on. If you cannot trace your data from source to model, you cannot confidently diagnose a failure, because you cannot separate a modelling problem from a data problem. You cannot demonstrate compliance, because compliance rests on evidence rather than assertion. And you cannot defend an automated decision to an auditor or an affected customer, because every explanation ends in the same place: we believe the data was correct, but we cannot show you how it was produced.

Think of lineage as the chain of custody document in a legal proceeding. Without an unbroken documented chain the evidence cannot be relied upon, not because it is wrong, but because nobody can verify it is intact. AI training data works the same way, and the parties asking the questions increasingly apply that standard. Gabriel's eight weeks were not spent fixing a bug; they were spent manufacturing, after the fact, evidence the pipeline should have produced continuously.

Why Lineage Breaks Down in Practice

Data lineage is rarely absent by design. It breaks down through a predictable accumulation of pragmatic shortcuts, each defensible when it is taken and expensive only much later. Recognising the patterns below in your own environment is more useful than any maturity score, because each has a cheap countermeasure if you apply it while the knowledge still exists in someone's head.

Undocumented source connections. Data pipelines are built quickly to meet a project deadline. The engineer building the pipeline knows exactly which source systems feed it, which tables, and which of the similarly named tables is the one actually maintained. None of that knowledge is written down, because it seems obvious to the person holding it and writing it down takes time the deadline does not contain. Six months later that engineer has moved to a different team, and the pipeline has become a black box that runs successfully every night and that nobody is willing to change.

Opaque transformations. Data passes through cleaning, filtering, normalisation, and feature engineering steps. The code exists and is usually version controlled, which creates a comforting but false sense that the transformations are documented. What the data actually looked like before and after each step, and what was decided about edge cases, is not recorded anywhere a non-engineer can reach. Reading the code tells you what the pipeline does today. It does not tell you what it did on the day the training set was assembled, and it does not tell you why.

Version drift. Source data changes over time, which is the entire point of an operational system. A customer record that was correct when the training set was assembled may since have been corrected, enriched, or superseded. Without a snapshot, a versioned extract, or at minimum a recorded extraction timestamp, there is no way to reconstruct the values the model actually trained on. Re-running the extraction query later gives you a dataset that looks superficially like the original, which is worse than having nothing, because it invites confident conclusions drawn from the wrong data.

Informal data sharing. An analyst exports a dataset to a spreadsheet, makes some modifications to fix an obvious problem, and emails it to the data science team. The attachment becomes the de facto training data source. There is no provenance and no audit trail, and the modifications exist nowhere except in the analyst's memory. This is the hardest pattern to detect, because the pipeline diagram still shows a clean path from the source system while the actual path went through an inbox.

Building a Lineage Framework

A data lineage framework for AI systems has four layers. You do not need to implement all four immediately, and most organisations should not try, but you do need to know which you have and which you are missing, because each absent layer removes a specific class of question from the set you can answer. The table below summarises what each layer records and what becomes impossible without it.

LayerWhat it recordsWhat you cannot answer without it
Source documentationSystems of record, owners, extraction method and timingWhere did this data originate, and who is accountable for its accuracy
Transformation loggingEach step applied, input and output state, and the rationaleWhat was done to this data, and why the non-obvious choices were made
Data quality gatesAutomated checks and their results at each pipeline stageWas the data within acceptable bounds at the point it entered the model
Visualisation and queryingNavigable maps of flows from source through transformation to modelWhich models does this source feed, and what is the full path back

Layer 1: Source Documentation

For every data source that feeds an AI system, document the system of record, the specific table or feed, the data owner, the extraction method and frequency, and the version or timestamp at which the data was extracted for each training or evaluation run. The data owner entry does the most work of these and is the one most often left blank. It should name the person or team accountable for that data's accuracy in the source system, not the team that built the pipeline, because when an investigation finds a suspect value the pipeline team can only tell you what they received.

This documentation belongs in a data catalog, a searchable inventory of data assets and their attributes. Tools such as Alation, Collibra, and Atlan provide catalog infrastructure, and most modern data platforms now include some catalog capability. For smaller organisations a well maintained spreadsheet is a perfectly viable start, and a better one than a procurement exercise that postpones all documentation until a platform is chosen. What matters is that the documentation exists, that it is maintained as pipelines change, and that people who are not data engineers can find and read it. The tool hosting it is a much smaller decision than it appears during selection.

Layer 2: Transformation Logging

Every transformation step applied to the data, including cleaning, filtering, joining, and feature engineering, should be logged with the transformation applied, the input data state or a reference to it, the resulting output state, the date and the person or process responsible, and the rationale for any non-obvious decision. Everything on that list except the rationale can and should be captured automatically by the pipeline framework. The rationale cannot, which is exactly why it is the element that disappears.

Consider two records of the same act. "We excluded records with null values in the risk score field" describes a transformation. "We excluded records with null values in the risk score field because preliminary analysis showed these records were disproportionately new accounts with insufficient history rather than random nulls" is a documented decision a future investigator can evaluate, challenge, and if necessary reverse. The second version also surfaces an assumption that may not hold in a later data period, which is exactly the kind of silent staleness that produces model drift nobody can explain. Version controlled code captures what was done; a decision log captures why. Neither substitutes for the other.

Layer 3: Data Quality Gates

Data quality gates are automated checks that run at defined points in the pipeline and halt it when the data falls outside agreed thresholds. They serve two purposes. They keep bad data away from the model, which is why they are usually funded, and they record the data's quality state at each checkpoint, which is why they matter for lineage. A gate that passes is not wasted computation; it is a dated assertion that the data met a stated standard at a stated point, and that assertion is evidence later.

Common gates include completeness checks, which measure what proportion of records carry non-null values for required fields; range checks, which confirm that numeric values fall within expected bounds; distribution checks, which compare the distribution of key features against a reference period; and referential integrity checks, which confirm that foreign keys resolve to valid records in the source system. Distribution checks are the ones most often skipped and the ones that catch the subtlest failures, because a transformation can produce values that are individually plausible and collectively wrong.

For Gabriel's insurance company the failure was eventually caught by a data quality gate, but the gate was written as part of the post-incident investigation rather than the original build. It found that the feature transformation was producing out-of-range values in roughly 0.4 percent of records. At the model's scale that was enough to systematically distort predictions in the affected policy category while leaving aggregate performance metrics looking healthy, which is why nobody noticed until the rejection rate moved.

Layer 4: Lineage Visualisation and Querying

The first three layers generate documentation and logs. The fourth makes that material usable by the people who need it under pressure: the engineer diagnosing a failure, the compliance officer answering a regulatory request, the data scientist checking whether a feature change will break something downstream. Documentation that cannot be queried in the shape the question arrives in will not be consulted, and a lineage programme nobody consults will not keep its funding.

Lineage visualisation capability, including the features built into modern data pipeline platforms such as dbt, Apache Atlas, and Microsoft Purview, produces navigable maps of data flows from source through transformation to model input. The test of whether the capability is real is whether it answers questions in both directions. Forward: if I change this source table, which models are affected? Backward: what is the complete path from this model's training data to its original source systems? Without querying in both directions, lineage documentation is a filing cabinet. With it, the documentation becomes a diagnostic instrument that shortens investigations rather than merely recording them.

Provenance for Third-Party and External Data

Many AI systems incorporate data from outside the organisation: purchased datasets, web-scraped content, public records, and partner data feeds. Third-party provenance needs extra attention precisely because you do not control the source system and cannot inspect it. Everything you know about that data you know because someone told you, so the record of what you were told, and when, is the only asset you have when the claim is later questioned.

For each third-party dataset, document the vendor or source organisation, the date of acquisition, the version or snapshot identifier, the licence terms governing permitted use, and any known limitations or quality issues the provider disclosed. That last item is worth chasing actively rather than accepting silence on, because a disclosed limitation you recorded is a defensible decision to proceed, while an undisclosed limitation you never asked about looks very different in hindsight.

Licence terms deserve particular emphasis. Many AI training use cases require rights that a standard data licence does not grant, and using data for model training when the licence permitted only analytics or research creates legal exposure that provenance documentation helps you find before the training run rather than after deployment. Whenever an existing dataset is proposed for a new AI purpose, keep the blunt version in mind: the data licence you agreed to before AI was part of your roadmap may not cover what you are now doing with the data. Document what you hold and check the terms before you train.

Making Lineage Work in Practice

The common failure mode in lineage programmes is building the documentation practice in isolation from the engineering workflow. If engineers have to separately document what they have just built, in a system they do not otherwise use, documentation lags and then atrophies, and the gap between the documented pipeline and the real one widens quietly until the documentation misleads. Embedded in the workflow, captured by the version control system, the pipeline framework, and the tools engineers already have open, it gets maintained as a by-product of the work rather than as a tax on it.

Two principles follow. The first is to automate whatever can be automated: transformation logging, quality gate results, and pipeline metadata should be captured by the pipeline itself, because manual documentation depends on discipline under time pressure and that is a losing bet over a long enough horizon. The second is to make lineage queryable within twenty-four hours of a production incident. If investigators wait longer than a day, they will build shadow systems, local copies, and personal spreadsheets to get answers faster, and those workarounds defeat the purpose of the lineage system while creating new undocumented data paths.

Gabriel's team spent three months after the incident building lineage infrastructure, an investment equivalent to roughly two engineer-months of effort. The first time a new model produced an anomalous output, eight months after the lineage system went live, the investigation team traced the cause to a change in a source system in ninety minutes. Measured against the earlier eight-week investigation, the infrastructure had paid for itself many times over on its first real use, and it will keep paying on every subsequent one.

Anti-Patterns

  • Treating version controlled code as lineage documentation. Code records what the pipeline does now. It does not record what the data looked like at each stage, what it looked like on the day a specific training set was built, or why an edge case was handled the way it was.
  • Documenting the pipeline team as the data owner. The pipeline team can tell you what they received. Accountability for accuracy belongs with the owner of the source system, and recording the wrong name sends every investigation to the wrong desk.
  • Logging transformations without rationale. The what is captured automatically and is the easy half. The why exists only in the head of the person who made the call, and it leaves when they do.
  • Adding quality gates only after an incident. A gate written during an investigation confirms the diagnosis. A gate written during the build would have prevented the deployment that required the investigation.

Practice Prompts

  • Pick one AI system in production and name, from memory, every source system that feeds it. Verify the list against the pipeline code and note the difference.
  • Take the most recent training run for that system and attempt to establish which version of the source data it used. Record how long the attempt took and where it stopped.
  • List the automated quality gates in one pipeline, classify each as completeness, range, distribution, or referential integrity, and identify which of the four categories is missing entirely.
  • Locate the licence for one external dataset in use, read the permitted use clause, and write one sentence stating whether it covers model training.

Reflection

Gabriel's team was not careless. Every shortcut behind their eight-week investigation was a reasonable decision made by a competent engineer against a real deadline, and each was cheaper than the alternative at the moment it was taken. The cost arrived later, concentrated, and landed on people who had made none of those decisions. Consider the pipelines your organisation depends on and ask which of them exist mainly in someone's memory. Then ask the harder question: if that person left, would you discover what you had lost before or after the next incident?

Glossary

  • Data lineage: The documented history of a piece of data, covering where it came from, what was done to it, and where it went.
  • Provenance: The origin and ownership history of data, establishing its authenticity and chain of custody, including the terms under which it was obtained.
  • Data catalog: A searchable inventory of data assets and their attributes, including systems of record, owners, and extraction details.
  • Decision log: A record of why non-obvious transformation choices were made, maintained alongside the version controlled code that records what was done.
  • Data quality gate: An automated check that runs at a defined pipeline stage and halts the pipeline when data falls outside agreed thresholds.
  • Distribution check: A quality gate comparing the distribution of a feature against a reference period, catching failures where individual values remain plausible.
  • Enterprise Data Architecture & Governance sets the structural context that determines how many source systems a lineage programme has to cover.
  • Data Quality & Master Data Management develops the quality dimension that Layer 3 gates measure and enforce.
  • Model & Data Integrity extends the chain of custody argument from training data through to the model artefacts themselves.
  • Documentation, Audit & Regulators covers what external parties ask for and the form the evidence has to take when they do.

Closing

Lineage work is unglamorous, and its value stays invisible right up to the moment it becomes the only thing that matters. Nobody is promoted for the investigation that took ninety minutes instead of eight weeks, because the eight-week version never happened and so never appeared in a report. That asymmetry is why these programmes stay underfunded until an incident forces the issue, and why the argument for building one has to be made in advance, in terms of the questions the organisation will eventually be required to answer. The organisations that handle their first serious model failure well are rarely the ones with the best models; they are the ones that can show, without a reconstruction project, exactly where their data came from.

Key Takeaways

  • Data lineage is the documented chain of custody for AI training data. Without it, failures cannot be diagnosed confidently, compliance cannot be demonstrated, and automated decisions cannot be defended.
  • Lineage breaks down through predictable shortcuts: undocumented source connections, opaque transformations, version drift, and informal data sharing. These accumulate through time pressure rather than negligence, so build documentation into the workflow.
  • A complete framework has four layers: source documentation held in a catalog, transformation logging that includes rationale, automated quality gates at each pipeline stage, and visualisation and querying that make the first three usable.
  • Document the why, not just the what. Version controlled code records the transformations applied. A decision log records why non-obvious choices were made, and future investigators need both to trust the data.
  • External data requires provenance documentation that includes licence terms. Data licensed for analytics may not be licensed for AI training, and the time to find that out is before the training run.
  • Automate lineage capture wherever possible. Manual documentation fails under time pressure, so transformation logs, gate results, and pipeline metadata should be captured by the pipeline and made queryable within twenty-four hours of any production incident.

Frequently Asked Questions

We have a large estate of pipelines. Where do we start? Start with the AI systems whose failure would be most expensive to explain, not with the pipelines that are easiest to document. For those systems complete Layer 1 first, because source documentation is cheap, requires no tooling decision, and immediately answers the question that opens every investigation. Layers 2 and 3 can then be applied to new pipeline work as a standard rather than retrofitted across the estate, which spreads the cost over normal delivery instead of requiring a separate programme.

Is a spreadsheet really acceptable, or is that just avoiding the problem? It is acceptable as a starting point and dangerous as an ending point. A maintained spreadsheet answers real questions today and costs nothing to begin, which is why it beats a catalog decision that has not yet been made. It fails at Layer 4, because it cannot say which models a given source feeds without a person reading the whole thing, and that limit is the signal to move to catalog infrastructure.

Who should own the lineage programme? Ownership works best where the pipelines are built, because lineage captured as a by-product of engineering work survives and lineage imposed from outside it does not. Governance and compliance functions are the customers of that output rather than its producers. What they should own is the standard: which systems require which layers, and what evidence has to exist before a model goes into production.