Data Infrastructure for Enterprise AI
Tom Becker, a program director at a county health department, green-lit a promising AI project to predict which clinics would run short on vaccines. The model design was sound. The pilot died anyway. The reason had nothing to do with AI: the data lived in eleven different systems, three of them spreadsheets emailed monthly, one a vendor database nobody had login credentials for. His data scientist spent four of the six pilot months just finding and cleaning data, and gave up before building a single model. Tom's hard lesson: "We didn't have an AI problem. We had a plumbing problem, and AI just made it visible."
That is the truth behind almost every stalled government AI effort. The model is the visible 10 percent. The other 90 percent is data infrastructure, the unglamorous plumbing that gets the right data, clean and current, to the place where a model can use it. Enterprise AI in government is ultimately a data problem. A capable model trained on poor, ungoverned or incomplete data produces decisions the agency will regret; a modest model trained on clean, well governed, well understood data produces decisions the agency can defend. The difference is rarely the algorithm.
As a program director you do not need to architect storage systems yourself. You need to understand the building blocks well enough to ask the right questions, fund the right things, and know when your "AI project" is actually a data project wearing a costume. You also need to know why federal data infrastructure is not commercial data infrastructure. Commercial firms assemble pipelines largely on their own terms. Public agencies build inside a stack of statutory, regulatory and policy constraints, and an engineer who does not know that stack will build something that cannot be authorized.
The three shapes data is stored in
Vendors throw three terms at you, often in the same sentence and often as if they were synonyms. They are not interchangeable, they solve different problems, and confusing them wastes money in a particular way: you buy the shape that fits the demo rather than the shape that fits the workload. A program director who can hold these three apart can follow almost any infrastructure conversation, because nearly every design argument in this space is really an argument about which of the three a given workload belongs in.
- Database. The workhorse that runs a live system, the one recording vaccine doses as they are given. Optimized for fast, small reads and writes. Great for running operations, poor for big analytical questions across years of history.
- Data warehouse. A separate store built for analysis. Data is cleaned, structured and organized so you can ask "how did vaccine demand trend across all clinics last winter?" without slowing down the live system. This is structured data with a fixed shape.
- Data lake. A large pool that holds raw data in many forms, structured tables, free-text notes, PDFs, images, before anyone decides how to use it. Flexible and cheap to fill, but without discipline it becomes a "data swamp" no one can navigate.
The plain-English rule: databases run the agency, warehouses help you understand it, and lakes hold everything until you figure out what you need. Modern AI usually draws from a warehouse or a well-governed lake, not directly from the live databases, which it would overwhelm. Those three shapes are the raw material. What turns them into infrastructure an agency can operate is the architecture wrapped around them.
The reference architecture, layer by layer
A defensible federal enterprise AI data architecture has seven layers, each hosted at an authorization level appropriate to the data it holds. Layer one is ingestion, bringing data from authoritative source systems into the AI environment under documented data use agreements and coverage by the relevant system of records notice. Layer two is the raw data lake, often called the bronze layer, storing data in its original form with immutable logs of what arrived and when. Layer three is the curated warehouse, the silver layer, where data is cleaned, conformed and governed with schemas, access controls and documented provenance.
Layer four is the data catalog, which makes everything in bronze and silver discoverable, carrying metadata about owner, lineage, quality, sensitivity and authorized uses. Layer five is the feature store, which exposes model-ready features consistently across training, online serving and batch inference. Layer six is the model training and evaluation harness, which takes features, runs training and captures evaluation metrics including accuracy, fairness and calibration; it usually integrates with an experiment tracking tool. Layer seven is serving, where the model is exposed to calling applications at the latency and throughput the use case requires, and serving must support model versioning, shadow deployment, rollback and audit logging.
Cutting across all seven layers are four concerns that cannot be assigned to any one of them. Security is the control implementation required by federal information security law and the associated control catalog at every layer. Privacy is compliance with the Privacy Act and agency-specific statutes, with records-notice coverage and computer matching agreements executed where needed. Quality is the data and feature validation pipeline that catches drift before it corrupts decisions. Observability is the metric, log and trace infrastructure that lets engineers diagnose failures. An agency can assemble this from authorized building blocks; it does not need to build everything from scratch. The design choice is which components to use, how to integrate them, and what governance to wrap around them.
The pipeline is the real product
Storage is where data rests. The pipeline is how it moves, the automated process that pulls data from source systems, cleans it, reshapes it, and delivers it where it is needed, on a schedule. Tom's pilot failed because his "pipeline" was a human emailing spreadsheets once a month. A real pipeline does this automatically, every night, and flags itself when a source breaks. The hard part is not moving data; it is keeping it trustworthy. A model is only as good as its freshest, cleanest input, and three pipeline qualities decide whether your AI can be trusted at all.
- Freshness. How current is the data when the model uses it? A vaccine-shortage model fed month-old data predicts last month's shortage.
- Completeness. Are records missing? If three clinics never report, the model is blind to a quarter of the county.
- Consistency. Do "Pfizer," "PFIZER," and "pfizer-biontech" all mean the same thing? Inconsistent values silently corrupt results.
An AI model is a very expensive way to discover that your data was broken all along. Fix the pipeline and a large share of what looked like AI problems turn out to have been data problems the model merely exposed.
Batch or streaming: how often the data has to move
One architectural question deserves an explicit answer before anyone chooses tools: does the mission decision happen on a schedule or continuously? Batch delivery, where the pipeline runs nightly or hourly and the model scores a whole set of records at once, is simpler to build, cheaper to run and easier to audit, because every run has a clean boundary. Tom's vaccine forecast is a batch problem; a daily refresh is more than current enough for a supply decision measured in weeks.
Streaming and real-time infrastructure is a different build. Event streaming platforms and agency-grade event mesh patterns exist to move records continuously as they are created, and they are the right answer when the decision cannot wait for the next scheduled run. They also cost more, fail in more ways, and complicate lineage, because there is no tidy batch boundary to point an auditor at. The discipline is to justify streaming from the decision, not from the architecture diagram. Most government AI use cases are batch problems that were quoted as streaming problems.
Catalog, lineage and provenance
Without a data catalog, nobody knows what data exists, and AI projects repeatedly rediscover or duplicate datasets that already sit in the estate. The catalog is the layer that answers "what do we have, who owns it, how good is it, how sensitive is it, and what are we allowed to do with it." It is the cheapest layer to stand up and the one agencies skip most often, because its payoff is measured in work that never has to happen twice.
Lineage and provenance are the audit trail. For any model decision that affects a citizen, the agency must be able to answer three questions: what data was used to train the model, what data was used at inference time, and how the training-to-inference distribution has changed over time. Without lineage infrastructure these cannot be answered except by labor-intensive manual reconstruction, which is slow, error-prone and indefensible under audit. Open lineage standards now exist for exactly this, and the common orchestration tools can emit lineage automatically as pipelines run. Model lineage extends the same idea to the model artifact, tracking which code, data and hyperparameters produced each model version.
Provenance is related but includes external context: which authoritative source produced the data, what agreements govern its use, and what statutes apply. Provenance is what an inspector general or a government auditor asks for when they want to understand whether an AI decision rested on lawful data. Maintain provenance metadata in the catalog layer, with explicit links between each dataset and the agreements, records notices and statutes that authorize its use. Infrastructure that cannot trace a result back to its sources will fail an audit under federal accountability expectations such as the GAO AI Accountability Framework.
Data quality and silent drift
Data quality is the active practice of detecting issues before they corrupt decisions. A federal data quality pipeline runs automated checks on every flow: schema conformance, value ranges, uniqueness constraints, referential integrity, and statistical drift from baseline distributions. Open-source validation frameworks and commercial data observability products both serve this purpose, and the choice matters far less than whether the checks actually run on a schedule with someone accountable for the alerts.
Drift deserves separate attention because it defeats every check that looks only at structure. Data can conform perfectly to schema while the distribution of values shifts underneath, and model performance degrades silently as a result. The source guidance for this lesson recommends alerting on drift at the 95th percentile or higher, with clear escalation paths tied into the agency's incident response. Treat that as a starting threshold to calibrate against your own baseline rather than a universal setting, and be honest about the limit of the whole discipline: a quality pipeline catches the failures you configured it to look for, and nothing else.
Feature stores and training-serving skew
A feature store is the layer that ensures consistency between training-time and inference-time data. Without one, a data scientist computes "average income over the last twelve months" one way during training, an engineer computes it a slightly different way at inference, and the result is silent training-serving skew that can substantially degrade model performance. The feature store centralizes the definition and the execution of each feature, so training and inference run the same code path. Open-source and commercial implementations both exist, and in federal environments the feature store must sit inside the authorization boundary, integrate with the agency catalog, and inherit the same access controls as the underlying data.
This is worth dwelling on because the failure is so quiet. There is no error, no outage, no alert. The model simply performs worse in production than it did in testing, and the team spends weeks suspecting the algorithm. Ensuring that the feature a model was trained on is identical to the feature it sees at inference removes one of the most common causes of silent model failure, and it is an infrastructure fix rather than a modeling fix.
The evaluation harness
The model evaluation harness is the layer that tells the agency whether a model is good enough to ship. A federal-grade harness includes offline evaluation on a held-out test set, backtesting on historical data, fairness evaluation with disaggregated metrics across protected classes, calibration analysis, robustness testing for distributional shift, and shadow deployment where the new model runs alongside the current one without affecting decisions. Results feed into the agency AI governance board's review and the Chief AI Officer's sign-off. Federal AI policy requires documented evaluation for rights-impacting uses, and the international AI management system standard adds ongoing operational evaluation expectations.
A common trap is to treat evaluation as a pre-production task only. Continuous evaluation matters just as much: recompute key metrics daily or weekly on recent production data, with alerts when they diverge from acceptance thresholds. That is how an agency catches model drift, data shift or a subtle regression before it produces harm. Be precise about what the harness buys you, though. An industrialized harness lets a team ship changes faster with better evidence. It does not certify that a model is safe, fair or correct; it tells you how the model behaved on the tests you chose to run, on the data you happened to have.
Government data standards: the rules that travel with the data
Public-sector data carries obligations a private company never faces, and the constraints are specific rather than general. The Privacy Act of 1974 and OMB Circular A-108 constrain how personally identifiable information moves through federal systems. The Evidence-Based Policymaking Act of 2018 and the DATA Act require particular data standards and spending transparency. Federal information security law and its control catalog impose security control families across every data layer. Agency-specific statutes narrow things further: 26 USC 6103 for tax data, HIPAA for health data. Three of these obligations bear directly on infrastructure design.
Classification and sensitivity. The same lake may hold public statistics and protected personal records. Your infrastructure must keep them separated, control who can reach the sensitive tiers, and log every access, because the log is what you will be asked for. You cannot pour regulated data into a lake and sort it out later; once mixed, the whole store inherits the sensitivity of its most protected record, and every subsequent access decision gets made at that level.
Retention and disposition. Records laws dictate how long data must be kept and when it must be destroyed. Your pipeline and storage need to enforce those schedules rather than fight them, which is harder than it sounds once a lake holds copies, derived tables and model training snapshots of the same underlying records. A disposition obligation attaches to the data, not to the original system, and it follows every copy the pipeline made.
Provenance and quality for accountability. When an AI decision affects a citizen, you must be able to show what data drove it and where that data came from. This is the same requirement lineage serves, restated as a legal obligation rather than an engineering convenience, and that restatement is what makes it fundable. Engineering conveniences lose budget arguments. Accountability obligations that an auditor will test do not.
Federal data governance for AI
Data governance is the unglamorous discipline that determines whether the architecture actually serves the mission. The baseline federal stack includes the interagency Chief Data Officer Council, the Federal Data Strategy published by OMB, the agency Chief Data Officer required by the Evidence-Based Policymaking Act of 2018, the agency Data Governance Board, and the Senior Agency Official for Privacy. For AI specifically, OMB M-24-10, the federal memorandum on advancing AI governance that agencies work under, adds obligations including documented data provenance, data minimization, and disaggregated monitoring for rights-impacting uses.
A strong agency AI data governance program runs five practices. First, a canonical data inventory maintained by the Chief Data Officer's office and cross-referenced to authoritative sources and agency records notices. Second, a Data Governance Board that reviews every new data use for legal basis, privacy impact, security posture and mission alignment. Third, a documented data minimization discipline, where AI projects must justify any personal or controlled unclassified information they ingest and confirm that no less sensitive alternative would serve. Fourth, disaggregated quality and performance monitoring for any rights-impacting use. Fifth, a data incident response plan naming specific actions when quality or privacy issues are detected, with coordination paths to the agency security office, privacy office and Chief AI Officer.
Cross-agency work amplifies all of this. The Chief Data Officer Council shares templates and practices; the Federal Data Strategy provides a shared frame. Agencies that participate actively avoid reinventing governance, while agencies that go it alone spend years on problems their peers already solved. Before designing anything, find out which interagency resources apply to your data domain and where your agency plugs in.
The hosting decision, framed for government
Tom's next question was where to put all this. The realistic government choice is rarely "cloud or not," it is which cloud and how, and the decisive factor is authorization. A cloud service that has already earned a federal security authorization through FedRAMP, the government-wide program that vets cloud providers for handling federal data, lets you stand on completed accreditation instead of repeating a one-to-two-year approval yourself. Be precise about what that buys: an authorization covers the assessed service's security posture within a defined boundary at a stated impact level. It does not authorize your system, your data use, or your model.
| Option | Best when | Watch out for |
|---|---|---|
| Authorized commercial cloud | You want scale, managed services, and fast standup on data the authorization covers | Confirm the authorization level matches your data's sensitivity |
| Government or agency cloud | Highly sensitive or classified data with stricter isolation needs | Fewer cutting-edge AI services; higher cost |
| On-premises | Legal mandate to keep data in your own facility | You own all the scaling, patching, and capacity pain |
| Hybrid | Sensitive data stays controlled, analysis uses cloud scale | Complexity at the boundary; data movement must be governed |
Impact level is the part people get wrong. FedRAMP Moderate is the common baseline for unclassified, non-sensitive data. The source guidance for this lesson states that FedRAMP High is required for controlled unclassified information, including personally identifiable information in many categories, and that Department of Defense Impact Level 5 is required for certain defense-sensitive unclassified workloads. Classified workloads require accredited environments outside the civilian marketplace entirely. Verify the level your own data owner and authorizing official require rather than assuming; the cost difference between tiers is real, and so is the consequence of guessing low.
The authorization inheritance model lets an agency reuse another agency's authorization for a specific service, cutting duplicate approval work. Inheritance requires the receiving agency to review the control implementations, accept the residual risks, and add agency-specific supplements where needed. It shortens the path; it is not a substitute for independent review. Agencies that set up authorization leveraging as a standing practice move faster than agencies that re-authorize every service from zero.
Classification boundary design matters especially for AI, because AI pipelines routinely combine data from sources that live at different levels. The discipline is to identify the highest classification any data will touch, architect the pipeline to that level, and provide explicit cross-domain solutions where lower-classification data flows up or higher-classification results must flow down with appropriate redaction. Moving data from High to Moderate is a potential spill of controlled unclassified information; the source guidance further states that moving data from Impact Level 5 to Impact Level 4 requires formal cross-domain solutions. Plan these boundaries at the architecture stage. Retrofitting them is the expensive way to learn this.
One more trap: the common mistake is matching the whole estate to your most sensitive data, paying premium isolation costs for public statistics that need none. Tier your data and match infrastructure to each tier.
Diagnosing which layer is actually weak
The reason to learn the layers is diagnostic. Weak data infrastructure does not announce itself as weak data infrastructure; it shows up as a specific, recognizable symptom, and each symptom points at a layer. Learn to read them backwards and you can fund the right repair instead of the loudest complaint.
When teams keep rediscovering or rebuilding datasets that already exist somewhere in the estate, the catalog is missing. When nobody can explain where a model's training data came from, and answering takes weeks of manual reconstruction, lineage is missing, and that is the finding that lands hardest in an inspector general or government auditor review. When every AI team invents its own version of the same input, producing inconsistent results across systems and bias nobody can locate, the feature store is missing. When silent data issues produce silent model degradation that surfaces only because a beneficiary complains, the quality pipeline is missing. When shipping a new model version feels like a leap of faith, the evaluation harness is missing.
Each of those is a concrete audit finding waiting to happen, and each is addressable with established tools and design patterns rather than research. That is the encouraging half of the message. The discouraging half is that none of them are visible from a project status report, which is why the diagnosis has to be run deliberately rather than waited for.
A data-readiness checklist before you fund a model
Tom now runs this gate before approving any AI project, and he runs it in a room with the people who would have to answer for each item rather than on his own. The point is not to score the proposal. It is to surface, before money moves, exactly which work has to happen and who owns it. If three or more answers are "no," it is a data project, and the model waits.
- Sources known? Can we name every system the data lives in and reach all of them?
- Automated pipeline? Does data flow on a schedule without a human emailing spreadsheets?
- Fresh enough? Is the data current enough for the decision the model supports?
- Complete and consistent? Do we know the gaps, and are values standardized?
- Properly classified and access-controlled? Is sensitive data separated, logged, and reachable only by the right people?
- Traceable? Can we trace any output back to the data and sources behind it?
- Authorized environment? Does the storage and compute environment carry the right security authorization for this data?
When Tom reran his vaccine project through this gate, six of the seven answers were "no." He spent the next quarter on plumbing: one automated pipeline, one tiered store, standardized vaccine names. The model that had been impossible took his data scientist three weeks the second time around. That is the whole argument for infrastructure in one anecdote. The investment did not make the model better. It made the model possible.
Anti-Patterns
- Funding the model before the plumbing. The most expensive version of this failure is the one that looks like success at the approval meeting. A model is easy to scope and easy to demo; a pipeline is neither. Agencies approve the visible 10 percent, discover the other 90 percent in month two, and spend the pilot budget on data archaeology. Run the readiness gate first and be willing to say the sentence out loud: this is a data project, and the model waits.
- Treating a security authorization as a data-use permission. An authorization speaks to the security posture of an assessed environment within a defined boundary at a stated impact level. It does not say what data you may put there, what your model may do with it, or whether your use is lawful. Data classification, records-notice coverage, minimization and use agreements are separate determinations by different officials, and each has to be made.
- Believing a quality pipeline means the data is good. Automated checks catch schema violations, out-of-range values and drift from a configured baseline. They do not catch data that is well formed and wrong. A field can be perfectly typed, perfectly populated and systematically incorrect because an upstream system changed a business rule nobody told you about. Green checks mean the checks passed, not that the data is right.
- Building the lake and calling it governance. Filling a lake is cheap, which is why it happens first and why it so often becomes a swamp. Without a catalog naming owners, sensitivity and authorized uses, and without lineage recording where everything came from, the lake is a liability with a storage bill. Stand up the catalog alongside the lake, not two years later under audit pressure.
- Letting every team invent its own features. Without a feature store, each AI team defines "income," "household" or "case age" its own way. The result is inconsistency across systems, quiet disagreement between training and serving, and hidden bias nobody can locate because no two definitions match. This is an infrastructure problem masquerading as a modeling problem for months at a time.
- Evaluating once, at acceptance. A pre-production evaluation describes a model on the day it was tested against the data you had. Conditions move. Without continuous evaluation on recent production data and thresholds that trigger alerts, drift accumulates unobserved and surfaces when a beneficiary complains or an auditor asks. The harness is a standing capability, not a milestone.
- Matching the whole estate to the most sensitive record in it. Paying premium isolation costs for open statistics is a budget failure dressed as caution, and it usually crowds out the catalog and lineage work that actually reduces risk. Tier the data, match infrastructure to each tier, and document why each tier sits where it does.
- Deferring lineage until an auditor asks. Reconstructing training data provenance after the fact is slow, error-prone and indefensible. If lineage is not captured automatically as pipelines run, it is not captured. Manual reconstruction produces a document, not evidence.
Practice Prompts
- Run the readiness gate on a live proposal. Take an AI project your agency is currently considering and answer all seven checklist questions in writing, with the name of the person who can confirm each answer. Count the "no" responses. Then write the one-paragraph recommendation you would actually deliver, including what the data work would cost in time before any model work begins.
- Map your seven layers. For one existing or planned AI use case, write down what currently plays the role of each layer: ingestion, raw store, curated store, catalog, feature store, training and evaluation harness, serving. Mark any layer whose answer is "a person does this manually" or "nothing." Those marks are your investment list, in order.
- Trace one decision backwards. Pick a single output from an AI or analytics system your agency runs and try to trace it to the exact datasets, versions and source systems behind it, plus the agreements or statutes authorizing each. Time how long it takes. Whatever that number is, it is also roughly what an audit response would cost you.
- Tier your data and price the difference. List the data categories one AI system touches, assign each to a sensitivity tier, and identify which tier actually drives your hosting decision. Then ask what you are paying to protect the categories that do not need it, and whether any of that budget belongs in catalog or lineage work instead.
- Design the drift alert. For one model, define the baseline distribution you would compare against, the checks you would run, the threshold that would fire an alert, who receives it, and what they are authorized to do in the first hour. Then list, honestly, the failure modes this alert would not catch.
Reflection
Think about the last AI effort in your organization that stalled or quietly disappeared. How much of the delay was modeling and how much was finding, cleaning, joining or getting permission to use data? If you had to name, today, every system holding data your next AI project needs, could you, and could you reach all of them? Who in your agency would be able to answer an auditor asking where a model's training data came from, and how long would they need? Which of the seven layers does your agency actually operate rather than improvise? And if a model you already run had been degrading quietly for six months, what specifically would have told you?
Glossary
- Data pipeline. The automated process that pulls data from source systems, cleans and reshapes it, and delivers it on a schedule to where it is needed, flagging itself when a source breaks.
- Data warehouse. A store built for analysis rather than operations, holding cleaned and structured data with a fixed shape so analytical questions do not slow the live system.
- Data lake. A large pool holding raw data in many formats before anyone decides how to use it. Cheap to fill; without a catalog and governance it degrades into a data swamp.
- Data catalog. The inventory layer recording what data exists, who owns it, where it came from, how good it is, how sensitive it is, and what uses are authorized.
- Lineage. The recorded path data travelled from source to model, captured automatically as pipelines run. Model lineage extends this to the code, data and hyperparameters behind each model version.
- Provenance. Lineage plus external context: which authoritative source produced the data, which agreements govern its use, and which statutes apply.
- Feature store. The layer that centralizes the definition and computation of model inputs so that training and inference use the same code path, preventing training-serving skew.
- Training-serving skew. The silent failure where a feature is computed one way in training and slightly differently at inference, degrading production performance with no error or outage.
- Drift. A shift in the distribution of input values over time. Data can pass every schema check while drifting, which is why drift needs its own detection.
- Evaluation harness. The standing capability that runs offline evaluation, backtesting, disaggregated fairness metrics, calibration, robustness testing and shadow deployment before and after a model ships.
- Shadow deployment. Running a new model alongside the current one on live traffic without letting it affect decisions, so its behavior can be compared before cutover.
- FedRAMP. The government-wide program that assesses and authorizes cloud services for handling federal data at defined impact levels, within a defined authorization boundary.
- Authorization inheritance. Reusing another agency's authorization for a service, after reviewing its control implementations, accepting residual risk, and adding agency-specific supplements.
- Cross-domain solution. The controlled mechanism for moving data between environments at different classification or impact levels, including redaction where higher-level results must flow down.
Related Lessons
This lesson sits between the pilot and the enterprise. Moving from Pilot to Production covers the operational transition this infrastructure has to support, and Scaling Capstone: From Your Pilot to Enterprise puts the whole progression together. Data Governance for AI goes deeper on the governance bodies and practices sketched here, while Data Quality and AI Performance and Data Sensitivity and Classification expand the quality and tiering material. FedRAMP and AI Cloud Authorization is the full treatment of the hosting and authorization decision, and Legacy System Modernization for AI addresses the source systems that make ingestion hard in the first place. For what happens once the architecture exists, see AI Metrics and KPIs for Government, Enterprise AI Risk Management, and Cross-Agency AI Coordination.
Closing
Data infrastructure is the least visible and most decisive part of government AI. It does not demo well. It does not appear in the press release. It is, however, the reason one agency ships a model in three weeks while another spends four months on data archaeology and gives up. The seven layers, the catalog and lineage, the quality checks and the authorization boundaries are not bureaucratic overhead sitting between you and the model. They are the conditions under which a model can exist at all in a public agency.
Tom Becker's second attempt succeeded for an unglamorous reason: he stopped treating the plumbing as something to work around and started treating it as the project. If you take one habit from this lesson, take the gate. Ask the seven questions before the money moves, count the "no" answers honestly, and be willing to fund the data project that is hiding inside the AI project. That single discipline will save more programs than any modeling decision you make.
Key Takeaways
- Most AI projects fail on plumbing, not models. Data infrastructure is the 90 percent of the work that decides whether the visible 10 percent can run at all.
- Know the three storage shapes. Databases run the agency, warehouses help you understand it, and lakes hold raw data. AI usually draws from a warehouse or governed lake, not the live system.
- Seven layers, four cross-cutting concerns. Ingestion, raw store, curated store, catalog, feature store, training and evaluation, serving, with security, privacy, quality and observability running through all of them.
- The pipeline is the real product. Automated, scheduled movement that keeps data fresh, complete and consistent matters more than where the data rests.
- Catalog and lineage are audit infrastructure. If you cannot say what data trained a model, what data it sees now, and how the two differ, you cannot answer an inspector general or an auditor.
- Feature stores prevent a silent failure. Training-serving skew produces no error and no outage, just a model that performs worse in production than it did in testing.
- Evaluate continuously, and know the limit. A harness tells you how a model behaved on the tests you chose to run. It does not certify that the model is safe, fair or correct.
- Government data carries obligations. Classification, retention schedules, minimization and provenance must be enforced by the infrastructure, not bolted on later.
- Let authorization drive the hosting choice, but do not over-read it. Standing on an authorized environment spares a one-to-two-year accreditation; it does not authorize your data use, and it does not license you to overpay for isolation on public statistics.
- Gate every project on data readiness. If sources, pipeline, freshness, quality, classification and traceability are not in place, it is a data project. Fund that first.
Frequently Asked Questions
How do I tell whether we have an AI problem or a data problem? Run the seven-question readiness gate and count the "no" answers. In practice the tell is simpler still: ask your technical staff how long it would take to assemble a clean, current, complete dataset for the decision the model is meant to support. If the answer is measured in months, or if it depends on a specific person's manual work, you have a data project. That is not a reason to abandon the AI ambition. It is a reason to sequence it honestly and fund the plumbing first.
Do we need all seven layers before we can do anything? No. The layers describe a mature target, not an entry requirement. What you do need is to know which layer a person is currently substituting for, because those are your fragile points. A small agency can run a credible AI use case with a simple curated store, a documented ingestion path, a lightweight catalog and a real evaluation routine. What it cannot do safely is skip lineage and evaluation, because those are the two layers whose absence you only discover under audit or after harm.
Does a FedRAMP authorization mean we can put our data in that service? Not by itself. The authorization tells you the service's security posture was assessed and authorized within a defined boundary at a stated impact level. Whether your particular data belongs there is a separate determination involving your data owner, your privacy office and your authorizing official, and it depends on classification, records-notice coverage, any applicable agency-specific statute, and the data-use terms of your contract. Treat the authorization as a necessary gate you have cleared, not as permission you have received.
Our data lake already exists and nobody uses it. What now? That is the swamp outcome, and the fix is a catalog rather than more storage. Start by inventorying what is actually in there, naming an owner for each significant dataset, recording sensitivity and authorized uses, and marking anything whose provenance cannot be established. Expect to find duplicates and abandoned copies. The datasets you cannot trace are the ones to quarantine first, because they are the ones you could not defend if a model were trained on them.
How often should we be checking for drift? Frequently enough that the gap between degradation and detection is shorter than the harm it would cause. For a model supporting a consequential decision, recomputing key metrics on recent production data daily or weekly is the pattern the source guidance describes, with alerting thresholds and an escalation path defined in advance. The threshold itself should be calibrated against your own baseline. What matters more than the number is that someone receives the alert, is authorized to act, and knows what the first hour looks like.
Who should own data infrastructure for AI, the AI team or the data office? The data office should own the shared layers and the AI team should consume them. When AI teams build their own ingestion, storage and feature definitions, you get duplicated datasets, inconsistent definitions and governance that has to be re-established per project. The federal pattern here is well established: the Chief Data Officer's office maintains the canonical inventory and the governance board reviews new uses, while program teams bring the mission requirements and the evaluation criteria. Where that split is unclear, write it down before the second AI project starts, not after.
Skill.re