←
AI for Government
Strategic · M22 · lesson 22 of 47 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Data Sharing Agreements for AI
📖
now learning

Data Sharing Agreements for AI

15 min

Niall Brennan manages interagency data programs for a state public health department: a staff of twelve, a surveillance system that ingests data from 47 reporting entities across five state agencies, and an AI pilot that had been stalled for six months because two of the necessary data providers could not agree on how to handle the sharing arrangement. The pilot was a disease outbreak early-warning system. It would integrate emergency department visit data from the state hospital association, prescription monitoring data from the pharmacy board, laboratory test results from the state lab, and demographic data from the vital statistics office. Each source was legally authorized to share data for public health purposes. Each source had privacy requirements, system compatibility constraints and institutional concerns about how the shared data would be used once it left their custody. Niall had managed data sharing arrangements before. He had never managed one that required a machine learning model to be trained on data drawn from separate legal custodians under separate statutory frameworks, two of which contained explicit restrictions on secondary use. He eventually produced a 34-page data sharing agreement that all parties signed. He says he learned more about government data law in those six months than in the preceding twelve years of his career.

Why AI Makes Data Sharing Harder

Traditional government data sharing agreements address a relatively simple transaction. Agency A transfers a defined dataset to Agency B for a specified purpose, under specified security controls, for a specified time period. The purpose is usually a particular analysis, a required report or a legal investigation. The data moves once or on a defined schedule. The downstream use is predictable and bounded, and when the period ends the arrangement can be wound up by deleting what was transferred.

AI training introduces complications those agreements did not anticipate. When data is used to train a machine learning model, it is transformed in ways that are difficult to specify in advance and difficult to reverse. The model learns from the data, embedding patterns from it into its parameters, in a way that may persist long after the original dataset is deleted. A model trained on prescription monitoring data retains the statistical patterns from that data even if every record is destroyed after training. Whether those embedded patterns constitute retained data for purposes of statutory use restrictions is a question that courts and legislators have not fully resolved.

That unresolved question is not an abstraction you can defer. It is the thing that determines what your agreement has to say, because an agreement drafted on the assumption that deletion ends the relationship will be silent exactly where the difficulty lives. Treat the uncertainty as a reason to be explicit rather than a reason to wait. Write down what the parties intend to happen to the model, note in the agreement that the parties recognise the legal position is unsettled, and revisit the provision when it becomes clearer. An agreement that records a considered position ages far better than one that never noticed the question.

The Questions Existing Templates Did Not Ask

The practical implication is that agencies entering into data sharing agreements for AI training must address questions their existing templates did not include. What happens to the trained model when the data sharing agreement expires? Can the model continue to be used? Must it be retrained on a dataset that does not include the expired source? Who has access to the model, and does that access constitute indirect access to the training data?

These are not rhetorical questions. They are the questions Niall's legal team spent six weeks answering for two of the data sources. Six weeks is worth sitting with, because it is the cost of answering them once, in advance, for a single arrangement. The alternative is answering them under pressure after a source has given notice, when the model is already in production and the honest answer may be that it has to come out of service.

Add one more question to the list, which experienced negotiators raise and templates rarely do: what happens to model outputs and derived artifacts? An evaluation dataset, a set of embeddings, a fine-tuned checkpoint and a published performance report are all downstream of the shared data and all outlive the training run. Decide explicitly which of them the agreement covers. Leaving them unmentioned does not put them outside the arrangement; it puts them in a place where the parties will disagree later about whether they were inside it.

A data sharing agreement for AI that will withstand legal scrutiny addresses seven distinct legal and governance components. They are listed in the order a negotiation should take them, because each later one depends on the earlier ones being settled.

Statutory authorization for sharing. Each data source must have a documented legal basis for sharing data for the specific AI purpose. That means identifying the specific statute or regulation that authorizes sharing, not sharing generally but sharing for the purpose of training an AI model; whether any additional consent or notification requirements apply to the individuals whose data will be included; and whether the AI purpose falls within the defined scope of the sharing authority or requires a new determination. For Niall's agreement, the vital statistics office required an opinion from the state attorney general on whether disease surveillance AI development was within the scope of its existing data sharing authority. It was, but that determination took nine weeks.

Purpose specification and use limitation. The agreement must specify exactly what the data will be used for and must include explicit restrictions on secondary use. A dataset licensed for training a disease outbreak early-warning model may not, without separate authorization, be used to train a model for a different purpose, even if the second purpose seems benign and public health related. Purpose limitation is the core privacy protection in most government data sharing frameworks. It cannot be waived by agreement between the data custodians, because it is established by the statute under which the data was originally collected.

Technical security requirements. Each data source may have its own technical security requirements: encryption standards, access control requirements, audit logging specifications, network isolation requirements. The shared AI environment must meet the most stringent applicable standard for any data it contains. Niall's agreement required the shared training environment to meet the pharmacy board's technical standards for sensitive health data, which were more stringent than the hospital association's standards. Both sets of standards had to be satisfied simultaneously.

Privacy impact assessment documentation. Before any personally identifiable information is transferred for AI training, the receiving agency should complete and document a privacy impact assessment (PIA), a formal analysis of what privacy risks the data sharing creates and what protections have been put in place. Many state and federal statutes require PIAs for new uses of personal information. The PIA documentation demonstrates due diligence to Inspector General reviewers and gives the agency a record of the analysis it performed if the sharing arrangement is later challenged.

Data minimization requirements. The agreement should specify that only the minimum data necessary for the specified AI purpose will be shared. Training datasets should be limited to the fields required for the model, the time period required for the model, and the geographic scope required for the model. Data minimization reduces privacy risk and reduces the potential liability surface if a security incident occurs during the period of sharing.

Retention and deletion requirements. The agreement must specify how long the training data will be retained, when and how it will be deleted, and what verification process confirms deletion. For model training purposes the relevant retention period is typically the training duration plus a defined period for validation and testing. Post-training retention of training data should be limited to what is required for audit and reproducibility purposes. The agreement should specify that the trained model continues to be governed by the AI governance framework of the receiving agency even after the original training data is deleted.

Incident notification and breach response. The agreement must specify how each party will be notified of a security incident involving the shared data, within what timeframe, by what channel, and what investigation and remediation obligations apply. A breach involving data from several custodians requires coordinated notification processes that must be pre-designed rather than improvised after an incident. Each custodian will also carry its own notification duties to individuals and to regulators, and the agreement should record how the parties will avoid duplicative or contradictory notices without any party delaying an obligation it holds independently.

What the Agreement Does, and What It Cannot Do

Niall's 34-page agreement is the visible artifact and it is easy to mistake it for the achievement. It is not. The achievement was the nine weeks of legal determination, the six weeks of model questions and the negotiation of standards between custodians. The agreement recorded those outcomes; it did not produce them, and it could not have.

State this plainly to anyone who asks whether the agreement makes the sharing lawful, because that is the question people actually mean when they ask whether it is signed. A data sharing agreement documents and constrains an authority that already exists in statute. It does not create authority, it cannot enlarge it, and a signature from every party does not cure a defect in the underlying legal basis. If the statutory determination is wrong, the agreement is a well-drafted record of an unlawful arrangement. This is why the components are ordered as they are: authorization first, and everything else conditional on it.

The same caution applies internally. A signed agreement tends to end the conversation inside the receiving agency, and the technical team reads the signature as clearance to proceed. What actually protects the arrangement is that the constraints in the document are implemented in the environment, monitored while the sharing is live, and re-examined when anything changes. An agreement with no compliance monitoring is a description of an arrangement rather than a control over one. Cross-Agency AI Coordination sets out the sequence in which the legal determination, the agreements, the technical safeguards and the monitoring should be established.

Privacy Protections and Their Limits

The privacy protections in the component list are real and they are not equivalent to each other, so it is worth being exact about what each one buys. Purpose limitation is the strongest, because it is imposed by the statute under which the data was collected rather than agreed between the parties, and no arrangement between custodians can relax it. Treat it as fixed and design the AI use inside it.

Minimization is the next strongest and it is genuinely load-bearing. Fewer fields, a shorter time period and a narrower geography reduce privacy risk and shrink what is exposed if something goes wrong. What minimization does not do is move the remaining data outside the scope of the rules that govern it. A minimized training set built from records covered by a statutory restriction is still covered by that restriction. Smaller is safer, and smaller is not exempt.

That distinction matters most when a team proposes to reduce a dataset until it is no longer personal. Removing direct identifiers reduces re-identification risk; it does not eliminate it, and the residual risk depends on what remains, how unusual the combinations are, and what else is available to be linked against. Treat data with identifiers removed as lower risk, never as anonymous, and never as a reason to stop applying the agreement's controls to it. The same holds for aggregation: aggregate figures can still expose individuals when the underlying groups are small. If a proposal's entire argument for lawfulness is that a transformation has been applied to the data, that is the point to bring in counsel rather than the point to proceed. PII and AI: The Bright Red Lines and Privacy Engineering for AI develop the technical side of this.

The PIA sits alongside these rather than above them. It is a formal analysis of the risks the sharing creates and the protections in place, it is required for new uses of personal information under many state and federal statutes, and it must be completed rather than initiated before data transfer occurs. Its value to a reviewer is that it shows the agency identified the risks and made deliberate choices about them. Its value is not immunity. A PIA that describes controls the environment does not actually implement is worse than no PIA at all, because it converts an oversight into a documented misstatement. Privacy Impact Assessments for AI Systems covers the artifact itself in depth.

Technical Requirements Across Custodians

The most-stringent-standard rule is easy to state and hard to implement, and the difficulty is rarely technical. Niall's shared environment had to satisfy the pharmacy board's standards for sensitive health data and the hospital association's standards simultaneously, and the pharmacy board's were stricter. That is the simple case, where one standard clearly dominates.

The awkward case is where two custodians' requirements are stringent in different directions rather than at different levels: one requires that access be logged to a particular level of detail while another restricts what may be recorded about queries, or one mandates a network posture that another's tooling cannot operate inside. These are not resolved by picking the higher standard, because there is no higher. They are resolved by the custodians talking to each other directly, early, with the technical staff in the room. The failure pattern is a program office that negotiates with each custodian separately, agrees to everything, and then discovers at build time that the commitments cannot coexist.

Two practical rules follow. First, get the technical requirements into the negotiation at the start rather than treating them as an implementation detail to be worked out after signature, because a requirement that turns out to be unbuildable is an agreement that has to be reopened. Second, write down which party is responsible for demonstrating that the environment meets each standard, and how. A standard that every party assumes another party is verifying is a standard that no one is verifying.

Negotiating Across Multiple Custodians

A multi-custodian arrangement is not several bilateral arrangements stapled together, and treating it as such is what stalled Niall's pilot for six months. Four sources were named for his early-warning system: the state hospital association, the pharmacy board, the state lab and the vital statistics office. Each was legally authorized to share for public health purposes, and each had its own privacy requirements, system compatibility constraints and institutional concerns about downstream use. The concerns were not obstacles to be overcome. They were the terms on which participation was possible.

What makes these negotiations tractable is sequencing and visibility. Establish each custodian's statutory position before negotiating any terms, because a term agreed with one party and then blocked by another party's authority has to be renegotiated with everyone. Bring the parties into the same conversation once the individual positions are known, so that the most stringent requirement is visible to everyone rather than arriving as a late surprise. And be specific in what you ask for. A request for a partner's dataset invites a slow refusal; a request for named fields, for a stated purpose, with a stated retention limit, invites a negotiation.

Expect the institutional concerns to be about control rather than about privacy law, and treat them seriously on their own terms. A custodian worried about how the shared data will be characterised in a published model, or about being associated with a decision its own leadership did not approve, is raising something an authorization opinion will not settle. Provisions that address it, such as review rights over publications, notice before the purpose changes, and a clean withdrawal mechanism, cost little and are frequently what converts a reluctant party into a signatory. Cross-Agency Governance Coordination covers the framework-level alignment that makes these conversations shorter next time.

Who Is Actually in the Negotiation

Niall runs a staff of twelve, which is a useful reminder that this work is rarely done by a dedicated legal function. It is done by a program manager who also has a system to run, supported by counsel who are answering the question for the first time. Knowing who has to be involved, and when, is most of what keeps the timeline from doubling.

Counsel is unavoidable and cannot be the only participant, because the questions that matter most are joint. Whether a purpose falls inside an existing authority is a legal question that depends on a technical description of what the model will actually do, and a lawyer given a vague description will return a cautious answer or an answer to the wrong question. The privacy function has to be present because the assessment has to be completed before transfer rather than assembled afterwards from decisions already made. Security has to be present because the environment has to satisfy the strictest applicable standard, and that constraint shapes what is buildable. And the program has to be present throughout, because every one of those constraints trades against model capability and only the program can say what the model can afford to lose.

The counterpart side matters as much. Each custodian brings its own counsel, its own privacy and security staff, and often an institutional stakeholder whose concern is reputational rather than legal. Identify those people early and by name. An arrangement that is agreed by program managers and then discovered by a custodian's counsel late is an arrangement that restarts, and the restart usually costs more than the original negotiation because positions have hardened in the meantime.

Monitoring While the Sharing Runs

The seven components describe what the agreement must say. What makes them protective is that somebody checks, while the sharing is live, that the environment still matches the document. This is the part that is almost always underspecified, because the negotiation exhausts everyone and signature feels like completion.

Decide who verifies what, and how often, in the agreement itself rather than afterwards. The access controls and audit logging can be evidenced from the environment. The minimization commitment can be checked against what was actually loaded rather than what was requested, and those two frequently diverge because an engineer pulled a convenient extract. The purpose restriction is the hardest to monitor because nothing about a second training run looks different from the first, which is an argument for requiring that new training runs against this data be recorded and reviewed rather than assumed to be in scope.

Give the providing custodians a defined way to see that compliance is real, whether through periodic reporting, a right to request evidence, or an audit provision. Custodians who can verify are custodians who will renew. Custodians who signed on trust and then heard nothing for a year tend to renegotiate hard at renewal, and the reason they give is almost never the reason they mean.

Winding Down: Expiry, Withdrawal and Incident

Three endings need provisions and most agreements draft only the first. Expiry is the easy one: the term runs out, the retention and deletion provisions execute, deletion is verified by the process the agreement names, and the model provisions determine what happens to what was trained. Draft it and the ending is administrative. Note that expiry has a habit of arriving unnoticed, because the people who negotiated the agreement have usually moved on and the system it governs is running quietly. Put the review date somewhere operational rather than only in the document, and give the receiving agency an internal owner whose job includes watching it, or the first sign that the term has lapsed will be a custodian's inquiry.

Withdrawal is harder, because it happens on someone else's schedule and usually for a reason unrelated to your program: a change in the custodian's own legal advice, a leadership change, an unrelated incident that makes the institution cautious. The agreement should state what notice a withdrawing party gives, what the receiving agency must do with data already transferred, and what happens to a model already trained on it. Getting a clean withdrawal mechanism into the document is also, counterintuitively, one of the better ways to get a reluctant custodian to sign at all, because a party that can leave is a party that is not making an irreversible commitment.

The third ending is an incident serious enough to suspend the arrangement. Specify who can suspend the sharing, what happens to processing already underway, and how the parties decide whether to resume. Every custodian retains its own notification duties to individuals and regulators regardless of what the coordination provisions say, and the agreement should make that explicit so that coordination is never read as permission for any party to wait.

Template Structures and Their Limitations

Several federal agencies, including HHS (the Department of Health and Human Services) and GSA (the General Services Administration), have published template data sharing agreement structures that address AI-specific provisions. State associations of public health officials and state CIO organizations have produced similar templates for state and local contexts. Confirm the current version and status of any template directly with the publishing body before building on it, because published structures in this area are being revised frequently.

Templates are useful starting points. They are not ready-to-sign documents. Every data sharing arrangement for AI training requires customization to the specific statutory authorities of the parties, the specific technical requirements of the sharing environment, and the specific governance requirements of the AI system being trained. Niall's 34-page agreement started from a 12-page template and required 22 additional pages to address the legal questions his multi-source arrangement generated. Those figures reconcile exactly, and the ratio is the useful part: the template supplied roughly a third of the final document.

The most important sentence in this section is the shortest. Template language that does not address a specific requirement does not satisfy that requirement. Teams under schedule pressure read a template's completeness as coverage, sign it, and discover later that the provision they needed was never there. Read a template as a checklist of questions somebody else found worth asking, not as a set of answers to yours.

Anti-Patterns to Avoid

  • Treating the signed agreement as the authority. An agreement documents and constrains an authority that exists in statute. It does not create one, and no number of signatures cures a defect in the underlying legal basis. If the statutory determination is wrong, the agreement is a well-drafted record of an unlawful arrangement.
  • Treating a transformation as a get-out. Removing direct identifiers reduces re-identification risk; it does not eliminate it, and aggregation can still expose individuals when groups are small. Data with identifiers removed is lower risk, not anonymous, and remains inside the agreement's controls.
  • Assuming minimization changes the applicable rules. Fewer fields and a shorter window genuinely reduce exposure. They do not move the remaining records outside the statute that governs them.
  • Building the model first and papering the data flow afterwards. By the time there is a working demonstration, a sunk cost is arguing against the lawyer. Establish each custodian's statutory basis before terms are negotiated, and terms before data moves.
  • Negotiating bilaterally in a multi-custodian arrangement. Agreeing terms separately with each party produces commitments that cannot coexist, discovered at build time. Get the parties into one conversation once their individual positions are known.
  • Leaving the model out of the agreement. An agreement that only governs the dataset is silent on what happens to the trained model, its checkpoints, its embeddings and its evaluation artifacts when the arrangement ends. Silence is not exclusion; it is a future disagreement.
  • Signing a PIA that describes controls the environment does not implement. That is worse than having none, because it turns an oversight into a documented misstatement in front of a reviewer.
  • Treating technical requirements as post-signature implementation detail. A requirement discovered to be unbuildable after signature reopens the agreement. Get technical staff into the negotiation and name who demonstrates compliance with each standard.
  • Reading template completeness as coverage. Language that does not address your requirement does not satisfy it, however thorough the surrounding document looks.
  • Stopping at signature. Constraints in a document protect nothing until they are implemented in the environment, monitored while sharing is live, and re-examined when the model, the purpose or the parties change.

Practice Prompts

  • Authority trace. For one dataset your agency wants to use in AI training, write down the specific statute or regulation that authorizes sharing, and then answer separately whether that authority covers training a model rather than performing an analysis. If you cannot answer the second question from the text, that is a question for counsel, not an assumption to make.
  • Model afterlife clause. Draft the provision governing what happens to a trained model when the sharing arrangement ends. Cover continued use, retraining obligations, who has access to the model, and whether checkpoints, embeddings and evaluation sets are in scope.
  • Minimization pass. Take a proposed training dataset and cut it three ways: fields, time period and geography. For each cut, record what model capability you lose. Bring both columns to the negotiation, because a custodian is far more likely to agree to a request that shows its own limits.
  • Standards collision check. List the encryption, access control, audit logging and network isolation requirements of each prospective custodian side by side. Mark where one is simply stricter and where two are stringent in incompatible directions. The second category is your early escalation list.
  • Seven-component audit. Take an existing data sharing agreement your agency has signed and mark which of the seven components it addresses adequately, partially or not at all. Most agreements written before AI training was contemplated will score poorly on the model provisions and on secondary use.
  • Coordinated breach walkthrough. Walk through a security incident affecting data from every custodian in one arrangement. Who tells whom, in what order, through what channel, and where do the custodians' independent notification duties overlap or conflict? Do this on paper before it happens.

Reflection Questions

  • For your most valuable potential training dataset, do you know the specific authority that would permit its use for training, or only that your agency is permitted to hold it?
  • If a data provider gave notice tomorrow, could you say what your agency would have to do with the models already trained on their data?
  • Who in your agency would notice if a training dataset was used for a purpose outside the one the agreement specifies, and what would tell them?
  • Which of your existing sharing agreements were written before AI training was contemplated, and what do they say about the model?
  • When a colleague proposes that a dataset has been de-identified enough to move outside the agreement's controls, what evidence would you ask for before agreeing?

Glossary

  • Data sharing agreement. The instrument that documents and constrains an existing statutory authority to share data between custodians. It records terms; it does not create authority.
  • Statutory authorization. The specific statute or regulation permitting a custodian to share data for a stated purpose. For AI, the question is whether that authority extends to training a model rather than to analysis alone.
  • Purpose limitation. The restriction of data use to the purpose for which it was collected. Established by statute rather than by agreement, and therefore not waivable by the custodians.
  • Secondary use. Any use of shared data beyond the purpose specified in the agreement, including training a different model, which requires separate authorization even where the new purpose appears benign.
  • Privacy impact assessment (PIA). A formal analysis of the privacy risks a data sharing arrangement creates and the protections put in place, required by many statutes for new uses of personal information and completed before transfer.
  • Data minimization. Limiting the shared dataset to the fields, time period and geographic scope required for the specified purpose, which reduces privacy risk without changing which rules apply to what remains.
  • Re-identification risk. The residual possibility that individuals can be identified from records whose direct identifiers were removed, depending on what remains and what it can be linked against.
  • Retention and deletion provisions. The terms specifying how long training data is held, when and how it is deleted, and what process verifies the deletion.
  • Custodian. The entity holding data under a legal duty, whose own statutory framework governs what it may share and on what terms.
  • Template agreement. A published starting structure for a data sharing agreement. A checklist of questions others found worth asking, requiring customization to the parties, the environment and the system.

Closing Thoughts

Niall says he learned more about government data law in six months than in the preceding twelve years, and the reason is that AI training forced questions the previous twelve years had let him leave implicit. Every one of those questions existed before the pilot. Traditional sharing simply never pressed hard enough on them, because a dataset that is transferred, analysed and deleted does not raise the problem of what persists.

The discipline this lesson asks for is not caution for its own sake. It is that the legal work happens first and visibly, that the agreement records determinations rather than substituting for them, and that every protection is described accurately as what it actually buys. Purpose limitation is fixed. Minimization reduces exposure without changing the rules. Removing identifiers lowers risk without producing anonymity. A PIA evidences analysis without conferring immunity. An agreement constrains an authority without creating one. Say each of those precisely and the arrangement will survive scrutiny. Blur any of them and you have built a program on a sentence that will not hold.

Key Takeaways

  • AI training creates legal questions traditional agreements did not anticipate. Whether patterns embedded in a trained model's parameters constitute retained data for statutory purposes is unresolved. Agreements must address what happens to the model, and to checkpoints, embeddings and evaluation artifacts, when the arrangement expires.
  • The agreement records authority; it does not create it. Signatures from every party do not cure a defect in the statutory basis. Establish authorization first and treat every other component as conditional on it.
  • Statutory authorization must be documented for the AI purpose specifically. General sharing authority may not extend to model training. Niall's vital statistics office needed a state attorney general opinion, which took nine weeks and came back affirmative.
  • Purpose limitation is established by statute, not by agreement. Custodians cannot waive it between themselves. The AI purpose must fall within the scope under which the data was originally collected, or a new legal authority must be established.
  • Minimization reduces exposure without changing which rules apply. Limit fields, time period and geography to what the model requires. The remaining records stay inside the statute that governs them.
  • Removing identifiers lowers risk; it does not produce anonymity. Residual re-identification risk depends on what remains and what it can be linked against, and aggregation can still expose individuals in small groups. Keep the agreement's controls applied.
  • The shared environment must meet the most stringent applicable standard. Where custodians' requirements are stringent in incompatible directions rather than at different levels, that is an early escalation, not an implementation detail.
  • PIAs must be completed, not initiated, before transfer. They evidence the analysis performed and demonstrate diligence to Inspector General reviewers. A PIA describing controls that were never implemented is worse than none.
  • Retention provisions must survive the model. Specify the retention period, the deletion method and the verification process, and state that the trained model remains governed by the receiving agency's AI governance framework after the training data is deleted.
  • Incident notification must be pre-designed across custodians. Coordinate the notices without any party delaying a duty it holds independently, and walk it through on paper before an incident tests it.
  • Templates are starting points. Niall's 34 pages began as a 12-page template plus 22 pages of customization. Language that does not address your requirement does not satisfy it, however complete the document looks.
  • Signature is not the finish line. Implement the constraints in the environment, monitor while sharing is live, name who verifies each standard, and re-examine when the model, purpose or parties change.

Frequently Asked Questions

Can we start training while the agreement is still in negotiation?

No, and the temptation to do so is exactly what the ordering of the seven components is designed to resist. Statutory authorization comes first because everything else is conditional on it, and training on data whose legal basis has not been settled creates a model you may have to destroy along with an exposure you cannot undo by deleting the dataset. Niall's pilot sat stalled for six months, which was expensive and recoverable. Building the model first and papering the flow afterwards is the failure that is not recoverable, because by then a working demonstration and a sunk cost are arguing against the lawyer.

What happens to the model when the agreement expires?

Whatever your agreement says, which is why the provision has to exist. The legal position on whether a trained model retains the data for purposes of statutory use restrictions has not been fully resolved by courts or legislators, so the parties should record their intended answer explicitly rather than leave it to be argued later. Cover continued use, any retraining obligation, who has access to the model and whether that access counts as indirect access to the training data, and extend the same treatment to checkpoints, embeddings and evaluation sets. Have counsel draft the provision and note in it that the parties recognise the legal position is unsettled.

If we de-identify the training data, do we still need an agreement?

Ask counsel, and go in with the right framing. Removing direct identifiers reduces re-identification risk; it does not eliminate it, and how much risk remains depends on what fields are left, how unusual the combinations are, and what external data could be linked against them. Treat de-identified data as lower risk, never as anonymous, and never as automatically outside the arrangement. If the entire argument for proceeding without an agreement rests on a transformation having been applied, that is the moment to stop and get a legal determination rather than the moment to save a step.

Whose security standard applies in a shared training environment?

The most stringent applicable standard for any data the environment contains, applied simultaneously rather than averaged. Niall's environment had to meet the pharmacy board's standards for sensitive health data, which exceeded the hospital association's. The harder case is requirements that are stringent in different directions, such as one custodian requiring detailed query logging while another restricts what may be recorded, and those cannot be resolved by picking a level. Get the custodians and the technical staff into the same conversation early, and write down which party demonstrates compliance with each standard.

Do we need a separate agreement for each data source?

Not necessarily separate documents, but you do need each source's statutory position established separately before any terms are negotiated. A multi-custodian arrangement is not several bilateral ones combined, and agreeing terms with each party in isolation reliably produces commitments that cannot coexist. Establish the individual legal positions, then bring the parties into one conversation so that the most stringent requirement is visible to everyone rather than surfacing late. Niall's four named sources produced a single agreement, and it was long because it had to reconcile them rather than list them.

Our agency already has a data sharing agreement with the source. Is that enough?

Probably not, and this is the most common way agencies get this wrong. An existing agreement was written for a purpose, and AI training is a different purpose unless the agreement says otherwise. Check whether its purpose specification covers model training, whether its secondary use restrictions permit it, and whether it says anything at all about what happens to a model. Agreements written before AI training was contemplated typically score poorly on exactly those provisions, which is a reason to amend rather than to interpret generously.

How long should we budget for this?

Longer than the technical work, and the figures in this lesson are illustrative rather than benchmarks. One statutory determination took nine weeks. The model-specific legal questions took six weeks for two of the sources. The pilot as a whole was stalled six months before the arrangement was settled. Your own timeline depends on the number of custodians, how novel the purpose is for each, and whether your agency has done anything similar before. Build the schedule around the legal path rather than fitting the legal path into a schedule set by the model.