AI Sandbox and Experimentation Frameworks
The day Marisol Cordova's team deployed the AI triage tool in her county's social services intake office, it had never processed a real case. It had been tested on synthetic data and a small batch of anonymized historical records. What it had never done was encounter a Spanish-speaking family with a housing instability case that also carried child welfare flags, a scenario that turned out to represent about 14 percent of their actual caseload and that the model handled poorly, misrouting cases to the wrong program queue at a rate nearly three times higher than its overall error rate. Marisol is the Deputy Director for Technology Innovation at a large urban county human services agency. Her agency processes roughly 4,500 new cases a month. The misrouting problem was caught by a frontline supervisor two weeks after deployment. The question Marisol asked afterward was not why the vendor had failed to test for this. The question was why her own agency had not. This lesson is about building the sandbox structure that catches those failures before they reach real families.
What a sandbox is, and what it is not
An AI sandbox is a controlled environment where tools can be tested, evaluated and iterated with reduced exposure to real operations, real data and real constituents. In government the term is often used loosely to describe any informal pilot, and that looseness is where the damage starts. A real sandbox has defined boundaries, defined risk limits, defined evaluation criteria and a defined graduation process into production. Without those four elements you do not have a sandbox. You have an uncontrolled experiment running on government data, with none of the review that an experiment on government data would normally attract.
The useful analogy is a food safety test kitchen. A test kitchen uses the same ingredients, the same equipment and the same recipes as the production kitchen, but nothing that comes out of it reaches a customer until it has been reviewed and approved. The rules about what may be tested, how long it may stay in testing, and what approval actually means are written down and applied consistently. A government AI sandbox works the same way: real tooling, realistic data structures under documented handling controls, and explicit criteria that any tool must meet before it is allowed to touch production systems or influence a decision about a person.
What a sandbox does not do
The most dangerous belief in this whole subject is that the sandbox label is itself a control. It is not. A sandbox reduces the probability and the blast radius of a failure. It does not guarantee that no real person is affected, and treating it as a guarantee is how agencies end up surprised. Sandbox environments are still connected to something: staff who form opinions from test outputs, data that was copied from somewhere real, cloud accounts that share credentials with production, and vendors who take away what they learned. A failure inside a sandbox can leak outward through any of those paths, and none of them appear on the architecture diagram that made the environment look isolated.
The same caution applies to the legal frame. Calling a system experimental does not, by itself, suspend the obligations the agency already carries around privacy, records, accessibility or civil rights. Whether a particular obligation attaches to a particular test is a question for your counsel, your privacy officer and your records officer, answered in writing before the test starts, not a question the project team should answer for itself by pointing at the word pilot. Where an exemption genuinely exists in your jurisdiction, get it stated in writing by the office that owns it. Where nobody will put it in writing, plan as though the obligation applies.
The four components of a government AI sandbox
Sandbox design: environment and data controls. The sandbox environment must be separated from production systems, and the separation must be verified rather than assumed. Data used in the sandbox should be synthetic, de-identified under a documented protocol, or drawn from a historical dataset explicitly cleared for testing use. The sandbox should log every interaction between test tools and data, and those logs should be retained for audit. At Marisol's agency the sandbox runs in a separate cloud environment with no direct connection to the county's production case management system. Access requires multi-factor authentication and is limited to named individuals on the testing team.
Risk boundaries: what can and cannot be tested. Not every AI capability is appropriate for sandbox testing at every agency. A county social services agency should be cautious about testing any tool that produces outputs capable of influencing child safety determinations, even against historical data, because the norms around those decisions are themselves part of the risk. Risk boundaries define which categories of tool may enter the sandbox, what kinds of output are permissible in a test environment, and which capabilities require escalated review before entry. Writing this down before the first vendor asks for access is far easier than designing it under pressure with a vendor waiting.
Evaluation criteria: how you decide a tool is ready to graduate. This is the component most commonly missing in government AI testing. Agencies run pilots without ever defining what success would look like, which means the pilot ends when someone runs out of patience rather than when a question is answered. A tool should graduate from sandbox to pilot to production only when it meets documented criteria on four dimensions: technical performance, including accuracy, latency and error rates by demographic subgroup; operational fit, including integration with existing workflows and staff usability; equity performance, including disparate impact analysis across protected classes; and compliance, including data handling, audit logging and privacy protections. Marisol's agency now requires equity performance broken out by language, race and case type before any tool reaches the live pilot stage.
Graduation to production: the formal gate. Graduation is not automatic when a tool clears the evaluation criteria. It requires formal sign-off from at least three functions: the program office that owns the affected workflows, IT security, and the agency's equity or civil rights officer. For tools affecting benefit eligibility or case routing, the highest-impact category, the agency director is added as a required approver. The graduation checklist is a public document inside the agency, and any staff member can ask to see the approval record for any deployed AI tool. That last provision costs nothing and does more for internal trust than any communications plan.
Those four components are easier to hold onto as a grid than as prose, and the grid is also the thing to hand a new gate owner on their first day. Each row names the question the component answers, the evidence that satisfies it at the gate, and the way it fails when nobody is watching.
| Component | Question it answers | Evidence at the gate | How it fails quietly |
|---|---|---|---|
| Sandbox design | Where does this run and on what data | Separation verified, data source cleared, access list and interaction logs | Environment drifts from production until results stop predicting behavior |
| Risk boundaries | Is this tool allowed in here at all | Written category limits and any escalated review decision | Boundaries written once and never revisited as capabilities change |
| Evaluation criteria | What would make this ready | Technical, operational, equity and compliance results, by subgroup | Aggregate results accepted in place of subgroup results |
| Graduation gate | Who authorized this to leave | Signed record from each required function, retained and available | Signature collected after the deployment decision was already made |
Data in the sandbox
Two beliefs about sandbox data cause most of the trouble. The first is that synthetic data removes privacy risk. It reduces it, and it is usually the right default, but synthetic data generated from a real dataset can carry statistical structure from that dataset, and a generator trained on records about a small population can reproduce more than intended. The second belief is that de-identification makes people unidentifiable. De-identification reduces the risk of re-identification; it does not eliminate it, particularly when a test dataset is small, geographically concentrated, or joinable against something else the tester already holds. Treat both techniques as risk reduction with a residual, and document who reviewed the residual and found it acceptable.
The practical consequence is that sandbox data deserves the same handling classification as the production data it came from, unless a privacy review says otherwise in writing. That is a stricter posture than most agencies start with, and it is the one that survives contact with an inspector general. It also has a useful side effect: when sandbox data carries the same handling rules as production data, teams stop copying entire tables into test environments because it is convenient, and start requesting the specific fields the test actually needs.
Assessing readiness before you build
Before designing a sandbox, assess where the agency actually stands. Four questions structure the diagnosis. On current capability: what testing capability exists today, what gaps are visible, and where is the largest opportunity to reduce risk. On organizational readiness: is the agency ready for tools to be told no at a gate, what barriers exist, and what would help. On stakeholder alignment: who are the key stakeholders, what are their interests and concerns, and how aligned are they on direction. On resources: what is available, what constrains you, and how do you work effectively inside those constraints rather than pretending they will lift.
The alignment question is the one that decides whether a sandbox survives. A sandbox exists to slow some things down on purpose. If the program executives whose projects will be slowed have not agreed in advance that this is the point, the first refused graduation becomes a political fight rather than a working control. Get that agreement while nothing is at stake, in a document that names the gate owners, and revisit it whenever leadership changes. A sandbox with no executive willing to defend a no is a documentation exercise.
Strategy, planning and risk
A sandbox is an initiative like any other, and it needs the same discipline. Goal clarity comes first: what is the sandbox for, why does it matter for this agency, and what does success look like in terms someone outside the technology office would recognize. Action planning follows: what specific steps move you toward that goal, in what sequence, and what resources does each require. Then risk management: what could go wrong with the sandbox itself, how will you mitigate it, and what is the contingency. Then stakeholder engagement: who needs to be involved, how will engagement be sustained past the launch, and how will concerns be addressed rather than logged.
Risks specific to the sandbox itself are worth listing explicitly, because they are easy to miss when the risk register is focused on the tools being tested. The environment can drift out of alignment with production until test results stop predicting production behaviour. Test data can go stale, so the sandbox measures last year's caseload. Access can accumulate as people join and nobody leaves. The gate can soften over successive approvals until it approves everything. Each of these has a cheap detection method, and none of them will be noticed by anyone whose job is the tool rather than the sandbox.
Running the sandbox
Good design fails without execution. Capability building asks what skills the sandbox requires, how they will be built, and how they will be sustained when the people who have them are recruited elsewhere. Process design asks what repeatable procedures support the sandbox, how consistency is ensured across different vendors and different tools, and how the procedures improve over time. Technology decisions ask what supports the sandbox, how those tools are selected, and how transitions are managed when a platform changes underneath you. Continuous improvement asks how progress is monitored, how improvement opportunities are identified, and how improvements actually get made rather than recorded.
The consistency question deserves particular attention in government, because inconsistency is where challenge finds purchase. If two vendors enter the same sandbox and face materially different evidence requirements, the one who was asked for more has a grievance and the one who was asked for less has a system in production that nobody examined properly. Publish the entry requirements, apply them uniformly, and record any deviation with a reason and an approver. That record is also the artifact that answers an audit question about why one tool graduated and another did not.
Measuring whether the sandbox works
A sandbox that is never measured becomes whatever the busiest person needs it to be. Decide in advance what you will count, who collects it, who receives the report and what the report is allowed to change. Useful measures are mostly about the gate rather than about the tools: how many tools entered, how many graduated, how many were refused and on what criterion, how long evidence assembly actually took, how many deviations were approved and by whom, and how many production incidents traced back to something the criteria should have caught. That last measure is the only real test of whether the criteria are the right ones.
Reporting matters as much as collection. Send the numbers to the executive who agreed to defend a refusal, not only to the technology office that runs the environment, because the sandbox erodes from the outside rather than the inside. And close the loop explicitly: each review should end with a decision to keep, change or retire a criterion, recorded with the reasoning. A criterion that has never refused anything is either preventing bad submissions or measuring nothing, and the difference between those two is worth an hour of somebody's attention rather than a permanent assumption.
Partners in the sandbox
Sandboxes frequently involve parties outside the agency: vendors under test, university researchers evaluating results, other agencies sharing the environment, and sometimes community organizations reviewing equity findings. Partnership design asks who the potential partners are, what value each brings, and what the basis of the relationship is. Governance asks how decisions get made across organizational boundaries, how conflicts are resolved, and how progress and results are managed when no single party is in charge. Benefit sharing asks how the value produced is distributed, whether the distribution mechanism is fair, and whether every partner is genuinely getting something.
Benefit sharing sounds abstract until you consider what a vendor gets from your sandbox. They get evaluation data about their product against real government workflows, which is commercially valuable and which they will use with other buyers. That is not improper, but it should be a negotiated term rather than an accident: who owns the evaluation results, what may be published, what may be said in a sales conversation, and what the agency receives in exchange. Agencies that leave this unstated frequently discover their sandbox findings quoted in a competitor's proposal.
Regulatory sandbox concepts for government
In some domains agencies are not only deploying AI, they are also responsible for regulating AI used by third parties. State insurance commissioners, financial regulators and public utilities commissions are increasingly fielding requests from regulated entities to deploy AI in rate setting, claims processing and service delivery. A regulatory sandbox in that context is a formal program allowing regulated entities to test AI applications under modified rules, with enhanced oversight, for a defined period and within defined parameters. It is a different instrument from the internal testing environment described above, and confusing the two produces bad policy.
The United Kingdom's Financial Conduct Authority pioneered this model for financial technology applications beginning in 2016, and several United States state regulators have adapted the concept for insurance AI. A regulatory sandbox does three things. It gives regulators direct visibility into how AI performs in real conditions before broad deployment is approved. It gives regulated entities a structured pathway to test innovation without facing a full enforcement action over a technical compliance gap. And it gives both parties the empirical basis needed to update regulatory requirements to reflect actual capabilities and actual risks rather than assumed ones.
For agency leaders overseeing regulated sectors, the question is not whether a regulatory sandbox is appropriate. The question is whether you have the staff capacity and the legal authority to run one. Both are solvable, and both require explicit executive and legislative support obtained before the first cohort, because a regulatory sandbox operating without clear authority is an enforcement discretion problem waiting to be litigated by somebody who was not admitted to it.
Scaling and sustaining the sandbox
Successful pilots have to scale and then survive. Scaling asks how a successful sandbox graduate becomes a production system, what changes at scale, and how quality is maintained when volume rises and the attentive project team disperses. Funding asks how the sandbox itself is funded, whether that funding is sustainable, and what happens to the gate when the funding source changes. Organisational embedding asks whether the sandbox is part of standard operations or a personality-dependent initiative, what happens when the people who built it leave, and how institutional knowledge is maintained.
The turnover question is the one that quietly ends most sandboxes. The gate holds because a specific person understands why each criterion exists and is willing to enforce it. When that person moves on, their successor inherits a checklist with no rationale attached, and criteria without rationale erode because every one of them is inconvenient to somebody. The countermeasure is to write down why each criterion exists, with the incident or the analysis that produced it, so the next person is inheriting reasoning rather than paperwork. Marisol's language and case type requirement is a criterion with a story attached, and it will outlast her for exactly that reason.
What Marisol's agency did next
After the misrouting problem, Marisol's agency rewound. The tool went back into the sandbox. The agency added a specific evaluation requirement: performance on Spanish-language cases, broken out by case type, had to meet the same accuracy threshold as English-language cases before graduation. The vendor took six weeks to retrain and retest. The revised tool met the threshold. It then went through the three-function graduation sign-off and was redeployed. Nobody in that sequence did anything clever. They simply refused to let a tool leave the sandbox until it had been measured on the population it had previously failed.
In the eighteen months since redeployment the tool has processed more than 54,000 cases. The overall misrouting rate is 2.3 percent, roughly equivalent to the pre-AI manual routing error rate for the same case types. Language-based disparities in error rates have been eliminated. No cases have been re-reviewed because of AI routing errors.
That outcome required the sandbox, and specifically it required the evaluation criteria. Without the requirement that equity performance be broken out by language and case type, the retraining would have addressed the visible symptom without ever measuring whether the fix closed the disparity. The team would have seen an improved overall error rate, declared victory and shipped. The sandbox is not a bureaucratic hurdle. It is the mechanism that separates deploying a tool from deploying a tool that works for everyone it touches.
Anti-patterns
- The sandbox as containment. Treating the environment label as proof that no real person can be affected. A sandbox reduces exposure; it does not remove it, and the leak paths run through staff, copied data, shared credentials and vendors rather than through the network diagram.
- Synthetic data as a privacy guarantee. Assuming generated data carries no risk. Synthetic data derived from a real dataset inherits its structure, and de-identification lowers re-identification risk rather than ending it, especially on small or geographically concentrated test sets.
- Experimental as an exemption. Deciding inside the project team that a pilot designation suspends privacy, records, accessibility or civil rights obligations. That determination belongs to counsel and the relevant officers, in writing, before the test.
- The undefined pilot. Starting with enthusiasm and no statement of what success means, so that months later nobody can say what was accomplished or whether it worked. Define measurable criteria before entry, and make sure stakeholders understand and support them.
- The private design. A small group designing the sandbox and its gates, with the wider community of affected staff, program offices and community stakeholders surprised by the result. Resistance then arrives disguised as a technical objection.
- The underfunded sandbox. Approving a sandbox without the staff, environment cost or evaluation capacity to run it. The gate becomes nominal because nobody has time to produce the evidence it asks for, and everything passes.
- No named owner. Multiple sponsors and no accountable owner for the sandbox itself, so decisions stall, criteria drift, and no one can say who is entitled to refuse a graduation.
- Overall accuracy as the whole story. Reporting a single aggregate error rate that conceals subgroup disparities. Marisol's tool looked acceptable in aggregate while failing one population at close to three times its own error rate.
- Gate erosion. Softening a criterion for one urgent project and never restoring it, until the graduation review is a signature ceremony. Record every deviation with a reason and an approver, and review the deviations as a set.
- The orphaned checklist. Handing a successor the criteria without the reasoning behind them. Criteria that nobody can justify are removed by the first person who finds them inconvenient.
Practice prompts
- For AI sandbox and experimentation in your agency, run the assessment: current testing capability, visible gaps, organizational readiness, stakeholder alignment and resource constraints. Write the answers down even where the answer is that nobody knows.
- Develop a sandbox strategy covering clear goals, an action plan with a sequence, risk management for the sandbox itself, and a stakeholder engagement approach that survives the launch week.
- Design the partnership terms for one vendor entering your sandbox: who they are, what value each side brings, how decisions get made, who owns the evaluation results, and what may be said about them publicly.
- Write the implementation plan: capabilities to build, processes to design, technology to select, and the continuous improvement mechanism that will catch environment drift and stale test data.
- Design the measurement approach for the sandbox itself. What metrics, collected how, reported to whom, and used how to change the way the sandbox runs.
- Take one AI tool already in production in your agency and reconstruct what its graduation record would have looked like. Note every criterion you cannot evidence today.
- Draft the four handling questions for one proposed sandbox dataset and take them to your privacy officer: what it contains, where it came from, what the residual re-identification risk is, and who accepted that risk.
Reflection
Marisol's tool did not fail because anyone was careless. It failed because everybody involved had implicitly agreed that testing was something you did to find out whether the tool worked, rather than something you did to find out who it worked worse for. Those are different questions and they need different evidence. Spend some time on your own agency's version of this. Think about the last tool your organization put into production, and ask what evidence existed at the moment of that decision about how it performed for the populations least able to complain when it got them wrong. If the honest answer is that nobody produced that evidence, the gap is not in your technology. It is in the definition of what your testing was for.
Glossary
- AI sandbox. A controlled environment with defined boundaries, risk limits, evaluation criteria and a graduation process, used to test AI tools with reduced exposure to real operations and real people.
- Risk boundary. A written statement of which categories of tool and which kinds of output are permitted in the test environment, and which require escalated review before entry.
- Evaluation criteria. The documented technical, operational, equity and compliance conditions a tool must meet before it moves to the next stage.
- Graduation gate. The formal, multi-function sign-off that authorises a tool to leave the sandbox, with named approvers and a retained approval record.
- Regulatory sandbox. A formal program through which a regulator allows regulated entities to test applications under modified rules, with enhanced oversight, for a defined period and within defined parameters.
- Synthetic data. Artificially generated records used in place of real ones. It lowers privacy risk substantially but can inherit structure from the dataset it was generated from.
- De-identification. Removing or transforming identifying fields to reduce the risk that a record can be traced to a person. It reduces re-identification risk rather than eliminating it.
- Disparate impact analysis. Measurement of whether a system's outcomes differ across protected classes, reported as subgroup results rather than as a single aggregate.
- Stakeholder. A person or group affected by or interested in an initiative, internal or external.
- Governance. The structure and process by which decisions are made and initiatives are managed, including who is entitled to say no.
Related lessons
- AI Pilot Program Design covers the pilot stage a sandbox graduate enters next, and how to scope it.
- Moving from Pilot to Production covers what changes when volume rises and the attentive project team disperses.
- Responsible Innovation: Speed and Safety is the wider argument about pace and caution that a sandbox operationalises.
- Testing and Validating AI Systems covers the evidence that the technical performance criterion should demand.
- Algorithmic Impact Assessments covers the structured equity and rights analysis that belongs at the graduation gate.
- Privacy Impact Assessments for AI Systems covers the review that should decide how your sandbox data is handled.
- AI Regulatory Design covers the regulator side, where the regulatory sandbox is one instrument among several.
- Building Innovation Ecosystems places the sandbox inside the wider set of structures an agency uses to try new things.
Closing
A sandbox is a promise about sequence: that certain questions get answered before certain consequences become possible. Everything else in this lesson, the environment controls, the boundaries, the criteria, the multi-function gate, exists to keep that promise enforceable when a schedule is slipping and someone senior wants the tool live. Build it while nothing is urgent, write down why each criterion exists, and give one person the authority to refuse. Then remember what the sandbox does not do. It does not make a bad idea safe, it does not make test data harmless, and it does not turn an obligation off. It buys you the chance to find out what you built before it finds out about somebody's family.
Key takeaways
- A sandbox is not an informal pilot. It requires four written elements: a controlled environment, risk boundaries, evaluation criteria and a formal graduation gate. Missing any one turns it into an uncontrolled experiment on government data.
- The label is not the control. A sandbox reduces exposure rather than eliminating it, and calling a system experimental does not by itself suspend privacy, records, accessibility or civil rights obligations. Get any claimed exemption in writing from the office that owns it.
- Equity performance must be a graduation criterion. Measure error rates and disparate impact by language, race and case type before production. Aggregate accuracy conceals the subgroup failures that land on real people.
- Require multi-function sign-off. Program office, IT security and the equity or civil rights officer at minimum, with the agency director added for tools affecting benefit eligibility or case routing, and the approval record available to any staff member who asks.
- Define risk boundaries before the first vendor arrives. Written limits on what may be tested and what outputs are permissible protect the agency and steer vendors toward productive testing rather than negotiation.
- Treat sandbox data as the data it came from. Synthetic and de-identified data reduce risk with a residual that someone must review and accept. The default handling classification should follow the source data unless a privacy review says otherwise in writing.
- Negotiate what the partners take away. Evaluation results are commercially valuable to a vendor under test. Decide ownership, publication and permitted claims in advance rather than discovering your findings in someone else's proposal.
- Regulatory sandboxes are a different instrument. For agencies overseeing regulated sectors they give regulators real performance data and give regulated entities a structured pathway, but they require legal authority and staff capacity secured before the first cohort.
- Write down why each criterion exists. Criteria without their reasoning erode at the first personnel change, because every one of them is inconvenient to somebody.
- When a tool fails evaluation, rewind rather than work around. Marisol's agency returned the tool to the sandbox instead of accepting worse performance for Spanish-language cases. Six weeks of retraining produced a deployment that has held for eighteen months.
Frequently Asked Questions
Does running a tool in a sandbox mean no real person can be harmed?
No. It means the most obvious paths to harm have been closed, which is genuinely valuable and is not the same claim. Test outputs shape what staff believe about a tool before it launches. Test data came from somewhere real and can be exfiltrated or over-shared. Environments share credentials and identity providers with production more often than their diagrams suggest. Vendors leave with knowledge. Treat the sandbox as substantial risk reduction with a residual you have identified and accepted, and the residual will stay small.
We used synthetic data. Do we still need a privacy review?
Ask your privacy officer rather than deciding in the project team. Synthetic data is usually the right default and it lowers risk considerably, but a generator fitted to a small or geographically concentrated real dataset can reproduce more of that dataset's structure than people expect. The review is cheap, it is the artifact that answers the question later, and if the answer is that no review was required you have gained a written record at almost no cost.
What if a vendor refuses to enter the sandbox and wants to go straight to a pilot?
That is a source selection signal, and it is worth reading it as one. The requirements should be published and applied uniformly, so a refusal is a refusal of the same terms every other vendor accepted. If the objection is about a specific criterion being unreasonable, that is a conversation worth having and possibly a criterion worth revising for everyone. If the objection is that their product should not have to be measured, you have learned something useful before award rather than after.
How long should a tool stay in the sandbox?
Long enough to produce the evidence the criteria demand, which is a different answer per tool and is the only defensible basis. Fixed durations create two failures at once: tools that could have graduated early sit idle, and tools that cannot meet the criteria graduate anyway because the clock ran out. If you need a scheduling estimate for planning purposes, derive it from how long it takes to assemble the required evidence, and record the estimate as an estimate.
Who should own the sandbox?
One named person with the authority to refuse a graduation, backed by an executive who has agreed in advance to be the one defending that refusal. Shared ownership across sponsors is the failure mode: decisions stall, criteria drift, and when a graduation is contested nobody is clearly entitled to say no. The owner does not need to be senior. They need to be named, and the person who will back them needs to have said so before the first contested case.
Our agency regulates an industry. Can we use one sandbox for both purposes?
Keep them separate. The internal testing environment answers whether a tool the agency intends to use is fit for that use. The regulatory sandbox is an exercise of regulatory authority toward parties you supervise, with its own legal basis, its own eligibility rules and its own oversight obligations. Running them as one program confuses the agency's role as a buyer with its role as a regulator, and that confusion is the first thing a party excluded from the program will raise.
Skill.re