Data Access & Security Governance
Elena Vargas manages data operations for a 4,000-person healthcare services company. Last year, a well-meaning analyst on her team connected a patient communication database to a commercial AI summarization tool to speed up intake notes. The analyst had checked that the tool was HIPAA-compliant. What she had not checked was whether the contract terms permitted using patient data for model training, a separate question from compliance. By the time the legal team reviewed it, 11,000 patient records had been processed. Nothing bad had happened. The potential fine, had it been found during an external audit, was in the seven-figure range. "She was trying to help and she got 80% of the way to the right answer," Elena told me. "The 20% she missed was the part that actually mattered."
Data access and security governance is how organizations make sure that the right data reaches the right people for the right purposes, and that the wrong data does not go anywhere it should not. In an AI context the challenge is amplified, because AI tools create new consumption points for data that existing access controls were never designed to cover. The controls in most organizations were built around a person opening a record. They were not built around a person routing a database through a service that will retain it.
The Three Access Decisions
Every data access situation involves three decisions that are usually treated as one, and separating them is the foundation of a workable governance approach. The first is who can see this data. That is the traditional access control question, answered by role-based access systems, network permissions, and authentication controls, and most organizations have it reasonably well covered for human users. It is also the question that gets all the attention, which is part of why the other two go unasked.
The second is what they can do with it, and this is where most AI governance gaps live. An employee may be permitted to view customer records as part of their job without being permitted to export those records into a commercial AI tool, use them to fine-tune a model, or synthesize them into outputs that could be traced back to individuals. Permissioned access is not the same as permissioned use, and the distinction rarely appears in the systems that grant the access.
The third is what happens to the data afterward. When data flows into an AI system, where does it go? Is it retained by the vendor, used for model training, or shared across accounts? The analyst in Elena's case answered the first question, whether she was allowed to use the tool, but never reached the third, which was what the vendor would do with what she sent. A complete access governance policy answers all three questions for every significant data category the organization works with, and the second and third are the ones that require new work.
Classifying Data by Sensitivity
Not all data needs the same level of protection. The governance overhead appropriate for a list of publicly available product specifications is not the overhead appropriate for personnel records or patient health information. Over-governing everything is as dangerous as under-governing it, because uniform friction creates compliance fatigue and drives people toward workarounds. A four-tier classification is practical for most organizations.
- Public: Data already available externally, or which would cause no harm if disclosed. Minimal access controls, and generally usable with any AI tool without special review.
- Internal: Data that is not public but not sensitive, such as internal process documents, general business data, and non-confidential analysis. Basic access controls, with AI tool use requiring that the tool meets a standard contract threshold, for example that data is not retained for training.
- Confidential: Data whose disclosure could cause significant harm, including financial projections, personnel data, customer personally identifiable information, and strategic plans. Restricted access, with AI tool use requiring explicit policy approval and usually a vendor data processing agreement.
- Restricted: Data subject to specific legal protections, including health information, payment card data, and certain government-regulated categories. Strict access controls, with AI tool use requiring legal review, specific contract terms, and often technical controls such as data masking or anonymization before processing.
The classification should be assigned at the data type level rather than case by case, because a scheme that requires judgment on every occasion will not be used on the occasions that matter. When Elena's organization built this framework they published a one-page reference any employee could consult before connecting data to a new AI tool. The analyst's mistake, in retrospect, would have been caught by a 90-second consultation with that reference, which is a useful measure of what the framework is for: not to make people think harder, but to make the right answer fast enough to be worth looking up.
Heightened Controls for Special Categories
The restricted tier deserves more than a line in a classification table, because health information, financial data, and personal information carry obligations that do not scale down with the size of the analysis. What makes these categories different is not only the potential penalty but the fact that the rules govern purpose and processing rather than access alone. A lawful basis to hold health records for care delivery is not a lawful basis to use them to train a model, and that gap is precisely the one the analyst fell into. The permission she had was real; it simply did not extend to what she did.
The practical controls that follow are worth naming explicitly, because teams tend to assume that restricted data is handled by someone else. Masking and anonymization before processing reduce what leaves the boundary in the first place, and where identifiers must be retained, tokenization keeps the linkage inside systems you control. Contract terms need to be read for what the vendor may do rather than for what certification they hold, since compliance with a standard and permission to train on your data are separate questions that the same sales conversation often conflates.
Approvals for this tier should require a named legal reviewer rather than a manager's sign-off, and the approval should record the purpose, because a purpose-bound approval is what makes later reuse a decision rather than a drift. Two further habits matter here. Keep the population as small as the analysis genuinely requires, since a query returning every patient when a cohort would do multiplies the exposure without improving the answer. And write down the retention position at the point of approval, covering both your systems and the vendor's, because the question of when data will be deleted is far easier to settle before processing than after an auditor asks it.
Role-Based Access Control
Role-based access control, usually shortened to RBAC, grants permissions based on job role rather than to individuals one at a time. Instead of approving 400 individual data access requests per year, you define 15 access roles and manage those, and the saving is not only administrative. A role is reviewable in a way that a pile of individual grants is not, because someone can look at a role definition and form a view about whether it is too broad.
RBAC works well where job roles are stable and well defined. It becomes harder in organizations where AI projects cross functional boundaries, since a data science team building a model needs data from multiple systems, each with its own role structure and its own owner. In these cases a project-level access provisioning process is required: a defined request and review process granting temporary, scoped access to the data a specific project needs.
The critical elements of such a request are who is requesting access, to which specific data, for what specific purpose, for what period, and who in the data owner's function has approved the use. That last element is what separates project access governance from simply having a request form, because the data owner is the only person positioned to say whether the stated purpose is one the data may serve.
The discipline that keeps this from degenerating is granting the narrowest access that lets the work proceed. Requests written as system-level access when the project needs three tables, or as indefinite access when the project has an end date, are the mechanism by which excessive permissions accumulate. Neither is caught by a form that asks only whether the requester is authorized. Both are caught by a reviewer who asks what the requester is actually going to do.
Monitoring and Anomaly Detection
Access permissions define who should have access. Monitoring tells you what is actually happening, and the gap between the two is where security incidents live. For AI-related data governance, a small number of signals carry most of the value.
- Unusual export volumes: A user who normally queries 50-100 customer records per day and suddenly queries 50,000 warrants investigation.
- Access outside normal hours or location: Legitimate business access usually follows predictable patterns, and deviations are worth reviewing even when they turn out to be innocent.
- New API connections: When a user connects a dataset to a new external service, that should generate an alert for review rather than silently succeeding.
- Data flowing to uncatalogued endpoints: If your data catalog shows data leaving to a service that is not on your approved vendor list, that is a governance failure in progress rather than a risk to be assessed later.
Elena's team implemented a simple rule after the incident: any API call connecting internal data to an external service that had not been pre-approved generated an automatic notification to the data owner and a 48-hour review window before the connection was permitted to complete at scale. This caught the next three similar situations before they became problems.
The design is worth noticing. The rule did not block the connection outright, which would have pushed people toward less visible routes. It made the connection visible to the person with standing to judge it, and it bought enough time for that judgment to happen.
Access Revocation
Access that should have been removed is often as dangerous as access that should never have been granted, and it is considerably easier to overlook because nothing about it looks like an event. The most common failure mode is an employee changing roles or leaving the organization while their access to sensitive systems persists. For AI systems specifically, two situations need revocation processes beyond the standard offboarding checklist.
The first is a role change within the organization. An employee who moves from a product team to a finance team should lose product roadmap data access and gain finance data access, but if your governance system tracks "employee with access" rather than "employee in role X with access", role changes are invisible to it and permissions only ever accumulate. The second is the end of a defined project. Where access was granted for a specific AI project, build the end date into the grant itself and require a review step before renewal. Project data access that is never revisited accumulates quietly over months and years, and it is usually discovered by an audit rather than by the organization.
Balancing Security and Usability
A security framework so restrictive that people work around it is not a security framework. It is a theater that provides false assurance while creating a shadow data environment nobody can see. The analyst who connected patient data to a commercial AI tool did so partly because the formal path to getting that analysis done was unclear and slow, and the workaround felt safer than it was because the risk was invisible while the delay was not.
Every governance policy should be tested against one question: if a well-intentioned employee wanted to do the right thing and still accomplish their goal, how difficult is the compliant path? If the answer is that it is very difficult, the policy will be circumvented, and the circumvention will be done by exactly the conscientious people the policy was written for. Streamlining the compliant path is as important as defining the rules, and it is usually cheaper. A published classification reference, a request process with a known turnaround, and a pre-approved list of tools for each tier remove most of the pressure that produces workarounds in the first place.
Anti-Patterns
- Treating a vendor's compliance certification as permission. Whether a tool meets a standard and whether its contract permits your data to be used for training are separate questions, and the second is the one that creates exposure.
- Answering only the first access question. Systems that record who may see data rarely record what they may do with it or what happens to it afterward, which is where AI governance gaps concentrate.
- Applying uniform controls to every data type. Governing public product specifications like patient records produces compliance fatigue and drives conscientious people toward workarounds.
- Classifying case by case. A scheme requiring fresh judgment on every occasion will not be consulted on the occasions that matter; assign classification at the data type level and publish it.
- Approving project access without the data owner. A request form that checks only whether the requester is authorized cannot tell whether the stated purpose is one the data may serve.
- Granting system-level access for table-level work. Over-broad and open-ended grants are how excessive permissions accumulate, and neither is visible to a process that asks only about authorization.
- Tracking access by person rather than by role. Internal role changes then become invisible to governance, and permissions only ever grow.
- Blocking rather than surfacing. A hard block on new connections pushes people toward routes you cannot see; a notification to the data owner with a review window keeps the activity visible.
Practice Prompts
- Take one significant data category and answer all three access questions for it in writing: who may see it, what they may do with it, and what happens to it afterward. Note which of the three your current systems can actually enforce.
- Find the contract for the AI tool most widely used in your organization and locate the clause covering whether your data may be used for model training. Time how long it takes.
- Draft the one-page classification reference an employee would consult before connecting data to a new AI tool, and test it on a colleague with a real example.
- List the project-based data access grants made in the last year and check how many carried an end date. For those that did, check what happened at that date.
- Trace what would happen today if someone connected an internal dataset to an unapproved external service. Identify who would find out and how long it would take.
- Pick a governance rule your organization enforces and walk the compliant path yourself end to end, recording every step and every wait.
Reflection
The analyst in Elena's story did everything her training had prepared her to do. She checked the tool's compliance status, satisfied herself that it was appropriate, and got on with helping her colleagues. The governance system failed her rather than the other way around, because it had told her which question to ask and had not told her the other two existed. Consider the last time someone in your organization connected data to a new tool without a review. Was that a person disregarding the rules, or a person following the only rule they had been given?
Glossary
- Permissioned access versus permissioned use: The distinction between being allowed to see data and being allowed to do a particular thing with it, such as exporting it to an external service or using it to train a model.
- Data classification: Assignment of data types to sensitivity tiers, typically public, internal, confidential, and restricted, so that controls can be proportionate rather than uniform.
- Role-based access control (RBAC): Granting permissions by job role rather than to individuals, which makes access reviewable as well as cheaper to administer.
- Project access provisioning: A defined request and review process granting temporary, scoped access for a specific project, valid only with the data owner's approval of the stated purpose.
- Data processing agreement: The contractual instrument setting out what a vendor may and may not do with data you send them, including retention and training.
- Tokenization: Replacing identifiers with substitute values so that the linkage back to individuals stays inside systems you control.
- Shadow data environment: The unmonitored set of data flows created when people route around a governance process they cannot practically follow.
Related Lessons
- Data Quality & Master Data Management covers the upstream discipline that determines what is in the datasets these controls govern.
- Data Privacy & Compliance Governance develops the regulatory obligations that sit behind the restricted tier described here.
- Enterprise Data Architecture & Governance provides the catalog and ownership structures that monitoring and project access depend on.
- AI-Specific Security Threats & Defenses explains what an attacker can do with training data and model interfaces once access controls are bypassed.
- Supply Chain Security & Third-Party Risk addresses the vendor side of the contract questions raised throughout this lesson.
Closing
Nothing bad happened in Elena's organization, and that is the least useful thing about the story. The controls did not catch the problem; the legal team did, later, by chance. What changed afterward was not a tightening of the rules but a clarification of them: three questions instead of one, a published classification anyone could check in a minute and a half, a data owner notified whenever data reached a new external service, and a compliant path fast enough that nobody needed to route around it. Access governance done well is not mainly about restriction. It is about making the right answer available at the moment someone is about to act.
Key Takeaways
- Permissioned access is not the same as permissioned use. Who may see data, what they may do with it, and what happens to it afterward are three distinct decisions, and AI governance requires explicit answers to all three.
- Classify data by sensitivity at the data type level. A four-tier framework of public, internal, confidential and restricted enables proportionate controls without creating the compliance fatigue that drives workarounds.
- Special categories need purpose-bound approval, not just access approval. Health, financial and personal data carry rules about processing and purpose, so a legal reviewer and a recorded purpose belong in the approval itself.
- RBAC handles stable roles; project requests handle cross-functional AI work. Project requests need the data owner's approval of the stated purpose to be meaningful, and the narrowest grant that lets the work proceed.
- Monitor what is actually happening, not just what is permitted. Unusual export volumes, new API connections, and data flowing to uncatalogued endpoints are the signals that catch problems before they become incidents.
- Build revocation into the grant. End dates for project access and role-based tracking for internal moves prevent the quiet accumulation of stale permissions that audits eventually find.
- Make the compliant path easy enough that workarounds are unnecessary. Security too difficult to follow creates a shadow data environment more dangerous than the original risk the policy addressed.
Frequently Asked Questions
The tool is certified compliant. Is that enough? No, and this is the specific trap in Elena's story. Certification tells you the vendor meets a standard. It does not tell you what the contract permits them to do with the data you send, and retention or use for model training is governed by contract terms rather than by certification. Both questions have to be answered, and only the second one is usually missing.
How granular should classification be? Coarse enough to be memorable and applied at the data type level rather than per request. Four tiers cover most organizations, and the test of the scheme is whether someone about to connect data to a new tool can reach the right answer in the time they are willing to spend looking. A scheme that is more precise but slower will not be consulted.
Who should approve access to restricted data? A named legal reviewer rather than the requester's manager, and the approval should record the purpose it was granted for. Purpose-bound approval is what makes later reuse of the same data a fresh decision instead of an unnoticed drift, which is the mechanism behind most incidents of this kind.
Should we block unapproved connections outright? Usually not as a first move. A hard block pushes determined people toward routes you cannot see, which is worse than the visible risk it removed. Notifying the data owner and holding a review window before the connection completes at scale keeps the activity visible while still preventing large-scale processing before anyone has looked at it.
How do we find stale access we have already granted? Start with project-based grants, since those had a defined end and are the easiest to adjudicate, then examine internal role changes over the last year and compare current permissions with current roles. If your systems record access by person rather than by role, that comparison is the exercise that tells you how much of a problem you have.
Skill.re