Data Privacy Basics: What You Share with AI
Yuna runs a small physical therapy practice in Minneapolis with two therapists and a front-desk coordinator. When she started using an AI assistant to draft patient communications, she thought she was being efficient. Then a colleague mentioned that the free tier of the tool could use her conversations to improve the model, which meant the names, appointment details and treatment notes Yuna had been pasting in could be used in ways she had never consented to. Yuna was not doing anything malicious. She simply had not read the privacy policy. That is exactly where most small-business owners stand, and in any business that handles sensitive customer information the gap has consequences.
The situation is easy to walk into. You paste your customer email list in to draft personalised marketing messages. You feed payroll data in to analyse salary patterns. You upload a client spreadsheet full of confidential financial information. You copy product descriptions in to have them refined. All four are ordinary business tasks, and in each of them you have just sent data to somebody else's servers. Do you know where it goes, whether the vendor uses it, whether it will be used to train the model, or whether it could surface in someone else's conversation?
What Happens After You Press Enter
When you type something into an AI tool, that text travels to the vendor's servers and is processed there to generate a response. What happens next depends entirely on the tool and the plan you are on. Four things can happen to the text you submitted, and on most consumer tiers of most tools, all four are permitted by the terms you agreed to when you signed up. None of this is hidden. It is written into the terms of service. It is simply written where almost nobody reads it.
- Training. Your conversations may be used to train and improve future versions of the model. What you type today can become part of the material the model learns from, which means your data is feeding the same assistant that other users interact with.
- Model improvement review. Engineers may read conversations to find examples of where the model failed or succeeded, and feed those findings into the next version. A human being can therefore read what you wrote.
- Safety monitoring. Conversations may be reviewed by automated systems or by people, to detect abuse, illegal activity or safety problems. This is legitimate and most vendors do it. It also means your data is not completely private.
- Retention. Consumer tools typically retain conversation data for periods ranging from thirty days to three years. It is eventually deleted, but not immediately, and not on your schedule.
Put those four together and the honest summary is this: on a free or consumer tier, you are trading data for access. You get cheap or free AI, the vendor gets material to improve the product. That is a business model, not a privacy oversight, and no vendor is being sneaky about it. The practical rule that follows is blunt. If you are not paying for a tier that carries privacy commitments in writing, assume your data is being used for model training and behave accordingly.
What the Paid Tiers Actually Promise
Most AI vendors offer paid, business or enterprise tiers where the data policy changes materially. The common pattern is that the vendor's terms state that data submitted under those tiers is not used to train models, that safety monitoring is narrowed or moved inside your own organisation, and that you gain administrative controls, audit logs and retention settings you can configure. Some tiers keep data within your organisation's boundary. Business agreements and API arrangements can also include custom data processing terms negotiated for your account.
Read that paragraph again and notice the phrasing. Their terms state that your data is not used for training. That is a contractual commitment, not a technical impossibility, and it applies to the specific tier and settings you are actually on. Check the terms for the tier you are actually on, not the tier the marketing page describes. On some paid consumer plans, for example, enabling conversation history in settings can bring some of your conversation data back into scope for product improvement. A plan name on your invoice is not a privacy guarantee. The written data policy attached to that plan is what governs.
You do not need to understand every line of a privacy policy. You need to know one thing: does this tool use my data to train its model? If yes, treat it like a public notice board.
Consumer Versus Enterprise: The Real Difference
The words "consumer" and "enterprise" describe different levels of privacy commitment, not just different price points. This distinction is the single most important privacy decision you will make about AI tools, and it is worth laying the two side by side.
| Dimension | Consumer AI tools | Enterprise AI tools |
|---|---|---|
| Data policy | Data may be used for training and improvement | Terms state data is not used for training; a data processing agreement specifies what happens to it |
| Retention | Data kept for extended periods | You control retention and can typically delete data immediately |
| Monitoring | Subject to safety monitoring that may include human review | Monitoring is automated or happens within your organisation rather than at the vendor |
| Recourse | Limited. You accept terms of service; if they are breached your options are few | Legal agreements, service level agreements and liability provisions give you contractual remedies |
| Compliance | Not designed for GDPR, CCPA, HIPAA or similar regimes | Built for those regimes, with privacy-by-design features |
| Audit trails | None provided | Full logs of who accessed what data and when |
| Typical cost | Free to about $20 per month | $50 to $500 and up per month depending on scale and features |
| Best suited to | Personal use, brainstorming, non-sensitive tasks, low-stakes content | Sensitive data, customer information, regulated industries, high-stakes processes |
Stepping a single seat up from a consumer plan to a business plan typically costs somewhere in the region of $20 to $30 per person per month, before you get to the larger enterprise arrangements in the table. For any business that regularly touches confidential or client data, that is not a budget question. It is an obligation question, and the answer is the same whether you have three staff or thirty.
Four Data Categories That Decide Everything
Not all data carries the same risk, and trying to make a fresh judgement every time you paste something is how mistakes happen. A four-level classification system lets you and your team make fast, consistent decisions. You classify your data once, you set rules about which classification may go into which tools, and then the daily decision becomes a lookup rather than a debate.
| Classification | Examples | AI tool rules | Risk if exposed |
|---|---|---|---|
| Public | Blog posts, published content, marketing materials, general industry information, business hours and services | Any AI tool, including free consumer tools | None, it is already public |
| Internal | Internal processes, team communications, non-sensitive strategy, meeting summaries without sensitive detail | Consumer tools are workable but consider the privacy tier; never share outside the organisation | Moderate, competitive disadvantage if disclosed |
| Confidential | Customer data, financial information, business strategy, proprietary processes, contracts, vendor pricing, salary detail | Only business or enterprise tools with strong privacy guarantees; requires a data processing agreement; anonymise first where you can | High, legal, financial or competitive harm |
| Restricted | Passwords, API keys, payment information, healthcare data, government identifiers, highly sensitive personal data | Never share with any AI tool; handle separately in specialised secure systems | Severe, legal liability, fraud, identity theft |
For Yuna this framework changed her practice straight away. Generic communication templates with no patient detail in them: public, fine on any tier. Appointment reminder wording with names removed: internal, fine on the business tier the practice pays for. Patient treatment notes: restricted, full stop. She drafts those manually or dictates them into software designed for handling protected health information. The classification did not slow her down. It removed the hesitation that used to sit in front of every paste.
What Never Goes Into an AI Tool
Some data should never be sent to any AI tool, on any tier, from any vendor. This is the restricted category, and it is worth naming the contents explicitly rather than leaving people to infer them. Every item below is on the list because exposure is either irreversible or legally consequential, and often both.
- Payment card information. Card numbers, expiry dates and CVV codes should never be sent to any AI tool. The PCI DSS, the payment card data security standard, strictly forbids this.
- Passwords and API keys. If someone pastes an API key in to get help with it, that key is compromised from that moment. Anyone with access to the conversation could use it, and every integration and data store reachable through that key is at risk.
- Personally identifying information for customers or employees. Full names alongside email addresses, phone numbers or home addresses. Even where the vendor does not train on it, you are exposing other people's information to risk they did not agree to.
- Healthcare information. Patient records, medical history, health conditions of customers or staff. HIPAA severely restricts what may be done with health data, and AI tools are not HIPAA-compliant in most cases.
- Government identification. Social security numbers, driving licence numbers, passport details. Too sensitive, and too valuable to bad actors, to risk.
- Proprietary formulas and trade secrets. If something is a genuine competitive advantage that depends on staying secret, it does not go in. Once it has entered training data you have lost control of it.
- Biometric data. Fingerprints, facial images, iris scans. These are permanently identifying. A password can be changed after a breach; a fingerprint cannot.
Before pasting anything into an AI tool, ask one question: if this data were published on the internet tomorrow, would it harm me, my customers, my employees or my business? If the answer is yes, it does not go in.
What GDPR Asks of You
The General Data Protection Regulation is European law. It reaches you if any of your customers are residents of EU countries, even when your business sits entirely in the United States. GDPR requires that you only process personal data for purposes the customer has consented to, and that you have appropriate agreements in place with any third-party tool that handles that data. Five of its requirements bear directly on how you use AI, and none of them are exotic.
- Data processing agreements. If you use an AI tool to process personal data belonging to EU residents, you need a data processing agreement, or DPA, with that vendor. The DPA specifies how the data will be handled, who may access it, where it is stored, and what your rights are if there is a breach. Not every AI vendor offers one, and several offer them only to enterprise customers, which is one more reason the tier you choose is a compliance decision rather than a budget decision.
- Data minimisation. You may collect and process only the personal data you actually need. Dumping an entire customer table into an AI tool on the theory that the extra columns might turn out to be useful is precisely what this rule forbids. Use only what the specific task requires.
- Purpose limitation. Personal data gathered for one purpose, such as fulfilling an order, cannot be used for a different purpose, such as training models, without additional consent. The purpose you collected under is the purpose you are bound to.
- Data subject rights. People have the right to know that their data is being processed. If you are using AI to process it, you should disclose that. They also hold rights to access, correct and delete what you hold about them, and those rights do not stop at the boundary of a tool you happen to use.
- Data localisation. Personal data belonging to EU residents should ideally be stored in the EU, or somewhere with equivalent privacy protections. Many consumer AI tools store data on United States servers, which may not satisfy this requirement. If it matters to you, ask the vendor where the data physically sits before you sign up rather than afterwards.
The practical reality: if you process EU personal data, use only business or enterprise AI tools that explicitly commit to GDPR compliance and will give you a DPA.
What CCPA Asks of You
The California Consumer Privacy Act applies to businesses that collect personal data from California residents and meet certain size or revenue thresholds. It is less demanding than GDPR but it still shapes your tool choices, and three of its requirements matter here.
- Consumer rights. California residents have the right to know what personal data is collected about them, to request its deletion, and to opt out of the sale of that data. If you are running their data through AI tools, you have to be able to honour those rights, which means you have to know where the data went.
- Disclosure. Your privacy policy must disclose what categories of data you collect and how you use them. If AI tools are part of how you use them, that belongs in the disclosure.
- Limitations on use. Data you collected for one purpose cannot be repurposed for something entirely different without consent, which is the same principle GDPR calls purpose limitation.
For most small businesses using AI internally to write emails, draft documents and summarise meetings, direct regulatory exposure is limited. It grows the moment customer personal data enters the tool: names, contact details, purchase history, or anything touching health or finances. Where a third-party vendor is processing that data on your behalf, most privacy regimes expect a data processing agreement to exist between you. Business tiers come with them. Consumer tiers generally do not.
The Paste-Time Checklist
Classification and regulation are the background. What your team needs in the moment is something short enough to remember while a customer is waiting. Write this on one page and post it where people can see it.
- Before entering anything into an AI tool, ask which category it is: public, internal, confidential or restricted.
- If it is restricted, it does not go into any AI tool. Draft it another way.
- If it is confidential, use only the approved business-tier tool. Not the free version, not a personal account.
- If it involves a customer's personal information, whether name, health or finances, remove or replace the identifying details before pasting.
- If you are not sure, ask before you paste.
That protocol prevents most of the data privacy problems that arise in AI tool use. What remains are edge cases that surface as your use evolves. Handle them as they arise, and fold the answer back into the page so the next person does not have to work it out again.
Rolling It Out Across Your Team
A classification system that lives in the owner's head protects nobody. Turning it into something the whole team follows takes five steps, and none of them require a consultant.
- Step one: classify your data. Walk through the main types of information your business handles and label each one public, internal, confidential or restricted. Produce a one-page reference with real examples from your own operation, not generic ones. People match against examples far faster than they reason from definitions.
- Step two: write the tool guidelines. Document which tools are approved for which classifications. Public data can go into any tool, including free tiers. Internal data is acceptable on consumer tiers, though paid tiers are the better habit, and self-hosted open-source models are an option if you have the infrastructure to run them. Confidential data goes only into business or enterprise tools covered by a data processing agreement. Restricted data goes nowhere.
- Step three: train the team. Fifteen minutes is enough for everyone who touches an AI tool. Cover the four categories, the approved tools, and the consequences of using unapproved ones: confidentiality breaches, liability for exposure, competitive damage. Make it concrete. "This customer email list is confidential, so it goes into the approved business tool, not into a free assistant on your phone."
- Step four: write a breach response plan before you need one. People will occasionally paste sensitive data into the wrong tool despite everything above. Decide now what happens next. Document the incident immediately, notify whoever manages the business, submit a data deletion request to the vendor, and review how it happened so the same route closes. A plan written calmly beats a plan improvised in a panic.
- Step five: audit periodically. Ask the team openly what kinds of data are going into which tools. Look for mismatches between sensitivity and tool. If a pattern appears, such as confidential material repeatedly going into consumer tools, the fix is either retraining or removing access to the consumer tools. Patterns are a system problem, not a person problem.
Anti-Patterns
- Treating a paid plan as a guarantee. A business tier is a contractual commitment on a specific plan with specific settings, not a technical wall. Settings such as conversation history can change what is in scope. Read the data policy attached to the plan you are actually on, and re-read it when the vendor changes terms.
- Assuming that removing names makes data anonymous. Stripping a name from a record reduces risk. It does not by itself make the record unlinkable to a person, and it does not on its own put the record outside the scope of privacy law. Treat de-identified customer data as still sensitive, and keep it on the same tier you would have used before you stripped it.
- Reasoning from the tool's friendliness. A conversational interface feels private in a way an email to a stranger does not. That feeling is an interface property, not a security property. The data has still left your building.
- Letting the policy live only with the owner. If the rules exist but the front desk has never seen them, the rules do not exist. Post them, train on them, and revisit them.
- Deciding case by case under time pressure. Fresh judgement at the moment of pasting, with a customer waiting, is where restricted data ends up in consumer tools. The classification exists precisely so that the decision is already made.
Practice Prompts
Use these with invented or already-public material only. The point of the exercise is the thinking, not the data.
- "Here is a list of the kinds of information my business handles, written generically with no real customer detail. Sort them into public, internal, confidential and restricted, and tell me which ones you are unsure about and why."
- "Draft a one-page data handling policy for a small team, covering four data classifications, which AI tools are approved for each, and what to do if someone realises they pasted the wrong thing."
- "I am about to sign up for an AI tool. List the questions I should be able to answer from its privacy policy before I submit any business data, covering training use, retention, human review, deletion and where data is stored."
- "Rewrite this customer complaint so that it keeps the substance and removes every detail that could identify the person, then tell me what identifying signals remain in your rewrite."
- "Write a fifteen-minute training outline that teaches a non-technical team the difference between confidential and restricted data, using examples from a service business."
Reflection
Think about the last few things you personally pasted into an AI tool. Which category did each of them fall into, and would you make the same choice now? Then widen the question: which tier is each member of your team actually using, on which account, and do you know the answer or are you assuming it? Finally, if a customer emailed you tomorrow asking what happens to their information when your business uses AI, could you answer briefly without checking anything? The gap between the answer you would give and the answer you could evidence is the work in front of you.
Glossary
- Data Processing Agreement (DPA) A legal contract with a vendor specifying how your data is handled, stored, protected, who may access it, and your rights if there is a breach.
- Personally identifying information (PII) Data that identifies a specific person, such as a full name combined with an email address, phone number or home address.
- Data minimisation The principle of collecting and processing only the personal data actually required for a specific purpose.
- Purpose limitation The principle that data gathered for one stated purpose may not be reused for a different purpose without additional consent.
- Data localisation Requirements about the physical jurisdiction in which personal data is stored.
- Retention period How long a vendor keeps your data before deleting it.
- SOC 2 A security standard applied to software-as-a-service companies, audited by a third party.
- PCI DSS The payment card data security standard, which governs how payment card information may be handled.
- HIPAA United States regulation governing the handling of health information.
- Anonymisation Removing or masking identifying details from a record. It reduces risk; it is not by itself a guarantee that a person cannot be identified.
Related Lessons
- Data Security in AI-Integrated Systems covers the controls that protect the data once you have decided what may enter a tool.
- Regulatory Compliance: GDPR, CCPA, and Industry Standards goes deeper into the obligations sketched here.
- Creating Your Business AI Use Policy turns the classification and the paste-time checklist into a written policy.
- Data Privacy Obligations for Small Businesses looks at the same ground from the obligations side rather than the tool side.
- Vendor Risk Assessment for AI Tools gives you the questions to put to a vendor before you sign.
- Intellectual Property Considerations for AI takes up the next ownership question: who owns the content an AI tool produces.
Closing
Data privacy with AI tools is not complicated once it is systematised. Classify your data into four levels. Know which tools are appropriate for each level. Train the people who use them. Watch what is actually happening rather than what you assume is happening. Have a plan for the day something goes wrong. Consumer tools are fine for public material, business tiers are required for anything sensitive, and restricted data stays out entirely. Yuna did not become a privacy expert. She wrote down four categories and five rules, and the anxiety left the building with the ambiguity.
Key Takeaways
- On free and consumer tiers, assume your conversations may be used for training, human review and safety monitoring, and retained for anywhere from thirty days to three years. That is the trade you accepted at signup.
- Business and enterprise tiers state in their terms that data is not used for training. That is a contractual commitment on a specific plan with specific settings, so check the terms for the tier you are actually on.
- Classify your data into four categories: public, internal, confidential and restricted. The classification turns every future paste decision into a lookup.
- Business tiers provide Data Processing Agreements, the legal documents most privacy regimes expect when a third-party tool handles customer data on your behalf.
- GDPR reaches you if any customers are EU residents; it requires processing only for consented purposes, data minimisation, purpose limitation, respect for data subject rights, and attention to where data is stored. CCPA applies to California resident data above certain thresholds.
- Payment card data, passwords and API keys, healthcare records, government identifiers, biometrics and trade secrets never enter any AI tool, regardless of vendor or plan.
- Removing identifying details before pasting reduces risk substantially, but it is a reduction rather than a guarantee, so keep de-identified customer data on the same tier you would have used anyway.
- A one-page protocol plus a fifteen-minute training plus a written breach response plan prevents most small-business AI privacy mistakes before they happen.
Frequently Asked Questions
Does my AI tool use my data to train its model? On free and basic tiers, the usual answer is that conversations may be used to improve the vendor's models. Paid, business and enterprise tiers generally state in their terms that they do not train on your data, and some paid consumer plans depend on a settings choice such as whether conversation history is enabled. There is no universal answer. Check the data policy for your specific plan and your specific settings, and check it again when the vendor updates its terms.
What should I absolutely never put into an AI tool? Payment card information, passwords and API keys, customer personal data without explicit consent, employee personal information, healthcare information, government identification numbers, biometric data, and proprietary trade secrets. Where that kind of material genuinely has to be processed, it belongs in specialised systems built for it under an appropriate agreement, not in a general-purpose AI tool.
Are enterprise AI tools really more private than consumer tools? Generally yes. They commit not to train on your data, they provide audit trails and data processing agreements, and they are built for GDPR, CCPA and HIPAA style requirements. But "enterprise" on a price list does not automatically mean completely private. Always verify the specific terms with the vendor rather than relying on the tier's name.
Do I need to worry about GDPR and CCPA when using AI tools? Yes, if you process personal data belonging to EU residents or California residents. In practice that means securing proper data processing agreements and confirming that the vendor commits to the relevant compliance standards. For sensitive personal data, use only tools with adequate privacy commitments in writing.
What is a simple classification system I can implement this week? Four levels: public can go anywhere, internal stays inside the organisation, confidential goes only into business or enterprise tools, and restricted goes into no AI tool at all. Classify your data, write the tool rules for each level, train your team on the categories, and audit usage periodically. That single page prevents most data privacy problems in AI use.
Skill.re