AI can help a business draft, classify, summarize, search, and respond. But every useful result begins with information. The real question is not simply whether an AI tool works. It is whether your business remains in meaningful control of the customer records, internal knowledge, creative work, commercial plans, and other proprietary information that make the tool useful.
That concern is reasonable. A team may be comfortable sharing a public product description with an AI service but unwilling to upload contract terms, source code, client correspondence, pricing logic, or an unreleased strategy. The difference is not resistance to AI. It is sound judgment about which information can leave a trusted boundary, who may receive it, and what happens after it is sent.
This guide provides a security-first way to assess AI data privacy. It explains what business control should mean in practice, which questions to ask third-party providers, when private or self-hosted options help, and how to design useful workflows without treating confidential information as an unlimited input.
Why trust is the real AI adoption question
Most businesses already entrust information to outside companies. Email, accounting, payments, file storage, customer support, and payroll commonly involve service providers. AI does not create the idea of third-party processing, but it can make the boundary harder to see. A person can paste a large document into a conversational interface in seconds, and the resulting convenience can conceal how much context has just crossed into another system.
Trust should therefore be specific. It is not a general feeling about whether a provider is well known. It is confidence, supported by evidence, that a particular service will handle a defined category of information for an agreed purpose under controls your business can accept. A provider may be suitable for public marketing copy and unsuitable for confidential acquisition documents. The same provider may offer consumer, developer, and enterprise services with materially different terms.
The US Federal Trade Commission has warned AI companies that privacy and confidentiality commitments must match their actual practices. Its guidance notes that model-as-a-service providers may receive customers’ internal documents and data, and that promises about purposes such as training matter wherever they are made, including terms and marketing. This is not evidence that every provider misuses information. It is a reason to require clear commitments and verify that the product, account type, and settings you use are covered by them. See the FTC’s guidance on privacy and confidentiality commitments for AI companies.
For a business buyer, the practical test is simple: can you explain to a client, employee, director, or regulator what information the system receives, why it receives it, who else can access it, and when it is removed? If the answer depends on assumptions or a generic statement that the service is “secure,” the assessment is incomplete.
Identify the information at stake before choosing a tool
“Business data” is too broad to guide a decision. Start by naming the information that a proposed workflow will use. A customer-reply assistant may receive names, contact details, order problems, account notes, and the company’s approved policies. A proposal assistant may receive pricing, margins, negotiation limits, staff biographies, previous bids, and a client’s confidential brief. A software assistant may receive source code, infrastructure details, bug reports, or access credentials if no boundary prevents it.
Some information belongs to the company in a commercial sense: product designs, operating methods, forecasts, unpublished research, formulas, code, or strategy. Some is information the business has been entrusted to handle: customer, employee, patient, student, supplier, or partner data. Those categories can overlap, but they should not be discussed as if a company simply “owns” every record. Privacy rights, contractual duties, professional obligations, and confidentiality promises may apply even when the data sits in a company system.
Proprietary information can lose value when control becomes unclear
The World Intellectual Property Organization explains that trade secrets can include algorithms, code, data, and confidential technical know-how when the information has commercial value because it is secret and the holder takes reasonable steps to protect it. Digital information is especially easy to copy and transmit. WIPO’s guidance on trade secrets and digital objects is useful here: keeping confidential information under control and limiting it to authorized people are practical parts of protecting its value.
Not every internal document is automatically a legal trade secret, and the consequences of disclosure depend on the facts and jurisdiction. The business lesson is broader. Information that provides an advantage deserves a deliberate boundary before it is supplied to an AI tool. Do not wait for a procurement questionnaire to discover that employees have already pasted sensitive material into unapproved services.
Classify by consequence, not by file type
A spreadsheet is not inherently sensitive, and an ordinary email can be extremely sensitive. Classify information according to what could happen if it were exposed, reused, altered, or made unavailable. A workable scheme might separate public information, internal routine information, confidential business information, restricted personal information, and crown-jewel material whose disclosure could cause serious commercial or legal harm.
Then connect each class to a rule. Public information may be suitable for approved general tools. Internal information may require a business account and defined retention. Restricted information may require a purpose-built workflow with minimization, contractual review, and named access. Crown-jewel material may stay outside external AI services unless leadership approves a design with stronger isolation. The labels matter less than having rules people can follow.
What meaningful control over AI data actually means
Control is often reduced to the location of a server. Location matters, but genuine control is the ability to make and enforce decisions throughout the information’s life. A useful assessment covers at least six dimensions.
- Purpose: your business decides the specific job the data may support and prevents unrelated reuse where required.
- Selection: the workflow sends only the fields and documents needed for that job.
- Access: named roles, services, and support processes determine who can reach the information.
- Use: you know whether inputs and outputs can be used for service delivery, safety monitoring, product improvement, or model training.
- Retention: you understand how long content, logs, backups, and derived records remain, and which settings or agreements govern them.
- Accountability: you can review activity, revoke access, investigate incidents, export necessary records, and obtain a credible answer when something changes.
These dimensions expose several misleading shortcuts. “Your data is not used to train models” does not necessarily mean the service never retains prompts, permits authorized human access, records diagnostic content, or sends data to subprocessors. “Encrypted” does not mean a provider cannot process information in readable form when the service requires it. “Stored in this region” does not by itself describe support access, backups, or every subprocessor. Each statement may be valuable, but each answers only part of the question.
Control also includes the right to say no. A team does not need to upload its most sensitive material merely because a tool can analyze it. A useful AI program can begin with lower-risk work, approved knowledge, redacted examples, or synthetic test data. Our guide to choosing a first AI business workflow explains how to start with a narrow, observable task instead of granting broad access at the outset.
Follow the complete data journey
Before approving an AI workflow, trace one real example from the moment a person selects information to the moment the final result is stored or acted upon. Include the user’s device, the application, databases, model provider, file storage, monitoring tools, identity provider, integrations, and any service that receives the output. This does not require an elaborate architecture diagram. A clear list of transfers, stores, and responsible parties is enough to reveal gaps.
Separate storage from processing. A product may keep its main database on infrastructure you control while sending selected text to an external model. Another service may process data temporarily and keep operational logs elsewhere. A support integration could copy an output into a ticket. A human approval step might prevent an email from being sent, but it happens after the text has already been processed by the model. Review before action and review before disclosure solve different problems.
Ask five questions at every handoff
- Exactly what information leaves the current system?
- Which organization and service receive it?
- What purpose authorizes that transfer?
- How long could the information or a derived record remain?
- Which technical, contractual, and human controls apply?
Imagine a consultancy using AI to draft a proposal. The model needs the prospect’s stated goals and an approved description of the consultancy’s services. It may not need the complete CRM history, another client’s proposal, staff salary data, or the company’s minimum acceptable margin. Selecting a smaller context can preserve the useful task while keeping unrelated proprietary information within the business boundary.
Data mapping also improves incident response. If a credential is exposed or a supplier changes its terms, the team can identify which workflows, information classes, and customers may be affected. Without that map, a business may know it uses “AI” but be unable to determine where or how.
How to evaluate a third-party AI provider
A third party is not automatically unsafe. Many providers can offer mature security operations, resilience, and independently reviewed controls that a small business could not reproduce alone. The decision should rest on evidence relevant to the service you will actually use. A certification logo, a famous brand, or a general security page cannot replace answers about your product tier, data flow, and agreement.
Purpose and model improvement
Ask whether customer content is used only to provide the requested service or may also support product improvement, evaluation, abuse monitoring, or training. Define “customer content”: does it include prompts, uploaded files, outputs, feedback, metadata, and support conversations? Check whether an opt-out is available, whether it is on by default, who can change it, and whether the promise is contractual.
Read current terms rather than relying on a screenshot or an answer about a different offering. The FTC has separately cautioned that quietly changing terms to permit new uses of previously collected data may be unfair or deceptive. That enforcement view reinforces a sensible procurement practice: record the terms and settings on which your approval depends, assign someone to monitor material changes, and provide an exit route.
Retention, deletion, and support access
Ask for the ordinary retention period and every important exception. Security or abuse logs, backups, legal preservation, failed requests, and support tickets may follow different schedules. Determine whether an administrator can set shorter retention, delete a conversation, delete an account, or request deletion of retained content. Ask what deletion means for backups and how long full removal may take.
Human access needs a precise answer. Can provider personnel see customer content for support, safety review, or incident investigation? If so, what triggers access, who authorizes it, how is it logged, and is the customer notified where appropriate? A claim that access is limited is more useful when the limiting process is described.
Subprocessors, location, and legal roles
Identify the organizations that help deliver the service, the processing locations that matter to your obligations, and the notification process for changes. If personal information is involved, establish the parties’ privacy roles for the relevant jurisdiction and document them. The UK Information Commissioner’s Office explains that a controller decides why and how personal data is processed, while a processor acts on the controller’s behalf; the real arrangement matters more than the label in a contract. See the ICO’s guide to controllers and processors. Obtain legal advice for obligations specific to your business or location.
Security evidence and incident response
Look for controls that address your risks: strong authentication, role-based access, encryption, secret management, vulnerability handling, isolation between customers, logging, recovery, and a tested incident process. Request the scope and date of independent assurance rather than assuming it covers every product. Ask how quickly the provider will notify you of an incident, what information it will provide, and how your data can be isolated or removed if the service relationship ends.
The European Data Protection Board has also cautioned that AI models should not simply be presumed anonymous. Its 2024 opinion on AI models and GDPR principles says anonymity must be assessed case by case, including whether personal data can be extracted from the model or obtained through queries. That is European guidance, but the underlying procurement lesson travels well: claims such as “anonymized” need evidence and a defined test.
Build security into the workflow, not around it
Provider review is only one layer. Your own application and operating process determine which users can submit information, what context is added, where outputs go, and what actions the AI can influence. Security should be present before a prompt is sent, while a result is handled, and after the task is complete.
Minimize and separate
Send the smallest useful context. Remove fields the task does not need, redact identifying details where practical, and keep unrelated customers or projects separate. Use an approved knowledge collection rather than connecting an assistant to every shared drive. If the system retrieves documents automatically, enforce the user’s existing permissions before retrieval so the model never receives material the user could not open directly.
OWASP lists sensitive information disclosure as a major risk for applications using large language models. Its guidance covers personal information, confidential business data, credentials, and other restricted material, and recommends controls including data sanitization, clear retention policies, user education, and least-privilege access to external sources. See OWASP’s Sensitive Information Disclosure guidance.
Keep secrets out of prompts and broad data stores
Passwords, private keys, tokens, and connection strings should use dedicated secret-management controls. Do not rely on a prompt telling a model never to reveal them. Likewise, do not place authorization rules only in instructions to the model. The application must decide whether a person may access a record or perform an action before giving the model the underlying data or tool.
Approve both disclosure and action
Human review is valuable when an AI draft could create a commitment, publish information, contact a customer, or change a record. But review at the end cannot undo an inappropriate upload at the beginning. High-risk workflows need two gates: an input policy that controls what may be disclosed, and an output or action policy that controls what may happen next.
That distinction is central to understanding AI agents versus conventional automation. The more tools, data sources, and actions a system can reach, the more important enforced permissions, narrow scopes, and explicit approval become. Capability should expand only when the evidence from a smaller deployment supports it.
Log enough to investigate without creating another archive
Logs can show who used the system, which workflow ran, which provider received a request, whether an approval occurred, and whether an action succeeded. They do not always need a permanent copy of every prompt and response. Decide which content is necessary for support and security, restrict access to it, and define an expiry period. Test whether deletion settings and account removal behave as documented.
Self-hosted and cloud AI are deployment choices, not security verdicts
Deployment remains relevant because it changes the people and systems you must trust. A managed AI application may place the application, database, and model access with one supplier. A self-hosted application can keep accounts, workflow records, and configuration under infrastructure your business controls while still sending selected prompts to an external model provider. A fully private arrangement may operate both application and model within a controlled environment.
Each pattern can be designed well or badly. A private server with weak administration, delayed security updates, excessive permissions, or untested recovery is not made safe by its location. A managed service with clear contractual limits, strong access controls, short retention, and independent assurance may be appropriate for some sensitive work. A hybrid design can retain core records internally and disclose only approved context to a selected provider.
The useful question is therefore not “cloud or self-hosted?” in isolation. Ask which parts of the system your business needs to control directly, which specialist responsibilities a provider can handle, what information crosses each boundary, and whether the remaining exposure is acceptable. The UK National Cyber Security Centre’s guidelines for secure AI system development place security across design, development, deployment, and ongoing operation. That lifecycle view is more reliable than treating a hosting decision as the end of the security work.
For especially sensitive material, private or local model operation can reduce disclosure to an external inference provider. It also creates responsibility for model selection, updates, access, monitoring, and infrastructure protection. Make that choice when the information risk justifies it and an accountable operator can maintain the environment. Do not promise absolute privacy merely because the model is local; integrations, telemetry, backups, and administrative access still need review.
What a security-first approach means at Waldok
For Waldok, security first means deciding the information boundary before optimizing the AI feature. We begin with the business task, identify the minimum data it requires, name the people and services that can access it, and keep consequential actions visible to a human. The goal is useful assistance with limits that a business can understand and operate.
Our published security approach and control principles emphasize data ownership, least-privilege access, encryption, logging, updates, and human approval. Those principles guide product and custom-system discussions; they are not a substitute for evaluating the exact deployment, provider, and information involved in a customer’s workflow.
Waldok products such as FacePrompt and ReplyPilot are described in the AI product catalog as self-hosted applications that use customer-supplied credentials for supported AI providers. That arrangement can give a customer direct control of the application environment and provider account. It does not mean information sent for model processing never reaches the selected provider. We prefer to state that boundary plainly so a customer can assess provider terms and decide what content is appropriate.
For custom systems, the Waldok delivery process can define data sources, access, approvals, hosting, retention, validation, and handoff as one design problem. A good solution may use a managed service, a controlled application with an external model, a private model, or a combination. The architecture follows the information risk and the business’s capacity to maintain it.
A practical AI data privacy checklist
Use this list before a pilot and revisit it before adding a new data source, provider, integration, or autonomous action.
- Purpose: Can we state the exact business task and the decision or action the AI will support?
- Information: Which fields, documents, messages, and metadata will the workflow receive?
- Classification: Does the input include personal data, client-confidential material, regulated records, credentials, or commercially sensitive knowledge?
- Minimum necessary: What can be omitted, redacted, summarized, or replaced with approved reference content?
- Recipients: Which application providers, model providers, subprocessors, integrations, and support teams can receive or access the data?
- Permitted use: Are inputs or outputs used for training, evaluation, product improvement, safety review, or any purpose beyond delivering the service?
- Retention: How long do primary records, logs, backups, and support records remain, and can the business change those periods?
- Access: Are strong authentication, named roles, least privilege, account removal, and administrative logging available?
- Action: Can the system publish, send, update, purchase, delete, or disclose anything, and which steps require human approval?
- Evidence: Do the contract, product documentation, settings, and independent assurance support the provider’s claims?
- Incident response: Who can suspend access, revoke credentials, identify affected data, contact the provider, and notify the right people?
- Exit: Can we export necessary records, remove integrations, delete content, and continue the business process elsewhere?
Document the decision in plain language. Record which uses are allowed, which information is prohibited, who owns the approval, and when the review expires. A one-page decision that staff can follow is often more protective than a long policy nobody can translate into daily behavior.
Common questions about security and AI data privacy
Is it safe to give business information to an AI service?
It depends on the information, the service, the account and settings, the agreement, and the controls around the workflow. Public or low-risk content may be suitable for an approved service. Confidential client information, trade secrets, credentials, and restricted personal data need a stronger assessment and may require minimization, a different service, or a decision not to disclose them.
Does “not used for training” mean our data is private?
It is an important commitment, but it addresses one use. Also ask about processing, retention, logs, human access, subprocessors, legal requests, deletion, and security incidents. Confirm that the commitment applies to the specific product and account your team uses.
Is self-hosted AI automatically more secure?
No. Self-hosting can give a business direct control over parts of the application and data environment. It also makes the business or its operator responsible for configuration, updates, access, monitoring, backups, and recovery. If the application calls an external model, approved inputs still leave the server for processing. Judge the complete data path.
Should employees be banned from using AI?
A blanket ban can be difficult to enforce and may drive use out of sight. A clearer approach defines approved tools and accounts, prohibited data, permitted tasks, escalation routes, and consequences for misuse. Training should use real examples from the business so staff can recognize the boundary. Some organizations may still prohibit particular tools or data classes when the risk warrants it.
Can anonymization solve the privacy problem?
Removing direct identifiers can reduce risk, but context may still identify a person or reveal confidential business facts. Effective anonymization depends on the dataset, available external information, and how the data can be queried or combined. Treat it as a control that needs testing, not a label that ends the assessment.
What should a small business do first?
Choose one narrow workflow, classify its information, approve one service and account type, minimize the input, and require review before any external action. Record provider terms and key settings. Run representative tests, including prohibited-data examples, before giving the workflow to a wider group. Expand only after the team can explain the data path and operate the controls.
Research references and further reading
- US Federal Trade Commission: AI companies must uphold privacy and confidentiality commitments.
- US Federal Trade Commission: changing terms for new data uses.
- World Intellectual Property Organization: trade secrets and digital objects.
- UK National Cyber Security Centre: guidelines for secure AI system development.
- OWASP: sensitive information disclosure in LLM applications.
- US National Institute of Standards and Technology: AI Risk Management Framework and Generative AI Profile.
- European Data Protection Board: opinion on AI models and GDPR principles.
- UK Information Commissioner’s Office: guide to controllers and processors.
Your business should not have to surrender control of its proprietary information to gain value from AI. Start with the data boundary, demand clear evidence from every provider, and build the workflow so people can see and enforce what the system is allowed to know and do. If you want to assess a real use case, share the workflow and data concerns with Waldok.
