Shadow AI is the use of artificial intelligence tools the company never approved or contracted. The risk is not the use itself: it is that corporate information leaves the perimeter with no contract, no record and no way to get it back. IBM links it to 20% of the breaches it studied in 2025 and to roughly $670,000 of extra cost per incident. The fix is a catalogue of approved tools, one clear rule about what data never leaves, and an internal alternative that is just as convenient.
What shadow AI is, and why it is not ordinary shadow IT
Shadow IT has been in every IT manager's vocabulary for twenty years: the employee who saves a file to a personal cloud because the corporate server is slow. Annoying, but bounded. The file moved; it was still the same file and it could be retrieved.
Shadow AI changes two things, and both of them matter:
- The data is not moved: it is handed over. The employee does not store the contract somewhere else, they give it to a third party for processing. What happens to that text afterwards depends on terms of service nobody at the company has read or negotiated.
- There is no purchase trail. Classic shadow IT left traces: an installation, a subscription, a card payment from some department. Here, opening a browser tab is enough. The indicators IT uses to spot unapproved tools simply never fire.
That is why industry surveys vary so much — depending on the study, between half and four fifths of workers admit to using unapproved AI — but agree on the essentials: the number is always far higher than management believes. And it is higher because the right question is not «how many people use AI?» but «how many people have found a shortcut that saves them an hour a day and have no intention of giving it up?».
Why it happens: this is friction, not rebellion
It is worth stripping the moral tone out of this, because the moral tone is what stops it being fixed. Nobody pastes a contract into a public chat to harm their employer. They do it because forty pages have to be summarised before six o'clock and the tool works.
The four frictions that generate shadow AI are almost always the same:
- There is no internal alternative. The company offers no AI tooling at all, or what it offers is worse than the free one outside.
- Asking permission takes longer than the work. If onboarding a tool means six weeks of committee, the tool gets used without being onboarded.
- The rule does not exist, or nobody has read it. «Use AI sensibly» is not a policy; it says nothing about which information may leave and which may not.
- The process forces copy and paste. When information lives in loose PDFs, email and spreadsheets, someone has to extract it by hand — and once it is out, where it gets pasted is anyone's guess.
The fourth is the interesting one, because it points at an architecture problem rather than a behaviour problem. A process that can only be completed by copying data from one screen to another manufactures shadow AI, and no internal memo is going to stop it.
What actually leaks
In documented cases, what leaves is not trivia: it is the asset. Here is the classification that is useful for writing a policy, in ascending order of severity.
| What gets pasted into the chat | What is at stake |
|---|---|
| Generic text, to improve the wording | Nothing material. This is the use you should explicitly allow, so it is not confused with the rest. |
| A commercial proposal or a price list | Trade secrets. No regulator will fine you, but competitive position is lost if it ends up in the wrong place. |
| A payslip, a medical record, a customer list | Personal data processed by a third party with no processor agreement and, frequently, no legal basis and no notice to the data subject. |
| A customer contract | Breach of the confidentiality agreement signed with that customer. The damage is contractual and immediate. |
| Proprietary source code, or credentials pasted «for review» | Intellectual property and attack surface. Credentials in a chat history are compromised credentials. |
The line you need to make explicit runs through the type of information, not the tool. A policy that says «AI is forbidden» is broken on day one; one that says «no personal data, no contracts and no code in tools outside the catalogue» fits in a sentence and can actually be enforced.
What it costs: the numbers that already exist
Discussions about shadow AI stay hypothetical until a number turns up. IBM's Cost of a Data Breach Report for 2025 supplies three worth keeping to hand:
- One in five breached organisations reported that shadow AI was involved.
- Those breaches cost on average around $670,000 more than breaches without that factor.
- 97% of organisations with an AI-related breach had no access controls over those systems.
In the same report, around two thirds of the organisations studied had no AI governance policy or were still drafting one. In other words: the problem is not that controls are failing, it is that in most places there is nothing yet that could fail. That is where an afternoon's work pays better than any tool.
What the law already requires
Even if nobody has knocked on your door yet, the European framework already binds you on two separate fronts.
GDPR. If an employee pastes personal data into an external tool, that tool acts as a processor. With no processor agreement under Article 28 of Regulation (EU) 2016/679, no record of processing activities and no guarantees about where processing takes place, the disclosure is unlawful. And liability sits with the controller — the company — not the employee.
EU AI Act. Regulation (EU) 2024/1689 introduced an AI literacy duty in Article 4, applicable since 2 February 2025: deployers must take measures to ensure a sufficient level of AI competence among their staff, including awareness of the risks. And since 2 August 2026 the transparency obligations of Article 50 apply. It is hard to train a workforce on tools you do not know exist. We covered the wider picture in what the EU AI Act requires from process automation.
For anyone operating under the NIS2 directive, there is one more layer: the duty to manage digital supply-chain risk is hard to sustain when half the workforce has onboarded suppliers that appear in no inventory.
Why banning it is the worst available policy
The instinctive response — block the domains, circulate a prohibition — fails for a simple reason: it does not remove the need. The report is still forty pages long and the deadline is still the same. All that changes is where the work happens and whether the company gets to know about it.
Blocking also carries a double cost that rarely gets counted:
- Usage moves to the personal phone, where there is no proxy, no logging and no possibility of control at all. A visible risk becomes an invisible one.
- You lose the conversation. Whoever found a genuinely good use case stops sharing it, and with it goes the information that would have turned it into an official process saving hours for a whole department.
The policy that works looks more like the one eventually applied to cloud storage: nobody banned saving files outside the server; a better place was provided, and it was made clear what could not leave it.
A 30-day plan to bring it back under control
You do not need a twelve-month governance programme to stop flying blind. Four weeks are enough to go from «we know nothing» to «we have a framework and it works».
- Week 1 — Look. Three sources: outbound proxy or firewall logs (which AI domains are visited and from which departments), the inventory of applications connected by OAuth to mail and office tooling, and a five-question anonymous survey about what people use it for. The survey explains what the logs cannot.
- Week 2 — Write a rule that fits in one sentence. What information never leaves (personal data, contracts, code, credentials), which tools are approved, and who to ask when in doubt. One page. If it runs to twelve, nobody reads it.
- Week 3 — Provide a better alternative. One or two tools under contract, processing inside the European Union, with no use of your data for training. Without this, the two previous steps are wallpaper.
- Week 4 — Open a fast lane. A form to request a new tool, with a commitment to answer within five working days. The speed of that channel determines next year's shadow AI level more than any other measure.
One note on week 3: the internal alternative does not win by being corporate, it wins by being convenient. If it takes three more clicks than opening a tab, it loses.
How Dokuflex solves it: the data never has to leave
Shadow AI grows out of one specific gap: the information is in one place and the intelligence is in another, so somebody has to bridge them by copying and pasting. The structural way to close it is to let the AI work where the document already lives.
- The AI runs inside the process. Classifying an invoice, extracting contract fields or summarising a case file happens as one more step of the flow in Dokuflex BPM low-code, on the document already held in document management. Nobody needs to export anything to ask for help.
- Permissions are inherited, not reinvented. The assistant sees exactly what the user invoking it sees: same authentication, same roles, same folders. That is the difference between AI wired into the corporate directory and an anonymous browser tab.
- It leaves a trail. Every AI intervention is recorded as an action on the case: which document, which operation, which user and when. That is what makes an answer auditable, and what lets you answer an inspection without reconstructing anything.
- Processing stays under control. The model is queried with the minimum necessary context and under contract, following the architecture we describe in LLM and RAG under the European GDPR. The company knows what leaves, where it goes and under what guarantees.
Bluntly: there is no feature that «removes shadow AI». What there is, is removing its reason to exist. When asking for a summary inside the case file is faster than opening another tab and pasting the PDF, the shortcut stops being worth it.
Frequently asked questions
What is shadow AI? +
It is the use of artificial intelligence tools inside a company that the organisation has never approved, contracted or assessed: a chat assistant open in a personal browser tab, an extension that summarises email, an online translator someone pastes a contract into. There is no contract, no record and no control over where the information ends up.
How is it different from ordinary shadow IT? +
In classic shadow IT an employee stored a file in an unapproved service: the data moved, but it was still the same file and it could be retrieved or deleted. With shadow AI the employee hands the content over to a third party to be processed, and that content may be retained, used to improve a service or exposed in a breach. Adoption also requires no purchase: opening a browser tab is enough, so the usual spend and installation signals never fire.
What does shadow AI cost a company? +
IBM's Cost of a Data Breach Report 2025 found that 20% of the organisations studied suffered a breach involving shadow AI, at an average of roughly 670,000 dollars more than incidents without that factor. The same report notes that 97% of organisations with an AI-related breach had no access controls over those systems.
Does banning AI solve the problem? +
No, it hides it. A ban does not remove the need that sent the employee looking for the tool: the forty-page report still has to be summarised before six o'clock. What changes is that they stop telling anyone, so the company loses the only source of information it had about what is being used and why. A short catalogue of approved tools, one clear rule about what data never leaves and a fast channel to request a new tool all work better.
What does the EU AI Act require here? +
Regulation (EU) 2024/1689 has imposed an AI literacy duty since 2 February 2025 under Article 4: deployers of AI systems must take measures to ensure a sufficient level of AI competence among their staff, including awareness of the risks. Since 2 August 2026 the transparency obligations in Article 50 also apply, covering interaction with AI systems and artificially generated content. A workforce using tools the company does not know about is, by definition, a workforce it has been unable to train.
Where do you start detecting shadow AI? +
With three sources that are almost always available: outbound logs from the firewall or proxy (which AI domains are visited, and from which departments), the inventory of extensions and applications connected to your mail and office environment through OAuth, and an anonymous internal survey asking what people use it for. Together they produce a reasonable map in a week, and the survey usually reveals the use cases the logs cannot explain.
What information should never go into an unapproved AI tool? +
Personal data about customers, patients or employees; contracts and pricing documents; proprietary source code; anything covered by trade secrecy or by a confidentiality agreement with a third party; and any data subject to a sector retention duty. The rule people actually remember is simpler: if you would not email it to a supplier with no signed contract, do not paste it into an AI chat.
Sources
- IBM — Cost of a Data Breach Report 2025: shadow AI involvement in breaches, associated extra cost, and the absence of access controls and governance policies.
- Regulation (EU) 2024/1689 (EU AI Act): Article 4, AI literacy; Article 50, transparency obligations.
- Regulation (EU) 2016/679 (GDPR): Article 28, processing by a processor.
Take away the shortcut's reason to exist
Book 30 minutes: we take one real process of yours, look at what information gets copied by hand today, and what it would take for AI to work on the document without it ever leaving home. No commitment, no product pitch.