
Shadow AI Is the ROI Nobody Measures.
MIT found 95% of enterprise AI pilots return nothing measurable. Workers in 90% of companies are still getting useful work done on personal tools.
95%
of enterprises in MIT NANDA's study reported no measurable return from GenAI pilots
90%
of companies where workers use personal AI tools for work, according to the handoff research
8.5%
of the 22.4 million prompts Harmonic analyzed contained sensitive corporate data
The paradox is attribution
The headline is not that companies bought AI and got nothing. It is that the company's official measurement often sees the purchase more clearly than it sees the work. MIT NANDA's 2025 State of AI in Business report found that 95% of organizations were getting no measurable return from the $30 billion to $40 billion enterprise GenAI investment it examined. Only about 5% of organizations had moved from pilots to material value extraction.
That result sits beside a different reality. McKinsey's State of AI research found 88% of organizations using AI in at least one business function. Harmonic Security analyzed 22,458,240 enterprise prompts from January through December 2025. The activity is not absent. It is split across programs with different owners, accounts, controls, and ways of counting value.
The mistake is to call that a simple adoption problem. A company can have high usage and low measured return when its approved program is not attached to a specific decision, while employees solve narrow tasks elsewhere. The return has not necessarily disappeared. It has become difficult to attribute, repeat, and defend.
The 95% figure should be read as a warning about the measurement unit, not as a claim that every pilot failed in the same way. Enterprise surveys use different samples and definitions. What the finding makes visible is the distance between an investment that can be named and a benefit that can be traced. If the unit is a license, the program can look adopted. If the unit is a changed decision with a known baseline, the proof gets much thinner.

Two AI programs, one company
Think of the official program as a visible road. Procurement bought it, IT configured it, and finance can point to the invoice. The road may be perfectly real. But its existence does not tell you whether it reaches the task employees need to finish, or whether anyone knows who owns the result.
The shadow program is a network of side paths. A marketer pays for a personal assistant to turn a messy interview into a first draft. A sales operator uses a public research tool to prepare for a call. A developer keeps a separate assistant because it handles a narrow code task faster than the approved environment. These are not the same use case, and they should not be flattened into one count of AI adoption.
Two programs. One P&L.
The visible program is not always the productive one.
Tap each route to separate the signal finance can see from the work employees actually use.
Official program
01The company can see the subscription, but not always the work it was meant to change.
This split explains why the official program can look busy and still fail an ROI review. Seat counts measure access. Prompt counts measure activity. Neither proves that a decision improved, that the output was checked, or that the workflow can be repeated without exposing data. The shadow path can have the opposite problem: a person has a concrete gain, but nobody records the task, the inputs, the quality check, or the boundary where the method stops working.
That makes attribution a design choice. A finance team reviewing an AI budget needs to know whether the program changed cycle time, error rate, conversion quality, or the number of handoffs. A usage dashboard cannot answer those questions by itself. The company has to connect the tool to a job before it can decide whether the return belongs to the tool, the person who adapted it, or a process change that would have happened anyway.
Why the official route stalls
Formal AI programs often start at the level of a platform rather than a job. An enterprise buys a general assistant, enables accounts, publishes a policy, and waits for transformation. The missing step is the operating question: which recurring decision should this tool improve, for whom, with what data, and how will the organization know the old process changed?
Governance can add friction without adding direction. A worker may need several approvals before trying a tool, but still have no approved template for the task. The result is a strange combination: high concern about the risk and low clarity about the work. The safe path becomes harder to choose because nobody has designed it around the person's actual problem.
The official stack also tends to reward what is easy to report. Renewals, enabled users, and usage dashboards travel upward. Time saved on a small task, fewer handoffs, or a better first pass can stay local. When the proof is not designed at the beginning, a weak result later gets blamed on the model even if the real failure was an undefined workflow.
There is an approval tax as well. If a worker needs permission to test a tool but has no clear route for describing the task, the safe process feels like a queue with no destination. That does not make the personal tool safe. It explains why the personal tool wins the first experiment. The approved alternative has to reduce uncertainty and friction at the same time: name the allowed data, supply a starting pattern, and say who reviews the output.

Why the shadow route works
Shadow AI is attractive for a reason that has little to do with rebellion. It starts with a task already in front of a person. The worker chooses a tool that fits the job, pays the small cost in time or money, and sees an immediate output. The loop is short enough to learn from.
The danger is that speed can hide the conditions that made the result useful. A personal tool may have received customer data, unreleased plans, source code, or a prompt that contains more context than the company realizes. Harmonic found that 8.5% of the prompts in its large 2025 sample contained sensitive corporate data. The figure does not mean every prompt created an incident. It shows why “people are using it successfully” is not a sufficient control.
The right question is not whether the shadow work should be celebrated or condemned. It is what the work reveals. Repeated personal use points to a job that employees value, a gap in the approved stack, or both. Treat it as discovery evidence. Then separate low-risk experimentation from workflows that need a governed home.
Discovery should happen at the level of patterns, not surveillance theater. Ask which tasks recur, what kind of input they require, what the worker checks before using the output, and what failure would cost. A first-draft workflow with public inputs is different from a workflow that summarizes customer records. The company needs that distinction before it decides whether to approve the tool, rebuild the process, or prohibit the use.
The bill for invisible work
The first bill is measurement. If the task lives outside the official account, the company cannot reliably compare the old process with the new one. It may know that employees feel faster, but not whether quality held, whether the time moved to review, or whether the gain belongs to one unusually skilled user.
The second bill is data exposure. IBM's 2026 study found that 70% of CIOs and CTOs said teams deploy technology faster than IT can track. That is a control gap, not a reason to pretend the deployments do not exist. The organization needs a route for people to disclose useful tools without turning discovery into punishment, alongside clear red lines for regulated or confidential data.
The third bill is duplication. Writer's 2026 adoption survey reports that super-users can make up roughly 40% of staff in key functions, with substantial time savings and some building their own tools. If that behavior is not mapped, every team reinvents its own method. The company pays for a platform and for the hidden labor of making the platform useful.
There is a fourth bill: an organization can lose the chance to learn from the people who are already adapting the work. The most capable users are often the first to build a workaround, but they are not automatically the right people to set policy. Their methods can reveal the demand and the missing capability. A separate review still has to test whether the method is safe, explainable, and transferable to someone who did not invent it.

Legalize, then measure
“Legalize” does not mean approve every tool or allow sensitive data anywhere. It means make the useful behavior reportable enough to evaluate. A worker should be able to say, “This is the task, this is the tool, this is the data class, and this is what improved,” without guessing whether disclosure ends their experiment.
Find the task
Ask teams which recurring jobs they already use AI to accelerate. Start with the work, not the vendor list.
Classify the data
Mark what can be used, what needs a managed environment, and what cannot leave a controlled system.
Test the approved path
Give the useful workflow a named owner, an approved tool, a quality check, and a clear stop condition.
Measure the old and new process
Record time, quality, rework, handoffs, and the decision that changed. Do not stop at seats or prompts.
Move the repeatable work
Bring validated tasks into a governed environment, then retire duplicates that add cost without improving the result.
Gravitee's State of AI Agent Security research describes shadow AI as a recurring security pattern rather than a one-time exception. That framing matters. The organization should not wait for a breach to discover that an unofficial workflow became infrastructure. Discovery, classification, and migration are operating work.
A workable policy therefore needs an exception path that is faster than avoidance. It can require a lightweight intake, a data classification, a named owner, and a time-boxed review. It should also state what cannot be entered into an external model and what happens when the output affects a customer, employee, financial forecast, or regulated decision. Clarity is part of control because an ambiguous policy pushes the real decision back onto an individual worker.
A scorecard that survives review
The test for an AI program is not whether it can produce an impressive demo. It is whether a skeptical reviewer can follow the work from task to result. Name the process. Name the owner. Preserve the input boundary. Show the quality check. Compare the result with the old method. Record what happens when the model is wrong.
This is how the 95% figure becomes useful instead of theatrical. It does not tell a leader to stop using AI. It tells them to stop treating formal spend as a proxy for working change. Find the productive paths, keep unsafe data out of them, and make the evidence specific enough to audit.
The counterfactual matters. If the team had kept the old workflow, what would the task have cost in time, review, error, or missed opportunity? If the AI-assisted process looks faster only because review disappeared, the apparent gain is deferred cost. If the result holds after a second person checks it and the task can be repeated within the same data boundary, the evidence is stronger. The scorecard should make that distinction visible.
When a pilot earns a place in the approved stack, keep the original discovery record with it. The reason employees adopted the workaround is part of the product requirement. If the managed version removes the speed, context, or flexibility that made the first version useful, the shadow path will return. Migration is complete only when the governed workflow still solves the job people were trying to solve.
That is also why the owner cannot sit in IT alone. IT can provide the boundary and the controls. The function doing the work has to define quality, exceptions, and the decision that counts as improved. Finance can challenge the baseline. Legal and security can set the red lines. The workflow needs all of those inputs, but it still needs one person accountable for whether it works.
The official program and the shadow program are not two separate futures. They are two present-tense operating conditions inside the same company. The first is visible but may be detached from the task. The second is close to the task but may be detached from governance. A credible AI program joins those facts without erasing either one.
That is the real ROI question: not “Did we buy an AI tool?” but “Which work changed, who can prove it, and what happens when the shortcut becomes important?”
FAQs
What is shadow AI?+−
Shadow AI is work-related use of AI tools outside the systems, accounts, or approvals an organization can reliably see. It can include a personal subscription, an unregistered browser tool, or a small workflow a worker built without an official owner. The term describes the company’s visibility, not whether the work is useful.
Why does formal AI show little ROI?+−
Formal programs often measure licenses, seats, and activity before they define the decision or workflow that should improve. A broad tool can be visible to finance while its intended benefit remains vague. That creates a reporting gap, not proof that every AI use case failed.
Should companies ban personal AI use?+−
A blanket ban can remove useful work without revealing why employees chose a different path. A better first move is to discover the recurring tasks, classify the data involved, test approved alternatives, and set clear boundaries for sensitive information. High-risk uses still need to be stopped or reviewed.
How can a company measure shadow AI safely?+−
Start with voluntary discovery and task-level measurement rather than collecting private prompts indiscriminately. Record the task, time saved, quality check, data class, tool, owner, and repeatability. Then move validated workflows into an approved environment with access controls and an audit trail.
