
Only 2% of AI Agents Actually Work
The problem is not that you have too few agents. It is that almost no one can tell which ones create a business result.
“Deployed” is a purchase order with better branding. It is not a result.
That distinction became hard to ignore in May 2026, when Prosus published field data from more than 60,000 AI agents built by 40,000 employees across its portfolio over eighteen months. The report does not say that 98% are worthless. It says the economic shape of an agent portfolio is brutally uneven: roughly 2% of active agents drive a disproportionate share of business impact.
That is not a weird exception to be engineered away. It is the central fact a serious operating model has to accept. Most productivity agents save small amounts of time. A middle tier earns its keep. A handful cross the line from useful tool to business system. The average agent is unremarkable, which makes averages the wrong lens. Nobody wins an agent strategy on the mean. They win by finding the tail.
There is an important caveat. Prosus is one global portfolio, weighted toward ecommerce and marketplaces, measuring its own systems with its own instrumentation. Treat the exact curve as one company's data, not universal law. Treat its message as transferable: when a company has a way to observe per-agent outcomes, the long tail becomes visible. When it does not, every agent looks equally plausible right up until the budget review.

The 20 boring use cases everyone builds
The most revealing Prosus finding is not the 2%. It is how often organizations converged on the same work without an instruction from headquarters. Across industries, countries, and languages, teams kept building roughly the same twenty use cases. Not campaign-strategy moonshots. Operational infrastructure.
Data analytics and market intelligence account for 18% of the agent tasks Prosus identified. Operations follows at 15%. These are agents that pull performance data, flag anomalies, summarize competitive moves, route recurring work, organize research, and draft the ordinary communications that keep a team moving. They are less exciting in a board deck precisely because they are close to actual work.
Then there is the number leaders should not ignore: 14% of agents sit outside a formal department. They are personal assistants made by employees because no supported tool met the need. That is shadow adoption. It is also a highly specific signal about where the organization has real friction. Treat it only as a security problem and you miss the opportunity. Inventory it, identify the useful repeat patterns, and harden the best ones into tools a team can measure and own.
Analyze: Compare the metrics people already collect but do not have time to connect.
Coordinate: Handle recurring handoffs, routing, summaries, and operational follow-through.
Surface: Find anomalies, research changes, and overlooked risks early enough for a person to act.
Support: Bring shadow workflows into a governed environment once their recurring value is clear.

Deployed is not the same as working
AI agent evidence in 2026 can look contradictory if you line up the headlines without looking at the instrument. MIT's State of AI in Business research found that 95% of enterprise GenAI pilots did not produce measurable P&L impact. Carnegie Mellon's TheAgentCompany benchmark found the best models completed only around 30% of simulated workplace tasks fully autonomously. Gartner predicts more than 40% of agentic AI projects will be canceled by the end of 2027 because of cost, unclear value, or risk controls.
None of that refutes Prosus. The benchmark measures broad autonomy in a simulated office. The pilot research measures value that organizations could not show in their financial statements. Gartner is describing the portfolio consequence of uncontrolled ambition. Prosus observed narrower agents connected to workflows, data, and per-agent outcomes. Same technology. Different scope and discipline.
This is why a demo is a poor procurement test. The question is not whether an agent can complete the happy path. It is whether the business can define a useful outcome, connect reliable inputs, see what the agent actually costs, and spot when its behavior stops earning the trust it was given. That is the measurement break explored in our guide to agentic AI measurement: outcome dashboards do not preserve the decision path.
The two quiet cost traps
The first trap is model overkill. Once a premium model makes a workflow work, teams are understandably reluctant to change it. But many recurring agent tasks do not need frontier reasoning forever. The good-enough frontier keeps moving down-market while the invoice quietly preserves yesterday's choice. If you cannot see cost and outcome together, you cannot know when a model swap is a safe optimization rather than a risky guess.
The second trap is prestige mismatch. Sophisticated, senior-sounding agents tend to get the executive attention. The consistent daily value often comes from the junior plumbing tier: pulling data, preparing context, checking a condition, moving a case to the right place. The switch is not to ban ambitious work. It is to fund the boring work that earns the right to make the ambitious work credible.
Portability matters here. Teams that preserve prompts, evaluations, inputs, and outcomes can move a workflow when price or capability shifts. Everyone else discovers that their real switching cost is the behavior they never recorded. That is why agentic AI switching costs are usually an observability problem before they are an API problem.
Premium model inertia
Keep a high-cost model only where it still produces a measured difference in outcome, not because no one wants to retest the behavior.
The keynote roadmap
A roadmap full of impressive autonomous agents can hide the small operational work that actually creates reliable daily leverage.
Unpriced switching
Assume models will change. Preserve enough evidence to re-evaluate a workflow deliberately instead of migrating by vibes.

How to find your 2%
The playbook is not exotic. Start by assigning every agent a metric at creation: hours saved, revenue touched, tickets resolved, cost avoided, quality improved, or risk reduced. An agent without a metric is not an experiment. It is an anecdote with a monthly invoice.
Build the boring repeatable cases first. They make the data, ownership, and review habit concrete. Then use the shadow agents as a discovery queue, not a hidden estate. Quarterly, review the portfolio as a whole: promote the agents that sustain value, merge the duplicates, and retire work that cannot clear a cost-and-outcome threshold.
Finally, price portability in from the beginning. Assume you will revisit models, tools, and workflows within a year. Store the information you need to prove a replacement works: the objective, inputs, policy, evaluation cases, cost, outcome window, and owner. The business that can do that has not eliminated uncertainty. It has made uncertainty manageable.
Source notes
FAQs
Where does the “only 2% of AI agents work” figure come from?+
It comes from The Coming Age of AI Colleagues, a Prosus report published in May 2026. After analyzing more than 60,000 agents built by 40,000 employees across its portfolio companies over eighteen months, Prosus found that roughly 2% of active agents drove a disproportionate share of business impact. It is a power-law finding, not a claim that every other agent has zero value.
Does that mean the other 98% of AI agents are useless?+
No. Most deliver small but real gains. Prosus reports that 82% of productivity agents save under 20 hours a month, while a middle tier saves between 20 and 173 hours. The lesson is that impact is heavily concentrated, so a business needs per-agent measurement and a portfolio review instead of treating every deployment as equally strategic.
How does this square with studies saying most AI pilots fail?+
The studies measure different things. MIT research focused on enterprise GenAI pilots that did not show measurable P&L impact. Carnegie Mellon measured full autonomy on simulated workplace tasks. Prosus observed scoped, instrumented agents in production. Narrow scope, reliable data, a named owner, and an outcome metric help explain why a small set of production agents can work while many broad pilots do not.
Which AI agent use cases should a marketing team build first?+
Start with repeatable operational work: data analytics and market intelligence, reporting, research, anomaly flagging, workflow coordination, and internal communications. These are less glamorous than autonomous campaign strategy, but they connect to real work, have legible inputs, and build the measurement discipline needed for more consequential automation.
How do I know which of my agents are in the top 2%?+
Give every agent a business metric at creation, including its full model and operational cost. Review actual results quarterly, then promote agents with durable impact, consolidate duplicates, and retire idle or unmeasurable work. The important comparison is not agent versus agent in a demo. It is business result versus full cost in a real workflow.
The agent is not the scarce resource.The discipline to prove what it changed is.
Got a project?