Skip to main content
A finance leader reviews a paper ledger and receipts at a rain-wet ferry terminal at blue hour.

Your AI ROI Is Invisible

The bill is precise. The business value often is not.

By Dellon S.12 min read

AI is not an investment case because it has a dashboard. It is an investment case only when a team can connect the full cost to a changed business outcome, and explain how it knows.

78%cannot tie spend to outcomes
22%can do it today
66%of boards want proof

The most dangerous AI number in a budget deck is the one everyone can see.

In May, Uber COO Andrew Macdonald described the problem with unusual candor: it was difficult to draw a line between rising AI costs and useful customer features. Fortune's reporting confirms the quote. The company has not disclosed a total spend, but the granular numbers explain the pressure: power users reportedly cost $500 to $2,000 a month each, including a reported $1,200 two-hour session, before Uber imposed a $1,500 monthly employee cap.

That correction matters. A dramatic, unsupported aggregate number can make a better story; it cannot make a better decision. The real story is more useful: finance could see a growing, auditable bill in granular detail, while the company still struggled to connect that bill to outcomes a customer or a board would recognize. That is an accounting gap, not a model-quality debate.

The paradox is simple. Token counts, model choice, timestamps, and usage caps arrive as a utility meter. The value of those inferences, less churn, more resolved tickets, higher conversion, faster approved work, or a stronger product, does not arrive in the same ledger. Precision on one side can create a false sense that the other side has also been measured.

A visual comparison of the metered AI bill and the proof required to establish business value.
A token invoice is evidence of consumption. It is not yet evidence of a return.

The number that changes the debate

The argument does not need an invented invoice total. It needs a number that describes the operating problem directly. A June 2026 CloudZero survey of 260 senior finance professionals found that 78% could not fully tie AI spending to business outcomes. Only 22% could do it today, even though 87% said they would need the capability within a year.

The detail is what turns a generic anxiety into a budget reality. Sixty-six percent of boards condition continued AI funding on proof of return. Forty-three percent of finance leaders are already being asked for a number they cannot produce. Three quarters of teams that cannot measure outcomes have held back further investment, while 35% have killed an initiative rather than keep funding an undefendable one.

As Josipa Majic put it, token billing made spend visible without making it legible. That distinction belongs in every AI review. Visible means the company can add a bill. Legible means a decision-maker can understand what the spend changed, compared with what would have happened without it.

Two people inspect a paper map, receipts, and a measuring tape outside a stone civic building.
An audit is a search for the comparison that makes a claim testable, not a tour of a dashboard.

If the baseline is missing, the story fills the gap.

The useful question is not whether a team can name activity. It is whether it can reconstruct the business decision, the old state, the full cost, and the result with enough care for someone else to challenge it.

Why attribution breaks as soon as AI reaches the customer

The cost side is straightforward. Current Anthropic pricing is published per token. A finance team can total a monthly bill exactly, flag a usage spike, and compare model tiers. That is necessary operating discipline. It just does not answer the customer question.

Imagine a support model that improves response quality. Did it reduce churn? The answer may depend on seasonality, the mix of ticket complexity, a product release, a pricing change, competitor activity, staffing, and which customers had the problem in the first place. Imagine a content model that produces five times the drafts. Did those drafts produce incremental qualified demand, or did they simply increase review work and publish more pages into a competitive search market? A high-quality output is not a causal chain.

Marketing and support teams often inherit the hardest part of the calculation after engineering has made a system work. Engineering can demonstrate a lower error rate or faster response. Those are real operational measures. But a CMO or board is entitled to a second question: did that improvement alter a business metric after the full implementation, review, and risk cost was included? Without a pre-defined baseline, the answer becomes a retrospective narrative built around whatever data is easiest to find.

This is why “ROI per token” is the wrong benchmark. One useful marketing intervention can be expensive in tokens and valuable in revenue. One cheap automated answer can create a costly mistake. The connection has to run through the actual workflow and the outcome it was designed to change.

When the cap arrives, it is usually too late.

Uber is not the only company that encountered cost discipline after deployment. The repeating pattern is that a usage cap becomes the first measurement system, even though it only manages spend, not value.

01

Uber

A per-employee cap followed unusually visible AI usage. The lesson is not the headline amount; it is that an auditable bill appeared before a defensible benefit model.

02

Microsoft

Reporting described the cancellation of many Claude Code licenses across a large engineering division after costs exceeded expectations. Cost control arrived after the workflow was already widespread.

Reporting
03

Walmart

Walmart capped use of its internal coding tool after employee demand outpaced the plan. A later control is sensible; a pre-set outcome and threshold would be stronger.

Reporting

These cases do not prove that AI spending is waste. They show that sophisticated organizations can reach a visible cost limit before they reach a shared definition of success. The better sequence is to set the decision rule first: what would make the initiative worth another 90 days, what would trigger instrumentation, and what would stop it?

The lock-in shadow is bigger than the contract.

Once value remains unmeasured for long enough, the question quietly changes. It stops being “is this model helping?” and becomes “can we afford to unwind the workflow?” That shift makes the provider look safer not because it has proven superior, but because the company has accumulated a cost of leaving it cannot clearly price.

The Register's reporting on a Zapier survey of 542 US executives describes the ingredients: tuned prompts, integrations, undocumented adaptations, and institutional memory. Only 42% of organizations that attempted an AI platform migration called it smooth; the rest reported failures or unexpected effort. That does not make switching impossible. It means the switching work needs to be counted before the company has no alternative.

This is lock-in by opacity. If no one has defined the business benefit of the current model, there is no honest comparison between a cheaper alternative, a smaller use case, or a different workflow. The organization is forced to compare the certain cost of changing with an imagined benefit of staying. That is a vendor advantage created by missing measurement.

An operations worker checks cable and route cards at a coastal railway platform.
The real migration bill is carried in prompts, integrations, and undocumented handoffs, not just in a contract.
A four-step visual showing how visible spend, vague value, emergency caps, and workflow dependency reinforce each other.
The work is to interrupt this loop before an emergency cap becomes the only governance mechanism.

The content generation trap is a measurement trap.

Consider an illustrative composite rather than a claim about a named company. A content team spends $50,000 to $150,000 a month on AI-assisted drafting. It publishes 100 posts instead of 20 and celebrates a traffic increase against human-written work from two years earlier. The team did produce more. It may even have produced better work. But the comparison does not control for algorithm changes, topic selection, link-building, competitive shifts, seasonality, or the extra editing effort required to keep the quality bar intact.

Three months later, traffic levels off. The team asks for better prompts. Finance sees the bill grow. Nobody can answer whether a different content mix, a smaller model, a human-first workflow, or a controlled test would have produced the same outcome. The organization has made a real operational change but has not earned the right to call it a business return.

The answer is not to demand laboratory perfection from every campaign. It is to decide what evidence is enough for the next action. A new workflow can start as an experiment, but it should not quietly become a permanent budget line just because nobody wrote down the threshold for saying no.

The scorecard before scale

A practical AI ROI review does not need a new executive dashboard. It needs the same five fields for every material deployment. The point is not to make the work look more controlled. The point is to make the next decision easier: scale a credible result, improve an ambiguous test, redesign a workflow, or stop spending.

A five-field scorecard for full cost, target outcome, baseline, 90-day result, and confidence.
Use the same scorecard for a support assistant, a content workflow, a recommendation model, or a coding tool. Consistency lets finance compare decisions instead of anecdotes.
01

01

Separate cost from value

Keep inference, integration, review, and vendor-management cost in one ledger. Track the business outcome in another.

02

02

Define the baseline first

Record the old state before the tool changes workflow, including the quality threshold that matters.

03

03

Run a 90-day comparison

Use a controlled split, a comparable cohort, or an explicit before-and-after window with known limitations.

04

04

Name the owner

One person should be accountable for the business question, not merely the platform adoption number.

05

05

Forecast the slope

Alert on volume and cost growth before a budget is exhausted, not after a surprise invoice.

06

06

Make a decision

At the review, choose to scale, instrument, redesign, consolidate, or stop. “Keep watching” is not a decision.

A business owner folds a handwritten checklist at a misty train station before sunrise.

A bill is not proof.

An invoice records what AI consumed. Proof records what it changed.

Questions readers ask

How much did Uber actually spend on AI?+

Uber has not disclosed a total dollar figure. Reporting confirmed per-engineer usage in the $500 to $2,000 monthly range for power users, including one reported $1,200 two-hour session, before a $1,500 monthly cap per employee. A specific $30 million aggregate figure is not supported by the reporting used for this article.

What percentage of companies can prove AI ROI?+

CloudZero's June 2026 survey of 260 senior finance professionals found that 22% could fully tie AI spending to business outcomes, while 78% could not. That is a useful signal, not a universal benchmark, but it captures why a visible AI bill is not enough to defend an investment.

Why does AI create vendor lock-in without a restrictive contract?+

The dependency accumulates in tuned prompts, integrations, workflow adaptations, data access, and institutional memory. If the company never defined the benefit it needs from the model, it cannot compare the value of staying with the real cost and risk of changing providers.

What is the fastest way to improve AI ROI measurement?+

For every material deployment, record the full cost, target business outcome, baseline, 90-day result, and confidence in the comparison. That turns an AI bill into a reviewable decision: scale it, instrument it, change it, or stop it.