Skip to main content
Theater technician checking power distribution and cables beneath a stage during setup
← All posts

AI Costs Don't End at Tokens

AI marketing costs include retries, review, shared systems, and unused capacity. Build a budget around accepted work and inspect the contract behind the price.

AI marketing costs are the resources required to produce and maintain usable work: model calls, software, data services, human review, and support. A low token price explains one input. A useful budget connects the complete workflow to an accepted result, then separates its economics from the supplier's financing story.

Provider debt isn't your invoice

A marketing director reading about data-center spending can reasonably wonder who will eventually pay for it. The answer won't come from a provider's debt headline alone. Your team buys a particular service under particular terms. The immediate budget risk is the combination of those terms, the work you send through the service, and the effort required to make its output usable.

Oracle's June 2026 results make the infrastructure scale concrete. It reported negative free cash flow of $23.7 billion for fiscal 2026 and $43 billion of debt financing raised that year. The same release reported $75 billion in prepaid and customer-supplied hardware portions of large AI contracts. These figures describe Oracle's investment and funding arrangements. They don't establish a coming price increase for your marketing software.

Maintenance engineer observing a large electrical substation at blue hour
Illustrative infrastructure photography. The economics of a supplier and a customer's contract are different questions.

There are several steps between financing a data center and changing a customer's bill. A supplier may change a price, alter an allowance, charge for an additional capability, or keep existing terms while finding savings elsewhere. Which possibility matters to your team depends on its agreement. A forecast about a provider's capital needs cannot settle that question by itself.

Read the actual renewal and usage conditions. Identify the commitment period, included capacity, overage treatment, notice provisions, and any services billed separately. Ask the account owner to obtain an answer where the wording is unclear. A list price on a public website may not describe an enterprise agreement, and a sales promise should be reconciled with the written terms.

Keep supplier continuity on the agenda without turning it into speculation about collapse. The operational questions are whether you can retrieve your inputs and outputs, preserve necessary records, and complete important work during an interruption. A workable alternative can be a manual process for a narrow task. It doesn't always require paying a second vendor for an identical stack.

The budget decision should start with what the team can observe and control. That includes the contract, the task volume, the quality requirement, and the cost of supervision. Broader infrastructure economics provide context. They don't replace those records.

AI marketing costs cross budget lines

A content workflow can generate charges in several places while appearing to have one owner. Marketing pays for the application. A shared technology account pays for model access or storage. An analyst spends time fixing the result. Procurement sees a subscription renewal, while nobody sees the combined cost of the work that subscription enables.

This is an accounting boundary the team needs to choose explicitly. Start with the costs that change when the workflow runs: model usage, paid searches, external tools, and human handling. Then show fixed commitments and shared services separately. The distinction makes it possible to answer both what another job will cost and what keeping the operation available costs over the month.

The FinOps Foundation's overview of AI cost management places AI across software, platforms, and infrastructure, with different charging models and ownership. Its relevance to marketing is organizational: the person requesting more output may not be the person who receives every associated bill. A useful review brings those records together around the use case.

Build a modest inventory before purchasing another management tool. Identify the application, the paying account, the responsible team, the usage measure, and the renewal date. Include tools purchased by individual employees where the company permits them. The goal is a complete view of the authorized workflow. Don't treat a missing line item as evidence that a service costs nothing.

Allocation needs a rule people can inspect. If several teams share a service, direct usage attribution is preferable when the data supports it. When it doesn't, use an agreed allocation basis and label the estimate. Charging every shared cost to the newest AI project can make it look worse than the old workflow. Leaving all shared costs out can make it look artificially cheap.

Keep cost estimates separate from the cash budget. An employee's review time has an economic cost even if payroll doesn't change this month. A prepaid annual license creates a cash commitment even when current usage is low. Finance may need both views, and combining them without explanation can produce an impressive but unusable return calculation.

Put accepted work under the total

The most useful unit is the result the business accepts. For one team, that might be a reviewed campaign brief. For another, it's a reconciled report delivered on time. Define the unit tightly enough that two people can agree whether the work qualifies. A generated draft and an approved draft represent different amounts of completed work.

J.R. Storment's explanation of token economics connects metered AI consumption to business value, describing it as “connected to business outcomes.” Tokens are small units of model input and output. They help explain consumption, but the marketing buyer still needs to know what that consumption produced. A large output count can include unusable work.

Divide by accepted work

Review adds $11. Total cost is $20; eight accepted drafts cost $2.50 each.

Invented teaching figures. Include the cost of all ten attempts, then divide by the eight accepted results.

Here is an illustrative batch, using invented costs to show the method. Ten drafts incur $4 in generation, $2 in retrieval and tools, $3 in retries, and $11 in review. The total is $20. Eight drafts meet the acceptance criteria, so the cost is $2.50 per accepted draft. Dividing by the ten initial drafts would report $2 and conceal the two failures.

Keep the numerator and denominator aligned. If the review cost covers the whole batch, the accepted-output count must cover that same batch and period. If a rejected draft later becomes acceptable after more work, add the recovery cost before changing the count. Reassigning the failure to another team doesn't remove the cost from the company's workflow.

The FinOps guidance on AI tools and services recommends evaluating total cost per use case outcome. It also explains why that cost can change as the configuration changes. The operational implication is to keep a stable definition of accepted work while comparing methods. Otherwise, a cheaper method can appear to improve economics simply by delivering less.

Avoid pretending one blended average answers every decision. Show high-volume routine work separately from rare specialist work when their costs and standards differ. A complex campaign brief may justify more review than a recurring report. Mixing the two can hide a deteriorating routine process or make a valuable specialist task look wasteful.

A cheaper call can cost more

A lower model price is useful only if the resulting workflow still meets its requirements. When a cheaper route produces more retries or longer review, the saving can disappear before the work is accepted. That doesn't mean expensive models are inherently better purchases. It means the comparison has to include what happens after the first response.

The comparison below uses two deliberately simple, illustrative routes. Route A costs $1 to generate a draft and $1 to review it. Route B costs $0.40 for the first attempt, another $0.40 after rejection, and $1.50 for review. Both eventually produce one accepted draft. Their totals are $2 and $2.30. These are teaching figures, not provider prices or measured performance claims.

Follow the rejected draft

With review included, Route A totals $2 and Route B totals $2.30 for one accepted draft each.

Invented costs, not provider prices. Compare the entire workflow at the same quality standard.

The extra attempt is the visible state change. The first draft on Route B fails the same quality check, returns for another attempt, and reaches acceptance only after additional handling. Watching that loop makes the cost mechanism easier to inspect. A production comparison should use your own observed attempts and review time, including jobs that never reach acceptance.

Agent architecture can also change consumption. Anthropic reported that its multi-agent research system used about 15 times as many tokens as chat interactions in its experience. That's evidence from a particular research workload, not a universal multiplier for agents or a direct 15-times cost estimate. Model mix, input prices, and output prices still matter.

Set a quality floor before testing cheaper routes. For a report, require correct figures and an accurate explanation. For publishable copy, include claim support and the receiving editor's acceptance. Then compare the complete cost, waiting time, and intervention burden on representative work. A small fast model may win many routine tasks. Let the evidence identify which ones.

Keep the difficult cases in view. An average can look attractive while the costliest failures land on an already overloaded specialist. Record that tail of the workload separately. If a routing change makes normal jobs cheaper but creates a difficult exception queue, decide who will staff it and whether the combined arrangement still pays.

Savings need somewhere to go

A team can lower cost per accepted result and still spend more overall. If the workflow becomes easier to use, people may run it more often or ask it to do more. That can be a good outcome when the additional work has value. It becomes a problem when the budget assumes savings will automatically appear as lower total expenditure.

Write the intended use of the gain into the plan. The business might keep output constant and reduce external spend. It might produce more approved work with the same team. It might reserve the freed capacity for tasks that previously went unfinished. Those choices lead to different measures of success and different expectations for the next budget cycle.

Ceramicist inspecting a finished bowl with accepted pieces and imperfect pieces separated on the bench
The denominator is usable work. Rejected output still consumed resources.

Time saved is especially easy to overstate. Ten minutes removed from each of several scattered tasks may make a workday less fragmented. It doesn't necessarily create a removable staff position or an equivalent amount of cash. Track whether the time becomes available in a form the team can use, then identify the work it enables.

Unused commitments complicate the picture. A seat remains paid for when its user leaves, and a minimum contract commitment can outlast the project that justified it. Review these costs with the same attention given to variable usage. Canceling an unnecessary workflow may produce no immediate cash saving if the agreement continues, although it can prevent new variable charges.

McKinsey's August 2026 survey found that about 20% of respondents reported operating costs constraining AI usage. That is a reported adoption constraint, not evidence that every marketing budget is being cut. It supports asking whether the next unit of usage earns its place in the plan.

Don't use generated volume as the destination for every saving. More drafts can simply move the bottleneck into editing. Budget for the amount of accepted work the organization can distribute, use, and maintain. An inexpensive backlog is still a backlog.

Set limits around the actual job

A monthly spending alert arrives too late to explain which job is going wrong. Pair the period budget with limits on individual workflows. Define the allowed number of retries, which paid tools can be called, and what happens when the job reaches its limit. The operator should receive the incomplete result and a useful reason, not a silent failure.

An alert and a hard stop serve different purposes. An alert asks someone to investigate. A hard stop prevents additional spending through the control that enforces it. Check what your particular product actually supports, how quickly usage is reported, and whether queued work can continue after a limit is reached. Don't promise a real-time cap based on a delayed billing screen.

Three costs need different controls

Per job

Limit retries and paid tool calls, with a visible fallback.

Per period

Track demand and total spend against the approved workload.

At renewal

Review commitments, unused seats, overages, and exit terms.

Choose a fallback that matches the task's consequences. A nonurgent research draft can wait. A scheduled report may need a partial result with missing fields identified. A consequential account change may need to remain unexecuted until a person reviews it. Cost controls should make the operational state clear so the team can decide what to do next.

Keep an explanation for spending changes. More eligible jobs, a longer source document, a model change, and a higher retry rate are different causes. They call for different responses. Without this distinction, a manager may restrict a useful workflow while leaving a broken configuration untouched.

Review the assumptions after material changes. A new model, a revised acceptance requirement, or a different data source can alter both cost and quality. Preserve a small comparison set and the earlier configuration so the team can investigate. A budget that was reasonable for last month's workflow may describe a system that no longer exists.

Take a complete case to renewal

The next renewal conversation should begin with a specific workload. Show how many eligible jobs arrived, how many produced accepted results, the full operating cost, and the work still requiring manual completion. Include the cost of keeping the system available. A vendor usage chart alone can't explain whether the workflow deserves a larger commitment.

Compare against a credible alternative. That may be the current manual method, a simpler automation, or a different service. Use the same quality requirement and task population. If one option has a setup cost, show it separately and state the period over which the decision expects to recover it. Avoid burying an optimistic payback assumption inside a monthly average.

Include uncertainty in a form someone can act on. A low-demand scenario tests whether a fixed commitment becomes wasteful. A high-demand scenario tests review capacity and variable spend. A failure scenario tests the cost of recovery. Use a small number of explicit assumptions that owners can change, rather than a precise-looking forecast built on unverified adoption promises.

The agent retention analysis asks whether ordinary operators keep obtaining useful work after launch. That operating record belongs beside the cost calculation. A cheap deployment that nobody chooses for eligible jobs has a different problem from an expensive workflow that delivers a valuable result consistently. The budget discussion should distinguish them.

Give the renewal owner a concrete recommendation with its boundary. Continue this task at this expected volume, fix this source of rework before expansion, or end this commitment when the agreement allows. The recommendation can be modest. It should be specific enough that the next review can tell whether the decision was right.

Before agreeing to a bigger allowance, bring one recent accepted job and its complete cost record into the meeting. If the team can't reconstruct that example, the next purchase should include better measurement.

FAQs

What belongs in an AI marketing cost estimate?

Include model usage, subscriptions, paid tools, data services, review, recovery, and maintenance. Show variable costs, fixed commitments, and allocated shared costs separately so the estimate can answer both marginal and total-cost questions.

Does provider debt mean AI prices will rise?

Provider financing alone does not establish a customer's future price. Inspect the agreement's renewal provisions, allowances, overage treatment, and notice terms. Use infrastructure financial results as context, not as a substitute for contract evidence.

How do you calculate cost per accepted output?

Divide the complete cost of a defined batch by the number of outputs that meet its acceptance criteria. Include failed attempts and recovery costs in the same period. Define what counts as accepted before comparing workflows.

Is a cheaper model always cheaper to operate?

No. Additional retries, longer review, or more failures can outweigh a lower first-call price. Test the complete workflow against a stable quality requirement and use your own observed costs.

Do AI time savings automatically reduce the budget?

No. Freed time may improve throughput or reduce fragmentation without changing payroll or committed software charges. Identify how the business will use the capacity and distinguish economic value from cash savings.

Tailor measuring navy fabric before making a planned cut

Price the work you can actually use.