Skip to main content
A maintenance worker beside the physical controls behind a roadside advertising billboard at dusk.
← All posts

AI Media Buying Needs a Brake

AI media buying can cut campaign costs while hiding weak business results. A guide to spending permissions, better signals, and independent measurement.

AI media buying lets software choose audiences, placements, bids, and sometimes creative within a campaign. The business still has to define a valuable customer and decide when the spending should stop.

AI media buying changes who decides

A campaign can hit its target while spending money your business shouldn't have spent. The conversion happens. The dashboard turns green. Later, someone notices that the customers were already buying, the discounted products barely covered their costs, or the sales team couldn't use the leads. Automation makes that disagreement arrive faster.

On August 6, 2026, Gartner forecast that more than 70% of global advertising spend would flow through AI-influenced self-service platforms by 2028. Its definition includes algorithms affecting delivery, pricing, and outcomes. It's a forecast about platform influence, not a finding that autonomous agents already control that share of spending.

Permission sits between the idea and the spend

An unanswered approval request stays stopped.

Gartner analyst Eric Schmitt puts the incentive problem plainly: “improved platform economics does not necessarily translate into lower costs for the advertiser.” A team can welcome faster execution while still requiring evidence that the financial benefit reaches its own business.

That scope matters. Automated bidding, a system that generates headlines, and an agent that can activate a campaign have different authority. Putting them under one AI label makes it harder to decide which controls each one needs. A useful inventory starts with the action the software can take and the account in which it can take it.

For a marketing leader, the immediate job is to draw that boundary. Can the system recommend a change, save it as a draft, or send it live? Can it move money between products? Can it change the event being counted as success? Those questions reveal more about exposure than the model name or the quality of a demonstration.

Before buying another tool, inventory the access already granted. An agency account, a feed service, and an optimization script may each be able to alter the same campaign. Record the owner of every path. Otherwise a human can correct a setting only to have another system restore yesterday's version during its next scheduled run.

There's a commercial reason to be precise. A platform earns money when advertisers buy media through it. Your company earns money when customers create enough value to cover the cost of reaching and serving them. Those interests can align. They don't become identical because both parties can see the same return-on-ad-spend chart.

Cheaper delivery needs a second test

The strongest case for automation is mundane and useful. Campaign setup takes time. Buying through unnecessary intermediaries costs money. People miss pacing issues while they work on something else. Software that removes those frictions can improve an operation without needing to become its strategist.

PubMatic's Butler/Till case study describes a campaign for Geloso Beverage Group across connected-TV and sports inventory. It reports approximately 80% lower buy-side costs, 40% more impressions than planned, and roughly 30% lower effective CPM, the cost per thousand impressions. The source is Butler/Till internal data from December 2025, published by the platform vendor.

Read the result at the level measured

Buying efficiency

Fees, setup time, effective CPM

Delivery quality

Eligible inventory, reach, completion

Business impact

Additional customers and contribution after costs

Those are meaningful execution results. The case doesn't provide a randomized estimate of additional sales or profit. Treating its impression gain as a sales gain would change what was measured. A team considering a similar workflow should keep the operational case and the demand-generation case in separate columns.

My test would start with the bill. Compare the full buying cost, including platform fees, agency time, setup, monitoring, and work required to fix mistakes. Then inspect inventory quality and delivery against the same definition used for the previous campaign. Changing the comparison midway makes an apparently efficient result hard to interpret.

Ask the supplier to explain the denominator behind every saving. Lower buying fees as a share of media spend are different from a reduction in total campaign cost. A faster setup measured after training may exclude the initial work required to connect systems. Neither distinction invalidates the result, but both affect what you should expect from your own first deployment.

The second test belongs with the business. Did the campaign bring in customers who wouldn't otherwise have purchased? Were their orders profitable after returns and fulfillment? The answer may take longer than the campaign itself. Until then, lower delivery cost is a valid finding with a limited meaning, and a reason to keep testing.

Your conversion event writes the brief

A machine follows the reward it receives more consistently than the intention someone described in a kickoff call. If the event is a form submission, the system can search for people likely to submit forms. Whether those people can afford the service, live in the delivery area, or represent a real buying opportunity is a separate question.

Consider a hypothetical home-services advertiser. One campaign brings forty inquiries and another brings twenty. If the first group contains duplicate requests and jobs outside the service area, optimizing toward its cheaper lead price teaches the wrong lesson. The useful event arrives later, when a request is accepted and becomes a completed, profitable job.

A retailer checks a returned jacket while fulfilled parcels leave the warehouse.
The sale has to survive returns and fulfillment.

Start with a definition the people doing the work will accept. For a lead, write down which conditions make it qualified and who confirms them. For an order, decide how cancellations and returns change its value. Give the measurement owner a way to correct records, because mistakes in the feedback will otherwise become instructions for the next campaign.

Google's AI Max documentation illustrates how much the execution can move. Its Search features include expanded search-term matching, text customization, and final URL expansion. The documentation describes separate controls for generated text and destination selection, and warns that pinned assets aren't used when a more relevant expanded URL is chosen. Review those settings individually.

My practical concern is the destination promise. A generated ad can be plausible while the selected page is obsolete, unsuitable for the offer, or impossible to track correctly. Test the actual route a customer takes, including mobile checkout or the lead form. An approved homepage doesn't establish that every page on the domain is fit for paid traffic.

Separate the arrival of an outcome from the date it's reported back to the platform. If accepted leads arrive several days after inquiries, a temporary gap can look like deterioration. The team needs a rule for waiting through that delay. Otherwise the optimizer can be interrupted just as the evidence needed to judge it becomes available.

A useful measurement model for AI-driven marketing keeps credited activity separate from commercial value. You don't need a perfect model before testing automation. You do need a visible account of what the system sees, what it misses, and how long the missing evidence takes to arrive.

Put the stop control beside the spend

A written policy can't stop a campaign if the software that activates it never checks that policy. The control has to sit where a proposed action becomes a live change. Otherwise the review process begins after the consequences, and a very detailed log becomes a record of spending you couldn't prevent.

In its August 5 governance announcement, PubMatic describes buyer-configured limits, pre-approved libraries, authenticated human approvals, and action logs in AgenticOS. It says missing required fields stop execution and decisions outside configured authority escalate. Rise, a Quad agency, is named as an early pilot. These are vendor-described capabilities to verify in a trial.

A hand turns a key in the physical isolator beneath a billboard.
A spending rule matters when execution can enforce it.

The acceptance test should be deliberately awkward. Ask the system to use an excluded destination, exceed a small trial limit, or select an unapproved asset. Use a safe test environment or a draft-only path. A reassuring explanation isn't a pass. Confirm that the forbidden change doesn't reach the live campaign and that the record identifies what blocked it.

Decide what happens when nobody answers an approval request. For a routine optimization, waiting may be harmless. For an expiring offer, the right outcome may be a pause. Write that choice before launch. An absent approver shouldn't silently become permission, and a backup approver should know the commercial rule they're being asked to apply.

Make changes expire when the reason expires. A clearance discount, a regional exception, or an emergency budget increase should have an owner and a review date. Permanent settings tend to outlive the conversation that justified them. Requiring a fresh decision prevents an old exception from becoming the default strategy for a different commercial situation.

Keep the ability to stop spending outside the agent's own access. The person responsible for the account should be able to revoke its permission or pause delivery directly. Record how long that action takes and which campaigns it affects. A stop button that only asks the agent to stop is another dependency on the system you're trying to control.

Attribution can't settle the argument

Attribution assigns credit for an outcome. Incrementality asks how much of that outcome the advertising caused. The distinction gets expensive when the same platform selects likely buyers, serves the ad, and reports the resulting purchase. A strong association can be commercially useful while still overstating how much the ad changed.

The warning predates generative AI. Tom Blake, Chris Nosko, and Steven Tadelis studied paid search at eBay using large field experiments. Their 2014 working paper found conventional estimates substantially overstated returns in that setting, with different effects across customer groups. It's an older study of one marketplace, not a current benchmark for every advertiser. Its experimental lesson still applies.

Give the campaign a fair comparison

Compare the same outcome and time window. Attribution alone cannot supply the missing comparison.

Reserve a comparison before the next optimization cycle. Where feasible, randomly hold back advertising from an eligible group and compare an agreed business outcome. Geography-based tests can sometimes help when individual randomization isn't available, but markets differ and campaigns can spill across boundaries. Have the design reviewed before treating a simple before-and-after chart as causal evidence.

Make the comparison fair to both sides. If the automated campaign receives fresher creative or a larger discount, the test mixes media buying with offer changes. If it also receives better product data, record that contribution. You may still choose the combined package. You just won't know which improvement earned the result.

Small budgets create a real constraint. A short trial may be too noisy to establish a profit difference. Report that uncertainty and cap the exposure instead of declaring victory after the first few orders. Delivery quality, qualified-lead rates, and customer complaints can still identify problems while the larger business outcome develops.

Don't discard the platform report just because it isn't causal. Use it to locate where delivery changed and which segment deserves investigation. Reconcile its totals against the business record with a shared time window. Differences can come from reporting delays, canceled purchases, or attribution rules, and understanding those differences is part of supervising the investment.

Watch the incentives around the report as well. A buying team judged only by platform return may resist evidence that reduces that number. Give someone outside daily campaign optimization responsibility for the comparison. Their task is to make the investment decision more accurate, including when the result supports spending more.

Give the buyer a smaller, harder job

The buyer's role becomes more demanding when routine execution gets cheaper. Someone has to choose which uncertainty is worth paying to resolve, which customers matter, and which restrictions are commercial commitments. Those decisions don't vanish when campaign setup becomes a conversation.

Begin with one workflow whose mistakes you can contain. Keep its current baseline, define the actions it can take, and choose a limited trial budget. Record the conversion definition and the expected delay before revenue quality becomes visible. Avoid changing the offer, measurement, and autonomy level in the same experiment unless you're explicitly testing the whole package.

A trial that can earn more authority

Define

Name the outcome, exclusions, owner, and spending limit.

Test

Prove the controls and compare with a credible baseline.

Expand

Increase authority only when the evidence supports it.

During the trial, review exceptions as a working list. A new product category, an unusual destination, or a sudden shift toward existing customers deserves an explanation. Check the actual settings and outcomes behind that explanation. Fluent generated reasoning can help investigation, but it shouldn't substitute for an observable change record.

After the trial, make the decision in plain language. Continue when the gain survives the full cost calculation and the controls work. Extend the test when the result is uncertain but the exposure remains acceptable. Stop when the workflow repeatedly crosses a boundary or when the additional operating burden consumes the delivery savings.

Keep what you learned somewhere the next buyer can use. Save the failed assumptions, the exceptions that mattered, and the setting changes that fixed them. That record becomes especially useful when staff, agencies, or platform defaults change. The same team shouldn't have to rediscover why an apparently efficient audience was excluded.

Put the cost of human review into the trial from the start. If every change needs someone to spend an hour reconstructing its context, the workload has moved rather than disappeared. Simplify the agent's remit until its actions can be inspected quickly. More restricted autonomy that the team understands can be more valuable than a broader system nobody can comfortably supervise.

Automation is worth using when it buys the team room to make these decisions well. If nobody has time to challenge the objective or reconcile the result, a faster buying process can deepen the original problem. The person approving next month's budget should be able to explain what was earned, what remains uncertain, and what would make the spending stop.

FAQs

What is AI media buying?

It's software choosing or adjusting advertising decisions such as bids, audiences, placements, and creative. Some tools only recommend changes. Others can execute them, so the permission scope matters.

Does cheaper CPM mean better advertising?

It means the reported cost per thousand impressions is lower. It doesn't establish additional sales, better customers, or higher profit. Check delivery quality and business outcomes separately.

Which controls should a small team start with?

Use a limited trial budget, approved destinations and assets, a clear conversion definition, and a named owner who can pause spending directly. Test what happens when a proposed change violates a rule.

How can we test impact with limited data?

Preserve a comparison where feasible, keep the offer stable, and allow time for customer quality to emerge. When the sample is too small for a reliable effect estimate, report uncertainty and limit exposure.

A woman looks toward a city billboard from a pedestrian bridge at dusk.

Keep the authority to stop spending.