AI agent abandonment needs better evidence
The agent worked beautifully in the demonstration. Someone selected a clean brief, connected the right accounts, and stayed nearby when it hesitated. That proves a useful capability under those conditions. It leaves a more expensive question unanswered: who will get acceptable work out of the system next month, with an ordinary brief and nobody coaching the run?
There isn't a defensible universal statistic showing that most AI agents disappear after 30 days. Gartner's June 2025 forecast predicted that over 40% of agentic AI projects would be canceled by the end of 2027. Analyst Anushree Verma described many as “early stage experiments or proof of concepts.” That's a prediction about projects, with a later deadline. It doesn't measure first-month retention.
The distinction matters because a dramatic failure rate can send a team toward the wrong repair. If the work stopped arriving, rewriting prompts won't restore demand. If employees still need the result but have returned to a spreadsheet, there's a workflow to investigate. If a pilot exposed unacceptable access requirements, ending it may be exactly the decision the business needed.
Adoption evidence also cuts against a blanket collapse story. In McKinsey's August 2026 survey, 40% of respondents at large organizations reported scaling agents, compared with 22% at smaller organizations. These are survey responses about organizational adoption, not longitudinal survival rates for individual agents. Both expansion and disappointing pilots can be happening at once.
I would judge a deployment by the work that survives the handoff to its regular operator. Can that person initiate the task, recognize a valid result, handle an interruption, and explain when the agent should stay idle? A product that needs the original builder for each of those steps has a support dependency that the demo may have hidden.
A month is a convenient review interval. It isn't a natural law of software adoption. A weekly campaign report offers several repeat opportunities in that period. A quarterly planning agent may offer none. Set the review window around the job's frequency before interpreting silence as rejection.
Count the next eligible job
A login is a poor proxy for useful agent work. Someone can open the product to inspect a failed run, or leave it closed while a scheduled job completes correctly. Count the underlying task opportunities first. That gives the team a denominator it can explain without relying on whatever the vendor happens to call an active user.
For a campaign reporting agent, an eligible job might be one scheduled report whose accounts, date range, and required data meet the pilot's entry conditions. Record whether the agent was chosen, whether its output was accepted, and how much intervention followed. Keep those decisions attached to the same job. Otherwise, a single report retried five times can look like five successful uses.
Follow all ten eligible jobs
One more is accepted after repair; one needs manual completion. Four never use the agent.
Consider an illustrative week with ten eligible briefs. The team sends six through the agent, accepts four without substantial correction, repairs one, and completes one manually after the run fails. The other four never enter the tool. That record exposes different questions about selection, quality, and recovery. It isn't a benchmark, and the numbers shouldn't be turned into a claim about other companies.
Follow the same cohort of operators and the same task definition through the review. Adding enthusiastic new users can conceal existing users leaving. Narrowing eligibility halfway through can improve the acceptance rate by removing difficult work. Sometimes either change is justified. Record it as a change in the experiment so the comparison remains honest.
Ask about skipped opportunities without making adoption a performance target for employees. A marketer may avoid the agent because it needs information the brief doesn't contain, because a client has unusual requirements, or because reviewing its answer takes longer than doing the job. Those explanations are product evidence. Pressure to improve usage numbers makes them harder to collect.
Retain a small sample of accepted results, too. A checked box alone won't reveal that standards drifted as the deadline approached. The sample lets a reviewer compare what counted as good work in the first week with what the team is accepting later. Keep sensitive campaign data within its existing access rules.
Check what actually left the system
The agent's last message can say the task is complete while the destination still contains a draft, an incorrect field, or nothing at all. Acceptance belongs where the work is used. For a marketing report, that means the right period, reconciled figures, usable commentary, and delivery to the intended place. Fluent explanation is only one part of that result.
Anthropic's guide to evaluating agents distinguishes an agent's recorded actions from the final state those actions produce. Its practical lesson for evaluation is to check the outcome in the environment. For a campaign workflow, the equivalent is inspecting the saved report or account state instead of accepting the agent's own completion claim.

Measure human effort across the whole attempt. Include preparing inputs, reviewing the result, correcting it, and recovering from failure. Keep elapsed waiting time separate from active work. A ten-minute run can save attention even when it takes longer on the clock, provided the operator can safely leave it alone. A fast answer that needs continuous supervision offers a different bargain.
Even rigorous measurements age. In its February 2026 update, METR explained why selection effects made its newer estimates of AI's effect on developer speed unreliable. Participants' willingness to work without AI had changed. That limits comparisons with its earlier results. The lesson for a marketing pilot is to inspect who participates and which tasks reach the test before carrying an old productivity estimate forward.
Klarna offers a useful counterexample to the assumption that every assistant is abandoned. Its 2025 annual report says its assistant handled 80% of customer-service chats during that year. That's a company-reported measure of a defined service workflow, not independent proof of quality or evidence about autonomous marketing agents. It does show why sustained use should be examined at the level of an actual operation.
Your own acceptance test can be much smaller. Have someone who receives the output review a sample without being told which method produced it, where that is feasible. Preserve the same criteria for manual and agent-assisted work. If the agent earns its place by making routine drafts cheaper while a person retains final judgment, measure that arrangement honestly.
Make interruptions recoverable
A recurring workflow has to survive an incomplete brief. Imagine a reporting agent that encounters a missing campaign identifier after collecting several valid inputs. Starting again from the beginning wastes the completed work. Guessing the identifier risks reporting on the wrong account. A useful interruption preserves progress and asks the responsible person for the specific missing decision.
The handoff should show what the agent was trying to do, which input it couldn't resolve, what it has already saved, and what will happen after the person responds. The reviewer needs a clear choice. A raw transcript of every tool call is useful for investigation, but it can bury the immediate task under details the operator doesn't need.
Resume from a known checkpoint
The reporting step resumes from the checkpoint. No account change is authorized.
The animation follows one illustrative job through a missing-input hold and a verified restart. The saved work stays attached to the same job. The human supplies the missing value, the system checks that value against the permitted scope, and execution resumes from the held step. The point is recoverability. Adding a chat window alone doesn't provide it.
Decide how long a held job remains valid. A report can become stale while waiting for approval. A campaign may close, or an account permission may change. Before resuming, recheck the conditions that matter to that particular action. Preserve the old context for explanation while using current authority for execution. The advertising agent governance article covers that boundary in more detail.
Also define the fallback when the operator doesn't respond. The appropriate result may be an explicitly incomplete draft, a reassigned exception, or a closed attempt with a reason. Automatic retries won't resolve a missing business decision. They can increase cost and produce a misleading impression that the system is making progress.
Test interruption handling with the person who will receive the request. Give them an ordinary workday and a realistic amount of context. If they need to phone the builder to understand the screen, improve the handoff before increasing volume. The time spent clarifying one confusing exception will recur every time the workflow encounters that condition.
Fit the work people already have
An agent can improve the middle of a task while making its beginning and end worse. A marketer exports a file, cleans it, uploads it to a separate tool, explains the brief again, then copies the answer into the team's existing system. The model may perform well. The surrounding work can still make the whole arrangement unattractive.
Map the actual entry and exit points before adding another capability. Where does the approved brief live? Who checks it? Which system receives the final output? Put the agent where it can receive the necessary context without asking someone to recreate it. Limit access to what the task requires, and keep a visible record of the version used.

Don't automate a process nobody can describe. If three team members disagree about what constitutes an approved brief, the agent will inherit that disagreement. A short working agreement may produce more value than another model upgrade. It can identify the required fields, the decision owner, and the conditions that send a job back for clarification.
Some tasks will be better served by a template or a simple rule. A predictable calculation doesn't need an agent to decide how to perform it each time. Reserve open-ended reasoning for parts of the job that benefit from it. This also makes faults easier to isolate because the team can distinguish a bad input, a deterministic processing error, and a poor judgment.
Keep the familiar route available during the trial and record when people use it. Removing the alternative can manufacture retention while concealing dissatisfaction. Once the agent has demonstrated enough quality and recoverability, the team can deliberately retire redundant steps. That decision should follow the operating evidence.
Budget for an ordinary Tuesday
The launch team often supplies labor that disappears from the pilot's cost estimate. Someone repairs a connector at night, curates a better input, or explains a confusing result in a direct message. That help makes the trial possible. It also creates an optimistic picture of what the agent can do without support.
Assign ongoing responsibility before the launch team leaves. One person owns the business result, someone operates the queue, and someone maintains the technical connection. In a small company, those may be the same person. The responsibilities still need names and an alternative when that person is unavailable.
A handoff someone can operate
Task owner
Defines acceptable work and decides which exceptions matter.
Operator
Checks the queue, resolves interruptions, and records recovery effort.
Maintainer
Repairs connections and rechecks behavior after a change.
Record maintenance time beside usage. A vendor update, expired permission, new source field, or revised campaign policy can change an established workflow. Keep a short set of representative jobs that can be replayed after a material change. The examples should include at least one exception the team already understands, so a successful routine run doesn't hide a broken recovery route.
Support should have a budget and a stopping condition. If one narrow task requires constant specialist attention, decide whether its value justifies that arrangement. Some high-value work will. Ordinary reporting may not. The decision becomes clearer when the team can see accepted output, active review time, and maintenance effort together.
Avoid treating every departure as resistance to change. A regular operator may have discovered a problem the launch team never encountered. Ask them to show the last job that made them choose another route. The saved brief and the resulting work usually provide a better starting point than a general satisfaction score.
Earn the second month
Give the pilot a written decision date and a limited task population. Before it starts, name the acceptance criteria, the manual comparison, the permitted access, and the person who will assess the result. Choose enough repeat opportunities to observe ordinary work. A date on the calendar alone won't create a useful sample.
At the review, separate evidence of useful output from evidence of continued interest. Enthusiasm can help a team learn a tool. The renewal decision needs to establish whether people can keep obtaining acceptable results within a support burden the business can afford. Record unresolved uncertainty instead of filling it with an adoption forecast.
A sensible next step may be to continue the narrow job unchanged. It may be to fix one recurring interruption and test again. It may be to retire the agent and retain the improved brief or evaluation method. A stopped experiment can leave the operation better understood, even when the software doesn't remain part of it.
Expansion deserves its own decision. Adding a new account, another task type, or a consequential action changes what has been demonstrated. Keep the successful use case available while testing the additional responsibility separately. That protects the value already earned and makes the new evidence easier to interpret.
Before paying for a broader rollout, ask the next operator to run a representative job while the builder stays out of the conversation. Watch where the operator hesitates, what they check, and how they recover. Those few minutes will tell you what still needs to be built.

