The Next Five Years of AI Will Be Slower Than the Sales Deck Says
By Callum Gracie, Founder, Otto Media
I run a 42-client agency on a skeleton crew and a stack of AI agents. So when someone tells me AI is about to change everything by Tuesday, I have a fairly direct way to check. I look at what my own agents did last week, then I look at what the researchers who actually measure this stuff are finding. Neither one supports the pitch.
Here is where I think the next two to five years genuinely land.
The Capability Curve Is Real. The Reliability Curve Is the Problem.
Start with the good news, because it is genuine. METR tracks how long a task an AI agent can complete, and that number has doubled roughly every seven months since 2019. Recently it moved faster still—by early 2026 it passed sixteen hours.
Then read the fine print. That figure describes a fifty percent success rate. Nobody runs a business on a coin flip, so the number that matters commercially is the eighty percent reliability horizon, and it sits far lower.
Agents also deteriorate once a task gets buried inside a long conversation—researchers have measured web agent success dropping from around forty percent to under ten percent once conversational history is factored in.
Capability keeps climbing. Dependability trails behind it. That gap is the whole story of the next five years.
The Money Arrives Years After the Demo
Economist Daron Acemoglu—no doomer—puts the total productivity gain from AI at no more than 0.66 percent over ten years. Consultancies project figures many times larger. Both cannot be right.
The gap comes down to method. Acemoglu counts measured cost savings on real tasks. The optimists extrapolate from what models could theoretically touch. The field data leans his way: MIT researchers found roughly 95 percent of generative AI pilots produced no measurable return. Gartner expects more than 40 percent of agentic AI projects to be cancelled by the end of 2027. Both blame the same culprit—not the models, but workflow, data, and ownership.
None of that makes the gains fake. It makes them late. Every general-purpose technology behaves this way. Firms have to rebuild their processes before any benefit reaches the accounts.
Where It Lands First—and Where It Stalls
Look at the wins that are already real and a pattern emerges quickly. High volume, verifiable, low liability: clinical note-taking, first-draft legal research, support triage, boilerplate code, marketing production. Those transform meaningfully inside this window.
Then look at what stays stubborn. Diagnosis, legal advocacy, anything a regulator signs off on, anything with a body attached. Legal research tools still hallucinate between 17 and 33 percent of the time under controlled testing. Drug candidates identified by AI sail through phase one and then fail phase two at much the same rate as everything else. Robots still need factories and capital cycles.
The rule of thumb that holds: where a human can check the output in seconds, AI moves fast. Where the check costs more than the task itself, it stalls.
The Shift That Already Hit the Business of Visibility
Pew tracked nearly 70,000 real Google searches and found people click a link 8 percent of the time when an AI summary appears, versus 15 percent when it does not. Around 1 percent click a link inside the summary itself.
That is not a forecast. It already happened—and it quietly guts the model of every agency still selling rankings as though rankings equal traffic.
So we changed what we sell. The job now is getting a client quoted inside the answer, then measuring success by booked jobs rather than sessions. If a client’s traffic drops 20 percent while their phone keeps ringing, we are winning. If leads fall with the traffic, that is a problem worth naming out loud.
What This Means If You Run a Small Shop
Two things happen simultaneously, and most people only notice one of them.
AI lowers the floor. Anyone can produce competent output now, which commoditises undifferentiated work and drags prices down. That part is real, and pretending otherwise helps nobody. Since everyone now buys from the same handful of models, owning the tooling stopped being a competitive edge.
But AI also raises the ceiling. Revenue per employee climbs fast for teams who know exactly where their agents fail. Klarna replaced the equivalent of 700 support agents, announced it proudly, then walked it back and started rehiring—because the averages hid what was breaking in the hard cases.
Research in software development says the same from another angle. Experienced developers in a controlled trial came out 19 percent slower with AI tools while believing they were 20 percent faster. The tool did not simply fail. It failed while feeling like it worked.
The Stance Worth Taking
Assume grinding, uneven change rather than a cliff edge. Watch a short list of signals: agent reliability at high success thresholds, hyperscaler capital spending, and your own clients’ lead volume measured against their traffic.
Build for the tasks nobody argues about. Missed calls. Quotes. Scheduling. Follow-up inside five minutes, because the odds of qualifying a lead fall sharply after that. Boring, checkable, valuable. Those are the jobs where a lean operator beats a slow incumbent every time.
The next five years will hand nobody a shortcut. But they will reward whoever gets closest to the messy edge where the software gives up and a person takes over.
That edge is the business now.
About the Author: Callum Gracie is the Founder of Otto Media, an SEO driven marketing team based in Canberra, dedicated to transforming the digital presence of small to medium-sized businesses looking to scale their business organically.

