Gartner says more than 40% of agentic AI projects will be cancelled by 2027. Most of them never had anyone checking whether they actually worked.
Written for anyone about to scale an agentic pilot past its own demo. The cancellation number is real and Gartner-sourced. The specific mechanism behind it is in Gartner's own research too — it just isn't the part getting repeated.
Gartner predicts that more than 40% of agentic AI projects will be cancelled by the end of 2027, and it names the reasons plainly: escalating costs, unclear business value, and inadequate risk controls. The prediction is built on a January 2025 poll of 3,412 people who attended a Gartner webinar: 19% said their organisation had made a significant investment in agentic AI, 42% a conservative one, 8% none at all, and the remaining 31% were still deciding or unsure. Gartner analyst Anushree Verma's own read on that spread is blunt — most agentic AI projects right now are early-stage experiments or proofs of concept, driven by hype and often misapplied. That's the line every outlet leads with. The more useful half of Gartner's own research is the part that explains why "misapplied" keeps happening even to teams that aren't chasing hype.
The gap Gartner actually named
Gartner's 2026 CIO and Technology Executive Survey and its accompanying Hype Cycle for Agentic AI put a number on that second half: only 17% of organisations have actually deployed an AI agent, against more than 60% who expect to within two years — the steepest adoption curve of anything the survey tracks. The space between those two numbers is what Gartner's own analysts describe as a capability-deployment gap: a pilot works cleanly in a controlled test, then stalls on the way to production because nobody checked its behaviour against real, live data before it was handed a wider scope, a bigger budget, or a customer-facing job. Cost overruns and unclear ROI are what that stall looks like once finance notices. The unverified gap between demo and production is what caused it in the first place.
What that looks like once real money is moving
Two examples from this year show the mechanism without ever using the word "agent." Uber ran an internal leaderboard rewarding teams for heavy use of AI coding tools, and its roughly 5,000 engineers burned through the company's entire 2026 AI budget in four months, mostly on Claude Code and Cursor, at $500 to $2,000 per engineer per month. Uber's CTO, Praveen Neppalli Naga, put the company "back to the drawing board" on AI budgeting; it now caps spend per tool at $1,500 a month and measures token cost directly against the cost of having a person do the same task, while Uber's own COO has publicly questioned whether the spend was worth it. Microsoft did the reverse: it began revoking employees' internal Claude Code licenses in May, after usage costs climbed well past what had been budgeted, and moved affected teams to GitHub's Copilot CLI — publicly framed as "toolchain unification," reported elsewhere as a cost decision. Neither company was running an agentic AI pilot in Gartner's sense. Both scaled a tool by incentive or mandate before anyone had priced what it was actually returning, which is the same missing step, one budget cycle later.
The arithmetic that should run first, not after
That's the argument I make about training spend generally, and it holds here without adjustment: ROI is arithmetic, not a multiplier. Measure one specific workflow, before and after, on its own terms — not a vendor benchmark multiplied across headcount, and not a company-wide token bill read as a verdict on the whole programme. Uber's own fix, once the bill forced the question, was exactly this: compare what a task costs to run against a model against what it costs to have a person do it, per workflow. That comparison was available in month one. It ran in month five, after the budget was already spent.
A pilot with no owner was never on a path to production
The other half of the fix is structural, not financial. Every workflow worth deploying should end with a named owner, a review gate, and a date it gets used again — not because it's tidy, but because that's the artifact that catches a confidently wrong answer before it reaches a customer. A pilot that scales from demo to production with no one specifically responsible for checking its output against real cases was never actually on a path to production. It was a demo that got a bigger budget.
Not every cancellation is the same failure
The 40% figure also flattens two very different diagnoses into one number, and it's worth sorting them before assuming a cancelled project was simply broken. Some genuinely couldn't be verified — the agent's behaviour on live data never got checked, and once someone looked, it wasn't ready. That one is fixable: run the workflow-level test before the wider rollout. Others get cancelled for close to the opposite reason: a team correctly decided the tool shouldn't be making a particular judgement call at all, and the cancellation gets filed as a technical failure rather than said out loud as a boundary. I've made a version of this case about AI adoption decisions generally — "couldn't verify it worked" and "could, and we were right not to let it" are different diagnoses, and treating them as one number hides which lesson a company actually needs to learn before its next pilot.
Before scaling any pilot past its own demo, run the check Gartner's own data says most companies skipped: what does this look like against a real case, who signed off on that specific check, and what would the wrong answer actually have cost. If the honest answer is "we don't know," that project isn't most of the way to production. It's a well-funded proof of concept, and the 2027 cancellation list is where those get counted.
This is a position, not a finding of my own — built on named, dated third-party research linked throughout. Uber and Microsoft's coding-tool spending decisions aren't agentic AI project cancellations in Gartner's technical sense; I've linked them because they show the same unmeasured-value pattern playing out with real budgets, one step removed from the survey data.
If you're about to scale a pilot past its demo, the methodology page walks through how I run the before/after measurement on one workflow before anyone touches the budget for the next ten.