Most enterprise AI pilots do not fail technically. They fail to produce a measurable business result, usually because nobody defined one before starting. The projects that reach production are narrow, attached to an existing process, owned by a named person, and measured against a number that existed before the project did.
First, my interest here, declared
I hold a contract role leading partner marketing for the global System Integrators at Amazon Web Services. I am not an AWS employee. Blue Harbour is not an AWS partner or reseller, and neither Blue Harbour nor I earn anything from whichever cloud or AI platform you land on.
What that role gives me is a particular vantage point rather than special access: I spend my week around the commercial reality of enterprise AI, which projects get funded, which get renewed, and which quietly stop being mentioned. Nothing here is written, reviewed or endorsed by AWS, and nothing here draws on anything confidential.
I am telling you this at the top rather than the bottom because an article about AI written by someone paid by a hyperscaler is worth reading differently, and you should be the one deciding how.
The 95% number, and what it measures
The figure everyone repeats comes from MIT’s Project NANDA, in a July 2025 paper called The GenAI Divide: State of AI in Business 2025. What it says, verbatim: “Despite $30 to 40 billion in enterprise investment into GenAI, this report uncovers a surprising result in that 95% of organizations are getting zero return.”
Read that carefully, because the common retelling is wrong in two ways.
First, it is 95% of organizations, not 95% of pilots. Second, and this is the part that matters, it does not say the technology failed to work. The report’s own phrasing is that the majority “remain stuck with no measurable P&L impact”, while “just 5% of integrated AI pilots are extracting millions in value”. A pilot that works and cannot be shown to have moved a number is inside the 95%. So is one nobody measured. So is one that solved a problem which was not costing anything.
Two caveats I would rather give you than have you find. It is labelled preliminary findings, not finished peer-reviewed work, and the copy in circulation is version 0.1. And the sample is modest: interviews with 52 organisations, survey responses from 153 senior leaders, and a review of 300 publicly disclosed initiatives, gathered between January and June 2025. Treat it as a strong signal rather than a settled fact.
The number in the same report I find more useful is the funnel. 60% of organisations evaluated enterprise-grade AI tools, 20% reached a pilot, and 5% reached production. The drop from evaluation to pilot is bigger than the drop from pilot to production, which tells you most of the loss happens before anybody builds anything.
Gartner, separately, forecasts that more than 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. That one is a prediction rather than a measurement and should be read as such.
Sources: MIT NANDA, The GenAI Divide, July 2025, executive summary. Gartner press release, 25 June 2025.
What I see from where I sit
The pattern is consistent enough that I can now usually guess the outcome from how a project is described.
The ones that die are described in terms of the technology. Someone is “exploring what AI could do for us”, or “building an AI capability”, or running a proof of concept to “see what is possible”. There is a committee. There is enthusiasm. There is no number that was going to move.
The ones that ship are described in terms of a job. Something specific and repetitive that a named person currently does, that takes a known amount of time, and that somebody has decided should take less. The technology is almost an afterthought in how they talk about it.
The projects that reach production are usually the least exciting ones in the building.
That is the whole pattern, and it holds at every size of company I have seen it in.
The MIT paper reaches the same place from the data. It reports that the core barrier to scaling is not infrastructure, regulation or talent, and that the buyers who succeed “evaluate tools based on business outcomes rather than software benchmarks”. It also found that external partnerships saw roughly twice the success rate of internal builds, which is worth knowing before you decide to build anything yourself.
Your position is different from theirs, in both directions
Where you have the advantage
A 40-person business has things a bank does not. You can decide in a week rather than two quarters. Your processes are visible: somebody in the building can describe exactly how a quote gets produced, because they produce it. And one person can own the whole thing end to end, which removes the single most common cause of enterprise stall, which is that nobody is accountable for the outcome as opposed to the project.
Where you do not
You have no slack. An enterprise can absorb a failed pilot inside a rounding error and a reorganisation. For you the money is real, the person running it has another job, and a project that consumes six months is six months you did not spend on something else. You cannot afford to be exploratory, which means you should not try to be.
How to read a vendor’s AI pitch
Five questions that separate real from theatre
- Which specific task does this do, and who does it today? If the answer is a category rather than a task, and a department rather than a person, the project has no owner and no baseline. It will end up in the 95%.
- What number is this supposed to move, and what is it today? If nobody can state the current figure, nobody will be able to state the improvement either. Measure before, not after.
- What happens when it gets something wrong? Every one of these systems is wrong some of the time. The question is whether that is caught, by whom, and what it costs when it is not. A vendor without a good answer has not deployed this to a business like yours.
- Show me a customer my size doing this in production. Not a pilot, not a logo slide, not an enterprise reference. In production, at your scale, for at least six months.
- What does this need from our data? The most common reason these projects stall is that the data was never in a state to support them. Find out what state it needs to be in, and who is doing that work, before you sign.
Where it pays at your size
I run a company as well as advising on them, and everything at PhoneStack, operations, research, our data, and the product itself, is built and run with AI day to day. So this is not theoretical for me either, including the parts that did not work.
What pays, reliably, is narrow and unglamorous: work that is high volume, low judgement, and currently done by a person who is too expensive to be doing it. Drafting the first version of something a human then edits. Reading and sorting inbound. Turning one piece of work into the six formats it needs to exist in. Answering the same question for the fortieth time.
What does not pay, at your size, is anything that requires you to become a technology company to run it. If a proposal needs a data platform before it needs a result, it is a project for somebody with a data team.
The answer is sometimes “not yet”, and that is a real answer
Some businesses are not ready, and the honest version of this advice says so. If your quotes still live in one person’s inbox, if nobody agrees what your customer list actually is, or if the last system you bought never got adopted, AI will not fix any of that. It will sit on top of it and make the mess faster.
Fix the process, then automate the process. In that order, every time. It is less interesting than the alternative and it is why the boring projects are the ones that ship.
If somebody has asked you what your AI plan is and you do not have one yet, that is a normal position to be in and it is one of the situations Navigator is built for. The useful output is rarely a list of AI tools. It is usually a short answer about which single thing is worth doing first, and an equally short list of what to ignore this year.
If someone has asked for your AI plan
Blue Harbour Navigator is a fixed-fee independent assessment for owner-led businesses. Seven to ten days, and you get a straight read on where the business is losing time and money, what would fix it, and what to leave alone this year.
I sell no software and I do not implement it, and I take no commissions from any vendor. $5,000 flat, in your own currency.
Book a 20-minute fit call Or tell me what you’re facing, no call needed
Two minutes to write, I read them myself. Or see how Navigator works first.
Sources. MIT NANDA, The GenAI Divide: State of AI in Business 2025, July 2025, for the 95% figure, the $30 to 40 billion investment estimate, the 60/20/5 funnel and the findings on what separates successful buyers. That paper is explicitly labelled preliminary findings and its sample is 52 organisation interviews, 153 survey responses and 300 public initiatives. Gartner’s press release of 25 June 2025 for the agentic-project forecast, which is a projection rather than a measurement. Observations about which projects ship are first-hand and general; no partner, customer or confidential information is described. All figures checked against the primary documents on 11 August 2026.
Disclosure. Trevor Jamieson holds a contract role in partner marketing at Amazon Web Services and is co-founder and CEO of PhoneStack. Blue Harbour Solutions is not an AWS partner or reseller and earns nothing from any cloud or AI platform a client selects. Nothing here is written, reviewed or endorsed by AWS.
Trevor Jamieson is the founder of Blue Harbour Solutions in Halifax, Nova Scotia, and runs independent technology and AI assessments for owner-led businesses.