95% of enterprise AI pilots fail. Inside a bank, the reason is worse than you think.
Everyone blames the models. The more I dug into why enterprise AI keeps failing in regulated industries, the more I think the real reason is something almost nobody is talking about.
There is a number I keep coming back to.
In its 2025 study, The GenAI Divide, MIT looked at 300 public enterprise AI deployments and found that 95% of them delivered no measurable business impact. Only one in twenty produced real value.
When I first read that, my instinct was the same as everyone’s: the models must not have been good enough. But the study is clear that this is wrong. The cause was not model quality. It was implementation. Generic tools bolted onto the side of real work, dazzling in a demo and useless in production a month later.
Lately, I have been reading everything I can find about one specific kind of enterprise, the regulated kind (also because I earn my bread through this): banks, fintechs, and payment processors. And the more I read, the more I think the 95% number hides a second, harder problem that almost nobody is talking about out loud.
Generic AI is built to always answer
Think about how a general-purpose model behaves. It is built to always produce something. Ask it a question it has never seen, and it will not stop and say, “I do not know this one.” It fills the gap with something that sounds right.
In most of life, that is a fine trade. A slightly wrong answer is a minor annoyance, and you move on.
Now put that same behaviour inside a bank.
A confident wrong answer about a regulatory deadline is not a typo. It is a fine. A misread threshold on a suspicious-transaction alert is not an inconvenience. It is a report that should have been filed and was not. A document that looks like it passes a check but quietly does not is not a rounding error. It is real money and real liability.
The work these institutions run on has a particular character that makes guessing intolerable:
The rules are published: Scheme rulebooks, central-bank timelines, verification requirements, and compliance thresholds. There is a correct answer, and it is written down somewhere.
The work is language-heavy and repetitive: Read a document, check it against a rule, decide, write the response, record why. Thousands of times over.
Every case has money and liability attached, which is exactly why nobody has been willing to let a generic chatbot anywhere near it.
This is high-stakes but low-judgment work. It follows the published rules almost every single time. And it is still done by hand, at enormous cost, because getting it wrong is expensive, and a general model cannot be trusted not to invent its way through the parts it does not know.
So the question changes
Here is the shift that I believe most of the industry has yet to make.
Everyone is asking, “Can AI do this work?” And technically, the answer is mostly yes. The models are capable of reading the document and drafting the response.
But that is not the question the person who actually has to approve it is asking. The compliance officer, the risk lead, the person whose name goes on the outcome, is asking something else entirely:
“Can AI do this in a way I am willing to sign?”
Those are not the same question. And I have come to think the gap between them is the whole story. It is why 95% of pilots die. It is why the impressive demo never becomes the thing that runs on Monday morning. The demo answers the first question. Nobody in the building can answer the second.
The uncomfortable part
If the problem were just “the models need to get smarter,” it would solve itself. Every lab on earth is working on exactly that, and the models do get better every few months.
But I do not think that is the problem. A smarter model that still guesses is still a model that guesses. Making it more fluent does not make a risk team trust it with a filing. If anything, a more confident wrong answer is more dangerous, because it is harder to catch.
Which leaves a genuinely interesting question, and it is the one I have been sitting with:
If the fix is not a smarter model, then what is it? What would AI have to actually be, structurally, before the person carrying the liability would put their name on its output? What has to be true about how it works, not how clever it is, before “interesting demo” becomes “we can run this”?
I have some strong opinions forming on this, from following how this work actually happens and where the published rules leave off. There is a shape to the answer, and it has very little to do with the things the AI headlines are about.
I am going to write about it properly in the next few pieces: what “trustworthy” has to mean when a wrong answer is a fine, why the deployment question (”where does our data even go?”) quietly kills more projects than the technology ever does, and why the same boring, rule-bound work shows up in department after department wearing a different uniform each time.
If that is a thread you want to follow, stick around. Follow along here, and I will take it apart one piece at a time.
The models are not the hard part. The hard part is everything the demos skip.
Some of my other articles:
💱 Dynamic Currency Conversion (DCC) - What & How 🤔
When you travel abroad and swipe your card, sometimes the POS asks if you’d like to pay in your home currency instead of the local one. That’s what’s called Dynamic Currency Conversion (DCC).
💡 How B2B Payments and Supply Chain Finance Are Quietly Powering Global Trade
I was reading Sam Boboev‘s article on B2B payment methods - it’s a simple but insightful breakdown of how businesses move money between each other.
🚧 Why Onboarding Friction Kills Merchant Growth
Every payment company talks about “merchant acquisition.”






