I read AI post-mortems for a living.
Everyone has a favorite failure number by now, and I have read the studies behind all of them. MIT found that 95% of enterprise generative AI pilots produce no measurable P&L impact. S&P Global measured 42% of companies abandoning most of their AI initiatives last year, up from 17% the year before. RAND puts AI project failure above 80%, twice the rate of ordinary IT projects.
Nobody disputes the failure anymore. It is consensus. The fight is over the diagnosis.
The diagnosis is missing a word
Read the post-mortems and the blame lands somewhere different every time. MIT blames a learning gap. The integrators blame data readiness. Gartner blames governance. The CFO blames the bill. All of it is real, and all of it is a symptom.
Look at what a production AI deployment is actually made of. Not a model. A stack: serving engines, gateways, vector stores, orchestration frameworks, guardrail services, observability platforms, cost dashboards, agent builders. By Bud's reference-architecture count, standing up full production capability takes 40-plus tools across seven layers. The outside view agrees on the shape of the problem: the average enterprise runs about 23 AI tools and only 38% can inventory them; across the wider estate, organizations average 897 applications, 2% of them integrated.
Every one of those tools is excellent. That is the trap. Nobody chose the boundaries between them. The stack grew one good decision at a time, and the cost of the boundaries grew with it. That cost has no name, so it has no line item, so no one owns it, so no one repeals it.
Naming things is my job. So we named it.
The AI Fragmentation Tax
The Fragmentation Tax is the compounding cost an enterprise pays at every boundary between the tools in its AI stack — in latency, accuracy, tokens, and auditability.
Itemized, from Bud's engineering analysis of production deployments:
- Latency. Each boundary adds 2–10 ms of serialization, authentication and queueing. Across one agent action that is 100–1,200 ms before any AI computation happens.
- Accuracy. A five-step workflow at 95% per-step reliability delivers 77% end to end. One completion in four carries a silent error, and the error is nobody's fault, which is why it never gets fixed.
- Tokens. Context re-serialized at every handoff means 40–60% of token spend is overhead, not answers.
- Oversizing. A fragmented stack cannot route by difficulty, so every request goes to a frontier model by default: 10–50× more model than the query needs.
None of this shows up in a demo. All of it shows up in production. That is why the demo works and the deployment doesn't, and why the same models produce millions for one enterprise and nothing for the next.
Proof the tax is real: the bill
Here is the strangest fact in enterprise technology right now. The price of intelligence collapsed, and the cost of using it went up. Stanford's AI Index shows inference prices for GPT-3.5-class performance fell 280× in two years. In the same period, roughly seven in ten enterprises ran over their AI budgets: 79% in DoiT's survey of finance leaders, 73% in the FinOps Foundation's State of FinOps, 68% in WitnessAI's polling.
Cheaper tokens made AI more expensive. Cost stopped being a procurement decision and became an architecture decision.
Agentic volume multiplies calls faster than prices fall, and every one of those calls crosses the boundaries above. Which is why a cheaper frontier API never fixes the CFO's problem.
What I will not claim
I am a marketer, so let me be precise about the claims I am not making.
- I will not argue about the bubble. Whether or not this is a bubble, wasted inference spend is real either way, and it is the one problem you can fix this quarter.
- I will not tell you your tools are bad. They are not. The architecture that connects them is the problem, and nobody designed it. It piled up one tool at a time.
The one claim I will make
The tax is optional.
The Bud Novaria AI OS is one native stack from silicon to agents, on your infrastructure. Training, inference, routing, guardrails, governance, agents and consumption under one control plane, on GPUs, CPUs, HPUs, NPUs or TPUs, in any cloud, on-prem or air-gapped. The boundaries come out, and the tax levied at each one goes with them. No boundaries, no tax.
A global fashion brand moved its customer-facing AI styling agent from a frontier-model stack to a domain-tuned small model on the Bud Novaria AI OS, served CPU-native on Intel Xeon. Monthly AI cost fell from $218,000 to $40,000. Accuracy held at 85.4%.
The saving did not come from a cheaper API. It came from hybrid AI: the right model on the right silicon, with the frontier held in reserve for the fraction of work that needs it.
Hybrid isn't a hedge. It's the architecture.
Not slideware. Production.
If you want to know what your stack is taxing you, we will measure it: a 30-day proof of concept on your infrastructure, with your data, against a metric you choose. A number, not a pitch.
The market is not missing intelligence. It is missing an operating system.
The guide Repeal the Tax. The Five Moves to AI That Scales The tax itemized line by line, and the five architecture moves that remove it →Failure rates: MIT Project NANDA, The GenAI Divide: State of AI in Business 2025 (95%, a preliminary report that has not been peer-reviewed); S&P Global Market Intelligence, 2025 enterprise AI survey (42%, up from 17% in 2024); RAND Corporation, The Root Causes of Failure for Artificial Intelligence Projects (above 80%). Tool counts: Bud reference architecture (40-plus tools, seven layers); external estate figures are industry survey ranges cited for illustration. Inference price decline: Stanford HAI, AI Index 2025 (280× for GPT-3.5-level performance, Nov 2022 to Oct 2024). Budget overruns: DoiT (79%), FinOps Foundation State of FinOps (73%), WitnessAI (68%). Latency, accuracy, token and oversizing figures are Bud engineering analysis of production deployments; the accuracy figure is arithmetic (0.955 ≈ 0.77). The fashion-brand figures are a customer deployment on Intel Xeon, as reported in the Bud Novaria launch release; the figure that heads this post ran in that release.
Kevin Johnson