For most of its history, software has carried a quiet flaw: the more people use it, the more expensive it becomes to run. More users bring more servers, more support tickets, more technical debt. Growth creates friction rather than leverage. Generative AI has inherited the flaw and sharpened it — every additional user is another stream of tokens billed at the same rate as the first.
The learning gap
A widely discussed study from MIT's Project NANDA, The GenAI Divide: State of AI in Business 2025, found that the vast majority of enterprise generative AI pilots were not producing a measurable return. What the researchers pointed to was not the models, the budget, or the talent. It was something they called the learning gap — a failure mode we have written about before, in seven reasons pilots stall between the sandbox and production.
Most enterprise AI systems simply do not learn from the way they are used. A model that answers its ten-thousandth question is usually the same model that answered its first. It repeats the same mistakes. It needs the same context supplied every time. And every correction a user makes quietly disappears the moment the conversation ends.
A model that answers its ten-thousandth question is usually the same model that answered its first.
That is the gap Bud Novaria AI Operating System (AIOS) is designed to close, and at the heart of it is a capability we call the Enterprise AI Flywheel. Instead of treating each interaction as a transaction that ends when the answer is delivered, Bud Novaria treats every interaction as a deposit — something that leaves your AI a little better than it was before.
One loop, running on every interaction
The flywheel is not a batch job or a quarterly retraining exercise. It is a loop that runs on every single interaction. The system answers the request, then captures what happened — the full trace, and whether the outcome was actually good. From there it extracts whatever is worth reusing, and deposits it back into the platform, so that the next request starts from a better place.
Four pools, improved at once
What makes the loop powerful is that it improves your AI in four different ways at the same time, from the same captured signal.
It improves the models themselves, because successful answers from a large model can be used to train a smaller, faster one that handles the same work at a fraction of the cost. It improves the infrastructure, since the serving layer learns the shape of real traffic and reuses what it has already computed, so responses arrive sooner. It sharpens the decisions around the model — the prompts, routing, memory and guardrails that every resolved case and every incident makes a little stronger. And it builds a growing library of reusable skills and tools, which means that by the time a company builds its seventh AI agent, much of the work is already done.
| Pool | What gets deposited | The metric that moves |
|---|---|---|
| 01 · Model capital | Distilled weights. Teacher traces become small-model training data; preference optimisation follows. | Marginal cost per request ↓ |
| 02 · Infrastructure capital | Serving substrate. KV-cache entries, radix trees, prefix caches, pod routing shaped by real traffic. | Latency ↓ |
| 03 · Policy capital | Prompts, routes, memory and tool patterns — the decision surfaces learned from trajectories. | Answer quality ↑ |
| 04 · Primitive capital | Skills and MCP tools, guardrails, and a registry of reusables the next agent inherits. | Time-to-agent ↓ |
Table 1 — the four pools. One captured interaction can deposit into more than one.
It already works — at the companies that built it themselves
This is not a thought experiment. Shopify published one of the clearest accounts of a production flywheel in August 2026.
Shopify's Sidekick assistant includes an agent that answers merchants' questions about their stores, serving up to two thousand requests every minute. Each day, Shopify takes the conversations that went wrong, repairs them, and feeds them back into training — so the model steadily learns from real merchant needs rather than from a static benchmark.
The result is a smaller specialised model that now outperforms the frontier model it started from. Shopify estimates it costs about a twenty-fifth as much to serve, runs noticeably faster, and needs fewer GPUs for the same traffic.
In a separate experiment highlighted by Shopify's CEO, a very small fine-tuned model — 0.8 billion parameters — beat a leading frontier model on a narrow task.
Daily output grew by more than thirty times in the process, from roughly 2 million to 72 million buyer profiles per day. Narrow task, small model, enormous throughput: that combination only becomes available once the loop is producing its own training data.
How fast the loop can turn
Cursor, the AI coding tool, shows the other end of the range. Its code-suggestion model handles hundreds of millions of requests a day, and every time a developer accepts or rejects a suggestion, that signal is used to retrain the model. Cursor now ships new versions several times a day, with each full cycle taking only an hour or two.
The interesting part is what the model learned. It did not simply get better at producing suggestions — it learned to make fewer of them, while getting noticeably more of them accepted. That is precisely the kind of improvement that only comes from learning in production, where the cost of an unwanted suggestion is real.
Why almost nobody runs one
There is a reason only a handful of companies operate loops like these today. Behind Shopify's results sits a long chain of specialised engineering — from defining what good looks like and calibrating the systems that judge quality, through repairing failures and retraining models, to optimising how those models are served.
On top of the engineering, the flywheel only turns if it is run with discipline. Nearly every interaction has to be captured, not a sampled few. The feedback has to be trustworthy rather than biased. The benchmarks need to keep getting harder, so that progress is real rather than an artefact of a stale test set. And the system has to keep seeing enough variety that it does not narrow itself into failure.
A loop that learns from the wrong signal does not stand still — it gets worse faster. Bias in the feedback compounds exactly as efficiently as quality does.
Four disciplines keep it honest: every interaction deposits · signals stay honest · benchmarks stay fresh · diversity is preserved. For most enterprises, building and operating all of this in-house would be a major undertaking in its own right — an AI platform team standing up before the first agent ships.
A 30-minute walkthrough of the flywheel against your hardware mix, governance constraints, and top use case.
Built in, not assembled
This is exactly why the Enterprise AI Flywheel is a native capability of Bud Novaria rather than something you assemble yourself. When your organisation runs its AI on Bud Novaria, the loop is already turning in the background — capturing outcomes, maintaining the quality checks, and feeding what it learns back into your models, infrastructure, policies and reusable tools.
The practical consequence is a staffing one. Standing up the pipeline in Figure 7 is a platform programme: quality engineering, evaluation infrastructure, a training stack, and a serving layer that can absorb a new model several times a week without an incident. Running it on Bud Novaria makes that an operating characteristic of the platform rather than a team you hire.
Why the gap widens
The difference between the two approaches is not a fixed percentage. It compounds. An organisation running static AI pays the same price for the same mistakes month after month. An organisation running on Bud Novaria gets cheaper, faster and more accurate with every interaction — and the improvements deposit into pools that the next agent inherits for free.
With the Enterprise AI Flywheel built into Bud Novaria, everyday usage turns into lasting value across your models, infrastructure, decisions and reusable tools. The result is enterprise AI that does not just run in production, but keeps getting better because it is there — without your teams having to engineer that improvement themselves.
- MIT Project NANDA attributes the failure of most generative AI pilots to a learning gap, not to models, budget or talent.
- One loop — serve, capture, extract, deposit — improves models, infrastructure, policy and reusable primitives at the same time.
- Shopify and Cursor prove the loop works in production; the seven-stage pipeline behind it is why almost nobody else runs one.
The 95% figure is from MIT Project NANDA, The GenAI Divide: State of AI in Business 2025 — a preliminary report that has not been peer-reviewed, and should be read as directional rather than definitive. Sidekick figures (serving cost, latency, GPU count, request rate) are Shopify Engineering's own published estimates from "Sidekick's continual learning loop" (5 Aug 2026); the ~$27M → ~$1M annual figures are estimates, not billed amounts. The 0.8B buyer-profile experiment is as reported by AlphaSignal from Shopify's CEO. Cursor figures are from "Improving Cursor Tab with online RL" (Sep 2025). Figures 1 and 9 are illustrative shapes, not measured data — substitute your own cost-per-resolved-request series for a firm number.