Self-improving AI: the enterprise AI flywheel

Most enterprise AI answers its ten-thousandth question exactly as badly as it answered its first. The flywheel turns every interaction into a deposit — compounding across models, infrastructure, policy and reusable tools.

SERVE CAPTURE EXTRACT DEPOSIT The Loop EVERY INTERACTION

For most of its history, software has carried a quiet flaw: the more people use it, the more expensive it becomes to run. More users bring more servers, more support tickets, more technical debt. Growth creates friction rather than leverage. Generative AI has inherited the flaw and sharpened it — every additional user is another stream of tokens billed at the same rate as the first.

The shape most enterprise software has
Cost keeps climbing Usage over time → Cost to serve
Figure 1 — illustrative. Scale adds servers, tickets and debt; nothing in the system converts usage back into efficiency.

The learning gap

A widely discussed study from MIT's Project NANDA, The GenAI Divide: State of AI in Business 2025, found that the vast majority of enterprise generative AI pilots were not producing a measurable return. What the researchers pointed to was not the models, the budget, or the talent. It was something they called the learning gap — a failure mode we have written about before, in seven reasons pilots stall between the sandbox and production.

95%
of enterprise generative AI pilots showed no measurable P&L impact (MIT Project NANDA, 2025)
~25×
cheaper to serve after one company closed the loop on a production assistant (Shopify, 2026)

Most enterprise AI systems simply do not learn from the way they are used. A model that answers its ten-thousandth question is usually the same model that answered its first. It repeats the same mistakes. It needs the same context supplied every time. And every correction a user makes quietly disappears the moment the conversation ends.

A model that answers its ten-thousandth question is usually the same model that answered its first.

That is the gap Bud Novaria AI Operating System (AIOS) is designed to close, and at the heart of it is a capability we call the Enterprise AI Flywheel. Instead of treating each interaction as a transaction that ends when the answer is delivered, Bud Novaria treats every interaction as a deposit — something that leaves your AI a little better than it was before.

One loop, running on every interaction

The flywheel is not a batch job or a quarterly retraining exercise. It is a loop that runs on every single interaction. The system answers the request, then captures what happened — the full trace, and whether the outcome was actually good. From there it extracts whatever is worth reusing, and deposits it back into the platform, so that the next request starts from a better place.

The four stages, in order
01
Serve
Respond to the query, under the same SLO as any production request.
02
Capture
Record the full trace, the outcome, and a quality label for it.
03
Extract
Distil the reusable essence — what generalises beyond this one case.
04
Deposit
Write it into one of four pools, where the next request can draw on it.
Figure 2 — the loop closes inside the platform. Nothing here depends on a user filing a ticket or an engineer noticing a pattern.

Four pools, improved at once

What makes the loop powerful is that it improves your AI in four different ways at the same time, from the same captured signal.

One turn of the loop, one deposit into every pool
01 MODEL marginal cost ↓ 02 INFRASTRUCTURE latency ↓ 04 PRIMITIVES time-to-agent ↓ 03 POLICY answer quality ↑ 01 SERVE 02 CAPTURE 03 EXTRACT 04 DEPOSIT
Figure 3 — one interaction, four deposits. The loop runs continuously; each turn leaves all four pools better than it found them.

It improves the models themselves, because successful answers from a large model can be used to train a smaller, faster one that handles the same work at a fraction of the cost. It improves the infrastructure, since the serving layer learns the shape of real traffic and reuses what it has already computed, so responses arrive sooner. It sharpens the decisions around the model — the prompts, routing, memory and guardrails that every resolved case and every incident makes a little stronger. And it builds a growing library of reusable skills and tools, which means that by the time a company builds its seventh AI agent, much of the work is already done.

PoolWhat gets depositedThe metric that moves
01 · Model capital Distilled weights. Teacher traces become small-model training data; preference optimisation follows. Marginal cost per request ↓
02 · Infrastructure capital Serving substrate. KV-cache entries, radix trees, prefix caches, pod routing shaped by real traffic. Latency ↓
03 · Policy capital Prompts, routes, memory and tool patterns — the decision surfaces learned from trajectories. Answer quality ↑
04 · Primitive capital Skills and MCP tools, guardrails, and a registry of reusables the next agent inherits. Time-to-agent ↓

Table 1 — the four pools. One captured interaction can deposit into more than one.

It already works — at the companies that built it themselves

This is not a thought experiment. Shopify published one of the clearest accounts of a production flywheel in August 2026.

Shopify's Sidekick assistant includes an agent that answers merchants' questions about their stores, serving up to two thousand requests every minute. Each day, Shopify takes the conversations that went wrong, repairs them, and feeds them back into training — so the model steadily learns from real merchant needs rather than from a static benchmark.

The result is a smaller specialised model that now outperforms the frontier model it started from. Shopify estimates it costs about a twenty-fifth as much to serve, runs noticeably faster, and needs fewer GPUs for the same traffic.

Sidekick — estimated annual cost to serve, before and after
Frontier model ~$27M / yr Distilled SLM ~$1M / yr Same workload — up to 2,000 requests per minute. Shopify's own estimate.
Figure 4 — roughly a 96% reduction in serving cost. The bars are to scale; the small one is genuinely that small.
What else moved, same deployment
Serving cost
−96%
Latency
−38%
GPUs needed
−14%
Figure 5 — reductions against the frontier model the system started from, at equal traffic.
A second data point

In a separate experiment highlighted by Shopify's CEO, a very small fine-tuned model — 0.8 billion parameters — beat a leading frontier model on a narrow task.

Daily output grew by more than thirty times in the process, from roughly 2 million to 72 million buyer profiles per day. Narrow task, small model, enormous throughput: that combination only becomes available once the loop is producing its own training data.

Our read on the Sidekick numbers, and why a loop like this belongs in the platform rather than in every team — 5 min. Nothing is requested from YouTube until you press play; open it in a new tab instead if you prefer.

How fast the loop can turn

Cursor, the AI coding tool, shows the other end of the range. Its code-suggestion model handles hundreds of millions of requests a day, and every time a developer accepts or rejects a suggestion, that signal is used to retrain the model. Cursor now ships new versions several times a day, with each full cycle taking only an hour or two.

The interesting part is what the model learned. It did not simply get better at producing suggestions — it learned to make fewer of them, while getting noticeably more of them accepted. That is precisely the kind of improvement that only comes from learning in production, where the cost of an unwanted suggestion is real.

Cursor Tab — after online reinforcement learning
21% fewer suggestions made less interruption per keystroke 28% more of them accepted the ones that survive are the useful ones 400M+ requests per day · a full retrain cycle in one to two hours
Figure 6 — fewer and better at the same time. Offline benchmarks do not reward this trade; production acceptance does.

Why almost nobody runs one

There is a reason only a handful of companies operate loops like these today. Behind Shopify's results sits a long chain of specialised engineering — from defining what good looks like and calibrating the systems that judge quality, through repairing failures and retraining models, to optimising how those models are served.

The pipeline behind one working flywheel
01
Quality rubric
A written definition of a good outcome.
02
Calibrated judges
Automated scorers checked against humans.
03
Harness tuning
Automated prompt and scaffold optimisation.
04
Failure repair
Broken conversations fixed into training data.
05
Fine-tuning
Distillation into the smaller served model.
06
Reinforcement learning
Preference signal from real outcomes.
07
Prompt compression
Token cost taken back out of the loop.
Figure 7 — seven specialised stages, each needing its own ownership, tooling and on-call. This is the part that does not appear in the headline numbers.

On top of the engineering, the flywheel only turns if it is run with discipline. Nearly every interaction has to be captured, not a sampled few. The feedback has to be trustworthy rather than biased. The benchmarks need to keep getting harder, so that progress is real rather than an artefact of a stale test set. And the system has to keep seeing enough variety that it does not narrow itself into failure.

The failure mode

A loop that learns from the wrong signal does not stand still — it gets worse faster. Bias in the feedback compounds exactly as efficiently as quality does.

Four disciplines keep it honest: every interaction deposits · signals stay honest · benchmarks stay fresh · diversity is preserved. For most enterprises, building and operating all of this in-house would be a major undertaking in its own right — an AI platform team standing up before the first agent ships.

See the loop on your own traffic

A 30-minute walkthrough of the flywheel against your hardware mix, governance constraints, and top use case.

Request a demo

Built in, not assembled

This is exactly why the Enterprise AI Flywheel is a native capability of Bud Novaria rather than something you assemble yourself. When your organisation runs its AI on Bud Novaria, the loop is already turning in the background — capturing outcomes, maintaining the quality checks, and feeding what it learns back into your models, infrastructure, policies and reusable tools.

Where the division of labour lands
Your team
Builds agents
Use cases, domain logic, the business problem worth solving.
→
Bud
Runs the flywheel
Capture, judging, repair, distillation, serving — managed natively, across all four pools.
Figure 8 — your teams do not have to build the pipeline, staff it, or keep it running.

The practical consequence is a staffing one. Standing up the pipeline in Figure 7 is a platform programme: quality engineering, evaluation infrastructure, a training stack, and a serving layer that can absorb a new model several times a week without an incident. Running it on Bud Novaria makes that an operating characteristic of the platform rather than a team you hire.

Why the gap widens

The difference between the two approaches is not a fixed percentage. It compounds. An organisation running static AI pays the same price for the same mistakes month after month. An organisation running on Bud Novaria gets cheaper, faster and more accurate with every interaction — and the improvements deposit into pools that the next agent inherits for free.

Cost per resolved request, over time
Static AI AI on Bud Novaria Same start Time in production → Cost per resolved request
Figure 9 — illustrative, not measured. The shape is the claim: the two systems start in the same place, and the shaded area is what static AI pays for standing still.

With the Enterprise AI Flywheel built into Bud Novaria, everyday usage turns into lasting value across your models, infrastructure, decisions and reusable tools. The result is enterprise AI that does not just run in production, but keeps getting better because it is there — without your teams having to engineer that improvement themselves.

In short
  • MIT Project NANDA attributes the failure of most generative AI pilots to a learning gap, not to models, budget or talent.
  • One loop — serve, capture, extract, deposit — improves models, infrastructure, policy and reusable primitives at the same time.
  • Shopify and Cursor prove the loop works in production; the seven-stage pipeline behind it is why almost nobody else runs one.
Related reading Why enterprise AI doesn't need another tool — it needs a platform that owns the stack The 40-tool stack, the seven layers, and why nobody owns the joins →

The 95% figure is from MIT Project NANDA, The GenAI Divide: State of AI in Business 2025 — a preliminary report that has not been peer-reviewed, and should be read as directional rather than definitive. Sidekick figures (serving cost, latency, GPU count, request rate) are Shopify Engineering's own published estimates from "Sidekick's continual learning loop" (5 Aug 2026); the ~$27M → ~$1M annual figures are estimates, not billed amounts. The 0.8B buyer-profile experiment is as reported by AlphaSignal from Shopify's CEO. Cursor figures are from "Improving Cursor Tab with online RL" (Sep 2025). Figures 1 and 9 are illustrative shapes, not measured data — substitute your own cost-per-resolved-request series for a firm number.

Get the next one by email

Product releases, benchmarks, and deployment patterns. Monthly, one email, unsubscribe any time.

Unsubscribe any time.

BN
Written by
Bud Newsroom
Bud Ecosystem
Get started with Bud

Put your data on it.

The fastest way to see what an integrated AI operating system does for your enterprise is a proof-of-concept on your infrastructure, with your data.

01 Identify a use case where complexity, cost, or governance is a known pain point.
02 Joint discovery — Bud maps your AI pain points to platform capabilities.
03 POC in days, on your hardware, with your data.