Partners/Cloud Service Providers
For Cloud Service Providers

Selling GPU-hours is a race to the bottom.

H100 rates have collapsed 64–75% in under two years. The differentiator was never the GPU — it's the 85% of the stack above it. Bud Novaria deploys that stack on the infrastructure you already own, in 30 days.

The squeeze

Three forces are closing at once.

Every one of them compresses the same line on your P&L: the price of a GPU-hour.

Prices collapsed
Market survey · Mar 2026

H100 $8.00 → $2.50–3.50/hr — a fall of 64–75% in under two years. Aggressive providers are already quoting $1.25–1.90. A100 went $3.10 → $1.50, −52%.

300+ new GPU clouds entered the market in 2025. Capacity is not scarce any more; it is inventory.

Margin of safety gone
McKinsey, Nov 2025 · public filings

Headline bare-metal-as-a-service EBITDA reads 57–62%. After depreciation and interest, EBIT lands near 8%. McKinsey's verdict: BMaaS is inherently commoditised.

CoreWeave: $5.13B 2025 revenue on a 4% Q3 operating margin, after an $863M net loss in 2024. Scale does not fix the shape of this business.

Platform captured
Synergy Research, Q3 2025 · Microsoft

AWS, Azure and GCP hold 66% of cloud infrastructure spend — $102.6B a quarter — and are committing $600B+ of combined AI capex in 2026.

Azure AI Foundry deploys 1,900+ models in five clicks, and 65% of Azure customers are already evaluating it. The layer above your GPUs is being sold without you.

Demand spiral
MIT / S&P Global, 2025

95% of GenAI pilots produce zero P&L impact. 42% of companies have abandoned most AI initiatives, up from 17% a year earlier.

Average loss per failed initiative: $7.2M. Your customers are not short of compute. They are short of outcomes.

$547B
of $684B in enterprise AI investment failed to deliver value in 2025. The gap isn't GPUs — it's the 85% of the stack above them.
Sources: McKinsey, Nov 2025 · Synergy Research, Q3 2025 · MIT / S&P Global, 2025 · public filings.
The opportunity

A $500 billion market is forming above the GPU.

Five tiers. The same racks underneath all of them. Find your rung, then read upward — every step adds revenue per customer, margin, and years of retention.

The value ladder Revenue & margin climb
T5 Outcome-as-a-Service Paid on business results $50–200K+Rev / cust / mo 65–80%Gross margin 5–8%Annual churn
T4 Agent-as-a-Service Production agents with SLOs $30–100K+Rev / cust / mo 55–70%Gross margin 8–12%Annual churn
T3 Model-aaS & Fine-tuning Hosted and tuned models $10–30KRev / cust / mo 45–60%Gross margin 12–15%Annual churn
T2 AI Infra · Inference-aaS Managed inference $6–15KRev / cust / mo 35–50%Gross margin 15–20%Annual churn
T1 GPU-as-a-Service Where most CSPs are — terminal $2–5KRev / cust / mo 15–25%Gross margin 25%+Annual churn
10–40×
more revenue per customer, bottom to top of the ladder.
$80B
sovereign-AI IaaS market in 2026, 35.6% CAGR — hyperscalers structurally can't enter.
$47–53B
agentic AI market by 2030 — zero neoclouds offer built-in agent orchestration today.
Analysys Mason · Gartner · MarketsandMarkets. Ladder pricing and margin bands: Bud Ecosystem, April 2026.
30×
lifetime value
A customer paying 10× more who stays 3× longer. That is the whole argument for climbing — and it is a retention story before it is a revenue story.
The alternatives

Every other path costs a year or leaves you renting.

Four ways to get a stack above your GPUs. Only one of them ships this quarter.

  Build in-house OpenShift / Nutanix / VMware Neoclouds Bud Novaria AIOS
Time to production18–24 months6–12 monthsn/a30 days
Cost$8–15M+/yr team + $2–5M infra$0.5–2M licence + $0.9–2.5M staffRevenue share
Managed inference
Managed RAG
Agent orchestration
Guardrails
Air-gapped deployment
Hardware freedom
Validated Year-5 revenueUnknownGPU-only economicsGPU-only economics$54.2M
3.2:1
AI talent demand to supply. You are hiring against every hyperscaler.
42%
of AI/ML roles sit unfilled for six months or longer.
$1.0–2.5M
a year of DIY tax, over a 3–18 month integration window.
Neoclouds assessed: CoreWeave, Together, Lambda, RunPod, Nebius. None ship guardrails, managed RAG, or air-gapped deployment. Orchestration and hardware freedom marked partial — Together and RunPod ship limited orchestration; RunPod supports AMD.
The stack

What Bud puts on your racks.

Seven layers, one deployment. This is the 85% of the stack that sits above the GPU and carries the margin.

600+
hardware SKUs
3.6×
throughput vs SGLang
160+
guardrail policies
1,000+
MCP integrations
60+
prebuilt agents
9
revenue streams live
07Consumption & adoption
Bud Studio
Natural-language to agent. Adoption moves from 5–10% of staff to enterprise-wide.
06Operations & FinOps
Bud Scaler
SLO-aware autoscaling with per-tenant cost attribution — billing you can defend.
05Governance & safety
Bud SENTRY + Bud Sentinel
23 CPU-native guardrails at 0.70ms p50 — $0.10 vs $24 per million, 239× cheaper.
04Agent orchestration
Bud Agent + MCP Foundry
60+ prebuilt agents and 1,000+ MCP integrations your tenants can wire themselves.
03Knowledge & data
Bud Runtime + MCP Foundry
200+ data sources and a managed vector database — RAG without a data team.
02Universal inference
Bud Runtime + Bud Latent
3.6× throughput vs SGLang, 12× faster cold start, under 1% embedding error.
01Hardware abstraction
Bud LayerZero + FCSP
600+ hardware SKUs. Utilisation moves from 35–50% to 80–90% on the fleet you own.
Cross-cutting Bud Model Foundry + ART — 120+ architectures, 420+ commercial models, and the self-improvement flywheel that keeps them current.

Hardware freedom

One control plane across NVIDIA, AMD, Intel, Gaudi and NPUs. You buy on price and availability, not on which vendor your software forces.

CPU-native guardrails

Governance that runs on the CPUs already in your racks — 0.70ms p50, $0.10 per million classifications. No GPU tax on safety.

Demand creation

Bud Studio puts agent-building in the hands of your customers' business teams. Consumption stops depending on their engineers.

The three claims a hyperscaler cannot answer.
The math

Same racks. Same power. A different business on top.

500 H100-equivalents. 50 enterprise customers at the start. $2.85/hr declining 12% a year. Every assumption is on the page.

Revenue per GPU
$6,288 $44,439
Roughly 7× more revenue from the same silicon — 42% utilisation on a rental model versus 85% on a software-defined one.
Metric
GPU-only
CSP + Bud
Basis
Year 1 revenue
$5.2M
$10.6M
500 H100-equivalents, 50 enterprise customers, $2.85/hr declining 12%/yr.
Year 3 revenue
$6.3M
$23.8M
Revenue mix shifts to 25% Model-aaS and 25% Agent-aaS.
Year 5 revenue
$7.7M
$54.2M
7.1× the GPU-only case; 90% of revenue software-defined by Year 5. Base case: 35% growth, 10% churn.
Year 5 operating margin
−7.9%
+28.2%
Bud-enabled is profitable from Year 1 at +2.5%.
Year 5 customers
27 eroded
120 grew
25% annual churn versus 10%.
Year 5 revenue / GPU
$6,288
$44,439
42% versus 85% utilisation.
5-year cumulative revenue
$32.0M
$138.5M
4.3× across the full curve.
5-year cumulative operating income
−$2.9M
+$31.2M
A $34M swing.
Customer lifetime value
$33,600
$1,478,340
$3.5K/mo at 20% margin over 4 years, versus $25.8K/mo at 47.8% margin over 10 years — 44×.
Sensitivity — Year 5
Conservative
$33.4M
27.9% margin
20% growth · 10% churn
Base
$54.2M
28.2% margin
35% growth · 10% churn
Optimistic
$87.6M
28.4% margin
50% growth · 10% churn
Pessimistic
$47.0M
28.1% margin
35% growth · 15% churn

Even the conservative case is 4.4× the GPU-only revenue — and the GPU-only model never reaches profitability across the five-year curve.

Bud Ecosystem five-year model, April 2026. Assumptions stated per row.
Proof

An agent that costs 80% less and answers 3.3× faster.

A global fashion brand's production agent, rebuilt on Bud. The workload did not change; the stack under it did.

Monthly cost
baseline−80%
Response latency
20s6s
Accuracy
85.44%85.46%
Go-to-market
baseline6.5×
Same fleet, higher tier
Cost fell four-fifths at equivalent accuracy.
The gain came from the layers above the GPU — routing, caching and CPU-native governance — not from more silicon.
What this means for you
The account started where you are: a fleet, a rental price sheet, and customers asking for outcomes. No new hardware was bought to move up the ladder.
The tier hyperscalers can't serve

Sovereignty is a market, not a feature.

Data residency, air-gap and jurisdiction are requirements a US hyperscaler cannot satisfy by policy. That exclusion is your addressable market.

$80B
sovereign cloud IaaS spend in 2026, growing at 35.6% CAGR to $169B by 2028. Gartner
20%
of workloads are migrating away from hyperscalers — the geopatriation shift.
51%
of German firms consider themselves US-dependent, with direct CLOUD Act exposure.
~50,000
GPU cap on Tier-2 procurement under the 2025 AI Diffusion Rule. Local capacity became strategic.
National commitments
India $250B+ · Saudi HUMAIN 600,000 GPUs · Korea $735B · the EU AI Act.
In production today — all CPU-native, $0 GPU, 4–8 weeks
India Income Tax Department 39 use cases
UAE Ministry of Finance In production
UAE Ministry of Health In production
South Korea Government In production

Four government deployments already shipped on this stack. Sovereign AI Consulting is the highest-margin line on the sheet at 65–80%, and it is the one hyperscalers are structurally excluded from bidding.

Activation

Four weeks from rack to revenue line.

Sequence matters. Each week turns on a tier of the ladder — nothing waits on a hardware order.

Week 1Foundation
LayerZero lands on the existing fleet and FCSP comes on. Hardware is abstracted, utilisation reporting starts, and the estate becomes addressable as one pool.
Week 2Serving
Runtime and Model Foundry go live. Inference-as-a-Service and Model-as-a-Service are sellable — this unlocks T2 and T3.
Week 3Agents
MCP Foundry, Bud Agent and Bud Studio switch on. Agent-as-a-Service and 1,000+ integrations reach your tenants — this unlocks T4.
Week 4Govern & bill
SENTRY, Scaler and FinOps complete the platform. Governance, multi-tenancy and per-tenant cost attribution are live — the prerequisites for outcome pricing at T5.
Get started with Bud

Put your GPUs on it.

A proof-of-concept on your infrastructure, with your workloads.

01 Audit your position on the ladder.
02 Joint discovery — map your fleet to the seven layers.
03 POC in days on your existing hardware.