Turn deployed GPU capacity into a platform that scales with demand.
Bud GPU Foundry is the GPU-as-a-service control plane. It runs the accelerators you already own as governed services — from bare metal to tokens — with one catalog, one permission model, one audit trail and one bill. Silicon to outcomes, fully integrated. Open source under Apache 2.0.
Silicon to outcomes, as one integrated stack.
GPU Foundry runs the accelerators you already own, across every cluster and site you register, and sells them as governed services — bare metal, virtual machines, Kubernetes, SLURM and serverless containers — to every tenant you serve, from a single node to a whole cluster, in one click. Then it keeps going, into inference, model services and token-based AI consumption on the same stack.
The platform is what was missing.
Accelerators arrive faster than the means to sell them well. Everyone who wants them wants something different — and serving each request separately leaves several stacks, several bills and no honest answer to who used what.
Bare metal
- A physical accelerator host, your OS, nobody else resident
- Provisioned in one workflow, power and break-fix as API calls
- Returned only after verified erasure, with evidence
Virtual machines
- Accelerators passed through at the hardware level
- Console that does not need the guest · snapshots · private images
- Placed by topology, not by count
Kubernetes
- Conformant upstream, a control plane you hold
- Accelerator node pools drawn from the same catalog
- Access brokered just in time — no standing kubeconfig
SLURM
- Login nodes, partitions and accounting, your accelerators behind them
- Keep your job scripts, module trees and workflows
- Accounting reconciled to the platform's own record
Serverless containers
- Declare an image, a command and a shape — the platform resolves the rest
- Stop is previewed before it happens; a queue is never silence
- Leases, schedules and idle thresholds release what is not working
Many tenants, nothing else shared
- Every unit names one isolation claim — about that unit, not the platform
- Service providers are offered only hardware- or fabric-enforced grades
- Resellers, sub-organizations and white-label on the same tree
Bare metal
Your service. Your invoice.
- Bud GPU Foundry is Layer 02 of the Bud Novaria AI OS stack — the GPU cloud layer. The capacity is installed; the platform was what was missing. A training team wants nodes and a scheduler, a platform team wants Kubernetes, an ISV wants a machine it can image, a research group wants a container and a notebook, a product team wants an endpoint and a token price. GPU Foundry serves all of them from one fleet across five silicon vendors — NVIDIA, AMD, Intel, Qualcomm, Huawei — with one catalog, one permission model, one audit trail and one invoice.
- One click, any shape: a tenant asks for a shape (cards, memory per card, interconnect); the platform filters to what that organization may buy, picks exactly one unit by a published rule and says which rule chose it, places the work by topology — the project's region, then clusters with placeable headroom, then the right silicon — and metering starts when the unit is healthy. Nothing before healthy is ever charged.
- Five service classes from one catalog — bare metal (sole occupancy, metered in node-hours), virtual machines (no shared kernel, allocation-seconds), Kubernetes (a conformant cluster of your own, access brokered just in time), SLURM (partitions, QoS and fair share, accounting reconciled to the platform's record) and serverless containers (scale to zero, wake on demand, stop previewed before it happens) — plus the tenancy model: five isolation grades, strongest first, with service providers offered only hardware- or fabric-enforced grades, and resellers, sub-organizations and white-label on the same tree.
- Operator controls make the fleet a business: isolation graded per unit, quota you can sell, metering and invoices read from the same phase log the console shows, resellers and white-label, an audit trail and evidence — for enterprises running chargeback by project and department, and for service providers and neoclouds serving mutually untrusting tenants on prepaid or contract billing.
- The payoff: your silicon, your service, your invoice. Silicon to outcomes, fully integrated. Open source under Apache 2.0.
Installs into your own clusters — on-premise, in colocation, in sovereign and air-gapped environments. No component needs the public internet to serve a tenant.
Idle accelerators become a governed service.
Everything between the accelerator and the invoice is one product: the ways tenants reach the platform, the control plane, the data plane, the site fabric, and the revenue record that ties them together. Adding a new way to consume the fleet never adds a second billing path or a second permission model. That is what keeps the platform governable as it grows from one rack to many sites — and what lets AI services run on top as a tenant, inheriting every guarantee.
Describe the shape your work needs. Receive exactly one unit.
The platform filters the catalog to what your organization may buy, picks one unit by a published rule, tells you which rule chose it, and places the work — on one node, or gang-scheduled across many with the fabric in mind.
no_offered_sku names the nearest units and what each one lacks; queued is a position in line with a reason, never silence.Filtered before it is displayed
What a project's organization may actually buy is decided by its tenant class and isolation ceiling before anything is shown. Units weaker than the ceiling are never offered — not even as alternatives.
Placed by topology, not by count
The project's region, then only clusters that could place this grade and shape right now, then the one with the most headroom, then cards that share a link and a socket at the bandwidth sold. Gang-scheduled work starts all-or-nothing.
Everything the console does, from a script or an agent
The API is the product. Console, command line, Python and TypeScript SDKs and an infrastructure-as-code provider are generated from one specification, and a tool server lets an autonomous agent use the same operations under the same audit trail.
Metering starts when what you bought answers. Not a second before.
State is observed from the hardware, never assumed from intent. Three records exist for everything the platform has ever accepted, kept apart on purpose — what was asked for, what was observed, what was billed — and the screen a tenant reads and the invoice they receive come from the same log.
A failure is a record, not a gap
Everything that reached admission produces an allocation record, including what failed. An absent record and a zero record are different facts, and the platform keeps the difference. Reasons come from one closed vocabulary:
start deadline exceeded,image pull failed,device never attached— each charged nothing.Published start deadlines
One deadline for a start whose image was already resident, a longer one for a start that had to fetch it. Missing it terminates the allocation, reports a machine-readable reason and charges nothing.
Replay cannot double-charge
A metering event's identity is a pure function of the allocation and its window bounds, so a replay after a broker failure is a no-op rather than a second row. Completeness is verified by an independent nightly recomputation rather than asserted.
What makes the fleet a business.
Schedulers share a cluster; a control plane sells one. These six are the difference between hardware your teams queue for and a service every tenant consumes — and every one of them is a property of the platform, not a process running alongside it.
Any silicon, every site
Operationalise NVIDIA, AMD, Intel, Qualcomm and Huawei accelerators, mixed vendors and generations, as a single governed pool. Register another cluster at runtime and quota, placement, addressing and billing extend to it without a second deployment.
Metering starts at healthy
Seven lifecycle phases are recorded; the charge starts at exactly one — healthy — and a failure is a zero-cost record with a reason from one closed vocabulary. Time spent provisioning, pulling images or failing to start is never charged, on any unit, under any circumstance.
Many tenants, nothing else shared
Every sellable unit names one isolation grade, ranked strongest first. An operator states once that an organization may buy nothing weaker than a given grade, and every request is measured against it. For service providers whose customers may be rivals, shared-kernel grades are withheld by the platform, not by guidance.
Serverless that scales to zero
Models and workloads appear when called, release when idle, and chain into compound AI systems. Leases, schedules and idle thresholds release capacity that is not working — each previewed like a manual stop, so nothing is lost that was not stated first.
From capacity into AI consumption
Model catalogs, inference endpoints, serverless models and token metering are delivered by Bud AI Foundry, which runs on GPU Foundry as a tenant like any other — no private interface, no exemption. GPU Foundry meters capacity; AI Foundry meters tokens. Two meters, one truth.
Open source, sovereign by design
Apache 2.0, so the stack that governs your fleet can be audited rather than trusted. Delivered as signed images and versioned charts into clusters you run; an air-gapped site is an ordinary deployment. Regions are declared by the operator, so data residency is a placement record, not a setting someone asserts.
Many tenants on one infrastructure, with nothing else shared.
Your own departments, a service provider's mutually untrusting customers, a reseller's sub-organizations and a sovereign customer can all run on the same fleet at once. What differs between them is recorded as a property of the organization, never as a separate deployment.
Isolation is a claim about the unit sold
Each sellable unit names exactly one service class, one isolation mechanism, one metering basis and the tenant classes it may be sold to. A unit binds to machines by what the hardware publishes about itself, so one catalog serves a fleet of mixed generations without a unit landing on the wrong silicon.
Two tenant classes, one platform
Enterprise tenants get every grade, chargeback against internal budgets and a quota ceiling. Service-provider tenants get only grades enforced in hardware or by the fabric, prepaid balance by default, and the full enforcement ladder. The API, the console, the lifecycle and the guarantees never change.
A usage feed for someone else's billing
Metered usage is a continuous, replayable feed keyed on the same windows the platform's own invoices read, so a reseller can rate it in their own system. What they charge is theirs; what was consumed is one number.
Five ways to consume the same fleet.
A machine, a guest, a cluster, a scheduler or a container. Each is an entry in the same catalog, carries one isolation mechanism and one metering basis, and is watched by the same lifecycle informer. Choose by what your work needs, not by which product your vendor sells.
| Service class | What you receive | Isolation | Metered on |
|---|---|---|---|
| Bare metal | A physical accelerator host provisioned to an operating system you choose, on your own network and storage, reclaimed with verified erasure. Power and break-fix are API calls. | Sole occupancynobody else resident · returned with evidence | node-hours |
| Virtual machine | A hardware-virtualized guest with its own kernel, accelerators passed through at the hardware level, a boot volume, a console and a power-state lifecycle. Snapshots and private images. | No shared kernelnever a shared kernel with another tenant | allocation-seconds |
| Kubernetes | A conformant control plane of your own, with accelerator node pools from the same catalog, your storage classes and your network, reached without a standing credential. | Dedicated control planeper-tenant node pools · brokered access | node-hours · per cluster |
| SLURM | An HPC cluster with login nodes, partitions, quality-of-service classes and fair share, whose accounting reconciles to the platform's own record. Keep your job scripts and module trees. | Per-tenant clusterlogin nodes through the same gateway | node-hours |
| Serverless container | A workload declared as an image, a command and a shape, placed on the grade the catalog resolved, reachable by shell, terminal or a published address. Scale to zero, wake on demand. | Per gradewhole node to hardware partition | allocation-seconds · GPU-seconds · VRAM-GB-hours |
Against the GPU-cloud platform you have already been shown.
Rafay is the name most sovereign operators, neoclouds and systems integrators meet first in this lane: an NVIDIA-Certified delivery layer for NVIDIA Run:ai, sold as a closed platform. The comparison is about platform risk before it is about features.
Rafay — an NVIDIA-anchored delivery layer
- Proprietary, closed-source platform; NVIDIA is its primary distribution partner
- NVIDIA Run:ai is the payload, Rafay the delivery vehicle — a Rafay customer runs NVIDIA's orchestration through Rafay's layer
- NVIDIA-Certified for HGX and NVL72 rack-scale systems; NVIDIA AI Cloud-Ready validated
- Kubernetes, VMs, SLURM and inference services across data center, cloud, hybrid and air-gapped — under policy, quota and audit
- Repositioned in 2026 around token-metered AI services on top of bare-metal GPU services
The objection to leaving is not feature parity. It is "why step off the path NVIDIA endorses" — and that is the question worth answering.
An open control plane for any silicon
- Open source under Apache 2.0 — the stack that governs your fleet can be audited, forked and kept
- Silicon-neutral: NVIDIA, AMD, Intel, Qualcomm and Huawei in one catalog, mixed vendors and generations in one fleet
- Bare metal, virtual machines, Kubernetes, SLURM and serverless from one catalog — then inference, models and tokens on the same stack
- The layer above is already there: Bud AI Foundry serves models as a tenant, Bud SENTRY governs, Bud Agent runs agents — one AI operating system, not one control plane per vendor
- Metering that starts when what you bought answers, with three records kept apart — asked, observed, billed
A stack welded to one vendor's orchestration cannot follow you onto AMD, Intel, Qualcomm or Huawei. A fleet that is silicon-neutral underneath keeps the choice of accelerator yours, generation after generation.
If your fleet is NVIDIA end to end, you want NVIDIA Run:ai specifically, and the NVIDIA certification matters to your buyer, Rafay is the shorter path today. It has the longer list of neocloud and telco references, the NVIDIA channel behind it, and a managed-service motion Bud does not sell directly.
Platform risk. The question to put to an NVIDIA-certified credential is what happens to the control plane when the next rack is not NVIDIA. GPU Foundry answers with an open stack you can audit and 600+ supported accelerator SKUs across five vendors.
What sits above. Rafay stops where tokens start. GPU Foundry is Layer 02 of Bud Novaria AI OS: models, routing, guardrails and agents run on the same fleet, under the same tenancy and the same audit trail.
Where you are starting from. Greenfield fleets run on GPU Foundry today. A migration path for Rafay-managed fleets is on the roadmap.
Comparison current as of Q4 2026 · Rafay facts from its own published materials · Talk to a solutions architect to go through it side by side
The full story, in depth.
The five service classes in full, the isolation grades and tenant classes, the seven lifecycle phases and the reason vocabulary, quota and placement, billing, the twenty-two things GPU Foundry will never do — and every headline number with its basis.
Product Brief
Bud GPU Foundry Product Brief
The deep-dive product reference
Service classes, shape resolution, tenancy and isolation grades, the observed lifecycle, quota and placement, network and storage, access and identity, billing, AI consumption on the same stack, and the full positioning against Rafay.
Read the product briefPlatform White Paper
The Enterprise AI Management Platform
Where GPU Foundry fits in the platform
The platform-level argument — why compute, training, serving, and governance belong on one plane, and the economics that follow.
Read the white paperGPU Foundry is released under Apache 2.0. The public repository link will appear here at the open-source release.
Put your fleet on it.
The fastest way to see what a GPU-as-a-service control plane does for the accelerators you already own is a proof-of-concept on your infrastructure, with your tenants.