Bud GPU Foundry
The GPU-as-a-service control plane. Run the accelerators you already own — NVIDIA, AMD, Intel, Qualcomm and Huawei, across every cluster and site you register — as governed services, from bare metal to tokens, with one catalog, one permission model, one audit trail and one bill. Open source under Apache 2.0.
Silicon to outcomes, as one integrated stack.
GPU Foundry runs the accelerators you already own, across every cluster and site you register, and sells them as governed services — bare metal, virtual machines, Kubernetes, SLURM and serverless containers — to every tenant and segment you serve, from a single node to a whole cluster, in one click. Then it keeps going, into inference, model services and token-based AI consumption on the same stack.
What became of Bud Pod. Bud GPU Foundry succeeds Bud Pod. Pod was the experimentation layer: researchers shared GPUs and went from notebook to production in one click. GPU Foundry keeps that and becomes the GPU-as-a-service control plane for enterprises and service providers — five service classes from one catalog, isolation per tenant, metering that starts when what was bought answers, and AI consumption as a tenant of the same platform.
Layer 02 — the governed fleet under the stack.
GPU Foundry turns the raw capacity LayerZero exposes into a served, multi-tenant fleet. Every layer above trains, serves, grounds, governs and builds on capacity it sells — and Bud AI Foundry runs on it as a tenant, which is how token-billed customers inherit its guarantees.
You pool — accelerators across your data centers, sites and regions become one governed fleet. Register another cluster at runtime and quota, placement, addressing and billing extend to it.
They consume — departments or paying tenants get bare metal, guests, clusters, schedulers and containers from one catalog: quota'd, metered from healthy, and isolated down to the fabric.
The capacity is installed. The platform is what was missing. Accelerators arrive faster than the means to sell them well. A training team wants nodes and a scheduler; a platform team wants Kubernetes; an ISV wants a machine it can image; a research group wants a container and a notebook; a product team wants an endpoint and a token price. Serving each separately leaves several stacks, several bills and no honest answer to who used what. GPU Foundry serves all of them from one fleet, one catalog and one invoice.
Six things one control plane has to get right.
The service classes, the shape resolution, the tenancy, the observed lifecycle, quota and placement, and the step from capacity into AI consumption — each expanded to the specifics an evaluator needs.
Five ways to consume the same fleet, governed by one control plane
A machine, a guest, a cluster, a scheduler or a container. Choose by what your work needs, not by which product your vendor sells. Adding a service class never adds a second billing path, a second permission model or a second audit trail.
| Service class | What you receive | Isolation | Metered on |
|---|---|---|---|
| Bare metal | A physical accelerator host provisioned to an operating system you choose, on your own network and storage, reclaimed with verified erasure. Power on, off, reset, rescue and maintenance state are API calls attributed to a person or token; out-of-band console access never crosses your network. | Sole occupancynobody else resident · returned only after verified erasure, with evidence | node-hours |
| Virtual machine | A hardware-virtualized guest with its own kernel, accelerators passed through at the hardware level, a boot volume, a console and a power-state lifecycle. A stopped guest keeps its boot volume and addresses and stops being metered. Snapshots and private images; confidential execution where the silicon offers it, sold as its own unit. | No shared kernelnever a shared kernel with another tenant · placed by topology, not by count | allocation-seconds |
| Kubernetes | A conformant upstream cluster of your own — three most recent minor releases, highly available control planes with encrypted state — with accelerator node pools drawn from the same catalog, your storage classes and your network. Credentials issued just in time, expiring on their own. | Dedicated control planeper-tenant node pools · no standing kubeconfig to leak | node-hours · per cluster |
| SLURM | Login nodes, named partitions, quality-of-service classes and fair share, with an accounting database and your own accelerators behind them. Keep your job scripts, module trees and workflows; containers run natively inside jobs; a POSIX filesystem with per-user quotas is mounted across every node. | Per-tenant clusteraccounting reconciled to observed allocations · divergence goes to human review | node-hours |
| Serverless container | A workload declared as an image, a command and a shape, placed on the grade the catalog resolved, reachable by shell, terminal or a stable published address. Creation returns a handle; replay is safe by construction; stop is previewed before it happens; leases, schedules and idle thresholds release what is not working. | Per gradewhole node to hardware partition | allocation-seconds · GPU-seconds · VRAM-GB-hours |
Any use case, any environment, one click
Describe the shape your work needs. The platform filters to what your organization may buy, picks exactly one unit by a published rule, tells you which rule chose it, and places the work — on one node, or gang-scheduled across many with the fabric in mind. Resolution is also a read-only call, so a cost model or a procurement script can ask what a shape would buy without launching anything.
When nothing fits, you are told exactly why
| What happened | Reason | What you are told |
|---|---|---|
| No offered unit fits the shape | no_offered_sku | The nearest offered units and what each one lacks: memory, card count, interconnect, grade or model |
| The combination is not in the catalog | off_catalogue | The combination and the catalog version, and nothing else |
| Not sold to your tenant class | profile_refused | The constraint that refused it, without naming units you may not buy |
| The unit exists but is not on sale | not_offered | The grade and the written reason. Never a waitlist for something that is not a capacity problem |
| The grade exceeds your organization's ceiling | above_ceiling | The ceiling. Alternatives the ceiling refuses are never offered as substitutes |
| Your organization's terms permit nothing | terms_exclude_all | A different audience: no edit to the request will help, and the operator must act |
| No quota right now | queued | A queued request with its position in line and a machine-readable reason |
Templates hold a reusable declaration and are copied into the resource at launch, so editing one never changes anything already running. A curated starter library can declare a memory floor, so a request that could never have loaded is refused at declaration rather than at runtime. Every template and image carries a trust tier — official, verified or community.
Many tenants and customer segments on one infrastructure, with nothing else shared
Every sellable unit names exactly one service class, one isolation mechanism, one metering basis and the tenant classes it may be sold to. Isolation is a claim about the unit sold, not about the platform in general — and every unit carries one isolation grade, ranked strongest first. An operator states once that an organization may buy nothing weaker than a given grade, and every request in every project is measured against it.
| Grade | What the tenant gets | Rank | Sold to |
|---|---|---|---|
| Bare-metal host | The machine itself, your OS, nobody else resident | rank 0 | both tenant classes |
| Whole node | Every accelerator on a node, yours for the duration | rank 0 | both tenant classes |
| Dedicated GPU in a guest | A whole device passed through to your own kernel | rank 1 | both tenant classes |
| Hardware partition | Dedicated memory and compute paths | rank 2 | both tenant classes |
| Memory-capped slice | A share with a hard memory ceiling | rank 3 | enterprise only |
| Enterprise tenant class | Service-provider tenant class | |
|---|---|---|
| Who the tenants are | Your own business units, departments or subsidiaries, sharing one fleet | Independent customers who are mutually untrusting and may compete with each other |
| Isolation posture | Every grade is available, including the memory-capped slice, because a hostile neighbor is not the threat | Only grades enforced in hardware or by the fabric. Shared-kernel grades are withheld by the platform, not by guidance |
| Billing | Chargeback by project and cost center, against internal budgets | Prepaid balance by default; contract billing in arrears after an identity and credit assessment |
| Enforcement | Quota ceilings and notifications; the eviction ladder exists but is rarely armed | The full ladder for prepaid tenants; contract tenants are never suspended automatically for balance reasons |
| What does not change | The API, the console, the command line, the service classes, the lifecycle, the permission model, the audit trail and the guarantees. A tenant class is a property of the organization and of the unit being sold, not a separate product and not a separate deployment. | |
Resellers and white-label. A reseller organization holds sub-organizations it creates, administers and is billed for. It is granted quota and divides it among its customers in whole instances; consumption aggregates back up the same tree. A reseller administrator has no route to a customer's workloads, sessions or logs — the audit trail records the boundary holding. Brand, logo, color, custom domain and language are properties of an organization, applied to the console, documentation links, invitations and notifications. Rate cards are drafted, compared and published with an effective date, and a published rate never applies retroactively.
Metering starts when what you bought answers. Not a second before.
State is observed from the hardware, never assumed from intent. Three records exist for everything the platform has ever accepted, kept apart on purpose: what was asked for, what was observed, and what was billed. The screen a tenant reads and the invoice they receive come from the same log, so they cannot disagree.
| Reason | What it means | Charged |
|---|---|---|
| start deadline exceeded | It did not reach healthy within the deadline for its start path. Terminated. | nothing |
| image pull failed | The image could not be retrieved: wrong reference, missing credential, or a registry that did not answer. | nothing |
| device never attached | It was scheduled and the accelerator never attached by the mechanism its grade requires. | nothing |
| provisioning failed | A machine, guest, cluster or scheduler did not complete provisioning; the underlying system's own words are carried alongside. | nothing |
| released before healthy | The allocation disappeared before it was ever usable. Recorded as a failure. | nothing |
| backfilled | Reconstructed after the fact from observed timestamps, for capacity that ran without going through the platform's own path. | nothing · counted |
| stopped | Ended because somebody asked it to stop. | up to the stop |
| failed | The workload itself failed, with the cluster's own words carried alongside. | up to the failure |
The log is append-only and keyed on allocation and phase, so a replay is a no-op rather than a second row. A metering event's identity is a pure function of the allocation and its window bounds, so a replay after a broker failure cannot double-charge. Completeness is verified by an independent nightly recomputation rather than asserted.
Scales with demand: quota you can sell, placement you can predict
Quota is the instrument an operator sells with. A grant is an organization's allowance for one unit in one region, in whole instances. Projects draw on it through named shares, idle headroom is borrowable inside the organization — never across organizations — and every placement follows the same published rule.
Network isolation, durable storage, and several ways in through one authorizer
Network. A tenant's network is programmed as part of provisioning them, and it is the same private routed scope whichever service class they bought. Multi-node units receive a dedicated fabric partition; a fabric the platform cannot confirm offers no capacity at all. Platform control surfaces default-deny from every tenant, and the orchestration substrate is not a network a tenant workload can see, let alone authenticate to. Egress is metered per tenant whether or not it is billed.
Storage. The platform drives the storage systems you already own — network volumes, a shared POSIX filesystem, boot volumes, S3-compatible object storage, versioned images and declared-ephemeral scratch disk. Each tenant receives its own namespace, access zone and object policy in the backing system; the enforced quota is the number billed; checkpoint writes are acknowledged only once durably stored; and a tenant with a contractual requirement supplies and controls the key protecting its data.
Access. SSH on your own key material or a short-lived certificate, a browser terminal and console on a single-use ticket, managed notebooks, brokered kubectl with credentials issued just in time, and SLURM login over the same gateway. Every session has a beginning, an end, an actor and, where recording is enabled, a replayable transcript. Revocation reaches live connections. Operator staff hold no standing access and cannot attach to a running tenant workload — the break-glass path is cordon and evacuate, refused and audited otherwise.
Billing an enterprise buyer recognizes, and spend caps that act
An allocation window with a provisioning phase priced at zero exists on the first record ever written. Metering, rating, invoices, chargeback, budgets and enforcement all read that one log, which is what makes a disputed charge answerable: a tenant receives a per-window explanation of what was billed and why, including the provisioning windows charged at zero.
From capacity to AI consumption: inference, model services and tokens on the same stack
Model catalogs, inference endpoints, serverless models, token metering and agent services are delivered by Bud AI Foundry, which runs on GPU Foundry as a tenant like any other. It receives no private interface and no exemption, so a token-billed customer inherits the isolation, quota and audit trail of the capacity beneath them. The capacity plane does not need to know what a token is; the model plane does not re-implement quota, isolation or metering. What connects them is specified, not assumed.
allocation-seconds · node-hours · capacity-hours — from the phase log, from healthy on. Margin per endpoint is a measurement, not an inference at the end of a quarter.
tokens served — per endpoint, per customer, per model, per agent service — recorded apart from the capacity. AI Foundry's own customers appear as sub-organizations, so a token-billed customer's consumption rolls up through the same tree as everything else.
Where several tenants are served by one model process, no tenant code executes and the isolation claim is made at the request boundary — stated as its own unit in the catalog rather than borrowing a compute grade's claim.
Everything between the accelerator and the invoice is one product.
The ways tenants reach the platform, the control plane, the data plane, the site fabric, and the revenue record that ties them all together. AI services sit on top as a tenant of the same platform, which is how they inherit its guarantees.
Identity, revenue and platform state live in separate domains that are never joined and never reference each other. A failure in one cannot take the others with it.
The API writes what was asked for. The operator writes the substrate. The lifecycle informer alone writes the phase log that metering reads. Nothing else can.
Signed images and versioned charts in a declarative deployment tree, pinned and upgraded together through a staging canary. Which version runs where is reviewable configuration.
Each cluster's region is stated by the operator who onboards it, and every tenant placement happens inside the project's region. Residency is evidence, not a preference.
Work that must have exactly one writer runs under a lease. Work that must scale with demand runs behind a load balancer. Neither is allowed to become the other.
The platform installs into clusters you run. No component needs the public internet to serve a tenant, so an air-gapped site is an ordinary deployment.
Six places in the console, and what each one is for
| Place | What it holds | Who |
|---|---|---|
| Operate | Fleet inventory, clusters, hosts and individual accelerators, with health and the reason anything is not sellable. Networks, address pools, storage systems. | operator |
| Sell | The sellable unit set with its reasons, quota granted per organization per region, rate cards, cost floors and margin. | operator |
| Build | Templates, images and the starter library; launch with resolve-before-submit for every service class; queues with a position and a reason. | tenant |
| Watch | Detail for a machine, guest, cluster or workload: phases with the metering marker, live logs, utilization, sessions, the embedded terminal and the stop preview. | both |
| Govern | Organizations, sub-organizations, projects, people, roles, invitations, domains, tokens, keys and the audit trail. | both |
| Spend | Usage, chargeback by cost center, invoices, budgets and caps, showing the same numbers the invoice will carry. | tenant |
The API is the product; the console, the command line and the SDKs are clients of it, generated from the same published specification. Every operation has a command-line command, the console may call only published endpoints, and no decision lives in a place the command line cannot reach.
Runs where you are: installed into your data centers, pinned as one set, upgraded through a canary.
Fonts and assets are self-hosted, images are mirrored into your own registry, and no component needs the public internet to serve a tenant. Register another cluster at runtime and the API, the console and quota span it; withdrawal cordons first and refuses a cluster still holding tenant capacity.
Silicon, environments and delivery
Silicon
Mixed vendors and generations in one fleet. A unit binds to machines by what the hardware publishes about itself, so one catalog serves a fleet of mixed generations without a unit landing on the wrong silicon. 600+ accelerator SKUs supported across the Bud Novaria AI OS.
Environments
An air-gapped deployment is a supported configuration rather than a special case. Each cluster's region is declared by the operator who onboards it, and work is placed inside its project's region, so data residency is a record you can show rather than a setting you assert.
Delivery & operations
Every pin records how mature that component is, so upgrade cadence is planned against real risk. The control plane is restored from backup within a published recovery time, verified by timed drill. Faulty accelerators drain without data loss and are quarantined individually. Every externally visible commitment is validated by a repeatable drill before it is published.
A tenant's life, both halves specified
Onboarding is usually designed carefully and exit usually is not. The exit half is what an enterprise procurement review asks about first, because it decides whether choosing the platform is reversible.
Directory deactivation revokes platform access, active sessions and issued keys within a published time bound. A tenant can move between organizations, resellers, plans or billing models without interruption to what is running and without published addresses changing.
Every headline number, with its basis.
The figures on this page are architectural properties and published commitments, not benchmark results. Each is paired with where it comes from — and the twenty-two guarantees below are stated in the negative so they can be held to.
Twenty-two things GPU Foundry will never do
- Bill for capacity that never became healthyMetering derives from the phase log, and starts at one phase and no other.
- Display capacity it cannot placeAvailability is derived from placeability per cluster, with a freshness stamp on every figure.
- Reclaim interruptible capacity on less notice than the tenant's declared checkpoint intervalThe notice period is a property of the unit and is delivered before reclamation begins.
- Discard logs when a workload crashesLines are read from the tenant log store, never from the thing that wrote them.
- Publish a best-case latency figure as a typical oneEvery published figure is a measured percentile with its measurement method stated.
- Require a tenant to detect a platform failure in order to be credited for itCredits are computed from platform telemetry and offered without a claim.
- Charge a stopped resource more than a running one without saying so at the point of stoppingThe stop preview and the unit's metering basis are both shown before the action.
- Withhold compliance documentation or security features behind a spend thresholdA published commercial policy, not a matter of sales discretion.
- Hold the only copy of a key protecting tenant data where the tenant has contracted otherwiseTenant-supplied keys, whose revocation renders the data unrecoverable.
- Suspend a contract-billed tenant automatically for balance reasonsFor that class, enforcement is a quota ceiling and a notification; suspension needs a recorded human decision.
- Present utilization telemetry as a charge basis on any unit where it cannot observe the tenantBlind classes are declared in the catalog in advance and render as a stated absence.
- Offer a fractional tier to mutually untrusting tenants unless the isolation is enforced in hardwareTenant classes are a property of the catalog row, and the generator refuses the combination outright.
- Expose a tenant to another tenant's residual accelerator memoryVerified memory clearing per hardware generation before a device is re-offered.
- Return a machine to the pool without verified erasure of the media that held tenant dataSanitization is part of the reclaim workflow, with evidence retained for the deletion period.
- Sell a dedicated fabric path that is in fact shared with another tenantFabric partitions are per tenant, and a fabric the platform cannot confirm offers no capacity at all.
- Issue a tenant credentials of any kind to the platform's own orchestration control planeAsserted against the API and against the access gateway's certificate issuer. Your cluster is yours; ours is never reachable.
- Require a tenant to hold a long-lived credential to reach what they boughtCertificates, tickets and cluster credentials are minted on demand and expire on their own.
- Let a reseller see inside its customers' workA reseller relation reaches quota, consumption and spend; there is no relation from it to a workload, a session or a log line.
- Allow a capacity request to bypass quota accounting by any pathAdmission policy checks that work is routed by the scheduler, and anything that slips past is counted.
- Promise live migration of accelerator-attached guestsStated as impossible rather than deferred, so no plan ever quietly assumes it.
- Market an isolation grade it does not enforce on the specific unit being soldRuntime pinning by provenance, and grade constraints checked against what the hardware publishes.
- Give operator personnel an interactive session into a running tenant workloadThe break-glass path is cordon and evacuate; the attempt is refused and audited.
How these claims are framed. Counts on this page — service classes, grades, phases, guarantees, capability groups — are properties of the catalog and the lifecycle, not benchmark results. Deployment-specific numbers (utilization gains, start times on your hardware, metering accuracy against your own accounting) are established in a proof-of-concept on your infrastructure. Independent assessment is completed before the first external tenant and repeated at least annually; the platform is validated against the accelerator vendor's own requirements for AI clouds — service delivery by API, hard and soft multi-tenancy, telemetry, metering accuracy and lifecycle operations.
Against the platform you have already been shown.
Rafay is the name most sovereign operators, neoclouds and systems integrators meet first in this lane. It markets its GPU platform as the way GPU clouds deliver NVIDIA Run:ai as a fully automated, multi-tenant managed service — Run:ai is the payload, Rafay the delivery vehicle. The comparison is about platform risk before it is about the feature table.
| Capability | Rafay | Hyperscaler GPU cloud | DIY — Kubernetes + scheduler | Bud GPU Foundry |
|---|---|---|---|---|
| License | Proprietary, closed source | Managed service | Open components, your integration | Apache 2.0 |
| Silicon | NVIDIA-anchored: Run:ai payload, NVIDIA-Certified for HGX and NVL72 | The provider's catalog | Per-vendor device plugins, your work | NVIDIA · AMD · Intel · Qualcomm · Huawei — one catalog |
| Service classes | Kubernetes, VMs, SLURM, inference | VMs, managed Kubernetes | Kubernetes, typically only | Bare metal · VMs · Kubernetes · SLURM · serverless |
| Multi-tenant metering & billing | Yes — policy, quota, audit | The provider's, not yours to sell | Build it | Phase-log metering · prepaid & contract · resellers & white-label |
| Air-gapped / sovereign | Yes | No | Yes | Yes — an ordinary deployment |
| From capacity to tokens | Token-metered services, repositioned 2026 | Separate AI services | Build it | AI Foundry as a tenant — two meters, one truth |
| Agents & governance in the same OS | No | Separate products | Build it | Bud Novaria AI OS — Model Foundry → AI Foundry → SENTRY → Agent on the same fleet |
| Auditable control plane | Closed | Closed | Yes | Yes — open source, signed, version-pinned |
If your fleet is NVIDIA end to end, you want NVIDIA Run:ai specifically, and the NVIDIA certification matters to your buyer, Rafay is the shorter path today. It has the longer list of neocloud and telco references, the NVIDIA channel behind it — NVIDIA is its primary distribution partner — and a managed-service motion Bud does not sell directly.
A Rafay customer is not running an alternative to NVIDIA's orchestration; they are running NVIDIA's orchestration through Rafay's delivery layer. The objection to leaving is "why step off the path NVIDIA endorses", not feature parity.
Platform risk. A stack welded to one vendor's orchestration cannot follow you onto AMD, Intel, Qualcomm or Huawei. The question to put to an NVIDIA-certified credential is what happens to the control plane when the next rack is not NVIDIA. GPU Foundry answers with an open stack you can audit and 600+ supported accelerator SKUs across five vendors.
What sits above. Rafay stops where tokens start. GPU Foundry is Layer 02 of Bud Novaria AI OS: models, routing, guardrails and agents run on the same fleet, under the same tenancy and the same audit trail, so an operator is not paying for one control plane per hardware vendor and one more per layer.
Where you are starting from. Greenfield fleets run on GPU Foundry today. A migration path for Rafay-managed fleets is on the roadmap.
Comparison current as of Q4 2026 · Rafay facts from its own published materials · hyperscaler and DIY columns are category-level, not a named vendor
Two kinds of operator. Every kind of tenant.
GPU Foundry is for every organization with accelerators in racks and demand in queues — whether the tenants are internal departments, paying customers, or AI services of its own.
Enterprise platform teams
Give every department governed self-service on a shared fleet — quota by department, chargeback by project and cost center, every isolation grade available, idle headroom borrowed inside the organization.
Service providers & neoclouds
Sell GPU-as-a-service to mutually untrusting customers on your own hardware — hardware- and fabric-enforced isolation only, prepaid or contract billing, spend caps that act, and a usage feed for your own rating.
Sovereign & regulated operators
Run a full GPU cloud in disconnected, air-gapped and regulated environments — regions declared not discovered, export-control screening and evidence as platform features, tenant-held keys.
Systems integrators & resellers
A reseller organization with its own brand, domain, price book and sub-organizations — quota divided in whole instances, consumption rolled up the same tree, no route into a customer's work.
Research & HPC groups
SLURM delivered as a governed tenancy — login nodes, partitions, QoS, fair share, a shared POSIX filesystem, topology-aware job placement — with accounting that reconciles to the platform's record.
Teams shipping inference
Serverless containers that scale to zero and wake on demand, and Bud AI Foundry's endpoints, model catalog and token metering running on the same fleet as a tenant.
The platform around the fleet.
This brief is the reference for Bud GPU Foundry. For the platform-level argument — why compute, training, serving, grounding and governance belong on one plane — read the whitepaper, or return to the product overview.
Platform White Paper
The Enterprise AI Management Platform
Where the governed fleet fits
The platform-level argument — why compute, training, serving, and governance belong on one plane, and the economics that follow.
Read the white paperBack to overview
Bud GPU Foundry
The GPU-as-a-service control plane
Return to the product page — the film, the figures, the service classes and the positioning at a glance.
Back to the product pagePut your fleet on it.
The fastest way to see what a GPU-as-a-service control plane does for the accelerators you already own is a proof-of-concept on your infrastructure, with your tenants.