Home/ Products/ Bud GPU Foundry/ Product Brief
Bud GPU Foundry overview
Product Brief · Layer 02 · GPU Cloud

Bud GPU Foundry

The GPU-as-a-service control plane. Run the accelerators you already own — NVIDIA, AMD, Intel, Qualcomm and Huawei, across every cluster and site you register — as governed services, from bare metal to tokens, with one catalog, one permission model, one audit trail and one bill. Open source under Apache 2.0.

Product reference v2.0 October 2026 ~16 min read
01At a glance

Silicon to outcomes, as one integrated stack.

GPU Foundry runs the accelerators you already own, across every cluster and site you register, and sells them as governed services — bare metal, virtual machines, Kubernetes, SLURM and serverless containers — to every tenant and segment you serve, from a single node to a whole cluster, in one click. Then it keeps going, into inference, model services and token-based AI consumption on the same stack.

Service classes from one catalog5
Isolation grades, scoped to the unit sold5
Observed lifecycle phases7
Platform credentials issued to a tenant0
the five service classes are detailed in §03 · every number's basis in §06
What it is
The GPU-as-a-service control plane: everything between the accelerator and the invoice — the ways tenants reach the platform, the control plane, the data plane, the site fabric, and the revenue record that ties them together — delivered as one signed set you install in your own data centers
One pooled fleet of NVIDIA, AMD, Intel, Qualcomm and Huawei accelerators, mixed vendors and generations, sold through one catalog with one permission model, one audit trail and one bill
Built for two operators: enterprises giving departments governed self-service on a shared fleet, and service providers, neoclouds and resellers selling GPU-as-a-service to mutually untrusting customers — with AI services on top as a tenant of the same platform
What it is not
✕Not a scheduler bolted onto a cluster — tenancy, metering, isolation and billing are the product, and adding a service class never adds a second billing path or a second permission model
✕Not an NVIDIA-only operating layer — the catalog binds units to what the hardware publishes about itself, so one fleet of mixed silicon is one catalog, and the control plane is open source under Apache 2.0
✕Not someone else's cloud — it installs into clusters you run; no component needs the public internet to serve a tenant, so an air-gapped site is an ordinary deployment

What became of Bud Pod. Bud GPU Foundry succeeds Bud Pod. Pod was the experimentation layer: researchers shared GPUs and went from notebook to production in one click. GPU Foundry keeps that and becomes the GPU-as-a-service control plane for enterprises and service providers — five service classes from one catalog, isolation per tenant, metering that starts when what was bought answers, and AI consumption as a tenant of the same platform.

02Where it fits

Layer 02 — the governed fleet under the stack.

GPU Foundry turns the raw capacity LayerZero exposes into a served, multi-tenant fleet. Every layer above trains, serves, grounds, governs and builds on capacity it sells — and Bud AI Foundry runs on it as a tenant, which is how token-billed customers inherit its guarantees.

You pool — accelerators across your data centers, sites and regions become one governed fleet. Register another cluster at runtime and quota, placement, addressing and billing extend to it.

They consume — departments or paying tenants get bare metal, guests, clusters, schedulers and containers from one catalog: quota'd, metered from healthy, and isolated down to the fabric.

The capacity is installed. The platform is what was missing. Accelerators arrive faster than the means to sell them well. A training team wants nodes and a scheduler; a platform team wants Kubernetes; an ISV wants a machine it can image; a research group wants a container and a notebook; a product team wants an endpoint and a token price. Serving each separately leaves several stacks, several bills and no honest answer to who used what. GPU Foundry serves all of them from one fleet, one catalog and one invoice.

03Capabilities, in full

Six things one control plane has to get right.

The service classes, the shape resolution, the tenancy, the observed lifecycle, quota and placement, and the step from capacity into AI consumption — each expanded to the specifics an evaluator needs.

01Five service classes, one catalogBare metal · virtual machines · Kubernetes · SLURM · serverless containers — each an entry in the same catalog with one isolation mechanism and one metering basis, watched by the same lifecycle informer5 classes
02Shape resolution — one click, any use caseDescribe the shape; the platform filters to what you may buy, picks exactly one unit by a published rule, says which rule chose it, and places the work on one node or gang-scheduled across many1 unit · 1 rule
03Tenancy — grades, classes, resellersFive ranked isolation grades · two tenant classes · reseller organizations with sub-organizations, white-label and price books per segment · a usage feed for someone else's billing5 grades
04Observed lifecycle & meteringSeven phases read from the hardware, never assumed from intent · metering starts at healthy · failures are zero-cost records with a reason from one closed vocabulary · three records kept apart7 phases
05Quota, placement & scaleQuota granted per organization, per unit, per region, in whole instances · borrowing inside the organization · a published placement rule · gang scheduling all-or-nothing · queues with a position and a reason1 placement rule
06From capacity into AI consumptionModel catalogs, inference endpoints, serverless models and token metering delivered by Bud AI Foundry as a tenant · GPU Foundry meters capacity, AI Foundry meters tokens · sub-tenancy carries through2 meters · 1 truth

Five ways to consume the same fleet, governed by one control plane

A machine, a guest, a cluster, a scheduler or a container. Choose by what your work needs, not by which product your vendor sells. Adding a service class never adds a second billing path, a second permission model or a second audit trail.

Service classWhat you receiveIsolationMetered on
Bare metalA physical accelerator host provisioned to an operating system you choose, on your own network and storage, reclaimed with verified erasure. Power on, off, reset, rescue and maintenance state are API calls attributed to a person or token; out-of-band console access never crosses your network.Sole occupancynobody else resident · returned only after verified erasure, with evidencenode-hours
Virtual machineA hardware-virtualized guest with its own kernel, accelerators passed through at the hardware level, a boot volume, a console and a power-state lifecycle. A stopped guest keeps its boot volume and addresses and stops being metered. Snapshots and private images; confidential execution where the silicon offers it, sold as its own unit.No shared kernelnever a shared kernel with another tenant · placed by topology, not by countallocation-seconds
KubernetesA conformant upstream cluster of your own — three most recent minor releases, highly available control planes with encrypted state — with accelerator node pools drawn from the same catalog, your storage classes and your network. Credentials issued just in time, expiring on their own.Dedicated control planeper-tenant node pools · no standing kubeconfig to leaknode-hours · per cluster
SLURMLogin nodes, named partitions, quality-of-service classes and fair share, with an accounting database and your own accelerators behind them. Keep your job scripts, module trees and workflows; containers run natively inside jobs; a POSIX filesystem with per-user quotas is mounted across every node.Per-tenant clusteraccounting reconciled to observed allocations · divergence goes to human reviewnode-hours
Serverless containerA workload declared as an image, a command and a shape, placed on the grade the catalog resolved, reachable by shell, terminal or a stable published address. Creation returns a handle; replay is safe by construction; stop is previewed before it happens; leases, schedules and idle thresholds release what is not working.Per gradewhole node to hardware partitionallocation-seconds · GPU-seconds · VRAM-GB-hours

Any use case, any environment, one click

Describe the shape your work needs. The platform filters to what your organization may buy, picks exactly one unit by a published rule, tells you which rule chose it, and places the work — on one node, or gang-scheduled across many with the fabric in mind. Resolution is also a read-only call, so a cost model or a procurement script can ask what a shape would buy without launching anything.

01You ask for a shape4 cards · 80 GB per card · high-bandwidth interconnect — or name a unit outright.
02Filtered to what you may buyYour tenant class, your isolation ceiling, your project's region. Units you may not buy stay hidden.
03Chosen by a published ruleSmallest memory per card that satisfies the request; ties go to the stronger grade. One sentence names the rule.
04Exactly one unitWhole node · 4 × 80 GB · dedicated fabric path · metered in allocation-seconds, fixed for the workload's life.
05Placed by topologySingle node, or gang-scheduled all-or-nothing across a cluster, with the fabric in mind.

When nothing fits, you are told exactly why

What happenedReasonWhat you are told
No offered unit fits the shapeno_offered_skuThe nearest offered units and what each one lacks: memory, card count, interconnect, grade or model
The combination is not in the catalogoff_catalogueThe combination and the catalog version, and nothing else
Not sold to your tenant classprofile_refusedThe constraint that refused it, without naming units you may not buy
The unit exists but is not on salenot_offeredThe grade and the written reason. Never a waitlist for something that is not a capacity problem
The grade exceeds your organization's ceilingabove_ceilingThe ceiling. Alternatives the ceiling refuses are never offered as substitutes
Your organization's terms permit nothingterms_exclude_allA different audience: no edit to the request will help, and the operator must act
No quota right nowqueuedA queued request with its position in line and a machine-readable reason

Templates hold a reusable declaration and are copied into the resource at launch, so editing one never changes anything already running. A curated starter library can declare a memory floor, so a request that could never have loaded is refused at declaration rather than at runtime. Every template and image carries a trust tier — official, verified or community.

Many tenants and customer segments on one infrastructure, with nothing else shared

Every sellable unit names exactly one service class, one isolation mechanism, one metering basis and the tenant classes it may be sold to. Isolation is a claim about the unit sold, not about the platform in general — and every unit carries one isolation grade, ranked strongest first. An operator states once that an organization may buy nothing weaker than a given grade, and every request in every project is measured against it.

GradeWhat the tenant getsRankSold to
Bare-metal hostThe machine itself, your OS, nobody else residentrank 0both tenant classes
Whole nodeEvery accelerator on a node, yours for the durationrank 0both tenant classes
Dedicated GPU in a guestA whole device passed through to your own kernelrank 1both tenant classes
Hardware partitionDedicated memory and compute pathsrank 2both tenant classes
Memory-capped sliceA share with a hard memory ceilingrank 3enterprise only
Enterprise tenant classService-provider tenant class
Who the tenants areYour own business units, departments or subsidiaries, sharing one fleetIndependent customers who are mutually untrusting and may compete with each other
Isolation postureEvery grade is available, including the memory-capped slice, because a hostile neighbor is not the threatOnly grades enforced in hardware or by the fabric. Shared-kernel grades are withheld by the platform, not by guidance
BillingChargeback by project and cost center, against internal budgetsPrepaid balance by default; contract billing in arrears after an identity and credit assessment
EnforcementQuota ceilings and notifications; the eviction ladder exists but is rarely armedThe full ladder for prepaid tenants; contract tenants are never suspended automatically for balance reasons
What does not changeThe API, the console, the command line, the service classes, the lifecycle, the permission model, the audit trail and the guarantees. A tenant class is a property of the organization and of the unit being sold, not a separate product and not a separate deployment.

Resellers and white-label. A reseller organization holds sub-organizations it creates, administers and is billed for. It is granted quota and divides it among its customers in whole instances; consumption aggregates back up the same tree. A reseller administrator has no route to a customer's workloads, sessions or logs — the audit trail records the boundary holding. Brand, logo, color, custom domain and language are properties of an organization, applied to the console, documentation links, invitations and notifications. Rate cards are drafted, compared and published with an effective date, and a published rate never applies retroactively.

Metering starts when what you bought answers. Not a second before.

State is observed from the hardware, never assumed from intent. Three records exist for everything the platform has ever accepted, kept apart on purpose: what was asked for, what was observed, and what was billed. The screen a tenant reads and the invoice they receive come from the same log, so they cannot disagree.

Fig. 1Each service class declares its own readiness: a container answering, a guest's agent responding, a cluster's control plane serving with its first node joined, a scheduler node online, a bare-metal host reachable on your own operating system. A unit with no declared readiness source cannot be published at all.
ReasonWhat it meansCharged
start deadline exceededIt did not reach healthy within the deadline for its start path. Terminated.nothing
image pull failedThe image could not be retrieved: wrong reference, missing credential, or a registry that did not answer.nothing
device never attachedIt was scheduled and the accelerator never attached by the mechanism its grade requires.nothing
provisioning failedA machine, guest, cluster or scheduler did not complete provisioning; the underlying system's own words are carried alongside.nothing
released before healthyThe allocation disappeared before it was ever usable. Recorded as a failure.nothing
backfilledReconstructed after the fact from observed timestamps, for capacity that ran without going through the platform's own path.nothing · counted
stoppedEnded because somebody asked it to stop.up to the stop
failedThe workload itself failed, with the cluster's own words carried alongside.up to the failure

The log is append-only and keyed on allocation and phase, so a replay is a no-op rather than a second row. A metering event's identity is a pure function of the allocation and its window bounds, so a replay after a broker failure cannot double-charge. Completeness is verified by an independent nightly recomputation rather than asserted.

Scales with demand: quota you can sell, placement you can predict

Quota is the instrument an operator sells with. A grant is an organization's allowance for one unit in one region, in whole instances. Projects draw on it through named shares, idle headroom is borrowable inside the organization — never across organizations — and every placement follows the same published rule.

01The project's regionAnd only that region. Residency is a placement record, not a preference.
02Clusters that can place it nowThose that could actually place this grade and shape — not those with a matching free-device count.
03The most headroomCapacity is spread, never packed into the first cluster that fits. Draining and cordoned clusters are never chosen.
04The right silicon by topologyCards that share a link and a socket, at the bandwidth sold. Re-run for queued work as headroom changes.
Gang scheduling starts all-or-nothing Whole instances, never fractions Queue position and a reason, never silence Preemption never shorter than the declared checkpoint interval

Network isolation, durable storage, and several ways in through one authorizer

Network. A tenant's network is programmed as part of provisioning them, and it is the same private routed scope whichever service class they bought. Multi-node units receive a dedicated fabric partition; a fabric the platform cannot confirm offers no capacity at all. Platform control surfaces default-deny from every tenant, and the orchestration substrate is not a network a tenant workload can see, let alone authenticate to. Egress is metered per tenant whether or not it is billed.

Storage. The platform drives the storage systems you already own — network volumes, a shared POSIX filesystem, boot volumes, S3-compatible object storage, versioned images and declared-ephemeral scratch disk. Each tenant receives its own namespace, access zone and object policy in the backing system; the enforced quota is the number billed; checkpoint writes are acknowledged only once durably stored; and a tenant with a contractual requirement supplies and controls the key protecting its data.

Access. SSH on your own key material or a short-lived certificate, a browser terminal and console on a single-use ticket, managed notebooks, brokered kubectl with credentials issued just in time, and SLURM login over the same gateway. Every session has a beginning, an end, an actor and, where recording is enabled, a replayable transcript. Revocation reaches live connections. Operator staff hold no standing access and cannot attach to a running tenant workload — the break-glass path is cordon and evacuate, refused and audited otherwise.

Billing an enterprise buyer recognizes, and spend caps that act

An allocation window with a provisioning phase priced at zero exists on the first record ever written. Metering, rating, invoices, chargeback, budgets and enforcement all read that one log, which is what makes a disputed charge answerable: a tenant receives a per-window explanation of what was billed and why, including the provisioning windows charged at zero.

rung 1NotifyThe tenant is told they are approaching the limit, with time to act. Nothing changes about what is running.
rung 2HoldNew work stops being admitted. Work already admitted keeps running. A hold never drains, because draining a training run to enforce a budget destroys more value than it protects.
rung 3Revoke accessInteractive access is withdrawn while the work continues, so nobody can start new work by hand around the hold.
rung 4EvictThe last rung, and only for tenant classes where it applies. A contract-billed tenant never reaches it automatically; suspension needs a recorded human decision.
Prepaid balance drawn down, forecast visible Contract in arrears, purchase-order references, cost-center breakdown Commitments backed by a reservation ledger Nightly reconciliation across balance, scheduler and invoice

From capacity to AI consumption: inference, model services and tokens on the same stack

Model catalogs, inference endpoints, serverless models, token metering and agent services are delivered by Bud AI Foundry, which runs on GPU Foundry as a tenant like any other. It receives no private interface and no exemption, so a token-billed customer inherits the isolation, quota and audit trail of the capacity beneath them. The capacity plane does not need to know what a token is; the model plane does not re-implement quota, isolation or metering. What connects them is specified, not assumed.

GPU Foundry meters capacity

allocation-seconds · node-hours · capacity-hours — from the phase log, from healthy on. Margin per endpoint is a measurement, not an inference at the end of a quarter.

AI Foundry meters tokens

tokens served — per endpoint, per customer, per model, per agent service — recorded apart from the capacity. AI Foundry's own customers appear as sub-organizations, so a token-billed customer's consumption rolls up through the same tree as everything else.

Where several tenants are served by one model process, no tenant code executes and the isolation claim is made at the request boundary — stated as its own unit in the catalog rather than borrowing a compute grade's claim.

04How it works

Everything between the accelerator and the invoice is one product.

The ways tenants reach the platform, the control plane, the data plane, the site fabric, and the revenue record that ties them all together. AI services sit on top as a tenant of the same platform, which is how they inherit its guarantees.

Fig. 2The planes of GPU Foundry. The revenue plane reads the same log the console shows, so a screen and an invoice cannot disagree.
Three state domains, kept apart

Identity, revenue and platform state live in separate domains that are never joined and never reference each other. A failure in one cannot take the others with it.

Exactly one writer per seam

The API writes what was asked for. The operator writes the substrate. The lifecycle informer alone writes the phase log that metering reads. Nothing else can.

Delivered as one signed set

Signed images and versioned charts in a declarative deployment tree, pinned and upgraded together through a staging canary. Which version runs where is reviewable configuration.

Regions are declared, not discovered

Each cluster's region is stated by the operator who onboards it, and every tenant placement happens inside the project's region. Residency is evidence, not a preference.

Handlers and reconcilers, separated

Work that must have exactly one writer runs under a lease. Work that must scale with demand runs behind a load balancer. Neither is allowed to become the other.

Your hardware, your data centers

The platform installs into clusters you run. No component needs the public internet to serve a tenant, so an air-gapped site is an ordinary deployment.

Six places in the console, and what each one is for

PlaceWhat it holdsWho
OperateFleet inventory, clusters, hosts and individual accelerators, with health and the reason anything is not sellable. Networks, address pools, storage systems.operator
SellThe sellable unit set with its reasons, quota granted per organization per region, rate cards, cost floors and margin.operator
BuildTemplates, images and the starter library; launch with resolve-before-submit for every service class; queues with a position and a reason.tenant
WatchDetail for a machine, guest, cluster or workload: phases with the metering marker, live logs, utilization, sessions, the embedded terminal and the stop preview.both
GovernOrganizations, sub-organizations, projects, people, roles, invitations, domains, tokens, keys and the audit trail.both
SpendUsage, chargeback by cost center, invoices, budgets and caps, showing the same numbers the invoice will carry.tenant

The API is the product; the console, the command line and the SDKs are clients of it, generated from the same published specification. Every operation has a command-line command, the console may call only published endpoints, and no decision lives in a place the command line cannot reach.

05Deployment & compatibility

Runs where you are: installed into your data centers, pinned as one set, upgraded through a canary.

Fonts and assets are self-hosted, images are mirrored into your own registry, and no component needs the public internet to serve a tenant. Register another cluster at runtime and the API, the console and quota span it; withdrawal cordons first and refuses a cluster still holding tenant capacity.

Silicon, environments and delivery

Silicon

NVIDIAAMDIntelQualcommHuawei

Mixed vendors and generations in one fleet. A unit binds to machines by what the hardware publishes about itself, so one catalog serves a fleet of mixed generations without a unit landing on the wrong silicon. 600+ accelerator SKUs supported across the Bud Novaria AI OS.

Environments

On-premiseColocationSovereignAir-gappedMulti-site

An air-gapped deployment is a supported configuration rather than a special case. Each cluster's region is declared by the operator who onboards it, and work is placed inside its project's region, so data residency is a record you can show rather than a setting you assert.

Delivery & operations

Signed imagesVersioned chartsStaging canaryDrilled restoreApache 2.0

Every pin records how mature that component is, so upgrade cadence is planned against real risk. The control plane is restored from backup within a published recovery time, verified by timed drill. Faulty accelerators drain without data loss and are quarantined individually. Every externally visible commitment is validated by a repeatable drill before it is published.

A tenant's life, both halves specified

Onboarding is usually designed carefully and exit usually is not. The exit half is what an enterprise procurement review asks about first, because it decides whether choosing the platform is reversible.

01SignupSignup to running capacity without human intervention, subject to the identity-verification tier.
02OnboardingFederated identity, directory provisioning, network scope, storage and contract terms configured before the first workload.
03OperatingProjects, people, quota, budgets, machines and services, all self-served, with an audit trail the tenant reads for its own scope.
04ClosureA final reconciled invoice to the moment of closure, not to the end of a billing period.
05ExitData exported within a published window at full rate, then deleted — backups included — and evidenced on request.

Directory deactivation revokes platform access, active sessions and issued keys within a published time bound. A tenant can move between organizations, resellers, plans or billing models without interruption to what is running and without published addresses changing.

06Proof & methodology

Every headline number, with its basis.

The figures on this page are architectural properties and published commitments, not benchmark results. Each is paired with where it comes from — and the twenty-two guarantees below are stated in the negative so they can be held to.

5
Service classes from one catalog
BasisThe machine-readable catalog. Every sellable unit is bound to exactly one service class — bare metal, virtual machine, Kubernetes, SLURM or serverless container — one isolation mechanism, one metering basis and one set of permitted tenant classes. Any combination not in the matrix is refused before capacity is reserved.
5
Isolation grades, ranked
BasisBare-metal host and whole node at rank 0, dedicated GPU in a guest at rank 1, hardware partition at rank 2, memory-capped slice at rank 3 (enterprise only). The claim is attached to the row being sold, never to the platform in general, and an organization's ceiling is compared by rank on every request.
7
Observed lifecycle phases
BasisAdmitted, scheduled, image ready, device attached, healthy, released, failed — each read from the capacity actually running the work. Metering starts at healthy, phase 5, and at no other phase. Phases 1–4 are never metered; failed is a zero-cost record with a reason.
0
Platform credentials issued to a tenant
BasisGuarantee 16, asserted against the API and against the access gateway's certificate issuer. Your Kubernetes cluster is yours; the platform's own orchestration control plane is never reachable from a tenant, in any form. Certificates, tickets and cluster credentials are minted on demand and expire on their own.
22
Guarantees, each closed by a check
BasisCapabilities stated in the negative, not statements of intent. Each is closed by an automated check that attempts the forbidden behavior and fails if it succeeds. Most describe something that happens routinely elsewhere in this market.

Twenty-two things GPU Foundry will never do

  1. Bill for capacity that never became healthyMetering derives from the phase log, and starts at one phase and no other.
  2. Display capacity it cannot placeAvailability is derived from placeability per cluster, with a freshness stamp on every figure.
  3. Reclaim interruptible capacity on less notice than the tenant's declared checkpoint intervalThe notice period is a property of the unit and is delivered before reclamation begins.
  4. Discard logs when a workload crashesLines are read from the tenant log store, never from the thing that wrote them.
  5. Publish a best-case latency figure as a typical oneEvery published figure is a measured percentile with its measurement method stated.
  6. Require a tenant to detect a platform failure in order to be credited for itCredits are computed from platform telemetry and offered without a claim.
  7. Charge a stopped resource more than a running one without saying so at the point of stoppingThe stop preview and the unit's metering basis are both shown before the action.
  8. Withhold compliance documentation or security features behind a spend thresholdA published commercial policy, not a matter of sales discretion.
  9. Hold the only copy of a key protecting tenant data where the tenant has contracted otherwiseTenant-supplied keys, whose revocation renders the data unrecoverable.
  10. Suspend a contract-billed tenant automatically for balance reasonsFor that class, enforcement is a quota ceiling and a notification; suspension needs a recorded human decision.
  11. Present utilization telemetry as a charge basis on any unit where it cannot observe the tenantBlind classes are declared in the catalog in advance and render as a stated absence.
  12. Offer a fractional tier to mutually untrusting tenants unless the isolation is enforced in hardwareTenant classes are a property of the catalog row, and the generator refuses the combination outright.
  13. Expose a tenant to another tenant's residual accelerator memoryVerified memory clearing per hardware generation before a device is re-offered.
  14. Return a machine to the pool without verified erasure of the media that held tenant dataSanitization is part of the reclaim workflow, with evidence retained for the deletion period.
  15. Sell a dedicated fabric path that is in fact shared with another tenantFabric partitions are per tenant, and a fabric the platform cannot confirm offers no capacity at all.
  16. Issue a tenant credentials of any kind to the platform's own orchestration control planeAsserted against the API and against the access gateway's certificate issuer. Your cluster is yours; ours is never reachable.
  17. Require a tenant to hold a long-lived credential to reach what they boughtCertificates, tickets and cluster credentials are minted on demand and expire on their own.
  18. Let a reseller see inside its customers' workA reseller relation reaches quota, consumption and spend; there is no relation from it to a workload, a session or a log line.
  19. Allow a capacity request to bypass quota accounting by any pathAdmission policy checks that work is routed by the scheduler, and anything that slips past is counted.
  20. Promise live migration of accelerator-attached guestsStated as impossible rather than deferred, so no plan ever quietly assumes it.
  21. Market an isolation grade it does not enforce on the specific unit being soldRuntime pinning by provenance, and grade constraints checked against what the hardware publishes.
  22. Give operator personnel an interactive session into a running tenant workloadThe break-glass path is cordon and evacuate; the attempt is refused and audited.

How these claims are framed. Counts on this page — service classes, grades, phases, guarantees, capability groups — are properties of the catalog and the lifecycle, not benchmark results. Deployment-specific numbers (utilization gains, start times on your hardware, metering accuracy against your own accounting) are established in a proof-of-concept on your infrastructure. Independent assessment is completed before the first external tenant and repeated at least annually; the platform is validated against the accelerator vendor's own requirements for AI clouds — service delivery by API, hard and soft multi-tenancy, telemetry, metering accuracy and lifecycle operations.

07Positioning · as of Q4 2026

Against the platform you have already been shown.

Rafay is the name most sovereign operators, neoclouds and systems integrators meet first in this lane. It markets its GPU platform as the way GPU clouds deliver NVIDIA Run:ai as a fully automated, multi-tenant managed service — Run:ai is the payload, Rafay the delivery vehicle. The comparison is about platform risk before it is about the feature table.

CapabilityRafayHyperscaler GPU cloudDIY — Kubernetes + schedulerBud GPU Foundry
LicenseProprietary, closed sourceManaged serviceOpen components, your integrationApache 2.0
SiliconNVIDIA-anchored: Run:ai payload, NVIDIA-Certified for HGX and NVL72The provider's catalogPer-vendor device plugins, your workNVIDIA · AMD · Intel · Qualcomm · Huawei — one catalog
Service classesKubernetes, VMs, SLURM, inferenceVMs, managed KubernetesKubernetes, typically onlyBare metal · VMs · Kubernetes · SLURM · serverless
Multi-tenant metering & billingYes — policy, quota, auditThe provider's, not yours to sellBuild itPhase-log metering · prepaid & contract · resellers & white-label
Air-gapped / sovereignYesNoYesYes — an ordinary deployment
From capacity to tokensToken-metered services, repositioned 2026Separate AI servicesBuild itAI Foundry as a tenant — two meters, one truth
Agents & governance in the same OSNoSeparate productsBuild itBud Novaria AI OS — Model Foundry → AI Foundry → SENTRY → Agent on the same fleet
Auditable control planeClosedClosedYesYes — open source, signed, version-pinned
Where Rafay wins

If your fleet is NVIDIA end to end, you want NVIDIA Run:ai specifically, and the NVIDIA certification matters to your buyer, Rafay is the shorter path today. It has the longer list of neocloud and telco references, the NVIDIA channel behind it — NVIDIA is its primary distribution partner — and a managed-service motion Bud does not sell directly.

A Rafay customer is not running an alternative to NVIDIA's orchestration; they are running NVIDIA's orchestration through Rafay's delivery layer. The objection to leaving is "why step off the path NVIDIA endorses", not feature parity.

Where the comparison gets specific

Platform risk. A stack welded to one vendor's orchestration cannot follow you onto AMD, Intel, Qualcomm or Huawei. The question to put to an NVIDIA-certified credential is what happens to the control plane when the next rack is not NVIDIA. GPU Foundry answers with an open stack you can audit and 600+ supported accelerator SKUs across five vendors.

What sits above. Rafay stops where tokens start. GPU Foundry is Layer 02 of Bud Novaria AI OS: models, routing, guardrails and agents run on the same fleet, under the same tenancy and the same audit trail, so an operator is not paying for one control plane per hardware vendor and one more per layer.

Where you are starting from. Greenfield fleets run on GPU Foundry today. A migration path for Rafay-managed fleets is on the roadmap.

Apache 2.0, auditable Silicon-neutral across five vendors Bare metal to tokens, one stack Metering from healthy, three records apart Air-gapped as an ordinary deployment

Comparison current as of Q4 2026 · Rafay facts from its own published materials · hyperscaler and DIY columns are category-level, not a named vendor

08Who it's for Optional

Two kinds of operator. Every kind of tenant.

GPU Foundry is for every organization with accelerators in racks and demand in queues — whether the tenants are internal departments, paying customers, or AI services of its own.

Enterprise

Enterprise platform teams

Give every department governed self-service on a shared fleet — quota by department, chargeback by project and cost center, every isolation grade available, idle headroom borrowed inside the organization.

Provider

Service providers & neoclouds

Sell GPU-as-a-service to mutually untrusting customers on your own hardware — hardware- and fabric-enforced isolation only, prepaid or contract billing, spend caps that act, and a usage feed for your own rating.

Sovereign

Sovereign & regulated operators

Run a full GPU cloud in disconnected, air-gapped and regulated environments — regions declared not discovered, export-control screening and evidence as platform features, tenant-held keys.

Channel

Systems integrators & resellers

A reseller organization with its own brand, domain, price book and sub-organizations — quota divided in whole instances, consumption rolled up the same tree, no route into a customer's work.

Research

Research & HPC groups

SLURM delivered as a governed tenancy — login nodes, partitions, QoS, fair share, a shared POSIX filesystem, topology-aware job placement — with accounting that reconciles to the platform's record.

Product

Teams shipping inference

Serverless containers that scale to zero and wake on demand, and Bud AI Foundry's endpoints, model catalog and token metering running on the same fleet as a tenant.

09Go deeper & next steps

The platform around the fleet.

This brief is the reference for Bud GPU Foundry. For the platform-level argument — why compute, training, serving, grounding and governance belong on one plane — read the whitepaper, or return to the product overview.

Get started with Bud

Put your fleet on it.

The fastest way to see what a GPU-as-a-service control plane does for the accelerators you already own is a proof-of-concept on your infrastructure, with your tenants.

01 Register one cluster — any vendor, any generation — and publish a catalog.
02 Onboard two tenants with different isolation ceilings and watch the catalog differ.
03 Read the first invoice against the phase log. They cannot disagree.