Home/Products/Bud GPU Foundry
Bud GPU Foundry · Layer 02 · GPU Cloud

Turn deployed GPU capacity into a platform that scales with demand.

Bud GPU Foundry is the GPU-as-a-service control plane. It runs the accelerators you already own as governed services — from bare metal to tokens — with one catalog, one permission model, one audit trail and one bill. Silicon to outcomes, fully integrated. Open source under Apache 2.0.

Overview

Silicon to outcomes, as one integrated stack.

GPU Foundry runs the accelerators you already own, across every cluster and site you register, and sells them as governed services — bare metal, virtual machines, Kubernetes, SLURM and serverless containers — to every tenant you serve, from a single node to a whole cluster, in one click. Then it keeps going, into inference, model services and token-based AI consumption on the same stack.

5 service classes from one catalog 0 platform credentials ever issued to a tenant 22 guarantees, each closed by a check Apache 2.0 · on-prem · sovereign · air-gapped
  • Bud GPU Foundry is Layer 02 of the Bud Novaria AI OS stack — the GPU cloud layer. The capacity is installed; the platform was what was missing. A training team wants nodes and a scheduler, a platform team wants Kubernetes, an ISV wants a machine it can image, a research group wants a container and a notebook, a product team wants an endpoint and a token price. GPU Foundry serves all of them from one fleet across five silicon vendors — NVIDIA, AMD, Intel, Qualcomm, Huawei — with one catalog, one permission model, one audit trail and one invoice.
  • One click, any shape: a tenant asks for a shape (cards, memory per card, interconnect); the platform filters to what that organization may buy, picks exactly one unit by a published rule and says which rule chose it, places the work by topology — the project's region, then clusters with placeable headroom, then the right silicon — and metering starts when the unit is healthy. Nothing before healthy is ever charged.
  • Five service classes from one catalog — bare metal (sole occupancy, metered in node-hours), virtual machines (no shared kernel, allocation-seconds), Kubernetes (a conformant cluster of your own, access brokered just in time), SLURM (partitions, QoS and fair share, accounting reconciled to the platform's record) and serverless containers (scale to zero, wake on demand, stop previewed before it happens) — plus the tenancy model: five isolation grades, strongest first, with service providers offered only hardware- or fabric-enforced grades, and resellers, sub-organizations and white-label on the same tree.
  • Operator controls make the fleet a business: isolation graded per unit, quota you can sell, metering and invoices read from the same phase log the console shows, resellers and white-label, an audit trail and evidence — for enterprises running chargeback by project and department, and for service providers and neoclouds serving mutually untrusting tenants on prepaid or contract billing.
  • The payoff: your silicon, your service, your invoice. Silicon to outcomes, fully integrated. Open source under Apache 2.0.

Installs into your own clusters — on-premise, in colocation, in sovereign and air-gapped environments. No component needs the public internet to serve a tenant.

Value proposition

Idle accelerators become a governed service.

Everything between the accelerator and the invoice is one product: the ways tenants reach the platform, the control plane, the data plane, the site fabric, and the revenue record that ties them together. Adding a new way to consume the fleet never adds a second billing path or a second permission model. That is what keeps the platform governable as it grows from one rack to many sites — and what lets AI services run on top as a tenant, inheriting every guarantee.

Service classes from one catalog5
Isolation grades, each scoped to the unit sold5
Observed lifecycle phases · metering starts at one7
Platform credentials ever issued to a tenant0
architectural properties, not benchmarks · each number's basis in the product brief
One click, any use case

Describe the shape your work needs. Receive exactly one unit.

The platform filters the catalog to what your organization may buy, picks one unit by a published rule, tells you which rule chose it, and places the work — on one node, or gang-scheduled across many with the fabric in mind.

Fig. 1A shape resolves to one offered unit by a published rule, and the rule is returned in a sentence. Resolution is also a read-only call, so a cost model or a procurement script can ask what a shape would buy without launching anything.
  • Filtered before it is displayed

    What a project's organization may actually buy is decided by its tenant class and isolation ceiling before anything is shown. Units weaker than the ceiling are never offered — not even as alternatives.

  • Placed by topology, not by count

    The project's region, then only clusters that could place this grade and shape right now, then the one with the most headroom, then cards that share a link and a socket at the bandwidth sold. Gang-scheduled work starts all-or-nothing.

  • Everything the console does, from a script or an agent

    The API is the product. Console, command line, Python and TypeScript SDKs and an infrastructure-as-code provider are generated from one specification, and a tool server lets an autonomous agent use the same operations under the same audit trail.

Metering

Metering starts when what you bought answers. Not a second before.

State is observed from the hardware, never assumed from intent. Three records exist for everything the platform has ever accepted, kept apart on purpose — what was asked for, what was observed, what was billed — and the screen a tenant reads and the invoice they receive come from the same log.

Fig. 2The seven observed phases of every allocation. Each service class declares its own readiness — a container answering, a guest's agent responding, a cluster's control plane serving with its first node joined — and a unit with no declared readiness source cannot be published at all.
  • A failure is a record, not a gap

    Everything that reached admission produces an allocation record, including what failed. An absent record and a zero record are different facts, and the platform keeps the difference. Reasons come from one closed vocabulary: start deadline exceeded, image pull failed, device never attached — each charged nothing.

  • Published start deadlines

    One deadline for a start whose image was already resident, a longer one for a start that had to fetch it. Missing it terminates the allocation, reports a machine-readable reason and charges nothing.

  • Replay cannot double-charge

    A metering event's identity is a pure function of the allocation and its window bounds, so a replay after a broker failure is a no-op rather than a second row. Completeness is verified by an independent nightly recomputation rather than asserted.

Key features

What makes the fleet a business.

Schedulers share a cluster; a control plane sells one. These six are the difference between hardware your teams queue for and a service every tenant consumes — and every one of them is a property of the platform, not a process running alongside it.

01

Any silicon, every site

Operationalise NVIDIA, AMD, Intel, Qualcomm and Huawei accelerators, mixed vendors and generations, as a single governed pool. Register another cluster at runtime and quota, placement, addressing and billing extend to it without a second deployment.

5 silicon vendorsone catalog · one fleet · any mix of generations
02

Metering starts at healthy

Seven lifecycle phases are recorded; the charge starts at exactly one — healthy — and a failure is a zero-cost record with a reason from one closed vocabulary. Time spent provisioning, pulling images or failing to start is never charged, on any unit, under any circumstance.

7 observed phasesmetering starts at phase 5 · failed = never billed
03

Many tenants, nothing else shared

Every sellable unit names one isolation grade, ranked strongest first. An operator states once that an organization may buy nothing weaker than a given grade, and every request is measured against it. For service providers whose customers may be rivals, shared-kernel grades are withheld by the platform, not by guidance.

5 isolation gradesbare-metal host → memory-capped slice · per unit sold
04

Serverless that scales to zero

Models and workloads appear when called, release when idle, and chain into compound AI systems. Leases, schedules and idle thresholds release capacity that is not working — each previewed like a manual stop, so nothing is lost that was not stated first.

0 idle costwake on demand · release when idle · stop previewed
05

From capacity into AI consumption

Model catalogs, inference endpoints, serverless models and token metering are delivered by Bud AI Foundry, which runs on GPU Foundry as a tenant like any other — no private interface, no exemption. GPU Foundry meters capacity; AI Foundry meters tokens. Two meters, one truth.

2 meters, one truthallocation-seconds beneath · tokens above · one tree
06

Open source, sovereign by design

Apache 2.0, so the stack that governs your fleet can be audited rather than trusted. Delivered as signed images and versioned charts into clusters you run; an air-gapped site is an ordinary deployment. Regions are declared by the operator, so data residency is a placement record, not a setting someone asserts.

Apache 2.0on-prem · colocation · sovereign · air-gapped
Tenancy

Many tenants on one infrastructure, with nothing else shared.

Your own departments, a service provider's mutually untrusting customers, a reseller's sub-organizations and a sovereign customer can all run on the same fleet at once. What differs between them is recorded as a property of the organization, never as a separate deployment.

Fig. 3One fleet, run by one operator, sold by several businesses. A customer of a reseller sees the reseller: brand, domain, price book and language are properties of the organization. The ladder on the right is what the catalog generator enforces — for service-provider tenants it refuses shared-kernel grades outright, so the choice is never left to a sales conversation.
  • Isolation is a claim about the unit sold

    Each sellable unit names exactly one service class, one isolation mechanism, one metering basis and the tenant classes it may be sold to. A unit binds to machines by what the hardware publishes about itself, so one catalog serves a fleet of mixed generations without a unit landing on the wrong silicon.

  • Two tenant classes, one platform

    Enterprise tenants get every grade, chargeback against internal budgets and a quota ceiling. Service-provider tenants get only grades enforced in hardware or by the fabric, prepaid balance by default, and the full enforcement ladder. The API, the console, the lifecycle and the guarantees never change.

  • A usage feed for someone else's billing

    Metered usage is a continuous, replayable feed keyed on the same windows the platform's own invoices read, so a reseller can rate it in their own system. What they charge is theirs; what was consumed is one number.

Featured components

Five ways to consume the same fleet.

A machine, a guest, a cluster, a scheduler or a container. Each is an entry in the same catalog, carries one isolation mechanism and one metering basis, and is watched by the same lifecycle informer. Choose by what your work needs, not by which product your vendor sells.

Service classWhat you receiveIsolationMetered on
Bare metalA physical accelerator host provisioned to an operating system you choose, on your own network and storage, reclaimed with verified erasure. Power and break-fix are API calls.Sole occupancynobody else resident · returned with evidencenode-hours
Virtual machineA hardware-virtualized guest with its own kernel, accelerators passed through at the hardware level, a boot volume, a console and a power-state lifecycle. Snapshots and private images.No shared kernelnever a shared kernel with another tenantallocation-seconds
KubernetesA conformant control plane of your own, with accelerator node pools from the same catalog, your storage classes and your network, reached without a standing credential.Dedicated control planeper-tenant node pools · brokered accessnode-hours · per cluster
SLURMAn HPC cluster with login nodes, partitions, quality-of-service classes and fair share, whose accounting reconciles to the platform's own record. Keep your job scripts and module trees.Per-tenant clusterlogin nodes through the same gatewaynode-hours
Serverless containerA workload declared as an image, a command and a shape, placed on the grade the catalog resolved, reachable by shell, terminal or a published address. Scale to zero, wake on demand.Per gradewhole node to hardware partitionallocation-seconds · GPU-seconds · VRAM-GB-hours
Fig. 4The five service classes. Adding one never adds a second billing path, a second permission model or a second audit trail — which is what keeps the platform governable as it grows from one rack to many sites.
Positioning · as of Q4 2026

Against the GPU-cloud platform you have already been shown.

Rafay is the name most sovereign operators, neoclouds and systems integrators meet first in this lane: an NVIDIA-Certified delivery layer for NVIDIA Run:ai, sold as a closed platform. The comparison is about platform risk before it is about features.

The incumbent

Rafay — an NVIDIA-anchored delivery layer

  • Proprietary, closed-source platform; NVIDIA is its primary distribution partner
  • NVIDIA Run:ai is the payload, Rafay the delivery vehicle — a Rafay customer runs NVIDIA's orchestration through Rafay's layer
  • NVIDIA-Certified for HGX and NVL72 rack-scale systems; NVIDIA AI Cloud-Ready validated
  • Kubernetes, VMs, SLURM and inference services across data center, cloud, hybrid and air-gapped — under policy, quota and audit
  • Repositioned in 2026 around token-metered AI services on top of bare-metal GPU services

The objection to leaving is not feature parity. It is "why step off the path NVIDIA endorses" — and that is the question worth answering.

Bud GPU Foundry

An open control plane for any silicon

  • Open source under Apache 2.0 — the stack that governs your fleet can be audited, forked and kept
  • Silicon-neutral: NVIDIA, AMD, Intel, Qualcomm and Huawei in one catalog, mixed vendors and generations in one fleet
  • Bare metal, virtual machines, Kubernetes, SLURM and serverless from one catalog — then inference, models and tokens on the same stack
  • The layer above is already there: Bud AI Foundry serves models as a tenant, Bud SENTRY governs, Bud Agent runs agents — one AI operating system, not one control plane per vendor
  • Metering that starts when what you bought answers, with three records kept apart — asked, observed, billed

A stack welded to one vendor's orchestration cannot follow you onto AMD, Intel, Qualcomm or Huawei. A fleet that is silicon-neutral underneath keeps the choice of accelerator yours, generation after generation.

Where Rafay wins

If your fleet is NVIDIA end to end, you want NVIDIA Run:ai specifically, and the NVIDIA certification matters to your buyer, Rafay is the shorter path today. It has the longer list of neocloud and telco references, the NVIDIA channel behind it, and a managed-service motion Bud does not sell directly.

Where the comparison gets specific

Platform risk. The question to put to an NVIDIA-certified credential is what happens to the control plane when the next rack is not NVIDIA. GPU Foundry answers with an open stack you can audit and 600+ supported accelerator SKUs across five vendors.

What sits above. Rafay stops where tokens start. GPU Foundry is Layer 02 of Bud Novaria AI OS: models, routing, guardrails and agents run on the same fleet, under the same tenancy and the same audit trail.

Where you are starting from. Greenfield fleets run on GPU Foundry today. A migration path for Rafay-managed fleets is on the roadmap.

Comparison current as of Q4 2026 · Rafay facts from its own published materials · Talk to a solutions architect to go through it side by side

Go deeper

The full story, in depth.

The five service classes in full, the isolation grades and tenant classes, the seven lifecycle phases and the reason vocabulary, quota and placement, billing, the twenty-two things GPU Foundry will never do — and every headline number with its basis.

GPU Foundry is released under Apache 2.0. The public repository link will appear here at the open-source release.

Get started with Bud

Put your fleet on it.

The fastest way to see what a GPU-as-a-service control plane does for the accelerators you already own is a proof-of-concept on your infrastructure, with your tenants.

01 Register one cluster — any vendor, any generation — and publish a catalog.
02 Onboard two tenants with different isolation ceilings and watch the catalog differ.
03 Read the first invoice against the phase log. They cannot disagree.