Home/Products/Bud Model Foundry
Bud Model Foundry · Layer 03 · Model Training

Train open models on the hardware you already own.

Build, fine-tune, post-train and agentic-train open models on your own infrastructure — research-grade control, production-grade operations, and no hardware tax. Bud DiLoCo trains across nodes over commodity Ethernet, so the SXM-and-InfiniBand dependency ends here.

Overview

The whole training spectrum, one platform.

Most platforms cover a slice of the training lifecycle. Bud Model Foundry covers all of it — data preparation, six training stages, step-level control, a built-in agentic-RL substrate, and a registry with lineage — behind one auth surface and one audit log, deployed inside your perimeter.

118+ supported open models 500× less inter-node bandwidth — typical multi-node case 4 GPU vendors — NVIDIA, AMD, Qualcomm, Intel 350+ platform REST APIs
  • Bud Model Foundry is Layer 03 of the eight-layer Bud stack — the training plane: the whole training spectrum on one sovereign platform, 118+ open models across 4 GPU vendors, deployed inside your perimeter.
  • Bud DiLoCo: conventional distributed training syncs gradients every step and demands 100+ Gbit/s InfiniBand; DiLoCo runs an inner AdamW loop per island and syncs one pseudo-gradient through an outer Nesterov optimizer every ~100 steps — 100–500× less inter-node bandwidth, standard Ethernet under 100 Mbit/s.
  • Six capability groups: full-spectrum training core; Bud DiLoCo; Bud Tinker; agentic RL & Simplified ART; data pipeline & flywheel; five interfaces.
  • The result: your models, your hardware, your perimeter — no hardware tax.
Value proposition

Full-spectrum training, on hardware you already own.

Everything between raw data and a registered, production-ready model — data preparation, six training stages, step-level control, a built-in agentic-RL substrate, and lineage — designed for sovereign deployment and engineered to perform on commodity hardware. No InfiniBand, no hosted dependency, no hardware tax.

Open models118+
Bandwidth reduction500×
Platform REST APIs350+
GPU vendors4
500× = inter-node, typical multi-node DiLoCo case · methodology in the product brief
Key features

Five commitments, built into the foundation.

Each one is a structural answer to a specific failure of the existing market — designed in, not bolted on. The sixth is the capability that breaks the hardware tax.

01

Commodity-Ethernet training

Bud DiLoCo islands train locally and sync one pseudo-gradient — no per-step gradient exchange, so the SXM-and-InfiniBand dependency ends here.

500× less bandwidthtypical multi-node · <100 Mbit/s sufficient
02

Sovereign by deployment

One-command install inside your perimeter — no outbound dependency, air-gapped as a first-class pattern, AES-256-GCM at rest.

30 minbud-install → running platform · no egress
03

Multi-vendor by design

NVIDIA, AMD, Qualcomm and Intel, with PCIe cards as first-class hardware — mixed fleets schedule into a single training job.

4 GPU vendorsPCIe first-class · mixed fleets, one job
04

Agentic-first by purpose

A full RL substrate ships in the box — and Simplified ART's teaching metaphor opens agentic training to domain experts, not just ML researchers.

10 graders3 modes · 4 environments · 5 recipes
05

End-to-end by scope

Data prep, training, RL, registry with lineage, and drift detection on one platform — every dataset version rebuildable, the loop closed inside the perimeter.

1 audit logdata → training → registry → drift
06

Production-grade from day one

bcrypt-hashed API keys, OAuth/OIDC, model-level RBAC, atomic quotas, four rate-limiting algorithms, structured audit, Prometheus metrics.

350+ REST APIsRBAC · quotas · audit · 5 interfaces
Featured components

What's inside the foundry.

The training plane is built from named engines — the distributed-training orchestrator, the step-level control surface, and the substrate that trains agents.

Distributed training

Bud DiLoCo

An inner AdamW loop per island, one pseudo-gradient through an outer Nesterov optimizer every ~100 steps — same total step count, a fraction of the traffic.

100–500× less inter-node traffic
Training control

Bud Tinker

The eight primitive operations of a training loop as REST + SDK — full training state preserved at every call, for custom RL and step-level debugging.

8 primitives · bit-exact pause/resume
Agentic training

Simplified ART

Student, coach, curriculum, grader — a teaching metaphor that compiles to a full RL run, with pre-built recipes for reasoning, code, support, tool use, and safety.

5 recipes · plateau auto-stop
Go deeper

The full story, in depth.

The training matrix in full, the DiLoCo mechanics, the Tinker primitives, the ART pipeline — and every headline number paired with how it was measured.

Get started with Bud

Put your data on it.

The fastest way to see what an integrated AI operating system does for your enterprise is a proof-of-concept on your infrastructure, with your data.

01 Identify a use case where complexity, cost, or governance is a known pain point.
02 Joint discovery — Bud maps your AI pain points to platform capabilities.
03 POC in days, on your hardware, with your data.