Bud Simulator · Inside Bud AI Foundry · Sizing Engine

Size the cluster before you buy it.

A serving deployment fixes six things at once — chip, count, engine, parallelism split, quantization and batching. Bud Simulator predicts the whole space, discards everything that cannot meet your SLO, and returns the cheapest survivor. It answers the same question for training runs, and models CPUs as carefully as GPUs. Open source.

Overview

Sizing is a search problem, not a spreadsheet.

vLLM alone carries around 400 million configuration permutations for a single node — and the chip is a variable too, across 72 hardware profiles spanning GPUs, TPUs, ASICs and CPUs. Nobody benchmarks their way through that. Bud Simulator predicts it instead: a closed-form roofline scores every candidate, learned regressors correct the residual, and an event-driven serving simulator settles p99 on the finalists.

400M configurations per single-node deployment 72 hardware profiles — GPU, TPU, ASIC and CPU 112 pre-built workload profiles 6× better TCO on a hybrid deployment it sized
  • Beat one — where it sits. Bud Simulator is the sizing engine inside Bud AI Foundry, Layer 04 of the eight-layer Bud stack. It fixes six things that are only correct together: chip, count, engine, parallelism split, quantization and batching. Three predictors are stacked by cost and fidelity — an analytical roofline, learned regressors, an event-driven serving simulator. Feasibility is gated before cost. The output is the deployment: the runtime applies the configuration it returns.
  • Beat two — one prediction pipeline. A candidate configuration passes a memory fit check (weights plus KV cache against aggregate device memory), an analytical roofline (prefill bound by compute, decode by memory bandwidth), learned regressors fitted to measured runs, an event-driven serving simulator covering queueing and continuous batching, and a compatibility check across model, hardware and engine. It comes out sized — the cheapest configuration that meets the SLO.
  • Beat three — the collapse. Roughly thirty billion candidate configurations narrow under one query of use case plus SLO to three finalists on the Pareto frontier, and then to one plan that is deployed. All 72 hardware profiles resolve to a single sized deployment. Cost per million tokens improves six-fold against a single-tier GPU deployment at equal service level. Feasibility is gated first, so cost is only ever optimised among plans that already meet the objective.
  • Beat four — the result. Guesswork becomes a search result.
Value proposition

Four hundred million configurations become one plan.

The reason enterprises staff a GenAI systems engineering team is largely that somebody has to guess well, because nobody can measure exhaustively. A benchmark answers for one configuration on hardware you already own. Bud Simulator answers for the whole space, including the chips you are still being quoted — and it answers before the purchase order, not after.

Configurations · one node400M
Hardware profiles modelled72
Workload profiles, pre-built112
Memory accuracy±10%
Throughput accuracy±15%
permutation count per single-node vLLM deployment · accuracy bands and the 900+ test suite as published in the repository, validated against MLPerf Training, DeepSpeed ZeRO and Megatron-LM · method in the technical deep dive
Key features

Nine reasons the answer holds.

A sizing tool is only worth the confidence you can place in it. These nine are what separate a prediction from a guess with better typography — across serving, training, and the silicon underneath both.

01

Three predictors, one funnel

A closed-form roofline scores every candidate at effectively zero cost, learned regressors correct the residual on hardware with measured history, and an event-driven serving simulator settles p99 on the finalists.

3 tierscost and fidelity, stacked
02

Feasibility before cost

Weights plus KV cache at the target concurrency and context must fit in aggregate device memory. That single inequality yields the minimum chip count and discards most of the space before anything expensive runs.

Min chips requireda yes-or-no gate, computed first
03

A frontier, not a weighted score

Multi-objective genetic search over cost, latency and throughput returns the Pareto frontier — every configuration you cannot improve on one axis without giving up another. No weights you do not actually know.

NSGA-IIcost · latency · throughput, together
04

The SLO is an input

Chat, code completion, summarisation and agentic traffic want different batch sizes, different cache blocks and often different chips. A pre-built profile carries that shape into the search instead of one latency number.

112 profilesup to 4.48× goodput vs untuned
05

Hardware you do not own yet

The most valuable sizing question is asked before purchase. Every accelerator reduces to four numbers — compute, memory, bandwidth, interconnect — so a chip that is quoted but not racked can still be compared honestly.

72 profilesabsent hardware on a confidence ladder
06

Training, not just serving

14 training stages from SFT through DPO, PPO, GRPO and KTO; six fine-tuning methods; 30+ optimizers; TP, PP, DP, EP and ZeRO 0–3. Memory is the peak of the forward, backward and optimizer phases — not their sum.

14 stagescluster ranking · generated configs
07

CPUs modelled properly

Not a fallback with a fudge factor — an operator framework. AMX, AVX-512, AVX2, SVE and NEON throughput, L1–L3 cache and KV placement, NUMA topology, threading and frequency scaling under thermal limits.

ISA-awareXeon · EPYC · Grace · Graviton
08

BudEvolve — the search, inverted

Hold the workload fixed and sweep the hardware instead: FLOPS, bandwidth, memory and interconnect as free variables. Plus Morris sensitivity ranking, what-if curves, and LLM-driven evolution of scheduling and cache-eviction algorithms.

Spec, not quotedesign-space exploration
09

Open source, and auditable

Every efficiency constant is sourced from a datasheet, ISA spec, published benchmark or microbenchmark — no constants fitted until the answer looked right. 900+ tests; validated against MLPerf Training, DeepSpeed ZeRO and Megatron-LM.

±10% memory±15% throughput · ±20% training time
Go deeper

The full story, in depth.

The roofline, the search, the calibration — and the code behind all three.

Get started with Bud

Put your data on it.

The fastest way to see what an integrated AI operating system does for your enterprise is a proof-of-concept on your infrastructure, with your data.

01 Identify a use case where complexity, cost, or governance is a known pain point.
02 Joint discovery — Bud maps your AI pain points to platform capabilities.
03 POC in days, on your hardware, with your data.