x86 is all you need for AI democratisation

The accelerator shortage is a distribution problem, not a physics problem. Most inference traffic fits on the CPUs an enterprise already owns — at 210 W against roughly 2,910 W.

x86 is all you need for AI democratisation

Market landscape

  1. AI Everywhere: More than 200 Billion CPUs exist currently, making it the only viable option to create an AI-enabled world.
  2. ROI-positive AI solutions: Current AI agents and systems mostly fail to meet ROI requirements because of high Opex and Capex costs.
  3. Sustainability & Environment Concerns: The current power infrastructure is not adequate enough for generative AI market needs.
  4. Vertical models & Agents: Enterprises are currently looking for SLMs with SOTA performance for their downstream tasks. Enterprise vertical Agents powered by vertical models market are deemed to be 300$ billion in value.

Technology landscape

  1. Larger Parameter Size Does Not Equal Better Accuracy: Model sizes are reducing while their accuracy is getting better. Scaling law has also plateaued, the only way forward is smaller, efficient and performant models.
  2. More Efficient Model Architectures: New more efficient architectures like liquid models, SSM models, Bitnet models, Hymba, Jamba etc are making models more compute, memory, power and bandwidth-efficient, indicating an AI future on CPUs.
  3. Inference Time Optimizations for Accurate SLM Agents: Inference time optimisations like MOA, Swarm of models, Textgrad etc can drastically improve SLM based Agent accuracy & Performance.
  4. Hybrid LLM Architecture: Augmenting Cloud based LLMs with client based SLMs can drastically reduce generative AI agent cost by federating inference across cloud-client-edge.
  5. Collaborative Inferencing Methods: Augmenting LLMs with CPU-based SLMs can drastically reduce the cost of even LLM inference.
  6. No More Expensive Finetuning or Pretraining: Model merging methods have evolved to become as effective as finetuning for certain applications, these methods can easily be done on a CPU. Model merging and SOTA RAG methods will allow for easy model personalisation on a consumer-grade CPU.
  7. Goodput instead of Throughput: The industry is moving towards use case based SLO metrics rather than generic ones.
  8. Batch processing or non realtime: Most of the generative AI applications are not chatbots that require real-time processing. While Xeons could be used for even chatbot applications, it can exponentially better TCO & ROI for non-realtime or batched requests.

Bottom Line: The market is currently seeking ROI positive, power, memory and compute efficient generative AI based solutions and Agents that can work on the existing infrastructure. Current progress in data preparation, model training, model architecture, inference time optimizations and model merging methodologies allows for generative AI agents that can run consumer-grade CPUs at scale.

Why CPUs suit inference

  1. Less CapEx: Currently, data centres need to be enormously changed in order to accommodate the networking, cooling and power requirements of using GPUs or accelerators.
  2. Less OpEx: GPUs require exponentially higher power, cooling and maintenance requirements, which means increased OpEx costs.
  3. Better TCO or ROI: CPUs can reduce CapEx and OpEx costs improving the ROI and TCO exponentially at scale. (Approximate calculations given in references)
  4. Scalability: Currently, there is an industry-wide shortage of accelerators and GPUs, this is only getting worse over time because of the rising demand. Now, companies can build a generative AI application as a POC or beta application for 1000 users, but it is extremely difficult to scale it to 100,000 users, It is almost impossible for a startup or developer.
  5. Easy to Adopt & Maintain: Building & maintaining a GPU-based application requires specific expertise both from a hardware and software perspective. These resources are scarce and extremely expensive. Almost all of the software, network, hardware and data centre engineers currently are well-versed with CPU-based systems at scale.
  6. The barrier to entry: Currently, generative AI is inaccessible to SMEs, Startups or traditional enterprises – With an Nvidia H1100 or AMD MI300 costing up to 85,000 USD/month with CSPs, While an SPR device only costs between 500-2000 USD/month. This drastically reduces the barrier to entry while promoting more adoption, experimentation and evolution of the technology.

Bottom Line: Democratise generative AI by commoditising it.

What Bud runs on x86

Bud Ecosystem develops a universal runtime, inference stack and SOTA GPU-free SLMs, MLLMs, & Diffusion models that can provide production-ready inference on consumer-grade CPUs with SOTA accuracy.

We are building a generative AI deployment agent (Agent for creating generative AI Agents) that finds the best models, finds the best configs for deployment, optimises the prompts, and initiates/manages end-to-end generative AI observability through a simple chat interface, allowing anyone to build SOTA generative AI infrastructure just by providing a few examples of their downstream tasks. We believe that this easy-to-use agent combined with cost-effective SPR/CPU devices can make generative AI applications as simple (create/deploy/manage) and cost-effective as a typical web application.

Bud already have worked with multiple large organizations & government agencies to move their GPU-based generative AI agents and solutions to CPUs with their proprietary runtime and GPU free models.

Example Models and TCO

ModelContext PrecisionFaithfulnessAnswer RelevancyContext RecallAverage
GPT-4o0.9000.91580.90160.89523.6126
Bud RAG Jr 1.7B0.95450.85300.82740.93943.5743

TCO – Xeon vs Nvidia A100 (Bud)

FeaturesIntel Xeon basedNvidia A100
Power Consumption210 W~810 W
Throughput615 tokens/sec1462 tokens/sec
Power cost Per 1M token0.094 kWh0.15 kWh
Cost / 1M tokens0.181 USD0.203 USD

TCO – Xeons (Bud) + Bud Jr Model vs Nvidia A100 vs Open AI

Power Consumption210 W~2910 W
Throughput1269 tokens/sec333 tokens/sec
Power cost Per 1M token0.045 kWh2.42 kWh
Cost / 1M tokens0.087 USD4.53 USD15-20 USD
Get the next one by email

Product releases, benchmarks, and deployment patterns. Monthly, one email, unsubscribe any time.

Unsubscribe any time.

BN
Written by
Bud Newsroom
Bud Ecosystem
Get started with Bud

Put your data on it.

The fastest way to see what an integrated AI operating system does for your enterprise is a proof-of-concept on your infrastructure, with your data.

01 Identify a use case where complexity, cost, or governance is a known pain point.
02 Joint discovery — Bud maps your AI pain points to platform capabilities.
03 POC in days, on your hardware, with your data.