Product 12 Feb 2025 3 min read

Introducing Bud Latent: High-Performance Inference Engine for Embeddings Models

A purpose-built embeddings inference engine designed to eliminate the high error rates competing engines hit at long context lengths — while running on virtually any hardware.

Bud Latent — high-performance inference engine for embeddings models

Bud Ecosystem has launched Bud Latent, a new inference engine designed to optimize embedding model performance at enterprise scale, with significant improvements over competing solutions.

Up to 90%
Faster inference vs. Hugging Face Text Embeddings Inference (TEI)
Up to 85%
Faster vs. the Infinity inference engine
<1%
Error rate at 8K tokens — vs. TEI's 94% and Infinity's 37%
16K
Token inputs processed where TEI crashed entirely

Performance and reliability

Bud Latent achieves up to 90% improvement in inference performance compared to Hugging Face's Text Embeddings Inference (TEI) and up to 85% improvement against the Infinity inference engine. Its error rate remains under 1% — substantially lower than TEI's 94% error rate and Infinity's 37% error rate when processing higher context lengths (8,000 tokens). During testing, TEI crashed entirely when handling 16,000-token inputs, while Bud Latent processed them successfully. Benchmarks used the gte-large-en-v1.5 model on Intel Xeon Platinum processors with 32 cores and 40GB memory.

Bud Latent inference performance benchmark vs. TEI and Infinity
Inference performance: Bud Latent vs. Hugging Face TEI and the Infinity engine.
Bud Latent error rate at long context lengths vs. competitors
Error rate by context length — Bud Latent stays under 1% where TEI and Infinity degrade sharply.

Core features and capabilities

Bud Latent supports diverse hardware platforms — NVIDIA CUDA, AMD ROCM, CPU, AWS Inferentia2, and Apple Silicon. It includes Flash Attention and Paged Attention optimizations, dynamic batching, and custom device-specific kernels.

  • Multimodal embeddings: text, image, and audio processing.
  • Precision options: both Int8 and FP8 formats for efficient large-model deployment.
  • Advanced CPU optimizations: AVX and AMX extensions, with NUMA node support.

Deployment and scaling

The engine integrates with Bud Simulator for automated configuration optimization to meet Service Level Objectives at minimal cost. It supports horizontal scaling across 16 different cloud platforms, enabling heterogeneous cluster deployments with automated hardware provisioning.

Cloud platforms supported by Bud Latent for horizontal scaling
Bud Latent scales horizontally across 16 supported cloud platforms.

It runs in cloud, on-premise, and client environments, supports Kubernetes and OpenShift production deployments, can launch multiple models simultaneously, and ships OpenTelemetry and Prometheus metrics support.

Enterprise applications

Use cases span AI agents, enterprise search and knowledge management, e-commerce personalization, financial fraud detection, and healthcare diagnostics. The sub-1% error rate positions Bud Latent as production-ready for reliability-critical applications.

Get started with Bud

Put your data on it.

The fastest way to see what an integrated AI operating system does for your enterprise is a proof-of-concept on your infrastructure, with your data.

01 Identify a use case where complexity, cost, or governance is a known pain point.
02 Joint discovery — Bud maps your AI pain points to platform capabilities.
03 POC in days, on your hardware, with your data.