Home/Products/Bud SENTRY/Bud Sentinel
Bud Sentinel · Inside Bud SENTRY · Guardrail Engine

Guardrails at the speed of the request.

Jailbreak detection, prompt-injection defence, and content moderation in single-digit milliseconds on commodity CPUs — powered by Resource Aware Attention, and built into Bud SENTRY.

Overview

A different curve, not another compression pass.

Safety classifiers run on every request — on the CPU fleets the application already runs on. Sentinel is designed for that envelope from the start: an attention mechanism shaped by the cache hierarchy, precision tier, and latency SLO of its deployment target, instead of a GPU transformer compressed until it fits.

0.70ms p50 at 10K concurrent connections 678× faster than baselines on the same silicon <20% ASR and FRR — the only evaluated model under both ~$0.10 per 1M classifications vs. ~$24 on GPU
  • Beat one — where it sits. Bud Sentinel is the guardrail engine inside Bud SENTRY, Layer 06 of the eight-layer Bud stack. CPU-first by design through Resource Aware Attention — a mechanism, not compression. Input guard and output guard on every inference call. 23 specialised models, 33 variants, one binary via the Bud Guardrail Gateway. Wrapped by SENTRY: policy, identity, FinOps, one immutable audit trace.
  • Beat two — one attention pass. A request rides a single Resource Aware Attention pass through five task heads: jailbreak (ASR 15.97%, FRR 14.92%), prompt injection (retrieved context vetted), content safety (7 categories, both edges), PII (11 regions, redaction), and long context (65,536 tokens native). Cleared at 0.70ms p50.
  • Beat three — inside the engine. Resource Aware Attention: the envelope is an input, no quadratic phase change, many heads over one pass. The deployable operating point: the only evaluated model with attack-success and false-refusal rates both under 20%. Full threat coverage: 23 models, 33 variants, 4.5M labelled samples. Long context, native: 65,536 tokens at 560ms p50 where baselines cap at 512. Production throughput: 4,400+ req/s on one Xeon node, p99 under 12ms, ~$0.10 per 1M classifications vs ~$24 on GPU. Faster than a $15,000 GPU: 8.39ms on a laptop CPU versus 18–19ms for every baseline on an A100 — 2.3× faster, air-gapped and sovereign.
  • Beat four — the result. Your laptop CPU beats a $15,000 GPU.
Value proposition

Hundreds of milliseconds become single digits.

The guardrail category sits at 334–3,855ms per classification on the server CPUs operators actually run, capped at 512 tokens. Sentinel is a different curve: single-digit milliseconds on the same silicon, 65,536 tokens natively, and the only evaluated operating point a product team can actually ship.

p50 latency · CPU0.70ms
ASR & FRR, both<20%
Req/s, one Xeon node4,400+
Tokens, natively65,536
measured end-to-end over the serving interface · methodology in the product brief
Key features

A different curve, honestly drawn.

The category forces a choice between guardrails that miss most attacks and guardrails that block most benign traffic — or a GPU in front of every classifier. These six are why Sentinel doesn't.

01

The deployable operating point

The only evaluated model with attack-success and false-refusal rates both under 20% — rivals with lower ASR refuse 82–89% of benign traffic.

<20% ASR & FRR15.97% / 14.92% · only occupant
02

Resource Aware Attention

A mechanism designed against the deployment envelope — cache, precision, and latency SLO declared before training — not a GPU transformer compressed until it fits.

678× fastervs transformer guards · same silicon
03

Long context, native

Transcripts, RAG context, documents, and code classified as one call — no quadratic phase change with length, no chunk-and-vote pipeline.

65,536 tokens560ms p50 · baselines cap at 512
04

Production throughput

A classifier on every request that will never be the bottleneck — sustained on one Xeon node at 512 tokens with p99 under 12ms.

4,400+ req/sone node · ~$0.10 per 1M classifications
05

Faster than a $15,000 GPU

8.39ms on a laptop CPU beats every baseline running on an A100 — so guardrails ship inside desktop apps, IDEs, and embedded agents with no server call.

2.3× fasterlaptop CPU vs A100 baselines · ~1,500 req/s fanless
06

One pass, many heads

Safety decision, PII spans, intent, and routing ride a single attention pass — a chain of five classifiers collapses into one shared cost.

1 attention passthe marginal head is nearly free
Go deeper

The full story, in depth.

The RAA architecture, the layered guardrail, and every benchmark table behind the claims.

Get started with Bud

Put your data on it.

The fastest way to see what an integrated AI operating system does for your enterprise is a proof-of-concept on your infrastructure, with your data.

01 Identify a use case where complexity, cost, or governance is a known pain point.
02 Joint discovery — Bud maps your AI pain points to platform capabilities.
03 POC in days, on your hardware, with your data.