The science behind the platform.
Papers, experiments, and perspectives from the Bud research team on making AI radically simpler, more efficient, and accessible to everyone.
Papers & experiments.
Peer-grade work spanning attention mechanisms, GPU virtualization, hybrid inference, datasets, and CPU acceleration. Each opens at its source.
Resource Aware Attention: a new attention mechanism designed for CPUs
We introduce Resource Aware Attention (RAA), a new attention mechanism designed from the ground up against the resource envelope of the deployment.
Read the paper (PDF)GPU-Virt-Bench: evaluating GPU virtualization methods
A comprehensive benchmarking framework that evaluates GPU virtualization systems across 56 performance metrics organized into 10 categories.
Read on arXivReward-based token modelling with selective cloud assistance
Reduces traffic to the cloud LLM — lowering cost — while allowing flexible, token-level control over response quality through reward-based routing.
Read on SciProfilesDataset for advancing academic knowledge and machine reasoning
11.53 billion tokens — integrating 8.01 billion tokens of synthetic data with 3.52 billion tokens of rich textbook data — to advance academic knowledge and machine reasoning.
Read on arXivInference acceleration for large language models on CPUs
We explore the utilization of CPUs for accelerating the inference of large language models — the foundation of the x86-first thesis.
Read on arXivPut your data on it.
The fastest way to see what an integrated AI operating system does for your enterprise is a proof-of-concept on your infrastructure, with your data.