Media Coverage
Lumai backs specialised hardware for AI prefill workloads
September 21, 2026
Lumai Editorial Team

Prefill Is the Bottleneck Nobody's Solving
The Prefill Bottleneck in Disaggregated LLM Inference

What Happens in Prefill?
Every request to a large language model runs through two very different phases: prefill and decode. Most infrastructure teams spend their time optimizing decode, the token-by-token generation everyone can see in a streaming response. Prefill gets less attention and for a growing share of workloads this is a mistake....