NeFut Logo NeFut
中 Admin Login

[CS.AI] FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving

Published at: 2026-10-07 22:00 Last updated: 2026-10-08 01:25
#Machine Learning #optimization #LLM

Prefill‑decode disaggregation is now common in LLM serving because it separates the two phases with distinct execution patterns and SLO goals. Existing systems fix the prefill‑to‑decode worker ratio and route requests across workers. Real workloads, however, show short bursts and sustained shifts in demand ratio, so a configuration that fits at one moment quickly becomes mismatched, causing latency SLO violations even when idle capacity exists elsewhere. Conventional autoscaling adds capacity slowly, requires spare GPUs, and does not address short‑timescale phase imbalance. FluidPD introduces in‑place elasticity through two complementary mechanisms. FluidToken offloads a bounded portion of prefill work to decode workers when decode‑side slack is available, handling transient imbalance. FluidRole reassigns running workers between prefill and decode roles without model reload or engine restart, handling sustained imbalance. Both rely on lightweight pressure indices that expose resource pressure before it manifests as SLO violations. Experiments on production Azure traces show FluidPD improves overall SLO attainment over static SGLang by up to 94.6 percentage points, demonstrating that SLO‑aware in‑place elasticity can boost service quality without provisioning extra workers.

Review: FluidPD’s approach cleverly leverages existing resources to adapt to both short‑term spikes and long‑term shifts, offering a practical path to more efficient LLM serving.

Original Source: https://arxiv.org/abs/2610.06917

[h] Back to Home