NeFut Logo NeFut
中 Admin Login

[CS.AI] Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#optimization #LLM #Artificial Intelligence

When serving demand for large language models exceeds training demand, serving efficiency becomes critical. Prefill‑Decode (P/D) disaggregation improves efficiency by specializing and isolating the two phases, but this benefit relies on a static partitioning of resources. In practice, phase demand is highly variable. In a large LLM fleet we observed that the ratio of uncached input to output tokens can reach a peak‑to‑mean of 4.7× on minute timescales, and in a public agentic trace the hourly ratio spans a median of 24.5× within a single day, while reassigning a replica takes tens of minutes. Agentic traffic sharpens this mismatch. Sizing each pool at the 95th percentile leaves up to 17% of cluster capacity idle; sizing below that creates queuing and unrealized throughput. Crossflow makes the boundary between the two phases elastic without changing node roles. Each decode node publishes a short‑lived, revocable lease that bounds local‑prefill compute, KV capacity, transfer work, and projected output. Across public and internal traces, Crossflow improves token throughput by 16.2‑17.4% on geometric mean over static P/D, and up to 43.4% at high load, while reducing mean time‑to‑first‑token (TTFT) at every evaluated point.

Review: By introducing a dynamic lease mechanism that keeps node responsibilities unchanged, Crossflow achieves adaptive scheduling of prefilling and decoding resources, markedly boosting efficiency and latency for agentic LLM serving.

Original Source: https://arxiv.org/abs/2609.27085

[h] Back to Home