NeFut Logo NeFut
中 Admin Login

[CS.AI] Bridging LLM Serving and CXL‑SSDs with Chunk‑Aware KV Cache Management

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#Machine Learning #LLM

LLM serving requires massive prefix caches, and NAND‑backed storage can provide the needed capacity. However, the block I/O path introduces CPU cache contention, host‑DRAM staging, and the inherent NAND latency. Even when DRAM replaces NAND as the storage medium, the interface overhead remains significant, motivating the use of CXL‑SSDs for byte‑addressable access to NAND‑backed capacity. Measurements show that a commodity CXL‑SSD is about three times slower than local DRAM and offers no advantage over an NVMe SSD, while generic prefetching yields little gain.

To address this, we propose LM‑CXD, a CXL‑SSD specialized for LLM prefix caching. LM‑CXD bridges the semantic gap between the serving engine, which knows which KV chunks will be consumed, and the device, which controls their placement and movement. It treats KV chunks as device‑visible I/O units, exposes the NAND‑to‑DRAM progress to the engine, and uses the device’s DRAM as a GPU‑accessible buffer. LM‑CXD also coordinates request scheduling with windowed prefetching and pipelines layer‑wise KV movement with GPU computation, hiding NAND latency under limited device DRAM.

Across five LLM models, LM‑CXD reduces average TTFT by up to 2.6× with compute‑asynchronous prefetching and 4.03× with layer‑wise prefetching, achieving TTFT within 1.5× of local DRAM on average.

Review

Original Source: https://arxiv.org/abs/2609.26828

[h] Back to Home