NeFut Logo NeFut
中 Admin Login

[CS.AI] EMA: Elastic and Performance Transparent Memory Across GPUs

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#Machine Learning #LLM #Artificial Intelligence

Multi‑GPU servers are now the building block of modern data centers, offering aggregated capacity via high‑bandwidth interconnects. At the same time, workloads such as LLM inference have highly dynamic memory demands, often causing one GPU to run out of local memory while others stay under‑utilized. EMA (Elastic Memory Allocation) addresses this mismatch by introducing an elastic memory‑sharing model: GPUs within the same server can borrow and later reclaim memory from each other, forming a unified elastic pool.

For borrowers, EMA guarantees performance transparency: a prefetching layer hides remote‑memory latency, making remote and local memory indistinguishable to the application. For lenders, borrowed memory remains reclaimable on demand, ensuring that performance never drops below the static partitioning baseline. Although the design focuses on memory, the same elastic principle can be extended to other GPU resources such as bandwidth or compute units.

Evaluation shows that EMA improves individual user throughput by up to 52% (152% of the baseline), achieves 96% of the throughput of a system provisioned with double capacity, and keeps latency comparable to the static‑local baseline.

Review

Original Source: https://arxiv.org/abs/2609.27040

[h] Back to Home