NeFut Logo NeFut
中 Admin Login

[CS.AI] BitNest: Bit-Nested Speculative Decoding for Memory-Efficient LLM Inference Acceleration

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#Machine Learning #optimization #LLM

Speculative decoding speeds up autoregressive generation by using a lightweight draft to propose multiple tokens for parallel verification. Existing methods typically require an extra draft model or separate weight representation, incurring notable memory overhead on resource‑constrained devices. Self‑speculative approaches reduce this overhead but still trade off draft quality, target quality, and storage efficiency.\ \ BitNest introduces a bit‑nested speculative decoding framework that embeds a low‑precision draft directly into the higher‑precision target representation. It first builds a strong low‑precision base and then recovers the high‑precision target via residual refinement, allowing both models to share a single physical weight storage. This progressive‑precision design is also extended to the KV cache for long‑context inference.\ \ Across several 7B‑8B edge‑friendly LLMs and diverse workloads, BitNest achieves an average speculative acceptance rate of 95.2%, closely preserving the quality of the higher‑precision model, and delivers a 1.48‑1.61× end‑to‑end speedup over FP16 autoregressive decoding. On LLaMA models covered by all representative self‑speculative baselines, BitNest consistently provides competitive or higher decoding speedups.\ \ Review

Original Source: https://arxiv.org/abs/2610.02800

[h] Back to Home