NeFut Logo NeFut
中 Admin Login

[CS.AI] Acoustic-to-Text KV Compression for Full-Duplex Speech Models

Published at: 2026-09-30 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Full‑duplex speech language models continuously accumulate acoustic key‑value (KV) states during interaction, which makes long‑running sessions memory‑intensive. While listening, the model often finishes processing an audio unit before the next one arrives; we call this interval listening‑time slack. We introduce acoustic‑to‑text KV compression that exploits the slack by adding a transcription side‑channel, converting incoming speech into a compact textual memory. During inference, if the KV cache exceeds a target budget, older acoustic states are evicted while transcripts and the most recent acoustic context are retained. The side‑channel is fine‑tuned with LoRA using cross‑entropy on transcription segments. To preserve the original model’s listening and speaking behavior, we apply knowledge distillation on token‑level output distributions at native prediction positions. In ten‑minute LongSpeech sessions, our MiniCPM‑o 4.5 implementation reduces peak streaming KV‑cache size by 64.6 % compared with the same model without eviction. The method also improves transcription, temporal question answering, and summarization over the baseline. Full‑Duplex‑Bench evaluations show comparable pause handling, turn‑taking, and interruption performance.

Review

Original Source: https://arxiv.org/abs/2609.31224

[h] Back to Home