NeFut Logo NeFut
中 Admin Login

[CS.AI] Trajectory-Retrieval Speculative Decoding: When Does a Model's Own History Help?

Published at: 2026-10-07 22:00 Last updated: 2026-10-08 01:25
#AI #Machine Learning #LLM

Long chain‑of‑thought reasoning greatly increases sequential decoding cost while creating a growing history of potentially reusable continuations. We investigate when this history can provide useful drafts that complement an existing drafter. Controlled source comparisons reveal trajectory‑specific reuse, motivating our Trajectory‑Local Adaptive Retrieval (TLAR) method. TLAR retrieves approximately matched continuations from the current generation trajectory and uses recent verification outcomes to adapt retrieval activation thresholds and candidate width. Retrieved continuations are merged with model‑generated drafts in a shared candidate tree, preserving the target model’s output distribution through exact verification. We evaluate on code debugging, mathematical reasoning, and open‑ended writing, linking source reuse, incremental acceptance, and execution cost. Experiments show that, under matched verification budgets, TLAR combined with strong retrieval baselines improves token acceptance and increases end‑to‑end throughput compared to a draft‑model‑only baseline. These findings support treating generated trajectories as runtime memory for adaptive inference.

Review: TLAR demonstrates that a model’s own generated trajectory can serve as an effective runtime memory, reducing inference overhead while boosting throughput, making it a compelling approach for adaptive speculative decoding.

Original Source: https://arxiv.org/abs/2610.07350

[h] Back to Home