NeFut Logo NeFut
中 Admin Login

[CS.AI] DynBranch: Speculative Subgraph Reuse for Dynamic Agentic LLM Serving

Published at: 2026-09-29 22:00 Last updated: 2026-09-30 01:41
#AI #optimization #LLM

Agentic LLM workflows decide their execution paths only at runtime. Downstream computation may be predictable or even have been executed before, but it cannot start until the model or the user resolves the branch, creating a branch‑resolution barrier. Conventional caching cannot hide this barrier because the key that identifies a reusable result is unknown until the branch resolves.

DynBranch’s key idea is to make an unresolved branch addressable before it resolves. It assigns each potential subgraph a stable coordinate, allowing candidate subgraphs to be launched speculatively during branch resolution and their completed results to be reused by later requests.

To keep resource usage in check, DynBranch employs a two‑level controller:

  1. Benefit predictor estimates the latency gain from reusing a subgraph;
  2. Load price estimator computes the cost of executing the subgraph under current system load. A subgraph is scheduled only when the expected benefit exceeds its cost.

DynBranch sits at the model‑API boundary and requires no changes to agent harnesses or model execution engines. Experiments on four Qwen3‑32B‑based agentic workloads using 4× H200 GPUs show that DynBranch reduces mean latency by up to 32% compared with each workload’s strongest prior system, and by 46%–66% against a no‑reuse baseline, while preserving workflow outputs. The advantage persists across backbone families (e.g., Qwen3‑8B) and on commodity hardware such as RTX 4090.

Review: By assigning reusable coordinates to unresolved branches, DynBranch enables speculative execution of potential computations, cutting end‑to‑end latency without sacrificing correctness. This approach offers a practical caching‑plus‑scheduling strategy for large‑model agentic services.

Original Source: https://arxiv.org/abs/2609.31047

[h] Back to Home