NeFut Logo NeFut
中 Admin Login

[CS.AI] JOVE: Joint Execution and Verification for Resource-Aware LLM Task Graphs

Published at: 2026-10-06 22:00 Last updated: 2026-10-08 01:25
#Machine Learning #optimization #LLM

Complex reasoning queries can be broken into directed acyclic task graphs and distributed across heterogeneous LLMs, which reduces latency via parallelism and enables smaller models to handle intricate tasks. In practice, the suitability of an LLM for a particular subtask is often unknown beforehand, and pure execution does not guarantee output correctness. We therefore introduce JOVE, an online framework that jointly assigns executor LLMs to subtasks and selects intermediate outputs for paid verification. Verification runs asynchronously; its feedback is used to improve future allocation decisions, forcing the system to balance current execution spending against future learning benefits. We study this trade‑off under a long‑term budget and a per‑query latency constraint, assuming stochastic and initially unknown LLM service quality, invocation costs, and execution times. JOVE makes execution and verification choices by solving a per‑query mixed‑integer linear program (MILP). An online learning component updates task‑dependent LLM quality estimates based on verification feedback, and an information‑gain bonus incorporates the value of learning into allocation decisions. Under natural assumptions we prove a sublinear quality‑learning regret for JOVE. Across four reasoning benchmarks, JOVE attains accuracy comparable to standard inference baselines while cutting average cost and latency by at least 3.17×.

Review

Original Source: https://arxiv.org/abs/2610.03296

[h] Back to Home