NeFut Logo NeFut
中 Admin Login

[CS.AI] Learning the Cost of Reliable Inference

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#AI #LLM

Benchmarking and routing platforms are increasingly acting as intermediaries between large language model (LLM) providers and end‑users. Existing platforms typically charge a fixed price per token, which prevents users from obtaining the most competitive rates for different workloads. To address this, we propose a procurement platform where token prices for each task are driven by competition among providers, allowing users to secure competitive pricing while guaranteeing quality.

The platform routes queries sequentially via a reverse second‑price auction, incentivizing model providers to truthfully bid their best estimate of the average cost to serve a query. As queries are routed, the platform learns the quality delivered by each provider and progressively directs queries to the most cost‑competitive provider that meets a desired quality threshold.

We validate the design with experiments on multiple LLMs from the \texttt{Llama} and \texttt{Qwen} families, using popular mathematical reasoning and question‑answering benchmarks. Results show that the pricing margin of the most cost‑competitive provider on our platform varies widely—between $10\%$ and $71\%$—depending on the task and quality threshold. This highlights a substantial inefficiency in the current fixed‑price market and demonstrates that our platform can enable users to capture maximum savings whenever competitive market conditions exist.

Review: The study introduces a market‑driven mechanism combining reverse second‑price auctions with quality learning, offering a practical approach to reduce LLM inference costs while maintaining reliability, and revealing significant savings potential in real‑world deployments.

Original Source: https://arxiv.org/abs/2609.28322

[h] Back to Home