NeFut Logo NeFut
中 Admin Login

[CS.AI] AMBER: Multi-View Adaptive Budget Allocation for Listwise Vision-Language Reranking

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #Neural

Vision‑language models (VLMs) serve as powerful listwise rerankers for multimodal retrieval, yet their high inference cost limits evaluation to a few local candidate views.\ \ Existing multi‑call strategies follow fixed schedules, wasting expensive VLM calls on uninformative candidate pairs or easy queries.\ \ We introduce Adaptive Multi‑view Budgeted Elo Reranking (AMBER), an online, budget‑constrained multi‑view reranking framework that dynamically optimizes global resource allocation. AMBER treats fragmented listwise VLM outputs as local tournaments and maintains a lightweight global ranking state via continuous Elo updates.\ \ Computation is allocated at two levels: dynamically constructing candidate views with high score ambiguity, and scheduling queries to maximize expected information gain. We show that each Elo update corresponds to a stochastic gradient ascent step on the Bradley‑Terry log‑likelihood, and we provide a submodular information‑theoretic justification for the query‑level allocation strategy.\ \ Experiments on CIRR, CIRCO, and PhotoBench demonstrate that AMBER achieves the strongest overall performance among compared multi‑call VLM reranking methods under comparable VLM‑call budgets, while remaining effective in low‑budget settings.\ \ The code is publicly available at https://github.com/wnlfc/AMBER\ \ Review: AMBER’s integration of Elo tournament dynamics with submodular information theory offers a fine‑grained, adaptive budgeting mechanism for VLM inference, presenting a reproducible pathway toward more efficient multimodal retrieval.

Original Source: https://arxiv.org/abs/2610.02831

[h] Back to Home