NeFut Logo NeFut
Admin Login

[CS.AI] Efficient Test-Time Adaptation through Human-AI Interaction

Published at: 2026-09-04 22:00 Last updated: 2026-09-05 12:23
#AI #Machine Learning #LLM

AI agents are trained on massive population data, endowing them with capabilities that span many practitioners. Yet their outputs often fall short of the personal standards professionals need to protect their reputation. In realistic, open‑ended tasks where success criteria are heterogeneous and under‑documented, individual expertise manifests as improvements and deviations from the average. In practice, iterative human‑agent interaction surfaces criteria that users cannot fully specify beforehand, yet apply repeatedly across tasks. We argue that this cross‑session interaction data is a rich, underused signal for bridging the gap to individual expertise. To this end, we propose Test‑time Adaptation through Human‑AI Interaction (TAHI), which injects interaction signals into the agent’s context and weights, and crystallizes each user’s training and evaluation criteria via an evolving rubric module. We adapt agents for 30 individuals across 600 tasks in two high‑utility domains: writing and visual creation. Results show that agents improve solo task success by 4.5%–20.9% after only tens of tasks. The evolving rubric serves as a scalable annotation tool, producing evaluation rubrics that catch 16.0%–22.3% more failures than those derived from LMs or humans alone. Although adaptation is personalized, these agents also yield up to 8.8% success gains that generalize across users.

Review: The study highlights the practical value of leveraging interaction signals for rapid personalization, offering a promising direction for real‑world AI deployment.

Original Source: https://arxiv.org/abs/2609.04141

[h] Back to Home