NeFut Logo NeFut
Admin Login

[CS.AI] Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

Published at: 2026-09-07 22:00 Last updated: 2026-09-08 00:37
#AI #Machine Learning #LLM

Evaluating language‑model agents has become increasingly difficult as the number of agentic benchmarks grows and many of them require intricate environments and integration pipelines. To address this, we introduce Harbor Adapters, a unified evaluation infrastructure that ports over 80 existing benchmarks into a common adapter format, validated through thorough code review and parity experiments.\ \ Using this infrastructure, we conduct a large‑scale study of eight models spanning multiple capability tiers across 54 benchmarks. Each model is run with Terminus‑2 and with three native harnesses, enabling a broader analysis of agent capabilities and failure modes than previously possible.\ \ We also present Harbor-Index, a curated collection of 82 difficult, diverse, high‑quality tasks drawn from 29 benchmarks. The selection process involves difficulty filtering, AI and human audits, and an audit‑and‑fix loop, preserving both challenge and affordability. No model‑harness configuration exceeds a 30% pass rate, and the strongest configuration (GPT‑5.5 with Codex) reaches only 28.0%.\ \ All adapters, evaluation results, detailed analyses, and the Harbor‑Index are released as open‑source artifacts to support more reliable and comprehensive evaluation of language‑model agents.\ \ Review

Original Source: https://arxiv.org/abs/2609.04298

[h] Back to Home