NeFut Logo NeFut
中 Admin Login

[CS.AI] Completed Pairs Hide Capped Failures: A ReVerPi Case Study of Selective Context Projection

Published at: 2026-09-29 22:00 Last updated: 2026-09-30 01:41
#algorithm #AI #Machine Learning

The phenomenon that completed pairs conceal capped failures is illustrated in a ReVerPi case study of selective context projection. Context projection replaces older tool observations with compact, addressable excerpts, reducing repeated input while potentially adding evidence‑retrieval turns. In an 86‑run source‑reading campaign we issued 641 model requests, yielding 15 completed pairs; both arms succeeded on 12 out of 15 cases. The remaining 12 boundary runs were stopped, with the runner suppressing the companion arm whenever the first arm failed to finish. Restoring all 27 boundary runs bounds the projected‑minus‑full success difference between –9 and +1 tasks.

One omitted, selector‑chosen projected continuation successfully retrieved archive text but exhausted twelve requests; its full counterpart answered in three. Eleven jointly correct pairs form a fully observed success stratum: projection cuts aggregate logical tokens by 25 % while raising the median pair’s token count by 29 % and increasing total suffix requests from 35 to 55.

Separating fitting from evaluation breaks the selector’s apparent tie: outside its four fitting pairs it incurs one extra failure and consumes 8.6 % more logical tokens over thirteen comparable runs. This methodological case study links stopping rules, known bounded failures, unexecuted companions, and resource aggregation. The findings pertain to the recorded campaign rather than to population non‑inferiority or superiority over unrestricted Pi. Future evaluations should retain every intervention boundary, execute both allocated arms independently of the first arm’s completion, and report completion alongside interaction counts and token expenditure.

Review

Original Source: https://arxiv.org/abs/2609.31381

[h] Back to Home