NeFut Logo NeFut
Admin Login

[CS.AI] Replication, Measurement Sensitivity, and Persistence in Hosted LLM Evaluations

Published at: 2026-09-23 22:00 Last updated: 2026-09-24 00:40
#Machine Learning #LLM #Artificial Intelligence

Behavioural evaluations of hosted language models can diverge because the service, the measurement instrument, or both differ across runs. We therefore distinguish three validation questions: (1) does a prior finding recur on fresh data under its historical configuration (replication); (2) does the endpoint change when the evaluation‑inference configuration is rebuilt under the same identifier (measurement sensitivity); and (3) does the finding persist across subsequently tested identifiers using a common instrument (persistence).

Regent Chess is a sequential environment that records a hidden mutable state exactly, allowing stated model beliefs to be scored against ground truth at action time; positive endpoint values indicate worse performance than a matched‑uniform comparator.

In the replication study, the previously reported Gemini 3.1 Flash‑Lite deficit reappears on fresh games under its historical configuration, with an endpoint of +0.0530 and a 95% confidence interval of [+0.0329, +0.0714].

In a back‑to‑back same‑day H/R comparison under the same public identifier, the model‑minus‑uniform endpoint is 0.0429 lower under the rebuilt configuration, 95% CI [+0.0182, +0.0667]. All six configuration components vary jointly, so no single component is isolated.

For persistence, under the rebuilt R configuration, a prospectively frozen, interleaved same‑window 4K comparison reverses sign between Gemini 3.1 and Gemini 3.7—identifiers that differ in release and product tier. Additional descriptive and exploratory cells show the same directional pattern. Any extra serving‑period contribution remains unresolved at -0.0166, CI [-0.0483, +0.0157].

Thus replication, measurement sensitivity, and persistence can lead to different conclusions within a single evaluation, motivating explicit indexing of hosted‑model behavioural claims by tested identifier, serving period, measurement instrument, and inference configuration.

Review

Original Source: https://arxiv.org/abs/2609.22478

[h] Back to Home