NeFut Logo NeFut
中 Admin Login

[CS.AI] Auditing Action Settlement in LLM Agent Environments: Order, Progress, and Replay

Published at: 2026-10-03 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

In large language model (LLM) agent environments, concurrent actions require arbitration even when each proposal is individually valid. We built a typed snapshot‑settlement contract and audited three properties: order sensitivity, useful progress, and replay consistency. Five settlement policies were evaluated across 28,800 exhaustive permutation trials and 2,160 scripted multistep episodes. Joint policies are spatially order‑invariant given fixed priorities, yet a conservative rejection completes only 31.25% of agents in a six‑agent doorway task compared to 90.28% for random tickets, a 59.03‑point improvement (95% bootstrap interval: 50.00‑68.06). All policies preserve spatial constraints, but priority arbitration still misses the independent small‑instance optimum. A separate full‑state journal audit exactly replays 156 checkpoints and rejects 1,332 constructed corruptions while retaining a terminal anchor. The evidence concerns execution semantics, not human realism or long‑run fairness.

Review

Original Source: https://arxiv.org/abs/2610.01138

[h] Back to Home