NeFut Logo NeFut
Admin Login

[CS.AI] Winning by Silence: Non-Monotonic Deletion in LLM Plan Evaluation

Published at: 2026-07-16 22:00 Last updated: 2026-07-17 08:46
#optimization #LLM #Artificial Intelligence

In LLM-generated plan evaluation, evaluators can reward plans that strategically become less explicit. This paper investigates this failure in a staged expected-value scorer.\n\nProposition 1 presents the score change from deleting an interior transition while retargeting its predecessor and retaining downstream value: \n$$\Delta_k = (\prod_{iBlogger's Review: This study highlights the misleading scoring phenomena in LLM plan evaluation due to the omission of elements, emphasizing the need to consider potential incentive issues when designing evaluation mechanisms. It has significant implications for improving the transparency and reliability of AI systems.\n

Original Source: https://arxiv.org/abs/2607.12986

[h] Back to Home