NeFut Logo NeFut
中 Admin Login

[CS.AI] From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Large‑language‑model (LLM) agents can generate and carry out actions, yet proposal, authority, dispatch, external effect verification, and service promotion are often treated as separate claims. Praxa makes these states explicit by means of deterministic admission, brokered execution, external read‑back, reconciliation, and reviewed promotion, forming a full execution pipeline. Four evidence lanes are reported. The first is an author‑run local repository audit on a pinned revision that passed all 1,027 unit tests and 89 Workerd tests, instrumented the expected 363 source files and satisfied four coverage thresholds; raw per‑test transcripts and independent reproductions are not available. The second lane is a provider‑backed Terminal‑Bench Core 0.1.1 pilot across 12 curated tasks, where baseline and reliability‑layer arms each succeeded in 17/36 strict trials. The reliability layer consumed 37.49% more input tokens and 50.73% more output tokens, providing no evidence of superiority. The third lane compares a post‑debug two‑order coordination‑proxy development: baseline and a source‑authored candidate both completed 180/180 trials with identical measured accuracy, full hermetic crash recovery, and zero protected violations. The candidate used 37.11% fewer tokens, 33.84% lower estimated endpoint cost, and 11.63% fewer steps, but this does not establish improvements in quality, latency, or production behavior. The fourth lane shows deployed source/configuration evidence of bounded reflection, recall accounting, memory compilation, and tool‑health paths, yet no lift in production outcomes is observed. Praxa’s main contribution is an evidence‑bound architecture that makes authority‑to‑effect transitions explicit and testable. Current evidence does not establish adversarial security, production safety, general specialist superiority, autonomous recursive optimization, or user benefit.

Review

Original Source: https://arxiv.org/abs/2610.00015

[h] Back to Home