Large language model (LLM) agents are moving beyond pure text generation to directly manipulate software systems, such as updating databases and invoking online services. Even when a database update is approved, the execution may produce unapproved notifications or other persistent side‑effects that downstream steps can mistakenly treat as successful.
Existing safeguards typically record the aftermath of a single action or approve an action before it runs, but they lack a unified check of all persistent changes that occur within a controlled execution boundary. To fill this gap we introduce EffectMatch, a runtime that captures every persistent modification inside a bounded execution context and compares it against the set of changes the application has explicitly approved for the current state.
EffectMatch operates in three stages: (1) intercept and log all writes to external persistent state during execution; (2) match the logged changes with the approved change set; (3) only when the match succeeds does it allow a commit and let dependent steps proceed, otherwise it aborts the commit and halts further execution. This guarantees that any unapproved persistent outcome cannot be silently accepted or propagated downstream.
In a comparative evaluation on 206 public business tasks, EffectMatch preserved all clean executions and blocked every tested incorrect commit. Six ablation studies, each run 20 times, revealed that removing any single mechanism (boundary interception, state matching, or commit control) caused the corresponding failures. Additionally, 80 task‑topology cases confirmed that EffectMatch maintains truthful handoffs while preventing invalid continuation.
These results demonstrate that EffectMatch effectively blocks the silent acceptance and downstream propagation of persistent outcomes that are inconsistent with application approval.
Review