Negation does not have a uniform interpretation across domains. In legal, regulatory, and medical reasoning, the intended reading depends on the semantics in force—open‑world vs. closed‑world, two‑valued vs. three‑valued, and credulous vs. skeptical inference. We investigate which negation semantics large language models (LLMs) adopt by default and whether they can override that preference when a different reading is explicitly specified.
To this end we introduce NAFBench, a procedural generator of solver‑certified instances spanning four semantic viewpoints: SLDNF, well‑founded semantics (WFS), and credulous and skeptical reasoning under stable‑model semantics. The generator emits ground normal logic programs with controlled depth, width, and cycle structure. Each program is solved with SWI‑Prolog, a WFS solver, and clingo, yielding up to four divergent labels. The programs are then verbalized into natural language under multiple framings and rule orderings that keep the answer invariant.
The results expose a consistent gap. Across open‑source models, following a specified negation semantics remains unsolved: the strongest models score 59%–74% across the four viewpoints, while the weakest score 31%–67%. More than half of logically identical rule shufflings cause order‑sensitivity, and the weaker models frequently over‑commit on WFS “undefined.” Two frontier models achieve 100% on the main fixed‑complexity evaluation set, and a third, o4‑mini, is near‑perfect, dropping only to 81% on WFS “undefined.” Delegating reasoning to a solver, fine‑tuning on certified traces, or forcing an explicit three‑valued verdict each partially closes the gap.
Review: This work provides the first systematic benchmark for LLM adherence to varied negation semantics, offering a reproducible suite (NAFBench) and highlighting substantial room for improvement in semantic consistency and order robustness. Future directions include deeper semantic alignment training and hybrid solver‑LLM pipelines to enhance reliability.