NeFut Logo NeFut
中 Admin Login

[CS.AI] ArgGYM: A Procedural, Engine-Verified Benchmark for Structured Defeasible Reasoning

Published at: 2026-10-01 22:00 Last updated: 2026-10-06 12:11
#Machine Learning #LLM #Artificial Intelligence

Recent advances in large language model reasoning have been driven by benchmarks and reinforcement‑learning environments that provide automatically verifiable rewards, especially in mathematics, code and formal logic. These settings simplify accuracy evaluation and optimization, yet it remains unclear how well success under fixed problem specifications and stable evaluation criteria transfers to broader reasoning contexts. Real‑world reasoning often proceeds with incomplete and revisable information: conclusions may be provisionally supported, defeated by counter‑evidence, reinstated by further arguments, or revised when stronger reasons emerge. This style is known as defeasible reasoning.

We introduce ArgGYM, a procedural benchmark and RLVR‑compatible training environment for structured defeasible reasoning. ArgGYM decomposes the reasoning process into twelve tasks and grounds task‑specific scoring in a symbolic argumentation engine that computes formal states for evaluating model outputs. The benchmark includes a frozen set of 1,440 verified instances across fifteen curriculum configurations, with two argument preference orderings (weakest‑link and last‑link) and two set orderings (elitist and democratic). The same generators and verifiers can also produce fresh instances for evaluation and verifiable‑reward training, reducing reliance on static test sets.

On the frozen benchmark, frontier models and open‑weight models exhibit sharply different reasoning profiles: they can recover substantial portions of structured answers without solving the full task, but performance declines in later curriculum stages that feature longer dependencies and more interacting structures. We release the benchmark, generators, and verifiers to enable reproducible evaluation and RLVR training.

Review

Original Source: https://arxiv.org/abs/2609.38409

[h] Back to Home