NeFut Logo NeFut
Admin Login

[CS.AI] From Atomic to Agentic: Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities

Published at: 2026-08-30 22:00 Last updated: 2026-09-01 02:31
#Math #LLM #Artificial Intelligence

Large Language Models (LLMs) are shifting from end‑to‑end mathematical reasoning toward agentic intelligence. Most existing math benchmarks only check the final answer, which offers limited diagnostic insight into step‑by‑step failures or logical rigor, and thus does not guide the transformation of LLMs into reliable mathematical agents. To address this gap, we introduce a process‑level benchmark that evaluates the intrinsic agentic reasoning abilities of LLMs. Our framework first defines a taxonomy of reusable atomic math capabilities and aligns problem‑solving behaviors with planning, action, and feedback stages. We then construct a comprehensive suite of tasks in both textual and multimodal settings, covering task decomposition, step execution, and result verification. An automated pipeline synthesizes high‑quality solution trajectories and employs controlled LLM rewriting to produce fine‑grained annotations for each step. Experiments reveal that models with comparable end‑to‑end accuracy can exhibit markedly different agentic capability profiles, underscoring the importance of process‑level evaluation for interpreting true potential and steering the development of next‑generation mathematical agents.

Blogger's Review: The paper convincingly argues that looking beyond final answers to the reasoning process is essential for truly assessing and advancing LLMs' mathematical intelligence.

Original Source: https://arxiv.org/abs/2608.26950

[h] Back to Home