NeFut Logo NeFut
Admin Login

[CS.AI] Crypto Accounting Bench: Evaluating Frontier and Open-Weight Models on Crypto-Asset Accounting Tasks

Published at: 2026-09-16 22:00 Last updated: 2026-09-18 00:46
#AI #LLM #Open Source

We introduce Crypto Accounting Bench (CAB) as a benchmark to test whether frontier and open‑weight language models can reconstruct the exact journal entry an organization posted for a crypto‑asset transaction. CAB comprises 118 evaluation tasks drawn from seven pseudonymized organizations; each task blends transaction mechanics, asset quantities, base‑currency values, wallet and legal‑entity context, counterparty evidence, related legs, recurrence, tax‑lot evidence, and the organization’s full chart of accounts. The target is a balanced structured entry that includes every required account, side, amount, currency, and full‑precision asset quantity. We evaluate twelve models, running three independent attempts per task for a total of 4,248 trajectories. Three metrics are reported: Mean Score, Best@3, and Pass@3, where Pass@3 is the fraction of tasks with at least one of three attempts satisfying all rubric criteria and required gates. The leading model reaches a 77.43% Mean Score and the best Pass@3 is 56.78%. Deterministic diagnostics, taken from each task’s best of three attempts and macro‑averaged across models, show a base‑amount agreement of 97.8% but a deciding‑account accuracy of 56.3%. Failure analysis indicates that account selection and complete‑entry composition remain the main challenges on CAB.

Review: CAB offers a rigorous, real‑world testbed for fine‑grained accounting reasoning in LLMs, highlighting the need for stronger account inference and holistic entry construction in future model development.

Original Source: https://arxiv.org/abs/2609.14811

[h] Back to Home