NeFut Logo NeFut
Admin Login

[CS.AI] CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition

Published at: 2026-09-21 22:00 Last updated: 2026-09-22 02:29
#AI #Machine Learning #Artificial Intelligence

CogGym is a scalable, unified framework grounded in cognitive science for systematically comparing model and human behavior on matched experimental trials. The system uses a semi‑automated, human‑in‑the‑loop pipeline to convert diverse experimental paradigms into a task‑agnostic Experiment Markup Language (EML), enabling reproducible and faithful large‑scale comparison.

In the initial release we curated and standardized 258 cognitive experiments from 100 papers, focusing on human commonsense reasoning, and evaluated 50 large language models on these tasks. Larger and newer models better reproduce human judgments, yet their improvement is markedly slower than gains observed on formal‑reasoning benchmarks such as math and coding. The best models achieve $R^2 = 0.59$ on text, $0.58$ on image, and $0.43$ on video experiments, compared with human split‑half reliability of $R^2 = 0.93$ (text), $0.95$ (image), and $0.92$ (video).

CogGym is intended as a living evaluation platform that will continuously incorporate new cognitive‑science experiments to map where model behavior aligns with or diverges from human behavior, and how these patterns evolve as models and experiments advance.

Review

Original Source: https://arxiv.org/abs/2609.21259

[h] Back to Home