NeFut Logo NeFut
中 Admin Login

[CS.AI] StudentBench: AI and Human Tutoring Yield Equivalent GRE Learning Gains

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

We introduce StudentBench, a suite of AI teaching evaluations and an open platform that has gathered over 175,000 student‑AI messages to examine whether large language models (LLMs) can achieve learning gains comparable to human tutoring. In Study 1, 2,383 participants were assigned to AI tutoring, human tutoring, or no tutoring and their score improvements on quantitative and verbal GRE items were measured. AI tutoring was statistically equivalent to expert human tutoring (p = .015), and in five of the seven GRE domains the top‑performing AI tutor outperformed the human tutor on average. Study 2 asked expert human tutors to perform 2,028 pairwise rubric evaluations of LLM‑generated lesson plans and practice problems, separating AI and human tutors across lesson planning, practice‑problem creation, conversational pedagogy, cost, and engagement. Remarkably, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918‑fold lower cost ($0.0052 per percentage‑point gain versus $4.81 for a human tutor). For quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with higher correct‑practice rates, and higher correct‑practice rates with larger learning gains (all p < .05).

Review: StudentBench offers a large‑scale, reproducible framework for evaluating AI tutoring, demonstrating that well‑tuned LLMs can deliver human‑level instructional outcomes in standardized test preparation at a fraction of the cost, opening new avenues for equitable and sustainable education.

Original Source: https://arxiv.org/abs/2609.28470

[h] Back to Home