In a compulsory first‑year university economics course, researchers used a two‑cohort difference‑in‑differences design to compare students who received AI‑generated study materials with those who did not. The AI resources—podcasts, FAQs, and quiz‑based guides—were produced by a source‑grounded model and then individually checked by a named graduate teaching assistant (170 students; 340 examination marks). Access to the verified AI materials yielded an average gain of 2.34 marks on a 50‑mark component. The share of students scoring below the upper‑second classification boundary fell by 24.7 percentage points. Significant effects appeared only between 23 and 31 marks, and roughly three‑quarters of the average gain originated from the bottom quintile of performers. When the lowest‑scoring students from the pre‑intervention cohort were removed, the threshold effect remained robust, whereas the overall average effect disappeared. Interviews with 36 students indicated that the “verified” label gave them a reason to engage with AI content while still scrutinising it. Evaluations that report only mean effects cannot reveal whether the resources truly help the intended student groups.
Review