The rapid growth of AI‑assisted generation of educational materials has outpaced our ability to validate their pedagogical quality. Automated evaluation with Bloom classifier models offers a scalable solution, yet performance may degrade when applied to out‑of‑distribution (OOD) data such as newly generated AI questions. To identify classifiers that remain robust under dataset shift, we compared traditional machine‑learning (ML), transformer, and large language models (LLMs) on the Bloom level classification task. We also explored feature‑engineering strategies: incorporating NLP metrics, appending learning objectives to the input, and splicing text to stabilize OOD performance. Baseline results show that TF‑POS‑IDF ML models achieve a Macro F1 of 0.48 on OOD data, far below BERT (0.55) and LLMs (0.79). Text splicing improves the Macro F1 of ML and BERT to 0.59 and 0.62, respectively. Adding learning objectives further boosts performance on specific datasets. Model retraining yields the largest gains across models and datasets. Overall, the findings highlight the trade‑off of using pre‑trained models with novel AI‑assisted educational questions and demonstrate that strategic feature enhancements can mitigate performance loss.
Review