Artificial intelligence in recruitment has moved from automating simple profile pairs and ranked lists to multi‑stage workflows that retrieve evidence, compare candidates, and support or execute decisions. Using a purposive search and coding protocol (updated to 23 July 2026 and further to 2 September 2026), we systematically organized 40 representative academic, industrial, and legal works.
Three coupled transitions emerge:
- From similarity‑based matching to reciprocal suitability, where models assess not only how a candidate fits a job but also how the job meets the candidate’s needs;
- From a single model to a compound workflow, embedding models across retrieval, ranking, assessment, interviewing, sourcing, and human hand‑off stages;
- From offline prediction to evidence‑ and productivity‑aligned evaluation, expanding metrics beyond accuracy to include evidence acquisition, decision uncertainty, and cost‑risk considerations.
Across document understanding, retrieval, ranking, assessment, interviewing, sourcing, and hand‑off, we distinguish evidence at the field, pair, list, case, trajectory, and outcome levels. Persistent gaps remain: behavioral labels conflate exposure, preference, and qualification; private or synthetic data limit external validity; final‑output scores hide pipeline failures; and within the coded set privacy is never directly evaluated, nor is any study jointly assessing utility, fairness, privacy, and security.
In response, we introduce a staged mapping that links evaluation evidence to the strongest defensible claim and outline an agenda for reciprocal, evidence‑grounded, temporally controlled, selective, and auditable systems. Progress should be judged by whether workflows retrieve the right evidence, preserve uncertainty, support contestable decisions, and improve outcomes under explicit cost and risk constraints.
Review