Before patients can rely on AI‑assisted psychiatric intake systems, health organizations need a practical routine to benchmark these tools against clinical standards for quality assurance. Because clinicians adopt diverse interviewing styles, the evaluation must satisfy three criteria: (1) enable cross‑style comparison, (2) minimize clinician workload, and (3) deliver performance metrics that matter to health systems deploying the technology. We built a clinician‑grounded evaluation platform around a memory‑augmented patient simulator called InterviewPlayground. Expert‑authored case vignettes were used to create interactive patients within InterviewPlayground, a simulated intake environment was assembled for open‑ended AI interviews, and evaluation modalities aligned with intake tasks were designed.
In a pilot involving six clinicians over a 25‑minute session, a GPT‑based large language model (LLM) acted as the intake interviewer. The LLM recovered 88.0% of clinically relevant items embedded in the vignettes, far surpassing clinicians at 38.9%. However, the LLM generated more clinical inferences not grounded in the interview (56.8% vs. 27.8%) and identified safety concerns less frequently (33.3% vs. 66.7%). These results establish a baseline for deployed quality assurance in this domain.
Review