NeFut Logo NeFut
Admin Login

[CS.AI] Risk Governance for Generative AI in Mental Health: A Breakthrough in Multi-Turn Safety Architecture

Published at: 2026-07-29 22:00 Last updated: 2026-07-30 03:24
#AI #Machine Learning #Open Source

Abstract

Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health risk. Existing safety approaches primarily detect risk but rarely shape how models respond as conversational risk unfolds. We developed a model-agnostic safety governance architecture that combines contextual risk detection, reasoning-based verification, and protocol-guided response generation for multi-turn mental health interactions.

Synthetic conversations grounded in real-world mental health narratives were used to evaluate the architecture's performance, tested with GPT-5-chat and Qwen3.5-27B, achieving high risk detection performance (specificity: 0.85 (95%CI: 0.78;0.91), sensitivity: 0.92 (95%CI: 0.88;0.95)) and increasing clinician-preferred escalation responses by 25.6--59.2pp while preserving rapport and connection. Performance remained stable across conversation length and generalized across both proprietary and open-source models. These findings demonstrate that clinically-grounded safety governance can extend beyond risk detection to improve how LLMs manage evolving mental health risk, providing a scalable framework for safer deployment across models.

Blogger's Review: This paper presents an innovative multi-turn safety architecture aimed at enhancing the safety of generative AI in mental health support. The focus on not only risk detection but also effective response to evolving risks offers significant implications for future AI applications in mental health. Its wide applicability and efficiency provide new perspectives and insights for related research.

Original Source: https://arxiv.org/abs/2607.22692

[h] Back to Home