NeFut Logo NeFut
中 Admin Login

[CS.AI] MASCRDM: Multi-Agent System for Compliance Risk Detection and Mitigation in LLM Training

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Large language models (LLMs) are deployed across many domains, yet ensuring their compliance and safety—especially avoiding discrimination and bias—remains challenging. Existing work mainly filters inputs and outputs after training, ignoring real‑time monitoring of the model’s internal architecture. Our analysis of the LLM training pipeline reveals two critical gaps: first, detection methods are largely static, offering only localized optimizations; second, there is no continuous risk detection and mitigation throughout training, limiting flexibility. To address these issues we introduce MASCRDM (Multi‑Agent System for Compliance Risk Detection and Mitigation), which embeds multiple agents into the training process. We first derive compliance rules from AI legislation and train a compliance‑focused LLM under the guidance of legal experts. Then we decompose the target LLM into components and use a compliance knowledge graph to identify key nodes. During training, the agents monitor these nodes, issue real‑time risk alerts, and suggest corrective actions to developers. Benchmarks on discrimination and bias show that MASCRDM improves compliance while preserving reasonable semantic performance, providing a systematic, internal pathway for mitigating compliance risk in LLMs.

Review

Original Source: https://arxiv.org/abs/2609.39107

[h] Back to Home