NeFut Logo NeFut
中 Admin Login

[CS.AI] Open Pipeline and Dashboard for Systemic‑Risk Evidence under the EU AI Act Code of Practice

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#AI #LLM #Open Source

We introduce the Systemic Risk Index, an open evaluation pipeline and interactive dashboard designed to make empirical AI safety evidence transparent and traceable for the public. The framework organizes 19 public benchmarks into four systemic‑risk categories defined by the EU GPAI Code of Practice: CBRN (chemical, biological, radiological, nuclear), cyber offense, harmful manipulation, and loss of control. Models are evaluated using harm‑preserving perturbations within simulated deployment contexts.

The dashboard lets users switch between average and worst‑case aggregation, adjust how model capability influences the aggregate score, and trace each risk rating back to its benchmark evidence. Across 18 models, worst‑case aggregation reduces scores by 14–37 points compared to average aggregation, highlighting information that can be hidden by a simple mean assessment.

LLM judges achieve agreement with human graders comparable to human‑human agreement ($\kappa = 0.78\text{--}0.82$). A blind audit finds that $83\%$ of sampled transformations preserve the original harm. In a survey (N = 21), most participants reported that the scores are easy to understand and that the dashboard encouraged them to view model evaluations under different settings.

Review: By exposing benchmark data, offering interpretable aggregation options, and providing traceable risk sources, this platform delivers a reproducible, auditable tool for systemic‑risk assessment of AI systems, which is valuable for regulators and public oversight.

Original Source: https://arxiv.org/abs/2609.28335

[h] Back to Home