We introduce a modular framework that evaluates large language models’ dangerous capabilities under a unified protocol. The framework consists of three orthogonal pipelines—Knowledge ($K$), Defense ($D$), and Harm ($H$)—and aggregates their outputs into a standardized dangerous‑capability profile $\phi$. Pluggable modules supply scenario seeds, knowledge banks, hazard queries, and judging rubrics, while the core evaluation engine stays unchanged across domains; a chemical‑biological (CB) pilot demonstrates protocol transferability. Using a CB module, we assess 12 commercial LLMs from four families. Horizontal comparison reveals sharply divergent profiles: models with similar $K$ differ markedly in refusal resilience, and strong defenders do not necessarily generate less harmful content when they comply; family‑level patterns further separate Claude, DeepSeek, and GPT models. Temporal analysis shows that dangerous capability has not declined monotonically: newer models deepen $K$ while only partially improving $D$, indicating that scaling and alignment progress do not uniformly translate into safety. Reliability is established via cross‑judge consistency (bootstrap $\rho = 0.79$, 4 of 5 judges) and pipeline orthogonality ($K$‑$D$‑$H$ inter‑correlations $\rho \in [0.32, 0.52]$).
Review