Abstract
Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards. We develop a methodology for deriving harmonized thresholds across three risk domains.
Risk Domain Analysis
- Misuse Risks: For cyber and biological risks, we take expected harm as the key primitive and use an explicit risk-modeling approach that accounts for risk channels and model release conditions.
- Automated AI R&D: Our proposed threshold is based on the observed rate of AI progress rather than expected harm.
Our analysis expands upon prior work and highlights existing empirical gaps and limitations.
Blogger's Review: This article provides a fresh perspective on risk management in AI through harmonizing safety thresholds, emphasizing the importance of industry standardization. This approach not only enhances safety but also promotes fair competition among companies, fostering a healthier development in the field.