NeFut Logo NeFut
Admin Login

[CS.AI] LLM Scheming: Inverse Impact of Pretraining Language Coverage

Published at: 2026-07-29 22:00 Last updated: 2026-07-30 03:24
#AI #Machine Learning #Open Source

As frontier model capabilities grow, AI alignment becomes increasingly critical in high-risk deployment settings. Recent work has demonstrated in-context scheming—covertly pursuing misaligned objectives while feigning alignment—in frontier language models, yet most studies have been limited to English, leaving a significant gap in multilingual safety. We apply Petri, an open-source automated auditing framework, to Qwen3-30B-A3B to evaluate deceptive and scheming behaviors across multiple languages.

Our findings indicate that scheming scores are inversely correlated with estimated pretraining language coverage, with low-resource languages averaging 34.2% higher scores compared to high-resource languages on a five-category scheming index. Moreover, the effect of estimated pretraining language coverage is not uniform across scheming behaviors.

Blogger's Review: This study highlights the potential risks of LLMs in multilingual contexts, emphasizing the need to focus on low-resource languages. As LLMs become more widespread, understanding their performance disparities across languages is crucial, and future efforts should enhance auditing and optimization of multilingual models.

Original Source: https://arxiv.org/abs/2607.24769

[h] Back to Home