As large language models (LLMs) are deployed on increasingly complex subjective tasks, aligning their intentions and behaviors with human values has become a critical scientific challenge. Current work, however, is hampered by a behavioral paradox: minor wording changes cause unpredictable “swing”, while explicit instructions to correct entrenched biases are met with “rigidity”. Resolving this duality is essential for reliable AI alignment.
To systematically understand and safely steer these latent subjective preferences, we focus on three core questions. First, do LLMs possess an intrinsic value system? By projecting responses from 106 LLMs (150 k queries per model) and 95 k human survey profiles into a shared sociological space, we empirically confirm that models exhibit value tendencies, yet their distribution is highly concentrated, lacking human diversity and forming an idealized value core. Second, how can these values be quantified? We introduce the Prior‑Environment‑Cognition (PEC) framework, mathematically defining value expression as the joint outcome of Prior (parameter weights), Environment (user prompts) and Cognition (chain‑of‑thought reasoning). Third, how can LLM values be aligned toward a target? Using PEC diagnostics, we devise an adaptive “Alignment Prescription” that avoids blind, resource‑intensive training and instead identifies the minimal effective intervention for each dimension, ranging from zero‑cost prompts to targeted parameter updates. Extensive experiments show that this approach validates the existence of LLM values, accurately quantifies their shifts, and achieves more efficient and precise steering than conventional blind training, all without degrading general capabilities.
Review