NeFut Logo NeFut
Admin Login

[CS.AI] Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models

Published at: 2026-09-21 22:00 Last updated: 2026-09-22 02:29
#AI #Machine Learning #LLM

We created a 12,000‑instance dataset of two‑option dilemmas that capture three pairwise value conflicts: Honesty vs. Justice, Justice vs. Autonomy, and Autonomy vs. Honesty. Each instance was translated into Hindi, Arabic, Spanish and Chinese to probe cross‑lingual behavior. Benchmarking GPT‑5‑mini revealed a consistent preference for Honesty over Autonomy across all five languages when no policy is supplied. Llama‑3.2‑1/3B models showed a strong first‑option bias, but both standard fine‑tuning and Direct Preference Optimization (DPO) fine‑tuning removed this bias, raising accuracy above 98%. To separate learned dataset correlations from abstract values, we introduced a task‑vector transfer experiment: after computing the task vector for a given value direction, we orthogonalize it against the general instruction‑following vector. Results demonstrate that this technique isolates the specific value preference direction and can be used in task arithmetic to obtain a model with the opposite stance.

Review

Original Source: https://arxiv.org/abs/2609.21094

[h] Back to Home