Personalizing large language models (LLMs) requires aligning generation behavior with user‑specific preferences rather than merely maximizing aggregate quality. Direct Preference Optimization (DPO) offers a stable framework for preference learning, yet its success in personalized settings hinges on how preference pairs are chosen. Existing methods often rely on heuristic criteria such as likelihood extremes, which decouple optimization from explicit user utility and can degrade personalization.\
We formalize personalized preference learning as a geometry‑aligned optimization problem by analyzing the first‑order interaction between the gradient of expected user utility and the DPO update direction. The analysis reveals that under off‑policy sampling, when preference margins are directionally aligned with utility gradients, the DPO update shifts from a purely error‑corrective signal to a reinforcement‑like update. This perspective treats pair selection as a geometric decision that determines whether preference optimization advances or hinders personalization.\
Motivated by this insight, we propose GAP‑DPO (Geometry‑Aligned Preference DPO), an iterative algorithm that performs utility‑aware, geometry‑aligned pair selection while controlling distribution shift via epoch‑wise regeneration. Experiments on personalized text generation benchmarks demonstrate that GAP‑DPO consistently improves stylistic fidelity, preference alignment, and generation quality compared to standard DPO variants.\
In summary, gradient alignment serves as a unifying principle for personalized preference optimization, and pair selection emerges as an intrinsic component of the optimization geometry rather than a heuristic preprocessing step.\
Review