Large Language Models (LLMs) have achieved impressive results on many benchmarks, yet aligning them to dominant normative values often yields homogenized answers that ignore diverse user preferences. Existing training‑free approaches rely on prompt engineering, consuming valuable context windows, while training‑based methods become static after fine‑tuning and cannot support continual improvement in real‑world deployments. To tackle these issues, we introduce COPE (Continual Optimization with Personalized embedding and self‑Evaluation), a framework designed for interaction settings with sparse user feedback. COPE assigns a learnable personalized embedding to each user and integrates preference capture, self‑evaluation calibration, and personalized response optimization in a single update step. The key novelty is using the model’s self‑evaluation to produce proxy rewards, enabling ongoing updates even when explicit feedback is absent. Experiments show that under sparse feedback COPE consistently outperforms strong baselines—both prompt‑based and fine‑tuned—and remains complementary to Retrieval‑Augmented Prompting (RAP). Additional analyses confirm COPE’s reliable self‑evaluation, meaningful preference patterns, stable general capabilities, and robustness to shifting preferences and alternative evaluators.
Review