NeFut Logo NeFut
中 Admin Login

[CS.AI] Traversing the Satisfaction-Diversity Frontier in Text-to-Image Diffusion

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#algorithm #AI #Machine Learning

Text‑to‑image generation lets users obtain several images from the same prompt. For these images to be useful, each must reflect user preferences—measured by a learned reward—and be visually distinct to preserve diversity. Existing approaches either handle reward and diversity separately or merge them into a single aggregate score, allowing high diversity to compensate for low reward.

This paper formulates generation as satisficing: every candidate image must meet a reward floor, and the whole batch must satisfy a diversity cutoff. The reward floor controls the trade‑off between the worst‑candidate reward and batch diversity; varying the floor traces a Pareto frontier.

To traverse this frontier at inference time without extra training, the authors introduce SatisDive. SatisDive uses a batch‑relative reward cutoff to separate lower‑reward candidates (which are pushed upward) from higher‑reward ones (which are encouraged to stay diverse), thereby jointly improving reward and diversity within the same batch.

Experiments on the Pick‑a‑Pic benchmark show that, with FLUX.1‑dev as the base model and HPSv3 as the reward, SatisDive raises the worst‑candidate reward by up to 0.43 at matched DreamSim scores; with SANA‑1.6B and ImageReward, the improvement reaches 0.70. Across all overlapping DreamSim ranges, SatisDive’s satisfaction‑diversity curve Pareto‑dominates the FK steering curve in every setting.

Review: By explicitly defining a reward floor and a diversity cutoff, SatisDive decouples the two objectives while still optimizing them together. Its training‑free, inference‑only nature offers a practical way to achieve a controllable balance between quality and diversity in multi‑image generation.

Original Source: https://arxiv.org/abs/2610.02372

[h] Back to Home