NeFut Logo NeFut
Admin Login

[CS.AI] Multi-Axis Max@K Reinforcement Learning for Enhanced Diversity in T2I

Published at: 2026-07-18 22:00 Last updated: 2026-07-22 01:03
#AI #Machine Learning #optimization

Abstract

Text-to-image (T2I) models can synthesize realistic, prompt-aligned images, yet samples generated for the same prompt often cover only a small subset of visually distinct modes. This limits the diversity of images, and for person-centric prompts, can reflect or amplify demographic skew. We formalize this problem as coverage of a predefined set of semantically specified modes, which we call target-mode coverage.

We propose multi-axis max@K, a group-based reinforcement learning objective for improving such coverage in diffusion-based T2I models. Given a group of samples and one score per target category, multi-axis max@K first takes the maximum score across samples for each category and then sums these category-wise maxima. The resulting credit assignment gives a sample positive weight on a category only when it increases that category's group-wise maximum, allowing different samples to contribute to different categories.

We first validate the credit-assignment mechanism on a synthetic mixture and on SD3.5-M using deterministic pixel-based color rewards. We then evaluate the same objective on perceived-appearance fairness. Across three automatic evaluators on held-out prompts, multi-axis max@K improves the Fairness Score by 0.23-0.36 relative to the base model, while maintaining image quality and text alignment.

Blogger's Review: The proposed multi-axis max@K reinforcement learning method significantly enhances diversity in text-to-image generation, particularly addressing potential biases in person-centric prompts. By optimizing the category contributions of samples, this method improves fairness in generated results while maintaining image quality, showcasing its vital application prospects.

Original Source: https://arxiv.org/abs/2607.14962

[h] Back to Home