Wrapping an image generation model with an agent that provides memory, skills, workflow orchestration, result verification, and iterative prompt refinement can markedly improve text‑to‑image performance. The agent continuously constructs and revises prompts to coax better images, yet these gains only manifest while the full harness runs; the diffusion model’s weights remain unchanged. To address this, we introduce Diffusion On‑Policy Context Distillation (D‑OPCD). The agent‑enhanced prompt is treated as privileged context, and the knowledge encoded in the agent harness is distilled into the diffusion model’s parameters, allowing the model to retain part of the harness’s benefit when conditioned solely on the original query. Using a text‑to‑image agent equipped with Auto Skill Evolver (ASE), D‑OPCD raises the average direct‑generation score from 60.52 to 65.09 across four benchmarks. Once this knowledge resides in the weights, the harness can drop its saturated skills and keep evolving: a second ASE round on the updated generator improves over a skill‑free harness by an additional 1.83 points, indicating a co‑evolutionary loop where harness and model continuously enhance each other.
Review