Commercial text‑to‑image services silently rewrite user prompts before image generation, a step that users cannot disable or even see. Existing cultural‑bias audits examine only the final images and treat generation as a single pipeline, so they cannot pinpoint the bias source. To address this, we introduce WORLDVIEW, a multilingual benchmark containing 8,960 prompts across 15 languages and 31 language‑context pairings.
We audit the revision layer in three systems (DALL‑E‑3, Imagen‑4, GPT‑Image‑1.5) using a three‑step analysis: (1) how heavily the layer marks each cultural context, (2) whether it flattens the context into a narrow vocabulary, and (3) whether that vocabulary is stereotypical. Compared with a no‑context English baseline, the United States is the least‑marked context, while non‑Western and non‑Anglophone contexts are marked far more heavily, reduced to narrow vocabularies applied across diverse topics, and collapsed into recognizable cultural stereotypes.
By comparing images generated from original versus revised prompts on models without a revision layer, we identify the revision layer itself as a previously undocumented causal source of this stereotyping. To locate cultural bias and fix it, we must audit the system as deployed, not just the model.
Review