NeFut Logo NeFut
中 Admin Login

[CS.AI] A Systematic Investigation of Bias in Large Language Models for Advertising Relevance

Published at: 2026-10-07 22:00 Last updated: 2026-10-08 01:25
#AI #Machine Learning #LLM

Large language models (LLMs) are increasingly employed to judge how well an advertisement matches a user query, yet the fairness of these judgments remains under‑explored. This work introduces a counterfactual framework to systematically assess the impact of advertiser identity, possible popularity, input language, and demographic wording on relevance judgments. We evaluate GPT‑4o as a categorical relevance judge and a Qwen‑7B model that has been fine‑tuned specifically for relevance prediction. The advertiser and language experiments draw query‑ad pairs from real advertising logs, while controlled synthetic queries are used to probe demographic associations in employment, housing, and credit domains. Results show that changing the advertiser identity or the input language can significantly alter relevance scores for both models; demographic comparisons reveal patterns consistent with common stereotypes, especially gender‑occupation links. Additional mitigation experiments indicate that the effectiveness of interventions depends on whether advertiser information is truly relevant to the query and on the distribution of advertiser labels in the training data. These findings help advertising practitioners spot fairness risks and guide the development of more equitable LLM‑based relevance systems.

Review: By combining real‑world logs with synthetic bias probes, the paper uncovers multi‑faceted bias in LLM‑driven ad relevance and offers practical inference‑ and training‑time mitigation strategies, providing a valuable roadmap for fair deployment.

Original Source: https://arxiv.org/abs/2610.07544

[h] Back to Home