As large language models (LLMs) become integral to software development, unintended bias in AI‑generated code has drawn attention. Evidence confirms such bias, yet systematic identification, categorization, and explanation remain underexplored. This paper introduces a taxonomy‑driven framework to detect and explain bias in generated code.
We expanded an existing biased Python code dataset and manually annotated each snippet with bias categories and human‑written justifications, creating a ground‑truth benchmark. Using this benchmark, we evaluated both proprietary and open‑source LLMs via in‑context learning (ICL), measuring classification performance and justification similarity.
Results show Gemini achieved 80.14% accuracy, 84.0% precision, and 95.7% recall. The best open‑source alternative, Qwen3‑coder, reached 82.45% accuracy, 68.64% precision, and 80.22% recall. For explanations, the models attained justification similarity scores of 80.4% and 80.14% relative to human reasoning, and code identification similarity scores of 86.0% and 87.82%, respectively. These findings indicate that LLMs can reliably detect biased logic in generated Python code and produce explanations closely aligned with expert interpretations.
Review