Whether reasoning steps improve fairness in large language models (RLMs) has been hotly debated: do they mitigate bias or amplify it? Prior work reports conflicting findings. To clarify, we performed a within‑model ablation comparing thinking vs. non‑thinking modes on three 32B‑parameter models—QwQ‑32B, DeepSeek‑R1‑Distill‑Qwen‑32B, and Qwen3‑32B—across three high‑stakes decision tasks: Adult, COMPAS, and Credit.
Our experiments reveal an asymmetric dual effect of thinking on counterfactual fairness. Thinking resolves some counterfactual flips generated by the non‑thinking baseline ("resolved flips"), yet it also creates new flips when model confidence is near saturation ("created flips"). Across all nine (model, dataset) combinations, created flips outnumber resolved flips by roughly a factor of five.
To explain this phenomenon, we treat the thinking trace itself as a measurable site of fairness change and introduce two dynamic analysis tools:
- Counterfactual Depth Probability Gap (CDPG) – tracks bias evolution as a function of reasoning depth. Results show bias propagates and amplifies with deeper thinking.
- Bias Transition Matrix (BTM) – captures how predictions for counterfactual pairs transition from non‑thinking to thinking. The asymmetric dual effect originates from joint state transitions of the pair, where some pairs become unfair under thinking while others become fair.
In summary, thinking is not a universal remedy for fairness; at high confidence it can generate additional unfair flips. Future work should incorporate explicit fairness constraints into reasoning pathways to curb bias amplification.
Review