Fair allocation of scarce, indivisible resources is a central challenge in many societal problems. Although several formal fairness theories exist, none can satisfy all scenarios simultaneously. As large language models (LLMs) are increasingly employed to support decisions and act as agents, new concerns about distributive justice arise: the models' judgments are not directly tied to any specific fairness framework and may violate key normative principles.
In this work we introduce a general method for evaluating fairness reasoning in LLMs. We collect first‑person fairness judgments across a broad set of models and compare them directly with human responses under matched scenarios and elicitation conditions. The experiments reveal that LLMs tend to adopt stricter fairness constraints than humans, exhibit more self‑interested behavior in competitive resource settings, are highly sensitive to how information is framed, and are difficult to align with human judgments through fine‑tuning on existing datasets.
These findings suggest that simple fine‑tuning is insufficient to address fairness alignment in LLM‑driven resource allocation; systematic evaluation and calibration mechanisms are required.
Review