Large language models (LLMs) are increasingly used to guide urban safety decisions such as where it is safe to walk, rent, or travel. This study asks whether these judgments reflect measured risk or are driven by the stigma attached to a neighborhood’s name. We evaluate seven instruction‑tuned models on 186 neighborhoods in Los Angeles and Chicago under three input conditions: coordinates‑only, name‑only, and name + coordinates, linking model scores to violent crime statistics and American Community Survey data.
Results show that for six of the seven models, safety scores are nearly flat when only coordinates are provided, with noticeable variation only at frontier geographic scales. In contrast, name‑only inputs generate substantial between‑neighborhood variation and are moderately calibrated to violent crime rates.
A deeper analysis reveals that names lower safety ratings more for neighborhoods with higher shares of the locally dominant marginalized group (percent Black in Chicago, percent Hispanic in Los Angeles). This name effect tracks demographic share across all models and both cities. In Los Angeles, where demographic share and crime are more separable, the effect persists after controlling for crime and income and is confirmed using crime‑matched neighborhood pairs.
An enforcement‑elasticity analysis indicates that over‑caution aligns with near‑fully‑reported homicides rather than discretionary, deployment‑driven offenses. Moreover, the effect scales with geographic knowledge: models that better distinguish real neighborhoods apply stronger demographic stereotypes. Because neighborhood names convey both genuine crime signals and demographic stereotypes, removing names reduces both bias and accuracy.
The paper discusses implications for deploying LLMs in advisory and decision‑support contexts, urging careful handling of place names to avoid amplifying location‑based stigma.
Blogger's Review: The work spotlights hidden biases in LLM‑driven safety advice and calls for a balanced approach that mitigates stigma without sacrificing useful risk information.