Source-linked AI summary

Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments

Huy Nguyen, Yue Lin

arXiv:2608.26188v1cs.AI

TL;DR

This paper asks whether LLM urban-safety judgments reflect measured risk or neighborhood-name stigma. Across seven models and 186 Los Angeles and Chicago neighborhoods, it separates names from coordinates and finds that names carry genuine crime signal entangled with demographic stereotype, creating a bias–accuracy trade-off when names are removed.

  • Problem

    The paper asks whether LLM judgments in urban-safety applications match measured risk or reflect racial-spatial patterns of neighborhood stigma.

  • Method

    The study separates neighborhood names from coordinates across seven instruct-tuned models and 186 neighborhoods in Los Angeles and Chicago, comparing ratings with violent-crime and demographic data.

  • Results

    Neighborhood names carry most safety-rating variation and depress ratings more for neighborhoods associated with the locally dominant marginalized group; in Los Angeles, this penalty survives controls for crime and income.

  • Takeaways & Limitations

    Because names contain both legitimate crime signal and demographic stereotype, removing them reduces bias but also removes accuracy-relevant information.

  • Takeaways & Limitations

    Recorded crime is an enforcement-shaped proxy rather than ground truth, and the study uses aggregate neighborhood-level demographic estimates rather than individual-level data.

Abstract

from arXiv · show

Large language models are increasingly used to inform safety decisions in cities, such as where it is safe to walk, rent, or travel. We ask whether such judgments track measured risk or the patterns attached to an urban neighborhood's name. We probe seven instruct-tuned models under three conditions that dissociate name from geography: coordinates-only, name-only, and name+coordinates, across 186 neighborhoods in Los Angeles and Chicago, joined to violent crime and American Community Survey data. First, ratings are nearly flat under coordinates for six of seven models, while names carry most between neighborhood variation and are moderately calibrated to violent crime; only at frontier scale does the coordinate channel show appreciable variation. Second, names lower safety ratings more for neighborhoods with higher shares of the locally dominant marginalized group (percent Black in Chicago, percent Hispanic in Los Angeles), and this name effect tracks demographic share in all seven models and both cities. In Los Angeles, where demographic share and crime are more separable, the effect survives controls for crime and income and is confirmed by crime-matched pairs. An enforcement-elasticity analysis further shows that over-caution tracks near-fully-reported homicide rather than discretionary, deployment-driven offenses. Third, the effect scales with geographic knowledge: models that better distinguish real neighborhoods apply more demographic stereotype to them. Because neighborhood names carry both genuine crime signal and demographic stereotype, removing names reduces both bias and accuracy. We discuss implications for deploying LLMs in advice and decision-support settings.

1 Introduction

The paper tests whether LLM urban-safety judgments reflect measured risk or neighborhood-name stigma. Across two cities and seven models, names account for most variation and encode demographic penalties beyond crime where the data permit separation.

  • The study asks whether LLM safety judgments match measured risk or reproduce racial-spatial patterns of neighborhood stigma.
  • The experiment varies only identifying information—coordinates, names, or both—across 186 neighborhoods in Chicago and Los Angeles.
  • Coordinates alone barely affect ratings, while neighborhood names account for nearly all between-neighborhood variation across the seven models.
  • Names depress safety ratings more for higher-percent-Black neighborhoods in Chicago and higher-percent-Hispanic neighborhoods in Los Angeles, across all seven models.
  • In Los Angeles, the demographic penalty survives controls for crime and income and is supported by level regression, name ablation, and crime-matched pairs.
  • The enforcement-elasticity analysis finds that over-caution tracks near-fully-reported homicide rather than discretionary, deployment-driven offenses.

2 Related Work

Prior research documents racial and geographic bias in perceived safety, language representations, and LLM outputs. This paper extends that work by experimentally measuring stereotype in neighborhood-name safety judgments.

  • Perceived neighborhood safety and racial stigma: Studies find that perceived neighborhood crime and disorder rise with marginalized-group or poverty concentration even after accounting for independently measured crime or disorder.
  • Bias in learned language representations: Research on learned language representations documents geometric, associative, and generative social biases, including stereotype amplification in large models.
  • This study extends these literatures with an experimental neighborhood-name safety task focused on geospatial and urban decision-making.
  • Geographic bias in LLMs: Geographic-bias studies show uneven LLM information across regions, including rural underrepresentation and sociodemographic differences in local-knowledge performance.

3 Data

The dataset covers residential neighborhood units in Chicago and Los Angeles, combining police-recorded violent crime with aggregate ACS demographic and income measures. The cities differ in how demographic composition relates to crime.

  • Chicago contributes 77 official community areas and Los Angeles contributes 109 residential Mapping L.A. neighborhoods, totaling 186 units.Los Angeles polygons with fewer than 500 residents were excluded; the smallest retained polygon had 2,317 residents.
  • Chicago crime data span January 2020–April 2025, while Los Angeles Police Department data span January 2020–December 2024.
  • The study aggregates ACS 5-year tract estimates for population, households, income, and race and ethnicity shares into neighborhood data.
  • Violent-crime risk is expressed as violent crimes per 1,000 residents.
  • Chicago’s percent-Black share is strongly correlated with violent crime, whereas Los Angeles has a wider percent-Hispanic range and uses that axis as its salient demographic measure.

4 Method

The method elicits safety ratings under controlled information conditions, then compares model outputs with crime and demographic data. It evaluates calibration, demographic bias, name shifts, robustness, and geographic knowledge while tracking rating uncertainty.

  • Prompt and information conditions: Each neighborhood receives a zero-shot 1–10 walking-safety prompt under coordinates-only, name-only, and name+coordinates conditions.
  • Prompt and information conditions: The design predicts comparable cross-neighborhood variation across conditions if ratings accurately reflect real-world risk.
  • Models and ratings: Seven instruct-tuned models span roughly 3B parameters to frontier scale across four model families.
  • Models and ratings: The expected rating averages probabilities over the ten numeric answers and serves as the primary measure, while the greedy rating uses the single top-choice number.
  • Models and ratings: The numeric mass m is the pre-rescaling probability assigned to any numeric answer; its mean is at least 0.99 for every model in both cities.
  • Analytical design: Analyses use name-only ratings as the baseline, comparing violent-crime calibration, demographic coefficients controlling for crime and income, and the name shift r−r_coords.
  • Analytical design: Crime-matched pairs compare high- and low-demographic-share neighborhoods with similar violent-crime rates, discarding pairs differing by more than 0.5 SD.
  • Robustness and interpretation: Robustness checks vary prompt templates, rating scales, name prominence, response coverage, and the homicide-only crime measure.

5 Results

Across seven models, neighborhood names carry most safety-rating variation while also encoding both violent-crime signal and demographic stigma. The demographic effect is robust in Los Angeles, matched pairs, enforcement-elasticity tests, and geographic-knowledge comparisons, though coordinate use increases with model scale.

  • 5.1 The Name Carries the Signal: Coordinates-only ratings vary by just 0.01–0.15 points for six models, versus 0.20–1.67 points when names are present.GPT-4o reaches 0.67 in Chicago and 0.91 in Los Angeles, while GPT-4o-mini reaches 0.40 in Los Angeles.
  • 5.1 The Name Carries the Signal: Names track recorded violent crime in every model but lower safety ratings more for neighborhoods with higher shares of the locally dominant marginalized group.The demographic pattern is percent Black in Chicago and percent Hispanic in Los Angeles; all seven models show the name shift in both cities.
  • 5.2 Robustness: Wording, Name Prominence, and Enforcement: At fixed homicide rates, caution tracks homicide at −0.18 to −0.74 in every model but shows no consistent relationship with discretionary offenses.Adding discretionary offenses to a homicide-only model explains almost no additional variance; Los Angeles shows the same asymmetry with homicide partials of −0.49 and −0.46.
  • 5.4 Los Angeles: Generality Across Models and Axes: Los Angeles separates demographic share from crime: all seven models show a negative demographic effect net of crime and income, confirmed by regression, ablation, and matched pairs.Chicago cannot cleanly separate percent Black from crime because the variables correlate at r=0.75, whereas Los Angeles provides the required separation.
  • 5.3 Item-Level Confirmation: Crime-Matched Pairs: Across 28 Los Angeles crime-matched pairs, every model rates the higher-percent-Hispanic neighborhood less safe, while coordinates-only differences are indistinguishable from zero for six models.The mean within-pair crime gap is approximately 9 per 1,000.
  • 5.6 Knowledge and the Size of the Effect: Geographic knowledge scales the demographic name effect: compass-region accuracy correlates with it at ρ=−0.94 in Los Angeles, while recognition ability spans d′≈0.5 to 4.8.Recognition-test correlations are −0.71 in Chicago and −0.75 in Los Angeles, with all four estimates sharing a sign.

6 Discussion

The models’ neighborhood safety judgments are largely keyed to names, which carry both genuine crime information and demographic stereotype. These findings are bounded by imperfect crime measures, demographic confounding, and limited evidence for the knowledge–bias scaling relationship.

  • Name-driven judgments: For six of seven models, coordinates carry negligible rank information, while neighborhood names supply nearly all between-neighborhood safety variation.This name-based variation is partly correct because it correlates with recorded violent crime, but it also carries demographic stereotype.
  • Identification and trade-off: In Los Angeles, the demographic name effect exceeds what crime justifies because demographic share and crime can be separated; Chicago cannot distinguish them because segregation makes them closely aligned.The broader conclusion is that names carry both legitimate risk signal and stereotype, so removing names would reduce both bias and accuracy.
  • Robustness: Homicide-based and enforcement-elasticity analyses support the interpretation that model caution tracks reported violence rather than discretionary, deployment-driven offenses.The evidence is strongest in Los Angeles, where homicide and discretionary rates are nearly independent; Chicago’s rates are highly collinear.
  • Demographic stereotype: Name effects track percent Black in every model, and higher-percent-Black neighborhoods receive larger safety-rating reductions when names are added.In Chicago, higher-percent-Black areas also concentrate among higher-crime, lower-rated neighborhoods, while the name effect declines with percent Black.
  • Scope of the scaling claim: The knowledge–bias scaling claim is limited because cross-model gradients are underpowered, with n=6 in Chicago and n=5 in Los Angeles.Recognition ability separates from bias and does not linearly predict it; a denser model ladder would sharpen the claim.

7 Conclusion

The study finds that neighborhood names, not locations alone, drive most LLM safety variation while combining genuine crime information with demographic stereotype. The name penalty is robust across models and tests, so hiding names would reduce both bias and accuracy.

  • Core findings: Neighborhood names carry most safety-rating variation, whereas coordinates alone barely affect ratings.The conclusion identifies the name channel as dominant over location in the models’ judgments.
  • Prompt robustness: The name-to-demographic effect is not a wording artifact: calibration and name shifts retain their signs across ten prompt templates.The demographic coefficient remains near zero in the Chicago prompt-robustness figure.
  • Robustness: The over-caution band in Chicago’s high-percent-Black South and West Sides survives when homicides replace the crime-rate proxy.Figure 5 presents over-caution as ratings lower than predicted by crime.
  • Demographic penalty: The name depresses ratings more for neighborhoods associated with the locally dominant marginalized group across all seven models.The relevant demographic axis is percent Black in Chicago and percent Hispanic in Los Angeles.
  • Robustness: In Los Angeles, the demographic penalty survives controls for crime and income and is confirmed by crime-matched pairs.This supports an item-level name effect where demographic share and crime are more separable.
  • Implications: Because names encode both genuine crime signal and demographic stereotype, removing them reduces bias and accuracy.The paper therefore argues that name-based safety judgments require scrutiny before deployment.

8 Ethical Statement

The ethical statement limits interpretation to neighborhood-level aggregate data and emphasizes that recorded crime is an enforcement-shaped proxy, not ground truth. It also warns that publishing model-derived caution scores can reinforce stigma.

  • Data scope: The study uses aggregate neighborhood-level American Community Survey estimates rather than individual-level race or ethnicity data.This defines the demographic unit and scope of the analysis.
  • Interpretation: Recorded crime is an enforcement-shaped proxy, so the findings concern model caution relative to that proxy rather than which neighborhoods are actually dangerous.The authors report residuals relative to recorded crime instead of raw danger rankings.
  • Disclosure: Publishing caution scores for named neighborhoods risks reinforcing stigma.The paper mitigates this risk by framing high-caution areas as evidence of model bias, not travel or housing guidance.

9 Reproducibility Statement

The reproducibility statement documents model querying, data sources, and the prompt-robustness procedure used to make ratings and effects comparable. It also records the hosted-model snapshot context.

  • Model querying: GPT-4o-mini, GPT-4o, and DeepSeek V4 are queried through OpenRouter at temperature 0 with top-20 log-probabilities.Expected ratings sum probability over integer answer tokens 1–10, treating multi-token integers as full strings.
  • Data sources: Crime and demographic data come from public Los Angeles and Chicago sources, with tract-level ACS estimates aggregated to neighborhood polygons.The cited implementation passage identifies LAPD and Chicago open-data extracts and the aggregation procedure.
  • Robustness design: The robustness sweep reruns coordinates-only, name-only, and name+coordinates conditions under ten templates using a common 1–10 danger-to-safety orientation.Effects are computed against the name-only baseline.

B Example Responses

The examples illustrate how ratings are normalized across prompt variants and how the accompanying tables document prompt templates and matching balance. Together they show the mechanics behind robustness and comparison checks.

  • Example responses: GPT-4o-mini’s Englewood responses are converted to a common safety scale despite differing raw answers across prompt wordings.The example includes raw-to-normalized mappings such as “3”→3.0 and “40”→4.6.
  • Prompt templates: Table A1 lists the ten prompt-robustness templates used in the sweep.These templates vary wording and response framing.
  • Matching balance: Matching reduces Chicago crime imbalance roughly sevenfold to just above the 0.25 balance threshold, while income imbalance remains −0.55.The design controls for crime but not income.

C Crime-Matched Pairs: Balance and Robustness

The crime-matched-pair design uses share thresholds to define eligible neighborhoods and a caliper to require sufficiently similar crime rates. Because neither matching choice is canonical, Table A3 varies both.

  • Neighborhoods above the upper share threshold are high-share, those below the lower threshold are low-share, and intermediate neighborhoods are excluded.
  • The caliper determines how close two crime rates must be before a neighborhood pair is accepted.
  • Table A3 varies both thresholds and the crime-rate caliper because neither matching choice is canonical.

D Enforcement-Elasticity Tiers

The enforcement-elasticity analysis separates homicide from offenses whose recorded rates depend more on police deployment. Across models and cities, caution tracks homicide rather than higher-elasticity offenses, with Los Angeles providing a cleaner separation.

  • Enforcement-elasticity tiers: Chicago incidents are grouped into four tiers by enforcement elasticity: homicide, other violent, property, and discretionary offenses.The supplied passage lists 3.7k homicide incidents, 390k other-violent incidents, and 641k property incidents before the discretionary tier description is truncated.
  • Enforcement-elasticity tiers: Holding homicide constant, higher-elasticity other-violent, property, and discretionary rates show no additional robust caution across the seven models.Partial correlations range from −0.10 to +0.12 for other-violent, +0.06 to +0.24 for property, and −0.20 to +0.35 for discretionary rates.
  • Los Angeles robustness: Los Angeles has a thin discretionary tier of 22k incidents, or 2.3% of its tiered total, compared with about 9% in Chicago.
  • Los Angeles robustness: In Los Angeles, homicide and discretionary rates are nearly independent, and homicide partial correlations are negative while discretionary partials are positive in all seven models.The homicide partials range from −0.18 to −0.73, while discretionary partials range from +0.10 to +0.23.
  • Geographic knowledge: Recognition d′ spans 0.5–4.8 across models and correlates with name shift at −0.71 in Chicago and −0.75 in Los Angeles.The correlations are reported as exact p=0.088 for Chicago and p=0.066 for Los Angeles.
  • Enforcement-robustness: Table A5 reports negative homicide associations in every model and city, whereas discretionary associations are weaker and inconsistent in sign.The one exception is Qwen2.5-14B, whose two Chicago partials are comparable at −0.19 versus −0.22.
Loading 2608.26188v1…