Source-linked AI summary

You Cannot Photograph the Same Street Twice: Reliability Limits in Vision-Language Measurement of Urban Change

Kaizhen Tan

arXiv:2609.00649v1cs.CV

TL;DR

The paper examines the poorly understood longitudinal reliability of vision-language perception scores from repeated street imagery. Using repeated Street View observations and controlled acquisition experiments, it finds that same-place photographic variation is large at individual locations, while aggregation and common-camera rendering recover more useful signals.

  • Problem

    Longitudinal reliability of vision-language perception scores from repeated street-level imagery is not well understood, especially at the individual-location reporting scale.

  • Method

    The paper builds a reliability ladder from consecutive-epoch Google Street View frame pairs and tests acquisition effects, image statistics, prompt order, encoding, and model dependence using controlled comparisons.

  • Results

    Same-place photographic variation is substantial, acquisition responses depend on the model, and common-camera rendering reduces spurious change classification in crowdsourced imagery while aggregation recovers redevelopment-related perception differences.

  • Takeaways & Limitations

    Vision-language measurement of urban change is usable for aggregated analyses over hundreds of paired observations, but not for interpreting individual sample-point changes without accounting for reliability limits.

  • Takeaways & Limitations

    The reliability estimate is based on five US cities and rendered Google Street View imagery, while control labels omit smaller persistent physical changes and the field-scoring models differ from the three models tested.

Abstract

from arXiv · show

Vision-language models are increasingly used to measure urban change from repeated street-level imagery, but their longitudinal reliability is not well understood. We test how much a perception score can change when the street itself does not undergo substantial redevelopment. Using 4,648 consecutive-epoch image pairs from 435 Google Street View standpoints across five US cities, we find that re-photographing the same street changes a perception score by 0.80 points on average, equivalent to 66.5% of the difference between two different streets in the same city. Repeated model calls contribute almost no variation, while image re-encoding and prompt-order changes each account for about one fifth of the between-street difference. Six image statistics describing scattering, contrast, colour, exposure, sharpness and specularity explain almost none of the remaining epoch-to-epoch variation. A small systematic drift of about 0.1 points remains and increases with the interval between captures, consistent with minor physical changes not recorded by redevelopment labels. Controlled experiments further show that acquisition conditions can shift scores when camera and image properties are allowed to vary, and that the direction of these shifts depends on the model. In crowdsourced imagery, camera geometry alone causes a model to report physical change in 45% of identical-scene pairs; normalising both images to a common virtual camera reduces this rate to 7.5%. Despite poor reliability at the individual-location level, aggregation recovers a coherent redevelopment signal: changed streets are judged wealthier, better maintained, more enclosed and less green. These results show that vision-language measurement of urban change is reliable at the scale of hundreds of paired observations, but not at the scale of individual sample points.

Highlights

Repeated photographs can substantially alter vision-language perception scores even when the street itself is unchanged, limiting reliable inference at individual locations.

  • Two-thirds of the difference between different streets is reproduced by re-photographing an unchanged street.
  • Six acquisition statistics explain almost none of the remaining epoch-to-epoch movement.
  • A tenth of a point is systematic, and its magnitude grows with the capture interval.
  • Camera geometry alone makes a model report physical change in 45% of identical-scene pairs.
  • The usable unit is a few hundred paired observations rather than a single sample point.

1 Introduction

The paper asks whether longitudinal perception scores remain interpretable when locations are repeatedly photographed, and tests this using unchanged Street View standpoints alongside controlled acquisition experiments. It finds substantial same-place variation, little explanatory power from six image statistics, a small interval-dependent drift, and a redevelopment signal recoverable through aggregation.

  • Longitudinal street-level studies compare repeated images to infer physical change and shifts in how streets are perceived.
  • A vacant Oakland lot changed substantially after eight townhouses were built, while lighting and foliage changes also shifted scores despite no redevelopment.
  • 4,648 consecutive-epoch frame pairs from 435 standpoints show that a second photograph moves scores by two-thirds of the between-street difference.
  • Six image statistics explain almost none of epoch-to-epoch variance, although they separate real weather conditions by up to 2.5 corpus standard deviations.
  • A signed drift reaches 0.12 points at most, is positive on almost every dimension, and grows with capture interval rather than calendar time.
  • Controlled degradation shifts scores, camera geometry causes reported change in 45% of identical-scene pairs, and acquisition-response signs depend on the model.
  • Aggregation over a few hundred paired observations recovers changed streets as wealthier, better maintained, more enclosed, and less green.
  • The paper contributes a reliability ladder, residual decomposition, a three-model acquisition test, and an estimate of the sample size needed for usable longitudinal measurement.

2 Related work

Related work established longitudinal street-level measurement, acquisition bias, and reliability evaluation for perception models. This paper extends those strands by directly testing repeatability, acquisition effects, and viewpoint-related camera geometry.

  • Prior longitudinal studies paired Street View imagery to measure physical change and its relationship to urban perception.
  • Earlier work addressed lighting, seasonality, learned structural-change embeddings, and change-detection networks.
  • Research has shown that capture conditions, including season and weather, can shift downstream urban-perception measurements.
  • Reliability research compares vision-language models with human annotations and argues that benchmarks should report inter-annotator reliability alongside model alignment.
  • Viewpoint-aware pairing tools filter candidates using feature matches and semantic-mask agreement, while this paper isolates residual camera-geometry effects after viewpoint control.

3 Data and design

The study combines a longitudinal Street View panel with experiments that progressively hold the scene, image, prompt, and camera conditions fixed or variable. Scores use twelve dimensions and a common between-standpoint yardstick to separate repeatability from cross-sectional variation.

  • Four experiments vary how much remains fixed: field conditions, model noise, controlled image perturbations, and virtual-camera optics.
  • A panorama is a Street View capture, a frame is one heading-specific rendering, and a frame pair joins consecutive epochs at one standpoint and heading.
  • The panel contains Google Street View records across 757 standpoints in five US cities from 2007 to 2023, with hand-labelled change points.
  • Each surviving panorama was requested twice at headings 0° and 180° using 640-by-640-pixel images.
  • The analysis accounts for differing proportions of control and labelled-change standpoints across the five cities.
  • The twelve perception dimensions include six Place Pulse-aligned measures and six physically auditable measures, each anchored at 0, 5, and 10.
  • The observational panel uses gemini-2.5-flash at temperature zero, while controlled experiments deliberately vary the model.
  • Differences are benchmarked against a 1.19-point mean absolute difference between randomly chosen different standpoints in the same city.

4 A second photograph of the same street

A second photograph of an unchanged street introduces substantial measurement variation, dominated by same-place photographic noise rather than repeated model sampling or the six measured acquisition statistics. A small interval-dependent drift remains, making individual-location estimates unreliable even though averaging can reduce random variation.

  • Repeated model calls differ by less than 1% of the between-place difference, so sampling variance contributes almost no observed change.
  • A second photograph differs by two-thirds of the between-place difference, while image re-encoding and prompt reordering each move scores by about one-fifth.The same-place difference is three times the prompt-order effect; averaging two standpoint headings reduces it to 0.64 points.
  • Six acquisition statistics explain almost none of the epoch-to-epoch variance, despite separating real weather conditions by up to 2.5 corpus standard deviations.The largest single correlation is 0.19, and the cross-fitted model has negative out-of-fold R2 on ten of twelve dimensions; greenery reaches 0.075.
  • Systematic drift reaches 0.118 points on wealthy and 0.110 on building condition, remaining below one-tenth of the between-place difference.Later photographs are judged marginally safer, wealthier, better maintained and more attractive.
  • Drift increases with capture interval, with same-place differences rising from 0.75 points under a year to 0.91 points at four to eight years.Averaging removes random variation but leaves about a tenth of a point of drift.
  • At a single location, standard error is comparable to the between-place standard deviation, so aggregation rather than point-level reporting is the usable scale.

5 What acquisition does when it is free to vary

Controlled perturbations show that acquisition conditions can shift perception scores when camera and image properties vary, complementing the flat result for rendered Street View imagery. The magnitude and direction of these shifts depend on the perturbation and the model.

  • Sixteen perturbation families are applied at 57 levels to 150 fixed, daylight, clear-weather, good-quality, non-panoramic images spanning six continents.The design holds the scene fixed by construction and reproduces axes along which the corpus varies.
  • 5.1 Magnitude: For qwen2.5-vl-72b-instruct, only 29 of 192 perturbation-by-dimension cells pass both significance and noise-floor tests.Each family is evaluated at the level producing its largest mean absolute response, so each displayed row is an upper bound for that family.
  • 5.1 Magnitude: At calibrated real-fog strength, atmospheric scattering shifts the most affected dimensions by about one-fifth of a between-place standard deviation, increasing monotonically with dose.
  • 5.2 Direction, and whose property it is: The direction of acquisition responses depends on the model: six of twelve dimensions have agreement across all three models, while no dimension is significant in all three.Beautiful reverses sign between gemini-3.7-flash and llama-4-maverick, and boring also reverses sign.
  • 5.2 Direction, and whose property it is: A directional correction cannot be applied without first fixing the model, because a correction estimated with one model may point the other way under another.

6 Camera geometry in crowdsourced imagery

Crowdsourced imagery can produce spurious change reports because images from the same place differ in camera geometry. Rendering paired images through a common virtual camera sharply reduces classification artefacts, but perception-score artefacts largely remain.

  • A connected-component metadata field is unusable as a co-location key: only 15.0% of cross-year pairs share one, while seconds-apart frames can carry 22 to 28 distinct values.Refined position and heading are available, but the connected-component field does not reliably identify repeated locations.
  • 47.7% of cross-year candidate pairs pass the 10 m and 20° geometric criterion, leaving camera differences as the main obstacle among visually inspected pairs.Dashcam, action-camera and panorama views can frame the same location very differently.
  • Resampling both frames into a shared virtual pinhole camera uses relative heading and a pair-specific field of view, reducing uncovered output content from 64% to under 1%.The method derives the field of view from what both sources can cover rather than fixing an 80° target.
  • 45% of identical-scene panorama pairs trigger a physical-change report under raw optics, whereas common-camera normalisation reduces the rate sixfold.The experiment uses two views rendered from a single panorama, so the scene is identical by construction.
  • Normalisation lowers the mean perception artefact only from 0.39 to 0.35, while median feature-match inliers change from 14 to 13.Residual yaw and field-of-view differences remain, and feature matching and perception scoring respond differently to the same image pair.

7 What survives aggregation

Aggregation recovers a coherent redevelopment signal despite noisy individual transitions. The signal survives city balancing, but small effects require many paired observations and reference labels omit some persistent physical changes.

  • 224 labelled-change transitions are compared with 2,367 control transitions to assess whether aggregation resolves redevelopment effects.The comparison spans 193 labelled-change standpoints and 435 control standpoints.
  • Four dimensions distinguish labelled change points from controls: streets are judged wealthier, better maintained and more enclosed, and less green.Three dimensions are physically auditable, and one aligns with Place-Pulse.
  • After re-weighting the five cities, greenery falls by 0.35, building condition rises by 0.24, wealthy rises by 0.20 and enclosure rises by 0.19, each at p<0.01.The effects shrink by one fifth to one third, but the pattern remains the same.
  • About 88 labelled-change transitions detect effects the size of the observed shifts, whereas a 0.1-point effect needs 1,010 and a 0.06-point effect needs 6,240.Below about 0.1 points, systematic drift rather than sample size becomes the binding constraint.
  • 7.1 Where the change threshold sits: Specificity rises ninefold across prompts while recall falls from 1.000 to 0.917, and model change calls cover 76.7%–97.7% of pairs against a 14.0% reference prevalence.The reference labels substantial redevelopment, whereas the prompts ask about persistent physical change, so the discrepancy partly reflects differing definitions.

8 Discussion

The discussion locates the measurement’s usable scale and boundaries. Aggregation can recover redevelopment patterns, but imagery correction is incomplete, model-dependent, and difficult to reproduce from unstable image identifiers.

  • A few hundred paired observations are the usable unit for perception-change measurement; individual-location maps display same-place photographic variation.The paper reports that re-photographing an unchanged street moves scores two-thirds as far as changing the street.
  • Common-camera rendering removes most spurious change classification in crowdsourced imagery but little perception artefact, so it is useful without being sufficient.The correction addresses camera heterogeneity, not the full perception-score instability.
  • Across three models, only six of twelve dimensions agree on acquisition-response signs, and a directional correction therefore requires fixing the model first.The sign for beautiful reverses between models with significant results.
  • Half of the panorama identifiers from a January 2024 dataset no longer resolved by August 2026, limiting retrieval-based reproducibility.The paper’s release consequently includes identifiers, retrieval dates, endpoint capture dates, prompts and derived scores.

9 Limitations

The reliability estimate is a lower bound because the panel uses Google Street View images normalised to a common virtual camera, unlike more heterogeneous crowdsourced imagery. Its control labels also miss some physical changes, and model choice can alter the direction of acquisition effects.

  • Scope and measurement: Google Street View images were rendered to a common virtual camera with fixed heading and field of view, making the reported same-place difference a lower bound for crowdsourced imagery.Crowdsourced camera heterogeneity adds a component measured separately in Section 6.
  • Scope and measurement: Control standpoints were labelled as lacking substantial redevelopment, but the labels do not record all street changes.Separating unrecorded physical change completely from instrument noise would require an explicit magnitude threshold in the control-set labelling scheme.
  • Model dependence: The three models used in the controlled experiments differ from the model scoring the field panel, so model-dependent sign disagreement does not provide a correction table for gemini-2.5-flash.This limits direct transfer of controlled acquisition effects to the field-panel scores.

Data and code availability

The paper’s primary data are openly documented across the underlying imagery and map sources, while Street View images themselves cannot be redistributed. Code, perturbation implementations, prompts, study materials and a content-addressed response cache will support regeneration of the results.

  • Data availability: Global Streetscapes and Mapillary imagery are distributed under CC BY-SA, while the CityPulse panorama index and OpenStreetMap history are also available through their stated sources.OpenStreetMap history is queried through the ohsome API.
  • Data availability: Street View imagery is not redistributable under Google Maps Platform terms, so the release provides panorama identifiers and derived scores instead of images.
  • Code availability: The release includes analysis code, sixteen perturbation implementations, five prompt variants, the study-design table and a content-addressed response cache.The cache allows every reported number to be regenerated without re-querying a model.

A Study design

The study combines four unit-specific experiments with controlled image perturbations and prompt manipulations, while defining change as substantial alteration to a street’s structure, surface or fixed elements. Its design also specifies how the perception and change judgments should discount acquisition artifacts and how results are aggregated and reported.

  • Unit of analysis: Frame pairs join consecutive epochs at one standpoint and heading, while transitions average the two headings of the same step.
  • Controlled perturbations: Sixteen perturbation families are applied across 57 total levels, covering atmospheric, optical, geometric, weather, image-quality, exposure and colour conditions.The design varies factors such as scattering, blur, yaw, rain, field of view, resolution, JPEG quality, exposure and white balance.
  • Prompt design: Five prompts manipulate perception scoring and change detection by adding an artefact block and, for one change prompt, a magnitude threshold.The two artefact blocks are the experimental manipulation behind Figure 4 and Table 1.
  • Perception scoring: The scoring instructions require judging the street rather than the photograph and discounting resolution, sharpness, compression, exposure, season, camera geometry and movable objects.The intended output is the perceived street condition rather than acquisition quality.
  • Change definition: A change is substantial when a structure, surface or fixed element is built, demolished, replaced, widened, narrowed or rebuilt, not merely altered in appearance.The protocol excludes seasonal plant changes, temporary barriers, repainting, litter, wear and individual street-furniture changes.
  • Reporting design: Figure 4 averages twelve dimensions in panel (a) and reports them individually in panel (b), with additional decompositions and condition profiles released as tables.The released tables include signed drift, variance decomposition, Figure 5b condition profiles, Figure 6 family-by-level responses and Figure 9 city contrasts.
Loading 2609.00649v1…