Source-linked AI summary

Vision Models Predict Urban Scene Appraisal with Limited Neural Alignment

Kaizhen Tan, Yuantao Deng

arXiv:2608.30964v1cs.CV

TL;DR

Predicting human appraisal from urban-scene embeddings does not establish that those embeddings represent scenes like the brain. Using public EEG, the paper separately measures neural-geometry correspondence and appraisal prediction across seventeen feature spaces, finding low correspondence despite strong rating prediction and no reliable relationship between the measures.

  • Problem

    Urban-scene embeddings are validated mainly by predicting human ratings, leaving open whether they organise scenes as human perception does.

  • Method

    The study compares seventeen image feature spaces with time-resolved EEG representational geometry from 63 adults viewing and rating Berlin street scenes.

  • Results

    Neural correspondence was low across models, while appraisal prediction reached r = 0.87 and showed no reliable relationship with brain alignment.

  • Takeaways & Limitations

    Appraisal prediction is weak evidence that a model represents a street as the brain does, although it remains appropriate for estimating appraisal outcomes.

  • Takeaways & Limitations

    The 55-scene set comes from one city and contrasts vegetation levels rather than sampling streets representatively, limiting claims about other cities and absent street types.

Abstract

from arXiv · show

Pretrained vision embeddings are increasingly used as general-purpose representations for modelling how people appraise urban scenes, and are validated almost entirely by how well they predict human ratings. High predictive accuracy does not establish that these embeddings organise scenes as human perception does. We test the two properties separately against brain data. Using openly released EEG from 63 adults who viewed and rated 56 Berlin street scenes, we estimate the representational geometry of the scenes over time, the proportion of that geometry that is explainable at all, and its correspondence with seventeen feature spaces spanning language-supervised, self-supervised, category-supervised and dense-prediction training, two orders of magnitude of scale, and interpretable controls. Correspondence is low throughout: the best representation, DINOv2 ViT-B, reaches 29.6% of the lower bound of the noise ceiling, the panel spans 11.0% to 29.6%, and a Gabor energy descriptor is indistinguishable from the best model while outperforming every language-supervised model tested. Within a model, deeper layers still match later neural responses, so the hierarchical correspondence found for object recognition survives even at this low overall level. The same embeddings predict held-out appraisal ratings well, up to r = 0.87, and the two measures do not track each other across models; reweighting features towards the neural geometry lowers appraisal prediction for every model tested, against a control of matched dimensionality. Predicting how a street is appraised is therefore weak evidence that a model represents the street as the brain does. The benchmark uses only public data and requires no training, so evaluating a new representation needs only its embeddings for 55 images.

1 Introduction

Urban-scene embeddings are usually validated by rating prediction, which does not show that they organise scenes like human perception. This paper separates predictive success from brain correspondence using EEG geometry and a broad representation benchmark.

  • Motivation: Held-out rating prediction tests a representation’s output, not whether its internal scene organisation matches human perception.A model can reproduce ratings while arranging scenes along different internal dimensions.
  • Approach: Representational similarity analysis compares brain and model systems through the pairwise arrangement of the same scenes.Time-resolved EEG or MEG can show how scene representations unfold.
  • Motivation: Urban-scene applications target evaluative properties such as safety, beauty and openness, but whether their learned representations resemble human neural representations had not been measured.Prior neural work related urban-scene responses to small sets of hand-specified image properties.
  • Research questions: The study asks how much neural urban-scene geometry current representations capture, how correspondence varies across model factors, and whether it predicts appraisal.The benchmark uses EEG from 63 adults viewing and rating 56 Berlin street scenes, comparing seventeen feature spaces.
  • Contributions: The benchmark uses public data without training and finds limited neural correspondence, while appraisal prediction remains strong and does not track brain alignment.A Gabor descriptor matches the strongest learned representations, and the best appraisal predictors are among the least aligned.

2 Methods

The study analyses public EEG and image features from a repeated urban-appraisal task, constructs cross-validated neural scene geometries, and compares them with seventeen model feature spaces. Noise ceilings, temporal analyses, scene bootstrapping and cross-validated appraisal prediction separate reliability, correspondence and predictive performance.

  • Dataset and appraisal task: EEG from 63 adults covered 56 Berlin scenes, each presented nine times under nine appraisal scales, yielding 31,752 trials.Scenes appeared for 3000 ms after scale and fixation cues; behavioural analyses used all 56 scenes.
  • Dataset and appraisal task: Image-feature analyses used 55 publicly available scenes because one presented image lacked public pixel data.Behavioural analyses retained all 56 scenes.
  • EEG preprocessing: Preprocessing filtered, interpolated, ICA-cleaned and re-referenced the EEG, then retained 94.0% of trials after epoch rejection.Fp1 and Fp2 served as electro-oculogram proxies because no dedicated ocular channel was recorded.
  • Neural scene geometry: Neural scene geometry was represented by cross-validated squared Mahalanobis RDMs computed after multivariate noise whitening.Three folds mixed trials across all nine appraisal scales, and the estimator was unbiased under null differences.
  • Neural scene geometry: Primary model comparisons used a less-noisy neural RDM averaged over 150–500 ms, while time-resolved analyses retained one RDM per time point.The window-level RDM has a higher noise ceiling than any individual time point within the window.
  • Vision representations: The benchmark compared seventeen feature spaces spanning supervised, self-supervised, category, dense-prediction, scene-description and interpretable low-level representations.Transformer and ResNet activations were sampled across depths, with model RDMs computed using correlation distance.
  • Model-brain correspondence: Correspondence was evaluated as participant-averaged Spearman correlation, with permutation tests, paired comparisons, scene bootstrapping and pre-stimulus controls.Depth analysis normalised layer position and measured peak latency across eleven multilayer models.

3 Results

The neural scene geometry was reliable and largely independent of evaluative goal, but current vision representations captured only a small fraction of it. Deeper layers matched later neural responses, while appraisal prediction and neural correspondence diverged across models.

  • Behavioural and neural reliability: 85.4% of appraisal variance was summarized by two principal components across the nine rating scales.The first component accounted for 74.1% and the second for 11.2%.
  • Behavioural and neural reliability: 2.04 mean crossnobis distance between 150 and 300 ms contrasted with 0.015 before image onset, while the lower noise ceiling reached 0.194.Neural scene dissimilarities rose after stimulus presentation and remained reliable thereafter.
  • Model–neural correspondence: 29.6% of the lower noise-ceiling bound was reached by the best representation, DINOv2 ViT-B, while the seventeen-space panel ranged from 11.0% to 29.6%.Every feature space correlated with neural geometry but remained far below the ceiling.
  • Model–neural correspondence: 28.1% of ceiling was reached by a Gabor energy descriptor, indistinguishable from the best foundation model and higher than every language-supervised model tested.The descriptor achieved ρ=0.079, compared with CLIP ViT-L/14 at ρ=0.051 and SigLIP SO400M at ρ=0.038.
  • Depth, scale, and training domain: +0.62 was the mean rank correlation between relative layer depth and peak latency in 10 of 11 models with at least three extracted layers.Deeper layers matched later neural responses, including ρ=1.00 for DINOv2 ViT-B/16 and ρ=0.98 for CLIP ViT-B/32.
  • Appraisal prediction: r=0.87 appraisal prediction for SigLIP SO400M contrasted with ρ=0.038 neural correspondence, and the two measures showed little cross-model correspondence.Reweighting features toward neural geometry lowered appraisal prediction from r=0.60 to r=0.15 for every one of 15 models tested.

4 Discussion

The results show that appraisal prediction and neural representation are distinct evaluation targets: models can predict ratings well while matching EEG geometry weakly. Neural correspondence depends on what EEG resolves and on the limited, nonrepresentative scene set.

  • Prediction and representation come apart: SigLIP and CLIP predicted held-out appraisal ratings up to r = 0.87 despite weak brain alignment, whereas DINOv2 and Gabor features aligned more strongly but appraisal prediction differed.Across models, appraisal prediction and neural correspondence did not track one another.
  • Prediction and representation come apart: Reweighting features toward neural geometry lowered appraisal prediction for every model and failed to improve held-out correspondence, indicating difficulty fitting a neural subspace from 55 scenes.The result does not establish a trade-off between the two objectives because correspondence also failed to increase.
  • Implications for urban analytics: Neural correspondence matters when embedding dimensions or distances are interpreted as properties of human perception, not merely when embeddings predict appraisal outcomes.The paper therefore separates predictive validation from representational validation.
  • Scale, supervision and depth: A Gabor energy descriptor matched the best foundation model, while larger variants did not improve alignment in matched families, suggesting shared correspondence may reflect low-level image structure.This comparison concerns geometry EEG resolves over the first second of viewing a street.
  • Scale, supervision and depth: Deeper layers matched later neural responses in 10 of 11 models, preserving hierarchical timing even though overall representational content remained poorly aligned.Depth behaved differently from scale.
  • Limitations: The appraisal geometry did not shift across nine evaluative prompts, although the null result bounds rather than excludes small or EEG-unresolved goal-dependent components.This makes task-set differences unlikely to account for the observed model-brain gap within the measured signal.
  • Limitations: The noise ceiling captures structure reproducible across participants, so model distance from it reflects more than measurement noise discounted by the ceiling.It bounds the neural structure available to a stimulus-computable model.
  • Limitations: The claims are limited to structure EEG resolves because sparse or deep neural populations may be underrepresented, although model comparisons remain matched to the same measurement.The limitation concerns the measurement scope rather than relative model evaluation.

5 Conclusion

The conclusion establishes a benchmark separating neural correspondence from appraisal prediction. It finds weak correspondence despite strong rating prediction and recommends direct neural validation before interpreting embedding geometry perceptually.

  • Conclusion: EEG from 63 adults viewing Berlin scenes was compared with seventeen image-feature geometries and calibrated against structure reproducible across participants.The benchmark uses public data and does not train models.
  • Conclusion: DINOv2 ViT-B reached 29.6% of the lower-bound noise ceiling, while the panel ranged from 11.0% to 29.6%; Gabor matched the best learned representation.Larger variants did not improve correspondence, but deeper layers matched later responses in 10 of 11 models.
  • Conclusion: Appraisal prediction reached r = 0.87, yet prediction quality showed no reliable relationship with neural-geometric correspondence across models.
  • Conclusion: Held-out predictive accuracy remains appropriate for estimating perceived safety or beauty, but rating performance does not justify interpreting embedding dimensions or distances as human perceptual properties.That perceptual property must be established directly.
  • Conclusion: The public, training-free benchmark can evaluate a new representation using only embeddings for 55 images, testing a question urban prediction performance cannot answer.

A Data and code availability

The study’s data, code, model weights, and analysis pipeline are publicly available, with preprocessing safeguards addressing scaling, channel labels, event decoding, and estimator validity.

  • Data and code availability: EEG, stimuli, experiment code, ratings, segmentation maps, descriptors, and official model weights are publicly available, and the pipeline regenerates every figure and table.
  • Preprocessing: Stored EEG values must be multiplied by 10^-6 because the BIDS headers omit channel types and values are in microvolts.Failing to rescale can empty the dataset through artifact thresholds without raising an error.
  • Preprocessing: The 66 recorded channels include 64 scalp electrodes plus ECG and electrodermal activity, and the latter two were dropped.The electrodermal channel is labelled in microvolts but records microsiemens.
  • Preprocessing: Bad electrodes were flagged using flatness, maximum absolute correlation below 0.4, or standard deviation above eight montage medians.Variance alone was rejected because blinks make working frontopolar electrodes high-variance.
  • Event decoding: Marker decoding normalised corrupted German, English, and control-byte spellings so beauty trials were not silently discarded.
  • Estimator validation: The crossnobis estimator returned near-zero mean dissimilarity under simulated null data and large positive distances when condition means were injected.The null simulation used 200 runs with 10 conditions, 20 channels, and 8 trials per condition.
  • Noise ceiling: The upper noise-ceiling bound is positively biased for unreliable RDMs, so model performance was expressed relative to the conservative lower bound.With 61 participants, pure noise predicts an upper bound near 0.128.

E An alternative to normalising by the ceiling

The paper checks ceiling-normalised correspondence with an independent attenuation-corrected correlation. The two normalisations agree closely, while the variance scale shows a smaller absolute neural-geometry match.

  • Alternative normalisation: The percentage-of-ceiling ratio has no clean statistical meaning because the ceiling is a plotting band rather than a well-defined denominator.An independent route was therefore used to recompute the same quantity.
  • Alternative normalisation: The corrected correlation estimates alignment between a deterministic model geometry and noise-free neural geometry using ρ_true = ρ_obs/√0.867.The correction is bounded by one and independent of participant count.
  • Results: DINOv2 ViT-B correlated at ρ = 0.233, or ρ = 0.250 after correction, while accounting for 6.3% of variance in the noise-free neural geometry.The panel spanned 1.0% to 6.3% on the variance scale.
  • Results: The two normalisations ranked models similarly, with rank correlation 0.97 across the panel.

F Layer selection

Layer selection was robust to participant-level cross-validation for nearly all multi-layer models, with negligible effect on the panel mean.

  • 12 of 14 multi-layer models retained the same best-matching layer when selection was cross-validated across participants.The layer was selected from 60 participants and evaluated on the held-out participant.
  • The panel mean shifted by only 0.002 under cross-validated layer selection.
  • Only CLIP ViT-B/32 (LAION) and SigLIP B/16 changed layers on some folds, and both ranked low either way.

G Validation of the task-set measure

Synthetic-data validation tested whether the task-set measure detects modulation when present and remains near zero when absent.

  • The ratio R of Equation 2 was validated on synthetic data matching the real analysis in participants, scenes, scales and channels.
  • −0.015 modulation index occurred with no modulation, compared with +0.558 under moderate and +0.913 under strong task-specific modulation.
  • The measure recovered task-specific modulation when present and reported none when absent.

H Decomposition of the neural geometry

The neural-RDM decomposition into early visual, semantic, spatial-layout and appraisal components was unresolved: estimates were small, intervals spanned zero, and the design was not identifiable.

  • H Decomposition of the neural geometry: The decomposition used early visual, semantic, spatial-layout and appraisal feature spaces on the group-mean RDM.A single noisy RDM produced upward-biased regression estimates, including pre-stimulus R^2 = 0.045.
  • H Decomposition of the neural geometry: R^2 = 0.044 was the pre-stimulus baseline for expressing all decomposition quantities.
  • H Decomposition of the neural geometry: 0.072 total explained variance above baseline occurred between 150 and 300 ms, with a 95% interval of [−0.041, 0.177].Every component and window had a scene-resampled interval spanning zero.
  • H Decomposition of the neural geometry: With 55 scenes and four correlated feature spaces, the decomposition was not identifiable and was reported as unresolved rather than as separate effects.The model comparison remained supported under scene resampling, distinguishing this limitation from the dataset as a whole.
  • H Decomposition of the neural geometry: Inference over scenes asks whether effects recur with new streets, whereas inference over participants asks whether they recur with new people viewing these streets.The two forms of inference can disagree sharply.
  • H Decomposition of the neural geometry: The model-panel gap from the neural ceiling persisted across analysis choices, although the ordering itself varied.

I The neural reweighting does not generalise

Reweighting model features toward the neural geometry reduced held-out appraisal prediction and failed to improve out-of-sample neural correspondence, indicating overfitting rather than generalisable alignment.

  • r = 0.60 fell to r = 0.15 across 15 models after neural reweighting, making the appraisal result alone ambiguous.
  • Weights were fitted on half the scenes, while correspondence was evaluated on scenes contributing to neither the weights nor the PCA basis.
  • 0.058 correspondence for the reweighted representation was below 0.113 for the matched control, improving in only 1 of 15 models.The manipulation was therefore reported as unsuccessful rather than as evidence about appraisal-carrying feature directions.
  • The main cross-model divergence relied on a comparison involving no fitting, rather than on the unsuccessful reweighting manipulation.

J Robustness to analysis choices

Robustness checks preserve the broad gap between the model panel and the neural ceiling, but model rankings are not stable across all analysis choices.

  • ρ=0.99 after excluding two participants, preserving the model ordering under the quality-control criterion.Restricting the latency window to 150–300 ms also preserves the ordering, with ρ=0.79.
  • The robust conclusion is the large gap between the whole model panel and the neural ceiling, not a stable ranking of individual models.This narrower claim is consistent with scene-bootstrap results and motivates not interpreting the ranking.
Loading 2608.30964v1…