Source-linked AI summary
Toward AI-Friendly Cartography: Understanding How Color Design Influences Foundation Model Spatial Reasoning on Sequential Choropleth Maps
Yonghe Sun, Zhenjia Liu, Hua Liao, Wenjia Xu, Nai Yang, Weihua Dong, Zhiwei Wei
TL;DR
It remains unclear whether cartographic color principles developed for human perception support foundation-model map reasoning. The study benchmarks controlled sequential choropleth variations and finds that ordering and sufficient lightness contrast matter more consistently than hue choice.
Problem
Systematic evidence remains limited on whether cartographic design principles developed for human cognition are effective for foundation-model spatial reasoning.
Method
The study evaluates 21 multimodal foundation models on 5,760 controlled choropleth maps and 28,800 tasks varying hue, sequential ordering, and lightness contrast.
Results
Hue effects are limited and inconsistent, whereas disrupting sequential ordering and reducing lightness contrast substantially impair performance, especially for comparison and ranking.
Takeaways & Limitations
Conventional sequential ordering and sufficient lightness contrast remain important considerations for machine understanding of sequential choropleth maps.
Takeaways & Limitations
The study focuses primarily on sequential choropleth maps, leaving other cartographic representations outside its evaluated scope.
Abstract
from arXiv · showhide
Foundation models (FMs) increasingly support multimodal and geospatial reasoning, yet it remains unclear whether cartographic principles designed for human perception are equally effective for machines. Focusing on sequential choropleth maps, we examine how hue palette, color ordering, and lightness contrast influence FM spatial reasoning. We construct a controlled benchmark of 5,760 maps and 28,800 questions spanning Attribute Identify, Spatial Recognition, Compare, Rank, and Pattern Delineate, and evaluate 21 open-source and proprietary multimodal FMs. Results show that hue choice has limited and inconsistent effects, whereas disrupting sequential color ordering substantially reduces performance, especially for comparison and ranking. Reduced lightness contrast also consistently impairs reasoning, while increasing contrast beyond sufficient separability provides only marginal gains. LoRA fine-tuning improves overall accuracy but preserves these relative sensitivities. Additional factorial experiments further indicate that errors arise from color-and-legend decoding, spatial reasoning, and the integration of thematic attributes with spatial structure. These findings show that conventional sequential ordering and sufficient contrast remain important for machine map understanding and provide empirical guidance for AI-friendly cartographic design.
1. Introduction
Foundation models are increasingly used for multimodal and spatial map reasoning, but it remains unclear whether human-oriented choropleth design principles support machine reasoning. This study addresses that gap through a systematic investigation of hue, sequential ordering, and lightness contrast in choropleth maps.
- Research context: Recent work applies foundation models to map understanding, structural and semantic extraction, and question-answering over map content.Examples include MapLayNet, MapReader, MapQA, CartoMark, and MapVerse.
- Research gap: Existing research emphasizes task performance and applications, leaving the effects of fundamental map design principles on FM reasoning unclear.Conventional choropleth design principles were developed primarily to optimize human visual perception, whereas FMs process visual information differently.
- Study focus: Choropleth maps encode quantitative attributes across geographic regions, with color serving as the most salient visual channel for conveying those attributes.Established cartographic frameworks such as ColorBrewer provide sequential, diverging, and qualitative color templates.
- Contributions: The study presents one of the first systematic investigations linking classical choropleth color design principles with FM spatial reasoning.The work frames this connection as a bridge between traditional cartography and AI-friendly cartography.
- Contributions: 5,760 choropleth maps and 28,800 spatial reasoning tasks form a controlled benchmark spanning hue palettes, sequential ordering, and lightness contrast.The benchmark covers five cognitive dimensions, including Attribute Identify and Spatial Recognition.
2. Related work
Prior work established perceptually grounded choropleth design and increasingly applied foundation models to map understanding and applications. However, systematic evaluation of human-oriented map design principles for machine reasoning remains largely unexplored.
- Human-centered choropleth design: Choropleth research developed empirical, perceptual, and standardized guidance for selecting color schemes, classification methods, and visually clear map designs.Brewer’s work emphasized perceptually grounded sequences, while later studies empirically evaluated color combinations, map-reading accuracy, and user preferences.
- Human-centered choropleth design: Existing cartographic developments progressed from manual expert-driven design toward semi-automated and AI-assisted map production and analysis.Automated contour-map representation and other tools extended design principles into machine-readable and scalable workflows.
- Foundation models for maps: FM map research broadly divides into map content understanding and map-based applications.Content-understanding work includes map-element recognition, map-based question answering, storytelling, geospatial question answering, and benchmark construction.
- Foundation models for maps: Map-based FM applications include AI-assisted symbol optimization, adaptive choropleth color generation, and multi-agent, user-guided color-scheme generation.These systems leverage foundation models to support map design and improve visual clarity, particularly for choropleth and administrative maps.
- Research gap: Systematic evaluation of human-derived map design principles for machine reasoning remains largely unexplored despite progress in FM map analysis and automated generation.Addressing this gap is presented as necessary for AI-friendly guidelines that support both human readability and machine reasoning.
3. Methodology
The methodology combines machine-centered hypotheses, controlled choropleth generation, benchmark construction, and unified evaluation of 21 multimodal foundation models. It isolates hue, sequential ordering, and lightness contrast while validating benchmark interpretability with human participants.
- Overall framework: The study follows four stages: hypothesis formulation, controlled map generation, benchmark construction with human validation, and evaluation of 21 multimodal FMs.The framework uses unified experimental settings to examine how cartographic color designs influence machine spatial reasoning.
- Machine-centered hypotheses: Three hypotheses target the effects of hue palettes, sequential color ordering, and color contrast on FM spatial reasoning.The hypotheses predict limited hue effects, potentially weaker-than-expected ordering effects, and improved performance with stronger contrast.
- Controlled map design: All controlled maps share identical geography, spatial structure, thematic distribution, rendering, and layout unless the target hypothesis is manipulated.This design isolates the effects of the specified color variables.
- Hue manipulation: The hue experiment uses all 18 representative ColorBrewer sequential palettes, including single-hue and multi-hue schemes.The benchmark is designed to test whether hue-palette variation influences FM spatial reasoning.
- Ordering manipulation: The ordering experiment compares Sequential Encoding, which preserves monotonic ordinal correspondence, with Randomized Encoding, which permutes the same colors across thematic classes.The two strategies retain identical color sets while differing in color–value assignment.
- Contrast manipulation: The contrast experiment compares standard, high, and low settings by manipulating lightness, with adjacent-class differences enlarged 1.5× or reduced to 50%.Only 4-class and 5-class sequential schemes are used so manipulated lightness intervals remain sufficiently distinguishable.
4. Experimental Results
Hue palette choice produced limited, model-dependent effects, whereas disrupting sequential ordering substantially harmed FM reasoning. Reduced lightness contrast consistently impaired performance, while additional contrast yielded only marginal gains.
- Hue: Hue palettes produced limited, non-systematic variation: average accuracy ranged from 51.2% (Purples) to 55.5% (YlGn), a maximum spread of 4.3 percentage points.No individual hue-palette pair differed significantly from the Blues reference (all |z| < 1.32, all p > .18).
- Sequential Ordering: Randomized color ordering reduced average accuracy from 53.7% to 44.5%, an absolute decline of 9.2 percentage points and relative decrease of 15.9%.The paired comparison was highly significant (p < 0.001) with Cohen’s d = 1.762.
- Sequential Ordering: FM accuracy gaps between sequential and randomized encoding were generally larger than the 4.0-percentage-point human gap, indicating greater relative disruption for models.Every evaluated model remained below human accuracy lines under both conditions.
- Lightness Contrast: Compared with Standard, High Contrast increased average accuracy from 55.3% to 55.8% (+0.5), whereas Low Contrast reduced it to 53.3% (-2.0).Seventeen of 21 models improved under High Contrast, while 19 of 21 degraded under Low Contrast; Low Contrast was significantly worse (odds ratio = 0.91, p < 0.001), whereas High Contrast showed a small improvement (odds ratio = 1.02, p = 0.016).
- Robustness: Robustness checks reproduced the direction and relative magnitude of H1–H3 effects after accounting for questions nested within maps and models nested within families.The ordering analysis estimated predicted accuracy of 54.7% versus 44.7% under randomized encoding, while task-specific checks appear in Sections 5.1 and 5.2.
5. Discussion
The discussion shows that foundation models rely on conventional sequential color ordering and sufficient lightness contrast, with sensitivity varying by reasoning task. Errors also arise from color-and-legend decoding, spatial reasoning, and cross-representational integration demands.
- Sequential ordering: 24.6 percentage points: disrupting sequential ordering reduced D4 accuracy from 58.4% to 33.8%, while D2 remained almost unchanged.D3 also decreased from 49.2% to 36.7% (-12.5).
- Sequential ordering: D4 > D3 > D1 ≈ D5 > D2: GLMM odds ratios were 3.62, 2.06, 1.24, 1.25, and 1.01, respectively.The condition × task-dimension interaction was significant (p < 0.001), with all non-null contrasts significant at p < 0.001; D2 was null (p = 0.656).
- Lightness contrast: D1 accuracy rose from 68.2% to 71.2% under high contrast and fell to 64.2% under low contrast, showing strongest sensitivity to lightness manipulation.The corresponding average changes were +3.0 percentage points and -4.0 percentage points.
- Robustness: Mild pixel-level noise preserved the sequential-randomized gap at 12.8 versus 12.7 percentage points, whereas rotation and resolution reduction caused 21–23-point accuracy losses.Gaussian noise changed accuracy from 59.1%→58.5% under sequential encoding and 46.3%→45.8% under randomized encoding.
- Error sources: 11.0 points: supplying both attributes and spatial information outperformed spatial information alone, at 68.9% versus 57.9%, revealing an integration bottleneck.The positive interaction was approximately 7.4 percentage points, indicating that errors cannot be attributed independently to color reading or spatial reasoning.
- Residual difficulties: D5 remained difficult across all four information conditions, with accuracy ranging from only 36.3% to 39.4%, while reversed color direction performed below randomized encoding for all three models.Explicit region values, centroids, and adjacency relations did not resolve global spatial-structure delineation; D1 and D2 were nearly indistinguishable across original, randomized, and reversed conditions.
6. Conclusion
The study evaluates how sequential choropleth color design influences multimodal foundation-model spatial reasoning, finding that sequential ordering and sufficient lightness contrast remain important. Its scope is limited to sequential choropleths and general-purpose multimodal models, leaving broader cartographic variables and map-specialized systems for future work.
- Study contribution: The benchmark contains 5,760 choropleth maps and 28,800 spatial reasoning tasks evaluated across 21 multimodal foundation models.The study examines sequential hue palettes, sequential ordering, and lightness contrast under unified experimental settings.
- Key findings: Disrupting sequential color ordering substantially reduces performance, while insufficient lightness contrast consistently impairs spatial reasoning.Hue choice has limited and inconsistent effects, whereas increasing contrast beyond sufficient separability provides only marginal gains.
- Key findings: LoRA fine-tuning improves overall accuracy but preserves models’ relative sensitivities to color ordering and lightness contrast.Additional factorial experiments attribute errors to color-and-legend decoding, spatial reasoning, and integrating thematic attributes with spatial structure.
- Limitations and future work: The benchmark focuses mainly on sequential choropleth maps and color-related variables, leaving other map types and cartographic elements unexplored.The authors identify broader cartographic variables, additional thematic map types, and machine-oriented map design principles as future directions.
- Limitations and future work: The evaluation primarily targets general-purpose multimodal foundation models rather than map-specialized systems.Future work should investigate machine-oriented map design principles and additional map-specialized systems.
Notes on contributor(s)
The contributors shared responsibilities across methodology, data curation, software, writing, supervision, project administration, funding acquisition, conceptualization, validation, and visualization.
- Yonghe Sun led methodology, data curation, software, validation, visualization, and original-draft preparation.
- Zhenjia Liu contributed to methodology, data curation, software, visualization, and review editing.
- Hua Liao handled review editing, supervision, project administration, and funding acquisition, while Wenjia Xu, Nai Yang, and Weihua Dong contributed review editing and supervision.
- Zhiwei Wei is credited with conceptualization and methodology.The supplied contributor note is truncated after “Methodolo”.
Appendix A. Control Baselines for Template and Answer-Position Bias
Three control baselines tested whether performance reflected genuine map-reading rather than template or answer-position shortcuts. The controls covered all five task dimensions and included no-image, blank-image, and shuffled-answer conditions.
- Control baselines: The study evaluated no-image, blank-image, and shuffled-answer controls on InternVL3.5-8B and Qwen3.5-9B across D1–D5.No-image provided only the question and options; blank-image replaced the map with white; shuffled-answer tested answer-position effects.
- Evaluation criterion: The baselines’ accuracy was averaged across all five task dimensions against a weighted chance level of approximately 27.8%.The reported average spans D1–D5 under the no-image, blank-image, and shuffled-answer conditions.
Appendix B. Accuracy by Individual Question Subtype
Appendix B supplements the aggregated D1–D5 results with full per-model, per-subtype Q1–Q12 heatmaps for H2 and H3. Randomized encoding most strongly harms ranking subtypes, while low- and high-contrast effects are concentrated in specific question subtypes.
- Appendix methodology: Figures B1 and B2 report full per-model accuracy for question subtypes Q1–Q12 under randomized encoding and contrast manipulation, respectively.The results complement the five-dimension D1–D5 aggregation and use heatmaps for readability.
- Randomized encoding: Randomized encoding most degrades Q6 (global rank) and Q7 (local rank), with a smaller effect on Q5 (attribute comparison).These subtype effects are consistent with the D3/D4 sensitivity reported in Section 5.1.
- Randomized encoding: Q3 (direction) and Q4 (adjacent) show negligible change across nearly all models under randomized encoding.Both subtypes correspond to D2.
- Contrast manipulation: Under contrast manipulation, low-contrast degradation and high-contrast improvement are concentrated in Q1.Figure B2 presents Standard minus Low Contrast for degradation and High Contrast minus Standard for improvement.
Appendix C. Example Maps: Source Data and Benchmark Rendering
Appendix C illustrates benchmark construction by showing maps before and after synthetic rendering. It covers real-world source thematic data and benchmark maps produced under hue, ordering, and contrast manipulations.
- Example Maps: The appendix presents example maps at two stages: source thematic data before synthetic rendering and resulting maps after controlled color manipulations.The manipulations vary hue, color ordering, and contrast.
C.1. Example Source Thematic Data
The benchmark draws attribute values from MapQA-derived real-world thematic data, whose varied US state-level spatial distributions precede controlled synthetic map rendering.
- Source data: MapQA-derived source data supplied attribute values for benchmark construction before controlled synthetic rendering.The source data included Kaiser Family Foundation health and healthcare indicators.
- Source data: Example maps represent health insurance coverage, healthcare expenditure, and mental health indicators across US states.These examples illustrate the diversity of real-world spatial distributions used to derive thematic value structures.
- Source data: Figure C1 shows these real-world thematic maps rendered at the US state level before synthetic map generation.
C.2. Example Benchmark Map Renderings
Figure C2 showcases CHROMA’s additional example choropleth maps across five thematic categories and multiple color-rendering variants. These examples illustrate the benchmark’s thematic diversity and controlled color manipulations.
- Benchmark Diversity: The showcased thematic content includes obesity prevalence, software usage rate, birth rate, and unemployment rate under controlled color manipulations.
- Thematic Coverage: Figure C2 spans five thematic categories: Obesity Population, Software Usage Rate, Birth Population, Unemployment Population, and Birth Rate.
- Color Variants: The example maps include multiple color-rendering variants used in the benchmark construction pipeline.
Appendix D. Value-in-Region Experiment
The value-in-region, color-free representation generally reduced FM accuracy relative to sequential color encoding, with the largest losses in D1 and D4. Numerical labels may be harder to recognize and compare, while color gradients can provide salient cues for attribute classes, extremes, and spatial patterns.
- Experimental control: The value-in-region condition additionally required numerical-text recognition and region–value association.The experiment used the same underlying maps, attribute values, questions, and answer choices as the sequential color-encoding condition, but replaced color patches with numerical labels.
- Accuracy comparison: Value-in-region encoding lowered accuracy on most task dimensions compared with sequential color encoding.D2 showed virtually no difference, at 68.8% versus 68.9%.
- Accuracy comparison: 24.2 percentage points was the largest decrease, with D1 accuracy falling from 80.1% to 55.9%.D4 also declined by 16.4 points, from 76.1% to 59.7%, while D5 declined by 5.2 points, from 39.4% to 34.2%.
- Interpretation: Color gradients may aid D1 and D4 by making attribute classes or extreme values perceptually salient, whereas numerical labels require comparing multiple distributed values.Color patches may also make spatial clusters and structural patterns more visually salient for D5.