Source-linked AI summary
TEMPLAR Wales: A georeferenced environmental and toponymic dataset of Welsh settlements
Oktay Karakuş, Can Eyupoglu
TL;DR
Quantitative environmental-toponymy research requires clear separation among mapped settlements, lexical annotations and environmental measurements. TEMPLAR Wales addresses this need with a versioned, georeferenced dataset linking these layers through deterministic procedures and provenance, while documenting important limits on interpretation and missing environmental values.
Problem
Quantitative reuse of place names requires a clear distinction between mapped settlement records, lexical annotations and environmental measurements.
Method
TEMPLAR Wales integrates mapped settlement records, deterministic lexical screening using a frozen registry, and environmental attributes through a versioned georeferenced dataset.
Results
1,350 lexical detections across 1,294 settlements are released as the lexical component, with exact- and prefix-token matches retained separately.
Takeaways & Limitations
The dataset provides a basis for enrichment with historical name forms, archival maps, geology, soils and hydrological information.
Takeaways & Limitations
Environmental attributes describe contemporary or product-specific geographic conditions and should not automatically be interpreted as historical landscape reconstructions.
Abstract
from arXiv · showhide
Place names provide persistent records of how landscapes have been described and organised, but their quantitative reuse requires explicit separation between mapped places, lexical annotations and environmental measurements. TEMPLAR Wales is a georeferenced environmental-toponymy dataset comprising 3,757 settlement records across Wales. The resource links a reproducible settlement frame to deterministic lexical screening and settlement-level environmental attributes through stable identifiers. It contains 1,350 lexical detections across 1,294 settlements, generated from a frozen registry of 24 Welsh place-name elements, while retaining exact- and prefix-token matches and their provenance separately. Environmental attributes describe river and coastal proximity, elevation and local terrain context at multiple spatial scales, land cover and neighbourhood woody cover, with parallel terrain measurements derived from independent elevation products. The dataset is distributed as four relational tables accompanied by a field-level data dictionary, source-provenance register and licensing metadata. Technical validation confirms relational integrity, deterministic lexical reconstruction, documented environmental coverage, strong agreement between independent terrain sources and reproducible reconstruction of the frozen release. TEMPLAR Wales provides a reusable foundation for research in toponymy, linguistic geography, historical and environmental landscape studies, GIS and spatial data analysis without treating computational lexical detections as verified etymologies or contemporary environmental measurements as historical landscape reconstructions.
1 Background & Summary
TEMPLAR Wales addresses the need for quantitative environmental-toponymy data that separates mapped settlements, lexical annotations, interpretations and environmental measurements. It provides a versioned, georeferenced and reusable resource linking these layers through stable identifiers while preserving their provenance and interpretative boundaries.
- 3,757 settlement records across Wales form a versioned, georeferenced environmental-toponymy resource.
- The resource separates mapped settlement identity, selected analytical names, deterministic lexical detections, registered lexical elements and environmental measurements.These distinctions are represented through separate layers, provenance fields and documented processing rules.
- 1,350 lexical detections across 1,294 settlements were generated from a frozen registry while exact- and prefix-token matches remain separate.Settlements without detections remain in the complete frame, allowing users to define their own lexical contrasts.
- Environmental attributes cover hydrological, coastal, terrain and land-cover conditions using multiple geospatial products and terrain scales.Independent OS Terrain 50 and Copernicus GLO-90 measurements support source-sensitivity assessment.
- Four relational tables, documentation and provenance metadata make the resource inspectable, reproducible and independently reusable.Automated checks cover identifier integrity, table relationships, schema, lexical reconstruction and environmental coverage.
- The dataset supports spatial, environmental and historical-landscape research without treating string matches as verified etymologies or contemporary measurements as historical reconstructions.Its stable structure also supports enrichment with archival, geological, soil, hydrological, climatic and historical land-cover information.
2 Methods
The dataset is constructed through a staged workflow that combines a Wales-wide settlement frame, deterministic lexical annotation and settlement-level environmental derivation. These components are integrated into four related records through stable identifiers and released with documented provenance and validation.
- The workflow begins with a Wales-wide settlement frame and one selected analytical name for each retained record.A frozen, source-audited registry is then applied through deterministic detection.
- Independent geospatial products provide hydrological, coastal, terrain and land-cover attributes linked to lexical and settlement records through stable identifiers.
- Four related records separate settlement frames, long-format lexical detections, environmental measurements and the frozen lexical registry.This structure preserves source identity, derived annotation and environmental measurement as independently reusable components.
- Validation treats the release as a relational resource and checks identifier uniqueness, one-to-one settlement–environment relationships and release-specific integrity.The construction also records source versions, provenance, licensing and lexical-source alignment.
- The settlement frame contains 3,757 source records: 1,361 villages, 1,198 hamlets, 866 suburban areas, 190 other settlements and 142 towns.
- The unit of observation is an upstream settlement record at a mapped point, so distinct source records sharing names or coordinates are retained.The final frame contains 3,328 normalised analytical-name groups and 3,755 coordinate clusters.
2.4 Analytical-name construction
Analytical names are selected and normalised deterministically before lexical screening, with source fields and language status retained for transparent reuse. The frozen lexical resource provides versioned, source-audited annotations rather than definitive etymological classifications.
- A single analytical name was selected deterministically for every settlement before lexical detection.
- 3,618 records used primary name1, including 203 explicitly Welsh-labelled names, while 139 used a Welsh-labelled name2 field.
- Every settlement retains a non-missing analytical name, its source field, and language status distinguishing unresolved metadata from positive classification.
- Normalisation lowercased text, decomposed diacritics, treated hyphens and apostrophes as boundaries, and removed non-a–z matching characters.
- The registry contains 24 source-audited Welsh place-name elements, each with documented alignment or limitation and version information.
- ETDE v1 applies frozen exact-token and prefix-token rules, preserving match type and provenance without establishing etymology, morphology, or language identity.
2.8 Terrain attributes
Terrain attributes represent settlement elevation and local terrain position across multiple neighbourhood scales, with independent Copernicus measurements retained for source-sensitivity assessment. Coverage, validity and provenance fields preserve the conditions needed for interpretation and reuse.
- Point elevation and neighbourhood terrain summaries were calculated at 1-km, 2-km and 5-km spatial scales.
- Local terrain position is the settlement elevation minus the mean valid terrain elevation within a circular neighbourhood buffer.
- Positive values indicate settlement points above the surrounding terrain reference, while negative values indicate points below it.
- The same frozen procedure generated all three scales, which support scale-aware reuse without defining a universally preferred neighbourhood.
- Raw terrain products are not redistributed; released fields retain settlement-level measurements, validity information, provenance and licensing metadata.
- Copernicus DEM GLO-90 supplies an independent 2-km elevation-based layer for source-sensitivity assessment alongside, not instead of, OS Terrain 50.
3 Data Records
TEMPLAR Wales is a relational, machine-readable release linking 3,757 settlement records to lexical detections, environmental attributes and registry metadata through stable identifiers. Its records document reproducible construction, coverage conditions and interpretive boundaries across the released tables.
- The dataset contains 3,757 settlement records organised into four primary tables linked by stable project identifiers.
- The primary key templar_id identifies each settlement and links one settlement and environmental row to zero or more lexical detections.
- Four coincident-coordinate records are retained because they represent distinct upstream settlement records.
- The lexical table contains 1,350 detections across 1,294 settlements: 378 exact-token and 972 prefix-token matches.
- 2,463 settlements have no registered detection, while 1,238 have one and 56 have two; no settlement has more than two.
- Lexical detections are reproducible string-level annotations, not validated morphological analyses, language classifications or individual-name etymologies.
- Core hydrological, coastal, terrain and point-level CORINE fields cover all 3,757 settlements, whereas woody-cover availability declines across neighbourhood scales.Woody-cover summaries are available for 3,619 records at 500 m, 3,463 at 1 km and 3,216 at 2 km.
- Missing environmental values preserve documented coverage or eligibility failures rather than being replaced with zero.
4 Technical Validation
Technical validation found the released resource internally consistent, reproducible, and complete under its documented construction rules. Independent terrain products showed strong agreement, while validation did not establish environmental ground truth.
- Settlement frame and relational integrity: 3,757 settlement records passed identifier, name, location and settlement-class checks, with unique templar_id values and one-to-one environmental links.Repeated names and coincident locations were retained as distinct source observations rather than treated as automatic errors.
- Environmental coverage: 3,757 settlements had complete principal hydrological, coastal and terrain coverage, while woody-cover eligibility declined from 3,619 at 500 m to 3,216 at 2 km.These correspond to 96.3%, 92.2% and 85.6% of the settlement frame, respectively.
- Independent terrain-source agreement: Pearson r = 0.9999158 and MAE = 1.18 m for 2-km terrain reference measurements, with RMSE = 1.63 m and mean signed difference = 1.05 m.For local terrain position, agreement remained strong at Pearson r = 0.9950185 and MAE = 2.70 m.
- Release reconstruction: All six post-freeze release tests passed, and repeated export produced identical scientific-table hashes without changing identifiers, annotations, values or schema.The validation assesses integrity and reproducibility rather than retesting environmental hypotheses or establishing either elevation product as ground truth.
5 Usage Notes
TEMPLAR Wales is designed for flexible settlement-level reuse across toponymic, linguistic and environmental analyses. Users must preserve the dataset’s observational units, lexical-screening scope, source-specific environmental meanings and missing-value conventions.
- Resource structure: The four-table structure supports joint or independent use of settlement records, lexical annotations and environmental measurements.Stable identifiers support enrichment with additional settlement-level information.
- Joins and observational units: Use templar_id for internal joins, treating settlements.csv as the reference table and lexical_detections.csv as one-to-many.Directly joining detections can duplicate settlements with multiple matches and omit settlements without detections unless handled explicitly.
- Joins and observational units: The 3,757 observations are retained OS Open Names settlement records, not 3,757 unique place names; aggregation changes the observational unit and should be reported.The frame contains 3,328 normalised analytical-name groups and 3,755 coordinate clusters.
- Names and lexical interpretation: The analytical name is a frozen, project-defined processing field, while lexical detections are reproducible ETDE v1 string matches rather than verified etymologies.Exact-token and prefix-token matches should remain distinguishable when lexical specificity matters.
- Environmental interpretation: Environmental variables describe contemporary or product-specific conditions, not automatically historical landscapes, channel positions or past vegetation.Historical interpretation requires independent historical environmental sources.
- Missing environmental values: Missing environmental values must not be recoded as zero: zero means an eligible measured fraction of zero, whereas missing means no eligible released measurement.Coverage, eligibility and quality fields should accompany analytical subsets, especially when comparing woody-cover scales.
6 Data Availability
TEMPLAR Wales v1.0.0 is publicly deposited on Zenodo with the released environmental attributes file listed among the dataset contents. The cited deposition information defines the public dataset availability boundary.
- Dataset deposition: TEMPLAR Wales v1.0.0 is publicly available from Zenodo.The passage supplies the DOI prefix but not the complete DOI.
- Dataset contents: The deposited dataset includes environmental_attributes.csv.
7 Code Availability
Code supporting deterministic construction and validation is publicly available through the project reproducibility repository and archived materials. These materials support reconstruction when used with documented upstream sources, which are not redistributed.
- Code repository: Code for deterministic TEMPLAR Wales construction and validation is publicly available from the project reproducibility repository.
- Reproducibility materials: The repository includes dataset-construction code, lexical detection implementation, configuration and focused validation tests.The archived materials are intended to reconstruct and verify the released products.
- Source-data boundary: Raw third-party geospatial sources are not redistributed and must be obtained from their providers under applicable licensing terms.
Funding
The authors received no specific funding for this work.
- No specific funding was received for this work.
- The work is reported without dedicated project funding.
- The funding statement records no specific financial support.
Use of generative AI and AI-assisted technologies
OpenAI ChatGPT/Codex and GitHub Copilot assisted with software development, analysis workflow support, and manuscript drafting and editing; the authors reviewed and verified all AI-assisted outputs.
- OpenAI ChatGPT/Codex and GitHub Copilot assisted software development, analysis workflow support, and manuscript drafting and editing.
- The authors reviewed and verified all AI-assisted outputs.
- Numerical results and bibliographic records were checked against project evidence and source records.
- Scientific decisions, interpretation and conclusions remain the responsibility of the authors.
Competing Interests
The authors declare no competing interests.
- No competing interests were declared by the authors.
- The authors report no conflicts of interest.
- The competing-interests statement identifies no competing interests.