Source-linked AI summary

HadISD: a quality-controlled global synoptic report database for selected variables at long-term stations from 1973--2011

Robert J. H. Dunn, Kate M. Willett, Peter W. Thorne, Emma V. Woolley, Imke Durre, Aiguo Dai, David E. Parker, Russ E. Vose

arXiv:1210.7191v1physics.ao-ph

TL;DR

HadISD addresses the need for accessible, quality-controlled sub-daily observations from the complex ISD archive, especially for studying local extreme events. The paper constructs and evaluates a global station dataset through compositing, reformatting, and staged automated quality control. The resulting archive contains over 6,000 station records, while important observational limitations remain.

  • Problem

    The ISD’s complex flatfile structure and variable prior quality control limit straightforward use of its high-frequency observations for local climate and extreme-event studies.

  • Method

    The authors composite duplicate station identifiers, convert selected observations to netCDF, and apply staged intra- and inter-station quality-control tests designed to preserve true extremes.

  • Results

    Over 6,000 individual station records and an audit trail are provided for HadISD, with most stations having less than 1 per cent of observations flagged.

  • Takeaways & Limitations

    The dataset supports investigation of high-frequency temperature, pressure, and humidity variations, individual extremes, and populations of extreme events over recent decades.

  • Takeaways & Limitations

    The distributional-gap and cloud-coverage checks rely on assumptions or conservative cross-checks that do not guarantee high-quality data for every variable or station.

Abstract

from arXiv · show

[Abridged] This paper describes the creation of HadISD: an automatically quality-controlled synoptic resolution dataset of temperature, dewpoint temperature, sea-level pressure, wind speed, wind direction and cloud cover from global weather stations for 1973--2011. The full dataset consists of over 6000 stations, with 3427 long-term stations deemed to have sufficient sampling and quality for climate applications requiring sub-daily resolution. As with other surface datasets, coverage is heavily skewed towards Northern Hemisphere mid-latitudes. The dataset is constructed from a large pre-existing ASCII flatfile data bank that represents over a decade of substantial effort at data retrieval, reformatting and provision. These raw data have had varying levels of quality control applied to them by individual data providers. The work proceeded in several steps: merging stations with multiple reporting identifiers; reformatting to netCDF; quality control; and then filtering to form a final dataset. Particular attention has been paid to maintaining true extreme values where possible within an automated, objective process. Detailed validation has been performed on a subset of global stations and also on UK data using known extreme events to help finalise the QC tests. Further validation was performed on a selection of extreme events world-wide (Hurricane Katrina in 2005, the cold snap in Alaska in 1989 and heat waves in SE Australia in 2009). Although the filtering has removed the poorest station records, no attempt has been made to homogenise the data thus far. Hence non-climatic, time-varying errors may still exist in many of the individual station records and care is needed in inferring long-term trends from these data. A version-control system has been constructed for this dataset to allow for the clear documentation of any updates and corrections in the future.

1 Introduction

HadISD addresses the difficulty of using a large, complex synoptic archive for sub-daily climate applications by providing accessible data and comprehensive quality control. The work targets reliable characterization of local extremes while recognizing unavoidable uncertainty in automated flagging.

  • The ISD is a rich global archive used to study climate variations, meteorological events, and historical climate impacts.
  • Automated quality control is essential because thousands of station records may each contain tens of thousands of observations.
  • Its annual ASCII flatfiles and synoptic-code format make the database difficult for non-experts to access and use.
  • HadISD applies a new, comprehensive quality-control process to support characterization of extreme events at specific locations.
  • Quality control must retain high- and low-frequency information for regional and local climate applications, while individual-observation homogenization remains a separate challenge.
  • The dataset covers 1973–2011, with future periodic updates planned and process outputs available alongside the final data.

2 Compositing stations

The station-compositing process resolves fragmented reporting identities by matching likely duplicates, manually assessing candidate sets, and merging records to improve completeness. The resulting composites remain subject to possible false merges.

  • Frequent station-identifier changes and simultaneous reporting under multiple identifiers fragment records and make compositing essential for record completeness.
  • Candidate station matches were assessed with hierarchical scores, graphical diagnostics, and manual review of anomalies, counts, and observing times.
  • 1504 likely composite sets were assigned as matches, comprising 3353 unique station IDs.
  • Assigned sets were merged by favoring longer records while using overlapping shorter records to infill missing elements.
  • Some assigned composites may incorrectly combine separate stations, particularly in densely sampled regions, although component identifiers are documented in metadata.

3 Selection and retrieval of an initial set of stations

The initial station set was selected to balance record length, temporal sampling, and global coverage, then converted from large ASCII files into hourly netCDF data. Reporting frequency strongly affects geographic coverage.

  • The ISD contains about 30,000 stations, but many report sparsely while almost 2,000 have records extending 60 or more years.
  • Station selection compared four climatology periods and four average reporting intervals to maximize spatial coverage.
  • Hourly reporting is concentrated essentially in northwest Europe and North America, whereas 3-hourly reporting provides a more globally complete distribution.
  • The source ASCII files were converted into hourly netCDF files containing selected mandatory and optional reporting variables.
  • Temperature and dewpoint reports containing both measurements were preferred for reliable humidity studies, even when they were not closest to the full hour.

4 Quality control steps and analysis

HadISD combines robust distributional methods, staged multi-level flagging, and inter-station checks to identify poor observations while reducing the risk of removing genuine extremes. The tests target diverse station and variable-specific failure modes.

  • The automated procedure was necessary because a fully sampled hourly record can exceed 340,000 observations and more than 6,000 candidate stations exist.
  • Robust IQR and MAD statistics were used instead of standard deviation where pervasive errors could inflate spread and make tests too conservative.
  • Flagging is multi-level: some checks remove observations immediately, while most retain flags and exclude flagged data from later threshold derivations.
  • Distributional gap checks can identify implausible secondary populations, including Yokosuka temperatures apparently recorded in Fahrenheit below 0 °C during winter.
  • The climatological outlier check examines observations farther from a distribution’s center than both a gap and a threshold value.
  • The ordered procedure progressively removes bad data, repeats selected intra- and inter-station checks, and allows tentative flags to be reinstated after neighbor comparisons.

4.1 QC tests

The QC pipeline combines duplicate detection, isolation and distribution checks, spike and streak tests, and anomaly screening to identify suspect observations while preserving plausible extremes.

  • Duplicate and isolation checks: 83 stations were removed after duplicate-pair and patchy-record checks, leaving 6103 stations for subsequent QC.Duplicate groups were assessed using match statistics, reporting frequencies, separation distance, and time series.
  • Duplicate and isolation checks: Short clusters of up to 6 hours within 24 hours, separated from other data by at least 48 hours, are flagged for key variables.The check applies separately to temperature, dewpoint temperature, and sea-level pressure.
  • Distributional and anomaly checks: Frequent-value screening can flag genuine observations near the distribution mean and miss pervasive anomalies confined to a few years.The method was retained because alternatives were considered computationally inefficient.
  • Distributional and anomaly checks: The distributional-gap check can capture mixed-unit data, including winter temperatures below 0 °C apparently recorded in Fahrenheit, but cannot choose between two equally sized populations.This extension compares calendar-month distributions across years to identify separated secondary populations.
  • Spike and streak checks: A repeated-streak test flags years with more than five times the median annual frequency of streaks longer than 10 consecutive elements.The criterion was added after development revealed unusually frequent short streaks in some station years.
  • Spike and streak checks: Spike thresholds are station- and month-specific, using first differences and six times the rounded-up interquartile range, with minimum thresholds of 1 °C or hPa.Separate critical values are calculated for one-, two-, and three-hourly differences to account for seasonal variability.

4.1.11 Test 11: temperature and dewpoint temperature cross-check

The QC tests cross-check related meteorological variables, logical cloud observations, temporal variability, neighbours, and station distributions to distinguish errors from real events.

  • Cross-variable checks: Humidity-related checks target supersaturation, wet-bulb reservoir drying, and dewpoint cutoffs at temperature extremes.The tests account for physically plausible exceptions and use precipitation, fog, and cloud-base information where available.
  • Cross-variable checks: If dewpoint exceeds temperature, dewpoint is flagged; if this affects at least 20 per cent of a month, the whole month is flagged.Extended dewpoint-equals-temperature streaks are also assessed with precipitation and fog information to reduce dry bias.
  • Cloud and distribution checks: Cloud checks use six logical tests, but ceilometer observations may under- or over-estimate cloud coverage and generally cannot guarantee a high-quality record.The conservative checks are intended to identify glaring inconsistencies rather than establish complete cloud-data quality.
  • Cloud and distribution checks: Whole months are flagged when normalized within-month variance exceeds the station’s historical median by more than 8 IQR for temperature and dewpoint or 6 IQR for pressure.The thresholds increase to 10 and 8 IQR respectively when reporting frequency or resolution is reduced.
  • Inter-station checks: Nearest-neighbour screening uses up to ten stations within 500 m elevation and 300 km, requiring at least three valid neighbours before application.Neighbour selection seeks representation from all four surrounding quadrants to reduce geographic bias.
  • Inter-station checks: Tropical low-pressure extremes can still be removed near coastlines when station spacing is insufficient to distinguish storm cores from neighbour outliers.The procedure reduces removal of real storms but remains vulnerable just after landfall in dense station networks.

4.2 Test order

Tests are ordered so early intra-station checks remove obvious errors before distributional analysis, followed by inter-station checks and a final cleanup pass.

  • Initial pass: Intra-station tests run before the neighbour check so distribution-based thresholds are not biased by glaring errors.The sequence is chosen for both computational convenience and the quality of later thresholds.
  • Initial pass: After flags are applied, suspect observations are masked while flagged values are retained separately for possible later retrieval.The main stream records flagged indicators rather than treating flagged observations as ordinary data.
  • Rerun and cleanup: Spike and odd-cluster tests are rerun on masked data, followed by another neighbour check and final bad-month cleanup.The rerun can identify new spikes or clusters exposed after earlier bad data are removed.

4.3 Fine-tuning

Fine-tuning used documented extreme events and known database issues to adjust the quality-control tests, while checking that valid extremes were retained and bad observations removed.

  • 167 British Isles stations and three documented extreme events were used to fine-tune critical and threshold values for the QC tests.The events were the August 2003 European heat wave and the October 1987 and January 1990 storms.
  • The revised tests retained the 1987 storm’s low-pressure minimum while removing two anomalous observations from one station.Previously, many valid observations around the minimum had been flagged; the two removals were identified by the spike test.
  • Suspect data discovered by users can be investigated, with amendments to the raw data or QC suite applied where possible.
  • Known issues affecting four included data cases were successfully identified and removed, while one reporting-accuracy error remained undetectable by the QC suite.A separate station-compositing issue was solved during the compositing process.

5 Validation and analysis of quality control results

Validation across global extreme events and station-level QC results indicates that HadISD generally retains useful signals while removing flagged observations, though isolated compositing and measurement issues remain.

  • The QC procedure was tested against global extreme events to assess extreme-value retention, identify limitations, and avoid over-tuning to one region.
  • The Katrina pressure-system passage remained characterisable even though the neighbour check removed observations where nearby stations recorded different simultaneous pressures.
  • The 1989 Alaskan cold snap remained clearly analyzable despite flagged and removed observations, with HadISD minima of −58.9 ◦C at McGrath and −46.1 ◦C at Fairbanks.The McGrath value was 0.5 ◦C warmer than the documented record, and sub-daily sampling may miss true minima.
  • The Australian heat-wave signal remained evident despite flagged observations, with HadISD maximum temperatures of 44.0 ◦C in Adelaide and 46.1 ◦C in Melbourne.The plots showed the exceptionally warm period in both average daily and synoptic-resolution data.
  • Most stations had less than 1 per cent of observations flagged, and the authors judged the QC adequate and unlikely to be over-aggressive.Flagging patterns were often geographically distinct and commonly followed geopolitical rather than physically plausible patterns.
  • Random checks found no obvious discontinuities in 20 composite stations, but isolated compositing cases may still degrade data quality.Users were encouraged to report such issues so amendments could be made where possible.

6 Final station selection

Final station selection applied additional completeness, reporting-frequency, and post-QC quality criteria to produce a long-term climate-monitoring network, while excluding many incomplete records.

  • HadISD’s “.clim” versions apply minimum temporal completeness, reporting-frequency, and QC-flagging criteria beyond the records available in “.all” versions.
  • 1234 stations failed record-completeness criteria, while 689 failed because their first or last observation occurred too late or early.Large data gaps caused a further 626 stations to fail.
  • Record-completeness criteria could not be relaxed without including records judged too incomplete for end-users.Other rejections reflected insufficient post-QC data for one or more variables.

7 Dataset nomenclature, version control and source code transparency

HadISD.1.0.0.2011f provides distinct all-station and climate-quality station versions, with versioning, metadata, and source code intended to document dataset changes and support reuse.

  • HadISD.1.0.0.2011f includes 6103 quality-controlled stations in .all and 3427 stations meeting climate-selection criteria in .clim.
  • Table 7 documents data precision and reporting intervals by month for the .all and .clim station sets.Months without data are excluded, while sparse months may have undetermined accuracy or reporting intervals.
  • Table 8 defines station inclusion ranges and notes that wind and cloud quality did not determine station exclusion.The filtering focused on temperature, dewpoint, and pressure data.
  • Version labels identify updates, specifications, preliminary or final status, and dataset types such as .all and .clim.Major releases are described in peer-reviewed publications, while smaller changes are documented through version numbers, websites, or readme files.
  • The IDL source code is released alongside HadISD for users to copy and use, although no support service is provided.

8 Brief illustration of potential uses

Examples demonstrate how HadISD’s sub-daily coverage supports analyses of diurnal temperature structure and globally evolving temperature patterns at specific observation times.

  • The examples illustrate capabilities of the sub-daily dataset beyond monthly or daily holdings.
  • Median diurnal temperature range: Daily diurnal temperature range was calculated from 24-hour maximum and minimum temperatures when at least four observations spanned 12 hours.
  • Median diurnal temperature range: Highest diurnal temperature ranges occur in arid or high-altitude regions, with altitude-related contrasts evident in southwestern China.
  • Median diurnal temperature range: Seasonal diurnal temperature-range differences are prominent in dense networks, including summer increases in Europe and central Asia and monsoon-linked decreases in India and sub-Saharan West Africa.
  • Temperature variations over 24 hours: Temperatures from 6103 quality-controlled stations on 28 June 2003 show the global progression of maxima through successive 00:00, 06:00, 12:00, and 18:00 UT snapshots.The example also identifies one Western Canada outlier missed by the quality-control suite.

9 Summary

The paper develops HadISD as a user-friendly, quality-controlled long-term station subset of ISD. It provides broad station coverage, a climate-quality subset, and versioned archiving for reproducible use.

  • HadISD converts the large ISD synoptic database into a user-friendly netCDF dataset with an alternative quality-control suite.The workflow composites duplicate stations, selects climate-applicable records, and applies intra- and inter-station checks.
  • The final dataset contains over 6000 station records from 1973 to 2011 with near-global coverage and over 3400 long-term climate-quality stations.
  • A version-control and archiving system identifies the HadISD version in use and records future methodological changes.

Copyright statement

The work is distributed under a Creative Commons Attribution 3.0 License with author copyright, and the paper acknowledges contributors and funding support.

  • HadISD is distributed under a Creative Commons Attribution 3.0 License together with author copyright.
  • The authors acknowledge Neal Lott and two anonymous referees for reviews that improved the manuscript and dataset.
  • Funding and institutional support came from the Joint DECC/Defra Met Office Hadley Centre Climate Programme, NCDC, PHEATS, NCAR, and related placements and assistance.
Loading 1210.7191v1…