Source-linked AI summary

Measuring the Installed Base: Nordic Health Dataset Catalogues Against HealthDCAT-AP Release 7

Fabio Rovai

arXiv:2608.27720v1cs.DLcs.CY

TL;DR

HealthDCAT-AP had been designed and validated on curated examples but not measured against live catalogues. This paper conducts that installed-base measurement across Nordic catalogues and finds pervasive gaps in mandatory metadata and vocabulary binding, within stated portal and temporal limits.

  • Problem

    The paper addresses the missing measurement of live catalogue conformance to the health metadata profile required for cross-border dataset discovery.

  • Method

    The study performs a dated census of 11 Nordic catalogues harvested by the European data portal, deriving mandatory requirements from official Release 7 shapes.

  • Results

    None of 2,811 health-themed descriptions satisfied all eight mandatory properties, and three properties were present on exactly zero records.

  • Takeaways & Limitations

    The comparison shows the health gap does not sit above sound generic practice: no Nordic catalogue reaches the MQA top band and 5 of 11 report zero per cent DCAT-AP compliance.

  • Takeaways & Limitations

    Coverage is limited to catalogues harvested by the European data portal, the theme filter makes 2,811 a floor, and every figure is a single dated observation.

Abstract

from arXiv · show

The European Health Data Space requires member states to publish machine readable descriptions of the health datasets available for secondary use, and the European Commission publishes HealthDCAT-AP as the metadata profile those descriptions are meant to satisfy. The profile has been designed and validated against curated examples, never against the catalogues already live. We report that measurement for the Nordic region. On 25 August 2026 the 11 Nordic national catalogues harvested by the European data portal held 2,811 dataset descriptions carrying the EU health theme, and none satisfies all eight properties HealthDCAT-AP Release 7 makes mandatory on a dataset. Three of the eight are present on exactly zero records across five countries. Set beside the portal's own quality assessment, which validates DCAT-AP and never mentions the health profile, this is not a health extension skipped on top of sound generic practice: no Nordic catalogue reaches the assessment's top rating band, 5 of the 11 are reported at zero per cent DCAT-AP compliant, and the properties surviving in both layers are the ones a human types into a form, not the ones needing a value bound to a controlled vocabulary. Two further results follow. Finland contributes 2,259 descriptions to the European portal of which 1,146 carry a theme, and not one uses the EU theme authority vocabulary, so a European health filter returns no Finnish dataset at all. Separately, the authority namespace answers HTTP 200 with a well formed empty document for terms it never defined, letting 1,238 datasets across the wider portal carry theme IRIs that resolve to nothing while passing any status code check. We publish the vocabulary, the shapes and the harvesters, record every verdict as a dated observation rather than a property of the dataset, and report the five errors this discipline caught before publication.

1 Introduction

This paper measures how far Nordic health dataset catalogues are from HealthDCAT-AP Release 7, rather than relying on design validation against curated examples. It finds widespread gaps in mandatory metadata, generic DCAT-AP conformance, vocabulary binding, and catalogue capability.

  • Research question and findings: 2,811 health-themed descriptions were assessed against HealthDCAT-AP Release 7, and none satisfied all eight mandatory properties.Three of the eight properties were present on exactly zero records.
  • Research question and findings: 1,146 Finnish descriptions carry a theme, but none is bound to the EU theme authority.Finland therefore contributes themed descriptions without the vocabulary binding used for European health discovery.
  • Research question and findings: 1,238 datasets across the wider portal carry theme IRIs that resolve to nothing while passing status-code checks.The authority service returns HTTP 200 with a well-formed empty RDF document for undefined terms.
  • Research question and findings: The paper publishes an OWL vocabulary, three SHACL layers, and dated conformance observations distinguishing absence from ungrounded or portal-computed values.The authors also report five errors caught before publication.

2 Background

HealthDCAT-AP extends DCAT-AP with health-specific metadata and SHACL validation, but prior work measured generic portal quality, DCAT-AP, or curated examples rather than live health-profile conformance. This paper fills that installed-base measurement gap.

  • Profile and prior work: DCAT-AP constrains DCAT dataset descriptions and binds selected properties to controlled vocabularies, while HealthDCAT-AP adds health-specific requirements.The health extension covers access responsibility, health category, structured-data status, and governing legislation.
  • Profile and prior work: Release 7 is the current target, aligned with DCAT-AP 3.0.1; measuring against deprecated Release 5 would target the wrong specification.The older specification location was decommissioned, and Release 6 was deprecated on 24 April 2026.
  • Profile and prior work: Release 7 makes eight properties mandatory on dcat:Dataset in both public and restricted layers.The listed public-layer requirements include access rights, applicable legislation, identifier, access body, distribution, health category, theme, and structured-data status.
  • Profile and prior work: Earlier large-scale studies assessed generic completeness, retrievability, licence openness, format accuracy, or DCAT-AP rather than domain-extension conformance.Their cross-portal generality is explicitly bounded by the inability to ask health-specific questions such as which access body governs a dataset.
  • Profile and prior work: The portal’s Metadata Quality Assessment scores five dimensions and four rating bands, but its published methodology does not mention HealthDCAT-AP.HealthDCAT-AP had been designed and validated on curated real-world examples, not measured against live national catalogues.
  • Profile and prior work: The paper’s specific contribution is measuring the health extension against the installed base of live catalogues.Generic portal quality and DCAT-AP conformance were already measured continuously and publicly.
  • Profile and prior work: The Commission’s validator checks one record before submission, whereas this work analyzes dated observations across many catalogues for bodies that can change harvest mappings.The two tools serve different operational purposes.

3 Method

The study measures what European users can see through the portal by harvesting Nordic catalogues, applying a conservative health-theme filter, and deriving requirements from official shapes. It verifies findings through independent computations and validation layers while recording scope limits.

  • Data and scope: 1,908,938 dcat:Dataset nodes formed the portal substrate, covering 11 harvested catalogues across Sweden, Denmark, Norway, Finland, and Iceland.The portal is used because it represents what the rest of Europe can actually see.
  • Data and scope: The mandatory requirement set was derived from published shapes rather than transcribed manually.A committed registry is checked for drift against the derivation procedure.
  • Data and scope: The health-themed selection used dcat:theme equal to the authority term HEAL, so 2,811 is a floor rather than an estimate of all Nordic health datasets.Datasets without that term or using local vocabularies are excluded by construction.
  • Data and scope: A national catalogue publishing no DCAT is invisible to the portal-based measurement, so Findata was harvested separately through public read endpoints.The study treats that structural invisibility as a finding and measures Findata directly.
  • Verification: 18 headline figures were computed twice, using harvested vectors and SPARQL over the emitted graph, and all agreed.The Swedish portal was independently checked at 2,418 health descriptions as a known-answer case.
  • Verification: 21,431 SHACL violations were reported, including 12,975 mandatory-property absences and 8,433 health-specific absences.The two classes overlap by design because health-specific absences are also mandatory absences.
  • Verification: An independent validation engine checked the OWL core, requirement registry, scheme registry, and each SHACL layer alongside the pyshacl run.

4 Results

Across the Nordic catalogues, HealthDCAT-AP Release 7 conformance is absent, while generic DCAT-AP quality is also incomplete. Missing vocabulary binding blocks Finland from European health discovery, and undefined authority terms undermine status-code validation.

  • HealthDCAT-AP conformance: 0 of 2,811 health-themed descriptions satisfy all eight mandatory HealthDCAT-AP Release 7 properties.Three properties are present on exactly zero descriptions across five countries and 11 catalogues.
  • HealthDCAT-AP conformance: Title and description appear on all 2,811 descriptions, and publisher appears on 2,785, whereas controlled-vocabulary, legal-identifier, and institutional-registration properties are largely absent.The four properties carrying values are those already known to DCAT-AP.
  • Theme vocabulary binding: Finland contributes 1,146 themed descriptions, but none uses the EU data theme authority, so the European health filter returns zero Finnish datasets.The catalogue’s own search returns 57 health results; the mismatch is a one-property harvest-mapping fix.
  • Portal quality assessment: The best Nordic MQA score is 300 of 405, none reaches Excellent, and 5 of 11 catalogues report 0% DCAT-AP compliance.Sweden scores 177 with 27% DCAT-AP compliance, while Denmark scores 287 with a 0% compliance indicator.
  • Authority-service defect: 1,238 datasets across the wider portal use 22 undefined theme terms that return HTTP 200 and empty RDF documents.A status-code-only validator therefore treats undefined vocabulary terms as healthy; the measurement instead parses the response content.
  • Findata schema capability: Findata has source fields for 6 of 8 mandatory properties, but exactly 2 lack sources: applicable legislation and the health data access body.Its catalogue contains usage conditions but no resolvable legal identifier, and cannot state that Findata itself is the responsible access body.

5 Recording verdicts as observations

The paper models conformance as dated, attributed observations rather than dataset properties, distinguishing absence, ungrounded values, and portal-computed values. Requirements and vocabulary rules are represented as data so profile changes regenerate validation inputs.

  • Missing healthdcatap:hdab is recorded as a false observation with a controlled reason, not inferred from silence.The model distinguishes absent, present but ungrounded, and publisher-unsupplied values because they require different remedies.
  • A boolean dataset-level conformance field cannot distinguish absence, ungrounded values, and publisher-unsupplied values.
  • Conformance observations bind a description, requirement, verdict, and date.
  • Vocabulary binding observations test whether supplied values belong to the profile’s declared scheme.This makes Finnish records with unresolvable theme values expressible as a distinct failure.
  • Requirements are generated from official shapes, while schemes declare their own namespace, closure, and member-count rules.Changing the profile release changes a generated file rather than a hard-coded string in the validation code.

6 Discussion

The discussion situates the Nordic measurement against proposed regional infrastructure, clarifies its narrow catalogue scope, and documents methodological corrections and boundaries. It argues that missing discovery metadata and partial generic conformance constrain implementation sequencing without implying weak national health systems.

  • 6.1 Two Nordic flagship documents, one missing layer: Neither 2026 Nordic infrastructure proposal names DCAT, DCAT-AP, or HealthDCAT-AP, focusing instead on dataset-internal standards and content models.The observation concerns the texts’ scope, not what their authors know or considered.
  • 6.2 What this does not say: The measurement covers machine-readable catalogues harvested by the European data portal, not national health-data capability or institutional performance.For Finland, the measured catalogue axis is zero, with a mapping change—not an institutional change—as the stated remedy.
  • 6.2 What this does not say: Five of eleven catalogues report zero per cent DCAT-AP compliance, so HealthDCAT-AP rollout cannot assume a sound generic foundation.
  • 6.3 Errors we made: A regex initially included five non-mandatory properties because SHACL property shapes are nested blank nodes; RDF parsing corrected the requirement set.
  • 6.3 Errors we made: Portal-computed metrics inflated one property from 0.96% to 98.40%, leading the study to exclude the portal’s metrics graphs.
  • 6.3 Errors we made: A cross-check reduced an erroneous undefined-theme count of 4 186 to the true figure of 1 238.
  • 6.3 Errors we made: HTTP 200 responses nearly served as vocabulary-membership evidence, which would have reported zero squatted IRIs.
  • Scope and limitations: Coverage excludes catalogues not harvested by the European data portal, and health-theme selection makes 2 811 a floor.The MQA comparison was retrieved two days later, scores whole catalogues, and uses the portal’s own points scheme.

7 Conclusion

The Nordic census finds no fully conformant HealthDCAT-AP Release 7 descriptions and shows that generic DCAT-AP quality is also weak. It additionally identifies Finnish theme binding and EU vocabulary dereferencing defects as distinct, potentially cheaper-to-fix discovery failures.

  • 2 811 health-themed descriptions across 11 Nordic catalogues yielded zero records satisfying all eight mandatory HealthDCAT-AP Release 7 properties.Three properties were present on exactly zero records across five countries.
  • Five of eleven catalogues report zero per cent DCAT-AP compliance, and none reaches the portal assessment’s top rating band.The properties surviving both layers are manually entered fields, while shared failures require controlled vocabularies, legal identifiers, or institutional registration.
  • Finland publishes 1 146 themed descriptions with none bound to the EU theme authority, removing Finnish datasets from European health discovery.The paper identifies this as a one-property harvest-mapping change.
  • 1 238 datasets carry undefined theme IRIs because the EU authority returns HTTP 200 and an empty document for unknown terms.
  • The vocabulary, SHACL shapes, harvesters, generated requirement registry, and full build report are publicly released for rerunning the dated measurement.

Competing interests

The authors disclose a commercial interest in ontology engineering and data governance consultancy while releasing the measurement artefacts openly, with derived mapping artefacts separately licensed for commercial use.

  • The author directs a company selling ontology-engineering and data-governance consultancy and training to public-sector buyers in the measured domain.The code, vocabulary, and shapes are open-licensed, while derived catalogue mappings have a separate commercial-use licence.

Data and code availability

The paper makes its vocabulary, validation shapes, harvesters, requirement registry, and build report publicly available. Results are generated from committed measurement artefacts, while harvested payloads are not redistributed.

  • The vocabulary, SHACL shapes, harvesters, generated requirement registry, and full build report are public.
  • Every table and figure is generated from committed measurement artefacts, allowing reruns to propagate into the paper without manual transcription.
  • Harvested payloads are not redistributed because the Findata metadata has no stated reuse licence, while portal vectors can be regenerated from its public SPARQL endpoint.The portal vectors are regenerable in about twenty minutes.

A The measurement, as pseudocode

The measurement pipeline derives its mandatory properties from RDF-parsed shapes, applies guarded census and vocabulary checks, and cross-validates results across representations. Its design targets shortcuts that previously produced wrong counts, false vocabulary membership, incomplete caches, or non-terminating checks.

  • The pipeline runs without credentials against three public endpoints: the European portal SPARQL service, the authority vocabulary host, and the Findata catalogue API.
  • Algorithm 1 counts datasets matching the health theme across catalogues and commits a catalogue only after every property query returns.The census also restricts the graph to publisher-supplied data, excluding the portal’s named quality-assessment graphs.
  • Algorithm 2 derives the mandatory property set from RDF-parsed SHACL shapes rather than hand transcription or line-oriented matching.Nested blank-node property constraints make simpler matching approaches assign cardinalities to the wrong paths.
  • Algorithm 3 tests whether a theme IRI is defined by parsing the authority response, not by treating HTTP 200 as proof of vocabulary membership.The authority can return a well-formed empty document for 22 undefined terms, so a body parse is regression-tested.
  • The pipeline parses emitted Turtle, recomputes headlines through set-based and SPARQL routes, and aborts when the routes disagree.Three shape layers run after these checks.
  • The implementation compares counts over a single node target because a self-join over 22,488 reified observations is quadratic and does not terminate at the reported scale.The reported run used Python 3.13, rdflib 7.x, pyshacl, and a graph of 221,431 triples.
Loading 2608.27720v1…