Source-linked AI summary
Gaia Early Data Release 3. Building the Gaia DR3 source list -- Cross-match of Gaia observations
F. Torra, J. Castañeda, C. Fabricius, L. Lindegren, M. Clotet, J. J. González-Vidal, S. Bartolomé, U. Bastian, M. Bernet, M. Biermann, N. Garralda, A. Gúrpide, U. Lammers, J. Portell, J. Torra
TL;DR
Gaia EDR3 requires cross-matching billions of onboard detections to physical sources while managing spurious events and evolving source identities. The paper describes this clustering pipeline and its improvements, including a reduction in spurious sources from 20.7% to 16.1%, while noting that close pairs remain unresolved in many cases.
Problem
The cross-match must identify onboard detections belonging to the same physical light source while handling spurious detections and difficult source configurations.
Method
The paper models the cross-match as clustering detection groups, assigning clusters to existing or new sources and preserving identifiers where possible.
Results
Spurious sources decreased from 20.7% to 16.1%, while the updated clustering improves results for high-proper-motion sources and variable stars.
Takeaways & Limitations
The XM supplies matched observations to Gaia’s astrometric and photometric pipelines and produces a cleaner, more stable source catalogue.
Takeaways & Limitations
The input precision is insufficient to detect and resolve close source pairs with many matched detections.
Abstract
from arXiv · showhide
The Gaia Early Data Release 3 (Gaia EDR3) contains results derived from 78 billion individual field-of-view transits of 2.5 billion sources collected by the European Space Agency's Gaia mission during its first 34 months of continuous scanning of the sky. We describe the input data, which have the form of onboard detections, and the modeling and processing that is involved in cross-matching these detections to sources. For the cross-match, we formed clusters of detections that were all linked to the same physical light source on the sky. As a first step, onboard detections that were deemed spurious were discarded. The remaining detections were then preliminarily associated with one or more sources in the existing source list in an observation-to-source match. All candidate matches that directly or indirectly were associated with the same source form a match candidate group. The detections from the same group were then subject to a cluster analysis. Each cluster was assigned a source identifier that normally was the same as the identifiers from Gaia DR2. Because the number of individual detections is very high, we also describe the efficient organising of the processing. We present results and statistics for the final cross-match with particular emphasis on the more complicated cases that are relevant for the users of the Gaia catalogue. We describe the improvements over the earlier Gaia data releases, in particular for stars of high proper motion, for the brightest sources, for variable sources, and for close source pairs.
1. Introduction
Gaia EDR3 extends the Gaia source list and presents a detailed cross-match process that links onboard detections to physical light sources. The process addresses large-scale processing, source identity changes, and complications affecting difficult source populations.
- Gaia EDR3 contains astrometry and photometry for more than 1800 million sources from the mission’s first 34 months.
- The cross-match identifies onboard detections belonging to the same physical light source and assigns each resulting cluster a unique source identifier.
- The full-scale EDR3 cross-match processes 78 billion detections associated with 2.5 billion sources over nearly three years.
- The shared transit-to-source match table supports the astrometric, photometric, and spectroscopic pipelines.
- Source identifiers normally persist from previous releases, but merge and split cases can require new identifiers as the detection set grows.
- The paper details spurious-detection rejection, sky partitioning, cluster modeling, source-list changes, validation, and future processing developments.
2. Input data
Gaia EDR3 cross-matching uses onboard detections, derived positions and brightnesses, and an evolving working source catalogue. The input stream is segmented and repeatedly processed, while detection unreliability and source-catalogue changes require tailored handling.
- Gaia EDR3 input data cover 1038 days from 25 July 2014 to 28 May 2017, divided into four processing data segments.
- The XM flow rejects spurious detections, partitions detections and sources into sky groups, clusters detections without source information, and assigns clusters to existing or new sources.
- The four segments contain 22.2, 8, 22.1, and 25.7 billion detections for DS-0 through DS-3, respectively.
- XM derives transit coordinates from the onboard AF1 reference acquisition pixel and encodes them in a unique transitId.
- Reference-pixel positions have errors of a few tenths of an arcsecond, depending on scan direction and magnitude.
- Because confirmation can be skipped for some onboard detections, spurious events—especially around bright sources and from cosmic rays—must be discarded first.
- Detection brightness, measured onboard in the sky mapper with an error of a few tenths of a magnitude, helps resolve close-object ambiguities and estimate magnitudes for new sources.
- The working catalogue is updated across processing cycles, and XM applies source-specific parameter treatments based on each source’s available solution and update history.
3. Detection classification
Gaia EDR3 first classifies onboard detections as genuine or spurious before cross-matching, using multiple criteria aimed at bright-source artefacts, cosmic rays, planets, odd profiles, and attitude problems. The cleaning removed 12 553 million detections, or 16.1% of all detections, while retaining safeguards for potentially genuine extended-source detections.
- Spurious detections can complicate clustering and create incorrect or artificial source assignments, so they are removed before cross-matching.
- Bright-source artefacts include diffraction-spike detections, phantom detections, cosmic rays, planetary-transit detections, odd flux profiles, and noisy-attitude detections.Diffraction-spike regions are identified using density maps and magnitude limits; other modules inspect saturation, AF-window signal, window samples, or attitude quality.
- 12 553 million detections, corresponding to 16.1% of all detections, were rejected in Gaia EDR3.The overall rejection percentage fell from 20.7% in Gaia DR2 to 16.1%; data segment 3 had a rejection fraction of about 5.1%.
- Spurious detections are concentrated along scanning-law caustics and dense Galactic-plane regions, whereas retained detections more closely trace dense source regions.The spurious-detection sky map shows narrow great circles and the Galactic plane; retained detections are concentrated in the Galactic plane and Magellanic Clouds, with less prominent scanning-law circles.
- Some genuine non-point-like detections may be classified as spurious, so external lists of extended objects, Solar System-nearby sources, and science alerts are admitted unconditionally.An additional image-quality module also evaluates surviving detections within clusters to prevent spurious detections from creating new sources.
4. Determination of isolated groups of detections
The pipeline converts preliminary detection-to-source associations into isolated, self-contained groups that can be clustered independently. A 5′′ candidate radius supports high-proper-motion recovery but increases ambiguity and group complexity, requiring recursive partitioning and source-candidate cropping.
- Sky partitioning and reprocessing with larger regions make the groups independently processable while avoiding direct-partition boundary effects.This organisation is essential because processing all Gaia detections as one job is computationally infeasible.
- Preliminary match candidates associate each detection with nearby working-catalogue sources whose positions are propagated to the detections’ mean epoch.Candidates are selected using a common angular-distance threshold.
- A 5′′ match radius is chosen to keep detections of high-proper-motion sources together while limiting excessively large groups.Later processing can use improved proper motions to regroup detections more precisely.
- 33% of detections have one source candidate, whereas 10% have more than seven source candidates within the 5′′ radius.Multiple candidates are more common in dense regions and around bright sources with spurious catalogue sources.
- A recursive link-traversal process groups match candidates and source candidates into closed groups or defers boundary-crossing candidates for later processing.Closed groups are self-contained; deferred candidates indicate links to candidates outside the current sky region.
- 90% of assembled-detection groups contain fewer than 95 detections, but groups in dense regions such as the Galactic centre can reach one million match candidates.A depth-first-search-based cropping algorithm discards selected links and recomputes groups to reduce the largest-group sizes.
5. Clustering model
The clustering model groups detections using catalogue-independent hierarchical clustering, extending the source model with proper motion when needed. Post-processing and classification address spurious detections and improve source solutions in challenging cases.
- Clustering model: The clustering model divides each detection group into subsets with similar characteristics, without using the existing source catalogue.Source-model parameters are determined solely from detections in each cluster.
- Clustering model: Ward dissimilarity minimises the increase in internal variance when two disjoint clusters are merged.It is based on the change in summed squared residuals and is non-negative definite.
- Clustering model: The model can include magnitude among weighted coordinates, alongside spatial components, when computing detection residuals.The cluster centre is chosen to minimise the weighted sum of squared residuals.
- Source model: Proper motion is incorporated with a linear time-dependent model, using mean position at the mean epoch and proper motion as its parameters.The time-dependent model is applied to clusters containing more than one detection.
- Source model: Parallax is not included in the current source model because it removes coordinate independence and is not considered essential for the current processing cycle.A parallax-inclusive model may be analysed in future releases with more detections.
- Examples and outcomes: The current-cycle solution merged detections into one high-proper-motion source and improved its proper-motion error by a factor of 27.The released source has µα∗ = 571.23 ± 0.04 mas yr−1 and µδ = −3691.49 ± 0.04 mas yr−1.
- Cluster classification: Rejected detections cannot enter subsequent pipelines, reducing new sources created from scratch and sources superseded by splits caused by spurious nearby clusters.This follows from applying a criterion similar to that used in image parameter determination.
- Cluster classification: About 162 million detections were removed by classifying about 96 million clusters for Gaia EDR3.About 80% of demoted clusters contained a single detection with several discarded windows.
6. Cluster-source assignment
Cluster-source assignment begins with candidate sources from the previous catalogue and resolves ambiguous links through a decision-tree search. Its priorities favour catalogue stability, controlled merging and distance-based fit quality.
- Assignment procedure: The clustering algorithm starts from scratch each cycle, using updated source parameters, attitude, calibrations, censoring and clustering models rather than previous match solutions.The source catalogue is not used during clustering itself.
- Source evolution: The new match solution can change source identification and astrometric or photometric parameters when new data alter the match.Some detections may also be assigned to different sources.
- Assignment procedure: Candidate sources are initially drawn from the common match-candidate subset across detections in each cluster.This produces the candidate cluster-source links for subsequent resolution.
- Contradiction resolution: A decision-tree algorithm resolves multiple candidate links, including contradictions where two clusters share the same source candidates.The tree explores alternative assignments to provide the final optimal match solution.
- Contradiction resolution: Candidate sources closer than 1′′.5 are considered during assignment, rather than the 5′′ radius used to create isolated detection groups.The reduced radius limits candidate cluster-source links at this stage.
- Resolution priorities: The resolution prioritises fewer new sources, more sources superseded by merging, and then lower accumulated cluster-to-source distance.These priorities stabilise identifiers, reduce potential spurious sources and select among remaining fits.
7. Results and validation
The Gaia EDR3 cross-match largely preserves and stabilizes the prior source list while resolving substantial source evolution, transit reassignment, and high-proper-motion cases. Validation shows improved treatment of faint and bright sources, close or ambiguous matches, and spurious detections, with remaining catalogue content also shaped by downstream processing.
- Source evolution: 65 255 million detections were matched to 2552 million sources, with 2286 million sources persisting and about 7.3% deleted.Of deleted sources, 57% corresponded to filtered spurious detections and 43% were removed during final cluster-source assignment.
- Validation scope: The Gaia DR2 source-list comparison is not determined by cross-matching alone, because astrometric, photometric, and internal-validation processing also affects which sources are released.This scope condition applies to the reported evolution of Gaia DR2 sources into Gaia EDR3 sources.
- Source evolution: 96.7% of detections matched existing sources, while about 11% of sources were new, largely because many new sources had only one matched transit.About 72.5% of these isolated new-source detections came from the new data segment.
- Transit evolution: 86.3% of matches were unambiguous, 11.9% had one ambiguous source, and 1.8% had more than one, stabilizing most cluster-source links.The source list retained most prior associations, although close sources, spurious sources, and dense regions could exchange linked transits and identifiers.
- Transit evolution: 84% of input sources retained all matched transits, whereas only 1.55% retained less than half and 0.95% lost all Gaia DR2 matched transits.Both very bright and faint sources lost more transits than medium-magnitude stars, although bright sources generally retained their assignments and gained transits from the added observation period.
- High proper motion sources: 2729 high-proper-motion sources remained in Gaia EDR3, about 26% fewer than in Gaia DR2, while faint-end counts decreased by about 85%.The improved clustering model produced more matched transits and more accurate parameters for high-proper-motion sources; source completeness also depends on later astrometric and filtering processes.
8. Conclusions and future improvements
Gaia EDR3's cross-match solution improves source-list stability and catalogue quality across several difficult cases, while retaining explicit scope limits and unresolved contamination from close neighbours. Future improvements focus on incorporating richer detection information and source parameters into clustering.
- Conclusions: 78 billion processed transits underpin the Gaia EDR3 cross-match solution, which supplies matched observations to downstream astrometric and photometric pipelines.The same cross-match solution supports subsequent Gaia DPAC processing.
- Conclusions: Bright-source detection classification was significantly improved, and cluster-quality classification prevents new sources when most cluster transits are dubious.These changes address problematic bright-source cases and low-quality transit clusters.
- Conclusions: A nearest-neighbour-chain clustering generalisation improves results for high-proper-motion sources and variable stars while reducing false parallaxes and close duplicated-source pairs.The separation limit for duplicated sources was reduced from 400 mas in Gaia DR2 to 180 mas in Gaia EDR3; some spurious sources remain, especially in dense regions.
- Conclusions: 97.5% of sources persist between Gaia DR2 and Gaia EDR3, indicating a more stable source list, although specific cases can still evolve because of processing improvements and spurious detections.Changes remain possible for high-proper-motion and variable sources, bright sources around G = 10, and very bright sources affected by multiple cluster-source links.
- Limitations: Source identifiers and surviving-source parameters can change between releases, so the authors recommend treating each release's identifiers as belonging to independent catalogues.A Gaia archive table is provided to trace Gaia DR2 sources into Gaia EDR3.
- Future improvements: Close source pairs remain difficult to resolve because the cross-match input precision is insufficient, leaving some spurious parallaxes and duplicated sources in Gaia EDR3.Future cross-matching may use multiple image peaks from IPD and incorporate parallax or other source parameters into clustering.
- Conclusions: 20.7% to 16.1%: the EDR3 cross-match reduces the number of spurious sources and is more stable than earlier solutions.The remaining spurious-source population is reduced more strongly in dense areas such as the Galactic centre, but it is not eliminated.
Appendix A: List of acronyms
The appendix lists acronyms used in the Gaia EDR3 cross-match paper, including mission, instrument, processing, release, and astrometric terminology.
- Acronyms: The acronym list expands Gaia processing and release terms such as DPAC, DR1, DR2, DR3, and EDR3.It also includes instrument and coordinate-direction abbreviations such as AF, BP, AC, and AL.
- Acronyms: The list defines Gaia observation terms including FoV, AF, BP, CCD, and IPD.These abbreviations describe fields of view, instruments, detector technology, and image-parameter determination.