Source-linked AI summary

ATLAS data quality operations and performance for 2015-2018 data-taking

ATLAS Collaboration

arXiv:1911.04632v2physics.ins-dethep-ex

TL;DR

ATLAS must scrutinise data from its large detector to identify hardware or software issues before physics certification. The paper presents Run 2 monitoring and assessment procedures, including online and offline checks, and reports that 95.6% of recorded 13 TeV pp collision data were certified for physics analysis.

  • Problem

    ATLAS data must be checked for hardware or software integrity issues before certification for physics analysis.

  • Method

    The paper presents ATLAS data-quality monitoring and assessment across detector monitoring, reconstructed event characteristics, conditions updates, and physics-data certification.

  • Results

    95.6% of recorded 13 TeV proton–proton collision data were certified for physics analysis.

  • Takeaways & Limitations

    Continuous review and development of data-quality and operating procedures yielded increasing efficiency for good-quality physics data during Run 2.

  • Takeaways & Limitations

    Some reconstruction stages can be delayed by problematic database uploads, excessive detector noise, or reconstruction failures.

Abstract

from arXiv · show

The ATLAS detector at the Large Hadron Collider reads out particle collision data from over 100 million electronic channels at a rate of approximately $100$ kHz, with a recording rate for physics events of approximately 1 kHz. Before being certified for physics analysis at computer centres worldwide, the data must be scrutinised to ensure they are clean from any hardware or software related issues that may compromise their integrity. Prompt identification of these issues permits fast action to investigate, correct and potentially prevent future such problems that could render the data unusable. This is achieved through the monitoring of detector-level quantities and reconstructed collision event characteristics at key stages of the data processing chain. This paper presents the monitoring and assessment procedures in place at ATLAS during 2015-2018 data-taking. Through the continuous improvement of operational procedures, ATLAS achieved a high data quality efficiency, with 95.6% of the recorded proton-proton collision data collected at $\sqrt{s}=13$ TeV certified for physics analysis.

1 Introduction

ATLAS must identify detector and software issues before data can be certified for physics analysis. During Run 2, continually improved data-quality monitoring and operational procedures increased the efficiency of good-quality data.

  • Over 100 million electronic channels span multiple detector subsystems, so missing or compromised subsystem data can affect physics-analysis quality.Examples include increased detector-electronics noise and high-voltage trips in detector modules.
  • Data-quality procedures identify affected time periods and mark them appropriately before data are provided for physics analysis.
  • During 2015–2018, ATLAS continually revised data-quality monitoring procedures and detector operational practices during Run 2.The paper presents the procedures and the resulting data-quality performance.
  • Data-quality efficiency increased year by year as monitoring and operational procedures improved.

2 ATLAS detector

The ATLAS detector combines tracking, calorimetry, and muon systems around the collision point, with trigger and reconstruction systems selecting and characterising recorded events. Physics-object monitoring therefore requires expertise across detector subsystems and reconstruction algorithms.

  • ATLAS covers nearly the entire solid angle with an inner tracking detector, calorimeters, and a muon spectrometer.
  • The inner detector provides charged-particle tracking for |η| < 2.5 in a 2 T axial magnetic field.It includes the pixel detector, insertable B-layer, and semiconductor tracking detector.
  • The calorimeter system covers |η| < 4.9 and combines electromagnetic and hadronic calorimetry using liquid-argon and scintillator technologies.
  • The muon spectrometer uses trigger and precision tracking chambers to measure muon deflection in superconducting-toroid magnetic fields.Precision chambers cover |η| < 2.7, while the muon trigger system covers |η| < 2.4.
  • The L1 trigger accepts events below 100 kHz, while the HLT reduces this rate to about 1 kHz for complete physics-event recording.Trigger selections use detector-subsystem information and physics-object or event-level criteria.
  • Monitoring reconstructed physics objects requires combined performance experts to work with dedicated subsystem and trigger experts.

3 Data quality infrastructure and operations

Run 2 data-quality operations combine detector monitoring, conditions management, event streams, and staged online/offline assessment. The workflow tracks data-taking conditions and supports certification of luminosity blocks for physics use.

  • Run 2 included standard 25 ns pp operation alongside special configurations such as heavy-ion, lower-energy, and low-µ periods.The standard 25 ns configuration followed initial 50 ns pp data-taking in July 2015.
  • A fill is an uninterrupted period with circulating beams, while a run is a continuously recorded dataset divided into luminosity blocks.A luminosity block generally lasts 60 s and has constant luminosity, detector and trigger configuration, and data-quality conditions.
  • The mean interactions per bunch crossing, µ, is calculated from per-bunch instantaneous luminosity using µ = Lbunch × σinel / fr.For 13 TeV collisions, σinel is taken as 80 mb and fr = 11245.5 Hz.
  • Detector conditions are stored with intervals of validity in the conditions database, linking status, configuration, and calibration information to data-taking periods.
  • Online monitoring uses DCS, GNAM, trigger and data-acquisition information, with monitoring data distributed through the Information Service and displayed by DQMD and OHP.
  • The DQ workflow starts with real-time monitoring, followed by prompt reconstruction of Express and calibration streams and then broader offline assessment.Conditions-database updates can be made after the first offline pass before bulk reconstruction.
  • The Debug stream preserves events that pass L1 but encounter HLT or DAQ errors, and represented less than 10^-7 of the recorded dataset from 2015 to 2018.These events are not available in other streams and are considered by physics analyses to avoid missing potentially interesting events.

4 Data quality monitoring and assessment

ATLAS combines real-time and offline monitoring with expert review, database defects, and calibration updates to assess data quality before physics use. The workflow supports early issue mitigation, final certification, and later improvement through reprocessing.

  • Certification: The Good Runs List certifies luminosity blocks for physics analyses and filters compromised data through integrated analysis tools.It is created by querying data-quality status in the defect database.
  • Defect recording: Primary defects document non-nominal detector conditions, while virtual defects determine whether those conditions are tolerated or reject data from analysis.Defects can evolve across processing iterations as detector understanding and cleaning improve.
  • Defect recording: Automatic defect assignment uses DCS archive data to flag suboptimal hardware status, including below-threshold magnet currents or incomplete RPC high-voltage ramp-up.The DCS defect calculator uploads primary defects for the corresponding luminosity-block intervals.
  • Online monitoring: Real-time control-room monitoring identifies serious detector issues quickly, while Express-stream checks enable corrective condition updates before physics-stream processing.Monitoring includes display tools and compatibility tests against reference histograms.
  • Calibration loop: The calibration loop combines database updates with primary data-quality review and typically lasts ∼48 hours, or 24 hours during busy periods.Updates include detector calibration, alignment, and beam-spot information before bulk reconstruction.
  • Reprocessing: Reprocessing repeats data-quality validation and can improve efficiency by applying updated conditions, such as more granular LAr-noise cleaning.Archived monitoring histograms and DQMF results support validation across processing iterations.

5 Data quality performance

ATLAS evaluated data quality using luminosity-weighted fractions of recorded data suitable for physics analysis, excluding nonstandard periods. Continuous monitoring, maintenance, software development, and incident response improved Run 2 performance, yielding 95.6% overall efficiency for the standard 13 TeV dataset.

  • Efficiency definition: DQ efficiency is the luminosity-weighted fraction of good-quality recorded data during stable beams intended for physics analysis.Machine commissioning, calibration runs, low-bunch fills, and bunch spacings greater than 25 ns are excluded.
  • Overall performance: 95.6% of the standard Run 2 √s = 13 TeV pp dataset was certified as good for physics analysis, corresponding to 139 fb−1.Yearly efficiencies increased from 88.8% in 2015 to 97.5% in 2018.
  • Sources of data loss: One-off incidents often produced the largest DQ inefficiencies, while recurring low-level issues were mitigated through infrastructure development and improved understanding.Countermeasures were introduced after significant incidents to reduce the risk of recurrence.
  • Subsystem performance: The pixel detector’s 2015 IBL cooling instability caused 0.2 fb−1 of rejected data and a 6% DQ inefficiency.Data were taken with the IBL powered off while cooling parameters were optimized.
  • Subsystem performance: Pixel read-out desynchronization affected 2015–2017, but firmware, software, and hardware improvements reduced its impact over time.SCT read-out problems became negligible after firmware improvements at the end of 2015, although an online-software incident affected about 130 pb−1 in 2018.
  • Subsystem performance: The muon system maintained at least 99.8% DQ efficiency except for a 2017 RPC hardware incident that rejected about 500 pb−1 and reduced RPC efficiency to 99.2%.Real-time DCS alarms enabled identification, and replacing a problematic high-voltage crate restored operation.
  • Trigger performance: A 2016 L1 trigger timing issue affected about 590 pb−1 and produced an L1 DQ efficiency of 98.3%.The issue caused a sharp drop in the 2016 cumulative DQ efficiency.
  • Run 2 evolution: Continuous detector maintenance, software development, and DQ monitoring increased standard 13 TeV pp efficiency from 88.8% in 2015 to 97.5% in 2018.Identified sources of loss were addressed through corrective actions intended to reduce subsequent occurrences.

6 Summary

ATLAS presents its Run 2 data quality operating and assessment procedures, reporting steadily improving efficiency for good-quality physics data. The combined 2015–2018 standard 13 TeV pp dataset achieved 95.6% data quality efficiency.

  • 97.5% data quality efficiency was achieved for the standard 13 TeV pp collision dataset collected during 2018.
  • 95.6% combined data quality efficiency was achieved for the standard Run 2 dataset collected between 2015 and 2018.
  • Continuous review and development of data quality and ATLAS operating procedures resulted in increasing efficiency for good-quality physics data during Run 2.
Loading 1911.04632v2…