Source-linked AI summary

The Data Reduction Pipeline for the Apache Point Observatory Galactic Evolution Experiment

David L. Nidever, Jon A. Holtzman, Carlos Allende Prieto, Stephane Beland, Chad Bender, Dmitry Bizyaev, Adam Burton, Rohit Desphande, Scott W. Fleming, Ana Elia Garcia Perez, Fred R. Hearty, Steven R. Majewski, Szabolcs Meszaros, Demitri Muna, Duy Nguyen, Ricardo P. Schiavon, Matthew Shetrone, Michael F. Skrutskie, Jennifer S. Sobeck, John C. Wilson

arXiv:1501.03742v2astro-ph.IMastro-ph.GA

TL;DR

APOGEE required an automated reduction pipeline because its large survey data volume and slightly undersampled spectra made spectral processing challenging. This paper documents the pipeline used for SDSS-III APOGEE public releases, which produced combined spectra and achieved the survey’s required S/N, while several limitations remained.

  • Problem

    APOGEE needed an automated reduction pipeline for its large survey data volume, while slightly undersampled spectra required observations at two detector-dither positions.

  • Method

    The paper documents the APOGEE reduction pipeline, including calibration data from continuum, ThArNe, UNe, and internal-flat sources and refined absolute radial-velocity determination against synthetic spectra.

  • Results

    The documented pipeline was used to produce the SDSS-III APOGEE DR10–DR12 public-release data products, and the spectra can achieve S/N∼100 required for the survey.

  • Takeaways & Limitations

    The pipeline provides the documented processing basis for APOGEE public-release spectra, while the paper identifies areas for improvement in persistence, sky subtraction, radial velocities, and binarity handling.

  • Takeaways & Limitations

    The DR12 pipeline contains no persistence correction because APOGEE persistence from multiple preceding stimuli was not yet fully characterized.

Abstract

from arXiv · show

The Apache Point Observatory Galactic Evolution Experiment (APOGEE), part of the Sloan Digital Sky Survey III, explores the stellar populations of the Milky Way using the Sloan 2.5-m telescope linked to a high resolution (R~22,500), near-infrared (1.51-1.70 microns) spectrograph with 300 optical fibers. For over 150,000 predominantly red giant branch stars that APOGEE targeted across the Galactic bulge, disks and halo, the collected high S/N (>100 per half-resolution element) spectra provide accurate (~0.1 km/s) radial velocities, stellar atmospheric parameters, and precise (~0.1 dex) chemical abundances for about 15 chemical species. Here we describe the basic APOGEE data reduction software that reduces multiple 3D raw data cubes into calibrated, well-sampled, combined 1D spectra, as implemented for the SDSS-III/APOGEE data releases (DR10, DR11 and DR12). The processing of the near-IR spectral data of APOGEE presents some challenges for reduction, including automated sky subtraction and telluric correction over a 3 degree diameter field and the combination of spectrally dithered spectra. We also discuss areas for future improvement.

1. INTRODUCTION

APOGEE acquired high-resolution near-infrared spectra for a large Milky Way stellar survey, requiring an automated pipeline because of several detector and observational challenges. This paper documents the reduction methods used for the SDSS-III/APOGEE public data releases.

  • Over 150,000 Milky Way stars were observed with a cryogenic near-infrared spectrograph producing 300 simultaneous spectra across 1.51–1.70 µm.
  • Telluric absorption and OH-dominated sky brightness vary spatially and temporally, complicating automated correction of the near-infrared spectra.
  • Undersampling at the short-wavelength end requires observations at two detector dither positions separated by approximately 0.5 pixel.
  • Detector persistence causes previous exposures to affect subsequent H2RG behavior, adding another complication to the reduction process.
  • The paper documents the pipeline used to produce the SDSS-III/APOGEE DR10, DR11, and DR12 data releases.
  • APRED reduces nightly plate observations through cube reduction, spectral extraction and wavelength calibration, then sky correction, dither combination, and initial radial-velocity estimation.

2. SURVEY OPERATIONS AND DATA TAKING

APOGEE survey operations combine repeated dithered exposures, dedicated calibration observations, and substantial data handling to support uniform spectroscopic measurements. Custom lossless compression reduces the volume of the large up-the-ramp data stream.

  • Each plate typically assigns 230 fibers to science targets, 35 to blue telluric standards, and 35 to sky regions across a roughly 3° field.
  • Routine observations use 500-second exposures at A or B dither positions, with approximately 0.5-pixel detector shifts and standard ABBA sequences.
  • The detectors are read non-destructively at slightly more than 10-second intervals, yielding 47 readouts during a standard exposure for monitoring and reduction.
  • Calibration sources include continuum, ThArNe, and UNe lamps, while internal-flat frames from infrared LEDs measure pixel-to-pixel sensitivity variations.
  • A full observing night produces roughly 100 GB of data because each exposure contains 47 readouts from three 2048×2048 detector chips plus bias information.
  • The custom compression workflow converts reads to difference images, removes their average difference, and applies lossless Rice compression in FITS files.
  • The custom files average a compression factor of approximately 2, compared with a practical best of approximately 2.2 for similar data.

3. PIPELINE OVERVIEW

The APOGEE reduction pipeline has two principal stages: APRED processes individual plate-night observations, while APSTAR combines multiple stellar visits and derives refined velocities. A separate ASPCAP stage determines stellar parameters and abundances.

  • APRED processes an individual plate observed on one night through AP3D, AP2D, and AP1DVISIT.
  • AP3D converts raw data cubes into calibrated 2D images, while AP2D extracts 300 well-sampled 1D spectra and determines wavelength calibration.
  • AP1DVISIT measures dither shifts, corrects sky emission and absorption, combines visit exposures, and estimates initial stellar radial velocities.
  • APSTAR resamples visits onto a constant-dispersion log(λ) grid, corrects visit-specific velocities, coadds spectra, and places radial velocities on an absolute scale.
  • ASPCAP separately determines stellar parameters and chemical abundances from the processed spectra.

4. AP3D: REDUCTION OF DATA CUBE TO 2D IMAGE

AP3D transforms non-destructive detector readouts into calibrated 2D images while constructing error and bad-pixel information. Its processing uses reference pixels and includes cosmic-ray handling, but some corrections remain unimplemented.

  • APOGEE arrays are read non-destructively approximately every 10.6 seconds in sample-up-the-ramp mode.
  • Sample-up-the-ramp readouts reduce read noise, detect cosmic rays, and can support saturated-pixel correction when at least three unsaturated reads remain and flux is assumed constant.
  • AP3D collapses raw data cubes into 2D images while applying detector calibration and generating error images and bad-pixel masks.
  • Linearity correction and saturated-pixel correction were not implemented in the listed AP3D processing steps.
  • Reference pixels correct electronic effects by subtracting reference arrays and removing vertical and horizontal bias ramps for each readout and detector.

4.2. Linearity

The pipeline tests detector nonlinearity but does not apply a linearity correction in DR10–DR12, because persistence is a larger effect in affected regions. Dark-current calibration is separately characterized and found to be stable.

  • 4.2. Linearity: Persistence in regions of two detectors is significantly larger than expected nonlinearity.Previous exposure affects the subsequent charge deposition in these regions.
  • 4.2. Linearity: The pipeline does not apply its available linearity correction to DR10–DR12 data.Initial tests suggested small nonlinearity, but low-light pixel behavior complicated its characterization.
  • 4.2. Linearity: Most pixels have dark rates below 0.5 counts/read, although a high-rate tail occurs and the middle array contains a higher-dark-current section.Very high-dark-current pixels show interpixel-capacitance effects, producing crosses through charge coupling to adjacent pixels.
  • 4.2. Linearity: The dark current is stable across observations separated by more than two years, with only a few pixels changing significantly.The pipeline uses a superdark formed from 20 long dark frames, with a separate calibration slice for each up-the-ramp readout.
  • 4.2. Linearity: Pixels exceeding 10 counts/read are marked bad, as are neighboring pixels exceeding 2.5 counts/read.

4.4. Cosmic Ray and Saturated Pixel Correction

The pipeline uses up-the-ramp data to detect cosmic rays and recover saturated-pixel signals, while flagging saturated corrections as unreliable and treating those pixels as bad downstream.

  • 4.4. Cosmic Ray and Saturated Pixel Correction: Cosmic rays are detected as positive jumps in successive-read difference counts after median filtering and robust local-scatter estimation.The method removes time-varying flux-rate structure before identifying anomalous jumps.
  • 4.4. Cosmic Ray and Saturated Pixel Correction: Up-the-ramp sampling records usable signal before saturation, allowing flux extrapolation when approximately 3–4 reads remain unsaturated.This assumes a stable count rate, which may fail under sub-optimal observing conditions.
  • 4.4. Cosmic Ray and Saturated Pixel Correction: Cosmic-ray histograms across the three detectors have median rates of 0.92, 1.00, and 0.99 per exposure time.Shorter exposures show slightly higher detected rates, likely because lower Poisson noise reveals weaker cosmic rays.
  • 4.4. Cosmic Ray and Saturated Pixel Correction: The pipeline currently corrects saturated pixels assuming constant flux, then flags them and treats them as bad in later reduction stages.The constant-rate assumption can be a poor approximation when count rates vary.
  • 4.4. Cosmic Ray and Saturated Pixel Correction: The detector-correction workflow includes Fowler and up-the-ramp sampling, with up-the-ramp used for all data except dome flats.Dome flats use simple CDS sampling because their short lamp-on interval produces a highly non-uniform count rate.

4.7. Gain, Noise Model and Bad Pixel Mask

The pipeline models detector gain and noise, propagates bad-pixel masks, and characterizes persistence as a major detector-specific concern in some regions.

  • Gain and noise: A gain of 1.9 e−/DN is adopted for all detector quadrants, with future recalibration planned.The gain is derived from flat-field variance measurements using multiple intensity bins.
  • Gain and noise: The error array combines Poisson noise from the image and dark exposure with sampling read noise.
  • Bad-pixel masking: Bad-pixel masks preserve why pixels were flagged and propagate those flags into reduced frames; affected pixels are excluded from subsequent analysis.Littrow-ghost regions are flagged but may remain usable when the effect is negligible.
  • Persistence: At 1800 s after a 25,000-count stimulus, normal persistence is about 60 counts, whereas superpersistence regions retain 10–20% of stimulus counts.High-persistence areas include roughly the top third of the blue array and the green-detector perimeter.
  • Persistence: The median relative blue excess in superpersistence regions is about 17.4%, with a long high-value tail, and the faintest stars are most affected.Persistence arises from preceding stars and calibration exposures, with dome-flat persistence generally featureless but capable of altering abundance measurements.
  • Persistence: Persistence from a single stimulus follows a double exponential with timescales near 120 s and 1700 s, but multiple stimuli remain insufficiently characterized.The DR12 pipeline contains no persistence correction because the behavior is complex; affected spectra are flagged.

4.9. Output

AP3DPROC outputs calibrated intermediate 2D detector products containing fluxes, errors, and bitwise pixel masks for each detector.

  • Output products: AP3DPROC writes one ap2D FITS file per red, green, and blue detector, each containing flux, errors, and bitwise pixel-mask extensions.Each extension has dimensions 2048×2048.

5. AP2D: EXTRACTION TO 1D SPECTRA

AP2D extracts 300 fiber spectra from detector images, corrects throughput, and wavelength-calibrates the resulting 1D products using spatial profiles and lamp and sky-line information.

  • Pipeline role: AP2D extracts 300 spectra, corrects fiber-to-fiber throughput variations, and applies wavelength calibration.
  • Extraction: The standard extraction uses empirically measured, visit-stable spatial PSFs for traces separated by roughly 6–7 pixels.Sparse-pack exposures provide widely separated spectra for accurate PSF determination.
  • Extraction: At each detector column, adjacent-fiber coupling is solved as a tridiagonal linear-algebra problem for 300 fiber fluxes.
  • Throughput correction: Dome flats provide plate-specific, moderately wavelength-dependent throughput corrections for the extracted spectra.
  • Throughput correction: Throughput correction reduces sky-line flux rms variation from about 12% to about 5%, and to about 1% around a smooth 2D spatial-polynomial fit.
  • Wavelength calibration: Wavelength solutions combine ThArNe and UNe lamp lines into separate fifth-order polynomial fits for each fiber, with detector-gap terms.Science spectra receive exposure-specific pixel zero-point corrections using sky emission lines.
  • Output products: AP2D outputs ap1D FITS files containing flux, errors, bitwise masks, and wavelength arrays, plus modeled 2D images.The main arrays have dimensions 2048×300.

6. AP1DVISIT: VISIT STAGE

AP1DVISIT corrects sky and telluric contamination, measures dither shifts, and combines separate exposures into well-sampled spectra while modeling the wavelength- and fiber-dependent LSF. The stage addresses undersampling, spatial and temporal variation, and detector resolution variation in APOGEE data.

  • Visit-stage processing: AP1DVISIT sky- and telluric-corrects exposures before combining different dither positions into one well-sampled spectrum per fiber.Sky and telluric corrections are performed exposure by exposure because sky conditions can vary on short timescales.
  • Dither measurement: ∼0.5 pixels separates the two spectral dither positions used to mitigate undersampling at bluer wavelengths.The actual shift is measured from the data because the dither positions are not known precisely enough beforehand.
  • Dither measurement: ∼0.005 pixels are the average formal errors from the pipeline’s cross-correlation dither-shift measurements.The current pipeline measures shifts fiber by fiber and array by array, then uses a robust mean of individual measurements.
  • Sky subtraction: The pipeline forward-models sky corrections using observed sky and telluric fibers plus an LSF varying with wavelength and fiber.This approach is designed for undersampled spectra and spectra whose LSF changes across the detector.
  • LSF characterization: ∼1–2% rms FWHM variation per fiber over three years indicates remarkably stable instrument resolution.The measured values show small-scale variations but remain stable over the monitoring interval.
  • LSF characterization: R=22,500 is the average traditional resolving power, with ∼10–20% spatial and spectral variation across the detectors.Because the APOGEE LSF is non-Gaussian, this traditional resolving-power value does not fully describe the resolution.
  • Sky subtraction: The current sky subtraction is temporary and sub-optimal, motivating investigation of two-dimensional airglow modeling and PCA.Sky continuum removal is especially important because residual continuum distorts normalized line depths; poorly removed airglow can affect radial velocities.
  • Telluric correction: The telluric correction fits hot-star spectra, spatially models species scaling, and constructs each science-fiber correction using the known LSF.Telluric absorption from H2O, CO2, and CH4 affects ∼20% of the APOGEE spectral range.

7. APSTAR: OBJECT STAGE

APSTAR places visit spectra on a common rest-frame logarithmic wavelength grid and combines them into object-level spectra. It uses interpolation, weighted coaddition, continuum handling, and LSF combination to produce high-S/N products, while the achieved S/N becomes systematics-limited for bright stars.

  • Object-stage combination: APSTAR resamples visit spectra onto a fixed log-wavelength grid after correcting each visit for its radial velocity, then coadds them.The output grid contains 8575 pixels spanning 15100.8–16999.8 Å at approximately three pixels per resolution element.
  • Object-stage combination: Sinc interpolation uses chip-dependent FWHM values of 5, 4.25, and 3.5 dithered pixels for the red, green, and blue chips.The interpolation accounts for sampling differences and filters noise at higher spatial frequencies.
  • Object-stage combination: The resampled visits are combined with weighted means using pixel- and spectrum-level inverse-square normalized-error weights.The combined spectrum is then multiplied by the average continuum of the individual visit spectra.
  • Radial velocities: APSTAR derives relative visit radial velocities by cross-correlation with the combined spectrum and places them on an absolute scale using a best-matching template.Radial-velocity determination and spectral combination are performed iteratively.
  • Output quality: S/N∼100 is readily achieved, although bright-star empirical S/N values are systematically below noise-model estimates because of ∼0.5% systematics.The comparison uses 9548 stars with six visits and S/N per half-resolution element.
  • Output products: A combined LSF is formed by weighted averaging of visit LSF arrays on the final apStar wavelength scale, with fitted Gauss–Hermite coefficients saved.The products include combined spectra, individual resampled visits, weighted LSFs, and summary field files.

8. RADIAL VELOCITY DETERMINATION

APOGEE derives radial velocities through successive visit- and object-level methods, combining observed spectra with synthetic templates and iterative relative shifts. DR12 improved the velocity zero point and dispersion, with internal precision near 70 m s−1 for high-S/N giants and long-term plate stability of 0.044 km s−1.

  • Visit-stage radial velocities: APOGEE first estimates visit radial velocities by cross-correlating each spectrum with a grid of synthetic spectra, but these estimates are not used later.The estimated RVs remain in apVisit files as a comparison product.
  • Object-stage radial velocities: Refined velocities iteratively determine relative shifts against the combined stellar spectrum, then place that spectrum on an absolute scale using a synthetic RV mini-grid.The combined spectrum generally provides a better match, while the synthetic grid supplies the absolute wavelength reference.
  • Synthetic templates and fitting: The RV mini-grid uses 96 synthetic templates spanning 3,500 < Teff < 25,000 K and 2.0 < log g < 5.0, followed by cross-correlation and χ2 minimization.Observed and template spectra are continuum-normalized before fitting.
  • Limitations: RV limitations include a temperature-dependent BT-Settl offset of about 1 km s−1, degraded relative-template performance for faint stars, and underestimated uncertainties that complicate binary identification.Future processing was expected to remove BT-Settl templates and rely more heavily on synthetic spectra for faint targets.
  • Quality checks: Synthetic RV cross-correlation functions check relative velocities when the library matches and can reveal double-lined spectroscopic binaries, although no automatic SB2 classifier is used.The pipeline stores SYNTH-SCATTER and flags combinations exceeding 1 km s−1.
  • Performance: DR12 improved the RV zero point by ∼0.25 km s−1 and reduced dispersion, while multiple-visit giants reached ∼70 m s−1 scatter compared with ∼110 m s−1 in DR11.The multiple-measurement sample has total S/N>20 and at least three visits.
  • Performance: Plate-to-plate RV differences across 4317 plate pairs have an rms scatter of 0.044 km s−1, indicating stable instrument performance over three survey years.The distribution is centered around zero for pairs with more than 50 stars in common.

9. DATA ACCESS

APOGEE data products are publicly accessible through the SDSS-III Science Archive Server and related web interfaces, with files organized according to the official data model.

  • The Science Archive Server provides web access to APOGEE data products from raw data cubes through reduced spectra and intermediate products.
  • The SDSS-III data model describes the directory structure and file organization used to navigate APOGEE products.
  • Reduced visit spectra are stored in apVisit files, while combined stellar spectra are stored in apStar files.
  • Multi-extension FITS products contain extension-specific data described by the data model, and APLOAD.PRO reads them into an IDL data structure.
  • Summary FITS tables store extracted parameters and object information for individual visits and combined spectra in allVisit and allStar files.

10. SUMMARY

The APOGEE pipeline automates reduction from raw up-the-ramp data cubes through extracted, calibrated, corrected, combined spectra and radial-velocity determination. Its distinctive procedures address cosmic rays, undersampling, telluric absorption, and iterative relative velocities, while future work targets remaining correction and low-S/N limitations.

  • 10. SUMMARY: The pipeline collapses 3D up-the-ramp cubes into 2D images, extracts 300 spectra, calibrates wavelengths and fluxes, applies sky and telluric corrections, and combines visits.
  • 10. SUMMARY: Up-the-ramp sampling enables detection and removal of most cosmic rays from APOGEE data.
  • 10. SUMMARY: Half-pixel spectral dithers are sinc-interlace interpolated into a single well-sampled spectrum because APOGEE spectra are slightly undersampled.
  • 10. SUMMARY: Telluric models fitted to hot-star spectra yield plate-dependent scaling values used to derive a model absorption spectrum for each science spectrum.
  • 10. SUMMARY: Iterative relative radial velocities use each star’s combined observed spectrum as its own template, with the velocities and combined spectrum determined together.
  • 10. SUMMARY: Future improvements include better sky subtraction, persistence correction, and updates to radial-velocity routines for low-S/N spectra.
Loading 1501.03742v2…