Source-linked AI summary

Planetary Candidates Observed by Kepler. VII. The First Fully Uniform Catalog Based on The Entire 48 Month Dataset (Q1-Q17 DR24)

Jeffrey L. Coughlin, F. Mullally, Susan E. Thompson, Jason F. Rowe, Christopher J. Burke, David W. Latham, Natalie M. Batalha, Aviv Ofir, Billy L. Quarles, Christopher E. Henze, Angie Wolfgang, Douglas A. Caldwell, Stephen T. Bryson, Avi Shporer, Joseph Catanzarite, Rachel Akeson, Thomas Barclay, William J. Borucki, Tabetha S. Boyajian, Jennifer R. Campbell, Jessie L. Christiansen, Forrest R. Girouard, Michael R. Haas, Steve B. Howell, Daniel Huber, Jon M. Jenkins, Jie Li, Anima Patil-Sabale, Elisa V. Quintana, Solange Ramirez, Shawn Seader, Jeffrey C. Smith, Peter Tenenbaum, Joseph D. Twicken, Khadeejah A. Zamudio

arXiv:1512.06149v2astro-ph.EP

TL;DR

Accurate planet-occurrence rates require uniformly vetting every Kepler threshold-crossing event, but manual review is impractical at the mission’s scale. This paper presents a fully automated robotic vetting catalog for the complete 48-month dataset, retaining most injected and confirmed planets while eliminating false-positive signals.

  • Problem

    Manual vetting is time-consuming, infeasible for roughly 20,000 threshold-crossing events, and can vary between human reviewers, complicating uniform evaluation for occurrence rates.

  • Method

    The catalog applies algorithmic tests in a staged robotic vetting procedure, using an “innocent until proven guilty” standard and artificial transit injections to evaluate retention and biases.

  • Results

    4,696 planet candidates are now included in the cumulative Kepler KOI catalog after uniformly vetting 20,367 Q1–Q17 DR24 threshold-crossing events.

  • Takeaways & Limitations

    The robotic vetting and transit-injection framework supports more accurate computation of Earth-size planet occurrence in the habitable zones of Sun-like stars.

  • Takeaways & Limitations

    KOI 1126.02 illustrates that the robovetter can incorrectly classify a signal produced by a subset of eclipsing-binary secondary eclipses as a new KOI.

Abstract

from arXiv · show

We present the seventh Kepler planet candidate catalog, which is the first to be based on the entire, uniformly processed, 48 month Kepler dataset. This is the first fully automated catalog, employing robotic vetting procedures to uniformly evaluate every periodic signal detected by the Q1-Q17 Data Release 24 (DR24) Kepler pipeline. While we prioritize uniform vetting over the absolute correctness of individual objects, we find that our robotic vetting is overall comparable to, and in most cases is superior to, the human vetting procedures employed by past catalogs. This catalog is the first to utilize artificial transit injection to evaluate the performance of our vetting procedures and quantify potential biases, which are essential for accurate computation of planetary occurrence rates. With respect to the cumulative Kepler Object of Interest (KOI) catalog, we designate 1,478 new KOIs, of which 402 are dispositioned as planet candidates (PCs). Also, 237 KOIs dispositioned as false positives (FPs) in previous Kepler catalogs have their disposition changed to PC and 118 PCs have their disposition changed to FP. This brings the total number of known KOIs to 8,826 and PCs to 4,696. We compare the Q1-Q17 DR24 KOI catalog to previous KOI catalogs, as well as ancillary Kepler catalogs, finding good agreement between them. We highlight new PCs that are both potentially rocky and potentially in the habitable zone of their host stars, many of which orbit solar-type stars. This work represents significant progress in accurately determining the fraction of Earth-size planets in the habitable zone of Sun-like stars. The full catalog is publicly available at the NASA Exoplanet Archive.

1. INTRODUCTION

Kepler searched roughly 170,000 stars for transit signals to measure the frequency of Earth-size planets in habitable zones around Sun-like stars. This catalog introduces fully automated, uniform robotic vetting of the complete Q1–Q17 dataset, with artificial-transit tests to quantify biases.

  • Kepler mission and motivation: ∼170,000 stars were monitored by Kepler to search for periodic brightness drops caused by transiting planets.The instrument observed across a 115-square-degree field and achieved ∼30 ppm noise on 12th-magnitude solar-type stars over six hours.
  • Prior catalogs: Previous Kepler catalogs added candidates as observations accumulated and supported studies of planetary occurrence rates and planet-confirmation techniques.These catalogs were also used in analyses of astrophysically variable systems.
  • Contribution: The catalog uses robotic vetting to evaluate every periodic signal from the entire 48-month Q1–Q17 DR24 dataset uniformly.The approach prioritizes uniform vetting and enables potential biases to be quantified with artificial transit injection and other tests.
  • Contribution: The authors caution that a pipeline veto flaw produced a non-uniform planet search, requiring care when using this catalog for planetary occurrence rates.The catalog's uniform vetting does not remove the upstream search non-uniformity.

2. Q1–Q17 DR24 TCEs

The Q1–Q17 DR24 catalog begins with uniformly processed pipeline data and 20,367 threshold crossing events. Its signal population includes astrophysical and instrumental false positives, while pipeline choices and excluded contact binaries define important completeness boundaries.

  • Dataset: 198,646 targets were observed over 48 months in Q1–Q17, with 112,001 observed in every quarter.DR24 processed all data with version 9.2 of the Kepler pipeline.
  • Dataset: 20,367 threshold crossing events formed the starting population for designating planet candidates and false positives.Each TCE has an associated KIC ID, period, epoch, depth, and duration.
  • False-positive population: Short-period TCE excesses primarily reflect quasi-sinusoidal variables and eclipsing binaries, while bright-variable contamination creates localized period spikes.Examples include rapid rotators, pulsating stars, RR Lyrae, V2083 Cyg, and V380 Cyg.
  • False-positive population: A spike at ∼459 days arose from edge effects involving three equally spaced data gaps and affected many stars across the field.This systematic corresponds to three equally spaced gaps in the Q1–Q17 data.
  • False-negative population: The statistical bootstrap veto removed many long-period false positives but also eliminated valid low-SNR transit-like signals and introduced a period-dependent, non-uniform search.Artificial-transit injections indicated that this complicates planetary occurrence-rate calculations.
  • False-negative population: 1,033 known contact eclipsing binaries were intentionally excluded because sinusoidal and quasi-sinusoidal signals were not considered transit-like for the mission search.The exclusion reduced processing demands but defines a boundary on the searched target population.

3. ROBOTIC VETTING

The robovetter replaces time-intensive, potentially inconsistent human inspection with uniform decision-tree vetting of Kepler threshold-crossing events. Quantitative tests classify signals as false positives or planet candidates while identifying contamination mechanisms and preserving injected transits.

  • Robotic Vetting: The robovetter uses simple decision trees to disposition every threshold-crossing event uniformly and provide a specific false-positive reason.It was developed from the Q1–Q16 catalog and refined through manual checks of the Q1–Q17 DR24 dataset.
  • Robotic Vetting: Each event is tested for transit likeness, significant secondary eclipses, centroid offsets, and ephemeris matches before classification as a false positive or planet candidate.The procedure aims to preserve at least ∼95% of injected transits while rejecting as many false positives as possible.
  • Ephemeris Matching: 1,910 Q1–Q17 DR24 threshold-crossing events are identified as false positives through ephemeris matching, including 189 found only by that method.The catalog matches all threshold-crossing events and lists likely parents and period ratios to support contamination studies.
  • Ephemeris Matching: 119 column-anomaly cases reveal contamination without visible centroid offsets, with 91.6% of children located at higher row numbers than their parents.The authors conjecture that decreasing charge-transfer efficiency, likely from cosmic-ray impacts, produces the contamination; an average depth ratio of ∼10^4 is consistent with ∼99.99% efficiency.

4. TCE DISPOSITIONING AND KOI MODELING

The robovetter processed all Q1–Q17 DR24 TCEs, federated transit-like signals with existing KOIs, assigned new KOI numbers, and dispositioned signals as planet candidates or false positives. Each TCE was evaluated using standardized pipeline metrics, while KOIs were modeled to obtain planetary parameters and uncertainties.

  • KOI federation: 5,992 TCEs were federated to existing KOIs using overlapping in-transit cadences between ephemerides.Federation linked transit-like signals across pipeline runs for consistent KOI tracking.
  • KOI federation: 90.2% of previously known transit-like KOIs were recoverable, compared with 81.5% for all previously known KOIs.The broader rate includes KOIs previously dispositioned as not transit-like false positives, which need not be recovered by the new pipeline run.
  • New KOI designation: 1,478 new KOIs increased the cumulative catalog to 8,826 KOIs.Twenty-five transit-like systems were not assigned KOI numbers because unusual or complicated signals did not accurately correspond to the underlying transit-like signals.
  • TCE dispositioning: False positives were identified using major flags for not-transit-like signals, significant secondary eclipses, centroid offsets, and related vetting tests.A significant secondary flag combination identifies TCEs corresponding to secondary eclipses; pre-existing KOIs can still federate with such TCEs.
  • KOI modeling: Every KOI was modeled consistently, using detrended DR24 light curves, a circular-orbit transit model, quadratic limb darkening, and MCMC uncertainties.Previously existing KOIs retained earlier fit parameters, whereas newly designated KOIs used DR24 light curves and updated stellar values.

5. ANALYSIS OF THE Q1–Q17 DR24 CATALOG

The Q1–Q17 DR24 catalog broadly agrees with prior and ancillary catalogs while using robotic vetting and injection tests to quantify performance and biases. It recovers most previously classified objects, identifies new or revised dispositions, and highlights potentially habitable candidates, while retaining important limitations for occurrence-rate studies.

  • Comparison to Past KOI Catalogs: 97.4% of pre-existing transit-like KOIs re-detected in DR24 were recovered as transit-like by the robovetter.The robovetter dispositioned 5,700 of 5,854 such KOIs as transit-like.
  • Comparison to Past KOI Catalogs: 237 KOIs changed from FP to PC, while 118 changed from PC to FP compared with past catalogs.The paper attributes many changes to the robovetter detecting small secondary eclipses and applying the transit-depth rule uniformly.
  • Comparison to Ancillary Catalogs: 94.5% of detached eclipsing binaries detected as TCEs were identified as false positives by the robovetter.The pipeline detected 894 of 933 cataloged systems, and the robovetter identified 805 specifically through significant secondaries, with additional systems flagged by other tests.
  • Comparison to Ancillary Catalogs: 99.1% of 985 confirmed Kepler planets federated with DR24 TCEs were designated planet candidates by the robovetter.The nine remaining cases were concluded to merit PC dispositions after manual examination, despite varied failure causes.
  • Artificial Transit Injection: 95.25% of 35,917 injected TCEs without centroid offsets passed the robovetter as planet candidates.Recovery increased with MES and decreasing period, while higher-temperature or more evolved stars showed lower recovery fractions.
  • Artificial Transit Injection: 98.3% of injected signals within 25% of Earth’s radius and insolation values were recovered as planet candidates.With an additional host-temperature constraint within 500 K of the Sun, the recovery fraction was 96.1%.
  • Catalog Population: The catalog contains 4,293 planet candidates in 3,355 systems, including 1,632 candidates in multi-candidate systems.Multi-KOI systems had an 8.6% FP rate versus 51.6% for single-KOI systems, providing a check against systematic rejection of multi-candidate systems.

6. DISCUSSION

The uniformly vetted catalog characterizes candidate and false-positive populations while documenting both its performance and important limitations for occurrence-rate studies.

  • Catalog outcomes: 4,298 Q1–Q17 DR24 TCEs were designated as PCs after the robovetter ruled 13,283 not transit-like and 2,786 transit-like false positives.Five designated PCs lacked KOI numbers, yielding 4,293 PCs in the Q1–Q17 DR24 catalog and 4,696 PCs cumulatively.
  • False-positive populations: 1,215 on-target eclipsing binaries appear in the Q1–Q17 DR24 KOI catalog, identified through false-positive dispositions caused by significant secondary eclipses.
  • Distribution comparisons: Figure 9 shows that short- and long-period excesses, local period spikes, and very small or large-radius TCE populations were generally eliminated.False-positive KOIs are represented by the difference between the green and blue distributions.
  • Validation and caveats: Artificial transit injection was used for the first time in developing and evaluating both the Kepler pipeline and the TCERT vetting process.The catalog warns that occurrence-rate calculations require care because the Q1–Q17 DR24 pipeline search was period-dependent.
  • Validation and caveats: The robovetter retains valid candidates while robustly identifying false positives, but full false-positive simulations remain needed to quantify false-positive rates across SNR and other parameters.Known failure areas include some eclipsing binaries, contaminated signals, strongly TTV-affected planets, and variable stars.

7. CONCLUSION

The paper presents a uniformly vetted catalog spanning the entire Kepler dataset and uses robotic vetting plus artificial-transit tests to improve occurrence-rate estimates.

  • Conclusion: The catalog is the first uniform planet-candidate catalog based on the entire 48-month Kepler dataset.
  • Conclusion: The robotic vetting eliminates most false-positive signals while retaining greater than 98% of injected Earth-sized, Earth-insolation planets and over 99% of confirmed planets.
  • Conclusion: Artificial transit injection and robotic vetting enable more accurate computation of the fraction of Earth-size planets in the habitable zones of Sun-like stars.The approach can also be applied to other large-scale photometric survey missions.

A. LIST OF ACRONYMS

This appendix defines the acronyms used throughout the Kepler catalog paper, covering data products, detection signals, catalog entities, and vetting groups.

  • Astrophysical terms: HZ means Habitable Zone, the region around a star where surface temperatures could allow liquid water.
  • Catalog terms: KOI means Kepler Object of Interest, a unique identifier for a signal consistent with a transiting or eclipsing system.
  • Signal and search terms: MES and SES denote the multiple-event and single-event signal-to-noise statistics used by the TPS module.
  • Signal and search terms: TCE means Threshold Crossing Event, a series of periodic flux decrements consistent with a transiting-planet signal.
  • Vetting terms: TCERT means Threshold Crossing Event Review Team, the committee that reviews TCEs for false-positive and planet-candidate designations.
  • Signal and search terms: TPS means Transiting Planet Search, the Kepler-pipeline module that searches for transits.

B. ROBOVETTER MNEMONIC FLAGS

The robovetter mnemonic flags record specific diagnostic outcomes and disposition overrides, especially for secondary eclipses and possible period misidentification.

  • Flag definitions: The mnemonic flags summarize the results of individual robovetter tests in the catalog comments column.
  • Flag definitions: ALT ROBO ODD EVEN TEST FAIL marks a TCE as a false positive because its alternate detrending fails the odd-even depth test.
  • Flag definitions: ALT SEC COULD BE DUE TO PLANET keeps a TCE as a planet candidate when a significant secondary may arise from planetary reflection or thermal emission.
  • Flag definitions: ALT SEC SAME DEPTH AS PRI COULD BE TWICE TRUE PERIOD overrides other flags when equal-depth primary and secondary eclipses suggest detection at twice the true period.

ALT SIG PRI MINUS SIG POS TOO

The alternate-detrending model-shift test rejects a TCE when the primary event is insufficiently distinct from the positive event, indicating a non-unique phased signal.

  • A primary–positive event significance difference below σ′FA marks the TCE as having a non-unique primary event.The alternate detrending is used in this model-shift test.

ALT SIG PRI MINUS SIG TER TOO

The alternate-detrending vetting evaluates primary–tertiary event separation, systematic-noise significance, centroid behavior, crowding, and secondary-event characteristics to disposition TCEs.

  • A primary–tertiary significance difference below σ′FA indicates that the primary event is not unique and yields a false-positive disposition.This criterion uses the alternate detrending in the model-shift test.
  • A primary-event significance below the σFA threshold after red-noise normalization indicates insufficient significance relative to systematic noise and yields a false-positive disposition.The metric divides primary-event significance by the red-noise to white-noise ratio.
  • Centroid uncertainty prevents confident false-positive disposition when only 3 or 4 low-SNR offset measurements are available.The centroid module cannot measure the offset significance precisely enough in this case.
  • A resolved offset star triggers a centroid-based false-positive disposition, while multiple potential stellar images flag the difference image as crowded.The CROWDED DIFF condition always sets the EYEBALL flag.
  • Odd-even depth failure marks a TCE false positive, whereas secondary eclipses possibly caused by planetary emission or equal-depth events compatible with twice the true period can preserve candidate status.The equal-depth, twice-period condition overrides the significant-secondary major flag when no other major flags are present.

DV SIG PRI MINUS SIG POS TOO

The DV-detrending model-shift test rejects TCEs whose primary event is insufficiently distinct from the positive event, indicating a non-unique phased signal.

  • A primary–positive event significance difference below σ′FA marks the TCE as having a non-unique primary event and yields a false-positive disposition.The test uses the DV detrending.

DV SIG PRI MINUS SIG TER TOO

The DV-detrending vetting combines primary-event uniqueness and significance tests with centroid, transit-shape, and transit-fit diagnostics to disposition TCEs, while retaining uncertain cases for scrutiny.

  • A primary–tertiary significance difference below σ′FA indicates a non-unique primary event and yields a false-positive disposition.This criterion is based on the DV detrending.
  • A primary-event significance below the σFA threshold after red-noise normalization indicates insufficient significance relative to systematic noise and yields a false-positive disposition.The metric uses the primary significance divided by the red-noise to white-noise ratio.
  • Centroid-boundary cases warrant further scrutiny, while failed transit fits do not automatically receive a centroid-offset false-positive disposition.Failed fits are typically associated with very deep eclipsing-binary transits.
  • Inverted difference images, usually associated with target variability, trigger candidate status requiring further scrutiny rather than a centroid-offset false-positive disposition.The inversion indicates that the difference image should not be trusted.
  • A KIC-referenced centroid offset is not by itself evidence of a statistically significant centroid shift when it is the only flag.The Kepler Input Catalog position is less accurate in sparse fields but more accurate in crowded fields.
  • High LPP values or a failed Marshall metric indicate non-transit-shaped signals or instrumental artifacts and yield false-positive dispositions.Both LPP detrendings can trigger this outcome, while the Marshall metric evaluates individual transit shapes.

OTHER TCE AT SAME PERIOD DIFF

A same-period TCE with a different epoch can identify the current signal as an eclipsing binary’s secondary eclipse. An ephemeris match can also identify a likely physical parent, though that parent is not guaranteed to be the true source.

  • A same-period, different-epoch TCE indicates the current signal is an eclipsing binary’s secondary eclipse.Unless planet-related secondary-eclipse flags are set, the TCE is dispositioned as a false positive.
  • When the planet-related secondary-eclipse flags are absent, the same-period, different-epoch signal receives a significant-secondary false-positive disposition.
  • An ephemeris match identifies the most likely parent or true physical source of a false-positive signal.The named parent is not guaranteed to be the true source given the available information.

PERIOD ALIAS IN ALT DATA SEEN

The vetting rules use period aliases, transit consistency, centroid availability, saturation, seasonal depth differences, and model-shift secondary-event tests to classify or flag TCEs. Some conditions trigger false-positive dispositions, while others remain informational or require further scrutiny.

  • PERIOD ALIAS IN ALT DATA SEEN: An X:1 model-shift pattern indicates the detected period may be X times longer than the true orbital period.This flag is informational only and does not declare a TCE a false positive.
  • PERIOD ALIAS IN ALT DATA SEEN: A same-period, same-epoch signal matching a previous transit-like TCE is treated as its residual artifact and dispositioned as a false positive.
  • PERIOD ALIAS IN ALT DATA SEEN: A TCE sharing a period with a prior not-transit-like false positive is attributed to the same signal and dispositioned likewise.
  • PERIOD ALIAS IN ALT DATA SEEN: Saturated stars violate centroid-module assumptions, so their TCEs receive an EYEBALL flag rather than a centroid-offset false-positive disposition.
  • PERIOD ALIAS IN ALT DATA SEEN: Seasonal depth differences indicate significant contamination, but alone cannot determine whether the signal is on-target and therefore remain informational.This rule is stated for both alternate-detrending and DV detrending light curves.
  • PERIOD ALIAS IN ALT DATA SEEN: Significant secondary events identified by alternate or DV model-shift tests trigger false-positive disposition unless the event could be due to a planet.The tests compare secondary-event significance against red-to-white noise and other event significances.
  • PERIOD ALIAS IN ALT DATA SEEN: Fewer than three high-SNR difference images limit centroid testing, and combined with a clear aperture may indicate a resolved neighboring source.
  • PERIOD ALIAS IN ALT DATA SEEN: A max ses in mes / mes ratio above 0.9 with a period longer than 90 days indicates domination by one systematic-like event and yields a not-transit-like false-positive disposition.
Loading 1512.06149v2…