Source-linked AI summary
Galaxy Zoo 2: detailed morphological classifications for 304,122 galaxies from the Sloan Digital Sky Survey
Kyle W. Willett, Chris J. Lintott, Steven P. Bamford, Karen L. Masters, Brooke D. Simmons, Kevin R. V. Casteels, Edward M. Edmondson, Lucy F. Fortson, Sugata Kaviraj, William C. Keel, Thomas Melvin, Robert C. Nichol, M. Jordan Raddick, Kevin Schawinski, Robert J. Simpson, Ramin A. Skibba, Arfon M. Smith, Daniel Thomas
TL;DR
Full morphological classification is valuable because common proxies have unknown and likely biased relationships to galaxy features, while detailed visual inspection is impractical at survey scale. GZ2 addresses this by using citizen-scientist classifications and finds good agreement with expert catalogues for several morphological features, especially medium to strong bars.
Problem
Common proxies for morphology have unknown and likely biased relationships with morphological features, motivating large-scale full morphological classification.
Method
GZ2 uses a web-based decision tree to collect crowd-sourced morphological votes for more than 300,000 SDSS galaxies selected by magnitude, size, and redshift criteria.
Results
GZ2 classifications show good agreement with expert catalogues for medium to strong bars, while weak or nuclear bars are identified less confidently.
Takeaways & Limitations
The release provides detailed crowd-sourced morphological data for large SDSS samples, including classifications of features beyond early- and late-type divisions.
Takeaways & Limitations
Debiased data for spectroscopic- and photometric-redshift samples should not be combined because redshift can strongly affect classification bias.
Abstract
from arXiv · showhide
We present the data release for Galaxy Zoo 2 (GZ2), a citizen science project with more than 16 million morphological classifications of 304,122 galaxies drawn from the Sloan Digital Sky Survey. Morphology is a powerful probe for quantifying a galaxy's dynamical history; however, automatic classifications of morphology (either by computer analysis of images or by using other physical parameters as proxies) still have drawbacks when compared to visual inspection. The large number of images available in current surveys makes visual inspection of each galaxy impractical for individual astronomers. GZ2 uses classifications from volunteer citizen scientists to measure morphologies for all galaxies in the DR7 Legacy survey with m_r>17, in addition to deeper images from SDSS Stripe 82. While the original Galaxy Zoo project identified galaxies as early-types, late-types, or mergers, GZ2 measures finer morphological features. These include bars, bulges, and the shapes of edge-on disks, as well as quantifying the relative strengths of galactic bulges and spiral arms. This paper presents the full public data release for the project, including measures of accuracy and bias. The majority (>90%) of GZ2 classifications agree with those made by professional astronomers, especially for morphological T-types, strong bars, and arm curvature. Both the raw and reduced data products can be obtained in electronic format at http://data.galaxyzoo.org .
1 INTRODUCTION
Galaxy Zoo 2 extends volunteer-based galaxy morphology classification beyond broad early-, late-, and merger categories to capture finer structural features at a scale impractical for expert inspection alone. Its large, multiply inspected sample supports detailed morphology studies and quantification of classification relationships and biases.
- Motivation: GZ2 extends the original Galaxy Zoo’s broad elliptical, spiral, and merger labels to classify bars, bulges, rings, arm structure, and other detailed features.These features trace processes including mergers, secular evolution, gas inflow, and bulge growth.
- Motivation: Modern surveys make expert visual classification of every galaxy impractical, while morphology proxies have unknown and likely biased relationships with the features being studied.Examples of proxies include colour, concentration, spectral features, surface-brightness profiles, and spectral energy distributions.
- Project scale: GZ2 contains more than an order of magnitude more systems than earlier detailed SDSS expert catalogues and gives each galaxy many independent inspections.The repeated inspections permit estimates of classification likelihood and, in some cases, feature strength.
- Prior results: Earlier GZ2 data have supported studies linking bars with galaxy colour, gas fraction, bulge prominence, and interactions, and identifying unusual AGN hosts and interaction signatures.The paper presents the data underlying these studies while adding bias corrections and comparisons with other catalogues.
- Paper scope: The paper describes GZ2 sample selection, classification collection, data reduction and debiasing, public tables, and comparisons with four additional SDSS-based morphology catalogues.The results are summarized in a dedicated conclusions section.
2 PROJECT DESCRIPTION
GZ2 combines selected SDSS Legacy and Stripe 82 galaxy images with a branching web-based decision tree to collect detailed volunteer morphology classifications. The release includes millions of classifications, multiple image depths, and repeated responses designed to estimate classification likelihoods.
- Sample selection: The main Legacy sample selects nearby, bright, large galaxies using r-band magnitude, angular-size, redshift, and image-quality criteria, yielding 245,609 original objects.The cuts require petroR90 r magnitude brighter than 17.0, petroR90 r > 3 arcsec, and, when available, 0.0005 < z < 0.25.
- Sample selection: Stripe 82 adds deeper imaging under the same general selection criteria but with a fainter limit of m_r < 17.77 and single- and coadded-exposure image sets.The primary analysis combines the original, extra, and Stripe 82 normal-depth images with m_r ≤ 17.0.
- Classification method: Each classification uses a 424 × 424 pixel gri colour composite, followed by a branching decision tree whose later tasks depend on earlier responses.The tree contains 11 tasks and 37 possible responses; Tasks 01 and 06 are answered for every classification.
- Participation and coverage: Main-sample galaxies received a median of 44 classifications, with more than 99.9% receiving at least 28; Stripe 82 coadd 2 galaxies had a median of 21 and more than 99.9% received at least 10.These repeated classifications support estimates of the likelihood of each morphology response.
- Participation and coverage: The project collected 16,340,298 classifications comprising 58,719,719 tasks from 83,943 volunteers over just over 14 months.Images with fewer responses were shown more often near the project’s end to improve coverage.
3 DATA REDUCTION
GZ2 reduces volunteer classifications by removing repeated votes, down-weighting low-consistency classifiers, and combining weighted responses into vote fractions. It then corrects redshift-dependent classification bias using local morphology baselines, producing debiased fractions that remain approximately stable across the corrected redshift range.
- Vote preparation: ∼1% of galaxies had repeated classifications by the same user, and removing them changed classifications for ≲0.01% of the sample.Only the last submission from repeat classifiers was retained.
- User weighting: User consistency κ compares each vote with task vote fractions, giving higher values to majority-agreeing votes and lower values to disagreeing votes.Each user's overall consistency is the mean across responses, and low-consistency classifiers are down-weighted iteratively.
- User weighting: ∼95% of classifiers receive weight w = 1, while only ∼1% receive w < 0.01; the weighting is recalculated twice to ensure convergence.The weighted votes and vote fractions are then used throughout the data release.
- Classification bias: Classification bias changes observed morphology fractions with redshift independently of true galaxy evolution, with finer-feature fractions generally decreasing at higher redshift.The strongest trend occurs for separating smooth from feature/disk galaxies, although nearly all tasks change to some degree.
- Classification bias: Bias corrections use local vote-fraction ratios from sufficiently sampled, spectroscopic-redshift galaxies and apply the resulting corrections to all sample galaxies.The correction is defined relative to a local baseline in bins of absolute magnitude and physical size, with task-specific vote thresholds.
- Classification bias: Debiased vote fractions are flat over 0.01 < z < 0.085, with early- and late-type fractions of 0.45 and 0.55 and a disk-galaxy bar fraction of approximately 0.35.These debiased type fractions agree with the corresponding GZ1 fractions for the same selection criteria.
4 THE CATALOGUE
The catalogue provides raw, weighted, and debiased vote data for GZ2 galaxies, with sample-specific corrections and conservative flags. Comparisons show that Stripe 82 coadded imaging changes classifications, while most coadd task results are consistent and morphology queries can select highly pure samples.
- Catalogue contents: GZ2 releases vote counts and raw, weighted, and debiased fractions for each classification-tree task, with clean-sample flags.The data are available electronically for five subsamples.
- Catalogue contents: Clean flags require preceding-task thresholds, at least 10 or 20 votes depending on sample, and a debiased fraction above 0.8.The thresholds are conservative and can be adjusted for different use cases.
- Bias correction: Spectroscopic and photometric-redshift samples should not have their debiased data combined because redshift strongly affects classification bias.Raw vote fractions may still be combined when galaxy count is the primary concern.
- Stripe 82: Coadded Stripe 82 images show more features-or-disk responses at all redshifts, increasing threshold-based unclassified galaxies and decreasing smooth classifications.Improved seeing and signal-to-noise likely help classifiers distinguish faint features and disks, so main-sample corrections cannot be reused.
- Stripe 82: 33/37 coadd tasks have |∆coadd| < 0.05 for galaxies with at least 10 responses, although bulge-prominence responses show the largest systematic difference.The mean just-noticeable-bulge fraction is 35% higher in coadd2, while the obvious-bulge fraction is 13% higher in coadd1.
- Using the classifications: A conservative three-armed-spiral query returns 308 galaxies, and visual inspection found no clear false positives.Lowering the p3 arms threshold may recover more genuine examples, but images are needed to distinguish radial symmetry from tidal-tail cases.
- Using the classifications: Debiased likelihoods can be used as probabilistic weights, retaining intermediate classifications that threshold-based samples would label unclassified.For example, vote fractions of 0.6, 0.3, and 0.1 contribute those weighted amounts to the three Task 01 responses.
5 COMPARISON OF GZ2 TO OTHER CLASSIFICATION METHODS
The paper compares GZ2 with four overlapping morphology catalogues derived from optical SDSS images. Agreement is evaluated using clean GZ2 likelihoods and majority or category assignments in the comparison catalogues.
- Comparison catalogues: GZ2 is compared with Galaxy Zoo 1, Nair & Abraham, EFIGI, and Huertas-Company catalogues.The comparison set includes citizen-science, expert visual, and automatic classifications.
- Comparison catalogues: The four comparison catalogues all use optical SDSS images and substantially overlap the GZ2 galaxy sample.
- Agreement measure: Table 4 counts agreement when GZ2 has likelihood at least 0.8 and the other catalogue has likelihood at least 0.5 or the relevant category assignment.It reports both the number of overlapping galaxies and the agreed fraction for each category.
- Agreement measure: The comparison section summarizes agreement between GZ2 and the other catalogues before discussing each comparison in detail.
5.1 Galaxy Zoo 1 and Galaxy Zoo 2
GZ2 broadly agrees with GZ1 for strong early- and late-type classifications while providing more detailed spiral and merger judgments. Differences are largest for intermediate spiral vote fractions and for tasks affected by classifier cross-talk.
- Ellipticals: 50.4% of 33,833 GZ1 ellipticals exceed 0.8 in both catalogues, while 97.6% exceed 0.5 in GZ2.Only about 1,200 GZ1 ellipticals have pfeatures/disk > 0.5 in GZ2; roughly 40% are barred and almost all show obvious bulges.
- Spirals: Among 83,956 GZ1 spirals, debiased GZ2 agreement is 38.1% at 0.8 and 78.2% at 0.5.Many low-pfeatures/disk cases are inclined disks or lenticular galaxies without spiral arms.
- Spirals: At intermediate GZ1 spiral vote fractions, GZ1 values exceed GZ2 by up to 0.25, while debiasing greatly reduces the difference but lowers correlation tightness at the extremes.
- Spirals: GZ2 is slightly more likely than GZ1 to identify galaxies as spirals using debiased likelihoods, with much lower spread in the joint clean samples.
- Mergers: More than 99% of 1,632 GZ1 merger systems are identified as odd galaxies in GZ2, and 77.7% have pmg > 0.5.Merger classification is complicated because multiple visible features can compete for a single response.
- Mergers: Task 08 suffers from cross-talk because participants must choose one response when galaxies exhibit multiple applicable odd features.For close pairs, merger votes increase as angular separation decreases while other odd-feature votes decline.
- Overall comparison: The matched catalogues contain nearly 250,000 galaxies, and early- versus late-type classifications are mostly consistent, especially above p > 0.8.GZ2 is more conservative than GZ1 in identifying spiral structure at intermediate vote fractions.
5.2 Expert visual classifications
Comparisons with expert catalogues show that GZ2 vote fractions reproduce expert classifications well for bars, while ring and merger classifications depend more strongly on thresholds, ring type, and task design.
- Bars: 78% of strong bars and 40% of intermediate bars have pbar > 0.8, compared with 9% of weakly barred galaxies.Using pbar > 0.5, these fractions rise to 94% and 80% for strong and intermediate bars, while 32% of weak bars exceed the threshold.
- Bars: ρ = 0.75 between EFIGI bar strength and GZ2 bar vote fraction, with 65% of galaxies having pbar < 0.3.Among galaxies with pbar < 0.3, 77% had EFIGI bar attributes of zero and 94% had attributes of 0.25 or less.
- Bars: GZ2 barred galaxies agree very well with EFIGI: less than 5% have EFIGI bar attributes of 0, with a mean attribute value of 0.56.The concentration around medium-length bars may reflect either selection preference or the prevalence of medium bars in disk galaxies.
- Bars: pbar ≥ 0.3 leaves less than 5% of galaxies classified as non-barred by both expert catalogues, while including 97% of strong and intermediate bars and 75% of weak bars.This threshold is more inclusive than the f ≥ 0.5 criterion used by Masters et al. (2011).
- Rings: Setting Nring > 5 raises ring-classification agreement to approximately 75%, whereas more than a third of galaxies with pring > 0.5 are ringless in NA10.Disagreement is concentrated among inner rings; 84% of 308 NA10 ringed galaxies with no GZ2 ring votes were inner rings.
- Rings: 89% of galaxies with pring > 0.5 and at least 10 “Anything odd?” votes were classified as rings in EFIGI, but inner rings had a lower mean vote fraction of 0.41.Mean ring vote fractions were 0.69 for outer rings and 0.71 for pseudo-rings.
5.3 Automated classifications
The paper compares GZ2 classifications with automated Huertas-Company et al. probabilities, finding broad agreement while revealing effects of bulge dominance and limited influence of bars.
- HC11 classifies 698,420 galaxies using an SVM, approximately twice the size of GZ2, but its sample requires z < 0.25, good photometry, and clean spectra.
- GZ2 robust classifications agree well with HC11 broad categories: median early-type probability is 0.85 for GZ2 ellipticals and late-type probability is 0.95 for GZ2 spirals.
- Some GZ2-smooth galaxies have low HC11 early-type probabilities, associated with fewer round and more cigar-shaped morphologies; their blue colours also contribute to the discrepancy.
- GZ2 bulge-dominance responses strongly affect HC11 late-type probabilities, with just-noticeable bulges increasing late-type probability and obvious bulges producing the opposite trend.
- 31% of oblique disk galaxies are classified as early-type and 69% as late-type by HC11 across GZ2 bar vote fractions, indicating bars do not strongly affect this broad split.
6 CONCLUSIONS
The GZ2 release provides large-scale, fine-grained citizen-science morphology measurements with calibrated vote products and comparisons showing strong broad-category agreement but feature-specific limits.
- GZ2 characterizes more than 300,000 SDSS galaxies using citizen-science votes, including normal and deeper Stripe 82 images selected by magnitude, size, and redshift.
- GZ2 extends beyond elliptical–spiral distinctions to identify bars, spiral structure, dust lanes, mergers, interactions, lenses, bulges, arm properties, and elliptica
- Classifier weighting, vote aggregation, and data-derived debiasing produce raw, weighted, and debiased vote fractions for each response.
- Comparisons find good agreement for broad types, medium-to-strong bars, and bulge dominance, while weak or nuclear bars, inner or nuclear rings, and clean pair or interaction matching are less reliable.
- GZ2 contains more than an order of magnitude more galaxies than the largest comparable expert catalogues while retaining detailed features not replicable by automated classifications.
- The public catalogue is intended as a reliable, rich dataset for studying galaxy morphology and its relationships with other galaxy properties.
APPENDIX A: GENERATING THE ABBREVIATION FOR A GZ2 MORPHOLOGICAL CLASSIFICATION
The appendix defines a compact gz2 class string that summarizes each galaxy’s consensus path through the GZ2 decision tree and records identified unusual features.
- The gz2 class is a convenient shorthand for the most common consensus classification, not a new classification system.
- The string selects the largest debiased vote fraction at Task 01 and then chooses the most common response at each subsequent decision-tree task.
- Smooth galaxies begin with ‘E’, with ‘r’, ‘i’, or ‘c’ encoding completely round, in-between, or cigar-shaped appearance.
- Feature/disk galaxies begin with ‘S’; later symbols encode edge-on bulge shape, oblique-disk bars, and bulge prominence.
- Parenthetical suffixes record user-identified odd features, including rings, lenses, disturbances, irregularities, mergers, and dust lanes.
- The appendix illustrates the twelve most common classifications with randomly selected galaxies at 0.050 < z < 0.055.