Source-linked AI summary

Galaxy Zoo: Reproducing Galaxy Morphologies Via Machine Learning

Manda Banerji, Ofer Lahav, Chris J. Lintott, Filipe B. Abdalla, Kevin Schawinski, Steven P. Bamford, Dan Andreescu, Phil Murray, M. Jordan Raddick, Anze Slosar, Alex Szalay, Daniel Thomas, Jan Vandenberg

arXiv:0908.2033v2astro-ph.COastro-ph.GA

TL;DR

The paper asks whether Galaxy Zoo’s visual morphology labels can be reproduced automatically for large SDSS samples. It trains an artificial neural network on human-classified objects, tests different image-parameter sets, and finds that a combined twelve-parameter set exceeds 90% agreement for all three classes. The results support using machine learning for future wide-field surveys, while highlighting incompleteness and rare-object misclassification as remaining concerns.

  • Problem

    The paper investigates whether visual Galaxy Zoo morphology classifications can be reproduced automatically for much larger datasets from future galaxy surveys.

  • Method

    An artificial neural network is trained on Galaxy Zoo classifications using SDSS-derived colours, profile-fitting, shape, concentration, and texture parameters.

  • Results

    Using twelve combined input parameters, the neural network achieves better than 90% agreement with human classifications for all three morphological classes.

  • Takeaways & Limitations

    The results indicate that machine-learning morphological classification is promising for next-generation wide-field imaging surveys, with Galaxy Zoo serving as a training set.

  • Takeaways & Limitations

    Further work is needed on training-set incompleteness, including sparse red-spiral and blue-elliptical examples that contribute to misclassification.

Abstract

from arXiv · show

We present morphological classifications obtained using machine learning for objects in SDSS DR6 that have been classified by Galaxy Zoo into three classes, namely early types, spirals and point sources/artifacts. An artificial neural network is trained on a subset of objects classified by the human eye and we test whether the machine learning algorithm can reproduce the human classifications for the rest of the sample. We find that the success of the neural network in matching the human classifications depends crucially on the set of input parameters chosen for the machine-learning algorithm. The colours and parameters associated with profile-fitting are reasonable in separating the objects into three classes. However, these results are considerably improved when adding adaptive shape parameters as well as concentration and texture. The adaptive moments, concentration and texture parameters alone cannot distinguish between early type galaxies and the point sources/artifacts. Using a set of twelve parameters, the neural network is able to reproduce the human classifications to better than 90% for all three morphological classes. We find that using a training set that is incomplete in magnitude does not degrade our results given our particular choice of the input parameters to the network. We conclude that it is promising to use machine- learning algorithms to perform morphological classification for the next generation of wide-field imaging surveys and that the Galaxy Zoo catalogue provides an invaluable training set for such purposes.

1 INTRODUCTION

Galaxy classification has traditionally relied on human inspection, but artificial neural networks have shown promise for reproducing visual classifications. This paper tests neural-network classification of SDSS objects into three morphological types.

  • Galaxy morphology classification has long been an astronomical research goal, with human inspection still commonly used.
  • Artificial neural networks have previously reproduced visual classifications and become prominent for astronomical applications.
  • Galaxy Zoo produced classifications for nearly 1 million SDSS DR6 objects through visual inspection by more than 100,000 users.
  • This paper explores whether artificial neural networks can classify SDSS objects as early types, spirals, or point sources/artifacts.

2 THE GALAXY ZOO CATALOGUE

The study uses a weighted Galaxy Zoo catalogue matched to SDSS photometric data, while retaining the original morphology classifications and applying quality cuts. The catalogue records class vote fractions and is subject to known classification biases.

  • Galaxy Zoo’s weighted catalogue contains classifications for 893,212 objects across ellipticals, spirals, mergers, and point sources/artifacts.
  • User votes are weighted by agreement with the majority, and each object’s final morphology is the weighted mean of its users’ classifications.
  • The catalogue records weighted vote fractions for elliptical, spiral, and point-source/artifact classifications, with residual votes assigned to mergers.
  • The catalogue is matched to SDSS DR7 PhotoObjAll to obtain neural-network inputs, after removing objects lacking g, r, or i detections or having problematic parameter values.
  • Faint disky objects may be classified as ellipticals, so the paper calls this potentially contaminated class early types and also defines a brighter r < 17 sample.

5 Note that the SDSS object IDs correspond for objects in DR6 and DR7

Figure 1 schematically depicts how human observers and machine-learning algorithms derive morphological classifications and image parameters from galaxy images.

  • Figure 1 compares human-eye and machine-learning routes from galaxy images to morphological classifications and extracted parameters.

3 ARTIFICIAL NEURAL NETWORKS

The paper uses an artificial neural network whose inputs are image-derived parameters and whose outputs represent probabilities for three morphological classes. Its architecture is selected using the number of inputs, with validation available to control over-fitting.

  • The neural network maps input parameters and Galaxy Zoo fractional votes to predicted probabilities for the morphological classes.
  • The network uses two hidden layers with 2N nodes each, giving an architecture of N:2N:2N:3 for N input parameters.
  • A validation set may be added to the training set to prevent over-fitting when the data are noisy or the network is highly flexible.
  • The three output nodes represent probabilities for early types, spirals, and point sources/artifacts.

4 INPUT PARAMETERS

The paper compares input-parameter sets for morphological classification, contrasting colours and profile-fitting features with concentration, adaptive shape parameters, and texture. The gold-sample distributions show that these parameters encode differences among early types, spirals, and point sources/artifacts.

  • Colours and profile fitting: The first parameter set uses dereddened (g-r) and (r-i) colours together with profile-fitting parameters.The colours are not k-corrected, and the profile-fitting inputs include quantities from the i-band images.
  • Observed parameter distributions: The gold sample retains objects with a Galaxy Zoo vote fraction greater than 0.8, reducing contamination but excluding many potentially well-classified early types and spirals.The threshold is described as arbitrary, and fractional vote differs from classification probability despite strong correlation.
  • Observed parameter distributions: In the gold sample, early types are redder and have larger de Vaucouleurs-fit axis ratios and log likelihoods than spirals, while point sources/artifacts span a wide colour range.Typical de Vaucouleurs-fit axis ratios are approximately 0.8 for early types and 0.3 for spirals; point sources/artifacts have a bimodal axis-ratio distribution.
  • Adaptive shape and texture: The second parameter set excludes colours and profile-fitting parameters, instead combining concentration, adaptive shape parameters, and texture.Concentration uses ratios of radii containing 90% and 50% of the Petrosian flux; adaptive moments are measured with an iteratively adapted radial Gaussian weight.
  • Adaptive shape and texture: Texture measures the ratio of surface-brightness fluctuation range to full dynamic range, vanishing for smooth profiles and becoming non-zero when structures such as spiral arms appear.
  • Observed parameter distributions: Adaptive ellipticity is large for spirals, small for early types, and slightly smaller still for point sources/artifacts, while concentration is larger for early types than spirals.

5 RESULTS

The neural network’s performance varies substantially with the input-parameter set: traditional parameters provide reasonable classification, adaptive parameters alone fail for point sources, and the combined set exceeds 90% agreement.

  • Input-parameter comparisons: 87% of early types, 86% of spirals and 95% of point sources/artifacts are correctly classified using colours and traditional profile-fitting parameters.The paper notes that colour information may bias morphology because colour and morphology trace related but distinct properties.
  • Input-parameter comparisons: 84% of early types and 87% of spirals are correctly classified using adaptive shape parameters, but only 28% of point sources/artifacts are.Adaptive parameters are similar for early types and point sources/artifacts, while spirals differ more clearly.
  • Input-parameter comparisons: 92% of early types, 92% of spirals and 96% of point sources/artifacts are correctly classified after combining the two parameter sets.The combined set contains twelve input parameters.
  • Evaluation: Figure 3 evaluates seven colour and profile-fitting parameters by plotting class probability against contaminants and discarded Galaxy Zoo objects.The panels cover early types, spirals, and point sources/artifacts.

Spiral

The study uses Galaxy Zoo as a human-classification reference and finds that carefully chosen, accessible parameters allow neural-network classifications to agree with it at better than 90%.

  • Spiral: The neural network agrees with Galaxy Zoo classifications to better than 90% using twelve readily available but not fully optimized parameters.The result concerns morphological classification across the paper’s three target classes.
  • Spiral: The results are summarized for three neural-network input-parameter sets and evaluated on a gold sample of Galaxy Zoo objects.The cited experiment passage states that the results are summarized in Tables 6, 7 and 8.
  • Spiral: Figure 4 examines adaptive moments, concentration and texture using class probabilities, contaminants, and discarded Galaxy Zoo objects.The five parameters are the second input set described in Table 2.

GALAXY ZOO

This section summarizes results for the entire sample using adaptive moments and distinguishes two point-source/artifact probability thresholds.

  • Table 4 summarizes results for the entire sample using the adaptive-moments input parameters specified in Table 2.
  • Objects with point source/artifact probability above approximately 0.2 belong to the point sources/artifacts class.
  • The section applies a stricter point source/artifact probability requirement of greater than 0.8.

5.3 The Bright Sample

The bright-sample analysis tests whether brighter objects classify better and whether training on a magnitude-limited subset affects classifications for the full sample.

  • 5.3 The Bright Sample: The bright sample is defined by r < 17 and uses the combined twelve-parameter input set.The analysis compares bright-sample performance with the entire sample and applies the bright-trained network to the full catalogue.
  • 5.3 The Bright Sample: Training on the bright sample allows the study to quantify magnitude incompleteness effects when classifying the entire sample.The results are summarized in Tables 9 and 10.
  • 5.3 The Bright Sample: The bright-sample classifications are slightly better than those for the entire sample.This comparison is made between Tables 5 and 9.
  • 5.3 The Bright Sample: Figure 5 evaluates the combined twelve-parameter set using class probability, contaminants, and discarded objects for early types, spirals, and point sources/artifacts.The figure has separate panels for the three morphological classes.

GALAXY ZOO

The study uses Galaxy Zoo classifications to train machine-learning morphological classifiers and examines their performance across samples. Results support applying these methods to future deep surveys.

  • Table 5 summarizes results for the entire sample using the input parameters specified in Tables 1 and 2.
  • More than 90% agreement with Galaxy Zoo users is retained when an incomplete bright training set classifies the full deeper sample.The training set uses r < 17, while classification covers objects with r < 17.77.
  • Galaxy Zoo classifications can serve as a training set for automated morphological classification in future deep surveys.

6 CONCLUSIONS

Artificial neural networks reproduce human classifications of SDSS objects across three morphological classes, with performance depending strongly on the chosen input parameters. The results support automated classification for future surveys, while highlighting unresolved training-set incompleteness and data-quality constraints.

  • 6 CONCLUSIONS: 87% of early-type, 86% of spiral, and 95% of point-source/artifact classifications agree with human classifications using colours and profile-fitting parameters.
  • 6 CONCLUSIONS: Adaptive weighted-fitting parameters alone cannot distinguish early types from point sources/artifacts because their parameter values are very similar.
  • 6 CONCLUSIONS: More than 90% agreement is achieved for early types, spirals, and point sources/artifacts using combined profile-fitting and adaptive weighted-fitting parameters.The network was trained on 75,000 objects and classified the remaining Galaxy Zoo objects.
  • 6 CONCLUSIONS: For the gold sample, early-type and spiral classifications match human classifications by better than 95%.
  • 6 CONCLUSIONS: A bright training sample still produces better than 90% agreement on a deeper sample because the selected input parameters are distance independent.
  • 6 CONCLUSIONS: Other sources of training-set incompleteness require investigation, including sparse representation of red spirals and blue ellipticals that are misclassified by the network.
  • 6 CONCLUSIONS: Future wide-field surveys can use Galaxy Zoo data to classify vast object populations, provided images have sufficient pixel size and resolution for the required photometric parameters.
  • 6 CONCLUSIONS: The Galaxy Zoo catalogue is promising for machine-learning morphology studies, but merger classification requires a more robust visually classified merger catalogue.
Loading 0908.2033v2…