Source-linked AI summary

A catalog of visual-like morphologies in the 5 CANDELS fields using deep-learning

M. Huertas-Company, R. Gravet, G. Cabrera-Vives, P. G. Pérez-González, J. S. Kartaltepe, G. Barro, M. Bernardi, S. Mei, F. Shankar, P. Dimauro, E. F. Bell, D. Kocevski, D. C. Koo, S. M. Faber, D. H. Mcintosh

arXiv:1509.05429v1astro-ph.GAastro-ph.CO

TL;DR

Galaxy morphology at different cosmic epochs is difficult to classify from massive imaging datasets while retaining the information present in galaxy pixels. This paper uses deep-learning ConvNets to produce visual-like H-band morphologies for about 50,000 galaxies across five CANDELS fields, achieving close agreement with expert classifications and releasing the resulting catalog.

  • Problem

    Traditional CAS-based parameters reduce galaxy images to a few global descriptors and neglect substantial pixel-level information needed for visual-like morphology classification.

  • Method

    A 5-layer ConvNet followed by 2 fully connected perceptron layers is trained on visual classifications of about 8,000 GOODS-S galaxies and applied to about 50,000 galaxies in five CANDELS fields.

  • Results

    ConvNets predict expert-classifier votes with < 10% bias and ∼10% scatter, while miss-classifications remain below 1%.

  • Takeaways & Limitations

    The released catalog expands public CANDELS morphologies by a factor of 5 and supports studies of merger rates, morphological evolution, environment, and morphology-AGN connections.

  • Takeaways & Limitations

    Classification agreement decreases for faint, distant, and low-mass galaxies, with magnitude—and therefore signal-to-noise—being the strongest limitation.

Abstract

from arXiv · show

We present a catalog of visual like H-band morphologies of $\sim50.000$ galaxies ($H_{f160w}<24.5$) in the 5 CANDELS fields (GOODS-N, GOODS-S, UDS, EGS and COSMOS). Morphologies are estimated with Convolutional Neural Networks (ConvNets). The median redshift of the sample is $<z>\sim1.25$. The algorithm is trained on GOODS-S for which visual classifications are publicly available and then applied to the other 4 fields. Following the CANDELS main morphology classification scheme, our model retrieves the probabilities for each galaxy of having a spheroid, a disk, presenting an irregularity, being compact or point source and being unclassifiable. ConvNets are able to predict the fractions of votes given a galaxy image with zero bias and $\sim10\%$ scatter. The fraction of miss-classifications is less than $1\%$. Our classification scheme represents a major improvement with respect to CAS (Concentration-Asymmetry-Smoothness)-based methods, which hit a $20-30\%$ contamination limit at high z. The catalog is released with the present paper via the $\href{http://rainbowx.fis.ucm.es/Rainbow_navigator_public}{Rainbow\,database}$

1. INTRODUCTION

Deep-field morphology classification is essential for studying galaxy evolution, but expanding datasets challenge human inspection and conventional automated methods. The paper extends deep learning to high-redshift CANDELS galaxies, producing a large visual-like morphology catalog.

  • Galaxy morphology helps investigate how bulges and disks form and evolve across cosmic time.
  • Human classification is difficult to scale as deep-field surveys produce increasingly large datasets.
  • CAS-based methods depend strongly on data quality and redshift, provide only rough classes, and reach ∼20−30% miss-classifications at high redshift.
  • Pixel-based deep learning can retain image information that compact parameters such as concentration and asymmetry discard.
  • The paper classifies ∼50.000 galaxies at median redshift <z>∼1.25 across five CANDELS fields and releases the resulting catalog.

2. DATASET

The dataset selects H-band galaxies consistently across the five CANDELS fields using the magnitude limit adopted for reliable visual morphology. It contains 50.000 galaxies, with substantial coverage of the 1 < z < 3 optical-rest-frame regime.

  • The sample selects all F160W galaxies with F160W<24.5 mag across the considered fields.
  • 50.000 galaxies comprise the resulting sample, increasing the existing public CANDELS visual catalog by a factor of 5.
  • About 50% of sources lie between 1 < z < 3, where CANDELS filters probe optical rest-frame morphologies.
  • The sample is ∼80% complete down to log(M∗/M⊙) ∼10.

3. CANDELS MORPHOLOGICAL CLASSIFICATION WITH DEEP LEARNING

The paper uses ConvNets to learn visual morphology from galaxy images and reproduce human vote fractions. Pre-processing adapts the smaller, high-redshift training set to the network, while held-out evaluation tracks convergence and error reduction.

  • ConvNet approach: Deep learning automatically extracts task-relevant features from raw data through nonlinear transformations.
  • Training and evaluation: The model is trained on ∼8.000 GOODS-S galaxies with expert visual classifications and evaluated using independent test data.
  • ConvNet approach: The ConvNet architecture uses five convolutional layers, two fully connected perceptron layers, and three max-pooling steps.
  • Morphology targets: The model reproduces five CANDELS vote fractions for spheroids, disks, irregularities, point sources, and unclassifiable objects.
  • Pre-processing: Because the training set is much smaller than the SDSS set, the workflow addresses over-fitting through interpolation, rotations, redundancy, and label noise.
  • Performance: The test-set RMSE is consistent with validation RMSE, while model averaging reduces RMSE by ∼10−3.
  • Performance: The average RMSE decreases from ∼0.25 before pre-processing to ∼0.13 after adding label noise, interpolation, redundancy, and noise.

4. ACCURACY

The ConvNet closely reproduces visual morphological vote fractions and dominant classes, with low bias and scatter across the test sample. Its accuracy remains broadly stable across magnitude and redshift, while uncertainty increases for faint, distant, and low-mass galaxies.

  • Recovering votes: The test-set predictions show clear one-to-one agreement with visual fractions, with typical bias and dispersion below 10%.Median bias ranges from 0–0.02 and scatter from 0.03–0.1 across morphological frequencies.
  • Recovering votes: Bias stays below 0.05 and scatter near 0.1 across the sampled magnitudes and redshifts, with larger bias only for objects near the PSF size or larger than four PSF sizes.This measures reproduction of the visual classification, including its eventual biases, rather than brightness-independent morphology assessment.
  • Recovering dominant classes and miss-classifications: For clear dominant classes, visual and automatic classifications agree at more than 95%, while CAS-based methods report about 20% dominant-class misclassification.The ConvNet also distinguishes unclassifiable and point-source objects that CAS methods may confuse with galaxy classes.
  • Recovering dominant classes and miss-classifications: The agreement parameter links classifier consensus to accuracy: less-consistent human classifications are harder for the model to recover.Agreement ranges from 0 to 1, with larger values indicating that most classifiers selected the same class.
  • Secondary classes - multi-component objects: For two-component galaxies, agreement is close to 95% for disk-spheroid and disk-irregular combinations but poor for the marginal spheroid-irregular class.The model therefore recovers secondary components well when galaxies contain two clear morphological structures.
  • Uncertain objects - Limitations: Fewer than 5% of galaxies have maximum visual or automatic frequency below 0.4–0.5, while uncertain classifications increase with magnitude, redshift, and lower stellar mass.The automated classifier reproduces the uncertainty trends encountered by human classifiers.
  • Uncertain objects - Limitations: The strongest limitation is signal-to-noise: classifier agreement declines for faint, distant, and low-mass objects, although median agreement remains above 0.4.The reported agreement corresponds to accuracy above 80% for all objects.

5. ACCURACY IN ALL CANDELS FIELDS

The authors extend automated morphology classification to CANDELS fields without existing visual inspection and test whether performance remains consistent across fields. Field distributions show no significant differences, while discrepancies in UDS are attributed to differences in visual-label sampling rather than algorithm behavior.

  • No significant field-to-field differences appear in morphological frequency distributions, suggesting similar algorithm behavior across CANDELS fields.The authors note that the machine smooths the visual distribution by removing gaps or abrupt changes.
  • UDS visual classification: A different visual-classification frequency distribution explains the weaker UDS comparison, because ∼90% of UDS galaxies were classified by only 3 people.The authors compare this effect with GOODS-S after restricting it to 3 classifiers per galaxy.
  • UDS visual classification: In the 3-classifier GOODS-S test, scatter and bias increase because the input distribution changes, while the automated output remains exactly the same.The similar trends between this test and UDS suggest that the UDS worsening is not caused by poor algorithm behavior in that field.
  • UDS visual classification: Fewer visual annotators increase the variance of class fractions, so noisier training labels produce noisier classifications.For binary labels, the fraction variance depends on the intrinsic probability and the number of annotations.
  • ∼95% agreement is obtained for the DS and DI two-component classes, indicating recovery of both primary and secondary morphology.Agreement is poor for the marginal SI class, which contains very few objects.
  • UDS visual classification: Automated classifications remain homogeneous across datasets, and similar UDS and three-classifier GOODS-S results support the absence of severe over-fitting.This is presented as an advantage over visual classifications when only a small number of classifiers is available.

6. CATALOG

The catalog provides continuous morphology indicators and optional threshold-based classes for CANDELS galaxies, alongside quality measures and examples of how these classes relate to galaxy properties.

  • Catalog contents: The catalog supplies five main indicators, agreement measures, dominant class, and maximum frequency for galaxies brighter than HF160W = 24.5.It is released through the Rainbow database.
  • Catalog representation: Each galaxy receives five continuous morphological parameters spanning 0 to 1, so class definitions depend on scientific purpose and galaxy properties.Threshold choices trade off sample purity against completeness.
  • Illustrative classes: Five illustrative classes separate pure bulges, pure disks, disk+spheroid systems, irregular disks, and irregulars or mergers using thresholds on fsph, fdisk, and firr.These thresholds are presented as one possible classification scheme rather than a unique definition.
  • Class interpretation: The class scheme distinguishes bulge-dominated, disk-dominated, mixed, and irregular systems, with thresholds calibrated through visual inspection.The DISKSPH class contains galaxies without a clearly dominant component.
  • Quality and uncertainty: Uncertain classifications become more common for fainter, higher-redshift, and lower-mass objects in both automatic and visual classifications.The figure compares uncertainty across fmax thresholds and relates agreement to magnitude, redshift, and stellar mass.
  • Galaxy properties: Morphological types show distinct but imperfect population trends: bulge-containing systems tend toward larger Sérsic indices and passive UVJ locations, while disk-dominated systems peak near n ∼1 and are star-forming.The bulge+spheroid class is split between passive and star-forming systems, illustrating why morphology can isolate populations that color selection divides.

7. SUMMARY AND CONCLUSIONS

This work applies a ConvNet to produce visual-like H-band morphologies for about 50,000 galaxies across the five CANDELS fields. The model reproduces expert classifications with low bias and scatter, and the resulting catalog is publicly released for further studies.

  • Dataset and method: The study classifies ∼50.000 galaxies with H < 24.5 across five CANDELS fields using a ConvNet trained on ∼8000 visually classified GOODS-S galaxies.The sample probes optical rest-frame morphologies at 1 < z < 3 and is ∼80% complete down to log(M∗/M⊙) ∼10.
  • Dataset and method: Each galaxy receives five frequencies representing spheroid, disk, irregularity, compact or point-source, and unclassifiable classifications.Images are resized, rotated, and randomly perturbed before being fed to the network.
  • Performance: < 10% bias and ∼10% scatter characterize the ConvNet predictions of expert-classifier votes.The authors describe this performance as making the classification almost equivalent to a visual one.
  • Performance: < 1% miss-classifications are reported for identifying non-galaxies, while the method distinguishes irregulars from disks and spheroids from disks at all redshifts.The comparison is made against generalized CAS methods such as galSVM.
  • Catalog release: The ∼50.000-galaxy morphology catalog is released through the Rainbow database and increases the existing public CANDELS morphologies by a factor of 5.The catalog is intended for applications including merger-rate evolution and morphology–environment studies.
  • Future directions: Future work targets optimization for EUCLID, WFIRST, and LSST-like data, deeper Hubble Frontier Fields data, and more detailed descriptors such as tidal features.
Loading 1509.05429v1…