Source-linked AI summary

A comprehensive review on convolutional neural network in machine fault diagnosis

Jinyang Jiao, Ming Zhao, Jing Lin, Kaixuan Liang

arXiv:2002.07605v1eess.SPcs.LGstat.ML

TL;DR

Machine faults can cause unscheduled downtime, economic loss, catastrophic accidents, and casualties. This review organizes CNFD around data collection, model construction, feature learning, and decision making, finding broad emphasis on classification while identifying prediction and transfer diagnosis as important directions.

  • Problem

    Machine faults can cause unscheduled downtime, economic loss, catastrophic accidents, and casualties.

  • Method

    The review summarizes the CNFD framework as data collection, model construction, and feature learning and decision making.

  • Results

    More than 65% of publications focus on fault classification, while transfer diagnosis approaches mostly target health condition monitoring.

  • Takeaways & Limitations

    Future studies should give more attention to degradation monitoring or remaining useful life prediction and quantitative early-damage detection.

  • Takeaways & Limitations

    Models developed with acquired or simulated data may be unsuitable for direct application to realistic industry because these data differ from industrial data.

Abstract

from arXiv · show

With the rapid development of manufacturing industry, machine fault diagnosis has become increasingly significant to ensure safe equipment operation and production. Consequently, multifarious approaches have been explored and developed in the past years, of which intelligent algorithms develop particularly rapidly. Convolutional neural network, as a typical representative of intelligent diagnostic models, has been extensively studied and applied in recent five years, and a large amount of literature has been published in academic journals and conference proceedings. However, there has not been a systematic review to cover these studies and make a prospect for the further research. To fill in this gap, this work attempts to review and summarize the development of the Convolutional Network based Fault Diagnosis (CNFD) approaches comprehensively. Generally, a typical CNFD framework is composed of the following steps, namely, data collection, model construction, and feature learning and decision making, thus this paper is organized by following this stream. Firstly, data collection process is described, in which several popular datasets are introduced. Then, the fundamental theory from the basic convolutional neural network to its variants is elaborated. After that, the applications of CNFD are reviewed in terms of three mainstream directions, i.e. classification, prediction and transfer diagnosis. Finally, conclusions and prospects are presented to point out the characteristics of current development, facing challenges and future trends. Last but not least, it is expected that this work would provide convenience and inspire further exploration for researchers in this field.

1. Introduction

Machine fault diagnosis is increasingly important as manufacturing becomes more complex, while existing physics-based, signal-processing, and classical machine-learning approaches face practical limitations. This paper therefore comprehensively reviews convolutional-network-based fault diagnosis (CNFD), following its data, modeling, feature-learning, and decision-making workflow.

  • Motivation: Traditional diagnosis approaches are constrained by complex machine mechanisms, noisy environments, manual feature engineering, shallow representations, and difficulty handling growing data diversity.Physics-based models can be inflexible and difficult to update, while classical machine-learning methods separate feature mining from decision making.
  • Motivation: Deep learning has accelerated intelligent fault diagnosis through data processing, feature learning, architecture innovation, industrial big data, and improved hardware.The paper identifies convolutional networks as a leading deep-learning architecture with strong benchmark performance.
  • Research gap: CNFD has rapidly expanded, creating a need for a systematic review that organizes current work and prospects for further research.Earlier reviews emphasized traditional machine learning, treated CNN alongside other deep models, focused on bearings, or omitted emerging transfer-learning studies.
  • CNFD framework: A general CNFD framework collects monitoring data, constructs convolutional models for task requirements, and jointly performs feature learning and decision making.The framework supports adaptive hierarchical feature learning and decisions such as fault classification and remaining useful life prediction.
  • Paper scope: The review covers public datasets, CNN theory and variants, and CNFD applications in classification, prediction, and transfer diagnosis.Its organization follows the CNFD workflow from data collection through model construction, applications, conclusions, and prospects.

2. Data preparation

Data collection is the foundation of CNFD, requiring suitable sensors, placement, sampling, and storage while coping with low-quality or scarce industrial fault data. Public datasets provide practical resources for evaluating diagnosis, transfer, and sensor-fusion methods.

  • Data collection: High-quality data is treated as the premise for successfully training convolutional neural networks.The collection process includes sensor selection and layout, sampling, and storage through acquisition systems and hard-disk or cloud platforms.
  • Sensor selection: Vibration, current, and built-in encoder sensors provide complementary monitoring data, but vibration measurements can suffer from interference and limited low-frequency sensitivity.Vibration sensing may also be impractical in high-temperature, high-pressure, or enclosed environments.
  • Sensor selection: Alternative sensors such as infrared imaging and built-in encoder signals can circumvent some vibration-sensing limitations.Infrared imaging offers non-contact measurement, while encoder signals provide better signal-to-noise ratio and low-frequency response.
  • Sensor selection: Sensor choice and placement should consider equipment, environment, monitoring target, and operating condition because location affects perceived health information and interference.The paper emphasizes selecting well-suited sensors and proper locations before sampling.
  • Data limitations: Real industrial data are difficult to acquire because fault operation is often prohibited and full life-cycle collection is time consuming, expensive, or prohibitive.Public datasets are introduced as resources for researchers evaluating diagnosis methods.
  • CWRU dataset: CWRU provides vibration data from accelerometers on a motor platform across four operating conditions and several simulated bearing-fault sizes.Its data support studies of sensor-position generalization and transfer diagnosis across fault sizes or operating conditions, although the faults differ from natural industrial damage.

2.2. PHM09 gearbox fault dataset

The PHM 2009 dataset captures gearbox health under varied speeds and loads using synchronized vibration and tachometer measurements. Its spur and helical gear conditions support multiclass, hybrid-fault, sensor-fusion, and transfer-diagnosis studies.

  • Dataset description: PHM 2009 contains an industrial gearbox with three shafts, four gears, and six bearings, tested using spur and helical gears.The dataset includes eight spur-gear health conditions and six helical-gear health conditions.
  • Dataset description: Vibration signals were sampled at 66.67 kHz from two accelerometers, while tachometer signals were collected at 10 pulses per revolution.Each data file contains two vibration columns and one tachometer column.
  • Dataset characteristics: The dataset supports multiclass diagnosis because it includes multiple health conditions for both spur and helical gears.The paper explicitly identifies it as suitable for multi-classification diagnosis.
  • Dataset characteristics: Two accelerometers enable double-sensor information fusion and transfer diagnosis across sensor positions.The dataset also includes multiple working conditions for transfer diagnosis under different speeds and loads.
  • Dataset characteristics: Hybrid faults involving gears, bearings, and shafts make PHM 2009 suitable for hybrid fault diagnosis.The database combines several component-fault types within the gearbox setting.

2.3. Paderborn dataset

The Paderborn dataset uses a dedicated bearing test bench to collect synchronized current and vibration signals from healthy and damaged bearings. Its artificial and real damages, multiple operating conditions, and fault-formation modes support diverse diagnosis and transfer studies.

  • Dataset description: Paderborn experiments used a rig containing an electric motor, torque shaft, rolling-bearing module, flywheel, and load motor.Bearings in different states were installed in the test module for data acquisition.
  • Dataset composition: The dataset includes 26 faulty bearings and 6 healthy bearings, comprising 12 artificial and 14 real damages.Artificial damages are described separately from real damages caused by accelerated lifetime testing.
  • Data modalities: Motor current and bearing-housing vibration signals were synchronously sampled at 64 kHz.This enables independent signal studies as well as multi-sensor information fusion.
  • Dataset characteristics: Four operating conditions vary rotational speed, load torque, and radial force, supporting transfer diagnosis across working conditions.The dataset is also suitable for transfer diagnosis between artificial and real fault-formation modes.
  • Dataset characteristics: The dataset is more comprehensive than CWRU for some studies because it includes artificial and realistic bearing damages simultaneously.Its damage-state diversity supports multi-fault classification and comparisons between current and vibration signals.

2.4. IMS bearing dataset

The IMS bearing dataset supports classification and RUL prediction through run-to-failure experiments with multiple bearing fault conditions, but its single operating condition and variable lifetimes constrain diversity and prediction difficulty.

  • Dataset setup: The dataset contains three test-to-failure experiments, with inner-race, roller-element, and outer-race failures occurring in different bearings.The first experiment produced inner-race and roller-element defects; the second and third produced outer-race failures.
  • Applications: Four health conditions—health, roller fault, outer race fault, and inner race fault—support bearing classification.
  • Applications: Each record describes a run-to-failure experiment, enabling bearing RUL prediction.
  • Characteristics and limitations: An “increase-decrease-increase” degradation trend reflects self-healing damage, making data selection during the decrease period more difficult.
  • Characteristics and limitations: Only one operating condition limits data diversity, while distinct unit lifetimes increase RUL prediction difficulty.

2.6. PHM 10 CNC Milling Machine Cutters Dataset

The PHM 2010 CNC milling cutter dataset provides multimodal signals for RUL prediction, but dry milling, limited labels, and one operating condition restrict realism and diversity.

  • Dataset setup: The dataset targets RUL estimation for high-speed CNC milling machine cutters and contains six individual cutter data records.Cutting forces, three-direction vibrations, and acoustic emission were measured during machining.
  • Data collection: Measurements combine three-dimensional cutting forces, three-dimensional vibration data, and acoustic emission signals.This supports single-sensor and multi-sensor fusion prediction scenarios.
  • Limitations: The dataset was collected under dry milling, which differs from real milling and restricts cross-validation prediction scenarios.
  • Limitations: Only three of the six cutters are labeled, potentially providing insufficient data for complex diagnostic networks.
  • Limitations: A single milling operating condition limits the diversity of the dataset.
  • Related dataset: The FEMTO dataset provides 17 full-life bearing records from healthy starts through degradation, using vibration and temperature sensors.Bearing life ends when vibration amplitude exceeds 20 g.
  • Related dataset: FEMTO is useful for bearing RUL prediction but has a small training set, widely varying lifetimes, and degradation differences across bearings.The dataset provides natural degradation without seeded defects, but no prior information about damage properties.

2.8. Epilog

The review summarizes seven popular public datasets by monitoring object, sensing information, and applicability to classification, RUL prediction, and transfer diagnosis, while noting that other datasets also exist.

  • Dataset summary: The review summarizes seven popular public datasets and presents their key characteristics in Table 9.
  • Dataset summary: Table 9 compares monitoring objects, multiple-sensor information, and applications in classification, RUL prediction, and transfer diagnosis.
  • Dataset summary: The listed datasets are representative rather than exhaustive, with additional public resources including MFPT, XJTU-SY, and the University of Connecticut gear fault dataset.
  • Dataset summary: PHM society and IEEE reliability society conferences also provide valuable datasets for researchers.

3. Convolutional neural network and its variants

The review introduces CNN fundamentals for machine fault diagnosis and then describes residual, densely connected, and generative adversarial variants. It covers convolutional feature extraction, regularization, decision outputs, and deeper-network designs.

  • Basic CNN: A basic CNN for one-dimensional mechanical signals contains input, convolution-pooling, fully connected, and output layers.Batch normalization and dropout may also be embedded to improve model performance.
  • Convolution and activation: Convolution uses sparse local connections and shared weights to learn comprehensive feature representations from mechanical data.Multiple kernels generate feature maps from local patches of preceding-layer outputs.
  • Convolution and activation: ReLU is widely used because it computes faster than sigmoid and tanh and can alleviate gradient vanishing, although inactive units may remain inactive.
  • Pooling: Pooling reduces feature dimensionality and improves robustness through maximum- or average-value aggregation over pooling regions.
  • Regularization and normalization: Batch normalization addresses internal covariance shift and promotes network training, while dropout randomly removes units during training to prevent overfitting.
  • Decision layer: CNN decision layers commonly output classification labels or single variables such as RUL, with Softmax producing positive class probabilities that sum to 1.
  • Residual network: ResNet uses shortcut connections and residual mappings to address degradation and gradient vanishing or exploding in deeper CNNs.A residual block combines the learned mapping f(x,w) with x through element-wise addition.
  • Densely connected network: DenseNet connects each layer to all preceding layers, concatenating earlier feature maps before applying BN, ReLU, and convolution.

4. Applications on CNFD

The review organizes recent CNFD applications into three directions: fault classification, health prediction, and transfer diagnosis.

  • CNFD applications are reviewed across fault classification, health prediction, and transfer diagnosis.

4.1. Applications on fault classification

Fault-classification studies are organized by convolutional-network structure, covering 2-D CNNs, 1-D CNNs, and CNN variants. Because mechanical data are usually 1-D time series, many methods transform signals into 2-D inputs or derive time- and frequency-domain representations.

  • Organization by network structure: Fault-classification studies are categorized as 2-D CNN, 1-D CNN, or convolutional-network-variant approaches.
  • Input representation: Mechanical data are usually 1-D time series, motivating transformations into 2-D forms for CNN processing.
  • Data matrix transformation: Data-matrix methods directly arrange raw mechanical data into 2-D inputs, including Hankel matrices and transformed vibration matrices.
  • Convolutional-network variants: Reviewed variants combine CNNs with optimization, fusion, recurrent, ensemble, autoencoder, residual, and adversarial-generation techniques.
  • Image transformation: Image-transformation methods align sensor signals or convert raw signals into vibration, gray, RGB, or other images before CNN classification.
  • Time or frequency domain transformation: Time- or frequency-domain methods use statistical measures, Fourier features, or spectral maps as convolutional-network inputs.
  • Reported results: Several reviewed approaches report higher accuracy, stronger feature learning, or superiority over traditional CNNs and other neural networks.
  • Reported results: GAN-generated samples improved diagnosing accuracy from 95% to 98% when the model was trained with the generated data.

4.2. Applications on health prediction

Health-prediction research tracks machinery degradation and supports early maintenance decisions. Reviewed CNN approaches estimate bearing, turbofan, tool-wear, and wind-turbine health using transformed or raw sensor data, often combined with temporal models and uncertainty methods.

  • Health prediction tracks machinery degradation before apparent failure, supporting early maintenance judgments and decisions.
  • Rolling bearings: Bearing RUL methods use raw signals, spectral features, STFT, wavelets, Hilbert-Huang representations, or time-frequency images as CNN inputs.
  • Rolling bearings: CNNs are combined with LSTM, recurrent, double-CNN, ensemble, and separable-convolution designs for feature learning and RUL estimation.
  • Uncertainty modeling: Bayesian multi-scale convolutional prognosis showed more accurate performance than point estimates, addressing uncertainty in health prediction.
  • Multi-task and uncertainty methods: Joint-loss and recurrent convolutional models address related diagnosis and prediction tasks while modeling temporal dependencies or quantifying uncertainty.
  • Turbofan engines: For turbofan engines, reviewed approaches construct 2-D matrices or time windows and use CNN, residual, DAG, or CNN-LSTM models for RUL estimation.
  • Other applications: CNN-based prediction methods also address tool wear, wind-turbine gearbox bearings, SCADA-based monitoring, and CNC machining health.
  • Section conclusion: The reviewed prediction approaches are presented as addressing weak generality and flexibility in existing methods.

4.3. Applications on transfer diagnosis

Transfer diagnosis addresses distribution differences between source and target machine conditions, which can degrade conventional models while making new labeled data costly or unavailable. The review covers parameter transfer, discrepancy measures, and adversarial domain adaptation for CNN-based diagnosis.

  • Motivation: Industrial changes in wear, operating conditions, environment, and human interference create source–target distribution differences.
  • Motivation: When source and target distributions differ, conventional models can degrade, while retraining requires many labeled target instances that may be costly or difficult to obtain.
  • Transfer strategies: The review introduces parameter transfer, moment matching strategies, and adversarial domain adaptation as common transfer-diagnosis techniques.
  • Parameter transfer: Parameter transfer fixes source-domain network parameters and fine-tunes task-specific layers using limited labeled target data.
  • Parameter-transfer applications: Applications transfer CNN features from ImageNet, annotated normal datasets, or other source tasks to bearing, gas-turbine, rotary-machinery, and unseen-condition diagnosis.
  • Parameter-transfer limitation: Parameter-transfer methods are limited because labeled data in the target domain remain necessary, creating obstacles when such labels are unavailable.
  • Discrepancy measures: Discrepancy-measure methods reduce domain differences by minimizing distances such as Maximum Mean Discrepancy or aligning second-order statistics.
  • Discrepancy-measure applications: Reviewed discrepancy-based applications combine MMD with feature clustering or multilayer adaptation for diagnosis under varying or noisy conditions.

5. Conclusions

The review finds that CNFD research is concentrated on fault classification, while health prognostics remain comparatively underrepresented. It identifies gaps in realistic-data validation, parameter selection, signal processing choices, and review coverage, and proposes directions for future work.

  • Scope: The review is incomplete because some papers are missing and some non-English publications were excluded due to language-proficiency limitations.Its observations and conclusions are therefore based on the literature covered in the review.
  • Research distribution: More than 65% of publications focus on fault classification, while mechanical health prognostic applications account for about 17.6%.The review attributes this imbalance partly to the greater ease and intuitiveness of classification compared with prognostics.
  • Research distribution: Health prediction, including degradation monitoring and RUL prediction, should receive more attention because action can precede final machine failure.The paper contrasts prognostic intervention during degradation with waiting for final fault conditions.
  • Generalization and validation: Most models are trained and tested experimentally or in simulation, so differences between experimental and industrial data limit direct real-world application.Some studies even use training and test data from the same experiment, allowing data similarity to produce excellent results.
  • Generalization and validation: Using reasonable experimental data together with diverse realistic industrial data is significant for training more powerful models.The review links broader data coverage to addressing the gap between experimental and industrial settings.
  • Model design: Network architectures and hyper-parameters are mainly selected subjectively, and no specific standard for appropriate parameter selection has formed.The paper identifies relationships between parameters and mechanical signal characteristics as a promising research direction.
  • Signal representation: Signal transformation and processing can increase framework complexity and reduce efficiency, whereas raw-signal models avoid domain-knowledge requirements but remain vulnerable to noise and interference.The review recommends objectively considering both approaches and organically integrating them for better performance.

6. Prospects

The review identifies unresolved challenges for CNFD involving interpretability, unseen faults, early damage, non-stationary operation, real-time diagnosis, fleet deployment, and industrial big data. It also highlights the need for stronger theoretical understanding, generalization, and robust use of diverse data.

  • Theoretical investigation: CNFD remains difficult to interpret because relations between network weights, mechanical features, and learned features are rarely theoretically explained.The black-box character may cause companies or factories to doubt these methods and reject realistic deployment.
  • Unseen damage and faults: Existing models generally recognize only fault categories represented in training data, making unseen-damage identification an open question.Building an all-encompassing training set is expensive or even impossible, while unfamiliar faults can arise as equipment and working environments change.
  • Early damage detection: Early degradation often lacks obvious fault types, motivating quantitative detection and analysis of weak damage before large or significant failures occur.The review notes that waiting for major failures is unreasonable in high-precision and vital industrial applications.
  • Variable operating conditions: Reliable diagnosis from directly processed non-stationary data remains urgent because prior studies commonly rely on smooth operation, stationary conditions, or expert-derived time-invariant features.Collecting stationary data is difficult or impossible in realistic continuously non-stationary environments.
  • Real-time diagnosis: CNFD training is slower than classical shallow algorithms, creating a need for acceleration to satisfy real-time diagnostic requirements.The review calls for novel technologies and techniques to accelerate and improve CNFD algorithms.
  • Fleet diagnosis and industrial big data: Future CNFD models should generalize beyond single-machine diagnosis to equipment fleets and exploit industrial big data, whose diversity and heterogeneity may support more robust models.Existing methods mostly target single machines, while data quantity is often constrained by subjective choices or experimental conditions.
Loading 2002.07605v1…