Source-linked AI summary
Long Horizon Transformer Quantile Fault Prediction for Multi Site Industrial Predictive Maintenance
David J Poland, Daniele Ravi, Na Helian
TL;DR
Long-horizon predictive maintenance needs models that distinguish gradual degradation from normal regime variation over day-scale planning windows. The paper proposes TQRNN30d, combining dual-stage conditional-quantile features with multi-stream temporal fusion for supervised fault-occurrence forecasting. It leads the evaluated baselines across 7-, 14-, and 30-day horizons, while evidence remains limited to held-out machines within a homogeneous nine-facility fleet.
Problem
Long-horizon PdM is less represented than short-horizon or asset-specific work, while gradual distributional change, heterogeneous sensing, regime transitions, and leakage complicate weeks-ahead forecasting.
Method
TQRNN30d combines dual-stage QRNN feature extraction with a multi-stream temporal-fusion classifier using quantile states, covariates, metadata, latent history, and heuristic instability-aware memory modulation.
Results
TQRNN30d led all 18 evaluated baselines at 7, 14, and 30 days; at 30 days it achieved 79.97% F1, 80.18% recall, 81.82% precision, 82.39% accuracy, and 0.820 ROC-AUC.
Takeaways & Limitations
The results support held-out-machine long-horizon fault forecasting within the observed homogeneous multi-facility industrial fleet.
Takeaways & Limitations
The evaluation does not establish transfer to unseen facilities, different machine types, sensing architectures, products, or industrial sectors, and the proprietary dataset limits independent reproduction.
Abstract
from arXiv · showhide
Long-horizon predictive maintenance requires models to distinguish slowly evolving degradation from normal operating-regime variation over planning windows measured in days rather than hours. This paper evaluates whether an explicit conditional-quantile representation provides an informative classifier interface for this problem. The proposed TQRNN30d framework combines a dual-stage quantile regression neural network (QRNN) feature extractor with a multi-stream temporal fusion classifier. Each hourly word of 81-channel machine behaviour is mapped to a 324-dimensional quantile-state representation, and 720 ordered hourly words form the 30-day document supplied to the long-horizon model. The classifier fuses quantile states with dynamic covariates, channel-level static metadata, and a 168-hour latent-history stream using gated residual processing, causal recurrent encoding, and metadata-conditioned cross-modal attention. A bounded instability-aware signal derived from sustained one-word-ahead prediction-error divergence provides auxiliary memory modulation at the longest horizon. Evaluation uses a machine-disjoint 43/14/15 train/validation/test allocation across 72 machines in nine manufacturing facilities. At 30 days, TQRNN30d achieves 79.97% F1, 80.18% recall, 81.82% precision, 82.39% accuracy, and 0.820 ROC-AUC. It leads all 18 evaluated baselines at the 7-, 14-, and 30-day fixed-threshold comparisons, with the largest F1 advantage at 14 days. The results support held-out-machine performance within the observed homogeneous nine-facility fleet, but do not establish unseen-site, cross-equipment, or cross-sector generalisation.
I. INTRODUCTION
Long-horizon industrial predictive maintenance addresses gradual distributional degradation, heterogeneous sensing, regime transitions, and evaluation leakage that complicate weeks-ahead forecasting. TQRNN30d treats the input representation as central, using conditional quantiles for supervised 7-, 14-, and 30-day fault-occurrence forecasting.
- Why long horizons, and why multiple sites: 30-day fleet planning motivates forecasting fault occurrence weeks ahead rather than issuing only hour-ahead alarms.The paper links long horizons to maintenance scheduling across geographically separated plants.
- Why long horizons, and why multiple sites: Long horizons intensify gradual distributional change, causal multi-rate sensor fusion, non-stationary regimes, and the limits of ordinary sequential classifiers.The cited challenges include variance growth, tail movement, envelope widening, incompatible sampling rates, regime transitions, and missing distributional information.
- Scope of the task: TQRNN30d forecasts whether at least one confirmed fault occurs within 7-, 14-, or 30-day future horizons from a completed observation document.The task is supervised Normal/Abnormal forecasting rather than direct RUL or degradation-index regression.
- Contributions: The framework combines an 81-channel quantile representation with a multi-stream Temporal Fusion Transformer and an auxiliary instability-aware memory pathway.The contributions distinguish the conditional-quantile interface, dual-stage representation, heuristic instability modulation, and machine-disjoint evaluation.
- Related work and gap: The literature gap is a long-horizon method centered on input representation and validated with partitioning strict enough to withstand leakage concerns.Prior work often uses short horizons, asset-specific settings, or benchmark datasets, while reported results are difficult to compare across protocols.
- A. Why a Quantile Encoding Rather Than Point Summaries: Quantile regression preserves distributional widening, asymmetric tails, and early degradation envelopes rather than reducing sensor behavior to point forecasts.Here quantile regression serves as a classifier interface, not as the final forecasting target.
D. Temporal Fusion and Attention Models
The paper distinguishes present-state anomaly detection from supervised prospective fault forecasting. Its temporal-fusion design uses sequential and contextual information while treating sustained error divergence as a heuristic instability cue.
- Temporal Fusion Transformer: TFT contributes gated residual processing, static covariate conditioning, temporal variable selection, and attention for multi-horizon sequence modelling.These properties are presented as suitable for machine evolution together with contextual operating factors.
- Task distinction: Fault forecasting asks whether a labelled abnormal event will occur within a future window, unlike anomaly detection, which evaluates present deviation.The forecasting task is supervised and evaluated on lead time as well as separability.
- Prospective prediction: TQRNN30d uses evolving distributional structure to predict future abnormal behaviour rather than merely detecting current deviation.The framework is prospective anomaly prediction within a future forecasting horizon.
- Comparative baselines: Autoencoder and GAN anomaly-oriented models are adapted to the same future fault-occurrence target for like-for-like baseline comparison.The comparison tests whether conventional anomaly-oriented representations remain competitive after shifting from present detection to long-horizon prediction.
- Instability Indicators: The instability-aware gate is explicitly heuristic and does not estimate a formal dynamical invariant such as a Lyapunov exponent.Classical diagnostics require reconstruction and stationary records unavailable in the described causal deployment.
- Deployment context: The deployment contains 72 machines across nine facilities within one organisation and a homogeneous production-equipment fleet.The sensor and machine context is multi-facility but not presented as a cross-organisation or cross-equipment setting.
2) Operational sensor-ablation context:
The operational sensor-ablation evidence indicates redundancy and resilience to isolated sensor loss, while the temporal interface converts multi-rate signals into ordered hourly quantile-state documents.
- Operational sensor-ablation context:: Approximately 94% or higher detection rates were retained for nonthermal channels under single-sensor removal.The thermal group was less central to the mechanically coupled fault-discriminative pathway.
- Operational sensor-ablation context:: Cross-group fusion is supported as resilient to isolated sensor loss, but the ablation is not an unseen-site generalisation test.It is distinct from the 15-machine held-out cohort used for formal classifier comparison.
- Temporal alignment: Causal zero-order hold projects asynchronous streams onto a fixed 20 ms reference grid without introducing future measurements.Channels retain each observed value until the next native sample becomes available.
- Word and document representation: Each 30-day document contains 720 ordered hourly words represented to the classifier as 324-dimensional quantile-state vectors.The classifier receives word-level vectors rather than the underlying 50 Hz sequence.
- Latent history: A 168-hour latent-history stream gives the classifier recent operating memory without reprocessing the complete raw trace.Documents refresh at one-hour intervals while temporal ordering is preserved.
- Quantile feature extraction: The dual-stage extractor refines four retained quantiles per channel, yielding the 324-dimensional word-level representation.The second stage adds interior quantiles around the median while retaining the interquartile envelope.
D. Label Construction
The study defines future fault labels from confirmed operational events within strictly future horizons while enforcing causal, machine-disjoint evaluation. Ground truth uses multiple operational sources, with explicit exclusions and scope boundaries for the long-horizon deployment.
- Confirmed fault events require support from at least two of operator logs, PLC fault codes, and maintenance records, with expert review for ambiguous cases.
- A window is positive when at least one confirmed fault occurs in (t0, t0 + H], using only information available before t0.
- Windows intersecting missing-data gaps longer than 300 s are excluded, while planned maintenance is treated as known context when available at prediction time.
- Machine-disjoint partitioning assigns every machine exclusively to training, validation, or testing, preventing shared measurements and overlapping documents.
- The allocation uses 43 machines for training, 14 for validation, and 15 for testing across the nine observed facilities.
- Long-horizon preprocessing and model-selection quantities are estimated from training/development data and fixed for held-out evaluation.
B. Dual-Stage Quantile Feature Extraction
The dual-stage extractor converts each hourly 81-channel word into a distribution-aware quantile state, which the classifier combines with contextual and historical streams. The design targets distributional changes while preserving causal long-horizon sequence processing.
- Dual-stage extraction: Each hourly word is transformed into a compact mid-tail quantile representation through two cascaded quantile-regression stages.
- Dual-stage extraction: QRNN1 estimates ten broad conditional quantiles, MLP2 compresses them, and QRNN2 refines the retained levels {0.25, 0.40, 0.60, 0.75}.
- Representation: The resulting sequence contains 720 hourly representations of 324 dimensions, corresponding to four retained quantiles for each of 81 channels.
- Multi-stream fusion: The classifier fuses quantile states with dynamic covariates, static metadata, and a 168-hour latent-history stream using temporal-fusion components.
- Multi-stream fusion: Causal recurrent encoding processes the document in temporal order, while metadata-conditioned cross-modal attention uses channel-indexed metadata keys and values.
- Prediction head: The output preserves structured hourly forecast-interval probabilities for 7-, 14-, and 30-day horizons before thresholding headline decisions.
D. Instability-Aware Memory Gate
The instability-aware memory gate derives a bounded auxiliary signal from sustained one-word-ahead prediction-error divergence and uses it to modulate causal temporal memory. It supplements, rather than replaces, the retained QRNN representation and final classifier.
- Signal construction: The pathway uses retained channel-level quantile states to predict each auxiliary channel summary one hourly word ahead.Predictions use only information strictly before word k.
- Persistence criterion: The gate identifies sustained local divergence by requiring three consecutive threshold exceedances before activation.The evaluated threshold is λ0 = 0.2.
- Signal aggregation: The aggregate divergence signal is restricted to degradation-sensitive channels, weighted by learned non-negative channel weights, and bounded before temporal gating.These steps limit the influence of contextual or instrumentation-related variation and constrain the modulation magnitude.
- Memory modulation: The bounded signal modulates the causal GRU update gate so persistent divergence biases memory toward retaining more of the previous hidden state.The classifier then combines the quantile, dynamic-covariate, and historical-latent streams through metadata-conditioned fusion.
- Role in the classifier: The instability pathway is auxiliary: the retained QRNN representation remains the primary predictive basis, and its independent contribution was not isolated by a dedicated ablation.The final head emits structured hourly forecast-interval probabilities before the horizon-level decision.
1) Interpretation and limitations of instability-aware gating:
Instability-aware gating detects sustained divergence in prediction-error trajectories but does not identify the physical cause. Its influence is constrained by channel selection, persistence, bounded modulation, and development-data calibration, yet sustained regime changes may still activate it.
- Interpretation: The gate may respond to progressive degradation, sensor drift, material transitions, or sustained operating-regime changes.It should therefore be interpreted as temporal modulation rather than an independent fault detector.
- Mitigations: Restricting aggregation to degradation-sensitive channels and bounding the modulation reduces sensitivity to non-degradation disturbances.The persistence criterion also suppresses isolated transient excursions.
- Residual limitation: Sustained non-fault regime changes may still activate the pathway despite the persistence requirement.This is especially relevant during material, lubrication, or operating-condition transitions without associated mechanical failure.
- Configuration: The implemented gate uses λ0 = 0.2, three consecutive hourly exceedances, ϵ = 10^-3, and retained quantile context α = 0.25, 0.40, 0.60, 0.75.These are implementation parameters rather than physical constants.
- Transferability: Threshold and persistence settings were selected on development data and are not claimed to transfer without recalibration to another machine population.The held-out test partition was excluded from parameter selection.
E. Training Objective
The training objective first constructs the upstream quantile representation, then trains the long-horizon classifier as a document-level multi-output binary predictor. Evaluation compares this framework with 18 complementary model families under common task and data foundations.
- Training objective: The two QRNN stages use pinball quantile loss, after which the TFT classifier is trained on the QRNN representation.The classifier predicts failure probabilities for each forecast interval.
- Classification loss: The classifier uses interval-wise binary crossentropy with clipped predicted probabilities for numerical stability.The target and predicted failure probability are defined for each training document and forecast interval.
- Class imbalance: Class-weighted BCE addresses sparse confirmed-failure intervals, with inverse-prevalence weights computed from training data only.The weights remain fixed during validation and testing.
- Baseline comparison: The comparison spans 18 baselines across kernel, instance-based, tree, recurrent, feed-forward, reconstruction, attention, and pretrained-transformer families.These families test temporal modelling, anomaly-detection, long-range attention, and pretrained sequence-classification alternatives.
- Baseline comparison: Attention baselines include Transformer, Transformer-XL, and TFTfull, while BERT, RoBERTa, DistilBERT, and ALBERT test pretrained sequence-model alternatives.TFTfull is identified as the direct architectural ancestor and most informative single comparator.
2) Adaptation of BERT-family sequence baselines:
The BERT-family rows are retained as architecture-family comparators under the same long-horizon task and held-out population, without unsupported claims about their adapter details. The broader evaluation fixes data-processing and threshold choices through training or validation data, while several transfer and repeated-seed analyses remain unreported.
- BERT-family baselines: BERT, RoBERTa, DistilBERT, and ALBERT are retained sequence-model baselines rather than newly specified token-level implementations.The experimental record does not provide sufficient adapter detail to justify new tokenisation or positional-embedding claims.
- Comparative fairness: All models use machine-disjoint training, validation, and held-out test assignments with preprocessing estimated from applicable development data only.This enforces a common data and evaluation foundation.
- Comparative fairness: Hyperparameters, stopping criteria, imbalance treatment, and decision thresholds are selected using training or validation data before held-out evaluation.The study does not claim identical search spaces or equal numbers of trials across model families.
- Scope of evidence: The completed programme excludes rolling-origin evaluation, leave-one-site-out transfer, additional long-sequence baselines, matched one-component ablations, and systematic multi-seed retraining.Headline results therefore come from one completed training and held-out evaluation run per configuration.
- TFT input regimes: TFTfull receives the complete MLP–QRNN quantile-state representation, whereas TFTflat removes that hierarchy and uses raw flat input for the ablation.They are the same network under different input regimes.
- Evaluation metrics: Reported performance uses thresholded F1-score, recall, precision, and accuracy, alongside ROC–AUC and a multi-site precision–recall profile area.The thresholded metrics are unweighted means across the corresponding held-out cohort.
D. Modelling Assumptions
The study states assumptions about labels, sensors, and deployment stability, while reporting performance under fixed evaluation and decision-threshold procedures. Across evaluated horizons, TQRNN30d outperforms RoBERTa, with the largest F1 difference at 14 days.
- The analysis treats the reported assumptions and limitations as conditions on interpreting the results.
- TQRNN30d outperforms RoBERTa at 7, 14, and 30 days, with the largest F1 difference of 6.27 percentage points at 14 days.
- 79.97% F1, 80.18% recall, 81.82% precision, and 82.39% accuracy were achieved at 30 days.
B. Component Ablation
The ablations separate day-scale pipeline evidence from TFT-side progression under flat-input conditions, while discrimination and computational comparisons assess the full model more broadly. The evidence supports gains from the integrated representation and fusion pathway without establishing single-component causal attribution.
- The ablation studies use complementary day-scale and TFT-side experiments rather than a one-component-at-a-time ablation of the complete model.
- 62.34% F1 was reached after adding environmental context, historical look-back, known-ahead covariates, and full TFT-side fusion to the flat-input setting.
- 0.820 ROC–AUC was achieved at 30 days, compared with 0.722 for RoBERTa and 0.717 for Transformer-XL.
- TQRNN30d achieved the largest mean multi-site precision–recall profile area, 0.638, ahead of RoBERTa at 0.581 and Transformer-XL at 0.552.
- TQRNN30d uses 11.4M parameters and 18 ms inference latency, remaining smaller than language-model baselines while exceeding their reported accuracy.
E. Operational Deployment Context
Deployment occurred across nine facilities within a broader maintenance programme, but the operational changes cannot be attributed independently to the predictive framework. The evaluation supports held-out-machine performance within a homogeneous fleet, not broader transfer or causal deployment claims.
- Production-output efficiency increased from 78.38% to 90.23% during the 24-month programme, alongside reported reductions in maintenance disruption and scrap-related losses.
- The observed operational and financial changes cannot be attributed independently to the framework because no contemporaneous untreated control facilities were available.
- The evaluation covers 72 machines across nine facilities from one homogeneous asset and product family with a common sensor configuration.
- The study does not establish transfer to unseen facilities, different machine types, sensing architectures, products, or industrial sectors.
- The instability-aware pathway may respond to sensor drift, material transitions, or operating-regime changes and requires recalibration for different machines or sites.
- The proprietary dataset limits independent reproduction and does not establish performance under substantially different conditions or longer periods of concept drift.