Source-linked AI summary
Prediction of remaining life of power transformers based on left truncated and right censored lifetime data
Yili Hong, William Q. Meeker, James D. McCalley
TL;DR
The paper develops reliability predictions for transformer fleets whose long lifetimes and incomplete records create left-truncated and right-censored data. It combines parametric lifetime modeling, stratification, and age-adjusted remaining-life distributions to produce prediction intervals for individual transformers and cumulative fleet failures. The results support maintenance prioritization and capital planning, although individual-transformer intervals are often wide.
Problem
Transformer owners need statistically based remaining-life and fleet-failure predictions for maintenance and capital planning despite long lifetimes, evolving designs, left truncation, and right censoring.
Method
The paper uses stratified parametric lifetime models and age-adjusted remaining-life distributions, with random weighted bootstrap and refined central-limit-theorem approximations for prediction intervals.
Results
The predictions provide useful rankings for maintenance priorities and special monitoring, while cumulative-failure prediction intervals support capital planning.
Takeaways & Limitations
The procedure provides a generic reliability-prediction framework for complicatedly truncated and censored data, with applications beyond transformer fleets.
Takeaways & Limitations
Individual-transformer prediction intervals tend to be wide because usage and environmental histories were unavailable for more informative modeling.
Abstract
from arXiv · showhide
Prediction of the remaining life of high-voltage power transformers is an important issue for energy companies because of the need for planning maintenance and capital expenditures. Lifetime data for such transformers are complicated because transformer lifetimes can extend over many decades and transformer designs and manufacturing practices have evolved. We were asked to develop statistically-based predictions for the lifetimes of an energy company's fleet of high-voltage transmission and distribution transformers. The company's data records begin in 1980, providing information on installation and failure dates of transformers. Although the dataset contains many units that were installed before 1980, there is no information about units that were installed and failed before 1980. Thus, the data are left truncated and right censored. We use a parametric lifetime model to describe the lifetime distribution of individual transformers. We develop a statistical procedure, based on age-adjusted life distributions, for computing a prediction interval for remaining life for individual transformers now in service. We then extend these ideas to provide predictions and prediction intervals for the cumulative number of failures, over a range of time, for the overall fleet of transformers.
1. Introduction.
The paper addresses reliability prediction for transformer fleets whose long lifetimes and incomplete historical records create truncated and censored data. It develops a stratified statistical workflow for individual remaining-life and fleet-failure predictions with calibrated prediction intervals.
- Transformer lifetime prediction matters for maintenance and capital planning because unexpected high-voltage transformer failures can cause large economic losses.
- The company’s records produce left-truncated and right-censored data because pre-1980 installations that failed before 1980 are unobserved, while surviving units remain censored.
- The procedure stratifies transformers into relatively homogeneous groups using manufacturer and installation date, then estimates group-specific lifetime parameters by maximum likelihood.
- Age-adjusted remaining-life distributions for at-risk transformers support predictions of individual remaining life and expected failures across future time intervals.
- Prediction intervals account for statistical uncertainty, while sensitivity analysis perturbs stratification and lifetime-distribution assumptions.
2. The transformer lifetime data.
The dataset contains 710 transformers with failures, censoring, and truncation, and reflects heterogeneous operating histories and failure mechanisms. Early failures are treated separately because they appear inconsistent with the predominant aging-related failure mode.
- The dataset contains 710 observations with 62 failures, alongside censored and truncated units summarized by manufacturer.
- Transformer failures generally arise when chemically degraded paper insulation loses dielectric strength under voltage stress, with degradation driven primarily by operating temperature.
- 2.2. Early failures.: Seven units failed within five years, and these early failures were believed to reflect a defect-related mechanism different from later failures.
- 2.2. Early failures.: The analysis treated early failures as right censored because including them implied an approximately constant hazard inconsistent with the known predominant aging failure mode.
- 2.3. Explanatory variables.: Insulation type and cooling class were studied as potentially important explanatory variables, while Figure 1 displays service-time events for a systematic data subset.
3. Statistical lifetime model for left truncated and right censored data.
The paper models transformer lifetimes with parametric log-location-scale distributions and constructs a likelihood that accounts for left truncation and right censoring. Parameters can vary by transformer or stratified group and are estimated by maximum likelihood.
- Transformer lifetimes are modeled with a log-location-scale distribution, including Weibull and lognormal families as common choices.
- The Weibull shape parameter determines hazard behavior: β > 1 indicates increasing wearout hazard, β = 1 constant hazard, and β < 1 decreasing hazard.
- Right censoring represents unfailed transformers still in service at the March 2008 data-freeze point.
- Transformers installed before 1980 are modeled with left-truncated lifetime distributions because units installed and failed before 1980 are unobserved.
- The likelihood uses lifetime, truncation, and censoring information, with indicators distinguishing truncated from untruncated and failed from censored transformers.
- Maximum-likelihood estimates are obtained by maximizing the likelihood, with parameter structure determined by the modeling context, including stratified groups or regression predictors.
4. Stratification and regression analysis.
The analysis stratifies transformers into relatively homogeneous groups, compares Weibull and lognormal lifetime models under truncation, and develops separate regression models for Old and New groups.
- 4.1. Stratification: The data were stratified by manufacturer and installation year because old transformers were overengineered relative to newer designs.The final partition used 1987 as the cutting year, with sparse groups combined where necessary.
- 4.1. Stratification: Figure 2 compares Turnbull nonparametric estimates with Weibull maximum-likelihood c.d.f. estimates for each group.The plotted nonparametric points correspond to observed lifetimes and midpoints of Turnbull c.d.f. steps, while censored units were not plotted.
- 4.1. Stratification: The parametric and nonparametric estimates disagree for Old groups because the Old data are heavily truncated, making the nonparametric estimator inconsistent in this setting.The maximum-likelihood estimator based on the truncated-data likelihood remains consistent.
- 4.1. Stratification: The fitted Old and New groups have slowly and rapidly increasing hazard rates, with estimated common Weibull shape parameters near 2 and 5, respectively.The common-shape assumption for each group is supported by the lifetime data and likelihood-ratio tests.
- 4.2. Distribution choice: Weibull distributions generally fit better than lognormal distributions, both visually and by loglikelihood, consistent with lifetimes governed by a minimum over multiple failure locations.The Weibull distribution is a limiting distribution of minima.
- 4.3. A problem with the MD group data: The MD group’s estimated Weibull shape parameter is 0.51, implying a decreasing hazard inconsistent with insulation aging; MD units were excluded from estimation but retained for prediction.The estimation problem is attributed to extremely heavy truncation, and engineering knowledge was used when assigning affected units for prediction.
- 4.4. Regression analysis: The regression models use categorical Manufacturer, Insulation, and Cooling variables, with Cooling selected for the Old group and Manufacturer for the New group.Likelihood-ratio tests found Manufacturer and Insulation unimportant for Old transformers, while Manufacturer was statistically important for New transformers.
5. Predictions for the remaining life of individual transformers.
The paper develops calibrated prediction intervals for the remaining life of individual transformers, conditioning on survival to each unit’s current age. Weibull-based results identify similar intervals among comparable younger units and short intervals for units assessed as near-term high-risk.
- 5. Predictions for the remaining life of individual transformers.: Calibrated prediction intervals quantify individual transformers’ future failure times conditional on survival to their present ages.The procedure uses age-conditioned lifetime distributions and accounts for parameter uncertainty through calibration.
- 5.1. The naive prediction interval procedure.: The naive plug-in interval ignores uncertainty in the estimated parameters and generally has coverage below its nominal confidence level.Calibration adjusts the procedure so its actual coverage is closer to the desired nominal level.
- 5.3. The random weighted bootstrap.: The random weighted bootstrap addresses sparse failures and heavy censoring that make traditional and parametric bootstrap approaches difficult to implement.Only about 9% of transformers had failed, and truncation-time and censoring-time distributions were unavailable.
- 5.3. The random weighted bootstrap.: Random weighted bootstrap intervals were insensitive to the tested weight distributions in this application.The study used Gamma(1,1) weights and also tested Gamma(1,0.5), Gamma(1,2), and Beta(2−1,1) alternatives.
- 5.5. Prediction results.: 90% Weibull prediction intervals are presented for a subset of at-risk transformers using a 1987 stratification cutoff.The years axis in the corresponding figure is logarithmic.
- 5.5. Prediction results.: Comparable younger transformers with the same explanatory values have similar intervals, while older units can have short intervals with lower endpoints near current age.The model therefore identifies some units as being at especially high risk of failure in the near term.
6. Prediction for the cumulative number of failures for the population.
The paper predicts cumulative transformer failures over the next 10 years and constructs calibrated pointwise prediction intervals using a Weibull regression model. It applies a skewness-corrected approximation and bootstrap procedure, reporting results for Old, New, combined, and manufacturer-specific groups.
- The procedure predicts monthly cumulative failures for the next 10 years with calibrated pointwise intervals that quantify statistical uncertainty and failure-process variability.The predictions support capital-expenditure planning for the transformer population.
- Cumulative future failures are modeled as a sum of independent, nonidentically distributed Bernoulli variables, with each transformer’s failure probability determined by its age-adjusted lifetime distribution.Different entry dates produce different transformer ages and failure probabilities.
- The estimated failure-count distribution uses Volkova’s refined central-limit approximation, which corrects for skewness rather than relying only on Poisson or ordinary normal approximations.Monte Carlo simulation is more computationally intensive when many nonidentically distributed components are present.
- Calibrated prediction intervals are obtained by repeatedly simulating Bernoulli outcomes, transforming them through the estimated failure-count distribution, and solving for the corresponding quantiles.Bootstrap resampling accounts for uncertainty in the fitted model parameters and estimated individual failure probabilities.
- The results use a Weibull distribution regression model stratified at 1987, with 449 Old and 199 New units analyzed separately and 648 units combined.Figures report 90% and 95% pointwise prediction intervals for the separate groups, the combined population, and selected manufacturers.
7. Sensitivity analysis and check for consistency.
Sensitivity analyses show that predictions are more affected by lifetime-distribution and cutting-year assumptions for newer transformers, while the model’s consistency was checked against nonparametric estimates.
- 7.1. Sensitivity analysis.: New-group predictions are somewhat sensitive to the lifetime-distribution assumption, whereas Old-group predictions are not highly sensitive.The difference is partly attributed to greater extrapolation for the New group over the next 10 years.
- 7.1. Sensitivity analysis.: Lognormal predictions are more optimistic than Weibull predictions when extrapolating future failures.
- 7.1. Sensitivity analysis.: Cutting-year changes have little effect in the Old group, but New-group predictions are more sensitive to this choice.
- 7.1. Sensitivity analysis.: Using 1990 as the cutting year widens prediction intervals as time increases because only one failure occurs in the relevant New subgroup.The resulting random weighted bootstrap samples have greater variability than those from other cutting years.
- 7.1. Sensitivity analysis.: The analysis uses 1987 as the cutting year because it lies on the pessimistic side of the New-group sensitivity results.
- 7.2. Check for consistency.: The model was checked by comparing parametric predictions with Turnbull nonparametric estimates, although insufficient data prevented holding out recent failures.
8. Discussion and areas for future research.
The paper presents a broadly applicable procedure for reliability prediction with truncated and censored data, while identifying practical uses and important limits of individual predictions and model assumptions.
- 8. Discussion and areas for future research.: The developed prediction procedure applies to reliability problems involving complicated censoring and truncation.The authors identify broader applications, including field reliability prediction for warranty data.
- 8. Discussion and areas for future research.: Population-level failure prediction intervals are useful for capital planning, while individual intervals are often too wide to determine replacement timing directly.
- 8. Discussion and areas for future research.: The analysis found that transformers from some manufacturers, including MA, tend to have shorter lives and should receive particular attention.Individual predictions can still rank maintenance priorities and identify transformers for special monitoring or more frequent inspection.
- 8. Discussion and areas for future research.: Individual prediction intervals could be narrowed with usage or environmental information, although suitable procedures would require further development.Examples include load, ambient-temperature history, and voltage spikes.
- 8. Discussion and areas for future research.: Bayesian approaches could use engineering knowledge about distribution shape or regression coefficients to narrow prediction intervals under alternative modeling assumptions.
- 8. Discussion and areas for future research.: Limited failure-cause information motivates separate analysis of failure modes when modes differ in behavior, cost, or engineering relevance.Such extensions must address dependence among failure-mode lifetimes in field data.
- 8. Discussion and areas for future research.: The remaining-life framework may also apply to aging aircraft, consumer products, and System Health Management, although modeling tools vary by application.
SUPPLEMENTARY MATERIAL
The supplement describes difficulties fitting the MD-group failure model because that group contains substantial truncation.
- SUPPLEMENTARY MATERIAL: The supplement documents model-fitting difficulties for the MD group caused by a large amount of truncation.