Source-linked AI summary

Semi-Supervised Classification with Informative Missing Labels in Weibull Mixture Models

Jinran Wu, You-Gan Wang, Geoffrey J. McLachlan

arXiv:2609.00774v1stat.MLcs.LG

TL;DR

Semi-supervised classification in Weibull mixtures is complicated by labels missing according to feature-dependent uncertainty, which may itself contain classifier information. The paper models this shared-parameter MAR mechanism, derives boundary-focused information and error results, and reports improved efficiency and boundary estimation in numerical and semi-synthetic studies.

  • Problem

    The paper addresses semi-supervised classification when feature observations are complete but some labels are missing through a feature-dependent mechanism that may inform the classifier.

  • Method

    It models missingness as shared-parameter MAR depending on classification uncertainty, then derives adjusted Fisher information and decision-boundary expansions for plug-in error.

  • Results

    The analyses show potential reductions in expected error and improved estimation of one- and two-boundary classification rules when feature-dependent missingness is modelled.

  • Takeaways & Limitations

    Classification gains depend on uncertainty in parameter directions that perturb decision boundaries, so improved estimation of every mixture parameter is not required.

  • Takeaways & Limitations

    The asymptotic boundary theory excludes double roots, and identifiability requires distinct components plus labelled observations from both classes.

Abstract

from arXiv · show

We consider semi-supervised classification from a partially classified sample arising from a two-component Weibull mixture. The feature is observed for all data, whereas some class labels are missing. The probability of a missing label is modelled as a function of classification uncertainty, giving a feature-dependent missing-at-random (MAR) mechanism that shares parameters with the Weibull-mixture classifier. The missing-label indicators can therefore provide information about the classifier in addition to the observed features and available class labels. Under a common Weibull shape, a Bayes' rule has at most one positive decision boundary, which is unique when the rule is nonconstant; under unequal shapes, it can have two. We characterise these decision regions, derive the Fisher information for the classifier after adjustment for nuisance parameters in the missingness model, and obtain a decision-boundary expansion of the expected error rate of the plug-in sample rule relative to the Bayes error. The expansion yields classification-specific asymptotic relative efficiency formulas for the one- and two-boundary cases and shows that a positive-definite increase in Fisher information is sufficient, but not necessary, for a smaller first-order expected error rate. Numerical studies and a semi-synthetic analysis based on hard-drive failure data illustrate potential reductions in expected error rate and improvements in decision-boundary estimation from modelling feature-dependent label missingness.

1 Introduction

The paper studies semi-supervised classification for two-component Weibull mixtures when label missingness depends on feature-based classification uncertainty. It develops decision-geometry and efficiency results to determine how informative missingness affects boundary estimation and expected classification error.

  • Motivation: Feature-dependent MAR missingness can make missing-label patterns informative about the Weibull-mixture classifier.The missingness probability shares parameters with the classifier, so the label-missingness mechanism is retained in the likelihood.
  • Decision geometry: Common-shape Weibull mixtures have at most one positive Bayes decision boundary, whereas unequal shapes can produce two.Unequal shapes can yield a bounded central decision region for one class and tail regions for the other.
  • Contributions: The paper derives adjusted Fisher information and a decision-boundary expansion for the plug-in rule’s expected error relative to Bayes error.The expansion isolates parameter directions that perturb decision boundaries and supports classification-specific asymptotic relative efficiency formulas.
  • Contributions: A positive-definite increase in Fisher information is sufficient, but not necessary, for reducing the first-order expected classification error.This links information comparisons to classification performance through boundary estimation rather than through overall parameter estimation alone.
  • Contributions: Numerical simulations and a hard-drive-failure semi-synthetic study examine expected error and decision-boundary estimation under informative missingness.The supplied contribution passage identifies both finite-sample evaluation settings and their classification-focused outcomes.

2 A two-component mixture of Weibull distributions for classification

The Weibull-mixture classifier uses posterior odds to define Bayes decisions and has distinct boundary geometries under common and unequal component shapes. Common shapes yield at most one positive boundary, while unequal shapes yield at most two and may create central-versus-tail decision regions.

  • 2.1 Model and Bayes’ rule: The model assigns a positive feature Y to one of two latent classes with prior probabilities π_i and component-specific Weibull shape and scale parameters.The mixture’s marginal feature distribution combines the two Weibull component densities, with θ denoting the mixture-parameter vector.
  • 2.1 Model and Bayes’ rule: Class-1 posterior probability is determined by the log-posterior odds δ(y; θ), with posterior ties occurring at its positive roots.At a boundary, δ(y_b; θ)=0 is equivalent to equal prior-weighted component densities and equal posterior class probabilities.
  • 2.1 Model and Bayes’ rule: The asymptotic boundary analysis restricts attention to simple roots that locally separate regions assigned to different classes.Double roots are treated as nonregular touching points rather than classification boundaries for this theory.
  • 2.1 Model and Bayes’ rule: The constrained parameters are transformed using ρ = log(π1/π2), α_i = log λ_i, and ν_i = log k_i so the transformed parameters range over the real line.These transformations give logit(π1)=ρ, λ_i=e^α_i, and k_i=e^ν_i.
  • 2.2 Common shape: For common shapes, the log-posterior odds are affine in X = Y^k, so a nonconstant Bayes rule has one unique positive boundary when AB < 0.The function is generally nonlinear in Y itself; if no boundary exists, the rule assigns every positive observation to one class.
  • 2.3 Unequal shapes: For unequal shapes, δ(y; θ_U) is nonlinear and has at most two positive roots, producing either no roots, a double root, or two simple roots.When k_1>k_2, class 1 occupies the interval between two roots; when k_1<k_2, class 2 occupies that interval and class 1 occupies both tails.
  • 2.3 Unequal shapes: Representative mixtures illustrate one boundary at y_b = 1.6088 for common shapes and two boundaries at y_1 = 0.6957 and y_2 = 2.9855 for unequal shapes.The figure’s dashed lines mark δ(y; θ)=0, and shaded bands indicate the associated Bayes decision regions.
  • 2.3 Unequal shapes: The unequal-shape asymptotic theory excludes double roots, while identifiability requires distinct components, positive class probabilities, and labelled observations from both classes.Partial labels resolve the mixture’s component-label permutation under these conditions.

3 Fisher information under the label-missingness model

The paper models feature-dependent label missingness through a shared-parameter MAR mechanism and decomposes the resulting information into label loss and missingness gains. It derives adjusted Fisher information for the classifier, including nuisance-parameter corrections and conditions for improved asymptotic covariance.

  • Likelihoods for the partially classified sample: The missingness mechanism is MAR because the missing-label probability depends on the observed feature through the classifier’s posterior-odds function.The model uses logit{q(y; θ, ξ)} = ξ0 + ξ1δ(y; θ)^2, with ξ1 < 0 making labels most likely to be missing near Bayes decision boundaries.
  • Likelihoods for the partially classified sample: The full likelihood combines the likelihood that ignores missingness with a Bernoulli likelihood for the missing-label indicators.The partially classified data contribute observed features and available labels, while the missingness indicators contribute through q(Yj; θ, ξ).
  • Information loss and gain: Unobserved labels contribute a per-observation information loss, while feature-dependent missingness can provide additional information about classifier parameters.The ignoring estimator’s information depends on ξ through the distribution of M, even though its criterion does not contain ξ explicitly.
  • Information loss and gain: When missingness parameters are estimated jointly, the useful missingness information about θ is the Schur complement after adjustment for nuisance parameters ξ.The resulting matrix is positive semidefinite and represents the missingness information about θ remaining after nuisance adjustment.
  • Information loss and gain: A positive-definite information gain under the stated regularity conditions yields a smaller asymptotic covariance matrix for the full-likelihood estimator than for the completely classified estimator.The comparison requires both full-likelihood and completely classified information matrices to be positive definite.
  • Computation: All required expectations, including Bayes error and asymptotic relative efficiency, can be evaluated using one-dimensional numerical integration.The information quantities require one-dimensional integration because posterior probabilities and missingness weights involve both mixture components.

4 Expected error rate of the plug-in sample rule and asymptotic relative efficiency

The section expands plug-in classification error around Bayes error through the estimated decision boundaries, then compares estimators using asymptotic relative efficiency. The comparison distinguishes one-boundary common-shape mixtures from two-boundary unequal-shape mixtures and shows that information gains need only affect boundary-relevant directions.

  • Expected error and expansion: Expected excess error is compared through E{R(bθs; θ)} − err(θ), with its leading coefficient determining asymptotic relative efficiency.The expansion has the form E{R(bθs; θ)}−err(θ) = Cs(θ)/n+o(n−1).
  • Decision-boundary cases: Under stable simple boundaries, the common-shape model has one boundary when AB < 0, whereas the unequal-shape model can have two locally stable boundaries.Tangent roots, root mergers, and constant Bayes rules are excluded from the expansion.
  • Expected error and expansion: The local error expansion represents boundary-estimation effects using the boundary density, local log-posterior-odds steepness, and covariance of estimated boundary scores.The weight pY(yj)/|δy(yj)| captures probability mass relative to local steepness, while the quadratic form captures uncertainty in estimated log-posterior odds.
  • Asymptotic relative efficiency: For the full estimator, the common-shape asymptotic relative efficiency compares one-boundary error coefficients after the common boundary factor cancels.The information matrices must come from the constrained common-shape likelihood rather than unequal-shape information followed by parameter substitution.
  • Asymptotic relative efficiency: A positive-definite information increase is sufficient but not necessary for AREs:CC > 1, because first-order error depends only on parameter directions that perturb Bayes boundaries.In the two-boundary case, the boundary weights generally do not cancel.

5 Numerical simulations

Simulations compare complete-case, missingness-ignoring, and full-likelihood estimators across common- and unequal-shape Weibull mixtures. Modeling feature-dependent label missingness improves relative efficiency, parameter accuracy, and decision-boundary estimation, especially with stronger MAR dependence and more missing labels.

  • Simulation setup: The study evaluates six Weibull-mixture designs with n = 500 and 500 Monte Carlo replications, spanning common-shape one-boundary and unequal-shape two-boundary settings.Labels were removed under mechanisms producing 10%, 30%, and 50% missingness across 72 simulation settings.
  • Simulation setup: The full estimator jointly models the classifier and feature-dependent missingness, whereas ig ignores missingness and CC uses completely classified data.Plug-in error rates were evaluated directly under each data-generating distribution.
  • Expected error-rate efficiency: When ξ1 = 0, full and ig perform similarly; stronger MAR dependence produces progressively larger gains for full over ig.The missingness probability concentrates labels near Bayes decision boundaries as ξ1 becomes more negative.
  • Expected error-rate efficiency: Under common-shape designs, full reaches empirical relative efficiencies above six in several strongly informative settings, while unequal-shape gains are generally more moderate.The ig estimator remains below one in nearly all settings and worsens as the missing-label proportion increases.
  • Decision-boundary estimation: Full generally lowers parameter RMSE and estimated-boundary RMSE relative to ig, with larger improvements under stronger MAR dependence and higher missingness.In unequal-shape models, full reduces RMSE for both boundaries, particularly the second boundary under strong MAR missingness.
  • Decision-boundary estimation: The results indicate that modeling missingness improves estimation of the classification rule, not merely estimation of mixture parameters.The full estimator often approaches or improves upon CC parameter RMSE when missingness concentrates near decision boundaries.

6 A semi-synthetic study using hard-drive failure data

A semi-synthetic hard-drive study emulates expert labeling by retaining easier labels and withholding ambiguous ones. The full likelihood produces boundaries closer to the complete-data benchmark and the lowest cross-validated error at every expert level.

  • Data and missingness construction: The analysis classifies failed drives from two hard-drive product models using positive-valued lifetimes as features and drive type as the class.The target is drive-type classification among failed drives, not prediction of future failures.
  • Data and missingness construction: Three nested regimes retain approximately 60%, 70%, and 80% of labels, with observations near the decision boundary more likely to remain unlabelled.These regimes represent stylised junior, intermediate, and senior expert-assessment levels.
  • Model comparison: BIC favoured the common-shape model, with values 2845.142 versus 2847.951, and the likelihood-ratio test did not reject it.The study therefore used the common-shape model with the MAR mechanism based on squared log-posterior odds.
  • Boundary estimation: The full estimates moved the decision boundary closer to the completely classified estimate, whereas ig estimates were substantially lower.Labels were more likely to be absent near the estimated Bayes boundary.
  • Predictive performance: The full estimator had the lowest mean estimated error at every expert level, reducing error relative to ig by 0.076, 0.075, and 0.053.Corresponding reductions relative to CC were 0.036, 0.036, and 0.030 for junior, intermediate, and senior experts.
  • Predictive performance: The example supports that modeling feature-dependent missingness can bring decision-boundary estimates closer to the complete-data benchmark while reducing cross-validated error.Errors were estimated using repeated stratified five-fold cross-validation with refitting in each training fold.

7 Discussion

The paper frames informative label missingness as useful classifier information in Weibull mixtures, while emphasizing that classification gains depend on boundary estimation rather than all-parameter accuracy.

  • Discussion: A common-shape Weibull Bayes rule has one positive boundary, whereas unequal shapes can produce two, shaping the relevant classification geometry.The Fisher-information and expected-error expressions make the missing-label pattern explicit.
  • Discussion: First-order classification error depends on parameter uncertainty only in directions that perturb decision boundaries, so improved estimation of every model parameter is unnecessary for a classification gain.The simulations and hard-drive analysis support this distinction.
  • Discussion: Useful extensions include nonregular boundary configurations, censored Weibull data, more flexible missingness models, and multiclass settings.These are identified as directions for further work.

A Decision-boundary proof

The proof analyzes the transformed log-posterior odds under unequal Weibull shapes, establishing the possible root counts and corresponding class regions. It also shows that any tangential root is exactly double.

  • Unequal-shape geometry: The transformed log-posterior odds has exactly one stationary point because its second derivative changes sign at most once.When k1 > k2, the function has a unique maximum; when k1 < k2, it has a unique minimum.
  • Unequal-shape geometry: When k1 > k2, the function tends to −∞ in both tails and can have zero, one double, or two simple zeros.With two simple zeros, it is positive only between them.
  • Multiplicity: A tangential zero at the stationary point has multiplicity exactly two, and this multiplicity is preserved under the transformation y = e^t.The proof uses the nonzero second derivative at the stationary point and the smooth one-to-one reparameterization.
  • Unequal-shape geometry: When k1 < k2, the function tends to +∞ in both tails and likewise permits zero, one double, or two simple zeros.The same tangency argument applies at the unique minimum.

B Derivation of the information identity

The information derivation decomposes the observed likelihood into the ignored likelihood and the missingness likelihood, then profiles out missingness-model nuisance parameters. This identifies how label indicators alter classifier information relative to ignoring missingness.

  • Ignored likelihood: The ignored-data score combines the marginal feature score with the observed-label conditional score.Conditional score identities make the conditional score contribution mean zero given the feature.
  • Ignored likelihood: The sensitivity and variability matrices of the ignoring score coincide, yielding the ignored estimator’s Fisher-information representation under the stated regularity conditions.The resulting sandwich covariance reduces accordingly.
  • Information decomposition: The full likelihood separates into the ignored likelihood and the Bernoulli missingness likelihood, so their information contributions add.The nuisance-parameter blocks arise only from the missingness likelihood.
  • Nuisance adjustment: Profiling out the nuisance parameter ξ produces the Schur complement of its information block.Substitution into the block-matrix expression gives the adjusted classifier information.
  • Nuisance adjustment: Under the stated positive-definiteness condition, the full-information comparison follows through Loewner-order inversion.The derivation separates information lost from unobserved labels and information contributed by missing-label indicators.

C Proof of the expected error-rate expansion

The proof expands excess classification error locally around each simple Bayes boundary and connects boundary displacement to estimator variability. Root-n consistency and uniform integrability justify taking expectations at first order.

  • Boundary perturbation: A simple Bayes boundary extends differentiably as the model parameter varies, by the implicit function theorem.The perturbed boundary satisfies δ{yj(eθ); eθ} = 0.
  • Boundary perturbation: Near the true parameter, disagreement between the plug-in and Bayes rules occurs only on disjoint intervals between corresponding boundaries.This localizes excess error to boundary displacement.
  • Expected error: Summing the local contributions over the finitely many boundaries gives the expected excess-error representation.The representation is obtained after applying the boundary expansion at each decision boundary.
  • Expected error: Root-n consistency makes the parameter error Op(n−1/2), while the remainder is negligible relative to the quadratic boundary-displacement term.Uniform integrability and the stated moment condition justify restricting the expansion to a local neighborhood.
  • Expected error: For regular maximum-likelihood estimation, substituting V = J(θ)^−1 yields the stated first-order result; the ignoring estimator uses its corresponding covariance.The distinction enters through the estimator’s information structure.

D Hard-drive expert-label construction and diagnostics

The semi-synthetic hard-drive study ranks observations by committee-based classification confidence and releases labels in three nested expert regimes. Diagnostics indicate feature-dependent missingness, while full estimation trades class-specific performance for a lower overall estimated error rate.

  • Expert-label construction: A seven-classifier committee score combines agreement, distance of mean class probability from 1/2, and between-classifier variability.The score ranks observations only for constructing the semi-synthetic missingness regimes.
  • Expert-label construction: Labels were released for the highest-ranked 60%, 70%, and 80% of observations, producing nested missing-label fractions of 0.402, 0.304, and 0.201.These regimes represent increasing expert-assessment coverage.
  • Missingness diagnostics: Estimated missingness slopes were negative in all three regimes, consistent with observations nearer the estimated Bayes boundary being more likely to remain unlabelled.Their magnitudes are not directly comparable because the regimes have different overall missing-label fractions.
  • Classification results: Relative to ig, full increased class 2 sensitivity but reduced class 1 specificity, so its lower overall estimated error rate was not uniform across classes.The class-specific rates were evaluated using five repetitions of five-fold cross-validation.
Loading 2609.00774v1…