Source-linked AI summary
Informative Label Missingness in Multiclass Classification Information Geometry and Excess Risk
Fariborz Setoudehtazang, Geoffrey J. McLachlan
TL;DR
Informative label missingness raises whether partially labelled data can outperform complete classification when missingness carries model information. The paper develops a likelihood-based multiclass information decomposition and a boundary-weighted excess-risk analysis, finding that partial classification can be favourable when information gains align with active Bayes-boundary directions, although the effect is regime-dependent and need not reflect global Fisher-information dominance.
Problem
The paper addresses the limited understanding of how informative label missingness affects classification risk in general multiclass problems with multiple active Bayes boundaries.
Method
The paper combines an efficient-information decomposition with a quadratic excess-risk expansion over active pairwise Bayes faces and a classification-weighted generalized-eigenvalue comparison.
Results
Informative partial classification can achieve smaller leading excess classification risk without globally dominating complete classification in Fisher information, depending on information alignment with active decision boundaries.
Takeaways & Limitations
The value of a partially labelled sample depends on where information gains and losses lie relative to decision-boundary geometry, not only on observed-label proportions or unweighted Fisher information.
Takeaways & Limitations
The theory assumes parametric regular likelihood asymptotics, smooth transversal Bayes-boundary geometry, conditional independence of missingness and latent class given features, and an adequately specified missingness model.
Abstract
from arXiv · showhide
Informative label missingness can change the usual efficiency ordering between completely and partially labelled classifiers because the pattern of missing labels may itself carry information about the classification model. We develop a general likelihood-based theory for this phenomenon in parametric multiclass classification. An efficient-information decomposition separates information lost through unavailable class memberships from information contributed by the missing-label mechanism. We then derive a quadratic expansion of plug-in excess risk over the active pairwise faces of the multiclass Bayes boundary, showing that classification efficiency depends on how information gains and losses align with directions that perturb the decision boundary. This yields a classification-weighted generalized-eigenvalue criterion under which informative partial classification may have smaller asymptotic classification risk without globally dominating complete classification in Fisher information. Near missing completely at random, with the marginal missing-label proportion fixed, redistribution of missing labels changes lost class-label information at first order, whereas efficient information from the missingness pattern appears only at second order. Three-class quadratic discriminant calculations, finite-sample experiments, and a semi-synthetic multiclass application illustrate the resulting regime-dependent behaviour.
1 Introduction
The paper develops a general multiclass likelihood framework for informative label missingness, where missingness can both remove class-label information and reveal information about the classifier. It connects these information effects to the geometry of active Bayes boundaries to explain when partial classification can improve asymptotic classification risk.
- Partially labelled data arise when features are observed but some class memberships are unavailable, forming a missing-data problem related to semi-supervised learning.
- When label availability depends on features and classification parameters, missing-label indicators can themselves inform estimation of the classification model.The mechanism is conditionally independent of the latent class given features but need not be ignorable when its distribution depends on the classification parameters.
- The classification-specific comparison weights information changes by their effects on active pairwise Bayes boundaries rather than requiring global Fisher-information dominance.A generalized-eigenvalue representation captures whether gains and losses align with parameter directions that perturb the decision boundary.
- Near MCAR with a fixed marginal missing-label proportion, redistribution of missing labels changes lost class-label information at first order, while information from missingness indicators appears only at second order.A weakly informative mechanism may initially worsen classification, and later improvement depends on global information geometry and boundary alignment.
- Three-class Gaussian quadratic discriminant calculations, finite-sample experiments, and a multiclass application illustrate regime-dependent classification-risk comparisons.The three-class model uses unequal covariance matrices to produce curved pairwise Bayes faces and multiclass junctions while retaining interpretable decision geometry.
- The efficient-information decomposition separates information lost from unavailable class memberships from efficient information contributed by the missingness indicators.The loss depends on where labels are missing in feature space, not only on the marginal missing-label proportion; under MCAR, the indicators contribute no classification-parameter information.
3 Efficient classifier information in the presence of nuisance parameters
Efficient information for classifier parameters must account for nuisance-parameter elimination and the geometry of the active Bayes boundary. The resulting risk expansion weights parameter directions by how strongly they perturb active decision faces.
- Nuisance parameters: Nuisance parameters are removed by efficient-score projection, while parameters determining the Bayes boundary remain classifier-relevant.For QDA, class probabilities, means, and covariance matrices are retained because classification risk depends on the full decision boundary.
- Nuisance parameters: Eliminating additional nuisance parameters introduces a nonnegative coupling penalty that can offset positive information changes in the classifier-relevant block.A gain in the classifier-relevant direction is insufficient unless it dominates coupling with the nuisance score.
- Excess-risk geometry: Only active pairwise Bayes faces contribute to the leading quadratic excess-risk term; generic triple and higher-order junctions contribute only at higher order.Active faces are portions of pairwise equality surfaces where the two classes jointly attain the Bayes maximum.
- Excess-risk geometry: The curvature matrix HR measures classification relevance through the first-order displacement of active Bayes faces under parameter perturbations.A direction is locally irrelevant exactly when it produces no first-order displacement on almost every active face.
- Information–risk link: Expected excess-risk differences connect efficient information matrices to boundary geometry, so information gains matter only in directions assigned substantial classification relevance.The criterion characterizes favourable informative missingness at leading n^-1 order.
5 Classification-weighted information and favourable missingness
The paper compares complete and informative partial classification through a classification-weighted information geometry. Favourable missingness occurs when information gains align with decision-boundary-sensitive directions, while local departures from MCAR need not help.
- Risk criterion: The leading excess-risk comparison uses complete and partial efficient information matrices together with the classification-risk curvature matrix HR.The asymptotic relative efficiency exceeds one exactly when the leading excess-risk difference is positive.
- Spectral characterization: Generalized eigenvalues identify directions with more or less information under partial classification, while classification weights measure their relevance to active Bayes faces.λj > 1 denotes an information gain and λj < 1 an information loss in direction vj.
- Spectral characterization: Favourable informative missingness does not require global Fisher-information dominance; gains in highly relevant directions can outweigh losses elsewhere.Neither J ⪰ A nor J ⪯ A is required in the general case.
- Departures from MCAR: Near MCAR with fixed marginal missingness, redistributing unavailable labels changes class-label information loss at first order, whereas missingness-pattern information enters at second order.Weak dependence between missingness and classification uncertainty therefore need not improve classification.
- Departures from MCAR: A local improvement analysis does not imply that classification advantage is monotone or eventually positive as dependence on uncertainty increases.Whether a favourable regime emerges at larger slopes depends on the global information geometry of the mechanism.
6 Three-class quadratic discriminant analysis
The numerical framework uses a three-class, two-dimensional QDA model to study classification-weighted information in a genuinely multiclass nonlinear geometry. Unequal covariances create curved boundaries and allow a transversal three-class junction.
- Model: The numerical analysis specializes the general theory to g = 3 classes and p = 2 features, the smallest configuration supporting curved active faces and a genuine three-class junction.The setting remains directly visualizable while retaining multiclass and nonlinear decision geometry.
- Bayes geometry: Each active face is a quadratic boundary segment where two class discriminants tie and exceed the third class discriminant.The face-wise construction supports attribution of classification gains and losses to individual pairwise boundaries.
- Bayes geometry: The three-class junction is treated as transversal and therefore contributes only a higher-order term to local excess risk.The junction is isolated under the stated gradient condition.
- Parameterization: QDA parameters include class probabilities, means, and covariance matrices, with positive-definite covariance optimization handled through log-Cholesky coordinates.Information and risk criteria are invariant under the smooth nonsingular reparameterization.
- Missingness mechanisms: Missingness mechanisms are driven by posterior uncertainty, with normalized Shannon entropy as the primary measure and normalized Gini uncertainty as a sensitivity check.Positive missingness slopes make labels more likely to be unavailable in regions of greater posterior uncertainty.
7 Numerical investigation
The numerical studies test when informative partial classification improves asymptotic classification risk in a three-class QDA model. Results show regime dependence on uncertainty strength, missingness proportion, class geometry, priors, and uncertainty measure.
- Reference configuration: 0.7791/n is the complete-classification leading excess-risk coefficient in the reference configuration.This is the first-order population quantity used for baseline comparison.
- Reference configuration: At a 30% marginal missing-label rate, informative partial classification is favourable when uncertainty dependence is sufficiently strong, whereas MCAR is detrimental.The same missing-label proportion produces opposite outcomes under sufficiently strong uncertainty dependence and MCAR.
- Phase transition: The critical uncertainty slope increases monotonically with missing-label proportion, so stronger dependence is required to offset greater class-label information loss.The phase boundary separates unfavourable and favourable regimes where ∆R = 0, equivalently ARER = 1.
- Phase transition: Near MCAR, the missingness mechanism first redistributes information loss before its own information contribution becomes large enough to compensate.This explains why sufficiently strong dependence is needed before the favourable regime appears.
- Geometric sensitivity: The advantage is nonmonotone in class separation: severe overlap is unfavourable, moderate separation can be favourable, and strong separation weakens the gain.At moderate separation, uncertainty remains concentrated near active boundaries, while strong overlap makes observed labels especially informative.
- Geometric sensitivity: Class overlap and prior imbalance can reverse the sign of ∆R, while covariance heterogeneity primarily changes the magnitude within the examined ranges.Along the covariance path, relative efficiency changes from 1.160 under common covariance matrices to 1.142 in the reference QDA model.
- Uncertainty sensitivity: Normalized Gini uncertainty yields the same qualitative phase structure as normalized Shannon entropy but reaches the favourable region at a smaller slope in the reference configuration.The comparison is presented as a robustness check rather than an ordering of uncertainty measures.
- Information–risk interpretation: Classification advantage reflects weighted directional alignment rather than uniform information improvement, because some parameter directions can lose information while relevant gains dominate.The generalized eigenvalues are not uniformly greater than one, so partial classification need not dominate complete classification in Loewner order.
8 Finite-sample validation
Finite-sample experiments test whether information-based population predictions, n^-1 excess-risk behavior, and the classification-weighted ordering persist empirically. In the reference configuration, IPC outperforms CC while MCAR performs worst, although application results remain regime-dependent.
- Design: The finite-sample design evaluates covariance convergence, n^-1 excess-risk behavior, and whether the population criterion predicts relative performance.Experiments use a three-class QDA configuration with n ∈ {250, 500, 1000} and B = 500 Monte Carlo replications.
- Design: With γ = 0.30, MCAR removes labels independently, whereas IPC concentrates missing labels in regions of high posterior uncertainty while preserving the marginal missingness rate.The IPC intercept is calibrated so that E{q(Y)} = 0.30, matching MCAR's marginal proportion of unavailable labels.
- Finite-sample results: The theoretical ordering is K_IPC < K_CC < K_MCAR, and this ordering is reproduced at every sample size considered.Thus the reference configuration predicts lower scaled excess classification risk for IPC than CC, with MCAR highest.
- Finite-sample results: At n = 1000, scaled risks are 1.5247 for CC, 2.0294 for MCAR, and 1.3682 for IPC, close to theoretical limits 1.5581, 2.0541, and 1.3645.The IPC estimate differs from its first-order limit by about 0.004.
- Finite-sample results: IPC-to-CC relative efficiencies are approximately 1.052, 1.034, and 1.114 for n = 250, 500, and 1000, compared with the population value 1.142.The favorable population comparison is already visible at the largest sample size.
- Validation: At n = 1000, classification-weighted covariance discrepancies are approximately 1.4% for CC, 0.5% for MCAR, and 0.7% for IPC.Quadratic-risk estimates are also close to directly evaluated scaled risks and theoretical limits.
- Application: The semi-synthetic application shows regime dependence: IPC reduced mean error relative to CC at γ = 0.20 and ξ1 = 6 log 3, but not at γ = 0.30 or γ = 0.40.Under the reference mechanism, IPC improved log loss and Brier score relative to IG, while CC had the lowest mean misclassification rate.
10 Discussion
The paper evaluates partial classification by how information gains and losses align with directions affecting the active Bayes boundary, rather than by Fisher information alone. Theory and experiments show that informative missingness can help or hurt depending on the regime.
- Discussion: Classification risk depends on information directions that perturb active Bayes-boundary faces, not merely on global Fisher-information ordering.The resulting criterion permits lower asymptotic classification risk without global information dominance.
- Discussion: Near MCAR, fixed marginal missingness redistributes lost class-label information at first order, while missingness-indicator information appears only at second order.Weak informativeness may initially worsen classification before a favourable regime emerges.
- Discussion: Three-class QDA results show that favourable missingness is reduced by severe class overlap, can reverse under strong class imbalance, and remains comparatively stable along the examined covariance-heterogeneity path.Entropy- and Gini-based mechanisms exhibit the same qualitative phase structure.
- Discussion: Finite-sample experiments reproduce population risk-coefficient orderings and approach theoretical excess-risk and risk-weighted covariance quantities as sample size increases.This supports the information–geometry explanation rather than only one numerical comparison.
- Discussion: In the Vertebral Column application, modelling uncertainty-dependent missingness generally improves prediction over treating the pattern as ignorable, but partial classification does not uniformly beat complete classification.Improvements are modest under weak or moderate dependence and appreciable under the strongest mechanism.
- Discussion: The theory is parametric and requires regular likelihood asymptotics, smooth transversal boundaries, an adequate missingness model, and conditional independence of missingness from latent class given features.Nonregular, singular, high-dimensional, directly class-dependent, or misspecified settings require further work.
- Discussion: The analysis treats the missingness mechanism as given, motivating future design of labelling or abstention policies under a labelling budget.The classification-weighted information criterion is proposed as a starting point for that extension.
S1 Proof of the information decomposition
The proof derives the observed-data likelihood and decomposes efficient information into feature, observed-label, and missingness contributions. Under MCAR, only the information loss from unavailable labels remains.
- S1 Proof of the information decomposition: The observed-data likelihood combines the feature model, observed class-label likelihood, and Bernoulli missingness mechanism.It is written for O=(Y,M,(1−M)Z) under M ⊥ Z | Y.
- S1 Proof of the information decomposition: The score for θ is the sum of the marginal feature score, observed-label score, and missingness-mechanism score.The nuisance parameter ξ contributes through the missingness score.
- S1 Proof of the information decomposition: Conditional mean-zero identities establish orthogonality among marginal-feature, observed-label, and missingness score contributions.This permits additive Fisher-information blocks.
- S1 Proof of the information decomposition: Partial classification loses the conditional class-label information associated with observations whose labels are unavailable.The proof compares the partial score contribution with the complete-classification score.
- S1 Proof of the information decomposition: Efficient information for θ is obtained by taking the Schur complement of the missingness-specific nuisance block.Positive semidefiniteness follows from the Bernoulli information matrix and its Schur complement.
- S1 Proof of the information decomposition: Under MCAR, qθ=0, so missingness indicators carry no information about θ and only class-label information loss remains.The conclusion follows after substituting the MCAR expressions into the decomposition.
S2 Proof of the nuisance-parameter information result
The proof removes nuisance parameters through block transformations and Schur complements. Sequentially eliminating missingness-specific and model nuisance parameters yields the same efficient information as joint elimination.
- S2 Proof of the nuisance-parameter information result: The complete-classification information matrix is partitioned into target parameter β and nuisance parameter λ, assuming the λ block is nonsingular.Its principal-block structure ensures positive definiteness needed for Schur complements.
- S2 Proof of the nuisance-parameter information result: The efficient score for β is formed by subtracting the projection onto the nuisance-score component, with covariance equal to the usual Schur complement.This is the standard efficient-information construction used in the proposition.
- S2 Proof of the nuisance-parameter information result: After eliminating ξ, the proof applies a nonsingular block-triangular transformation that replaces the β score by its nuisance-adjusted version.Information transforms by congruence while the nuisance-score space remains unchanged.
- S2 Proof of the nuisance-parameter information result: The transformed information matrix has the efficient β information in its upper-left block after nuisance adjustment.The resulting block expressions are given for both information and perturbation matrices.
- S2 Proof of the nuisance-parameter information result: Sequential elimination of λ and ξ is equivalent to their joint elimination from the full observed-data information matrix.The equivalence follows from the quotient identity for Schur complements.
S3 Proof of the multiclass excess-risk expansion
The proof localizes classifier disagreement near Bayes boundaries and expands excess risk face by face. Regular active pairwise faces determine the quadratic term, while higher-order tie neighborhoods are negligible.
- S3 Proof of the multiclass excess-risk expansion: A small parameter perturbation can change the classifier only in a shrinking neighborhood of the Bayes boundary.Away from the boundary, the winning class remains separated by a positive margin.
- S3 Proof of the multiclass excess-risk expansion: The excess loss is obtained by integrating the local disagreement-strip contribution over regular portions of each active pairwise face.Summing these facewise terms yields the quadratic expansion.
- S3 Proof of the multiclass excess-risk expansion: At an interior active face, competing classes other than the leading pair remain separated, reducing the local comparison to a binary boundary.The pairwise contrast supplies a normal coordinate to the boundary.
- S3 Proof of the multiclass excess-risk expansion: Tubular coordinates and a Taylor expansion quantify the perturbed boundary displacement and the volume of the disagreement strip.The coarea formula supplies the corresponding volume element.
- S3 Proof of the multiclass excess-risk expansion: For an r-way transversal tie, the affected neighborhood has volume O(ε^(r−1)) and excess-risk contribution of order O(ε^r).Triple and higher-order tie neighborhoods therefore contribute only o(ε^2).
- S3 Proof of the multiclass excess-risk expansion: Only regular active pairwise faces contribute to the leading quadratic term; higher-order tie strata affect lower-order remainders.The result extends to unbounded supports under additional tail, exhaustion, and integrability conditions.
S4 Noncompact boundaries and tail conditions
The section extends the quadratic excess-risk expansion to noncompact active Bayes boundaries by imposing regular-exhaustion, boundary-integrability, and tail conditions. Truncation and limiting arguments then recover the expansion and spectral criterion.
- Assumptions: Noncompact active faces require an increasing sequence of regular truncation radii, together with boundary-integrability and tail-control assumptions.These conditions ensure regular face intersections, finite curvature integrals, and negligible tail contributions.
- Limiting argument: For each fixed truncation radius, the local remainder is o(∥h∥2) as h → 0.
- Limiting argument: Choosing the truncation radius sufficiently large controls boundary and tail terms, while the local remainder vanishes as the perturbation tends to zero.
- Conclusion: The quadratic excess-risk expansion remains valid for noncompact active Bayes boundaries under the stated regular-exhaustion and boundary-integrability conditions.
- Spectral criterion: The spectral proof separates information gains and losses, yielding the classification-weighted gain–loss criterion for the sign of excess-risk difference.
S6 Proof of the local departure from MCAR result
The local analysis near MCAR calibrates the missingness mechanism while holding the classification model fixed. It shows that redistributed missingness changes label-information loss at first order, whereas mechanism information begins at second order.
- Local expansion: At MCAR, the first-order perturbation of the missingness probability depends on centered uncertainty, U − E(U).
- Calibration: The missingness probability is calibrated to preserve the marginal missing-label proportion along a path varying only the mechanism.The classification model remains fixed at θ0, while the intercept adjusts locally through an implicit-function argument.
- Efficient information: The intercept and uncertainty slope are treated as nuisance parameters in the efficient-information calculation.
- Regularity: The assumed variance and moment conditions ensure nonsingularity of the nuisance information block and justify differentiation under expectations.
- Result: Label-information loss changes at first order in t, whereas efficient information from missing-label indicators begins at order t2.
S7 Quadratic discriminant calculations
The QDA calculations provide the derivatives and information matrices needed to evaluate classification efficiency under informative missingness in a three-class model. They express boundary sensitivity through active pairwise log contrasts.
- Model parameterization: The three-class QDA parameterization includes class probabilities, class means, and covariance parameters, using class 3 as the baseline.
- Boundary geometry: On an active face, density contrasts can be represented through log contrasts, whose derivatives supply the surface-integral contribution.The direct-contrast representation does not require positivity, whereas the log-contrast form is convenient for QDA calculations.
- Complete-classification information: Gaussian mean, covariance, and prior score blocks are orthogonal under the stated class structure and conditional-mean properties.
- Posterior and boundary derivatives: Posterior probabilities are differentiated with respect to spatial coordinates and class-specific parameters to obtain boundary-sensitivity quantities.
- Missingness information: The logistic missing-label mechanism uses posterior uncertainty, and its derivatives generate the Bernoulli information blocks combined with conditional label-information loss.
S8 Additional population robustness results
The robustness calculations show that informative partial classification is regime-dependent across covariance geometry, class-prior imbalance, uncertainty measures, and finite samples. Its advantage can persist across controlled covariance paths but reverse for some imbalanced configurations.
- Covariance heterogeneity: 1.160 under common covariance matrices versus 1.142 in the reference QDA model: informative partial classification remains favourable along the covariance-heterogeneity path.The advantage decreases modestly as covariance heterogeneity increases, while ignoring informative missingness performs worse throughout the path.
- Prior imbalance: At π3 = 0.05, ARE12 = 1.040, ARE13 = 0.968, and ARE23 = 0.856, so losses on rare-class faces outweigh the gain on the common-class face.
- Prior imbalance: At π3 = 0.90, ARE12 = 0.931, ARE13 = 0.938, and ARE23 = 1.051, leaving only the 2-versus-3 boundary favourable.
- Geometry-dependent phase boundaries: The favourable regime depends jointly on missing-label proportion and classification geometry; stronger uncertainty dependence alone does not guarantee a crossing.
- Uncertainty measures: Entropy- and Gini-based mechanisms share the same qualitative phase structure, although Gini reaches the favourable region at a smaller normalized slope in the reference QDA configuration.
- Finite-sample validation: At n = 1000, classification-weighted covariance discrepancies are approximately 1.4%, 0.5%, and 0.7% for CC, MCAR, and IPC, respectively.The finite-sample calculations approach the theoretical coefficients and directly evaluated excess risks as sample size increases.