Source-linked AI summary
Classification with Asymmetric Label Noise: Consistency and Maximal Denoising
Gilles Blanchard, Marek Flaska, Gregory Handy, Sara Pozzi, Clayton Scott
TL;DR
The paper addresses classification when random label noise is asymmetric, unknown, and present despite overlapping class distributions. It derives identifiability conditions based on majority-correct labels and mutual irreducibility, connects estimation to mixture proportion estimation, and develops a consistent discrimination rule. The proposed approach is supported by benchmark and nuclear particle experiments, while the CPE implementation requires further development.
Problem
Existing label-noise theory commonly assumes separable classes, symmetric noise, or known noise proportions, leaving more general settings insufficiently addressed.
Method
The paper characterizes identifiable contamination models using majority-correct labeling and mutual irreducibility, estimates noise through mixture proportion estimation, and applies the estimates to surrogate-loss discrimination.
Results
The proposed discrimination rule is universally consistent for the maximally denoised distributions, and experiments report accurate noise estimation and good ROC-method accuracy across three datasets.
Takeaways & Limitations
The framework handles overlapping classes and asymmetric, unknown noise while identifying a unique maximally denoised pair under its conditions.
Takeaways & Limitations
The CPE implementation still requires further development, and reported findings warrant further investigation because competing CPE results differ.
Abstract
from arXiv · showhide
In many real-world classification problems, the labels of training examples are randomly corrupted. Most previous theoretical work on classification with label noise assumes that the two classes are separable, that the label noise is independent of the true class label, or that the noise proportions for each class are known. In this work, we give conditions that are necessary and sufficient for the true class-conditional distributions to be identifiable. These conditions are weaker than those analyzed previously, and allow for the classes to be nonseparable and the noise levels to be asymmetric and unknown. The conditions essentially state that a majority of the observed labels are correct and that the true class-conditional distributions are "mutually irreducible," a concept we introduce that limits the similarity of the two distributions. For any label noise problem, there is a unique pair of true class-conditional distributions satisfying the proposed conditions, and we argue that this pair corresponds in a certain sense to maximal denoising of the observed distributions. Our results are facilitated by a connection to "mixture proportion estimation," which is the problem of estimating the maximal proportion of one distribution that is present in another. We establish a novel rate of convergence result for mixture proportion estimation, and apply this to obtain consistency of a discrimination rule based on surrogate loss minimization. Experimental results on benchmark data and a nuclear particle classification problem demonstrate the efficacy of our approach.
1 Introduction
The paper studies binary classification when training labels are randomly corrupted, including asymmetric and unknown noise with overlapping class supports. It establishes identifiability and consistency under majority-correct labeling and mutual irreducibility, then evaluates the approach experimentally.
- Problem: Random label noise corrupts training labels, while the classes may overlap and the noise may be asymmetric, unknown, and independent of features.The contamination model mixes each true class distribution with the other class, with proportions π0 and π1.
- Identifiability: The paper introduces necessary and sufficient conditions that uniquely determine true class-conditional distributions and noise proportions from contaminated distributions.The conditions include a majority of labels being correct on average and mutual irreducibility of the class-conditional distributions.
- Identifiability: Mutual irreducibility excludes expressing either true class distribution as a nontrivial mixture of the other class and another distribution.This condition limits the similarity between the two class-conditional distributions.
- Maximal denoising: The identified solution corresponds to maximal denoising of the observed distributions, maximizing total compatible label noise and total variation separation under the majority-correct constraint.The paper characterizes the population geometry of all solutions and argues that the selected pair is unique.
- Consistency: The authors develop a universally consistent discrimination rule that estimates unknown noise proportions and targets optimal performance for the maximally denoised distributions.The approach connects the problem to mixture proportion estimation and surrogate loss minimization.
- Experiments: Experiments on benchmark datasets and nuclear particle classification indicate that the proposed methodology can accurately estimate noise proportions and is practically viable.The motivating application concerns distinguishing neutron from gamma-ray pulses for nuclear safeguards and nonproliferation.
2 The Challenge of Label Noise
Label noise preserves some discrimination structure but can make classifiers optimized on contaminated performance measures suboptimal for clean evaluation. The paper analyzes this mismatch across several performance criteria.
- Shared ROC: Under the majority-correct condition, true and contaminated likelihood-ratio tests generate the same receiver operating characteristic.The thresholds differ, but sweeping either threshold produces the same ROC.
- Performance mismatch: A classifier optimized for contaminated performance generally fails to optimize uncontaminated performance, except in special cases.The paper examines probability of error, Neyman–Pearson, minmax, and balanced error criteria.
- Neyman–Pearson: For Neyman–Pearson classification, matching contaminated and clean Type I error rates generally fails when the negative class is contaminated.Equality occurs only when π0 = 0 or when the classifier effectively random-guesses under the stated conditions.
- Minmax: For minmax classification, contaminated and clean balancing coincide only under symmetric noise or a degenerate condition associated with random guessing.Thus asymmetric label noise changes the operating point selected by contaminated minmax optimization.
- Balanced error: Balanced error is reported as the only function of the two class errors whose corrupted and clean optimization objectives remain equivalent regardless of noise proportions or class priors.This does not remove the need to estimate true performance for other criteria.
3 Alternate Contamination Model
An alternate contamination representation decouples the two observed distributions and reduces estimation of transformed noise proportions to mixture proportion estimation. Under the majority-correct condition, this representation is uniquely linked to the original model.
- Alternate representation: The alternate contamination model separates estimation of each transformed noise proportion by expressing each contaminated distribution using one observed distribution and one transformed mixture weight.This decoupling avoids involving P1 when estimating one proportion and P0 when estimating the other.
- Alternate representation: If P0 ≠ P1 and the majority-correct condition holds, the contaminated distributions are distinct and have unique transformed proportions 0 ≤ ˜π0, ˜π1 < 1.The transformed proportions satisfy ˜π0 = π0/(1−π1) and ˜π1 = π1/(1−π0).
- Identifiability: Conversely, distinct contaminated distributions imply unique original contamination proportions under the alternate representation.The resulting expressions recover π0 and π1 from ˜π0 and ˜π1.
- Identifiability: The two contamination representations establish an explicit one-to-one correspondence between original and transformed noise proportions.This correspondence holds for known distinct uncontaminated distributions under the majority-correct constraint.
- Mixture proportion estimation: Estimating transformed proportions reduces the label-noise problem to mixture proportion estimation, supporting subsequent estimation of clean classifier errors.The paper uses this framework to estimate contaminated and then uncontaminated Type I and Type II errors.
4 Irreducibility and Mixture Proportion Estimation
This section defines irreducibility and mixture proportion estimation to resolve otherwise non-identifiable mixture decompositions. It connects the maximal mixture proportion to ROC geometry, density ratios, and practical checks for mutual irreducibility.
- Identifiability: Without additional assumptions, a mixture proportion is not identifiable because multiple decompositions of the same observed distribution are valid.The ambiguity arises because an unknown component can absorb part of the other component while preserving the mixture.
- Irreducibility: Mutual irreducibility requires that neither distribution contain a nonzero mixture component of the other.Equivalently, each distribution cannot be decomposed into the other distribution plus a residual probability distribution.
- Mixture Proportion Estimation: For F = (1 −γ)G + γH with G irreducible relative to H, the identifiable mixture proportion equals the maximal proportion κ∗(F|H) of H present in F.This maximal proportion is the largest α for which F can be written as (1 −α)G′ + αH.
- ROC Interpretation: κ∗(F|H) is linked to the slope of the optimal ROC at its right endpoint and motivates both universal and practical estimators.The ROC interpretation treats H(S) as the false-positive rate and F(S) as the true-positive rate.
- Examples and Characterizations: Mutual irreducibility can be checked through density ratios: for continuous distributions, their essential infimum and supremum must be 0 and ∞.It holds for examples including distributions with non-contained supports and equal-variance Gaussians with unequal means, although κ∗(P1|P0) can decrease rapidly as means separate.
5 Mutual Irreducibility: Sufficiency, Necessity, and Maximal Denoising
Mutual irreducibility is necessary and sufficient for uniquely identifying the contamination model, and selects a unique maximally denoised solution among generally many decompositions.
- Sufficiency: Mutual irreducibility of P0 and P1 is necessary and sufficient for identifying the contamination proportions and true class-conditional distributions.It is equivalent to irreducibility relative to the contaminated opposite class when the majority-correct-label condition holds.
- Sufficiency: Given distinct observed distributions, the mutually irreducible solution is unique and can be recovered from explicit contamination-proportion functions.The recovered proportions are obtained from mixture-proportion quantities, after which P0 and P1 follow from the contamination identities.
- Necessity: Without additional conditions, the population contamination equations generally admit multiple solutions, including the trivial decomposition and class-label swaps.Thus, decontamination is not well-defined solely from the observed contaminated distributions.
- Necessity: Mutual irreducibility is forced by universality, symmetry, continuity of recovered weights, and stability of recovered sources.Theorem 11 states that any decontamination operator satisfying these requirements must return the unique mutually irreducible solution.
- Maximal Denoising: The feasible contamination-proportion region is a closed quadrilateral, with the mutually irreducible solution as its unique extremal point where both constraints are active.The same solution uniquely maximizes the total variation distance between P0 and P1.
- Maximal Denoising: The unique mutually irreducible solution maximizes both total label noise and total variation separation between the recovered source distributions.This gives mutual irreducibility its interpretation as maximal denoising or maximal source separation.
6 Mixture Proportion Estimation and a Rate of Convergence
The paper connects mixture proportion estimation to label-noise decontamination and establishes conditions yielding consistency and a known convergence rate. These results support estimating contamination proportions from observed distributions.
- Mixture proportion estimation: Without assumptions, mixture proportion decompositions are non-identifiable; the paper reviews a universally consistent estimator for the maximal mixture proportion.The estimator is based on empirical distributions and infima over unions of VC classes.
- Estimator properties: The estimator remains an upper bound on κ∗ with high probability, using VC concentration and a union-bound argument.The stated bounds apply to empirical true-positive and false-positive probabilities over rejection regions.
- Caveat: The paper notes that an earlier consistency statement contains an error corrected in an appendix, without affecting the present work.The correction is attributed to the consistency result reviewed from Blanchard et al. (2010).
- Distributional assumptions: Assumption (D) requires a support relationship that makes one distribution irreducible with respect to the other and identifies the target mixture proportion.Under (D), γ equals κ∗(F|H).
- Distributional assumptions: Assumption (AP2) requires a set in the approximation sequence with zero probability under G and positive probability under H.This support-separating condition is sufficient for the rate analysis.
- Rate result: Under (AP2) and (D), the paper proves a finite-sample convergence-rate result for the mixture proportion estimator.Theorem 14 supplies a constant C > 0 for sufficiently large sample sizes.
7 Consistent Classification with Unknown Label Noise Proportions
The paper uses mixture proportion estimates to recover unknown label-noise proportions and construct consistent classifiers. It then makes the procedure computationally tractable through surrogate-risk minimization, including a clippable-loss variant.
- Unknown-noise classification: Estimated mixture proportions and contaminated risks can be plugged into risk identities to obtain uniformly consistent estimates over VC classifier classes.Growing the VC class with sample size yields empirical-risk-minimization consistency for performance measures defined through R0 and R1.
- Surrogate-risk approach: Empirical risk minimization over general VC classes is computationally intractable, motivating a computationally tractable rule based on surrogate risk minimization.The paper develops the surrogate approach as an alternative to direct VC-class empirical risk minimization.
- Risk transformation: Under label noise independent of X conditional on Y, minimizing an appropriately cost-sensitive noisy risk is linked to minimizing the clean cost-insensitive risk.The relationship uses the transformed conditional probability ˜η(x) and the parameter α.
- Problem formulation: The target problem is to construct a computationally tractable classifier that does not know α, π0, or π1 and whose excess P-risk converges to zero.The proposed algorithm instead drives the corresponding cost-sensitive contaminated-risk excess toward zero.
- Estimating noise parameters: When (A) and (C) hold, mixture proportion estimation provides consistent estimates of the unknown contamination proportions and therefore of α.Known-rate convergence of these estimates requires the stronger two-direction assumption (C’).
- Consistency results: Theorems 18 and 19 establish consistency under universal bounded kernels and Lipschitz losses, with Theorem 19 using clippable losses under the milder condition (C).The clippable-loss result avoids requiring a convergence rate for α.
8 A More General Analysis of Co-Training
The co-training analysis converts predictions from one feature view into noisy labels for the other view. Under conditional independence, mutual irreducibility, and a weakly-useful classifier, consistent learning is possible from unlabeled data.
- Setup: Co-training partitions features into two views and assumes the views are conditionally independent given the class label.The paper considers a known weakly-useful classifier based on view A and unlabeled data.
- Setup: A weakly-useful classifier has nontrivial prediction frequency and satisfies q0(h) + q1(h) < 1.This condition ensures that the induced noisy labels have total contamination below one.
- Consistency result: If the class-conditional distributions of view B are mutually irreducible, an algorithm achieves Bayes-risk consistency using iid unlabeled data.The result is stated as RQ(bfn) → R∗.
- Reduction to label noise: Under the co-training assumption, predictions from view A create a label-noise problem for learning from view B.The induced contamination probabilities are πi = qi(hA), and the sample counts for both noisy labels grow with n.
- Interpretation: The analysis weakens deterministic class-label assumptions to mutual irreducibility; a small amount of labeled data could provide the weakly-useful classifier.This identifies a route to satisfying the co-training premise without requiring deterministic labels.
9 Mutual Irreducibility and Class Probability Estimation
The section connects mutual irreducibility to class probability estimation and shows that posterior extrema characterize the contamination proportions. It also presents class-probability and ROC-based routes to estimating these proportions.
- Mutual irreducibility: Mutual irreducibility holds exactly when the posterior class probability has essential infimum 0 and supremum 1.This connects a distributional condition to observable extremes of the class probability function.
- Mixture proportion estimation: A supremum-norm-consistent posterior estimator produces a consistent estimate of κ∗(P1|P0), although distributional assumptions are then required.The distribution-free estimator of Blanchard et al. (2010) is described as a more general alternative.
- Class probability estimation: Estimates of the contaminated posterior minimum and maximum, together with the contaminated class prior, directly yield estimates of π0 and π1.The contaminated class prior is readily estimated from the fraction of examples carrying the apparent label 1.
- Empirical comparison: The class-probability approach to estimating label-noise proportions is compared experimentally with the ROC-based estimator.The section also relates posterior extrema to formulas for the contamination proportions.
10 Implementation of Estimators
The proposed mixture-proportion estimator uses sample splitting, a universally consistent classifier, and conservative ROC estimates. An alternative estimator uses class-probability extrema from a separately trained model.
- ROC estimator: The ROC-based algorithm splits each contaminated sample, trains a classifier on one portion, and constructs a full ROC.The implementation uses kernel logistic regression with a Gaussian kernel and cross-validated bandwidth and regularization.
- ROC estimator: Conservative ROC estimates are built on held-out data using direct binomial-tail inversion rather than empirical errors plus VC bounds.The procedure uses one-sided exact Clopper–Pearson confidence intervals.
- Class probability estimator: The class-probability estimator trains kernel logistic regression on one split and plugs held-out minimum and maximum posterior estimates into formulas for π0 and π1.This reverses the split emphasis used by the ROC estimator in the reported implementation.
- Implementation choices: The reported splits were 20/80 for the ROC estimator and 80/20 for the class-probability estimator.These ratios were selected because they appeared to give the best results.
11 Experiments
Experiments evaluate mixture-proportion estimation on waveform, digit, and nuclear-particle data, including overlapping classes and realistically contaminated measurements. The ROC method is generally accurate, while the class-probability implementation is unstable and label-noise correction changes nuclear-data evaluation.
- Datasets: The experiments use waveform, MNIST digits, and nuclear-particle data, with the waveform task having overlapping classes and approximately 10% Bayes risk.The waveform binary problem uses two classes, specified noise proportions, and n0 = n1 = 1000.
- ROC estimator: The ROC estimator provides reasonably accurate noise-proportion estimates in the four settings with known true proportions.The results also suggest that mutual irreducibility can be a reasonable practical assumption.
- Class probability estimator: The class-probability estimator is sometimes accurate but sometimes incurs considerable error.Its nuclear-data estimate also reverses the expected ordering π0 < π1.
- Class probability estimator: Class-probability estimates can be conservative or overly optimistic when empirical posterior values fail to cover the full interval [˜ηmin, ˜ηmax].Percentiles far from 0 or 1 indicate over-optimism, while exact endpoint percentiles can indicate conservatism.
- Method comparison: The ROC method’s advantage is attributed to uncertainty quantification and the concavity constraint used when estimating the ROC endpoint slope.The authors state that analogous uncertainty quantification could improve the class-probability method.
- Nuclear classification: Correcting nuclear-test probabilities for label noise changes the ROC from uncorrected errors to corrected errors that account for contaminated test labels.The paper notes that some apparently incorrect predictions are actually correct classifications with erroneous labels.
12 Conclusion
The paper identifies majority-correct labels and mutual irreducibility as conditions supporting consistent classification with asymmetric, unknown label noise. It develops consistent estimators and finds the ROC implementation effective across three datasets, while the class-probability implementation needs further development.
- Conclusion: Consistent classification is argued possible when a majority of labels are correct on average and P0 and P1 are mutually irreducible.Under these conditions, mixture-proportion estimation supplies consistent estimators of the noise proportions.
- Conclusion: Mutual irreducibility is argued necessary for population decontamination satisfying universality, symmetry, continuity, and stability.It can also be interpreted as maximum denoising or maximum separation of the unknown sources.
- Conclusion: The ROC implementation exhibits good accuracy across three datasets, including the nuclear-particle classification problem.The class-probability implementation still requires further development.
A Mixture Proportion Consistency Result
The mutually irreducible decontamination operator is characterized as the unique operator satisfying the stated domain, symmetry, stability, and continuity conditions. The consistency argument establishes its behavior under contamination, while related work requires an additional sample-size growth condition for almost-sure convergence.
- Related consistency result: Blanchard et al.’s almost-sure consistency result additionally requires log max(n0, n1) = o(min(n0, n1)).The supplied discussion notes that this qualification is needed for the original almost-sure convergence argument.
- Operator characterization: Mutual irreducibility makes the decontamination operator’s sources uniquely determined, up to label swapping.The proof uses source symmetry and stability to identify the mutually irreducible pair as the only admissible source pair.
- Consistency range: The operator returns the mutually irreducible solution for contamination levels below the threshold defined by ϵ∗.ϵ∗ is the supremum contamination range over which the defining identity remains satisfied.
B.3 Proof of Theorem 12
The proof maps feasible contaminated-mixture decompositions to a decoupled parameterization, then uses this correspondence to establish a unique mutually irreducible solution. It also derives ROC and learning-theoretic consequences, including convergence in probability for surrogate-loss minimization.
- Existence and uniqueness: Feasible quadruples in the original contamination model correspond one-to-one with feasible quadruples in the decoupled formulation.The feasible proportion region is obtained by mapping a rectangle of decoupled contamination parameters through this correspondence.
- Existence and uniqueness: The unique feasible quadruple with maximal decoupled contamination parameters yields the maximum total label noise π0 + π1 in the original model.The transformation makes total noise strictly increasing in both decoupled contamination parameters.
- Maximal denoising: The maximum total variation separation ∥P1 − P0∥T V is attained by the unique mutually irreducible solution.Subtracting the contamination equations links total variation separation to the total noise maximum.
- Mixture proportion estimation: κ∗(F|H) equals the left derivative of the likelihood-ratio-test ROC at (1, 1), when that derivative exists.The result identifies mixture proportion estimation with the ROC slope at the endpoint under the stated differentiability assumption.
- Surrogate-loss consistency: Rademacher-complexity bounds and regularization control show that surrogate-loss risk converges to its optimum in probability.The proof bounds empirical-process, approximation, and regularization terms, with the estimation-error term vanishing under the stated rate condition.
- Surrogate-loss consistency: The convergence argument requires the estimation error in the noise parameter to vanish relative to λn.The supplied proof explicitly states that |bα − α|/λn must tend to zero, except on a vanishingly small event.