Source-linked AI summary

Tight Majorizations and Convergence Rates of Nuclear Norm Minimization IRLS

Christian Kümmerle, Tomas Masak, Dominik Stöger

arXiv:2608.23765v1cs.LGmath.NAmath.OC

TL;DR

IRLS methods for constrained nuclear norm minimization lacked a clear account of how weight-operator choice affects convergence rates. This paper proves harmonic-mean majorization and optimality, then establishes global linear convergence and a dimension-independent local rate for harmonic-mean IRLS under a null space property. The main scope boundary is that the analysis requires this property, which fails in some settings such as low-rank matrix completion.

  • Problem

    IRLS convergence rates and the role of the weight operator for constrained nuclear norm minimization have remained poorly understood.

  • Method

    The paper analyzes smoothed nuclear-norm quadratic models, proves harmonic-mean majorization and power-mean tightness, and studies IRLS under a Schatten-1 null space property.

  • Results

    Harmonic-mean IRLS has global linear convergence and a locally linear convergence rate independent of matrix dimension, while several weight operators receive global linear-rate guarantees.

  • Takeaways & Limitations

    Harmonic-mean reweighting is supported as a principled, tight choice for obtaining faster local convergence in nuclear-norm IRLS.

  • Takeaways & Limitations

    The convergence analysis requires a suitable null space property, an assumption that does not hold in some important settings such as low-rank matrix completion.

Abstract

from arXiv · show

Iteratively reweighted least squares (IRLS) methods constitute a natural approach to nuclear norm minimization, but their convergence rates and the role of the weight operator have remained poorly understood. This paper establishes sharp convergence rates for IRLS methods for constrained nuclear norm minimization in low-rank recovery. A central ingredient is a new majorization analysis for the smoothed nuclear norm: we prove that the harmonic-mean weight operator defines a valid global quadratic majorizer. Furthermore, we show that this weight operator is optimal within the family of power-mean weights, clarifying why it improves over classical one-sided reweighting schemes that use only row- or column-space information. Under a Schatten-1 null space property, we prove global linear convergence of IRLS algorithms using a variety of weight operators, including the harmonic-mean weights. For IRLS with harmonic-mean weights, we prove a dimension-independent, locally linear convergence rate. We provide a counterexample showing that this dimension-independent local rate cannot in general be obtained for IRLS algorithms using one-sided weight operators, which predominate in the literature. Numerical experiments corroborate the theoretical results and illustrate the practical advantage of harmonic-mean reweighting across square, rectangular, and adversarially initialized recovery problems.

1 Introduction

Nuclear norm minimization is statistically mature but computationally challenging, while IRLS design and convergence for spectral objectives remain insufficiently understood. This paper establishes majorization, optimality, and convergence-rate results centered on harmonic-mean reweighting.

  • Motivation: Nuclear norm minimization recovers low-rank matrices from underdetermined measurements, but specialized solvers can have sublinear rates and require repeated full singular value decompositions.Under standard random sensing, m ≳ r⋆(d1+d2) Gaussian rank-one measurements suffice up to constant factors.
  • Motivation: IRLS replaces the nonsmooth spectral objective with a sequence of weighted least-squares problems, but its weight operator and convergence guarantees remain unsettled.Most existing spectral IRLS methods use one-sided operators encoding only row- or column-space information.
  • Contributions: IRLS objective gaps decrease linearly for several weight operators, with harmonic-mean reweighting achieving a notably improved linear rate in the representative rank-two experiment.The experiment uses a 140×140 rank-2 ground truth and m = 2520 rank-one measurements.
  • Contributions: The harmonic-mean weight operator globally majorizes the smoothed nuclear norm and is the tightest member of the globally valid power-mean family.The majorization proof uses a spectral comparison based on a Sylvester-equation characterization and iterative pinching.
  • Contributions: Under an order-r null space property, IRLS has global linear convergence, while harmonic-mean IRLS has a local linear rate independent of d1 and d2.The global contraction factors retain matrix-dimension dependence; the local result applies after entering a specified neighborhood of the ground truth.
  • Contributions: The paper extends a workshop result by adding proofs and the majorization, optimality, and fast local-rate results developed here.The preliminary workshop version contained only a global linear convergence result.

2 Related Work

Prior low-rank IRLS work established global convergence or limiting-iterate guarantees without convergence rates, while related methods and sparse IRLS provide partial precedents. The paper positions harmonic-mean nuclear-norm IRLS as addressing this gap with a distinct majorization and rate analysis.

  • Low-Rank Matrix Recovery: Nuclear norm minimization is a convex estimator that does not fix rank and can support global linear-rate algorithms under suitable sharpness or null-space conditions.Existing first-order schemes are typically analyzed with sublinear rates, whereas the paper studies global linear rates for IRLS variants.
  • Low-Rank IRLS: Foundational low-rank IRLS methods used one-sided reweighting and proved global convergence under the null space property, but did not establish convergence rates.Later work provided error bounds or stability guarantees for limiting iterates, while a claimed local linear-rate analysis was not substantiated.
  • Low-Rank IRLS: Harmonic-mean IRLS for Schatten-p quasi-norms achieved local superlinear convergence for 0 < p < 1, but that result does not apply at the nuclear-norm endpoint p = 1.The present paper uses a p = 1 variant with a different smoothing-parameter schedule and supplies a new majorization proof.
  • Related Spectral Methods: Recursive feature machines provide a related derivation of one-sided IRLS methods for log-determinant and nuclear norm minimization in linear models.This connection is presented as a related line of work rather than as a convergence-rate result for the present algorithms.
  • IRLS for Sparse Recovery: Sparse IRLS benefits from separability, which reduces weight-operator design freedom compared with spectral IRLS and has no analogue of the matrix weight core’s off-diagonal entries.The paper links tight quadratic-model design to achieving comparable convergence-rate results for MatrixIRLS.

3 The MatrixIRLS Algorithm and Its Majorization Properties

The MatrixIRLS framework smooths the nuclear norm and minimizes quadratic reweighted least-squares models whose design depends on spectral information. The harmonic-mean operator provides global majorization and is the tightest power-mean choice, supporting sharper convergence analysis.

  • 3.1 IRLS for Low-Rank Recovery and Basic Properties: IRLS mitigates nuclear-norm nonsmoothness by minimizing ε-smoothed objectives through quadratic models updated with the reference matrix, smoothing parameter, and weight operator.The quadratic model is minimized as a reweighted least-squares objective once the reference point and smoothing parameter are fixed.
  • 3.1 IRLS for Low-Rank Recovery and Basic Properties: Left-sided, right-sided, and harmonic-mean operators share diagonal entries but differ in their off-diagonal dependence on row and column singular values.The harmonic-mean core uses both row- and column-associated singular values, whereas one-sided cores use only one side.
  • 3.2 Majorization of Quadratic Model Implied by Harmonic-Mean Weights: The quadratic model must majorize the smoothed objective pointwise to connect weighted least-squares updates with objective progress and convergence-rate analysis.Algebraic properties of the weight construction alone do not establish this progress relationship.
  • 3.2 Majorization of Quadratic Model Implied by Harmonic-Mean Weights: The harmonic-mean quadratic model globally majorizes the ε-smoothed nuclear norm, resolving the previously missing result for this operator.Existing proofs had established majorization for one-sided models but not for harmonic-mean models.
  • 3.4 Optimality of Harmonic-Mean Weight Operator: q ≥ −1 exactly characterizes power-mean models that majorize the smoothed objective, making the harmonic mean the tightest valid choice in that family.Any power-mean operator smaller in Loewner order violates majorization locally; the harmonic-mean model has the smallest pointwise gap among majorizing power-mean variants.
  • 3.4 Optimality of Harmonic-Mean Weight Operator: The harmonic-mean guarantees discussed here concern convex nuclear-norm minimization; for nonconvex Schatten-p quasi-norms with 0 < p < 1, its quadratic model is not a valid global majorizer.That nonconvex setting lies outside the scope of the paper.

4 Linear Convergence Rates of IRLS for Nuclear Norm Minimization

The section establishes global linear convergence for IRLS under a null space property, largely independent of the admissible weight operator, and a sharper dimension-independent local rate for harmonic-mean weighting. It also explains why this local rate is unavailable for one-sided operators in general.

  • Global linear convergence: Theorems 3 and 4 provide the first global linear convergence rates for IRLS, covering low-rank and approximately low-rank ground truths under null space property assumptions.The global results apply across an admissible family of weight operators, with the choice affecting only a convergence constant.
  • Global linear convergence: The global convergence factor retains a 1/d dependence for all admissible weight operators, so no scheme is singled out by the global rates.The weight-dependent constant lies between 1 and 3, while the dimension dependence remains in the global factor.
  • Fast local linear convergence: Theorem 5 establishes a fast local linear rate for harmonic-mean MatrixIRLS that is independent of the ambient dimension d.For ηr = 1/10, the local decrease factor is approximately 1 − 0.0037, and the bound becomes sharper than the global rate when d > 273 under the stated additional assumption.
  • Fast local linear convergence: The dimension-independent local rate is driven by a sharp upper bound on the weighted quadratic form available for harmonic-mean weights.This bound is tied to the quadratic model used by IRLS and supports the local convergence analysis.
  • Weight-operator scope: Fast dimension-free local convergence also remains available for q-power means with q ∈ [−1, 0), whereas the paper’s results do not establish it for q ∈ [0, ∞].Numerical experiments suggest that the fast local rate should not be expected for larger-q power means, including arithmetic means.
  • One-sided weighting: A counterexample shows that one-sided weight operators cannot generally achieve the same dimension-independent fast local rate, with analogous behavior for right-sided operators.For sufficiently large dimensions relative to rank, the weighted quadratic form admits a dimension-dependent lower bound for a left-sided operator.

5 Numerical Experiments

The experiments compare IRLS weight operators, dimensions, initialization schemes, rectangularity, and smoothing updates. Harmonic-mean reweighting consistently delivers the strongest convergence behavior, while alternatives become more dimension-sensitive or sampling-sensitive.

  • Linear Convergence Rate Factors: IRLS variants all show linear convergence, but harmonic-mean weights converge fastest across the considered dimensions, followed by arithmetic and then one-sided variants.For d = 30, harmonic-mean IRLS reaches an error threshold of 10^-5 after 16 iterations.
  • Linear Convergence Rate Factors: Harmonic-mean decrease factors remain stable or improve after initialization, whereas left-sided, right-sided, and arithmetic variants deteriorate as dimension grows.This pattern is consistent with a sharper local linear rate for harmonic-mean IRLS than for the other variants.
  • Rectangular Recovery: In tall rectangular recovery, harmonic-mean IRLS reaches 10^-5 after 19, 28, and 38 iterations for d = 20, 40, and 80, respectively.The best alternative, right-sided IRLS, requires 40, 113, and 324 iterations, while left-sided IRLS fails to reach the threshold for d = 80 within 500 iterations.
  • Rectangular Recovery: For d = 80 rectangular recovery, harmonic-mean limiting decrease factors are about 0.62, compared with 0.96 to 0.98 for the other variants.The widening gap suggests that the other variants do not share the same fast local linear rate.
  • Adversarial Initialization: Under adversarial initialization, 1/(1 − µ(0)) grows approximately linearly with dimension for all IRLS variants, making the first algorithmic decrease slower at larger d.The reported fit is d/7 − 3 for the setup considered.
  • Adversarial Initialization: For harmonic-mean IRLS, terminal 1/(1 − µ(k)) stays below 3.0 through d ∈ {100, 110, 120, 130, 140}, while other variants exceed 40 or 50.The harmonic-mean transition is therefore transient under the extended initialization, whereas alternatives retain stronger dimension dependence.
  • Smoothing Parameter Update Rules: With smoothing-parameter updates, the ℓ∞-tail rule fails at C = 3 but converges at C = 4, while harmonic-mean convergence remains faster across update rules.The ℓ1-tail and ℓ2-tail variants track nuclear norm minimization recovery closely, whereas ℓ∞-tail recovery becomes consistent only around C = 2.9, 3.4, and 3.7 as d increases.
  • Smoothing Parameter Update Rules: IRLS can achieve order-of-magnitude speedups over generic SDP solvers, especially farther above the sample-complexity phase transition.The comparison uses SCS solve times of 25.37 and 29.32 seconds.

6 Conclusion

The paper resolves central design and convergence questions for IRLS nuclear norm minimization, establishing majorization and linear-rate results for multiple weight operators. Its analysis and experiments identify harmonic-mean reweighting as especially effective, while delimiting assumptions and open extensions.

  • The paper frames IRLS design around weight operators and provides convergence-rate analysis for nuclear norm minimization.
  • Global linear convergence holds from any initialization for several weight operators, including traditional and harmonic-mean choices.
  • A dimension-free linear rate applies after iterates enter a neighborhood of the low-rank solution when the correct weight operator is used.
  • Experiments across square and rectangular instances show faster empirical linear rates for harmonic-mean reweighting than for one-sided or arithmetic-mean variants.
  • The convergence theory depends on a suitable null space property, and the certified local neighborhood still depends on dimension.
  • Majorization analysis currently covers the smoothed nuclear norm, leaving other spectral objectives and computationally economical updates for future work.

A.1.1 Proof of Lemma 2 via Iterative Pinching

The proof of the harmonic-mean lower bound combines a Sylvester-equation characterization with iterative pinching and operator majorization. This transforms noncommuting matrix structure into an explicitly solvable diagonal setting.

  • The proof first characterizes the harmonic-mean weight operator through a Sylvester equation.
  • Unlike one-sided weights, harmonic-mean weights require a more involved lower-bound argument because their weighted inner product lacks a simple separable form.
  • The harmonic-mean operator solves LUW + WLV = 2Z uniquely because LU and LV are positive definite.
  • Iterative pinching replaces LU and LV with simpler matrices aligned to Z's singular vectors.
  • Once both matrices are diagonal in the singular-vector coordinates, the Sylvester equation can be solved explicitly to derive the lower bound.
  • The abstract majorization argument uses orthogonal projections, self-adjoint positive-definite operators, and matching projected actions.

A.2.2 Proofs of Lemmas 8 and 9 (Second Order Necessary Condition for Majorization)

The proof of the second-order majorization condition reduces perturbations to a two-dimensional singular block and applies a spectral-function Hessian formula. Local majorization then imposes a necessary constraint on weight-core entries.

  • The proof uses symmetrization and antisymmetrization operators to evaluate the Hessian along the skew-symmetric perturbation.
  • A Hessian formula for spectral functions supplies the second-order expansion when the relevant singular values exceed ε.
  • The argument reduces perturbations involving two singular values to a 2 × 2 singular block.
  • The resulting necessary condition applies to distinct indices i and j with σi > ε and σj > ε, without assumptions on remaining singular values.
  • The reduction cancels all unchanged singular-value contributions, leaving only σi and σj in the local analysis.
  • The constructed perturbation has zero first derivative at the origin, so local majorization forces its second derivative to be nonpositive.

A.2.3 Proof of Theorem 2

Theorem 2 proves that harmonic-mean weights are optimal within the power-mean family for majorizing the smoothed nuclear norm. The proof establishes failure of local, and therefore global, majorization below the harmonic-mean threshold.

  • Thus, the harmonic mean provides the tightest majorization among power-mean weight operators.
  • For power-mean weights, majorization holds for every exponent q ≥ −1 by harmonic-mean majorization and monotonicity of power means.
  • When q < −1, distinct singular values above ε violate the necessary second-order majorization condition.
  • For q < −1, the quadratic model fails to majorize the smoothed objective locally around X.
  • Local failure implies that global majorization also fails for q < −1.

A.3.1 Preliminaries and General Proof Strategy

The proof develops the inequalities needed to relate the smoothed nuclear-norm objective to the nuclear-norm error, then uses quadratic majorization and weight-operator norm bounds to analyze IRLS decrease. Harmonic, power-mean, and one-sided weight operators are treated within this framework.

  • Objective bounds: Lemma 11 relates Jε(X)−∥X⋆∥∗ to rank-r approximation and null-space-property quantities under A(X⋆)=A(X).The argument covers 0≤ε≤βr(X)∗/d, with ε=0 interpreted as J0=∥·∥∗.
  • Majorization strategy: The proof optimizes a scalar step t after separately bounding the two terms governing the quadratic-model decrease.Global and local convergence use the same estimate for term (a), but the local proof obtains a sharper estimate for term (b).
  • Majorization strategy: Quadratic majorization permits decrease estimates for harmonic, q-power-mean, and one-sided weight operators.The harmonic case corresponds to q=−1; power means are covered for q∈[−1,∞].
  • Weight-operator bounds: Norm bounds for Hadamard products control quadratic forms involving the weight operator core matrix and the error X⋆−X.These bounds are established through auxiliary lemmas and then applied to the quadratic terms in the convergence proof.
  • Weight-operator bounds: The harmonic core matrix has entries 1/(λi+λj) and is positive semidefinite, enabling direct application of the general unitarily invariant norm bound.Its positive semidefiniteness follows from a Gram-matrix representation.
  • Weight-operator bounds: For q∈[−∞,1], the weight-dependent constant is cq=1, including the harmonic case q=−1; for q>1, it satisfies 1<cq≤3.The resulting norm bound is λmax|||B||| for q∈[−∞,1] and 3λmax|||B||| for q∈(1,∞].

A.3.3 Proof of Theorem 3

Theorem 3 is proved by combining the per-iteration decrease estimate with the null-space-property-based objective bounds. This yields global linear convergence when the ground truth X⋆ is exactly low-rank.

  • Decrease estimate: Proposition 2 quantifies the decrease of Jεk across an IRLS iteration for admissible weight operators under the NSP.It applies with arbitrary positive definite initial weights and rank estimate er=r while εk>0.
  • Decrease estimate: The proof applies quadratic majorization and the optimality of X(k+1) to compare the next iterate with points along X(k)+t(X⋆−X(k)).The scalar t is chosen to minimize the resulting upper bound.
  • Global convergence: The weight-dependent constant is cq=1 for one-sided, harmonic-mean, and power-mean weights with q∈[−1,1].For power means with q∈(1,∞], the constant satisfies 1<cq≤3.
  • Global convergence: Theorem 3 establishes global linear convergence for exactly low-rank X⋆ by chaining the per-iteration decrease estimate.The proof uses ηr<3/5 to obtain the stated contraction condition.

A.3.4 Proof of Theorem 4

Theorem 4 extends the global linear convergence analysis to approximately low-rank ground truths. The proof handles the transition to a rank-r approximation and derives bounds after a threshold iteration.

  • Approximate low-rank case: The proof of Theorem 4 requires a more involved argument because X⋆ is only approximately low-rank.It introduces a threshold index ˆk associated with the tail quantity βr(X⋆)∗.
  • Pre-threshold convergence: Before ˆk, chaining Proposition 2 yields a geometric bound on the smoothed objective gap.When βr(X⋆)∗=0, ˆk=∞ and the stronger γ=3/4 case applies throughout.
  • Post-threshold convergence: After ˆk, the proof combines null-space-property estimates, monotonicity, and Eckart–Young bounds to control the approximation tail.These steps establish the inequality corresponding to the post-threshold convergence estimate.
  • Post-threshold convergence: The post-threshold estimate requires ηr<1/3.The proof uses this assumption to ensure positivity before taking logarithms and deriving the upper bound on ˆk.

A.4 Proofs of Theorems 5 and 6 (Local Linear Convergence and Counterexample)

The local analysis sharpens the quadratic-term bound for harmonic and related power-mean weights near a rank-r ground truth, yielding a faster rate. A counterexample shows that one-sided weights do not generally admit the same dimension-independent local rate.

  • Local convergence: Global rate proofs use quadratic-term estimates scaling with 1/ε, whereas the local analysis removes this ε-dependence near X⋆.The sharper estimate applies to harmonic-mean and related q-power-mean weight operators.
  • One-sided counterexample: A counterexample shows that the dimension-independent local rate cannot generally be obtained for left- or right-sided weight operators.This failure persists even in arbitrarily small neighborhoods of X⋆ relative to the stated local neighborhood.
  • Local convergence: The local estimate is established under an NSP of order r with ηr≤3/5 and closeness of the iterate to the rank-r ground truth.The proof uses order-one NSP control, Weyl’s inequality, and singular-subspace estimates.
  • Local convergence: The improved estimate yields a convergence-rate improvement by a factor of d compared with Theorem 3.The improvement follows because the relevant local bound no longer depends on ε.
  • Local convergence: For q=−1 and ηr=1/10, the local quadratic-form bound is approximately 27.20∥X−X⋆∥∗.For ηr=1/2, the corresponding bound is 298∥X−X⋆∥∗.

A.4.2 Proof of Theorem 5 (Dimension-Free Fast Linear Rate of MatrixIRLS)

Theorem 5 proves that MatrixIRLS with harmonic-mean weights converges linearly at a dimension-free rate once iterates enter a neighborhood of the rank-r solution, under NSP-based assumptions.

  • Assumptions and contraction: Under an NSP of order r, Proposition 4 provides the one-step contraction needed for the local analysis.The theorem assumes 0 < ηr < 3/5 and uses fixed q-power mean weights with q ∈ [−1, 0), including q = −1.
  • Harmonic-mean specialization: For q = −1, the harmonic-mean choice, the general contraction estimate specializes to the dimension-free rate used in Theorem 5.The proof identifies q = −1 with harmonic-mean reweighting and verifies the associated contraction factor.
  • Local convergence: Theorem 5 establishes a dimension-free local linear convergence rate for MatrixIRLS with harmonic-mean weight operators.The result follows from the one-step local contraction in Proposition 4 and an induction argument.
  • Proof structure: The induction proof propagates both the error bound and the local-neighborhood condition from iteration k to k + 1.It applies Proposition 4 at each step and uses the induction hypothesis, NSP estimates, and the rank-r assumption.

A.4.3 Proof of Theorem 6 (Counterexample for One-Sided IRLS)

Theorem 6 constructs a counterexample showing that one-sided IRLS can have dimension-dependent local behavior, unlike the dimension-free harmonic-mean result.

  • Counterexample consequence: The next one-sided IRLS iterate cannot reduce nuclear norm error by more than a factor of (1 − cr/d) for some constant c > 0.Thus the local contraction deteriorates with the ambient dimension d in the constructed example.
  • Right-sided analogue: The same obstruction applies to right-sided weight operators after transposing the underlying matrices.The argument is presented for left-sided weights and transfers by matrix transposition.
  • Construction: The construction uses d ≥ 220r and a measurement operator whose kernel is spanned by a carefully chosen perturbation Δ.The perturbation is built from orthogonal rank-one terms and the operator is defined using an orthonormal basis of Δ’s Frobenius-orthogonal complement.
  • One-sided weighting mechanism: For left-sided weights, directions orthogonal to the iterate’s column space receive weight 1/ε, producing a dimension-dependent lower bound on the quadratic form.The lower bound is ⟨X − X⋆, W_X,ε(X − X⋆)⟩F ≥ d^2/(20r) ∥X − X⋆∥∗.

B.3 Prior Art of Majorization Proofs for Low-Rank IRLS Algorithms

Prior work established global majorization for one-sided IRLS quadratic models, but corresponding harmonic-mean majorization arguments had not been available.

  • Established one-sided results: Proposition 5 restates that left- and right-sided quadratic models globally majorize the ε-smoothed nuclear norm.These majorization results were established in earlier work and are used in the IRLS analysis.
  • Role in IRLS analysis: The one-sided majorization property implies the required model inequality used by IRLS updates.This follows by choosing the current and preceding iterates as the arguments in the global majorization inequality.
  • Gap addressed by the paper: Existing proof strategies for one-sided weights do not extend directly to harmonic-mean weights.The paper notes that these strategies rely on representing the operator as left- or right-matrix multiplication, which is unavailable for harmonic weights.

B.3.1 Proof of Proposition 5 Using Concavity Arguments

The one-sided majorization proof uses concavity after a matrix-variable transformation and an auxiliary variational formulation, while harmonic weights require a different strategy.

  • Concavity argument: Concavity of the spectral trace function Gε yields a linear upper bound that proves one-sided quadratic majorization.The proof applies concavity to ZZ⊤ or Z⊤Z after expressing the smoothed nuclear norm through a spectral trace function.
  • Scope of the proof: The concavity-based proof strategy remains specific to one-sided operators because harmonic weights cannot be written as simple left- or right-matrix multiplication.The paper therefore treats harmonic-mean majorization separately.
  • Auxiliary formulation: For left-sided weights, the auxiliary minimizer has diagonal entries 1/max(σi(Z), ε), matching the one-sided weight construction.The scalar minimization selects the unconstrained critical point when σi(Z) ≥ ε and the boundary ε^-1 otherwise.
  • Recovery of the surrogate: At the auxiliary minimizer, scalar contributions reproduce the smoothed nuclear-norm terms, with an additional ε/2 contribution for extra zero eigenvalue directions.The extra constant occurs only when the row dimension exceeds the column dimension.
  • From auxiliary bound to majorization: The resulting operator representation converts the auxiliary bound into the quadratic model majorization inequality for left-sided weights, with the right-sided case following by transposition.The proof uses the identified left-sided operator action and then applies the same reasoning to transposed matrices.

B.3.4 Challenges for Harmonic-Mean Majorization Proofs

The proposed variational route to harmonic-mean majorization has substantive gaps, including unjustified matrix differentiation and unsupported singular-vector alignment. The paper therefore uses a direct weighted-inner-product lower bound to establish the needed inequality, while restricting its focus to p = 1.

  • The claimed variational minimizer is aligned with the singular vectors of Z, but its optimality is not completely proved.The missing proof concerns minimization of Φε(Z, ·) with respect to the auxiliary matrix M.
  • The harmonic-mean weight couples left and right singular spaces through an inverse Kronecker sum, preventing scalar decoupling in the auxiliary minimization.This differs from one-sided constructions, which can be handled after a change of basis.
  • Matrix square-root differentiation cannot use the scalar chain rule for noncommuting perturbations, invalidating the cited derivation of the stationarity equation.The correct Fréchet derivative involves divided differences or an equivalent Sylvester-type operator.
  • The commutation relation used in the proof does not by itself force BB⊤ to be diagonal without additional nondegeneracy assumptions.Consequently, the later singular-vector alignment argument is unsupported as stated.
  • The paper replaces the incomplete variational argument with a direct lower bound for ⟨W X,ε(Z), Z⟩F to obtain the essential majorization inequality.This approach avoids relying exclusively on a variational envelope.
  • The authors focus on the nuclear-norm case p = 1 and leave majorization for nonconvex Schatten-p surrogates with 0 < p < 1 to future work.They believe the cited statement is correct for p = 1 but likely incorrect for 0 < p < 1.
Loading 2608.23765v1…