Source-linked AI summary

Quantum Federated Learning Based on Bures--Uhlmann Geometry for Heterogeneous Noisy Clients

Haruki Emori, Masaki Uchihara, Yuuki Tokunaga

arXiv:2608.28379v1quant-phcs.LG

TL;DR

Noisy quantum federated learning requires geometry and aggregation rules that account for mixed states and heterogeneous client reliability. The paper uses mixed-state Bures–Uhlmann geometry for local preconditioning and precision-weighted aggregation, with theoretical variance guarantees and strong emulator results. On a trapped-ion emulator, QFedQGT achieved the best round-averaged accuracy across all four heterogeneity conditions and outperformed federated averaging overall.

  • Problem

    Quantum federated learning must aggregate updates from noisy, heterogeneous devices, while pure-state and diagonal QGT approaches discard mixed-state and incompatibility information.

  • Method

    The method uses the mixed-state Bures metric for local natural-gradient preconditioning and mean Uhlmann curvature for achievable-precision client weighting, with exact block-diagonal estimation.

  • Results

    QFedQGT achieved the highest round-averaged accuracy in all four conditions, averaging 0.815 versus 0.790 for QFedAvg and 0.759 for QFedFisher.

  • Takeaways & Limitations

    The proposed aggregation reduces residual variance under device heterogeneity and maintains higher accuracy and worst-round performance than the evaluated baselines.

Abstract

from arXiv · show

Quantum federated learning enables collaborative model training across quantum devices without sharing raw data, and it faces the data and hardware heterogeneity inherent to noisy quantum devices. Utilizing the quantum geometric tensor is a natural remedy, yet pure-state approaches and diagonal approximations discard the correlations that encode parameter incompatibility. To address this, we extend the parameter-space geometry to the mixed states that noisy clients actually prepare. The real part of the resulting mixed-state geometric tensor is the Bures metric, which measures how fast the physical state changes under parameter variation, and the imaginary part is the mean Uhlmann curvature, which quantifies the incompatibility of estimating multiple parameters simultaneously. Accordingly, we employ the Bures metric as a local preconditioner and use the mean Uhlmann curvature to develop an achievable-precision aggregation rule that dynamically down-weights unreliable clients. Furthermore, we establish theoretical guarantees by proving a convergence theorem and a variance-dominance proposition. Empirical evaluations on a trapped-ion quantum emulator demonstrate that the proposed method maintains high accuracy across diverse device-heterogeneity conditions and outperforms standard federated averaging, whose accuracy degrades under strong noise.

I. INTRODUCTION

Federated learning keeps private data local but faces non-IID data, communication costs, and, in QFL, heterogeneous noisy quantum hardware. The paper addresses these limitations by replacing pure-state geometry with mixed-state Bures–Uhlmann geometry for local optimization and aggregation.

  • Motivation: Federated learning trains a shared model from siloed private data by exchanging parameter updates rather than raw records.Its structural difficulties include non-IID client data and costly update communication.
  • Motivation: Quantum federated learning uses client variational quantum circuits, quantum-state data encoding, expectation-value measurements, and server-side parameter aggregation.Quantum resources motivate QFL through high-dimensional state spaces and potentially secure quantum communication.
  • Problem: Different gate errors, coherence times, and shot budgets make otherwise similar quantum clients produce updates with different reliability.Under depolarizing rates of 0.4–0.6, naive averaging can be strongly affected by unreliable devices.
  • Problem: Pure-state QGT preconditioning reproduces the quantum natural gradient, while diagonal metric approximations discard off-diagonal correlations carrying incompatibility information in mixed states.These limitations motivate changing the geometry rather than merely inserting pure-state QGT into federated learning.
  • Contribution: The mixed-state QGT uses the Bures metric for local preconditioning and mean Uhlmann curvature for achievable-precision aggregation.The Bures metric interpolates between Fubini–Study and classical Fisher geometries, while Uhlmann curvature captures multiparameter incompatibility.
  • Contribution: The proposed framework also provides an exact block-diagonal estimator with O(L) overhead, a depolarizing correction, convergence analysis, and a variance-dominance proposition.These contributions target scalable estimation and theoretical guarantees for heterogeneous noisy clients.

C. Bures metric and mean Uhlmann curvature for mixed states

The paper extends quantum parameter-space geometry from pure states to the mixed states prepared by noisy clients. Its Bures and mean Uhlmann components have distinct optimization and statistical roles.

  • Mixed-state geometry: Noisy hardware causes client k to prepare a mixed state ρk(θ) rather than an ideal pure state.The mixed-state geometry is constructed using symmetric logarithmic derivatives.
  • Mixed-state geometry: The mixed-state QGT has the Bures metric as its real part and mean Uhlmann curvature as its imaginary part.These are the mixed-state counterparts of the pure-state QGT components.
  • Bures metric: The Bures metric is minimal among monotone density-matrix metrics, reduces to Fubini–Study geometry for pure states, and equals one quarter of classical Fisher information for commuting families.At O(1) noise it coincides with neither limiting geometry.
  • Mean Uhlmann curvature: Mean Uhlmann curvature vanishes for pure or commuting families and is nonzero when [ρk, ∂iρk] ≠ 0.It measures the incompatibility of estimating multiple parameters and has no classical analogue.
  • Separate roles: The Bures component determines the real-parameter descent direction, whereas the Uhlmann component quantifies the excess covariance associated with incompatible estimation.The imaginary component is therefore excluded from the gradient step and used for client reliability instead.
  • Algorithm: The proposed scheme combines a Bures-metric mixed-state natural-gradient update with achievable-precision server weights, estimated without a diagonal approximation.Both ingredients use an exact block-diagonal procedure.

A. Local update and its reparametrization invariance

QFedQGT uses Bures-preconditioned local updates and precision-based server aggregation to account for heterogeneous client noise. The weighting rule is reparametrization-consistent, variance-dominant over QFedAvg, and adaptable to reliability differences.

  • Local update: Clients estimate gradients with parameter-shift measurements, model the noise as unbiased fluctuation with covariance Vk, and apply a regularized block-diagonal Bures update.Metric normalization by its mean diagonal makes the local step scale invariant.
  • Reparametrization invariance: The local update is invariant to first order under smooth invertible reparametrization, unlike the Euclidean gradient.This lets clients with different circuit parametrizations or encodings combine their updates consistently.
  • Precision aggregation: Client update precision incorporates gradient covariance, shot noise, depolarizing attenuation, and mean Uhlmann incompatibility.The incompatibility correction is tied to the quantum Cramér–Rao and Holevo bounds.
  • Precision aggregation: The server solves a generalized least-squares combination problem on the tangent space and applies a damped first-order Bures barycenter update.The matrices Ak are selected to preserve unbiasedness for the consensus direction while minimizing aggregated variance.
  • Guarantees: The aggregated variance of the proposed rule never exceeds QFedAvg’s, with equality when client precisions are proportional to data shares.Under device homogeneity, the rule degenerates to QFedAvg; under heterogeneity, it reduces aggregated variance.
  • Implementation: The scalable scalar weighting increasingly favors high-shot, low-noise, compatible clients as heterogeneity grows.A nonzero floor prevents any client from being discarded, while the floor adapts between reliable and strongly heterogeneous pools.

C. Shared model with quantum feature map and Pauli-expectation readout

The shared model is a dressed variational classifier combining a quantum feature map with a classical readout. Bures geometry acts only on the quantum parameters, while all methods use the same model and readout.

  • Quantum feature map: The feature map re-uploads data through Ry rotations, applies trainable RzRyRz blocks, and uses linear CNOT entanglers across L layers.The quantum parameter count is dq = 3nL.
  • Pauli-expectation readout: The model measures expectation values of a compact Pauli-observable set rather than all computational-basis probabilities.For n = 6, the feature count is Mf = 33 and the observables require only three measurement settings.
  • Classical readout: A trainable classical readout head maps the measured features to class scores using parameters θh = (W1, b1, W2, b2).The single tanh layer avoids the approximate 0.85 ceiling associated with a linear readout on the ten-class digit task.
  • Method comparison: The Bures geometry is applied only to the quantum block θq, while the classical head remains Euclidean and identical across QFedAvg, QFedFisher, and QFedQGT.Thus performance differences are attributed to geometry and aggregation rather than model changes.

D. Hardware estimation without diagonal truncation

The protocol estimates mixed-state geometry without diagonal truncation by retaining exact intra-layer correlations at linear-in-depth measurement cost. Depolarizing noise attenuates client precision, motivating reliability-aware aggregation that suppresses noisy updates.

  • Hardware estimation: The diagonal approximation discards off-diagonal elements that encode Uhlmann incompatibility, so the method retains exact intra-layer blocks instead.The neglected inter-layer blocks are bounded separately.
  • Hardware estimation: Layerwise Bures blocks are estimated from parameter-shifted fidelities, preserving intra-layer correlations at O(L) measurement cost.This avoids the O(d_q^2) cost of estimating the full matrix.
  • Noise correction: For global depolarizing noise, G_k = κ(p_k)G_Pure, with κ → 1 as p → 0 and κ ≈ 1 − p for N ≫ 1.The scalar correction leaves the natural-gradient direction unchanged while attenuating achievable precision.
  • Noise correction: Noisy clients are suppressed by κ(p_k)^2(1 − p_k)^2 and their shot budgets M_k, making depolarizing rate the dominant reliability determinant.The suppression is approximately quartic in (1 − p) for N ≫ 1, while dependence on M_k is linear.
  • Noise correction: Local depolarizing or coherent errors require the full matrix-valued correction because U_k ≠ 0.For the coherent miscalibration used experimentally, the method uses a leading-order proxy r_k ≃ p^2.

F. Protocol

The protocol combines local Bures preconditioning with precision-weighted aggregation under assumptions controlling smoothness, noise, heterogeneity, metric conditioning, and descent alignment. Its convergence analysis separates geometric decay from a variance-driven error floor, while the aggregation variance is no larger than QFedAvg's.

  • Protocol: Each client performs fixed local minibatch steps, estimates gradients and Bures blocks, computes precision-based weights, and sends updated parameters to the server.Client-side randomness is deterministically seeded by round and client for reproducibility.
  • Protocol: QFedAvg, QFedFisher, and QFedQGT share the same VQC and readout, differing only in local optimization and aggregation weighting.This isolates geometry and aggregation as the sources of performance differences.
  • Assumptions: The analysis assumes smoothness, Polyak–Łojasiewicz structure, unbiased bounded gradient noise, bounded heterogeneity, metric regularity, and descent alignment.The heterogeneity parameter ζ is zero in the IID case, while metric regularity keeps the preconditioner well conditioned.
  • Assumptions: The assumptions are motivated by finite trigonometric VQC objectives, unbiased parameter-shift gradients, independent hardware executions, and controlled Dirichlet heterogeneity.These arguments support the stated regime rather than removing the assumptions from the theorem.
  • Convergence analysis: Theorem 1 gives geometric convergence with round complexity T_ϵ ≃ log(1/ϵ)/(cηµ), while the Bures preconditioner improves conditioning and enlarges the admissible step.Precision weighting raises the alignment constant c by removing mis-aligned unreliable contributions.
  • Variance analysis: The residual error floor decreases as total achievable precision grows, and Proposition 1 states that the proposed aggregate variance never exceeds QFedAvg's.Under device homogeneity the rule reduces to QFedAvg; under heterogeneity it reduces aggregated variance.

V. NUMERICAL EXPERIMENTS

On MNIST-10 with nine heterogeneous clients, QFedQGT achieved the strongest accuracy and robustness across four emulator conditions, with its advantage largest under strong noise and scarce shots. Its Bures–Uhlmann weighting reduced variance and avoided the noise sensitivity observed for QFedFisher.

  • Accuracy and convergence: QFedQGT attained the highest round-averaged accuracy in every condition, averaging 0.815 versus 0.790 for QFedAvg and 0.759 for QFedFisher.Its final-round average was 0.859, compared with 0.832 for QFedAvg and 0.823 for QFedFisher.
  • Robustness and stability: QFedQGT’s worst evaluated round ranged from 0.733 to 0.753 and exceeded both baselines in every condition.Its round-to-round standard deviation was also smallest in three of the four conditions.
  • Heterogeneity dependence: The QFedQGT advantage was largest under strong noise and scarce shots, consistent with trust ratios falling to 1.1 × 10^-3 in condition (a) and increasing with shot budget.Increasing depolarizing rates reduced client trust by more than an order of magnitude, whereas a fivefold shot increase recovered about a factor of five.
  • Accuracy and convergence: 0.040, 0.036, 0.015 and 0.009 were QFedQGT’s round-averaged advantages over QFedAvg in conditions (a)–(d), respectively.The advantage decreased monotonically as devices became more reliable.
  • Robustness and stability: The proposed aggregation’s variance was at most that of the best client alone and was never inflated by bad clients, unlike QFedAvg.This supports the observed lower fluctuations and the variance-dominance proposition.
  • Comparison with QFedFisher: QFedFisher averaged 0.759 across conditions and remained below plain averaging in every condition because its noisy-client weighting inverted the achievable-precision prescription.In condition (a), QFedFisher assigned noisy clients 1.46 times the reliable-client weight, while QFedQGT assigned 0.042 times.

C. Contribution of the individual components

The paper isolates the contributions of mixed-state geometry, block structure, and incompatibility-aware aggregation, then evaluates QFedQGT across heterogeneous trapped-ion conditions.

  • Component comparisons: Diagonal truncation lowers accuracy and removes the Uhlmann incompatibility signal, while the block-diagonal estimator remains close to the full metric at O(L) cost.The block-diagonal estimator retains exact intra-layer matrix elements and the full antisymmetric incompatibility term.
  • Component comparisons: Including the incompatibility term C(ωUhl_k) improves robustness under local and coherent noise with U_k ≠ 0.
  • Experimental evaluation: QFedQGT attained the highest round-averaged accuracy in every evaluated condition on ten-class MNIST using the reimei-E trapped-ion emulator.Across four device-heterogeneity conditions and three seeds, it also achieved the highest worst-round accuracy throughout.
  • Experimental evaluation: QFedQGT reached 0.81 round-averaged accuracy and 0.86 final-round accuracy, versus 0.79 and 0.83 for federated averaging.
  • Outlook: The framework's next tests include larger-qubit, larger-class hardware demonstrations and regimes dominated by coherent or correlated errors.The proposed extensions also include inter-layer geometry and quantum-error-correction integration.

Appendix A: Derivation of κ(p) for global depolarizing

Appendix A derives the depolarizing attenuation κ(p) for a globally depolarized pure state by evaluating the mixed-state SLD quantum Fisher information in an eigenbasis of the density matrix.

  • State decomposition: The globally depolarizing state is diagonalized in an eigenbasis containing |ψ⟩ and N−1 vectors spanning its orthogonal complement.Both eigenvalues are strictly positive for p ∈ (0,1), so every relevant eigenvalue pair contributes to the SLD expression.
  • Derivative structure: Differentiating the state shows that its parameter derivative connects |ψ⟩ only with the orthogonal complement under the chosen gauge.The only nonzero derivative matrix elements are the complementary off-diagonal terms.
  • SLD evaluation: Each nonzero derivative pair has eigenvalue sum 1−p+2p/N, and inserting the pair contributions into the SLD definition yields Eq. (26).The two orderings of each pair provide identical real contributions.
  • Closed form and limits: The resulting attenuation is κ(p)=(1−p)^2/(1−p+2p/N), with κ(0)=1 and κ(p)^2(1−p)^2≈(1−p)^4 for N≫1.For the experiments with N=64, κ(p) decreases from 0.897 at p=0.1 to 0.382 at p=0.6.

Appendix B: Block-diagonal truncation error

Appendix B bounds the error from replacing the regularized full Bures metric with a block-diagonal layerwise approximation in terms of neglected inter-layer coupling.

  • Metric decomposition: The full and block-diagonal regularized metrics are X=G+εI and Y=B+εI, with E=G−B collecting neglected inter-layer blocks.G and B are positive semidefinite Gram matrices, while B retains exact intra-layer blocks.
  • Inverse perturbation: The second-resolvent identity expresses X^−1−Y^−1 as X^−1(Y−X)Y^−1, enabling an operator-norm error bound.
  • Direction error: The resulting natural-gradient direction difference is controlled by the norm of the inter-layer coupling E.
  • Approximation regime: For geometrically local entanglement, the inter-layer coupling norm decays with layer separation, making the approximation error small.Unlike diagonal truncation, the block-diagonal estimator retains all intra-layer matrix elements and the full antisymmetric incompatibility term exactly.

Appendix C: Proof of Theorem 1

Appendix C proves Theorem 1 by deriving a one-step expected-loss recursion from smoothness, bounded client variation, and independent zero-mean noise, then unrolling the contraction.

  • Proof setup: The proof starts from the one-step update and applies the L-smoothness inequality to obtain a conditional expected-descent bound.
  • Noise analysis: Independent zero-mean client noises eliminate cross terms when decomposing the second-order update term into mean and fluctuation components.
  • Gradient bounds: The client-gradient deviation term is bounded using the assumptions on client variation and weighted Cauchy–Schwarz inequalities.
  • Optimal aggregation: The optimal aggregation matrices arise from a convex matrix least-squares problem under the unbiasedness constraint, with stationarity yielding Eq. (18).The stationary point is the minimum because the objective is convex in the aggregation matrices.
  • Convergence recursion: The recursion contracts with q=1−cηµ∈[0,1), and unrolling gives Δ_T≤q^TΔ_0+Lη/(2cµ).The constants are those introduced in the stated assumptions and equations, and the proof does not use its own conclusion.
  • Convergence recursion: The final bound is Eq. (29), with Δ_0=F(θ_0)−F*.

Appendix D: Experimental configuration, evaluation, and emulation

The experiments evaluate three aggregation methods under a controlled heterogeneous-client setup using a trapped-ion emulator, with accuracy measured from reconstructed Pauli expectations and readout predictions. Table III additionally compares client weighting and round-to-round variability in condition (a).

  • Experimental configuration: The shared classifier uses 6 qubits, 3 layers, 54 quantum parameters, 33 Pauli observables, and a 256-unit readout head for ten-class MNIST.Each layer re-uploads data, applies trainable single-qubit rotations, and uses a linear CNOT entangler.
  • Experimental configuration: Nine clients participate fully each round; four reliable clients receive balanced class-covering data, while five unreliable clients receive the remaining data under a Dirichlet(α = 0.5) partition.The supplied configuration identifies the reliable and unreliable client groups and the non-IID partition parameter.
  • Evaluation: Table III reports mean weights assigned to noisy and reliable clients in condition (a), together with mean per-client round-to-round standard deviations.A ratio below one denotes down-weighting of noisy clients.
  • Training and noise: Clients perform 20 local minibatch steps per round with batch size 32 and up to 50 epochs, while training runs for 60 rounds with specified local and server learning rates.The local learning rate decays as ηLoc(t) = ηLoc/(1 + 0.02 t), and the server step is ηSrv = 0.7.
  • Training and noise: Noisy clients combine coherent gate miscalibration, readout-induced label corruption, finite-shot noisy gradients, and a device-noise model applied on top of exact reverse-mode differentiation.The coherent miscalibration has magnitude 0.3 pk, label randomization affects a fraction 0.5 pk, and reverse-mode gradients match parameter-shift gradients to 2 × 10^-15.
  • Evaluation: All methods use identical ansatz, readout head, data partition, client noise profiles, random seeds, initial parameters, and evaluation set, differing only in local updates and aggregation weights.QFedAvg uses data-size weights, QFedFisher uses layerwise diagonal classical-Fisher weights, and QFedQGT uses the paper’s specified equations.
  • Evaluation: Accuracy is computed by executing the trained global model on the emulator, reconstructing Pauli expectations from counts, forming class scores, and comparing the arg-max prediction with the ground-truth label.The reported accuracy is averaged over three evaluation runs according to the supplied passage.
  • Emulation: The reimei-E emulator reproduces the trapped-ion processor’s native gates, all-to-all connectivity, and calibrated error model while accepting the physical machine’s job specification.Circuits are constructed and compiled with pytket and submitted through Quantinuum Nexus.
Loading 2608.28379v1…