Source-linked AI summary

Learning Rate Optimization for Federated Learning Exploiting Over-the-air Computation

Chunmei Xu, Shengheng Liu, Zhaohui Yang, Yongming Huang, Kai-Kit Wong

arXiv:2102.02946v3cs.ITeess.SP

TL;DR

Wireless federated learning needs efficient aggregation, but AirComp suffers aggregate distortion from fading and noise. The paper proposes dynamic learning rates jointly with wireless-resource optimization, extending closed-form MISO analysis to iterative MIMO methods; simulations on MNIST and CIFAR10 report reduced distortion and improved learning performance.

  • Problem

    AirComp-based federated learning can suffer aggregate distortion from fading and noise, which can degrade learning performance.

  • Method

    The paper adapts dynamic learning rates to wireless channels and jointly optimizes them with wireless resources in MISO and MIMO AirComp systems.

  • Results

    Simulations on MNIST and CIFAR10 validate reduced MSE and improved testing accuracy, while MISO admits a closed-form solution and MIMO uses an iterative method.

  • Takeaways & Limitations

    Dynamic learning rates provide a channel-adaptive way to mitigate AirComp distortion while preserving learning performance in the evaluated federated-learning settings.

Abstract

from arXiv · show

Federated learning (FL) as a promising edge-learning framework can effectively address the latency and privacy issues by featuring distributed learning at the devices and model aggregation in the central server. In order to enable efficient wireless data aggregation, over-the-air computation (AirComp) has recently been proposed and attracted immediate attention. However, fading of wireless channels can produce aggregate distortions in an AirComp-based FL scheme. To combat this effect, the concept of dynamic learning rate (DLR) is proposed in this work. We begin our discussion by considering multiple-input-single-output (MISO) scenario, since the underlying optimization problem is convex and has closed-form solution. We then extend our studies to more general multiple-input-multiple-output (MIMO) case and an iterative method is derived. Extensive simulation results demonstrate the effectiveness of the proposed scheme in reducing the aggregate distortion and guaranteeing the testing accuracy using the MNIST and CIFAR10 datasets. In addition, we present the asymptotic analysis and give a near-optimal receive beamforming design solution in closed form, which is verified by numerical simulations.

I. INTRODUCTION

The paper addresses communication and privacy challenges in federated learning by combining AirComp with dynamic learning rates that adapt to wireless fading. It develops optimization methods for MISO and MIMO settings and evaluates their effects on distortion and learning performance.

  • Motivation: AirComp enables efficient wireless aggregation in federated learning, but fading and noise can distort the aggregate and degrade training or inference.FL uploads local models from distributed devices for server aggregation, making communication efficiency and aggregate distortion central concerns.
  • Proposed approach: Dynamic learning rate (DLR) adapts between minimum and maximum boundaries to reduce aggregate error caused by fading and noisy channels.The paper distinguishes DLR from approaches that optimize only wireless resources.
  • Optimization framework: The work jointly optimizes DLR ratios and wireless resources, deriving a closed-form solution for MISO and an iterative algorithm for MIMO.The MISO and MIMO formulations target MSE minimization through AirComp.
  • Asymptotic analysis: The paper also derives a closed-form asymptotic beamforming solution for massive antenna deployments and verifies it numerically.The asymptotic analysis covers MISO, SIMO, and MIMO settings, with a near-optimal receive beamformer formed by summing channel vectors.

III. DLR FOR CHANNEL ADAPTION

The paper proposes adapting local learning rates to time-varying wireless channels, targeting fading-induced aggregate distortion in AirComp-based federated learning. The resulting formulation minimizes MSE subject to distortion-elimination constraints and bounded learning-rate ratios.

  • III. DLR FOR CHANNEL ADAPTION: Dynamic learning rates adapt each device’s local learning rate to the wireless environment, unlike prior approaches that optimize wireless resources.The objective is to mitigate fading-related distortion through learning-process hyperparameters.
  • III. DLR FOR CHANNEL ADAPTION: The fading-induced aggregate error is mitigated by selecting learning-rate ratios that satisfy the conditions derived from the channel-dependent error expression.The proof states that mitigation requires both terms in the fading-error expression to vanish.
  • III. DLR FOR CHANNEL ADAPTION: The aggregate AirComp error contains fading-related and noise-related components, with severe distortion degrading the global model and learning performance.The residual noise-related error can motivate retransmission when aggregate error is present.
  • III. DLR FOR CHANNEL ADAPTION: The DLR optimization minimizes MSE while enforcing fading-error elimination and learning-rate ratios r_k within [r_min, r_max].SISO and SIMO are treated as special cases of the MISO and MIMO formulations, respectively.

A. MISO

The MISO formulation models each device’s channel and transmit coefficient while the aggregator uses a single receive antenna. Its MSE objective supports a constrained transmit design and extends naturally to MIMO notation.

  • A. MISO: In MISO, each multi-antenna device transmits to a single-antenna aggregator, with channel vector h_k and maximum power P_k.The channel vector and power constraint determine the device-side transmission design.
  • A. MISO: The MISO aggregate error is expressed through the channel vectors, transmit coefficients, and noise, and its optimization removes fading-induced error under the stated equality constraints.The equality constraints are identified as guaranteeing elimination of the fading-related error.
  • A. MISO: The optimal transmitting coefficient vector is designed using the channel and power constraints, while the learning-rate ratios satisfy Σ_k r_k/K = 1.The same ratio relation is used when reformulating the MIMO problem.
  • A. MISO: MIMO generalizes the model by giving devices and the aggregator multiple antennas, with H_k as the channel matrix and m as the receive beamforming vector.The MIMO variables reduce to the corresponding MISO or SIMO forms in special cases.

V. DLR OPTIMIZATION

The DLR optimization develops a convex MISO solution with an equivalent linear-program formulation and derives lower bounds and equality conditions for the resulting MSE.

  • V. DLR OPTIMIZATION: The MISO DLR problem is convex and admits a closed-form solution.The broader algorithmic treatment contrasts this convex case with the nonconvex MIMO problem.
  • V. DLR OPTIMIZATION: The optimization is equivalent to minimizing the maximum c_k l_k under the constraints, yielding a typical linear programming problem.The equivalence follows because minimizing max_k(c_k l_k)^2 has the same solution as minimizing max_k c_k l_k.
  • V. DLR OPTIMIZATION: The MISO MSE has a lower bound attained only when c_i l_i = c_j l_j for all device pairs.This equality condition characterizes the coefficients that achieve the bound.
  • V. DLR OPTIMIZATION: Algorithm 1 searches over the maximizing device index and uses bisection to solve each learning-rate coefficient with complexity O(log2(1/δ)).The outer search requires K iterations.

B. MIMO

The MIMO DLR problem is nonconvex, so the paper alternates between receive beamforming and learning-rate optimization. The resulting procedure uses feasibility checks, DC programming, and bisection.

  • B. MIMO: The MIMO MSE has a lower bound, and when the no-DLR solution already attains it, DLR cannot improve performance and equals 1.This special case occurs when the relevant channel norms are equal.
  • B. MIMO: The MIMO problem is difficult because both its objective and one constraint are nonconvex.An auxiliary variable τ is introduced to reformulate the problem.
  • B. MIMO: The proposed iterative method alternately fixes the DLR ratios and receive beamforming vector, solving the resulting sub-problems in turn.With fixed beamforming, the DLR sub-problem reduces to the MISO case through an equivalent channel vector.
  • B. MIMO: Receive beamforming is obtained through feasibility reformulation with M = mm^H, relaxation or DC programming for the rank-one constraint, and bisection over τ.The DC formulation has complexity O(N^3 t), while bisection terminates when the interval width is below δ.
  • B. MIMO: Algorithm 3 initializes the DLR ratios, alternates Algorithms 2 and 1, and returns the receive beamformer, DLR ratios, and τ.Each iteration solves the beamforming and DLR sub-problems with respective accuracy parameters.

VI. ASYMPTOTIC ANALYSIS AND RECEIVE BEAMFORMING DESIGN

The section analyzes MSE and dynamic learning-rate ratios asymptotically across MISO, SIMO, and MIMO settings, then uses the analysis to design a near-optimal closed-form receive beamformer.

  • As antenna counts grow, the analysis covers MSE and DLR ratios in MISO, SIMO, and MIMO scenarios.
  • The asymptotic results support a near-optimal receive beamforming solution in closed form.

A. MISO

For MISO channels with infinitely many device antennas, independent Rayleigh fading yields equal effective channel coefficients, enabling the lower-bound MSE.

  • A. MISO: Independent Rayleigh channels lead to equal effective coefficients across devices in the asymptotic MISO regime.
  • A. MISO: The achieved MSE is inversely proportional to K^2 and N_d.

B. SIMO

In SIMO, increasing receive antennas makes device channels asymptotically orthogonal, enabling simple receive beamforming and an MSE that decreases with device count and antenna count.

  • B. SIMO: As N_t increases, device channels become asymptotically orthogonal, enabling a simple receive beamforming design.
  • B. SIMO: The resulting beamforming construction guarantees the lower-bound MSE in the asymptotic SIMO setting.
  • B. SIMO: The achieved MSE is inversely proportional to K and N_t.

C. MIMO

The MIMO analysis uses asymptotic channel orthogonality to construct receive beamformers from dominant eigendirections, yielding closed-form MSE scaling in two antenna-growth regimes.

  • C. MIMO: When N_d grows beyond N_t, the achieved MSE is inversely proportional to K^2 and N_d, irrespective of N_t.
  • C. MIMO: As N_t grows beyond N_d, device channel subspaces become orthogonal and a receive beamformer can be formed from one eigenvector per device.
  • C. MIMO: The asymptotic beamforming construction relies on equal singular values and orthogonal channel subspaces to equalize effective-channel power.
  • C. MIMO: For N_t →∞ with N_t > N_d, the achieved MSE is inversely proportional to K and N_t, irrespective of N_d.

D. Observations

The asymptotic analysis characterizes how antenna and device counts shape MSE and the benefit of dynamic learning rates (DLR). With many antennas, simple normalized-channel combining reaches the MSE lower bound while DLR contributes less.

  • As antennas increase, signal alignment improves, reducing DLR’s performance gain and pushing the DLR ratio closer to 1.
  • Summing normalized channel vectors achieves the MSE lower bound as antenna count tends to infinity because channel vectors become orthogonal with equal power.
  • MSE is inversely proportional to K and N_t in SIMO, but to K^2N_d in MISO.
  • In MIMO, N_d →∞ yields the MISO MSE σ^2/(PK^2N_d), whereas N_t →∞ yields the SIMO MSE σ^2/(PKN_t), independent of the other antenna count.

VII. SIMULATION RESULTS

The simulations evaluate DLR and closed-form receive beamforming against methods without DLR across wireless scenarios. They also assess aggregate-error behavior, learning performance, and asymptotic antenna effects.

  • Simulations compare the proposed DLR design and near-optimal closed-form receive beamforming with existing approaches without DLR.
  • The evaluation spans SISO, MISO, SIMO, and MIMO scenarios, including MSE, MNIST and CIFAR10 learning performance, and massive-antenna beamforming.
  • The compared baselines include NDLR in SISO/MISO and SDR or DC methods in SIMO/MIMO.

A. Performance on MSE using DLR

Across wireless scenarios, DLR reduces aggregate error and slightly improves learning performance relative to fixed learning rates, while the proposed beamforming design approaches theoretical bounds as antennas increase.

  • A. Performance on MSE using DLR: Aggregate error decreases with more devices, and adding DLR further reduces it under SISO, MISO, SIMO, and MIMO settings.
  • A. Performance on MSE using DLR: The iterative DLR and receive-beamforming algorithm achieves performance relatively close to the near-optimal MIMO solution.
  • B. Performance of Learning Task Using DLR: DLR assigns smaller learning rates to devices with higher channel gain and larger rates to devices with lower gain.
  • B. Performance of Learning Task Using DLR: A larger DLR-ratio range decreases MSE, but overly broad ranges can increase training-loss and test-accuracy variance.
  • B. Performance of Learning Task Using DLR: The proposed DLR slightly improves learning and inference performance over fixed-learning-rate methods on MNIST and CIFAR10.
  • C. Performance of the Proposed Closed-form Receive Beamforming Solution: Increasing antenna count reduces aggregate error, and the proposed closed-form receive beamforming MSE approaches the theoretical bound in massive-antenna settings.
Loading 2102.02946v3…