Source-linked AI summary

Stable Self-Modulating Quantum Fast-Weight Programmers with Bounded Memory Gates

Kuo-Chung Peng, Jiun-Cheng Jiang, Chun-Hua Lin, Yifeng Peng, Junghoon Justin Park, Huan-Hsin Tseng, Hsin-Yi Lin, Kuan-Cheng Chen, Chen-Yu Liu, Shinjae Yoo, Samuel Yen-Chi Chen

arXiv:2607.02363v1quant-phcs.AIcs.ETcs.LGcs.NE

TL;DR

Recurrent quantum models inherit a training limitation: nonlinear recurrence requires propagating gradients through a time-ordered nonlinear state evolution. The paper applies a sign-preserving tanh bound only to the recurrent old-state gate, leaving additive updates and new-update modulation unchanged. Old-state modulation provides the most consistent long-sequence improvement over Standard QFWP, while the bounded gate completes all long-sequence cells and reduces mean final MSE for both old-state variants across both quantum datasets.

  • Problem

    Recurrent quantum models inherit a training limitation: nonlinear recurrence requires propagating gradients through a time-ordered nonlinear state evolution.

  • Method

    The paper applies a sign-preserving tanh bound only to the recurrent old-state gate, leaving additive updates and new-update modulation unchanged.

  • Results

    Old-state modulation provides the most consistent long-sequence improvement over Standard QFWP, while the bounded gate completes all long-sequence cells and reduces mean final MSE for both old-state variants across both quantum datasets.

  • Takeaways & Limitations

    Accumulated-memory modulation is the key mechanism of Self-Modulating QFWP, and bounded old-state gating is a targeted stabilization strategy when unbounded recurrence becomes unstable.

  • Takeaways & Limitations

    The Milan SMS benchmark evaluates the original unbounded model rather than the bounded modification because all tested configurations completed training without the bounded gate.

Abstract

from arXiv · show

Quantum Fast-Weight Programmers (QFWPs) store temporal information in dynamically programmed variational-circuit parameters rather than in nonlinear recurrent hidden states, offering a practical route to quantum sequence modeling. Self-Modulating QFWP improves this framework by using input-dependent gates for both new fast-weight updates and the accumulated fast-weight state, but its unbounded old-state multiplier can diverge in long-sequence regimes. We propose a bounded old-state modulation rule that applies a sign-preserving tanh gate only to the recurrent memory branch while leaving the additive update and new-update modulation unchanged. We evaluate standard QFWP, full Self-Modulating QFWP, Only-New, and Only-Old variants on two CUDA-Q quantum-dynamics forecasting tasks and on Milan SMS telecommunication activity prediction. The quantum-dynamics results show that old-state modulation is the most consistent source of improvement over Standard QFWP, and that bounding the old-state gate removes long-sequence divergence while improving aggregate robustness. On Milan SMS forecasting, the original unbounded Self-Modulating QFWP converges across the tested grid and shows its clearest gains at longer input windows, with behavior close to the Only-Old ablation. These findings identify accumulated-memory modulation as the key mechanism of Self-Modulating QFWP and bounded old-state gating as a targeted stabilization strategy.

1 Introduction

QFWP stores temporal memory in dynamically programmed circuit parameters, avoiding nonlinear recurrent hidden-state evolution. Self-Modulating QFWP controls both new updates and accumulated memory, while bounded old-state gating targets divergence in long sequences.

  • 1 Introduction: QFWP moves temporal memory from a nonlinear recurrent hidden state into dynamically programmed variational-circuit parameters, avoiding BPTT across a quantum recurrent cell.A classical slow programmer generates fast-weight updates for the variational quantum circuit.
  • 1 Introduction: Self-Modulating QFWP uses input-dependent multiplicative gates to control both newly generated updates and previously accumulated fast-weight state.The old-state gate can retain, suppress, amplify, or sign-reverse past memory.
  • 1 Introduction: Bounded old-state modulation applies a sign-preserving tanh gate only to recurrent memory, preserving additive updates and new-update modulation.This isolates stabilization to the accumulated fast-weight branch.
  • 1 Introduction: Old-state modulation is the most consistent source of improvement over Standard QFWP, while bounding it removes long-sequence divergence and preserves low-error behavior.The paper therefore identifies accumulated-memory control as central to Self-Modulating QFWP.

2 Related Work

Quantum sequence modeling builds on hybrid quantum–classical architectures and variational circuits for temporal learning. QFWP instead stores context in dynamically programmed fast-network parameters rather than nonlinear hidden-state recurrence.

  • 2 Related Work: Quantum sequence models use hybrid quantum–classical architectures to learn temporal dependencies, with QLSTM and related models applying variational circuits to sequence tasks.These studies report promising empirical behavior while retaining recurrent training challenges.
  • 2 Related Work: Fast Weight Programmers store temporal information in context-dependent parameters written by a slow network into a fast network.This differs from nonlinear hidden-state recurrence and connects to linear self-attention and recurrent fast-weight variants.
  • 2 Related Work: QFWP transfers the fast-weight principle to hybrid quantum learning by using a classical slow programmer to update variational quantum-circuit parameters.Prior demonstrations include time-series prediction and reinforcement learning.

3 Model

The model uses a classical controller to generate additive fast-weight updates and input-dependent modulation of new and accumulated weights. The proposed bounded variant constrains only recurrent old-state multiplication with a sign-preserving tanh gate.

  • 3.1 QFWP baseline: A classical slow controller maps each scalar input to a hidden state, whose affine heads generate layer and qubit vectors for the fast-weight update.Their outer product forms the raw update, and the quantum circuit uses the resulting parameters for prediction features.
  • 3.2 Self-modulation inherited from the previous model: Self-Modulating QFWP adds separate input-dependent modulation matrices for the new update and accumulated old state, with additive and multiplicative ablations.Only-New leaves the old state unmodulated, while Only-Old isolates accumulated-memory modulation.
  • 3.2 Self-modulation inherited from the previous model: Old-state modulation is targeted because it directly controls whether past fast-weight updates are retained, suppressed, amplified, or sign-reversed.This mechanism motivates changing the recurrent branch without modifying the full model.
  • 3.3 Bounded old-state modulation: The bounded variant replaces unconstrained old-state gates with a sign-preserving tanh bound, while leaving M_new unchanged and applying no bound to Standard QFWP or Only-New.The tanh is smooth, close to identity near zero, and prevents large recurrent multipliers from becoming unbounded.
  • 3.3 Bounded old-state modulation: The bounded recurrence prevents geometric amplification of stored updates, so growth comes from additive new updates rather than repeated old-state multiplication.This is the intended stability property of the bounded Only-Old recurrence.

4 Quantum-Dynamics Prediction Tasks

The evaluation uses two CUDA-Q quantum-dynamics trajectories as univariate next-step forecasting tasks. The benchmarks cover an open Jaynes–Cummings system and a closed dispersive transmon–resonator system with distinct observables and time ranges.

  • 4 Quantum-Dynamics Prediction Tasks: Each CUDA-Q benchmark extracts one scalar observable from a simulated quantum trajectory and predicts its next value from chronological sliding windows.Each trajectory has 3000 normalized time steps, with chronological 80% training and 20% test splits.
  • 4 Quantum-Dynamics Prediction Tasks: The open Jaynes–Cummings benchmark couples a two-level qubit to a single cavity mode truncated to five Fock levels and includes photon loss.The target is qubit excitation probability over t ∈[0, 50].
  • 4 Quantum-Dynamics Prediction Tasks: The closed dispersive transmon–resonator benchmark models a two-level transmon with a resonator truncated to 20 Fock levels.The prediction target is the resonator position quadrature over t ∈[0, 25] ns.

5 Telecommunication Activity Prediction

The study applies Self-Modulating QFWP to SMS activity forecasting across 100 Milan spatial-cell time series, treating each cell as a separate univariate sequence.

  • Motivation: The application evaluates whether the sequence-modeling framework transfers to realistic urban telecommunication dynamics with both short- and longer-range temporal patterns.The setting is intended as a practical benchmark for sequential forecasting models.
  • Dataset and task: The Milan dataset supplies spatially gridded telecommunication activity, and this study focuses on SMS signals as univariate forecasting tasks.The data include SMS, call, and Internet modalities, but only SMS activity is modeled here.
  • Dataset and task: Each spatial cell becomes an independent time series whose historical SMS values are used to predict future activity.The experiments cover 100 spatial cells, providing 100 SMS forecasting series.

6 Experimental Protocol

The experiments compare Standard QFWP, Self-Modulating variants, and bounded-old modifications across controlled quantum-dynamics and long-sequence stability settings. Evaluation uses final test MSE, with medians emphasized when divergent cells distort means.

  • Compared models: The protocol compares Standard QFWP, full Self-Modulating QFWP, Only-Old, Only-New, and bounded-old variants across the experimental grid.The bounded-old modification is evaluated separately in the long-sequence stability study.
  • Training trajectories: The trajectory figure compares Standard and full Self-Modulating QFWP at H = 10 and N = 32 across epochs 10, 20, 50, and 100 on two quantum-dynamics benchmarks.Columns represent the two model variants and the two benchmarks, while rows represent training epochs.
  • Metrics and completion: Final test MSE is the primary accuracy metric, while medians are emphasized when isolated divergent cells distort mean performance.A cell is complete only if it reaches epoch 100 without a cancellation sidecar.
  • Comparison metrics: Relative improvement is measured against Standard QFWP, with positive values indicating lower MSE; Relative Strength and Synergy separate single-sided and combined modulation effects.The bounded-versus-unbounded stability analysis is kept separate from unbounded reproduction aggregates because the bounded rule changes the architecture.

7 Results

Across the quantum-dynamics benchmarks, old-state modulation most consistently improves long-sequence prediction, while tanh-bounding the recurrent gate removes divergence and improves aggregate robustness. On Milan SMS forecasting, self-modulation is strongest for longer windows and behaves similarly to Only-Old.

  • Quantum-dynamics prediction: Across completed unbounded cells, modulated variants generally converge faster and reach lower test MSE than Standard QFWP, especially at longer sequence lengths.Only-New often remains close to Standard QFWP at N = 32 and N = 64, whereas Only-Old and full Self-Modulating QFWP recover low-MSE solutions across most long-sequence configurations.
  • Quantum-dynamics prediction: Full Self-Modulating QFWP and Only-Old each improve over Standard QFWP in 21 of 28 completed Jaynes–Cummings cells, while Only-New improves in 15 of 28.For transmon-resonator dynamics, full Self-Modulating QFWP improves in 25 of 26 completed cells, while Only-Old and Only-New each improve in 22 of 26.
  • Modulation diagnostics: Relative-strength diagnostics favor old-state over new-update modulation in many configurations, but synergy is mixed because the full model does not uniformly exceed the better single-sided variant.The results therefore identify accumulated fast-weight-state modulation as the most consistent performance contributor rather than establishing uniform complementarity between both gates.
  • Bounded versus unbounded stability: Bounded old-state modulation completes long-sequence cells and reduces mean final MSE for both old-state multiplicative variants on both quantum-dynamics datasets.The bounded gate also improves the median in three of four dataset–variant pairs; the exception is Jaynes–Cummings full Self-Modulating QFWP, changing from 1.24 × 10^-5 to 1.35 × 10^-5.
  • Milan SMS forecasting: On Milan SMS forecasting, full Self-Modulating QFWP does not consistently beat baselines for short windows but shows clearer gains as sequence length increases and remains close to Only-Old.Only-New is weaker in the long-sequence regime, supporting the relevance of modulating memory-bearing existing QFWP parameters beyond simulated quantum dynamics.

8 Conclusion

The study identifies accumulated fast-weight modulation as the main source of Self-Modulating QFWP’s long-sequence gains and proposes bounded old-state gating to stabilize recurrence.

  • 8 Conclusion: Old-state modulation provides the most consistent long-sequence improvements over Standard QFWP on two CUDA-Q quantum-dynamics benchmarks.Only-Old and full Self-Modulating QFWP deliver the strongest consistent improvements in the reported quantum-dynamics experiments.
  • 8 Conclusion: The unbounded old-state multiplier can diverge in difficult long-sequence regimes, particularly when recurrent memory carries more history.The instability is associated with multiplicative old-state modulation rather than the new-update branch.
  • 8 Conclusion: The proposed sign-preserving tanh gate bounds only the recurrent memory branch while leaving the new-update branch unchanged, removing the identified failure mode.This targeted modification is presented as a stabilization strategy for unbounded recurrence.
  • 8 Conclusion: On Milan SMS forecasting, original Self-Modulating QFWP is most useful at longer input windows and behaves similarly to the Only-Old ablation.The reported conclusion extends accumulated-memory modulation beyond simulated quantum dynamics to noisy real-world time series.
Loading 2607.02363v1…