Source-linked AI summary
Estimation of effective temperatures in quantum annealers for sampling applications: A case study with possible applications in deep learning
Marcello Benedetti, John Realpe-Gómez, Rupak Biswas, Alejandro Perdomo-Ortiz
TL;DR
Boltzmann sampling is important for machine learning, but quantum annealers may produce instance-dependent effective temperatures that differ from their physical temperature. The paper estimates this temperature from two sets of annealer samples and evaluates the method on a hardware-embedded Chimera-RBM. For the studied case, instance-dependent temperature estimation approaches CD-100, whereas fixed temperatures outperform only CD-1.
Problem
Quantum annealers may sample Boltzmann-like distributions at unknown, instance-dependent effective temperatures, limiting their use for Boltzmann sampling.
Method
The paper estimates effective temperature from two sample sets using linear regression and tests it in Chimera-RBM learning on quantum hardware.
Results
Instance-dependent effective-temperature estimation achieves performance close to CD-100 in the studied case, while fixed effective temperatures outperform CD-1.
Takeaways & Limitations
Effective-temperature estimation provides a practical route for using quantum-annealer samples in the studied Boltzmann-machine learning setting.
Takeaways & Limitations
The study uses a less powerful sparse Chimera-RBM that struggles to reproduce the 4 × 4 Bars And Stripes dataset faithfully.
Abstract
from arXiv · showhide
An increase in the efficiency of sampling from Boltzmann distributions would have a significant impact on deep learning and other machine-learning applications. Recently, quantum annealers have been proposed as a potential candidate to speed up this task, but several limitations still bar these state-of-the-art technologies from being used effectively. One of the main limitations is that, while the device may indeed sample from a Boltzmann-like distribution, quantum dynamical arguments suggest it will do so with an {\it instance-dependent} effective temperature, different from its physical temperature. Unless this unknown temperature can be unveiled, it might not be possible to effectively use a quantum annealer for Boltzmann sampling. In this work, we propose a strategy to overcome this challenge with a simple effective-temperature estimation algorithm. We provide a systematic study assessing the impact of the effective temperatures in the learning of a special class of a restricted Boltzmann machine embedded on quantum hardware, which can serve as a building block for deep-learning architectures. We also provide a comparison to $k$-step contrastive divergence (CD-$k$) with $k$ up to 100. Although assuming a suitable fixed effective temperature also allows us to outperform one step contrastive divergence (CD-1), only when using an instance-dependent effective temperature do we find a performance close to that of CD-100 for the case studied here.
I. INTRODUCTION
The paper addresses whether quantum annealers can support Boltzmann sampling and how instance-dependent effective temperatures affect their use in learning Boltzmann machines. It proposes estimating that temperature from device samples and evaluates the approach on a Chimera-RBM.
- Motivation: Quantum annealers may sample from Boltzmann-like distributions, but their effective temperature can depend on the problem instance rather than equal the physical temperature.This uncertainty complicates their use for sampling and machine learning.
- Motivation: Boltzmann-machine learning is generally intractable because sampling methods such as MCMC can require long equilibration times.RBMs reduce this difficulty through factorized conditional distributions and serve as building blocks for deeper architectures.
- Problem and approach: Estimating control-parameter shifts directly is impractical because it resembles the distribution-learning problem itself.The paper therefore focuses on estimating effective temperature without additional instance-dependent shifts.
- Problem and approach: The proposed algorithm estimates an instance-dependent effective temperature using two sample sets and linear regression.Those samples are also used for the eventual sampling application, unlike the many log-likelihood-gradient evaluations required by the earlier approach.
- Evaluation: The study tests quantum-assisted learning on a special Chimera-RBM, a hardware-compatible restricted architecture intended as a building block for deep-learning systems.The implementation uses a Chimera-RBM on the Bars And Stripes dataset with D-Wave hardware.
B. Quantum annealing
Quantum annealing slowly transforms an easily prepared initial ground state into a problem-encoding final Hamiltonian. D-Wave hardware controls this process through qubit fields, couplings, and time-dependent annealing schedules on a Chimera topology.
- Quantum annealing: Quantum annealing maps an optimization cost function into a physical-system energy function while incorporating quantum fluctuations.The algorithm aims to maintain the system near its lowest-energy solution space.
- Quantum annealing: The annealing process transforms the ground state of an initial quantum system into the ground state of a final problem Hamiltonian.D-Wave implements this paradigm for quadratic unconstrained optimization on binary variables.
- Hardware implementation: D-Wave control parameters include a field h_i for each qubit and a coupling J_ij for each interacting pair.The interactions follow the Chimera graph, composed of coupled 4 × 4 complete bipartite cells.
- Sampling behavior: Under certain conditions, quantum annealers can sample approximately from a Boltzmann distribution at an effective temperature.This behavior is supported by theoretical arguments and experimental evidence, despite the device being designed for ground-state finding.
C. Quantum annealing for sampling applications
Quantum annealers may assist sampling for Boltzmann-machine learning, but their dynamics can freeze before the anneal ends, producing an effective temperature distinct from the physical temperature. The paper studies this mechanism and focuses on estimating the temperature rather than establishing a general quantum advantage.
- Sampling motivation: Estimating model averages over Boltzmann distributions is computationally hard, especially for models with slow mixing under MCMC.This motivates investigating quantum annealers as possible sampling resources.
- Effective-temperature mechanism: Quantum annealers may sample from classical Boltzmann distributions because thermalization and decoherence can freeze the dynamics near the end of the anneal.In a quasistatic regime, the final distribution can reflect the classical problem Hamiltonian.
- Effective-temperature mechanism: The freezing point rescales the physical temperature into an effective temperature that can exceed it and depend on the instance.For the DW2X, the physical fridge temperature is 0.033 in the paper’s dimensionless units.
- Scope: Quantum tunneling may assist thermalization before freezing, but any advantage depends on the energy landscape and available quantum resources.The paper leaves the question of general quantum advantage for future work.
- Application: The work applies effective-temperature estimation to a machine-learning problem involving a Chimera-RBM embedded on quantum hardware.The hardware-compatible model is obtained by removing RBM links absent from the D-Wave topology.
D. Chimera restricted Boltzmann machines
Restricted Boltzmann machines simplify learning through factorized conditional distributions, while Chimera-RBMs adapt this structure to quantum annealers with limited connectivity. This adaptation trades generality for hardware-compatible representation.
- RBMs use a complete bipartite graph, with visible-hidden interactions but no within-layer interactions.This structure makes the conditional distributions factorize into single-variable marginals.
- Factorized conditional distributions allow data averages to be computed exactly in one shot, while model averages can be approximated using CD-k.CD-k alternates conditional sampling for k steps before estimating model averages.
- CD-k is not guaranteed to produce correct results or follow a log-likelihood gradient.The passage motivates better sampling methods for learning Boltzmann-machine models.
- Embedding an RBM on quantum hardware requires substantially more physical qubits and couplings because of limited device connectivity.The hardware representation can therefore be much larger than the original logical model.
- Chimera-RBMs remove RBM links absent from the D-Wave topology, producing models that are naturally represented on the device.This model class is used to accommodate the hardware interaction graph.
III. RELATED WORK
Related work identifies restricted connectivity, parameter noise, and control-parameter bounds as important hardware limitations for quantum-assisted Boltzmann-machine learning. Other approaches estimate gradients or embed logical variables using qubit chains.
- Limited connectivity was identified as the most relevant limitation for Chimera-RBM learning.Chimera-RBMs are sparse relative to complete bipartite RBMs, with localized connections that may hinder higher-level correlations.
- Parameter noise was reported as the next most relevant limitation, with weight noise more important than bias noise in standard RBMs.The relative importance can differ for Chimera-RBMs because their bias-to-weight proportions differ.
- Upper bounds on model-parameter magnitudes appeared to have little impact in the studied context.The passage notes that lower bounds may matter more for sampling applications, especially when control-parameter noise is present.
- Denil and De Freitas optimized one-step reconstruction error as a black-box function and estimated its gradient empirically.Their approach used simultaneous perturbation stochastic approximation to bypass instance-dependent corrections.
- Adachi and Henderson embedded logical variables as strongly coupled qubit strings and averaged effects of control-parameter noise.They used the annealer to estimate model averages for pre-training a two-layer neural network.
IV. QUANTUM-ASSISTED LEARNING OF BOLTZMANN MACHINES
The QuALe approach incorporates an instance-dependent effective temperature into Boltzmann-machine learning while using the annealer to estimate model averages. It initializes and iteratively updates model parameters using temperature estimates from device samples.
- QuALe assumes the annealer samples from a Boltzmann distribution whose effective temperature may depend on the instance.The learned parameters are ratios of the device control parameters to this effective temperature.
- The procedure initializes small control parameters, samples from the device, and estimates an initial effective temperature.That estimate is used to compute the initial model parameters.
- At each iteration, samples and current model parameters are used to estimate temperature before updating the model parameters.The update follows the learning rules referenced by the paper.
- Device samples estimate model ensemble averages, while data ensemble averages use samples with visible units clamped to data points.The same learning step therefore combines unconstrained and data-conditioned device sampling.
- The effective temperature for the next iteration cannot be estimated using the not-yet-known updated parameters, creating a learning-step dependency.The method resolves this by using an estimate available from the current step.
- Temperature-estimation error introduces noise that can make learning deviate from the intended update rules.The paper identifies comparison with finite-sample gradient-estimation noise as an open issue.
A. Extracting temperature from two sample sets
The temperature-estimation method compares energy-level probabilities from two sample sets generated at original and rescaled control parameters. A linear regression across many populated energy-level pairs estimates the instance-dependent effective temperature.
- At inverse temperature β, the energy-level probability combines degeneracy, a Boltzmann factor, and the partition function.This relationship motivates using probability log-ratios to estimate temperature.
- The log-ratio between two energy-level probabilities is estimated from their sampled frequencies, with binning used when more robust statistics are needed.The method uses energy levels E1 and E2 and their difference ΔE.
- Rescaling control parameters by x is treated as setting β = xβeff, assuming Teff remains approximately invariant under small rescalings.Plotting the log-ratio against x should yield a line whose slope contains βeffΔE.
- The improved method samples at the original parameters and one rescaling, then differences log-ratios across all populated energy-level pairs.This uses information from more energy levels than the earlier two-level approach.
- The resulting plot of Δℓ against ΔE is expected to be linear, with slope (x−1)βeff.Matching bin intervals across both histograms is required for valid overlap comparisons.
- Using K bins for R samples per set yields O(K^2) = O(R) regression data points.Raw energies are evaluated using the original control parameters before binning.
- The rescaling factor x must balance informative changes against noise, unpopulated levels, and possible violation of effective-temperature invariance.Both excessively small and excessively large rescalings reduce the method’s reliability.
B. A rule of thumb for the scaling factor
The scaling factor x is chosen so two Boltzmann distributions remain close yet distinguishable with the available samples. The rule uses KL-divergence distinguishability and a second-order Fisher-information approximation.
- Choosing x: The KL divergence characterizes whether two distributions are distinguishable with a given sample size R.For sufficiently large R, distinguishability requires D_KL(P_β′||P_β) > ln(C/P0)/R.
- Choosing x: When β and β′ are sufficiently close, the KL divergence can be expanded to second order using Fisher information.Here the Fisher information is essentially the specific heat, or generalized susceptibility.
- Choosing x: The prescribed x makes the two distributions as close as possible while maintaining the target distinguishability dKL/R.The criterion is 1/2χ(β)(1−x)^2β_eff^2 = dKL/R.
- Practical considerations: Equation (18) is a rule of thumb because β_eff appears in the expression used to choose x.The initialization can use a reasonable guess or a pseudo-likelihood estimate.
VI. A FEW GADGETS TO IMPROVE PERFORMANCE
The learning procedure supplements quantum sampling with bias correction, suitable initialization, and samples from rescaled control parameters. These additions target hardware biases, noise, and effective-temperature estimation.
- Performance improvements: Persistent and random control-parameter biases can significantly impair quantum-annealer performance, motivating explicit bias correction.A technique is cited for determining and correcting persistent biases.
- Performance improvements: Running CD-1 briefly before QuALe supplies initial control parameters above the device noise level.This workaround is attributed to the current state of quantum annealing technologies.
- Temperature estimation: Effective-temperature estimation uses two sample sets: the actual control parameters and versions rescaled by x.The scaling is chosen so the resulting distributions are close yet distinguishable.
- Implementation: QuALe includes the three performance-improvement gadgets unless otherwise specified.
VII. LEARNING OF A BOLTZMANN MACHINE ASSISTED BY THE D-WAVE 2X
The study applies quantum-assisted learning to a Chimera-RBM on BAS and evaluates temperature estimation, bias correction, importance sampling, and initialization. Instance-dependent temperature estimation approaches CD-100 more closely than fixed-temperature alternatives.
- Experimental setup: The Chimera-RBM uses 16 visible and 16 hidden units, with learning rate η = 0.03, and temperature estimation uses R = 1000 samples with dKL = 500.The experiment runs on the D-Wave 2X and compares actual with rescaled control parameters.
- Temperature estimation: Teff ≈0.095 and Rcoeff ≈−0.95 are obtained from a linear regression of log-likelihood-ratio differences against energy differences.The two histograms use actual and x = 0.72-rescaled control parameters.
- Performance improvements: Persistent bias correction outperforms QuALe without correction, while rescaled samples and brief CD-1 initialization also improve performance.These effects are evaluated through average log-likelihood across repeated hardware runs or a single chip location.
- Comparison with CD-k: QuALe@Teff outperforms CD-1 after about 300 iterations and CD-10 after about 1500 iterations, but does not surpass CD-100 within 5000 iterations.The results show a clear trend toward CD-100, while CD-k reaches peak average performance earlier and QuALe@Teff rises more steadily.
- Fixed versus estimated temperature: Using the physical temperature TDW2X = 0.033 gives Lav < −14, while fixed temperatures 0.08, 0.16, and Tav perform below instance-dependent estimation.The average fixed temperature is QuALe@Tav ≈0.1.
- Temperature variation: The effective temperature varies beyond finite-sampling error during an 80-iteration QuALe@Teff window, despite being estimated only once during execution.The variation is compared with repeated estimates at each iteration using medians and interquartile error bars.
VIII. CONCLUSIONS AND FUTURE WORK
The study finds that instance-dependent effective-temperature estimation improves quantum-assisted learning on a small Chimera-RBM, approaching CD-100 more closely than fixed-temperature variants. It also identifies model expressiveness, scalability, and dataset size as boundaries for future work.
- Conclusions: QuALe estimates effective temperature and the model-distribution term in the log-likelihood gradient from quantum-annealer samples during learning.The approach was evaluated on a Chimera-RBM with 16 visible and 16 hidden units.
- Limitations: The Chimera-RBM has about 31% of the weight parameters of a corresponding dense RBM and struggles to reproduce the 4 × 4 BAS dataset faithfully.This reduced expressiveness limits conclusions about more powerful Boltzmann-machine models.
- Limitations: The study uses a moderately small dataset and 16 visible plus 16 hidden units because exact log-likelihood becomes intractable for larger systems.Reconstruction and cross-entropy errors are described as rougher proxies for log-likelihood.
- Comparison with CD-k: Only instance-dependent effective-temperature estimation shows a steady performance increase close to the largest tested CD value, k = 100.Fixed suboptimal temperatures can outperform CD-1, but not the higher-k comparisons.
- Future work: Future work includes testing larger and more complex datasets, determining how temperature-estimation sample counts scale, and exploiting Chimera models with lateral connections.The authors also mention extensions to deep architectures, discriminative models, and momentum-based learning.
Appendix A: Comparison to alternative temperature estimation techniques
The appendix compares pseudo-likelihood and linear-regression estimates of instance-dependent effective temperature. On the 4×4 BAS task, pseudo-likelihood performs better initially, but linear regression later reaches higher likelihood values.
- Comparison rationale: The authors report that their linear-regression method produces superior results to the alternative techniques considered on the BAS dataset.Pseudo-likelihood is identified as the state-of-the-art technique for estimating Ising-model parameters, motivating the comparison.
- Mean-field alternative: The Bethe mean-field approach yields real-valued parameter estimates only during approximately the first hundred learning iterations, indicating it is unsuitable for this regime.This feasibility test preceded development of a method targeted specifically at Teff estimation.
- Pseudo-likelihood estimation: The pseudo-likelihood estimator is obtained by maximizing average pseudo-likelihood over Teff using second-order Newton iteration.The procedure starts at Teff = 1 and stops when the update is below 10^-5.
- Temperature-estimation methods: The study compares pseudo-likelihood maximization with linear regression for estimating the instance-dependent effective temperature Teff.Both techniques are evaluated within quantum-assisted learning of a Chimera-RBM on the BAS dataset.
- Performance comparison: Pseudo-likelihood performs better during roughly the first 1000 iterations, whereas linear regression performs better afterward and reaches higher average log-likelihood.The figure measures average log-likelihood every 50 iterations across five runs, with bands representing one standard deviation.
- Temperature trajectories: Pseudo-likelihood estimates are consistently smaller and less variable than the effective temperatures estimated by linear regression along one learning path.Temperature estimation begins at iteration 100 after restarting from CD-1.