Source-linked AI summary

Many-body quantum state tomography with neural networks

Giacomo Torlai, Guglielmo Mazzola, Juan Carrasquilla, Matthias Troyer, Roger Melko, Giuseppe Carleo

arXiv:1703.05334v2cond-mat.dis-nncond-mat.quant-gasphysics.comp-phquant-ph

TL;DR

The paper addresses the difficulty of reconstructing complete many-body quantum states and their entanglement-related properties from limited experimental measurements. It trains neural-network wave-function representations on measurements in multiple bases, using RBMs and gradient-based optimization. The approach is presented for quantum states and models including highly entangled systems, with entanglement entropy and magnetic observables evaluated from the reconstructed representation.

  • Problem

    Quantum-state tomography seeks complete quantum descriptions from limited accessible measurements, while multi-qubit entanglement is difficult to probe directly.

  • Method

    The approach trains RBM neural-network wave-functions on measurement distributions from multiple bases, using staged amplitude and phase learning with gradient-based optimization.

  • Results

    The reconstructed neural-network state is used to compute entanglement entropy and magnetic observables for quantum many-body systems.

  • Takeaways & Limitations

    The approach reconstructs experimentally relevant many-body information from measurement data without requiring direct access to all target quantities.

Abstract

from arXiv · show

The experimental realization of increasingly complex synthetic quantum systems calls for the development of general theoretical methods, to validate and fully exploit quantum resources. Quantum-state tomography (QST) aims at reconstructing the full quantum state from simple measurements, and therefore provides a key tool to obtain reliable analytics. Brute-force approaches to QST, however, demand resources growing exponentially with the number of constituents, making it unfeasible except for small systems. Here we show that machine learning techniques can be efficiently used for QST of highly-entangled states, in both one and two dimensions. Remarkably, the resulting approach allows one to reconstruct traditionally challenging many-body quantities - such as the entanglement entropy - from simple, experimentally accessible measurements. This approach can benefit existing and future generations of devices ranging from quantum computers to ultra-cold atom quantum simulators.

Appendix A: RBM Quantum State Tomography

The scheme reconstructs many-body wave-functions with neural networks trained on experimentally accessible measurements across multiple bases. The resulting compact representation can support calculations of observables, overlaps, and other state information.

  • The method approximates a many-body wave-function using an artificial neural network whose auxiliary neurons increase its expressive power.The network is optimized using experimentally accessible information.
  • Measurements in multiple bases encode information about amplitudes and phases, and training minimizes the total Kullback–Leibler divergence between measured and reconstructed distributions.The divergence reaches zero when reconstruction is perfect in every basis.
  • A sufficiently large set of measurement bases may be needed to estimate phases, but for most states of interest its size scales polynomially with system size.
  • After training, the neural network provides a compact representation that can be used to compute observables, overlaps, and information not directly accessible experimentally.

1. The RBM wave-function

The paper represents the many-body wave-function with a restricted Boltzmann machine containing visible physical variables and hidden stochastic neurons. Separate parametrizations support sampling amplitudes and modeling the phase.

  • An RBM uses visible neurons for physical variables and hidden stochastic binary neurons connected through weighted edges.Its expressive power is characterized by α = M/N, the ratio of hidden to visible neurons.
  • The RBM probability over visible variables is obtained by marginalizing over hidden degrees of freedom.The parameters include interlayer weights and visible and hidden biases.
  • The wave-function parametrization uses two parameter sets, λ and µ, with φµ(σ) = log pµ(σ) defining the phase component.
  • Sampling configurations uses only the amplitude distribution pλ(σ)/Zλ and can be performed efficiently through block Gibbs sampling.The bipartite architecture allows all units in one layer to be sampled simultaneously.

2. Gradients of the total divergence

Training begins with datasets of measurements in multiple bases and optimizes the neural-network parameters against their statistical distributions. Amplitudes and phases are learned in separate stages using different measurement information and optimization requirements.

  • Training datasets contain independent measurements from basis-dependent distributions Pb(σ[b]) ∝|Ψ(σ[b])|2.
  • The negative log-likelihood objective is constructed from the measurement datasets, while rotated neural-network states model observations in auxiliary bases.Basis transformations are evaluated efficiently when they act non-trivially on only a limited number of qubits.
  • The total divergence is differentiated with respect to network parameters using gradients derived from the RBM distributions and pseudo-averages.
  • Averages requiring the intractable normalization constant are approximated using samples generated by Markov-chain Monte Carlo.
  • The simplified training first learns amplitudes from the reference basis, then fixes them and learns phase parameters from auxiliary-basis measurements without Monte Carlo sampling from the neural network.

3. Training the neural network

The training procedure uses stochastic gradient descent for amplitude parameters and natural gradient descent for phase parameters. The latter accounts for the nonlinear parameter dependence of the RBM distribution through the Fisher information matrix.

  • Each parameter is updated by stochastic gradient descent using a learning rate and gradients averaged over randomly sampled mini-batches.
  • Natural gradient descent is used for phase learning because it is more effective than ordinary gradient descent, at increased computational cost.
  • The Fisher information matrix adapts the effective learning rate to nonlinear changes in the RBM distribution and can accelerate optimization.

4. Training datasets

The study benchmarks NN-QST using artificial datasets of independent measurements and exact or Monte Carlo sampling of quantum ground-state distributions. For TFIM and XXZ systems, path-integral Monte Carlo maps the quantum problem to a higher-dimensional classical one for data generation.

  • Dataset construction: NN-QST is benchmarked on artificial datasets composed of independent measurements from projections of the wave-function into multiple bases.Exact sampling is used when the system is sufficiently small or the wave-function is simple, such as for the W state.
  • Quantum sampling: Path-integral Monte Carlo samples the exact ground-state density distribution |Ψ(σ)|2 for the TFIM and XXZ models.The method maps d-dimensional quantum spin systems onto (d + 1)-dimensional classical systems.
  • Quantum sampling: Classical Metropolis Monte Carlo on the enlarged system collects samples in the {σ} basis at sufficiently large inverse temperature.The simulations use β = 10−20 and converged Mτ = 1024−2048, with independent samples collected beyond the autocorrelation time.

Appendix B: Cases of Study

This appendix section introduces details of the training procedures and measurements for the physical systems studied in the main paper.

  • Training datasets: The section describes training details for the physical systems investigated in the main paper.
  • Measurements: The section describes measurement details for the physical systems investigated in the main paper.
  • Scope: The appendix focuses on implementation details concerning training and measurements rather than introducing a new physical system.

1. W state

The W-state case studies test RBM tomography for amplitude-only and phase-bearing states, using specialized measurement bases and overlap-based performance evaluation. For phase-shifted W states, the full wave-function is learned from supplementary bases.

  • 1. W state: Because the W-state coefficients are real and positive, the RBM learns only amplitudes using one parameter set.
  • 1. W state: The training performance is quantified by computing the overlap between the W-state wave-function and the RBM wave-function.
  • 1. W state: For locally phase-shifted W states, the RBM learns both amplitudes and phases using the full wave-function.
  • 1. W state: The phase-bearing W-state construction uses supplementary bases containing neighboring X-X and X-Y measurements.
  • 1. W state: The X-X and X-Y bases encode phase differences through cosine- and sine-dependent probability relations, respectively.

2. Magnetic observables of local Hamiltonians

For TFIM and XXZ ground states, whose wave-functions are real and positive, QST restricts RBM learning to amplitudes and evaluates reconstruction through magnetic observables. Diagonal and sparse off-diagonal observables can be estimated from RBM samples.

  • 2. Magnetic observables of local Hamiltonians: The TFIM and XXZ ground-state wave-functions are real and positive, so QST learns amplitudes with one RBM parameter set.
  • 2. Magnetic observables of local Hamiltonians: Diagonal operators are evaluated by sampling configurations directly from the RBM distribution.
  • 2. Magnetic observables of local Hamiltonians: Sparse off-diagonal operators remain estimable through a local estimate computed from their matrix representation in the sampling basis.
  • 2. Magnetic observables of local Hamiltonians: For TFIM, the transverse-field magnetization ⟨σx⟩ is compared with a QMC estimate obtained using the path-integral formulation for non-diagonal operators.

3. Unitary evolution

The approach also reconstructs quantum states produced by unitary time evolution, including their complex-valued phases. It trains the full RBM wave-function from spin-density measurements in selected measurement bases.

  • The method studies quantum quenches, in which an initial state evolves under a Hamiltonian to a time-dependent state.
  • For a fixed time, spin-density measurements are used to train the RBM to learn the evolved wave-function.
  • Because time evolution makes the state complex-valued, the full RBM wave-function is used for quantum-state tomography.
  • Measurements in bases with one local X or Y rotation, alongside Z measurements, suffice to reconstruct the wave-function phases.

4. Entanglement entropy

The method estimates second Rényi entanglement entropy from an RBM wave-function using two copies and Swap-operator expectation values. An improved ratio trick reduces sampling noise for larger subregions.

  • For a bipartition into subregion A and complement A⊥, the entropy is obtained from the reduced density matrix of A.
  • The replica construction represents basis states separately for A and A⊥ before evaluating the Swap expectation value.
  • Second Rényi entropy is evaluated using a replica trick based on two copies of the physical system and the Swap operator.
  • The standard Swap expectation value becomes very small for larger subregions, producing high sampling noise.
  • The improved ratio trick computes entropy for a one-dimensional subregion by combining ratios of Swap expectation values for successively larger regions.
  • Monte Carlo samples two-copy spin configurations from a defined probability distribution to estimate the required expectation values.

Appendix C: Overfitting

The appendix examines whether RBMs overfit when learning the W state by varying training-set size, model capacity, and training progress. It uses overlap and held-out likelihood to assess reconstruction quality and generalization.

  • Overfitting can occur when an RBM is excessively powerful or trained on statistically small datasets.
  • For small α, RBMs poorly approximate the W state, while increasing α helps once the training dataset becomes sufficiently large.
  • Figure 6(a) measures W-state/RBM wave-function overlap against training-set size Ns for different hidden-unit densities α at N = 20.
  • The held-out likelihood approaches the W-state theoretical optimum near training’s end, with no observed evidence of overfitting.
Loading 1703.05334v2…