Source-linked AI summary

Active learning machine learns to create new quantum experiments

Alexey A. Melnikov, Hendrik Poulsen Nautrup, Mario Krenn, Vedran Dunjko, Markus Tiersch, Anton Zeilinger, Hans J. Briegel

arXiv:1706.00868v3quant-phcs.AIstat.ML

TL;DR

The paper asks whether intelligent machines can contribute to scientific research by designing quantum experiments, addressing the poorly understood reachability of entanglement classes. It applies projective simulation to simulated photonic laboratories, where an agent learns to assemble setups for high-dimensional entangled states. The agent improves experiment efficiency, discovers varied states and useful techniques, and outperforms automated random search in finding interesting experiments.

  • Problem

    The reachability and structure of multipartite high-dimensional entangled states in quantum experiments remain insufficiently understood, while the broader research role of artificial intelligence beyond data analysis is largely unexplored.

  • Method

    The paper uses a projective simulation reinforcement-learning agent that interacts with simulated optical tables, placing elements and receiving rewards for producing target entangled states or new interesting experiments.

  • Results

    The agent learns to construct high-dimensional entangled states, improves experiment efficiency, discovers useful experimental techniques and sub-setups, and finds significantly more interesting experiments than automated random search.

  • Takeaways & Limitations

    The results support a role for learning machines in autonomously designing quantum experiments and contributing emergent, potentially creative experimental tools.

Abstract

from arXiv · show

How useful can machine learning be in a quantum laboratory? Here we raise the question of the potential of intelligent machines in the context of scientific research. A major motivation for the present work is the unknown reachability of various entanglement classes in quantum experiments. We investigate this question by using the projective simulation model, a physics-oriented approach to artificial intelligence. In our approach, the projective simulation system is challenged to design complex photonic quantum experiments that produce high-dimensional entangled multiphoton states, which are of high interest in modern quantum experiments. The artificial intelligence system learns to create a variety of entangled states, and improves the efficiency of their realization. In the process, the system autonomously (re)discovers experimental techniques which are only now becoming standard in modern quantum optical experiments - a trait which was not explicitly demanded from the system but emerged through the process of learning. Such features highlight the possibility that machines could have a significantly more creative role in future research.

I. INTRODUCTION

The introduction asks whether intelligent machines can move beyond data analysis to help create and understand quantum experiments. It proposes projective simulation as a learning-agent framework for discovering photonic setups and useful experimental substructures.

  • I. INTRODUCTION: Machine learning is established as a physics data-analysis tool, but the broader role of artificial intelligence in scientific research remains less explored.The paper frames intelligent agents as devices that interact with a laboratory environment and learn from it.
  • I. INTRODUCTION: The study applies reinforcement learning to an agent that interacts with simulated optical tables and learns to generate novel experiments.The task is formulated within the projective simulation framework.
  • I. INTRODUCTION: The projective simulation agent builds a memory network of correlations between optical components and exploits it to design targeted experiments efficiently.The network represents combinations of components that work well together.
  • I. INTRODUCTION: The agent autonomously discovers useful sub-setups or gadgets as a byproduct of learning, although this capability was not explicitly demanded.The paper presents this emergent behavior as relevant to a more creative role for machines in research.
  • I. INTRODUCTION: The target problem is high-dimensional multipartite entanglement, whose reachability and structure in optical experiments remain incompletely understood.Such states matter both for fundamental quantum mechanics and for quantum communication and computation.

II. RESULTS

The simulated laboratory lets a projective simulation agent assemble optical experiments, receive task-dependent rewards, and learn which actions produce desired entangled states. The setup uses a constrained photonic toolbox and reward criteria based on entanglement properties.

  • II. RESULTS: The agent sequentially places optical elements on a simulated table, while an analyzer evaluates the resulting quantum state and supplies rewards.Experiments end when a reward is obtained or when the maximum number of elements is reached.
  • II. RESULTS: The reward initially targets high-dimensional multipartite entangled states characterized by a Schmidt-Rank Vector and additional stated properties.The Schmidt-Rank Vector records the ranks of the reduced density matrices of the subsystems.
  • II. RESULTS: The initial photonic state comes from double spontaneous parametric down-conversion and is modeled as two pairs of low-order OAM-entangled photons.Higher-order OAM terms with |m| > 1 are neglected because their amplitudes are significantly smaller.
  • II. RESULTS: The toolbox contains beam splitters, mirrors, shift-parametrized holograms, and Dove prisms, yielding 30 distinct element-placement choices.Each element can be placed in one or, for beam splitters, two of the four optical arms.
  • II. RESULTS: Projective simulation represents percepts and actions as clips connected by weighted directed edges whose probabilities are adjusted during learning.The agent’s memory network links observed optical tables to possible element-placement actions.

A. Designing short experiments

The first experiment tests whether learning can produce short optical setups for specified entanglement classes. Prior learning on a simpler target substantially improves subsequent construction of a more complex target, while the agent also shortens experiments over time.

  • A. Designing short experiments: The agent first designs setups producing SRV (3, 3, 2), then is asked to construct SRV (3, 3, 3) states within 6 × 10^4 experiments.The maximum experiment length is L = 8.
  • A. Designing short experiments: Without prior training on SRV (3, 3, 2) states, the agent shows no significant progress toward constructing an SRV (3, 3, 3) experiment.The comparison uses simulations beginning with and without prior learning of the simpler target.
  • A. Designing short experiments: With prior training, the agent learns to design an SRV (3, 3, 3) setup very quickly, with probability 0.5 in the reported simulations.The authors suggest that the earlier knowledge may benefit the second phase, while also noting that the more complicated state might be easier to generate.
  • A. Designing short experiments: The projective simulation agent continually improves by constructing shorter setups while searching for rewarded experiments.The figure tracks average experiment length and success probability across the learning process.

B. Designing new experiments

The second experiment rewards discovery of distinct high-dimensional three-photon entangled states. Projective simulation discovers significantly more interesting experiments than automated random search, including when action composition is available.

  • B. Designing new experiments: A database of generated experiments could help reveal the structure of optical setups that produce high-dimensional entangled states and identify recurring useful gadgets.The connection between entangled states and the optical-table structures generating them is not well understood.
  • B. Designing new experiments: The agent receives a reward for each new implementation of an interesting experiment whose Schmidt-Rank Vector was not previously reached in that experiment.This prevents trivial extensions of already-known implementations from receiving rewards.
  • B. Designing new experiments: Action composition lets the agent form composite actions that place multiple elements in fixed configurations, effectively enhancing its toolbox.The mechanism is associated with a primitive notion of creativity in the paper.
  • B. Designing new experiments: The projective simulation model discovers significantly more interesting experiments than automated random search, both with and without action composition.The comparison is shown in Fig. 2(b) using solid and dashed curves for the respective methods.

C. Ingredients for successful learning

Successful learning is supported by structure in the optical-setup environment: experiments form a maze-like graph, and interesting setups cluster in regions of that graph.

  • The PS agent succeeds because the task environment contains hidden structure that learning can exploit.The paper presents successful learning as relying on structure hidden in the task environment or dataset.
  • Optical experiments form a directed graph in which placing elements is a walk, while multiple setups can produce the same quantum state.Because setups are not unique, the graph resembles a maze rather than a tree.
  • 45,605 nodes from 1.6 × 10^4 experiments included 67 interesting setups in the explored setup space.The graph used up to six optical elements from a toolbox of 30 elements.
  • Interesting experiments appear clustered in high-density regions, so previously learned setups can help the agent tackle similar experiments.The paper connects this clustering to reinforcement learning’s usefulness in handling situations similar to those previously encountered.

D. The potential of learning from experiments

The learning agent extends beyond solving specified targets by reusing and recombining optical techniques, including parity sorters and dimensionality-increasing devices, while discovering experimentally relevant configurations.

  • The paper asks whether machines can gain research-relevant insight beyond designing experiments for precisely specified tasks.The comparison is framed around whether a machine can identify useful techniques, not merely search all conceivable setups.
  • The PS agent composed and extensively used an interferometer corresponding to a parity sorter, an essential component in many high-dimensional entanglement experiments.The parity sorter is especially associated with experiments involving more than two photons.
  • Rewarding only novel experiments pushed the agent away from repeatedly selecting the original parity sorter and toward different-looking configurations.These configurations were described as similarly useful.
  • The agent frequently selected a nonlocal parity sorter rather than the original Mach–Zehnder form, with the two setups equivalent in the Klyshko wave-front picture.The wave-front transformation reduces the four-photon nonlocal setup to a two-photon experiment.
  • In simulations of 100 agents, setups corresponding to the local sorter, nonlocal sorter, and dimensionality-increasing device appeared among highest-weighted substructures 11, 22, and 43 times, respectively.Other substructures were highest-weighted in 24 cases.
  • The discovered setups were modern quantum-optical gadgets with existing applications or potential use in experiments generating high-dimensional entanglement from lower-dimensional entanglement.The paper characterizes these devices as either already used in state-of-the-art experiments or potentially useful individually.

III. DISCUSSION

The discussion presents automated laboratories as a route toward machines contributing to research, combining improved experiment-search methods with behavior that can extend beyond explicit training tasks.

  • Smart laboratories could reduce human involvement in tedious or hazardous laboratory tasks as laboratory automation expands.The paper places this development within the broader growth of smart technologies.
  • The work asks whether machines could not only assist research but perhaps genuinely perform it.This question frames the paper’s discussion of automated laboratories.
  • More sophisticated learning agents improve on search algorithms for finding special optical setups, while future PS extensions may further improve autonomous experiment design.The paper specifically mentions generalization and meta-learning as possible extensions.
  • The basic PS reinforcement-learning machinery with action-clip composition may tackle problems that it was not directly instructed or trained to solve.The paper presents this as support for AI methods contributing to research.

1. Projective simulation

Projective simulation represents percepts and actions as a weighted clip network, selecting actions probabilistically and updating transition weights from environmental feedback while incorporating memory dynamics.

  • The PS model uses episodic and compositional memory: a weighted network of clips representing remembered percepts, actions, or short sequences.In this paper, clips include observed optical tables and actions corresponding to placing optical elements.
  • Each directed edge connects a percept to an action, and its h-value determines the transition probability from that percept to that action.The network contains percepts s_i and actions a_j, with time-dependent edge weights h_ij.
  • Initially, all h-values equal 1, so the agent chooses actions uniformly and behaves randomly.The initial action probability is p_ij = 1/K.
  • Environmental feedback updates the h-matrix after each interaction to improve the probability of receiving a nonnegative reward.The learning rule changes transition weights based on feedback from the environment.
  • The glow matrix redistributes reward toward recent decisions, with traversed edges receiving glow and older decisions receiving less internal credit.Glow values are initialized at zero and set to 1 for edges traversed during the latest decision-making process.
  • The damping parameter γ controls forgetting, while η is the glow parameter governing glow decay.In this paper, γ and η were fixed at time-independent values without extensive optimization.
  • Action composition and clip deletion dynamically alter the structure of the PS network in addition to the basic model.

a. Action composition

Action composition lets the agent turn rewarded optical-element sequences into reusable single-step actions, modestly improving assigned-task performance while supporting creative behavior.

  • Action composition: Rewarded sequences of optical elements can be added to the toolbox as composite actions.These actions let the agent reach the corresponding rewarded experiments in one decision step instead of placing elements individually.
  • Action composition: Composite actions can be placed at any time during an experiment.
  • Action composition: Reward rescaling compensates for the growing toolbox and preserves transition-probability changes as the action space expands.Rewards from the analyzer are rescaled proportionally to the toolbox size K(t), relative to K(0).
  • Action composition: Action composition produces only a modest improvement on the assigned tasks but is vital for the agent’s creative features.

b. Clip deletion

Clip deletion keeps the projective-simulation network manageable as percept and action spaces grow, while preserving exploration of new successful sequences.

  • Clip deletion: Clip deletion removes infrequently used clips so the agent can handle large percept and action spaces.Percept and action creation use different underlying mechanisms, so their deletion mechanisms are treated separately.
  • Clip deletion: O(K^L) possible table configurations make storing the full percept space impractical.For K = 30 and L = 6, the simplest studied case contains more than 0.5 billion configurations.
  • Clip deletion: Unsuccessful experiments are deleted after L elements without positive reinforcement, removing their associated clips and edges.This keeps the PS network compact, with no more than 10^4 percept clips on average in the most complicated scenario described.
  • Clip deletion: Action deletion balances exploiting successful composite sequences against exploring new actions as the learned toolbox grows.Without deleting action clips, the action count can exceed K = 100, making basic elements less likely to be selected.
  • Clip deletion: Action deletion is triggered after rewards and uses incoming-edge weights to identify weakly supported actions for removal.The deletion probability is activated after a reward and deactivated until another reward is obtained.
  • Clip deletion: Newly created actions receive 10 updates of immunity because their initial deletion probability is high.Composite actions start with h-values of 1, making the initial quantity NR(t) tend to be below one.
  • Clip deletion: Automated random search uses the same composition and deletion mechanisms, but learns no additional actions because it has no other learning component.
  • Clip deletion: Figure 5 averages learning-performance curves over 100 agents across efficiency, interesting-experiment rates, and SRV diversity.Panels compare PS with automated random search as functions of prior discoveries, experiment count, and maximum length L.

2. Details of learning performance

The PS agent becomes increasingly effective at discovering useful photonic experiments, finding more diverse entangled states and shorter implementations than random search. Its exploration also reveals clustered structure among interesting setups.

  • Learning efficiency: Around 350 experiments are needed initially to find a (3, 3, 2) state, with this number decreasing as more such states are discovered.The agent also finds progressively shorter implementations of the same state.
  • Discovery rate: The PS agent finds many more interesting experiments than random search after 1.2×105 time steps.
  • Discovery rate: Unlike random search, PS agents sustain an increasing rate of discovering new interesting experiments for longer.Random search maintains an approximately constant rate before its fraction slowly decreases as previously found experiments are encountered.
  • Discovery rate: Random-search curves converge to a maximum rate of obtaining new useful experiments before decreasing again.
  • Diversity of outcomes: PS finds more different state representations (SRVs) on average than automated random search, although finding diverse SRVs was not an assigned task.
  • Exploration structure: The PS exploration graph contains 34155 nodes after 1.6 × 104 experiments, about one quarter fewer than the random-search maze, with interesting experiments clustered by shared SRVs.The setups use a 30-element toolbox and contain at most six optical elements; same-colored nodes share at least one SRV.
Loading 1706.00868v3…