Source-linked AI summary

Supervised learning in physical networks: From machine learning to learning machines

Menachem Stern, Daniel Hexner, Jason W. Rocks, Andrea J. Liu

arXiv:2011.03861v2cond-mat.softcond-mat.dis-nncond-mat.stat-mech

TL;DR

Physical networks are usually designed or tuned using global information, whereas autonomous local learning remains difficult for some over-constrained systems. This paper introduces coupled learning, using free and clamped states to derive local rules for flow and elastic networks. The framework is physically plausible and can support autonomous adaptation, including flow-network image classification and learning with large nudges.

  • Problem

    Existing directed-aging methods fail in some over-constrained physical networks, while global supervised tuning requires nonlocal information and microscopic intervention.

  • Method

    Coupled learning derives local physical update rules by comparing network responses in free and clamped steady states.

  • Results

    Coupled learning was applied to flow and mechanical networks, including flow-network MNIST classification and successful learning with large nudges η ∼1.

  • Takeaways & Limitations

    Because the rules use local responses, coupled learning can in principle enable scalable, in-situ autonomous adaptation without full microscopic characterization.

Abstract

from arXiv · show

Materials and machines are often designed with particular goals in mind, so that they exhibit desired responses to given forces or constraints. Here we explore an alternative approach, namely physical coupled learning. In this paradigm, the system is not initially designed to accomplish a task, but physically adapts to applied forces to develop the ability to perform the task. Crucially, we require coupled learning to be facilitated by physically plausible learning rules, meaning that learning requires only local responses and no explicit information about the desired functionality. We show that such local learning rules can be derived for any physical network, whether in equilibrium or in steady state, with specific focus on two particular systems, namely disordered flow networks and elastic networks. By applying and adapting advances of statistical learning theory to the physical world, we demonstrate the plausibility of new classes of smart metamaterials capable of adapting to users' needs in-situ.

I. INTRODUCTION

The paper contrasts globally designed physical networks with autonomous local learning and introduces coupled local supervised learning to train flow and mechanical networks through physically plausible responses. The framework addresses failures of directed aging in over-constrained networks while targeting scalable, in-situ adaptation.

  • I. INTRODUCTION: Global supervised learning optimizes a cost function using microscopic system information, whereas local learning adapts network parts from information available in their immediate vicinity.In physical networks, global tuning can require external microscopic intervention, while local learning is autonomous.
  • I. INTRODUCTION: The target task is to produce tailored responses at target edges or nodes when external constraints are applied at source edges or nodes.Examples include target strain caused by source strain in mechanical networks and target pressure drops caused by source pressure drops in flow networks.
  • I. INTRODUCTION: Directed aging can train some physical networks but fails particularly in flow and highly coordinated mechanical networks because it minimizes the cost of a desired state rather than the desired-function cost.This mismatch explains why a successful local method does not generalize to these over-constrained systems.
  • I. INTRODUCTION: Coupled local supervised learning derives local update rules for flow and mechanical networks that are intended to be as successful as global supervised learning.The rules adjust learning degrees of freedom using local conditions such as spring tension or pipe current.
  • I. INTRODUCTION: Coupled learning compares free and clamped steady states, applying source constraints alone versus source and target constraints together.The approach is inspired by contrastive Hebbian learning and equilibrium propagation.
  • I. INTRODUCTION: The framework is presented as experimentally plausible, including approximate rules, while addressing measurement noise, drifting learning degrees of freedom, and network-size effects.The authors connect these properties to scalable training and metamaterials that adapt in situ.

A. Coupled learning in flow networks

Coupled learning trains disordered flow networks through a strictly local conductance-update rule, using free and slightly nudged clamped states to make target pressures approach desired responses. The networks generally converge successfully, while convergence depends on network size and task geometry.

  • A. Coupled learning in flow networks: Strictly local conductance updates train flow networks to produce desired target pressures from constrained source pressures.Each pipe changes according to local physical responses, without requiring the network to encode the desired function beforehand.
  • A. Coupled learning in flow networks: The free state equilibrates source-constrained pressures, while the clamped state nudges target pressures toward desired values before conductances are updated from the two states’ power difference.The update uses a small target nudge and a learning rate; derivatives of physical variables cancel in the small-nudge limit.
  • A. Coupled learning in flow networks: Target-pressure error decreases exponentially by many orders of magnitude to machine precision, while conductance changes also decay exponentially during training.The shrinking update magnitude indicates that learning slows as the network approaches a good solution.
  • A. Coupled learning in flow networks: Because flow networks implement only non-negative-conductance linear mappings, some tasks remain unrealizable and non-zero final errors are expected.Larger networks are expected to be more likely to succeed because they provide more degrees of freedom, but source-target geometry also matters.
  • A. Coupled learning in flow networks: Uniform conductance initialization k_j = 1 can train successfully, with mildly randomized initial conductances producing qualitatively similar accuracy and training-time behavior.The tested random initialization has variance σ^2 with 0 < σ ≤ 0.5.

B. Coupled Learning in general physical networks

The coupled-learning framework extends from flow networks to general athermal physical networks whose states minimize a physical cost. A small supervisor nudge yields a learning rule based on direct edge-level cost differences, guaranteeing locality while leaving the desired response external to the network.

  • B. Coupled Learning in general physical networks: For any athermal physical network with a known cost function, coupled learning defines free and clamped states and derives an iterative learning rule for edge parameters.Physical variables equilibrate by minimizing the cost, while source and target constraints distinguish the two states.
  • B. Coupled Learning in general physical networks: The desired response need not be encoded in the network: an external trainer supplies it by nudging target variables toward their desired values at each iteration.Learning continues iteratively until the desired function is achieved.
  • B. Coupled Learning in general physical networks: The resulting rule is local because the cost decomposes into edge terms depending only on each edge’s parameter and its two attached nodes.Only direct derivatives with respect to learning parameters are retained; physical-state derivatives cancel between free and clamped states.

C. Elastic networks

The coupled learning framework trains elastic networks to produce desired target strains by locally modifying spring parameters. Both spring constants and rest lengths support successful learning across network sizes.

  • C. Elastic networks: Spring networks minimize elastic energy while adapting edge lengths to map source strains onto desired target strains.The network uses spring constants and equilibrium lengths as adjustable learning degrees of freedom.
  • C. Elastic networks: Elastic networks learn desired target strains by modifying spring constants, rest lengths, or both through local spring-specific updates.Each spring changes according to the strain on that spring.
  • C. Elastic networks: Training reduces target-length error by orders of magnitude for both stiffness-based and rest-length-based learning.The result holds across multiple source-target choices and network sizes.
  • C. Elastic networks: N = 512 and N = 1024 networks achieve similar learning success, although larger networks require slightly more iterations.Rest-length learning produces more consistent results, reflected by smaller error bars across network and source-target choices.

D. Supervised classification

The authors apply coupled learning to classify MNIST images of handwritten 0s and 1s using a flow network with source pressures and two target nodes. The trained network generalizes to test images with high accuracy.

  • D. Supervised classification: The network learns multiple input-output mappings by presenting one randomly selected training image per iteration and nudging target pressures toward the desired digit label.Each image is represented by its top 25 principal components applied as source pressures.
  • D. Supervised classification: The flow network reaches 98% training accuracy and 95% test accuracy when classifying MNIST digits 0 and 1.Logistic regression on the same data reaches 100% training and 98% test accuracy.
  • D. Supervised classification: Classification error decreases on both training and test sets, demonstrating generalization beyond the presented training images.The predicted label is the target node with the larger pressure.

III. EXPERIMENTAL CONSIDERATIONS

The paper examines what physical implementation would require for coupled learning in flow and elastic networks. It identifies material changes, experimental approximations, noise, drift, and relaxation as practical considerations.

  • III. EXPERIMENTAL CONSIDERATIONS: Autonomous learning in elastic networks requires physical changes to bond stiffnesses or rest lengths.Materials including EVA and thermoplastic polyurethane are identified as adapting stiffness in response to applied strain.
  • III. EXPERIMENTAL CONSIDERATIONS: Existing adaptive materials provide possible physical mechanisms for changing elastic-network stiffness during learning.EVA has previously been used to train materials for auxetic behavior through directed aging.
  • III. EXPERIMENTAL CONSIDERATIONS: The authors organize implementation considerations around an experimentally simpler flow-network rule and the effects of noise, drifting learning variables, and network size.These issues are treated as potential difficulties in realizing coupled learning physically.

A. Approximate learning rules

The paper develops an experimentally simpler approximation to the coupled learning rule for flow networks. Although less effective than the full rule, the approximation remains correlated with optimal tuning and successfully trains networks.

  • A. Approximate learning rules: The approximation replaces the full free-versus-clamped power difference with a hybrid δ-state and a simpler sign-based update.Further restriction can enforce conductance changes in only one direction.
  • A. Approximate learning rules: Full and approximate rules train networks on similar time scales, despite the approximate rule being more restricted.The approximate rule often saturates at a higher error level across random source-target sets.
  • A. Approximate learning rules: The sign-based intuition decreases a pipe’s conductance when its free-state pressure drop indicates excessive flow through that pipe.This rule uses local pressure information to guide the conductance update.
  • A. Approximate learning rules: The approximate rule has lower correlation with optimal tuning than the full rule, yet its correlation remains significant across η values.The comparison uses dot products with the gradient or optimal local tuning direction.
  • A. Approximate learning rules: The approximate rule successfully reduces flow-network error by at least an order of magnitude, but typically plateaus above the full rule’s final error.This result is reported for networks with 512 nodes, 10 sources, and 3 targets.

B. Noise in physical or learning degrees of freedom

Measurement noise and drift in learning parameters can impede coupled learning, but suitable choices of nudge amplitude and learning rate mitigate these effects.

  • Increasing η improves training when measurement noise makes small nudges ineffective.The free–clamped signal scales as η^2, so small nudges are more easily dominated by measurement noise.
  • Measurement errors create a trade-off because small nudges improve rule accuracy but can spoil learning when noise dominates.
  • Drift in the learning degrees of freedom acts as additive noise in the update rule and is expected to impede learning.
  • Higher learning rates improve training under constant diffusion of the learning degrees of freedom.The paper notes that α cannot be increased without regard to convergence.

C. Physical relaxation

Coupled learning is studied in the quasi-static limit, where networks equilibrate before each update, but finite relaxation and information-propagation times constrain training speed and scaling.

  • The quasi-static protocol waits for stable equilibrium or steady state before computing free and clamped responses and updating learning variables.This assumes learning is much slower than the physical dynamics.
  • Small nudges do not generally accelerate physical relaxation because relaxation time does not scale with perturbation amplitude in the linear-response regime.
  • Relaxation requires local perturbations to propagate through the network, with propagation speed set by the speed of sound.In flow networks this depends on fluid compressibility; in elastic networks it depends on node masses and spring constants.
  • Training larger networks requires longer waiting times, with τ′ ≈ τf(L′/L) and f faster than linear.

IV. COMPARISON OF COUPLED LEARNING TO OTHER LEARNING APPROACHES

Coupled learning provides a local alternative to globally optimized tuning: it matches tuning closely in flow networks, remains substantially aligned in nonlinear elastic networks, and generalizes directed aging by retaining free-state information.

  • Local versus global supervised learning: Global tuning is not an autonomous physical learning rule because its updates depend on network-wide information and the desired response.Computing the gradient and modifying conductances therefore requires external intervention.
  • Local versus global supervised learning: Coupled learning differs from tuning because target values enter through clamped boundary conditions rather than explicitly into each local update.After boundary conditions are applied, each conductance uses only its own local response, while the physics propagates their effect through the network.
  • Flow and elastic network comparison: Coupled learning is as efficient as gradient-descent tuning for training flow networks and can converge faster on a demonstrated task.The comparison uses N = 512-node networks with random sources and targets.
  • Flow and elastic network comparison: For sufficiently small η, coupled-learning and tuning modification vectors are nearly identical even in the high-dimensional conductance space.The normalized dot product is close to unity for flow networks, although fluctuations occur during training.
  • Flow and elastic network comparison: In nonlinear elastic networks, coupled-learning and tuning updates remain significantly aligned, especially when stiffnesses are trained.Alignment decreases near low-error solutions, particularly when rest lengths are trained, while network training remains successful.
  • Relation to directed aging: Coupled learning generalizes directed aging by including both free and clamped terms, enabling training of network classes where directed aging fails.Directed aging omits the free term and is unsuccessful in many flow and highly coordinated elastic networks.

V. DISCUSSION

Coupled learning provides physically plausible local rules for training flow and mechanical networks, with potential advantages in scalability, in-situ adaptation, and user-directed task selection. The framework also extends beyond one source-target function, although multi-function training remains dependent on training-example order and frequency.

  • V. DISCUSSION: Coupled learning uses local responses to train physical networks, offering rules that can be implemented without directly optimizing a global cost function.The framework is applied to flow and elastic networks and is described as physically plausible because updates respond to local conditions.
  • V. DISCUSSION: Local responses may make training scalable because networks of different sizes can require approximately similar training-step times.The discussion contrasts this with gradient computation for collective cost functions, whose time can increase rapidly with system size.
  • V. DISCUSSION: In-situ training can avoid detailed knowledge of network geometry, topology, and microscopic physical or learning degrees of freedom.This is particularly valuable for experimental systems that need not be fully characterized before training.
  • V. DISCUSSION: A user can train the network for desired tasks without individually manipulating every edge's learning degree of freedom.The authors suggest that the supervisor can therefore be an end user rather than an expert designer.
  • V. DISCUSSION: Coupled learning can train a flow network to distinguish MNIST digits, but multi-function performance may depend strongly on the order and frequency of training examples.The authors leave the dynamics of learning multiple functions for subsequent work.
  • V. DISCUSSION: Physical networks are not intended to outperform computational neural networks, but symmetric linear-response mappings may support generative models.The paper leaves physical learning of such generative models for future study.

Appendix A: Nudge amplitude η

Small nudges make coupled learning more consistent and efficient because they keep free and clamped states close enough to approximate gradient descent. Large nudges can slow or disrupt learning, while non-negative learning parameters impose an additional architectural constraint shared by coupled learning and gradient descent.

  • Appendix B: Effective cost function: η ≪1 keeps free and clamped states close, allowing the local rule to approximate derivatives and mimic global cost-function optimization.The small-nudge rule can be interpreted through an effective cost function whose gradient is substantially aligned with the target cost gradient.
  • Appendix A: Nudge amplitude η: At η ≈0.6, one 512-node mechanical network undergoes an abrupt slowdown by orders of magnitude as learning crosses between attractor basins.At lower η, the network behaves more like a linear flow network; at higher η, the free and clamped states can occupy different basins.
  • Appendix A: Nudge amplitude η: η ≪1 yields superior results, whereas η ∼1 usually deteriorates learning efficiency in both linear and nonlinear networks.Small nudges support continuous learning, while large nudges can produce intermittent unlearning events.
  • Appendix D: Non-negative learning parameters: Non-negativity excludes some linear source-target mappings, so linear flow networks may be less successful than general linear models for certain tasks.Pipe conductances and spring parameters are constrained to non-negative values.
  • Appendix D: Non-negative learning parameters: Non-negative learning parameters can reach cutoff values during training, and harder tasks with more simultaneous targets induce more vanishing edges.A hard N = 512, M_S = 10, M_T = 10 task reduced the initial error by only a factor of 3, with the limitation shared by coupled learning and gradient descent.
  • Appendix D: Non-negative learning parameters: The non-negative constraint slows learning in both coupled learning and standard gradient descent.The shared slowdown suggests that this difficulty arises from the network's physical or architectural constraints rather than uniquely from the coupled-learning rule.
Loading 2011.03861v2…