Source-linked AI summary
The thermodynamic freedom of a thermodynamic computer
Stephen Whitelam
TL;DR
Thermodynamic computers face speed-limit constraints linking runtime, computational progress, and thermodynamic cost. This paper evaluates a trained stochastic MNIST classifier with the Wasserstein speed limit and tests alternative inference protocols. The computer retains comparable accuracy while approaching 40% of the thermodynamic efficiency bound or operating faster at fixed efficiency.
Problem
Thermodynamic computers require evaluation of how computational progress, runtime, and thermodynamic cost constrain their operation.
Method
The paper applies the Wasserstein speed limit to a simulated thermodynamic computer trained for MNIST classification and evaluates two inference protocols.
Results
The computer reaches 97.25 ± 0.09% MNIST accuracy using one trajectory per image, while protocol choice changes efficiency, speed, heat, and runtime.
Takeaways & Limitations
A thermodynamic computer trained for a particular task can perform it across a wide range of thermodynamic conditions.
Abstract
from arXiv · showhide
Thermodynamic computers are stochastic physical devices designed to perform calculations at the thermal energy scale. Their operation is constrained by the equations of stochastic thermodynamics, among which are a set of bounds, known as speed limits, that relate a thermodynamic computer's run time to its computational progress and the heat it dissipates. Using the Wasserstein speed limit we assess the thermodynamic efficiency of a simulation model of a thermodynamic computer trained to perform a standard machine-learning classification task. On this task the thermodynamic computer is as capable as a simple multilayer perceptron. We show that different inference protocols allow the computer to operate within 40\% of the thermodynamic limit of efficiency without loss of accuracy, or to perform inference increasingly rapidly at fixed accuracy and thermodynamic efficiency. These results indicate that a thermodynamic computer designed for a particular task retains considerable freedom in its thermodynamic operation.
I. INTRODUCTION
Thermodynamic computers use stochastic physical dynamics near the thermal energy scale, whose runtime, computational progress, and thermodynamic cost are constrained by speed limits. The paper applies the Wasserstein speed limit to a trained MNIST classifier and compares alternative inference protocols.
- Motivation: Thermodynamic computers perform calculations with physical devices operating near the energy scale of thermal fluctuations.Their dynamics can often be approximated by a Langevin equation.
- Speed limits: Speed limits constrain stochastic devices by relating process duration to computational progress and thermodynamic cost.For thermodynamic computing, the process duration can be interpreted as program runtime.
- Speed limits: The Wasserstein distance measures computational progress by quantifying how far the system’s probability distribution moves between initial and final states.Larger W2 indicates greater movement in probability space.
- Paper scope: The study uses the speed limit to assess a nonlinear thermodynamic computer trained to classify MNIST images.The computer is trained to operate at a predetermined observation time.
- Paper scope: A single nonequilibrium trajectory per image achieves multilayer-perceptron-comparable accuracy without repetition or averaging.Different inference protocols then trade efficiency, computation rate, heat emission, and runtime while preserving accuracy.
II. A MODEL THERMODYNAMIC COMPUTER
The simulated computer has 74 real-valued units governed by a programmable nonlinear potential and trained to classify MNIST images. Inference turns on the trained parameters, evolves the system for a fixed duration, and reads the largest-valued output unit as the prediction.
- Architecture: The model contains N = 74 real-valued degrees of freedom, including 64 hidden units and 10 outputs.The units are classical and represent the computer’s dynamical variables.
- Potential: A quartic on-site term gives each unit a nonlinear response, enabling computations that are not linearly separable.The model sets J2 = J4 = kBT.
- Parameterization: Adjustable biases, symmetric inter-unit couplings, and input couplings from 282 pixels encode the trained MNIST program.The inter-unit couplings connect all N(N − 1)/2 pairs of units.
- Inference protocol: Inference begins with independent equilibration, then activates trained parameters and evolves the computer until τ = 1/5 before reading its 10 outputs.The output with the largest value xc(τ) determines the predicted class.
III. DYNAMICS OF INFERENCE
The thermodynamic computer classifies MNIST images using single stochastic trajectories, achieving high overall accuracy while relying on a designated nonequilibrium readout time. Thermal noise affects individual outcomes, but training makes aggregate accuracy stable.
- Classification performance: 97.25 ± 0.09% accuracy is achieved on MNIST using one trajectory per image, comparable to a simple multilayer perceptron.Removing the quartic interaction reduces accuracy to 90.5%, indicating that nonlinear dynamics provide essential computational power.
- Stochastic reliability: Thermal noise changes which images are classified correctly, but repeated operation changes overall accuracy by only 0.09% standard deviation.In the noise-free limit, accuracy is 97.70%, only 0.45 percentage points above the value at kBT = 1.
- Classification performance: 86.3% of test images are always correctly classified, 0.45% are always misclassified, and 13.2% vary across stochastic runs.The variable subset has a mean success probability of 82.6%.
- Readout dynamics: Accuracy peaks at the trained observation time τ = 1/5 and diminishes at shorter and longer readout times.The computer remains functional at longer times, including after reaching steady state, but is most accurate at its trained readout time.
- Nonequilibrium operation: At the designated readout time, the number of locally unstable directions is still changing, so the computer is out of equilibrium.The system begins with about 34 locally unstable directions, and this number decreases continuously during operation.
IV. STOCHASTIC THERMODYNAMICS OF INFERENCE
The paper evaluates inference efficiency using a Wasserstein speed-limit formulation that compares data-averaged transport with thermodynamic cost. Two protocols expose a trade-off between efficiency, computation rate, and heat emission.
- Efficiency measure: η measures thermodynamic efficiency as data-averaged transport divided by data-averaged thermodynamic cost.It can also be interpreted as the ratio of the thermodynamic minimum runtime or entropy production to the computer’s actual value; η = 1 is ideal.
- Efficiency measure: The efficiency calculation averages thermodynamic quantities first over noise for each image and then over the dataset.The resulting ratio measures the efficiency of the computer’s workload.
- Operating protocols: Finite-rate loading increases efficiency η at the cost of a decreased computation rate.This protocol generalizes inference by turning on trained couplings over a finite time.
- Operating protocols: Rescaling the potential and adding controlled noise increases computation rate at fixed efficiency and entropy production, but increases heat emission.The two protocols therefore support operation across different thermodynamic regimes.
A. Under a loading protocol
Gradual loading reduces dissipation while preserving classification accuracy, creating a trade-off between thermodynamic efficiency and runtime. The optimal loading time reaches ηG = 0.61, within 40% of the thermodynamic limit.
- A. Under a loading protocol: Slower loading reduces entropy production by more than an order of magnitude without impairing classification accuracy.Accuracy changes very little and increases slightly as the loading rate decreases.
- A. Under a loading protocol: Gradual loading shifts energy expenditure toward work returned to the loading source rather than heat dissipation.Finite-time loading shifts the work distribution to W < 0, whereas a quench has zero-mean work.
- A. Under a loading protocol: Slower loading reaches a fixed accuracy at smaller mean energy loss, showing that accuracy is not fixed by the computer’s internal-energy change.The accuracy-versus-energy-loss curves for tℓ≤0.1 almost coincide.
- A. Under a loading protocol: ηG = 0.61 at the optimal loading time tℓ= 0.5, placing the computer within 40% of the thermodynamic efficiency limit.For slower loading, longer runtime outweighs reduced entropy production; for faster loading, increased entropy production outweighs shorter runtime.
B. Under the clock-acceleration procedure
Clock acceleration rescales the computer’s dynamics so inference occurs faster without changing trajectory statistics, accuracy, entropy production, or efficiency. Its thermodynamic cost is increased heat emission and mean dissipated power.
- B. Under the clock-acceleration procedure: Uniform potential rescaling and added Gaussian noise increase the computer’s basic rate without changing its program or stationary trajectory ensemble.The procedure is equivalent to renormalizing the mobility parameter as µ→λµ.
- B. Under the clock-acceleration procedure: The accelerated and original computers produce the same total entropy and have the same mobility-runtime product, so their efficiencies are equal.The instantaneous entropy-production rate increases by λ while runtime decreases by the same factor.
- B. Under the clock-acceleration procedure: Faster inference at fixed accuracy and efficiency requires increased mean power and heat emission.The heat emitted at the rescaled readout time increases in proportion to λ, while total entropy production remains unchanged.
- B. Under the clock-acceleration procedure: Clock acceleration compresses the dynamics in time while preserving the trajectory distribution and inference accuracy.Numerically, λ = 1, 2, 4, and 8 produce the same accuracy at observation times rescaled by 1/λ.
V. CONCLUSIONS
The thermodynamic computer maintains comparable accuracy while its efficiency and runtime vary substantially with the inference protocol. Gradual loading favors efficiency, whereas clock acceleration favors speed at increased heat and power cost.
- V. CONCLUSIONS: The computer classifies MNIST using one nonequilibrium dynamical trajectory per image with accuracy comparable to a simple multilayer perceptron.Its thermodynamic efficiency can vary through the inference protocol without significant change in operational accuracy.
- V. CONCLUSIONS: Gradual loading reduces heat and entropy production and brings efficiency within 40% of the thermodynamic limit, at the cost of greater runtime.The protocol changes thermodynamic operation without significant loss of accuracy.
- V. CONCLUSIONS: Clock acceleration enables faster operation at the same efficiency and entropy production, but increases dissipated heat and mean power.The speed-up changes the heat cost rather than total entropy production.
- V. CONCLUSIONS: A computer trained for one task can perform that task across a wide range of thermodynamic conditions.This complements reported energy-time-accuracy trade-offs and ranges of dissipative operation in thermodynamic logic gates.
Appendix A: Simulation details
The simulations use numerical integration and sampling procedures designed to estimate thermodynamic quantities and test robustness to numerical and sampling choices. Reported efficiency estimates are stable under the tested changes.
- Appendix A: Simulation details: The Langevin dynamics were integrated with the Euler–Maruyama method using times measured in units of µ^-1.The timestep was chosen separately for the standard and clock-acceleration simulations.
- Appendix A: Simulation details: Efficiency components were estimated from R = 2000 trajectories for 50 weighted test images, while accuracy and work/heat distributions used all 10^4 test images.Finite-R bias in G2s was removed by linear extrapolation in 1/R.
- Appendix A: Simulation details: Quadrupling the number of images, doubling R, or halving ∆t changes ηG by at most 0.003 and ΣGs and G2s by less than 0.6%.These checks were performed at the most accurate readout times for tℓ= 0, 0.1, and 0.5.
- Appendix A: Simulation details: The modified per-image efficiency agrees with the original data-averaged efficiency to within 1.1 × 10^-3 under loading and clock acceleration.This comparison tests whether averaging before or after taking the efficiency ratio materially changes the result.