Source-linked AI summary

Data Driven Computing with Noisy Material Data Sets

Trenton Kirchdoerfer, Michael Ortiz

arXiv:1702.01574v2physics.comp-ph

TL;DR

Material modelling from noisy, high-dimensional data can introduce error and uncertainty, while distance-minimizing Data Driven Computing can be dominated by outliers. The paper develops max-ent Data Driven Computing, which uses clustering, maximum-entropy relevance, and free-energy minimization under physical constraints; it generalizes the distance-minimizing scheme and reports convergence properties for its solvers and solutions.

  • Problem

    Material data can contain outliers that may unduly dominate distance-minimizing Data Driven solutions, while classical material modelling introduces error and uncertainty.

  • Method

    Max-ent Data Driven Computing assigns data points variable relevance through distance and maximum-entropy estimation, then minimizes free energy over phase space under compatibility and equilibrium constraints.

  • Results

    The max-ent paradigm generalizes distance-minimizing Data Driven Computing, is robust with respect to outliers, and has convergence properties assessed through numerical tests.

  • Takeaways & Limitations

    Distance-minimizing Data Driven schemes are recovered in the zero-temperature limit, linking the robust formulation to the earlier Data Driven approach.

  • Takeaways & Limitations

    The implementation computes sums over entire data sets at each material point, and the annealing schedule may still admit iteration and convergence-rate improvements.

Abstract

from arXiv · show

We formulate a Data Driven Computing paradigm, termed max-ent Data Driven Computing, that generalizes distance-minimizing Data Driven Computing and is robust with respect to outliers. Robustness is achieved by means of clustering analysis. Specifically, we assign data points a variable relevance depending on distance to the solution and on maximum-entropy estimation. The resulting scheme consists of the minimization of a suitably-defined free energy over phase space subject to compatibility and equilibrium constraints. Distance-minimizing Data Driven schemes are recovered in the limit of zero temperature. We present selected numerical tests that establish the convergence properties of the max-ent Data Driven solvers and solutions.

1 Introduction

The paper develops Data Driven Computing for scientific computing to use material data directly instead of calibrated empirical material models. It extends distance-minimizing schemes to noisy data with outliers through maximum-entropy clustering and free-energy minimization.

  • Classical material modelling calibrates empirical laws from observations, introducing error and uncertainty through imperfect laws, phase spaces, and noisy experimental data.
  • Data Driven Computing formulates initial-boundary-value problems directly from material data, bypassing the empirical material-modelling step.
  • Distance-minimizing Data Driven Computing selects the material-data point closest to satisfying the problem’s field equations.
  • Max-ent Data Driven Computing assigns variable relevance using solution distance and maximum-entropy estimation, then minimizes free energy under compatibility and equilibrium constraints.
  • The max-ent scheme is designed to remain robust to outliers through clustering, while recovering distance-minimizing schemes at zero temperature.

2 The Data Driven Science paradigm

Data Driven Computing separates universal field constraints from material-specific information supplied by finite data sets. Its distance-minimizing formulation relaxes empty intersections between data and constraints by selecting the closest admissible data point, with convergence to classical solutions under increasingly faithful sampling.

  • Conservation and localization laws impose universal, material-independent constraints, while material laws provide empirically determined material specificity.
  • Finite material data sets may not intersect the constraint set, so the exact solution set must be replaced by a relaxation.
  • Distance-minimizing Data Driven Computing selects the data point closest to the constraint set, equivalently solving a phase-space distance minimization.
  • 2.3 An elementary example: In the bar example, the phase space is the strain–stress plane and the constraint set is determined by device stiffness and applied displacement.
  • 2.4 Uniform convergence: As sampling density increases along a material graph, distance-minimizing solutions converge to the corresponding classical solution.

3 Probabilistic Data Driven schemes

The probabilistic scheme ranks material data points by relevance using maximum entropy and distance, then minimizes a free energy over the constraint set. This clustering-based formulation reduces the influence of outliers while recovering distance-minimizing Data Driven Computing as β→+∞.

  • Motivation: Distance-minimizing solvers can be dominated by a single outlier that lies unusually close to the constraint set.This motivates ranking data points by relevance and treating random material data sets probabilistically.
  • Data clustering: The scheme assigns weights p_i to quantify each material data point’s relevance to a phase-space point z.Relevance is represented by weights in [0,1].
  • Data clustering: Maximum-entropy estimation provides an unbiased relevance distribution while distance-based cost assigns less weight to points farther from z.The two objectives are combined through Pareto optimality.
  • Probabilistic Data Driven scheme: The resulting max-ent solver minimizes a free energy over the constraint set C.The free energy incorporates the relevance-weighting formulation.
  • Data clustering: β^-1/2 sets the effective neighborhood width, so distant points and outliers have negligible influence on the solution.The solution is dominated by the local cluster within that neighborhood.
  • Data clustering: As β→+∞, max-ent Data Driven Computing recovers the distance-minimizing scheme, while finite β retains influence from all data points with distance-dependent weights.A marginally closer outlier therefore does not significantly alter the solution.

4 Numerical implementation

The numerical implementation uses fixed-point iterations and simulated annealing to minimize the generally non-convex free energy. Convexity at small β and contractivity conditions guide the annealing schedule and establish local convergence behavior.

  • Fixed-point iteration: The max-ent problem is solved by minimizing free energy F(z) over the constraint set C, with optimality expressed through a fixed-point formulation.For constraint sets represented by f(z)=0, the formulation can use closest-point projection and a Lagrange multiplier.
  • Numerical implementation: The free energy is generally strongly non-convex, with multiple wells that can cause iterative solvers to fail or return local minimizers.This is the central numerical difficulty addressed by simulated annealing.
  • Simulated annealing: For β→0, the Hessian approaches the identity and the free energy is convex.This provides the starting regime for the annealing strategy.
  • Simulated annealing: Simulated annealing starts with sufficiently small β and increases it according to a schedule intended to guide the solver toward the absolute minimizer.The schedule is selected to maintain local contractivity of the fixed-point mapping.
  • Fixed-point iteration: Fixed-point convergence follows when the mapping is contractive; convexity of C makes the closest-point projection contractive.The paper gives conditions under which the remaining mapping is contractive in a neighborhood of the solution.
  • Simulated annealing: As β increases, the solution becomes controlled by an increasingly smaller local data cluster, eliminating outlier influence.Initially, the schedule allows all data points to influence the solution.

5 Numerical tests

Numerical truss tests examine max-ent Data Driven Computing under annealing and noisy-data conditions. The solver converges toward reference or classical solutions, with accuracy depending on annealing aggressiveness and noise distribution.

  • Test case: The truss test case contains 1,246 members, with reference states obtained from a nonlinear stress-strain model and Newton-Raphson solution.Member states are superimposed on the reference stress-strain curve to visualize phase-space coverage.
  • Annealing schedule: The simulated-annealing solver uses β schedules whose iteration counts increase with data-set size as larger sets require broader exploration.For λ = 0.01, β grows roughly linearly before rapidly diverging at a data-set-size-dependent iteration.
  • Annealing schedule: Larger λ accelerates convergence but can prematurely freeze the solution in a non-optimal cluster, whereas smaller λ improves exploration and accuracy.The trade-off is reported by comparing λ = 0.01 and λ = 0.1 schedules.
  • Annealing schedule: For coarse data sets, solution error is relatively insensitive to λ; with larger data sets, conservative annealing more effectively identifies accurate local clusters.A distance-minimizing iteration after quenching further improves the solution compared with stopping immediately at quenching.
  • Uniform convergence of a noisy data set towards a classical material model: For capped normal noise shrinking as N^-1/2, sufficiently small λ yields max-ent solution convergence of order N^-1, faster than the distance-minimizing N^-1/2 rate.Error histograms are compiled from 100 randomly generated data-set samples for each size.
  • Random data sets with fixed distribution about a classical material model: With fixed-width normal noise, the mean error converges to zero at order 0.22, slower than the linear rate observed for capped normal noise.The slower rate may be attributable to the wider spread of data about its mean, although the precise convergence–uncertainty trade-off remains unresolved.

6 Summary and discussion

The paper summarizes max-ent Data Driven Computing as an outlier-robust extension of distance-minimizing schemes, with free-energy minimization constrained by compatibility and equilibrium. It also discusses non-locality, convergence, implementation limits, and adaptive data coverage.

  • Max-ent Data Driven Computing generalizes distance-minimizing schemes and improves robustness to outliers through clustering analysis and maximum-entropy relevance weighting.Data-point relevance depends on distance to the solution and maximum-entropy estimation.
  • The resulting problem minimizes a free energy jointly over driving forces and fluxes subject to compatibility and equilibrium constraints.The formulation is non-standard because its free energy is defined over both local states and corresponding fluxes.
  • Distance-minimizing Data Driven schemes are recovered in the limit of zero temperature, while simulated annealing provides a solver through a quenching schedule.Selected numerical tests establish convergence properties of the max-ent solutions and solvers.
  • 6.1 Irreducibility to classical material laws: The effective energies are global and generally do not correspond to classical local material laws; reduced formulations remain non-local.This non-locality persists after eliminating displacement or Airy-potential variables.
  • Related work: Existing Material Informatics and model-calibration approaches differ because Data Driven Computing explicitly incorporates physical field equations while avoiding prespecified material-model forms.The approach relies on fundamental phase-space data rather than parameterized data tied to particular material models.
  • 6.5 Implementation Improvements: Data Driven distances identify poorly covered local states, enabling adaptive expansion of material data sets toward better application-specific phase-space coverage.The proposed implementation also leaves room for more efficient annealing schedules, subsampling, summarized data, and cutoff-based truncation.
Loading 1702.01574v2…