Source-linked AI summary

Neural networks meet hyperelasticity: A guide to enforcing physics

Lennart Linden, Dominik K. Klein, Karl A. Kalina, Jörg Brummund, Oliver Weeger, Markus Kästner

arXiv:2302.02403v2cs.CE

TL;DR

Compressible hyperelastic neural-network models have struggled to satisfy all common constitutive requirements simultaneously. The paper constructs a physics-augmented neural network using invariant-based convex potentials, analytical growth terms, and normalization terms, then evaluates it on analytical and noisy data. The resulting PANN satisfies the targeted conditions by construction and demonstrates accurate, robust prediction, extrapolation, and finite-element applicability, while potential non-negativity is examined numerically.

  • Problem

    Existing neural-network constitutive models do not generally enforce all common conditions of compressible hyperelasticity exactly, especially polyconvex stress normalization.

  • Method

    The PANN uses invariant-based convex neural potentials with analytical growth and polyconvex normalization terms to encode constitutive requirements by construction.

  • Results

    Accurate and robust predictions were obtained for isotropic and transversely isotropic analytical-potential data, including highly multiaxial and noisy cases, with good extrapolation and finite-element applicability.

  • Takeaways & Limitations

    Including constitutive physics gives the model physically sensible predictions, improved generalization and extrapolation, and supports calibration with small engineering datasets.

  • Takeaways & Limitations

    For compressible models, non-negativity of the neural-network potential is examined numerically rather than established analytically in general.

Abstract

from arXiv · show

In the present work, a hyperelastic constitutive model based on neural networks is proposed which fulfills all common constitutive conditions by construction, and in particular, is applicable to compressible material behavior. Using different sets of invariants as inputs, a hyperelastic potential is formulated as a convex neural network, thus fulfilling symmetry of the stress tensor, objectivity, material symmetry, polyconvexity, and thermodynamic consistency. In addition, a physically sensible stress behavior of the model is ensured by using analytical growth terms, as well as normalization terms which ensure the undeformed state to be stress free and with zero energy. In particular, polyconvex, invariant-based stress normalization terms are formulated for both isotropic and transversely isotropic material behavior. By fulfilling all of these conditions in an exact way, the proposed physics-augmented model combines a sound mechanical basis with the extraordinary flexibility that neural networks offer. Thus, it harmonizes the theory of hyperelasticity developed in the last decades with the up-to-date techniques of machine learning. Furthermore, the non-negativity of the hyperelastic neural network-based potentials is numerically examined by sampling the space of admissible deformations states, which, to the best of the authors' knowledge, is the only possibility for the considered nonlinear compressible models. For the isotropic neural network model, the sampling space required for that is reduced by analytical considerations. In addition, a proof for the non-negativity of the compressible Neo-Hooke potential is presented. The applicability of the model is demonstrated by calibrating it on data generated with analytical potentials, which is followed by an application of the model to finite element simulations. In addition, an adaption of the model to noisy data is shown and its [...]

1 Introduction

The introduction frames the challenge of combining all constitutive requirements in compressible hyperelasticity and motivates physics-augmented neural networks as a flexible alternative to black-box models. The proposed PANN enforces these requirements by construction and is evaluated through calibration, extrapolation, and finite-element simulations.

  • Research challenge: Fulfilling polyconvexity, objectivity, and other constitutive conditions simultaneously remains a major challenge in hyperelastic material modeling.The difficulty increases as more restrictions are imposed, despite the conditions not being mutually contradictory.
  • Motivation: Black-box neural networks fit stress-strain training data but may extrapolate poorly because they do not encode physical principles.Physics-augmented approaches are introduced to combine neural-network flexibility with constitutive structure.
  • Research gap: Existing invariant-based and convex-neural-network approaches enforce selected properties, but compressible models generally lack simultaneous exact enforcement of all common conditions.Prior normalization terms are either non-polyconvex or restricted to nearly incompressible behavior.
  • Contribution: The proposed PANN enforces thermodynamic consistency, stress symmetry, objectivity, material symmetry, polyconvexity, growth, and energy-stress normalization for compressible hyperelasticity.The approach is presented as a systematic construction rather than an application-specific post hoc correction.
  • Validation: The model is calibrated on analytical-potential data, tested in finite-element computations, adapted to noisy data, and accompanied by a proof of non-negativity for compressible Neo-Hooke energy.The manuscript also numerically examines non-negativity of the neural-network potentials.

2 Fundamentals of hyperelasticity

This section establishes the continuum-mechanics basis of hyperelastic constitutive modeling, linking deformation to stress through an energy potential. It specifies the requirements used later to construct admissible models, including objectivity, material symmetry, polyconvexity, growth, normalization, and non-negativity.

  • Hyperelastic formulation: Hyperelasticity relates strain and stress through an energy potential whose derivatives generate the constitutive stress response.The potential is identified with the Helmholtz free-energy density, making the stress a gradient field and ensuring path independence.
  • Stress symmetry: Invariant-based hyperelastic potentials imply symmetric Cauchy stress, satisfying the angular-momentum balance.Stress symmetry follows from the constitutive representation and stress transformations.
  • Polyconvexity: Polyconvexity represents the potential as convex in F, cof F, and det F, and it implies ellipticity associated with material stability.Every construction step and invariant must preserve polyconvexity.
  • Growth and normalization: The volumetric growth condition prevents compression to zero volume or expansion to infinite volume, while normalization makes the undeformed state stress-free with zero energy.These requirements constrain physically sensible behavior beyond symmetry and thermodynamic consistency.
  • Symmetry requirements: Formulating the potential with complete, irreducible invariants of the material symmetry group enforces objectivity and material symmetry without redundant arguments.The invariants must also be selected so that their use preserves polyconvexity.
  • Scope boundary: A complete non-polyconvex invariant set may not exist for every symmetry group or arbitrary anisotropic microstructure.In such cases, model construction requires choosing which constitutive property to prioritize.
  • Energy non-negativity: Non-negativity of the compressible Neo-Hooke potential is proved, whereas non-negativity for the broader analytical model is verified numerically.The Neo-Hooke proof is notable because the potential is polyconvex but not convex in F or C.

3 Physics-augmented neural network constitutive model

The PANN model uses invariant-based neural networks and architecture constraints to enforce key hyperelastic constitutive conditions, while normalization and growth terms address reference-state stresses and compressible behavior. Its energy non-negativity is examined numerically because polyconvexity alone does not guarantee it.

  • Basic conditions: Taking the neural-network output as a hyperelastic potential and constructing stress from its gradient enforces thermodynamic consistency, objectivity, material symmetry, and stress symmetry.The invariant formulation automatically yields a symmetric Cauchy stress tensor.
  • Model formulation: The PANN potential uses invariants as neural-network inputs, providing flexibility beyond analytical formulations while incorporating physics into the model.The invariant set may include additional quantities to satisfy physical conditions or improve approximation quality.
  • Polyconvexity: Convex, non-decreasing activations with non-negative weights preserve polyconvexity when polyconvex invariants are used as inputs.The model uses Softplus as a convex, non-decreasing activation function.
  • Energy non-negativity: Although the PANN fulfills the other constitutive conditions exactly, energy non-negativity requires numerical testing over admissible deformation states.Polyconvexity gives convexity in F, cof F, and det F, but not directly in F or C, so non-negativity does not follow automatically.
  • Growth and normalization: Normalization terms make the undeformed state a local minimum and stress-free, while the proposed invariant-based terms preserve all common constitutive conditions with less restrictive modifications.The construction addresses isotropic and transversely isotropic material behavior.
  • Growth and normalization: For compressible models, suitable normalization activations are difficult to construct because invariants lack a global minimum at the identity and must satisfy convexity, monotonicity, smoothness, and slope constraints.Related approaches may be restricted to incompressibility or fail to represent simple models such as Neo-Hooke.
  • Model calibration: The model parameters are calibrated after selecting an adequately flexible architecture, with the growth term included during calibration because high volumetric compressions can make it influential.Energy normalization may instead be added after calibration.

4 Numerical examples

The numerical studies show that physics-augmented neural networks preserve accurate interpolation while substantially improving physically plausible extrapolation. Across isotropic and transversely isotropic cases, enforcing constitutive conditions maintains high prediction quality, though complexity increases errors for transverse isotropy.

  • Interpolation behavior: All three models accurately approximate ideal uniaxial training data, while the PANN additionally reproduces the elastic energy and exactly enforces normalization.The basic-conditions model also reproduces the energy accurately, but only the PANN fulfills the normalization condition by construction.
  • Interpolation behavior: The PANN predicts approximately zero stress in the undeformed state and remains physically meaningful near λ1 = 1, whereas unconstrained models can fit offset data more closely.For offset data, the physically constrained models have MSEs several orders of magnitude larger than the direct stress model.
  • Extrapolation behavior: The PANN markedly improves extrapolation: in uniaxial loading its extrapolated MSE is two orders of magnitude lower than the basic-conditions model and remains accurate up to λ1 = 4.The improvement is attributed to polyconvexity, growth, and energy and stress normalization.
  • Extrapolation behavior: For biaxial loading and simple shear, the PANN predicts unseen states almost perfectly, while the direct stress model fails and the basic model deviates beyond λ1 > 1.4 or γ ≈ 0.8.The direct model can predict an implausible near-zero shear stress because it learned that diagonal stress components vanish.
  • Overall prediction quality: Isotropic models achieve median relative errors below 0.003%, and adding all physical principles does not deteriorate their approximation quality.Restricting weights to positive values worsens intermediate polyconvex models, but normalization restores performance close to the basic model.
  • Energy non-negativity: Numerical sampling detected only positive energies for isotropic and transversely isotropic PANNs over the tested admissible deformation ranges.The isotropic test used 1/10 ≤ λ ≤ 10, while the transverse-isotropy test varied stretches from 1/10 to 10 and rotations from 0 to π/2.
  • Overall prediction quality: Transversely isotropic models retain median errors below 0.4%, although enforcing physical principles increases errors by approximately one order of magnitude relative to the basic architecture.The authors note that larger or deeper networks could represent the more complex material behavior more accurately.
  • Finite-element validation: A three-dimensional torsion example is used to validate the NN constitutive approaches for a nonlinear multiaxial reference law.The study considers torsion of a prismatic sample and compares predictions against the reference constitutive behavior.

5 Conclusion

The paper proposes PANN, a neural-network constitutive model for compressible finite-strain hyperelasticity that enforces common constitutive conditions exactly. Calibration studies show accurate, robust prediction and extrapolation, while future work targets broader anisotropy and experimental-data applications.

  • PANN enforces thermodynamic consistency, stress symmetry, objectivity, material symmetry, polyconvexity, growth, and energy and stress normalization exactly.
  • The approach uses invariant sets as inputs to a convex neural network and is calibrated on isotropic and transversely isotropic analytical-potential data.
  • Highly accurate and robust predictions are reported for multiaxial deformation states and noisy stress-strain data, alongside extremely good extrapolation capability.
  • The model is applied straightforwardly in finite-element simulations.
  • Future extensions include polyconvex normalization for further material symmetry groups, tensor-basis networks for anisotropy discovery, and identification from experimental data.

CRediT authorship contribution statement

The authors contributed across conceptual, methodological, computational, validation, visualization, writing, funding, and resource activities.

  • Lennart Linden contributed to conceptualization, formal analysis, investigation, methodology, visualization, software, validation, and manuscript writing.
  • Dominik Klein contributed to conceptualization, formal analysis, methodology, visualization, software, validation, and manuscript writing.
  • Karl A. Kalina contributed to conceptualization, formal analysis, methodology, visualization, software, and manuscript writing.
  • Jörg Brummund contributed to methodology and manuscript review and editing.
  • Oliver Weeger and Markus Kästner contributed funding acquisition and resources, with Oliver Weeger also contributing conceptualization and manuscript review and editing.

A Multilayered neural networks

The multilayered PANN architecture generalizes the single-hidden-layer network while preserving polyconvexity through convex, non-decreasing activations and nonnegative weights.

  • The model uses sets of invariants as inputs to feed-forward neural networks with scalar-valued outputs representing the hyperelastic potential.
  • Increasing hidden-layer width or depth provides flexibility, and the multilayered architecture extends the single-layer formulation.
  • The multilayered network takes polyconvex, irreducible, independent invariants and additional adapted invariants as inputs.
  • Convex, non-decreasing activations in every hidden layer and nonnegative weights preserve the polyconvexity of the invariants.
  • The Softplus activation is convex and non-decreasing, yielding a polyconvex neural network under the stated weight conditions.

B Derivatives of invariants

This appendix provides invariant derivatives and identifies the multilayered PANN illustration for the material symmetry group under consideration.

  • Table 4 provides derivatives of isotropic and transversely isotropic invariants with respect to the right Cauchy-Green deformation tensor C.
  • The invariant derivatives support the construction of the model’s constitutive quantities from deformation measures.
  • Figure 11 illustrates the multilayered PANN constitutive model for the material symmetry group under consideration.

C Non-negativity of the strain energy density

This section characterizes physically admissible deformation states through restrictions on the principal invariants of the right Cauchy-Green tensor. These restrictions support the subsequent non-negativity analysis of strain-energy functions.

  • Admissible deformation states: A deformation state is physically admissible when the right Cauchy-Green tensor is positive definite, requiring all eigenvalues κ1, κ2, and κ3 to be positive.The eigenvalues are squared principal stretches.
  • Invariant restrictions: The tensor’s characteristic equation has real eigenvalues if and only if its invariant-based discriminant restriction is satisfied.This result follows from Cardano’s formula.
  • Invariant restrictions: Physical admissibility is equivalent to the discriminant restriction together with positivity of the principal invariants I1, I2, and I3.The proof uses the invariant relations to establish positive real eigenvalues.
  • Boundary characterization: When the discriminant vanishes, at least two eigenvalues coincide, allowing the invariants to be parameterized by one distinct eigenvalue x and one repeated eigenvalue y.The resulting relations are I1 = x + 2y, I2 = 2xy + y^2, and I3 = xy^2.

C.1 Neo-Hooke model

The compressible isotropic Neo-Hooke potential is shown to be non-negative over all physically admissible deformation states. Its unique minimum occurs at the undeformed state, where the energy is zero.

  • Main result: The isotropic Neo-Hooke strain-energy density is non-negative for all physically admissible deformation states.The result is established analytically on the admissible invariant domain.
  • Boundary analysis: The potential’s boundary behavior excludes local minima as invariants approach zero because the term −ln I3 diverges to infinity.This applies to the relevant boundaries where I1, I2, or I3 approaches zero.
  • Extremum analysis: On the discriminant boundary, the analysis reduces to positive eigenvalues x and y representing one distinct and one repeated eigenvalue.Partial derivatives with respect to these variables are used to locate extrema.
  • Main result: The point x = y = 1 is the global minimum, and the Neo-Hooke potential there equals zero.Convexity with respect to the spherical eigenvalue rules out other real non-negative solutions.

C.2 Isotropic PANN

The isotropic PANN analysis narrows possible local minima to volumetric deformation states and confirms the normalized undeformed state as a local minimum. Unlike the Neo-Hooke case, uniqueness of this minimum is not generally established.

  • Main result: Possible local minima of the isotropic PANN can occur only at volumetric deformation states C = λ^2 1.This reduces the search to positive triple eigenvalues of spherical tensors.
  • Extremum analysis: For the isotropic PANN, local extrema are analyzed by setting invariant derivatives expressed through the eigenvalues x and y to zero.The derivative equations are reduced using the repeated-eigenvalue boundary representation.
  • Extremum analysis: The alternative derivative solution is contradictory, leaving x = y as the only solution of the coupled extremum equations.The contradiction uses y > 0 together with positive invariant derivatives.
  • Limitation: Normalization makes x = 1 a local minimum with zero energy, but the minimum need not be unique because the PANN is not generally convex in the triple eigenvalue.Additional triple eigenvalues x ≠ 1 may therefore satisfy the stationary conditions.

D Stochastic

The stochastic study reports error distributions for isotropic and transversely isotropic models across progressively stronger constitutive constraints. Each model configuration was evaluated over 300 training runs using repeated training on reduced data.

  • Experimental design: 300 training runs were completed for each model configuration, with 30 neural-network trainings performed within each run.The histograms report the best of the 30 trainings for each run.
  • Error distributions: Figure 12 compares error-measure histograms for isotropic and transversely isotropic invariant-based models.The figure includes medians and the 25th and 75th percentiles.
  • Model configurations: The configurations range from basic conditions through polyconvexity and growth conditions to PANNs satisfying all conditions including normalization.Calibration uses the reduced dataset Dred,□.
Loading 2302.02403v2…