Source-linked AI summary
Polyconvex anisotropic hyperelasticity with neural networks
Dominik K. Klein, Mauricio Fernández, Robert J. Martin, Patrizio Neff, Oliver Weeger
TL;DR
The paper addresses how to build flexible constitutive models for finite-deformation anisotropic materials while enforcing polyconvexity and related stability requirements. It uses ICNN-based invariant and deformation-gradient formulations, with symmetry methods and mechanically motivated data augmentation, and finds strong performance for challenging metamaterial and transversely isotropic data. The authors also note limitations associated with polyconvexity’s mathematical scope and the interpretability of data-driven models.
Problem
Constructing flexible machine-learning constitutive models that simultaneously satisfy hyperelasticity, anisotropy, objectivity, material symmetry, polyconvexity, and stability remains difficult.
Method
The paper develops two ICNN-based polyconvex hyperelastic models: one using objective anisotropic invariants and another using F, Cof F, and det F with group symmetrization and mechanically motivated data augmentation.
Results
The deformation-gradient model reproduces and predicts challenging cubic metamaterial behavior very well, while both models perform excellently for transversely isotropic data.
Takeaways & Limitations
The models provide flexible polyconvex constitutive representations applicable to anisotropic materials and can generalize from fairly small training datasets.
Takeaways & Limitations
Polyconvexity is a mathematical stability framework rather than a fundamental physical requirement, and data-driven models require many parameters and are difficult to interpret.
Abstract
from arXiv · showhide
In the present work, two machine learning based constitutive models for finite deformations are proposed. Using input convex neural networks, the models are hyperelastic, anisotropic and fulfill the polyconvexity condition, which implies ellipticity and thus ensures material stability. The first constitutive model is based on a set of polyconvex, anisotropic and objective invariants. The second approach is formulated in terms of the deformation gradient, its cofactor and determinant, uses group symmetrization to fulfill the material symmetry condition, and data augmentation to fulfill objectivity approximately. The extension of the dataset for the data augmentation approach is based on mechanical considerations and does not require additional experimental or simulation data. The models are calibrated with highly challenging simulation data of cubic lattice metamaterials, including finite deformations and lattice instabilities. A moderate amount of calibration data is used, based on deformations which are commonly applied in experimental investigations. While the invariant-based model shows drawbacks for several deformation modes, the model based on the deformation gradient alone is able to reproduce and predict the effective material behavior very well and exhibits excellent generalization capabilities. In addition, the models are calibrated with transversely isotropic data, generated with an analytical polyconvex potential. For this case, both models show excellent results, demonstrating the straightforward applicability of the polyconvex neural network constitutive models to other symmetry groups.
1 Introduction
The paper addresses the challenge of constructing anisotropic constitutive models that satisfy shared physical and mathematical requirements, especially polyconvexity, while retaining flexibility for complex materials. It introduces two ICNN-based polyconvex models and evaluates them on challenging metamaterial data.
- Motivation: Advanced manufacturing has produced flexible and functional metamaterials with mechanical characteristics unlike those of classical materials.Their microstructures can contain beams and shells and exhibit lattice instabilities.
- Motivation: Constitutive models must satisfy requirements including ellipticity, thermodynamic consistency, objectivity, and material symmetry.These requirements remain common across specialized material formulations.
- Polyconvexity challenge: Polyconvexity is useful because it implies ellipticity, which ensures material stability and is easier to impose than ellipticity directly.This is particularly relevant for numerical methods such as finite element analysis.
- Polyconvexity challenge: Anisotropic polyconvex formulations historically struggled to satisfy polyconvexity, objectivity, material symmetry, and a stress-free reference configuration simultaneously.The simultaneous fulfillment of these requirements is identified as a major open problem.
- Polyconvexity challenge: Machine-learning constitutive modeling remains challenged by the difficulty of enforcing multiple physical requirements within flexible data-driven models.Prior approaches considered convexity properties, but the authors identify no known method satisfying polyconvexity.
- Proposed models: The paper proposes two ICNN-based polyconvex models for hyperelastic, anisotropic behavior under finite deformations.The first uses objective, symmetry-aware invariants; the second uses deformation-gradient variables with data augmentation and group symmetrization.
- Evaluation: Both models are calibrated with moderate synthetic homogenization data for cubic beam-lattice metamaterials featuring finite deformations and lattice instabilities.The models are compared with each other and with a conventional polyconvex model.
2 Basics of material theory
This section introduces finite-elasticity requirements for physically sensible and mathematically well-posed constitutive models. It explains objectivity, material symmetry, stress-free reference states, polyconvexity, coercivity, and their relation to ellipticity and stability.
- Constitutive requirements: The strain-energy density depends on the deformation gradient and must satisfy restrictions needed for physically sensible and mathematically well-posed formulations.The section focuses on stress-free reference states, objectivity, material symmetry, polyconvexity, and related conditions.
- Stress-free reference: A stress-free reference configuration requires the first Piola–Kirchhoff stress to vanish at the identity.Equivalently, the potential has its unique global minimum at the identity with W(1)=0 and W(F)≥0.
- Objectivity: Objectivity requires the potential to remain invariant under transformations by rotations of the observer.For stress, the corresponding transformation is S(QF)=QS(F).
- Objectivity: When formulated as W(C) with C=F^T F, material objectivity is fulfilled automatically.This is an advantage over formulations depending directly on the deformation gradient.
- Material symmetry: For anisotropic materials, material symmetry requires invariance under transformations FQ with Q in the material symmetry group G.The condition is W(FQ)=W(F).
- Polyconvexity and stability: Polyconvexity requires the energy to be representable by a function convex in F, Cof F, and det F.With ξ=(F,Cof F,det F), polyconvexity implies ellipticity, and ellipticity ensures material stability.
- Polyconvexity and stability: Existence of minimizers requires polyconvexity together with an additional coercivity condition.The present work does not consider coercivity because its focus is on the polyconvex formulation.
3 Polyconvex constitutive models based on FFNNs
The paper develops two ICNN-based constitutive approaches that combine hyperelasticity, anisotropy, objectivity, material symmetry, and polyconvexity for finite-deformation modeling. One uses polyconvex invariants, while the other uses deformation-based inputs with symmetry enforcement and data augmentation.
- ICNN architectures impose convexity through convex, non-decreasing layers and suitable weight restrictions, unlike unrestricted neural networks.The first hidden layer is convex, later layers are convex and non-decreasing, and the output layer is convex and non-decreasing.
- Two polyconvex ICNN-based approaches are introduced for anisotropic hyperelastic constitutive modeling at finite deformations.The models extend earlier approaches while using convex neural-network architectures to construct polyconvex potentials.
- Model based on invariants: The invariant-based model uses objective invariants built from the right Cauchy-Green tensor and cubic structural tensors.For cubic materials, the invariant set includes I1, I2, I3, J7, and J11; their convexity is distributed across F, Cof F, and det F.
- Model based on invariants: The invariant model is polyconvex, objective, and materially symmetric because its inputs are convex invariants and its network uses convex non-decreasing layers.An analytical volumetric term is added because ICNNs are not necessarily coercive and therefore do not by themselves ensure the volumetric growth condition.
- Model based on invariants: The invariant model does not enforce a stress-free reference configuration exactly; calibration only approximates S(1) = 0.An exact projection approach is incompatible with the present objectivity and material-symmetry constructions.
- Model based on the deformation gradient: The deformation-based approach approximates objectivity through mechanically motivated data augmentation and incorporates anisotropy through group symmetrization.The augmented dataset uses rotated observers and requires no additional experimental or simulation data; explicit objectivity penalty terms are not added.
4 Application to cubic metamaterials
The models are evaluated on homogenized cubic beam-lattice metamaterials using moderate calibration data spanning common experimental deformation modes and challenging lattice instabilities. The deformation-gradient model W_F consistently outperforms the invariant-based model W_I, while W_C can approximate data but may lose ellipticity.
- 4.1 Homogenized behavior of soft beam-lattice metamaterials: Finite-element homogenization supplies deformation gradients, effective stresses, and strain-energy densities for cubic BCC and X lattice cells.The calibration set contains 905 triplets covering uniaxial, equibiaxial, planar, shear, and volumetric deformations; test cases include biaxial and tension-shear loading.
- 4.3 Model evaluation for BCC cell: W_C approximates the BCC responses well but loses ellipticity even at small deformations, whereas W_F satisfies the volumetric growth condition as J approaches 0+.For these metamaterials, the study therefore treats polyconvexity and volumetric growth as important safeguards against instability and problematic extrapolation.
- 4.3 Model evaluation for BCC cell: W_I performs acceptably mainly for deformation gradients with dominant diagonal entries but fails for shear and the associated mixed test case.Its errors in S33 also transfer from uniaxial and equibiaxial calibration cases to the biaxial test case.
- 4.3 Model evaluation for BCC cell: W_F reproduces all calibration and test deformation modes with only small deviations in S12 for the mixed test case, and can perfectly fit both datasets after recalibration.It also yields excellent MSE and MD values, combining approximation quality with consistent behavior across model instances.
5 Application to transverse isotropy
The proposed polyconvex neural-network models are applied to transversely isotropic data generated from an analytical polyconvex potential. Both models reproduce the data well, with the invariant-based model showing especially strong agreement.
- Analytical data generation: The analytical potential generates transversely isotropic stress data using a preferred x1-axis and structural-tensor invariants J4 and J5.J4 and J5 are convex in F and Cof F, respectively, and are combined with isotropic invariants.
- Datasets and tests: The calibration data combine uniaxial, equibiaxial, and shear tests, while the test data use biaxial and mixed tension-shear cases.The calibration set contains 650 tuples and the test set contains 200 tuples.
- Model construction: The invariant-based model uses six invariant inputs, whereas the deformation-gradient model uses (F, det F) in R10.Convex Softplus activations and restricted network parameters enforce convexity in the models.
- Model construction: Objectivity and transverse isotropy are incorporated for the deformation-gradient model through rotation-based data augmentation and group symmetrization.The calibration uses 128 random rotations for objectivity and six rotations around the x1-axis for transverse isotropy.
- Results: Both models agree excellently with the transversely isotropic analytical data, although the deformation-gradient model shows slight observer dependence in shear and test cases.The invariant-based model performs especially well, and the results support applicability to other symmetry groups.
6 A critique of machine learning in nonlinear elasticity theory
The paper critiques machine-learning constitutive models for nonlinear elasticity as flexible but difficult to interpret, calibrate reliably, and validate beyond observed data. It therefore advocates embedding physical and mathematical structure and using such models where classical approaches lack sufficient accuracy.
- Shortcomings: Machine-learning constitutive models lack intuitive interpretations linking their parameters and predictions to material behavior.Their high parameter counts and absence of simple closed-form expressions make this relation especially difficult to understand.
- Shortcomings: Parameter optimization can depend on initialization and experimental data, so fitted behavior is not determined by measurements alone.Strong nonlinear dependence of stress-strain relations on parameters contributes to this instability.
- Shortcomings: Validation cannot establish applicability outside the range of prior experimental observations.Testing and validation remain based on available data, leaving extrapolation assumptions unsupported.
- Position: The authors recommend incorporating hyperelasticity, anisotropy, objectivity, and polyconvexity, while reserving machine learning for cases where classical models lack acceptable accuracy.The lattice microstructures studied here are cited as an example because of their strong nonlinearity.
- Interpretability: Even accurate trained algorithms do not explain why their predictions are accurate, unlike models developed from explicit geometrical and mechanical considerations.The paper contrasts this black-box limitation with the interpretability of Hencky elasticity.
7 Conclusion
The proposed ICNN constitutive models enforce polyconvexity for anisotropic finite elasticity, with the deformation-gradient model performing especially well on challenging metamaterials and both models succeeding for transversely isotropic data.
- The two hyperelastic models use input convex neural networks to enforce polyconvexity, which implies ellipticity and ensures material stability.They are formulated for finite deformations and anisotropic material behavior.
- The invariant-based model uses polyconvex anisotropic invariants that satisfy objectivity and material symmetry by construction.Its neural-network core represents highly nonlinear functions of the invariants.
- The deformation-gradient model uses the deformation gradient, cofactor, and determinant, while approximating objectivity through mechanically motivated data augmentation.The augmentation extends the calibration dataset without additional simulation or experimental data.
- The deformation-gradient model reproduces and predicts cubic metamaterial behavior well, including test scenarios excluded from training, whereas the invariant-based model fails for shear deformations.The benchmark includes lattice instabilities and finite-deformation responses.
- Both models perform excellently on data generated from an analytical transversely isotropic polyconvex model, supporting applicability to other anisotropy classes.The deformation-gradient model is preferable for highly challenging metamaterials, while the invariant-based model is effective for moderately challenging behavior.
- Application to finite element simulations remains future work, although the models' smooth neural-network cores and ellipticity support straightforward adaptation.The models can use small networks for moderately challenging materials and larger networks for complex behavior.
A Input convex feed-forward neural networks
This appendix defines feed-forward and input convex neural networks and gives sufficient composition conditions, activation choices, and weight constraints for constructing smooth convex network cores.
- A feed-forward neural network recursively composes vector-valued functions, with vector components called nodes or neurons and each node using an activation function.
- An input convex neural network is a feed-forward network whose scalar output is convex with respect to its vector-valued input.
- Function composition preserves convexity when the outer function is convex and non-decreasing, while the inner function is component-wise convex.These are sufficient conditions for convexity with respect to the original input.
- A convex feed-forward network requires a component-wise convex first hidden layer, subsequent component-wise convex non-decreasing layers, and a convex non-decreasing scalar output.The result follows by recursively applying the composition theorem.
- Convexity is preserved under affine transformations, so biases may be arbitrary and the input may be multiplied by any constant matrix.
- Log-Sum-Exp and Softplus activations are smooth and convex for arbitrary weights and biases, but are non-decreasing only when their weights are non-negative.For invariant-based models, non-negative first-layer weights ensure the required non-decreasing dependence on nonlinear invariants.
- The numerical investigations use Softplus-based cores because the manually implemented Log-Sum-Exp formulation converged too slowly to produce satisfying results.The authors attribute this to the implementation and optimization approach rather than a general defect of Log-Sum-Exp.