Source-linked AI summary
Data-Driven Design for Metamaterials and Multiscale Systems: A Review
Doksoo Lee, Wei Wayne Chen, Liwei Wang, Yu-Chin Chan, Wei Chen
TL;DR
Metamaterial design must address vast design spaces, intricate structure-property relationships, and costly evaluations. This review synthesizes data-driven methods through modules for data acquisition, machine-learning-based unit-cell design, and multiscale design. It organizes approaches by shared principles, compares their applicability, and identifies open research opportunities.
Problem
Metamaterial design involves high-dimensional topology, multiscale structure-property mappings, local optima, absent analytical gradients, and expensive evaluations.
Method
The review uses a design-centered, holistic synthesis organized into data acquisition, machine-learning-based unit-cell design, and multiscale optimization.
Results
The review provides a standardized, methodological organization of data-driven metamaterial design practices across modules and domains.
Takeaways & Limitations
Data-driven design can extract patterns unavailable or difficult to obtain from physical models and incorporate them into metamaterial design workflows.
Takeaways & Limitations
The review is limited to machine-learning-based design methods for metamaterials and their multiscale systems, while spanning multiple physics domains.
Abstract
from arXiv · showhide
Metamaterials are artificial materials designed to exhibit effective material parameters that go beyond those found in nature. Composed of unit cells with rich designability that are assembled into multiscale systems, they hold great promise for realizing next-generation devices with exceptional, often exotic, functionalities. However, the vast design space and intricate structure-property relationships pose significant challenges in their design. A compelling paradigm that could bring the full potential of metamaterials to fruition is emerging: data-driven design. In this review, we provide a holistic overview of this rapidly evolving field, emphasizing the general methodology instead of specific domains and deployment contexts. We organize existing research into data-driven modules, encompassing data acquisition, machine learning-based unit cell design, and data-driven multiscale optimization. We further categorize the approaches within each module based on shared principles, analyze and compare strengths and applicability, explore connections between different modules, and identify open research questions and opportunities.
1 Introduction
Metamaterial design offers broad architectural control but must navigate complex, expensive, and often non-differentiable structure-property spaces. This review organizes data-driven design into modules and synthesizes methods across domains from data acquisition through multiscale optimization.
- Metamaterial architectures can provide broad or spatially varying properties without relying solely on precise material composition and processing control.
- Design is challenging because it combines high-dimensional topology, multiscale structure-property mappings, local optima, absent analytical gradients, and expensive evaluations.
- Data-driven methods support high-throughput prediction, dimensionality reduction, accelerated exploration and optimization, and faster solutions to ill-posed inverse-design problems.
- Their design freedom can extend to heterogeneous metamaterial system design beyond conventional approaches.
- The review adopts a design-centered, cross-domain synthesis organized around data acquisition, machine learning-based unit-cell design, and multiscale optimization.
- Its scope includes machine-learning-based metamaterial and multiscale design across optical, acoustic, mechanical, thermal, and magneto-mechanical physics, including methods without prior training data.
2 Preliminaries
The review establishes a hierarchy from unit cells to microstructures and metamaterials, then introduces the machine-learning concepts used to design their properties. It covers supervised, unsupervised, semi-supervised, and reinforcement-learning approaches.
- Key Concepts: A unit cell is the smallest representative material unit used to control properties, while microstructures assemble unit cells and metamaterials assemble microstructures.
- Key Concepts: Multiscale design targets desired properties across multiple length scales, and effective properties arise from collective microstructure behavior.
- Key Concepts: The review treats data acquisition, machine learning-based unit-cell design, and multiscale optimization as reusable modules in an ordered design framework.
- Key Concepts: Design space contains all combinations of design variables, whereas shape space and property space describe unit-cell geometries and responses, respectively.
- Key Concepts: Evaluation obtains system responses under an architecture and loading condition, and is treated interchangeably with labeling in this review.
- Machine Learning Methods: Supervised learning models shape-property relations and can replace resource-intensive unit-cell evaluation with surrogate models.
- Machine Learning Methods: Unsupervised methods learn representations or distributions from unlabeled metamaterial geometries, commonly using autoencoders, variational autoencoders, or GANs.
- Machine Learning Methods: Semi-supervised learning combines labeled and unlabeled data, while conditional generative variants can support inverse design or property-related latent representations.
3 Data Acquisition
Data acquisition determines which finite structure-property pairs represent the downstream design space, making it foundational but challenging. The review organizes approaches by shape-centric generation, property-aware acquisition, and data assessment.
- Overview: Unit-cell datasets enable data-driven multiscale design but face high-dimensional exploration, expensive evaluation, many-to-one mappings, distributional bias, and compounding data-quality effects.
- Overview: The review proposes a standardized taxonomy that bridges diverse acquisition strategies through methodological comparison.
- Overview: Shape-centric generation uses heuristics and domain knowledge to create large datasets without handcrafting every shape.
- Overview: Property-aware acquisition accounts for target properties to improve sampling efficiency and customize data for specific design tasks.
- Overview: Data assessment is included as a supporting practice for data-quality assurance and data sharing.
- Shape-Centric Data Generation Method: Shape generation addresses which unit cells to specify and how to expand sparse data into sufficiently large datasets.
- Representation of Unit Cells: Unit-cell representation is selected early because it projects high-dimensional shapes into lower-dimensional descriptions and determines resulting data distributions.
A. Parametric Multiclass
Parametric multiclass representations organize unit cells around geometric classes while combining explicit parameters that support exploration within or across classes. Mixed-variable forms can include both qualitative and quantitative design variables.
- A. Parametric Multiclass: Geometric classes serve as pivots for shape generation and are typically equipped with low-dimensional explicit parameterizations.
- A. Parametric Multiclass: Parameters such as length, thickness, volume fraction, entity angle, and rotational angle enable exploration within or across geometric classes.
- A. Parametric Multiclass: Mixed-variable representations combine qualitative variables such as building-block type with quantitative variables such as scaling factor for multiclass topology optimization.
B. Implicit Function
Implicit-function representations encode unit-cell shapes through surface functions, supporting geometric families with smooth topological variation and relatively few tunable parameters. Common examples include TPMS, spinodoid, spectral, and pixel/voxel approaches.
- Implicit-function representations describe shape instances through surface functions and are widely used to generate geometric families.
- Triply Periodic Minimal Surfaces are widely used because they have zero mean curvature and large surface areas.
- Spinodoid representations provide smooth, aperiodic variations of complex topologies with tunable anisotropy.
- Spectral representations combine Fourier transforms with level-set functions, enabling topologically rich unit cells, inverse-Fourier reconstruction, and efficient symmetry handling.
- Pixel/voxel representations treat shapes as spatial aggregates of solid or void elements, offering freeform designs and a direct connection to inverse topology optimization.
- Inverse topology optimization has generated hundreds of freeform unit cells matching prescribed effective properties or force-displacement responses.
D. Parametric Curve/Surface
Parametric curve/surface representations describe shapes using ordered control points or tunable analytical boundary functions. In metamaterial design, they support exploration beyond canonical families, particularly for wave-based structures.
- Boundary-based representations encode shapes as ordered sequences of control points on curves or surfaces.
- These representations are primarily used for wave-based metamaterials seeking design exploration beyond canonical shape families.
- For metagratings, boundary representations have specified shape instances with 16 boundary parametric curves or surfaces.
E. Constructive Solid Geometry
Constructive Solid Geometry creates solids by composing primitives through set-theoretic operations, yielding interpretable, CAD-compatible parameterizations. In metamaterial design, it has been applied to synthesize plasmonic and dielectric metasurface structures.
- Constructive Solid Geometry composes primitives such as rectangles, cylinders, and spheres through operations including union and intersection.
- Its semantic construction makes instances highly interpretable and provides a seamless connection with Computer-Aided Design.
- Photonic metasurface studies have used parameterized primitives and union operations to synthesize plasmonic nanostructures.
- After selecting a representation, reproduction strategies grow large shape collections and influence the resulting data distribution and downstream-task quality.
A. Parametric Sweep
Parametric Sweep samples a low- or moderately dimensional descriptor space to enlarge sparse shape datasets. It is widely used, but its freedom, sampling density, and property-space coverage deteriorate as design dimensionality and complexity increase.
- Parametric Sweep explores a low- or moderately dimensional descriptor space with a finite set of approximately uniform samples.
- Figure 3a illustrates a sweep across six lattice classes by varying volume fraction.
- Parametric Sweep is the most widely used reproduction strategy and has been combined with diverse low-dimensional unit-cell representations.
- As dimensionality increases, space-filling sampling density drops sharply because of the curse of dimensionality.
- Sweeping within one selected pivot offers little design freedom, cannot bridge multiple unit-cell classes, and can bias property-space coverage.
B. Multiclass Blending
Multiclass blending grows shape libraries by combining seed classes, potentially creating unseen inter-class instances and continuous shape manifolds. Its implementation choices and data-acquisition relationships require careful specification and analysis.
- Blending interpolates across multiple seed classes to generate new instances that can merge classes into a unified landscape.
- Boolean image operations and deep generative models synthesize freeform inter-class unit cells and distill continuous shape manifolds.
- Blending requires specifying the operation type, number of classes, and weighting-factor rules.
- Perturbation-based reproduction expands datasets by repeatedly modifying near-boundary or boundary instances to increase property coverage.
- Perturbation has mainly been applied to Pixel/Voxel representations but may explore new instances in lower-dimensional lattice representations.
- Shape collections serve as downstream design elements because they determine the landscape explored and influence resulting property distributions.
A Trade-Off between Dimensional Compactness and Expressivity
Metamaterial representations trade dimensional compactness against expressivity, while class choice, randomness, and attribute handling shape their suitability for data-driven design.
- Compact representations simplify generation but can restrict exploration, whereas expressive representations support freer topologies while complicating exploration and manufacturability.
- Class-centric representations use predefined seed classes, while class-free approaches such as Pixel/Voxel generally do not.
- Seed classes can be selected using shape attributes, material properties, and shape diversity to secure broad shape-space coverage.
- Biologically inspired acquisition uses motifs from natural systems, but reported work has mainly emphasized proof-of-concept demonstrations.
- Deterministic representations dominate multiscale datasets, while stochastic representations address intrinsically random microstructures and deployment uncertainties.
- Explicit low-dimensional representations generally simplify enforcing symmetry, periodicity, invariance, volume, manufacturability, and connectivity constraints.
3.3 Property-Aware Data Acquisition Strategy
Property-aware acquisition addresses the limits of exhaustive and property-agnostic sampling by steering evaluation toward useful, diverse, or task-relevant datasets. Active learning and subset selection provide complementary routes for sequential expansion and bias reduction.
- Exhaustive sampling becomes impractical for costly simulations, datasets exceeding 100k instances, or spaces above 50 dimensions.
- Active learning iteratively selects unlabeled unit cells for evaluation, enabling systematic acquisition from large pools with few labels.
- Sequential active learning monitors dataset growth and can address how much data is needed while supporting progressive, task-aware generation.
- Diversity-based subset selection uses Determinantal Point Processes to find small representative subsets with tunable shape, property, or joint diversity.
- Diversity-based acquisition handles high-dimensional inputs through pairwise kernels but requires O(N^2) storage and O(N^3) matrix-inversion time.
- Property-agnostic acquisition can create biased output distributions because uniformity in the input space may not transfer to the property space.
- Task-aware acquisition targets user-preferred properties such as negative Poisson’s ratio or broadband reflectivity rather than treating all data as equally useful.
- Property supervision is difficult because unseen properties require costly evaluation and regression-distribution control remains under-explored.
3.4 Data Assessment
Data assessment combines quantitative metrics and visualization to compare metamaterial datasets across shape and property spaces. The review emphasizes coverage, uniformity, diversity, task relevance, and interpretable latent representations.
- Large datasets are assessed with distributional summaries and visualization because inspecting individual samples is intractable.
- Assessment protocols must balance data size, uniformity, property coverage, task relevance, and manufacturability, so subjectivity cannot be fully eliminated.
- Space-filling and diversity metrics evaluate whether datasets broadly and uniformly cover design or property spaces.
- Task-related metrics check whether regions associated with goals such as high anisotropy or broadband reflectivity are sufficiently represented.
- Diversity-driven subsets showed larger shape diversity than random sampling, while sequential acquisition quantified diversity gain and task-awareness.
- The 21,684-instance multiclass dataset was more uniform in Young’s modulus and volume fraction, while the 924-instance TPMS dataset uniquely covered some E-vf values.
- Dimensionality reduction makes high-dimensional properties visualizable and can encode resonance-frequency shifts in latent spaces.
- PCA, t-SNE, UMAP, and related projections support exploratory assessment of high-dimensional shape manifolds.
3.5 Discussion
The review identifies dataset construction, periodicity assumptions, limited full-field modeling, data-size selection, and data-sharing standards as central challenges for data-driven multiscale design.
- Data acquisition: Most current multiscale studies create task-specific datasets from scratch, motivating versatile datasets that support disparate design tasks.The review links current practice to substantial trial-and-error and computational costs during data creation.
- Modeling assumptions: Periodic boundary conditions underpin many structure-property datasets but may misrepresent fully aperiodic systems and other nonperiodic settings.The assumption can fail for systems with large deformation, strong neighbor coupling, or long-range interactions.
- Structure-property mapping: Surrogate models that predict homogenized properties are easier to learn but discard full-field information, prompting interest in physics-informed and operator-learning approaches.Field-based methods must address high-dimensional outputs, overfitting, and physics-related priors such as field smoothness.
- Data efficiency: An optimal data size is difficult to determine because it depends on model complexity, representation dimensionality, and simulation cost or fidelity.The review calls for more research in small-data regimes, especially where simulations and computational resources are limited.
- Data sharing: Open data platforms such as MetaMine can consolidate structure-property data, but reusable and reproducible sharing requires domain-specific, extensible protocols.Generic FAIR principles leave discipline-specific standards unresolved across metamaterials domains.
- Beyond acquisition: Data augmentation, consolidation, bias reduction, domain adaptation, and exploratory analysis complement acquisition by improving data generality, customizability, and reusability.Augmentation can increase training data without new evaluations, encode operational invariances, and mitigate overfitting.
4 Data-Driven Unit Cell Design of Metamaterials
Unit-cell metamaterial design uses data-driven methods to address costly evaluations, high-dimensional geometry, and nondifferentiable physics across multiple domains. Reviewed approaches include surrogate prediction, representation learning, reinforcement learning, generative mappings, and physics-informed optimization, each with distinct applicability and trade-offs.
- Overview: High-fidelity simulations can be costly, geometric design spaces can be high-dimensional, and nondifferentiable physics limits gradient-based optimization.Complex numerical nanophotonic simulations may take hours or days, while advanced fabrication enables broad design freedom.
- Overview: The review covers optical, acoustic/elastic, mechanical, thermal, and magneto-mechanical metamaterials because their properties depend on geometry and are expensive to compute.Random geometries are comparatively cheap to generate, supporting data-driven property modeling across these physics domains.
- Method categories: Machine learning commonly accelerates property evaluation, learns efficient representations, supports sequential decisions, or generates solutions using physical information.These roles organize the reviewed unit-cell optimization methods.
- Accelerated optimization: CNNs typically serve high-dimensional pixelated designs, whereas MLPs are used for lower-dimensional parametric or shape representations.These models dominate reviewed ML-accelerated optimization studies based on Figure 8.
- Accelerated optimization: Gaussian processes suit relatively small datasets because uncertainty estimates support adaptive sampling and Bayesian optimization, but standard models scale as O(N^3) in time and O(N^2) in memory.This complexity limits large datasets and higher-dimensional problems, though scalable variants have been proposed.
- Sequential decision making: Reinforcement learning can outperform genetic algorithms in larger action spaces, but exploration costs rise rapidly and reviewed designs remained below 50 dimensions.Its practical use requires finding an action-space sweet spot that balances performance against computational burden.
- Physics-based learning: Physics-informed methods can guide design without training data; hPINNs matched adjoint optimization objectives while producing simpler, smoother solutions with faster convergence.Their metamaterials-design applications remain relatively limited despite broad attention to physics-informed machine learning.
- Generative design: Conditional generative models can map target optical images or dispersion curves to corresponding metasurface or elastic-metamaterial designs.These mappings use physical operation mechanisms or target-property information during training.
4.4 Discussion and Future Opportunities
The discussion compares data-driven metamaterials methods by cost, design dimensionality, generality, and trustworthiness, while identifying unresolved challenges in accuracy, uncertainty quantification, interpretability, and novel-design discovery.
- Cost-Benefit: Neural-network models dominate complex metamaterials design, but their large-data and interpretability requirements remain important trade-offs.DGMs, MLPs, and CNNs are the most frequently used models, while data collection usually dominates machine-learning cost.
- Cost-Benefit: MLP methods are most common below 200 design dimensions, whereas DGMs are most common above 1,000 dimensions.Dimensionality reduction enables MLPs and GPs to address substantially higher-dimensional problems, while DGMs support representation learning and one-to-many inverse mappings.
- Cost-Benefit: Iteration-free inverse design reduces computation time relative to iterative optimization by trading off accuracy, with near-optimal warm starts offering a refinement strategy.The proposed combination uses inverse design for initialization and a relatively small number of optimization iterations for refinement.
- Cost-Benefit: Generality depends on applicability across problem settings, while physics-based models requiring no training data may need retraining when constraints or operating conditions change.Generality can be evaluated through the similarity between training and test problems and other task characteristics.
- Trustworthiness: Fabrication uncertainty can substantially change metamaterial responses, motivating uncertainty quantification that preserves high-dimensional geometric variation rather than assuming uniform boundary changes.A small geometric perturbation in a metasurface can significantly alter its absorbance spectrum.
- Trustworthiness: Interpretability remains under-explored for complex, high-dimensional metamaterial designs, despite existing interpretable and physics-assisted approaches.Prior work includes linear models, decision trees, PDE-constrained models, and physics-based solvers integrated with neural networks.
- Future Opportunities: Novel-design discovery beyond interpolation remains under-studied because reinforcement learning has so far targeted low-dimensional problems under computational-cost constraints.Classic machine-learning approaches commonly rely on independent and identically distributed training and testing data.
5 Data-driven Multiscale Metamaterial System Design
Data-driven multiscale design addresses heterogeneous, spatially varying requirements by coupling macroscale structure optimization with microscale unit-cell design. The review contrasts bottom-up and top-down frameworks, examines orientation, topology, homogenization, and task-specific data, and highlights unresolved applicability challenges.
- Motivation: Heterogeneous property distributions support spatially varying requirements and complex functions that homogeneous designs cannot meet.An invisibility cloak is cited as an example requiring heterogeneous properties around an object.
- Multiscale design framework: Multiscale design jointly optimizes macroscale topology and properties while selecting microscale unit cells that realize the required local responses.The two scales should ideally be designed concurrently rather than sequentially.
- Bottom-up and top-down frameworks: Bottom-up frameworks use microscale variables with surrogate structure-property models, whereas top-down frameworks optimize macroscale targets before retrieving compatible building blocks.Top-down assembly can use microscale ML models to represent or generate unit cells efficiently.
- Design flexibility: Single-class graded designs are efficient but can be sub-optimal because fixed topology and orientation cannot represent requirements such as multi-loading responses.Oriented rank-2 and rank-3 materials are needed for certain single- and multi-loading compliance problems.
- Design flexibility: Broader topology sets improve design flexibility but require descriptors that represent diverse unit cells without substantially increasing dimensionality.This challenge motivates variants beyond low-dimensional graded representations.
- Assumptions and challenges: First-order homogenization is reliable only under scale-separation and periodicity assumptions, which many manufacturable, aperiodic designs violate.Smoothly varying or compatible neighboring cells can still yield relatively satisfying results, but full-scale performance should also be reported.
- Task specificity: Task-specific property distributions show that data utility depends on the target deformation, so acquisition may require deliberate task-informed bias rather than uniform coverage alone.The smiley-face task favors clustered and anisotropic samples, whereas bridge-like deformation benefits more from wide coverage; negative-Poisson-ratio samples remain unused there.
- Assumptions and challenges: The review cautions that black-box ML models can have questionable applicability when their assumptions and physical constraints are ignored.Mean-error losses may still produce infeasible outputs, including singular stiffness matrices, disconnected cells, discontinuities, and non-differentiable responses.
6 Conclusion
The review organizes data-driven metamaterial design into a cohesive framework and evaluates current practices across data acquisition, unit-cell learning, and multiscale design. It identifies methodological gaps and research opportunities, while emphasizing that the field remains early-stage and must better connect data-driven and physics-based approaches.
- The review categorizes prior research into data acquisition, unit-cell learning and optimization, and multiscale system design.It examines shape generation, property-aware sampling, data assessment, machine-learning-enabled unit-cell design, and multiscale optimization strategies.
- Multiscale studies primarily replace homogenization with surrogate modeling while accommodating diverse and oriented unit cells.
- Current research is biased toward downstream products and lacks principled data-acquisition methods, benchmark datasets, and standard assessment protocols.The review identifies these needs as important for rigorous and robust deployment of data-driven design frameworks.
- Unit-cell design research under-studies cost-benefit trade-offs, model trustworthiness, creativity, uncertainty quantification, and generalization.The review highlights interpretable and physics-informed machine learning as opportunities for addressing these gaps.
- The field remains in its early stages, with a major challenge being the disconnect between data-driven and physics-based approaches.