Source-linked AI summary
Simultaneous inference of environmental and interaction forces in collective dynamics
Nipuni de Silva, Ming Zhong, James M. Greene
TL;DR
The paper addresses how to infer collective dynamics when both inter-agent interactions and environmental forces shape observed trajectories. It extends variational learning with semi-parametric and fully nonparametric force representations, and reports accurate mechanism recovery and prediction across diverse benchmark systems. The study also identifies a limitation: semi-parametric recovery depends strongly on correctly specifying the environmental-force form.
Problem
Learning collective dynamics requires identifying local interaction mechanisms while avoiding arbitrary high-dimensional vector-field inference and accommodating unknown governing forms.
Method
The framework jointly learns a nonparametric interaction kernel and environmental force using either semi-parametric or fully nonparametric variational representations, with trajectory-based hypothesis-space estimation.
Results
Across benchmark systems involving synchronization, attraction/repulsion, alignment, and external forcing, the methods accurately recover governing mechanisms and predict collective behavior beyond the training horizon.
Takeaways & Limitations
Structured low-dimensional force learning yields mechanistic estimators that are computationally tractable and interpretable for collective dynamics.
Takeaways & Limitations
Semi-parametric learning depends strongly on the prescribed environmental-force form and can estimate forces and trajectories poorly when the true force is not representable in that form.
Abstract
from arXiv · showhide
Collective dynamics arise in a wide range of physical, biological, and engineering applications. Examples include cell migration, swarm robotics, social dynamics, and animal behavior. A defining characteristic of these systems is the emergence of large-scale coordination from local interactions among agents; a fundamental question is thus to understand the local interactions that give rise to the observed emergent dynamics. We are interested in methods for learning interactions generally, which can describe a wide class of physical systems exhibiting collective dynamics defined by an interaction kernel, without a priori assumptions on the analytical form of this kernel (i.e. it is nonparametric). The advantage of this kernel-based approach is that it incorporates the underlying physics of the model (i.e. collective dynamics), which more general equation-learning approaches may ignore, potentially limiting their effectiveness for model accuracy and predictions. In this work, we extend existing variational learning approaches to collective systems with both interaction kernels and environmental/intra-agent forces. The proposed framework simultaneously infers the interaction kernel non-parametrically while learning the environmental force using either semi-parametric or fully nonparametric representations. The methodology is validated on several benchmark models exhibiting synchronization, alignment, attraction-repulsion, and external environmental forces. We also introduce a model-selection procedure based on our nonparametric learning framework to identify models that optimally explain a given set of trajectory observations. By exploiting the feature-identification capability of the learned models, the proposed procedure can distinguish among different collective dynamics frameworks and recover mechanistic interaction mechanisms directly from trajectory data.
1 Introduction
The paper develops structured, data-driven methods for learning collective dynamics while addressing the high dimensionality and unknown functional forms of governing mechanisms. It extends variational inference to jointly identify interaction kernels and environmental forces, with model selection for distinguishing mechanisms.
- Learning governing equations from trajectories can support prediction by generating forecasts from an inferred mechanistic model rather than directly extrapolating dynamical data.
- Treating collective dynamics as an arbitrary full-state vector field creates a high-dimensional, computationally demanding regression problem whose dimension grows with agent number.
- Collective interaction structure reduces inference from a full vector field to a low-dimensional interaction kernel depending on pairwise distance.
- Because agents may experience self-propulsion, friction, and other intra-agent effects, the framework simultaneously identifies interaction kernels and environmental forces.
- The paper introduces semi-parametric and fully nonparametric approaches, evaluates feature recovery and trajectory prediction, and studies sample complexity and observational-noise robustness.
- A model-selection procedure uses learned mechanistic features to distinguish collective-dynamics frameworks directly from trajectory observations.
2 Methods for learning collective systems
The methods infer environmental forces and interaction kernels from observed trajectories within a structured first-order collective-dynamics model. They represent unknown functions in finite-dimensional hypothesis spaces and estimate their coefficients by minimizing a trajectory-based least-squares functional.
- The framework infers both environmental and interaction forces from trajectory data, with analogous procedures available for second-order systems.
- The first-order model assumes identical environmental forces across agents and an interaction strength depending only on inter-agent distance.
- The observed data comprise replicated trajectories sampled at discrete times, with velocities required by the learning algorithm and pairwise distances computed from agent states.
- Unknown environmental and interaction functions are approximated in finite-dimensional hypothesis spaces, whose coefficients define the estimators.
- The error functional measures average squared prediction error over time and replicates, and minimizing it becomes a quadratic least-squares problem in the basis coefficients.
3 Evaluation of estimators
The evaluation measures recovery of mechanistic features, consistency with governing equations, and trajectory-prediction accuracy, while examining data requirements and noise robustness. Metrics emphasize regions explored by the dynamics and include both feature and trajectory errors.
- The evaluation assesses estimated environmental and interaction mechanisms alongside trajectory predictions using multiple performance measures.
- Feature-recovery metrics quantify how accurately variational estimators recover underlying interaction kernels and environmental forces in synthetic experiments.
- Weighted measures ρR and ρX bias accuracy assessment toward pairwise-distance and state-space regions highly explored by the dynamics.
- Trajectory RMS error captures average prediction fidelity and is less dominated by outliers than maximum pointwise trajectory error.
- Expanding training data while fixing other settings provides empirical evidence about the amount required to recover interaction mechanisms.
- The noise analysis notes that independent position, velocity, and acceleration noise is generally unrealistic because experiments typically measure position and derive derivatives.
4 Model selection
The model-selection procedure evaluates candidate collective-dynamics frameworks built from different dynamical orders and force mechanisms, then selects the best-supported parsimonious model from trajectory data. Candidates are first screened for compatibility and numerical stability, compared on validation trajectories, and ranked by complexity and residual error.
- Candidate frameworks: The procedure enumerates candidate frameworks representing specific combinations of dynamical order and interaction or environmental mechanisms.Candidates may include first- or second-order systems, with environmental, energy-based interaction, and alignment forces included or omitted.
- Data partitioning: Each candidate is fit using training trajectories and assessed using separate validation trajectories before final testing on unseen data.The test subset does not participate in model fitting or selection.
- Compatibility screening: Candidates are rejected when they fail trajectory or residual-error compatibility checks, cannot be integrated numerically, or exceed the regression system’s condition-number threshold of 10^12.The compatibility gate uses reconstruction and residual criteria before ranking admissible candidates.
- Failure condition: If no candidate passes compatibility checks, the procedure declares that no proposed first- or second-order framework adequately explains the dataset.The algorithm outputs either a selected candidate with learned interaction laws or a declaration of incompatibility.
- Validation: Among admissible candidates, validation trajectory error identifies the competitive set using absolute and relative tolerances of εabs = 10^-3 and εrel = 0.05.The trajectory threshold is defined relative to the minimum admissible validation error while also enforcing an absolute tolerance.
- Final ranking: If multiple candidates remain statistically tied, the procedure selects lower model complexity first and lower training residual error second.A unique validation candidate is selected directly; ties are resolved by parsimony and then residual error.
5 Results
The experiments show that both semi-parametric and fully nonparametric approaches recover collective-dynamics mechanisms and predict trajectories accurately, while model selection identifies active mechanisms across benchmark systems. Recovery benefits from correct parametric structure and diverse training trajectories, and remains robust to moderate observation noise.
- Kuramoto model: Both methods estimate Kuramoto interaction kernels with relative errors on the order of 10^-2, while their trajectory errors remain small and statistically similar.The semi-parametric method nearly exactly recovers the environmental force.
- Self-propelled particle (SPP) model: Both formulations closely recover the SPP interaction kernel where pairwise distances are well sampled, but semi-parametric environmental-force recovery improves by nearly two orders of magnitude.The kernel is difficult to estimate near the singular interaction distance of zero, where few trajectories provide data.
- Trajectory accuracy: Correct parametric structure yields approximately one order of magnitude smaller environmental-force error, while both methods retain trajectory errors of order 10^-6 to 10^-5.This indicates accurate reconstructions and predictions despite differences in feature recovery.
- Scalability: Across Kuramoto, SPP, and phototaxis systems, learned interaction mechanisms generalize to larger populations and reproduce qualitative collective behavior, with modestly smaller prediction errors for semi-parametric models.The models also predict over a horizon five times longer than training while scaling the number of agents by a factor of ten.
- Dependence on training data: Increasing trajectory count or temporal resolution improves mechanism recovery, but temporal resolution alone cannot compensate for insufficient replicate diversity.With M = 1, increasing L has little apparent effect because new time samples do not explore additional initial conditions or state-space regions.
- Robustness to noise: The algorithm remains capable of accurate extrapolated trajectory prediction at 10% relative observation noise, although mechanism estimates deviate most in poorly sampled regions.Well-supported regions remain largely unaffected as noise increases.
- Model selection: Model selection identifies active mechanisms across first- and second-order systems, including alignment, environmental, and energy interactions, while suppressing inactive terms.For the fully active second-order system, candidate S6 uniquely achieves the minimum validation trajectory error and matches the true generating framework.
6 Discussion and conclusions
The framework learns interaction and environmental forces through low-dimensional mechanistic representations, accurately recovering and predicting diverse collective dynamics while supporting model selection. Its scope is bounded by trajectory coverage, semi-parametric specification, and current focus on homogeneous agents and relatively low-dimensional laws.
- The framework learns interaction kernels and environmental forces directly, producing computationally tractable and interpretable mechanistic estimators.It exploits collective-dynamics structure instead of treating the system as an arbitrary high-dimensional vector field.
- Two formulations support environmental-force learning: semi-parametric estimation with a known functional form and fully non-parametric recovery from trajectories.The semi-parametric formulation estimates unknown parameters, whereas the fully non-parametric formulation is used when the environmental structure is unavailable.
- Across benchmark systems, the methods recover governing mechanisms and predict collective behavior beyond the training horizon.Benchmarks include synchronization, attraction/repulsion, velocity alignment, and externally driven motion.
- Feature recovery is most accurate in well-sampled state-space and pairwise-distance regions, while sparse trajectory coverage can degrade recovery.Semi-parametric estimates can also become biased or misleading when the assumed functional form is misspecified.
- The model-selection procedure identifies active mechanistic components, including environmental, energy-based, alignment, and combined interaction structures.It is intended for cases where the relevant dynamical structure is not known a priori.
- The present study primarily covers homogeneous-agent systems and relatively low-dimensional interaction laws.Future directions include heterogeneous agents, higher-dimensional dependencies, stochastic dynamics, and more complex environmental couplings.
A Variational methods for learning collective and environmental forces: fully non-parametric approach (second-order)
This second-order formulation infers interaction and environmental forces jointly from particle trajectories. It assumes identical environmental forces across agents and, for the studied systems, removes spatial dependence from that force.
- The algorithm targets second-order collective-dynamics models with both interaction and environmental forces inferred from observed trajectories.
- The environmental force is modeled as an identical function f of each agent’s velocity for all systems studied.Although initially defined as f(x_i,v_i), the force is simplified to f(v_i) because no spatial dependence was found.
A.1 Trajectory data
The trajectory-data procedure collects positions, velocities, and acceleration information from repeated discrete-time observations, then constructs empirical quantities and solves a joint variational estimation problem for the two forces.
- Trajectory observations include the system state and velocity at discrete times across repeated experiments.The observation times satisfy 0=t_1<t_2<⋯<t_L=T.
- The observed trajectories provide the data used to estimate both the environmental force f and interaction kernel ϕ.
- The second-order algorithm requires acceleration observations or approximations in addition to positions and velocities.Acceleration is defined through the derivative of velocity and may be approximated from velocity data.
- The estimator uses localized basis functions for environmental-force and interaction hypotheses, with second-order environmental bases depending on state and velocity.
- The variational objective reduces to a quadratic problem in basis coefficients, solved through normal equations.The coefficient solution exists because the objective is quadratic.
- Algorithm 3 constructs pairwise quantities, estimates observed supports, assembles the linear system, and solves for the coefficients.
B Hypothesis spaces and empirical measures
The appendix specifies hypothesis spaces and trajectory-induced empirical measures used by the estimation algorithms. For the benchmark systems, visualizations document basis functions and induced measures for Kuramoto, phototaxis, and self-propelled-particle models.
- The appendix provides details on hypothesis spaces and trajectory-induced empirical measures for the model systems studied.
- For the Kuramoto model, linear B-spline basis functions define the interaction-kernel and environmental-force hypothesis spaces.The interaction-kernel basis uses n_ϕ=10 functions.
- The benchmark visualizations show basis functions and induced measures for the Kuramoto, phototaxis, and self-propelled-particle models.These quantities define the spaces and empirical measures used by the learning algorithms.
C Metric evaluations
The metric evaluations quantify interaction-kernel and environmental-force estimation accuracy, while empirical-measure comparisons assess whether training trajectories represent the explored state regions.
- Metric evaluations: Kuramoto training-data measures closely agree with reference measures for phase-difference and phase variables.The comparison uses ˆρ∆θ versus ρ∆θ and ˆρθ versus ρθ.
- Metric evaluations: Linear B-spline basis functions construct hypothesis spaces for interaction kernels and environmental forces in the phototaxis and SPP models.The velocity space uses 10 basis functions per dimension, yielding nf = 100 tensor-product interaction-kernel basis functions; the environmental-force space uses nϕ = 8.
- Metric evaluations: Phototaxis training-data measures closely agree with reference measures for position and velocity variables.The comparison uses ˆρR versus ρR and ˆρV versus ρV .
- Metric evaluations: The SPP training-data measures closely agree with reference measures for position and velocity variables.The comparison uses ˆρR versus ρR and ˆρV versus ρV .
C.1 Dependence on training data
Performance is evaluated across numbers of replicates and time samples, with reported errors summarized as means and sample standard deviations over repeated learning trials.
- C.1 Dependence on training data: For the Kuramoto model, reported entries include environmental-force errors of 8.103 × 10−2 ± 6.689 × 10−2 at M = 4 and L = 51, and 6.180 × 10−3 ± 1.905 × 10−3 at M = 256 and L = 801.These values come from the corresponding rows in the reported error table.
- C.1 Dependence on training data: For the reported M = 64 settings, environmental-force errors range from 1.150 × 10−2 ± 4.800 × 10−3 at L = 201 to 1.389 × 10−2 ± 5.573 × 10−3 at L = 401.The entries correspond to the rows with M = 64 and the indicated time-sample counts.
- C.1 Dependence on training data: For the reported M = 256 settings, interaction-kernel errors remain near 1.2 × 10−2 while environmental-force errors are approximately 6–7 × 10−3 across several L values.The rows report interaction-kernel and environmental-force errors for L = 51, 101, 201, 401, and 801.
- C.1 Dependence on training data: Tables 19–21 report relative interaction-kernel and environmental-force errors for varying M and L across the three benchmark models.M denotes the number of replicates and L the number of time samples.
C.2 Noise robustness
The framework’s robustness to observational noise is assessed through mechanism recovery, trajectory reconstruction, and prediction across the benchmark models.
- C.2 Noise robustness: Figure 23 visualizes how observational noise affects recovery of interaction and environmental mechanisms.The noise-robustness experiments concern the models discussed in Section 5.2.
- C.2 Noise robustness: Figure 24 shows the corresponding effect of observational noise on trajectory reconstruction and prediction.The two figures jointly evaluate noise effects on learned mechanisms and downstream trajectories.
D Limitations
Learning accuracy depends strongly on how broadly the initial-condition distribution samples the state space.
- D Limitations: Broader initial-condition coverage provides more informative observations and more accurate recovery of the environmental force.Narrowly concentrated initial conditions leave portions of the state space poorly sampled and reduce learning accuracy.
- D Limitations: A narrowly concentrated initial-condition distribution limits environmental-force learning because parts of the state space remain poorly sampled.The limitation follows from the dependence of performance on the support of µ0.
D.2 Basis functions
The variational algorithm’s recovery accuracy depends on the number of basis functions used for the interaction kernel and environmental force. For the SPP system, approximately ten basis functions provide the lowest average recovery errors and a balance between approximation accuracy and numerical stability.
- Recovery accuracy: Too few basis functions cause underfitting and larger feature-recovery errors.Increasing the basis dimension improves approximation quality until performance saturates.
- Numerical stability: Beyond saturation, additional basis functions may degrade performance because the least-squares equation becomes increasingly ill-conditioned for fixed training data.Thus, the approximately ten-function choice balances approximation accuracy and numerical stability.
- Robustness and initialization: The noise-robustness figures compare learned interaction kernels and environmental forces, along with trajectory recovery and prediction, across multiplicative noise levels.Figures 23 and 24 cover feature learning and trajectory behavior for the benchmark systems, while Figure 25 compares initial-condition distributions for SPP.
- Basis-dimension sensitivity: Figure 26 varies one B-spline basis dimension at a time while averaging recovery errors over tested values of the other dimension.The solid curves show mean relative weighted L2 error, and shaded regions show one standard deviation.
- Recovery accuracy: Approximately ten basis functions achieve the lowest average recovery error for both the interaction kernel and environmental force in the SPP system.This result applies to the kernel basis count nϕ and environmental-force basis count nf.
E.1 Data generation parameters
The numerical experiments use separate datasets for estimating forces and kernels, validating trajectories, and constructing reference probability densities. Interaction kernels use radial B-splines, while environmental forces use tensor-product B-splines with system-dependent environmental variables.
- Data roles: Training data estimate interaction kernels and environmental forces, testing data validate trajectories, and Mρ data construct reference probability densities.
- Basis representations: The framework approximates interaction kernels with radial B-spline bases and environmental forces with tensor-product B-splines.For first-order systems the environmental variable is the state; for second-order systems it is generally velocity.
E.3 System initial conditions
Initial conditions are system-specific and are designed to cover relevant state or velocity domains for learning. Phototaxis and SPP velocity initialization uses perturbed tensor grids, while parametric recovery is evaluated over independent trials.
- System initialization: For first-order Kuramoto dynamics, the initial condition is u0 = θ0; for second-order systems, it is u0 = (x0, v0).
- Phototaxis: Phototaxis velocities come from a perturbed tensor grid on [−1, 1]2, with random agent assignment and independent Gaussian perturbations.The construction provides approximately uniform velocity-domain coverage and improves tensor-product basis conditioning.
- SPP: SPP spatial initial conditions use uniformly populated spatial clumps, while velocities come from a perturbed tensor grid on [−2veq, 2veq]2.The velocity domain is tied to the equilibrium speed of the self-propulsion model and supports environmental-force learning.
- Trial structure: Parametric coefficients are reported as means and standard deviations over Tr = 10 independent trials.Coefficients for mechanisms absent from a system are expected to be close to zero.
E.6 Index of notation
The notation index defines the dimensions, hypothesis spaces, basis parameters, empirical measures, error metrics, time horizons, noise variables, and candidate-model selection quantities used throughout the framework.
- Hypothesis spaces and bases: The interaction-kernel domain is the observed distance interval [Rmin, Rmax], partitioned into P sub-intervals with nϕ = P × nP localized basis functions.
- Dimensions: The stacked state-space dimension is D = dN for first-order systems and D = 2dN for second-order systems.
- Estimators: The combined coefficient vector θ = (α, β) contains the coefficients for the environmental force and interaction kernel, with n = nf + nϕ total basis functions.
- Model variables: The environmental-force input is Z = X for first-order systems and Z = (X, V) for second-order systems.
- Error metrics: Feature-recovery errors use weighted L2 norms, while residual errors compare observed and predicted velocities or accelerations using RMS-based normalization.
- Time horizons: The training horizon is [0, T], and the prediction horizon is (T, Tf] beyond the training window.
- Noise robustness: Multiplicative observational noise is independently sampled from Unif[−ζ, ζ], with ζ controlling perturbation magnitude.