Source-linked AI summary

Data-Driven Control of Complex Networks

Giacomo Baggio, Danielle S. Bassett, Fabio Pasqualetti

arXiv:2003.12189v3eess.SYmath.OCphysics.soc-ph

TL;DR

Large complex networks are difficult to control reliably when their dynamics are unknown or their models are inaccurate. The paper develops data-driven optimal control formulas from finite experimental data, including arbitrary or random inputs, and reports favorable numerical and computational properties relative to model-based approaches. The framework is analytically established for linear dynamics and its behavior is also characterized with noisy data and nonlinear dynamics.

  • Problem

    Accurate and tractable network-dynamics models are difficult to obtain for large or time-varying complex networks, although model-based control depends on them.

  • Method

    The paper constructs closed-form and approximate point-to-point control inputs directly from finite offline experiments without requiring knowledge of the network dynamics.

  • Results

    The data-driven controls are reported to be more accurate and computationally efficient than model-based controls and can analyze and manipulate network controllability.

  • Takeaways & Limitations

    Finite experimental data can support optimal or approximate control of complex networks and provide an alternative way to study their controllability.

  • Takeaways & Limitations

    The main analytical results assume linear network dynamics, whereas many real-world networks are inherently nonlinear.

Abstract

from arXiv · show

Our ability to manipulate the behavior of complex networks depends on the design of efficient control algorithms and, critically, on the availability of an accurate and tractable model of the network dynamics. While the design of control algorithms for network systems has seen notable advances in the past few years, knowledge of the network dynamics is a ubiquitous assumption that is difficult to satisfy in practice, especially when the network topology is large and, possibly, time-varying. In this paper we overcome this limitation, and develop a data-driven framework to control a complex dynamical network optimally and without requiring any knowledge of the network dynamics. Our optimal controls are constructed using a finite set of experimental data, where the unknown complex network is stimulated with arbitrary and possibly random inputs. In addition to optimality, we show that our data-driven formulas enjoy favorable computational and numerical properties even compared to their model-based counterpart. Although our controls are provably correct for networks with linear dynamics, we also characterize their performance against noisy experimental data and in the presence of nonlinear dynamics, as they arise when mitigating cascading failures in power-grid networks and when manipulating neural activity in brain networks.

I. INTRODUCTION

Modern sensing and data availability create opportunities to understand and control complex networks, but model-based control remains difficult because accurate large-network models are challenging to construct and sensitive to uncertainty. The paper addresses this gap by learning finite-time point-to-point optimal controls directly from experimental data, including arbitrary or random experiments.

  • Motivation: Accurate models of large-scale networks are challenging to construct, and model uncertainties can make model-based network controls unreliable and highly sensitive.Uncertainties include missing or extra links and incorrect link weights.
  • Research gap: Existing data-driven control approaches commonly target closed-loop stabilization or tracking rather than finite-time point-to-point control.The paper identifies this as the gap motivating its control framework.
  • Contribution: The paper learns optimal controls that steer selected network nodes from an initial state to a desired final state within a finite horizon.The setting uses linear dynamics, quadratic costs, and offline control experiments.
  • Contribution: The data-driven formulas do not require optimal experiments and can use arbitrary, possibly random, control experiments.This allows the framework to work with experimental data that are already available rather than requiring a specially designed dataset.
  • Results: The paper establishes closed-form and suboptimal computationally simple expressions, characterizes the minimum experiments needed for exact reconstruction, and compares their numerical properties with model-based methods.Applications include restoring power-grid operation after faults and analyzing controllability in functional brain networks.
  • Implications: The resulting expressions also provide a computationally reliable and efficient way to analyze controllability in large network systems where Gramian-based methods are limited.The cited limitation concerns investigations restricted to small and well-structured topologies.

II. RESULTS

The paper formulates finite-horizon optimal point-to-point control for linear network dynamics and replaces model-dependent computation with expressions based solely on offline experimental data. With arbitrary or random experiments, finite data can reconstruct optimal inputs under rank conditions, while fewer independent trials can still achieve the target suboptimally.

  • Problem formulation: The control task steers the network output from y(0) = y0 to y(T) = yf in T steps while minimizing a quadratic combination of control effort and trajectory locality.The formulation assumes output controllability and uses tunable matrices Q and R to penalize output deviation and input usage.
  • Data-driven formulation: The proposed expressions use experimental input and output data collected with an unknown network matrix A, rather than requiring explicit knowledge of the network dynamics.The data are organized into matrices of inputs, intermediate outputs, and final outputs.
  • Experimental data: The experiments need not be optimal or informative by design; they may be arbitrary, random, or carefully chosen.This distinguishes the framework from approaches that require experiment-design procedures.
  • Exact reconstruction: Finite data can exactly reconstruct the optimal control input when the input data contain mT linearly independent experiments, equivalently when U0:T−1 has full row rank.Linear independence is described as a mild condition normally satisfied by randomly generated experiments.
  • Suboptimal regime: With at least p but fewer than mT independent trials, the data-driven control still reaches yf in T steps but generally incurs a higher cost.The control becomes optimal for any target when the collected independent trials are themselves optimal.

2. Data-driven minimum-energy control

The minimum-energy formulation yields data-driven controls that can converge to or exactly reproduce model-based optima under stated data conditions, while simpler expressions trade optimality for computational efficiency. These controls can outperform sequential identification-and-control procedures and remain computationally favorable for large networks.

  • Minimum-energy control: Setting Q = 0 and R = I recovers a data-driven minimum-energy control, whose infinite-data limit equals the minimum-energy solution under zero-mean finite-variance random inputs.The convergence result is asymptotic in the amount of data.
  • Approximate control: The approximate control is a simple suboptimal sequence that correctly steers the network to yf in T steps when p independent data are available.Noise analysis links the cost deviation to the worst-case control energy needed to reach a unit-norm target.
  • Data requirements: The exact data-driven expression becomes optimal with N = mT data, while its approximate counterpart approaches the optimum only asymptotically.Both controls achieve zero final-state error after N = p data in the reported comparison.
  • Numerical comparison: Data-driven strategies significantly outperform the standard sequential identification-and-control approach for both dense and sparse network topologies.The standard approach is more vulnerable to round-off sensitivity because it requires more operations.
  • Computational efficiency: For large networks, data-driven computation is normally faster because it operates on matrices typically smaller than the network matrix A instead of computing powers of A.The approximate control has the most favorable performance because of its particularly simple expression.

4. Data-driven controls with noisy data

The paper extends data-driven control analysis to noisy data and demonstrates the framework on power-grid dynamics. In the grid case, data-driven control restores operation after a fault, while noisy-data formulas require correction terms for asymptotic correctness.

  • Noisy data: With noisy data, the original data-driven controls are typically inconsistent, but modified formulas can become asymptotically correct by adding noise-variance correction terms.If noise affects only output or only input data, the corresponding original control formula requires no correction; small noise causes slight deviations.
  • The paper presents two applications to demonstrate the relevance and applicability of its data-driven control formulas.
  • Power-grid application: In the New England power grid, a line fault can desynchronize generators and trigger cascading failures or major blackouts if not mitigated promptly.The case study uses a 39-node network with 29 load nodes and 10 generator nodes.
  • Power-grid application: Using N = 4000 input/state experiments, the computed short control input steers generator phases and frequencies back toward steady-state operation after fault clearance.The control horizon is T = 400 samples, corresponding to 0.1s.
  • Power-grid application: The grid study finds that the data-driven input recovers correct operation, while its computation uses only pre-collected data and is optimal for linearized dynamics.The numerical study suggests the strategy can control complex nonlinear networks around an operating point.

2. Controlling functional brain networks via fMRI snapshots

The study uses task-based fMRI data to construct a data-driven control framework for functional brain networks without requiring known network dynamics. Compared with model-based control, it achieves nearly identical behavior for highly controllable targets, while using less energy but producing larger errors for less controllable targets.

  • Data and network construction: Task-based fMRI experiments encode six motor-task and visual-cue stimuli as binary input signals for brain-network control.The study includes finger tapping, toe squeezing, and tongue movement tasks, with BOLD time series as outputs.
  • Data and network construction: The brain is parcellated into p = 148 regions, and its functional-network dynamics are approximated with a low-dimensional n = 20 linear model.The approximation provides the model-based baseline used in the comparison.
  • Control comparison: For the most controllable targets, data-driven and model-based inputs exhibit almost identical behavior.The comparison evaluates both the final-state error and the norm of the control input.
  • Control comparison: For less controllable targets, data-driven control yields larger final-state errors but requires less control energy.The authors describe the lower-energy strategy as potentially more feasible in practice.
  • Interpretation and scope: Because the actual brain dynamics are unknown, final-state errors are computed using the identified linear model and may not reflect inaccuracies in the real brain.The study therefore treats the brain-network application as a numerical evaluation under an approximate model.
  • Interpretation and scope: The framework is presented as a viable alternative for inferring brain-network controllability and enforcing functional configurations non-invasively without requiring a network model.The paper also reports applications to analyzing and manipulating controllability properties in network systems.
  • Interpretation and scope: The study is restricted by its assumption of linear network dynamics, despite the inherent nonlinearity of many real-world networks.The authors note that linear models can nevertheless approximate nonlinear dynamics around desired operating points.

IV. METHODS

The paper contrasts model-based optimal control and data-driven identification for network systems, then describes applications to power-grid dynamics. Model-based formulas rely on controllability Gramians, while subspace methods estimate network matrices from input-output data.

  • Model-based control: Model-based optimal control uses a batch expression based on the network’s output controllability matrix and Gramian.The output controllability matrix collects the effects of inputs across T steps, and the Gramian is invertible exactly when the network is target controllable.
  • Model-based control: The classic Gramian-based minimum-energy expression is numerically unstable even for moderately sized systems.This motivates data-driven alternatives that avoid relying directly on an identified network model.
  • Subspace identification: With noiseless data, controllability in T−1 steps, and full-row-rank input data, the procedure correctly estimates A and B.These conditions provide the stated guarantee for the subspace-based identification procedure.
  • Subspace identification: The subspace procedure estimates the T-step controllability matrix by minimizing a Frobenius-norm residual, then extracts B and estimates A.The method first solves an optimization problem for the controllability matrix and uses its first m columns to estimate B.
  • Power-grid dynamics: The power-grid application models generator phases and frequency deviations with swing equations and discretizes them using forward Euler integration.The simulations perturb generator frequencies and initial conditions around steady state before applying data-driven control.

D. Task-fMRI dataset, pre-processing pipeline, and identification setup

The experimental sections preprocess HCP motor-task fMRI data, identify an approximate low-dimensional linear model, and evaluate data-driven controls in network simulations and applications.

  • Task-fMRI dataset: The brain-network study uses HCP motor-task fMRI data with six binary task inputs and BOLD outputs from 148 brain regions.The task channels are CUE, LF, LH, RF, RH, and T, while outputs follow the Destrieux 2009 atlas.
  • Pre-processing pipeline: BOLD measurements undergo minimal preprocessing, band-pass filtering to 0.06–0.12 Hz, and regression-based removal of physiological signals.The removed signals include cardiac, respiratory, and head-motion effects.
  • Identification setup: The data matrices are generated with sliding windows of length T = 100, assuming inputs and states are zero at times less than or equal to 10.The implementation uses singular-value-decomposition pseudoinverses with threshold 10^-8.
  • Identification setup: The fMRI input-output dynamics are approximated with a linear model of state dimension n = 20 identified from data in the interval [0, 150].Unstable identified matrices are stabilized by dividing A by ρ(A) + 0.01.
  • Network-control experiments: Figure 2 evaluates data-driven and model-based controls on Erdős–Rényi networks over 500 realizations, using Q = R = I and n = 100.The figure varies the number of data points and reports control cost and final-state error.
  • Network-control experiments: Figure 3 compares minimum-energy data-driven controls with model-based and identification-based strategies across data volume, network size, and computation time.Its curves average over 500 realizations in the main performance comparisons.
  • Applications: The study also includes data-driven fault recovery in the New England power grid and control of functional brain networks.These applications connect the framework to nonlinear generator dynamics and experimentally recorded neural activity.

I. EXPRESSION OF OPTIMAL DATA-DRIVEN CONTROLS FOR ARBITRARY

The data-driven framework represents control inputs as linear combinations of recorded experiments, allowing optimal controls to be reconstructed from finite input-output data under rank and controllability conditions.

  • Data-driven representation: By linearity, combining recorded input experiments with coefficients α produces the same combination of their output trajectories.This yields yT = YTα while using the input sequence U0:T−1α.
  • Optimal control reconstruction: If an optimal coefficient vector α⋆ exists for the target yf, the corresponding data-driven input is optimal for the stated control cost.The construction enforces yf = YTα while minimizing the quadratic objective.
  • Optimal control reconstruction: The unique optimal coefficient vector is obtained by projecting a pseudoinverse solution onto the feasible kernel component.The expression uses a basis of the kernel of YT and a matrix associated with the quadratic cost.
  • Minimum data requirements: At least p linearly independent optimal data are required to reconstruct the optimal control for arbitrary targets.Linear combinations of optimal controls remain optimal when their reached outputs are combined correspondingly.
  • Minimum data requirements: With p linearly independent data and full row rank of YT, the reconstructed input reaches any target yf, although it may be suboptimal.Thus data can guarantee target steering before they guarantee recovery of the optimal input.
  • Nonzero initial states: When experiments have varying initial states, the data-driven construction separates free and forced responses and uses the kernel of the initial-state matrix.The resulting condition permits reconstruction using input and initial-state data together.
  • Nonzero initial states: A sufficient rank condition with measured initial states requires at most mT + n linearly independent experiments.Under this condition, the optimal control can be reconstructed from a single uninterrupted trajectory partitioned into segments.

IV. CLOSED-FORM EXPRESSION FOR Q = 0 AND R = I (MINIMUM-ENERGY

For minimum-energy control, the data-driven expressions simplify substantially and can be related to pseudoinverse formulas; their approximation improves with data quantity and network excitability.

  • Minimum-energy formula: For Q = 0 and R = I, the general data-driven expression reduces to a compact minimum-energy control formula.The reduced expression is connected to the standard minimum-energy input through pseudoinverse identities.
  • Minimum-energy formula: The kernel-projection identity establishes equivalence between alternative data-driven expressions, even when the optimal input cannot be reconstructed directly.The equivalence follows from the relation between the data matrix and the output controllability matrix.
  • Approximate controls: The data-driven input generally differs from the minimum-energy input for finite data but approaches it as the number of experiments grows.This convergence is stated for independently generated Gaussian input experiments.
  • Approximation quality: Larger σmin(CT) yields lower approximation error because more excitable networks provide more favorable conditioning.Here σmin(CT) is the smallest nonzero singular value of the output controllability matrix.
  • Approximate controls: With random independent Gaussian input experiments, the data-driven input reaches yf when sufficiently many linearly independent experiments are available.The result assumes output controllability and full row rank of the experimental input matrix.

A. Data corrupted by small noise

Small perturbations preserve the correctness of the data-driven expressions under rank conditions, but i.i.d. measurement noise can produce persistent bias. Variance-corrected controls restore asymptotic optimality under suitable assumptions.

  • Robustness to small noise: Small perturbations in full-rank data matrices produce small deviations in the data-driven expressions.Continuity of the Moore–Penrose pseudoinverse supports this stability result.
  • Robustness to small noise: Pseudoinverse truncation with ε > 0 preserves the rank needed for stability of the optimal control.This applies because LKYT is not typically full rank.
  • Persistent bias under noise: Noisy data typically bias the original data-driven controls, which do not converge to the true input as N →∞.The approximate minimum-energy control exhibits this failure even in a scalar system with T = 1.
  • Noise correction: The corrected control sequence converges almost surely to the optimal control input as N →∞ for sufficiently small ε and full-row-rank input data.The result relies on i.i.d. noise with zero mean and finite variance.
  • Noise correction: Corrected minimum-energy expressions are likewise asymptotically correct for Q = 0 and R = I.The paper gives corrected versions for the full, compact, and approximate data-driven controls.
  • Noise correction: Correction terms depend on the noise source: the full control corrects input and output noise, while compact forms correct only one source.With output noise only, one compact corrected expression coincides with the original control.
Loading 2003.12189v3…