Source-linked AI summary
Optimal experimental design and some related control problems
Luc Pronzato
TL;DR
The paper asks how experimental design can support estimation, prediction, optimization, and control across parametric and nonparametric settings. It surveys mathematical foundations, asymptotic estimator behavior, optimal inputs, and designed perturbations. Its conclusions identify open challenges in adaptive control, nonparametric models, robust control, and practical DOE methods.
Problem
Experimental design must support differing estimation, prediction, optimization, and control objectives, while adaptive control and nonparametric settings retain unresolved difficulties.
Method
The paper surveys optimal experimental design, information-based parameter estimation, sequential designs, optimal inputs, and designed perturbations for control.
Results
The survey relates DOE to control and reports that suitable perturbations and estimation strategies can support consistency, while direct nonlinear-feedback control is not sufficient under random disturbances.
Takeaways & Limitations
DOE should be driven by the intended application of identification and extended toward nonparametric models, robust control, and non-standard situations.
Takeaways & Limitations
The paper’s treatment of optimal input design depends on unknown model parameters, and the survey is not exhaustive.
Abstract
from arXiv · showhide
This paper traces the strong relations between experimental design and control, such as the use of optimal inputs to obtain precise parameter estimation in dynamical systems and the introduction of suitably designed perturbations in adaptive control. The mathematical background of optimal experimental design is briefly presented, and the role of experimental design in the asymptotic properties of estimators is emphasized. Although most of the paper concerns parametric models, some results are also presented for statistical learning and prediction with nonparametric models.
1 Introduction
The paper presents optimal experimental design as a bridge linking estimation, prediction, optimization, and control. It surveys unresolved challenges, especially for nonparametric models, robust control, adaptive control, and non-standard designs.
- Motivation: DOE extracts useful information from data to estimate unknown quantities and connects optimization, estimation, prediction, and control.The paper presents DOE as relevant whenever unknown quantities must be estimated and the estimation method remains selectable.
- Objectives: The paper explains inherent difficulties when estimation is combined with optimization or control, including why adaptive control is intrinsically difficult.It also discusses tentative remedies and possible future developments.
- Research directions: Classical DOE relies on persistence of excitation, but many issues remain open in other situations.The paper argues that design criteria should reflect the intended application of the identified model.
- Research directions: Prospective directions include DOE for nonparametric models and robust control, while algorithms and practical methods for non-standard situations remain missing.These directions are presented as part of the paper’s main message concerning DOE and control.
- Scope: The survey is not exhaustive: nonparametric modelling is treated briefly and only for static systems, among other stated scope limits.The authors state that none of the results is really new, while their collection in one document may be useful.
2 Examples of applications of DOE
Examples show how DOE improves parameter estimation by choosing informative measurements, while model discrimination and optimization require sequential designs that adapt to competing predictions and objectives.
- A weighing problem: Hadamard weighing reduces each estimate’s variance to σ2/8 using eight observations, whereas method a needs 64 observations for the same precision.Method a weighs objects successively; method b uses eight orthogonal configurations.
- A weighing problem: In linear models, DOE seeks design vectors that make the information matrix well-conditioned by minimizing a scalar function of M_N^-1 or maximizing one of M_N.The weighing problem becomes combinatorial when design components are restricted to −1, 0, or 1.
- Dynamical parameter estimation: The pharmacokinetic example compares conventional and D-optimal sampling designs, both using eight observations, with the optimal design repeating selected times.The optimal times are t*= (1, 1, 10, 10, 74, 74, 720, 720) minutes.
- Dynamical parameter estimation: The optimal pharmacokinetic design yields much more precise parameter estimates than the conventional design, but four sampling times cannot test model validity.The limitation follows because the optimal design uses only four distinct sampling times for four parameters.
- Model discrimination: For model discrimination, the next design point is placed where fitted competing models’ predictions differ most.With more than two structures, the procedure uses the best- and second-best-fitting models.
- Optimization: Sequential optimization designs create a dual-control problem because each input must both estimate parameters and seek high responses.Feedback from observations enters the sequence of design points, giving the initially static problem a dynamical aspect.
3 Statistical learning, nonparametric models
The paper contrasts model-free and model-based DOE for nonparametric prediction, emphasizing local uncertainty reduction and the distinct design requirements of prediction versus parameter estimation.
- Statistical learning: Statistical learning uses training data to predict unsampled responses with methods including kernel approaches, SVM regression, radial basis functions, and Kriging.The paper focuses on Kriging because of its flexibility and interpretability.
- Kriging: Kriging represents observations with a stationary random process and produces a best linear unbiased predictor subject to an unbiasedness constraint.The predictor is written as a linear combination of observed responses with weights chosen to minimize prediction error.
- Kriging: Without measurement errors, Kriging has zero MSE at observed sites and therefore perfectly interpolates deterministic computer experiments.The random-process representation supplies model uncertainty even for deterministic systems.
- Nonparametric DOE: Nonparametric designs may use maximin distance, minimax distance, maximum MSE, integrated MSE, or maximum entropy criteria.Space-filling designs maximize separation or distribute observations across the design space.
- Nonparametric DOE: In nonparametric prediction, adding an observation mainly reduces MSE near its input, so asymptotically optimal designs should distribute points throughout U.This local influence differs from the situation described for parametric models.
- Model-based design: Estimating covariance and variance parameters from data affects prediction precision, and space-filling designs are not appropriate for precise estimation of those parameters.This issue is identified as having received little attention.
4 Parametric models and information matrices
This section develops parametric regression and estimation frameworks in which asymptotic estimator behavior is determined by the experimental design. It also connects information-based estimation with estimating functions and adaptive or dynamical-system settings.
- Asymptotic estimation: Design points may be random or generated with empirical measures converging to a limiting design measure ξ, which characterizes the estimator’s asymptotic distribution.The two cases are i.i.d. random design and deterministic designs with finitely supported limiting proportions.
- Asymptotic estimation: Under continuity, boundedness, smoothness, and estimability conditions, the LS estimator is consistent, while its asymptotic covariance depends on the design and weighting function.The relevant covariance expression is C(w, ξ, θ̄)/N in nonlinear regression.
- Weighted least squares: For WLS, choosing w(u) = c σ^-2(u) yields C(w, ξ, θ̄) = M^-1(ξ, θ̄), and this weighting is asymptotically optimal among WLS estimators.For any weighting function, C(w, ξ, θ̄) − M^-1(ξ, θ̄) is non-negative definite.
- Likelihood-based estimation: With normal errors, ML estimation coincides with optimally weighted LS, whereas other error distributions produce different ML estimators and corresponding Fisher information.For i.i.d. errors, the Fisher information for location is constant across design points.
- Estimating functions: Estimating functions provide broadly applicable parameter-estimation tools, including simpler dynamical-system estimators and instrumental-variable methods when regressors and errors are correlated.In the discrete-time example, substituting noisy observations for the state produces a computationally simpler estimator that is less precise than LS.
- Limitations: DOE becomes more difficult for nonlinear models because local criteria depend on a guessed parameter value and standard covariance results are asymptotic, while finite-sample approximations are complicated.Additional complications arise for correlated observations, autoregressive models, and some recursive estimators.
5 DOE for parameter estimation
This section formulates optimal DOE for parameter estimation through scalar criteria of the Fisher information matrix. Classical criteria connect information optimization to confidence-ellipsoid geometry and prediction precision.
- Information-based design: Optimal experiments are designed by minimizing or maximizing scalar functions of the average-per-sample Fisher information matrix.Finite designs specify N observation points, while the associated information matrix summarizes their estimation content.
- Inference: Information-based confidence regions for parameters can be transformed into simultaneous confidence regions for functions of those parameters.The stated coverage property is asymptotically valid in nonlinear situations.
- Classical criteria: A-optimality minimizes the trace of M^-1, E-optimality maximizes the minimum eigenvalue of M, and D-optimality maximizes det(M).These criteria respectively target confidence-ellipsoid axis lengths, the longest axis, and the determinant-based volume.
- Classical criteria: D-optimality is invariant under re-parameterization and often produces designs that replicate a small number of experimental conditions.The section illustrates this pattern with duplicated sampling times in a D-optimal design.
- Finite-design optimization: For finite designs, local optimization selects support points in the admissible set, but standard and exchange-type algorithms generally provide only locally optimal solutions because multiple local optima may exist.The design problem becomes optimization over N × d variables subject to admissibility constraints.
5.3 Approximate design theory
Approximate design theory represents experiments as probability measures over support points and exploits convexity, information-matrix geometry, and equivalence theorems. Sequential algorithms can converge to optimal designs, although practical speed and nonlinear dependence create limitations.
- Design measures: A finite design with replicated support points induces weights equal to replication proportions; relaxing these constraints yields approximate designs and, further, arbitrary design measures ξ on U.The information matrix is a convex combination of rank-one information contributions.
- Design measures: Every information matrix from a design measure can be represented using at most p(p + 1)/2 + 1 support points, including an optimal design.A discrete design with repetitions approximates the support weights by choosing replication counts whose proportions approach λi.
- Multiple-output models: The same framework extends from scalar to multiple-output observations with covariance-weighted information matrices, and the support-point bound remains m ≤ p(p + 1)/2 + 1.The multiple-output generalization uses the inverse output covariance matrix as the weighting structure.
- Optimality properties: For D-optimality, the Kiefer-Wolfowitz equivalence theorem states that optimality is equivalent to maxu∈U dθ(u, ξD) = p and to minimizing the maximum sensitivity function.Every support point of ξD satisfies dθ(ui, ξD) = p.
- Optimality properties: D-optimality is equivalent to G-optimality because it minimizes the maximum asymptotic prediction variance over the experimental domain.This connects parameter-space information optimization with worst-case prediction precision.
- Sequential algorithms: Sequential steepest-ascent methods add the point maximizing dθ(u, ξk); Wynn’s algorithm converges under standard step-size conditions, while Fedorov’s variant ensures monotonic convergence toward a D-optimal measure.These methods guarantee convergence, but basic steepest-ascent procedures can be slow because existing support points are never fully removed.
6 Control in DOE: optimal inputs for parameter estimation in dynamical models
The paper connects input design for dynamical-system identification with control, using Fisher information to shape informative inputs and control-oriented uncertainty criteria. It also highlights parameter dependence, finite-spectrum optimal inputs, and unresolved challenges for adaptive or robust applications.
- Control in DOE: optimal inputs for parameter estimation in dynamical models: Input design for dynamical systems treats the control input as part of the experiment and uses Fisher information to characterize asymptotic estimator precision.The presentation focuses on single-input single-output systems, with extensions to multi-input multi-output systems.
- Control in DOE: optimal inputs for parameter estimation in dynamical models: For Box–Jenkins models, the input sequence affects estimation precision for parameters in F but not for parameters in G under the stated open-loop assumptions.The result assumes F and G have no common parameters.
- Control in DOE: optimal inputs for parameter estimation in dynamical models: In the frequency-domain framework, the experimental domain becomes frequency and the design measure becomes the input power spectral density; an optimal input with discrete spectrum exists.Such inputs can be sought among finite combinations of sinusoidal components, with frequencies and associated input powers as support points and weights.
- Control in DOE: optimal inputs for parameter estimation in dynamical models: Control-oriented criteria can constrain frequency-dependent transfer-function uncertainty through H∞-related formulations, with special cases connected to E-, G-, D-, and minimax-optimal design.Uniform weighting and white noise yield the stated connection to G-optimal and D-optimal design.
- Control in DOE: optimal inputs for parameter estimation in dynamical models: Unknown model parameters make optimal input criteria parameter-dependent, motivating designs averaged or minimized over nominal values, or constructed sequentially.Sequential design generally uses many observations per step and only a few steps to achieve suitable performance.
- Control in DOE: optimal inputs for parameter estimation in dynamical models: Adaptive control raises difficulties because its input serves objectives beyond estimation, so the stationarity or persistence-of-excitation conditions underlying asymptotic results may fail.The paper identifies robust-and-adaptive controller construction as an open issue.
7 DOE in adaptive control
The paper examines adaptive control as a setting where input design conditions used for asymptotic estimation may break down because the input has objectives beyond parameter estimation.
- DOE in adaptive control: The asymptotic results for optimal design rely on estimator properties under stationarity-like conditions, including random design or persistence of excitation.Adaptive control can violate these conditions because its input has another objective than estimation.
- DOE in adaptive control: Adaptive control therefore motivates examples of estimation difficulties arising from the sequential construction of the design.The paper introduces these examples after noting that the required conditions may fail.
7.1 Examples of difficulties
The examples show that sequential or performance-driven inputs can make information matrices asymptotically singular or estimators inconsistent, creating a conflict between learning and control objectives.
- 7.1 Examples of difficulties: Least-squares asymptotic normality may fail when information matrices remain nonsingular at finite sample sizes but converge to singular matrices.The relevant condition is λmin[MF(U_N)] → 0 as N → ∞.
- 7.1 Examples of difficulties: In a feedback-controlled regression example, the LS estimator is non-consistent because future design points depend on previous measurement errors.Although λmin(M_N) → ∞, the information does not grow fast enough and M_N/N becomes singular.
- 7.1.1 ARX model and self-tuning regulator: In self-tuning regulation, certainty-equivalent optimal control can make the information matrix singular and render the true parameter non-estimable.When the true parameter is known, the optimal controller produces regressors orthogonal to the parameter, yielding r_k^Tθ̄ = 0.
- 7.1.1 ARX model and self-tuning regulator: Additional perturbations can ensure sufficiently rapid information growth, but persistent excitation conflicts with the global-convergence objective.The paper states this conflict through the relationship ||θ̄||^2λmin(M_N) < R_N.
- 7.1.2 Self-tuning optimizer: For self-tuning optimization, fixing the control at the optimum makes the information matrix singular, while perturbations are needed for consistency and oppose the performance objective.With periodic disturbance magnitude α, the output exponentially converges to an O(α^2) neighborhood of the extremum.
7.2 Nonlinear feedback control is not the answer
The nonlinear-feedback example demonstrates that stability of a closed loop does not guarantee consistent parameter estimation under random measurement disturbances. Combining nonlinear feedback with classical estimation and suitably exciting perturbations is presented as a promising direction.
- 7.2 Nonlinear feedback control is not the answer: Nonlinear-feedback control stabilizes systems with unknown parameters, but the paper distinguishes stability from consistency under random disturbances.The section argues that direct application of NFC fails in the presence of random disturbances.
- 7.2 Nonlinear feedback control is not the answer: The paper contrasts noisy NFC estimation with more classical estimation techniques, which have satisfactory behavior in the same setting.This comparison motivates combining NFC with traditional estimation methods and suitably exciting perturbations.
- 7.2 Nonlinear feedback control is not the answer: The NFC construction uses an auxiliary parameter estimator and certainty-equivalent feedback, with Lyapunov analysis showing state convergence and parameter-error convergence in the noiseless case.The estimator follows θ̂˙ = x(x + 1), while the controller uses u = −(a + θ̂)x − θ̂.
- 7.2 Nonlinear feedback control is not the answer: Under measurement noise with σ = 0.5, substituting observations for the state causes the parameter estimates not to converge and the state not to be driven to zero.The noisy trajectories are shown with dash-dotted and dotted curves in Figure 2.
7.3 Some consistency results
Consistency of parameter estimates depends on excitation conditions, and regulation objectives can conflict with the perturbations needed for identification. Bayesian imbedding weakens these requirements, while adaptive optimization can balance estimation and optimization through input selection.
- Consistency and excitation: Regulation objectives can conflict with consistent parameter estimation because inputs that achieve regulation asymptotically vanish and provide insufficient excitation.Perturbations may therefore be required, with their minimal necessary amount remaining an important question.
- Least-squares consistency: Least-squares consistency has different conditions for deterministic, independent, and observation-adapted regressors.The supplied results distinguish sufficient, necessary-and-sufficient, and almost-sure conditions across these settings.
- Least-squares consistency: Persistence of excitation is stronger than the logarithmic eigenvalue condition sufficient for consistency in the stated stochastic setting.The condition [log λmax(MN)]^1+δ = o[λmin(MN)] a.s. is described as much weaker than requiring MN to grow at the same speed as N.
- Bayesian imbedding: Bayesian imbedding can establish consistency under conditions as weak as those for non-random regressors, but only almost surely for parameter values drawn from the prior.Exceptional parameter values may still lack consistency.
- Adaptive regulation: In self-tuning regulation, Bayesian imbedding can achieve global convergence without perturbations, unlike least squares with forced certainty equivalence control.The least-squares approach requires perturbations and has control objective growth of at least log n.
- Adaptive optimization: A weighted input strategy compromises optimization and D-optimal estimation, yielding strong consistency while inputs concentrate at the optimal location.The stated Bayesian result requires αk →∞ and αk/k →0 with i.i.d. normal errors.
7.4 Finite horizon: dynamic programming and dual control
Finite-horizon adaptive optimization is a stochastic dynamic programming problem because each input affects both current reward and future parameter uncertainty. Practical control strategies simplify this dual effect, often with limited loss relative to more active approaches.
- Dynamic programming: Finite-horizon self-tuning optimization is formulated as stochastic dynamic programming with inputs chosen to maximize expected cumulative response.The objective is non-additive in the design setting but remains of stochastic dynamic programming type.
- Dual control: Each control input has a dual effect: it changes the current objective value and the future uncertainty represented by posterior measures.This coupling makes the optimization difficult because it embeds maximizations and expectations.
- Approximate strategies: Forced certainty equivalence and open-loop-feedback-optimal control simplify the problem by approximating future posteriors without fully propagating future observations.Their passive treatment of future information explains their frequent use.
- Active control: Active-control strategies account for future decisions' effects on posterior precision, but active alternatives generally provide only marginal improvements over passive strategies.The active strategy based on linearization and extended Kalman filtering is described as relatively complex and little used.
- Approximate strategies: For linear-Gaussian models, a small-noise approximation yields a control strategy within O(σ^4) of the optimal unknown strategy.This approximation exploits linearity to propagate an approximation of the expected future maximum.
8 Sequential DOE
Sequential DOE adapts design points to parameter estimates and can connect estimation with optimization. Full-sequential designs may converge to optimal designs, whereas more active finite-horizon strategies often improve only marginally over passive ones.
- Full-sequential design: Full-sequential DOE introduces one design point after each observation and adapts subsequent points using the current parameter estimate.For D-optimality, the next point maximizes the estimated variance-reduction function dθ(u, ξ).
- Active sequential design: Active sequential DOE can be formulated as stochastic dynamic programming because future design choices affect both the objective and information gathered.The resulting active strategies are difficult to solve exactly and generally offer only marginal gains over FCE control.
- Motivation: Sequential design is natural when each sampled response should be large, because estimation-focused designs may select points far from the unknown optimum.The sequential approach alternates observing responses and updating the parameter estimate.
- Asymptotic behavior: Under suitable conditions, sequential designs converge to the optimal design, and Bayesian sequential designs can be strongly consistent when an identifiability condition holds.The result is stated for full-sequential design and extends through Bayesian estimation under the specified condition.
9 Concluding remarks and perspectives in DOE
The concluding discussion identifies open DOE problems spanning correlated and nonlinear models, nonparametric learning, robust optimization, and control under stability constraints. It emphasizes designing experiments around their intended purpose while acknowledging limited theory in several settings.
- Scope and limitations: DOE with correlated observations has relatively few results, although correlated errors are classical in adaptive control.The paper points to recent developments in DOE and strong-law results from the control literature.
- Scope and limitations: Nonlinear dynamical models permit Fisher information construction through simulation, but few results address their nonlinear dependence on model parameters.The main difficulty is parameter nonlinearity rather than constructing the information matrix itself.
- Future directions: Non-stationary experiments with asymptotically vanishing inputs are proposed as a possible alternative to random perturbations when persistence of excitation is unavailable.The discussion connects this direction with exploration–exploitation compromises in adaptive control.
- Nonparametric models and active learning: DOE for nonparametric models remains at an early stage because its design problems are complex, while active strategies address decisions that jointly affect information and performance.The paper connects this challenge with active learning and reinforcement learning.
- Design criteria: DOE criteria should reflect the intended application, including robust-control objectives that motivate minimax design.Minimax design can also address dependence of locally optimal designs on unknown model parameters.
- Global optimization: Expensive global optimization motivates sequential surrogate-based sampling that balances exploration and exploitation through expected improvement.The expected-improvement procedure updates a Kriging model after each evaluation and supplies a stopping criterion when expected improvement becomes small.
- Adaptive control perspectives: In the NFC example, the estimator is inconsistent, whereas least squares converges quickly to the true parameter and an estimating-function estimator converges more slowly with lower computational cost.A switching strategy combining estimators can drive the state to zero, but its encouraging simulations also expose further unresolved questions.
- Adaptive control perspectives: A central open problem is designing informative, possibly vanishing input sequences subject to stability constraints while combining stability-oriented control with simple predictors.The paper also identifies robust-and-adaptive controller design as a prospective direction.