Source-linked AI summary
Concurrent learning for parameter estimation using dynamic state-derivative estimators
Rushikesh Kamalapurkar, Ben Reish, Girish Chowdhary, Warren E. Dixon
TL;DR
Large transient state-derivative estimation errors challenge concurrent-learning parameter estimation. The paper combines a dynamic derivative estimator with history-stack purging and models the resulting system as switched, establishing convergence under persistent excitation and ultimate boundedness under finite excitation.
Problem
Large transient derivative-estimation errors challenge concurrent-learning parameter estimation because recorded data are reused from the history stack.
Method
The estimator uses an observer-based dynamic state-derivative estimate and a novel purging algorithm that removes possibly erroneous history-stack data, with the closed-loop system modeled as switched.
Results
Under persistent excitation, error states converge asymptotically to the origin; under finite excitation, they decay to an ultimate bound that can be reduced by increasing learning gains.
Takeaways & Limitations
The developed estimator outperforms numerical differentiation-based techniques in the reported comparison and provides bounded or convergent error behavior under the stated excitation conditions.
Abstract
from arXiv · showhide
A concurrent learning (CL)-based parameter estimator is developed to identify the unknown parameters in a linearly parameterized uncertain control-affine nonlinear system. Unlike state-of-the-art CL techniques that assume knowledge of the state-derivative or rely on numerical smoothing, CL is implemented using a dynamic state-derivative estimator. A novel purging algorithm is introduced to discard possibly erroneous data recorded during the transient phase for concurrent learning. Since purging results in a discontinuous parameter adaptation law, the closed-loop error system is modeled as a switched system. Asymptotic convergence of the error states to the origin is established under a persistent excitation condition, and the error states are shown to be ultimately bounded under a finite excitation condition. Simulation results are provided to demonstrate the effectiveness of the developed parameter estimator.
I. INTRODUCTION
The paper develops concurrent learning for parameter estimation without numerical differentiation, using a dynamic state-derivative estimator and purging of erroneous history data. It establishes asymptotic convergence under persistent excitation and ultimate boundedness under finite excitation.
- Parameter convergence matters because closed-loop stability and controller performance depend critically on estimates reaching their ideal values.
- Existing concurrent-learning methods require known or estimated state derivatives, while numerical smoothing introduces unquantified error and extra data processing.Smoothing also requires storing data over a time window containing the point of interest.
- The paper uses an observer whose dynamic derivative estimate converges exponentially to a neighborhood of the actual state derivative.This replaces numerical smoothing in the concurrent-learning estimator.
- A novel purging algorithm removes possibly erroneous transient data from the history stack, addressing large parameter errors caused by derivative-estimation errors at recorded points.Because purging makes the adaptation law discontinuous, the resulting closed-loop error system is modeled as switched.
- With enough data to repopulate the history stack after each purge, the error states converge asymptotically to the origin under persistent excitation.Persistent excitation is sufficient to ensure enough data can be recorded after each purge.
- Under sufficiently long finite excitation, the error states decay to an ultimate bound that can be made arbitrarily small by increasing the learning gains.Simulations demonstrate the method under measurement noise.
II. SYSTEM DYNAMICS
The system is an uncertain, nonlinear, control-affine system whose unknown component is linearly parameterized. The estimator must identify the unknown parameters while the state derivative remains unavailable.
- The dynamics are modeled as nonlinear, uncertain, and control-affine, with locally Lipschitz system functions.
- The uncertain dynamics are linear in an unknown constant parameter vector through a known regressor.The parameter vector has a known norm bound.
- The objective is to design a parameter estimator for the unknown parameters, assuming a stabilizing controller keeps the state, state derivative, and input bounded.
- The state is available for feedback, but its derivative is assumed to be unknown.
III. CL-BASED ADAPTIVE DERIVATIVE ESTIMATION
The estimator combines concurrent learning with a dynamically generated state-derivative estimate, avoiding numerical smoothing while adapting parameters using current and recorded data. Its parameter error is tied to derivative-estimation error, motivating an estimator designed to drive that error toward zero.
- The paper uses an adaptive observer to generate the state-derivative estimates required for concurrent learning.
- Unlike numerical smoothing, the proposed method uses a dynamically generated state-derivative estimate for concurrent learning.Numerical smoothing requires additional processing and storage over a time window.
- Concurrent learning computes parameter-estimation information at recorded data points from the state, control, and state-derivative estimate.The history stack contains recorded tuples of these quantities.
- The parameter update drives estimation error to a ball whose size is of the order of the state-derivative estimation error.Reducing derivative-estimation error therefore reduces the attainable parameter-estimation error.
- The state-derivative estimator uses state-estimation-error feedback and constant learning gains to generate the derivative estimate.Its update includes an auxiliary signal and gains k1, α1, and γ1.
A. Purging of history stacks
The purging strategy addresses erroneous history-stack data collected during transients by replacing older data with newer data as the derivative estimate improves. This keeps concurrent learning data more representative while producing a discontinuous adaptation law.
- Transient state-derivative estimation errors can make stored data produce large parameter-estimation errors even when parameters converge.The resulting parameter estimates may converge to a neighborhood rather than the ideal values.
- Because the derivative estimator converges exponentially to a neighborhood of the true derivative, newer data is guaranteed to represent the system better than older data.
- Purging makes the parameter adaptation law discontinuous, requiring the closed-loop error system to be treated as a switched system.
B. Algorithm to record the history stack
The history-stack algorithm maintains informative data by collecting candidate points, monitoring rank quality through the minimum singular value, and replacing the active stack when a threshold is met. Its analysis establishes convergence under persistent excitation and arbitrarily small error under finite excitation with suitable design choices.
- Algorithm to record the history stack: The history stack is initialized full rank, while newly collected data populate a separate candidate stack.
- Algorithm to record the history stack: Once the candidate stack is full rank and exceeds a threshold, it replaces the active history stack, after which the candidate stack is purged.
- Algorithm to record the history stack: The dynamic threshold is a fraction of the highest minimum singular value previously encountered for the active history stack.The threshold fraction ξ lies in (0, 1).
- Algorithm to record the history stack: The algorithm maintains a positive lower bound on the minimum singular value of the active history-stack matrices.This supports the rank condition used in the stability analysis.
- Analysis: Under persistent excitation, the parameter-estimation error asymptotically decays to zero.
- Analysis: Under finite excitation, the parameter-estimation error can be made as small as desired by selecting the dwell time T and learning gains according to theorem conditions.
A. Asymptotic convergence with persistent excitation
Under persistent excitation, the purged history-stack estimator is analyzed as a switched nonlinear system and its error states converge asymptotically to the origin.
- Switched-system model: Purging makes the closed-loop error dynamics a switched nonlinear system, with each history stack defining a subsystem.Switching occurs when the history stack is purged; the algorithm avoids Zeno behavior through a minimum dwell time.
- Stability analysis: The analysis uses filtered tracking and parameter errors together with a candidate Lyapunov function to establish boundedness across switching intervals.The error variables include the parameter error, state error, and filtered tracking error.
- Stability analysis: Increasing the learning gains can make the state-derivative estimation error arbitrarily small before the next switching instance.This result is used alongside boundedness estimates in the inductive convergence argument.
- Persistent excitation result: The complete error state, including parameter, state, and filtered tracking errors, converges asymptotically to the origin.The proof combines the preceding Lyapunov bounds through an inductive argument over switching intervals.
- Persistent excitation result: Under persistent excitation, the history stack can be repeatedly purged and replenished with full-rank data, enabling asymptotic convergence of the state-derivative and parameter estimates.The result assumes sufficient data can be recorded after each purge and the required gain conditions are satisfied.
B. Ultimate boundedness under finite excitation
Without persistent excitation, the switched-system analysis instead establishes ultimate boundedness when the history stack is updated under the stated finite-excitation and dwell-time conditions.
- Finite excitation condition: The finite-excitation result removes the persistent-excitation hypothesis while retaining the other hypotheses of the convergence theorem.The result is formulated for the same purging-based switched system.
- Boundedness result: The Lyapunov function remains bounded after the final history-stack update when the stack is not updated thereafter.An inductive argument extends the bound over the remaining time interval.
- Finite excitation condition: A minimum dwell-time condition is imposed while the history stacks are updated and populated using the purging algorithm.The dwell time controls the intervals over which the Lyapunov envelopes decay.
VI. SIMULATION RESULTS
Simulations on a noisy two-link robot-manipulator model compare the dynamic state-derivative estimator with numerical differentiation-based concurrent learning and report improved performance for both noise levels.
- Simulation setup: The developed estimator is evaluated on a nonlinear two-link robot manipulator with four unknown parameters and additive Gaussian measurement noise.Low-noise and high-noise cases use state-measurement noise variances 0.005 and 0.1, respectively.
- Simulation setup: The comparison baseline uses polynomial regression over collected-data windows and moving-average filtering to estimate numerical derivatives.Multiple runs vary gains, window sizes, and thresholds, with settings selected by steady-state RMS error over five runs.
- Comparative results: The developed technique outperforms numerical differentiation-based concurrent learning in both low-noise and high-noise cases.Table I reports the comparison for the two noise conditions.
- Convergence behavior: The parameter estimates converge to a neighborhood of their true values, while the state and state-derivative estimation errors converge to neighborhoods of the origin.The figures also show noisy state evolution and the transient behavior motivating history-stack purging.
- Purging behavior: The history stack is purged faster initially, then its purging rate levels off approximately to a constant as the system reaches steady state.The minimum singular value increases under the thresholding algorithm in the reported simulation.
VII. CONCLUDING REMARKS
The paper develops a concurrent-learning parameter estimator using an adaptive observer for state-derivative estimates and validates it in noisy nonlinear-system simulations. The method performs better than numerical differentiation-based concurrent learning as measurement-noise variance increases, although the theoretical development does not account for measurement noise.
- The estimator uses an adaptive observer to generate the state-derivative estimates required for concurrent learning.
- Simulation validation uses a nonlinear system whose state measurements are corrupted by additive Gaussian noise.
- The developed technique yields better results than numerical differentiation-based concurrent learning, especially as additive-noise variance increases.
- The theoretical development does not account for measurement noise, and a provably noise-robust extension is left for future research.The paper notes that noise-rejecting observers and inherently noise-robust function approximators could address the two identified error sources.
- Measurement noise affects both the observer-generated state-derivative estimates and the history-stack data used for parameter estimation.The paper identifies separate noise effects in the derivative estimates and recorded state measurements.