Source-linked AI summary

Functional linear regression from sparse to dense designs: a pooling-ridge method and minimax optimality

Shunxing Yan, Fang Yao

arXiv:2608.25468v1stat.MEcs.LGmath.STstat.ML

TL;DR

The paper tackles the unresolved problem of minimax-optimal functional linear regression from discretely and noisily observed functions. It introduces pooling-ridge estimation, which combines pooled operator estimation with RKHS methods, and proves optimal prediction rates across sparse-to-dense designs for scalar-on-function and function-on-function regression. The resulting rates reveal one phase transition in the scalar-on-function case and up to three in the function-on-function case.

  • Problem

    Minimax-optimal linear regression with discretely observed functional data remains unresolved, although mean and covariance estimation under discrete sampling have been studied extensively.

  • Method

    Pooling-ridge estimation pools discretely observed measurements across subjects to obtain unbiased RKHS operator estimators for scalar-on-function and function-on-function regression.

  • Results

    Matching upper and lower bounds establish minimax-optimal prediction rates, with one phase transition for scalar-on-function regression and up to three for function-on-function regression.

  • Takeaways & Limitations

    The theory characterizes how discrete sampling frequencies affect prediction rates across sparse, dense, predictor-sparse, and response-sparse regimes.

  • Takeaways & Limitations

    The function-on-function analysis leaves unclear whether the proposed estimator achieves minimax optimality in L2 error.

Abstract

from arXiv · show

Functional data analysis is an important statistical field that treats data as random functions. In practice, the random functions are often not fully observed but instead measured at discrete times. While simpler problems, such as mean and covariance estimation, have been widely studied for discretely observed data, optimal estimation of linear regression for this data type has remained unsolved for over two decades. To tackle this fundamental challenge, we propose a novel approach, referred to as pooling ridge estimation, which combines the advantages of pooling strategy and RKHS-based method by incorporating the unbiased estimation of operators based on discretely observed measurements from all subjects. This unified estimation framework enables us to achieve minimax optimality in prediction risk in arbitrary sampling schemes ranging from sparse to dense designs, for both scalar-on-function and function-on-function regression models. Such methodological and theoretical advances are obtained for the first time and accurately reveal the influence of discrete sampling. For scalar-on-function regression, the phase transition occurs once, separating the convergence behavior into two distinct regimes. Remarkably, for function-on-function regression, up to three phase transitions may occur, determined by the sampling frequencies of the predictor/response functions. Finally, simulation experiments and two real data examples provide empirical support for the proposed methods.

1. Introduction.

The paper addresses minimax-optimal functional linear regression when functions are discretely and noisily observed, proposing a unified pooling-ridge framework for sparse-to-dense designs. It characterizes how sampling affects prediction rates and phase transitions in scalar-on-function and function-on-function regression.

  • 1. Introduction.: Optimal minimax estimation for linear regression with discretely observed functional data has remained unresolved, despite progress for mean and covariance estimation.The difficulty combines an infinite-dimensional inverse problem with discrete and noisy measurements.
  • 1. Introduction.: Existing RKHS-based regression methods generally assume fully observed trajectories, while pre-smoothing reconstructs each curve individually and may be unsuitable for sparse observations.The paper motivates pooling observations across subjects as a more realistic alternative.
  • 1. Introduction.: The pooling-ridge estimator pools discrete measurements from all subjects to construct unbiased RKHS operator estimators for both scalar-on-function and function-on-function regression.It combines the pooling strategy used for sparse functional data with RKHS-based estimation.
  • 1. Introduction.: For scalar-on-function regression, sampling creates one phase transition: dense designs attain the full-observation rate, whereas sparse designs have a rate depending on the total number of observation points nmx.The transition occurs at mx ≍n2rc/(2rb+2rc+1), with dense rate n−(2rb+2rc)/(2rb+2rc+1) and sparse rate (nmx)−(2rb+2rc)/(2rb+4rc+1).
  • 1. Introduction.: For function-on-function regression, discretization of predictor and response functions can generate up to three convergence phases, including dense, predictor-sparse, and response-sparse regimes.The phase structure depends on rbx + rc relative to rby and rby(1 + 2rc), with sampling frequencies mx and my determining regime boundaries.
  • 1. Introduction.: The proposed estimators achieve minimax optimal prediction rates through matching upper and information lower bounds, while numerical studies support their advantages.The theoretical results are stated for both scalar-on-function and function-on-function models.

2. Model and methodology.

The paper formulates scalar-on-function and function-on-function regression for noisy, discretely observed functional data and develops pooling-ridge estimators for sparse-to-dense sampling. The method combines pooled operator estimation with RKHS regularization to avoid sparse-design bias while covering both regression settings.

  • Regression models: Scalar-on-function regression models a scalar response as an intercept plus the integral of the predictor against an unknown slope function and random noise.
  • Observation scheme: Discrete observations are noisy measurements collected at subject-specific time points, with within-subject dependence and independence across subjects.
  • Observation scheme: Pre-smoothing each curve before regression creates unignorable bias when measurements are sparse, making fully observed-data procedures unsuitable in that regime.
  • Regression models: Function-on-function regression extends the model to a functional response with an intercept function and a bivariate slope function.
  • Pooling-ridge estimation: Pooling discrete measurements across subjects yields unbiased estimators of the operators used in RKHS formulations, while measurement error affects convergence bounds only through multiplicative constants.
  • Pooling-ridge estimation: The resulting estimator applies across arbitrary sampling frequencies and remains representable in the span of kernel evaluations, even when the true slope lies outside the RKHS.

3. Theoretical results: minimax optimality and phase transition.

Theoretical upper and lower bounds match for both scalar-on-function and function-on-function regression with discrete observations, establishing minimax-optimal rates across sparse-to-dense designs. Sampling frequencies create one phase transition for scalar responses and up to three regimes for functional responses.

  • Minimax optimality: Matching upper and lower bounds establish minimax-optimal convergence rates for discretely observed scalar-on-function and function-on-function regression.The results cover arbitrary sampling schemes ranging from sparse to dense designs.
  • Scalar-on-function regression: For scalar-on-function regression, the rate has one phase transition at mx ≍ n^(2s/(2(1−α)r+1)), after which more measurements do not improve convergence.Sparse sampling yields (nmx)^−2(1−α)r/(2(1−α)r+2s+1), while denser sampling reaches n^−2(1−α)r/(2(1−α)r+1).
  • Function-on-function regression: For function-on-function regression, the minimax rate includes separate terms for sample size, predictor sampling, and response sampling.The second and third terms reveal the effects of discrete predictor and response observations, respectively.
  • Function-on-function regression: When ry ≥ rx, only one phase transition occurs, with dense sampling reaching n^−2(1−α)rx/(2(1−α)rx+1) and sparse sampling governed by mx.The response-discretization term is dominated by the first term in this case.
  • Function-on-function regression: When ry < rx, up to three regimes arise: dense, predictor-sparse, and response-sparse dominance, depending on mx, my, and which sparsity dominates.In the dense regime, further increases in predictor and response sampling do not improve the sample-size-dependent rate.
  • Special cases: In special cases, finite sampling frequencies give rate n^−2(1−α)r̃/(2(1−α)r̃+1), where r̃ = min{rx/(2s+1), ry}, while α = 0 corresponds to the well-specified model.For sufficiently large my, the convergence rate becomes independent of the marginal complexity of Ky.

4. Simulation.

The simulations evaluate pooling-ridge estimation across basis designs, sample sizes, sampling frequencies, and competing methods for scalar-on-function and function-on-function regression. Results show particularly strong performance under sparse sampling, with convergence toward fully observed RKHS performance as sampling becomes denser.

  • Parameter tuning: A two-step cross-validation strategy combines the stability of biased empirical criteria with the unbiasedness of the proposed risk estimate.Candidates are first filtered using a thresholded biased criterion, then selected by minimizing the unbiased criterion.
  • Simulation setup: The simulations use cosine and Legendre predictor designs, Gaussian measurement noise, Matérn 3/2 kernels, and sample sizes n = 50,100,200.The Legendre design tests robustness to misalignment between the predictor basis and the slope function.
  • Compared methods: The study compares pooling-ridge estimation with plug-in, FPCA-based approximated least-squares, and fully observed RKHS methods.FPCA scores are estimated using integral approximation, PACE, and another competing technique.
  • Scalar-on-function results: Pooling-ridge estimation achieves the lowest average errors across settings, especially when the sampling frequency is small.Its standard deviations remain relatively small, indicating stable performance in the reported simulations.
  • Scalar-on-function results: As sampling frequency increases, the difference between pooling-ridge estimation and fully observed RKHS estimation becomes smaller.This pattern is consistent with the theoretical expectation that denser sampling reduces the effect of discrete observation.

5. Real data.

The real-data examples evaluate the proposed methodology in scalar-on-function and function-on-function regression under dense, sparsified, and longitudinal sampling. Across both applications, the proposed method performs favorably against alternative approaches.

  • 5.1. The wheat dataset.: The wheat dataset contains NIR spectra at 701 wavelengths and protein concentrations for 100 individuals, supporting scalar-on-function regression.The spectra range from 1100 nm to 2500 nm with 2 nm spacing.
  • 5.1. The wheat dataset.: Random sparsification selects m = 5,10,20 training measurements per subject, with experiments repeated 100 times across 80-subject training and 20-subject testing splits.The procedure evaluates performance under multiple sampling frequencies despite dense original observations.
  • 5.1. The wheat dataset.: Across sampling frequencies m_x, the proposed approach achieves the lowest empirical mean error and significantly improves over five alternative methods on the wheat data.FullyRKHS also performs relatively well in this sparse setting, possibly because the spectral patterns are smooth.
  • 5.2. The CONTENT dataset.: The CONTENT study predicts BMI Z-scores from days 1–150 to days 151–300 using irregular, sparse measurements in a function-on-function regression model.Predictor observations range from 5 to 17 per subject, with median 14; response observations range from 2 to 11, with median 7.
  • 5.2. The CONTENT dataset.: The CONTENT comparison evaluates OPFFR-S, RKHS-PCA, and approximated least-square baselines using multiple score estimators after centering by the estimated mean function.The full data are randomly divided into training and test sets.
  • 5.2. The CONTENT dataset.: The proposed pooling-operator method consistently outperforms and is more stable than the alternative methods across the reported evaluation metrics.The comparison includes curve-recovery, FPCA-based, and RKHS-PCA approaches.

SUPPLEMENTARY MATERIAL

The supplementary material provides proofs for the main theoretical results and technical lemmas, along with additional simulation results.

  • SUPPLEMENTARY MATERIAL: The supplement contains proofs of the main theoretical results and technical lemmas, together with additional simulation results.
Loading 2608.25468v1…