Source-linked AI summary

Why not to use the Gaussian kernel

Toni Karvonen, Chris J. Oates

arXiv:2608.26974v1stat.MLcs.LGmath.NAmath.ST

TL;DR

The paper makes a mathematically rigorous case against using the Gaussian kernel as a default in Gaussian process regression. It shows that analyticity produces unrealistically small conditional variances and severe numerical ill-conditioning, motivating broader caution about analytic kernels.

  • Problem

    The paper addresses whether the widely used Gaussian kernel should serve as a default when predictive uncertainty and numerical stability matter.

  • Method

    The paper develops two mathematically rigorous results relating Gaussian-kernel analyticity to overconfident conditional variances and numerical ill-conditioning, and extends the concern to analytic kernels.

  • Results

    The Gaussian kernel can assign essentially zero conditional variance away from observed inputs and has a condition number that blows up super-exponentially with sample size.

  • Takeaways & Limitations

    The authors recommend avoiding the Gaussian kernel as a default and considering finitely smooth Matérn kernels, such as orders ν = 3/2 or ν = 5/2, without strong prior information.

  • Takeaways & Limitations

    The analysis focuses on Gaussian processes and does not claim that the Gaussian kernel never works; its risks depend on the sampling configuration and application setting.

Abstract

from arXiv · show

Kernels measure similarity or correlation in tasks such as regression and classification. The Gaussian kernel, other names of which include squared exponential and radial basis function kernel, is one of the most popular in Gaussian process regression. We argue that the Gaussian kernel is best avoided and should never be used as a default. The argument rests on two results demonstrating that the Gaussian kernel is extremely brittle. First, the Gaussian kernel gives rise to a conditional variance that is unrealistically small. If the variance is used to quantify predictive uncertainty, catastrophic overconfidence is almost inevitable. Second, a small variance goes hand in hand with numerical ill-conditioning, so that to use the Gaussian kernel in practice requires tricks such as nugget terms that effectively modify the underlying regression or classification model. These problems are caused by the unnatural smoothness of the Gaussian kernel, a fact we are far from the first to take notice of. The problem is not the Gaussian form itself but the analyticity of the kernel: Our argument is more broadly that analytic kernels are best avoided. For stationary kernels analyticity is essentially equivalent to an exponential decay of the spectral density.

1. Introduction

The paper argues that the widely used Gaussian kernel is brittle because its analyticity produces overconfident uncertainty and severe numerical ill-conditioning. It therefore recommends avoiding analytic kernels as defaults and considering less smooth alternatives such as low-order Matérn kernels.

  • The Gaussian kernel is also called the squared exponential, exponentiated quadratic, or radial basis function kernel, and is among the most popular kernels in Gaussian processes.
  • Its analyticity assumes that the modeled function is completely determined by its values in a neighborhood, an assumption the authors describe as rarely true.
  • When observations occupy only [1, 2]2, the Gaussian kernel assigns essentially zero conditional variance throughout [0, 1]2, unlike the Matérn kernels.This creates a risk of overconfident uncertainty quantification when predictions are required outside the sampled region.
  • The paper generalizes these failures to stationary analytic kernels whose spectral densities decay exponentially or faster.It identifies sub-exponentially decaying Gevrey kernels and finitely smooth Matérn kernels as alternatives, while noting that Gevrey kernels are rarely available in closed form.
  • The authors recommend a low-order Matérn kernel, such as ν = 3/2 or ν = 5/2, by default when there is no strong prior information.Inverse multiquadrics and Gevrey kernels are suggested when infinite differentiability is desired without using the Gaussian kernel.

2. Gaussian processes

Gaussian processes use positive-definite covariance kernels to define distributions over functions, while kernel smoothness is inherited by sampled functions. Regression conditions this process on noisy or noiseless observations to obtain conditional means and variances.

  • Gaussian processes: A Gaussian process is a stochastic process whose finite-dimensional distributions are multivariate Gaussians determined by its mean and covariance kernel.For any finite collection of inputs, the function values have a multivariate Gaussian distribution.
  • Kernels: Strictly positive-definite kernels produce invertible kernel matrices when the input points are pairwise distinct.The kernel matrix is therefore suitable for the conditioning formulas used in Gaussian process regression.
  • Kernels and priors: The covariance kernel determines how Gaussian-process samples fluctuate, and sample smoothness is inherited from kernel smoothness.Figure 5 contrasts non-differentiable, once-differentiable, and infinitely differentiable samples from different kernels.
  • Regression and interpolation: Gaussian process regression models an unknown latent function and conditions it on observations collected at pairwise-distinct input points.Observations may be noisy, with yi = f0(xi) + εi and εi independently distributed as N(0, σ^2).
  • Regression and interpolation: The conditional mean provides point estimates, while the conditional variance quantifies uncertainty through quantities such as pointwise credible intervals.When σ = 0, the conditional mean interpolates the data and the conditional variance is zero at every observed input.

3. Why not to use the Gaussian kernel

The Gaussian kernel’s analyticity produces two forms of brittleness: conditional variances can collapse globally despite local data, and covariance matrices become numerically ill-conditioned. These effects are especially severe in interpolation, motivating caution about analytic kernels and default kernel selection.

  • Overconfident uncertainty quantification: For higher-dimensional interpolation, Theorem 3.2 requires the input set to contain a Cartesian product of sets of m distinct real points.The proof uses a tensor-product argument, while related results show variance collapse whenever points are dense in an open subset, though without a convergence rate for d > 1.
  • Overconfident uncertainty quantification: The Gaussian kernel can make predictive uncertainty overconfident by driving conditional variance toward zero across the domain, even when inputs occupy only a small sub-domain.In one dimension this occurs for any sequence of sampling points; in higher dimensions, tensor-grid configurations suffice, with factorial-rate decay in the stated theorems.
  • Overconfident uncertainty quantification: Local information becomes effectively global under the Gaussian kernel, so predictions can receive little or no uncertainty in regions lacking observations.The paper notes that this behavior conflicts with the expectation that accurate prediction over a whole domain requires observations covering that domain.
  • Overconfident uncertainty quantification: Hyperparameter estimation cannot remove the Gaussian kernel’s global variance decay; changing the lengthscale or scaling parameter only preserves the problematic decay pattern.Choosing a lengthscale so variance decays more modestly in the sampled region forces the same rate everywhere on the domain.
  • Numerical ill-conditioning: The strongest failures occur in interpolation, whereas regression with σ > 0 is less dangerous because the noise term regularizes the covariance matrix.The paper nevertheless warns that regression does not eliminate the broader concerns associated with analytic kernels.
  • Numerical ill-conditioning: The Gaussian covariance matrix becomes numerically unusable as its condition number grows factorially or super-exponentially with the number of points.The cited threshold cond(Kn) ≈ 10^16 was exceeded with n = 8 points in double precision; nugget terms improve stability by effectively introducing observation noise.

4. Smoothness matters

For stationary kernels, spectral-density decay controls smoothness and conditional-variance behavior. Polynomial decay requires space-filling inputs for variance reduction, whereas exponential decay makes analytic kernels vulnerable to variance collapse away from observations.

  • The spectral density’s decay rate determines kernel smoothness: faster decay permits more differentiability, while Gaussian decay is faster than any polynomial.The Gaussian kernel is therefore infinitely differentiable.
  • Analytic functions: Analytic functions are determined by their values on any open connected set, explaining why Gaussian-process models with analytic kernels can learn globally from local observations.This follows from the identity theorem for real-analytic functions.
  • Matérn and finite smoothness: Polynomially decaying spectral densities prevent conditional variance from vanishing everywhere unless input points become space-filling.Theorem 4.1 applies when the spectral density has a polynomial lower bound and the nearest-input distance controls the variance.
  • Analytic kernels: Exponentially decaying spectral densities make conditional variance tend to zero whenever inputs cover some open set, even if predictions extend beyond observed points.Theorem 4.3 formalizes this behavior for stationary strictly positive-definite kernels.
  • Infinitely smooth non-analytic kernels: Gevrey kernels are infinitely differentiable but non-analytic, placing them between finitely smooth and analytic kernels.Their conditional variance behaves more like Matérn kernels: it requires space-filling inputs to vanish.

5. Apology of the Gaussian kernel

The section explains why the Gaussian kernel’s simplicity, smoothness, separability, and universality have sustained its popularity, while emphasizing that these properties do not justify default use.

  • Reasons for popularity: The Gaussian kernel remains popular because its familiar algebraic form avoids advanced special functions and simplifies implementation, differentiation, and integration.Its easy derivatives also support modelling with derivative data.
  • Reasons for popularity: The Gaussian kernel factorises into one-dimensional versions, making coordinate-specific lengthscales and automatic relevance determination straightforward.Under mild assumptions, no other isotropic kernels share this coordinate-wise product property.
  • Smoothness and analyticity: Its analytic smoothness produces simple series expansions, but this extreme regularity is also connected to the broader concern about analytic kernels.The Gaussian is infinitely differentiable, analytic with infinite radius of convergence, and the Matérn family approaches it as smoothness increases.
  • Universality: Universality is not a useful everyday criterion for choosing among common kernels because every commonly used kernel discussed here is universal.Universality means the associated reproducing kernel Hilbert space is dense in C(D) on compact domains.
  • Practical conclusion: The section’s practical conclusion is to reserve the Gaussian kernel for phenomena believed to be extremely smooth and consider other infinitely smooth kernels instead.The paper distinguishes regression from interpolation, where it describes the Gaussian as more dangerous.

6. What about kernel learning?

The section limits the paper’s direct claims to standard parametric kernels rather than flexible kernel learning, while identifying implications for deep kernel learning architectures.

  • Kernel-learning scope: Tuning only a few parameters of a standard parametric kernel family is often too inflexible for kernel learning.The paper identifies function or kernel composition and spectral-domain learning as two main flexible approaches.
  • Kernel-learning scope: The paper’s results do not directly apply to kernel learning, but they motivate research into architectures that account for analyticity.The discussion uses deep kernel learning as a concrete example of this research direction.
  • Deep kernel learning: Because compositions of analytic functions remain analytic, Gaussian base kernels and analytic activation functions may preserve the analyticity that motivates the paper’s warning.The paper gives the swish function as an example of an analytic activation function.

7. Proofs

The proofs establish variance identities and monotonicity, characterize Gaussian-kernel function spaces spectrally and analytically, and derive variance behavior from interpolation and approximation arguments.

  • RKHS characterization: For stationary kernels, the RKHS has a spectral characterization, and faster spectral decay yields a smaller RKHS.This connects kernel smoothness to the size of the associated function space.
  • Variance identities: The reproducing property rewrites conditional variance as a worst-case error, while adding data cannot increase the variance.The variance also increases with the noise parameter σ.
  • Analyticity: The Gaussian RKHS contains only real-analytic functions, as shown by bounding derivatives and applying Stirling’s approximation.The proof explicitly concludes that every function in the Gaussian RKHS is real analytic.
  • Variance bounds: Tensor-product structure reduces multivariate Gaussian-kernel variance calculations to coordinatewise terms on a hypercube.The proof uses Kronecker-product representations for the kernel matrix and evaluation vector.
  • Variance bounds: The proofs conclude that Gaussian-kernel conditional variance tends to zero as the number of observations grows, with uniform convergence on compact sets.The result follows from arbitrarily small bounds together with variance monotonicity and Dini’s theorem.
  • Variance bounds: Bump-function constructions support variance bounds based on the distance from a prediction point to the design, using spectral membership and RKHS norm estimates.The argument constructs a function that vanishes on observed points but remains nonzero at the prediction point.
Loading 2608.26974v1…