Source-linked AI summary

A Tutorial on Fisher Information

Alexander Ly, Maarten Marsman, Josine Verhagen, Raoul Grasman, Eric-Jan Wagenmakers

arXiv:1705.01064v2math.ST

TL;DR

Mathematical psychologists need an accessible account of how Fisher information operates across major statistical paradigms. The tutorial explains its use in frequentist inference, Bayesian prior specification, and MDL model-complexity measurement, while noting parameterization dependence and the tutorial’s limited scope.

  • Problem

    An accessible introduction to Fisher information for mathematical psychologists is lacking despite its relevance across frequentist, Bayesian, and minimum description length analyses.

  • Method

    The tutorial defines and calculates Fisher information, then illustrates its applications to frequentist inference, Bayesian default priors, and minimum description length model selection.

  • Results

    Fisher information is used to construct frequentist hypothesis tests and confidence intervals, specify a default parameterization-invariant Bayesian prior, and measure model complexity in MDL.

  • Takeaways & Limitations

    The tutorial provides concrete examples intended to give psychologists more insight when proposing, developing, and comparing mathematical models of psychological processes.

  • Takeaways & Limitations

    The account is not comprehensive, focuses on one-dimensional parameters, and notes that generalizations to vector-valued parameters are provided in the appendix.

Abstract

from arXiv · show

In many statistical applications that concern mathematical psychologists, the concept of Fisher information plays an important role. In this tutorial we clarify the concept of Fisher information as it manifests itself across three different statistical paradigms. First, in the frequentist paradigm, Fisher information is used to construct hypothesis tests and confidence intervals using maximum likelihood estimators; second, in the Bayesian paradigm, Fisher information is used to define a default prior; lastly, in the minimum description length paradigm, Fisher information is used to measure model complexity.

1. Introduction

Mathematical psychologists use quantitative models across frequentist, Bayesian, and minimum description length paradigms, which are connected by Fisher information. This tutorial introduces Fisher information and illustrates its uses for these different purposes.

  • Mathematical psychologists use quantitative models to describe behavior and understand latent psychological processes.
  • Frequentist, Bayesian, and minimum description length analyses address different questions but are connected through Fisher information.
  • The tutorial fills an accessible-introduction gap by explaining Fisher information for graduate students and researchers interested in cognitive modeling and mathematical statistics.
  • The paper develops notation and key concepts before defining Fisher information and demonstrating its calculation.
  • For iid binary data, the number of heads summarizes the trials because the conditional distribution of the raw sequence given that count does not depend on θ.

2. The Role of Fisher Information in Frequentist Statistics

In frequentist statistics, Fisher information characterizes the large-sample behavior of maximum likelihood estimators and supports experiment design, hypothesis tests, and confidence intervals. The resulting normal approximations provide tractable inference when exact sampling distributions are difficult to compute.

  • 2.1. Maximum likelihood estimation: The MLE selects the parameter value under which the observed data are most likely; for 7 heads in 10 trials, it is θ̂obs = 0.7.The observation has 26.7% likelihood under θ = 0.7 versus 11.7% under θ = 0.5.
  • 2.2. Sampling distributions: For iid data and sufficiently large n, the MLE error is approximately normally distributed around the true parameter value.The approximation applies under general conditions and characterizes the sampling distribution using Fisher information.
  • 2.1. Using Fisher information to design an experiment: The MLE standard error is 1/√(nI_X(θ∗)), so it decreases when sample size or unit Fisher information increases.Researchers can increase n to obtain more information because the true parameter value is not under experimental control.
  • 2.1. Using Fisher information to design an experiment: For a Bernoulli experiment, requiring an MLE distance of 0.1 from the true value with 68% chance is solved by n = 25 in the least informative case θ = 1/2.
  • 2.2. Null hypothesis tests and confidence intervals: The normal approximation supports null hypothesis tests and confidence intervals based on the Fisher-information-derived sampling distribution.A 95% confidence procedure is expected to contain the true value in 95 of 100 simulated intervals.
  • 2.2. Null hypothesis tests and confidence intervals: Asymptotic normality is especially useful when an exact MLE sampling distribution is unavailable or difficult to compute.For the Laplace model, the exact distribution is unwieldy, whereas the MLE is the sample median and unit Fisher information is b^-2.

3. The Role of Fisher Information in Bayesian Statistics

In Bayesian statistics, Fisher information motivates Jeffreys’s prior as a parameterization-invariant alternative to uniform parameter priors. The tutorial shows geometrically why uniform priors can produce different posteriors, whereas Jeffreys’s prior corresponds to uniformity in model space.

  • Failure of the uniform distribution on the parameter as a noninformative prior: Uniform priors can yield different posteriors and conclusions for equivalent models when assigned to different parameterizations.The tutorial contrasts uniform priors on the propensity θ and angle φ, whose induced posteriors on θ differ notably.
  • A default prior by Jeffreys’s rule: Jeffreys’s prior is parameterization-invariant and therefore produces the same posterior knowledge whether constructed on θ or on φ.The corresponding posterior probability for Jθ = (0.6,0.8) is 0.53 under either updating route.
  • Uniform on the parameter space versus uniform on the model space: The model-space geometry explains why parameter-space uniformity is not equivalent to uniformity over probability mass functions.The Bernoulli model can be represented as a line or the positive part of a circle, and equal parameter increments correspond to unequal model-space distances.
  • Uniform on the parameter space versus uniform on the model space: Neither a uniform prior on θ nor one on φ yields an equiprobable assessment of probability mass functions in model space.A uniform prior on φ favors pmfs with φ close to zero, while the uniform prior on θ also fails to partition model space uniformly.
  • Uniform prior on the model: Jeffreys’s prior is uniform on model space because Fisher information compensates for geometric differences between parameterizations.With yobs = 7 out of n = 10, the interval Jm accounts for 14% of model length before updating and 53% posterior probability afterward.

4. The Role of Fisher Information in Minimum Description Length

Within MDL, Fisher information contributes a geometrical complexity penalty to model description length, balancing goodness-of-fit against model volume. In the MPT example, FIA selects the model closest to the empirical pmf while favoring the simpler model when fit is comparable, with this complexity preference diminishing as sample size grows.

  • FIA and model complexity: All three criteria select the model with the lowest criterion value, but FIA additionally accounts for model volume beyond dimensionality.AIC and BIC penalize complexity through dimensionality terms, whereas FIA also incorporates the models’ volumes.
  • FIA and model complexity: FIA approximates model description length by combining goodness-of-fit with a penalty based on Fisher-information-derived model complexity.The complexity term incorporates model volume, while the fit term evaluates the observed data under the model’s MLE.
  • MPT model comparison: The two MPT models are one-dimensional, so their dimensionality penalties are equal and their comparison isolates Fisher information’s role in MDL selection.The models are neither saturated nor nested and overlap only at one pmf.
  • Generalization and overfitting: Model-dependent generalization determines which observations are treated as structural versus noise, while overfitting is controlled by complexity penalties.Under M2, an empirical pmf on the model produces no noise in the example; AIC, BIC, and FIA penalize complexity to guard against poor generalizability.
  • MPT model comparison: The individual-word strategy has the larger model volume and incurs an additional penalty of 1/2 log(2) relative to the only-mixed strategy.Its larger volume represents greater capacity to fit data, which FIA penalizes as greater complexity.
  • Finite-sample selection: For n = 30, FIA generally selects the model closest to the empirical pmf, but near the models’ overlap it favors the simpler only-mixed strategy M2.Empirical pmfs between the non-decision curves select M2; observations near the cross are assigned to M2 despite comparable fit.
  • Large-sample behavior: As n increases, FIA’s preference for the simpler model decreases, and for extremely large samples selection is governed by goodness-of-fit proximity.In the large-sample limit, the additional penalty for M1’s greater volume becomes irrelevant, yielding the same selection as BIC in this comparison.

5. Concluding Comments

The tutorial illustrates Fisher information across frequentist, Bayesian, and information-theoretic settings. It presents concrete examples while acknowledging that the account is introductory rather than comprehensive.

  • Fisher information supports hypothesis tests and confidence intervals in frequentist statistics, a default parameterization-invariant prior in Bayesian statistics, and model-complexity measurement in information theory.
  • The Bayesian treatment fixes the first argument at observed data, producing a likelihood function of the parameters.
  • The information-geometric treatment lets both arguments vary in f(⋅∣⋅), with observed data and the maximum likelihood plugged in.
  • The tutorial considers one-dimensional parameters in the exposition, points readers to the appendix for vector-valued parameters, and is not comprehensive.

Appendix A: Generalization to Vector-Valued Parameters: The Fisher Information Matrix

For vector-valued parameters, Fisher information is represented by a symmetric positive semidefinite matrix whose entries are expectations of score-function products or, under regularity conditions, negative expected second derivatives.

  • For a d-dimensional parameter vector, Fisher information is a d×d symmetric positive semidefinite matrix.
  • Each matrix entry is formed from partial derivatives of the log probability mass function, with the derivatives evaluated at the same parameter vector used for weighting.
  • Under mild regularity conditions, matrix entries can equivalently be calculated as negative expectations of second-order partial derivatives.
  • The expectation or sum is taken over the outcomes x of X.
  • Off-diagonal entries need not vanish; parameters are called orthogonal when the corresponding entry is zero.
  • For iid trials, Fisher information is additive: I_Xn(θ) = nI_X(θ), including for vector-valued parameters.

B.1. Asymptotic normality of the MLE for vector-valued parameters

For regular models, the vector-valued MLE is asymptotically multivariate normal, with covariance governed by the inverse Fisher information matrix. At finite sample sizes, replacing the true sampling distribution by this approximation can be inaccurate.

  • For regular parametric models, the MLE for vector-valued parameters converges in distribution to a multivariate normal distribution.
  • The asymptotic approximation uses the inverse Fisher information matrix evaluated at the true parameter value.
  • The normal approximation is used at fixed n, and its approximation error is negligible only when n is sufficiently large.
  • Because the true data-generating probability mass function is unknown, the adequacy of the approximation depends on the underlying distribution.
  • A 5% test based on the asymptotic normal distribution could have a 42% type 1 error rate, while a nominal 95% interval could cover the true parameter only 20% of the time.

B.2. Asymptotic normality of the MLE and the central limit theorem

The CLT gives broad normal approximations for sample means, whereas MLE asymptotic normality uses the model’s functional form and can yield more efficient estimators. The examples show exactness for Gaussian data and differing finite-sample requirements for Laplace and Cauchy data.

  • CLT versus MLE asymptotic normality: The CLT applies under finite population mean and variance without requiring a particular distributional form, but its finite-sample accuracy depends on n being large enough.
  • CLT versus MLE asymptotic normality: Computing the MLE requires knowledge of the functional relationship linking parameters to outcomes, enabling stronger statements when that model is correct.
  • Gaussian distribution: For Gaussian observations with known variance, the MLE is the sample mean and the normal approximation holds exactly for every finite n.
  • Gaussian distribution: When σ^2 = 1 and n = 100, P(|θ̂ − θ*| ≤ 0.196) = 0.95 holds exactly for the Gaussian case.
  • Cauchy distribution: For Cauchy data, the CLT cannot be used, and the sample mean’s Cauchy sampling distribution does not improve as n increases.
  • Cauchy distribution: For the Cauchy case, achieving the stated approximate 95% precision requires n = 25π^2 ≈247 for the sample median and n = 200 for the MLE.

B.3. Efficiency of the MLE: The H´ajek-LeCam convolution theorem and the Cram´er-Fr´echet-Rao information lower bound

For regular estimators, the H´ajek-LeCam theorem bounds asymptotic variance below by inverse Fisher information, and the MLE attains this bound. The Cramér-Fréchet-Rao bound extends Fisher-information efficiency to unbiased estimators at finite sample sizes, with important restrictions.

  • H´ajek-LeCam convolution theorem: The H´ajek-LeCam convolution theorem decomposes a regular estimator’s limiting distribution into independent components, so its asymptotic variance is bounded below by inverse Fisher information.The bound follows because the additional component has nonnegative variance.
  • H´ajek-LeCam convolution theorem: The MLE attains the lower bound because its additional convolution component has zero variance, making it best among regular estimators.This result applies when the sample size is large enough and the model’s functional relationship generates the data.
  • Practical limitations: Normal approximations used for confidence intervals and hypothesis tests can perform poorly when n is small, while model misspecification can undermine MLE efficiency.The paper identifies both issues as important concerns for hard inferential decisions.
  • Cramér-Fréchet-Rao information lower bound: The Cramér-Fréchet-Rao information lower bound states that an unbiased estimator’s variance cannot be below inverse Fisher information, and it also holds for finite n.Thus, Fisher information is useful beyond large-sample approximations.
  • Cramér-Fréchet-Rao information lower bound: The finite-sample bound is limited because unbiased estimators are restrictive, generally exclude the MLE, and cannot be attained for parameters with more than one dimension.For vector-valued parameters, the bound therefore does not determine whether a better estimator exists.

Appendix C: Bayesian use of the Fisher-Rao Metric: The Jeffreys’s Prior

The tutorial interprets Jeffreys’s prior geometrically as a uniform prior over model space. Fisher information supplies the metric that supports this interpretation and the associated parameterization-invariant construction.

  • Geometric interpretation: Jeffreys’s prior is presented as a uniform prior on the model MΘ rather than merely a prior uniform in parameter space.The tutorial develops this interpretation using geometric model-space arguments.
  • Geometric interpretation: The geometric construction uses Fisher information to measure displacement in model space and connect model-space intervals with parameter-space intervals.The explanation reduces to arc-length computation by integration by substitution.

C.1. Tangent vectors

Tangent vectors approximate how parameter changes displace probability mass functions in model space. Taylor expansions provide the component displacements, whose combined length is linked to Fisher information.

  • Parameterization: A model-space interval of pmfs is mapped to a parameter-space interval through a parameter functional that uniquely assigns parameter values.Distinct pmfs correspond to distinct parameter values under this assignment.
  • Tangent vectors: In Fig 12, the full tangent arrow combines horizontal and vertical component arrows associated with outcomes x = 0 and x = 1.The caption places the expansions at θa = 0.8 in the left panel and φa = 0.6π in the right panel.
  • Taylor approximation: Tangent vectors approximate the displacement between nearby pmfs by replacing the parameterization with its first-order Taylor expansion.The approximation treats the displacement as a straight line when the parameter distance is small.
  • Tangent vectors: For the φ example, changing φ by 0.2π at φa = 0.6π produces horizontal and vertical model-space displacements of 0.17 and 0.14.The combined displacement is represented by the full arrow in the right panel.
  • Tangent-vector length: The length of a tangent vector is computed as the root of the sum of squared component displacements and can equivalently be obtained from the square root of Fisher information.More generally, the square root of Fisher information converts a small parameter distance into model-space displacement.

C.2. The Fisher-Rao metric

The Fisher-Rao metric measures distance in model space using infinitesimal pmf displacements. Integrating this metric yields the length of model-space intervals and defines their total volume.

  • Metric construction: The parameter functional maps a model-space interval of pmfs to a parameter-space interval, enabling integration of model-space length through Fisher information.The resulting metric is based on the displacement dmθ(X).
  • Metric construction: Choosing the entire model MΘ as the integration domain yields the normalizing constant V, interpreted as the total length of the model.This interpretation follows from using dmθ(X) as the distance measure in model space.
  • Terminology: The distance measure based on Fisher information is also known as the Fisher-Rao metric.The name recognizes Calyampudi Radhakrishna Rao’s contribution to the theory.

C.3. Fisher-Rao metric for vector-valued parameters

A categorical model with three outcomes can be parameterized in multiple ways, and these parameterizations describe the same model space. The β and γ coordinates are related through distinct parameter spaces and inverse mappings.

  • Three-outcome probability mass functions are parameterized by β=(β1,β2), with probabilities [β1,β2,1−β1−β2].
  • The β parameter space is constrained by β1,β2≥0 and β1+β2≤1, whereas the γ parameter space is [0,1]×[0,1].
  • The γ parameterization uses a stick-breaking construction, with γ1=pL and γ2=pM/(1−pL).
  • The β and γ parameterizations are isomorphic because both map onto the complete model of probability mass functions.

C.3.3. Multidimensional Jeffreys’s prior via the Fisher information matrix and orthogonal parameters

The Fisher information matrix characterizes orthogonality between parameters and supports comparing parameterizations geometrically. In the categorical example, γ is preferred because its parameters are orthogonal and its volume calculation decouples.

  • The multidimensional Jeffreys prior is invariant to parameterization, with its normalization constant obtained from the model volume using the Fisher information.
  • The complete model is easier described by γ because its parameters are orthogonal, meaning the corresponding off-diagonal Fisher-information entries are zero.
  • At [1/3,1/3,1/3], γ produces orthogonal tangent vectors and a rectangle, whereas β produces a diamond formed by its tangent vectors.
  • Nonzero off-diagonal Fisher-information terms couple the β parameters, while γ orthogonality allows the parameters to be treated independently and the double integral to decouple.
  • Orthogonality is relevant in Bayesian analysis because it supports choosing a prior on a vector-valued parameter that factorizes.
  • The geometric visualization uses three outcomes; the same ideas extend to more general random variables under the stated regularity conditions.

D.1.3. Entropy, cross entropy, log-loss

Entropy is the average code length under the true coding system, while cross entropy measures coding with a postulated model. Log-loss replaces the unknown population quantity with a sample-based average for inference.

  • The true coding system yields an average code length of 1.5 bits per trial when p∗(X)=[0.25,0.5,0.25].
  • Cross entropy is the population average code length when data generated by p∗(X) are encoded with the postulated model f(X∣β).
  • 2.97 bits per trial results when f(X∣β)=[0.01,0.18,0.81] encodes data generated by p∗(X)=[0.25,0.5,0.25].
  • Shannon–Fano coding assigns outcome x an ideal code length of −log2 p∗(x) bits, subject to rounding in actual prefix codes.
  • Cross entropy cannot be smaller than entropy because it decomposes into entropy plus a nonnegative Kullback–Leibler divergence.
  • Because the true pmf is unknown, the population average is approximated by replacing it with the empirical pmf and using a sample average.

D.2. Data compression and statistical inference

Data compression connects statistical inference with minimizing coding length: the MLE selects the model minimizing sample log-loss. This connection is developed for parameterized models under regularity assumptions and remains bounded by model specification choices.

  • Finding the shortest average code is equivalent to finding the true data-generating process within the candidate model space.
  • Minimizing negative log-likelihood is equivalent to maximizing likelihood, so log-loss is minimized by the coding system associated with the MLE.
  • The cross-entropy decomposition makes log-loss minimization equivalent to minimizing KL divergence, whose lower bound is zero.
  • A positive KL divergence indicates that the empirical pmf associated with the observations does not reside on the model.
  • Normal-model assumptions can be stringent because symmetry around the mean and limited accommodation of outliers increase the risk of misspecification.
  • Regular parametric models require an open parameter domain, differentiability, nonsingular Fisher information, and continuity of the score mapping.
Loading 1705.01064v2…