Source-linked AI summary

Application of the hierarchical bootstrap to multi-level data in neuroscience

Varun Saravanan, Gordon J Berman, Samuel J Sober

arXiv:2007.07797v2q-bio.NC

TL;DR

Neuroscience datasets often contain nested measurements, motivating methods that account for hierarchical structure rather than treating all observations as independent. The paper proposes the hierarchical bootstrap as an accessible, scalable approach that better balances intended Type-I error and statistical power than traditional or summarized methods, while noting assumptions about the data structure.

  • Problem

    Neuroscience studies commonly collect multiple samples within categories and hierarchical structures, creating a need for analyses that account for these relationships.

  • Method

    The paper proposes the hierarchical bootstrap as a powerful, easy-to-implement method that can scale to large and complicated datasets and be applied across hierarchical levels.

  • Results

    The bootstrap was not statistically different from the expected 5% false positive rate and performed better than summarized measures without sacrificing as much statistical power, especially at low sample sizes and small effect sizes.

  • Takeaways & Limitations

    The hierarchical bootstrap offers a practical way to analyze complex nested datasets while balancing intended Type-I error control and statistical power.

  • Takeaways & Limitations

    The discussion notes that the bootstrap assumes relationships underlying the latent variables that define the hierarchical structure, including settings with multiple trials from very few neurons.

Abstract

from arXiv · show

A common feature in many neuroscience datasets is the presence of hierarchical data structures, most commonly recording the activity of multiple neurons in multiple animals across multiple trials. Accordingly, the measurements constituting the dataset are not independent, even though the traditional statistical analyses often applied in such cases (e.g., Students t-test) treat them as such. The hierarchical bootstrap has been shown to be an effective tool to accurately analyze such data and while it has been used extensively in the statistical literature, its use is not widespread in neuroscience - despite the ubiquity of hierarchical datasets. In this paper, we illustrate the intuitiveness and utility of this approach to analyze hierarchically nested datasets. We use simulated neural data to show that traditional statistical tests can result in a false positive rate of over 45%, even if the Type-I error rate is set at 5%. While summarizing data across non-independent points (or lower levels) can potentially fix this problem, this approach greatly reduces the statistical power of the analysis. The hierarchical bootstrap, when applied sequentially over the levels of the hierarchical structure, keeps the Type-I error rate within the intended bound and retains more statistical power than summarizing methods. We conclude by demonstrating the effectiveness of the method in two real-world examples, first analyzing singing data in male Bengalese finches (Lonchura striata var. domestica) and second quantifying changes in behavior under optogenetic control in flies (Drosophila melanogaster).

1 | INTRODUCTION

Neuroscience datasets often contain nested measurements that violate independence assumptions, making common tests underestimate uncertainty and p-values. The paper motivates hierarchical bootstrap as a structure-agnostic approach and evaluates it against traditional, summarized, and mixed-model methods.

  • Motivation: Student’s t-test and ANOVA treat all data points as independent, underestimating uncertainty and inflating apparent significance through pseudoreplication.The resulting bias includes underestimated p-values and inappropriate handling of within- and between-group variance.
  • Motivation: Nested measurements from spines, neurons, and animals are statistically dependent because samples within the same lower-level unit tend to be more similar.This dependence creates hierarchical datasets rather than collections of independent observations.
  • Related methods: Linear Mixed Models can account for hierarchical variance but may be unreliable with few clusters and sensitive to modeling choices.The paper also notes that LMMs assume a linear hierarchical structure and can produce biased or unstable fits.
  • Hierarchical bootstrap: The hierarchical bootstrap is relatively agnostic to underlying structure and has performed better than traditional statistics at quantifying uncertainty and identifying signal.The authors note concerns that bootstrap estimates can be excessively conservative in a limited subset of cases.
  • Paper scope: The paper uses simulations to show inflated Type-I error, then applies hierarchical bootstrap to singing data in finches and optogenetic behavior data in flies.These examples illustrate the method on strongly hierarchical neuroscience datasets.
  • Paper scope: The real-world analyses emphasize using statistical tests appropriate to hierarchical datasets in neuroscience.The conclusion links the examples to the broader need for appropriate analysis of nested data.

2 | MATERIALS AND METHODS

The paper compares traditional, summarized, bootstrap, and mixed-model analyses for nested neuroscience data. It develops sequential resampling across hierarchy levels, uses bootstrap distributions for uncertainty and hypothesis support, and addresses the trade-off between error control and power.

  • Methods overview: The study compares four methods: Traditional, Summarized, Bootstrap, and LMMs for hierarchical neuroscience data.The methods differ in how they handle dependence among neurons, subjects, and experimental conditions.
  • Traditional statistics: Traditional analysis treats every neural measurement as independent and applies a Student’s t-test to group means.It computes the mean and SEM across all data points without accounting for hierarchical structure.
  • Summarized statistics: Summarized analysis first averages measurements within each subject, then compares subject-level means between groups with a t-test.This approach acknowledges dependence among neurons within a subject but can reduce sample size and power when subjects are few.
  • Bootstrap inference: A 67% confidence interval, equivalently the standard deviation, of bootstrap means estimates uncertainty in the group mean.The bootstrap can also quantify support for hypotheses by evaluating proportions of resampled means beyond specified thresholds.
  • Design effect: Traditional analyses can underestimate standard errors and inflate Type-I error as within-cluster sampling increases.The design-effect discussion connects clustering and within-cluster sample size to the required correction of standard errors.

3 | RESULTS

Simulations and two real-world reanalyses show that hierarchical bootstrap controls false positives while retaining power and better matching observed biological signals than traditional or summarized methods.

  • Simulations: 46% to almost 96%: the traditional method’s false positive rate increased with trials per neuron despite identical group means.The summarized, bootstrap, and LMM methods remained near the expected 5% rate.
  • Simulations: 0.66 ± 0.08%: the bootstrap produced a conservative false positive rate when all simulated data points were independent.Its error bars were roughly 1.4 times larger than those of the other methods, partially accounting for the reduced significance rate.
  • Power and error control: Around or above 80%: the traditional method’s false positive rate remained high for hierarchical data with group differences, whereas summarized statistics stayed near 5%.At very small subject numbers, LMM and bootstrap false positive rates were also elevated, with LMM reaching 16 ± 1% at N=1.
  • Power and error control: The bootstrap and LMM retained more statistical power than summarized methods while remaining sensitive to the Type-I error rate, with bootstrap marginally better.Traditional statistics had the highest power but an unacceptably high false positive rate; summarized statistics had the lowest power.
  • Real-world reanalyses: In songbird data, hierarchical bootstrap supported a target-syllable shift but not anti-adaptive generalization across different syllables.The different-type effect was too small to detect with the original sample size and appeared driven largely by a small number of individual birds.
  • Real-world reanalyses: In fly behavior maps, hierarchical bootstrap identified a concise head-grooming region, unlike traditional statistics’ extra false-positive region and summarized statistics’ null result.LMM was more conservative and likely produced many false negatives, although its false positive rate was low.

4 | DISCUSSION

The discussion presents hierarchical bootstrap as a practical alternative for neuroscience data, aiming to control false positives without sacrificing as much power as summarized analyses. Simulations and two real-world examples show its statistical advantages and implementation trade-offs.

  • Motivation: The hierarchical bootstrap addresses shortcomings of common statistical analyses for hierarchically structured neuroscience datasets.The paper motivates the method through dependence among observations and concerns about inappropriate statistical tests.
  • Simulations: The bootstrap kept false-positive rates within the intended bound while retaining more power than summarized statistics.This advantage was observed especially at low sample sizes and small effect sizes.
  • Methodological comparison: Unlike LMMs, the bootstrap does not require assumptions about relationships among latent variables defining the hierarchy.However, resampling may need adjustment when sampling many trials from very few neurons.
  • Real-world examples: In the songbird example, bootstrap analysis rejected statistical significance for anti-adaptive generalization that had appeared significant with summarized statistics.The authors suggest the apparent effect was driven largely by a subset of birds rather than a population-wide effect.
  • Real-world examples: In the fly optogenetics example, the hierarchical bootstrap isolated the true signal more effectively than traditional, summarized, and LMM analyses.Traditional analysis included likely false positives, summarized analysis found no significant areas, and LMM was more conservative than the bootstrap.
  • Implications: The authors argue that widespread bootstrap use could reduce false positives and improve statistical-test selection across neuroscience data.They also emphasize scalability, direct hypothesis-support probabilities, and implementation checks through shared analysis code.
Loading 2007.07797v2…