Source-linked AI summary
Power contours: optimising sample size and precision in experimental psychology and human neuroscience
Daniel H. Baker, Greta Vilidaite, Freya A. Lygo, Anika K. Smith, Tessa R. Flack, Andre D. Gouws, Timothy J. Andrews
TL;DR
Experimental power depends not only on how many participants are tested but also on how precisely each participant is measured through repeated trials. This paper develops joint power contours, applies them across eight psychology and neuroscience paradigms, and finds that trial number can materially affect power, while noting limitations from non-independent trials and speculative prospective variances.
Problem
The paper addresses how participant sample size and repeated-trial number jointly determine statistical power.
Method
It develops two-dimensional power contours, applies subsampling to eight real datasets, and provides an online tool and variance-estimation code.
Results
Across the examined paradigms, power contours generally allowed fewer participants when each participant completed more trials, although some designs reached an asymptote.
Takeaways & Limitations
Power contours can guide choices about participants and testing time within the study-design stage.
Takeaways & Limitations
The methods assume random, independent trials, whereas practice and fatigue can limit gains from additional testing.
Abstract
from arXiv · showhide
When designing experimental studies with human participants, experimenters must decide how many trials each participant will complete, as well as how many participants to test. Most discussion of statistical power (the ability of a study design to detect an effect) has focussed on sample size, and assumed sufficient trials. Here we explore the influence of both factors on statistical power, represented as a two-dimensional plot on which iso-power contours can be visualised. We demonstrate the conditions under which the number of trials is particularly important, i.e. when the within-participant variance is large relative to the between-participants variance. We then derive power contour plots using existing data sets for eight experimental paradigms and methodologies (including reaction times, sensory thresholds, fMRI, MEG, and EEG), and provide example code to calculate estimates of the within- and between-participant variance for each method. In all cases, the within-participant variance was larger than the between-participants variance, meaning that the number of trials has a meaningful influence on statistical power in commonly used paradigms. An online tool is provided (https://shiny.york.ac.uk/powercontours/) for generating power contours, from which the optimal combination of trials and participants can be calculated when designing future studies.
Introduction
Statistical power is usually framed around participant sample size, but repeated trials provide a second design dimension when participant-level measurements are noisy. The paper motivates power contours that jointly represent participants and trials.
- Neuroscience studies have estimated power values of 8%-30%, below the desired level of ≥80%.
- Power is commonly increased by recruiting more participants, although the number of repeated trials is another available design choice.
- When within-participant variance is large, increasing trials improves the accuracy of each participant’s estimated mean.
- Cohen’s d depends on the sample mean and sample standard deviation, so trial number can affect power through measurement precision.
- The paper introduces two-dimensional power contour plots, an online generator, and code for estimating within- and between-participant variance.
Power contours
Power contours show how statistical power changes jointly with participant number and trial number. They reveal trade-offs, diminishing returns, and design-specific benefits from allocating effort across both dimensions.
- When within-participant variance is low relative to between-participant variance, power depends mainly on sample size and repeated trials provide little benefit.
- When within-participant variance exceeds between-participant variance, sample standard deviation decreases with more trials and power depends on both trials and participants.
- An 80% power level can be achieved by multiple combinations of sample size and trial number, allowing designs to reflect participant and testing-time constraints.
- Around k = 50 trials, the example contour asymptotes, so further trials provide no additional benefit in that design.
- The authors reanalysed eight studies spanning reaction times, proportional choices, sensory thresholds, EEG, MEG, and fMRI using subsampling.
Reaction times
Reaction-time data produced curved power contours, confirming that power can be maintained through different combinations of participants and trials. The estimated knee-point balanced both resources.
- The reaction-time study included N = 38 participants with k = 600 congruent and k = 200 incongruent trials per participant.
- The estimated within-participant reaction-time standard deviation was σw = 151 ms.
- A curved 80% power contour allowed high power with N > 20 and k < 10, or with k > 50 and N = 8.
- The contour’s knee-point was approximately N = 10 participants with k = 20 trials each.
Proportional choices in the Iowa Gambling Task
In the Iowa Gambling Task, power depended on both participants and trials, but trial order mattered because participants learned the deck contingencies over time. Order-preserving analysis revealed changing power across learning.
- The Iowa Gambling Task dataset contained N = 504 participants choosing among four decks with different overall payoffs.
- The average probability of choosing a good deck was 0.54 versus the chance baseline of 0.5, with d = 0.24.
- Random-trial resampling showed that power depended on both sample size and trial number.
- Participants initially selected bad decks more often, then shifted toward good decks after learning the task contingencies.
- With trial order retained, power was near zero around 60 trials and reached 80% by approximately 80 trials using the full sample.
Sensory thresholds
The psychophysics analysis estimates detection thresholds from fitted psychometric functions and shows how trials and participants jointly determine power. Increasing trials can substantially reduce the participants needed for high power.
- Threshold estimation: Thresholds were estimated by fitting psychometric functions to binary responses across stimulus intensities, typically at 75% correct.The analysis used cumulative Gaussian or Weibull functions to interpolate the criterion threshold.
- Observed effect: The binocular summation effect had d = 1.8, with a mean effect of 6.6dB and sample SD σs = 3.6dB.
- Power analysis: Power contours were generated by subsampling participant-specific trial percentages and refitting each psychometric function.Participants completed approximately 225 trials per condition on average, with adaptive staircases producing different trial counts.
- Variance estimates: The best-fitting variance estimates were σw = 33.5dB within participants and σb = 1.3dB between participants.
EEG: event-related potentials
The ERP analysis examines peak voltages and latencies across three component windows, using repeated-measures comparisons and power contours based on resampled trials and participants. Voltage effects were robust, but latency effects were weaker and varied by component.
- ERP measures: ERP peaks were estimated in 100–150ms, 200–300ms, and 500–700ms windows corresponding to P100, P200, and N600 components.The study analyzed responses from N = 22 participants who each completed k = 600 stimulus pairs.
- Voltage effects: Voltage-difference effect sizes were d = 1.18, 1.11, and 1.32 across the P100, P200, and N600 windows.
- Latency effects: Latency effect sizes were d = 0.21, 0.04, and 0.47 across the same three windows and were not considered further.
- Power contours: Power increased across the tested sample-size and trial-number ranges for P100, whereas N600 power was mainly determined by sample size.For N600, adding trials materially reduced required sample size only when k < 200.
- Variance estimates: Within-participant standard deviations ranged from 12µV to 21µV, compared with between-participant values from 1.1µV to 5.3µV.The relative contribution of trials and participants therefore varied with the effect being studied.
EEG: steady-state evoked potentials
The steady-state EEG analysis compares coherent and incoherent averaging of frequency-tagged responses. Coherent averaging produced larger effects and substantially greater power, with both trials and participants contributing across much of the tested range.
- Data structure: Each participant contributed k = 80 one-second observations per condition after dividing 10-second EEG segments into epochs.The stimuli flickered at 7Hz, producing responses at the fundamental frequency and its second harmonic.
- Averaging method: At 8% contrast, coherent averaging increased the effect size from d = 0.2 to d = 0.68 relative to the 0% contrast baseline.The improvement reflects phase-locked stimulus responses whose noise averages out across repetitions.
- Power comparison: Power contours confirmed substantially greater statistical power for coherent than incoherent averaging.
- Power contours: With coherent averaging, halving the sample from N = 100 to N = 50 required increasing trials from approximately k = 20 to k = 40 to maintain 80% power.
- Variance estimates: The fitted variance estimates for coherent averaging were σw = 3.1µV within participants and σb = 0.19µV between participants.
fMRI: event-related design
The event-related fMRI analysis models stimulus responses in a V1 region of interest using GLMs and estimates power from resampled trials and participants. Sufficient trials can support high power across a wide range of sample sizes, whereas too few trials can severely underpower the design.
- Dataset and preprocessing: The analysis used event-related fMRI data from N = 625 participants, each viewing k = 124 repeated checkerboard stimuli.The study used a V1 region of interest mapped to individual functional data.
- GLM analysis: General linear models separated randomly allocated trials into target and non-target conditions, with a third auditory-only condition.A canonical double gamma haemodynamic response function was convolved with each condition.
- Power analysis: Power contours were generated from beta-value effect sizes using 10,000 resamplings of trials and participants.
- Power contours: 80% power could be maintained for sample sizes from N = 20 to N = 600 by varying the number of trials.
- Design boundary: Fewer than k < 60 trials produced a severely underpowered design, despite the flexibility available from increasing trials.The fitted standard deviations were σw = 515 and σb = 32.2 in beta units.
fMRI: blocked design
Blocked-design fMRI data show stimulus-driven BOLD responses, with power contours indicating diminishing benefits from additional trials in this paradigm.
- Design context: Blocked designs generally have greater power than event-related designs because their stimulus timing better matches sluggish haemodynamic activity.Longer stimulus presentations also contribute to this alignment.
- Dataset and design: The blocked-design dataset included 83 participants viewing image stimuli in repeated 6-second blocks separated by 9-second blank intervals.Each block contained five sequential images, and fMRI volumes were acquired every 3 seconds.
- Observed responses: BOLD responses showed stimulus-driven modulations with a 15-second cycle and peaked 9 seconds after stimulus onset.The same pattern appeared in the example participant and across the population.
- Power contours: Power approximately asymptoted above k = 15 trials, unlike event-related fMRI power, which continued increasing across the assessed range.For larger effects, power was high even with samples below 20 participants, although the visual V1 response produced unusually large effects.
MEG: evoked responses
MEG evoked responses displayed greater within-participant than between-participant variance, and power contours showed that additional trials could offset reduced sample size at later time points.
- Data and analysis: The MEG analysis used 120 trials and selected 50, 54, and 58 ms time points from a dataset of 637 participants.The time points were chosen to examine effects comparable to small response differences in typical experiments.
- Observed effects: Effect sizes increased from d = 0.17 at 50 ms to d = 0.51 at 58 ms when all trials and participants were included.Evoked responses began polarising around 50 ms and showed a larger opposite-polarity peak at 130 ms.
- Variance: Within-participant standard deviations ranged from 8.25−11.77 pT/m, compared with 0.87−6.61 pT/m between participants.The within-participant variance was therefore clearly greater than the sample variance across the analysed time window.
- Power contours: At 54 ms, power could be maintained when reducing N = 400 to N = 200 by increasing trials from k = 20 to k = 60.Power reached 80% at 50 ms only when the full dataset was used.
Discussion
The discussion presents power contours as a joint framework for choosing participants and trials, while showing that trial benefits depend on variance ratios and study context.
- Power contour approach: Power contour plots represent statistical power jointly as a function of sample size N and trials per participant k.The paper provides an online tool and uses subsampling across common psychology and neuroscience paradigms.
- Design trade-offs: Most iso-power contours showed that fewer participants could maintain power when each participant completed more trials.Some paradigms reached a trial-number asymptote, whereas model-fit paradigms continued improving beyond the available data range.
- Assumptions and caveats: Prospective predictions are limited by how well estimated effects and variance parameters generalise to the new experiment.The paper specifically cautions about differences in experimental setups, laboratories, and participant groups.
- Variance estimation: For paradigms lacking directly available within-participant standard deviations, the paper estimated them by fitting simulated power-contour surfaces.For SSVEP and event-related fMRI, an imaginary σb from equation 2 led the authors to set σb = σs for Table 1.
- Variance patterns: Across the considered paradigms, within-participant variance was substantially greater than between-participant variance, although this pattern is not guaranteed universally.Fano-factors clustered for some methods, but generic methodology-level values require evidence across studies, laboratories, and equipment.
- Design trade-offs: When σw > σb, additional trials materially reduce standard error; when σw < σb, increasing participants is more profitable.The σw/σb ratio indicates how strongly changing k is likely to influence power, with blocked fMRI showing relatively small trial gains.
- Assumptions and caveats: The simplifying assumption of one within-participant standard deviation was reasonable in simulations, while empirical MEG power was lower because high-variance outliers contributed disproportionately.The authors note that analysis pipelines may reject such participants or trials.
- Repeated measures: Power contours for repeated-measures designs shift toward the origin as covariance between measures increases, without changing their overall shape.With zero covariance, repeated measures provided no benefit beyond conventional factors.
Application to other statistical tests and approaches
The subsampling approach and power contours extend beyond t-tests to ANOVA and other statistical methods, while also motivating Bayesian and adaptive alternatives. Their usefulness depends on the study design, variance structure, and assumptions about repeated trials.
- Extension to other statistical tests: The subsampling method can extend to nonparametric tests, ANOVA, correlations, and regression without specific data-form requirements beyond each test’s assumptions.Related contour approaches have also been used for reliability and structural equation modelling.
- ANOVA applications: In repeated-measures ANOVA, power contours show how participant number and trial number jointly constrain detection of effects.For the smallest effect in the SSVEP example, an 80% contour could be reached with 100 participants completing 40 trials or 75 participants completing 80 trials.
- Limitations and variance modeling: The methods assume trials are random and independent, so practice and fatigue can limit the gains from additional trials.More accurate variance modeling, including intra-class correlations, can improve power-analysis accuracy; task-based fMRI test-retest reliability is typically low, with mean intra-class correlation below 0.4.
- Bayesian approaches: Bayes factor contours could analogously estimate the trials and participants needed to reach a specified level of evidence for competing hypotheses.The paper presents this as a potential alternative to traditional power analysis as Bayesian methods become more widespread.
- Computational modeling and individual differences: Power contours may be less helpful for computational-modeling studies that treat each participant as an independent replication and emphasize many trials for data quality.Within-participant standard deviation can still inform decisions about the number of trials.
Conclusions
The paper argues that trial count should be incorporated alongside participant count when planning power in psychology and human neuroscience studies. Power contours and subsampling can inform these design choices, although effect sizes and variances remain uncertain before data collection.
- Conclusions: Power contours incorporate trials and participants jointly, helping researchers choose how many people to test and how long to test each one.They can be generated by subsampling existing data sets or using an online tool.
- Conclusions: A priori power calculations remain speculative because true effect sizes and variances are unknown until data are collected.
- Conclusions: The paper presents trial count as a design factor in statistical-power calculations for experimental psychology and human neuroscience.