Source-linked AI summary
Across-Design Uncertainty in Short Pricing Panels: Evidence from Simulated Price Trajectories
Pedro Cadahia Delgado
TL;DR
Short pricing panels can have many observations but few independent price movements, raising questions about what conventional intervals measure. Using a synthetic pricing process, the paper separates uncertainty conditional on one trajectory from estimator variation across trajectories and finds that the latter dominates baseline error variance, motivating designs with independent identifying variation.
Problem
Short pricing panels may offer many rows but only a small number of distinct price movements, limiting evidence about inferential performance across alternative pricing trajectories.
Method
The paper uses a synthetic pricing environment to compare repeated shocks within fixed price trajectories with repeated price trajectories, and evaluates a variance component across independently priced units.
Results
97.6% of gradient-boosted estimation-error variance is associated with design-specific conditional mean error in the baseline simulation.
Takeaways & Limitations
The findings shift practical emphasis toward generating independent treatment variation rather than relying solely on fixed passive panels.
Takeaways & Limitations
Quantitative results come from a synthetic generator and do not establish that between-design dispersion dominates every short pricing panel or estimator.
Abstract
from arXiv · showhide
Short observational pricing panels can contain many observations while offering only a small number of distinct price movements. This paper studies the inferential consequences of that distinction in a synthetic data-generating process calibrated to a sparse pricing regime. We separate uncertainty conditional on a realised price trajectory from variation in estimation error across alternative trajectories generated by the same pricing process. In the baseline simulations, the latter component accounts for 97.6% of the variance of estimation error for the gradient-boosted specification. Within-panel resampling procedures use the information of one realised trajectory and do not identify this across-design component. Three results organise the analysis. First, across-design dispersion is well described by the empirical relation sigma_hat approx 0.182 V^(-0.271), where V equals moves times magnitude squared. Second, adding regions sharing a common price path reduces outcome noise but does not create independent price trajectories; conversely, averaging across units with independent design-specific errors reduces dispersion at the standard square root rate. Third, a Paule-Mandel variance component estimated across independently priced units substantially increases empirical coverage in homogeneous simulations, from 0.469 to 0.931. The broader implication is a shift toward designing data-generating processes that create independent identifying variation rather than relying solely on fixed passive panels.
1 Introduction
Short pricing panels may contain many rows but little independent price variation. The paper uses simulations to separate within-trajectory shock uncertainty from across-trajectory design uncertainty and argues that additional identifying variation matters for inference.
- Motivation: Many region–week observations can reflect only a handful of common list-price changes, so row count may overstate available treatment variation.The paper studies what conventional intervals represent when the observed price trajectory is one draw from a broader pricing process.
- Inferential target: The simulations distinguish shock uncertainty conditional on one realised design from estimator displacement across alternative price trajectories.Repeated shocks hold the design fixed, whereas repeated designs redraw the pricing trajectory and other non-shock features.
- Inferential target: Within-panel resampling uses one realised design and therefore does not identify the distribution of estimator displacement across unrealised price histories.This does not imply that bootstrap or clustered procedures are invalid for conditional questions supported by the observed panel.
- Baseline evidence: 97.6% of gradient-boosted estimation-error variance is associated with design-specific conditional mean error in the baseline simulation.The corresponding share is 99.3% for the sieve specification, and the bootstrap standard error is much smaller than across-design dispersion.
- Data design: Adding rows under a common price path can reduce conditional outcome noise but does not create independent identifying trajectories.Under approximately independent design-specific errors, averaging independently priced units reduces standard deviation at the square-root rate; common design components limit that gain.
- Data design: The empirical relation sigma_hat approx 0.182V^(-0.271) describes simulated design dispersion as a function of V = nmoves × magnitude^2.The exponent is reported as a simulation regularity rather than a general convergence rate.
3 Within- and between-design uncertainty
The paper defines two components of estimation-error dispersion by holding pricing designs fixed while resampling shocks and then comparing alternative designs. This distinction separates conditional inference from repeated-design performance.
- Definitions: The variance decomposition separates within-design shock variance from between-design variance in the design-specific conditional mean error.The law of total variance provides the algebraic basis for measuring both components.
- Coverage targets: Conditional coverage holds the realised design fixed, whereas repeated-design coverage redraws the price trajectory from the simulation process.Neither target dominates by definition; the appropriate target depends on the decision problem.
- Baseline decomposition: 97.6% of gbr estimation-error variance and 99.3% of sieve estimation-error variance comes from between-design centring dispersion in the baseline DGP.For gbr, within-design standard deviation is sigma_w = 0.0756, while conditional mean error varies substantially across designs.
- Baseline decomposition: The mean gbr bootstrap standard error is 0.1595, about 2.1 times sigma_w and about one third of sigma_b.Thus the bootstrap is not simply estimating within-design shock dispersion, but it also does not reproduce the measured across-design distribution.
- Coverage targets: A typical gbr interval half-width of about 0.305 still yields mean conditional coverage of 0.565 across designs.The result points to variation in the estimator’s conditional centre, rather than interval width alone, as an important source of undercoverage.
4 No evaluated within-panel construction achieves nominal coverage
The paper compares eight within-panel interval constructions on identical fits and finds that none reaches nominal coverage in the finite simulation experiment. Wider intervals improve coverage but do not by themselves resolve the shortfall under the stated precision criterion.
- Simulation comparison: None of the eight evaluated interval constructions reaches 0.95 coverage over 200 independently drawn panels.The comparison uses the same point estimator and identical fits across constructions.
- Simulation comparison: Multiway clustering has the highest measured coverage among the evaluated procedures for both learners.This ranking is a finite simulation comparison, not an impossibility result for all within-panel methods.
- Simulation comparison: The reference repeated-design width is 3.92 sigma_hat, where sigma_hat is the realised across-panel standard deviation of estimation error.The table labels this quantity as a descriptive benchmark without a formal finite-sample coverage guarantee.
- Precision trade-off: Under the application-specific criterion W <= 0.6, the higher-coverage intervals remain too wide to meet the paper’s operational precision requirement.The resulting coverage–width trade-off motivates the across-unit variance-component exercise.
5 An empirical dispersion pattern over the simulation grid
Across the simulated grid, larger heuristic price-movement variation is associated with lower but still substantial across-design dispersion. Aggregation reduces dispersion when design-specific errors are weakly correlated, while persistent common bias and covariance limit what pooled estimates can achieve.
- 576 configurations were evaluated across 288 grid points, with each learner replicated eight times and dispersion measured across replications.The exercise varied price-move counts, magnitudes, promotion confounding, and pass-through.
- The fitted relation is ˆσb = 0.182 V^-0.271, with standard error 0.023 and R2 = 0.86 for gbr in the selected simulation column.Here V = n · magnitude^2 summarizes the amount and size of list-price movement.
- The exponent −0.271 is an empirical regularity of this simulation column, not a theoretical convergence rate or asymptotic identification result.The fitted relationship is descriptive and should not be mechanically generalized beyond the simulated support.
- None of the 576 learner–configuration combinations achieved both absolute bias below 0.15 and design dispersion below 0.20.The result applies to the enumerated grid and finite replication design, not to configurations outside it.
- The median ratio of repeated-design dispersion to reported standard-error scale was 2.26 for gbr and 8.03 for sieve.These are simulation diagnostics rather than generic correction factors for minimum detectable effects.
- Averaging k independently priced units reduces design-error dispersion under weak covariance, but shared price paths preserve common design components and aggregation does not remove common mean bias.The simulated cross-unit reductions were close to the independent-design benchmark, while learner rankings and coverage changed with aggregation level.
7 A variance-component interval estimated across realised designs
The paper uses a Paule–Mandel variance component to estimate excess dispersion across independently realised pricing designs, then evaluates intervals that incorporate it. In homogeneous simulations this improves coverage substantially, but the added width creates a precision trade-off and heterogeneous truths confound interpretation.
- Working model: A single realised price trajectory cannot nonparametrically identify estimator displacement across counterfactual trajectories; independently realised pricing designs provide information only under additional structure.The working model assumes approximately independent design draws, adequate within-unit variances, exchangeable displacement, and common truths for a clean design-variance interpretation.
- Estimation: Paule–Mandel estimates excess between-unit dispersion as a working variance component rather than introducing a new variance-component method.The paper adapts a standard meta-analysis estimator to independently realised pricing designs.
- Results: Coverage increased from 0.469 for the bootstrap percentile interval to 0.931 for the posterior-centred variance-augmented interval in homogeneous simulations.The estimated τ = 0.372 exceeded the mean bootstrap standard error of 0.144, indicating substantial excess dispersion across unit-level designs.
- Results: Shrinkage alone produced 0.426 coverage because it reduced conditional variance while leaving design-specific displacement largely unmodelled.This result is specific to the simulated DGP and should not be generalized to hierarchical procedures that explicitly model the design process or bias component.
- Coverage–precision trade-off: Variance-augmented intervals were 2.70 and 2.76 times as wide as the bootstrap, exceeding the paper’s 0.6 operational precision threshold.Representing across-design dispersion improved coverage but materially reduced decision precision in this DGP.
8 The variance-component interval at higher aggregation levels
The paper tests whether aggregating units with independently generated price paths makes variance-component intervals sufficiently narrow for decisions. For the evaluated portfolio and threshold, coverage remained strong for sieve but no aggregation level met both coverage and precision criteria.
- Portfolio exercise: The portfolio exercise asks whether aggregation makes the wider variance-component interval narrow enough to support a decision.The construction is evaluated at unit, brand, and category levels using 24 units across six categories, two brands per category, and two units per brand.
- Overall result: For the evaluated portfolio and threshold, no aggregation level simultaneously achieved the paper’s coverage and precision criteria.This negative result is specific to the learners, scenarios, portfolio, and width threshold studied, not an impossibility claim for other settings.
- Sieve: Sieve coverage remained near or above nominal across aggregation levels, reaching 0.953, 0.958, and 0.962 in the homogeneous scenario.In the heterogeneous scenario, sieve coverage ranged from 0.974 to 0.993, while its estimated τ declined toward the independent-design benchmark.
- Gradient-boosted specification: For gbr, mean bias persisted as aggregation increased and homogeneous-scenario coverage declined from 0.899 to 0.665.The heterogeneous-scenario decline was milder, but τ also included genuine effect heterogeneity.
- Precision threshold: The narrowest covering interval had 0.861 coverage and remained 1.43 times the 0.6 operational precision threshold.None of twelve combinations achieved 0.90 coverage within width 0.6; the closest coverage was 0.993 at width 1.250.
- Interpretation: With heterogeneous truths, estimated τ mixes genuine heterogeneity with design dispersion and should be interpreted conservatively rather than as either component alone.The reported cell achieved 0.993 coverage under this mixture interpretation.
9 Implications for data design
The simulations distinguish adding observations from adding independent identifying variation. They support prioritising price trajectories with low design-error covariance, including through randomisation or plausibly exogenous variation, while treating the fitted dispersion relation as heuristic.
- Independent variation: Adding regions under a common national price path can reduce conditional outcome noise but does not create an independent price trajectory.Aggregation gains depend on covariance among design-specific errors rather than raw row counts.
- Price-path design: Across the simulation grid, larger V = nmoves × magnitude^2 was associated with lower across-design dispersion.Doubling V was descriptively associated with a 1.21-fold reduction in σb under the fitted relation, not a general forecast.
- Independent variation: Independently generated price trajectories produced approximately square-root reductions in design-specific dispersion when their relevant design errors had low covariance.Distinct product labels alone do not guarantee independence; pricing rules, common cost shocks, and synchronised promotions must be assessed.
- Design recommendation: Randomised regional price assignment can create known treatment variation that supports design-based reasoning, while natural experiments and staggered policy changes may provide similar identifying variation.The recommendation is to create or exploit independent identifying variation, not to treat randomisation as the only admissible design.
10 Limitations
The evidence is limited to a synthetic generator and several working assumptions, so its quantitative magnitudes and estimator-specific conclusions do not automatically transport to empirical pricing panels or other estimators.
- Synthetic evidence: All quantitative results come from the study’s synthetic data-generating process, limiting external validity across pricing panels and estimators.The simulations establish that the mechanism can be important under the stated DGP, not that it dominates universally.
- Measurement: The decomposition must subtract the design-specific truth θ(D); otherwise variation in the estimand is combined with variation in estimation error.Reproducibility reports should state explicitly which object is used to compute σb.
- Variation index: V = nmoves × magnitude^2 is a heuristic simulation-grid summary, not a derived information measure or proven general scaling law.The exponent −0.271 is descriptive, and extrapolations from it are illustrative only.
- Variance-component assumptions: Interpreting Paule–Mandel τ^2 as design variance requires approximate independence, adequate within-unit variance estimates, exchangeability, and common unit-level truths.With heterogeneous truths, τ^2 mixes genuine heterogeneity and design uncertainty; with eight units, its own estimation uncertainty may also be material.
- Estimator scope: The simulations use a DML-motivated estimator that is not the canonical pooled DML2 estimator, so findings should not be attributed to Double Machine Learning as a class.Analogous dispersion for instrumental-variable, structural, Bayesian, or alternative orthogonal-score estimators remains unresolved.
- External validation: The W ≤0.6 precision threshold and identification filter are simulation-specific operational criteria rather than externally validated pricing cutoffs.A natural next step is testing variation summaries and dependence diagnostics on real histories with multiple quasi-independent or randomized price paths.
11 Conclusion
The simulations show that short-panel inference is dominated by variation across alternative price trajectories, motivating data designs that create independent identifying variation. Across independently priced units, variance-component methods can improve coverage, but their interpretation depends on homogeneity.
- Conclusion: 97.6% of gradient-boosted estimation-error variance comes from variation in the design-specific conditional mean error across price trajectories.For the sieve specification, the corresponding share is 99.3%.
- Conclusion: Within-panel resampling does not identify estimator displacement across unobserved price histories, even when conventional standard errors capture conditional shock variation.The moving-block bootstrap standard error exceeds measured within-design dispersion in the baseline, but none of eight evaluated within-panel constructions reaches nominal coverage.
- Conclusion: Independent design draws reduce the standard deviation of averaged design-specific errors, whereas additional observations under a common price path do not create equivalent identifying variation.The gain from aggregation depends on covariance across design-specific errors.
- Conclusion: 0.469 to 0.931: Paule–Mandel variance augmentation raises empirical coverage for the posterior-centred construction in homogeneous simulations.The resulting intervals are much wider, and with heterogeneous true effects the component mixes heterogeneity with design uncertainty.
- Conclusion: σ̂b ≈ 0.182V^-0.271 summarizes lower across-design dispersion with more independent price movement within the explored simulation grid.The exponent is descriptive rather than a general scaling law.
A.1 Exogenous series
The exogenous-series generator constructs prices and related covariates through stochastic regional and time processes, with configurable promotion confounding and competitor-price collinearity. It also records realised pass-through because the estimand depends on realised series behaviour.
- Exogenous series: The inflation index, calendar variables, weather, costs, promotions, and competitor prices are generated as components of the exogenous series.Weather includes region-specific temperature draws and seasonal variation; costs receive random step shocks.
- Exogenous series: The list price is a step function with randomly timed moves and signed magnitudes, while promotion windows create persistent within-series discounts.Price moves are drawn without replacement from eligible weeks.
- Exogenous series: Promotion confounding is controlled by discarding moves on promotion weeks when promo_conf = 0 and relocating moves onto promotion weeks with positive probability otherwise.This mechanism produces orthogonal and confounded pricing regimes.
- Exogenous series: In the collinear regime, list price becomes a near-deterministic function of competitor price, while competitor prices also follow a random walk with optional own-price reaction.Discrete competitor jumps are retained in experiments because otherwise the cross-price channel is not identified by construction.
- Exogenous series: Pass-through varies across regions and depends on the retailer’s margin state, so the paper uses realised series pass-through rather than configured ρ as ground truth.Comparing against βi or configured ρi would confound generator dispersion with estimator bias.
A.4 Demand
The demand construction combines price, promotion, competitor, calendar, weather, income, and inflation inputs, with optional forward buying and order refusal. The estimator targets the realised convolution of consumer elasticity and pass-through under split pricing.
- Demand: The demand equation uses a cash-pressure index and an elasticity profile that can remain constant or include a structural break.The cash-pressure index combines calendar timing with standardised regional income.
- Demand: Sell-in volume adds forward buying before list-price increases and optional order-refusal noise to sell-out volume.Forward buying applies in the two weeks preceding qualifying price rises.
- Demand: ∂log Q/∂log P_si = βρ: the generated sell-in elasticity is the consumer elasticity convolved with pass-through.This identity is the convolution discussed in the paper’s demand remark.
- Demand: Nominal prices are stored while demand is generated from deflated prices, with log inflation included among controls so deflation cancels in the partialled-out regression.This implementation detail affects replication of estimated bias.
- Demand: The baseline configuration used for the tables is defined separately from the demand equations, with entries outside the indicated baseline reserved for other experiments.The caption identifies the table as the baseline configuration.
A.5 Regimes, portfolio and baseline parameters
The paper defines pricing regimes and portfolios for controlled simulation comparisons, then specifies cross-fitting, estimator aggregation, bootstrap, replication, and pre-registration procedures. These choices determine how baseline inference and coverage results are evaluated.
- A.5 Regimes, portfolio and baseline parameters: The five regimes are clean, weak, confounded, collinear, and round point, spanning different amounts of price variation and treatment confounding.The round-point regime pins shelf prices and sets θ = 0.
- A.5 Regimes, portfolio and baseline parameters: The default portfolio contains eight units in four brands and two categories, with elasticities from −0.90 to −1.40 and configured pass-through from 0 to 0.85.Controlled-topology experiments substitute portfolios tailored to their aggregation design.
- A.5 Regimes, portfolio and baseline parameters: Nuisance functions use five contiguous temporal blocks without an embargo, allowing a week to appear in training through another region.An embargoed variant exists but is not adopted because it would alter the residualiser in two ablations.
- A.5 Regimes, portfolio and baseline parameters: −0.0911 for gbr and +0.0400 for sieve: paired absolute-error contrasts support the fold median for gbr but not for sieve.Against the across-fold mean, gbr improves by −0.1974 while sieve is statistically tied at +0.0108.
- A.5 Regimes, portfolio and baseline parameters: The moving-block bootstrap uses seven-week blocks at W = 120, carrying every regional observation from a resampled week together.The hierarchical variant additionally resamples cells within replicated weeks.
- A.5 Regimes, portfolio and baseline parameters: 35 experiments and 69.5 hours of aggregate compute support the reported simulation program, with independent panels defining replication-level standard errors.The paper reports only part of the broader run.
- A.5 Regimes, portfolio and baseline parameters: 13 of 30 pre-registered rules pass, 16 fail, and one does not apply, forming the empirical core of the reported sections.Selection and reporting use disjoint seeds to avoid inheriting benchmark selection bias.