Source-linked AI summary
On Binscatter
Matias D. Cattaneo, Richard K. Crump, Max H. Farrell, Yingjie Feng
TL;DR
As datasets grow denser, classical scatter plots become harder to use for assessing functional form, while correct covariate adjustment remains subtle. The paper develops formal and visual binscatter tools for improved conditional-mean estimation, uncertainty quantification, and specification assessment, while showing that incorrect adjustment can mislead conclusions.
Problem
As datasets grow denser, scatter plots become increasingly difficult to interpret, and assessing functional form correctly is particularly subtle.
Method
The paper introduces formal and visual binscatter tools centered on a partially linear model, including improved conditional-mean estimation, formal inference, and principled covariate control.
Results
The tools restore and in some dimensions surpass classical scatter-plot visualization benefits through variance visualization, precise uncertainty quantification, and formal tests of substantive specifications.
Takeaways & Limitations
Incorrect covariate adjustment in binscatter applications can mislead practitioners assessing linearity or other hypothesized parametric or shape specifications.
Takeaways & Limitations
The paper identifies covariate adjustment as a methodological constraint: incorrect adjustment can produce misleading specification assessments.
Abstract
from arXiv · showhide
Binscatter is a popular method for visualizing bivariate relationships and conducting informal specification testing. We study the properties of this method formally and develop enhanced visualization and econometric binscatter tools. These include estimating conditional means with optimal binning and quantifying uncertainty. We also highlight a methodological problem related to covariate adjustment that can yield incorrect conclusions. We revisit two applications using our methodology and find substantially different results relative to those obtained using prior informal binscatter methods. General purpose software in Python, R, and Stata is provided. Our technical work is of independent interest for the nonparametric partition-based estimation literature.
Introduction
Binscatters simplify bivariate visualization by displaying bin-level conditional means, but they omit distributional information and require careful covariate adjustment. The paper develops formal estimation, uncertainty quantification, specification testing, and corrected adjustment methods to make binscatter analysis more reliable.
- Binscatter foundations: Binscatters partition x into bins and display the average outcome for observations within each bin.The resulting points represent conditional averages rather than individual observations.
- Limitations: A binscatter is not an exact substitute for a scatter plot because averaging masks observations, bunching, anomalies, and other features of the conditional distribution.The paper addresses this information loss by augmenting binscatters with formal uncertainty quantification.
- Binscatter foundations: Binscatters help assess conditional-mean shape, including linearity, monotonicity, convexity, and concavity, and can guide regression analyses.The paper provides formal results supporting this use in a principled way.
- Contributions: The paper develops a binscatter toolkit for estimating conditional means, visualizing variance, quantifying uncertainty, and testing hypotheses such as linearity or monotonicity.Its framework centers on principled covariate control and includes an extensive theoretical analysis of partition-based methods.
- Covariate adjustment: Incorrect residualization of covariates is formally justified only under a linear conditional-mean structure and can distort the conditional mean’s shape and support.The authors show that this problem can mislead assessments of linearity and other hypothesized specifications.
- Applications: In an application, corrected methods produce a more informative plot and support a linearity conclusion consistent with the original linear regression analysis.The paper also reports an application where linearity is not supported but the methods reinforce and extend an empirical conclusion.
I Canonical Binscatter and Covariate Adjustments
The paper formalizes binscatter as a nonparametric estimator and shows that common covariate-adjustment practices can distort both interpretation and empirical conclusions.
- Canonical binscatter: Binscatter partitions the support of x_i into J bins and plots within-bin averages of y_i.Quantile-based bins contain roughly equal numbers of observations.
- Canonical binscatter: A fixed number of bins can target conditional means within bins rather than the underlying conditional-mean function, producing different statistical and economic interpretations.The discrepancy matters especially when the conditional expectation varies within bins.
- Canonical binscatter: For continuously distributed x_i, binscatter is a piecewise-constant nonparametric approximation to the conditional mean υ_0(x)=E[y_i|x_i].The approach characterizes approximation error and uncertainty while allowing polynomial extensions within bins.
- Covariate adjustment: Residualized binscatter is not recommended because linear residualization of outcomes and regressors can yield difficult-to-interpret or incorrect shapes and supports.The method can make nonlinear regression functions appear linear and can produce incorrect empirical findings.
- Covariate adjustment: Correct covariate adjustment can materially change the visual evidence, producing a clearer estimate and stronger support for the linear regression used by AGNS.The corrected analysis contrasts with apparent nonlinearity under the original binscatter.
II Choosing the Number of Bins
The number of bins J is a tuning parameter that controls the bias–variance trade-off and the interpretation of a binscatter. The paper develops an IMSE-optimal, data-driven selector for J.
- Choosing J: Increasing J reduces approximation bias but increases variance because each bin contains roughly n/J observations.Small J can oversmooth, whereas large J can undersmooth the conditional mean.
- Choosing J: Consistent nonparametric estimation requires J to grow with sample size, but neither too rapidly nor too slowly.As J increases, bin widths shrink and within-bin averages better approximate the conditional mean.
- Choosing J: The paper develops a selector for J that is optimal in integrated mean squared error and incorporates heteroskedasticity, clustering, and covariates.The selector balances asymptotic variance against squared bias.
- Choosing J: J = 11 is the feasible IMSE-optimal choice in the AGNS application.Figure 4 compares J = 5, J = 50, and the optimal J = 11 using two-way cluster-robust variance estimation.
- Choosing J: The data-driven choice of J can benchmark fixed-J implementations and help applied researchers discipline their bin selections.Choosing J much larger than the data-driven benchmark is expected to produce more variability than bias.
III Quantifying Uncertainty
The paper adds uniform confidence bands and formal shape testing to binscatter, addressing uncertainty in estimated conditional-mean functions rather than showing point estimates alone.
- Confidence bands: Confidence bands display uncertainty around the estimated conditional mean uniformly over the support of x_i.They allow readers to assess functional-form and shape hypotheses from the plotted band.
- Shape testing: Uniform confidence bands support formal tests of linearity, monotonicity, and other shape restrictions.A shape is consistent with the data when a corresponding function lies within the band.
- Confidence bands: The bands target Υ_0(x), the function of interest, rather than only the fixed-J parameter Ξ_0.This requires accounting for approximation error from the binscatter representation.
- Confidence bands: Valid nonparametric bands use IMSE-optimal binning together with debiasing and variance estimation that includes the uncertainty introduced by bias correction.The construction is a major technical contribution of the paper.
- Empirical illustration: Figure 5 contrasts fixed-J and large-J inference, with the large-J band ruling out horizontal functions in the application.The fixed-J band cannot rule out that the five conditional means are equal, while the large-J band rejects no relationship.
IV Another Empirical Illustration
Reanalysis of Moretti’s application shows that correcting covariate adjustment and bin selection changes the visual and econometric conclusions, revealing a nonlinear relationship between patents and high-tech cluster size.
- Original analysis: The original binscatter is misleading because its dots estimate conditional means rather than representing individual observations.The raw scatter plot is dense and uninformative with close to one million observations.
- Corrections: Incorrect residualization and too many bins make the original estimate visually distorted and substantially undersmoothed.The corrected approach restores the proper scale and uses a more appropriate bin count.
- Corrected estimates: With the IMSE-optimal choice J_IMSE = 18, the conditional expectation appears roughly flat for smaller cluster sizes and rises sharply for larger clusters.This pattern suggests a nonlinear relation between productivity and high-tech clusters.
- Formal inference: The confidence band rejects both no relationship and linearity but fails to reject convexity.The same conclusions arise in the specification with 11 fixed effects.
- Implications: The revised analysis adds nuance to Moretti’s results by indicating that incentives may need to be especially generous for small inventor clusters to grow.The paper contrasts this nonlinear interpretation with the reported elasticity of 0.0676.
V Theoretical Foundations
The paper develops theoretical foundations for canonical and covariate-adjusted binscatter estimators, including flexible polynomial bases, random partitions, and inference. Its results characterize estimation, variance, approximation, optimal binning, and uniform distributional behavior.
- Estimator construction: The theoretical framework studies canonical and covariate-adjusted binscatter estimators with polynomial basis functions and smoothness constraints across bins.The basis permits piecewise polynomial fitting, derivative estimation, and continuity restrictions across bin boundaries.
- Large-sample theory: The appendix derives technical results for random empirical-quantile partitions, their basis functions, Gram matrices, variance, approximation error, and uniform convergence.These results explicitly account for randomness in the binning scheme.
- Inference: Theoretical results establish point estimation, variance estimation, distributional approximation, and inference for the extended binscatter estimator.The framework includes robust bias correction and explicitly incorporates covariate adjustment and random binning.
- Optimal binning: A density-weighted IMSE expansion yields an optimal bin count J_IMSE(p, s, v) by optimizing the leading terms.The paper also discusses feasible implementation of IMSE-optimal tuning.
- Theoretical implications: The rate conditions support canonical binscatter with p = s = 0 while allowing bias and variance to be controlled simultaneously.Under a subexponential moment restriction, the text states that J/n → 0, up to log(n) terms, suffices.
Appendix treats the general case). We study the t-statistic
The paper develops conditional strong approximations for binscatter t-statistic processes despite the discontinuities and randomness created by data-dependent binning. These approximations underpin feasible confidence bands and hypothesis tests, with robust bias correction improving tuning-parameter robustness.
- Approximation strategy: The analysis targets a distributional approximation for the entire t-statistic process uniformly over x.This process-level focus supports inference on functionals such as suprema.
- Feasible inference: The feasible strong approximation proceeds from stochastic linearization and conditional coupling to a data-dependent Gaussian process that can be simulated.The resulting construction uses Gaussian draws and already-computed estimator and variance components.
- Approximation strategy: Random binning creates sharp indicator discontinuities that obstruct uniform convergence of the basis, so the method retains binning randomness in a conditional Gaussian coupling.The coupling circumvents the technical lack of uniformity in the random basis.
- Feasible inference: The simulated process supplies critical values for confidence bands and hypothesis tests, including tests based on suprema of the t-statistic process.The paper describes the plug-in implementation as relying only on Gaussian draws and computed binscatter quantities.
- Tuning and bias correction: Robust bias correction based on an IMSE-optimal binscatter is more robust to the choice of J than directly using the IMSE-optimal estimator for inference.The direct IMSE-optimal choice can be too small to remove enough bias for correct t-statistic centering and may require additional smoothness assumptions.
- Theoretical implications: The rate conditions are weak enough to accommodate canonical binscatter, unlike stronger restrictions required by some prior strong-approximation results.The paper states that its theoretical improvements have direct practical consequences for the p = 0 estimator.
VI Conclusion
The conclusion presents binscatter as a useful but previously insufficiently validated visualization and specification-testing tool. The paper supplies formal methods for estimation, uncertainty, covariate adjustment, and testing, and applies them to empirical reanalyses where incorrect practice can mislead conclusions.
- Motivation: Binned scatter plots offer a simple way to display relationships between an outcome and a covariate, but their accuracy had not been formally established.The paper links this gap to concerns about reliability and usability in applications.
- Contributions: The paper introduces formal and visual tools for conditional means, covariate adjustment, variance visualization, uncertainty quantification, and tests of linearity or monotonicity.These tools are intended to improve and, in some cases, correct empirical practice.
- Applications: Two empirical applications show pitfalls from using binned scatter methods incorrectly and that correct implementation can strengthen reported empirical findings.The applications revisit published work by Acemoglu et al. and Moretti.
- Scope: The methods extend to multivariate x_i and can support heat maps, although the paper itself focuses on conditional means.Conditional quantiles and other nonlinear features are treated in related work rather than here.
- Applications: Figure 6 compares raw, residualized, semi-linear, and optimally binned representations of productivity and high-tech cluster size, with 95% cluster-robust confidence bands.The figure uses log patents per inventor per year as the dependent variable and log cluster size as the independent variable.
Online Appendix
The online appendix supplies additional methodological and theoretical results, including proofs and broader results for partitioning-based semi-linear estimation and inference. Companion software and replication files are available online.
- Supplement contents: The supplement presents additional methodological results, general theoretical results, and all technical proofs.It broadens the theoretical treatment beyond the results reported in the main paper.
- Theoretical scope: The new theoretical results concern least-squares partitioning-based semi-linear series estimation and inference.The appendix describes these results as independently relevant to the broader literature.
- Resources: Companion general-purpose software and replication files are available through the binsreg project website.The supplement identifies the project URL as https://nppackages.github.io/binsreg/.
Additional Methodological Results
The appendix develops two analytical examples of incorrect covariate adjustment and examines how the evaluation point w affects visualization, estimation, and inference.
- The analysis addresses both covariate-adjustment problems and evaluation-point choices as distinct methodological issues.
- Two stylized examples characterize how incorrect covariate adjustment affects binscatter results.
- The evaluation point w plays a role in visualization, estimation, and inference for Υ0(x, w) = E[yi|xi = x, wi = w].
Bias of Residualized Binscatter
Incorrect residualization can transform the relationship displayed by binscatter, producing functions that differ from the original conditional relationship and can appear misleadingly linear.
- For a quadratic true function, residualization produces a vertical shift of the true function.
- For a cubic true function, residualization adds a linear covariate component that can visually dominate when |ρσx| is large.
- The resulting visual pattern can incorrectly suggest a linear model even when the underlying relationship is cubic.
- Residualized binscatter generally yields a polynomial relationship that may differ from the original polynomial function.
- Residualizing the covariate changes its support in both location and length, further separating the transformed relationship from the original one.
- In the Bernoulli example, the original relationship is quadratic and heterogeneous across groups, whereas residualization produces a linear function in zi.
Impact of Evaluation Point w
The evaluation point w affects binscatter levels, uncertainty, and specification testing, even when the central interest is how the outcome relates to x after controlling for w.
- The function µ0(x) is defined relative to how the controls wi are coded, making evaluation-point choice consequential.
- Specification-test results can depend on w because the estimated control coefficients and implied intercepts differ across evaluation points.
- Changing w shifts the visualization and estimator vertically and changes their comparison with parametric fits.
- For categorical controls, changing the baseline category can produce different test conclusions even when the hypothesis is unchanged.
- With continuous controls, some evaluation point can lead to rejection when the semiparametric and parametric control estimates differ.
- The authors propose testing derivatives of µ0(x) rather than levels to avoid evaluation-point dependence in the central binscatter specification question.
- Confidence bands also vary with w: the curve's shape is unchanged, but its level and band size can change.
General Setup and Notation
The framework defines binscatter estimation using partition-based polynomial bases, control adjustment, and evaluation points, with assumptions supporting estimation and inference.
- The data consist of a scalar outcome yi, scalar covariate xi, and vector controls wi under a partially linear regression framework.
- The setup imposes smoothness, density, moment, and conditional-design assumptions on the data-generating process and least-squares model.
- Binscatter partitions the support of xi into J intervals using empirical quantiles, with J serving as a tuning parameter that diverges with sample size.
- The estimator uses piecewise polynomial bases of degree p, optionally transformed to impose smoothness across bins.
- The standardized rotated basis represents the same linear function space as the original piecewise polynomial basis while scaling local polynomials by bin length.
- The evaluation point estimator is assumed either fixed or generated from the controls W.
- The paper uses an IMSE-optimal bin choice for estimation and a higher-degree basis for inference to support bias-corrected confidence bands.
Notation
The paper establishes notation for norms, asymptotic orders, empirical processes, partitions, bases, matrices, and common data objects. It distinguishes random empirical-quantile partitions from their population counterparts and defines associated binscatter approximations.
- Functions and asymptotics: The paper defines L2 and L∞ norms, derivatives, asymptotic order notation, convergence in probability, and equivalence up to constant factors.
- Empirical processes: Empirical-process notation includes empirical averages and covering numbers for measurable function classes relative to L2(Q) and an envelope function.
- Partitions and bases: Partitions are indexed by empirical and population quantiles, with bin widths defined as differences between consecutive quantile cutoffs.
- Partitions and bases: The binscatter basis is a piecewise-polynomial basis of degree p with smoothness controlled by s, constructed from a specified partition.
- Projections and errors: Population least-squares projections and approximation errors are defined separately for random empirical partitions and nonrandom population partitions.
Theoretical Results
This section focuses its main theoretical analysis on an estimator of a conditional target and notes that the paper’s estimator for the target evaluated at the mean covariate is covered as a special case.
- The main theoretical results analyze the estimator bΥ(v)(x, bw).
- The framework therefore encompasses both the general covariate-indexed estimator and the version evaluated at the mean covariate.
- The estimator bΥ(x) for Υ(v)0(x, E[wi]) is covered as a special case of the analyzed framework.
Properties of Quantile-Based Partition and Binscatter Basis
The paper characterizes empirical-quantile partitions and the associated binscatter basis, emphasizing quasi-uniformity and explicitly accounting for partition randomness in the approximation analysis.
- Quantile-based partitions: A partition is quasi-uniform when its bin lengths do not differ excessively, a property needed for partitioning-based asymptotics.
- Quantile-based partitions: Quantile-spaced partitions possess quasi-uniformity with probability approaching one under the stated data-generating-process assumptions.
- Binscatter basis: The empirical binscatter basis is linked to a piecewise-polynomial basis through a transformation matrix based on the empirical-quantile partition.
- Binscatter basis: The paper establishes local basis bounds and characterizes approximation errors in sup norm under its regularity conditions.
- Contribution: Relative to prior spline analyses, the results formally incorporate randomness from empirical-quantile partitions.
Preliminary Technical Lemmas
The preliminary lemmas establish the technical ingredients for binscatter estimation, inference, variance estimation, optimal bin selection, and distributional approximation. Across these results, the analysis accounts for empirical-quantile partition randomness and semi-linear covariate adjustment.
- Core estimation lemmas: The lemmas characterize Gram-matrix properties, asymptotic variance, variance-driven uniform convergence, projection errors, and the parametric component of the estimator.
- Covariate adjustment: The covariate-adjustment lemma supports treating the estimation error of γ as negligible for large-sample inference on the target functions.
- Improvements and applications: The results improve prior work by accounting for empirical-quantile partition randomness, semi-linear regression structure, and multiple inference problems while exploiting the locally bounded binscatter basis.
- Linear approximation: The stochastic linear approximation theorem yields a sharp Bahadur-type approximation rate for binscatter bases with or without random binning.
- Inference: Uniform convergence and consistent variance estimation follow as consequences of the main approximation results.
- Pointwise inference: Pointwise inference uses higher-order binscatter estimators with robust bias correction and an IMSE-optimal bin choice under stated moment and rate conditions.