Source-linked AI summary

A Practical Introduction to Regression Discontinuity Designs: Foundations

Matias D. Cattaneo, Nicolas Idrobo, Rocio Titiunik

arXiv:1911.09511v1stat.MEecon.EMstat.APstat.CO

TL;DR

Because many social-science treatments cannot be randomly assigned, the authors provide a practical guide to canonical Sharp RD designs and illustrate its methods. The guide focuses on a single continuous score, one cutoff, and perfect compliance, while emphasizing common practices for analyzing and interpreting RD evidence.

  • Problem

    Many social-science treatments cannot be randomly assigned, while treated and untreated units may differ systematically in ways related to outcomes.

  • Method

    The authors develop an accessible guide to canonical Sharp RD analysis with one continuous score, one cutoff, and perfect treatment compliance.

  • Results

    2.92708 is the estimated treatment-indicator coefficient in the illustrative regression, matching the difference between separate-regression intercepts.

  • Takeaways & Limitations

    The guide encourages common practices for analyzing and interpreting canonical Sharp RD designs and supports the accumulation of RD-based empirical evidence.

  • Takeaways & Limitations

    Covariate adjustment generally cannot repair an RD design with discontinuous predetermined covariates without additional parametric assumptions or redefining the parameter.

Abstract

from arXiv · show

In this Element and its accompanying Element, Matias D. Cattaneo, Nicolas Idrobo, and Rocio Titiunik provide an accessible and practical guide for the analysis and interpretation of Regression Discontinuity (RD) designs that encourages the use of a common set of practices and facilitates the accumulation of RD-based empirical evidence. In this Element, the authors discuss the foundations of the canonical Sharp RD design, which has the following features: (i) the score is continuously distributed and has only one dimension, (ii) there is only one cutoff, and (iii) compliance with the treatment assignment is perfect. In the accompanying Element, the authors discuss practical and conceptual extensions to the basic RD setup.

1 Introduction

The introduction presents Regression Discontinuity (RD) as a credible non-experimental design for causal analysis and motivates a practical guide to improve consistency in RD implementation and interpretation. It focuses on the canonical Sharp RD design and the continuity-based framework, illustrated through practical methods and applications.

  • Motivation: RD assigns treatment according to whether a unit’s continuously measured score exceeds a known cutoff, making it a prominent strategy for studying causal effects without randomized assignment.The introduction describes RD as one of the most credible non-experimental strategies for causal analysis.
  • Contribution: The guide responds to growing RD use across disciplines alongside substantial disparities in how RD analyses are implemented, interpreted, and evaluated.The authors aim to encourage common practices and facilitate the accumulation of RD-based empirical evidence.
  • Scope and framework: This Element focuses on the canonical RD setup and the standard continuity-based framework, which relies on smoothness conditions for regression functions and is most commonly used in practice.The accompanying Element covers extensions and the alternative local-randomization framework.
  • Practical approach: The presentation covers assumptions, target-parameter interpretation, graphical analysis, treatment-effect estimation, statistical inference, and strategies for evaluating design plausibility.The discussion is intentionally geared toward practitioners and emphasizes conceptual clarification.
  • Illustration and implementation: Methods are illustrated through Meyersson’s study of Islamic political representation and women’s educational attainment in Turkey, using the margin of victory as the RD score.The example is suited to illustrating continuity-based methods in this Element and local-randomization methods in the accompanying Element.

2 The Sharp RD Design

The Sharp RD design identifies a local causal effect by comparing units with nearly identical scores on opposite sides of a cutoff, under continuity and perfect compliance. The Turkish mayoral-election example illustrates how this design addresses systematic differences between municipalities and estimates the effect of Islamic local rule.

  • Sharp RD foundations: Sharp RD requires that the assigned treatment condition be identical to the treatment actually received.Non-compliance introduces complications and typically requires stronger assumptions to identify treatment effects.
  • Sharp RD foundations: Because treatment and control units cannot share the same running-variable value, RD analysis relies on local extrapolation toward the cutoff.This is an extreme case of lack of common support.
  • Sharp RD foundations: At the cutoff, the Sharp RD effect is identified by the vertical difference between the potential-outcome regression curves, where both curves are almost observed.The treatment effect at a score value is the difference E[Yi(1)|Xi = x]−E[Yi(0)|Xi = x].
  • Sharp RD foundations: Continuity of the treated and untreated regression functions at the cutoff makes units with very similar scores on opposite sides comparable.This comparability assumption is the fundamental concept underlying RD designs.
  • Turkish mayoral-election example: In the Turkish example, the score is the Islamic margin of victory in close 1994 mayoral elections, and the design studies Islamic-party control of municipalities.The unit of analysis is the municipality, and the score compares the largest Islamic party with its largest secular opponent.
  • Turkish mayoral-election example: Under appropriate assumptions, comparing municipalities where the Islamic party barely wins versus barely loses isolates the causal local effect of Islamic rule on women’s educational attainment.The design addresses systematic differences that could otherwise confound comparisons between Islamic and secular municipalities.

3 RD Plots

RD plots make analyses more transparent by displaying observations and summarizing empirical findings, but raw scatter plots often obscure discontinuities. Binning the data and adding separate polynomial fits reveals cutoff behavior, regression-function shape, and observation density, with bin choices selected using data-driven criteria.

  • RD plots enhance transparency by displaying estimation and inference observations while summarizing empirical findings and other features of the analysis.
  • A typical RD plot combines local sample means with separate fourth- or fifth-order polynomial fits above and below the cutoff using the raw data.
  • Raw scatter plots are often ineffective for detecting RD discontinuities, whereas binned means with a global polynomial fit reveal cutoff jumps and regression-function shape.In the Meyersson application, the raw scatter plot contains 2,629 municipality-level observations, yet shows no visible discontinuity; the binned plot does.
  • Quantile-spaced bins contain approximately equal numbers of observations and quickly display observation density, whereas evenly-spaced bins have equal lengths.
  • Data-driven RD plots select the numbers of bins on each side after choosing quantile-spaced or evenly-spaced bins, using criteria such as IMSE or mimicking variance.For the Meyersson data, IMSE selects 11 below and 7 above the cutoff for evenly-spaced bins, versus 21 and 14 for quantile-spaced bins; mimicking variance selects 40 and 75 evenly-spaced bins.

4 Continuity-based RD Approach · 4 The Continuity-Based Approach to RD Analysis

The continuity-based RD framework estimates the cutoff treatment effect by locally approximating regression functions with polynomial methods under continuity assumptions. Its implementation requires choices about localization, polynomial order, kernel, and bandwidth, with practical examples illustrating sensitivity to these choices.

  • 4 The Continuity-Based Approach to RD Analysis: The continuity-based RD framework defines τSRD as the parameter of interest and estimates it using local polynomial approximations based on continuity and differentiability assumptions.These methods also support falsification and validation of the RD design.
  • 4.1 Local Polynomial Approach: Overview: Because continuous running variables generally provide no observations exactly at the cutoff, RD estimation necessarily relies on local extrapolation.The approach estimates average control and treatment responses at the cutoff from nearby observations.
  • 4.1 Local Polynomial Approach: Overview: Global fourth- or fifth-order polynomial fits can produce unreliable boundary estimates and misleading conclusions, so they are discouraged for formal RD analysis.Local methods instead use low-order polynomials near the cutoff, improving robustness to boundary and overfitting problems.
  • 4.2 Local Polynomial Point Estimation: Local polynomial estimation fits separate weighted regressions on observations within [c − h, c + h], using a chosen polynomial order p, kernel K(·), and bandwidth h.The estimated intercepts on the two sides identify the cutoff responses, and their difference estimates τSRD.
  • 4.2.2 Bandwidth Selection and Implementation: The triangular kernel is recommended because, with an MSE-optimal bandwidth, it yields an estimator with optimal properties while assigning greater weight to observations nearer the cutoff.The bandwidth determines the neighborhood and balances approximation bias against variance; smaller bandwidths reduce bias but increase variance.
  • 4.2.1 Choice of Kernel Function and Polynomial Order: Researchers generally prefer local linear approximations because higher polynomial orders increase variability and can overfit near the boundary, while constant fits have undesirable boundary properties.An appropriately selected bandwidth can make the linear approximation reliable.
  • 4.2.3 Optimal Point Estimation: Given p and K(·), selecting a common or separate MSE-optimal bandwidth produces a consistent and MSE-optimal RD point estimator.The MSE criterion balances squared bias and variance through a data-driven choice of h.
  • 4.2.4 Point Estimation in Practice: 2.927 percentage points is the estimated increase in high-school completion for women aged 15 to 20 when the Islamic party barely won rather than barely lost.With fixed h and p, changing from uniform to triangular weights changes the estimate only slightly, from about 2.9271 to 2.9373; changing p from 1 to 2 changes it from 2.937 to 2.649.

4.3 Local Polynomial Inference

Valid inference in local polynomial RD analysis must account for approximation bias and misspecification rather than treating the regression as correctly specified. Two approaches are proposed: robustly modify inference at hMSE or use a separate inference bandwidth.

  • Inference and misspecification: OLS inference is inappropriate because it treats the local polynomial regression as parametric and disregards its non-parametric approximation nature.Selecting bandwidths through a bias-variance trade-off while assuming zero bias is methodologically incoherent.
  • Inference and misspecification: Valid inference must account for misspecification because hMSE, hMSE,−, and hMSE,+ are not small enough to remove the leading bias term.These MSE-optimal bandwidths produce estimators that are consistent and MSE-optimal, but create problems for standard distributional approximations.
  • Inference approaches: Two approaches address the problem: modify the t-statistic at hMSE or use hMSE for point estimation and select a different bandwidth for inference.The modified t-statistic accounts for large-bandwidth misspecification and the additional sampling error introduced by the modification.

4.3.1 Using the MSE-Optimal Bandwidth for Inference

Inference using the MSE-optimal bandwidth must account for the local polynomial estimator’s asymptotic bias, which reflects non-parametric misspecification. The authors recommend robust bias correction as a theoretically valid and practically effective approach.

  • Using the MSE-Optimal Bandwidth for Inference: The local polynomial RD estimator’s large-sample distribution includes asymptotic bias and variance when using hMSE or a data-driven implementation.The variance can be estimated in weighted least-squares settings accounting for heteroskedasticity and clustered data, with formulas implemented in rdrobust.
  • Using the MSE-Optimal Bandwidth for Inference: 95% confidence intervals for τSRD depend on the unknown bias term B.Ignoring B produces incorrect inference unless the bias is negligible, such as when the local linear model is close to correctly specified.
  • Using the MSE-Optimal Bandwidth for Inference: The bias arises because local polynomial regression provides a non-parametric approximation rather than assuming the underlying regression functions are pth-order polynomials.This distinguishes the approach from OLS estimation, which would impose polynomial structure directly.
  • Using the MSE-Optimal Bandwidth for Inference: The authors recommend robust bias correction because it is theoretically valid, has optimality properties, and performs well in practice.The recommendation is presented as a response to misspecification bias in asymptotic inference procedures.

Conventional Inference and Undersmoothing

Conventional confidence intervals ignore misspecification bias and are generally invalid unless approximation error is negligible, while undersmoothing offers a theoretically sound but ad hoc alternative. The undersmoothing bandwidth lacks clear selection criteria, reducing transparency and encouraging specification searching.

  • Conventional Inference: An MSE-optimal bandwidth cannot be selected while assuming zero misspecification error, and standard OLS inference is invalid when that bandwidth is used.Ignoring bias with an MSE-optimal bandwidth is described as both invalid and methodologically incoherent.
  • Conventional Inference: Conventional inference ignores misspecification bias and is invalid unless approximation error is negligible.It treats the local polynomial as parametric within the cutoff neighborhood and de facto ignores the bias term.
  • Conventional Inference: Conventional confidence intervals assume the chosen polynomial exactly approximates both conditional outcome functions, an unverifiable and rarely credible assumption.Using CIus when approximation error is non-negligible undermines inference.
  • Undersmoothing: Undersmoothing constructs conventional confidence intervals using a bandwidth smaller than the MSE-optimal choice used for the point estimator.The procedure first selects the MSE-optimal bandwidth, then shrinks it before constructing CIus.
  • Undersmoothing: Undersmoothing has no clear, transparent rule for shrinking the bandwidth below its MSE-optimal value, making choices ad hoc and potentially encouraging specification searching.Researchers may divide the MSE-optimal bandwidth by two, by three, or use another arbitrary reduction.

Standard Bias Correction

Standard bias correction permits inference at the MSE-optimal bandwidth by estimating and removing the induced bias. Its confidence intervals accommodate wider bandwidth choices but can perform poorly because they ignore variability from bias estimation.

  • Standard Bias Correction: Bias correction estimates the bias term B and removes it from the distributional approximation, allowing inference with the MSE-optimal bandwidth.The bias estimate is already computed for MSE-optimal bandwidth selection.
  • Standard Bias Correction: The RD estimate uses bandwidth h, whereas the bias estimate uses an additional bandwidth b to estimate regression-function derivatives of order p+1 or higher.The ratio ρ = h/b relates the variability of the bias estimate to the point estimate.
  • Standard Bias Correction: Bias-corrected confidence intervals allow a wider range of h and valid inference at the MSE-optimal bandwidth, but typically perform poorly because their variance ignores bias-estimation variability.CIbc uses the same variance as CIus despite including the estimated bias term B.

Robust Bias Correction

Robust bias-corrected confidence intervals provide improved finite-sample coverage and shorter average length than CIus and CIbc. They remain valid with MSE-optimal bandwidths, require no undersmoothing, and permit the same observations for point estimation and inference.

  • Robust Bias Correction: Robust bias correction yields smaller coverage error and shorter average length than CIus or CIbc.The approach is described as theoretically sound and delivering demonstrably superior inference procedures in finite samples.
  • Robust Bias Correction: Robust bias-corrected intervals remain valid with the MSE-optimal point-estimation bandwidth, so no undersmoothing is necessary.Validity also holds when ρ = h/b = 1, meaning h = b.
  • Robust Bias Correction: CIrbc subtracts the estimated bias from the local polynomial estimator and uses a new variance formula that incorporates bias-correction variability.Unlike CIbc, the bias estimate may converge to a random variable and contribute to the distributional approximation.
  • Robust Bias Correction: CIrbc is both recentered and rescaled relative to the conventional interval, while CIus is invalid when h = hMSE.CIrbc centers on the bias-corrected estimate and uses a larger variance when the same bandwidth h is used.
  • Robust Bias Correction: Using CIrbc with hMSE allows the same observations with score Xi ∈[c −hMSE, c + hMSE] for optimal point estimation and valid inference.This is identified as the interval’s most important practical feature.

4.3.2 Using Different Bandwidths for Point Estimation and Inference

The section recommends separating bandwidth choices for point estimation and inference in RD designs. Practitioners should use hMSE for estimating τSRD and choose either hMSE or the CER-optimal hCER for robust bias-corrected confidence intervals.

  • Bandwidth choice: Decoupling point estimation and inference uses hMSE to estimate τSRD and hCER to construct confidence intervals that minimize coverage-error approximations.The coverage-error criterion measures the discrepancy between empirical coverage and the nominal confidence level.
  • Bandwidth choice: Using hCER for point estimation remains consistent but is not MSE-optimal because its estimator has too much variability relative to its bias.The recommendation is therefore to retain hMSE for point estimation of τSRD.
  • Inference: For robust bias-corrected inference, CIrbc is valid with hMSE and valid and CER-optimal with hCER.Practitioners may use either bandwidth for constructing CIrbc, depending on whether CER optimality is desired.

4.3.3 RD Local Polynomial Inference in Practice

The section demonstrates how `rdrobust` reports conventional and robust bias-corrected inference for local linear RD estimates, including their centers, standard errors, confidence intervals, and bandwidth choices. It also contrasts MSE-optimal and CER-optimal bandwidths and explains the corresponding changes in estimates and intervals.

  • Robust bias-corrected inference: The robust bias-corrected 95% confidence interval is [-0.309, 6.276], includes zero, and is not centered at the conventional estimate.Robust inference uses a bias-corrected center and a robust standard error rather than the conventional interval’s center and standard error.
  • Confidence-interval construction: For a fixed common bandwidth, the robust bias-corrected interval is always longer than the conventional interval, but this need not hold when the intervals use different bandwidths.The comparison concerns interval lengths under a common bandwidth; the passage explicitly limits the generalization when bandwidths differ.
  • CER-optimal bandwidth inference: Using the CER-optimal bandwidth changes the common bandwidth from 17.239 to 11.629, the RD estimate from 3.020 to 2.430, and the robust interval from [-0.309, 6.276] to [-1.158, 5.979].The CER-optimal results have a conventional p-value of 0.149 and a robust p-value of 0.186.
  • Bandwidth selection: The `rdbwselect` output distinguishes MSE-optimal bandwidths, which minimize the RD estimator’s MSE, from CER bandwidths, which optimize confidence-interval coverage error rates.It reports common-bandwidth and two-bandwidth variants, along with combinations of the MSE-optimal choices.

4.4.1 Adding Covariates to the Analysis

Covariate adjustment in RD designs should use predetermined covariates and a linear, additive-separable specification without treatment interactions. Under covariate balance, this approach estimates the standard RD treatment effect, whereas posttreatment or imbalanced covariates change the parameter being estimated.

  • 4.4.1 Adding Covariates to the Analysis: Predetermined covariates can augment RD analysis through conditioning or subsetting, or through partialling out with local polynomial methods.Conditioning is most natural for a few discrete covariates, while partialling out accommodates discrete or continuous covariates and can improve efficiency.
  • 4.4.1 Adding Covariates to the Analysis: The recommended specification adds predetermined covariates linearly and additively to a fully interacted local polynomial regression without interacting covariates with treatment.The resulting estimator captures the outcome jump at the cutoff after partialling out covariate effects and reduces to standard RD without covariates.
  • 4.4.1 Adding Covariates to the Analysis: Under mild regularity conditions, equal covariate means under treatment and control at the cutoff is sufficient for the adjusted estimator to consistently estimate τSRD.This zero treatment effect on covariates is analogous to covariate balance in randomized experiments and holds naturally for truly predetermined covariates.
  • 4.4.1 Adding Covariates to the Analysis: The covariate-adjusted estimator estimates the standard RD treatment effect when covariates are included linearly, additive-separably, and without treatment interactions.Interacting covariates with treatment means zero RD treatment effects on covariates are no longer sufficient for consistency.
  • 4.4.1 Adding Covariates to the Analysis: To estimate τSRD, adjustment should include only predetermined covariates, because posttreatment or imbalanced covariates change the parameter being estimated.Imbalanced covariates cannot generally be included to repair an RD design where discontinuities challenge the required continuity assumptions.

Practical Implementation of Covariate-Adjusted RD Analysis

Covariate-adjusted RD analysis requires bandwidth selection that accounts for the included covariates. In the Meyersson application, adjustment left the estimated effect roughly unchanged while shortening the confidence interval and lowering the robust p-value.

  • Implementation: Covariate-adjusted local polynomial RD requires choosing polynomial order, kernel, and bandwidth, with bandwidth selection accounting for the included covariates.The recommended approach uses optimal data-driven MSE- or CER-optimal bandwidth methods adapted to covariate adjustment.
  • Application: The Meyersson application uses seven predetermined 1994-election and geographic covariates, excluding i89 because of many missing values.The covariates are vshr islam1994, partycount, lpop1994, merkezi, merkezp, subbuyuk, and buyuk.
  • Bandwidth selection: 14.409 is the MSE-optimal bandwidth with covariates, versus 17.239 without adjustment, showing that covariate adjustment can change optimal bandwidths and point estimates.The implementation uses a first-order polynomial, triangular kernel, common bandwidth on both sides, and the mserd option.
  • Estimation and inference: 3.108 is the covariate-adjusted RD estimate versus 3.020 unadjusted, while the robust p-value falls from 0.076 to 0.037.The confidence-interval length decreases from 6.585 to 5.938, a 9.82% reduction.

4.4.2 Clustering the Standard Errors

Cluster-robust variance estimators accommodate within-group error correlation in local-polynomial RD designs, but they also alter bandwidth selection and inference. The Meyersson application shows how clustering, and then covariate adjustment, change estimates and confidence intervals.

  • Clustering the Standard Errors: Cluster-robust variance estimators are appropriate when observations are grouped and errors may be correlated within, but not across, groups.Examples include individuals within households, municipalities within counties, and households within villages.
  • Clustering the Standard Errors: Cluster-robust variance estimators change CER- and MSE-optimal bandwidth selection in local-polynomial RD designs.The bandwidth selectors depend on the variance estimators.
  • Clustering the Standard Errors: 2.969 percentage points is the cluster-robust point estimator, versus 3.020 without clustering, with bandwidths of 19.035 and 17.239, respectively.The cluster-robust estimator uses the larger MSE-optimal bandwidth.
  • Clustering the Standard Errors: 3.146 is the point estimate after adding covariates and cluster-robust variance estimation, with a different optimal bandwidth of 15.675 and confidence interval [0.243,6.154].The application clusters municipality observations by province.
  • Clustering the Standard Errors: The cluster-robust covariate-adjusted 95% confidence interval excludes zero because covariate adjustment shortens and shifts it rightward relative to the unadjusted, unclustered case.The reported interval is [0.243,6.154].

5 Validation and Falsification of the RD Design

RD validation combines institutional and empirical checks because the assignment rule alone does not guarantee the identifying assumptions. The five empirical tests examine covariate continuity, score-density continuity, placebo cutoffs, exclusion of observations near the cutoff, and bandwidth sensitivity.

  • Motivation and validation tests: Institutional review should assess appeals and potential score manipulation, including informal mechanisms that may invalidate continuity assumptions.Known cutoffs can encourage units to change scores, while the absence of formal appeal mechanisms does not rule out informal manipulation.
  • Predetermined covariates: 0.999 is the robust p-value for the population-size covariate, providing no evidence of a systematic treated-control difference at the cutoff.The point estimate is very close to zero, and the same estimation and inference procedure should be repeated for all important covariates.
  • Predetermined covariates: All predetermined-covariate estimates are small, with 95% robust confidence intervals containing zero and p-values ranging from 0.333 to 0.999.The results provide no empirical evidence that the covariates are discontinuous at the cutoff.
  • Score-density continuity: 0.6173 is the p-value from a simple sorting test, finding no evidence that treated and control observations are unusually concentrated around the cutoff.The observed counts are consistent with assignment by the flip of an unbiased coin.
  • Score-density continuity: −1.394 is the continuity-based density statistic, with p-value 0.1633 and no rejection of equal treated-control density at the cutoff.The accompanying graphical analysis displays the data histogram, density estimate, and shaded 95% confidence intervals.
  • Placebo cutoffs: Placebo-cutoff estimates generally support the design: all artificial-cutoff p-values exceed 0.4, and all but one absolute estimates are below the true-cutoff estimate of 3.020.At the artificial 1% cutoff, the robust p-value is 0.787, indicating no outcome jump.
  • Excluding observations near the cutoff: Excluding observations with |X_i| < 0.3 and repeating the exercise for different exclusion amounts leaves the conclusions unchanged.After exclusion, the analysis uses 2307 observations left of the cutoff and 309 observations right of it.
  • Bandwidth sensitivity: Bandwidth sensitivity is limited: h_CER = 11.629 and h_MSE = 17.239 yield similar point estimates, while h_CER = 11.629 produces a longer confidence interval that includes zero.The doubled bandwidths, 2·h_CER = 23.258 and 2·h_MSE = 34.478, produce broadly consistent findings.

6 Final Remarks

The section concludes that this Element develops foundational methods for identification, estimation, inference, and falsification in the canonical Sharp RD design. The accompanying Element extends that framework to departures from the canonical setting and together they offer a practical template for principled, rigorous, transparent analysis.

  • Foundations: The Element covers identification, estimation, inference, and falsification for the average treatment effect at the cutoff in the Sharp RD design.It focuses on the simplest case: one running variable, one cutoff, and perfect compliance with treatment assignment.
  • Extensions: The accompanying Element considers departures from the canonical Sharp RD design, including a local-randomization interpretation based on an as-if random treatment window around the cutoff.This contrasts with the continuity-based approach adopted in the present Element.
  • Practical guidance: Together, the Elements aim to provide applied researchers with a practical template for analyzing and interpreting RD designs in a principled, rigorous, and transparent way.The template combines the discussion in this Element with the additional methods in its companion.
Loading 1911.09511v1…