EconBase
← Back to paper

Regression discontinuity aggregation, with an application to the union effects on inequality

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

127,939 characters · 21 sections · 127 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Regression Discontinuity Aggregation, with an Application to the Union Effects on Inequality

abstractWe extend the regression discontinuity (RD) design to settings where each unit's treatment status is an average or aggregate across multiple discontinuity events. Such situations arise in many studies where the outcome is measured at a higher level of spatial or temporal aggregation (e.g., by state with district-level discontinuities) or when spillovers from discontinuity events are of interest. We propose two novel estimation procedures — one at the level at which the outcome is measured and the other in the sample of discontinuities — and show that both identify a local average causal effect under continuity assumptions similar to those of standard RD designs. We apply these ideas to study the effect of unionization on inequality in the United States. Using credible variation from close unionization elections at the establishment level, we show that a higher rate of newly unionized workers in a state-by-industry cell reduces wage inequality within the cell.

\global\long\def\expec#1{\mathbb{E}\left[#1\right]} \global\long\def\var#1{\mathrm{Var}\left[#1\right]} \global\long\def\cov#1{\mathrm{Cov}\left[#1\right]} \global\long\global\long\global\long\global\long\global\long\def\one#1{1\left[#1\right]} \global\long

Introduction

Regression discontinuity (RD) designs are a popular tool for causal inference. In canonical sharp RD designs, the treatment is determined by whether a “running variable” exceeds a fixed cutoff. For instance, one may study whether unionization affects establishment outcomes, leveraging the cutoff of 50% votes in a union election dinardo2004economic,frandsen_2021. A local causal parameter is typically estimated via local linear regression under the assumption that potential outcomes are continuous around the cutoff. For standard RD methods to apply, each unit has to be exposed to only one discontinuity.

In this paper, we extend the RD design to settings in which a unit is exposed to multiple discontinuity events because the outcome is defined at a higher level of aggregation (e.g., in space or time) than the elections.\footnote{We use the term “elections” broadly to refer to events that generate RD variation, as the majority of RD papers we cite are indeed about elections of some sort.} The endogenous explanatory variable (“treatment”) in such settings is typically a simple or weighted average or sum of the RD shocks across multiple elections. We propose two new estimators in these “RD aggregation” (RDA) settings. The “upper-level” estimator is based on an instrumental variable (IV) regression at the level at which the outcomes are measured, with a weighted average or sum of the RD shocks from narrow elections as the instrument, while controlling for similar weighted sums of the controls used in local linear estimation (e.g., of the running variables). The “lower-level,” or “stacking,” estimator is a fuzzy RD estimator in the sample of narrow elections, with the unusual feature that the same values of the outcome and endogenous variable are repeated for the elections affecting the same unit. We show that, under assumptions similar to those of standard RD, both estimators identify the same local average treatment effect. Moreover, Monte Carlo simulations confirm that these estimators inherit attractive bias-reduction properties of conventional RD methods. The setting we consider nests many empirical literatures, and our ideas yield practical insights for them. We then apply the proposed methods to estimate the effects of unionization on earnings inequality at the level of state-by-industry cells, with typical cells exposed to multiple establishment-level unionization events.

The primary use cases for the RDA approach are empirical settings in which RD events are aggregated across space, time, or in other ways to construct a treatment variable at the level at which the outcome is measured. Such settings are very common. First, a large group of papers considers legislatures that consist of politicians elected in single-seat constituencies. The researcher is interested in the impact of the seat share of politicians who possess a specific characteristic: e.g., genderclots2011women,clots2012female,bhalotra2014health,ben-porath2023, religion bhalotra2014religion,bhalotra2021religion, or affiliation with a specific party folke2014shades,nellis2016parties,nellis2018secular,ben-porath2023. For instance, a state-level outcome is regressed on the overall seat share of women; the seat share of women who won close races is used as an IV. Second, in many cases a unit can be exposed to multiple RD events over time (cellini2010value,dell2018nation,biasi2024works). In the RDA approach, the researcher can specify the treatment as the total fraction of the period during which the unit is treated. Other studies aggregate discontinuities across investors in a region, across university courses taken by a student, across politicians that a firm is linked to, across destinations that a city is connected to by a direct flight, and across unionization elections to which we return below.\footnote{Specifically, ater2021real construct a regional housing supply shock by aggregating across real estate investors. They leverage a discontinuity in the capital gains tax law that pushes some investors to sell their property. tan2023consequences considers the effects of the number of high letter grades of a university student, as aggregating across individual courses. akey2015valuing studies firms that donate to multiple politicians and measures the effects of a firm's political connectedness, as aggregating election outcomes of those politicians. campante2018long aggregate discontinuities in flight availability between city pairs with respect to distance to construct a shock to a city's overall connectivity. Kolerman2023 and the application in this paper aggregate establishment-level unionization data from close elections to estimate the effects of unionization on political outcomes and inequality, respectively.}

Our results also extend naturally to estimating spillover effects from discontinuity events, such that each unit is exposed to multiple elections of their neighbors (in a spatial or network sense). isen2014local, for instance, leverages close referenda on tax changes in Ohio counties to study the effects on neighboring counties. Similarly, dechezlepretre2023tax study the effects of tax breaks, which have discontinuous eligibility rules, on technologically similar firms.

While the two estimators we propose are natural extensions of conventional RD to RDA settings and they possess attractive statistical properties, they differ from most ad hoc estimation approaches used in the existing empirical literature. Specifically, papers estimating upper-level specifications do not include all local linear controls aggregated in the way suggested by our upper-level solution. As a result, they generally forfeit the reduced bias advantages of local linear estimation that made it standard in conventional RD designs. Conversely, studies using lower-level specifications include local linear controls but, in many cases, use reduced-form estimation which does not define a coherent causal model for the upper-level outcome. Some of these papers also impose restrictions limiting the sample that is not necessary with the RDA approach. Finally, some papers resort to difference-in-differences and other non-RD methods in settings where RD would be a standard tool for credible identification, if not for the aggregated outcomes. The RDA approach makes RD identification feasible in those settings.

We apply the RDA approach to study the union effects on earnings inequality. This question is one of the oldest in labor economics, dating back to at least lewis1963unionism. It is also of central policy importance in modern days: Barack Obama, for instance, noted that “as union membership has fallen, inequality has risen” Obama2016— a correlation observed in the data by card2001effect. Yet, establishing a causal link between unionization and inequality has proved difficult. Recently, fortin2022right point out that “empirical evidence on the distributional impact of {[}union{]} threat effects is limited by the challenge of finding exogenous sources of variation in the rate of unionization.” Similarly, farber2021unions for the most part use a selection-on-observables designs. We make progress by leveraging close union elections in an RDA design.

We find that a higher rate of unionization leads to a decline in various measures of inequality at the level of state-by-industry cells: the Gini index, the 90/10 ratio, the share of labor income by top 10% earners, and the variance in log wages. These impacts result primarily from the decline of wages at the top of the wage distribution rather than an increase at the bottom. While this result calls for further investigation, we find evidence for one mechanism that can help explain it: pension coverage increases in cells with higher unionization rates. In a back-of-the-envelope calculation, we find that around 35% of the growth in the key within-cell inequality measures between 1970 and 2010 can be explained by the declining rates of unionization relative to the 1960s. Our estimate of the union contribution is moderately larger than the most comparable estimates from farber2021unions.

We believe our empirical strategy for estimating the impacts of unionization on aggregated outcomes can be helpful beyond the effects on inequality. Some papers evaluated such effects by comparing neighboring counties across the border of right-to-work laws, which were mostly passed in the 1940s and 1950s and weakened unions: holmes1998effect, feigenbaum2019bargaining, and makridis2019right studied the impacts on productivity, manufacturing activity, and voting, respectively. Other papers leveraged specific case studies that generate exogenous variation in union strength when looking at outcomes such as wages biasi2021labor and the gender pay gap biasi2022flexible. In contrast, the RDA method allows researchers to leverage credible RD variation from multiple decades and the whole common, which is standard when studying the impacts on establishment-level outcomes. Complementary work by Kolerman2023 has used the RDA method to show that unionization makes local politics more Democratic in the US.

Methodologically, our paper bridges the literatures on shift-share instruments and RD designs. Our theoretical analysis rests on the observation that both the treatment and the instrument in the RDA setting are shift-share aggregates (i.e., weighted sums) of RD shocks. We then leverage the results from borusyak2022quasi to show how upper-level IV regressions with the aggregated instrument and controls are numerically equivalent to particular fuzzy RD regressions at the lower level. This equivalence is useful in two ways: on one hand, it brings upper-level regressions closer to standard fuzzy RD and motivates our choice of upper-level controls as aggregations of the standard controls from local linear estimation. On the other hand, it motivates our lower-level estimator by showing that an appropriate lower-level regression can identify a parameter at the upper level, at which the causal effect is naturally defined.

The setting we study is distinct from “multi-score” RD settings which also feature multiple running variables per unit (see cattaneo2022regression for a review). For instance, in spatial RD designs black1999better,dell2010persistent, the latitude and longitude are the two running variables, and a geographic boundary separates the treated and control areas. Another example is when a student is treated if several tests scores (e.g., for math and English) simultaneously exceed some cutoffs. Such designs feature a single binary treatment (or instrument, in the fuzzy case), and a simple transformation of the running variables reduces the problem to a standard RD, in which the signed distance to the boundary is the scalar running variable. In contrast, in our setting the instrument aggregates RD shocks. While still a function of the underlying running variables, it does not feature a single boundary, making existing methods inapplicable.

A recent literature has examined certain RD settings with interference. torrione2024regression extend the multi-score approach, based on the minimum Euclidian distance from the set of neighbors' running variables to the boundary at which the spillover treatment, e.g. the number or share of treated neighbors, switches between two prespecified values. This allows them to identify more detailed spillover effects, e.g. of having all treated neighbors vs. all untreated neighbors conditionally on the unit itself being treated. We propose a very different estimation approach which identifies a causal spillover parameter in a simpler way, does not depend on the discreteness of the treatment, and applies to a variety of settings beyond just spillovers. Other papers have examined the case of interference occurring among units with similar values of the running variable, as may be natural in spatial RD designs (keele2018geographic, auerbach2024regression). In our setting, even if the structure of spillovers is based on geography, each unit has its own running variable, leading to different issues and solutions.

While our RDA design considers treatments and instruments that are shift-share aggregates of RD shocks, a recent literature has considered more complex instruments constructed from RD variation, via the idea of recentering. As borusyak2023nonrandom pointed out, if within some bandwidth around the cutoff the RD shocks are viewed as completely random (as in the local randomization approach to RD; see cattaneo2015randomization), one can compute the expectation of the instrument over the distribution of RD shocks and control for it to isolate the exogenous variation. kott2022income and narita2021algorithm compute the expected instrument by simulating counterfactual realizations of RD shocks, while abdul2022breaking and angrist2024credible compute it analytically. Our paper focuses on the settings where the instrument is linear in the RD shocks — which still covers a wide variety of empirical applications. Thanks to this focus, we are able to leverage the local linear estimation approach to RD and develop estimators with better statistical properties than those based on local randomization.

The rest of this paper is organized as follows: in Section 2 we present our theoretical framework and results in the context of cross-sectional RD aggregation, as well as Monte Carlo simulations for the proposed estimators. In Section 3, we discuss how our proposals differ from the current empirical practice. In Section 4, we extend the analysis to settings with spillovers and temporal aggregation. In Section 5, we apply the RDA approach to estimate the impact of unions on US inequality. Section 6 concludes.

Theoretical Framework

In Section (ref), we introduce the setting and compare it with conventional RD designs. Sections (ref) and (ref) propose two estimators that we put forward and provide intuition why they may provide valid causal estimates. Section (ref) formally shows how, under conditions similar to those of standard fuzzy RD, these estimators identify a local average causal effect, and Section (ref) establishes their attractive bias-reduction properties in a Monte Carlo simulation.

Setting

We measure outcome $Y_{i}$ for a set of units $i=1,\dots,N$, which we refer to as the “upper level.” Each unit consists of a set $\mathcal{J}_{i}$ of “lower-level” subunits $j$ (with $\left|\mathcal{J}_{i}\right|=J_{i}>0$). A subunit is characterized by a running variable $r_{j}$ and treatment $z_{j}=\one{r_{j}\ge0}$ that happens whenever the running variable exceeds the cutoff normalized to zero.\footnote{We use capital and small letters to denote unit- and subunit-level variables, respectively.} We are interested in estimating the causal model

equation[equation omitted — 73 chars of source]

where $\varepsilon_{i}$ is the untreated potential outcome and $\beta$ is the causal effect (assumed homogeneous for now) of some upper-level treatment $X_{i}$. Our benchmark case is when $X_{i}$ simply aggregates $z_{j}$ across all relevant subunits,

equation[equation omitted — 82 chars of source]

but other treatments can be considered, too (see, e.g., Section (ref) for an example). Here $s_{j}>0$ are known importance weights which may or may not be the same for all $j\in\mathcal{J}_{i}$ and may or may not add up to one.\footnote{We assume that the sets are non-overlapping to simplify notation. The setting naturally extends to the case where different units can be exposed to the same subunits, as when estimating the spillovers of RD events; see Section (ref).} In the election example, $X_{i}$ may be the fraction of seats in the state $i$ legislature represented by candidates from the focal party. Here $z_{j}$ indicates the focal party's victory in district $j$ and $s_{j}$ measures the seat share of the district $j$ in the legislature. In the common case of one seat per district, $s_{j}$ is the inverse of the total number of seats.

The conventional sharp RD setting corresponds to the case where the outcomes are measured at the same level as the running variables, i.e. $i$ and $j$ are the same objects: $\mathcal{J}_{i}=\left\{ i\right\} $, $s_{j}=1$, and $X_{i}=z_{i}$. In that scenario, the standard local linear estimation approach restricts the sample to observations near the cutoff ($\left|r_{j}\right|\le h$ for bandwidth $h$) and estimates by ordinary least squares (OLS) the specification

equation[equation omitted — 133 chars of source]

where the controls $q_{j}=(1,r_{j},r_{j}^{+})$ include the intercept, the running variable, and its interaction with being on the right of the cutoff (i.e., $r_{j}^{+}=r_{j}\cdot\one{r_{j}\ge0}$). Sometimes this regression is weighted by “kernel” weights that are larger when $r_{j}$ is closer to the cutoff; importance weights are also allowed.

But how can local variation in $r_{j}$ be leveraged to obtain a consistent estimate of $\beta$ in ((ref)) when each unit can be exposed to multiple subunits, potentially with more than one near the cutoff? We next propose two strategies, one at the upper level and the other at the lower level, focusing on the intuitions before proceeding to formal and simulation-based analysis of consistency and bias.

\protectAn Upper-Level Solution

Our first proposal, at the level of upper-level of observations $i$, involves instrumenting $X_{i}$ with its component arising from the subunits whose running variables are near the cutoff:

equation[equation omitted — 71 chars of source]

where $\mathcal{C}_{i}=\left\{ j\in\mathcal{J}_{i}\colon\ \one{\left|r_{j}\right|\le h}\right\} $ is the set of $i$'s subunits near the cutoff. In the election example, $Z_{i}$ is the fraction of state legislature's seats represented by a candidate from the focal party who won in a close election, out of all seats.

We further propose to aggregate the vector of standard RD controls $q_{j}$ to the $i$ level in the same way as the instrument is constructed,

equation[equation omitted — 204 chars of source]

We label these three aggregated variables the RDA controls. The first of them is the total weight of subunits near the cutoff. For instance, in the election example, it is the fraction of legislature seats determined in a close election. The second is the average running variable in narrow elections, rescaled by their total weight. And the last one is the average running variable in narrow wins, rescaled by the total weight of those narrow wins.

As a result, the upper-level RDA estimator corresponds to the instrumental variable (IV) specification

equation[equation omitted — 238 chars of source]

instrumented by $Z_{i}$, where $\tilde{W}_{i}$ includes the $i$-level intercept and possibly additional variables that are continuous around the cutoff. Its reduced-form parallels ((ref)):

equation[equation omitted — 264 chars of source]

While aggregating both the instrument $z_{j}$ and the covariates $q_{j}$ in the same way is intuitive, why do we expect RDA to identify a causal parameter? We provide intuition by showing that the IV estimator from ((ref)) is numerically equivalent to a standard fuzzy RD estimator from a transformed regression, and can therefore be expected to inherit some its properties.

Indeed, $Z_{i}$ can be viewed as a shift-share instrument constructed from shifts $z_{j}$ and exposure shares $s_{j}$.\footnote{More precisely, the shares are $s_{j}$ for $j\in\mathcal{C}_{i}$ and zero for $j\not\in\mathcal{C}_{i}$. While shift-share instruments usually involves shifts that are common to all observations, this is not a requirement borusyak2024practical.} We can therefore leverage the numerical equivalence result of borusyak2022quasi to represent $\hat{\beta}$ as an IV estimator at the subunit level:

propFor any unit-level variable $V_{i}$ let $V_{i}^{\perp}$ denote the residual from a projection of $V_{i}$ on $W_{i}=\left(Q_{i},\tilde{W}_{i}\right)'$. Let $\textbf{i}(j)$ denote the unit to which subunit $j$ belongs. Then $\hat{\beta}$ from ((ref)) equals the IV estimator of $\beta$ from a subunit-level specification \begin{equation} Y_{i(j)}^{\perp}=\beta X_{i(j)}^{\perp}+\lambda'q_{j}+error_{j}, \end{equation} on the sample of $j\in\mathcal{C}\equiv\cup_{i}\mathcal{C}_{i}$, where $X_{\textbf{i}(j)}^{\perp}$ is instrumented by $z_{j}$ and the regression is weighted by $s_{j}$.
proofFollows from Proposition 1 in borusyak2022quasi by observing that the shock-level weight defined in that paper, $\sum_{i}\sum_{j\in\mathcal{C}}s_{j}\one{j\in\mathcal{C}_{i}}$, simplifies to $s_{j}$, and that shock-level averages simplify to $\bar{V}_{j}=\sum_{i}s_{j}\one{j\in\mathcal{C}_{i}}V_{i}/\sum_{i}s_{j}\one{j\in\mathcal{C}_{i}}=V_{\textbf{i}(j)}$.

The specification ((ref)) is artificial but it is useful because it maps directly to the standard fuzzy RD design: the instrument is the subunit-level RD shock $z_{j}$ and the covariates are $q_{j}=(1,r_{j},r_{j}^{+})$.\footnote{While in many fuzzy RD designs the treatment is binary, this is not necessary almond2010estimating.} One can use the equivalent $j$-level representation of the upper-level specification to choose bandwidth and to perform statistical inference on the estimator, as in calonico2014robust

\protectRD Stacking: A Lower-Level Solution

Our second proposal is to estimate a fuzzy RD specification on the sample that stacks all subunits near the cutoff:

equation[equation omitted — 193 chars of source]

with $X_{\textbf{i}(j)}$ instrumented by the subunit-level treatment $z_{j}$ and using $s_{j}$ as importance weights. Kernel weights may be additionally used, too.\footnote{Kernel weights should not be used when constructing $Z_{i}$ for the upper-level specification. Correspondingly, they were not introduced in Section (ref).} Since ((ref)) is a conventional fuzzy RD specification, it is clear it uses local variation around the cutoff.

The specification ((ref)) is peculiar: it provides multiple expressions for $Y_{i}$ of the same unit. One may also find it odd that the instrument $z_{j}$ varies among multiple observations — subunits of the same unit — that mechanically have the same value of the treatment $X_{\textbf{i}(j)}$.

Despite this peculiarity, Proposition (ref) demonstrates why specification ((ref)) may be helpful. It only differs from specification ((ref)) by having the outcome and treatment initially residualized on $W_{i}$. These controls, including $Q_{i}$, are continuous around the cutoff and therefore may be expected not to affect identification. Moreover, since local linear controls $q_{j}$ enter both specifications, one may hope that both approaches inherit the bias properties of standard RD. Below we show that such hope is justified.

The stacking approach has an additional advantage over the upper-level solution: since it is just a fuzzy RD specification with raw outcome and treatment, it naturally lends itself to standard RD plots for the reduced-form and first-stage of ((ref)).

\protectIdentification of a Local Causal Effect

To describe the local parameter identified by RDA, we introduce the following causal model with heterogeneous causal effects. For $\boldsymbol{z}_{i}=\left(z_{j}\right)_{j\in\mathcal{J}_{i}}$, let $X_{i}(\boldsymbol{z}_{i})$ denote the potential treatment of unit $i$ depending on the treatments of its subunits. When equation ((ref)) holds, it defines $X_{i}(\boldsymbol{z}_{i})$, but we allow for general $X_{i}$. We further make an exclusion restriction that $\boldsymbol{z}_{i}$ only affects $Y_{i}$ via $X_{i}$, with $Y_{i}(x)$ denoting potential outcomes.\footnote{Here we assume for clarity that both the first and second stages of ((ref)) are causal. Both assumptions can be relaxed as in small2017instrumental and Appendix A.1 of borusyak2022quasi.} We further impose:

assumption[Sampling of units] The data are a random sample of $J_{i},\boldsymbol{r}_{i},\boldsymbol{z}_{i},\boldsymbol{s}_{i},\boldsymbol{q}_{i},X_{i},X_{i}(\cdot),\tilde{W}_{i},Y_{i},Y_{i}(\cdot)$ where bold symbols indicate $J_{i}$-dimensional vectors containing similarly named subunit-level variables, and consistency conditions hold: $z_{j}=\one{r_{j}\ge0}$, $X_{i}=X_{i}(\boldsymbol{z}_{i})$, and $Y_{i}=Y_{i}(X_{i})$.
assumption[Irrelevance of labels] For any $J$ in the support of $J_{i}$, the distribution of $\left(\pi(\boldsymbol{r}_{i}),\pi(\boldsymbol{z}_{i}),\pi(\boldsymbol{s}_{i}),\right.$\\ $\pi(\boldsymbol{q}_{i}),X_{i},X_{i}(\pi(\cdot)),\tilde{W}_{i},Y_{i},Y_{i}(\cdot))\mid J_{i}=\bar{J}$ is the same for all permutations $\pi$ of $\bar{J}$ elements.

Under these assumptions, the distribution of unit-level data induces a distribution of the relevant variables over subunits, such as the joint distribution of $z_{j}$ and $ $$\boldsymbol{z}_{\textbf{i}(j)-j}$ where $\left(z_{j},\boldsymbol{z}_{\textbf{i}(j)-j}\right)\equiv\boldsymbol{z}_{\textbf{i}(j)}$. We then impose a generalization of the RD continuity assumption:

assumption[Continuity] (a) Expected reweighted potential treatment and outcome, $\expec{s_{j}X_{\textbf{i}(j)}(z,\boldsymbol{z}_{-j})\mid r_{j}=r}$ and $\expec{s_{j}Y_{\textbf{i}(j)}(X_{\textbf{i}(j)}(z,\boldsymbol{z}_{\textbf{i}(j)-j}))\mid r_{j}=r}$ are continuous in $r$ at $r=0$ for $z=0,1$. (b) Expected importance weights $\expec{s_{j}\mid r_{j}=r}$ and the cumulative distribution function of $\boldsymbol{z}_{\textbf{i}(j)-j}\mid r_{j}=r$ are continuous in $r$ at $r=0$.

While this assumption is high-level, it will be satisfied when the distributions of raw potential outcomes $Y_{\textbf{i}(j)}(x)$, potential treatments $X_{\textbf{i}(j)}(z,\boldsymbol{z}_{\textbf{i}(j)-j})$, importance weights, and leave-out treatments $\boldsymbol{z}_{\textbf{i}(j)-j}$ are continuous around the cutoff for $r_{j}$. The last assumption is satisfied when the vector of running variables $\boldsymbol{r}_{i}$ has joint density, in particular ruling out the case where multiple running variables are perfectly correlated.\footnote{However, when multiple running variables are perfectly correlated, they can be combined into one without loss.}

assumption[Density] The density of $r_{j}$ is positive and continuous at 0.
assumption[Monotonicity and first stage] $X_{\textbf{i}(j)}(1,\boldsymbol{z}_{\textbf{i}(j),-j})\ge X_{\textbf{i}(j)}\left(0,\boldsymbol{z}_{\textbf{i}(j),-j}\right)$ a.s. and\\ $\expec{s_{j}\left(X_{\textbf{i}(j)}(1,\boldsymbol{z}_{\textbf{i}(j),-j})-X_{\textbf{i}(j)}(0,\boldsymbol{z}_{\textbf{i}(j),-j})\right)\mid r_{j}=0}>0$.

Assumption 5 is trivially satisfied when the treatment is defined by equation ((ref)). For simplicity, we abstract away from additional controls, setting $\tilde{W}_{i}=1$; calonico2019regression show that adding predetermined covariates in a linearly separable way does not affect the RD estimand. The following proposition shows that under these assumptions both of our proposals identify the same local average treatment effect which puts more weight on units with more numerous and more important subunits near the cutoff, as well as subunits with a stronger first-stage impact on the treatment.

propSuppose Assumptions 1–5 hold. Then, as $h\to0$, the estimand $\beta_{h}^{u}\equiv\cov{Y_{i}^{\perp},Z_{i}}/\cov{X_{i}^{\perp},Z_{i}}$ of ((ref)) and the similarly defined estimand $\beta_{h}^{\ell}$ of ((ref)) converge to the same convexly-weighted average of potential outcome slopes: \begin{equation} \beta_{0}\equiv\frac{\expec{s_{j}\cdot\left(Y_{i(j)}\left(X_{i(j)}(1,\boldsymbol{z}_{i(j)-j})\right)-Y_{i(j)}\left(X_{i(j)}\left(0,\boldsymbol{z}_{i(j)-j}\right)\right)\right)\mid r_{j}=0}}{\expec{s_{j}\cdot\left(X_{\textbf{i}(j)}(1,\boldsymbol{z}_{\textbf{i}(j)-j})-X_{\textbf{i}(j)}\left(0,\boldsymbol{z}_{\textbf{i}(j)-j}\right)\right)\mid r_{j}=0}}. \end{equation}
proofSee Appendix (ref).

We note that the proof makes it clear that, among the covariates, only $\sum_{j\in\mathcal{C}_{i}}s_{j}$ in the upper-level specification and the intercept in the lower-level specification are necessary for identification. A parallel to the local randomization approach to RD designs provides intuition for this conclusion: if the running variables are viewed as drawn from a uniform distribution, controlling for $\sum_{j\in\mathcal{C}_{i}}s_{j}$ is equivalent to controlling for the expected value of $Z_{i}$, as in the borusyak2023nonrandom recentering procedure. This control isolates a unit's “luck” in getting a higher-than-expected vs lower-than-expected value of the instrument. While our identification result (or the local randomization approach) do not provide a justification for including the aggregated local linear controls, the next subsection shows that they are important for reducing bias, as in standard RD designs based on the continuity assumption.

We also note that the effective weights on causal effects in ((ref)) diverge from the importance weights the regressions are estimated with. Consider the case where $X_{i}$ is defined as in equation ((ref)), such that the first-stage effect of $z_{j}$ equals $s_{j}$, and the causal effects are heterogeneous across units but linear in $x$, $Y_{i}(x)=Y_{i}(0)+\beta_{i}x$. Then the total weight placed on $\beta_{i}$ in ((ref)) is $\sum_{j\in\mathcal{C}_{i}}s_{j}^{2}$. When $s_{j}=1/J_{i}$, as when the treatment is the share of legislature seats won by certain candidates (e.g., female or Democratic), the effective weight is smaller in larger legislatures, as the law of large numbers eliminates much of the variation in the instrument. If, however, $s_{j}=1$, as when the treatment is the number of such legislature seats, the effective weight is larger in larger legislatures, as they will tend to have more individual seats with narrow elections. Importance weights in both upper- and lower-level regressions can be adjusted if other averages of causal effects are of interest.

\protectBias and Efficiency: Monte Carlo Simulations

The key advantage of the two estimators we propose is that they inherit the attractive bias and variance properties of standard RD estimators. In particular, when the bandwidth shrinks to zero, their bias shrinks at a quadratic rate or faster. We leave theoretical results establishing this property to future drafts; for now, we illustrate the performance of alternative estimators via a Monte Carlo simulation.

We report the performance of the upper-level IV estimator with specification ((ref)) and the lower-level IV estimator with specification ((ref)).\footnote{We do not include any additional controls besides the intercept, setting $\tilde{W}_{i}=1$.} We contrast them with a benchmark upper-level IV estimator which includes the control necessary for identification, $\sum_{j\in\mathcal{C}_{i}}s_{j}$, but not the other controls aimed at reducing bias. Section (ref) we show that the latter estimator is commonly used in practice, and it also aligns with the recentering approach based on the local randomization view of RD.

In each simulation, we draw 1,000 units (cells), each with 5 subunits (elections) with equal importance weights $s_{j}=1/5$. For each subunit, we generate the running variable as $r_{j}=\rho\zeta_{\textbf{i}(j)}+\sqrt{1-\rho^{2}\cdot}\xi_{j}$, where $\zeta_{i}$ and $\xi_{j}$ are i.i.d. standard normal variables and $\rho=0.5$ captures the correlation of running variables among subunits of the same unit. The treatment variable is defined as the share of treated subunits, $X_{i}=\sum_{j\in\mathcal{J}_{i}}s_{j}z_{j}$ for $z_{j}=\one{r_{j}>0}$. To focus on the bias and variance, we assume the true causal effect is zero for all observations.

The top row of Figure (ref) reports the median bias of the three estimators for a range of bandwidth values, along with the bootstrap 95% confidence interval for this median.\footnote{The range $h\in[0.25,1.25]$ corresponds to approximately $[20\%,80\%]$ of subunits included in the close-elections sample. The confidence interval is hard to see on some panels because the number of simulations is large enough for the median to be very precisely estimated.} The bias depends crucially on how potential outcomes vary with the running variables of the subunits. We therefore consider several data-generating processes for the outcome; focusing here on the median bias, we do not include an independent error. Panel (a) considers a unit-level outcome that is linear in the running variables of all subunits, $Y_{i}=\sum_{j\in\mathcal{J}_{i}}s_{j}r_{j}$. In standard RD, that would eliminate bias completely, as long as the local linear controls are included. In the RDA setup, this is not the case, as the same unit can include both close and non-close elections. Yet, the qualitative behavior is the same: the bias of our proposed upper- and lower-level estimators is negligible compared to the benchmark upper-level estimator with fewer controls (and the 95% bootstrap confidence interval includes zero for most bandwidths).\footnote{Another way to see this is that the reduced-form of our lower-level IV regression is exactly unbiased, as $\expec{Y_{i}\mid r_{j}}$ is linear in this simulation.} Panel (b) considers an outcome which depends on the running variables nonlinearly but with the same curvature on either side of the cutoff, $Y_{i}=\sum_{j\in\mathcal{J}_{i}}s_{j}r_{j}^{2}$. Here all three estimators achieve negligible bias; note the scale of the vertical axis as well as the 95% confidence interval for the bias estimate based on our 2,500 simulations. Finally, panel (c) considers the case where curvature changes around the cutoff, $Y_{i}=\sum_{j\in\mathcal{J}_{i}}s_{j}r_{j}^{2}\one{r_{j}>0}$. This is the case studied in the standard RD setup by, e.g., imbens2012optimal who show the importance of local linear controls in achieving quadratic decay of the bias. Inheriting this property, the bias of our proposed estimators is not only smaller than that for the benchmark one for all bandwidths; its slope also approaches zero for smaller $h$, which does not happen for the benchmark estimator.

figure[figure omitted — 1,567 chars of source]

In unreported simulations, we check robustness of these findings. First, we generate heterogeneous importance weights across subunits of the same unit (adding up to one). Second, while all outcomes in our main simulation depend symmetrically on all running variables, we also consider the outcome which equals the running variable of one arbitrarily picked subunit. Third, we increase the number of units from 1,000 to 2,000 or the number of subunits per unit from 5 to 10. The bias patterns are unchanged.

The results so far have shown that, at least for the data-generating processes we considered, the bias of the upper- and lower-level regressions is essentially identical. There are, however, some differences in efficiency. To assess them, we repeat the simulation for the same outcomes adding i.i.d. standard normal noise to capture the idiosyncratic component of the error. The second row of Figure (ref) reports standard deviations of the estimators across simulations. It shows that the upper-level regression is more efficient than the lower-level one, albeit the difference is small. The intuition is that aggregated controls better capture the correlation between potential outcomes and the running variables. This is not an artifact of our outcome definitions; rather, it follows from the irrelevance of labels (Assumption (ref)): to the extent the running variable of one subunit predicts the unit outcome, the appropriate average of them is at least as predictive.

To sum up, the Monte Carlo simulations confirm that including the aggregated local linear controls or estimating the effects at the lower-level in a standard fuzzy RD both succeed at substantially reducing bias. The upper-level estimator is also slightly more efficient, as it is better able to reduce residual variance.

Common Practices

Many published articles feature designs suitable for RDA. We identified over 50 such papers listed in Appendix Table (ref). In this section, we describe the empirical strategies employed by the authors of those articles. We detail the main differences between their approaches and ours, highlighting the potential gains from using the estimators from Section (ref). Section (ref) discusses studies that, in the absence of the RDA methodology, did not use RD variation at all. We then turn to papers that used some RD specifications at the aggregated or disaggregated levels in Sections (ref) and (ref), respectively. Throughout the discussion, we generically refer to the subunits as elections.

Difference-in-Differences Specifications in RDA Settings

The first group of studies, listed in Panel A of Appendix Table (ref), uses difference-in-differences and similar designs in settings where RDA estimation is feasible and could provide more credible identification. While the parallel trends assumption underlying difference-in-differences analyses is often hard to justify ex ante, continuity assumptions underlying conventional RD and inherited by RDA are usually more convincing.

besley2003political consider the effects of legislature composition — one of the prime applications of RDA. To estimate the effects on policy choices of the fractions of female and Democratic legislators in U.S. state houses, where most seats are elected in single-seat races, they employ the two-way fixed effects OLS specification with a time-varying treatment: \[ Y_{it}=\beta X_{it}+\tilde{\gamma}'\tilde{W}_{it}+FE_{i}+FE_{t}+\text{error}_{it}. \] A credible instrument for $X_{it}$ could be obtained by leveraging the structure of the treatment as aggregating discontinuity events.

Similar examples are found in difference-in-differences studies of the effects of political connections on industries and firms. akcigit2023connecting estimate how a larger fraction of politically connected firms in an industry — defined as firms whose employee has been elected in a local council — affects industry outcomes. They use OLS estimation, despite using conventional RD for firm-level outcomes. cooper2010corporate similarly consider the association between the number of politicians supported by a firm who get elected and firm performance.

Finally, saez2019payroll estimate the wage effects of a payroll tax cut that applied to workers below the age of 25. Their difference-in-differences specification compares firms with different shares of young workers, under a parallel trends assumption. The RDA approach would instead leverage the age cutoff, recognizing that a firm's share of young workers is an aggregate of worker-level discontinuities. The RDA strategy would thus entail comparing firms with workers age 25 or just above to firms with workers just below 25, providing credible RD variation.

Upper-Level Outcome and Upper-Level Specification

We now consider 19 papers from Panel B of Appendix Table (ref) where RD variation is used in some way and the regression analysis is conducted at the upper-level ($i$-level in the notation of Section (ref)). These specifications can generally be represented as

equation[equation omitted — 126 chars of source]

where $X_{i}$ is constructed from all subunits and is instrumented by $Z_{i}=\sum_{j\in\mathcal{C}_{i}}s_{j}z_{j}$ constructed from the subunits near the cutoff only. Furthermore, $\tilde{Q}_{i}$ represents controls constructed in some way from subunit-level variables and $\tilde{W}_{i}$ are other controls and fixed effects.

The key difference between our proposal and what is found in the literature is in the set of controls $\tilde{Q}_{i}$. We discuss the three RDA controls $Q_{i}$ in equation ((ref)) in turn. The majority of papers include the first control: the total weight of narrow elections.\footnote{Instead of controlling for it, some papers (e.g., folke2014shades, azoulay2019public), in the language of borusyak2023nonrandom, recenter the instrument $Z_{i}$ by subtracting 0.5 times the total weight of narrow elections.} One exception is valentim2024does who instead control for the weight of narrow elections below the cutoff; this produces a biased coefficient.\footnote{To see this bias, notice that $\beta\sum_{j\in\mathcal{C}_{i}}s_{j}z_{j}+\gamma_{0}\sum_{j\in\mathcal{C}_{i}}s_{j}=\left(\beta+\gamma_{0}\right)\sum_{j\in\mathcal{C}_{i}}s_{j}z_{j}+\gamma_{0}\sum_{j\in\mathcal{C}_{i}}s_{j}(1-z_{j})$. The left-hand side here corresponds to ((ref)), while the right-hand side to the specification of valentim2024does. The reduced-form coefficient therefore identifies $\beta+\gamma_{0}$ instead of $\gamma_{0}$. akey2015valuing (Table 7, column 2) employs a similar specification, except he interprets the second coefficient as a causal effect of losing.}

The second RDA control — the aggregated running variable in narrow elections — is not found in any of the papers we have reviewed. While some papers have no analog to it (e.g., campante2018long,nellis2018secular), several variations have been employed. azoulay2019public take an unweighted average of running variables near the cutoff, while the instrument is a sum of RD shocks weighted by some dollar amounts. The same is true for folke2014shades and merilainen2022political whose average additionally includes non-close elections. clots2012female and bhalotra2014religion,bhalotra2021religion use a different strategy: they control for the running variables of each subunit (including with non-narrow elections) separately. This requires ordering the subunits in an arbitrary way (such that changing the order would change the IV estimate) and filling missing values as zeros for units that have fewer subunits than others. Neither of these strategies appears to have a clear justification.

Finally, we are not aware of any study that included any version of the running variable aggregated among narrow wins only. Absent the second and third RDA variables, the estimators used in the literature do not inherit the properties of standard RD and therefore may not enjoy the low bias advantages of local linear estimation.

We conclude by noting that, while most of the papers estimate ((ref)) by IV, akey2015valuing considers its reduced-form instead. We see several disadvantages to this approach. It is a less coherent specification than IV since the explanatory variable depends on the bandwidth choice, and thus is not an economic variable with independent meaning. Practically speaking, if the fraction of narrow victories is correlated with the fraction of non-narrow victories (as may happen when the bandwidth is not infinitesimal), the reduced-form specification is biased. And even if the correlation is small or absent, the reduced-form specification is likely less efficient, as the error term includes the effects of non-narrow victories.

Upper-Level Outcome and Lower-Level Specification

The final group of 8 papers (Panel C of Appendix Table (ref)) considers outcomes measured at the more aggregated level than the discontinuity events but estimates certain regressions at the lower level. To structure the discussion, recall that the stacking approach of Section (ref) involves the sample of all narrow elections and a fuzzy RD specification with the reduced-form

equation[equation omitted — 97 chars of source]

with $X_{\textbf{i}(j)}$ as the endogenous variable in the structural equation, and $s_{j}$ as weights. The current practice deviates in terms of the sample restrictions and the use of OLS vs IV.

First, while some papers (e.g., de2020politics,de2024partisanship) keep the full stacked sample of narrow elections, others restrict the sample in different ways. beach2017gridlock only keep one narrow election for each unit, which has the running variable closest to the cutoff (although this does not affect too many units in their setting). dell2018nation only keep the narrow election that happened chronologically first. These approaches limit the amount of identifying variation, which our paper shows not to be necessary.

Second, while some papers (carozzi2022political and dell2018nation) clearly introduce the upper-level treatment and use the IV specification ((ref)), the majority of papers estimate the reduced-form ((ref)) by OLS. This has both conceptual and practical issues. Conceptually, ((ref)) is not a coherent causal model: it provides multiple, mutually inconsistent, expressions for the outcome, and the Stable Unit Treatment Value Assumption (SUTVA) is clearly violated. The causal effects of an extra election victory are also expected to vary substantially with $s_{j}$, limiting external validity of the OLS coefficient. For instance, de2020politics find the effect of an extra Democrat to be much larger in small legislatures. This issue would be avoided by an IV specification with the share of Democrats in the legislature as the treatment. Practically, in finite samples where the bandwidth is non-negligible, $z_{j}$ may be non-trivially correlated with the treatment statuses of other subunits (both narrow and others). This potentially results in a form of asymptotic bias that doesn't arise in IV analyses.

Extensions: Spillovers and Temporal Aggregation

So far, we have considered settings in which outcomes are measured at a higher level of cross-sectional (e.g., spatial) aggregation than the RD events. In this section, we discuss two related settings to which our RDA ideas extend. Section (ref) considers estimation of spillovers in RD settings, while Section (ref) focuses on aggregating RD events over time. We describe how the two estimators from Section (ref) apply in those settings and compare them with current practice.

Spillovers from RD Events

A number of studies consider spillovers from discontinuity events. Instead of units and subunits, we now have outcome units $i$ and intervention units $j$.\footnote{This terminology is common when analyzing experiments on bipartite networks (e.g., zigler2021bipartite).} We first discuss how our framework and solutions extend to this setting, and then compare it with the current empirical practice.

In spillover cases, the endogenous variable $X_{i}$ is often the number or fraction of treated neighbors, corresponding to the exogenous (“contextual”) peer effects model. Thus, it still follows equation ((ref)), with $\mathcal{J}_{i}$, reinterpreted as the set of $i$'s peers. Alternatively, the endogenous linear-in-means peer effects model is sometimes used, in which $X_{i}$ is the average outcome of the peers manski1993identification. In either case, the instrument is still defined by ((ref)), with $\mathcal{C}_{i}$ reinterpreted as the set of $i$'s peers with narrow elections.

Our upper-level solution applies with no change, while the lower-level solution can now take two forms. The stacked equation is now defined at the bilateral level: on a sample of pairs $\left\{ (i,j)\colon j\in\mathcal{C}_{i}\right\} $, estimate

equation[equation omitted — 127 chars of source]

with $z_{j}$ as the IV for $X_{i}$ and $s_{j}$ as weights. This is still a standard fuzzy RD specification with a continuous treatment, albeit with both $i$ and $j$ potentially observed multiple times. Absent additional $i$-level controls, i.e. if $\tilde{W}_{i}=0$, one can also obtain the same estimate $\hat{\beta}$ by collapsing ((ref)) to the level of intervention units, i.e. to the sample of narrow elections $j$, by averaging across outcome units that are $j$'s neighbors: i.e., one can estimate

equation[equation omitted — 116 chars of source]

where $\bar{Y}_{j}=\frac{1}{N_{j}}\sum_{i\in\mathcal{I}_{j}}Y_{i}$ and $\bar{X}_{j}=\frac{1}{N_{j}}\sum_{i\in\mathcal{I}_{j}}X_{i}$ for $\mathcal{I}_{j}=\left\{ i\colon j\in\mathcal{J}_{i}\right\} $ and $N_{j}=\left|\mathcal{I}_{j}\right|$, instrumenting $\bar{X}_{j}$ with $z_{j}$ and using $s_{j}N_{j}$ as weights.\footnote{Note that $\bar{X}_{j}$ is a convoluted treatment: it is the average treatment $X_{i}$ of $j$'s neighbors where $X_{i}$ itself is an average or sum of its neighbors' RD shocks. The first-stage captures that, when $j$ is treated, the average treatment exposure $X_{i}$ of $j$'s neighbors mechanically goes up.} Versions of all three approaches have been used previously; yet, like in Sections (ref) and (ref), none of them exactly coincides with our proposal.

Surprisingly, we have found only one paper that performs estimation at the $i$ level mora2023peer. This stands in contrast not only with the many papers using the upper-level specification in the aggregation setting of Section (ref), but more importantly with how peer effects are usually estimated in economics, via either exogenous or endogenous peer effects models manski1993identification. Ultimately, $i$-level specifications are, in our view, the most natural and coherent way to study $i$-level outcomes. Similar to papers in Section (ref), mora2023peer do not include the third RDA control of equation ((ref)) and only include the second RDA control as a robustness check.\footnote{In addition, they limit the sample of outcome units to those outside the bandwidth and, in the main analysis, define the treatment by averaging only across the peers within the bandwidth. Our theory would apply even with the full sample and a more natural definition of the treatment based on all peers.}

All other papers use specifications at the bilateral level or averaged by intervention unit $j$, and they are subject to the issues raised in Section (ref). The first issue involves unnecessarily restricting the sample. We find this in dahl2014peer who entirely drop all units with more than one neighbor near the cutoff. The second issue is that most papers consider reduced-form OLS regressions on $z_{j}$, rather than introducing the treatment at the level of outcome units (or perhaps as in ((ref))). The only exception we are aware of is dube2019fairness who estimate a bilateral IV specification similar to what we propose, with the average status of the peers as the endogenous variable. Other papers — e.g., baskaran2018does and dechezlepretre2023tax at the bilateral level and santoleri2022causal and bhalotra2018pathbreakers at the $j$ level — are subject to the conceptual and practical concerns of reduced-form specifications discussed in Section (ref).\footnote{These concerns also apply to papers that use IV estimation with a $j$-level treatment, rather than our proposed $i$-level treatment. isen2014local, for instance, estimates the effect of $Y_{j}$ on the average outcome of $j$'s neighbors. This is in contract to a conventional endogenous peer effects model that would estimate the effect of the average outcome of $i$'s neighbors on $Y_{i}$. See clark2009performance and dahl2014peer for related examples.}

RD Aggregation over Time

Another setting in which RDA ideas can be useful is when multiple RD events are aggregated over time. For instance, one may specify a state-level treatment as the share of years in the last decade when a Democratic governor is in power (in a cross-section of states or a panel with a rolling window). Our upper-level solution would entail instrumenting this treatment with the share of years in which the current governor is a Democrat who won in a narrow election, and controlling for the share of years determined by narrow elections, along with the other RDA controls. In turn, the lower-level solution would involve stacking all narrow elections of the previous decade, repeating the same outcome and aggregated treatment, and including standard RD controls. An unusual feature of the dynamic setting is that earlier elections can influence the outcomes of later elections, for instance via incumbency advantage lee2008randomized. Our solutions allow for this, under a standard exclusion restriction.

A more complex version of this setting is when elections are not regular and the previous election's result can affect whether election is held in a future year. This is the case in the literature that studies the effects of public school funding on, e.g., house prices and educational outcomes. This literature uses RD variation from local referenda for school district bonds that authorize capital spending (see jackson2024impacts for a review). If the bond is approved in one year, it is less likely that another referendum will be held the following year. In our lower-level solution, there will be endogenous selection of the set of narrow referenda included in the stacked sample. Still, nearly random outcomes of those referenda should make the estimator valid under an appropriate exclusion restriction (that precludes, for instance, a direct effect of holding a referendum on the outcome through an increase in media attention). Similarly, in our upper-level solution, RDA controls, such as the share of years with narrow referenda, may now resemble “bad controls,” as they can be affected by the results of earlier referenda while also related to the error term. Yet, it is nearly by chance whether a narrow election results in the bond passing or not, conditional on prior referendum margins. Based on Proposition (ref), we cautiously conjecture that the upper-level solutions also continues to be valid.

The current practice has followed a different approach, proposed by cellini2010value in the school bond context. They offer several estimators; the “one-step” dynamic RD specification most frequently adopted in later studies is as follows:

equation[equation omitted — 200 chars of source]

Here the set of $\beta_{h}$ coefficients for a set $\mathcal{H}$ of lags ($h\ge0$) and possibly leads ($h<0$) aims to capture the dynamic effects of school spending $x_{i,t-h}$. The spending $h$ periods ago is instrumented with the indicator of the bond approved in a referendum, $z_{i,t-h}$ (and reduced-form specifications are also considered). The set of controls includes the indicators $M_{i,t-h}$ that there was a referendum in year $t-h$ and global polynomials $P(r_{i,t-h};\theta_{h})$ in the vote margin $r_{i,t-h}$ with coefficients $\theta_{h}$ (where the vote margin is normalized to zero in years without an election), as well as additional controls $\tilde{W}_{it}$. cellini2010value point out that their specification does not allow earlier elections to affect the occurrence of future elections or the future vote margin: if that happened, $M_{i,t-h}$ and $P(r_{i,t-h};\theta_{h})$ would be bad controls biasing the coefficients of interest $\beta_{h}$.\footnote{They also propose a “recursive” estimator which allows for some of these endogenous responses, although hsu2021dynamic clarified that this is only true with homogeneous effects. The recursive estimator has not been popular in later work, perhaps because it tends to be noisier.}

We view our RDA solutions as complementary to the dynamic RD specification. The key advantage of the cellini2010value specification is that it allows to trace the dynamics of the effects (and also do placebo testing by looking at the lead coefficients). We note, however, that the effect of cumulative spending is also of interest: for instance, biasi2024works, and baron2022school similarly report the average effect of the bond approval across different horizons. The key disadvantage of specification ((ref)) is the potential bias from bad controls, which may not arise with RDA solutions, especially our lower-level specification. Additionally, ((ref)) is based on global polynomial estimation, which is generally avoided in more recent RD analyses for being noisy and sensitive to the order of the polynomial gelman2019high, while the RDA solutions follow the prevalent local linear estimation approach.\footnote{A local polynomial version of the IV specification ((ref)) would replace $M_{i,t-h}$ with a dummy of holding a close referendum at $t-h$, replace $P\left(r_{i,t-h};\theta_{h}\right)$ with a local polynomial (filled in with a zero if there was no referendum at all or the referendum was not close), and replace $z_{i,t-h}$ with dummies of the bond approved in a close referendum as the set of instruments. We are not aware of any paper that used this estimation procedure or studied its properties formally.}

Application: The Impact of Unions on US Inequality

We now apply the RDA identification strategy to estimate the effect of unions on wage inequality within state-by-industry cells in the US. Section (ref) introduces the empirical setting and adapts the RDA design to this application. Section (ref) describes the data construction and presents summary statistics. Section (ref) validates our research design in a series of balance tests. Section (ref) presents the empirical results, and Section (ref) discusses what these results mean for the contribution of declining unions to growing inequality in the US.

\protectEmpirical Setting and Specifications

Trade unions are significant labor market institutions in all Western countries, including the US. Over the last 100 years, union density in the US has followed an inverse U-shaped pattern: there was a sharp increase in unionization during the Roosevelt era (1930–45), a relative steady state of high unionization rates from 1945 to 1960, and a continuous decline since the 1960s farber2021unions. The strong negative correlation between this trend and the U-shaped pattern of U.S. income inequality piketty2018distributional contributes to the widespread perception of a relationship between the two phenomena.

Recent studies have explored the relationship between unionization and inequality callaway2018unions,collins2019unions,fortin2021labor,farber2021unions. Although these studies report a significant negative association between unions and inequality in a cross-section or a panel, a causal interpretation of those results requires strong assumptions, as illustrated by the quote in the Introduction. Notably, these studies have not adopted the RD approach that is common in the analyses of the union effects on establishment-level outcomes dinardo2004economic,sojourner2015impacts,frandsen_2021.

Unionization in the US occurs at the level of a bargaining unit, defined as “a group of two or more employees who share community of interest and may reasonably be grouped together for purposes of collective bargaining” nlrb1997basic. In most cases, a bargaining unit corresponds to a single establishment, representing on average close to half of the establishment's workforce frandsen_2021.\footnote{To be precise, frandsen_2021 reports an average of 93 votes and 254 total employees in his sample covering most unionization elections between 1980–2009. Using data from all unionization elections in this period, we found an average turnout rate of 85% (as a fraction of eligible voters), suggesting a typical bargaining unit size of approximately 109 employees (93/0.85). This indicates that bargaining units represented about 43% (109/254) of the total workforce in establishments holding elections.} The election typically offers two choices: for or against union formation. To form a union, a majority of 50% plus one out of the cast votes is necessary, lending itself to an RD design.

We exploit the discontinuities from narrow union elections to estimate union effects at the more aggregate level of state-industry cells. Specifically, we define the outcome $\Delta Y_{sit}$ as the change in an economic outcome (e.g., a measure of wage inequality) in state $s$ and industry $i$ during decade $t$,\footnote{We use differenced outcome variables both to increase estimation power and as a way to reduce potential bias from manipulated unionization election results, as suggested by frandsen2017party,frandsen_2021. Additionally, using differenced outcome variables aligns with our treatment—the rate of new unionization—representing most of the change in the share of unionized workers.} and the endogenous treatment $\text{NewUnions}_{sit}$ as the share of newly unionized workers. The latter is measured as the ratio of the number of workers in the region-industry cell who join unions through unionization elections (both narrow and not) during the decade and the workforce in the cell at the beginning of the decade.

Looking at outcomes more aggregated than establishments has two advantages. First, it allows us to study the effects on outcomes, such as inequality, that are typically defined at more aggregated levels. Second, even for outcomes—such us average wages—that can be measured by establishment, our estimates capture the effects due to the spillovers of unionization on other establishments within the same cell. Establishments in the same state and industry often compete for the same workers and therefore may adjust wages or even enter or exit the market in response to another establishment's unionization. Defining cells as a combination of state and industry follows the approach of fortin2021labor,fortin2022right. Aggregating the analysis to purely geographic units — either states farber2021unions or economic areas similar to commuting zones (such as Census State Economic Areas in collins2019unions) — could help capture cross-industry spillovers, but such aggregation is not feasible for our study. State-level aggregation would substantially reduce statistical power, as typical cells would contain many close elections, and the law of large numbers would eliminate most identifying variation in our RDA approaches. Meanwhile, data limitations do not allow us to achieve spatial resolution smaller than states.\footnote{For metropolitan areas, Census public files report geographic units less aggregated than states. However, both the number of these units and their boundaries change each decade, hindering the construction of panel data and the use of differenced outcomes.}

Applying our upper-level IV solution from Section (ref) to this setting, we estimate the following structural equation and first-stage:

align[align omitted — 262 chars of source]

Here the instrument $Z_{sit}$ is the employment share of the bargaining units unionized in a narrow election in the state-industry workforce. Formally, $Z_{sit}=\sum_{j\in\mathcal{C}_{sit}}s_{j}z_{j}$ where $\mathcal{C}_{sit}$ is the set of close elections in the $si$ cell during decade $t$, $z_{j}$ is the dummy that the vote share for the union strictly exceeds 50% in workplace $j$, and $s_{j}$ is the number of eligible voters in election $j$ as a fraction of the cell workforce at the beginning of the decade.\footnote{We distinguish between winning and obtaining a majority of the votes, as in a very small fraction of cases the final result was revised due to an appeal process. Our setting is “fuzzy” in this sense, and we handle it similarly to the standard fuzzy RD design: our instrument is defined based on dummies for crossing the 50% cutoff, while the treatment is based on dummies for union victory.} The covariates $Q_{sit}$ consist of the three RDA controls defined by ((ref)); in this context they represent the share of workforce involved in narrow elections ($\sum_{j\in\mathcal{C}_{sit}}s_{j}$), a rescaled average of the running variable $r_{k}$, defined as the union vote share minus 0.5, in closed elections ($\sum_{j\in\mathcal{C}_{sit}}s_{j}r_{j}$), and a rescaled average of the running variable in close wins ($\sum_{j\in\mathcal{C}_{sit}}s_{j}r_{j}^{+}$). Finally, $FE_{st}$ and $FE_{it}$ are state-by-decade and industry-by-decade fixed effects, respectively, included solely to increase estimation power, as the RDA design alone should be sufficient to obtain causal estimates.\footnote{We do not include state-by-industry fixed effects both because the outcome is already measured in differences and because including such fixed effects would result in a substantial reduction of the sample: our industry definitions change after 2000, leading to only one post-2000 state-by-industry observation (see Section (ref)).} We also explore specifications with different sets of fixed effects. To obtain more efficient and economically meaningful estimates, we weight each observation by the cell employment at the beginning of the decade.

We also apply the stacking approach of Section (ref), which involves estimating the following structural equation and first-stage in the sample of narrow unionization elections $j$:

align[align omitted — 299 chars of source]

where $s(j),i(j),t(j)$ respectively indicate the state, industry, and decade in which election $j$ took place and $q_{j}$ consist of the standard RD controls: a constant, the vote margin, and that margin interacted with the union getting a majority. The structural equation is identical to the upper level specification, except for the controls and the error term. Per Proposition (ref), we weight observations by the number of eligible workers in each workplace $j$.

\protectData

We collect data on unionization and labor market outcomes by state, industry, and decade over five decades, from 1960s to 2000s.\footnote{This period was chosen based on the availability of unionization data, which is accessible from 1961 onwards and contains industry variables until the beginning of 2009.} The data on establishment-level union elections is from the National Labor Review Board (NLRB), maintained by Henry Farber and John-Paul Ferguson. For each election, we observe the state, industry, and year, as well as the total number of eligible workers and the votes cast for and against. We construct the treatment variable $\text{\text{NewUnions}}_{sit}$ using the results of all elections. The instrument, in turn, aggregates only the elections that satisfy three criteria. First, it only includes close elections, defined by the vote share for the union of $50\pm10$ percent (the smallest bandwidth employed in recently published papers that employ the RD strategy in the context of US unions; knepper2020fringe,frandsen_2021).\footnote{The 10p.p. bandwidth also closely aligns with the optimal bandwidth for local linear estimation. Applying the calonico2014robust procedure to our lower-level specification across our main five inequality outcomes, we find optimal bandwidths ranging from 9.5% to 13.7%, with an average of 10.9%. We also check robustness to using a 15% bandwidth which is also used by knepper2020fringe and frandsen_2021.} Second, we exclude tied elections and elections with a 1-vote margin, as there is evidence of manipulation in such elections frandsen_2021. Finally, we drop elections with less than 20 votes, as is customary in the unionization literature dinardo2004economic,lee2012long,frandsen_2021.\footnote{During the period of our analysis, there are two types of unionization elections: certification elections, where workers attempt to form a new union, and decertification elections, where workers or employers try to dissolve an existing union. The rules for both types of elections are identical, as are the consequences: if the union receives a strict majority of votes, it will be certified or re-certified; if not, it will be decertified. The vast majority of elections are certification elections, accounting for 88.2% of elections and covering 90.2% of the workers. We do not distinguish between these two types of elections in our analysis; e.g., $NewUnions_{sit}$ counts workers in both newly certified and re-certified unions. Our results are robust to restricting the set of elections used for the treatment and instrument construction to certification elections only.}

Our wage inequality measures and other wage-related outcomes by state, industry, and year are based on the decennial population census samples from 1960 to 2000, supplemented by the American Community Survey for 2009–2013 that represents 2010. With some variation over the years, these samples represent roughly 5% of the US population. We include observations from all 50 states, as well as from the District of Columbia, defined as an additional state. We only include workers with a strong attachment to the labor market, defined as working all 52 weeks during the year and at least 20 hours per week on average.\footnote{We exclude from the sample workers who worked fewer than 52 weeks, as they have a higher likelihood of having switched industries during the past year, in which case we could attribute their annual income to the wrong industry.} Two of our main inequality outcomes are computed using total labor income: the Gini index and the top-10% income share. For the other measures, we calculate the average annual hourly pre-tax wage of each worker by dividing the total annual wage income by the product of the number of weeks employed and the average hours worked per week last year. We then compute the college wage premium, the log 90-10 wage ratio, and the variance of log wages, again by state-industry cell. The construction of the inequality measures is detailed in the Data Appendix, following autor2008trends and farber2021unions with minor modifications. For the inequality measures to be meaningful, we restrict the sample to include only cells with more than 1,000 workers after applying census projection weights.\footnote{This entails dropping 24.9% of the cells, but only 0.8% of total employment and 1.3% of close unionization elections. We verify there are no remaining cells with fewer than 10 raw observations. For measures based on smaller samples, such as for the average college wage, we further exclude cells with fewer than 10 relevant observations at the beginning or at the end of the decade.}

We additionally use the Current Population Survey (CPS) Monthly and Annual Social and Economic Supplement samples. We measure union density in each cell and year based on the self-reported union coverage question. We also construct two measures of worker benefits: indicators of a pension plan and employer-sponsored health insurance coverage. These questions are available prior from 1977–2014; to increase sample size we combine CPS data from several years to represent each decennial census year, as detailed in the Data Appendix. The cleaning process of the CPS data is otherwise the same as for the census data.

A challenge in the data construction process was to harmonize industry definitions over a long period and between different data sources. The NLRB data is based on the SIC and NAICS schemes, depending on the year, while the population census, ACS, and CPS use a different classification. We created two distinct coding schemes. Until 2000, our industry coding is based on the initial two digits of the SIC87 codes; from 2000 onwards, it relies on the initial three digits of the NAICS97 industry classification. In both cases, we merged a few categories, resulting in 70 industries categories pre-2000 and 73 (different) categories post-2000 that cover all private sectors.\footnote{The NLRB data do not cover public sector unionization and, thus, those industries are excluded from our analysis. Our classification is more detailed than 10–11 industries used by fortin2021labor,fortin2022right.} To calculate differenced outcomes for each decade, we estimate all measures for the year 2000 based on both schemes.\footnote{Since the 2000 census industry question was based on the NAICS industry scheme, we relied on the IPUMS crosswalk to match observations to the SIC-based census industries. While this could introduce noise in the 1990–2000 differences, we find it reassuring that employment changes do not exhibit substantially higher volatility than in the decade before or after.} In total, we were able to match 93% of unionization elections to our industry codes and 99% of individual census observations.

table[table omitted — 5,216 chars of source]

Table (ref) presents summary statistics for the key variables at the state-industry-census year level, both in levels and first-differences. Panel A includes outcomes, while Panel B details other cell-level statistics. One can see, in particular, that all within-cell inequality measures show an increasing trend, mirroring national inequality trends. Panel C reports the endogenous variable, the instrument, and the RDA controls described above, as well as the average numbers of all and narrow unionization elections. In an average decade, there are 55 union elections per cell resulting in 2.5% of the workforce joining the union. An average of 8 elections per cell, involving 0.6% of the workforce, results in a narrow win that enters our instrument construction. Appendix Table (ref) reports summary statistics at the election level, showing in particular that the number of elections has seen a continuous decline since 1970s (which matches the negative trend in union membership reported in Table (ref), Panel A, with the CPS data since 1980s).

While we only look at inequality within state-industry cells, this type of inequality is central to both the level and growth in overall inequality. To show this, Appendix Table (ref) decomposes the overall variance of log wages into within- and between-cell components in each census year, broadly following helpman2016trade. Throughout the period, the within component of the variance is the dominant one, accounting for 73–80% of total inequality. Moreover, the within component experienced a continuous rise, contributing significantly to the overall increase in inequality. In contrast, the between-group component remained relatively stable, decreasing in some periods and only experiencing a notable rise from 2000 to 2010.\footnote{A concern in this analysis is that some of the changes in the variance components are due to the changes in the definitions of industries in the year 2000 from SIC based to NAICS based. To address this, we conducted an identical analysis on with annual CPS samples in Appendix Figure (ref). It indicates no discernible changes in the within and between components between 2002 and 2003, the year of the CPS industry scheme change, providing strong evidence that this transition does not drive the trends in the components. This analysis also replicates the findings from Appendix Table (ref) using different data, confirming the key role of within-cell inequality.}

\protectBalance

We now perform balance tests to illustrate how the RDA instrument and controls help isolate idiosyncratic variation in the treatment. Column 1 of Table (ref) regresses the share of all newly unionized workers (i.e., our treatment of interest) on pre-determined covariates and lagged inequality measures from Table (ref) (excluding those which are not available for all decades). The treatment is correlated with many observables, as evidenced by the partial $R^{2}$ of 13.5% and a decisively rejecting $F$-test, which suggests that the treatment is also likely correlated with unobservables. In column 2 we estimate the same regression for our instrument: the share of workers who join a union through close elections; the conclusion is unchanged. In column 3, the inclusion of the first RDA control that represents the share of workers involved in narrow union elections vastly changes the statistics. The partial $R^{2}$ decreases massively to 0.1%, the F-statistic of the covariates' joint significance shrinks from 185.6 to 3.5, and most coefficients become orders of magnitude smaller. Yet, covariates are still jointly significant (with a p-value of 0.01%). When we finally add the two other RDA controls to the regression in column 4, the partial $R^{2}$ becomes essentially zero and, while one out of ten covariates is individually significant, jointly they are not (p-value of 12.3%). The lack of correlation between the residual variation in our instrument and the observables speaks in favor of the validity of our IV strategy. In Appendix Table (ref) we replicate this analysis with a 15% bandwidth defining close elections, and the same patterns emerge.

table[table omitted — 4,185 chars of source]

Appendix Table (ref) performs a related pre-trend and placebo analysis, using predetermined covariates and lagged outcomes, both in levels and differences, as dependent variables when estimating the reduced-form of our upper-level IV specification:

align[align omitted — 110 chars of source]

None of the 15 coefficients are significant at the 5% level using either conventional standard errors or robust bias-corrected confidence intervals of calonico2014robust.

\protectResults and Mechanisms

sidewaystable\begin{centering} \caption{The Impact of New Unionization on 10-Year Changes in Inequality} \end{centering} \begin{centering} { \begin{tabular}{lccccc} \toprule & {$\Delta$Log college premium} & {$\Delta$Log 90/10} & {$\Delta$Gini coeff.} & {$\Delta$Top 10% share} & {$\Delta$Var(log wage)}\tabularnewline & { (1)} & { (2)} & { (3)} & { (4)} & { (5)}\tabularnewline \midrule \multicolumn{6}{c}{{\smallPanel A: Upper-Level Estimator}}\tabularnewline { Share newly unionized} & { -0.314} & { -0.459} & { -0.176} & { -0.143} & { -0.255}\tabularnewline & { (0.285)} & { (0.225)} & { (0.061)} & { (0.055)} & { (0.084)}\tabularnewline & {{[}-0.921,0.416{]}} & {{[}-0.933,0.122{]}} & {{[}-0.325,-0.035{]}} & {{[}-0.279,-0.020{]}} & {{[}-0.463,-0.054{]}}\tabularnewline & & & & & \tabularnewline { Mean outcome} & { 0.515} & { 1.318} & { 0.348} & { 0.282} & { 0.311}\tabularnewline { Mean treatment} & { 2.65%} & { 2.76%} & { 2.76%} & { 2.76%} & { 2.76%}\tabularnewline { RDA controls} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$}\tabularnewline { Industry-decade and state-decade FE} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$}\tabularnewline { Observations} & { 10,512} & { 12,910} & { 12,910} & { 12,910} & { 12,910}\tabularnewline & & & & & \tabularnewline \multicolumn{6}{c}{{\smallPanel B: Lower-Level Estimator}}\tabularnewline { Share newly unionized} & { -0.258} & { -0.266} & { -0.133} & { -0.115} & { -0.206}\tabularnewline & { (0.284)} & { (0.185)} & { (0.052)} & { (0.048)} & { (0.073)}\tabularnewline & {{[}-0.896,0.443{]}} & {{[}-0.735,0.140{]}} & {{[}-0.284,-0.035{]}} & {{[}-0.256,-0.026{]}} & {{[}-0.427,-0.075{]}}\tabularnewline & & & & & \tabularnewline { Mean outcome} & { 0.492} & { 1.189} & { 0.305} & { 0.251} & { 0.254}\tabularnewline { Mean treatment} & { 9.00%} & { 9.54%} & { 9.54%} & { 9.54%} & { 9.54%}\tabularnewline { RD controls} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$}\tabularnewline { Industry-decade and state-decade FE} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$}\tabularnewline { Observations} & { 29,146} & { 31,256} & { 31,256} & { 31,256} & { 31,256}\tabularnewline \bottomrule \end{tabular}} \end{centering} {\footnotesizeNotes: }{ Panel A reports IV estimates of ((ref)). Observations are defined as state-industry cells by decade (for 1960–2010). The endogenous variable is the share of workforce unionized through unionization elections, relative to the beginning-of-decade cell workforce. The instrument is the share of workforce unionized through close elections, defined by the 10% bandwidth for the vote share around the cutoff of 50%. RDA controls are as in equation ((ref)), and state-year and industry-year FEs are included. See Data Appendix for the outcome definitions and details of the sample construction. Observations are weighted by employment at the beginning of the decade. Panel B shows the coefficients obtained from the lower-level IV specification ((ref)), in the sample of close elections. The specification includes standard RD controls and is weighted by the number of eligible workers. In both panels, heteroskedasticity-robust standard errors are in parentheses, and robust bias-corrected confidence intervals of calonico2014robust are in brackets (in Panel A, constructed via the equivalent lower-level specification as in Proposition (ref)).} {\footnotesizeSources:}{ Decennial Census and NLRB unionization data; authors' calculations.}

Table (ref) examines the impact of unionization on five key measures of income inequality: the college premium, the 90-10 wage ratio, the Gini coefficient, the top-10% income share, and the variance of log hourly wages. Panel A shows the estimates using the upper-level approach, while Panel B shows the coefficients obtained at the lower level. Along with each estimate, we report robust bias-corrected confidence intervals following calonico2014robust, obtained using the RDrobust package in R.\footnote{To obtain confidence intervals for our upper-level approach, we use the numerically equivalent lower-level specification, per Proposition (ref). We do not use clustered inference (e.g., by state or by industry) for several reasons: theoretically, election shocks should be mutually independent; empirically, we confirm that clustered standard errors do not substantially exceed heteroskedasticity-robust ones; and pragmatically, the implementation of clustering in the RDrobust package does not allow elections on the two sides of the cutoff to be part of the same cluster.}

The results reveal a negative effect of unionization on all measures of inequality, with 95% confidence intervals rejecting zero effects for the last three inequality measures and with large but statistically insignificant negative effects for the college premium and the 90-10 wage ratio. Specifically, upper-level estimates suggest that an increase of 1 percentage point in new unionization reduces the college premium by 0.31 log-points, the 90-10 ratio by 0.46 log-points, the Gini coefficient by 0.018, the top-10 income share by 0.14 percentage points, and the variance of log wages by 0.0025 (which is around 1% of its mean). Lower-level estimates are moderately smaller in magnitude but show the same statistical significance. Figure (ref) visualizes these results with standard RD plots for the first-stage and reduced-form lower-level specifications (excluding the fixed effects, to stay closer to the raw data).

figure[figure omitted — 1,466 chars of source]

Appendix Table (ref) reports several robustness tests for these results, focusing on the upper-level analyses. First, we replace our baseline bandwidth of 10% with 15% bandwidth, a common ad hoc choice in the literature (e.g., frandsen_2021,knepper2020fringe). Second, we exclude union decertification elections from the construction of both the treatment and the instrument (along with the RDA controls). Third, we replace state-year and industry-year fixed effects with separate state, industry, and year and FEs, or with no FEs at all. Finally, we perform estimation without initial employment weights. All results are similar to those from Table (ref), with coefficients from the unweighted estimation generally slightly larger, possibly indicating a larger union effect on inequality in smaller cells.

While we specified the treatment variable to be the rate of new unionization, the results would also be similar if we considered changes in union density, i.e. in the overall share of unionized workers. The latter measure would take into account differential employment growth of unionized and non-unionized establishments, as well as establishment entry and exit. Unfortunately, union density is only observed in the CPS since 1977. Thus, in the spirit of two-sample IV angrist2009mostly, we estimate the effect of new unionization rates (from the NLRB data) on union density (from the CPS monthly samples) for the shorter period of 1980–2010.\footnote{One may think of this analysis as a first-stage, except that this is an IV specification leveraging close elections for the instrument as before.} The coefficient is reported in the first column of Table (ref). This coefficient may be smaller than one due to various types of measurement error: both on the right-hand side (discrepancies arising from out-of-state workers, errors in coding the industry of unionization, or lower employment growth in unionized establishments) and on the left-hand side (misclassification of the union coverage variable in the CPS leading to a non-classical measurement error).\footnote{Based on a 1977 validation survey, the estimate for the misclassification rate in the union variable is 2.7% card1996effect.} Conversely, if unionized workplaces have higher employment growth, the first-stage parameter might exceed one. In the data, we find estimates ranging from 0.63 to 1.27, suggesting that our main IV estimates are robust to using the change in the self-reported share of total unionized workers as the endogenous variable.

table[table omitted — 2,887 chars of source]
sidewaystable\begin{centering} \caption{The Impact of Unions on 10-Year Wage Changes} \end{centering} \begin{centering} { \begin{tabular}{lcccccccc} \toprule & { All workers} & { College} & { High school} & { P90} & { P50} & { P10} & { Mgmt} & { Non-Mgmt}\tabularnewline & { (1)} & { (2)} & { (3)} & { (4)} & { (5)} & { (6)} & { (7)} & { (8)}\tabularnewline \midrule \multicolumn{9}{c}{{\smallPanel A: Upper-Level Estimator}}\tabularnewline { Share newly unionized} & { -0.348} & { -0.449} & { -0.105} & { -0.282} & { -0.129} & { 0.176} & { -0.917} & { -0.181}\tabularnewline & { (0.150)} & { (0.349)} & { (0.142)} & { (0.180)} & { (0.145)} & { (0.166)} & { (0.399)} & { (0.134)}\tabularnewline & {{[}-0.762,-0.032{]}} & {{[}-1.221,0.468{]}} & {{[}-0.442,0.227{]}} & {{[}-0.735,0.138{]}} & {{[}-0.533,0.188{]}} & {{[}-0.283,0.497{]}} & {{[}-1.854,0.058{]}} & {{[}-0.561,0.091{]}}\tabularnewline & & & & & & & & \tabularnewline { Mean outcome} & { 2.832} & { 0.545} & { -0.019} & { 3.305} & { 2.637} & { 1.986} & { 0.507} & { 0.106}\tabularnewline { Mean treatment} & { 2.76%} & { 2.65%} & { 2.76%} & { 2.76%} & { 2.76%} & { 2.76%} & { 2.71%} & { 2.76%}\tabularnewline { RDA controls} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$}\tabularnewline { Industry-decade and state-decade FE} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$}\tabularnewline { Observations} & { 12,910} & { 10,541} & { 12,873} & { 12,910} & { 12,910} & { 12,910} & { 10,638} & { 12,910}\tabularnewline & & & & & & & & \tabularnewline \multicolumn{8}{c}{{\smallPanel B: Lower-Level Estimator}} & \tabularnewline { Share newly unionized} & { -0.303} & { -0.380} & { -0.119} & { -0.158} & { -0.173} & { 0.108} & { -0.924} & { -0.130}\tabularnewline & { (0.128)} & { (0.351)} & { (0.121)} & { (0.151)} & { (0.127)} & { (0.136)} & { (0.386)} & { (0.111)}\tabularnewline & {{[}-0.725,-0.100{]}} & {{[}-1.237,0.433{]}} & {{[}-0.44,0.134{]}} & {{[}-0.584,0.146{]}} & {{[}-0.547,0.061{]}} & {{[}-0.241,0.395{]}} & {{[}-1.946,-0.127{]}} & {{[}-0.492,0.048{]}}\tabularnewline & & & & & & & & \tabularnewline { Mean outcome} & { 2.809} & { 0.525} & { -0.010} & { 3.253} & { 2.658} & { 2.064} & { 0.468} & { 0.057}\tabularnewline { Mean treatment} & { 9.54%} & { 9.00%} & { 9.54%} & { 9.54%} & { 9.54%} & { 9.54%} & { 9.20%} & { 9.54%}\tabularnewline { RDA controls} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$}\tabularnewline { Industry-decade and state-decade FE} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$} & {$\checkmark$}\tabularnewline { Observations} & { 31,256} & { 29,148} & { 31,250} & { 31,256} & { 31,256} & { 31,256} & { 30,171} & { 31,256}\tabularnewline \bottomrule \end{tabular}} \end{centering} {\footnotesizeNotes: }{ This table reports IV estimates from the same specifications as in Table (ref) but for different outcomes. The outcome in column 1 is the change in log average wage in the sample of workers. In columns 2–3 and 7–8, the outcome is the change in the log average wage among specific groups of workers: those with exactly a college degree, exactly a high school degree, managerial occupations, and non-managerial occupations, respectively. State-industry cells with fewer than ten observations for the relevant group at the beginning or end of the decade are excluded. In columns 4–6, the outcome is the change in the weighted 90th, 50th, and 10th percentile of the distribution of log hourly wages within the cell. Observations are weighted by total cell-level employment at the beginning of the decade in all columns.} {\footnotesizeSources:}{ Decennial Census and NLRB unionization data; authors' calculations.}

Table (ref) investigates the sources of the inequality reduction due to unionization, finding that it originates primarily from the decline in earnings among high-earners, rather than an increase at the bottom of the earnings distribution. Column 1 reports a large negative and marginally significant effect on average wages.\footnote{This large negative effect is unusual relative to the other estimates of a union wage premium at the worker or establishment level. card1996effect and farber2021unions find a positive effect using a selection-on-observables design at the individual level. Staggered difference-in-differences estimates from de2020two also at the individual level and RD estimates by dinardo2004economic and frandsen_2021 at the establishment level point out to a null or slightly negative effect. Our estimate may differ for several reasons. First, we incorporate spillover effects across establishments in the same state and industry. Second, our estimates incorporate spillovers on non-unionized employees, such as management (including even managers located at a different establishment), while worker-level estimates do not. Finally, like any RD-based estimate, we only use local variation from close elections.}$^{,}$\footnote{A caveat to this finding is that it may partially result from worker selection, as unionization has unequal effects on employment across worker groups with different earnings potential, as studied in Appendix Table (ref). There is no clear negative effect on overall employment (Column 1), nor on the employment of high-school graduates and non-college workers more broadly (Columns 2–3). In contrast, employment of workers with exactly a college degree or those with college or more education declines substantially in response to unionization (Columns 4–5, marginally insignificant). This finding aligns with frandsen_2021 who documents similar compositional effects at the establishment level, where high-paid workers tend to leave following unionization. Such composition effects can appear as a reduction in average wages. One may expect the composition effects to be smaller for columns 2–3 and 7–8 of Table (ref) that investigate less heterogeneous worker populations.}The following columns of Table (ref) measure the impacts of unionization on average wages of different worker groups and parts of the wage distribution. Columns 2 and 4 indicate a negative (albeit not statistically significant) impact of unions on wages among the higher-earning segments of the labor force: college graduates and 90th percentile earners, respectively; column 7 indicates a large and marginally significant negative effect on managers' earnings. In contrast, columns 3, 5, and 8 show negative but small and statistically insignificant wage impacts on the lower segments: high school graduates, 50th percentile earners, and non-managers, respectively. Column 6 finds a small and insignificant increase in wages at the 10th percentile of the distribution. By construction, the more negative effects at the top are consistent with our main finding of the negative impact of unions on inequality.\footnote{Appendix Table (ref) conducts several robustness checks for these results. The overall conclusion of a negative effect on average wages driven by the high-earners is preserved, although the negative effect on the managerial wages tends to be weaker in the robustness analyses.}

The null, or even negative, effects on median and average wages, is puzzling and may raise questions about the motivations for unionization and the resistance of employers to unionization. We shed light on a possible explanation by considering the effects of unionization on fringe benefits—health insurance and pensions coverage—in columns 2 and 3 of Table (ref). Since both the outcomes and the endogenous variable are measured in percentages, the coefficient is interpreted as the number of covered workers induced by one new union member. For health insurance, the effect is small and insignificant. However, for pension coverage, we estimate that each new union member corresponds to an increase of 1.48 pension holders (standard error of 0.78) using the upper-level instrument and 0.86 (SE of 0.50) when the lower level-instrument is employed. Although the coefficients are noisy, the fact that they are around one or even larger suggests a substantial spillover effect of unionization across establishments within state-industry cells, given that some of the unionized establishments had pension plans prior to unionization. This is consistent with the recent findings by knepper2020fringe that unionization of establishments within firms has significant impacts on employer pension contributions on other establishments of the same firm. Our analysis may also suggest spillovers to other firms within the same state and industry.

\protectHow Much Did Declining Unions Contribute to Growing Inequality?

Rates of new unionization in the private sector in the US experienced a massive and continuous decline since the 1970s, as can be seen in the first panel of Figure (ref). In the 1960s, the net flow of unionized workers through unionization elections was 7.8%. This is the result of the share of workers joining unions through unionization elections (8.3%), minus some workers leaving unions through union decertification elections. The net flow declined consistently, reaching less than 1% in the 2000s.

figure[figure omitted — 1,476 chars of source]

Using the results in Table (ref), we conduct a back-of-the-envelope calculation of how much of the growth in US inequality during the period of our analysis can be attributed to this sharp decline. To this end, we compute counterfactual inequality measures that may have occurred if new unionization rates had remained at their 1960s levels. Specifically, we add to each actual inequality measure the product of the relevant treatment effect and the cumulative shortfall in new unionization resulting from the sharp decline after the 1960s.

Like any back-of-the-envelope calculation, ours requires some caveats. The estimates from Table (ref) capture a short-term effect and are identified using local variation from narrow elections. Here, we extrapolate these effects to the long term, assuming that the effect of new unionization during a decade on the level of inequality by the end of the decade equal the effect on inequality at later dates, too. Our analysis is also limited to assessing the effects of unions on inequality within state-industry cells. However, as shown above, within-cell inequality accounts for the majority of overall inequality both in levels and in changes.\footnote{Another limitation is that we assess only the effect of new unionization, not the full impact of changes in union density, which may also result from the closure or growth of unionized establishments. However, this concern is mitigated by our earlier finding that the effect of new unionization on union density is approximately equal to 1.}

With these caveats, the second panel of Figure (ref) presents the actual and counterfactual average within-cell Gini coefficients by census year. The gap between the actual and counterfactual inequality is small in 1980 and 1990 but becomes substantial by 2000 and 2010. Appendix Figure (ref) reports similar results for the other inequality measures.

Table (ref) summarizes all results of this back-of-the-envelope exercise showing a sizable contribution of declining unions to the growing within-cell inequality between 1970 and 2010. Unions account for 34-38% of the changes in the Gini coefficient, the 90/10 ratio, the top 10% share, and the variance of log wages. The contribution to the log college premium is even higher (at 61%) but imprecisely estimated. These estimates are moderately larger than the farber2021unions estimates of the contribution of unions to within-state inequality in the US between 1968 and 2014.

table[table omitted — 2,181 chars of source]

Conclusion

We introduced a novel extension to the regression discontinuity design, termed Regression Discontinuity Aggregation (RDA), which aims at identifying local average causal effects in contexts where each unit may be exposed to multiple discontinuity events. The RDA approach is applicable when the outcome is defined at a higher level of aggregation (e.g., in space or time) than the discontinuity events or when spillovers of discontinuity events are of interest. Our theoretical analysis rests on the observation that an aggregation of RD events can be viewed as a special case of a shift-share variable. This connection allowed us to build two estimators for causal effects in RDA settings that build on local linear estimation in conventional sharp and fuzzy RD designs and inherit their bias-reduction properties.

To illustrate the usefulness of this approach, we applied the proposed estimation techniques to the effects of unionization on earnings inequality at the level of state-by-industry cells. We found that increased rates of new unionization significantly reduce inequality, primarily by lowering top-tier wages rather than raising lower-end wages. As a result, the decline of unions since 1970s can explain a substantial fraction of increased inequality. We additionally provide evidence that unionization leads to increased pension coverage, potentially in non-unionized establishments, too. Our findings provide causal evidence supporting common perceptions about the effects of unionization on inequality that were previously difficult to obtain as credibly.