Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
110,013 characters · 13 sections · 82 citation commands
Identification of Causal Effects with a Bunching Design
\onehalfspacing
In this paper, we show that bunching in the distribution of a treatment variable can be used to identify causal effects. In particular, although the treatment may be endogenous, we do not rely on instrumental variables, regression discontinuity designs, panel data, functional form, or distributional assumptions.
The setting is a standard causal model in which both the treatment and the outcome variables are observed. The treatment variable has a mass point and is continuously distributed near this bunching point. The outcome variable is continuously distributed near and at the bunching point. The example in our application is a useful benchmark: the treatment is the number of cigarettes a woman smokes during pregnancy (with 81% of the observations bunching at zero), and the outcome is the baby's birth weight. caetano2015 showed strong evidence of endogeneity in this application. To identify treatment effects, we need to eliminate selection bias, the part of the outcome variation that is due to confounders.
The key insight that makes identification possible is that the change-of-variables theorem from integration theory can be used to write the magnitude of selection bias as the ratio of the probability density of the treatment variable and the probability density of the part of the outcome that is due to confounders. As a consequence, we do not need to observe the values of those confounders; instead, it is sufficient to identify the distribution of the part of the outcome that varies with confounders. This is where bunching is useful: at the bunching point, the distribution of the outcome reflects only the variation of unobservables, since the treatment stays fixed. We use this to identify the selection bias.
We identify the average per unit effect of marginally increasing treatment among those near the bunching point (equivalently, among the observations at the bunching point that are most similar to those near it). In our application, this quantity represents the expected rate of birth-weight loss if a woman who is currently not a smoker but is very similar to the women who smoke very little were to start smoking. Alternatively, minus this quantity can be interpreted as the expected rate of birth-weight gain if the women who currently smoke little were to quit smoking.
The approach relies on four conditions. First, the treatment effects must be sufficiently smooth near the bunching point. In our application, this condition means that smoking is not so harmful that a marginal amount could (on average) cause a discrete birth-weight change. Second, selection must also be sufficiently smooth at the bunching point. In our application, this means that the mothers who smoke very little (“marginal smokers”) are comparable to the mothers who do not smoke, but are indifferent between not smoking and smoking a small positive amount (“marginal nonsmokers”). Third, the selection bias must maintain the same sign in a neighborhood near the bunching point. In our application, this means that if selection into smoking is negative among mothers who smoke little (i.e. smoking less is associated with higher untreated birth weights), then selecting into not smoking must be associated with even higher birth weights. Finally, the outcomes at the bunching point are determined not only by confounders, but also by idiosyncratic (i.e., unconfounded) variation, which we must then separate from the distribution of outcomes at the bunching point by deconvolution. By construction, the unconfounded variation is mean independent of the confounders, but the deconvolution step requires full independence at the bunching point (or alternatively the weaker subindependence condition in schennach2019convolution).
If the function that characterizes selection is real analytic, then our results also allow for the identification of $\text{ATT}$s near the bunching point. If some bounds on the selection bias derivatives hold, then the $\text{ATT}$s can be identified further away from the bunching point. Specifically, we can identify the effect among those who take a given treatment value as compared to a counterfactual in which they take the bunching point value of treatment. In our application, it is thus possible to identify the birth-weight gains (or losses) if mothers who smoke a given amount were to quit smoking.
We propose estimators that use well-known building blocks. We estimate expectations and derivatives near the bunching point with local linear estimators fangijbels1992 and boundary densities using the estimator of pinkse2023estimates, both of which achieve interior rates of convergence at the boundary. The deconvolution step follows standard nonparametric methods, equivalent to a standard kernel density estimator using a special kernel. All building blocks can be implemented straightforwardly using off-the-shelf packaged software.
In a supplementary appendix, we also explore how control variables may be used to study heterogeneous treatment effects as well as to weaken the identification assumptions, which are then required to hold only conditional on controls. In particular, this allows the sign of selection bias near the bunching point to differ across different groups of observations. We discuss estimation when controls are discrete, continuous, or a mixed vector of both.
We apply our approach to data on smoking and birth weight from almond2005costs. After correcting for selection bias, we find that the effect of the first daily cigarette is a small loss of about 8 grams in birth weight (less than one-third of an ounce). Smoking five cigarettes per day causes a loss of 1.4 ounces (compare this to the average weight of a full-term newborn, which is 122 ounces in our analysis sample). These estimates confirm and strengthen the qualitative point in almond2005costs that smoking is not an important determinant of birth weight, but under different and weaker identification assumptions.
Bunching is a common phenomenon. It is often found at zero in variables naturally constrained to be non-negative, such as consumption goods,\footnote{E.g., number of tobacco products, alcoholic beverages, caffeinated drinks, sugary drinks, fast food meals, dining out meals, subscription services, supplements and vitamins, public transportation rides, books read, gym visits, doctor visits, trips, fuel usage amounts, expenditures on health, fitness, travel, vacations, education, childcare.} financial variables,\footnote{E.g., credit access, bequests, savings, emergency fund levels, retirement account contributions, mortgage balance, credit card debt, student loan debt, income from investments, expenditures on ads, charitable donations, HSA and FSA balances, life insurance coverage, number of trades.} time use,\footnote{Bunching is found for almost all time uses. Some examples include exercising, working, watching TV, using digital devices, doing homework, doing chores, volunteering, and commuting.} and neighborhood characteristics.\footnote{E.g., number of public transportation options or stops, retail stores, coffee shops, rental units, affordable housing units, vacant units, electric vehicle charging stations; length of biking lanes or walking paths; areas of green space, commercial districts, sports fields, parking lots.} Artificial constraints also can generate bunching, such as regulatory minimums\footnote {E.g., schooling time, wages, 401(k) contributions, coverage for auto insurance, nutritional standards for school meals, bank capital, bank deposit insurance, age started working, age started withdrawing from retirement accounts, age retired.} and maximums.\footnote{E.g., contribution size in 401(k), Roth IRA, HSA, FSA accounts, untaxed gifts, FHA loans, FDIC insurance, carbon emissions, liquor licenses, lot coverage, contributions to political campaigns, data usage, grades, absences from school, class size, commissions on sales.} Bunching also occurs at interior points, often due to kinks or notches in budget sets, social norms, and other restrictions.\footnote{E.g. income at tax brackets, hours worked at overtime rules, multiples of 5 or 10, 40 hours per week, car speeds at ticket thresholds, financial reporting around profit targets, energy consumption around utility billing tiers, pricing below psychological points (\$0.99), doctor visits at medical protocol numbers, hospital stay length at insurance payment thresholds.} Of course, small samples, coarse measurements, or attrition could make it impossible to implement this method in some of the examples above.
The use of bunching phenomena for identification began with saez10, followed by a large applied literature interested in identifying the effects of policies that lead to bunching in a manipulable variable at a policy threshold. Theoretical treatments of these approaches may be found in blomquist2021bunching, bertanha2023better, goff2020treatment and lu2024identifying. Our approach is more related to the literature initiated by caetano2015, where bunching for any reason on the treatment variable of a reduced-form causal model allows the testing of the model's identification conditions (see caetano2016discontinuity, caetano2018identifying, CCFN_Dummy, and khalil2022test). ccn_metrics developed the first strategy for identification of treatment effects under endogeneity in this setting, followed by CCNT, CCT and CMS_Placebo. Surveys of the bunching literature include kleven2016bunching,jalesyu,blomquist2023econometrics, and bertanha2024bunching. In this paper, we show for the first time that nonparametric identification with bunching is possible.
Section (ref) introduces our setting and presents a novel approach to identification of the average marginal treatment effect using the change-of-variables theorem from integration theory. We detail how the average marginal treatment effect at the bunching point can be identified in Section (ref), and how treatment effects can then be identified away from the bunching point in Section (ref). We turn to estimation in Section (ref), and in Section (ref) we present the application to the effects of smoking on birth weight. We conclude in Section (ref). Appendices contain proofs, examples and generalizations referred to in the text. A supplementary appendix contains extensions for the use of control variables as well as further plots and proofs.
In this section, we set up the identification problem in a neighborhood of the bunching point. Then, we relate the selection bias to the ratio of the density of the treatment and the density of the part of the outcome that varies with confounders.
Our setting is the standard potential outcomes framework, where observation $i$'s outcome depends on the value $x$ of a multivalued scalar treatment variable through the potential outcome function $Y_i(x).$ We observe the treatment value $X_i$ and the outcome $Y_i=Y_i(X_i)$. The support of $X_i$ includes a nondegenerate interval whose left boundary $\bar{x}$ exhibits bunching. In many relevant applications, $X_i$ is continuously distributed on the positive real line with the bunching point $\bar{x}=0$. The right panel of Figure (ref) depicts this case, while the left panel depicts a case in which $\bar{x}$ is in the interior of the support of $X_i$. Bunching on the right boundary of the support of $X_i$ can be accommodated by redefining $X_i$ as $2\bar{x} - X_i$.
We adopt the following notational conventions. For an arbitrary function $v\mapsto g(v),$ we denote the $k$-th derivative at $\tilde{v}$ as $g^{(k)}(\tilde{v})$. For the first derivative, we also use the notation $g'(\tilde{v}):=g^{(1)}(\tilde{v}).$ For a function $g(x)$, let $g(\bar{x}^+):= \lim_{x \downarrow \bar{x}} g(x)$. We let $g'(\bar{x}^+):=\lim_{x\downarrow \bar{x}}(g(x)-g(\bar{x}^+))/(x-\bar{x})$ denote a right-derivative at $\bar{x}$ defined with respect to the limit $g(\bar{x}^+)$. We show that this definition is equivalent to $g'(\bar{x}^+)=\lim_{x\downarrow \bar{x}} g'(x)$, provided that $g(\bar{x}^+)$ exists and $g$ is differentiable in a neighborhood above $\bar{x}$. We define higher order limit derivatives $g^{(k)}(\bar{x}^+)$ analogously. For other composite functions $(g\circ h)(x) := g(h(x))$, we let $g(h(\bar{x}^+)):= (g\circ h)(\bar{x}^+)$. For an arbitrary random variable $V_i,$ let $F_{V}(v)=\mathbb{P}(V_i\leq v)$ and $f_{V}(v)=F'_{V}(v)$ if the derivative exists. Analogously, for a set $S \subseteq \mathbbm{R},$ $F_{V|S}(v)=\mathbb{P}(V_i\leq v|S)$ and $f_{V|S}=F'_{V|S}(v)$ if and where the derivative exists. For example, $f_{V|S}(x)$ denotes $\frac{f_V(v)}{P(S)}$ for any $v \in S$. Define the sign of $v$ as $\text{sgn}(v)=\bm{1}(v\geq 0)-\bm{1}(v\leq 0)$. For a set $S \subseteq \mathbbm{R}$, we denote the image of the function $g$ over $S$ as $g(S)$. For simplicity, we refer to the “support of $V_i$” when we mean the support of the distribution of $V_i$.\\
Remark: (Friction around $\bar{x}$) The examples in Figure (ref) depict distributions with “perfect” bunching, in the sense that bunching is at a point mass in the distribution of $X_i$. By contrast, interior bunching is often somewhat diffuse around the bunching point because of optimization frictions, or because $X_i$ is measured with error. We abstract from this issue and assume that the researcher has a means of identifying the “bunched” observations $X_i=\bar{x}$. This is generally not problematic in settings with boundary bunching, and for interior bunching, measurement error can sometimes be eliminated by using administrative data (see goff2020treatment for an example). For general discussions of optimization frictions in bunching settings, see kleven2016bunching and bertanha2024bunching.
To define treatment effect parameters, we let the outcome $Y_i(\bar{x})$ that would occur if treatment were equal to $\bar{x}$ play the role of the “untreated” state. When bunching is at zero, this yields the familiar notation $Y_i(0)$. Let $\text{ATT}(x)$ denote the average effect of changing treatment from the bunching point to $x$, among those treated $X_i=x$: $$\text{ATT}(x):=\mathbbm{E}[Y_i(x)-Y_i(\bar{x})|X_i=x] = \mathbbm{E}[Y_i|X_i=x]-\mathbbm{E}[Y_i(\bar{x})|X_i=x],$$ where we have used $Y_i = Y_i(X_i)$ in the rightmost expression. This expression highlights the challenge of identifying the causal quantity $\text{ATT}(x)$, since $\mathbb{E}[Y_i(\bar{x})|X_i=x]$ is counterfactual and not directly observed for $x \ne \bar{x}$.
In our empirical application, $x$ measures cigarettes per day, and the bunching point is $\bar{x}=0.$ So, $\text{ATT}(x)$ measures the average birth-weight loss (or gain) mothers who smoke $x$ cigarettes incur for smoking that amount in comparison to a counterfactual where they do not smoke at all. About 80% of mothers in this application smoke zero cigarettes, leading to a case of boundary bunching as in the right panel of Figure (ref). Here, $\mathbb{E}[Y_i(\bar{x})|X_i=x]$ is the expected birth weight if mothers who smoke $x$ cigarettes per day were to quit. Due to confounders (e.g. drinking during pregnancy), this quantity can be expected to vary with $x$.
Local effects of the treatment around a value $x$ can be obtained by inspecting the derivative of the function $\text{ATT}(x)$. We define the average marginal effect near the bunching point, $\text{AME}_{\bar{x}}^{+},$ as the right derivative of $\text{ATT}(x)$ as $x$ approaches the bunching point from above: $$\text{AME}_{\bar{x}}^{+}:=\lim_{x\downarrow \bar{x}}\frac{\text{ATT}(x)}{x-\bar{x}}=\lim_{x\downarrow \bar{x}}\mathbb{E}\left[\frac{Y_i(x)-Y_i(\bar{x})}{x-\bar{x}}\middle| X_i=x\right].$$ When the bunching point is interior as in the left panel of Figure (ref), or the bunching point is on the right boundary of the support of $X_i$, one could define a similar $\text{AME}_{\bar{x}}^-$ parameter describing the left limit at $\bar{x}.$ We use the right derivative to define the treatment effects of interest because it fits our application and many others. In our application, $\text{AME}_{\bar{x}}^{+}$ is the birth weight loss (or gain) incurred by those who smoked just a little versus the counterfactual where they would not have smoked at all, expressed as a per-cigarette rate.
If the individual potential outcome functions $Y_i(x)$ are differentiable, and regularity conditions permitting the exchange of limits and expectations hold, one can interpret $\text{AME}_{\bar{x}}^{+}$ in terms of an average of the derivatives $Y_i'(x)$ of these dose-response functions, i.e. $$\text{AME}_{\bar{x}}^{+} = \lim_{x\downarrow \bar{x}} \text{AME}(x),$$ with $\text{AME}(x):=\mathbb{E}[Y_i'(x)|X_i=x]$. It is for this reason that we use the term marginal effect to refer to $Y_i'(x)$, in line with e.g., hoderleinmammen,imbensnewey,chiangsasaki. Other authors use the term partial effect (see e.g., sasaki2015, katosasaki2017), emphasizing the interpretation of $Y_i(x)$ as $g(x,U_i)$ for an underlying structural function $g$ over heterogeneity $U_i$ in potential outcomes, in which case $Y_i'(x)$ denotes the partial derivative of $g(x,U_i)$ with respect to $x$. A related quantity is the average causal response on the treated (ACRT) function, studied in callaway2024difference. Despite our use of the term marginal effect, our results do not require $Y'_i(x)$ to exist. Rather, we maintain weaker differentiability assumptions at the level of expectations of outcomes, which are sufficient to ensure that $\text{AME}_{\bar{x}}^{+}$ is well-defined.
Like with $\text{ATT}(x)$, identifying average marginal effects is challenging because the regression derivative of $Y_i$ on $X_i$ generally confounds the causal effect of treatment with a bias due to endogeneity:
where the first term above is, under regularity conditions, equal to the average marginal effect at $x,$ $AME(x)$. This equation depicts the fundamental problem of causal inference: the observed differences in mean outcomes among those with different treatment values combines both the causal effect of $x$ and selection bias. An analogous decomposition for $\text{AME}_{\bar{x}}^{+}$ implies that:
The first term above is the right limit of the derivative of the regression function, and the second term reflects endogeneity: those with different values of $X_i$ may have different mean values of $Y_i(\bar{x})$.
The following definitions will be useful in simplifying exposition throughout our analysis. First, we define $m(x)$ as the mean difference in observed outcomes at $x$ versus the boundary as $x \downarrow \bar{x}$:
Similarly, define:
which denotes the comparison of the counterfactual outcomes $Y_i(\bar{x})$ relative to the boundary as $x \downarrow \bar{x}$. Then, we write
where the equalities follow from Theorem (ref) below. Both $m$ and $s$ are defined relative to the limit at the bunching point, so that $m(\bar{x}^+)=s(\bar{x}^+)=0$. Note as well that since only the first term in the definitions of $m$ and $s$ depends on $x$, we have $m'(x)=\frac{d}{dx} \mathbb{E}[Y_i|X_i=x]$ and $s'(x)=\frac{d}{dx} \mathbb{E}[Y_i(\bar{x})|X_i=x],$ whenever these derivatives exist.
Note that the function $m$ and the quantity $m'(\bar{x}^+)$ are identified directly from observables. Thus, the identification challenge for the causal parameters in (ref) comes from the $s(\cdot)$ terms, which depend on unobserved counterfactuals.
The main insight of this paper is established in this section. Specifically, it is the observation that to identify the derivative of the counterfactual function $\mathbb{E}[Y_i(\bar{x})|X_i=x]$ for values of $x$ near the bunching point $\bar{x}$, it is sufficient to identify the distribution of the expected counterfactuals $E[Y_i(\bar{x})|X_i]$ for $X_i$ near $\bar{x}$. We then show in Section (ref) that bunching in $X_i$ makes it possible to identify this distribution, even though the counterfactual outcome $Y_i(\bar{x})$ is never observed for those with $X_i \ne \bar{x}$. The bias term $\lim_{x\downarrow \bar{x}}\frac{d}{dx} \mathbb{E}[Y_i(\bar{x})|X_i=x]$ in (ref) can then be identified.
To establish our first results, we assume the following:
Part (i) of Assumption (ref) states that $X_i$ is continuously distributed with a density on an interval to the right of $\bar{x}$, though this density need not exist everywhere (e.g. there can be multiple bunching points provided that they are well-separated). Part (ii) of Assumption (ref) says that a regression derivative exists, and the regression function and its derivative have a right limit at the bunching point. Both parts (i) and (ii) of Assumption (ref) are restrictions on the observable data that can in principle be verified empirically.
Part (iii) of Assumption (ref) requires that the selection function $s(x)$ be differentiable in a neighborhood above the bunching point. It also rules out the case in which there is no endogeneity when one approaches $\bar{x}$ from above, although, in this case, no correction for endogeneity is necessary. Note that caetano2015's test can be used in this setting to diagnose the presence of endogeneity needing a correction.
The first statement in part (iv) of Assumption (ref) requires that $\mathbb{E}[Y_i|X_i=x^+]=\mathbb{E}[Y_i(\bar{x})|X_i=\bar{x}^+]$. The second statement in part (iv) of Assumption (ref) states that on average, the treatment effects are sufficiently smooth near the bunching point. Among those with treatment levels near the bunching point, marginally small doses of the treatment should have, on average, only marginally small effects. In the smoking example, this condition states that among mothers who smoke very little, smoking is not so harmful that a small amount can cause a stark decline in the baby's health.
Part (iv) of Assumption (ref) is sufficient to guarantee that the $\text{AME}_{\bar{x}}^+$ is well defined, since $\text{ATT}(\bar{x}^+)=0$ implies that $\text{AME}_{\bar{x}}^+=\text{ATT}'(\bar{x}^+)$. A sufficient condition (though stronger than necessary) for part (iv) is that the $Y_i(x)$ are differentiable and uniformly bounded with probability one near the bunching point, which furthermore implies that $\text{AME}_{\bar{x}}^+=\mathbb{E}[Y_i'(\bar{x})|X_i=\bar{x}^+]$. The following proposition establishes the connection between the parameters of interest and the functions $m$ and $s,$ as presented in Equation (ref). All proofs are found in Appendix (ref).
We now move to the first main result of this paper. Part (iii) of Assumption (ref) implies that the counterfactual function $s(x)$ is differentiable and locally monotonic for $x$ in some neighborhood above the bunching point (without loss, we can take $I=(\bar{x},\bar{x}+\varepsilon),$ for some $\varepsilon < \min\{\varepsilon_1,\varepsilon_2,\varepsilon_3,\varepsilon_4\},$ and we note that $\varepsilon$ need never be known for our identification strategy). The local monotonicity and differentiability in $I$ allow us to apply the well-known change-of-variables formula from integration theory to derive the following result.
Note the absolute value $\left|s'(x)\right|$ in (ref): the right-hand side is always positive, but $s'(x)$ will be negative if there is negative selection. The textbook change-of-variables formula (see e.g. fremlin2011measure for a general formulation) states that $f_{u(X)}(t) = f_{X}(u^{-1}(t))/|u'(u^{-1}(t))|$ for any $t \in u^{-1}(I)$, given any function $u(x)$ that is differentiable and strictly increasing on an interval $I$ (and analogously for a strictly decreasing $u$). Then, the claim follows with $u(X)=s(X)$ and $t=s(x)$. Though the change-of-variables formula is a standard tool, a proof of Theorem (ref) is provided in Appendix (ref), which also establishes the local monotonicity of $s$ required for the change-of-variables result. Figure (ref) provides a visual illustration of the change-of-variables theorem for scenarios with high and low selection.
Let $\theta:=\text{sgn}\left(\lim_{x\downarrow \bar{x}}\frac{d}{dx}\mathbb{E}[Y_i(\bar{x})|X_i=x]\right)=\text{sgn}(s'(\bar{x}^+))$ be the sign of the selection bias as $x$ approaches $\bar{x}$ from the right. As an application of Theorem (ref), we have the following expression for the $\text{AME}_{\bar{x}}^{+}$:
The previous section established that identification of treatment effects may be possible if we can identify the sign of the selection bias at the boundary point, $\theta=\text{sgn}(s(\bar{x}^+))$, and the limit of the density of the selection bias variable, $f_{s(X)}(s(x)),$ as $x \downarrow \bar{x}$. In this section, we show how information at the bunching point may be used to obtain these quantities.
We start by noting that the treatment variable plays two distinct roles. First, it describes the dose taken by an individual, i.e. the number of cigarettes smoked in our application. Second, it tells us something about that individual, namely the fact that those with that treatment value are of the “type” that selected (or was selected into) that amount. Thus, for example, in the parameter $\text{ATT}(x)=\mathbb{E}[Y_i(x)-Y_i(0)|X_i=x],$ the $x$ in “$Y_i(x)$” describes the dose, and the $x$ in “$X_i=x$” describes the group that selected it. Henceforth, we separate the notation of the two concepts: the dose is the treatment variable, denoted $X_i$ as before, and the selection variable is denoted $X_i^*$.
In most cases, the selection variable $X_i^*$ is identical to $X_i.$ Indeed, this is precisely what gives rise to endogeneity: if $X_i^*$ is correlated with the potential outcomes $Y_i(x),$ then the correlation between $Y_i=Y_i(X_i)$ and $X_i=X_i^*$ reflects both the causal effect (i.e. the part that refers to the variation of the function $Y_i(x)$ with $x$ for a given $i$), and the selection function (i.e. the part that refers to the variation of $Y_i(x)$ across the $i$ with different values of $X_i^*$). This is the case here as well for values away from the bunching point because, when $X_i^*>\bar{x},$ there is no constraint on the treatment value, and so $X_i^*=X_i$. However, the bunching point is special in that multiple values of the selection variable $X_i^*$ occur simultaneously at the same treatment value $X_i=\bar{x}$.
A separation between “types” and their value of a variable that exhibits bunching is a common feature of the bunching literature saez10, klevenwaseem, blomquist2021bunching, bertanha2024bunching,bertanha2023better,ccn_metrics,goff2020treatment,CCNT,pollinger. The separation arises from constraints on individuals' choices that cause different types to all choose the common bunching point. The idea is that at the bunching point, selection breaks away from the dose, and observations with diverse selection values receive the exact same dosage amount. Thus, the bunching point affords the possibility of learning about the relationship between the selection variable and other variables without any interference from dose variation.
We write
which fits the case of bunching at the left boundary of the support of $X_i$. The right boundary case also fits this description, by redefining the treatment to be $2\bar{x}-X_i$. Interior bunching resulting from kinks in the budget function can also be adapted to fit Equation (ref), as we describe in Example (ref) (Section (ref)).
In general, one can think of $X_i^*$ as an index summarizing all observable and unobservable individual characteristics that determine the treatment. We return to $X_i^*$ in Section (ref), where we examine its possible economic interpretation in terms of structural primitives or reduced form quantities in specific models, and discuss its (non-)uniqueness. To aid comprehension in the following sections, we provide here a brief intuitive discussion of $X_i^*$ in the context of our empirical application to maternal smoking. In that setting, it is natural to think that, although all nonsmoking mothers share a value $X_i=0$ of the treatment, they may differ in the intensity of their preference towards not smoking or other factors that influence their choice. For instance, if $\rho_i$ is a parameter that governs the person's relative preference towards smoking, there may exist a value $\bar{\rho}$ and a smooth function $h$ such that $X_i = h(\rho_i)$ when $\rho_i \ge \bar{\rho}$ (i.e. when the person's preference towards smoking is sufficiently high), and $X_i=0$ otherwise, where $P(\rho_i < \bar{\rho})>0$ (i.e. some mothers strictly prefer not to smoke).
Intuitively, in our method $X^*_i$ allows us to track observations at $X_i=\bar{x}$ based on how selected they are relative to the observations near the bunching point on the positive side. The key assumptions about $X_i^*$ will be that the smoothness conditions from Section (ref) can be extended to values of $X_i^*$ around $\bar{x},$ so that those observations away from the bunching point with $X_i$ near $\bar{x}$ are comparable to the observations with $X_i^*=\bar{x}$. This allows us to substitute the limit of the density of the selection above the bunching point in Theorem (ref) into the density of the selection at the bunching point. The following sections then show that the sign of the selection bias is identified (Section (ref)), and that the density may be obtained by a deconvolution of the density of the outcome near the bunching point from the density of the outcome at the bunching point (Sections (ref) and (ref)). Finally, in Section (ref) we discuss the nature of $X_i^*$ when it arises from choice models, how it may be artificially constructed under some conditions, and the invariance of the identification results to monotonic transformations of $X_i^*$.
We begin by extending the definition of $s(x)$ to leverage the notation $X_i^*,$ as $s(x):=\mathbb{E}[Y_i(\bar{x})|X_i^*=x]-\mathbb{E}[Y_i(\bar{x})|X_i^*=\bar{x}^+]$. Since $X_i=X_i^*$ when $X_i > \bar{x}$, this coincides with the function $s$ defined in Section (ref) for all $x > \bar{x}$. For $x\leq \bar{x},$ $X_i=\bar{x}$ but $s(x)$ may vary with the following restriction:
Assumption (ref) states that if $s(x)$ is increasing in a positive neighborhood around the bunching point, then $s(x)<0$ for all values of $x<\bar{x}$. Conversely, if $s(x)$ is decreasing in a positive neighborhood around the bunching point, then $s(x)>0$ for all values of $x< \bar{x}$. Specifically, either $\mathbb{E}[Y_i(\bar{x})|X_i^*=x']<\mathbb{E}[Y_i(\bar{x})|X_i^*=\bar{x}^+]<\mathbb{E}[Y_i(\bar{x})|X_i^*=x'']$ for all $x'< \bar{x}$ and all $x''>\bar{x}$ in a neighborhood right above the bunching point, or the reverse ordering is true. Intuitively, the selection function $\mathbb{E}[Y_i(\bar{x})|X_i^*=x]$ remains uniformly either above or below its value at the bunching point $\mathbb{E}[Y_i(\bar{x})|X_i^*=\bar{x}]$, for all $x$ to the left of the bunching point, and is unrestricted to the right of the bunching point.
In the smoking example, suppose that among mothers smoking positive amounts, those who smoke more are negatively selected relative to those who smoke less (i.e., smoking more is associated with worse untreated outcomes). Then, Assumption (ref) states that the nonsmoking mothers would have even higher birth weights. This needs to hold only on average for every selection value, so some individual nonsmoking mothers could have lower birth weights than some smoking mothers.
Assumption (ref) holds trivially if $x\mapsto\mathbb{E}[Y_i(\bar{x})|X_i^*=x]$ is monotonic, in which case $s(x)$ is also monotonic. However, the assumption is weaker than monotonicity. Figure (ref) illustrates two examples of non-monotonic functions that satisfy Assumption (ref), one case with $s'(\bar{x})>0,$ and another with $s'(\bar{x})<0,$ respectively. Note that $s(\bar{x}^+)=0$ by definition, and Assumption (ref) does not constrain the behavior of $s(x)$ for positive $x,$ outside of a neighborhood of the bunching point.
The following lemma establishes that the sign of the selection bias is equal to the sign of the discontinuity of the expected outcome at the bunching point.
Lemma (ref) relates to caetano2015's test of exogeneity, which is based on the discontinuity of the outcome at the bunching point. If the sign of the discontinuity is not zero, then the test rejects the exogeneity of $X_i$. However, in our setting, we can do more than just test exogeneity, as the same discontinuity also allows us to sign the selection bias within an interval of $\bar{x}$.
The intuition of Lemma (ref) can be obtained from Figure (ref). Suppose that we observe a positive discontinuity of the outcome at the bunching point. Then, $\mathbb{E}[Y_i(\bar{x})|X_i^*=x]$ must be below $\mathbb{E}[Y_i(\bar{x})|X_i^*=\bar{x}^+]$ for at least some $x<0$. By Assumption (ref), we then know that all $\mathbb{E}[Y_i(\bar{x})|X_i^*=x]$ must be below $\mathbb{E}[Y_i(\bar{x})|X_i^*=\bar{x}^+]$ for all $x<\bar{x},$ so $s(x)<0$ for all $x<\bar{x}$. We must therefore be in a situation akin to the left plot. It follows that $\mathbb{E}[Y_i(\bar{x})|X_i^*=x]$ for $x$ slightly above $\bar{x}$ are all above $\mathbb{E}[Y_i(\bar{x})|X_i^*=\bar{x}^+],$ and thus $s'(\bar{x}^+)>0$.
We next concern ourselves with the elimination of the unknown interval $I$ from the formula for $\text{AME}_{\bar{x}}^{+}$ in Theorem (ref). Precisely, we need to substitute the density $f_{s(X)|I}$ with the density $f_{s(X^*)|X=\bar{x}}.$ In the following section, we then show that $f_{s(X^*)|X=\bar{x}}$ may be identified.
In Theorem (ref), we established the existence of $f_{s(X^*)}(v)$ on $I$. Part (i) of Assumption (ref) extends this result, requiring that the density exists also to the left of the bunching point. It is sufficient (but not necessary) for Assumption (ref) (i) that $x\mapsto \mathbb{E}[Y_i(\bar{x})|X_i^*=x]$ is monotonic and $X_i^*$ has a density for $(-\infty,\bar{x}]$. More generally, if $X_i^*$ has a density on $(-\infty,\bar{x}]$, then part (i) holds provided that the set of $x$ such that $s(x)=t$ has Lebesgue measure zero, for any $t \in s^{-1}((-\infty,\bar{x}])$. Thus part (i) could be thought of as a consequence of $X_i^*$ being continuously distributed on the left side of $\bar{x}$, requiring no restrictive assumptions on selection. We state condition (i) above because it is weaker than this.
Item (ii) of Assumption (ref) implies that the observations with $X_i^*=\bar{x}$ are comparable on average to the observations with $X_i^*=\bar{x}^+$. In the smoking example, it is equivalent to saying that if the mothers who smoke very little (i.e. those with $\rho_i$ slightly larger than $\bar{\rho}$) were to stop smoking, their outcomes would be very similar to the outcomes of nonsmoking mothers who are indifferent between not smoking and smoking a bit (i.e., those with $\rho_i=\bar{\rho}$). Note that the claim is not that mothers near the bunching point are comparable to mothers at the bunching point, since the mothers at the bunching point can also include those who strictly prefer not to smoke.
Assumption (ref) (ii) states that $\mathbb{E}[Y_i(\bar{x})|X_i^*=\bar{x}] = \mathbb{E}[Y_i(\bar{x})|X_i^*=\bar{x}^+],$ and thus it implies that $s(\bar{x})=0$. Moreover, it allows us to extend part (iii) of Assumption (ref) to the interval $[\bar{x},\bar{x}+\varepsilon_3),$ and note that, since $\lim_{x \downarrow \bar{x}} \frac{d}{dx}\mathbb{E}[Y_i(\bar{x})|X_i=x] \ne 0,$ we have $s'(\bar{x})\ne 0$. We then obtain the following extension of Theorem (ref):
Equation (ref) shows that, under Assumptions (ref)-(ref), identifying $\text{AME}_{\bar{x}}^{+}$ reduces to the problem of identifying $\theta$ and $f_{s(X^*)|X= \bar{x}}(0).$ Since $\theta$ is identified by Lemma (ref) in Section (ref), the only remaining piece is the identification of the density of $s(X_i^*)$ at the bunching point, which we tackle in the following section.
Figure (ref) illustrates how we apply the change-of-variables theorem in Theorem (ref). The left panel is identical to the left panel in Figure (ref), except that the densities are with respect to the selection variable $X^*_i,$ and dashed lines depict the parts of these densities that are not identifiable. The densities can both be identified only at $x=\bar{x}$. The right panel shows how identification is achieved. The red line is $f_{s(X^*)|X=\bar{x}},$ which on $(-\infty,\bar{x}]$ corresponds to the solid red line in the left panel divided by the probability of bunching, $F_X(\bar{x})$. Equivalently, we rescale the solid part of the blue density in the left panel by the same amount.
Finally, we turn to the identification of $f_{s(X^*)|X=\bar{x}}(0)$. Define the random variable $$\epsilon_i = Y_i - \mathbbm{E}[Y_i|X^*_i]=Y_i-\text{ATT}(X_i)+\mathbb{E}[Y_i(\bar{x})|X_i^*],$$ which is the unconfounded variation in the outcome, i.e. the idiosyncratic part of the outcome that remains after we eliminate the mean effect of treatment and the part of the outcome determined by selection. Note that for $X_i>\bar{x},$ $\epsilon_i=Y_i - \mathbbm{E}[Y_i|X_i]$ is identified.
We can write with probability one that $Y_i=y_{\bar{x}}+\text{ATT}(X_i)+s(X_i^*)+\epsilon_i,$ where the constant $y_{\bar{x}}:=\mathbbm{E}[Y_i(\bar{x})|X_i^*=\bar{x}]$ (using that $s(X_i^*)=\mathbbm{E}[Y_i(\bar{x})|X^*_i]-y_{\bar{x}}$ by Assumption (ref)). This shows that when $X_i > \bar{x}$, there is variation in $Y_i$ coming both from $\text{ATT}(X_i)$ and from $s(X_i^*)$. However, exactly at $X_i=\bar{x},$ $\text{ATT}(\bar{x})=0,$ and thus
with probability one. Thus, the outcome variation at the bunching point reflects the variation of $s(X_i^*)$ unconfounded by the variation of $\text{ATT}(X_i)$. Unfortunately, the distribution of the outcome at the bunching is a convolution of the density we want to identify, $f_{s(X^*)|X=\bar{x}},$ and the distribution of the idiosyncratic remainder, $y_{\bar{x}}+\epsilon_i$ (i.e., $f_{Y|X=\bar{x}}=f_{s(X^*)|X=\bar{x}}\ast f_{y_{\bar{x}}+\epsilon|X=\bar{x}}$). The following conditions allow us to deconvolve these distributions.
Item (i) of Assumption (ref) states that the outcome has a density at the bunching point. Item (ii) says that the idiosyncratic variation in the outcome near the bunching point is distributed similarly to the idiosyncratic variation in the outcome of those at the bunching point with $X_i^*=\bar{x}$. If $Y|X=x$ also has a density for all $x$ near the bunching point, then item (ii) of Assumption (ref) is equivalent to $Y_i(\bar{x})|X^*_i=x\rightarrow_{d}Y_i(\bar{x})|X_i^*=\bar{x}$ as $x\downarrow \bar{x}$, i.e. that the distribution of $Y_i(\bar{x})$ does not change abruptly as $X_i^*$ approaches $\bar{x}$.
Note that $\epsilon_i=Y_i-\mathbb{E}[Y_i|X_i^*]$ is mean independent of $X_i^*$ by construction. Part (iii) of Assumption (ref) extends this into full independence, at least at the bunching point. In our application, this condition says that, among the nonsmoking mothers, after removing the mean of the birth weight that is due to the selection variable $X_i^*$, the remainder is independent of the relative preference for smoking (or whatever else determines $X_i^*$). Part (iii) of Assumption (ref) may be substituted with a weaker but less intuitive condition known as subindependence, which has long been used in the deconvolution literature (see discussion in e.g. Hamedanisubindependence), and was formalized in schennach2019convolution. In our context, the relevant subindependence condition translates to: for all $t \in \mathbbm{R},$ $$\mathbb{E}[e^{\textbf{i}t(s(X_i^*)+\epsilon_i)}|X_i=\bar{x}]=\mathbb{E}[e^{\textbf{i}t s(X_i^*)}|X_i=\bar{x}]\cdot \mathbb{E}[e^{\textbf{i}t\epsilon_i}|X_i=\bar{x}],$$ where $\textbf{i}=\sqrt{-1}$. schennach2019convolution shows that subindependence is no “stronger” than mean independence, in the sense that subindependence imposes the same number of restrictions on the data-generating process as mean independence does. Nevertheless, mean independence does not imply subindependence.
Item (iii) of Assumption (ref) implies that $f_{y_{\bar{x}}+\epsilon|X=\bar{x}}=f_{y_{\bar{x}}+\epsilon|X^*=\bar{x}}.$ Item (ii) of Assumption (ref) implies that $f_{y_{\bar{x}}+\epsilon|X^*=\bar{x}}=f_{Y|X=\bar{x}^+}.$ Therefore, we can write $f_{Y|X=\bar{x}}=f_{s(X^*)|X=\bar{x}}\ast f_{Y|X=\bar{x}^+},$ which is a standard convolution problem with a well-known closed form solution using the Fourier representation (see e.g. schennachmeasurement),
Evaluating the density at zero yields the desired identification result.
Figure (ref) illustrates the components of the deconvolution. The outcome distributions were specified as a sequence of normal distributions $f_{Y|X=x},$ depicted in solid blue. Since, for $X_i>0,$ $Y_i=\text{ATT}(X_i)+y_{\bar{x}}+\epsilon_i,$ the observable solid blue densities $f_{Y|X=x}$ converge to $f_{y_{\bar{x}}+\epsilon|X=\bar{x}^+}$ as $x\downarrow \bar{x}$. So, the dotted blue density $f_{Y|X=\bar{x}^+}=f_{y_{\bar{x}}+\epsilon|X=\bar{x}^+}.$ The dashed red line is the unobserved density of $s(X_i^*)$ at the bunching point, $f_{s(X^*)|X=\bar{x}}$, which is specified here as normal as well. The solid black line is the observed density of the outcome at the bunching point, which is the convolution of the red dashed density ($f_{s(X^*)|X=\bar{x}}$) and the blue dotted density ($f_{\epsilon|X=\bar{x}^+}$). Since the images were produced using the actual convolution of the depicted densities, the dimensions illustrate the real relationship between these densities.
Our final identification result is derived from the combination of Equation (ref) with Lemmas (ref) and (ref):
The expression is familiar in that the treatment effect is calculated by correcting the outcome variation with a scaled inverse Mills Ratio term, as is usually seen in the censoring and sample selection literatures, where some of the model components are truncated or missing below a certain threshold. In our setting the scaling factor makes use of the full distribution of the outcome observable from the data both at the bunching point and as one approaches it.
Note that when there is no endogeneity at $X_i=\bar{x}$, $s'(\bar{x})=0$. Therefore, $\theta=0$ and Equation (ref) still holds. This means that the right-hand side of (ref) can be used to identify $\text{AME}_{\bar{x}}^+$ while remaining agnostic about endogeneity
In the Supplementary Appendix (ref), we show how identification may be obtained in the presence of controls. Specifically, all the assumptions required for identification may be done conditional on a vector of controls $Z_i,$ so that the requirements may effectively be weaker. In particular, the treatment and selection effects may change direction for different subgroups, and we can identify the $\text{AME}_{\bar{x}}^+(Z_i)$ to study heterogeneous treatment effects. We propose estimators in the case with discrete controls (Section (ref)), continuous controls (Section (ref)), and when the vector of controls is large and with mixed continuous and discrete controls (Section (ref)).
While the existence of a selection variable $X_i^*$ satisfying Equation (ref) is without loss of generality, the identification result of Theorem (ref) relies upon assumptions made about $X_i^*$. To guide researchers in assessing the plausibility of these assumptions, in this section we illustrate how $X^*_i$ relates to well-defined empirical quantities.
We begin with a parametric choice model that leads to a particularly simple expression for $X_i^*$.
In Appendix (ref), we extend this example beyond the isoelastic family. We show that in a class of utility maximization models that feature a scalar preference parameter $\rho_i$, we can write $X_i^*=h(\rho_i)$ where $h$ is a strictly increasing and differentiable function. The intuition is that with scalar heterogeneity, if $\rho_i$ pins down $i$'s marginal rate of substitution between $X_i$ and the numeraire good when $X_i=0$, then it also pins down $X_i$.
In Appendix (ref), we show that any monotonic differentiable transformation of $X_i^*$ yields precisely the same constructive estimand that identifies the parameter $AME_{\bar{x}}^+$ in Theorem (ref). Specifically, we prove that if $X_i^*=h(\rho_i)$ for any strictly increasing and differentiable $h$, the assumptions that establish Theorem (ref) can be made directly on $\rho_i=h^{-1}(X_i^*)$, rather than on $X_i^*$. Our result does not require that $\rho_i$ be observable, nor that the function $h$ be known to the econometrician, and holds whether or not one posits a utility maximization model. Thus, to make use of Theorem (ref) to identify $AME_{\bar{x}}^+$, one needs only to believe that there exists some $\rho_i$ that satisfies the Assumptions (ref)-(ref), and some monotonic and differentiable transformation $h$ such that $X_i = \max\{\bar{x},h(\rho_i)\}$. This result allows the researcher to make assumptions about $\rho_i$ rather than about the more abstract object $X_i^*$ without any loss of generality, while retaining the same identification equations and estimators.
In Appendix (ref), we consider a general utility maximization framework in which heterogeneity in individuals' choices are not necessarily explained by a single scalar parameter. We show that $X_i^*$ can be defined for those individuals with $X_i=\bar{x}$ as the marginal utility (after substituting budget constraints and profiling the utility function over any additional choice variables), evaluated at the bunching point. Those individuals who are indifferent between $X_i=\bar{x}$ or an amount slightly larger than the bunching point have $X_i^*=\bar{x}$ exactly, while those who strictly prefer $X_i=\bar{x}$ to a value just above it have $X_i^* < \bar{x}$. One can then translate the key aspects of Assumptions (ref)-(ref) in terms of comparisons of observations with different degrees of indifference towards veering away from the bunching point.
Finally, in Appendix (ref), we show that $X_i^*$ can be constructed without the need for any underlying choice model. As discussed in the introduction, bunching has been observed in some examples where the treatment variable is not a clear function of individual choices (e.g., caetano2018identifying study of the effects of neighborhood crime), and this type of construction can be useful in such cases. Concretely, if $\epsilon_i \mathbin{ \mathpalette{\@indep}{} } X_i $ for $X_i>\bar{x},$ and $s(x)$ is sufficiently smooth, then $s$ can be extrapolated to $x\leq \bar{x}$. In this case, an $X_i^*$ satisfying $\epsilon \mathbin{ \mathpalette{\@indep}{} } X_i^*|X_i=\bar{x}$ can always be constructed to satisfy the identification restrictions.
While the above results are motivated by settings in which bunching occurs at the boundary of the support of the treatment variable, the following example discusses how $X_i^*$ emerges naturally in settings with interior bunching at a kink. In such settings, there is no hard constraint that $X_i \ge \bar{x}$. Rather, Equation (ref) emerges from a discontinuous change in individuals' incentives at $X_i=\bar{x}$.
Given identification of $\text{AME}^+_{\bar{x}}$ demonstrated in Section (ref), we consider now how global effects $\text{ATT}(x)$ may be identified by extrapolating the information available near the bunching point. The extension requires a sufficient degree of smoothness of the counterfactual function near $\bar{x},$ which is guaranteed by the following assumption.
Functions that are real analytic on $I$ are infinitely differentiable functions whose Taylor series around any point $x \in I$ converges pointwise to the function in a neighborhood of $x$. This class includes all functions that locally behave like polynomial, exponential, trigonometric, hyperbolic, logarithmic, or inverse trigonometric functions, as well as compositions, ratios, and roots of these. A sufficient condition for Assumption (ref) is that the observable function $\mathbbm{E}[Y_i|X_i=x]$ and the $\text{ATT}(x)$ function are both analytic on $I$. This may be a more appealing argument than reasoning about the properties of the selection function $s(x),$ since $\mathbb{E}[Y_i|X_i=x]$ is identifiable, and it may be plausible to assume that the dose-response function $\text{ATT}(x)$ is sufficiently smooth in some applications.
Recall that $\text{ATT}(x)=m(x)-s(x)$. The following theorem establishes the desired local extrapolation.
The second part of Theorem (ref) offers a practical strategy for finite approximations to Equation (ref). Explicitly, for a suitably large $K,$ $\text{ATT}(x)\approx m(x)-\sum_{k=1}^K s^{(k)}(\bar{x}^+)\cdot \frac{(x-\bar{x})^k}{k!} $. The error of the $K$-th degree approximation does not exceed $\sup_{x \in I_{\varepsilon}}|s^{(K+1)}(x)|\cdot \varepsilon^{K+1}/(K+1)!$. Since $(K+1)!$ has supra-exponential growth, even if $\varepsilon$ and the high-order derivatives are very large, the approximation error decays quickly.
Equation (ref) indicates that if all the derivatives $s^{(k)}(\bar{x}^+)$ are identifiable, then the $\text{ATT}(x)$ is identifiable in a neighborhood of $\bar{x}$. This implies that extrapolations near the bunching point are possible, provided $\theta$ is identified and the $s^{(k)}(\bar{x}^+)$ are known for all $k \ge 1$. By differentiating Equation (ref) with respect to $x$, we can see that this is indeed possible since the density $f_{s(X)|I}$ (and hence its derivatives) is identified on $s(I_{\epsilon})$.
Corollary (ref) guarantees the identification of the $\text{ATT}(x)$ in a neighborhood of the bunching point. In practice, finite approximations may be used to approximate the value. For example, a first-degree approximation is simply $$ \text{ATT}_{\bar{x}}(x)\approx \mathbb{E}[Y_i|X_i=x]-\mathbb{E}[Y_i|X_i=\bar{x}^+]-2\pi\theta \left(\int \frac{\mathbb{E}[e^{\textbf{i}\xi Y_i}|X_i=\bar{x}]}{\mathbb{E}[e^{\textbf{i}\xi Y_i}|X_i=\bar{x}^+]}d\xi\right)^{-1} \frac{f_X(\bar{x}^+)}{F_X(\bar{x})}(x-\bar{x}),$$ and a second-degree approximation adds the term
where $s'(0)$ is the correction term in Equation (ref), and $f'_{s(X^*)|X=\bar{x}}(0)/f_{s(X^*)|X=\bar{x}}(0)$ is equal to $\left(\int \mathbb{E}[e^{\textbf{i}\xi Y_i}|X_i=\bar{x}^+]^{-1}{\mathbb{E}[e^{\textbf{i}\xi Y_i}|X_i=\bar{x}]}d\xi\right)^{-1}\left(\int \textbf{i}\xi\mathbb{E}[e^{\textbf{i}\xi Y_i}|X_i=\bar{x}^+]^{-1}\mathbb{E}[e^{\textbf{i}\xi Y_i}|X_i=\bar{x}]d\xi\right)$. This expression may seem complex but, in practice, one would have already estimated $f_{s(X^*)|X=\bar{x}}(0)$ for a first-degree approximation, and standard deconvolution packages often automatically provide the first derivative $f_{s(X^*)|X=\bar{x}}'(0)$ at the same time.
One limitation of Theorem (ref) and Corollary (ref) is that the interval $I_\varepsilon$ could in principle be quite small. A sufficient condition for the ATT to be defined far away from the bunching point is that the higher order derivatives decay suitably fast with $k$.
Corollary (ref) specifies “how far” one can extrapolate from the derivatives of $s^{(k)}(\bar{x})$ to obtain $\text{ATT}(x)$. Specifically, if the $|s^{(k)}(\bar{x})|$ are bounded by $M^{-k}\cdot k!$ uniformly over $k$ for some $M,$ then the $\text{ATT}(x)$ identification can be extrapolated as far as $\bar{x}+M$.
For a sample $\{(Y_i,X_i)', i=1,\dots, n\},$ the average marginal effect at the bunching point may be estimated following Equation (ref). Specifically, we use the following formulas:
where $\hat{m}'(\bar{x}^+)$ is an estimator of $\lim_{x\downarrow\bar{x}}\frac{d}{dx}\mathbb{E}[Y_i|X_i=x],$ and $\hat{\theta}=\text{sgn}(\hat{\mathbb{E}}[Y_i|X_i=\bar{x}^+]-\hat{\mathbb{E}}[Y_i|X_i=\bar{x}])$. A first-degree approximation following Corollary (ref) uses the estimator
and analogously for a second-degree approximation. All the components of these formulas are standard objects frequently studied in econometrics. We discuss next how each component can be estimated.
The terms $\hat{F}_X(\bar{x})$ and $\hat{\mathbb{E}}[Y_i|X_i=\bar{x}]$ may be estimated with simple averages:
The terms $\hat{\mathbb{E}}[Y_i|X_i=\bar{x}^+]$ and $\hat{m}'(\bar{x}^+)$ are standard non-parametric regression boundary quantities. Estimation of these objects has been extensively researched in the statistics literature on local polynomial estimators, and in the Regression Discontinuity Design and Regression Kink Design literatures in economics. In line with classical methods in this literature and with the vast majority of applications in boundary regression estimation, we propose using a local linear regression of $Y_i$ onto $X_i$ at $X_i=\bar{x},$ using only observations such that $X_i>\bar{x},$ for its superior properties of bias reduction and variance control at the boundary over other methods.\footnote{See ruppert1994multivariate and fan2018local, and also cheruiyot2020local. See imbens2019optimized and citations therein for recent proposals which may be superior to local linear estimators.} The intercept coefficient of this regression is $\hat{\mathbb{E}}[Y_i|X_i=\bar{x}^+],$ and the slope coefficient is $\hat{m}'(\bar{x}^+).$ This may be executed using any package for local linear regression available in standard statistical software (R, STATA, etc.).
Explicitly, for a bandwidth $h_1>0$ and a kernel function $k_1,$\footnote{The triangular kernel $k_1(\nu)=(1-|\nu|),$ where $|\nu|\leq 1$ is recommended for boundary regressions such as this (cheng1997automatic).} solve the problem
then $\hat{\mathbb{E}}[Y_i|X_i=\bar{x}^+]=\hat{b}_0,$ and $\hat{m}'(\bar{x}^+)=\hat{b}_1.$ This estimator has a closed-form expression, which is commonly found in nonparametric econometrics textbooks, e.g. li2007nonparametric. Note that the term $\hat{E}[Y_i|X_i=x]$ is not a boundary quantity, but it may be estimated analogously, with a local linear regression of $Y_i$ onto $X_i$ at $x,$ using only observations such that $X_i>\bar{x}.$
The term $\hat{f}_X(\bar{x}^+)$ is a boundary density. As in the case of nonparametric boundary regression discussed above, the tendency for higher bias in this scenario necessitates the use of corrective methods, such as the use of local polynomial estimators. We recommend the approach recently proposed in pinkse2023estimates,\footnote{Other estimators of boundary densities include hjort1996locally, loader1996local, cheng1997automatic, zhang1998kernel, bouezmarni2010nonparametric and cattaneo2020simple.} which has two important properties which are of great value in our case, and which are not found in other estimators currently available. First, this estimator achieves the same rates of bias convergence at the boundary that is normally achieved in interior points. Second, the density estimator is never negative, a situation which would be complicated to address in our case. Additionally, the estimators have simple closed-form expressions, requiring only the choice of a bandwidth tuning parameter, $h_2$.
Following pinkse2023estimates, let $L_{X}(x)=\log f_{X}(x).$ We begin by estimating $L'_{X}(\bar{x}^+)$ as\footnote{This estimator is derived from applying Example 1 with $z=0$ to Equation (2) in pinkse2023estimates.}
This, then, allows us to estimate the density at the boundary as
which can be calculated for many standard positive kernel functions $k_2$. For example, as in Example 5 of pinkse2023estimates, when $k_2$ is the Epanechnikov kernel $k_2(\nu)=3/4(1-\nu^2)$ (which is the kernel recommended for boundary estimation in that paper), the denominator is equal to
This estimator is available in packaged form in standard statistics software and can be implemented by simply restricting the sample to observations such that $X_i>\bar{x}$ and then using the package to estimate the density of $X_i$ at $X_i=\bar{x}.$ Incidentally, the same package also provides the estimator of the derivative $f_{X}'(\bar{x}^+),$ that can be used in the second-order approximation of the $\text{ATT}(x)$ (see Section (ref)).
The final term $\hat{f}_{s(D^*)|X=0}(0)$ is a standard deconvolution estimator. We follow the estimator described in schennachmeasurement, which is the focus of an extensive literature, although there are many alternative proposals which are also referenced therein.
We first write $\hat{\mathbb{E}}[e^{\textbf{i}\xi Y_i}|X_i=\bar{x}^+]$ as a local linear regression of $e^{\textbf{i}\xi Y_i}$ onto $X_i$ at $X_i=\bar{x}$ using only observations such that $X_i>\bar{x}.$ To do this, for a matrix $\textbf{x}$ with rows $(1,(X_i-\bar{x}))'$ and a diagonal matrix $\textbf{k},$ with diagonal elements $k_3((X_i-\bar{x})/h_3)\bm{1}(X_i>\bar{x}),$ where $k_3$ is the triangular kernel, and $\textbf{e}_1=(1,0)',$ define the vector $A(\xi)=(e^{\textbf{i}\xi Y_1},\dots,e^{\textbf{i}\xi Y_n})',$ and program the function
This is then imputed into a standard convolution estimator such as, for example:
with
where $\phi_{k_4}(h_4\xi)=\int k_4(\nu)e^{\textbf{i}h_4\xi \nu}d\nu$ is the Fourier transform of the kernel $k_4$ evaluated at $h_4\xi.$
The nonparametric estimators just described require the choice of the bandwidth tuning parameters: $h_1,h_2, h_3$ and $h_4$, which modulate the bias-variance trade-off. This choice is rather important, and the subject of a great deal of interest in the nonparametrics estimation literature. At this stage, our recommendation is that if an optimal method for bandwidth selection exists for the specific estimator used at a given step, then it should be used.\footnote{For the selection of $h_1$ and $h_3$, ruppert1995effective propose an optimal bandwidth estimator for the local linear regression, and this or similar approaches for bandwidth selection are usually offered in standard local linear regression packages. There are many proposals for improvement of bandwidth selection in the Regression Discontinuity Design literature which may be adapted to this context, see, e.g. imbens2012optimal, arai2016optimal, arai2018simultaneous and calonico2020optimal. For $h_2,$ the optimal bandwidth is $h = (72/(nf_X(0)^+\beta_2^2))^{1/5}$, which may be calculated following Example 6 in pinkse2023estimates. $\beta_2$ can be estimated using a pilot estimate of $\hat{f}_X(0)_+,$ and both these terms are then added to the formula of the optimal bandwidth. Nevertheless, although theoretically sound, this method has not yet been studied. Thus, in practice, we recommend testing several bandwidths around this benchmark and looking for robustness of the results. For $h_4,$ consider the several approaches studied in delaigle2004practical.} However, it is possible that the optimal bandwidths for $\widehat{\text{AME}}_{\bar{x}}^+$ are not the optimal bandwidths for each of the separate components.
Additionally, there is an interest in the use of bias correction techniques for inference in the Regression Discontinuity Design literature which may have relevance in this context as well (e.g. calonico2014robust, noack2024bias, he2020wild, armstrong2020simple and citations therein). This is because, if optimal bandwidths are used, $\widehat{\text{AME}}_{\bar{x}}^+$ will likely be asymptotically biased. We leave these questions for future research.
We apply our method to estimate the marginal effect of maternal smoking during pregnancy on birth weight. This question is important for both economics and epidemiology, given that maternal smoking during pregnancy is recognized as a critical modifiable risk factor for low birth weight (almond2005costs). Low birth weight not only results in immediate societal costs, but also has significant long-term consequences for children's later-life outcomes (black2007cradle).
Our application uses the data set from almond2005costs, which is also used in caetano2015. These data, from the U.S. National Center for Health Statistics, include both maternal cigarettes smoked daily during pregnancy (our treatment) and birth weight in grams (our outcome) for over 430,000 mother-child pairs. These data also include many additional covariates/controls which almond2005costs use to estimate an effect of maternal smoking on birth weight of around -200 grams, assuming selection on observables. caetano2015 uses these data to illustrate the discontinuity test, showing that selection-on-observables does not seem to be a valid assumption using almond2005costs's very detailed control specification.
We drop premature births (gestation < 36 weeks) as well as birth weight outliers (a small subset of observations with very high, >6 kg, or very low, <1 kg, full-term birth weights) from the data.\footnote{Our estimates barely change when we extend the sample to include premature babies and outliers.} About 81% of the mothers in our analysis sample smoke zero cigarettes daily, about 11% smoke between 1 and 10 cigarettes, and 99.95% smoke 40 cigarettes or less. Figure (ref) shows $\mathbb{E}[Y_i|X_i=x]$, the average birth weight among mothers smoking different amounts in our sample. The evidence of a discontinuity in $\mathbb{E}[Y_i|X_i=x]$ at $x=0$ is clear. While the average birth weight among mothers who smoke zero cigarettes is 3,499 grams, the average birth weight for mothers who smoke one cigarette is 3,338 grams, with the analogous quantity for mothers smoking 2--5 cigarettes ranging between 3,278 and 3,330. The rich controls in almond2005costs can account for only 55 out of the 161 grams (3,499-3,338) difference in birth weight between the children of mothers who smoke zero versus one cigarette. Thus, there remains a lot of “room”---106 grams---for both the treatment
effect of cigarettes and selection on unobservables to explain this birth weight discontinuity.\footnote{The numbers cited in this paragraph are very similar to those reported in caetano2015. That paper removes neither premature nor “outlier” births from the analysis sample. Similarly, including premature and outlier births in the causal analysis yields similar estimates of $\text{AME}_{0}^{+}$ and $\text{ATT}(x).$ Indeed, the full-sample results tend to be smaller in magnitude than those reported in Table (ref) and Figure (ref), further supporting our qualitative claims.}
We estimate the marginal effect $\text{AME}_{0}^{+}$ of cigarette smoking on birth weight using the procedure outlined in Section (ref). In particular, we increase efficiency by assuming that $Y_i|X_i=\bar{x}^+$ is normally distributed (see Remark (ref)). This assumption is empirically well-supported in our setting; Figure (ref) makes clear that the conditional distributions $Y_i|X_i=x$ are all very close to normal with nearly the same variance for $0<x\leq 5$.\footnote{Figure (ref) in the Supplementary Appendix (ref) presents QQ-plots as additional graphical evidence that these conditional distributions are approximately normal.} In addition to simplifying estimation, Figure (ref) can also be interpreted as indirect evidence for item (iii) of Assumption (ref), that $\epsilon_i \mathbin{ \mathpalette{\@indep}{} } X_i^*|X_i=0$ (cf. Proposition (ref) in Appendix (ref)). To see this, note that the distribution of $\epsilon_i|X_i=x$ is just a horizontal shift of the distribution of $Y_i|X_i=x.$ Therefore, the fact that the $f_{Y|X=x}$ look like simple horizontal shifts for $0<X_i\leq 5$ is direct evidence that $\epsilon_i \mathbin{ \mathpalette{\@indep}{} } X_i|0<X_i\leq 5$ as well as indirect evidence that this pattern may continue for $X_i^*\leq 0$ (though, of course, this cannot be directly verified).
For context, in Figure (ref) we also show the conditional distribution $Y|X=0$. While we estimate the variance of $Y_i|X_i=\bar{x}^+$ allowing for a trend in the analogous variances of $Y_i|X_i=x$ for $x>0$, we obtain nearly identical results simply using the variance of $Y|X=1$ as our estimator. This is not surprising given the evident homoskedasticity in the figure.
Estimation requires that we select several different bandwidths. We use a bandwidth of 4 cigarettes for both $\hat{\mathbb{E}}[Y_i|X_i=0^+]$ and $\hat{f}_X(0^+).$ However, we find very similar results using alternative, reasonable bandwidth choices for these objects. The selection of a bandwidth for $\hat{m}'(0^+)$ is more consequential for the standard error of our estimates, thus we report $\widehat{\text{AME}}_{0}^{+}$ for several bandwidth choices.
Table (ref) presents our main results. We estimate the marginal effect of smoking on birth weight at zero to be around -8 grams. We estimate $s'(0^+),$ the selection bias around zero, to also be around -8 grams. The endogeneity term is quite precisely estimated in our application -- most of the sampling variation in $\widehat{\text{AME}}_{0}^{+}$ comes from sampling variation in the estimated slope of $\mathbb{E}[Y_i|X_i=x]$ as $x\downarrow 0$.
As shown in Section (ref), if we can make a local extrapolation, then we can further recover $\text{ATT}(x)$ for positive $x$. We use the first-degree ATT approximation formula in Section (ref), and present the estimates in Figure (ref) for $x \leq 5.$
Our estimates show small negative $\text{ATT}$s. For instance, if mothers who currently smoke one cigarette were to quit smoking, their babies would gain about 8 grams at birth. The effects of quitting smoking add up to a gain of only about 40 grams (1.4 ounces) for mothers who smoked 5 cigarettes per day. The estimates for $x\leq 3$ are not significant at standard levels, while the $x=4$ and $x=5$ estimates are marginally significant at 10% and 5%, respectively. Figure (ref) thus suggests that the effect of maternal smoking on birth weight is small, and we can rule out effects larger than 85 grams (3 ounces) at standard levels of significance. Indeed, these estimates are much smaller than the effect implied by the discontinuity of the outcome at $X=0$ shown in the dashed line of Figure (ref). Our estimates support the qualitative point in almond2005costs that smoking seems to have only small effects on birth weight, although our findings suggest even smaller effects than in that paper. For reference, the average birth weight for a full-term birth in our sample is 3,458 grams (7 pounds and 10 ounces).
When a treatment variable has bunching, this paper presents a new design for identification of the average marginal treatment effects at the bunching point. This is the first identification approach leveraging bunching which does not make assumptions on functional forms or on the shape of the distribution of the unobservables. Since the method does not rely on exclusion restrictions or special data structures, it provides a new avenue for the identification of treatment effects when well-established methods are not applicable.
The approach requires that the treatment be continuously distributed near the bunching point, and it relies on the continuity of the selection function at the bunching point selection value. Intuitively, those who chose the bunching point as an interior solution (i.e., not as a corner solution) are comparable to those right above the bunching point. Besides this and other regularity conditions, the method also requires two other conditions that are hard to explain succinctly, but are implied if the selection equation is monotonic in the selection variable and if the idiosyncratic errors, which are by definition mean independent from the selection variable, are independent from the selection variable at the bunching point.
Identification is achieved by the comparison of the density of the treatment near the bunching point (observed on the positive side) and the density of the selection function at the bunching point (identifiable on the negative side, thanks to a deconvolution of the distribution of the outcome at the bunching point to eliminate the noise from the idiosyncratic error). The ratio of these is exactly the magnitude of the selection bias.
The approach results in the identification of the average marginal effect as a closed-form expression of identifiable quantities which are fairly standard well-known quantities in the econometrics literature, including the limits as the treatment approaches the bunching point of (1) the density of the treatment, (2) the expected outcome, and (3) the derivative of the expected outcome. The final term is the density of the selection variable at the bunching point, which is obtained through a deconvolution of the outcome near the bunching point from the outcome at the bunching point. All the terms in the identification equations can be estimated with off-the-shelf methods readily available in package form in all standard statistical software.
We apply the method to the estimation of the effect of smoking during pregnancy on the baby's birth weight. Our results show that the effects are rather small, strengthening the qualitative results in the previous economics literature.
There is ample opportunity for further technical advancements that would enhance the applicability of this method. Of note, there seems to be a scarce supply of options for the estimation of boundary derivatives in the literature. Even for local polynomial estimators fangijbels1996, the optimal degree, kernel, and bandwidths for estimation of boundary derivatives remain unknown. Deconvolution estimators are ubiquitous in other fields, but still rare in economics, with the notable exception of the measurement error literature (see e.g. schennachmeasurement). The application of deconvolution to bunching is more aligned with the classical setting, where the distributions of the “recorded signal” and the “distortion” are identified, but the behavior of deconvolution estimators with boundary plugins is largely unexplored.