EconBase
← Back to paper

Better Bunching, Nicer Notching

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

120,366 characters · 20 sections · 89 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Better Bunching, Nicer Notching

\onehalfspacing

\ifid \fi

\thispagestyle{empty}

abstractThis paper studies the bunching identification strategy for an elasticity parameter that summarizes agents' responses to changes in slope (kink) or intercept (notch) of a schedule of incentives. We show that current bunching methods may be very sensitive to implicit assumptions in the literature about unobserved individual heterogeneity. We overcome this sensitivity concern with new non- and semi-parametric estimators. Our estimators allow researchers to show how bunching elasticities depend on different identifying assumptions and when elasticities are robust to them. We follow the literature and derive our methods in the context of the iso-elastic utility model and an income tax schedule that creates a piece-wise linear budget constraint. We demonstrate bunching behavior provides robust estimates for self-employed and not-married taxpayers in the context of the U.S. Earned Income Tax Credit. In contrast, estimates for self-employed and married taxpayers depend on specific identifying assumptions, which highlight the value of our approach. We provide the Stata package bunching to implement our procedures.

JEL: C14, H24, J20 \\ Keywords: partial identification, censored regression, bunching, notching

\onehalfspacing

Introduction

Estimating agents' responses to incentives is a central objective in economics and many other social sciences. Piecewise-linear schedules provide identifying variation in incentives to estimate these responses. A continuous distribution of agents that face a piecewise-linear schedule of incentives results in a distribution of responses with mass points located where the slope or intercept of the schedule changes. For example, a progressive schedule of marginal income tax rates induces a mass of heterogeneous individuals to report the same income at the level where marginal rates increase. Many studies in economics use mass points in the response distribution to recover primitive parameters that govern agents' responses to incentives.

Pioneering work by Saez2010, chetty2011, and KlevenWaseem2013 develop bunching estimators to use mass points in response distributions to recover primitive parameters. These estimators are widely applied in economics and rely on the idea that a mass point is larger, the more responsive agents are to incentives. The size of the mass point, however, also depends on the unobserved distribution of agents' heterogeneity. Current methods are only able to map the size of mass points to primitive parameters because they make specific assumptions about the unobserved distribution.

This paper places bunching estimators on a statistical foundation and makes three contributions on the identification of a primitive parameter that summarizes agents' responses to incentives. First, we clarify how the mapping of observed variables to an elasticity parameter depends on assumptions about the unobserved distribution of heterogeneity. The elasticity parameter captures the log percentage change of a response to a log percentage change in an incentive. A change in the intercept of the incentive schedule admits nonparametric point identification of the elasticity but a change in slope does not. Second, we examine the assumptions made by current bunching methods and propose weaker assumptions for partial and point identification of the elasticity. Third, we revisit the original empirical application of the bunching estimator, which is in the literature that examines the largest means-tested cash transfer program in the United States \textemdash the Earned Income Tax Credit (EITC). Our weaker assumptions about the unobserved distribution of heterogeneity result in meaningful changes in estimates of individual responses to taxes.

Our first contribution is to clarify the importance of assumptions about unobserved heterogeneity for the identification of the elasticity. Many existing estimates are based on an agent optimization problem with a piece-wise linear constraint that has one change in slope or intercept. Slope changes in the constraint are often referred to as “kinks” while intercept changes are often called “notches.”\footnote{We generalize the constraint of the agent's problem to a schedule with multiple changes in intercepts and slopes because agents typically encounter a combination of both kinks and notches. The general problem and solution are in Sections (ref) and (ref) of the supplement. For ease of exposition, we keep the problem with one kink or notch in the main text (Section (ref)).} The literature began with an iso-elastic utility model and an income tax schedule that creates a piece-wise linear constraint. We demonstrate our methods in this context and note that our results extend to other contexts with piece-wise linear constraints.

We highlight two insights about identification with kinks and notches assuming a nonparametric family of distributions for unobserved heterogeneity that have continuous probability density functions (PDFs). First, if the constraint has at least one notch, it is possible to point identify the elasticity. Identification comes from using the empty interval in the support of the observed distribution that is created by agents' responses to a notch. Second, point identification is impossible if the incentive schedule only contains kinks. Identification is impossible because there always exists an unobserved distribution that reconciles any elasticity with the observed distribution of responses.

Our second contribution is to propose three novel identification strategies for the elasticity if the incentive schedule has kinks but no notches. Each of these strategies relies on weaker assumptions than those implicit in current implementations of the bunching estimator. Our first strategy identifies upper and lower bounds on the elasticity \textemdash partially identifies the elasticity \textemdash by making a mild shape restriction on the nonparametric family of heterogeneity distributions. The other two strategies point identify the elasticity using covariates and semi-parametric restrictions on the distribution of heterogeneity.

The first strategy partially identifies the elasticity by assuming a bound on the slope magnitude of the heterogeneity PDF, that is, Lipschitz continuity. Intuition for identification of the elasticity in this setting is as follows. We observe the mass of agents who bunch, which equals the area under the heterogeneity PDF inside an interval. The length of this bunching interval depends on the unknown elasticity. The maximum slope magnitude of the PDF implies upper and lower bounds for all possible PDF values inside the bunching interval that are consistent with the observed bunching mass. This translates into lower and upper bounds, respectively, on the size of the bunching interval, which corresponds to lower and upper bounds on the elasticity. These bounds allow researchers to examine the magnitude of the impossibility result in their empirical context. Depending on the data, it might take an unreasonably high slope magnitude on the heterogeneity PDF to produce bounds that include all possible elasticity values. In other settings, the difference between upper and lower bounds may be economically large even for small slope magnitudes.

The next two strategies rely on the fact that bunching can be rewritten as a censored regression model with a middle censoring point. We stress that while these strategies necessarily add structure to point identify the elasticity, they do not require fully parametric assumptions, such as normality, on the unconditional distribution of heterogeneity.

The second strategy identifies the elasticity by estimating a maximum likelihood mid-censored model, using data truncated to a window local to the kink. The likelihood function assumes that the unobserved distribution conditional on covariates is parametric, but we demonstrate that correct specification of the conditional distribution is not necessary for consistency, as long as the unconditional distribution is correctly specified. For example, conditional normality yields a mid-censored Tobit model, which has a globally concave likelihood and is easy to implement. Nevertheless, consistency only requires that the unobserved distribution is a semi-parametric mixture of normals, and that the unconditional distribution implied by the Tobit model matches that; conditional normality is not necessary. Truncating the sample around the kink point improves the fit of the model and further weakens these distribution assumptions.

The third strategy restricts a quantile of the unobserved distribution, conditional on covariates, and point identification follows existing theory for censored quantile regressions powell1986,chernozhukov2002,chernozhukov2015.

Both of the two semi-parametric methods are censored regression models that incorporate covariates. These approaches extend bunching estimators to control for observable heterogeneity for the first time. Observable individual characteristics generally account for substantial variation across agents and leave less heterogeneity unobserved. This fact suggests that identification strategies that utilize covariates should be preferred over identifying assumptions that only restrict the shape of the unobserved distribution without covariates. In addition, covariates generally allow for more precise estimates.

Our third contribution is to illustrate the empirical relevance of our methods by revisiting Saez2010's original, influential application of bunching in the distribution of U.S. income caused by kinks in the EITC schedule. That approach implicitly assumes the unobserved PDF of agents that bunch is linear and uses a trapezoidal approximation to compute the bunching mass. This assumption fits poorly when the true density is non-linear or the interval of agents that bunch is large. We compare elasticity estimates based on our identification assumptions with estimates based on the trapezoidal approximation using annual samples of U.S. federal tax returns from the Internal Revenue Service (IRS).

Our partial identification method indicates that households adjust their reported income in response to marginal tax rates by a considerable amount. Placing a conservative limit on the slope magnitude, the lower bound for the elasticity among self-employed married individuals is 0.48 \textemdash that is, a one percent increase in the marginal tax rate results in a reduction in reported income of at least 0.48 percent. This estimate contrasts with the estimate of 0.77 using the trapezoidal approximation. The difference in these estimates matters. For example, saez2001 shows that the optimal top marginal tax rate for an economy with an elasticity of 0.48 is 48%, while the optimal tax rate for an elasticity of 0.77 is 37%\textemdash a difference of 11 percentage points.

The truncated Tobit model with covariates fits well the observed distribution of income making our semi-parametric consistency result operative. Elasticity estimates from this model differ substantially from estimates based on the trapezoidal approximation for some categories of U.S. taxpayers. For example, we estimate an elasticity of 0.55 versus a trapezoidal estimate of 0.77 for self-employed and married individuals. This large difference highlights the sensitivity of estimates to functional form assumptions, as well as the need for methods that rely on weaker assumptions.

Our three new methods provide a suite of ways to recover elasticities from bunching behavior. Each method differs in the assumptions they make about the unobserved distribution to achieve identification. There is no way to determine which assumption is correct because the unobserved distribution is not fully identified. Nevertheless, estimates that are stable across many methods indicate that different identifying assumptions do not play a major role in the construction of those estimates. On the contrary, estimates that are sensitive to different assumptions are dependent on the validity of those assumptions. Therefore, we recommend that researchers examine the sensitivity of elasticity estimates across all available methods as a matter of routine.

Bunching estimators are widely applied in settings including fuel economy regulations Sallee2012, electricity demand ito2014, real estate taxes kopczuk2015, labor regulations Garicano2016, goff2022, prescription drug insurance EinavFinkelsteinSchrimpf2017, marathon finishing times Allen2017, attribute-based regulations Ito2018, education Dee2019, caetano2020should, minimum wage Jales2018,Cengiz2019, charitable giving hungerman2021, and air-pollution data manipulation ghanem2019, among others. Kleven2016 and bertanha2023 provide a recent review of the many applications and branches of the bunching literature; JalesYu2017 relates bunching to regression discontinuity design (RDD).\footnote{ Variation in the size of the mass point across groups of individuals has also been used as a first stage in a two stage approach to control for endogeneity chetty2013,Caetano2015,Grossman2019 An additional complication in many applications arises when the bunching mass is spread over a range instead of being a mass point. blomquist2019 provide a discussion about the potential sources for this complication and cattaneo2018 propose a filtering method to resolve it.}

In the context of kinks, blomquist2017a were the first to prove the impossibility of point identification and the possibility of partial identification in the iso-elastic quasi-linear utility model \textemdash and an earlier paper provides intuition for the impossibility result blomquist2015. We derive partial identification bounds by assuming the PDF has a bounded slope, whereas blomquist2017a assume the PDF of heterogeneity is monotone. We developed our partial identification result independently of theirs. Our partial identification approach has three valuable features that make it novel: closed-form solutions, observed bunching always implies a positive elasticity, and nesting of the original bunching estimator. blomquist2018a explain that a notch can identify the elasticity and a formal proof of identification appears contemporaneously in an earlier version of our paper, BMS2018. To the best of our knowledge, ours is the first paper to demonstrate point identification using censored regression models, covariates, and semi-parametric assumptions on the distribution of heterogeneity. More generally, the theory demonstrating that a kink fails to point identify the elasticity relates to the literature on impossible inference reviewed by bertanha2019.

The paper proceeds with an utility maximization model subject to a piecewise-linear budget constraint in Section (ref). Section (ref) investigates the identification of the elasticity in the case of kinks and notches. We propose the three identification strategies for the elasticity in Section (ref) and illustrate these methods empirically in the context of the EITC in Section (ref). Section (ref) concludes. Appendix (ref) contains all proofs, and supplemental Appendix (ref) collects auxiliary results and examples. Finally, we developed the Stata command bunching that implements our procedures. The Stata package is presented by bertanha2022stata and available for download from the website of the authors or the Statistical Software Components (SSC) online repository.\footnote{ Type ssc install bunching in Stata to install the package. }

Utility Maximization Subject to Piecewise-Linear Constraints

Firms' and individuals' optimization problems often face piecewise-linear constraints. The nature of constraints is dictated by differential tax rates, insurance reimbursement rates, or contract bonuses. A budget set is fully characterized by a sequence of intercepts and slopes that change at known points. A change in the intercept is referred to as a notch, and a change in the slope is referred to as a kink.

Model Setup

We start with the labor supply characterization employed by the vast majority of the literature, which follows the seminal work of Saez2010 and KlevenWaseem2013. Agents maximize an iso-elastic quasi-linear utility function and choose consumption and labor subject to a piecewise-linear budget set. For ease of exposition, we focus on budget sets with one kink or one notch in the main text and generalize to multiple kinks and notches in the supplemental appendix, Sections (ref) and (ref)

Consider a population of agents that are heterogeneous with respect to a scalar variable $N^{\ast}$, referred to as ability. Ability is distributed according to a continuous probability density function (PDF) $f_{N^{\ast}}$, with support $(0,\infty)$, and a cumulative distribution function (CDF) $F_{N^{\ast}}$. Agents know their $N^{\ast}$, but the econometrician does not observe the distribution of $N^{\ast}$.

Agents maximize utility by jointly choosing a composite consumption good $C$ and labor supply $L$. Utility is increasing in $C$ and decreasing in $L$. These variables are constrained by a budget set, where the agent may consume all of its labor income net of taxes plus an exogenous endowment $I_0$. For simplicity, we assume the price of labor and consumption are equal to one, such that taxable labor income $Y$ is equal to $L$.

In the budget constraint with a kink, the tax rate increases from $t_0$ to $t_1$ as income increases above the kink value $K$. The budget constraint has a notch when the agent is charged a lump-sum tax of $\Delta>0$ as income crosses $K$. Agent type $N^\ast$ maximizes utility $U(C,Y;N^\ast)$ as follows,

eqnarray[eqnarray omitted — 317 chars of source]

where $\mathbb{I}\{ \cdot \}$ is the indicator function; the budget line has intercept $I_0$ and slope $1-t_0$ if $Y\leq K$, but intercept $I_1 = I_0+K(1-t_0) - \Delta $ with slope $1-t_1$ if $Y > K$; and $\varepsilon$ is the elasticity of income $Y$ with respect to one minus the tax rate when the solution is interior. In the case of a kink, $\Delta =0 $, and the budget frontier is continuous; otherwise, in the case of a notch, it has a jump discontinuity of size $\Delta$ at $Y=K$. The solution is always on the budget frontier in Equation (ref).

Model Solution

The solution for $Y$ in Problem (ref) is well known in the literature, when $K$ is a kink Saez2010 and when $K$ is a notch KlevenWaseem2013:

equation[equation omitted — 341 chars of source]

where the expressions for the thresholds $\underline{N}$ and $\overline{N}$ are given below. Section (ref) in the supplement presents the proof of this result and the more general solution in the case of multiple kinks and notches. We provide a graphical analysis of these solutions in Section (ref) in the supplement.

In the case of a kink, $\underline{N} = K(1-t_0)^{-\varepsilon}$, and $\overline{N} = K(1-t_1)^{-\varepsilon}$. The budget frontier is continuous, but its slope suddenly decreases at $Y=K$. For values of $N^*$ inside the bunching interval $[\underline{N},\overline{N}]$, the agent's indifference curve is never tangent to the budget frontier, and we have the non-interior solution $Y=K$. For values of $N^*$ outside of the bunching interval, the indifference curve is always tangent to some point on the budget frontier.

In the case of a notch, the solution is interior for $N^* < \underline{N} = K(1-t_0)^{-\varepsilon}$, but there are no tangent indifference curves for $N^* \in [K(1-t_0)^{-\varepsilon}, K(1-t_1)^{-\varepsilon}]$, just as in the case of a kink. Although tangency occurs for $N^* > K(1-t_1)^{-\varepsilon}$, some of the resulting utility levels are lower than the utility at the notch point. The budget frontier with a jump-down discontinuity at $Y=K$ has an interval of income values $(K,Y^I]$ that no agent ever chooses. The value $Y^I>K$ corresponds to the interior solution of the agent with $N^*=N^I$; that is, the smallest $N^*$ such that the agent's utility is equal to the utility of the agent choosing $Y=K$. Thus $\overline{N} = N^I$, and the solution is at $Y=K$ for $N^* \in [ \underline{N}, \overline{N}] $. As the ability $N^*$ increases above $N^I$, the utility gets larger than the utility at $K$, and again there is an interior solution. Section (ref) in the supplemental appendix has a formal definition of $N^I$ in Equation (ref).

To make the solution more tractable, we take the natural logarithm of all variables. Define $y= \log(Y)$, $n^*=\log(N^*)$, $\underline{n}=\log(\underline{N})$, $\overline{n}=\log(\overline{N})$, $k=\log(K)$, $s_0 = \log(1-t_0)$, and $s_1=\log(1-t_1)$.

equation[equation omitted — 324 chars of source]

As ability $n^*$ increases, the optimal choice of $y$ increases, except when $n^*$ falls inside the bunching interval $[\underline{n},\overline{n}]$, in which $y$ remains constant and equal to $k$.

Bunching and the Counterfactual Distribution of Income

The solution in the previous section expresses income as a function of the model parameters and $n^*$. For given values of $(t_0,t_1,k,\varepsilon)$, the continuously distributed $n^*$ maps into a mixed continuous-discrete distribution for $y$. The model predicts bunching in the distribution of $y$ at a kink or notch point (i.e. $\mathbb{P}(y=k)>0$), but a continuous distribution of $y$ otherwise. The amount of bunching depends on the elasticity $\varepsilon$ and the unobserved distribution $n^{\ast}$,

equation[equation omitted — 301 chars of source]

where the length of the interval $[\underline{n},\overline{n}]$ varies with $\varepsilon$.

The literature typically defines $B$ in terms of the counterfactual distribution of income in the scenario without any kinks or notches. Let counterfactual income be $y_0$ in such case. The solution to Problem (ref) is simply $y_0=n^* + \varepsilon s_0$ for every value of $n^*$. The variable $y_0$ has continuous PDF $f_{y_0}$ and CDF $F_{y_0}$. The bunching mass is derived as

gather[gather omitted — 157 chars of source]

where $\Delta y = \varepsilon (s_0 - s_1)$. Figure (ref), Panels a and b, illustrate the distributions of $y$ and $y_0$, and how they relate to each other, to $B$, and to $f_{n^*}$.

Saez2010's insight is that the mass of agents bunching $B$ is increasing in the elasticity $\varepsilon$ for a given distribution of $y_0$. In other words, the more agents shift income to the kink-point $k$, the more sensitive they are to changes in tax rates. All current bunching and notching estimators use this insight to identify the elasticity. First, the researcher obtains an estimate of the counterfactual distribution of ${y_0}$ and the bunching mass $B$. Plugging these into Equation (ref) allows us to solve for an estimate of the elasticity.

In reality, instead of $y$, researchers typically observe the distribution of $\widetilde{y}=y+e$, where $e$ is a random variable accounting for optimization and friction errors. We focus on the identification problem associated with identifying the counterfactual distribution of $y_0$ in the absence of errors $e$. In work in progress, cattaneo2018 show how to solve this problem with a deconvolution method, which is necessary to recover the distribution in the absence of these errors. Common strategies such as the polynomial strategy, first proposed by chetty2011, fails to solve this issue. We provide a simple counterexample in the supplemental appendix Section (ref) where the polynomial strategy fails to recover the true distribution of $y.$ A practical solution that works under limited assumptions is provided by bertanha2022stata and implemented in Section (ref).

Identification

This section investigates identification with one notch or one kink. We show that identification is possible with one notch without any restriction on the distribution of $n^*$. On the other hand, identification in case of a kink is impossible, unless the researcher imposes restrictions on the distribution of $n^*$. The general solution to Problem (ref) with multiple kinks and notches is found in Section (ref) of the supplement. That section discusses interesting insights to the identification of the elasticity that arise in that context.

Identification from Gaps in the Distribution

We show that identification in the case of a notch is possible using the additional information from the gap in the distribution that is not used in previous studies. Specifically, the gap in the distribution of $Y$ is $(K, Y^I],$ where $Y^I = N^I (1-t_{1})^{\varepsilon}$, and $N^I$ is defined above. Once $Y^I$ is identified from the support of the distribution of $Y$, we numerically solve for $\varepsilon$ that satisfies the indifference condition in Equation (ref) below.

theoremSuppose the support of $N^*$ is equal to $(0,\infty)$, that $K$ is a notch, and that the upper limit of the empty interval in the support of $Y$ to the right of $K$ is equal to $Y^I$. Then the indifference condition that defines $Y^I$ is equivalent to \begin{gather} Y^I + \varepsilon K \left( \frac{K }{Y^I} \right)^{\frac{1}{ \varepsilon}} = \left( {1+\varepsilon } \right) \left( \frac{C + I_{1} + K(1 - t_{1})} {1 - t_{1}} \right), \end{gather} where $C$ is the consumption value on the budget frontier at the notch point. Moreover, there exists an unique $\varepsilon$ that solves Equation (ref) as a function of $Y^I$, $K$, $C$, $I_{1}$, $t_{1}$. Therefore the elasticity is identified.

This and all other proofs are given in Appendix (ref).\footnote{ Similar arguments were given by blomquist2018a around the same time this result appeared in an earlier version of our paper, BMS2018.} Theorem (ref) and its proof consider the case of a notch from a lump-sum tax, $ \Delta > 0 $ in Equation (ref). A minor change to that proof shows the elasticity is also nonparametrically identified in the case of a notch from a lump-sum subsidy, $ \Delta < 0 $. Another case where the elasticity is nonparametrically identified is that of a concave kink, that is, a kink arising from a decrease in tax rates, $t_0>t_1$. This is formally demonstrated in Section (ref) of the appendix and is the only part of the paper where we refer to a kink caused by $t_0>t_1$. When the rest of the paper refers to a kink, we mean a kink generated by $t_0<t_1$.

The identification for these three cases (notch with $\Delta >0$, notch with $\Delta<0$, and kink with $t_0>t_1$) does not depend on the bunching mass, but instead on the size of the region with missing mass (size of the gaps in the distribution). In all the cases we consider, the distribution of income cannot have optimization frictions, which we discuss in detail in Section (ref) of the supplement.

Lack of Identification With One Kink

Although bunching is increasing in the elasticity for a fixed distribution of $y_0$ or $n^*$, it is also true that, for a fixed elasticity, bunching increases as $f_{n^*}$ becomes more concentrated between $\underline{n}$ and $\overline{n}$. If all we know about $f_{n^*}$ is that it is continuous with full support and that its integral over $[\underline{n},\overline{n}]$ equals $B$, then there is no way to identify both the elasticity and $f_{n^*}$ using only Equation $\ref{eq:bunch_j}$; equivalently, there is no way to identify both the elasticity and the distribution of $y_0$ using only Equation $\ref{eq:bunch_j_saez}$. Intuitively, identification using only $\eqref{eq:bunch_j}$ or $\eqref{eq:bunch_j_saez}$ is impossible because each uses one equation to solve for two unknowns. This was first shown in Theorem 1 by blomquist2017a, although the idea was first presented by blomquist2015. We discuss the impossibility result in this section and present our novel identification strategies in the next sections.

Figure (ref) provides intuition behind this impossibility result. It illustrates that the observable PDF $f_y$ in Figure (ref) is generated by applying Equation (ref) to two different combinations of latent variable distributions and elasticities, $f_{n^*,\varepsilon}$ and $f_{n^*,\varepsilon'}$ in Figures (ref) and (ref), respectively. In fact, for any value of the elasticity $\varepsilon>0$, there exists a continuous PDF of $n^*$ that justifies the observed distribution $f_y$ according to Equation (ref). The model assumptions imply no restrictions on $\varepsilon$ over $(0,\infty)$. Theorem 1 by blomquist2017a clarifies that current bunching methods are either implicitly restricting $\mathcal{F}_{n^*}$ or simply inconsistent for the true elasticity. Section (ref) in the supplement details the implicit restrictions on $\mathcal{F}_{n^*}$ made by the original bunching methods, namely, the affine PDF assumption of Saez2010 and the uniform PDF assumption of chetty2011. A direct consequence of the impossibility result is that restrictions on $\mathcal{F}_{n^*}$ are untestable. One may argue that the affine or uniform assumption is a good approximation to any potentially non-linear density $f_{n^*}$ if the bunching interval $\left[k - \varepsilon s_0, k - \varepsilon s_1 \right]$ is small. The problem with this argument is that the size of the interval is itself a function of the elasticity. It is impossible to state that the interval is small and the linear approximation is a good one without a priori knowledge of the elasticity.

We conclude this section by stating a sufficient condition on how flexible $\mathcal{F}_{n^*}$ may be for point identification of $\varepsilon$ to be possible.

assumptionLet $\mathcal{F}_{n^*}$ be a set of all possible CDFs of $n^*$ that are continuously differentiable. For any $F_{n^*} \in \mathcal{F}_{n^*}$, consider the possible values of $e \geq 0 $ and $G_{n^*} \in \mathcal{F}_{n^*}$ that satisfy the following system of equations: \begin{align} G_{n^*}( u - e s_0 ) & = F_{n^*}( u - \varepsilon s_0 ) for \forall u < k, \\ G_{n^*} ( u - es_1 ) & = F_{n^*} ( u - \varepsilon s_1 ) for \forall u \geq k. \end{align} The set of distributions $\mathcal{F}_{n^*}$ is restricted to be such that the only values of $e \geq 0 $ and $G_{n^*} \in \mathcal{F}_{n^*}$ that satisfy Equations (ref) \textendash (ref) are $e=\varepsilon$ and $G_{n^*}=F_{n^*}$.

Intuitively, Assumption (ref) restricts $\mathcal{F}_{n^*}$ in such a way that knowledge of the tails of an unknown CDF in $\mathcal{F}_{n^*}$ is enough to reconstruct that entire CDF; and no other CDF in $\mathcal{F}_{n^*}$ has the same tails. The assumption is easily verified in parametric families, e.g., $\mathcal{F}_{n^*} = \{G_{n^*}(n;\theta) ~,~ \theta \in \Theta \}$, for some set of parameters $\Theta \subseteq \mathbb{R}^p$. In Section (ref), we show how to verify Assumption (ref) using an example of the Gaussian family.

Solutions

The rest of the paper focuses on methods that identify the elasticity in the kink case. We present three types of identification assumptions on the distribution of ability, from less restrictive to more restrictive. We start with a nonparametric shape restriction that bounds the slope magnitude of $f_{n^*}$, which leads to partial identification of $\varepsilon$. Next, we connect bunching to the literature on censored regressions, where $n^*$ is the regression error. It becomes natural to use covariates to explain $n^*$, and we propose two types of semi-parametric restrictions on the distribution of $n^*$ that point-identify the elasticity. The first restricts the distribution of $n^*$, conditional on covariates; and the second restricts a quantile of the distribution of $n^*$, conditional on covariates. In general, more data variation and structure are needed to provide any information about the elasticity.

Nonparametric Bounds

Our partial identification approach relies on restricting the class $\mathcal{F}_{n^*}$ to distributions with PDFs, $f_{n^*}$, that are Lipschitz continuous with constant $M \in (0,\infty)$. In other words, the slope magnitude of any $f_{n^*}$ in this class is bounded by $M$. The following theorem gives the partially identified set for $\varepsilon$ as a function of identified quantities and the maximum slope magnitude $M$.

theoremAssume $\mathcal{F}_{n^*}$ contains all distributions with PDF $f_{n^*}$ that are Lipschitz continuous with constant $M \in (0, \infty)$. Then the elasticity $\varepsilon \in \Upsilon$, where \begin{gather*} \Upsilon = \left\{ \begin{array}{ll} \emptyset & , if B < \frac{ \left|f_{y}(k^+) - f_{y}(k^-) \right| \left[ f_{y}(k^+) + f_{y}(k^-) \right]}{2M} \\ \left[\varepsilon , \overline{\varepsilon}\right] & , if \frac{ \left|f_{y}(k^+) - f_{y}(k^-) \right| \left[ f_{y}(k^+) + f_{y}(k^-) \right]}{2M} \leq B < \frac{ f_{y}(k^+)^2 + f_{y}(k^-)^2 }{2M} \\ \left[\varepsilon , \infty \right) & , if \frac{ f_{y}(k^+)^2 + f_{y}(k^-)^2 }{2M} \leq B \end{array} \right., \end{gather*} where $\emptyset$ is the empty set, and \begin{gather*} \varepsilon = \frac{2 \left[f_{y}(k^+)^2/2 + f_{y}(k^-)^2/2 + M B \right]^{1/2} - \left( f_{y}(k^+) + f_{y}(k^-) \right) } {M(s_0 - s_1)} \\ \overline{\varepsilon} =\frac{-2 \left[f_{y}(k^+)^2/2 + f_{y}(k^-)^2/2 - M B \right]^{1/2} + \left( f_{y}(k^+) + f_{y}(k^-) \right) } {M(s_0 - s_1)}. \end{gather*}

Figures (ref) and (ref) provide the intuition behind the derivation of the bounds in $\Upsilon$. For a fixed value of $\varepsilon$, the length of the interval $[\underline{n},\overline{n}]$ is fixed. If the magnitude of the derivative of $f_{n^*}$ is bounded by $M$, we obtain maximum and minimum areas under $f_{n^*}$ over $[\underline{n},\overline{n}]$. We repeat this exercise for every value of $\varepsilon$ to get a range of possible areas associated with each $\varepsilon$. Given the probability of bunching $B$ is the area under the true $f_{n^*}$ over $[\underline{n},\overline{n}]$, the partially identified set has all values of $\varepsilon$ whose range of possible areas contains $B$. The partially identified set is empty if $M$ is not big enough to allow for the existence of a continuous function $f_{n^*}$ which connects $f_{y}(k^-) = f_{n^*}(k-\varepsilon s_0)$ to $f_{y}(k^+) = f_{n^*}(k-\varepsilon s_1)$. The partially identified set is unbounded if $M$ is large enough to allow $f_{n^*}$ to be zero inside the interval $[\underline{n},\overline{n}]$.

The uniform approximation made by one of the original estimators (Example (ref) in Section (ref) of the supplement) says that $f_{n^*}$ has zero slope inside the bunching interval, that is, $M=0$. The trapezoidal approximation (Example (ref) in that same section) implicitly chooses $M=m_0$ such that $m_0$ is the smallest value of $M$ for which we have bounds that are well defined. Formally, $m_0$ solves $B= \left|f_{y}(k^+) - f_{y}(k^-) \right| ~ \left[ f_{y}(k^+) + f_{y}(k^-) \right]/ 2m_0$, which makes $\underline{\varepsilon} = \overline{\varepsilon}$ and point-identifies $\varepsilon$. Thus the exercise of computing bounds necessarily involves assumptions weaker than the uniform and trapezoidal approximations.

Smoothness assumptions such as slope restrictions are now common in the partial identification literature. For example, kim2018 study partial identification of average treatment effects under smoothness conditions on the treatment response function; rambachan2023 derive bounds on treatment effects in difference-in-difference designs under smoothness conditions on the class of deviations of the parallel trend assumption. A common issue in this literature is that the researcher must choose the smoothness assumption. In the case of Theorem (ref), the researcher must specify the value of $M$. The impossibility of identifying the elasticity without structure on $\mathcal{F}_{n^*}$ (Section (ref)) implies that it is impossible to identify, and thus estimate, the value of $M$.

We recommend researchers to conduct a sensitivity analysis by plotting the bounds in Theorem (ref) as a function of $M$, for a range of values of $M$ that is considered reasonable given the empirical context. From above, we know that $M=m_0$ yields point identification. A useful reference for $M$ comes from the maximum slope magnitude of the continuous part of $f_y$, say $m_1$. The PDF $f_y$ is identified and is the shifted PDF of $n^*$. Thus, the maximum slope of $f_{n^*}$ outside of the bunching interval is identified and equal to $m_1$. If we assume that the slope of $f_{n^*}$ inside the bunching interval is never bigger than outside, then $M=m_1$. Thus, a rule of thumb for the range of values of $M$ is to start at $m_0$ and go up to at least $m_1$. One may also estimate the worst case PDFs of Figures (ref) and (ref) to evaluate the visual effect of $M$ on the shape of the latent distribution. Sensitivity analysis of this kind are not new in the partial identification literature. The idea is to report what can be learned under a sequence of progressively weaker assumptions. For examples, we refer the reader to kim2018 and rambachan2023.

Theorem (ref) is important to quantify the magnitude of the impossibility problem of identification using kinks. If the bounds plotted for a range of $M$ values admit elasticities that are too different in economic terms, then the identifying assumptions play a critical role in determining the elasticity. We give full details and implement this sensitivity analysis in the empirical section using our bunching Stata package (Section (ref)).\footnote{It is important to clarify that the problem of choosing $M$ is different than the typical problem of choosing a tuning parameter, e.g., a bandwidth or polynomial order in nonparametric estimation. The value of $M$ represents a choice of functional form assumption, while in nonparametric estimation, you typically choose the tuning parameter to achieve desirable properties of the estimator for a given functional form assumption.}

We developed Theorem (ref) independently of blomquist2017a, who were the first to present a partial identification result for $\varepsilon$. While we assume the PDF has bounded slope, blomquist2017a partially identify the elasticity by assuming the PDF of heterogeneity is monotone. Our approach has three valuable properties that make it novel. The first is that the bounds of our partially-identified set have closed-form solutions. Second, an observed mass point implies a positive elasticity even for large values of the slope $M$, which is in line with the theoretical prediction that agents respond to a change in incentives. Third, it nests and is easily comparable to the original bunching estimator based on the trapezoidal approximation.

We end this subsection with the case of a budget set with several kinks $k_j$, $j=1,\ldots, J$, but no notches. One may ask whether the existence of several kinks helps identify the elasticity. As noted above, the bunching intervals do not overlap across kinks, that is, $\overline{N}_{j} = K_{j}(1-t_{j-1})^{-\varepsilon} < K_{j}(1-t_{j})^{-\varepsilon}= \underline{N}_{j}$. Multiple kinks do not necessarily point-identify $\varepsilon$, because the distribution of ${n^*}$ may be very different across different bunching intervals.

Multiple kinks do help with the identification of $\varepsilon$, as long as the researcher restricts the slope of $f_{n^*}$ and believes the model in Equation (ref) applies to all individuals. This arises from the fact that every individual is assumed to have the same elasticity parameter $\varepsilon$, and that the bounds of Theorem (ref) vary in length as $B_j$, $f_y(k_j^\pm)$, $s_j$ vary across cutoffs $j=1,\ldots, J$. The partially identified set is narrowed down by the intersection of bounds specific to each one of the multiple kinks.

corollaryAssume the conditions of Theorem (ref) for each kink $k_j$, $j=1, \ldots, J$. Then the elasticity $\varepsilon \in \bigcap_{j=1}^{J} \Upsilon_j$, where $\Upsilon_j$ is the partially identified set of Theorem (ref) applied to kink $k_j$.

Semi-parametric Identification with Covariates

Identification with kinks is impossible when the distribution of ability $n^*$ belongs to the nonparametric class of all continuous distributions. Parametric functional form assumptions identify the elasticity, but identification relies on fitting such functional form to non-bunching individuals and extrapolating the functional form to bunching individuals.

This section considers alternative identification assumptions that rely on the existence of additional covariates in the dataset. There is strong empirical evidence suggesting that ability is well explained by individual characteristics, such as age, demographics, filing status, etc. For example, the ability distribution of young workers may have a very different mean and variance, compared to that of older workers. Extrapolations based on covariates that predict $n^*$ are much more reasonable than extrapolations solely based on the shape of the PDF of $n^*$. The key assumption is that covariates that help explain the distribution of $n^*$ for non-bunching individuals also help explain the distribution of $n^*$ for bunching individuals.

We start by connecting bunching to censored regression models. This allows us to relate to the vast econometrics literature in this area. Consider again the data generating process given by Equation (ref). The model for $y$ is a mid-censored model, where the error term is $n^*$, the intercept to the left of the kink is $\varepsilon s_0$, the intercept to the right of the kink is $\varepsilon s_1$, and the censoring point is $k$. The main difference between (ref) and a typical censored regression model is that the latter has the censoring point at either the minimum or maximum of the distribution of $y$ (see Equation (ref) in the next subsection). Identification, estimation, and inference in these models have been widely studied in econometrics since Tobin1958.

There are many advantages of framing the estimation of $\varepsilon$ as estimation of a censored model. Surveys of censoring models and their applications are provided by Maddala1983, Amemiya1984, Dhrymes1986, Long1997, Demaris2005, and Greene2005. There are straightforward extensions that account for optimizing frictions. Moreover, censored models are easily estimated with a number of different techniques that are available in many computer packages. Most importantly, it becomes extremely practical to add covariates as explanatory factors for the distribution of $n^*$.

Assume the researcher has access to a vector of covariates $X \in \mathbb{R}^{1 \times (d+1)}$, where $X$ contains one intercept variable and $d$ slope variables, $\mathbb{E}[X'X]$ has full rank, but otherwise the distribution of $X$ is unrestricted. We build on censoring models with covariates to identify the elasticity by imposing two types of semi-parametric assumptions on the distribution of $n^*$.

The first type of assumption states that the distribution of $n^*$ is a certain mixture of normal distributions averaged over the distribution of covariates. This assumption does not imply conditional normality of $n^*$ given $X$ but it is implied by conditional normality of $n^*$. Although the Tobit model assumes normality of the unobserved distribution conditional on covariates, we demonstrate that the Tobit estimator remains consistent under our semi-parametric class of normal mixtures \textemdash as long as the unconditional distribution for $y$ implied by the Tobit model matches the true distribution of $y$. In addition, the researcher may estimate a truncated Tobit model on data in a small neighborhood of the kink point, which requires even weaker distribution assumptions for consistency. Our main motivation to study the robustness of Tobit to lack of normality comes from practical reasons: Tobit is extremely popular and easy to implement due to its likelihood function being globally concave and software being ubiquitous.

The second type of assumption imposes a parametric functional form on a quantile of the conditional distribution of $n^*$ given $X$. Sufficient variation in covariates yields point-identification of the elasticity, which is consistently estimated by mid-censored quantile regressions.

Tobit Regression

The first type of assumption is formally stated in Assumption (ref) below. In the meantime, we construct the Tobit estimator by simply assuming that $F_{n^*|X}(n,x) = \Phi\left(\frac{n-x\beta}{\sigma} \right)$ for $\beta \in \mathbb{R}^{(d+1) \times 1 }$ and $\sigma >0$, where $F_{n^*|X}$ denotes the true CDF of $n^*$ conditional on $X$, and $\Phi(\cdot)$ is the CDF of a standard normal distribution. Again, $X$ contains one intercept variable and $d$ slope variables, and $\mathbb{E}[X'X]$ has full rank. Our goal is to first relate bunching to Tobit, which is the most popular censoring model. Conditional normality is assumed for ease of exposition in defining the mid-censored Tobit estimator; it will be relaxed in Assumption (ref) below.

Define the error term $U = n^\ast - X \beta$, the latent variables $y_{0}^* = {\varepsilon} s_0 + X \beta + U$ and $y_{1}^* = {\varepsilon} s_1 + X \beta + U$, where $y_{1}^* < y_{0}^*$, since $\varepsilon>0$ and $s_0>s_1$. Using Equation (ref), we see that $y$ follows a mid-censored Tobit model:

equation[equation omitted — 312 chars of source]

This is different from the classic Tobit model, where the censoring point is either at the minimum or at the maximum of the distribution of $y$. A possible estimation strategy is to adapt the two-step Heckit estimator to our setting Heckman1976,Heckman1979. In the first step, we estimate a binary outcome for bunching and not bunching individuals including covariates. In the second step, we regress income of not bunching individuals on covariates and the equivalents of the inverse Mills ratio. It is useful to relate the mid-censored Tobit model to two classic Tobit models (left\textendash \: and right\textendash censored). To see that, construct the variables $y_{0}=\min\{y,k \}$ and $y_{1}=\max\{k , y \}$. It turns out that $y_{0}$ follows a right-censored Tobit with intercept $\varepsilon s_0 + \beta_0$, slope coefficients $\beta_1, \ldots, \beta_d$, where $\beta = (\beta_0, \beta_1, \ldots, \beta_d)$. Similarly, $y_{1}$ follows a left-censored Tobit with intercept $\varepsilon s_1 + \beta_0$, and slope coefficients $\beta_1, \ldots, \beta_d$. Thus, the elasticity is consistently estimated by the difference of both intercepts $(\varepsilon s_1 + \beta_0) - (\varepsilon s_0 + \beta_0) $ divided by $(s_1-s_0)$. Although readily implementable in most statistical packages, this estimation strategy does not constrain the slope coefficients and variances to be equal on both sides of the kink, which translates into loss of efficiency. The mid-censored Tobit likelihood naturally takes these constraints into account and provides the most efficient estimates. It is therefore our preferred implementation.

Let $(y_i, X_i)$, $i=1,\ldots,n$, be an iid sample of observations. The maximum likelihood estimator (MLE) for the true parameters $(\varepsilon, \beta, \sigma)$ is constructed by maximizing the log-likelihood function with respect to $(e,b,s)$ given the sample data,

align[align omitted — 699 chars of source]

Regardless of what the true distribution $F_{n^*|X}$ is, the MLE based on (ref) is consistent for the parameters that maximize the population average of the log-likelihood function, that is, $(e^*, b^*, s^*) = \arg \max_{e,b,s} \mathbb{E}[\ell_i(e, b, s)]$, where the solution is unique (see, e.g., hayashi2000, Section 8.3.). We say the elasticity is identified by a mid-censored Tobit when $e^* = \varepsilon$. This occurs when $F_{n^*|X}(n,x) = \Phi\left(\frac{n-x\beta}{\sigma} \right)$, but we would like to relax the normality assumption and still have $e^* = \varepsilon$. We pursue this exercise in the next paragraphs. The Tobit estimator is extremely practical to implement, and we believe it is worth investigating a set of assumptions on $F_{n^*|X}$ that are weaker than normality and yet sufficient for consistency. We start by describing these conditions on the set of possible distributions $\mathcal{F}_{n^*|X}$.

assumptionThe set of true conditional CDFs of $n^*$ given $X$ is denoted $\mathcal{F}_{n^*|X}$. We assume $\mathcal{F}_{n^*|X}$ is such that: \begin{enumerate} • for every ${F}_{n^*|X} \in \mathcal{F}_{n^*|X}$, the unconditional CDF satisfies $F_{n^*}(n) = \mathbb{E}\left[ {F}_{n^*|X}(n,X) \right] = \mathbb{E}\left[ \Phi\left(\frac{n-X b}{s} \right) \right]$ for some $b \in \mathbb{R}^{(d+1) \times 1}$ and $s>0$. The set of all unconditional CDFs $\mathcal{F}_{n^*} = \left\{ F_{n^*}(n) = \mathbb{E}\left[ {F}_{n^*|X}(n,X) \right] \text{ for } {F}_{n^*|X} \in \mathcal{F}_{n^*|X} \right\}$ satisfies Assumption (ref); • recall that $\underline{n} = k-\varepsilon s_0$, $\overline{n} = k-\varepsilon s_1$, $B = \mathbb{E}\left[ {F}_{n^*|X}(\overline{n}, X) - {F}_{n^*|X}(\underline{n}, X) \right]$; define $D=\mathbb{I}\{n^* \geq \overline{n} \}$ and $B_N(X,\delta,\theta,s) = \Phi\left(( \overline{n} - \delta - X \theta )/s \right) - \Phi\left(( \underline{n} - X \theta)/ s \right)$; for every ${F}_{n^*|X} \in \mathcal{F}_{n^*|X}$, we have $\mathbb{E}\left[ {F}_{n^*|X}(n,X) \right] = \mathbb{E}\left[ \Phi\left(\frac{n-X \beta}{\sigma} \right) \right]$ for $\beta$ and $\sigma$ that satisfy \begin{align} (0,\beta,\sigma) = \arg \min_{\delta, \theta, s} & \frac{ 1-B }{2} \left\{ \log(s^2) + \frac{1}{s^2} \mathbb{E}\left[ \left(n^* - D \delta - X \theta \right)^2 | n^* \not \in [n, \overline{n}] \right] \right\} \notag \\ & - B \mathbb{E}\left[ \log\left( B_N(X,\delta,\theta,s) \right) | n^* \in [n, \overline{n}] \right]. \end{align} \end{enumerate}

Assumption (ref)(i) says that if we take any conditional distribution of $n^*$ given $X$ from the set $\mathcal{F}_{n^*|X}$ and integrate it over $X$ we obtain a marginal distribution of $n^*$ which is a mixture of normals. The mixture is averaged over the distribution of $X$ and the more variation in covariates one has, the richer the set $\mathcal{F}_{n^*}$ is. Assumption (ref)(i) also assumes the set of mixtures of normals satisfies Assumption (ref), so it is sufficient for point identification of the elasticity. To see that, note that each mixture of normals in $\mathcal{F}_{n^*}$ is characterized by $b \in \mathbb{R}^{(d+1) \times 1}$ and $s>0$, that is, $G_{n^*}^{b,s}(n) = \mathbb{E}\left[ \Phi\left(\frac{n-X b}{s} \right) \right]$. That along with a value of the elasticity $e \geq 0$ imply a distribution of $y$,

equation*[equation* omitted — 192 chars of source]

If we search for a mixture of normals in $\mathcal{F}_{n^*}$ and an elasticity value that matches the observed distribution of $y$, we will solve uniquely for the true elasticity and a distribution $G_{n^*} \in \mathcal{F}_{n^*}$ (Assumption (ref)). Uniqueness of $G_{n^*}$ does not necessarily mean that $G_{n^*}$ is indexed by a unique set of parameter values $(b,s)$. The assumption still allows for cases where different values of parameters $(b,s)$ lead to the same CDF, $G_{n^*}$. For example, let $X=[1,W]$ and $(n^*,W)$ be distributed as a standard bivariate normal with correlation $\rho$. The CDF of $n^*$ conditional on $X$ equals $\Phi( n-\rho W )$. The unconditional CDF of $n^*$ is equal to $\Phi(n)$, which is the same as $\mathbb{E}[\Phi( n-\rho W )]$ or $\mathbb{E}[\Phi( n )]$. Thus, $\Phi(n)$ belongs to $\mathcal{F}_{n^*}$ with $s=1$ and either $b=(0,\rho)$ or $b'=(0,0)$. Coming back to Assumption (ref), part (i) says that $G_y^{e,b,s}=F_y$ for some value of $(e,b,s)$ where $e$ is always equal to $\varepsilon$.

The mid-censored Tobit estimator produces a value for the elasticity $e^*$ and picks the mixture of normals characterized by $(b^*, s^*);$ both of these imply a distribution of $y$: $G_y^{*} = G_y^{e^*,b^*,s^*}$. The Tobit best-fit distribution $G_y^*$ may or may not match the observed distribution of $y$, even if Assumption (ref)(i) is true; in case it does match, the equality $F_y=G_y^*$ implies $e^*=\varepsilon$ by virtue of Assumption (ref)(i). In general, we may not have $F_y=G_y^*$ because the Tobit MLE does not necessarily minimize the distance between $F_{y}$ and $G_{y}^{e,b,s}$ for a choice of parameters $(e,b,s)$. In fact, the Tobit MLE minimizes the Kullback-Leibler divergence between two conditional distributions of $y$ given $X$ averaged over $X$, where the two conditional distributions are $F_{y|X}$ and $G_{y|X}^{e,b,s}$ for a choice of parameters $(e,b,s)$. These are two different minimization problems in general. Assumption (ref)(ii) states a condition on $F_{n^*|X}$ that guarantees that minimizing the average Kullback-Leibler divergence is equivalent to minimizing the distance between $F_{y}$ and $G_{y}^{e,b,s}$.

To interpret Assumption (ref)(ii), note that the objective function of the minimization problem in (ref) is a known function of $(\delta, \theta, s)$ once we fix $F_{n^*|X} \in \mathcal{F}_{n^*|X}$ plus knowledge of the distribution of $X$, which is identified. That objective function is the average of two terms weighted by the bunching mass $B$ and $1-B$. The term on the LHS is the objective function of the least-squares problem conditional on $n^* \not \in [\underline{n}, \overline{n}]$, and the term on the RHS is an average of the log of the bunching mass conditional on $X$ in case the distribution of $n^*$ given $X$ were normal. Thus, the objective function is a “penalized” least-squares problem. To gain further insight, assume $\varepsilon \approx 0$, so that $B \approx 0$ and $\underline{n} \approx \overline{n}$. In this case, (ref) is approximately \[ (0,\beta,\sigma) = \arg \min_{\delta, \theta, s} \left\{ \log(s^2) + \frac{1}{s^2} \mathbb{E}\left[ \left(n^* - D \delta - X \theta \right)^2 \right] \right\}. \]

This is equivalent to saying that the population regression of $n^*$ on $(D,X)$ must produce $(0,\beta)$ as coefficients, and that the variance of the regression error must be $\sigma^2$. This translates to linear restrictions on the first two moments of the joint distribution of $\left( n^*, \mathbb{I}\{n^* \geq \overline{n} \}, X \right)$, but that distribution is otherwise unrestricted. In particular, a Gaussian distribution ${F}_{n^*|X}(n,X) = \Phi\left(\frac{n-X b}{s} \right)$ belongs to the set $ \mathcal{F}_{n^*|X}$, but not every element of $\mathcal{F}_{n^*|X}$ is Gaussian. We give examples of $F_{n^*|X}$ that satisfy Assumption (ref) and are not Gaussian in the simulation experiments at the end of this section (Figures (ref)\textendash (ref)) and in Section (ref) of the supplemental appendix. Lemma (ref) below shows that the mid-censored Tobit MLE identifies the elasticity under Assumption (ref), and the proof is in Section (ref) of Appendix (ref).

lemmaLet $F_y$ be the CDF of the observed distribution of $y$ and $G_y^*$ be the CDF of the Tobit best-fit distribution for $y$ as defined above. Suppose Assumption (ref)(i) holds. If $G_y^* = F_y$, then $e^*=\varepsilon$. Moreover, suppose Assumption (ref)(ii) holds. Then, $G_y^* = F_y$ and $e^*=\varepsilon$.

If the Tobit best-fit distribution of $y$ matches the true distribution of $y$, Lemma (ref) guarantees that the elasticity estimated by the Tobit is consistent for the true elasticity, regardless of whether $F_{n^*|X}$ is normal. We provide an example of this in the second simulation experiment at the end of this section (Figure (ref)), as well as in Section (ref) of the supplemental appendix (Figure (ref)). Standard quasi-MLE asymptotic inference procedures apply here. Namely, the MLE $(\hat \varepsilon, \hat \beta, \hat \sigma)$ obtained from (ref) and centered at $(e^*, b^*, s^*)$ is asymptotically normal, with zero mean and the usual variance-covariance matrix in the “sandwich form.”

One of the features of bunching estimators is the reliance on data local to the kink point. With the mid-censored Tobit model, the researcher may also restrict the sample to observations of $y$ lying in a small neighborhood of $k$ and estimate a truncated Tobit.\footnote{ The truncated Tobit model has log-likelihood that is slightly different from (ref). Instead of the log-likelihood of $y|X$, we maximize the log-likelihood of $y|X,k-\delta<y<k+\delta$ for $\delta>0$, which has a truncated normal distribution censored at $k$.} The truncated Tobit is an attractive estimation strategy, because consistency of $\hat \varepsilon$ has weaker requirements in terms of Assumption (ref) and the $G_y^* = F_y$ condition. Assumption (ref) only needs to hold for the distribution of $n^*$ conditional on $n^*$ being in a small interval containing $[\underline{n}, \overline{n}]$. Moreover, the smaller the truncation window, the easier it is to fit the unconditional distribution of $y$ with a Tobit, and the stronger is the robustness result of Lemma (ref).

As a matter of routine, we recommend researchers estimate a truncated Tobit model for various window sizes around the kink point and examine two things: first, the plot of the estimated elasticity as a function of the size of the truncation window; second, the plot of the best-fit Tobit distribution of $y$ compared to the histogram of $y$ for various sizes of truncation windows. The distribution fit tends to improve as the size of the window decreases. The better the fit, the more likely the conditions of Lemma (ref) are met, and the closer is the elasticity to the truth. We illustrate this exercise with simulated data below and with real data in Section (ref).

To end this section, we carry out two simulation experiments to illustrate the robustness property of our Tobit estimator to lack of normality. In both experiments, we use the Skewed Generalized Error Distribution (SGED) with parameters $\mu$, $\sigma$, $k$, and $\lambda$ (see theodossiou2000, Section 5A). We denote the PDF $f_{n^*}(n)$ of a SGED as $SGED(n;\mu,\sigma,k,\lambda)$. Parameters $\mu \in \mathbb{R}$ and $\sigma \in \mathbb{R}_+$ equal the mean and standard deviation of the distribution, respectively. The parameter $k \in \mathbb{R}_+$ relates to kurtosis, while $\lambda \in (-1,1)$ regulates skewness. SGED nests the Laplace distribution ($\lambda=0$, $k=1$), the normal distribution ($\lambda=0$, $k=2$), and the uniform distribution ($\lambda=0$, $k \to \infty$) as special cases. The data generating processes of these experiments have some empirical features that resemble those of the EITC data in Section (ref), e.g., location of the kink and range of the distribution.

Experiment 1 has two goals. First, the experiment shows that a mixture of normals can be almost any distribution as long as the distribution of $X$ is rich enough. The unconditional distribution of $n^*$ does not need to be “locally normal” or “locally symmetric” at the kink. The second goal of Experiment 1 is to show that truncation without covariates may require a much smaller truncation window to fit the distribution of $y$ compared to truncation with covariates. The key parameters of Equation (ref) are $\varepsilon=1$, $k=2.0794$, $s_0=0.2624$, and $s_1=-0.1054$. Figure (ref) displays the distribution of $n^*$, which is a mixture of two SGEDs: $f_{n^*}(n)= (1/2)SGED(n; 1.6,0.75,4,-.5) + (1/2)SGED(n; 6,0.75,1,.5)$. The random variable $X$ is a scalar and $\beta=1$. We specify $F_{n^*|X}(n^*) = \Phi\left((n-X)/0.0717\right)$ and solve numerically for the distribution of $X$ that satisfies $F_{n^*}(n) = \mathbb{E}\left[\Phi\left((n-X)/0.0717\right) \right]$ (Figure (ref)). We generate 50,000 observations of $(y,X)$ according to this model, and the histogram of $y$ is displayed in Figure (ref). Although $f_{n^*|X}$ is normal, it is clear from the figures that $f_{n^*}$ is very far from Gaussian, even within smaller truncation windows.

We estimate two different Tobit models using data from Experiment 1. The first model is correctly specified with covariate $X$. We start with the full sample of simulated data and produce estimates for truncation windows that are symmetric around the kink point and shrink in size. For example, Figures (ref)\textendash (ref) show the histogram of simulated data for $y$, and the best-fit Tobit distributions for three truncation sizes, 100%, 60%, and 20%. Figure (ref) displays the elasticity estimate as a function of the percentage of data used in each truncated estimation. As expected, the elasticity estimate is stable over all truncation windows, because the model is correctly specified. The Tobit fits the distribution of $y$ perfectly for all truncation windows, and the estimated elasticity is approximately equal to the truth.

The second model we estimate with data from Experiment 1 omits the covariate $X$. The Tobit model is misspecified because $f_{n^*}$ is not normal. As expected, estimation using all of the data does not fit the distribution of $y$ (Figure (ref)). As a general rule, the smaller the truncation window, the better the fit to the distribution of $y$ (Figures (ref)\textendash (ref)), although a perfect fit is not always guaranteed for reasonable sample sizes. It is only in the last feasible truncation window of 20% that the fit becomes reasonable and the elasticity estimate reaches the true value (Figure (ref)). This experiment shows the importance of covariates for the fit of the distribution and that truncation does not always yield the same fit that the inclusion of relevant covariates does.

We move to Experiment 2, which again has two goals. First, not only normality of $n^*$ is not required for consistency; conditional normality of $F_{n^*|X}$ is not required either. Second, consistency of the Tobit elasticity and perfect fit of the distribution of $y$ do not require truncation in models without conditional normality (i.e., models with $F_{n^*|X}$ misspecified). Figure (ref) plots the PDF of $n^*$, which is approximately the mixture of two SGEDs, $(1/2)SGED(n;1,0.75,4,-.5) + (1/2)SGED(n;6,0.75,1,.5)$. The distribution of scalar $X$ is discrete with 20 mass points (Figure (ref)) and is chosen such that $F_{n^*}(n) = \mathbb{E}\left[\Phi\left((n-X)/0.1919\right) \right]$ approximates the CDF of the mixture of two SGEDs. We solved numerically for non-normal conditional distributions of $n^*$ given $X$ that satisfy Equation (ref). Figure (ref) displays the true PDFs $f_{n^*|X=x}$ in black and the normal PDFs $g_{n^*|X=x}$ assumed by the Tobit in gray, for all values of $x$. We clearly see that $f_{n^*|X}$ is not normal. We then generate 50,000 observations of $(y,X)$ and fit our Tobit model with the covariate $X$ to the entire sample. Despite the lack of conditional normality and truncation, the Tobit model fits the distribution of $y$ (Figure (ref)) and estimates the elasticity at $\ha\varepsilon = 1.0083$ (S.E. 0.0073). Section (ref) in the supplement repeats this experiment for the case $n^*$ has uniform distribution.

This section gives a practical rational for using the truncated Tobit model but other more flexible model assumptions, such as index models or the semi-parametric Tobit model of chen2011 are also possible. For example, and in terms of our notation, Section 7.3 of chen2011 assumes $n^* = X \beta + \sigma(X) \nu$, with $G$ being the CDF of the distribution of $\nu$ conditional on $X$. Both $\sigma$ and $G$ are unknown but smooth functions, and $\varepsilon$ and $\beta$ are partially identified. The more structure one imposes on $\sigma$ and $G$, the narrower the partially identified set gets, eventually getting to point identification as in our case. chen2011 then construct confidence regions for $\varepsilon$ and $\beta$ that are valid regardless of point or partial identification using a sieve MLE bootstrap. chen2018 provide a computationally attractive alternative to the sieve MLE bootstrap that is based on Monte Carlo simulations. Overall these methods constitute helpful approaches to performing sensitivity analysis of model assumptions and suggest a rich area for future work.

Censored Quantile Regressions

Another type of semi-parametric assumption on the ability distribution consists of restricting a quantile of the distribution of $n^*$, conditional on $X$. Namely, for $\tau \in (0,1)$, we assume that there exists $\beta(\tau) \in \mathbb{R}^{1 \times (d+1)}$ such that

equation[equation omitted — 90 chars of source]

where $Q_{\tau}$ denotes the $\tau$-th quantile of a distribution. A common choice in applied work is $\tau=1/2$ or the median regression. The restriction in (ref) may be a flexible one if one includes transformations of $X$ on the right-hand side, e.g., polynomials and interaction terms.

Equation (ref) leads to $y= \min\{ \varepsilon s_0 + n^* ; ~ \max \{k ;~ \varepsilon s_1 + n^* \} \}$, which is an increasing and continuous function of $n^{\ast}$. The quantile of an increasing and continuous function of $n^*$ is equal to that same function evaluated at the quantile of $n^*$. Using Equation (ref),

equation[equation omitted — 170 chars of source]

For those observations such that $X \beta(\tau)< k - \varepsilon s_0$ or $X \beta(\tau)> k - \varepsilon s_1$, the quantile $Q_{\tau} \left(y \mid X \right)$ varies linearly with $X$; otherwise, it is constant and equal to $k$. Intuitively, if there is enough variation in $X$ for uncensored observations, then the slope coefficients and the intercepts are identified. This leads to identification of $\varepsilon$.

lemmaDefine $\tilde{X}= \left[X, ~ \mathbb{I}\left\{ Q_{\tau} \left(y \mid X \right) > k \right\} \right]$, a random vector in $\mathbb{R}^{1 \times (d+2) }$. Assume \\ $\mathbb{E} \left[ \mathbb{I}\left\{ Q_{\tau} \left(y \mid X \right) \neq k \right\} \tilde{X}' \tilde{X} \right] $ has full rank and that Equation (ref) holds. Then $\varepsilon$ is point identified.

The quantile method does not contradict the impossibility of point identification discussed in Section (ref), that is, we do need restrictions on the distribution of $n^*$ conditional on $X$ to point identify the elasticity. These restrictions are Equation (ref) and the rank condition of Lemma (ref). To see that, note that in the absence of covariates, the rank condition is never satisfied. When we have covariates, we are not free to specify $Q_{\tau} \left(n^* \mid X \right)$ as flexible as desired. For a fixed distribution of $X$, the rank condition eventually fails as we increase the flexibility of $Q_{\tau} \left(n^*\mid X \right)$.

We illustrate this fact with a simple example. Suppose the researcher has two dummy variables, $W_1$ and $W_2$, and wants to be fully flexible. An unrestricted $Q_{\tau} \left( n^* \mid W_1,W_2 \right)$ contains four parameters, because the conditional quantile function takes at most four different values. That is, $Q_{\tau} \left( n^* \mid W_1,W_2 \right) = \beta_0 + \beta_1 W_1 + \beta_2 W_2 + \beta_3 W_1 W_2$. In terms of Lemma (ref), $X=[1,W_1,W_2,W_1 W_2]$ is $1 \times 4$, $\widetilde{X}$ is $1 \times 5$, which implies $d=3$. In the best case scenario for identification, four values of $Q_{\tau} \left( n^* \mid W_1,W_2 \right)$ translate to four values of $Q_{\tau} \left( y \mid W_1,W_2 \right)$ that are all different from $k$, so that $\mathbb{I}\left\{ Q_{\tau} \left(y \mid X \right) \neq k \right\}=1$. The matrix $\mathbb{E} \left[ \mathbb{I}\left\{ Q_{\tau} \left(y \mid X \right) \neq k \right\} \tilde{X}' \tilde{X} \right] $ $=\mathbb{E} \left[ \tilde{X}' \tilde{X} \right] $ is $5 \times 5$ but has rank equal to $4$ at most. Thus, $Q_{\tau} \left(n^* \mid X \right)$ must be restricted to fewer parameters for point identification to be possible.

As seen in the example with dummies, the amount of variation in covariates trades off with the degree of flexibility in specifying the parametric functional form for $Q_{\tau} \left( n^* \mid X \right)$. The more variation researchers have (e.g., continuous vs. discrete $X$), the more flexible the parametric functional form may be, and the less damaging misspecification errors will be. Unfortunately, our data in Section (ref) only have binary covariates, which severely limits our ability to obtain sensible estimates using the quantile identification method. An interesting extension to Lemma (ref) relates to the work by hong2003 and khan2009, which connects censored quantile regressions to moment inequalities and partial identification. Although increasing the flexibility in the specification of $Q_{\tau} \left( n^* \mid X \right)$ may violate the rank condition of Lemma (ref), researchers may still use those methods to partially identify the elasticity. For inference, researchers may use the simulation-based confidence regions of chen2018, which are valid regardless of point or partial identification.

Theoretical work on estimation and inference of parameters in censored quantile regression (CQR) models dates back to the 1980s (powell1984, powell1986). Recent advances include the computationally attractive three-step estimator by chernozhukov2002, and CQR with endogeneity by hong2003, khan2009, and chernozhukov2015. In the simpler case of $Q_{\tau} \left(y \mid X \right) = X \beta(\tau)$, koenker1978 show that a consistent estimator for $\beta(\tau)$ is obtained by the solution to the problem

equation[equation omitted — 113 chars of source]

where $(y_i,X_i)$ $i=1,\ldots,n$ is an iid sample and $\rho_{\tau} \left(u \right)= \left(\tau - 1 \left(u \le 0 \right)\right)u$ is the so-called “check function.” In our case, the parametric conditional quantile function $Q_{\tau} \left(y \mid X \right)$ is given in Equation (ref). The slope and intercept coefficients are estimated by

equation[equation omitted — 252 chars of source]

where $\hat{b}(\tau)$ is consistent for $\beta(\tau) + [\varepsilon s_0, ~ 0, ~ \ldots, ~ 0]'$, and $\hat{\delta}(\tau)$ is consistent for $\varepsilon(s_1 - s_0)$. Therefore the elasticity is consistently estimated by $\hat\varepsilon = \hat\delta/(s_1-s_0)$ and is asymptotically normal.

The optimization problem in Equation (ref) is computationally difficult. For the left (or right) censored case, chernozhukov2002 proposed a fast and practical estimator that consists of three steps. Our case of middle censoring requires a straightforward modification of their method. We delineate practical steps to obtain $\hat\varepsilon$ and its standard error using CQR in Section (ref) of the supplemental appendix.

Application to EITC

We demonstrate and compare our new methods using bunching behavior created by kinks in the earned income tax credit (EITC). Each method differs in the assumptions they make about the unobserved distribution to achieve identification. There is no way to determine which assumption is correct because the unobserved distribution is not fully identified. Nevertheless, estimates that are stable across many methods indicate that different identifying assumptions do not play a major role in the construction of those estimates. On the contrary, estimates that are sensitive to different assumptions are dependent on the validity of those assumptions. PatelSeegertSmith2016 provide an empirical illustration of this sensitivity.

First, we use our nonparametric bounds to provide initial information about how sensitive the elasticity estimate is to different shapes of the underlying ability distribution. When the bounds are tight, then the shape of the underlying distribution is not critical. But when the bounds are wide, then the shape is critical. In this case, reducing the range of possible elasticities requires either stronger restrictions on the shape of the ability distribution or additional data on determinants of ability.

Second, we combine observed determinants of ability with our semi-parametric approach to point identify the elasticity. We compare the resulting best-fit Tobit income distribution to the observed distribution for alternative samples that range from using all observations to using only data local to the kink. When the best-fit Tobit distribution coincides with the observed distribution, the estimated elasticity is consistent (Lemma (ref)). Furthermore, if the Tobit elasticity is within narrow nonparametric bounds, then the identifying assumptions are inconsequential; if within wide bounds, then the identifying assumptions are not contradictory and the covariates provide point identification. In contrast, if the Tobit elasticity is outside of the bounds, then the elasticity estimate is not robust to the two alternative identifying assumptions. Finally, when the best-fit Tobit distribution does not coincide with the observed distribution, the determinants of ability used for estimation are uninformative or the semi-parametric assumption is inappropriate.

We recommend that researchers examine the sensitivity of elasticity estimates across all available methods as a matter of routine. We illustrate these steps in the context of the EITC in the rest of this section.

Data

We use data from the Individual Public Use Tax Files, constructed by the IRS. The annual cross-section for each year 1995 to 2004 includes sampling weights which allow interpretation of any estimates as being based on the population of U.S. income tax returns. This data was initially used by Saez2010 to demonstrate how to use bunching to estimate an elasticity.

The income distribution for individuals with one child demonstrates clear bunching around the \$8,580 kink (year 2008 dollars) in the EITC schedule (Figure (ref)). Because the marginal tax rate increases from $-34$ percent to $0$ percent at \$8,580, individuals have strong incentives to report more income up to the kink point.

Observed bunching in the distribution of income suggests that people do respond to changes in tax rates. To effectively set tax rates, however, it is imperative to quantify this response precisely. Small variation in elasticity estimates imply large differences in optimal tax rates saez2001. For example, a true elasticity of 0.2 implies the optimal top marginal tax should be 69%. If instead, the true elasticity is 0.3, the optimal top marginal tax should be 60%.\footnote{ This example comes from saez2001. In particular, Equation 9 states $\bar{\tau} = (1 - g)/(1 - g + \varepsilon^{u} +\varepsilon^{c}(a -1)),$ where $g$ is defined as the value the government has for the marginal consumption of high income earners (often set to 0), $a$ is the Pareto parameter (with baseline value of 2), and $\varepsilon^{c}$ and $\varepsilon^{u}$ are the compensated and uncompensated elasticities of taxable income. For the calculation in the text, we utilize $\varepsilon^{u}= \varepsilon^{c}$, a Pareto parameter of 2, and a $g$ value of 0.1.} As demonstrated above, identifying the elasticity requires information on the amount of bunching and the income distribution. The following sections show how different methods leverage different types of variation to identify the elasticity.

The methods of this paper are designed for data without friction errors or sharp bunching. For examples of sharp bunching data, see Figure 4 by glogowsky2018, Figure 1 by goncalves2018, and Figure 8 by goff2022. Nevertheless, the IRS data do have friction errors as the excess mass due to bunching is visibly dispersed in a small interval near the kink (for example, Figure 5 by weber2016). Therefore, to apply our procedures to the IRS data, we first need to filter reported income out of friction error. A proper deconvolution theory must be developed to tackle this problem, but it is beyond the scope of this paper. \footnote{There are several works in progress investigating frictions. See, for example, aronsson2018,alvero2020,cattaneo2018, and McCallumNavarrete2022. } For now, we simply need a practical way of removing friction error before applying the different bunching estimators, so that they may be properly compared.

Following the intuition of chetty2011, we fit a seventh-order polynomial to the empirical CDF of reported income with friction errors $\tilde{y}$. As does Saez2010, we exclude observations that lie within \$1,500 of the kink and allow an intercept change at the kink. The extrapolation of the fitted polynomial to the excluded region results in a CDF with a jump discontinuity at the kink. This is an estimate for the CDF of income without friction error, that is, $F_{y}(y)$. The size of the discontinuity equals the bunching mass. We then rely on the fact that $y = F_{y} \left( F_{\tilde{y}}^{-1}(\tilde{y}) \right)$ and use the estimated CDFs to transform $\tilde{y}$ into $y$.

Our filtering procedure is different from the polynomial strategy discussed in Section (ref) and Section (ref) of the supplement. We simply aim at removing the friction error from the sample, while the the polynomial strategy aims to remove friction error and recover the counterfactual distribution of income, which requires much stronger restrictions according to the discussion in Section (ref). Our filtering procedure consistently estimates the CDF of $y$ under the following conditions: (i) frictions only affect bunching individuals additively; (ii) friction error is independent of unobserved heterogeneity $n^*$ and has known support containing zero, i.e., $[-1500,+1500]$; (iii) the CDF of $y$ is represented by a seventh order polynomial with an intercept change at the kink. Section (ref) of the supplement has the proof of this claim. A more general filtering method is deferred to future work.\footnote{Section (ref) of the supplement recomputes our estimates using the filtering procedure employed by Saez2010.}

Estimates Across Methods

Table (ref) reports estimates of the elasticity of taxable income using a classic bunching method, nonparametric bounds, and Tobit models with covariates. Each of these estimates relies on a different set of assumptions to identify the elasticity of taxable income, and together they provide insights into which assumptions are most defensible in the context of the EITC.

Column 1 reports our estimates of the elasticity of taxable income using a trapezoidal approximation (Example (ref)).\footnote{ We estimate the PDF of the variables in logs rather than in levels, which simplifies the elasticity formula based on the trapezoidal approximation in Example (ref). } This method assumes the unobserved PDF is linear in the bunching region, which prior literature believed to approximate non-linear distributions well. In practice, the appropriateness of this approximation depends on the true distribution and length of the bunching region, which are both unobserved. Linearity may be inappropriate if the distribution is sufficiently non-linear or the bunching region is wide.

Column 1 demonstrates substantial heterogeneity in estimates across different subsamples. In particular, the elasticity estimate is 0.32 for the all filers sample, 0.81 for self-employed individuals, 0.77 for self-employed married individuals, and 0.84 for self-employed not married individuals.

The guidelines for implementation of our nonparametric bounds in Section (ref) utilizes a range of values for $M$ that includes the maximum slope magnitude of $f_y$. We reiterate that $M$ is unidentified and that the slope of $f_y$ provides a starting point. The bunching Stata package consistently estimates the maximum slope of $f_y$ by taking the maximum slope in the histogram of $y$ across all consecutive bins. We find that the slope is never bigger than $0.5$ across our subsamples. For a more conservative view, we report our nonparametric bounds using $M=0.5$ and $M=1$ in Table (ref), Columns 2 and 3, and plot bounds for $M$ up to $2$ in Figure (ref). The vertical lines in these figures designate the minimum and maximum slope, such that both the upper and lower bounds are finite numbers. The first line is the smallest slope that allows a continuous PDF to be consistent with both the bunching mass and observed income distribution. At the minimum slope, both lower and upper bounds are equal to the estimate based on the trapezoidal approximation, reported in column 1.

As $M$ increases, the set of possible PDF shapes in the bunching region becomes richer. The second line is the maximum slope before the set of possible distributions allows for a PDF that touches zero in the bunching interval. In that case, the bunching mass remains constant for arbitrarily large $\varepsilon$, and the upper bound is infinity (Theorem (ref)).

A large range between lower and upper bounds in Figure (ref) suggests the estimates change substantially with the shape of the unobserved distribution. For example, the bounds are uninformative for the self-employed married sample, even for small values of $M$. This indicates that the data will not provide precise information on the elasticity unless the researcher imposes further functional form restrictions on the distribution of $n^*$. In contrast, we learn the most in the case of all filers and self-employed not married, where the bounds are narrower than in other subsamples for $M=0.5$. The lower bound is always defined for larger choices of $M$, which gives partial information on the elasticity without the need of being precise with the choice of $M$. For the exceedingly high value of $M=2$, the lower bound is about 0.25 for all filers and at least 0.4 for the other three subsamples.

Columns 4\textendash7 report our estimates of the Tobit model using the full sample and truncated samples at 75%, 50%, and 25% of the data. Figures (ref)\textendash(ref) complement these estimates by graphing the actual distribution and the implied distribution from the Tobit estimates at different levels of truncation in panels a through e.

These estimates incorporate a list of indicator variable covariates for whether a tax return filer: filed as married, used a tax preparer, claimed a real estate interest deduction, claimed mortgage interest as a business expense, received unemployment compensation, contributed to charity, received social security benefits, took the educator expenses deduction, paid into an Individual Retirement Account (IRA) for the primary filer, claimed a student loan interest deduction, claimed exemption for children living away from home, received pension income, capital gains, farm profit or loss, a positive tax credit, positive wages, owed positive income tax before tax credits but zero after, took a self-employed health insurance deduction, paid into a Keogh retirement plan, received income from a business or farm, filed a 1040 form instead of simpler form, claimed secondary taxpayer exemption, and paid into an IRA for a secondary filer. We include year dummies variables with 1995 as the excluded year.

The fit of the Tobit model generally improves as we truncate the sample closer to the kink, which implies that the semi-parametric assumption of mixed normals is more reasonable locally than globally. The minimum truncation necessary for a reasonable fit varies by subsample. For example, for self-employed not married, the fit seems reasonable using 60% or less of the data, but for all filers, the fit only becomes reasonable at around 20%. It is interesting to observe that the Tobit with covariates fit the distribution better in narrower cuts of the data than for all filers.

Panel f in Figures (ref)\textendash(ref) graph the elasticity estimate as a function of the percentage of data used. The estimates tend to plateau as the distribution fits improve. For example, in the self-employed not married sample depicted in Figure (ref), the estimates are all around 0.75, using less than 60% of the data. It is worth pointing out that truncated samples with less than 20% of the data lead to numerical issues, such as perfect collinearity of covariates and lack of convergence in the likelihood maximization. This leads to imprecise estimates, as indicated by an upward bend in left extremity of the curves depicted in panel f of Figures (ref)\textendash(ref).\footnote{ We note that the filtering procedure that removes optimizing errors dictates the shape of the distribution close to the kink. Researchers should keep in mind, as part of their robustness checks, that different filtering methods may yield different shapes of the distribution close to the kink, which may in turn lead to different elasticity estimates. }

Comparisons Across Methods

Comparisons across methods provide insights into the reasonableness of different assumptions used to estimate the elasticity. The trapezoidal approximation is always within the bounds, because its estimate is based on a linear interpolation of the PDF in the bunching region. The slope of such line equals the minimum slope for which the bounds are defined. In contrast, the Tobit model using 100% of the data is often below the lower bounds, but the Tobit distribution fails to fit the observed distribution of income globally. Truncated Tobit estimates generally enter the bounds as the truncation window decreases and, as a result, the fit of the Tobit distribution improves. For the all filers sample, an M larger than 0.5 is needed for the bounds to cover the Tobit estimate truncated at 25%. This reiterates our previous discussion that the Tobit fit for all filers is poor until we use 20% or less of the data.

Consider self-employed married and self-employed not married filers. Figures (ref) and (ref) demonstrate that the bunching mass is approximately 6 times larger for self-employed not married individuals than for self-employed married individuals. This difference in bunching mass might lead a researcher to conjecture that the elasticity is much larger for self-employed not married individuals. Whether this conjecture is true depends, however, on differences in the underlying distribution of heterogeneity. Estimates based on the trapezoidal approximation in column 1 of Table (ref) indicate a higher elasticity for self-employed not married individuals and so do the lower bounds of our partial identification method and Tobit estimates. They disagree in the magnitudes. The trapezoidal estimate in column 1 says the elasticity for self-employed not married individuals is 9% higher compared to self-employed married individuals; conservative lower bounds in column 3 say it is 53% higher while the truncated Tobit estimates in column 7 point to a 46% difference. The disagreement across methods for these subsamples indicates that assumptions on the distribution of heterogeneity are critical to obtain informative elasticity estimates.

Conclusion

We show how to use bunching from piecewise-linear budget constraints to identify elasticities, under conditions weaker than those used in the literature on kinks and notches. The key theoretical point is that bunching is determined by the elasticity parameter and the shape of an unobserved distribution. Additional assumptions or data are needed to identify the elasticity.

We propose a suite of estimation techniques that allows researchers to tailor their estimation to different assumptions and data variation. These include nonparametric bounds and semi-parametric censored models with covariates. The nonparametric bounds are the least restrictive method and also nest estimators from the previous literature.

These techniques have wide applicability, because piecewise-linear budget constraints are common across fields, from public finance and labor, to industrial organization and accounting. Our estimation strategies also provide a foundation for future advances in techniques that will account for different empirical hurdles. Of particular interest are extensions that consider optimization and friction errors, extensive margin responses, and panel data methods.

\ifid {

Acknowledgements

The views expressed in this paper are those of the authors and do not necessarily reflect the views of the Federal Reserve Board or the Federal Reserve System. We would like to thank Matias Cattaneo, Bill Evans, Roger Gordon, Jim Hines, Dan Hungerman, Michael Jansson, Henrik Kleven, Brian Knight, Erzo Luttmer, Byron Lutz, Dayanand Manoli, Magne Mogstad, Marcelo Moreira, Whitney Newey, Andreas Peichl, Emmanual Saez, Dan Silverman, and Joel Slemrod for valuable comments and discussions. The paper also benefited from feedback received from seminar participants at the UCSD Workshop on Bunching Estimators, Econometric Society, International Association for Applied Econometrics, International Institute of Public Finance, National Tax Association, Dartmouth College, Federal Reserve Board, and University of Michigan. Jessica C. Liu, Michael A. Navarrete, and Alexis M. Payne provided excellent research assistance. All remaining errors are our own. Bertanha acknowledges financial support received while visiting the Kenneth C. Griffin Department of Economics, University of Chicago.} \fi

\singlespacing

{\color{white}

Figures and Tables

}

\singlespacing

figure[figure omitted — 2,872 chars of source]

\newgeometry{headheight=1mm,tmargin=10mm,bmargin=15mm, lmargin=10mm,rmargin=10mm}

landscape\begin{figure}[tbp] \caption{Robustness of Tobit Estimates to Lack of Normality\textemdash Experiment 1} \begin{subfigure}[b]{0.30\linewidth} \caption{Probability Density Function of $n^*$} \end{subfigure} \begin{subfigure}[b]{0.30\linewidth} \caption{Probability Mass Function of $X$} \end{subfigure} \begin{subfigure}[b]{0.22\linewidth} \caption{100% of the data used} \end{subfigure} \begin{subfigure}[b]{0.22\linewidth} \caption{60% of the data used} \end{subfigure} \begin{subfigure}[b]{0.22\linewidth} \caption{20% of the data used} \end{subfigure} \begin{subfigure}[b]{0.22\linewidth} \caption{Elasticity by percent used} \end{subfigure} \begin{subfigure}[b]{0.22\linewidth} \caption{100% of the data used} \end{subfigure} \begin{subfigure}[b]{0.22\linewidth} \caption{60% of the data used} \end{subfigure} \begin{subfigure}[b]{0.22\linewidth} \caption{20% of the data used} \end{subfigure} \begin{subfigure}[b]{0.22\linewidth} \caption{Elasticity by percent used} \end{subfigure} \caption*{ Notes: This simulation experiment illustrates that the mid-censored Tobit model is able to fit non-normal distributions of $n^*$ and retrieve the right elasticity using covariates and truncation. We generate 50,000 observations of $y$ and a scalar $X$ following Experiment 1 detailed in Section (ref). The variable $n^*$ is distributed as a mixture of two Skewed Generalized Error Distributions (Panel a), $n^*|X$ is Gaussian, the kink point is at $k=2.0794$, and $\varepsilon=1$. Panels c\textendash f correspond to estimates from a Tobit model that is correctly specified with covariate $X$. Panels c\textendash e show the histogram of simulated data for $y$ and the best-fit Tobit distributions for three truncation sizes. Panel f displays the elasticity estimate as a function of the percentage of data used in each truncated estimation, along with 95% confidence intervals. Similarly, Panels g\textendash j correspond to estimates from a Tobit model that is incorrectly specified without covariate $X$. } \end{figure}

\restoregeometry

\newgeometry{headheight=1mm,tmargin=10mm,bmargin=15mm, lmargin=10mm,rmargin=10mm}

landscape\begin{figure}[tbp] \caption{Robustness of Tobit Estimates to Lack of Normality\textemdash Experiment 2} \begin{minipage}{0.2\linewidth} \begin{subfigure}[b]{\linewidth} \caption{PDF of $n^*$} \end{subfigure} \begin{subfigure}[b]{\linewidth} \caption{PMF of $X$} \end{subfigure} \begin{subfigure}[b]{\linewidth} \caption{100% of the data used} \end{subfigure} \end{minipage} \begin{minipage}{0.67\linewidth} \begin{subfigure}[c]{\linewidth} \caption{Conditional Probability Density Functions $n^* | X=x$} \end{subfigure} \end{minipage} \caption*{ Notes: This simulation experiment illustrates that the mid-censored Tobit model is able to fit non-normal distributions of $n^*$ and retrieve the right elasticity even when the conditional distribution $n^*|X$ is not Gaussian and there is no truncation. We generate 50,000 observations of $y$ and a scalar $X$ following Experiment 2 detailed in Section (ref). The variable $n^*$ is approximately a mixture of two Skewed Generalized Error Distributions (Panel a), the distribution of $X$ is discrete (Panel b), the kink point is at $k=2.0794$, and $\varepsilon=1$. Panel c shows the histogram of simulated data for $y$ and the best-fit Tobit distribution using covariate $X$ and no truncation. Panel d displays the true conditional PDFs of $n^*|X$ in black along with the Gaussian PDFs in gray that are assumed by the Tobit model. The elasticity is estimated at $\ha\varepsilon = 1.0083$ (S.E. 0.0073). } \end{figure}

\restoregeometry

landscape\begin{table} \caption{Estimates Using U.S. Tax Returns 1995--2004} \scalebox{0.94}{ \begin{tabular}{lccccccc|lc} \hline\hline & (1) & (2) & (3) & (4) & (5) & (6) & (7) & \multicolumn{2}{c}{(8)} \\ Statistical Model & Trapezoidal & Theorem (ref) & Theorem (ref) & Tobit & Tobit & Tobit & Tobit & \\ & Approximation & Bounds & Bounds & Full Sample & Trunc. 75% & Trunc. 50% & Trunc. 25% & \multicolumn{2}{c}{Sample} \\ & & M = 0.5 & M = 1 & & & & & \multicolumn{2}{c}{details} \\ \hline All & & & & & & & & Obs. & 188.3m \\ Elasticity $\left(\varepsilon\right)$ & 0.323 & $\left[ 0.297, 0.364 \right]$ & $\left[ 0.276, 0.448 \right]$ & 0.163 & 0.242 & 0.250 & 0.278 & Avg. & \$53.5k \\ & & & & (0.0001) & (0.0002) & (0.0002) & (0.0002) & Std. & \$64.6k \\ & & & & & & & & \\ Self-employed & & & & & & & & Obs. & 33.4m \\ Elasticity $\left(\varepsilon\right)$ & 0.811 & $\left[0.686,1.183\right]$ & $\left[0.612,\infty \right]$ & 0.568 & 0.752 & 0.746 & 0.750 & Avg. & \$60.7k \\ & & & & (0.0005) & (0.0007) & (0.0008) & (0.0008) & Std. & \$77.2k \\ Self-employed, & & & & & & & & & \\ married & & & & & & & & Obs. & 23.9m \\ Elasticity $\left(\varepsilon\right)$ & 0.770 & $\left[0.563,\infty \right]$ & $\left[0.475, \infty \right]$ & 0.304 & 0.464 & 0.540 & 0.554 & Avg. & \$73.6k \\ & & & & (0.0006) & (0.0009) & (0.0010) & (0.0010) & Std. & \$84.4k \\ Self-employed, & & & & & & & & \\ not married & & & & & & & & Obs. & 9.5m \\ Elasticity $\left(\varepsilon\right)$ & 0.842 & $\left[0.777, 0.936 \right]$ & $\left[0.728, 1.107\right]$ & 0.875 & 0.774 & 0.745 & 0.807 & Avg. & \$28.3k \\ & & & & (0.0010) & (0.0009) & (0.0010) & (0.0015) & Std. & \$39.6k \\ & & & & & & & & \\ \hline \end{tabular} } \captionsetup{font=footnotesize,justification=raggedright,singlelinecheck=false}\caption*{\textit{Notes:} The table shows estimates of the elasticity for four different subsamples of the IRS data, and using three different approaches discussed in the paper. The first approach (column 1) uses the trapezoidal approximation to point-identify the elasticity (Example (ref)). We obtained non-parametric estimates of the side limits of $f_y$ at the kink using the method of cattaneo2019. The estimate for the bunching mass equals the sample proportion of $y$ observations that equals the kink point (see discussion on friction errors in Section (ref)). The second approach (columns 2 and 3) uses the same estimates of the bunching mass and side limits to compute partially identified sets for the elasticity (Theorem (ref)). Upper and lower bounds are calculated for two choices of M, that is, the maximum slope of the PDF of the unobserved heterogeneity $n^*$. Column 4 has Tobit MLE estimates of the elasticity that utilizes the full sample of data, along with robust standard errors in parentheses. Columns 5 through 7 report truncated Tobit MLE estimates. As we move from column 5 to column 7, we restrict the estimation sample to shrinking symmetric windows around the kink that utilizes 75% to 25% of the data. The set of covariates that enters the Tobit estimation is kept constant across different truncation windows and are listed in Section (ref). } \end{table}
figure[figure omitted — 1,864 chars of source]
landscape\begin{figure}[tbp] \caption{Truncated Tobit - All Filers} \begin{subfigure}[b]{0.455\textwidth} \caption{100% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{80% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{60% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{40% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{20% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{Elasticity by percent used} \end{subfigure} \caption*{ Notes: the figure displays best-fit Tobit distributions and elasticity estimates for various choices of a symmetric truncation window around the kink point. The set of covariates that enters the Tobit estimation is kept constant across different truncation windows and are listed in Section (ref). Panels a through e show the histogram of income for all filers (bars), along with the best-fit Tobit PDF for each truncation window (line). The best-fit PDF is constructed using the truncated Tobit likelihood averaged over covariate values in the sample. Panel f displays the Tobit elasticity estimate as a function of the percentage of data used in estimation. } \end{figure}
landscape\begin{figure}[tbp] \caption{Truncated Tobit - Self-employed Filers} \begin{subfigure}[b]{0.455\textwidth} \caption{100% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{80% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{60% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{40% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{20% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{Elasticity by percent used} \end{subfigure} \caption*{ Notes: the figure displays best-fit Tobit distributions and elasticity estimates for various choices of a symmetric truncation window around the kink point. The set of covariates that enters the Tobit estimation is kept constant across different truncation windows and are listed in Section (ref). Panels a through e show the histogram of income for self-employed filers (bars), along with the best-fit Tobit PDF for each truncation window (line). The best-fit PDF is constructed using the truncated Tobit likelihood averaged over covariate values in the sample. Panel f displays the Tobit elasticity estimate as a function of the percentage of data used in estimation. } \end{figure}
landscape\begin{figure}[tbp] \caption{Truncated Tobit - Self-employed and Married Filers} \begin{subfigure}[b]{0.455\textwidth} \caption{100% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{80% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{60% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{40% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{20% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{Elasticity by percent used} \end{subfigure} \caption*{ Notes: the figure displays best-fit Tobit distributions and elasticity estimates for various choices of a symmetric truncation window around the kink point. The set of covariates that enters the Tobit estimation is kept constant across different truncation windows and are listed in Section (ref). Panels a through e show the histogram of income for self-employed and married filers (bars), along with the best-fit Tobit PDF for each truncation window (line). The best-fit PDF is constructed using the truncated Tobit likelihood averaged over covariate values in the sample. Panel f displays the Tobit elasticity estimate as a function of the percentage of data used in estimation. } \end{figure}
landscape\begin{figure}[tbp] \caption{Truncated Tobit - Self-employed and Not Married Filers} \begin{subfigure}[b]{0.455\textwidth} \caption{100% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{80% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{60% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{40% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{19% of the data used for estimation} \end{subfigure} \begin{subfigure}[b]{0.455\textwidth} \caption{Elasticity by percent used} \end{subfigure} \caption*{ Notes: the figure displays best-fit Tobit distributions and elasticity estimates for various choices of a symmetric truncation window around the kink point. The set of covariates that enters the Tobit estimation is kept constant across different truncation windows and are listed in Section (ref). Panels a through e show the histogram of income for self-employed and not married filers (bars), along with the best-fit Tobit PDF for each truncation window (line). The best-fit PDF is constructed using the truncated Tobit likelihood averaged over covariate values in the sample. Panel f displays the Tobit elasticity estimate as a function of the percentage of data used in estimation. } \end{figure}