EconBase
← Back to paper

Nonparametric Difference-in-Differences in Repeated Cross-Sections with Continuous Treatments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

97,713 characters · 24 sections · 41 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonparametric Difference-in-Differences in Repeated Cross-Sections with Continuous Treatments

tabular[tabular omitted — 561 chars of source]

}

abstract{0.45cm} This paper studies the identification of causal effects of a continuous treatment using a new difference-in-difference strategy. Our approach allows for endogeneity of the treatment, and employs repeated cross-sections. It requires an exogenous change over time which affects the treatment in a heterogeneous way, stationarity of the distribution of unobservables and a rank invariance condition on the time trend. On the other hand, we do not impose any functional form restrictions or an additive time trend, and we are invariant to the scaling of the dependent variable. Under our conditions, the time trend can be identified using a control group, as in the binary difference-in-differences literature. In our scenario, however, this control group is defined by the data. We then identify average and quantile treatment effect parameters. We develop corresponding nonparametric estimators and study their asymptotic properties. Finally, we apply our results to the effect of disposable income on consumption. \newline Keywords: identification, repeated cross-sections, nonlinear models, continuous treatment, random coefficients, endogeneity, difference-in-differences. \newline ${}$

Introduction

Differences-in-Differences (DID) is arguably one of the most popular methods for policy evaluation. In its standard version, it allows to identify the causal effect of a binary treatment on a given outcome, even when units are not allocated randomly to the treatment. The idea is to compare the evolution of the average outcome of the treatment group, which receives a treatment after a certain date, with that of the control group, which remains untreated. The DID strategy builds on the so-called common trend assumption, viz. the assumption that the changes in $Y(0)$, the potential outcome absent the treatment, are identical between the treatment and the control group. One way to see this condition is that treatment changes should be exogenous in that they are not related to changes in $Y(0)$. Under common trends, the average time trend on $Y(0)$ can be identified using the control group, as this group does not experience the effect of the treatment. Once this time trend is accounted for, we can identify the average treatment effect by a simple before-after comparison on the treatment group.

A crucial limitation of the standard DID framework is that the treatment is required to be binary. Yet, in many cases, units experience various treatment intensities, and not just a zero-one treatment. Examples include, among many others, unemployment benefits, specific public expenditures (e.g., hospitals expenditure per capita, teacher's wages), changes in prices, in income etc. Usual solutions in such cases are to either consider a linear model or to discretize the treatment. Both solutions are problematic. In the first case, the model cannot account for, e.g., unobservable terms affecting both treatment intensity and treatment effects, effectively assuming that the treatment has the same effect for every level of treatment intensity. In the second case, discretization introduces arbitrariness and leads to an information loss. E.g., after discretization, even vastly different income changes would be assumed to have the same effect, because all that matters is the fact that they change. This makes the DID framework not useful to study, e.g., causal marginal propensities to consume out of income, because individuals at low levels of income are more likely to face liquidity constraints than individuals at high levels, and the size of the income change arguably matters hsieh2003consumers.

In this paper, we propose a solution that circumvents these issues. Specifically, we show identification of several treatment effect parameters, allowing for nonlinear and heterogeneous effects of the treatment without imposing functional form restrictions (e.g., linearity), nor any discretization. We also allow for heterogeneous time trends on potential outcomes. This is important, since assuming the same time trends for all units may be overly restrictive. Firms or individuals with different productivities may be affected differently by macroeconomic shocks, for instance.

The idea behind our identification strategy retains the spirit of the DID approach. We use the fact that the distribution of the treatment, in our case a continuous random variable, changes over time for some exogenous reason, yet some units remain at the same level of treatment. In our approach, the latter units form the control group, which allows us to identify the (heterogeneous) time trend on potential outcomes. Once this time trend has been removed, the distribution of the appropriately modified potential outcome does not vary over time any longer. Then, any difference over time in the distribution of the modified, observed outcome should be solely due to the treatment. With this insight, we can identify causal effects of the treatment.

To make this strategy operational, we rely on three main assumptions. First, we assume that units sharing the same rank in the distribution of the treatment at two different periods have the same distribution of the unobservables governing potential outcomes. This assumption is related to the aforementioned exogeneity of the change in the distribution of the treatment over time. It first ensures that groups, similar to the control and treatment group with binary treatments, may be defined through the ranking of the treatment. With this construction of groups, the condition becomes almost the same as Assumptions 1 and 3 in Athey06, on which our paper builds to identify (heterogeneous) time trends. Second, we assume that a unit with stable unobservables in two different periods will also have the same ranking in the distributions of potential outcome in these two periods. Again, this assumption is identical to Assumption 2 in Athey06. Third, we suppose that the change in the distribution of the treatment is heterogeneous, in the sense that the cumulative distribution functions (cdfs) of the treatment variable between the two periods cross. This crossing point defines the control group and allows us to identify the heterogeneous time trends.

Despite some similarities with the nonlinear difference-in-differences setting of Athey06, our continuous treatment set-up exhibits several important distinct features. First, Athey06 focus on the binary treatment case, while we consider a continuous treatment. Second, our control group is determined by the data rather than fixed ex ante. While our paper also shares some similarities with the paper by Chaisemartin14, the framework and identification strategy is nonetheless very different: In particular, Chaisemartin14 focus on binary (or ordered) treatments. With a continuous treatment, their strategy would require to have a control group for which the whole distribution of the treatment variable would remain unchanged over time, an assumption unlikely to be satisfied in practice. In contrast, we only require the distribution of the treatment to change in such a way that there exists a crossing point. A change in the mean and the variance of a normal distribution, for instance, satisfies this requirement.

We consider several extensions to our main setup. First, we show how covariates can be included into our analysis. Second, we establish that our model extends in a straightforward way to multidimensional continuous treatments. Third, we show that while a number of parameters cannot be point identified, basically because time only provides us with limited exogenous variations, they can be partially identified using weak local curvature conditions. Finally, we prove that under functional form restrictions that still allow for ample heterogeneity, we can point identify all marginal effects.

Based on this extensive identification analysis, we also develop nonparametric sample counterparts estimators. While our estimators of the average and quantile effects involve several nonparametric steps, each one of these steps is straightforward to perform. We show the asymptotic normality of the estimators. When the \textquotedblleft control group\textquotedblright\ corresponds to a single point of support of the continuous treatment, the estimators are not root-n consistent, but converge at standard univariate nonparametric rates.

Finally, we apply our methodology to analyze the marginal propensity to consume in the US. We exploit for that purpose a change in the schedule of the Earned Income Tax Credit (EITC) between 1987 and 1989. We argue that this change generates exactly the crossing condition we require. Applying our method, we obtain an estimated time trend that displays heterogeneity, underlying the need to go beyond mere additive time trends. Moreover, our estimates of the marginal effects suggest that low-income individuals increase substantially their consumption (by around 50%), while medium-income individuals would not significantly adjust their consumption. This is in line with many findings in the literature, see e.g., johnson2006household and kaplan2014model.

The paper is organized as follows. In Section 2, we introduce the model formally, discuss the parameters of interest and provide our main identification results. The extensions considered above are discussed in Section 3. Section 4 is devoted to estimation. Section 5 presents the application, and Section 6 concludes. All proofs are gathered in the appendix.

Model and Main Identification Results

Assumptions

We consider a potential outcome framework with a continuous treatment. The potential outcome at period $t$, corresponding to a treatment $x\in \mathcal{ X}\subset \mathbb{R}$, is denoted by $Y_{t}(x)$, with $Y_t(x)\in\mathbb{R}$. We observe, at each period $t\in \{1,...,T\}$, the actual treatment $X_{t}$ and the corresponding outcome, $Y_{t}\equiv Y_{t}(X_{t})$. We are particularly interested in the average and quantile treatment on the treated effects:

eqnarray*[eqnarray* omitted — 219 chars of source]

for any $x$ and $x^{\prime }$ in the support $\text{Supp}(X_T)$ of $X_{T}$. Here, $F_{A|B}(a|b)$ denotes the conditional cdf of a random variable $A$ at $a$, given that a random vector $B$ takes the value $b$, and $ F_{A|B}^{-1}(\tau |b)$ denotes its inverse, the $\tau $-conditional quantile function. We henceforth focus on the effects at period $T,$ because they are the most natural to compute in general, but we can identify similar effects at any date.

The main issue in identifying the parameters above is endogeneity of the actual treatment, i.e., $X_{t}$ may depend on $(Y_{t}(x))_{x\in \mathcal{X}}$. In such a case, naive estimators do not coincide with the average and quantile treatment effects defined above. For instance, $ E(Y_{T}|X_{T}=x^{\prime })-E(Y_{T}|X_{T}=x)\neq \Delta ^{ATT}(x,x^{\prime })$ . Our idea for identifying these causal parameters, then, is to use exogenous changes in $X_{t}$ (due to, e.g., a policy change), and apply a difference-in-difference type strategy. To make this idea operational, we restrict the way time affects both observed and unobserved variables by imposing three main restrictions. The first restriction is a stationarity condition on the observed and unobserved determinants of the outcome. The second restriction limits the way time affects the outcome itself. The third restriction affects the way the distribution of $X_{t}$ changes over time. We discuss them in turn using the notation $V_{t}=F_{X_{t}}(X_{t})$ to denote the rank of an individual in the distribution of the treatment.

assumption(stationarity of unobservables) For all $t \in \{1,...,T\}$, $Y_t(x)=g_t(U_t(x))$ where for all $(x,v)\in \mathcal{X}\times [0,1]$ the distribution of the unobserved variable $U_t(x)|V_t=v$ does not depend on $t$.

We can interpret this assumption as follows. First, it defines implicitly groups, similar to control and treatment groups with binary treatments, through the rank variable $V_t$. Then, we assume that within each group, unobserved terms related to potential outcomes have a time-invariant distribution. This latter condition is similar to Assumptions 3.1 and 3.3 in Athey06, where the authors also assume that within both the control and the treatment group the distribution of the unobserved term related to $ Y_{t}(0)$ is constant over time.

Importantly, Assumption (ref) does not restrict the cross-sectional dependence between $U_{t}(x)$ and $V_{t}$, which is at the core of the endogeneity problem we face in this scenario. On the other hand, it rules out changes in the type of endogeneity, as the distribution of $(U_{t}(x),V_{t})$ is supposed to be time invariant. In our application below, $X_{t}$ corresponds to disposable income. The tax rate affects disposable income, but a change in the tax rate is unlikely to change the ranking of individuals in the income distribution, holding other characteristics constant (e.g., the number of household members). In other applications, this condition may be more restrictive. We discuss this point further when we draw a parallel with instrumental variable models in Section (ref) below.

The following assumption specifies the second requirement mentioned above:

assumption(rank invariance on the time trend) For all $(x,t)\in \mathcal{X}\times \{1,...,T\}$, $U_t(x)\in \mathbb{R}$ and $g_t$ is strictly increasing. Without loss of generality, we let $g_T(y)=y$ for all $y\in\text{Supp}(Y_T)$ .

Assumption (ref) is again materially identical to Assumption 3.2 in Athey06. Combined with Assumption (ref), it states that an individual that has the same unobservable in two different periods (i.e., $U_{t}(x)\equiv U_{t^{\prime }}(x)$) will also have the same ranking in the distributions of potential outcomes $Y_{t}(x)$ and $Y_{t^{\prime }}(x)$ in the same two periods. Assumption (ref) generalizes the standard translation model $g_{t}(u)=\delta _{t}+u$ to allow for heterogeneous time trends. This can be important in some applications. For instance, macroeconomic shocks may have different effects on high- and low-wage earners. Note that given the strict monotonicity condition, we can always make the normalization $ g_{T}(y)=y$, by just redefining $U_{t}(x)$ as $g_{T}(U_{t}(x)) $ and $ g_{t}(y)$ as $g_{t}\circ g_{T}^{-1}(y)$.

assumption(crossing points) For all $t\in \{1,...,T-1\}$, there exists $x_{t}^{\ast }\in \mathbb{R}$ such that $F_{X_t}(x_{t}^{\ast })=F_{X_T}(x_{t}^{\ast })\in (0,1) $.

Contrary to Assumptions (ref)-(ref), Assumption (ref) only involves observables, and is therefore directly testable in the data. It means, roughly speaking, that the exogenous change (induced by, e.g., a policy change) affects individuals' treatment in a heterogeneous way. Requiring time to have a heterogeneous effect on the treatment is also required in the usual difference-in-difference strategy, and a similar condition is required in fuzzy settings considered by Chaisemartin14.

Note that $x_{t}^{\ast }$ can be identified and estimated using the data. We consider such an estimator in Section (ref) below. However, sometimes the value of the crossing point may also be inferred from the design of the policy change. Comparing the theoretical crossing point with the crossing point obtained from the data then constitutes a check for the hypothesis that the policy change has not changed the distribution of the unobservables. We refer to the application below for more details about this.

Two additional remarks on Assumption (ref) are in order. First, this assumption holds if $F_{X_{t}}$ remains constant with $t$ . In this case, however, we identify only the trivial parameters $\Delta ^{ATT}(x,x)=\Delta ^{QTT}(p,x,x)=0$. Second, we assume for simplicity crossings between the cdf of $X_{T}$ and all other cdfs, but actually, $T-1$ crossings are sufficient, provided that we can \textquotedblleft relate\textquotedblright\ them to each other, for instance if the cdf of $ X_{t}$ crosses that of $X_{t+1}$ for $1\leq t<T$. Also, with only one crossing between $F_{X_{s}}$ and $F_{X_{t}}$, we still identify some treatment effects at periods $s$ or $t$, following the same logic as in Section (ref) below. So even if Assumption (ref) does not make this apparent, adding periods help because it increases the odds of having at least one crossing point, which is sufficient for identifying some causal parameters.

The last assumption we impose is a regularity condition:

assumption(regularity conditions) For all $t\in\{1,...,T\}$, $E(|Y_t|)<\infty$ and $ F_{X_t}$ is continuous on $\text{Supp}(X_t)$, which is an interval included in $\mathcal{X}$. For all $x^{\prime }\in \text{Supp}(X_t)$ and $u\in \text{ Supp}(U_t(x^{\prime }))$, there exist versions of $E[Y_t|X_t]$ and $ P^{U_t(x^{\prime }) |V_t}$ such that $x\mapsto E[Y_t|X_t=x]$ and $v \mapsto F_{U_t(x^{\prime })|V_t}(u|v)$ are continuous.

The continuity conditions are mild, yet important to define properly conditional expectations or cdfs (e.g., $E[Y_{t}|X_{t}=x_{t}^{\ast }]$ or $ F_{Y_{t}|X_{t}}(y|x_{t}^{\ast })$).

Examples

To better understand the types of data generating processes which our assumptions permit, we consider two examples of workhorse models.

Simple Linear Systems

Let us suppose that

align[align omitted — 144 chars of source]

where $(\alpha _{t},\beta ,\gamma _{t},\delta _{t})$ are constants and the marginal distribution of $(U_{t},\eta _{t})$ is assumed constant over time. Suppose also that $\text{Supp}(\eta _{t})=\mathbb{R}$ and $\delta _{t}\neq \delta _{T}$ for all $t\neq T$. Then, Assumptions (ref)-(ref) hold with $U_{t}(x)=x\beta +U_{t}$, $g_{t}(u)=\alpha _{t}+u$ and $x_{t}^{\ast }=(\gamma _{t}-\gamma _{T})/(\delta _{T}-\delta _{t})$. Assumption (ref) holds under mild restrictions on the distribution of $(U_{t},\eta _{t})$. Note that the model allows for any dependence between $U_{t}$ and $\eta _{t}$. Thus, $X_{t}$ is endogenous in general in the outcome equation, and we cannot recover $\beta $ directly by the OLS. Note, moreover, that we cannot use time as an instrumental variable in the outcome equation either, because it has a direct effect on $Y_{t}$, so none of the standard tools work.

As mentioned above, if the policy change has a pure location effect on $X_t$, so that $\delta_t=\delta_T$ for all $t$, then Assumption (ref) is not satisfied. We require to have individuals unaffected by the change, and this holds with a change in scale in (ref).

Note that we did not impose any condition on the dependence between $(U_{s},\eta _{s})$ and $(U_{t},\eta _{t})$. Hence, the model allows for any form of serial dependence of the unobservables. On the other hand, models with a lagged dependent variable are typically ruled out by our assumptions. To see this, suppose that we replace (ref) by

equation[equation omitted — 88 chars of source]

Then,

equation*[equation* omitted — 75 chars of source]

with $\widetilde{\alpha }_{t}=\alpha +\sum_{k=1}^{\infty }\rho ^{k}\left[ \alpha _{t-k}+\beta \gamma _{t-k}\right] $ and $\widetilde{U} _{t}=U_{t}+\sum_{k=1}^{\infty }\rho ^{k}[\beta \delta _{t-k}\eta _{t-k}+U_{t-k}]$. Therefore, the distribution of $\widetilde{U}_{t}$ depends on $t$ in general, unless $\rho =0$. Then,

equation*[equation* omitted — 206 chars of source]

which depends on $x$ in general when $\beta\neq 0$ and $\rho\neq 0$. On the other hand, Assumptions (ref)-(ref) imply that for any $(t,t^{\prime })$, $F_{Y_{t}(x)}^{-1}\circ F_{Y_{t^{\prime }}(x)}(y)$ does not depend on $x$. In other words, Assumptions (ref)-(ref) are violated in general when $ \rho \neq 0$.

This feature is not specific to our assumptions. A similar issue arises in the standard difference-in-differences setup. To see this, consider model (ref) again, but now with a binary treatment, for which $X_{t}=0$ for all $t\leq 1,$ and $X_{2}=G$, the dummy of being in the treatment group as opposed to the control group in the second period. Suppose, moreover, that $E(U_{t}|G)$ does not depend on $t$. In the case of $ \rho =0$, the common trend condition is satisfied with $ E(Y_{2}(0)|G)-E(Y_{1}(0)|G)=\alpha _{2}-\alpha _{1}$. But if $\rho \neq 0$, then the common trends assumptions fails to hold in general, since

equation*[equation* omitted — 146 chars of source]

which depends on $G$.

Quantile Regression Type Models

The previous model does not allow for heterogeneous treatment effects or heterogeneous time trends on potential outcomes. We may, however, analyze models with heterogeneous features as they are compatible with our assumptions. The following model exemplifies this:

align*[align* omitted — 103 chars of source]

where we assume that the marginal distribution of $(U_{t},\eta _{t})$ is constant over time, $\text{Supp}(\eta _{t})=\mathbb{R}$, and for all $t\neq T $, there exists $e_{t}$ such that $h_{t}(e_{t})=h_{T}(e_{t})$. We also assume that $f_{t}$ and $e\mapsto \alpha (e)+x\beta (e)$ are strictly increasing. In this scenario, Assumptions (ref)-(ref) are satisfied, with $U_{t}(x)=f_{T}(x\beta (U_{t}))$ and $ g_{t}(y)=f_{t}\circ f_{T}^{-1}(y)$. Contrary to the previous example, this model allows for both heterogeneous treatment effects, through the random coefficient $\beta (U_{t})$, and an heterogeneous time trend, through the function $f_{t}$. In the special case where $f_{t}(y)=y+\gamma_t$, the model is a linear correlated random coefficients model. Note that even with such a restriction on $f_{t}$, the treatment effect function $e\mapsto \beta (e)$ cannot be identified through standard quantile regression of $Y_{t}$ on $ X_{t}$, because of the dependence between $X_{t}$ and $U_{t}$.

Main Identification Results

Our identification strategy works in two steps: In the first step, we identify the effect of time on the outcome, i.e., the function $g_{t}$. This implies that we identify $\widetilde{Y}_{t}=g_{t}^{-1}(Y_{t})$, whose distribution does not depend on time anymore (conditional on $V_{t})$. Then, in a second step, we can use time as an instrument to recover specific causal effects. For ease of exposition, we first outline our method in the case of $T=2$.

Step 1: Identification of the Time Trend

To recover $g_{1}$, we rely on observations at the crossing point, i.e. observations for whom $X_{1}=x_{1}^{\ast }$. Under Assumptions (ref)-(ref), the following is true:

align[align omitted — 551 chars of source]

where we indicate the respective assumptions employed by superscripts upon equalities. As a result, $g_{1}$ is identified by

equation[equation omitted — 127 chars of source]

Hence, under our assumptions, the time trend $g_{1}$ can be identified using observations for which $X_{1}=x_{1}^{\ast }$ and $X_{2}=x_{1}^{\ast }$. These two sets of observations, though distinct as we use repeated cross sections, have the same distribution of unobservables and the same value of the treatment. Therefore, differences between the distributions of outcomes can only stem from the effect of time itself. This idea is very similar to that used in difference-in-differences, where the control group permits the identification of the (common) time trend. For this reason, in what follows we classify all observations satisfying $X_{1}=x_{1}^{\ast }$ to form the \textquotedblleft control group\textquotedblright .

Note that our model allows for heterogeneous time trends. As Athey06, we therefore recover a whole function $g_{1}$ rather than a single coefficient for the time trend, as in the standard difference-in-differences model. Also as Athey06, we identify $g_{1}$ by a quantile-quantile transform. When it comes to the identification of the time trend, the main difference between our approach and that of Athey06 lies in how the control group is defined. While it is defined ex ante in Athey06, it is data-driven and defined through the crossing points here.

Beyond the identification of $g_{1}$, (ref) reveals that the model is testable, if there are several crossing points between $ F_{X_{1}}$ and $F_{X_{2}}$, say $x_{1}^{\ast }$ and $x_{1}^{\ast \ast }$. In such a case, our model implies indeed that for all $y$,

equation*[equation* omitted — 192 chars of source]

which is a testable restriction. Related to this, if the true set of crossing points is an interval $I$, say, we have $F_{X_1|X_1\in I} = F_{X_2|X_2\in I}$. Then, integrating (ref) over $x_1^*\in I$ , we obtain

equation*[equation* omitted — 101 chars of source]

Therefore, $g_1(y)=F^{-1}_{Y_1|X_1\in I}\left[F_{Y_2|X_2\in I}(y)\right]$, which implies that $g_1$ could be in principle estimated at a parametric rather than nonparametric rate in this case.

Step 2: Identification of ATT and QTT

Next, we consider the identification of the treatment effects $ \Delta^{ATT}(x,x^{\prime })$ and $\Delta ^{QTT}(p,x,x^{\prime })$. Start out by considering the transformed potential and observed outcomes $\widetilde{Y} _t(x)=g_t^{-1}(Y_t(x))$ and $\widetilde{Y}_t=g_t^{-1}(Y_t)$ for $t=1,2$. By virtue of Assumption (ref), $F_{\widetilde{Y}_1(x)}=F_{ \widetilde{Y}_2(x)}$. Time can thus be seen as an instrument for the treatment: while it affects the treatment (or its distribution, to be precise), it has no direct effect on potential outcomes. The same idea is used in a different DID framework by Chaisemartin14.

To proceed with the identification of our model, let $ q_t=F_{X_t}^{-1}\circ F_{X_T}$. Thus, $q_t(x)$ denotes the value of $X_t$ (say, income in period $t$) for an individual at the same rank as another individual whose period $T$ income is $X_T=x$. Then,

eqnarray*[eqnarray* omitted — 205 chars of source]

By the normalization $g_2(y)=y$, the latter is the mean counterfactual outcome at period 2 for individuals with $X_2=x$ if $X_2$ was moved exogenously to $q_1(x)$. We can therefore identify $\Delta ^{ATT}(x,q_1(x))$ , the average effect of moving $X_2$ from their initial value $x$ to $q_1(x)$ , by

eqnarray*[eqnarray* omitted — 187 chars of source]

This means that we can obtain $\Delta^{ATT}(x,x^{\prime })$ for any pair $ (x,x^{\prime })$ such that $x^{\prime }=q_1(x)$.

Similarly, we have, for any $p\in (0,1)$,

eqnarray*[eqnarray* omitted — 255 chars of source]

This implies that

equation*[equation* omitted — 120 chars of source]

Theorem (ref) summarizes our findings so far, and generalizes it to any value of $T$.

theoremUnder Assumptions (ref)-(ref), we identify, for all $x\in \text{Supp}(X_T)$, $p \in (0,1)$ and $t\in \{1,...,T-1\}$, the functions $g_t$ and the average and quantile treatment effects $\Delta ^{ATT}(x,q_t(x))$ and $\Delta ^{QTT}(p,x,q_t(x))$.

Note that if $x\mapsto Y_T(x)$ is differentiable, we have, by the mean value theorem,

equation*[equation* omitted — 82 chars of source]

for some random term $\widetilde{X}\in [x, q_t(x)]$. As a result, by Theorem (ref), we identify

equation[equation omitted — 156 chars of source]

In other words, $\Delta^{ATT}(x,q_t(x))/(q_t(x) - x)$ may be interpreted as an average marginal effect for units at $X_T=x$. Contrary to usually, however, the derivative of $Y_T(.)$ is not evaluated at the current treatment value $x$, but at another point $\widetilde{X}\in (x, q_t(x))$. If $q_t(x)$ is close to $x$ or $Y_T$ is close to being linear, we can nevertheless expect $Y_T^{\prime }(\widetilde{X})$ to be close to the usual term $Y^{\prime }_T(x)$. As shown in Appendix (ref), we actually exactly identify the usual average marginal effect $ \Delta^{AME}(x)\equiv E[Y_T^{\prime }(x) | X_T=x]$ at some particular values of $x$.

Equation (ref) also implies that we can identify average marginal effects on larger subpopulation. Specifically, let $ I_c=\{x\in \text{Supp}(X_T):|q_t(x)-x|>c\}$ for some $c>0$. Then, we identify

equation*[equation* omitted — 148 chars of source]

The advantage of considering this object is statistical accuracy, as we average over the subpopulation such that $X_T\in I_c$.

With $T>2$, more periods produce more variations and thus allow one to identify more treatment effects. Also, while Theorem (ref) establishes the identification of treatment effects at period $T$, the same reasoning yields the identification of treatment effects at any other periods. To see this, note that

align*[align* omitted — 345 chars of source]

The right-hand side is identified, since $g_{t}$ is identified, as outlined above. Hence, we identify all period $t$-parameters of the form $E\left[ Y_{t}(q_{t}^{-1}(x))-Y_{t}(x)|X_{t}=x\right] $ and $ F_{Y_{t}(q_{t}^{-1}(x))|X_{t}=x}^{-1}-F_{Y_{t}(x)|X_{t}=x}^{-1}$.

If $F_{X_{t}}$ does not vary over time, then $q_{t}(x)=x$ and Theorem (ref) boils down to the identification of the trivial parameters $\Delta ^{ATT}(\xi ,\xi )=0$ an $\Delta ^{QTT}(\xi ,\xi )=0$. As mentioned above, the distribution of $X_{t}$ needs to vary for our method to have non-trivial identification power. Finally, we cannot point identify from Theorem (ref) the parameters $\Delta ^{ATT}(x,x^{\prime })$ and $\Delta ^{QTT}(x,x^{\prime })$ if $x^{\prime }\neq q_{t}(x)$ for some $t\in \{1,...,T-1\}$. We show however in Subsection (ref) that we can at least set identify these parameters under plausible curvature restrictions, and in Subsection (ref) that we can point identify them under stronger conditions.

Relationship to other Approaches

Comparison with Panel Data Models

While we rely on time variation to identify causal effects, as in panel data, our assumptions contrast with those typically used in panel data. First, our stationarity condition is different from the condition

equation[equation omitted — 94 chars of source]

commonly assumed in panel data Manski87,Honore92,Hoderlein12,Graham12,Chernozhukov13,chernozhukov2015nonparametric . To understand the differences between the two, consider two polar cases. In the first, endogeneity stems from contemporaneous simultaneity between $ U_{t}(x)$ and $V_{t}$, as is often the case with variables that are jointly determined, while $(U_{t}(x),V_{t})_{t=1...T}$ are i.i.d. across time. If so, Assumption (ref) is satisfied. On the other hand, (ref) does not hold, unless $U_{t}(x)$ is independent of $ V_{t}$, because the distribution of $U_{s}(x)$ conditional on $ (X_{1},...,X_{T})$ is a function of $X_{s}$ only, i.e., $ f_{U_{s}(x)|X_{1},...,X_{T}}(a|x_{1},...,x_{T})=f_{U_{s}(x)|X_{s}}(a|x_{s})$ , while the conditional distribution $U_{t}(x)$ is a function of $X_{t}$ only, and they do not coincide in general if $x_{s}\neq x_{t}$. Assuming $ (U_{s}(x),V_{s})$ independent of $(U_{t}(x),V_{t})$ is of course often unrealistic, but the same conclusion would hold with, say, a vector autoregressive structure.

In the second case, $U_{t}(x)=(A(x),U_{t})$ where $A(x)$ is an individual effect potentially correlated with $X_{1},...,X_{T}$ and $ (U_{t})_{t=1}^{T}$ are i.i.d. idiosyncratic shocks that are independent of $ (A(x),X_{1},...,X_{T})$. In this case, the condition (ref) is always satisfied. On the other hand, Assumption (ref) holds only under a special correlation structure between $A(x)$ and $ (X_{1},...,X_{T})$: $A(x)|V_{t}=v\sim A(x)|V_{s}=v,$ which for instance imposes Cov$(A(x),V_{t})=$ Cov$(A(x),V_{s}),\;s\neq t$. While this still allows for arbitrary contemporaneous correlation between $A(x)$ and $V_{t}$, it does not allow for any time-varying covariance.

Another difference with panel data models lies in the type of variations that we require on $X_{t}$. With panels, we require the individual value of the treatment to vary over time, the fixed effects absorbing any variable that is constant across time. Such a requirement is not needed here, since the distribution of $X_{t}$ can change over time even if $X_{t}$ is constant for each individual, provided new generations are involved at date $t$ compared to date $s$. On the other hand, compared to panel data, we do not identify anything here, apart from the time trend $ g_{t}$, when the treatment changes at an individual level but the distribution of $X_{t}$ remains constant over time. This is one key aspect that distinguishes our identification strategy from panel data based strategies.

Comparison with instrumental variable models

Our result is also related to the literature on identification of triangular models with instruments and cross-sectional data Imbens09. Such models take the following form:

align*[align* omitted — 46 chars of source]

where $Z$ denotes the instrument, $V\in\mathbb{R}$ $h(z,\cdot)$ is increasing and $(U, V) \perp\!\!\!\perp Z$. We rely on a similar structure here, with $Z$ playing the role of the instrument. Assumption (ref) then corresponds to the condition $(U, V) \perp\!\!\!\perp Z$. Our model still has one distinctive feature from this model: time may have a direct effect on the outcome variable, though this effect has to be restricted through Assumption (ref).\footnote{ Another difference with Imbens09 is that the instrument is discrete in our setup. As a result, some common parameters such as the overall average marginal effects are not identified without further restrictions.} The role of the crossing condition, then, is to pin down this effect, so that we can modify the outcome in such a way that time becomes a valid instrument.

This parallel also illustrates some possible limitations of our approach. In particular, Kasy11 showed that if in reality $V$ is multidimensional (with still $(U, V) \perp\!\!\!\perp Z$), then in general $ U $ is not independent of $Z$ conditional on $F_{X|Z}(X|Z)$. In our context, this means that Assumption (ref) fails to hold if $X_t$ depends on a multiple unobserved terms. Consider for instance returns to schooling. In the model of card2001estimating, schooling $X_t$ depends on individual marginal cost $c_t$ and individual returns $r_t$ through the relationship $X_t=(c_t - r_t)/k$ for some constant $k>0$. Suppose that returns $r_t$ are time invariant but exogenous variations in tuition fees affect marginal costs multiplicatively, so that $c_t=\alpha_t \widetilde{c}_t$ and $(\widetilde{c}_t, r_t)$ is time invariant. The results of Kasy11, and in particular his Section 2, then imply that Assumption (ref) would fail in this example.

Extensions

Including Covariates

We consider here the case where exogenous covariates $Z_{t}$ also affect the outcome variable. Specifically, let $Y_{t}(x,z)$ denote the potential outcome associated with the values $x$ and $z$ (of random variables $X_{t}$ and $Z_{t}$, respectively). We observe $Y_{t}\equiv Y_{t}(X_{t},Z_{t})$. We still focus on the effect of $X_{t}$ hereafter. In this case, the preceding analysis can be conducted conditionally on $Z_{t}$. We briefly discuss this extension here, by considering only the discrete average and quantile effects

eqnarray*[eqnarray* omitted — 270 chars of source]

The marginal effects can be handled similarly. We first restate our previous conditions in this context. The rank variable is now defined conditionally on $Z_{t}$, i.e., $V_{t}=F_{X_{t}|Z_{t}}(X_{t}|Z_{t})$.

\begin{assumption_nonumber} {\normalfont 1C}\; $\text{Supp} ((V_t,Z_t))$ does not depend on $t$. For all $t \in \{1,...,T\}$, $ Y_t(x,z)=g_t(z,U_t(x,z))$ where for all $(x,v,z)\in \mathcal{X}\times \text{ Supp}((V_t,Z_t))$, the distribution of $U_t(x,z)|V_t=v,Z_t=z$ does not depend on $t$. \end{assumption_nonumber}

\begin{assumption_nonumber} {\normalfont 4C}\; For all $ (t,z)\in\{1,...,T\}\times \text{Supp}(Z_t)$, $E(|Y_t|)<\infty$ and $ F_{X_t|Z_t}(.|z)$ is continuous and strictly increasing on $\text{Supp} (X_t|Z_t=z)$. For all $(x^{\prime },z)\in \text{Supp}((X_t,Z_t))$ and $u\in \text{Supp}(U_t(x^{\prime },z))$, there exist versions of $E[Y_t|X_t,Z_t]$ and $P^{U_t(x^{\prime }) |V_t,Z_t}$ such that $x\mapsto E[Y_t|X_t=x,Z_t=z]$ and $v \mapsto F_{U_t(x^{\prime })|V_t,Z_t}(u|v,z)$ are continuous. \end{assumption_nonumber}

Next, we consider two versions of Assumptions (ref) and (ref), namely Assumptions 2C-3C and 2C'-3C' below. The trade-off between these two versions is basically between the generality of the model and the requirements on the data. In the first version, we allow for a more general time trend (i.e., Assumption 2C' is a particular case of Assumption 2C). However, the crossing condition in Assumption 3C is more demanding than in Assumption 3C', because the former requires to observe a crossing point for each value of $z$.

\begin{assumption_nonumber} {\normalfont 2C} \; For all $(z,t)\in \text{Supp}(Z_T) \times \{1,...,T\}$, $U_t(x,z)\in \mathbb{R}$ and $g_t(z,.)$ is strictly increasing. Without loss of generality, we let $g_T(z,y)=y$ for all $(y,z)\in \text{Supp}((Y_T, Z_T))$. \end{assumption_nonumber}

\begin{assumption_nonumber} {\normalfont 3C} \; For all $(z,t)\in \text{Supp}(Z_T) \times \{1,...,T-1\}$, there exists $x^*_t(z)$ such that $F_{X_T|Z_T}(x^*_t(z)|z) = F_{X_t|Z_t}(x^*_t(z)|z) \in (0,1)$. \end{assumption_nonumber}

\begin{assumption_nonumber} {\normalfont 2C'} \; For all $(t,x,z) \in \{1,...,T\}\times \text{ Supp}((X_t,Z_t))$, $g_t(z,U_t(x,z))=h_t(U_t(x,z))$, with $U_t(x,z)\in \mathbb{R}$ and $h_t(.)$ strictly increasing. Without loss of generality, we let $h_T(y)=y$ for all $y\in\text{Supp}(Y_T)$. \end{assumption_nonumber}

\begin{assumption_nonumber} {\normalfont 3C'} \; For all $t\in \{1,...,T-1\}$, there exists $ (x^*_t,z^*_t)$ such that $F_{X_T|Z_T}(x^*_t|z^*_t) = F_{X_t|Z_t}(x^*_t|z^*_t) \in (0,1)$. \end{assumption_nonumber}

Both sets of the assumptions lead to the same results, which are qualitatively very similar to those of Theorem (ref). In what follows, we let $ q_{t}(x|z)=F_{X_{t}|Z_{t}}^{-1}(F_{X_{T}|Z_{T}}(x|z)|z) $. The proof of Theorem (ref)C is a straightforward extension of the proof of Theorem (ref), and hence is omitted.

\setcounter{theorem}{0}

theorem{\normalfont C} \, Suppose that Assumptions 1C and 4C and either Assumptions 2C-3C or Assumptions 2C'-3C' hold. Then, for almost all $(x,z)\in \text{Supp}((X_T,Z_T))$, all $p \in (0,1)$ and all $ t\in \{1,...,T-1\}$, the functions $g_{t}$ and the average and quantile treatment effects $\Delta ^{ATT}(x,q_t(x|z),z)$ and $\Delta ^{QTT}(p ,x,q_t(x|z),z)$ are identified.

Here again, we can relate $\Delta ^{ATT}(x,q_t(x|z),z)$ with average marginal effects. If $Y_T(.,z)$ is differentiable, by the mean value theorem,

equation*[equation* omitted — 99 chars of source]

for some $\widetilde{X}_z\in (x,q_t(x|z))$. Then,

equation*[equation* omitted — 187 chars of source]

This equation implies that we can also average over $x$ and $z$ to gain statistical power. Specifically, let $I_c=\{(x,z)\in \text{Supp}((X_T,Z_T)): |x-q_t(x|z)|>c\}$ for some $c>0$. Under the conditions behind Theorem (ref), we can identify

equation*[equation* omitted — 181 chars of source]

Multivariate Treatment

Our framework directly extends to multivariate treatments, $ X_{t}=(X_{1t},...,X_{kT})\in \mathbb{R}^{k}$, $k\geq 2$, by just making a few changes. First, we now define $V_t$ to be $V_t = (F_{X_{1t}}(X_{1t}),...,F_{X_{kt}}(X_{kT}))$. Second, we replace Assumption (ref) by the following condition:

\begin{assumption_nonumber} {\normalfont 3M}\; For all $(j,t)\in \{1,...,k\}\times \{1,...,T\}$ , there exists $x_{jt}^{\ast }\in \mathbb{R}$ such that $F_{X_{jt}}(x_{jt}^{ \ast })=F_{X_{jT}}(x_{jt}^{\ast })\in (0,1)$. \end{assumption_nonumber}

Finally, we now define $q_{t}$ as $q_{t}(x_{1},...,x_{k})=\left( F_{X_{1t}}^{-1}\circ F_{X_{1T}}(x_{1}),...,F_{X_{kt}}^{-1}\circ F_{X_{kT}}(x_{k})\right)$. Then, we obtain the same point identification result as before.

\setcounter{theorem}{0}

theorem{\normalfont M} \, Suppose Assumptions (ref), (ref), 3M and (ref) hold. Then, for all $(t,x)\in \{1,...,T-1\}\times \text{Supp}(X_T)$, the function $ g_t$ and $\Delta^{ATT}(x,q_t(x))$ and $\Delta^{QTT}(p,x,q_t(x))$ are identified.

\setcounter{theorem}{2}

Partial Identification of Other Treatment Effects

Theorem (ref) implies that we can point identify some but not all average treatment effects $\Delta ^{ATT}(x,x^{\prime })$. Similarly, we point identify the average marginal effects only at some particular points. We show in this subsection that with three or more periods of observation, we can get bounds for many other points under a weak local curvature condition. Let us consider average marginal effects, for instance. The idea is that if $x\mapsto U_{T}(x)$ is locally concave (say) and $ q_{t}(x)<x$, then $[U_{T}(q_{t}(x))-U_{T}(x)]/[q_{t}(x)-x]$ is an upper bound for $dU_{T}/dx(x)=dY_{T}/dx(x)$. By integration, $\Delta^{ATT}(x,q_{t}(x))/(q_{t}(x)-x)$ is therefore an upper bound for $\Delta ^{AME}(x)$. Similarly, we obtain a lower bound for $ \Delta ^{AME}(x)$ if $q_t(x)>x$. Figure illustrates this idea with $T=3$ and $q_2(x)<x<q_1(x)$. Note the same idea can be used to obtain bounds $\Delta ^{ATT}(x,x^{\prime })$ for $x^{\prime }\not\in \{q_{t}(x),t=2,...,T\}$.

figure[figure omitted — 176 chars of source]

The above argument works even if we do not know a priori whether $U_{T}(.)$ is concave or convex. Using the minimum and the maximum of the local discrete treatment effect will be sufficient to obtain bounds, provided that $U_{T}(.)$ is locally concave or locally convex around $x$. We therefore adopt the following definition.

definition$x\mapsto U_T(x)$ is locally concave or convex on $[\widetilde{x},\widetilde{ x}^{\prime }]$ if, almost surely (a.s.), it is twice differentiable and \begin{equation*} \frac{\partial ^2 U_T}{\partial x^2}(x) \leq 0 \; \forall x\in [\widetilde{x} ,\widetilde{x}^{\prime }] \; a.s. or \; \frac{\partial ^2 U_T}{ \partial x^2}(x) \geq 0 \; \forall x\in [\widetilde{x},\widetilde{x}^{\prime }] \; a.s. \end{equation*}

Let us introduce, for all $(x,x^{\prime })\in \text{Supp}(X_{T})$, $( \underline{x}_{T}(x^{\prime }),\overline{x}_{T}(x^{\prime }))$ defined by

eqnarray*[eqnarray* omitted — 260 chars of source]

If the sets are empty, we let $\underline{x}_{T}(x^{\prime })=-\infty $ and $ \overline{x}_{T}(x^{\prime })=+\infty $.

theoremSuppose that Assumptions (ref)-(ref) aresatisfied. For any $x<x^{\prime }$, if $U_{T}$ is locally concave or convex on $[\min (x,\underline{x}_{T}(x^{\prime })),\overline{x} _{T}(x^{\prime })]$, then \begin{align*} & (x^{\prime }-x)\min \left\{ \frac{\Delta ^{ATT}(x,x _{T}(x^{\prime }))}{x_{T}(x^{\prime })-x},\frac{\Delta ^{ATT}(x, \overline{x}_{T}(x^{\prime }))}{\overline{x}_{T}(x^{\prime })-x}\right\} \leq \Delta ^{ATT}(x,x^{\prime }) \\ \leq & (x^{\prime }-x)\max \left\{ \frac{\Delta ^{ATT}(x,x _{T}(x^{\prime }))}{x_{T}(x^{\prime })-x},\frac{\Delta ^{ATT}(x, \overline{x}_{T}(x^{\prime }))}{\overline{x}_{T}(x^{\prime })-x}\right\} . \end{align*} If $U_{T}$ is locally concave or convex on $[\underline{x}_{T}(x),\overline{x }_{T}(x)]$, then \begin{align*} & \min \left\{ \frac{\Delta ^{ATT}(x,x_{T}(x))}{x _{T}(x)-x},\frac{\Delta ^{ATT}(x,\overline{x}_{T}(x))}{\overline{x}_{T}(x)-x} \right\} \leq \Delta ^{AME}(x) \\ \leq & \max \left\{ \frac{\Delta ^{ATT}(x,\underline{x}_{T}(x))}{\underline{x }_{T}(x)-x},\frac{\Delta ^{ATT}(x,\overline{x}_{T}(x))}{\overline{x}_{T}(x)-x }\right\} . \end{align*} The bounds are understood to be infinite when either $\underline{x} _T(x^{\prime })=-\infty$ or $\overline{x}_T(x^{\prime })=+\infty$ (whether $ x^{\prime }>x$ or $x^{\prime }=x$).

Both bounds are finite, provided that there exists $t,t^{\prime }$ such that $q_{t}(x)<x<q_{t^{\prime }}(x)$, which implies that $T\geq 3$. More generally, the bounds improve with $T$, because $(\underline{x} _{T}(x^{\prime }))_{T\in \mathbb{N}}$ and $(\overline{x}_{T}(x^{\prime }))_{T\in \mathbb{N}}$ are by construction increasing and decreasing, respectively. Also, the local curvature condition becomes less restrictive as $T$ increases, because the interval on which $U_{T}$ has to satisfy this condition decreases. This condition is particularly credible if $ q_{t}(x)\mapsto \Delta (x,q_{t}(x))/(q_{t}(x)-x)$ is monotonic, because such a pattern is implied by global concavity or global convexity.

Two other remarks on Theorem (ref) are in order. First, we do not establish that the bounds are sharp, though we conjecture that they are. Second, similar to the point identification results of Theorem (ref), the partial identification results of Theorem (ref) can be extended to the multivariate setting. Specifically, we can use the system of inequalities

equation*[equation* omitted — 105 chars of source]

which hold for all $t=1...T-1$ if $U_{T}$ is locally convex (inequalities are reverted if $U_{T}$ is locally concave). These inequalities imply some bounds on $E[\partial U_{T}(x)/\partial x]$. A necessary condition for the bounds to be finite on each component of $E[\partial U_{T}(x)/\partial x]$ is that $T-1\geq 2\dim (X_{t})$. This condition generalizes the above restriction $T\geq 3$. It makes intuitive sense that more time periods are required when the endogenous treatment is multivariate.

To illustrate Theorem (ref), we consider the following example:

eqnarray*[eqnarray* omitted — 115 chars of source]

where $V_{t}\sim U[0,1]$ and $U_{t}|V_{t}\sim \mathcal{N}(V_{t},1)$. We also suppose that

eqnarray*[eqnarray* omitted — 214 chars of source]

In this example, Assumptions (ref), (ref) (with $g_{t}(y)=1-\exp (-0.5\delta _{t})(1-y)$) and (ref) are satisfied, the latter because $\sigma _{t}\neq \sigma _{T}$ almost surely. The local curvature condition also holds, since $u\mapsto 1-\exp (-0.5u)$ is concave. Figure (ref) displays the bounds on $ \Delta_{1}^{AME}(x)$ for $T=3,4,5$ and 6. Note that the bounds coincide for $ T-1$ points. This simply reflects the point identification result of Theorem (ref). We also see that in the interval where we get finite bounds, i.e., the interval for which $-\infty <\underline{x}_{T}(x)<\overline{x}_{T}(x)<\infty$ , the bounds are quite informative even for $T=3$. Figure (ref) also shows that as $T$ increases, both the bounds shrink and the interval on which we get finite bounds increase. For $T=6$, we get informative bounds for $x\in \lbrack 1,\;3.85]$, which corresponds roughly to 85% of the population. This means that we could also obtain finite bounds for the average partial effect for this large fraction of the total population.

figure[figure omitted — 695 chars of source]

Point Identification with a Correlated Random Coefficient Model

As we have established in Theorem (ref), we can point identify several treatment effect parameters under Assumptions (ref)-(ref), but these are by no means all possible causal effects one may be interested in. Many more treatment parameters can be set identified under often plausible curvature restrictions, in particular average marginal effects and effects of the kind $\Delta ^{ATT}(x,x^{\prime })$. However, these bounds may be wide in some applications, conducting inference on the corresponding parameters may be cumbersome or even impractical. Hence it makes sense to search for additional assumptions that yield point identification of average structural effects over the entire population.

We suggest here a possible route for extrapolation, based on a random coefficient linear model of the form:

equation[equation omitted — 82 chars of source]

Therefore, we impose a linear structure on $g_t$ ($g_t(u)=\delta_t+u$ and $ U_t(x)$ ($U_t(x)=U_{0t}+x U_{1t}$). The model still allows for a rich, non-scalar heterogeneity pattern through the two unobserved terms $U_{0t}$ and $U_{1t}$. Under this structure, we have, for any $(x,x^{\prime })\in \text{Supp}(X_T)^2$, $x\neq x^{\prime }$,

equation[equation omitted — 185 chars of source]

By Theorem (ref), $\Delta ^{ATT}(x,q_t(x))$ is point identified under Assumptions (ref)-(ref). This implies that $\Delta^{AME}(x)$ and $\Delta ^{ATT}(x,x^{\prime })$ are identified as well, whenever $q_t(x)\neq x$. As a result, the average marginal effect over the whole population, $\Delta^{AME}=E\left[ \Delta^{AME}(X_T)\right]$, is also point identified if $q_t(X_T)\neq X_T$ almost surely. We summarize this finding in the following theorem.

theoremUnder Assumptions (ref)-(ref) and Equation (ref), for all $t<T$ and $(x,x^{\prime })\in \text{ Supp}(X_T)^2$, $q_t(x)\neq x$, $(\delta_t)_{t<T}$, $\Delta^{ATT}(x,x^{\prime })$ and $\Delta^{AME}(x)$ are identified. If $q_t(X_T)\neq X_T$ almost surely, $\Delta^{AME}$ is point identified as well.

Several remarks on this result are in order. First, we recover the same parameter as Graham12, who also consider a random coefficient linear model similar to (ref). They obtain identification with panel data, relying on first-differencing. Compared to them, we rely on variations in the cdf of $X_{t}$ rather than on individual variations. We rely on a different, non-nested, restriction on the distribution of the error term. In particular, for the same individual, $U_{1t}-U_{1s}$ could be correlated with $X_{t}$ in our framework.

Second, Theorem (ref) readily extends to a multivariate treatment, by just replacing the condition $q_t(x)\neq x$ by a rank condition. Specifically, let, as in Section (ref), $q_t(x)= (q_{1t}(x_1),...,q_{kt}(x_k))^{\prime }$ and define the matrix $\mathbf{Q} (x) $ by

equation*[equation* omitted — 138 chars of source]

Then $\Delta^{AME}(x)$ and $\Delta^{ATT}(x,x^{\prime })$ are identified if $ \mathbf{Q}(x)$ is full column rank. Note that the rank condition implies that $T-1 \geq k$. It also implies that the distribution of $X_t$ differs at each date, so that $q_s(x) \neq q_t(x)$. It makes sense that with several endogenous variables, more time variation on $X_t$ is needed to identify causal effects.

Third, coming back to the univariate case, Theorem (ref) ensures that all parameters of interest are identified with only two time periods. This suggests that the model can be either tested or enriched when $T>2$. To see why the linearity assumption is testable when $T>2$, note that Equation (ref) implies

equation*[equation* omitted — 129 chars of source]

which can be checked in the data. With more than two time periods, we can also identify treatment effects in the more general random coefficient polynomial model of order $T-1$:

equation[equation omitted — 106 chars of source]

With the same arguments as above, we recover not only average marginal effect, but actually $E(U_{kt}|X_{t}=x)$ for all $k=1,...,T$ and all $x$ such that $(x,q_1(x),...,q_{T-1}(x))$ are all distinct. Identification of a model similar to (ref) was studied before by Florens08, with cross-sectional data and under assumptions that typically rule out discrete instruments Heckman98. Here, we rely only on a finite number of time periods, which would be equivalent to a discrete instrument, and allow for time trend, which would correspond to a direct effect of the instrument in Florens08.

Alternatively, we can use additional periods to identify higher moments of the distribution of the coefficients in the linear model (ref). For instance, with $k=1$, $V(U_{01}|X_{T}=x)$ , $V(U_{1T}|X_{T}=x)$ and Cov$(U_{01},U_{1T}|X_{T}=x)$ can be shown to be identified with $T=3$ as soon as $x,q_1(x)$ and $q_2(x)$ are distinct.

Estimation of Average and Quantile Treatment Effects

We consider in this section estimators of the parameters $\Delta^{ATT}(x, q_t(x))$ and $\Delta^{QTT}(p, x, q_t(x))$ that are shown to be identified in Theorem (ref). We suppose for that purpose to observe two independent samples corresponding to the periods $1$ and $T=2$. For simplicity, we suppose hereafter that the two corresponding sample sizes are identical.

assumptionWe observe the two independent samples $(Y_{i1}, X_{i1})_{i=1...n}$ and $(Y_{i2}, X_{i2})_{i=1...n}$, which are both i.i.d. random variables drawn from the distributions $F_{Y_1,X_1}$ and $F_{Y_2,X_2}$ , respectively.

Our estimator follows closely our identification strategy. Let us define

equation*[equation* omitted — 73 chars of source]

where $\widehat{F}_{X_2}$ (resp. $\widehat{F}_{X_1}$ ) denotes the empirical cdf of $X_2$ (resp. $X_1$). We first estimate $x^*_1$ by

equation[equation omitted — 330 chars of source]

where $\widehat{F}_{X_1}^{-1}$ denotes the empirical quantile function and $ 0<\underline{p}<\overline{p}<1$ are two given constants used to avoid reaching the boundaries of the support of $X_1$. Note that the minimum in (ref) is well defined because $\Psi_n$ is a right-continuous step function.

Next, we estimate $q_{1}(x)=F_{X_{1}}^{-1}\circ F_{X_{2}}(x)$ by its empirical counterpart $\widehat{q}_{1}(x)=\widehat{F}_{X_{1}}^{-1}\circ \widehat{F}_{X_{2}}(x)$. We then estimate $g_{1}$ using an empirical counterpart of (ref). For that purpose, we estimate the conditional cdf $F_{Y_{t}|X_{t}}$, for $t\in \{1,2\}$, by

equation*[equation* omitted — 189 chars of source]

where $K$ is a kernel function and $h_{n}$ denotes the bandwidth. We then let $\widehat{F}_{Y_{t}|X_{t}}^{-1}(.|x)$ denote the generalized inverse of $ \widehat{F}_{Y_{t}|X_{t}}(.|x)$. We estimate $g_{1}$ by

equation*[equation* omitted — 159 chars of source]

Now, let us recall that $\Delta ^{ATT}(x,q_{1}(x))$ and $\Delta ^{QTT}(p,x,q_{1}(x))$ satisfy, under Assumptions (ref)-(ref),

align*[align* omitted — 190 chars of source]

We then estimate these two parameters by

align*[align* omitted — 485 chars of source]

For notational simplicity, we chose here the same kernels and bandwidths for each nonparametric terms, though we could obviously consider different ones. We establish below that $\widehat{\Delta }^{ATT}(x,q_{1}(x))$ and $\widehat{ \Delta }^{QTT}(p,x,q_{1}(x))$ are consistent and asymptotically normal. Our result is based on the following conditions.

assumption(Conditions for the root-n consistency of $ \widehat{x}^\ast_1$ and $\widehat{q}_1(x)$) \newline (i) There exists a unique $x^*_1$ satisfying $F_{X_1}(x^*_1)=F_{X_2}(x^*_1) \in (0,1)$. Moreover, $F_{X_1}(x^*_1) \in (\underline{p},\overline{p})$. \newline (ii) For $t \in \{1,2\}$, $X_t$ admits a continuous density $f_{X_t}$ satisfying, for all $x$ in the interior of $\mathcal{X}$, $f_{X_t}(x)>0$. Moreover, $f_{X_1}(x^*_1) \neq f_{X_2}(x^*_1)$.
assumption(Regularity conditions on $(X_t,Y_t)$) \newline (i) For $t \in \{1,2\}$, $\text{Supp}(X_t,Y_t)= \mathcal{X} \times \mathcal{Y }$ with $\mathcal{Y}=[\underline{y},\overline{y}]$ with $-\infty<\underline{y }<\overline{y}<+\infty$.\newline (ii) For $(t,x) \in \{1,2\}\times \mathcal{X}$, $F_{Y_t|X_t}(.|.)$ is continuously differentiable and $\inf_{y \in \mathcal{Y}} f_{Y_t | X_t}(y | x) > 0$. \newline (iii) For all $(t,y) \in \{1,2\}\times \mathcal{Y}$, $F_{Y_t|X_t}(y|.)$ and $ f_{X_t}$ are twice differentiable. $f_{X_t}$, $|f^{\prime }_{X_t}|$ and $ |f^{\prime \prime }_{X_t}|$ are bounded. $\sup_{(y,x)\in \mathcal{Y} \times \mathcal{X}}|\partial_x F_{Y_t|X_t}(y|x)|<\infty$ and $\sup_{(y,x)\in \mathcal{Y} \times \mathcal{X}}|\partial_{xx} F_{Y_t|X_t}(y|x)|<\infty$.
assumption(Conditions on the kernels and bandwidths) \newline (i) $nh^3_n/|\log(h_n)| \rightarrow +\infty$, $nh_n^5 \rightarrow 0$. \newline (ii) $K$ has a compact support, is differentiable with $K^{\prime }$ of bounded variation and satisfies $K(y)\geq 0$ for all $y$. Besides, $\int K(y)dy=1$ and $\int y K(y)dy=0$.

Assumption (ref)-(i) strengthens Assumption (ref) by assuming the uniqueness of the crossing point. We make this assumption for the sake of simplicity. We could also consider the case where the set of crossing points is an interval. As discussed in Section (ref) above, we would actually expect a parametric rather than a nonparametric rate of convergence for $\widehat{g}_1$, so Theorem (ref) below should still hold in this more favorable case. Assumption (ref)-(ii) is a mild regularity condition on $F_{X_{2}}$ and $F_{X_{1}}$. As Lemmas (ref) and (ref) in Appendix A show, these two restrictions ensure that $\widehat{x}_{1}^{\ast }$ and $\widehat{q}_{1}(x)$ are root-n consistent. Assumption (ref) provides a set of conditions ensuring that $\widehat{g}_{1}$ is consistent and asymptotically normal. Conditions (i) and (ii) are also made by Athey06, without any $X_{t}$ in their case, in another context where quantile-quantile transforms must be estimated. Condition (iii) is required as well here because we deal with nonparametric estimators of conditional cdfs rather than usual empirical cdfs, as Athey06 do. Finally, Assumption (ref) is a standard condition on the bandwidths and the kernels appearing in the nonparametric estimators. We impose $nh_{n}^{5}\rightarrow 0$ in order to avoid any asymptotic bias on $ \widehat{\Delta }^{ATT}(x,q_{1}(x))$ and $\widehat{\Delta } ^{QTT}(p,x,q_{1}(x))$.

theoremSuppose that Assumptions (ref)-(ref) and (ref)-(ref) are satisfied. Then, for any $x\in \mathcal{X}$ such that $F_{X_1}$ is differentiable at $q_1(x)$ with $F_{X_1}^{\prime }(q_1(x))>0$, \begin{align*} \sqrt{n h_n}\left(\widehat{\Delta}^{ATT}(x,q_1(x)) - \Delta^{ATT}(x,q_1(x))\right) & \overset{d}{\longrightarrow} \mathcal{N} (0,V_1) \\ \sqrt{n h_n}\left(\widehat{\Delta}^{QTT}(p, x,q_1(x)) - \Delta^{QTT}(p, x,q_1(x))\right) & \overset{d}{\longrightarrow} \mathcal{N}(0,V_2), \end{align*} for some $V_1,V_2$.

We do not display the asymptotic variances here, as they involve many terms due to the multiple compositions of nonparametric estimators -- see Appendix (ref) for details as well as a proof. In practice, we suggest to rely on bootstrap, as we do in the application below. We conjecture that the bootstrap is consistent in our setting, though a formal proof of its validity is beyond the scope of this paper. The main issue for establishing its validity would be to prove the (conditional) weak convergence of the process

equation*[equation* omitted — 128 chars of source]

where $\widehat{F}^*_{Y_t|X_t}$ is the bootstrap counterpart of $\widehat{F} _{Y_t|X_t}$. Up to our knowledge, such a result is not available in the literature yet.

Application to the Marginal Propensity to Consume

In this section we provide an application to a substantive economic question: The magnitude of the marginal propensity to consume out of current disposable income. When analyzing this question, we focus in particular on how results obtained through our approach compare to those obtained in the literature. In order to facilitate this comparison, we first briefly review the literature on this question, before explaining the policy experiment we are using, and detailing the data. We then outline how our methodology is employed, and finally close by comparing our results with those in the literature.

The Economic Question

A crucial question for the classical theory of consumption is the marginal propensity to consume (MPC) out of income. Given its implications for the business cycle, taxes, and government policy, the importance of the MPC can hardly be overstated, and thus this quantity was, and still is, at the center of a very active debate jappelli2010consumption. An upshot of the rational expectations revolution which, since the seminal paper of hall1978stochastic, tried to answer questions about the effect of a marginal change in income on consumption, is that expectations about the change matter.

In the absence of liquidity constraints (and precautionary saving motives at very low income levels), the following is the key insight in the literature about the effect of a marginal income change on the nondurable consumption of a rational consumer, see, e.g., deaton1992understanding : If the income change is anticipated, i.e., not related to new information, then consumption does not respond to the income change. For an income change that is not anticipated, if the change is viewed as transitory, then the rational consumer is predicted to use very little of the income increase immediately, as the transitory change in income is distributed over the life-cycle, and its small quantity (relative to life-cycle income) does not alter fundamentally the trade-off between consumption today and saving for the future. Conversely, if the income change is expected to be permanent, the individual is expected to essentially increase her consumption by the amount of the change. This means that we only observe a substantial change in consumption in response to an income change, if the change is surprising and considered to be permanent.

The empirical evidence on the hypothesis of a rational consumer is rather mixed, and has spurned an active debate. Perhaps the most problematic evidence comes from studies involving one time transfers, see e.g., johnson2006household and parker2013consumer (2013, PSJM). In these studies, consumers are given what is clearly an expected and transitory income shock (PSJM actually documenting aspects of the Obama era stimulus package), yet the effect on consumption is not zero. Instead, typical estimates for the marginal effects of an anticipated income change range between 15% and 25%.

There are a number of counterarguments in defense of the rational consumer. First, consumers could be credit constrained. PSJM find indeed lower responses for older and high-income households, who are less likely to be constrained. Second, consumers may exhibit a form of bounded rationality. There are significant costs associated with computing the optimal consumption path. If an income change is small relative to the level of income, the benefits from adapting the optimal path in light of the changes are small relatively to the costs associated with it, and individuals simply avoid optimizing completely, as they would in the case of a large income change. Evidence that individuals indeed smooth large anticipated income changes is provided by browning2001response and hsieh2003consumers, among others. Another counterargument is that some of the changes considered in the literature are not just small, but also outside the \textquotedblleft usual\textquotedblright\ consumer experience. As such, they are not representative of the typical real-world surprise income shocks individuals deal with (a distinction that is reminiscent to the question of whether individuals are able to assign probabilities to these events).

In this section, we use our econometric method in conjunction with an experiment involving the Earned Income Tax Credit (EITC) to analyze the causal effect of increase in income on consumption for households in 1987. We believe that this natural experiment is very insightful for the above debate. While it provides exactly the type of variation we require for our method, it provides (at least for a good number of households) a significant and anticipated change in their income. Finally, the fact that our procedure allows for nonlinearities, i.e., for the marginal effect to vary with income, is going to be crucial to shed light on the question of the existence of liquidity constraints.

Policy Background: The EITC

In the following, we provide more background on the policy experiment that provides the exogenous variation: The Earned Income Tax Credit (EITC) is an income support program which started in 1975 in the United States for the purpose of mitigating poverty. The EITC provision schedule varies from year to year, exhibiting interesting non-linearities. This feature of the program has been used for economic analysis before, e.g., by dahl2012impact, and a detailed documentation of the EITC can be found in falk2014earned. In most of the past years, the change to the EITC schedule has been monotone to match increasing price levels. However, the change in the schedules between 1987 and 1989 exhibits a specific pattern which, as we will now demonstrate, generates a crossing of the cdfs of (deflated) total income in the respective years.

Figure (ref) displays the EITC schedules in 1987 (solid line) and 1989 (dotted line) in terms of thousands of Year 2000 US dollars for families with two or more children. Note that for individuals with income between 9K USD and 10.75K USD, the 1987 EITC provision was higher than the provision in 1989, whereas the reverse is true for individuals with income above 10.75K USD. This is exactly the type of variation which generates a crossing, if everything else is held constant.

figure[figure omitted — 341 chars of source]

To see this more precisely, consider the left graph in Figure (ref). The graph shows total income, obtained as the sum of the pre-aid income and the EITC amount, for each of the years 1987 (solid line) and 1989 (dotted line) plotted against that of year 1989, i.e., the solid line is the 45-degree line. The right graph in the same figure (Fig. (ref)) focuses on this difference. As these figures suggest, we expect a crossing at 12K USD, computed as the sum of 10.75K USD (the cut-off for the change in the schedule) and 1.25K USD for the corresponding EITC amount, provided that total pre-EITC income does not change substantially. Note that these figures are solely derived from the known policy schedules, but we will confirm our expectation with real data below. Before we detail this, however, we first give an overview of the data.

figure[figure omitted — 716 chars of source]

Data: The CEX

For our analysis, we use repeated cross-sectional data from the Consumer Expenditure Survey (CEX) for the calendar years 1987 and 1989. The treatment variable is, more precisely, total disposable family income measured in thousands of Year 2000 US dollars. The outcome (dependent) variable is non-durable household consumption, defined as the sum of expenditures for food at home, apparel, health, entertainment, personal care, and readings, measured in thousands of Year 2000 US dollars. Since the policy described above applies only to families with two or more children, we use the sub-sample of individuals with two or more children. Table (ref) shows summary statistics for our sub-sample.

Note that after controlling for inflation (i.e., in year 2000 prices), the mean of total family disposable income does not change substantially between 1987 and 1989 (roughly 2%). Indeed, this modest increase from 1987 to 1989 is quite consistent with the EITC policy change and an otherwise pretty stationary environment, strengthening the case that we should expect to have the type of variation in cdfs our method requires \footnote{ As a caveat, we remark that not all families take up the aid even if eligible, and that only a part of the population of families is eligible, which together accounts for the modest 2% increase in mean total family income from 1987 to 1989.}.

table[table omitted — 1,343 chars of source]

Turning to our nondurable consumption measure, we first notice that it only captures a little less than half of disposable income. Within the subsample that we focus on, the average ratio of nondurable consumption to total disposable income (which is different than the ratio of averages) is around 50%. This may be due to the fact that the large category of rent and mortgage payments are excluded as are large and durable and nondurable consumption items (e.g., TVs, cars, phones). However, we also suspect a certain modest degree of underreporting in the data. Like in the standard Diff-in-Diff approach, our analysis would be invalidated if the evolution of this underreporting is systematically different between treatment and control group. We believe this to be unlikely and certainly have no evidence of this difference in effects. Moreover, since the overall degree of underreporting seems to be tolerable as well (e.g., food and clothing account for a budget share of 50% in the British FES as well, see Hoderlein (2011)), we hence proceed with our analysis.

One thing that stands out is that the nondurable consumption measure increased more than proportionally to the change in disposable income in both the sample we focus on (average share increase from 49.1% to 53.6%), but also in the population at large. This may be due to changes in economic outlook and general optimism in 1989 at the end of the cold war. Because of this observation, we definitely want to include a time trend $ g_{t}$ in the empirical analysis, as our method warrants. Indeed, our model identifies an increase in nondurable consumption in particular at higher levels of the consumption distribution even if our policy experiment would not have taken place.

figure[figure omitted — 591 chars of source]

Analysis and Results

First, we use our data to confirm that the policy change in the EITC described above indeed induces a crossing in the cdfs. In particular, we want to study whether there is a divergence from 12.0 K to 26.8 K USD of cdfs of total family income between 1987 and 1989. The left panel of Figure (ref) displays the two empirical cdfs. The solid vertical lines indicate the limit points of the range inside which the policy change matters; these lines correspond to those displayed in Figure (ref). To check that the distributions of income are in line with this policy change, we made one-sided test of $F_{1}(x)\leq F_{2}(x)$ for all $x\in \lbrack 12.0K,26.8K]$. We find that at the 5% level, $F_{1}(x)>F_{2}(x)$ for at least some $x\in \lbrack 12.0K,26.8K]$. We take this as strong evidence that the change in the EITC was, at least for this subpopulation of households, the main driving force in the change of the empirical cdfs between the two years. Moreover, the direction of the crossing is what we expect from the design of the policy change: the families falling within the range where we expect an increase in total disposable income due to the change in EITC experience a positive change in total family income between 1987 to 1989.

Using these two empirical cdfs, we next compute the empirical quantile-quantile plot of the total family income from 1987 to 1989 in terms of Year 2000 US dollars. The right panel of Figure (ref) displays the plot. Observe how well this data-based figure resembles Figure (ref), which is constructed using the policy formulas. This provides further evidence that the data follows our research design, and that there are no other major unaccounted sources of change in disposable income. Recall, moreover, that this quantile-quantile plot, which is mathematically represented by $q_{1}$ in our framework, is the main building block for our identification results.

figure[figure omitted — 983 chars of source]

After having confirmed that the change in the distribution of the treatment is in line with our modeling assumption and largely driven by the policy change, we proceed to use our framework and estimate the time trend $ g_{1}(.)$. In line with the theoretical design, Figure (ref) shows that we have more than a single point $x^{\ast }$ as a control group. We can use the whole set $\mathcal{S}=[10K;12K]\cup \lbrack 26,8K;50K]$, where the two cdfs overlay. This results in more precise estimates of $g_{1}(.)$ and marginal effects, because we can use the whole set $\mathcal{S}$ instead of a single point $x^{\ast }$. Specifically, we can use $g_{1}(y)=F_{Y_{1}|X_{1}\in \mathcal{S}}^{-1}\left[ F_{Y_{2}|X_{2}\in \mathcal{S}}(y)\right] $ instead of $ g_{1}(y)=F_{Y_{1}|X_{1}}^{-1}\left[ F_{Y_{2}|X_{2}}(y|x^{\ast })|x^{\ast } \right] $. Figure (ref) displays the estimate of $g_{1}^{-1}$, which corresponds to the (heterogeneous) time trend between 1987 and 1989. As mentioned before, we observe an increase in the upper tail of the distribution of nondurable consumption, corresponding with an improved overall economic outlook, in particular for middle and upper class households.

figure[figure omitted — 417 chars of source]

To come to the main purpose of this application, we estimate average marginal effects of the total family income in 1987 in terms of Year 2000 US dollars on various expenditures in terms of Year 2000 US dollars. Specifically, we estimate $\Delta _{app}^{AME}(x)=\Delta ^{ATT}(x,q_{1}(x))/(q_{1}(x)-x)$ instead of $\Delta ^{ATT}(x,q_{1}(x))$. The former quantity has the advantage over the latter of being interpretable as an average marginal effect. By the mean value theorem (and under mild regularity conditions), indeed, $\Delta _{app}^{AME}(x)=E\left[ dY_{2}/dx( \widetilde{X})|X_{2}=x\right] $, for some random $\widetilde{X}\in \lbrack q_{1}(x),x]$. To the extent that $q_{1}(x)$ is close to $x$, we then interpret $\Delta _{app}^{AME}(x)$ as the average marginal effect at $ X_{2}=x $. Note, on the other hand, that by dividing by $\widehat{q} _{1}(x)-x $, the estimator of $\Delta _{app}^{AME}(x)$ is more volatile than that of $\Delta ^{ATT}(x,q_{1}(x))$, especially when $q_{1}(x)-x$ is close to zero. To obtain more precise estimates, we rely hereafter on a piecewise linear estimator of $q_{1}(x)-x$. Such a constrained estimator is consistent with the policy design and fits well the data. We refer to Appendix (ref) for more details on its construction.

Figure (ref) presents the estimated average marginal effects. The estimates are displayed on the interval $[15.2,23.8]$, namely the interval on which the EITC policy change is supposed to be pronounced. We focus on this region because elsewhere the denominator of $ \Delta _{app}^{AME}(x)$ is either close or equal to zero. The solid line represents the point estimate of the average marginal effect. Specifically, the line shows how much out of one dollar increase is spent on our nondurable consumption bundle. Our results are very much in line with the literature, with values ranging from 0.5 for disposable income just below \$16K to virtually zero for incomes above \$22K. Our point estimate also suggests that the average marginal effect decreases with income. This is in line with previous findings in the literature, in particular those of PSJM. Such a pattern is also consistent with rational consumers facing credit constraints. Indeed, credit constraints are likely to be less severe for households with higher income, as such consumers are on average more able to use parts of their wealth as collateral to get new credits more easily.

figure[figure omitted — 485 chars of source]

Several remarks are in order. The first concerns significance: While the results for low levels of disposable income are borderline pointwise significant at the 90% level, most of the estimated effect is insignificant (as are results based on 95% significance). This is in particular regrettable at income levels around \$ 18K where there is probably a substantive nonzero effect, but the evidence is slightly too weak to conclude this with statistical certainty. As already outlined above, there is significant noise in the data that complicates our analysis and the instrumental variation used to identify the model is only moderately strong. Having said that, given the borderline significance at lower income levels, we are confident that if we were to consider an estimator for the average marginal effect across the region between \$16K and \$19K we would find a strongly significant effect, because average derivatives are much more accurately estimable than pointwise derivatives. Developing such a formal test is quite involved and thus left for future research. Note that the monotonically declining shape is very much in line with the literature which finds the strongest evidence for the failure of intertemporal smoothing at lower income. While certainly not as precise as we had hoped for, we feel that our estimates lend support to the recently found evidence of excessively large effects of an anticipated shock to income.

The second remark concerns our modeling assumptions. As mentioned above, the stationarity assumption Assumption (ref) together with the modeling assumption limits the degree of unobserved heterogeneity. In particular, individual households might have heterogeneous preferences both for consumption and leisure that enter in a complicated fashion resulting in a multivariate $A_{t}$. While we acknowledge the possibility of these effects biasing our results, we do not think that they are large in absolute size. Labor supply of the main breadwinner, especially in families in the 1980s, has proven to be very inelastic to the degree that wages are frequently used as an instrument in consumer demand studies, see blundell1993we. This is less true for secondary income (e.g., part time work by the spouse). However, given the relatively small magnitude of the change, we would be surprised if the effect on labor supply be large (which would be the main channel for misspecification impacting our estimates). Still, we do acknowledge that a cautionary remark is in order at this point, also with respect to our omission of potentially complex dynamics as would arise, e.g., with habit formation.

The third remark concerns our omission of observable heterogeneity.While clearly important, as the paper does not develop the associated theory we leave this for future research. Having said, note that we work with the subsample of families with two or more children with at least (and typically in 1987 also at most) one bread winner of a low income level which is a fairly homogeneous population. A similar stratification strategy to deal with observed heterogeneity is very common in the consumer demand literature hoderlein2011many.

The last remark concerns the magnitude of the effect. Here, it is instructive to compare the marginal effect with the average expenditure share of our nondurable consumption measure. This share is roughly equal to 0.5 for the levels of income we consider.\footnote{ It is also very mildly decreasing with income levels, as one could expect.} Similarly, our results imply that at a disposable income level of 16.5., the consumers spend roughly 50 cent out of an additional dollar on nondurable consumption. This is compatible with a model where low income households, when receiving an (anticipated) additional dollar of income, consume it entirely and in roughly equal proportions on our set of nondurable consumption goods as well as on the remaining (mostly durable) consumption items. This points clearly to a violation of the hypothesis of rational consumers. The marginal effect diminishes to near zero for higher income levels. For incomes lower than 16.5, we find effects that are even larger than 0.5, meaning that households spend a larger fraction of every additional dollar on nondurable consumption than its income share. Since durable consumption is illiquid, we view such an effect as entirely conceivable, though we want to voice caution given the aforementioned large level of noise in the data.

In sum, we interpret our evidence as favoring the recent findings in the literature that low (disposable) income households spend large parts of an anticipated and possibly transitory real world shock on consumption. Conversely, they do not engage in intertemporal smoothing to the degree that the theory of rational consumer behavior would predict. Again, very much in parallel to recent findings, we also observe that this effect decreases with increasing disposable income, meaning that the driver for the higher effects at low levels is either liquidity constraints or a precautionary savings motive.

Conclusion

We consider in this paper an extension of the change-in-change model of Athey06 to continuous treatments. We impose similar restrictions as theirs on time effect and a crossing condition on the cdfs of the treatment variable. This crossing condition may be seen as a generalization of the existence of a control group in both the usual difference-in-difference and change-in-change settings. Importantly, our framework can allow for heterogeneous time trends and treatment effects. We show that under these conditions, some average and quantile treatment effects are point identified. We propose nonparametric multistep estimators of these treatment effects and show their asymptotic normality. Finally, we apply our method to the effect of disposable income on consumption. Our results suggest large effects for low-income households, in line with recent empirical findings.