Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
66,630 characters · 7 sections · 0 citation commands
Nonparametric Identification in Panels using Quantiles
Keywords: Panel data, nonseparable model, average effect, quantile effect, Engel curve
A frequent object of interest is the ceteris paribus effect of $x$ on $y,$ when observed $x$ is an individual choice variable partly determined by preferences or technology. Panel data holds out the hope of controlling for individual preferences or technology by using multiple observations for a single economic agent. This hope is particularly difficult to realize with discrete or other nonseparable models and/or multidimensional individual effects. These models are, by nature, not additively separable in unobserved individual effects, making them challenging to identify and estimate.
A fundamental idea for using panel data to identify the ceteris paribus effect of $x$ on $y$ is to use changes in $x$ over time. In order for changes over time in $x$ to correspond to ceteris paribus effects, the distribution of variables other than $x$ must not vary over time. This restriction is like \textquotedblleft time being randomly assigned\textquotedblright\ or "time is an instrument."\ In this paper we consider identification via such time homogeneity conditions. They are also the basis of many previous panel results, including Chamberlain (1982), Manski (1987), and Honore (1992). Recently time homogeneity has been used as the basis for identification and estimation of nonseparable models by Chernozhukov, Fernandez-Val, Hahn, Newey (2013), Evdokimov (2010), Graham and Powell (2012), and Hoderlein and White (2012). Because economic data often exhibits drift over time, we also allow for some time effects, while maintaining underlying time homogeneity conditions.
In this paper we give identification and estimation results for quantile effects with time homogeneity and continuous regressors. The effects of interest are derivatives of quantile structural functions of the model. We find that these derivatives are identified with two time periods for \textquotedblleft stayers", i.e. conditional on $x$ being equal in two time periods. Time homogeneity is too strong for many econometric applications where time trends are evident in the data. We weaken homogeneity by allowing for location and scale time effects. Allowing for such time effects makes identification and estimation more complicated but more widely applicable. We also give analogous results for conditional mean effects under weaker identification conditions than previously.
Quantile identification under time homogeneity is based on differences of quantiles. It is also interesting to consider whether quantiles of differences can help identify effects of interest. We do not find that time homogeneity alone can lead to identification from quantiles of differences. We do give quantile difference identification results that restrict the distribution of individual effects conditional on $x$, similarly to Chamberlain (1980), Altonji and Matzkin (2005), and Bester and Hansen (2009). In our opinion these added restrictions make quantiles of differences less appealing. We therefore focus for the rest of the paper, including the application, on differences of quantiles.
To illustrate we provide an application to Engel curve estimation. The Engel curve describes how demand changes with expenditure. We use data from the 2007 and 2009 waves of the Panel Study of Income Dynamics (PSID). Endogeneity in the estimation of Engel curves arises because the decision to consume a commodity may occur simultaneously with the allocation of income between consumption and savings. In contrast with the previous cross sectional literature, we do not rely on a two-stage budgeting argument that justifies the use of labor income as an instrument for expenditure. Instead, we assume that the Engel curve relationships are time homogeneous up to location and scale time effects, which leads to identification of structural effects from panel data.
An alternative approach to identification in panel data is to impose restrictions on the conditional distribution of the individual effect given $ x$. This approach leads to nonparametric generalizations of Chamberlain's (1980) correlated random effects model. As shown by Chamberlain (1984), Altonji and Matzkin (2005), Bester and Hansen (2009), and others, this kind of condition leads to identification of various effects. In particular, Altonji and Matzkin (2005) show identification of an average derivative conditional on the regressor equal to a specific value, an effect they call the local average response (LAR). In this paper we take a different approach, preferring to impose time homogeneity rather than restrict the relationship between observed regressors and unobserved individual effects. We refer to Hsiao (2003) for a broader perspective of panel data models.
Section 2 describes the model and gives an average derivative result. Section 3 gives the quantile identification result that follows from time homogeneity. Section 4 considers how quantiles of differences can be used to identify the effect of $x$ on $y$. Section 5 explains how we allow for time effects. Estimation and inference are briefly discussed in Section 6, and the empirical example is given in Section 7. The Appendix contains the proofs of the main results.
The data consist of $n$ observations on ${\boldsymbol{Y}}_{i}=(Y_{i1}, \ldots,Y_{iT})^{\prime }$ and ${\boldsymbol{X}}_{i}=[X_{i1}^{\prime },\ldots,X_{iT}^{\prime}] ^{\prime }$, for a dependent variable $Y_{it}$ and a vector of regressors $X_{it}$. Throughout we assume that the observations $ ({\boldsymbol{Y}}_{i},{\boldsymbol{X}}_{i})$, $(i=1,\ldots,n)$ , are independent and identically distributed. The nonparametric models we consider satisfy
We focus in this paper on the two time period case, $T=2$, though it is straightforward to extend the results to many time periods. The vector $ A_{i} $ consists of time invariant individual effects that often represent individual heterogeneity. The vector $V_{it}$ represents period specific disturbances. Altonji and Matzkin (2005) considered models satisfying Assumption 1. As discussed in Chernozhukov et. al. (2013), the invariance of $\phi $ over time in this Assumption does not actually impose any time homogeneity. If there are no restrictions on $V_{it}$ then $t$ could be one of the components of $V_{it},$ allowing the function to vary over time in a completely general way. The next condition together with Assumption 1 imposes time homogeneity on the model.
This is a static, or "strictly exogenous" time homogeneity condition, where all leads and lags of the regressor are included in the conditioning variable ${\boldsymbol{X}}_{i}.$ It requires that the conditional distribution of $V_{it}$ given ${\boldsymbol{X}}_{i}$ and $A_{i}$ does not depend on $t,$ but does allow for dependence of $V_{it}$ over time. This assumption rules out dynamic models where lagged values of $Y_{it}$ are included in $X_{it}.$
Setting $U_{it}=(A_{i}^{\prime },V_{it}^{\prime })^{\prime }$, an equivalent condition is
Thus, the time invariant $A_{i}$ has no distinct role in this model. As further discussed in Chernozhukov et. al. (2013), this seems a basic condition that helps panel data provide information about the effect of $x$ on $y.$ It is like the time period being "randomly assigned" or "time is an instrument," with the distribution of factors other than $x$ not varying over time, so that changes in $x$ over time can help identify the effect of $ x$ on $y$.
Although they seem useful for nonlinear models, the time homogeneity conditions are strong. In particular they do not allow for heteroskedasticity over time, which is often thought to be important in applications. We partially address this problem below by allowing for location and scale time effects.
For notational convenience we shall drop the $i$ subscript and let $T=2$ in the following. Our focus in this paper is on the case where the regressors ${ \boldsymbol{X}}$ are continuously distributed. We will be interested in several effects of ${\boldsymbol{X}}$ on ${\boldsymbol{Y}}$. For $ u=(a^{\prime },v^{\prime })^{\prime }$ we let $\phi (x,u)=\phi (x,a,v)$. We will let $x$ or $x_{t}$ denote a possible value of the regressor vector $ X_{t}$ and ${\boldsymbol{x}}=(x_{1}^{\prime },x_{2}^{\prime })^{\prime }$ a possible value of ${\boldsymbol{X}}=(X_{1}^{\prime },X_{2}^{\prime })^{\prime }$. Let $\partial _{x}\phi (x,u)$ denote the vector of partial derivatives of $\phi $ w.r.t. the coordinates of $x$. One effect we consider is a conditional expectation of the derivative $\partial _{x}\phi (X_{t},U_{t})$ given by
This is the object considered in Hoderlein and White (2012) and is similar to the local average response considered in Altonji and Matzkin (2005). It gives the local marginal effect for individuals with regressor value $x$ in both periods. This effect is related to the conditional average structural function (CASF):
through
under the conditions that permit interchanging the derivative and expectation.
The other effects we consider are similar to this effect except that we also condition on certain values of $Y_{t}$. One of these is given by
where $q(\tau ,x)$ is the $\tau ^{th}$ conditional quantile of $\phi (x,U_{t})$ given $X_{1}=X_{2}=x.$ This is a quantile derivative effect, similar to the local average structural derivative in Hoderlein and Mammen (2007). It gives the local marginal effect for individuals with regressor value $x$ in both periods and at the quantile $q(\tau ,x)$. This effect is also related to the conditional quantile structural function (CQSF), $ q_{\tau }(x|{\boldsymbol{x}})$, that gives the $\tau $-quantile of $\phi (x,U_{t})$ conditional on ${\boldsymbol{X}}={\boldsymbol{x}},$ through
We also consider linking quantiles of arbitrary linear combinations of the dependent variables $Y_{1}$ and $Y_{2}$ to conditional expectations of the form
These are dependent variable conditioned average effects. One intended direction is to compare the derivative of the quantiles of the differences $ Y_{2}-Y_{1}$ to the differences of the derivative of the quantiles of $Y_{2}$ and $Y_{1}$ in terms of objects they identify. In what follows we carry out the comparison.
To set the stage for the quantile results we first discuss mean identification. We first give an explanation of identification of the mean effect and then give a precise result with regularity conditions.
Consider the identified conditional mean
Together these conditional expectations are a nonparametric version of Chamberlain's (1982) multivariate regression model for panel data. Derivatives of them can be combined to identify the conditional mean effect. Let $f(u|{\boldsymbol{x}})$ denote the conditional density of $U_{t}$ given $ {\boldsymbol{X}}={\boldsymbol{x}},$ that does not depend on $t$ by Assumption (ref). Assume that $\phi (x,u)$ and $f(u|{ \boldsymbol{x}})$ are differentiable in $x$ and ${\boldsymbol{x}}$ respectively and that differentiation under the integral is permitted. For ${ \boldsymbol{x}}=(x_{1}^{\prime },x_{2}^{\prime })^{\prime }$ we let $ \partial _{x_{s}}M_{t}({\boldsymbol{x}})$ and $\partial _{x_{s}}f(u|{ \boldsymbol{x}})$, $s,t=1,2$, denote the vector of partial derivatives w.r.t. the coordinates of $x_{s}$. Then for $s,t=1,2$,
where the first term is the conditional mean effect of interest and the second term is the analog to Chamberlain's (1982) heterogeneity bias. Subtracting and using Assumption (ref) gives
Evaluating at ${\boldsymbol{x}}=(x^{\prime },x^{\prime })^{\prime }$ we find that
where
It also follows similarly that
Thus, the conditional mean effect is identified from the derivative of the conditional expectation of the difference with respect to the leading time period for individuals where $X_{t}$ is the same in both periods. We note here that this means the conditional mean effect is overidentified. Introducting time effects, as we do below, will lead to exact identification. Thus, testing for the presence of time effects is one way of testing this overidentifying restriction.
The importance of conditioning on the event $x=X_{1}=X_{2}$ can be seen from equation ((ref)), where setting $X_{1}=X_{2}$ eliminates heterogeneity bias. Thus, one can think of the conditioning on $X_{1}=X_{2}$ as a device to eliminate the heterogenity bias in nonseparable models under time stationarity. In contrast, if $\phi (x,u)$ were additively separable with $\phi (x,u)=\mu (x)+u,$ the heterogeneity bias would be zero for all $ X_{1}$ not necessarily equal to $X_{2}$ because $\int \partial _{x_{2}}f(u|{ \boldsymbol{x}})du=0$. Hence the derivative effect of interest would be $ \partial_{x_2} \Delta M(\boldsymbol{x})$ for each value of $x_{1}$ and one could estimate that derivative more precisely by averaging over its first argument. Also, one could test for whether the model is additively separable by testing whether $\Delta M(\boldsymbol{x})$ varies with its first argument, though it is beyond the scope of this paper to analyze such tests.
Conditioning on $x=X_{1}=X_{2}$ does restrict the set over which the structural derivative is averaged but this can correspond to an interesting set of individuals. For example, in the Engel curve application we give $x$ is total expenditure so the restriction $X_{1}=X_{2}$ corresponds to individuals whose total expenditure was the same in the two time periods. This seems mostly likely to occur for middle aged individuals, which is an interesting though special group to focus on.
Altonji and Matzkin (2005) are able to identify derivative effects without conditioning on $X_{1}=X_{2}$ but they also restrict the distribution of $ U_{t}$ conditional on $\boldsymbol{X}.$ We do not impose such type of assumptions but instead require time stationarity of the distribution of $ U_{t}$ conditional on $\boldsymbol{X}$. The different assumptions make it hard to compare results. We prefer to focus on time stationarity in this paper, where we do not yet know whether it is possible to identify interesting effects for continuous regressors without imposing $X_{1}=X_{2}.$
Graham and Powell (2012) consider a linear model with individual specific coefficients where $\phi (x,u)=\beta _{1}(u)+\beta _{2}(u)x$ in the scalar $x $ case. In this case
Here we find that average slope for the stayer subpopulation with $ X_{1}=X_{2}$ is identified. Graham and Powell (2012) use linearity of $\phi (x,u)$ in $x$ to identify the average slope $E\big[\beta _{2}(U_{t})\big]$ over the whole population using the movers with $X_{1}\neq X_{2}$. We identify an average slope over a smaller population for a fully nonlinear, nonparametric specification $\phi (x,u)$.
The following result makes the previous derivation precise, including conditions for differentiating under integrals.
This result has slightly weaker conditions than that of Hoderlein and White (2012). Here we drop their assumption that $V_{t}$ is independent of $X_{1}$ conditional on $A$. The result given here allows for $X_{1}$ to be correlated with $(V_{1},V_{2}),$ as long as the marginal distribution of $ V_{t}$ conditional on $(X_{1},X_{2},A )$ does not vary with $t.$ We maintain these weaker conditions as we consider identification of quantile effects in the next Section.
Turning now to the identification of the quantile effects given above, let $ Q_{t}(\tau \mid {\boldsymbol{x}})$ denote the $\tau ^{th}$ conditional quantile of $Y_{t}$ conditional on ${\boldsymbol{X}}={\boldsymbol{x}} =(x_{1}^{\prime },x_{2}^{\prime })^{\prime }$. It will be the solution to
The pair $[Q_{1}(\tau | {\boldsymbol{x}}),Q_{2}(\tau | {\boldsymbol{x}})]$ is a quantile analog of Chamberlain's (1982) multivariate regression for panel data. We can identify a quantile analog of the Hoderlein and White (2012) average derivative effect. We first describe how these multivariate panel quantiles can be used to identify an average derivative effect, then give a precise interpretation of the effect. This description helps explain the source of identification as well as the precise nature of the identified effect.
To describe how identification works, differentiate both sides of the previous identity with respect to $x_{s}$, treat the derivative of an indicator function as a dirac delta, and assume the order of differentiation and integration can be interchanged. This calculation gives
Let $g_{t}(\tau \mid {\boldsymbol{x}})=\int_{\phi (x_{t},u)=Q_{t}(\tau \mid { \boldsymbol{x}})}f(u|{\boldsymbol{x}})du$ and note that
Solving for $\partial _{x_{s}}Q_{t}(\tau |{\boldsymbol{x}})$ we find that,
Note that at $X_{1}=X_{2}=x$, $Q_{1}(\tau \mid x,x)=Q_{2}(\tau \mid x,x)=q(\tau ,x)$ and $g_{1}(\tau \mid x,x)=g_{2}(\tau \mid x,x)$ by time homogeneity. Then differencing the conditional quantile derivatives gives
where the last term does not depend on $t$ due to time homogeneity. The equation ((ref)) is a panel version of the Hoderlein and Mammen (2007) identification result. It is interesting to note that, unlike in the mean case, differences of derivatives of quantiles generally differ from derivatives of quantiles of differences. Below we will consider identification from derivatives of quantiles of differences.
To make the above derivation precise we need to formulate conditions that allow differentiation under the integral. The following regularity condition is one approach to this, in particular for the dirac delta argument given above.
The boundedness conditions on the derivatives of $\phi (x,u)$ could further be weakened at the expense of much more complicated notation and conditions.
For fixed $x$ let $f_{Y_{x}|{\boldsymbol{X}}}(y|{\boldsymbol{x}})$ denote the conditional pdf of $Y_{x}=\phi (x,U_{t})$ given ${\boldsymbol{X}}={ \boldsymbol{x}} = (x_1^{\prime }, x_2^{\prime })^{\prime }.$ The following lemma shows differentiability of $\mathbb{P} (\phi (x,U_{t})\leq y|{\boldsymbol{X}}= {\boldsymbol{x}})$ with respect to $x$ and $y $ for given ${\boldsymbol{x}}$ , and computes the derivatives.
With this result in hand we can now make precise the quantile effect sketched above.
To illustrate the previous result, consider the familiar linear model with additive heterogeneity $Y_t = X_{t}^{\prime }\theta + U_t,$ where $U_t = A + V_t.$ Let $\overline{Q}_{\tau}(\cdot \mid {\boldsymbol{X}})$ denote the linear $\tau$-quantile regression on $\text{vec}({\boldsymbol{X}}),$ a quantile version of the panel multivariate regression of Chamberlain (1982). Under time homogeneity
where $\gamma_{\tau1}$ and $\gamma_{\tau2}$ do not depend on $t$. Taking derivatives and differencing over time, for $s \neq t,$
Here the result holds for sequences ${\boldsymbol{x}}$ with $x_1 \neq x_2$ because the heterogeneity is additive.
In this section we answer the question whether we can relate quantiles of the first difference of the dependent variable to causal effects. In fact, the same arguments and assumptions that are used for first differences can also be employed for arbitrary functions of the dependent variables which map the $T$-vector of dependent variables ${\boldsymbol{Y}}$ (in our case for simplicity $T=2$) into a scalar \textquotedblleft index". However, as it turns out, if we restrict ourselves to using only two time periods of the covariates $X_{t} $, we have to strengthen the assumptions significantly to make statements about causal effects. This is related to the fact that we do not have an auxiliary equation at our disposal that allows us to correct for the heterogeneity bias that arose from the correlation of $X_{t}$ and $ U_{s}. $
To be more specific about the assumptions: While still considering the model specified in Assumption (ref), instead of time homogeneity assumption (ref), in this section we shall use independence assumptions.
The first part of this assumption states that the transitory error component is independent of covariates, given the persistent fixed effect, which is a notion of strict exogeneity. The second part of this assumption is more restrictive as it rules out the case where $A$ is arbitrarily correlated with the $X_{t}$ process. This is a special case of the sufficient statistic type assumptions in Altonji and Matzkin (2005). Assumption (ref) does not restrict the relationship between $X_t$ and $A$ and allows for $X_t$ and $V_t$ to be correlated, but it is not formally nested within Assumption (ref). To see this, consider the example of the panel multivariate quantile regression in the additive linear model of Section (ref). Without time homogeneity,
Assumption (ref) imposes time homogeneity on the coefficients, i.e., $\gamma_{\tau1,t} = \gamma_{\tau1},$ and $\gamma_{\tau2,t} = \gamma_{\tau2},$ whereas Assumption (ref) imposes the exclusion restrictions $\gamma_{\tau2,1} = 0$ and $\gamma_{\tau2,2} = 0$ but lets $\gamma_{\tau1,t}$ vary with $t$. In our view these exclusion restrictions are stronger than time homogeneity in most economic applications.
To adopt a similar framework as above, we rewrite
and note that the independence and strict exogeneity assumptions imply that:
As already mentioned above, we consider now quantiles of differences and other transformations of the dependent variables. To this end, let $\psi (y_{1},y_{2})$ be an arbitrary (differentiable) function and note that
so that for ${\boldsymbol{u}}=(v_{1},v_{2},a)$, we have that $g(x_{1},x_{2},{ \boldsymbol{u}})=\psi (\phi (x_{1},v_{1},a),\phi (x_{2},v_{2},a))$. Denote by $\tilde{q}(\tau ,x_{1},x_{2})$ the conditional quantile of $\tilde{Y}$ given $X=x_{1},X_{2}=x_{2}$, so that
For convenience, we first formulate and prove a result along the lines of Hoderlein and Mammen (2007) for a general model of the form
in terms of regularity assumptions similar to Assumption (ref), and then specialize it to ((ref)).
These assumptions are by and large regularity conditions, akin to those employed in Hoderlein and Mammen (2007), e.g., differentiability conditions. They do not restrict the model significantly, and we therefore do not discuss them at length. Together with the independence condition, they allow us to establish an extension to the Hoderlein and Mammen (2007) result:
We now specialize this general result to the setup of this paper, and discuss it below in this specialized setup. To this end, we modify the regularity conditions accordingly:
These preliminaries lead to the expected corollary:
This result is very similar in spirit to the results in the previous section, again an LAR for a subpopulation (or a derivative for an ASF) is identified. The advantage, however, is now that we can look at subpopulations that are characterized by arbitrary combinations of $Y_{1}$ and $Y_{2}$. If we confine ourselves to linear combinations, i.e., $\tilde{Y} =\lambda Y_{1}+\pi Y_{2},$ we can consider conditioning on arbitrary weights $\lambda ,\pi $. Since we can vary $\lambda ,\pi $ freely, this means that we can use the entire joint distribution in the sense of the Cramer-Wold device, by looking at any linear combination, and hence use multivariate information through repeated use of one regular regression quantiles. It allows to construct subpopulations where we put different weights on the outcome in different periods. For instance, if $X$ is schooling, and $Y_{t}$ is labor income in different periods, we may think of $\tilde{Y}$ as some long run or average income. And when computing this long run income, we could either discount future income stronger or emphasize it more when characterizing the subpopulations, depending on the intention of the researcher. Of course, one should always remember that the strength in statements we can make always comes at the expense of the structure we impose on the dependence between $A$ and $X_{t}.$
This result covers important special cases:
Note that the first special case answers one of the questions posed in the introduction:\ should we consider the difference of the quantiles or the quantiles of the differences, when talking about causal effects in panels. In terms of the strength of the assumptions, the verdict has to be clearly differences of quantiles. However, two remarks are in order: First, it also happens to be the case that under the additional structure on the dependence the quantiles of the difference yield a new effect that we could not have obtained through differences in quantiles. In particular, for targeted policy measures it may be sensible to use subpopulations that are defined by, e.g., first differences $\Delta Y$. More precisely, since individuals are often assumed to exhibit a pronounced loss aversion, i.e., they are more much sensitive towards a negative change in their status than a positive, it is conceivable that a policy maker would be much more interested in the subpopulation for which the effect $\Delta Y$ is negative. Similarly, measures that focus on the subpopulation exhibiting large values of $\Delta Y $ may be of interest, as high variance of $Y$ over time may not be a desirable feature for an individual.
Second, with more time periods we could weaken the restrictive independence assumptions. In particular, if three periods are available and only effects on supopulations defined by, say, first differences between two periods are of interest, we may allow for more correlation between the unobservables and the $X_{t}$ process, and use the third period to perform an analogous correction as in the previous section. Since this involves a simple combination of arguments, we do not elaborate on this further, and we still want to point to the difference in assumptions in the two periods case.
The time homogeneity assumption is a strong one that often seems not to hold in applications. In this section we consider one way to weaken it, by allowing for additive location effects and multiplicative scale effects. Allowing for such time effects leads to effects of interest being exactly identified, unlike the overidentification we found in Sections 2 and 3.
We allow for time effects by replacing Assumption 1 with the following condition.
The time effects $\mu_t$ and $\sigma_t$ are not separately identifiable from $\phi$ without location and scale normalizations because
for $\tilde \mu_t(x) = \mu_t(x) + \sigma_t(x) \Delta_{\mu} (x),$ $\tilde \sigma_t(x) = \Delta_{\sigma}(x )\sigma_t(x),$ $\tilde \phi(x,u) = [\phi(x,u) - \Delta_{\mu}(x)] /\Delta_{\sigma}(x)$, and $\Delta_{\sigma}(x) \neq 0$.
In this model the effects of interest vary with time. We consider the time-averaged conditional mean effect:
and the time-averaged conditional quantile effect:
where $\bar \mu(x) = [\mu_1(x) + \mu_2(x)]/2,$ $\bar \sigma(x) = [\sigma_1(x) + \sigma_2(x)]/2,$ and $q(\tau, x)$ is the $\tau^{th}$ conditional quantile of $\phi(x,U_t)$ given $X_1 = X_2 = x.$
The conditional mean effect is related to the time-averaged CASF:
through
under the conditions that permit interchanging the derivative and expectation. Similarly, the conditional quantile effect is related to the time-averaged CQSF, $\bar q_{\tau}(x \mid {\boldsymbol{X}} = {\boldsymbol{x}} )$, that gives the $\tau$-quantile of $\bar \mu(x) + \bar \sigma(x) \phi(x,U_t)$ conditional on ${\boldsymbol{X}}= {\boldsymbol{x}}$, through
Let $V_t({\boldsymbol{x}}) =\text{Var}[Y_t \mid {\boldsymbol{X}} = { \boldsymbol{x}}],$ and $\sigma(x) = \sigma_2(x)/\sigma_1(x).$
This theorem shows that the time effects are identified up to location and scale normalizations. For example, if we set $\mu_1(x) = 0$ and $\sigma_1(x) =1,$ then $\sigma_2^2(x) = V_2(x,x)/V_1(x,x)$ and $\mu_2(x) = E[Y_2 - \sigma_2(x) Y_1 \mid X_1 = X_2 =x]$. The identification of the conditional mean effect does not require any normalization. Note that we now have just one equation for identifying the conditional mean effect.
We find a similar result for quantiles.
As in Theorem (ref), the time effects are identified up to location and scale normalizations, whereas the conditional quantile effects are identified without any normalization. Here, however, instead of conditional mean and variance restrictions, we use quantile restrictions to identify the time effects up to the normalizations. These effects are over identified by many possible quantiles $\tau _{1},$ $\tau _{2}$ and $\tau _{3} $. For example, for $\tau _{1}=.9,$ $\tau _{2}=.1$ and $\tau _{3}=.5,$ the scale is identified by a ratio of conditional interdecile ranges across time and the location is identified by a difference of conditional medians across time.
We note that Graham and Powell (2012) allowed for random time effects in location and slope rather than location and scale effects that could depend on $X$.
The conditional mean and quantile effects of interest are identified by special cases of the functionals:
and
respectively, where $h_m$ and $h_q$ are known smooth functions, $\mathcal{X}$ is a region of regressor values of interest, and $\mathcal{W}$ is a region of regressor values and quantiles of interest. We consider the estimators of $\theta_m$ and $\theta_q$ based on the plug-in rule:
and
where $\widehat M_t(x,x)$, $\widehat Q_t(\tau \mid x,x)$, and $\widehat V_t(x,x)$ are nonparametric series estimators of $M_t(x,x)$, $Q_t(\tau \mid x,x)$, and $V_t(x,x)$.
To describe the series estimators, let $P^K({\boldsymbol{x}}) = (p_{1K}({ \boldsymbol{x}}), \ldots, p_{KK}({\boldsymbol{x}}))^{\prime }$ denote a $K \times 1$ vector of approximating functions, such as tensor products of univariate polynomial or spline series terms of the components of $ \boldsymbol{x}$, and let $\boldsymbol{P}_i = P^K({\boldsymbol{X}}_i)$. Then,
where $A^-$ denotes any generalized inverse inverse of the matrix $A$;
is a series version of the (kernel) conditional variance estimator of Fan and Yao (1998); and $\widehat Q_t(\tau \mid x,x) = P^K(x,x)^{\prime }\widehat \beta_t(\tau),$ where $\widehat \beta_t(\tau)$ is the Koenker and Bassett (1978) quantile regression estimator
Following Praestgaard and Wellner (1993), Hahn (1995), and Chamberlain and Imbens (2003), we use weighted bootstrap for inference.\footnote{See also Ma and Kosorok (2005) and Chen and Pouzo (2009, 2013) for other applications of weighted bootstrap; we are grateful to a referee for pointing out the latter references.} To describe this method, let $ (w_{1},\ldots,w_{n})$ be an i.i.d. sequence of nonnegative random variables from a distribution with mean and variance equal to one (e.g., the standard exponential distribution), independent of the data. The weighted bootstrap uses the components of $(w_{1},\ldots,w_{n})$ as random sampling weights in the construction of the bootstrap version of the series estimators. Thus, the bootstrap versions of $\widehat \theta_m(w)$ and $\widehat \theta_q(w)$ are
and
where
is the bootstrap version of $\widehat M_t(x,x),$
is the bootstrap version of $\widehat V^*_t(x,x),$ and $\widehat Q^*_t(\tau \mid x,x) = P^K(x,x)^{\prime }\widehat \beta^*_t(\tau)$ is the bootstrap version of $\widehat Q_t(\tau \mid x,x)$, with
Belloni, Chernozhukov, Chetverikov, and Kato (2013) and Chernozhukov, Lee, and Rosen (2013) developed functional distributional theory and bootstrap consistency for series estimators of functionals of the conditional mean function, and Belloni, Chernozhukov, and Fernandez-Val (2011) developed similar theory for series estimators of functionals of the conditional quantile function. We can use these results to construct analytical or bootstrap confidence bands for the effects that have uniform asymptotic coverage over regressor values and quantiles. For example, the end-point functions of a $1-\alpha$ confidence band for $\theta_q$ have the form
where $\widehat \Sigma_q(w)$ and $\widehat t_{q,1-\alpha}$ are consistent estimators of the asymptotic variance function of $\sqrt{n} [\widehat{\theta} _q(w) - \theta_q(w)]$ and the $1 - \alpha$ quantile of the Kolmogorov-Smirnov maximal $t$-statistic
The following algorithm describes how to obtain uniform bands for quantile effects using weighted bootstrap:
The validity of Algorithm (ref) follows from the results in Belloni, Chernozhukov, and Fernandez-Val (2011) and the delta method. We can construct uniform bands for the conditional mean effects with a similar algorithm replacing $\theta_q(w)$ by $\theta_m(x)$, adjusting all the steps accordingly, and relying on the results of Belloni, Chernozhukov, Chetverikov, and Kato (2013) and Chernozhukov, Lee, and Rosen (2013).
In this section, we illustrate the results with an empirical application on estimation of Engel curves with panel data. The Engel curve relationship describes how a household's demand for a commodity changes as the household's expenditure increases. Lewbel (2006) provides a recent survey of the extensive literature on Engel curve estimation. We use data from the 2007 and 2009 waves of the Panel Study of Income Dynamics (PSID). Since 2005, the PSID gathers information on household expenditure for different categories of commodities. The PSID does not collect information on total expenditure. We construct the total expenditure on nondurable goods and services by adding all the expenses in housing, utilities, phone, child care, food at home, food out from home, car, transportation, schooling, clothing, leisure, and health. We exclude expenses in mortgage, home insurance, car insurance, and health insurance because these categories have many missing values. Our sample contains $968$ households formed by couples without children, where the head of the household was 20 to 65 year-old in 2009, and that provided information about all the relevant categories of expenditure in 2007 and 2009. We focus on the commodities food at home and leisure for comparability with recent studies (e.g., Blundell, Chen, and Kristensen (2007), Chen and Pouzo (2009, 2013), and Imbens and Newey (2009)). The expenditure share on a commodity is constructed by dividing the expenditure in this commodity by the total expenditure in nondurable goods and services.
Endogeneity in the estimation of Engel curves arises because the decision to consume a commodity may occur simultaneously with the allocation of income between consumption and savings. In contrast with the previous cross sectional literature, we do not rely on a two-stage budgeting argument that justifies the use of labor income as an instrument for expenditure. Instead, we assume that the Engel curve relationships are time homogeneous up to location and scale time effects, and rely on the availability of panel data. Specifically, we estimate
where $Y$ is the observed share of total expenditure on food at home or leisure, $X$ is the logarithm of total expenditure in dollars of 2005, $ \mu_t(X)$ and $\sigma_t(X)$ are location and scale time effects, $U$ is a vector of unobserved household heterogeneity that satisfies time homogeneity and captures both differences in preferences and idiosyncratic household shocks, $t=1$ corresponds to 2007, and $t=2$ corresponds to 2009.\footnote{ To deflate total expenditure, we use a price index for personal consumption expenditures in nondurable goods constructed from Tables 2.4.4 and 2.4.5 of the Bureau of Economic Analysis.} The inclusion of time effects might be important to account for temporal changes in preferences and relative prices across commodities. For example, the price index of nondurable goods increased by $7\%$ between 2007 and 2009, whereas the price indexes for food and leisure increased by $10\%$ and $6\%$ during the same period.\footnote{ Source: Tables 2.4.4.U and 2.4.5 of Bureau of Economic Analysis.} We allow these time effects to vary with total expenditure, what gives flexibility to the model. This model does put some restrictions on interactions between prices and heterogeneity, implying that price changes only shift the location and scale of the distribution of demand.
Table (ref) reports descriptive statistics for the variables used in the analysis. Both total expenditure and expenditure shares display within and between household variation, with means and standard deviations that remain stable between 2007 and 2009. The low percentage of within variation in expenditure indicates that there might be a substantial number of households with zero or little change in expenditure across years. Figure (ref) plots histogram and kernel estimates of the density of the change in expenditure between 2007 and 2009. The kernel estimates are obtained using a Gaussian kernel with Silverman's rule of thumb for the bandwidth. The estimates confirm that there is a high density of households with zero change in expenditure. Our methods will identify mean and quantile effects for these households with $X_{i1} = X_{i2}$.
We estimate the location time effects, scale time effects, conditional mean effects, and conditional quantile effects using sample analogs of the expressions in Theorems (ref) and (ref). In particular, we estimate the conditional expectation, variance, and quantile functions by the nonparametric series methods described in Section (ref). We consider two different specifications for the series basis in all the estimators: a quadratic orthogonal polynomial and a cubic B-spline with three knots at the minimum, median and maximum of total log-expenditure in the data set. Both specifications are additively separable in the total log-expenditures of 2007 and 2009.\footnote{ We select these specifications by under smoothing with respect to the specification selected by cross validation applied to the estimators of the conditional expectation function.} We also compute cross sectional estimates that do not account for endogeneity. They are obtained by averaging the nonparametric series estimates in 2007 and 2009 that use the same specification of series basis as the panel estimates but only condition on contemporaneous expenditure. For inference, we construct 90% confidence bands around the estimates by weighted bootstrap with exponential weights and 499 repetitions. These bands are uniform in that they cover the entire function of interest with 90% probability asymptotically.
Figures (ref) and (ref) show the estimates and confidence bands for the time effects functions:
based on Theorem (ref) with $\tau_1 = .9,$ $\tau_2 = .1,$ and $ \tau_3 = .5$, where $\mathcal{X}$ is the interval of values between the 0.10 and 0.90 sample quantiles of log-total expenditure. We find that we cannot reject the hypothesis that there are no location and scale time effects for food at home, whereas we find significant evidence of time effects for leisure with both series specifications. In results not reported, we find similar estimates and confidence bands for the time effects functions based on conditional means and variances using Theorem (ref).
Figure (ref) plots the estimates and confidence bands for the time-averaged conditional quantile effects or CQSF derivates integrated over the values of $x$:
for $\tau \in \mathcal{T},$ where $\mu$ is the empirical measure of log-expenditure, and $\mathcal{T} = [0.1,0.9].$ Here we find heterogeneity in the Engel curve relationship across the distribution. The pattern of the effect is increasing with the quantile index for both food at home and leisure, although the estimates are not sufficiently precise to distinguish these patterns from sampling noise. The cross sectional estimates plotted in dashed lines lie outside the confidence band for leisure, indicating significant evidence of endogeneity. We do not find such evidence for food at home.
In figures (ref) and (ref), we show that the panel estimates of the CQSF as a function of expenditure are decreasing for food at home and increasing for leisure at low values of expenditure. Imbens and Newey (2009) and Chen and Pouzo (2009, 2013) found similar patterns in their estimates of the QSF and the quantile Engel curves, respectively. Figure (ref) plots the estimates and confidence bands for the time-averaged conditional mean effects or CASF derivatives:
We also again evidence of endogeneity for leisure in the mean effects, but not for food at home. As in Blundell, Chen and Kristensen (2007), the conditional ASF is decreasing in expenditure for food at home, whereas it is increasing for leisure. We find that the curve is convex for food at home and concave for leisure. Note, however, that we should interpret the shape of our panel estimates with caution because they formally correspond to multiple conditional QSFs and ASFs as the conditioning set $X_1=X_2 =x$ changes with $x$ along the curve.
Overall, the empirical results show that our panel estimates of the Engel curves are similar to previous cross sectional estimates based on IV methods to deal with endogeneity. Thus, the Engel curve relationship is decreasing for food at home and increasing for leisure. Moreover, we find evidence of the presence of time effects and endogeneity for leisure, but not for food at home. These finding are consistent with consumer preferences where food at home is a necessity good with little effect on the marginal allocation of income between consumption and savings. Leisure, on the other hand, is a superior good that affects the marginal allocation of income between consumption and savings. The Engel curve relationship is stable over time for food at home, whereas it is sensitive to changes over time in preferences and relative prices for leisure.