Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
134,566 characters · 13 sections · 0 citation commands
Average and Quantile Effects in Nonseparable Panel Models
Interesting empirical questions are often formulated in terms of the ceteris paribus effect of $x$ on $y,$ when observed $x$ is an individual choice variable partly determined by preferences or technology. Panel data holds out the hope of controlling for individual preferences or technology by using multiple observations for a single economic agent. This hope is particularly difficult to realize with discrete or other nonseparable models and/or multidimensional individual effects. These models are, by nature, not additively separable in unobserved individual effects, making them challenging to identify and estimate. There are some simple solutions, such as the conditional MLE for the slope parameter of a binary-choice logit model with an individual location effect. However these are rare and dependent on specific models or distributions. For example, the slope parameter of the binary-choice model with a time dummy is identified only for logit as shown by Chamberlain (2010), and the average treatment effect is not identified even for logit without a time dummy, as shown below.
A fundamental idea for using panel data to identify the ceteris paribus effect of $x$ on $y$ is to use changes in $x$ over time to estimate the effect. In order for changes over time in $x$ to correspond to ceteris paribus effects, the distribution of variables other than $x$ must not vary over time. This condition is like \textquotedblleft time being randomly assigned\textquotedblright\ or \textquotedblleft time is an instrument.\textquotedblright\ In this paper we consider identification via such time homogeneity conditions. They are also the basis of many previous panel results, including Chamberlain (1982), Manski (1987), and Honore (1992). Here we consider the identifying power of time homogeneity for nonseparable models, i.e. for models that are not additively separable in unobserved factors. We allow for multidimensional heterogeneity, as motivated by models where effects of interest, such as price and income elasticities, are distributed among individuals in unrestricted ways; see Altonji and Matzkin (2005), Browning and Carro (2007), and Fernandez-Val and Lee (2010), among others. We also weaken the strict time homogeneity conditions to allow some time effects.
Models with discrete regressors have many applications and are the subject of most of this paper. With discrete regressors, time homogeneity only leads to partial identification of many effects, though some conditional effects are identified. This paper considers partial identification and estimation of average and quantile effects, under static or dynamic conditions, in fully nonparametric and in semiparametric models, with time effects.
For the nonparametric, static model we give simple estimators of the identified average effect of $x$ on $y$, conditional on $x$ varying over time. These estimators extend Chamberlain (1982, pp. 10-17) to multiple regressors with location and scale time effects. We also find that linear, fixed-effects estimate a variance-weighted average effect instead of the average effect. For bounded $y$ we move beyond the analysis of identified effects and give simple estimators of sharp bounds for average effects. These bounds provide nonparametric, partial-identification estimates of average effects in important cases, such as binary choice in panel data.
The quantile estimators given here are more novel than the average-effect estimators. They provide simple estimators of the effect of $x$ on quantiles of $y,$ conditional on $x$ varying over time, that allow for location and scale time effects. Estimators of sharp bounds are also provided for the unconditional, overall quantile effect. The estimators allow for multidimensional heterogeneity, for example for both location and slope to vary across individuals in an unrestricted way. In this way we provide a solution for the important problem of nonparametric quantile regression in panel data with individual effects, for discrete regressors. Graham, Hahn and Powell (2009) also consider quantile effects in linear, heterogenous coefficients models, but impose conditions which essentially restrict the heterogeneity to be one-dimensional, and focus on identification of the distribution of coefficients.
Dynamics is often an important feature of economic models with intertemporal choice. Here we give a dynamic, nonseparable, panel model that nests the static one. Simple estimators of bounds on average and quantile effects are provided. We show that these results provide a partial-identification solution to the important problem of distinguishing state dependence from heterogeneity.
This paper shows the impact of the number of time periods $T$ on identification. We find that the identified set of effects shrinks to a point exponentially quickly as $T$ grows, when individual effects are bounded and time period disturbances are not, and that the rate is some power of $T^{-1}$ more generally. In a nonparametric, dynamic, binary-choice model we find that the rate is faster the larger the variance of the period-specific disturbance relative to the variance of the individual effect.
In numerical examples we find that the nonparametric bounds can be quite wide, motivating more informative models. Semiparametric models that specify the distribution of the outcome given regressors and individual effect is an important class of more informative models. Here we describe both static and dynamic semiparametric models. When restrictions are imposed on the heterogeneity, like only some coefficients varying across individuals, semiparametric models can have substantially tighter bounds than nonparametric models. We find that in the important binary-logit model with just a location effect the average effect bounds shrink exponentially quickly as $T$ grows, in both dynamic and static models, even when the nonparametric bounds shrink slowly. This result quantifies the gain in information of a semiparametric model with just a location effect over the nonparametric model. We also find quite tight bounds for semiparametric models relative to nonparametric ones in numerical examples.
We show that semiparametric, discrete-choice models have finite dimensional parameterizations. This reduces bounds calculation and estimation to a finite-dimensional problem, albeit a large dimensional, highly nonlinear, and computationally difficult one. To make computation more feasible we use grids of fixed values for individual effects, so that average choice probabilities are finite-dimensional, linear combinations. We combine this with minimum squared distance fitting of data cell probabilities to obtain a quadratic programming approach for estimating the individual-effect distributions. This approach is computationally convenient and overcomes problems with previously proposed methods, as further discussed below. We also allow the grid to grow in order to approximate the true support points. It turns out that because the model is finite dimensional there is no need to limit the number of grid points. Mathematically, a richer fixed grid simply corresponds to a bigger submodel of the finite-dimensional model.
The semiparametric bounds build on Honor\'{e} and Tamer (2003, 2006) and Chernozhukov, Hahn, and Newey (2004). Both papers gave results for bounds in semiparametric, nonlinear, panel-data models. Honore and Tamer (2006) proposed linear programming, minimum distance, and maximum likelihood methods for dynamic models. Chernozhukov, Hahn, and Newey (2004) proposed sieve likelihood estimation of bounds for static models. These approaches are not very useful for estimation. Plugging in sample frequencies in place of cell probabilities in the linear-programming algorithm produces empty identification regions because the frequencies need not satisfy constraints imposed by the model. Also, the minimum-distance objective function is computationally difficult, as is sieve maximum likelihood, given the dimensionality of the individual-effect distributions. Honore and Tamer (2006) also assumed a fixed known grid for true individual effects, while we consider an approximation to an unknown grid.
The inferential problem for the semiparametric models is also rather challenging. The models impose data-dependent constraints that are often infeasible in finite samples or under misspecification, which produces empty confidence regions. We overcome these difficulties by projecting these data-dependent constraints onto the model space using the quadratic-programming approach mentioned above, thus producing an always-feasible, data-dependent constraint set. We then suggest linear and nonlinear programming methods that use these new modified constraints. Our inference procedures have the appealing justification of targeting the true model under correct specification and targeting a best approximating model under incorrect specification. We also develop two novel inferential procedures, one called the perturbed bootstrap, that is described in the paper, and another called modified projection, that is described in the Supplementary Material. These methods produce uniformly valid inference in large samples and may be of substantial independent interest.
We give two empirical illustrations. One is to estimate the effect of unions on earnings quantiles. There we find that a decline in the union effect as the quantile increases can be attributed to individual heterogeneity. The other illustration is to estimate the effects of fertility on women's labor force participation. There we compare nonparametric and semiparametric estimates.
Recent research has considered nonseparable panel models with time homogeneity and continuous regressors. Graham and Powell (2011) give estimators of the average effect in a linear model with heterogeneous slopes. Hoderlein and White (2011) give estimators of the average derivative conditional on equality of regressors across time periods.
Chamberlain (1980, 1984), Altonji and Matzkin (2005), Bester and Hansen (2008), and others have used control functions for panel data estimation. We focus instead on time homogeneity with unrestricted dependence between individual effects and regressors. Bias-corrected, fixed-effects estimation of semiparametric models has been proposed by Hahn and Kuersteiner (2002), Alvarez and Arellano (2003), Woutersen (2002), Hahn and Newey (2004), and Fern\'{a}ndez-Val (2009). These estimators depend on large $T$ for consistency while we estimate identified effects and bounds for fixed $T$.
Section 2 describes the models and effects we consider. Section 3 discusses estimation of identified effects. Sections 4 and 5 derive bounds for the static and dynamic nonparametric models respectively. Section 6 describes the impact of $T$. Section 7 describes and gives results for semiparametric, discrete-choice models. Section 8 gives computationally convenient methods for semiparametric models and numerical examples. Section 9 considers estimation and inference for semiparametric models. Section 10 gives empirical examples. The Supplementary Material Chernozhukov et. al. (2012) includes a variety of omitted discussions and results along with the proofs of results stated in the paper.
The data consist of $n$ observations on $Y_{i}=(Y_{i1},...,Y_{iT})^{\prime }$ and $X_{i}=[X_{i1},...,X_{iT}]^{\prime }$, for a dependent variable $Y_{it}$ and a vector of regressors $X_{it}$. Throughout we assume that the observations $(Y_{i},X_{i})$, $(i=1,...,n)$, are independent and identically distributed. The nonparametric models we consider satisfy
Assumption 1: There is a function $g_{0}(x,\alpha ,\varepsilon )$\ and vectors $\alpha _{i}$\ and $ \varepsilon _{it}$\ of random variables such that
The vector $\alpha _{i}$ consists of time invariant individual effects that often represent individual heterogeneity. The vector $\varepsilon _{it}$ represents period-specific disturbances. Altonji and Matzkin (2005) considered models satisfying Assumption 1. The invariance of $g_{0}$ over time in this Assumption does not actually impose any time homogeneity. If there are no restrictions on $\varepsilon _{it}$ then $t$ could be one of the components of $\varepsilon _{it},$ allowing the function to vary over time in a completely general way. The next condition, together with Assumption 1, imposes time homogeneity on the model.
Assumption 2: $\varepsilon _{it}|X_{i},\alpha _{i}\overset{d}{=} \varepsilon _{i1}|X_{i},\alpha _{i}$, for all $t.$
This is a static, or \textquotedblleft strictly exogenous\textquotedblright\ time homogeneity condition, where all leads and lags of the regressor are included in the conditioning variable $X_{i}.$ It requires that the conditional distribution of $\varepsilon _{it}$ given $X_{i}$ and $\alpha _{i}$ does not depend on $t,$ but does allow for dependence of $\varepsilon _{it}$ over time. An equivalent condition is $\tilde{\varepsilon}_{it}|X_{i} \overset{d}{=}\tilde{\varepsilon}_{i1}|X_{i}$ for $\tilde{\varepsilon} _{it}=(\alpha _{i},\varepsilon _{it}).$ Thus, the time invariant $\alpha _{i} $ has no distinct role in this model. The condition is just that whatever the unobserved disturbances are, their conditional distribution given $X_{i}$ does not depend on $t$.
This seems a basic condition that helps panel data provide information about the effect of $x$ on $y.$ It is like \textquotedblleft time is randomly assigned\textquotedblright\ or \textquotedblleft time is an instrument\textquotedblright\ with the distribution of factors other than $x$ not varying over time, so that changes in $x$ over time can help identify the effect of $x$ on $y$. Assumption 2 also turns out to be a natural strengthening of linear model conditions, as shown in Theorem A1 and the associated discussion in the Supplementary Material.
A dynamic model can be obtained by only including current and lagged $X_{is}$ in the conditioning set for each $t,$ as in the following condition:
Assumption 3: $\varepsilon _{it}|X_{it},...,X_{i1},\alpha _{i} \overset{d}{=}\varepsilon _{i1}|X_{i1},\alpha _{i}$,\ for all $t.$
This is a \textquotedblleft predetermined\textquotedblright\ version of time homogeneity that is nested within the static model of Assumptions 1 and 2, as shown in Theorem A2 of the Supplementary Material. Here the conditional distribution given only current and lagged regressors must be time invariant. It also implies that the conditional distribution of $\varepsilon _{it}$ given current and lagged regressors only depends on $X_{i1}$. Here $ \varepsilon _{it}$ can be thought of as additional information that is independent of the past regressors. A conditional-mean version of this condition arises in rational-expectations models that implies disturbances have mean zero conditional on past information. Here the stronger conditional independence restriction is imposed as seems needed for a nonseparable model. The conditioning on $X_{i1}$ is a way to account for the initial conditions of this dynamic model. Bhargava and Sargan (1983) adopted this approach in a linear model as have Honore and Tamer (2006) and Browning and Carro (2007) in a likelihood setting.
If $X_{it}$ includes lagged $Y_{it}$ then Assumption 3 specifies that the model is \textquotedblleft dynamically complete,\textquotedblright\ ruling out $Y_{it}=g_{0}(X_{it},\alpha _{i},\varepsilon _{it})$ as one equation of a dynamic system. For instance, $X_{it}$ could be $Y_{i,t-1},$ in which case $Y_{it}=g_{0}(Y_{it-1},\alpha _{i},\varepsilon _{it})$ is an explicit nonseparable dynamic model with $\varepsilon _{it}$ being time shocks that are independent of $Y_{it-1},...,Y_{i1}$. An important example is one where $ Y_{it}\in \{0,1\}$ is binary, representing state dependence, with $\alpha _{i}$ representing unobserved heterogeneity. This example is treated in Section 5.
We will focus in the nonparametric model on two objects, the average structural function (ASF) of Blundell and Powell (2003) and the quantile structural function (QSF) of Imbens and Newey (2009). The ASF is
where throughout the paper $F$ denotes the cumulative distribution function (CDF) of a random vector that appears as the arguments of $F$. This object is useful for quantifying the effect of $x$ on the mean of the outcome $ Y_{it}$. In the treatment-effects literature the average treatment effect (ATE) of changing $x$ from $x^{b}$ (before) to $x^{a}$ (after) is
The QSF $q(\lambda ,x)$ is the $\lambda ^{th}$ quantile of $g_{0}(x,\alpha _{i},\varepsilon _{it}).$ Under conditions specified below the QSF will equal the inverse of the CDF of $g_{0}(x,\alpha _{i},\varepsilon _{it})$,
In the treatment-effects literature the $\lambda ^{th}$ quantile treatment effect (QTE) of changing $x$ from $x^{b}$ to $x^{a}$ is
as in Lehmann (1974). This effect does not give the quantile of the treatment effect but does quantify the shift in the distribution of $Y_{it}$ that is due to a change in $x.$ It accounts for multidimensional individual effects that may be correlated with $x$.
The static model implies a conditional-mean model that has been considered by Chamberlain (1982), Hahn (2001), Wooldridge (2005), and Chernozhukov et. al. (2007). This conditional-mean model specifies that there is an $\alpha _{i}$ and $m_{0}(x,\alpha )$ such that $E[Y_{it}|X_{i},\alpha _{i}]=m_{0}(X_{it},\alpha _{i}).$ A conditional mean ATE, as in Wooldridge (2005), is $\int [m_{0}(x^{a},\alpha )-m_{0}(x^{b},\alpha )]dF(\alpha )$. This model and effect differ from those we consider in specifying conditional-mean restrictions, while we specify conditional distribution restrictions. In Theorem A3 of the Supplementary Material we show that the conditional-mean model is implied by Assumptions 1 and 2, or 1 and 3, and that the conditional mean ATE is equal to the ATE we consider. Thus all results we give for the ATE, including bounds, apply to the conditional mean models, such as that of Chernozhukov et. al. (2007).
To help explain the relationship between the conditional-mean model and the model of our paper, and to illustrate other results, it is useful to consider examples. Binary choice is a very important model for panel data, as it has many applications. For this reason we use binary choice as a main example. The most common model has been one with a scalar individual effect that is an additive shift to a linear combination of $X_{it}$, where
for scalar $\varepsilon _{it}$ and an unknown parameter vector $\beta ^{\ast }$. In this example $g_{0}(x,\alpha ,\varepsilon )=1(x^{\prime }\beta ^{\ast }+\alpha \geq \varepsilon )$ and the ATE is
This is an unusual object in the binary choice literature but is equal to a conditional mean ATE. In particular, if $\varepsilon _{it}$ is independent of $(X_{i},\alpha _{i})$ with CDF $H(\varepsilon )$ for each $t$. Then $ E[Y_{it}|X_{i},\alpha _{i}]=\Pr (Y_{it}=1|X_{i},\alpha _{i})=H(X_{it}^{\prime }\beta ^{\ast }+\alpha _{i})$ and
Thus the ATE is also the effect of changing $x$ on the choice probabilities averaged over the individual effect, i.e. the conditional mean ATE.
Our model also includes binary choice with individual-specific slopes as a special case. Economic motivation for varying slopes is provided by Browning and Carro (2007, 2009) who point out that with constant slopes the sign of the treatment effect is the same for every individual and give empirical examples where varying slopes are important. A general model with varying slopes is $Y_{it}=1(X_{it}^{\prime }\alpha _{i}\geq \varepsilon _{it})$ where $X_{it}$ now includes a constant and $\varepsilon _{it}$ is independent of $(X_{i},\alpha _{i})$ with CDF $H(\varepsilon ).$ In this model
accounting for individual specific slopes. When $X_{it}$ is discrete and fully saturated (e.g. consists of a full set of dummies, one for every discrete outcome) this model is actually equivalent to the general static model. It will be more restrictive when the distribution of $\alpha $ is restricted in some way, such as having some components of $\alpha $ be constant. In the semiparametric analysis described below we show how to impose such restrictions.
Time effects are clearly important in practice but identification of treatment effects will preclude including $t$ among the regressors $X_{it}$ in the nonparametric model of Assumptions 1 - 3. Identification will be based on variation over time in $X_{it}$, and if $t$ is a regressor then $ g_{0}(X_{it},\alpha _{i},\varepsilon _{it})$ has unrestricted variation over time, precluding identification of the effect of any other regressor. Some time effects can be allowed for by restricting the way $t$ enters $g_{0}.$ Below we will describe how this is done in semiparametric, discrete-choice models. With continuous $Y_{it}$ one can allow for location and scale time effects that are relatively easy to estimate.
Assumption 4: There is a function $g_{0}(x,\alpha ,\varepsilon ),$\ vectors $\alpha _{i}$\ and $\varepsilon _{it}$\ of random variables, and constants $\tau _{t},s_{t},(t=2,...,T)$\ such that for $\tau _{1}=0,$ $s_{1}=1,$
This condition allows the mean and variance of $Y_{it}$ to vary over time in an unrestricted way. The condition could be generalized to allow for other time effects, but we leave that to future work. It does not apply to $Y_{it}$ with fixed, discrete support because Assumption 4 does not make sense in that case. There $t$ must be included \textquotedblleft inside\textquotedblright\ $g_{0}$, as we do in the semiparametric analysis described below.
With these time effects the ASF and QSF can depend on $t$. The ASF and QSF for the first period will be $\mu (x)$ and $q(\lambda ,x)$ as given above, and for the other periods are
Corresponding period-specific and time-averaged ATE and QTE are given by
where $s_{1}=1$.
In the rest of this paper we will focus on discrete regressors, imposing the following condition from here on:
Assumption 5: The support of $X_{i}$\ is finite$.$
With discrete $X_{it}$ the model can also be written as a multiple regression with random coefficients, though we find it convenient to use the notation given here.
The analysis of identification in the static model is quite simple. This simplicity is a virtue, leading to estimators of identified effects and bounds on unidentified effects that are easy to calculate in a very general model. For example, this approach gives a simple solution to the important problem of identification of quantile treatment effects in panel data. The idea is based on Assumption 2, which states that, conditional on $X_{i},$ the distribution of unobservables does not vary over time. Therefore, conditional on $X_{i}$ where both $x^{b}$ and $x^{a}$ occur for some time periods, one can identify effects from the changes in $Y_{it}$ across those time periods. For the ATE, the identified conditional effects can be averaged to identify effects conditional on $X_{i}$ being in subsets where both $x^{b}$ and $x^{a}$ occur for some time period. This idea is a slight extension of Chamberlain (1982, pp. 10-17) to discrete regressors that are not binary. For the QTE the distribution functions can be averaged and inverted to identify corresponding quantile effects. This idea appears to be novel.
There is a simple approach to allowing for covariates. Suppose $ x=(x_{1},x_{2}),$ and one is interested in the effect of $x_{1}$ holding $ x_{2}$ fixed. Then one can take $x^{b}=(x_{1}^{b},x_{2})$ and $ x^{a}=(x_{1}^{a},x_{2}),$ so that the effect of changing from $x^{b}$ to $ x^{a}$ is then the effect of interest. Furthermore, one could average these effects over $x_{2}$ to identify an effect that is averaged over covariates. We explicitly allow for covariates in the semiparametric models given below. Because we are already attempting to cover so much ground here, we leave averaging over covariates in the nonparametric model to future work.
To describe identified effects and their estimators we will focus on the ATE and QTE conditional on both $x^{a}$ and $x^{b}$ appearing in $X_{i}$ for some time period. We could also consider effects conditional on smaller subsets of $X_{i}$ but postpone this until later in order to keep the exposition relatively simple. We need a little more notation to give a precise description. Let $1(X_{it}=x)$ denote the indicator function that is equal to one when $X_{it}=x$ and zero otherwise and let $T_{i}(x)= \sum_{t=1}^{T}1(X_{it}=x).$ Here we let the subscript $i$ denote a random variable that may depend on $X_{i}$ and $Y_{i}$. Let $ D_{i}=1(T_{i}(x^{a})>0)1(T_{i}(x^{b})>0)$ be the indicator for the event that $X_{i}$ includes both $x^{a}$ and $x^{b}$ for some time period. Define
This $\delta $ is the ATE for those individuals where both $x^{b}$ and $ x^{a} $ occur for some time period. This effect may be of interest in many settings. For example, when $Y_{it}$ is log earnings and $X_{it}\in \{0,1\}$ represents union status, $\delta $ would be the average effect of union status on earnings for those who changed union status over the time periods we observe. For a given number of time periods $T,$ this is all one could hope to identify nonparametrically. However, we may be interested in other effects too. We might be interested in union effects for those who ever changed union status at some time. \ This is $\delta $. \ Or we might even be interested in the effect for those who were ever in a union. Bounds for such an effect are described below.
A simple estimator of the conditional ATE $\delta $ is
Consistency of this estimator results from
see Lemma A5 of the Supplementary Material. Intuitively, this equation follows from time being randomly assigned, so that we can estimate the effect by comparing $Y_{it}$ where $X_{it}=x^{a}$ with $Y_{is}$ where $ X_{is}=x_{b}$.
Since $\bar{Y}_{i}(x^{a})-\bar{Y}_{i}(x^{b})$ is a difference of means it can be interpreted as a coefficient of $1(X_{it}=x^{a})$ in a regression of $ Y_{it}$ on that dummy and on $1(X_{it}=x^{a})+1(X_{it}=x^{b})$. Thus, $\hat{ \delta}$ is an average of least-squares estimates for each $i$ with $D_{i}=1$ . From this interpretation we see that $\hat{\delta}$ extends Chamberlain's (1982, p. 12) estimator to discrete regressors that are not binary. A consistent estimator of the asymptotic variance of $\sqrt{n}(\hat{\delta} -\delta )$ is $n^{-1}\sum_{i=1}^{n}\hat{\psi}_{i}^{2}$ where $\hat{\psi} _{i}=nD_{i}[\bar{Y}_{i}(x^{a})-\bar{Y}_{i}(x^{b})-\hat{\delta} ]/\sum_{i=1}^{n}D_{i}$. For brevity we leave the asymptotic theory to the Supplementary Material (see Theorem A6) and efficiency results to future work.
We can also identify and estimate a conditional QTE. Let $G(y,x|D_{i}=1)=\Pr (g_{0}(x,\alpha _{i},\varepsilon _{i1})\leq y|D_{i}=1)$ denote the CDF of $ g_{0}(x,\alpha _{i},\varepsilon _{i1})$ conditional on $D_{i}=1$. The QTE conditional on $D_{i}=1$ is
An estimator of this effect can be constructed using a CDF $\Phi (u)$ and a scalar bandwidth $h$. An estimator of $G(y,x|D_{i}=1)$ is given by
In this estimator the indicator function $1(Y_{it}<y)$ has been replaced by a smoothed approximation $\Phi (\frac{y-Y_{it}}{h})$, as suggested by Yu and Jones (1998) for estimating a conditional CDF. An estimator of $\delta _{\lambda }$ is then
Note here that we first average, then invert, and then difference. This estimator solves an important problem of estimating panel quantile effects and appears to be novel.
A consistent estimator of the asymptotic variance of $\sqrt{n}(\hat{\delta} _{\lambda }-\delta _{\lambda })$ is $n^{-1}\sum_{i=1}^{n}\hat{\psi}_{\lambda i}^{2}$ for
where $\hat{G}^{\prime }(y,x|D_{i}=1)=\partial \hat{G}(y,x|D_{i}=1)/\partial y$. Here the denominator terms are actually kernel density estimates. For this reason one might use different bandwidths $h$ in the numerator and denominator, with the denominator chosen to be appropriate for density estimation. Asymptotic theory for this estimator is given in the Supplementary Material (see Theorem A8). Alternatively, one could simply use the bootstrap to construct a confidence interval for $\hat{\delta}_{\lambda }.$
A helpful example is the binary regressor case where $X_{it}\in \{0,1\}.$ Here $X_{it}$ could be thought of as a treatment variable where $X_{it}=1$ for treated and $X_{it}=0$ for untreated. Let $Y_{it}(0)=g_{0}(0,\alpha _{i},\varepsilon _{it})$ and $Y_{it}(1)=g_{0}(1,\alpha _{i},\varepsilon _{it})$. Assumption 2 is equivalent to the assumption that the conditional distribution of $(Y_{it}(0),Y_{it}(1))$ given $X_{i}$ does not vary with $t$ . This is the key assumption that identifies treatment effects from time variation in treatment. In this context $\delta =E[Y_{it}(1)-Y_{it}(0)|D_{i}=1]$ is the ATE for individuals where both treatment and nontreatment occurs during the observation period. Similarly, $ \delta _{\lambda }$ is the difference between the $\lambda $ quantile of the distribution of $Y_{it}(1)$ and the $\lambda $ quantile for $Y_{it}(0)$ conditional on $D_{i}=1$. The ATE and QTE are not identified for those individuals that either receive treatment in every time period or receive no treatment in every time period.
In general the usual panel data within (linear fixed effects) estimator is not a consistent estimator of $\delta .$ This inconsistency results because the within estimator constrains the slope coefficient to be the same for each $i$ when the slope is actually varying with $i$. For simplicity we demonstrate this inconsistency in the binary $X_{it}$ example. The within estimator $\hat{\delta}_{w}$ is given by
Let $\sigma _{i}^{2}=(T-1)^{-1}\sum_{t=1}^{T}(X_{it}-\bar{X}_{i})^{2}$ be the sample variance over time of $X_{it}$.
Theorem 1: If Assumptions 1 and 2 are satisfied, $X_{it} \in \{0,1\}$, $ E[Y_{it}^{2}]<\infty ,$\ $(t=1,...,T)$, and $E[D_{i}\sigma _{i}^{2}]>0$, then $\delta =E[D_{i}\{\bar{Y}_{i}(1)-\bar{Y} _{i}(0)\}]/E[D_{i}]$ and
Note that the limit of the within estimator is a weighted average of individual, least-squares estimates $\bar{Y}_{i}(1)-\bar{Y}_{i}(0)$ from equation ((ref)). If $T\geq 4$ then the weights $\sigma _{i}^{2}$ vary over the positive $\sigma _{i}^{2}$ and so the limit $\delta _{w}$ of $\hat{ \delta}_{w}$ is not the identified conditional ATE $\delta $.
Theorem 1 is different than Yitzhaki (1996) and Angrist (1998), who gave weighted average interpretations of least squares in other, non-panel settings. Theorem 1 is also different from Hahn (2001), who found that $\hat{ \delta}_{w}$ consistently estimates the ATE. Hahn (2001) considered $T=2$ and assumed $X_{i}=(0,1)^{\prime }$. As noted by Hahn (2001), those conditions are quite special. Theorem 1 is also different from Wooldridge (2005), who showed that if $b_{i}=E[Y_{it}(1)-Y_{it}(0)|\alpha _{i}]$ is mean independent of $X_{it}-\bar{X}_{i}$ for each $t$ then linear fixed effects is a consistent estimator of $\delta $. The problem is that the mean-independence assumption is very strong when $X_{it}$ is discrete. For instance, if $T=2$, $X_{i2}-\bar{X}_{i}$ takes on the values $0$ when $ X_{i}=(1,1)$ or $(0,0)$, $-1/2$ when $X_{i}=(1,0)\,,$ and $1/2$ when $ X_{i}=(0,1)$. Thus mean independence of $b_{i}$ and $X_{i2}-\bar{X}_{i}$ actually implies that
This is quite close to independence of $b_{i}$ and $X_{i}$, which is not very interesting if we want to allow the treatment effect to vary with $ X_{i} $.
The conditional ATE and QTE estimators can easily be modified to accommodate the time effects of Assumption 4. The changes in $Y_{it}$ over time for fixed $X_{it}$ can be used to identify and estimate the time effects that can then be included in the estimation of the ATE and QTE. To describe this approach, let $\hat{m}_{t}=\sum_{i=1}^{n}1(X_{it}=X_{i1})Y_{it}/ \sum_{i=1}^{n}1(X_{it}=X_{i1})$ and
This $(\hat{\tau}_{t},\hat{s}_{t})^{\prime }$ is an instrumental variables estimator where the residual is $Y_{it}-\tau _{t}-s_{t}Y_{i1}$, the instruments are $(1,X_{i1})^{\prime }$, and the estimation is done on the subsample where $X_{it}=X_{i1}$. These estimators will be consistent and asymptotically normal as long as $Cov(X_{i1},Y_{i1}|X_{it}=X_{i1})\neq 0$ for each $t=2,...,T.$ One could also use other functions of $X_{i1}$ as instrumental variables to improve efficiency. We focus on just $X_{i1}$ as an instrument for simplicity. Graham and Powell (2011) use a similar approach to identify time effects in a linear model with continuous regressors.
The time effects are accounted for in ATE and QTE estimation by removing time location and scale effects from all periods when estimating the first period effect, and then putting the scale effects back for other periods. Note first that under Assumption 4 $\delta $ is the conditional ATE for the first time period. Let $\tilde{Y}_{it}=(Y_{it}-\hat{\mu}_{t})/\hat{s}_{t}$ be the $t^{th}$ period observation with estimated location and scale removed. Replacing $Y_{it}$ by $\tilde{Y}_{it}$ in the formula for $\hat{ \delta}$ gives
The conditional ATE for the $t^{th}$ time period is given by $s_{t}\delta $ and a time average by $(\sum_{t=1}^{T}s_{t}/T)\delta $ for $s_{1}=1$, analogously to equation ((ref)). These can be estimated by $\hat{s}_{t} \tilde{\delta}$ and $\bar{s}\tilde{\delta},$ respectively for $\bar{s} =\sum_{t=1}^{T}\hat{s}_{t}/T$ and $\hat s_{1}=1.$ These estimators will be consistent and asymptotically normal. Because of their multistage nature the bootstrap may provide the easiest approach to carrying out inference on these estimators, where one resamples from the empirical distribution of $ (Y_{i},X_{i}),$ $(i=1,...,n)$ to form confidence intervals for the true parameter. For brevity we omit explicit results.
An analogous approach can be followed to account for time effects in the QTE. The interpretation of $\delta _{\lambda }$ now becomes QTE for the first time period conditional on $D_{i}=1$. An estimator of $ G(y,x|D_{i}=1)=\Pr (g_{0}(x,\alpha _{i},\varepsilon _{i1})\leq y|D_{i}=1)$ that adjusts for time, location and scale is given by
Let $\tilde{q}_{\lambda }^{a}=\tilde{G}^{-1}(\lambda ,x^{a}|D_{i}=1)$ and $ \tilde{q}_{\lambda }^{b}=\tilde{G}^{-1}(\lambda ,x^{b}|D_{i}=1).$ Estimators for the conditional QTE for the first period, other periods, and a time average are given by $\tilde{\delta}_{\lambda }=\tilde{q}_{\lambda }^{a}- \tilde{q}_{\lambda }^{b}$, $\hat{s}_{t}\tilde{\delta}_{\lambda },(t=2,...,T), $ and $\bar{s}\tilde{\delta}_{\lambda }$, respectively. Here again the bootstrap provides a convenient method for inference. One could also use quantiles to estimate the time effects, but we avoid that for simplicity.
When $g_{0}(x,\alpha _{i},\varepsilon _{it})$ is bounded we can estimate bounds for the ASF and corresponding bounds for the ATE. For the QSF and QTE we can also estimate bounds without any restriction on $g_{0}$, using the fact that there are known upper and lower bounds for the indicator function $ 1(g_{0}(x,\alpha _{i},\varepsilon _{it})\leq y).$ The idea of the bounds is an extension of the estimation of identified effects discussed in the previous Section. Time homogeneity allows us to use time averages to estimate the identified parts of the ASF or QSF when $x$ is an element of $ X_{i},$ i.e. $X_{it}=x$ for some $t$, and apply the lower or upper bounds when $x$ does not appear in $X_{i}$.
We first describe bounds estimation for the ASF. These bounds depend on bounds on $g_{0}$ imposed in the following condition:
Assumption 6: $B_{\ell }\leq g_{0}(x,\alpha _{i},\varepsilon _{it})\leq B_{u}$ for constants $B_{\ell }$\ and $B_{u}$ and all $x.$
For example, in the binary-choice model, where $Y_{it}\in \{0,1\}$, upper and lower bounds are $B_{u}=1$ and $B_{\ell }=0$ respectively. We could allow $B_{\ell }$\ and $B_{u}$ to depend on $x$ and using that information could tighten the ATE bounds given below. To avoid further complication we do not allow this.
Let $T_{i}(x)$ and $\bar{Y}_{i}(x)$ be as in Section 3 and $\bar{P} (x)=\sum_{i=1}^{n}1(T_{i}(x)=0)/n$ be the sample frequency of $x$ not occurring in any time period. Estimated lower and upper bounds for $\mu (x)$ are
Here $\bar{Y}_{i}(x)$ estimates the identified part of the ASF, corresponding to $T_{i}(x)>0$, and the upper and lower bounds are applied for observations where $T_{i}(x)=0\,.$ Corresponding estimated lower and upper bounds for the ATE are $\hat{\Delta}_{\ell }=\hat{\mu}_{\ell }(x^{a})- \hat{\mu}_{u}(x^{b})$ and $\hat{\Delta}_{u}=\hat{\mu}_{u}(x^{a})-\hat{\mu} _{\ell }(x^{b}).$ The width of these estimated bounds is
For example, for binary choice with a binary regressor, where $B_{u}=1$ and $ B_{\ell }=0,$ the width of the estimated bounds for the ATE\ is $\bar{P}(0)+ \bar{P}(1),$ where $\bar{P}(0)$ and $\bar{P}(1)$ are the sample proportions of $X_{i}$ with $X_{it}=1$ for all $t$ and $X_{it}=0$ for all $t$, respectively
These estimators will be jointly asymptotically normal under i.i.d. $ (Y_{i},X_{i})$. The asymptotic variance can be estimated by $\hat{\Sigma} =\sum_{i=1}^{n}\hat{\Psi}_{i}\hat{\Psi}_{i}^{\prime }/n$, where
Confidence intervals for the identified set can then be formed using results of Chernozhukov, Hong, and Tamer (2007) or Beresteanu and Molinari (2008, pp. 779-781) on estimators of intervals where the upper and lower endpoints are jointly asymptotically normal.
Turning to the bounds for the QSF, lower and upper estimated bounds for the $ G(y,x)=\Pr (g_{0}(x,\alpha _{i},\varepsilon _{i1})\leq y)$ are $\hat{G} _{\ell }(y,x)=\sum_{i=1}^{n}\bar{G}_{i}(y,x)/n$ and $\hat{G}_{u}(y,x)=\hat{G} _{\ell }(y,x)+\bar{P}(x)$ respectively. The idea of these bounds is similar to the ASF, with a known lower bound of $0$ and upper bound of $1$ for $ 1(g_{0}(x,\alpha _{i},\varepsilon _{i1})\leq y)$. To obtain quantile bounds we need to invert these functions of $y$. For a strictly increasing function $G(y)$ with range contained in $[0,1]$ let
This is a function with domain $[0,1]$ and range equal to the extended real line that can be used to invert $\hat{G}_{u}(y,x)$ and $\hat{G}_{\ell }(y,x). $ Estimators of lower and upper bounds on the QSF are given by
Corresponding lower and upper bounds for the QTE are $\hat{\Delta}_{\lambda \ell }=\hat{q}_{\ell }^{a}-\hat{q}_{u}^{b}$ and $\hat{\Delta}_{\lambda u}= \hat{q}_{u}^{a}-\hat{q}_{\ell }^{b}$ where $\hat{q}_{\ell }^{a}=\hat{q} _{\ell }(\lambda ,x^{a}),$ $\hat{q}_{u}^{a}=\hat{q}_{u}(\lambda ,x^{a}),$ $ \hat{q}_{\ell }^{b}=\hat{q}_{\ell }(\lambda ,x^{b}),$ and $\hat{q}_{u}^{b}= \hat{q}_{u}(\lambda ,x^{b}).$ The width of these bounds depends on the shape of the empirical distribution of $Y_{it}$ and on $\bar{P}(x).$ The width of the bounds will be finite when
and otherwise they are infinitely wide.
The bounds will be joint asymptotically normal under the following regularity condition:
Assumption 7: $\Pr (g_{0}(x,\alpha _{i},\varepsilon _{i1})\leq y|X_{i})$ is twice continuously differentiable in $y$ with uniformly bounded derivatives and $G_{\ell }(y,x)=E[E[1(T_{i}(x)>0)|X_{i}]\Pr (g_{0}(x,\alpha _{i},\varepsilon _{i1})\leq y|X_{i})]$ is strictly increasing in $y$\ on the interior of its range for all $x$. Also $nh^{4}\longrightarrow 0$ and $nh^{2}\longrightarrow \infty $.
For $\lambda $ satisfying equation ((ref)) the asymptotic variance can be estimated by $\hat{\Sigma}_{\lambda }=\sum_{i=1}^{n}\hat{\Psi }_{\lambda i}\hat{\Psi}_{\lambda i}^{\prime }/n$, where
As in estimation of the conditional quantile effect, one might want to use different bandwidths for numerators and denominators, or just bootstrap to estimate the asymptotic variance.
Here is a result for both ATE and QTE bounds:
Theorem 2: Suppose that Assumptions 1, 2, and 5 are satisfied.\ If Assumption 6 is satisfied then there are $\Delta _{\ell },$\ $\Delta _{u},$\ and $\Sigma $ such that \textit{\ }
\ where $\Delta _{\ell }\leq \Delta \leq \Delta _{u}$, and these bounds are sharp. If Assumption 7 is satisfied then there are $\Delta _{\lambda \ell },$\ $\Delta _{\lambda u},$\ and $\Sigma _{\lambda }$ such that
\ where $\Delta _{\lambda \ell }\leq \Delta _{\lambda }\leq \Delta _{\lambda u}$. If $G_{\ell }(y,x)$ is also everywhere strictly increasing in $y$ then these bounds are sharp.
The sharpness conclusion of Theorem 2 for the ATE depends on being able to let $g_{0}(x,\alpha _{i},\varepsilon _{it})$ take any value between $B_{\ell }$ and $B_{u}.$ That is not possible for binary choice, where the outcome is restricted to zero or one. Nevertheless the bounds can still be shown to be sharp.
Similarly to the treatment-effects literature, we may be interested in the ATE or QTE, conditional on $X_{i}\in S$ for some set $S$. For example, if $ X_{it}\in \{0,1\}$ represents treatment then we might be interested in the effect of treatment conditional on ever treated, i.e. conditional on $ X_{i}\neq (0,...,0)^{\prime }$. Tighter bounds for such effects can be formed and in some cases the effects may be identified. These bounds can be estimated by replacing $1(X_{it}=x)$ by $1(X_{i}\in S)1(X_{it}=x)$ in the definition of $\bar{Y}_{i}(x)$ and $\bar{G}_{i}(y,x)$, $1(T_{i}(x)=0)$ by $ 1(X_{i}\in S)1(T_{i}(x)=0)$ in the definition of $\bar{P}(x)$, and dividing through by $\sum_{i=1}^{n}1(X_{i}\in S)/n.$ If $1(X_{i}\in S)\leq D_{i}$ for $D_{i}$ from Section 3 the corresponding effects will be identified, and the upper and lower estimated bounds will be identical.
Time effects can easily be allowed for in quantile-effect bounds by adapting the approach used earlier. It is not clear that allowing for time effects in that way makes sense for bounds on the ATE, e.g. for binary choice models where the support of $Y_{it}$ is fixed. Therefore we focus just on time effects in quantile bounds. For QTE bounds we can replace $Y_{it}$ by $ \tilde{Y}_{it}=(Y_{it}-\hat{\mu}_{t})/\hat{s}_{t}$ in the formula for $\hat{G }_{\ell }(y,x)$ given above, and interpret $\hat{\Delta}_{\lambda \ell }$ and $\hat{\Delta}_{\lambda u}$ as estimators of the first period bounds. Estimators of $t^{th}$ period lower and upper bounds for the QTE are then given by $\hat{s}_{t}\hat{\Delta}_{\lambda \ell }$ and $\hat{s}_{t}\hat{ \Delta}_{\lambda u}$ respectively. Estimators of time average bounds are $ \bar{s}\hat{\Delta}_{\lambda \ell }$ and $\bar{s}\hat{\Delta}_{\lambda u},$ where $\bar{s}=\sum_{t=1}^{T}\hat{s}_{t}/T.$ These upper and lower bounds will be joint asymptotically normal, and their asymptotic variance can be estimated by the bootstrap.
Analysis of the dynamic model is more challenging than that of the static one. In the dynamic model of Assumption 3 only the first-period regressor is common to the conditioning sets for each time period. Consequently location and scale time effects are not identified, because the conditioning set is different for every time period. For this reason we do not consider time effects in the nonparametric dynamic model. Also, the identification and bounds analysis is limited to objects that are conditional on the first period or are unconditional. For example, we cannot identify or bound the ATE conditional on $X_{it}$ changing over time because that event involves information about all time periods. We can bound unconditional objects and ones that are conditional on just $X_{i1}.$ These bounds are simple and novel, for example in providing partial-identification results for the average effect of state dependence with heterogeneity in both location and slope when $Y_{it}$ is binary and $X_{it}=Y_{it-1}$.
The model with a binary, lagged dependent variable has $ Y_{it}=g_{0}(Y_{i,t-1},\alpha _{i},\varepsilon _{it})$, and under Assumption 3,
where $F(\varepsilon |\alpha _{i},Y_{i0})$ denotes the conditional CDF of $ \varepsilon _{it}$ given $\alpha _{i}$ and $Y_{i0}.$ Here $\Pr (Y_{it}=1|Y_{i,t-1},\alpha _{i},Y_{i0})$ does not vary with $t$, and the model places no other restrictions on $\Pr (Y_{it}=1|Y_{i,t-1},\alpha _{i},Y_{i0})$. Conditioning on $Y_{i0}$ is present to account correctly for the initial condition, as in Honore and Tamer (2006) and Browning and Carro (2007, 2009). The probabilities can be distributed across individuals in any way at all through the individual effect $\alpha _{i}$. That is we can think of the four conditional probabilities,
as having an unrestricted distribution. Here the ATE is
This object quantifies the effect of state dependence in the presence of individual heterogeneity, an important problem posed by Feller (1943) and Heckman (1981). The dynamic bounds here provide a simple, estimable, identified set for this object. This model is considered by Browning and Carro (2007, 2009), who derive properties of various estimators and restrictions on $\alpha _{i}$ that lead to identification. We give nonparametric bounds.
A partition of $X_{i}$ values that preserves the dynamic structure of Assumption 3 is used to obtain bounds for the ASF and QSF. For each $x$ we partition $X_{i}$ into realizations where the first occurrence of $x$ is at time $t$ and the set where $x$ never occurs. This partition is given by $\{ \mathcal{\bar{X}}(x),\mathcal{X}_{1}(x),...,\mathcal{X}_{T}(x)\}$ where
Define $\hat{Y}_{i}(x)=\sum_{t=1}^{T}1(X_{i}\in \mathcal{X}_{t}(x))Y_{it}$, which picks out the $Y_{it}$ for the time period where $x$ first occurs. Estimated lower and upper ASF bounds are
Corresponding lower and upper bounds for $\Delta $ are $\hat{\Delta}_{\ell }= \hat{\mu}_{\ell }(x^{a})-\hat{\mu}_{u}(x^{b})$ and $\hat{\Delta}_{u}=\hat{\mu }_{u}(x^{a})-\hat{\mu}_{\ell }(x^{b}).$ A joint asymptotic-variance estimator $\hat{\Sigma}$ can be constructed exactly as for the static case with $\hat{Y}_{i}(x)$ replacing $\bar{Y}_{i}(x).$
It is interesting to note that the width $\bar{P}(x)(B_{u}-B_{\ell })$ of the estimated ASF bounds is the same for the dynamic and static models. Because the static model is a special case of the dynamic one we conjecture that the bounds for the dynamic model are sharp like the bounds for the static one, but have not yet been able to show this.
To construct estimated lower and upper bounds for the CDF of $g_{0}(x,\alpha _{i},\varepsilon _{it})$ let $\hat{G}_{i}(y,x)=\sum_{t=1}^{T}1(X_{i}\in \mathcal{X}_{t}(x))\Phi (\frac{y-Y_{it}}{h}).$ The estimated CDF bounds are
Estimated lower and upper bounds for the QSF are then given by
Corresponding lower and upper bounds for the QTE are $\hat{\Delta}_{\lambda \ell }=\hat{q}_{\ell }(\lambda ,x^{a})-\hat{q}_{u}(\lambda ,x^{b})$ and $ \hat{\Delta}_{\lambda u}=\hat{q}_{u}(\lambda ,x^{a})-\hat{q}_{\ell }(\lambda ,x^{b}).$ A joint asymptotic variance estimator $\hat{\Sigma}_{\lambda }$ can be constructed just as for the static case with $\hat{G}_{i}(y,x)$ replacing $\bar{G}_{i}(y,x).$
Theorem 3: Suppose that Assumptions 1, 3, and 5 are satisfied.\ If Assumption 6 is satisfied then there are $\Delta _{\ell },$\ $\Delta _{u},$\ and $\Sigma $ such that
\ where $\Delta _{\ell }\leq \Delta \leq \Delta _{u}$. Also if Assumption 7 is satisfied with $X_{i1}$ replacing $X_i$ then there are $\Delta _{\lambda \ell },$ \ $\Delta _{\lambda u},$\ and $\Sigma _{\lambda }$ \textit{ such that}
\ where $\Delta _{\lambda \ell }\leq \Delta _{\lambda }\leq \Delta _{\lambda u}$.
Similarly to the static model we may be interested in effects conditional on $X_{i1}\in S_{1}$ for some set $S_{1}$. For example, if $X_{it}\in \{0,1\}$ represents treatment then we might be interested in the effect of treatment conditional on being treated in the first period, i.e. conditional on $ X_{i1}=1$. Tighter bounds for such effects can be estimated by replacing $ 1(X_{i}\in \mathcal{X}_{t}(x))$ by $1(X_{i1}\in S_{1})1(X_{i}\in \mathcal{X} _{t}(x))$ in the definition of $\hat{Y}_{i}(x)$ and $\hat{G}_{i}(y,x)$, $ 1(T_{i}(x)=0)$ by $1(X_{i1}\in S_{1})1(T_{i}(x)=0)$ in the definition of $ \bar{P}(x)$, and dividing through by $\sum_{i=1}^{n}1(X_{i1}\in S_{1})/n.$
In the binary, lagged-dependent-variable example we have $B_{\ell }=0$ and $ B_{u}=1$, so the bounds on the ATE are
Here $\bar{P}(1)+\bar{P}(0)$ estimates the width of the bounds, providing a very simple measure of the severity of the problem of identifying state dependence in the presence of heterogeneity. The bounds will tend to be wide in short panels but more informative in long ones.
Figure 1 shows the width of corresponding population bounds in a numerical example based on a dynamic probit model where
We consider different DGPs indexed by $\beta ^{\ast }\in \lbrack -2,2]$ and compute the width of the bounds for $T\in \{2,4,8,16,32,64\}$. The width is asymmetric with respect to $\beta ^{\ast }=0$ because $\Pr (X_{i}=(1,...,1)^{\prime })$ grows with $\beta ^{\ast }$, whereas $\Pr (X_{i}=(0,...,0)^{\prime })$ does not depend on $\beta ^{\ast }$. The width growing with $\beta ^{\ast }$ may therefore be explained by having fewer switches of $ Y_{it}$ between one and zero when $\beta ^{\ast }$ is larger. It is presumably the changes that help identify the ATE. We find that the bounds can be substantially wide for high values of $\beta ^{\ast }$ even for large $T$, consistent with the width of the nonparametric bounds shrinking only at rate $1/T,$ as shown in the next Section. Semiparametric bounds for this model that impose the constancy of $\beta ^{\ast }$ across individuals, will shrink much faster at $T$ grows, as shown in Section 7.
Increasing $T$ improves identification, shrinking the estimated and population-identified sets for the objects of interest. The rate at which the identified set shrinks quantifies this improvement. Here we give rates for the ASF and, for brevity, leave the quantile results to the Supplementary Material.
The width of the population bounds for the ASF is $(B_{u}-B_{\ell })\mathcal{ \bar{P}}(x)$ where
Thus, the rate at which the identified set shrinks, that we will refer to as the identification rate, is the same as the rate at which $\mathcal{\bar{P}} (x)$ shrinks. Factors that determine this rate can be seen when $X_{it}$ is i.i.d. conditional on $\alpha _{i}$. In that case
The rate at which $\mathcal{\bar{P}}(x)$ goes to zero will be determined by how much probability mass of $\Pr (X_{it}\neq x|\alpha _{i})$ is close to one. If $ \Pr (X_{it}\neq x|\alpha _{i})=1$ with positive probability then $\mathcal{ \bar{P}}(x)$ does not go to zero. This corresponds to nonidentification of the ASF, where $x$ does not occur for some individuals as indexed by $\alpha _{i}$ (see Theorem A11 of the Supplementary Material). On the other hand, if $ \ \Pr (X_{it}\neq x|\alpha _{i})$ is bounded away from one then the identified set will shrink exponentially quickly, since $\Pr (X_{it}\neq x|\alpha _{i})^{T}\leq (1-\varepsilon )^{T}$ for some $\varepsilon >0$. In between the nonidentified and exponential rate cases there are a range of rates depending on how much of the distribution of $\Pr (X_{it}\neq x|\alpha _{i})$ is close to $1$. The following result shows the range of rates.
Theorem 4:\ Suppose that Assumptions 1, 3, 5, and 6 are satisfied and $(X_{i1},X_{i2},...)$\ is stationary and Markov of order $J$\ conditional on $\alpha _{i}$. If for some $ \varepsilon >0$,\ $\Pr (X_{it}=x|X_{i,t-1},...,X_{i,t-J},\alpha _{i})\geq \varepsilon $\textit{\ a.s. then }$\mu _{u}(x)-\mu _{\ell }(x)\leq (B_{u}-B_{\ell })(1-\varepsilon )^{T-J}.$\textit{\ If }$X_{it}$\textit{\ is i.i.d. conditional on }$\alpha _{i},$\textit{\ }$\Pr (X_{it}\neq x|\alpha _{i})$\textit{\ is continuously distributed with pdf }$f_{P}(p),$\textit{\ and\ }
then $\mu _{u}(x)-\mu _{\ell }(x)=O(T^{-v}).$
The upper bound on the rate at which the pdf $f_{P}(p)$ of $\Pr (X_{it}\neq x|\alpha _{i})$ grows or converges to zero as $p\longrightarrow 1$ provides an upper bound on the rate at which the identified set shrinks. For example, if $v=1$ so that $f_{P}(p)$ is bounded as $p\longrightarrow 1,$ then the identified set shrinks at rate $1/T.$ All of the rates implied by this result are slower than the exponential rate, reflecting how having $\Pr (X_{it}\neq x|\alpha _{i})$ close to $1$ affects the rate. Also, $\gamma $ has no effect on the convergence rate because that rate is determined by closeness of $\Pr (X_{it}\neq x|\alpha _{i})$ to $1$, and not to $0$.
The dynamic, binary-choice model is an example where more explicit conditions can be given. Suppose $Y_{it}=1(\alpha _{i1}+(\alpha _{i2}-\alpha _{i1})Y_{i,t-1}\geq \varepsilon _{it})$ and $\varepsilon _{it}$ is i.i.d. and independent of $\alpha _{i}=(\alpha _{i1}, \alpha _{i2})$ with CDF $H(\varepsilon )$. Here $\Pr (Y_{it}=1|Y_{i,t-1}=0,\alpha _{i})=H(\alpha _{i1})$ and $\Pr (Y_{it}=1|Y_{i,t-1}=1,\alpha _{i})=H(\alpha _{i2}).$ Unbounded $\alpha _{i}$ and bounded $\varepsilon _{it}$ will correspond to the unidentified case. Bounded $\alpha _{i}$ and unbounded $\varepsilon _{it}$ lead to an exponential convergence rate. The following result covers the in-between case. Let $f_{\varepsilon }(\varepsilon ),$ $f_{\alpha _{1}}(\alpha )$, and $ f_{\alpha _{2}}(\alpha )$ denote the pdfs of $\varepsilon _{it},$ $\alpha _{i1},$ and $\alpha _{i2}$ respectively, all are assumed to be continuously distributed.
Theorem 5:\ If $Y_{it}=1(\alpha _{i1}+(\alpha _{i2}-\alpha _{i1})Y_{i,t-1}\geq \varepsilon _{it}),$ where $\varepsilon _{it},(t=1,...,T)$ is i.i.d. and independent of $(\alpha _{i1},\alpha _{i2})$\ and there is $v,C>0$ such that for all $\varepsilon $
then $\Delta _{u}-\Delta _{\ell }=O(T^{-v}).$
Here we see that the identification rate in the nonparametric dynamic model is related to the tail thickness of the distribution of $\alpha _{i1}$ and $ \alpha _{i2}$ relative to the distribution of $\varepsilon _{it}$. The thinner the tail of $f_{\varepsilon }(\varepsilon )$ relative to the tails of $f_{\alpha _{1}}(\alpha _{1})$ and $f_{\alpha _{2}}(\alpha _{2})$ the smaller $v$ will need to be to satisfy the inequality in Theorem 5 and the slower the identification rate will be. In this way the identification rate is slower the less strong the signal provided by $\varepsilon _{it}$ relative to the individual effects. Here there is no $\gamma $ present because both left and right tails matter, in order to bound the rate for the ATE, and not just for the ASF at a particular $x$.
For a specific example consider $\alpha _{i1}$ and $\alpha _{i2}$ as $ N(0,\sigma _{\alpha }^{2})$ and $\varepsilon _{it}$ as $N(0,\sigma _{\varepsilon }^{2})$ where $\sigma _{\varepsilon }^{2}\leq \sigma _{\alpha }^{2}.$ Then for constants $C_{1},$ $C_{2},$ and $v=\sigma _{\varepsilon }^{2}/\sigma _{\alpha }^{2}$ we have $f_{\alpha _{j}}(\varepsilon )=C_{1}[f_{\varepsilon }(\varepsilon )]^{v}.$ Also, as is well known for the Gaussian distribution, $f_{\varepsilon }(\varepsilon )\geq C_{2}F_{\varepsilon }(\varepsilon )[1-F_{\varepsilon }(\varepsilon )],$ where $F_{\varepsilon }(\varepsilon )$ denotes the CDF of $\varepsilon $. It follows by $v\leq 1$ that
Thus equation ((ref)) is satisfied with $v=\sigma _{\varepsilon }^{2}/\sigma _{\alpha }^{2}$ so that
Hence the width of the bounds shrinks at a rate no larger than $T^{-1}$ and the rate is slower the smaller $\sigma _{\varepsilon }^{2}/\sigma _{\alpha }^{2}$ is. It can also be shown that convergence is faster than $T^{-1}$ when $\sigma _{\varepsilon }^{2}>\sigma _{\alpha }^{2}$ and increases with $ \sigma _{\varepsilon }^{2}/\sigma _{\alpha }^{2}$. Thus we see that the stronger the signal provided by $\varepsilon $ relative to that provided by $ \alpha ,$ in the sense that the higher $\sigma _{\varepsilon }^{2}$ is relative to $\sigma _{\alpha }^{2}$, the faster will be the identification rate.
One can obtain analogous results in a static model. If $X_{it}=1(\alpha _{i}\geq \eta _{it})$ is a binary regressor where $\eta _{it}$ is i.i.d. over time then the identification rate will be $T^{-v}$ when the inequality in Theorem 5 is satisfied with the pdf $f_{\eta }(\eta )$ of $\eta _{it}$ replacing the pdf $f_{\varepsilon }(\varepsilon ).$ If $\alpha _{i}$ and $ \eta _{it}$ are distributed as $N(0,\sigma _{\alpha }^{2})$ and $\eta _{it}$ as $N(0,\sigma _{\eta }^{2})$ respectively with $\sigma _{\eta }^{2}\leq \sigma _{\alpha }^{2}$, then the identified set shrinks at rate $T^{-\sigma _{\eta }^{2}/\sigma _{\alpha }^{2}}$. For brevity we omit the details.
The nonparametric bounds are informative but may be quite wide for small $T$ . They can be tightened by imposing additional structure on the model. One way to do this is to specify a parametric model for the conditional distribution of $Y_{i}$ given values for $(X_{i},\alpha _{i}).$ We focus here on multinomial choice models. In those models $Y_{i}$ is one of a finite number of outcomes, denoted here by $\{Y^{1},...,Y^{J}\}.$ The parametric part of the model are the known conditional probabilities $ \mathcal{L}_{j}^{k}(\alpha ,\beta )$ of $Y_{i}=Y^{j}$\ given $ \alpha _{i}$ and $X_{i}\in \mathcal{X}^{k},(k=1,...,K),$ where $\beta $ is a parameter vector with true value $\beta ^{\ast }$, and $\mathcal{X}^{k}$ is the set of $X_{i}$ values being conditioned on. Formulating the model in this way allows for $X_{i}$ that are lagged dependent variables. The nonparametric part of the model will be the unknown CDF's $F_{k}^{\ast }(\alpha ),(k=1,...,K)$ of $\alpha _{i}$ conditional on $X_{i}$ in each $ \mathcal{X}^{k}.$ The model then satisfies
Assumption 8: $\Pr (Y_{i}=Y^{j}|X_{i}\in \mathcal{X}^{k})=\int \mathcal{L}_{j}^{k}(\alpha ,\beta ^{\ast })dF_{k}^{\ast }(\alpha ),(j=1,...,J;k=1,...,K)$.
Some examples may be helpful. An important example is a binary choice model where $Y_{it}\in \{0,1\}$, $\alpha $ is a scalar location individual effect, $\Pr (Y_{it}=1|X_{i},\alpha _{i},\beta ^{\ast })=H(X_{it}^{\prime }\beta ^{\ast }+\alpha _{i})$ for a CDF $H(\varepsilon ),$ and $Y_{i1},...,Y_{iT}$ are mutually independent conditional on $X_{i}$ and $\alpha _{i}$. In this case we would let $\mathcal{X}^{k}$ be a singleton given by the $k^{th}$ value $X^{k}$ in the finite support of $X_{i}$ and
Time effects can be included in this model by specifying that some components of $X_{t}^{k}$ only depend on $t.$ This model can also be generalized to allow for some slopes to vary across individuals by specifying that
This model allows the coefficients of $X_{t2}^{k}$ to vary with individuals, which will include a location effect when some element of $X_{t2}^{k}$ does not vary with $t$ or $k.$
This set up also allows for dynamic models. For example, consider a binary choice model with a lagged dependent variable where $\Pr (Y_{it}=1|Y_{i,t-1},...,Y_{i0},\alpha _{i},\beta ^{\ast })=H(Y_{i,t-1}\beta ^{\ast }+\alpha _{i}).$ Here $X_{i}=(Y_{i,T-1},...,Y_{i0})$ and we take $ K=2, $ with $\mathcal{X}^{k}=\{X_{i}:X_{i1}=Y_{i0}=k-1\}.$ The parametric part of the model is
This model could be generalized to allow individual specific coefficients for the dynamic effect, time effects, and other covariates, including the model of Browning and Carro (2009). For brevity we omit this generalization.
The ATE and its bounds can be decomposed into a weighted average of conditional ATE and corresponding bounds, weighted by the identified $\Pr (X_{i}\in \mathcal{X}^{k})$. The semiparametric model may restrict the conditional bounds so we focus first on them. We will assume that a conditional ATE takes the form
where $\Delta (\alpha ,\beta )$ denotes a treatment effect conditional on $ \alpha $. For example, in the model of equation ((ref)) we could take $\Delta (\alpha ,\beta )=H(x^{a\prime }\beta +\alpha )-H(x^{b\prime }\beta +\alpha ),$ in which case
is the ATE conditional on $X_{i}=X^{k}$. One could also consider the ASF conditional on $X_{i}=X^{k},$ that would be $\int H(x^{\prime }\beta ^{\ast }+\alpha )dF_{k}^{\ast }(\alpha )$ in this example.
Neither $\Delta ^{k}$ nor $\beta ^{\ast }$ need be identified. Instead, there may be sets of $\beta ^{\ast }$ and ATE values that are consistent with the distribution of the data. To describe the identified sets let $ \mathcal{P}=(\mathcal{P}_{1}^{1},...,\mathcal{P}_{J}^{1},...,\mathcal{P} _{J}^{K})^{\prime }$ denote the vector of population choice probabilities with $\mathcal{P}_{j}^{k}=\Pr (Y_{i}=Y^{j}|X_{i}\in \mathcal{X}^{k})$ and
where $\mathcal{F}_{k}(\beta ,\mathcal{P})$ may be empty. The identified set for $\beta ^{\ast }$ is
That is, $B$ is the set where there exist individual effect distributions such that integrals of model probabilities equal population choice probabilities. Sharp upper and lower bounds $\Delta _{u}^{k}$ and $\Delta _{\ell }^{k}$ for $\Delta ^{k}$ are given by
This characterization of bounds for the ATE extends that of Honore and Tamer (2006) from a finite dimensional $F_{k}$, where $\alpha $ is restricted to a known fixed grid, to infinite-dimensional $F_{k}$ where any distribution for $\alpha $ is allowed.
For purposes of comparison with the nonparametric results we consider models without trends, where the semiparametric models in equations ((ref) ) and ((ref)) are nested in the nonparametric static or dynamic model. In those models $\Delta ^{k}$ will be identified if it is also identified in the nonparametric model. In the static case $\Delta ^{k}$ is nonparametrically identified if $X_{t}^{k}$ takes on the values $x^{b}$ and $ x^{a}$ for some time periods$.$ This follows similarly to the identification of the conditional effect $\delta $ in Section 3. Therefore, in static models obtaining a smaller identified set by imposing the restrictions of a semiparametric model is limited to those $\Delta ^{k}$ where at least one of $x^{b}$ or $x^{a}$ does not appear in any time period. In what follows we focus on these $\Delta ^{k}$.
When slopes vary across individuals the semiparametric bounds may be no tighter than the nonparametric ones. To illustrate consider a binary-choice model with a single binary regressor $X_{it},$ where $Y_{it}=1((\alpha _{i2}-\alpha _{i1})X_{it}+\alpha _{i1}>\varepsilon _{it}),$ $\varepsilon _{it}$ is independent of $(X_{i},\alpha _{i2},\alpha _{i1}),$ and $ \varepsilon _{it}$ has known CDF $H(\varepsilon )$ that is strictly increasing on the entire real line. The joint distribution of $H(\alpha _{i1})$ and $H(\alpha _{i2})$ conditional on $X_{i}=X^{k}$ is entirely unrestricted. Therefore when $X^{k}=(0,...,0)^{\prime }$ the fact that $ E[H(\alpha _{i1})|X_{i}=X^{k}]=E[Y_{it}|X_{i}=X^{k}]$ for every every $t$, and so is identified gives no information about $E[H(\alpha _{i2})|X_{i}=X^{k}].$ Thus, $E[H(\alpha _{i2})|X_{i}=X^{k}]$ can be anything in the unit interval. \ Therefore, the width of the bound for $\Delta ^{k}=E[H(\alpha _{i2})-H(\alpha _{i1})|X_{i}=X^{k}]$ will be equal to the width in the nonparametric case, $\Delta _{u}^{k}-\Delta _{\ell }^{k}=1$. More generally, in the panel binary choice model of equation ((ref)), when there are no time effects, every coefficient of $X_{it}$ varies across individuals, and $X_{it}$ is fully saturated (e.g. is a complete set of dummies, one for every possible value of $X_{it}$), the semiparametric bounds will equal the nonparametric ones.
In the binary-regressor case the width of the overall bound on the ATE is given by
where we assume $X^{1}=(0,...,0)^{\prime }$ and\ $ X^{2}=(1,...,1)^{\prime }$. The semiparametric bounds will be smaller than the nonparametric bounds if and only if $\Delta _{u}^{1}-\Delta _{\ell }^{1}$ or $\Delta _{u}^{2}-\Delta _{\ell }^{2}$ are smaller than the nonparametric values of $1.$ This decomposition also shows that the semiparametric identification rate will be determined by the nonparametric rate, which governs how fast $\mathcal{\bar{P}}(0)$ and $\mathcal{\bar{P}}(1)$ shrink, and the rate that the conditional bounds converge. When the slope does not vary across individuals it turns out that the conditional bounds can converge very rapidly. The following result shows this in static and dynamic, binary-choice logit models with binary regressors.
Theorem 6: Suppose that $H(v)=e^{v}/(1+e^{v}),$ $\Delta (\beta ,\alpha )=H(\beta +\alpha )-H(\alpha ),$\ and either equation ((ref)) is satisfied with, $X_{it}\in \{0,1\}$, and $ X^{1}=(0,...,0)^{\prime }$ and $X^{2}=(1,...,1)^{\prime },$ or equation ((ref)) is satisfied with $k\in \{1,2\}$. \textit{Then there are }$C>0$ \textit{and }$1>\varepsilon >0$\textit{\ such that}
This fast rate occurs because $T$ conditional moments of a one-to-one transformation of $\alpha _{i}$ are identified from probabilities of various $Y$ values, and these moments lead to a fast approximation of the conditional ATE. For example, $\Pr (Y_{i}=(1,...,1)^{\prime }|X_{i}=X^{1})=E[H(\alpha _{i})^{T}|X_{i}=X^{1}]$, and other conditional moments of $H(\alpha _{i})$ can be similarly identified. For the logit $ H(\alpha )$, identification of these moments leads to fast approximation of $ \Delta ^{1}=E[H(\beta ^{\ast }+\alpha _{i})-H(\alpha _{i})|X_{i}=X^{1}]$ and hence to fast shrinkage of the conditional bound.
From equation ((ref)) we see that the semiparametric identification rate in this example will be at least exponential, and may be even faster, depending on the nonparametric rate. This result illustrates how imposing a single, additive individual effect can speed up the identification rate. We expect that this type of improvement will extend beyond the logit model with binary regressors.
In this section we discuss computation of population bounds, give examples, and present theoretical results. A challenge for computation and for estimation is the dimensionality of the unknown parameters and the nonlinearity of the probabilities in those parameters. A useful feature of multinomial panel models is that they are finite dimensional, in spite of the presence of distributions. The following lemma shows that one only need consider discrete distributions with $J$ unknown support points in the specification of the likelihood and the bounds for the ATE. Let $\Upsilon $ denote the set of possible values for the individual effect and $\mathbb{B}$ the set of parameters for $\beta .$
Lemma 7: If Assumptions 5 and 8 are satisfied and $ \mathcal{L}_{j}^{k}\left( \alpha ,\beta \right) $ is a measurable function of $\alpha $ for each $\beta \in \mathbb{B},$ then for each $\beta $\ and every CDF $F_{k}$\textit{\ on $ \Upsilon $ there is a discrete distribution }$F_{k}^{J}$\textit{\ with no more than }$J$\textit{\ support points such that }$\int \mathcal{L} _{j}^{k}\left( \alpha ,\beta \right) dF_{k}^{J}(\alpha )=\int \mathcal{L} _{j}^{k}(\alpha ,\beta )dF_{k}(\alpha )$\textit{\ }$(j=1,...,J).$ \textit{ If, in addition, }$\Delta (\alpha ,\beta )$ \textit{is bounded for each }$ \beta $ \textit{then }$\Delta _{u}^{k}$ \textit{and }$\Delta _{\ell }^{k}$ \textit{\ are not affected by restricting attention to }$F_{k}\in \mathcal{F} _{k}(\beta )$\textit{\ that are discrete with no more than }$J$ \textit{ support points.}
Thus, no matter what the dimension of $\alpha $ is, the multinomial panel model is finite dimensional, with the number of parameters given by $\dim (\beta )+(2J-1)^{K}.$ Another implication of this result is that the distribution of the individual effect is generally not identified in multinomial models. For example, if the true distribution $F_{k}^{\ast }$ were continuous then Lemma 7 would imply that there is a discrete distribution that gives exactly the same likelihood. The proof of this result is similar to Lindsay's (1983) result that the maximum likelihood estimator of a mixture model has a finite support. It is interesting that the model takes a discrete mixture form, although the finite-dimensional nature of the model is expected because the data have finite support.
Although the individual-effect distribution can be taken to be finite dimensional, the dimension can be large, and the probabilities depend nonlinearly on the support points for the individual effect. We overcome this challenge by using an approximation with a fixed but large number of support points for the individual effects. This approximation makes approximate probabilities and the ATE linear in parameters, simplifying computation. Honore and Tamer (2006) used a similar approach, but assumed that the true distribution of individual effects had known support points. We explicitly allow for approximation of unknown support points.
To describe how the approximation can be used to calculate the identified set, let $M$ denote a number of support points for the individual effect and $\Upsilon _{M}\mathcal{=}$$(\bar{\alpha}_{1M},...,\bar{\alpha} _{MM})^{\prime }$ be a grid of fixed values for the individual effect. Also let $\pi =(\pi ^{1\prime },...,\pi ^{K\prime })^{\prime }$ denote a $ MK\times 1$ vector of possible probabilities, with each $\pi ^{k}$ an element of the $M$ dimensional unit simplex $\mathcal{S}_{M}$. Approximate model probabilities are
Consider the function
where $w_{j}^{k}$ are positive weights, such as the chi-square ones $ \mathcal{P}^{k}/\mathcal{P}_{j}^{k},$ for $\mathcal{P}^{k}=\Pr (X_{i}\in \mathcal{X}^{k})$, and $\lambda _{M}>0$ is a penalty multiplier that controls the impact of the penalty term $\lambda _{M}\pi ^{\prime }\pi $. This term is present to help regularize the objective function and ensures a nonsingular Hessian matrix. Let $\tilde{T}_{\lambda }(\beta ,M)=\min_{\pi \in \mathcal{S}_{M}^{K}}T_{\lambda }(\beta ,\pi ,M)$ and let $\epsilon _{M}>0 $ be a positive scalar. We approximate the identified set for $\beta $ by
The use of $\epsilon _{M}$ here in allowing a range of values of the objective function is analogous to Manski and Tamer's (2002) estimation method. A positive $\epsilon _{M}$ ensures that the set sequence $ (B(M))_{M=1}^{\infty }$ is lower hemi-continuous and that $B(M)$ need not be smaller than the identified set, even though the individual effect distributions are restricted by fixing their support points for each $M.$
We calculate the identified set by letting $M$ grow and $\lambda _{M}$ and $ \epsilon _{M}$ shrink until there is little change in $B(M)$. Calculation of $\tilde{T}_{\lambda }(\beta ,M)$ is straightforward because it is the minimum of a quadratic function. In practice we have found that $B(M)$ changes little as $M$ increases even when $M$ is quite small. As $M$ grows and $\epsilon _{M}$ shrinks the set $B(M)$ will converge to the identified set under conditions given below.
For the ATE bounds, note
is the set of possible conditional ATE (given $X \in \mathcal{X}^{k})$ that are consistent with $\tilde{T}_{\lambda }(\beta ,M)\leq \epsilon _{M}$. Approximate lower and upper bounds are
As $M$ grows and $\epsilon _{M}$ shrinks these bounds will converge to $ \Delta _{\ell }^{k}$ and $\Delta _{u}^{k}$ respectively, under conditions given below.\
Computation of these ATE bounds is challenging because it requires searching over a large dimensional set of possible $\pi $. In practice we start with a smaller set of probabilities and then try others. Specifically, let $\tilde{ \pi}(\beta )\in \arg \min_{\pi \in \mathcal{S}_{M}^{K}}T_{\lambda }(\beta ,\pi ,M),$ $\tilde{S}^{k}(\beta )=\{\pi ^{k}:P_{j}^{k}(\beta ,\pi ,M)=$ $ P_{j}^{k}(\beta ,\tilde{\pi}(\beta ),M),$ $j=1,...,J\},$ and
For each $\beta $ these bounds are easy to calculate by linear programming. We have done so and then checked to see if other values $\pi $ violate these bounds. We have not found this to be so for values of $M$ that we use to compute $\beta $. We conjecture that these bounds also converge to the population bounds as $M\longrightarrow \infty $ although we have not yet been able to prove this (because we have not been able to show that the ATE bounds are continuous in the true probabilities).
We carry out some numerical calculations for the probit model where
We consider different DGPs indexed by $\beta ^{\ast }\in \lbrack -2,2]$ and $ T\in \{2,3\}$. Figures 2 and 3 show nonparametric bounds for ATEs and semiparametric bounds for $\beta ^{\ast }$ and ATEs for $T=2$ and $T=3$, respectively. The semiparametric bounds are obtained using the computational algorithm described above with $M=100$ and $\lambda _{M}=1.3\times 10^{-8}$. The elements of the fixed grid $\Upsilon _{M}$ are located at the percentiles of the standard normal distribution. We find that $\beta ^{\ast } $ is not identified for $T=2$, extending the result of Chamberlain (2010) to this example without time dummy. This result also holds for $T=3$, although it is difficult to appreciate in the figure because the identified set $B$ is very small. The nonparametric bounds for the ATEs (NP-bounds) can be very wide, even when we impose monotonicity (NPM-bounds) as described in the Supplementary Material. The semiparametric bounds for the ATEs (SP-bounds) are tighter than the nonparametric bounds and shrink very fast with $T$. In the Supplementary Material we report similar results for the logit, including nonidentification of the ATEs, except that $\beta ^{\ast }$ is identified, as is well known. Honore and Tamer (2006) also found tight bounds for the coefficient of a dynamic model.
To show that the approximate sets converge to the identified set as $M$ grows we impose some conditions. Let $d(\alpha ,\tilde{\alpha})$ denote a metric on the set $\Upsilon $ of possible values for $\alpha $.
Assumption 9: (i) $\Upsilon $ is a compact metric space with metric $d(\alpha ,\tilde{\alpha})$; ii) $\eta (M)=\sup_{\alpha \in \Upsilon }\min_{\tilde{\alpha}\in \Upsilon _{M}}d(\alpha ,\tilde{\alpha})$\ $\longrightarrow 0$ as $ M\longrightarrow \infty ;$ (iii) $\mathbb{B}$ is a compact subset of $\Re ^{b}$\textit{; (iv) there is }$C$ \textit{such that for all } $(\alpha ,\beta ),(\tilde{\alpha},\tilde{\beta})\in $\textit{$\Upsilon $}$ \times \mathbb{B}$, $\left\vert \mathcal{L}_{j}^{k}\left( \tilde{\alpha}, \tilde{\beta}\right) -\mathcal{L}_{j}^{k}\left( \alpha ,\beta \right) \right\vert \leq C[d(\tilde{\alpha},\alpha )+\left\Vert \tilde{\beta}-\beta \right\Vert ];$ \textit{and v) }$\Delta (\alpha ,\beta )$ \textit{is continuous on $\Upsilon $}$\times \mathbb{B}$\textit{.}
Although condition (i) seems restrictive, unbounded individual effects may be allowed if $\Upsilon $\ is chosen appropriately. For example, in the binary-choice model of equation ((ref)) this condition will be satisfied if $\Upsilon $ is taken to be a two-point compactification of the real line and $d(\alpha ,\tilde{\alpha})$ is specified appropriately, as shown in the following result.
Lemma 8: If Assumptions 5 and 8 and equation ((ref) ) are satisfied, where $H(v)$\ is strictly monotonic on $\Re $ with bounded continuous derivative, and $\mathbb{B}$\ is a compact subset of $\Re ^{b},$\ then there is a metric $d(\alpha , \tilde{\alpha})$\textit{\ and for each }$M$\textit{\ there is }$\Upsilon _{M}=\{\bar{\alpha}_{1M},...,\bar{\alpha}_{MM}\}$\textit{\ such that Assumption 9 is satisfied with }$\eta (M)=1/(M-1).$\textit{\ }
For the convergence results for the identified set we use the Hausdorff set metric,
\
Theorem 9: If Assumptions 5, 8, and 9 are satisfied, $ \epsilon _{M}\longrightarrow 0$, and $\left( \eta (M)+\lambda _{M}\right) /\epsilon _{M}\longrightarrow 0$\ then as $ M\longrightarrow \infty ,$
Under Assumptions 5 and 8 the complete description of the data-generating process is provided by the parameter vector $(${$P_{X}^{\prime }$}$,${$P$}$ ^{\prime })^{\prime },$ where $P_{X}=(P^{k},k=1,...,K)^{\prime }$ and $ P=(P_{j}^{k},j=1,...,J,k=1,...,K)^{\prime }.$ The true value of the parameter vector is $\Pi =({\mathcal{P}}_{X}^{\prime },\mathcal{P}^{\prime })^{\prime }$, where $\mathcal{P}_{X}=(\mathcal{P}^{k},k=1,...,K)^{\prime }$ and $\mathcal{P}=(\mathcal{P}_{j}^{k},j=1,...,J,k=1,...,K)^{\prime },$ and the empirical estimate is $\hat{\Pi}=({\hat{P}}_{X}^{\prime },\hat{P} ^{\prime })^{\prime }$, where $\hat{P}_{X}=(\hat{P}^{k},k=1,...,K)^{\prime }$ and $\hat{P}=(\hat{P}_{j}^{k},j=1,...,J,k=1,...,K)^{\prime }.$
The estimation method is like the computational one in using linear-in-parameters approximations to the probabilities. Here we describe the estimation method and give a consistency result, and in the Supplementary Material we provide the implementation details. We follow the same steps as the computational one except that we use estimated weights $\hat{w}_{j}^{k}$ and estimated probabilities $\hat{P} _{j}^{k}$. Let $\hat{M}$ be a choice of $M$ that may depend on the data and sample size, and
Let $\hat{T}_{\lambda }(\beta )=\min_{\pi \in \mathcal{S}_{M}^{K}}\hat{T} _{\lambda }(\beta ,\pi )$ and $\epsilon _{n}$ $>0$ be a positive scalar. We estimate the identified set for $\beta $ by
where $\mathbb{B}$ is the parameter space and $\epsilon _{n}$ is a cut-off parameter that shrinks to zero with the sample size, as in Manski and Tamer (2002) and Chernozhukov, Hong, and Tamer (2007). The ATE bounds can be estimated by
This approach to estimation (and computation) can be easily modified to handle the case where the distribution of the individual effect is restricted to be the same across some values of $k$. Such a modification could be implemented by imposing equality of $\pi _{m}^{k}$ across those values of $k.$ An example would be a model where the distribution of $\alpha _{i}$ did not depend on some component of $X_{it}.$ That restriction could be imposed setting $\pi _{m}^{k}$ to be equal across $k$ where the other components of $X_{it}$ do not vary. Or in a case with a lagged dependent variable we could restrict the distribution of $\alpha $ to only depend on the initial condition by imposing equality of $\pi _{m}^{k}$ across all $k$ where $Y_{i0}$ takes on a particular value.
The following is a consistency result.
Theorem 10: If Assumptions 5, 8, and 9 are satisfied, $ \hat{w}_{j}^{k}\overset{p}{\longrightarrow }w_{j}^{k}>0,$ $\hat{P}_{j}^{k} \overset{p}{\longrightarrow }\mathcal{P}_{j}^{k}$, $\epsilon _{n}\longrightarrow 0$, and $\left( n^{-1}+\eta (\hat{M})+\lambda _{n}\right) /\epsilon _{n}\overset{p}{\longrightarrow }0$,\ then $ d_{H}(\hat{B},B)\overset{p}{\longrightarrow }0,\hat{\Delta}_{\ell }^{k} \overset{p}{\longrightarrow }\Delta _{\ell }^{k},\hat{\Delta}_{u}^{k}\overset {p}{\longrightarrow }\Delta _{u}^{k}.$
It is interesting to note that no upper limit is placed on $M$ in this result or in Theorem 9. The reason for this is that the model is finite dimensional, so there is no need for such a limit. Mathematically, a richer, fixed grid simply corresponds to a bigger submodel of the finite-dimensional model.
Turning now to the inference for the semiparametric models, we note that it is rather challenging. The estimators of parameters and ATE are obtained by nonlinear programming subject to data-dependent constraints that are modified to respect the constraints of the model. The distributions of these highly-complex estimators are not tractable, and are also non-regular in the sense that the limit versions of these distributions do not vary with perturbations of the DGP in a continuous\ fashion. This implies that the usual bootstrap is not consistent. To overcome all of these difficulties we will rely on a variation of the bootstrap, which we call the perturbed bootstrap. We also give an alternative inference method based on a modified projection in the Supplementary Material.
The usual bootstrap computes the critical value -- the $\alpha $-quantile of the distribution of a test statistic -- given a consistently-estimated data-generating process (DGP). If this critical value is not a continuous function of the DGP, the usual bootstrap fails to consistently estimate the critical value. We instead consider the perturbed bootstrap, where we compute a set of critical values generated by suitable perturbations of the estimated DGP and then take the most conservative critical value in the set. If the perturbations cover at least one DGP that gives a more conservative critical value than the true DGP does, then this approach yields a valid inference procedure.
The approach outlined above is most closely related to the Monte-Carlo inference approach of Dufour (2006); see also Romano and Wolf (2000) for a finite-sample inference procedure for the mean that has a similar spirit. In the set-identified context, this approach was first applied in the MIT thesis work of Rytchkov (2007); see also Chernozhukov (2007).
We consider the problem of performing inference on a real parameter $\theta ^{\ast }$. For example, $\theta ^{\ast }$ can be an upper (or lower) bound on the conditional ATE $\Delta ^{k}$ such as
where $P^{\ast }$ denotes the projection of $P$ onto the model space $\Xi =\{P:\exists \beta \in \mathbb{B}$ with $\mathcal{F}_{k}(\beta ,P)\neq \varnothing ,\forall k=1,...,K\}$, i.e.
and $B^{\ast }(P)$ is the corresponding projection for the identified set of the parameter, i.e.
Alternatively, $\theta ^{\ast }$ can be an upper (or lower) bound on a scalar functional $c^{\prime }\beta ^{\ast }$ of the parameter $\beta ^{\ast }$. Then we define
In both cases we project $P$ onto the model space in order to address the problem of infeasibility of constraints defining the parameters of interest under misspecification or sampling error. Under misspecification, we interpret our inference as targeting the parameters of interest in a best approximating model; see the Supplementary Material on the modified projection method for further details. Under correct specification, our inference targets the parameters of interest in the true model.
In order to perform inference on the true value $\theta ^{\ast }=\theta ^{\ast }(\mathcal{P})$ of the parameter, we use the statistic
where $\hat{\theta}=\theta ^{\ast }(\hat{P})$. Let $G_{n}(s,P)$ denote the distribution function of $S_{n}(P)=\hat{\theta}-\theta ^{\ast }(P)$, when the data follow the DGP $P$. The goal is to estimate the distribution of the statistic $S_{n}$ under the true DGP $P=\mathcal{P}$, that is, to estimate $ G_{n}(s,\mathcal{P})$.
The method proceeds by constructing a confidence region $CR_{1-\gamma }( \mathcal{P})$ that contains the true DGP $\mathcal{P}$ with probability $ 1-\gamma $, close to one. For efficiency purposes, we also want the confidence region to be an efficient estimator of $\mathcal{P}$, in the sense that as $n\rightarrow \infty $, $d_{H}(CR_{1-\gamma }(\mathcal{P}), \mathcal{P})=O_{p}(n^{-1/2}),$ where $d_{H}$ is the Hausdorff distance between sets. Specifically, in our case we use
where $c_{1-\gamma }(\chi _{K(J-1)}^{2})$ is the $(1-\gamma )$-quantile of the $\chi _{K(J-1)}^{2}$ distribution and $W$ is the goodness-of-fit statistic:
Then we define the estimates of the lower and upper bounds on the quantiles of $ G_{n}(s,\mathcal{P})$ as
where $G_{n}^{-1}(\alpha ,P)=\inf \{s:G_{n}(s,P)\geq \alpha \}$ is the $ \alpha $-quantile of the distribution function $G_{n}(s,P)$. Then we construct a $(1-\alpha -\gamma )\cdot 100\%$ confidence region for the parameter of interest as
where, for $\alpha =\alpha _{1}+\alpha _{2}$,
This formulation allows for both one-sided intervals (either $\alpha _{1}=0$ or $\alpha _{2}=0$) or two-sided intervals ($\alpha _{1}=\alpha _{2}=\alpha /2)$.
For the inference results we condition on the observed distribution of $X$ and thus set $P_{X}=\mathcal{P}_{X}=\hat{P}_{X}.$ We make the following assumption about the data-generating process.
Assumption 10: \ $\Pi \in $\thinspace $\mathbb{P=} \{(P_{X},P):P^{k}>\varepsilon ,P_{j}^{k}>\varepsilon ;j=1,...,J,k=1,...,K\}$ for some $\varepsilon >0$.
The following theorem shows that this method delivers (uniformly) valid inference on the parameter of interest.
Theorem 11: If Assumptions 5, 8, and 9 are satisfied then for any sequence of data-generating process $\Pi = \Pi_n$ satisfying Assumption 10,
In practice, we use the following simulation approach to compute the confidence intervals.
Algorithm: Perturbed Bootstrap
We illustrate the estimation and inference results with two empirical examples. One estimates identified effects and calculates bounds for the effect of unions on earnings quantiles. The other compares nonparametric and semiparametric bounds for the effect of fertility on women's labor force participation.
We revisit the empirical question of how unions impact wage structure using panel data. Our major contribution here is to estimate the effect without imposing the assumption that unobserved heterogeneity is some additive term that can be simply differenced out. In our model unobserved heterogeneity can have an almost unrestricted impact on the structural/causal response functions, with the time homogeneity serving as the only restriction.
Our analysis is motivated by previous empirical studies that find differences in unobservables between union and nonunion workers. For instance, in an influential study, Chamberlain (1982) finds strong evidence of heterogeneity bias in the estimation of the union effect by comparing estimates of cross-sectional models and panel data models with additive heterogeneity. This finding demonstrates the important need of controlling for unobserved heterogeneity. Also, Angrist and Newey (1991) reject the hypothesis that the unobserved heterogeneity acts solely in an additive fashion, motivating the need to control for more general unobserved heterogeneity. Card (1996) found differences in the union and selection effect across skill levels. Here we account fully for differences across individuals in the union effect while allowing correlation of that effect with union status, thus accounting for selection. Recently Frandsen (2011) focused on quantile union effects using a regression-discontinuity design that estimates union effects for those near a union election discontinuity rather than for those whose union status changes. We find a flatter quantile profile than he does, consistent with his theoretical results that suggest a flatter profile away from the discontinuity.
We use data from the National Longitudinal Survey (Youth Sample). The sample consists of full-time, young, working males, 20 to 29 years old in 1986, followed over the period 1986 to 1993. We exclude individuals who failed to provide sufficient information for each year, were in the active armed forces or were students any year, or who reported too high (more than \$500 per hour) or too low (less than \$1 per hour) wages. The final sample includes 2,065 men followed over 8 years. We use the union membership and the log-hourly wage rate in 1980 dollars as the covariate and the outcome variables. The union membership variable reflects whether or not the individual had his wage set by a collective bargaining agreement. Vella and Verbeek (1998) also used data from the NLSY for different years and found evidence of important union effect heterogeneity with a random effects model.
We begin by imposing the stationarity condition that income with and without union membership has the same distribution in each time period but also will allow for location and scale time effects. It turns out that time effects are not important in this data. Some covariates are also allowed for since time-invariant covariates are absorbed in the individual effects. Insensitivity to time effects also suggests that time-varying covariates may not be important though a fuller exploration would be useful. For brevity we focus on the case without covariates.
In our analysis, we focus on estimating the union quantile effect for the subpopulations of workers that ever became unionized within the sample (47% of the sample) or that were unionized in the first year (20% of the sample). For these subpopulations, the union effect is not point-identified, since there are 13% of the ever-unionized workers that always stayed unionized between 1986 and 1993, and there are 32% of the workers unionized in 1986 that remained unionized until 1993. However, we hope to construct informative bounds on the union effect. We consider both a static model that allows for the union membership decisions to be strictly exogenous with respect to wage-setting decisions, and a dynamic model that allows for the union-membership decisions to be only predetermined with respect to wage-setting decisions. We shall also report the estimates of the union effect for the subpopulation of workers who change their union status at least once within the sample. For this subpopulation, the effect is point-identified in the static model, that is, the bounds on the union effect collapse to a point. We shall not estimate the union effect for the entire population of workers, since the bounds are completely uninformative in this case. This happens because more than half of the workers are never unionized within the sample (see Table 1).
All the results are reported in Table 1 and Figure 4. Table 1 assesses the plausibility of the time-homogeneity assumption by comparing moments and quantiles of the cross-sectional distributions of log-wages across years for workers that do not change union status. Under time homogeneity, these cross sectional distributions should remain time invariant in the static model. In the table we observe distributional changes across years, but most of the variation can be captured by additive location effects for both always-unionized and never-unionized workers.
Panels A and B of fig. 4 present the estimates of the union effect in the static model for the subpopulation of workers who change their union status at least once within the sample. In panel A we compare our panel data estimates of quantile effects that control for individual heterogeneity with pooled estimates that do not control for individual heterogeneity. In the pooled estimates, we see that the quantile effect of union membership is positive but declines sharply at the upper end of the distribution, which agrees with previous cross-sectional findings (Chamberlain, 1994). A common explanation for this phenomenon is that the high-skill workers at the lower end of the earning distribution tend to join the union, whereas the high-skill workers at the high end of the earning distribution tend not to join the union. The estimated quantile effect in the cross-section therefore captures this selection effect of unobserved skills. In the panel-data estimates, which control for unobserved skills, we see that the quantile effects of union membership become very flat across the quantile indices. Thus, by controlling for individual heterogeneity, we have eliminated the selection effect. Panel B shows that the results are not sensitive to the inclusion of location and scale effects.
Panel C presents estimated bounds on the union effect for the subpopulation of workers that ever became unionized within the sample using the static model with time effects. The bounds are informative, and show that the effect is positive for most of the quantile indices. The panel also shows bounds obtained using the assumption of monotonic and positive union effect on earnings described in the Supplementary Material. These bounds are also informative, and in fact are substantially tighter than the bounds obtained without the monotonicity assumption. Panel D presents similar bounds on the union effect for the subpopulation of workers unionized in the first period using the dynamic model. The bounds in this case are not informative, even after imposing monotonicity.
All the panels include 90% uniform confidence bands for the quantile union effects constructed by bootstrap with 200 repetitions. These bands allow us to make visual simultaneous inference on the entire quantile functions. For example, we cannot reject that the identified union effect is constant and positive for all the quantiles. For the ever unionized, the quantile union effect is positive for a large range of quantiles.
For an application of the semiparametric bounds we consider a binary choice panel model of female labor force participation. We focus on the relationship between participation and the presence of young children in the household. Other studies that estimate similar models of participation in panel data include Heckman and MaCurdy (1980, 1982), Chamberlain (1984), Hyslop (1999), Chay and Hyslop (2000), Carrasco (2001), Carro (2007), and Fern\'{a}ndez-Val (2009).
The empirical analysis is based on a sample of married women from the National Longitudinal Survey of Youth 1979 (NLSY79). The sample consists of 1,587 married women. Only women continuously married, not students or in the active forces, and with complete information on the relevant variables in the entire sample period are selected from the survey. Descriptive statistics for the sample are shown in Table 2. The labor force participation variable ($LFP$) is an indicator that takes the value one if the woman's employment status is \textquotedblleft in the labor force\textquotedblright\ according to the CPS definition, and zero otherwise. The fertility variable ($kids$) indicates whether the woman has any children younger than 3 years. We focus on very young, preschool children as most empirical studies find that their presences have the strongest impact on the mother's participation decision. $LFP$ is stable across the years considered, whereas $kids$ is decreasing. The proportion of women that change fertility status grows steadily with the number of time periods of the panel, but there are still $49\%$ of the women in the sample for which the effect of fertility is not identified after 3 periods.
The empirical specification we use is similar to Chamberlain (1984). In particular, we estimate the following equation
where $\alpha _{i}$ is an individual-specific effect. The parameters of interest are $\beta ^{\ast }$ and the ATE of fertility on participation. We compute nonparametric and semiparametric probit and logit bounds for these parameters. We also obtain linear and nonlinear fixed effects estimates, together with large-$T$ analytical bias corrected estimates and conditional fixed effects logit estimates.\footnote{ The analytical corrections use the estimators of the bias based on expected quantities in Fern\'{a}ndez-Val (2009).} The nonparametric bounds impose monotonicity on the effects. For the semiparametric bounds, we use the method described in Section 9 with penalty $\lambda _{n}=1/(n\log n)$ and iterate the quadratic program 3 times with initial weights $\hat{w}_{j}^{k}= \hat{P}^{k}$. This iteration makes the estimates insensitive to the penalty and weighting. We search over discrete distributions with $\hat{M}=23$ support points at $\{-\infty ,-4,-3.6,...,3.6,4,\infty \}$ for the parameter $\beta ^{\ast }$, and with $\hat{M}=163$ support points at $\{-\infty ,-8,-7.9,...,7.9,8,\infty \}$ for the ATE. The estimates are based on panels of 2 and 3 time periods, both of them starting in 1990.
Table 3 reports estimates and 95% confidence regions for the parameters of interest. The confidence regions for the nonparametric bounds are constructed using the normal approximation $(95\%\ N)$ and nonparametric bootstrap with 200 repetitions $(95\%\ B)$. The confidence regions for the semiparametric bounds are obtained using the procedures described in Section 9 and the Supplementary Material. For the perturbed bootstrap method $ (95\%\ PB)$ we use $R=100$, $\gamma =.01$, $\alpha _{1}=\alpha _{2}=.02,$ and 200 simulations from each DGP to approximate the distribution of the statistic. For the modified projection method $(95\%\ MP)$, the confidence interval for $\mathcal{P}$ in the first stage is approximated by 5,000 DGPs drawn from the empirical multinomial distributions that pass the goodness-of-fit test. Together the modified projection and the perturbed bootstrap took several days to compute on a personal computer. We also include confidence intervals obtained by a canonical projection method $ (95\%\ CP)$ less robust to model misspecification than the modified projection method, that intersects a nonparametric confidence interval for $ \mathcal{P}$ with the space of probabilities compatible with the semiparametric model $\Xi $:
For the fixed-effects estimators, the confidence regions are based on the asymptotic normal approximation. The semiparametric estimates are shown for $ \epsilon _{n}=0$, i.e., for the solution that gives the minimum value in the quadratic problem.
Overall, we find that the nonparametric bound estimates and confidence regions are too wide to provide informative evidence about the relationship between participation and fertility. The semiparametric bounds offer a good compromise between producing more informative results without adding too much structure to the model. Thus, these estimates are always inside the confidence regions of the nonparametric model and do not suffer important efficiency losses relative to the fixed-effects estimates. Another salient feature of the results is that the misspecification problem of the canonical projection method clearly arises in this application. Thus, this procedure gives empty confidence regions for the panel with 3 periods. The perturbed bootstrap and modified projection methods produce similar (non-empty) confidence regions for the model parameters and ATEs.
The semiparametric intervals for the ATE cover the -9.6% estimate of Chamberlain (1984) for the expected effect of having an additional young child on the participation probability. He obtained this estimate from a correlated, random-coefficient probit model, a richer specification that includes education and fertility covariates, and a different sample from the PSID.
\setcounter{section}{0}
In this supplemental material we provide omitted discussions, results, and proofs by Section in the same order they are referred to in the paper. Let w.p.a.1 denote "with probability approaching one" and $C$ denote a generic constant that may be different in different uses.