EconBase
← Back to paper

Average and Quantile Effects in Nonseparable Panel Models

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

134,566 characters

Average and Quantile Effects in Nonseparable Panel Models



\title{Average and Quantile Effects in Nonseparable Panel Models\thanks{
We thank J. Angrist, G. Chamberlain, D. Chetverikov, B. Frandsen, B. Graham, J. Hausman, and
many seminar participants for comments. Brad Larsen and Seongyeon Chang provided capable
research assistance. Parts of this paper were given at the 2007 CEMMAP
Microeconometrics: Measurement Matters Conference, the Shanghai Lecture of
the 2010 World Congress of the Econometric Society, and conferences in
between. We gratefully acknowledge research support from the NSF.}}
\author{Victor Chernozhukov \\
MIT \and Iv\'an Fern\'andez-Val \\
BU \and Jinyong Hahn \\
UCLA \and Whitney Newey \\
MIT}
\maketitle

\begin{abstract}
Nonseparable panel models are important in a variety of economic settings,
including discrete choice. This paper gives identification and estimation
results for nonseparable models under time homogeneity conditions that are
like \textquotedblleft time is randomly assigned\textquotedblright\ or
\textquotedblleft time is an instrument.\textquotedblright\ Partial
identification results for average and quantile effects are given for
discrete regressors, under static or dynamic conditions, in fully
nonparametric and in semiparametric models, with time effects. It is shown
that the usual, linear, fixed-effects estimator is not a consistent
estimator of the identified average effect, and a consistent estimator is
given. A simple estimator of identified quantile treatment effects is given,
providing a solution to the important problem of estimating quantile
treatment effects from panel data. Bounds for overall effects in static and
dynamic models are given. The dynamic bounds provide a partial
identification solution to the important problem of estimating the effect of
state dependence in the presence of unobserved heterogeneity. The impact of $
T$, the number of time periods, is shown by deriving shrinkage rates for the
identified set as $T$ grows. We also consider semiparametric,
discrete-choice models and find that semiparametric panel bounds can be much
tighter than nonparametric bounds. Computationally-convenient methods for
semiparametric models are presented. We propose a novel inference method
that applies in panel data and other settings and show that it produces
uniformly valid confidence regions in large samples. We give empirical
illustrations.
\end{abstract}

\section{Introduction}

Interesting empirical questions are often formulated in terms of the \textit{
ceteris paribus} effect of $x$ on $y,$ when observed $x$ is an individual
choice variable partly determined by preferences or technology. Panel data
holds out the hope of controlling for individual preferences or technology
by using multiple observations for a single economic agent. This hope is
particularly difficult to realize with discrete or other nonseparable models
and/or multidimensional individual effects. These models are, by nature, not
additively separable in unobserved individual effects, making them
challenging to identify and estimate. There are some simple solutions, such
as the conditional MLE for the slope parameter of a binary-choice logit
model with an individual location effect. However these are rare and
dependent on specific models or distributions. For example, the slope
parameter of the binary-choice model with a time dummy is identified only
for logit as shown by Chamberlain (2010), and the average treatment effect
is not identified even for logit without a time dummy, as shown below.

A fundamental idea for using panel data to identify the \textit{ceteris
paribus} effect of $x$ on $y$ is to use changes in $x$ over time to estimate
the effect. In order for changes over time in $x$ to correspond to \textit{
ceteris paribus} effects, the distribution of variables other than $x$ must
not vary over time. This condition is like \textquotedblleft time being
randomly assigned\textquotedblright\ or \textquotedblleft time is an
instrument.\textquotedblright\ In this paper we consider identification via
such time homogeneity conditions. They are also the basis of many previous
panel results, including Chamberlain (1982), Manski (1987), and Honore
(1992). Here we consider the identifying power of time homogeneity for
nonseparable models, i.e. for models that are not additively separable in
unobserved factors. We allow for multidimensional heterogeneity, as
motivated by models where effects of interest, such as price and income
elasticities, are distributed among individuals in unrestricted ways; see
Altonji and Matzkin (2005), Browning and Carro (2007), and Fernandez-Val and
Lee (2010), among others. We also weaken the strict time homogeneity
conditions to allow some time effects.

Models with discrete regressors have many applications and are the subject
of most of this paper. With discrete regressors, time homogeneity only leads
to partial identification of many effects, though some conditional effects
are identified. This paper considers partial identification and estimation
of average and quantile effects, under static or dynamic conditions, in
fully nonparametric and in semiparametric models, with time effects.

For the nonparametric, static model we give simple estimators of the
identified average effect of $x$ on $y$, conditional on $x$ varying over
time. These estimators extend Chamberlain (1982, pp. 10-17) to multiple
regressors with location and scale time effects. We also find that linear,
fixed-effects estimate a variance-weighted average effect instead of the
average effect. For bounded $y$ we move beyond the analysis of identified
effects and give simple estimators of sharp bounds for average effects.
These bounds provide nonparametric, partial-identification estimates of
average effects in important cases, such as binary choice in panel data.

The quantile estimators given here are more novel than the average-effect
estimators. They provide simple estimators of the effect of $x$ on quantiles
of $y,$ conditional on $x$ varying over time, that allow for location and
scale time effects. Estimators of sharp bounds are also provided for the
unconditional, overall quantile effect. The estimators allow for
multidimensional heterogeneity, for example for both location and slope to
vary across individuals in an unrestricted way. In this way we provide a
solution for the important problem of nonparametric quantile regression in
panel data with individual effects, for discrete regressors. Graham, Hahn
and Powell (2009) also consider quantile effects in linear, heterogenous coefficients models,
but impose conditions which essentially restrict the heterogeneity
to be one-dimensional, and focus on identification of the distribution of
coefficients.

Dynamics is often an important feature of economic models with intertemporal
choice. Here we give a dynamic, nonseparable, panel model that nests the
static one. Simple estimators of bounds on average and quantile effects are
provided. We show that these results provide a partial-identification
solution to the important problem of distinguishing state dependence from
heterogeneity.

This paper shows the impact of the number of time periods $T$ on
identification. We find that the identified set of effects shrinks to a
point exponentially quickly as $T$ grows, when individual effects are
bounded and time period disturbances are not, and that the rate is some
power of $T^{-1}$ more generally. In a nonparametric, dynamic, binary-choice model
we find that the rate is faster the larger the variance of the
period-specific disturbance relative to the variance of the individual
effect.

In numerical examples we find that the nonparametric bounds can be quite
wide, motivating more informative models. Semiparametric models that specify
the distribution of the outcome given regressors and individual effect is an
important class of more informative models. Here we describe both static and
dynamic semiparametric models. When restrictions are imposed on the
heterogeneity, like only some coefficients varying across individuals,
semiparametric models can have substantially tighter bounds than
nonparametric models. We find that in the important binary-logit model with
just a location effect the average effect bounds shrink exponentially
quickly as $T$ grows, in both dynamic and static models, even when the
nonparametric bounds shrink slowly. This result quantifies the gain in
information of a semiparametric model with just a location effect over the
nonparametric model. We also find quite tight bounds for semiparametric
models relative to nonparametric ones in numerical examples.

We show that semiparametric, discrete-choice models have finite dimensional
parameterizations. This reduces bounds calculation and estimation to a
finite-dimensional problem, albeit a large dimensional, highly nonlinear,
and computationally difficult one. To make computation more feasible we use
grids of fixed values for individual effects, so that average choice
probabilities are finite-dimensional, linear combinations. We combine this
with minimum squared distance fitting of data cell probabilities to obtain a
quadratic programming approach for estimating the individual-effect
distributions. This approach is computationally convenient and overcomes
problems with previously proposed methods, as further discussed below. We
also allow the grid to grow in order to approximate the true support points.
It turns out that because the model is finite dimensional there is no need
to limit the number of grid points. Mathematically, a richer fixed grid
simply corresponds to a bigger submodel of the finite-dimensional model.

The semiparametric bounds build on Honor\'{e} and Tamer (2003, 2006) and
Chernozhukov, Hahn, and Newey (2004). Both papers gave results for bounds in
semiparametric, nonlinear, panel-data models. Honore and Tamer (2006)
proposed linear programming, minimum distance, and maximum likelihood
methods for dynamic models. Chernozhukov, Hahn, and Newey (2004) proposed
sieve likelihood estimation of bounds for static models. These approaches
are not very useful for estimation. Plugging in sample frequencies in place
of cell probabilities in the linear-programming algorithm produces empty
identification regions because the frequencies need not satisfy constraints
imposed by the model. Also, the minimum-distance objective function is
computationally difficult, as is sieve maximum likelihood, given the
dimensionality of the individual-effect distributions. Honore and Tamer
(2006) also assumed a fixed known grid for true individual effects, while we
consider an approximation to an unknown grid.

The inferential problem for the semiparametric models is also rather
challenging. The models impose data-dependent constraints that are often
infeasible in finite samples or under misspecification, which produces empty
confidence regions. We overcome these difficulties by projecting these
data-dependent constraints onto the model space using the
quadratic-programming approach mentioned above, thus producing an
always-feasible, data-dependent constraint set. We then suggest linear and
nonlinear programming methods that use these new modified constraints. Our
inference procedures have the appealing justification of targeting the true
model under correct specification and targeting a best approximating model
under incorrect specification. We also develop two novel inferential
procedures, one called the \textit{perturbed bootstrap}, that is described
in the paper, and another called \textit{modified projection,} that is
described in the Supplementary Material. These methods produce uniformly
valid inference in large samples and may be of substantial independent
interest.

We give two empirical illustrations. One is to estimate the effect of unions
on earnings quantiles. There we find that a decline in the union effect as
the quantile increases can be attributed to individual heterogeneity. The
other illustration is to estimate the effects of fertility on women's labor
force participation. There we compare nonparametric and semiparametric
estimates.

Recent research has considered nonseparable panel models with time
homogeneity and continuous regressors. Graham and Powell (2011) give
estimators of the average effect in a linear model with heterogeneous
slopes. Hoderlein and White (2011) give estimators of the average derivative
conditional on equality of regressors across time periods.

Chamberlain (1980, 1984), Altonji and Matzkin (2005), Bester and Hansen
(2008), and others have used control functions for panel data estimation. We
focus instead on time homogeneity with unrestricted dependence between
individual effects and regressors. Bias-corrected, fixed-effects estimation
of semiparametric models has been proposed by Hahn and Kuersteiner (2002),
Alvarez and Arellano (2003), Woutersen (2002), Hahn and Newey (2004), and
Fern\'{a}ndez-Val (2009). These estimators depend on large $T$ for
consistency while we estimate identified effects and bounds for fixed $T$.

Section 2 describes the models and effects we consider. Section 3 discusses
estimation of identified effects. Sections 4 and 5 derive bounds for the
static and dynamic nonparametric models respectively. Section 6 describes
the impact of $T$. Section 7 describes and gives results for semiparametric,
discrete-choice models. Section 8 gives computationally convenient methods
for semiparametric models and numerical examples. Section 9 considers
estimation and inference for semiparametric models. Section 10 gives
empirical examples. The Supplementary Material Chernozhukov et. al. (2012) includes a variety of omitted
discussions and results along with the proofs of results stated in the paper.

\section{The Models and Effects}

The data consist of $n$ observations on $Y_{i}=(Y_{i1},...,Y_{iT})^{\prime }$
and $X_{i}=[X_{i1},...,X_{iT}]^{\prime }$, for a dependent variable $Y_{it}$
and a vector of regressors $X_{it}$. Throughout we assume that the
observations $(Y_{i},X_{i})$\textit{, }$(i=1,...,n)$\textit{, }are
independent and identically distributed. The nonparametric models we
consider satisfy

\bigskip

\textsc{Assumption 1: }\textit{There is a function }$g_{0}(x,\alpha
,\varepsilon )$\textit{\ and vectors }$\alpha _{i}$\textit{\ and }$
\varepsilon _{it}$\textit{\ of random variables such that}
\begin{equation*}
Y_{it}=g_{0}(X_{it},\alpha _{i},\varepsilon _{it}),(i=1,...,n;t=1,...,T).
\end{equation*}

The vector $\alpha _{i}$ consists of time invariant individual effects that
often represent individual heterogeneity. The vector $\varepsilon _{it}$
represents period-specific disturbances. Altonji and Matzkin (2005)
considered models satisfying Assumption 1. The invariance of $g_{0}$ over
time in this Assumption does not actually impose any time homogeneity. If
there are no restrictions on $\varepsilon _{it}$ then $t$ could be one of
the components of $\varepsilon _{it},$ allowing the function to vary over
time in a completely general way. The next condition, together with
Assumption 1, imposes time homogeneity on the model.

\bigskip

\textsc{Assumption 2:} $\varepsilon _{it}|X_{i},\alpha _{i}\overset{d}{=}
\varepsilon _{i1}|X_{i},\alpha _{i}$\textit{, for all }$t.$

\bigskip

This is a static, or \textquotedblleft strictly exogenous\textquotedblright\
time homogeneity condition, where all leads and lags of the regressor are
included in the conditioning variable $X_{i}.$ It requires that the
conditional distribution of $\varepsilon _{it}$ given $X_{i}$ and $\alpha
_{i}$ does not depend on $t,$ but does allow for dependence of $\varepsilon
_{it}$ over time. An equivalent condition is $\tilde{\varepsilon}_{it}|X_{i}
\overset{d}{=}\tilde{\varepsilon}_{i1}|X_{i}$ for $\tilde{\varepsilon}
_{it}=(\alpha _{i},\varepsilon _{it}).$ Thus, the time invariant $\alpha
_{i} $ has no distinct role in this model. The condition is just that
whatever the unobserved disturbances are, their conditional distribution
given $X_{i}$ does not depend on $t$.

This seems a basic condition that helps panel data provide information about
the effect of $x$ on $y.$ It is like \textquotedblleft time is randomly
assigned\textquotedblright\ or \textquotedblleft time is an
instrument\textquotedblright\ with the distribution of factors other than $x$
not varying over time, so that changes in $x$ over time can help identify
the effect of $x$ on $y$. Assumption 2 also turns out to be a natural
strengthening of linear model conditions, as shown in Theorem A1 and the
associated discussion in the Supplementary Material.

A dynamic model can be obtained by only including current and lagged $X_{is}$
in the conditioning set for each $t,$ as in the following condition:

\bigskip

\textsc{Assumption 3:} $\varepsilon _{it}|X_{it},...,X_{i1},\alpha _{i}
\overset{d}{=}\varepsilon _{i1}|X_{i1},\alpha _{i}$,\textit{\ for all }$t.$

\bigskip

This is a \textquotedblleft predetermined\textquotedblright\ version of time
homogeneity that is nested within the static model of Assumptions 1 and 2,
as shown in Theorem A2 of the Supplementary Material. Here the conditional
distribution given only current and lagged regressors must be time
invariant. It also implies that the conditional distribution of $\varepsilon
_{it}$ given current and lagged regressors only depends on $X_{i1}$. Here $
\varepsilon _{it}$ can be thought of as additional information that is
independent of the past regressors. A conditional-mean version of this
condition arises in rational-expectations models that implies disturbances
have mean zero conditional on past information. Here the stronger
conditional independence restriction is imposed as seems needed for a
nonseparable model. The conditioning on $X_{i1}$ is a way to account for the
initial conditions of this dynamic model. Bhargava and Sargan (1983) adopted
this approach in a linear model as have Honore and Tamer (2006) and Browning
and Carro (2007) in a likelihood setting.

If $X_{it}$ includes lagged $Y_{it}$ then Assumption 3 specifies that the
model is \textquotedblleft dynamically complete,\textquotedblright\ ruling
out $Y_{it}=g_{0}(X_{it},\alpha _{i},\varepsilon _{it})$ as one equation of
a dynamic system. For instance, $X_{it}$ could be $Y_{i,t-1},$ in which case
$Y_{it}=g_{0}(Y_{it-1},\alpha _{i},\varepsilon _{it})$ is an explicit
nonseparable dynamic model with $\varepsilon _{it}$ being time shocks that
are independent of $Y_{it-1},...,Y_{i1}$. An important example is one where $
Y_{it}\in \{0,1\}$ is binary, representing state dependence, with $\alpha
_{i}$ representing unobserved heterogeneity. This example is treated in
Section 5.

We will focus in the nonparametric model on two objects, the average
structural function (ASF) of Blundell and Powell (2003) and the quantile
structural function (QSF) of Imbens and Newey (2009). The ASF is
\begin{equation*}
\mu (x)=E[g_{0}(x,\alpha _{i},\varepsilon _{it})]=\int g_{0}(x,\alpha
,\varepsilon )dF(\alpha ,\varepsilon ),
\end{equation*}
where throughout the paper $F$ denotes the cumulative distribution function
(CDF) of a random vector that appears as the arguments of $F$. This object
is useful for quantifying the effect of $x$ on the mean of the outcome $
Y_{it}$. In the treatment-effects literature the average treatment effect
(ATE) of changing $x$ from $x^{b}$ (before) to $x^{a}$ (after) is
\begin{equation*}
\Delta =\mu (x^{a})-\mu (x^{b}).
\end{equation*}

The QSF $q(\lambda ,x)$ is the $\lambda ^{th}$ quantile of $g_{0}(x,\alpha
_{i},\varepsilon _{it}).$ Under conditions specified below the QSF will
equal the inverse of the CDF of $g_{0}(x,\alpha _{i},\varepsilon _{it})$,
\begin{equation*}
q(\lambda ,x)=G^{-1}(\lambda ,x),G(y,x)=E[1(g_{0}(x,\alpha _{i},\varepsilon
_{it})\leq y)].
\end{equation*}
In the treatment-effects literature the $\lambda ^{th}$ quantile treatment
effect (QTE) of changing $x$ from $x^{b}$ to $x^{a}$ is
\begin{equation*}
\Delta _{\lambda }=q(\lambda ,x^{a})-q(\lambda ,x^{b}),
\end{equation*}
as in Lehmann (1974). This effect does not give the quantile of the
treatment effect but does quantify the shift in the distribution of $Y_{it}$
that is due to a change in $x.$ It accounts for multidimensional individual
effects that may be correlated with $x$.

The static model implies a conditional-mean model that has been considered
by Chamberlain (1982), Hahn (2001), Wooldridge (2005), and Chernozhukov et.
al. (2007). This conditional-mean model specifies that there is an $\alpha
_{i}$ and $m_{0}(x,\alpha )$ such that $E[Y_{it}|X_{i},\alpha
_{i}]=m_{0}(X_{it},\alpha _{i}).$ A conditional mean ATE, as in Wooldridge
(2005), is $\int [m_{0}(x^{a},\alpha )-m_{0}(x^{b},\alpha )]dF(\alpha )$.
This model and effect differ from those we consider in specifying
conditional-mean restrictions, while we specify conditional distribution
restrictions. In Theorem A3 of the Supplementary Material we show that the
conditional-mean model is implied by Assumptions 1 and 2, or 1 and 3, and
that the conditional mean ATE is equal to the ATE we consider. Thus all
results we give for the ATE, including bounds, apply to the conditional mean
models, such as that of Chernozhukov et. al. (2007).

To help explain the relationship between the conditional-mean model and the
model of our paper, and to illustrate other results, it is useful to
consider examples. Binary choice is a very important model for panel data,
as it has many applications. For this reason we use binary choice as a main
example. The most common model has been one with a scalar individual effect
that is an additive shift to a linear combination of $X_{it}$, where
\begin{equation*}
Y_{it}=1(X_{it}^{\prime }\beta ^{\ast }+\alpha _{i}\geq \varepsilon _{it}),
\end{equation*}
for scalar $\varepsilon _{it}$ and an unknown parameter vector $\beta ^{\ast
}$. In this example $g_{0}(x,\alpha ,\varepsilon )=1(x^{\prime }\beta ^{\ast
}+\alpha \geq \varepsilon )$ and the ATE is
\begin{equation*}
\Delta =\int [1(x^{a\prime }\beta ^{\ast }+\alpha \geq \varepsilon
)-1(x^{b\prime }\beta ^{\ast }+\alpha \geq \varepsilon )]dF(\varepsilon
,\alpha ).
\end{equation*}
This is an unusual object in the binary choice literature but is equal to a
conditional mean ATE. In particular, if $\varepsilon _{it}$ is independent
of $(X_{i},\alpha _{i})$ with CDF $H(\varepsilon )$ for each $t$. Then $
E[Y_{it}|X_{i},\alpha _{i}]=\Pr (Y_{it}=1|X_{i},\alpha
_{i})=H(X_{it}^{\prime }\beta ^{\ast }+\alpha _{i})$ and
\begin{equation*}
\Delta =\int [H(x^{a\prime }\beta ^{\ast }+\alpha )-H(x^{b\prime }\beta
^{\ast }+\alpha )]dF(\alpha ).
\end{equation*}
Thus the ATE is also the effect of changing $x$ on the choice probabilities
averaged over the individual effect, i.e. the conditional mean ATE.

Our model also includes binary choice with individual-specific slopes as a
special case. Economic motivation for varying slopes is provided by Browning
and Carro (2007, 2009) who point out that with constant slopes the sign of
the treatment effect is the same for every individual and give empirical
examples where varying slopes are important. A general model with varying
slopes is $Y_{it}=1(X_{it}^{\prime }\alpha _{i}\geq \varepsilon _{it})$
where $X_{it}$ now includes a constant and $\varepsilon _{it}$ is
independent of $(X_{i},\alpha _{i})$ with CDF $H(\varepsilon ).$ In this
model
\begin{equation*}
\Delta =\int [H(x^{a\prime }\alpha )-H(x^{b\prime }\alpha )]dF(\alpha ),
\end{equation*}
accounting for individual specific slopes. When $X_{it}$ is discrete and
fully saturated (e.g. consists of a full set of dummies, one for every
discrete outcome) this model is actually equivalent to the general static
model. It will be more restrictive when the distribution of $\alpha $ is
restricted in some way, such as having some components of $\alpha $ be
constant. In the semiparametric analysis described below we show how to
impose such restrictions.

Time effects are clearly important in practice but identification of
treatment effects will preclude including $t$ among the regressors $X_{it}$
in the nonparametric model of Assumptions 1 - 3. Identification will be
based on variation over time in $X_{it}$, and if $t$ is a regressor then $
g_{0}(X_{it},\alpha _{i},\varepsilon _{it})$ has unrestricted variation over
time, precluding identification of the effect of any other regressor. Some
time effects can be allowed for by restricting the way $t$ enters $g_{0}.$
Below we will describe how this is done in semiparametric, discrete-choice
models. With continuous $Y_{it}$ one can allow for location and scale time
effects that are relatively easy to estimate.

\bigskip

\textsc{Assumption 4:} \textit{There is a function }$g_{0}(x,\alpha
,\varepsilon ),$\textit{\ vectors }$\alpha _{i}$\textit{\ and }$\varepsilon
_{it}$\textit{\ of random variables, and constants }$\tau
_{t},s_{t},(t=2,...,T)$\textit{\ such that for }$\tau _{1}=0,$ $s_{1}=1,$
\begin{equation*}
Y_{it}=g_{t0}(X_{it},\alpha _{i},\varepsilon _{it}),\text{ }g_{t0}(x,\alpha
,\varepsilon )=\tau _{t}+s_{t}g_{0}(x,\alpha ,\varepsilon
),(i=1,...,n;t=1,...,T).
\end{equation*}

This condition allows the mean and variance of $Y_{it}$ to vary over time in
an unrestricted way. The condition could be generalized to allow for other
time effects, but we leave that to future work. It does not apply to $Y_{it}$
with fixed, discrete support because Assumption 4 does not make sense in
that case. There $t$ must be included \textquotedblleft
inside\textquotedblright\ $g_{0}$, as we do in the semiparametric analysis
described below.

With these time effects the ASF and QSF can depend on $t$. The ASF and QSF
for the first period will be $\mu (x)$ and $q(\lambda ,x)$ as given above,
and for the other periods are
\begin{equation*}
\mu _{t}(x)=\tau _{t}+s_{t}\mu (x),q_{t}(\lambda ,x)=\tau
_{t}+s_{t}q(\lambda ,x),(t=2,...,T).
\end{equation*}
Corresponding period-specific and time-averaged ATE and QTE are given by
\begin{eqnarray}
\mu _{t}(x^{a})-\mu _{t}(x^{b}) &=&s_{t}[\mu (x^{a})-\mu
(x^{b})],q_{t}(\lambda ,x^{a})-q_{t}(\lambda ,x^{b})=s_{t}[q(\lambda
,x^{a})-q(\lambda ,x^{b})],  \label{tqte} \\
&&\left( \frac{\sum_{t=1}^{T}s_{t}}{T}\right) [\mu (x^{a})-\mu
(x^{b})],\left( \frac{\sum_{t=1}^{T}s_{t}}{T}\right) [q(\lambda
,x^{a})-q(\lambda ,x^{b})],  \notag
\end{eqnarray}
where $s_{1}=1$.

In the rest of this paper we will focus on discrete regressors, imposing the
following condition from here on:

\bigskip

\textsc{Assumption 5:} \textit{The support of }$X_{i}$\textit{\ is finite}$.$

\bigskip

With discrete $X_{it}$ the model can also be written as a multiple
regression with random coefficients, though we find it convenient to use the
notation given here.

\section{Identified Effects in the Nonparametric Static Model}

The analysis of identification in the static model is quite simple. This
simplicity is a virtue, leading to estimators of identified effects and
bounds on unidentified effects that are easy to calculate in a very general
model. For example, this approach gives a simple solution to the important
problem of identification of quantile treatment effects in panel data. The
idea is based on Assumption 2, which states that, conditional on $X_{i},$
the distribution of unobservables does not vary over time. Therefore,
conditional on $X_{i}$ where both $x^{b}$ and $x^{a}$ occur for some time
periods, one can identify effects from the changes in $Y_{it}$ across those
time periods. For the ATE, the identified conditional effects can be
averaged to identify effects conditional on $X_{i}$ being in subsets where
both $x^{b}$ and $x^{a}$ occur for some time period. This idea is a slight
extension of Chamberlain (1982, pp. 10-17) to discrete regressors that are
not binary. For the QTE the distribution functions can be averaged and
inverted to identify corresponding quantile effects. This idea appears to be
novel.

There is a simple approach to allowing for covariates. Suppose $
x=(x_{1},x_{2}),$ and one is interested in the effect of $x_{1}$ holding $
x_{2}$ fixed. Then one can take $x^{b}=(x_{1}^{b},x_{2})$ and $
x^{a}=(x_{1}^{a},x_{2}),$ so that the effect of changing from $x^{b}$ to $
x^{a}$ is then the effect of interest. Furthermore, one could average these
effects over $x_{2}$ to identify an effect that is averaged over covariates.
We explicitly allow for covariates in the semiparametric models given below.
Because we are already attempting to cover so much ground here, we leave
averaging over covariates in the nonparametric model to future work.

To describe identified effects and their estimators we will focus on the ATE
and QTE conditional on both $x^{a}$ and $x^{b}$ appearing in $X_{i}$ for
some time period. We could also consider effects conditional on smaller
subsets of $X_{i}$ but postpone this until later in order to keep the
exposition relatively simple. We need a little more notation to give a
precise description. Let $1(X_{it}=x)$ denote the indicator function that is
equal to one when $X_{it}=x$ and zero otherwise and let $T_{i}(x)=
\sum_{t=1}^{T}1(X_{it}=x).$ Here we let the subscript $i$ denote a random
variable that may depend on $X_{i}$ and $Y_{i}$. Let $
D_{i}=1(T_{i}(x^{a})>0)1(T_{i}(x^{b})>0)$ be the indicator for the event
that $X_{i}$ includes both $x^{a}$ and $x^{b}$ for some time period. Define
\begin{equation}
\delta =E[g_{0}(x^{a},\alpha _{i},\varepsilon _{i1})-g_{0}(x^{b},\alpha
_{i},\varepsilon _{i1})|D_{i}=1].  \label{ATE cond}
\end{equation}
This $\delta $ is the ATE for those individuals where both $x^{b}$ and $
x^{a} $ occur for some time period. This effect may be of interest in many
settings. For example, when $Y_{it}$ is log earnings and $X_{it}\in \{0,1\}$
represents union status, $\delta $ would be the average effect of union
status on earnings for those who changed union status over the time periods
we observe. For a given number of time periods $T,$ this is all one could
hope to identify nonparametrically. However, we may be interested in other
effects too. We might be interested in union effects for those who ever
changed union status at some time. \ This is $\delta $. \ Or we might even
be interested in the effect for those who were ever in a union. Bounds for
such an effect are described below.

A simple estimator of the conditional ATE $\delta $ is
\begin{equation}
\hat{\delta}=\frac{\sum_{i=1}^{n}D_{i}[\bar{Y}_{i}(x^{a})-\bar{Y}_{i}(x^{b})]
}{\sum_{i=1}^{n}D_{i}},\bar{Y}_{i}(x)=\left\{
\begin{array}{c}
T_{i}(x)^{-1}\sum_{t=1}^{T}1(X_{it}=x)Y_{it},T_{i}(x)>0 \\
0,T_{i}(x)=0
\end{array}
\right. .  \label{ideff}
\end{equation}
Consistency of this estimator results from
\begin{equation*}
E[D_{i}\{\bar{Y}_{i}(x^{a})-\bar{Y}_{i}(x^{b})\}]=E[D_{i}\{g_{0}(x^{a},
\alpha _{i},\varepsilon _{i1})-g_{0}(x^{b},\alpha _{i},\varepsilon _{i1})\}],
\end{equation*}
see Lemma A5 of the Supplementary Material. Intuitively, this equation
follows from time being randomly assigned, so that we can estimate the
effect by comparing $Y_{it}$ where $X_{it}=x^{a}$ with $Y_{is}$ where $
X_{is}=x_{b}$.

Since $\bar{Y}_{i}(x^{a})-\bar{Y}_{i}(x^{b})$ is a difference of means it
can be interpreted as a coefficient of $1(X_{it}=x^{a})$ in a regression of $
Y_{it}$ on that dummy and on $1(X_{it}=x^{a})+1(X_{it}=x^{b})$. Thus, $\hat{
\delta}$ is an average of least-squares estimates for each $i$ with $D_{i}=1$
. From this interpretation we see that $\hat{\delta}$ extends Chamberlain's
(1982, p. 12) estimator to discrete regressors that are not binary. A
consistent estimator of the asymptotic variance of $\sqrt{n}(\hat{\delta}
-\delta )$ is $n^{-1}\sum_{i=1}^{n}\hat{\psi}_{i}^{2}$ where $\hat{\psi}
_{i}=nD_{i}[\bar{Y}_{i}(x^{a})-\bar{Y}_{i}(x^{b})-\hat{\delta}
]/\sum_{i=1}^{n}D_{i}$. For brevity we leave the asymptotic theory to the
Supplementary Material (see Theorem A6) and efficiency results to future
work.

We can also identify and estimate a conditional QTE. Let $G(y,x|D_{i}=1)=\Pr
(g_{0}(x,\alpha _{i},\varepsilon _{i1})\leq y|D_{i}=1)$ denote the CDF of $
g_{0}(x,\alpha _{i},\varepsilon _{i1})$ conditional on $D_{i}=1$. The QTE
conditional on $D_{i}=1$ is
\begin{equation*}
\delta _{\lambda }=G^{-1}(\lambda ,x^{a}|D_{i}=1)-G^{-1}(\lambda
,x^{b}|D_{i}=1).
\end{equation*}
An estimator of this effect can be constructed using a CDF $\Phi (u)$ and a
scalar bandwidth $h$. An estimator of $G(y,x|D_{i}=1)$ is given by
\begin{equation*}
\hat{G}(y,x|D_{i}=1)=\frac{\sum_{i=1}^{n}D_{i}\bar{G}_{i}(y,x)}{
\sum_{i=1}^{n}D_{i}},\bar{G}_{i}(y,x)=\left\{
\begin{array}{c}
T_{i}(x)^{-1}\sum_{t=1}^{T}1(X_{it}=x)\Phi (\frac{y-Y_{it}}{h}),T_{i}(x)>0,
\\
0,T_{i}(x)=0.
\end{array}
\right. .
\end{equation*}
In this estimator the indicator function $1(Y_{it}<y)$ has been replaced by
a smoothed approximation $\Phi (\frac{y-Y_{it}}{h})$, as suggested by Yu and
Jones (1998) for estimating a conditional CDF. An estimator of $\delta
_{\lambda }$ is then
\begin{equation*}
\hat{\delta}_{\lambda }=\hat{q}_{\lambda }^{a}-\hat{q}_{\lambda }^{b},\hat{q}
_{\lambda }^{a}=\hat{G}^{-1}(\lambda ,x^{a}|D_{i}=1),\hat{q}_{\lambda }^{b}=
\hat{G}^{-1}(\lambda ,x^{b}|D_{i}=1).
\end{equation*}
Note here that we first average, then invert, and then difference. This
estimator solves an important problem of estimating panel quantile effects
and appears to be novel.

A consistent estimator of the asymptotic variance of $\sqrt{n}(\hat{\delta}
_{\lambda }-\delta _{\lambda })$ is $n^{-1}\sum_{i=1}^{n}\hat{\psi}_{\lambda
i}^{2}$ for
\begin{equation*}
\hat{\psi}_{\lambda i}= - \frac{nD_{i}}{\sum_{i=1}^{n}D_{i}}\left[ \frac{\bar{G}
_{i}(\hat{q}^{a},x^{a})-\lambda }{\hat{G}^{\prime }(\hat{q}
^{a},x^{a}|D_{i}=1)}-\frac{\bar{G}_{i}(\hat{q}^{b},x^{b})-\lambda }{\hat{G}
^{\prime }(\hat{q}^{b},x^{b}|D_{i}=1)}\right] ,
\end{equation*}
where $\hat{G}^{\prime }(y,x|D_{i}=1)=\partial \hat{G}(y,x|D_{i}=1)/\partial
y$. Here the denominator terms are actually kernel density estimates. For
this reason one might use different bandwidths $h$ in the numerator and
denominator, with the denominator chosen to be appropriate for density
estimation. Asymptotic theory for this estimator is given in the
Supplementary Material (see Theorem A8). Alternatively, one could simply use
the bootstrap to construct a confidence interval for $\hat{\delta}_{\lambda
}.$

A helpful example is the binary regressor case where $X_{it}\in \{0,1\}.$
Here $X_{it}$ could be thought of as a treatment variable where $X_{it}=1$
for treated and $X_{it}=0$ for untreated. Let $Y_{it}(0)=g_{0}(0,\alpha
_{i},\varepsilon _{it})$ and $Y_{it}(1)=g_{0}(1,\alpha _{i},\varepsilon
_{it})$. Assumption 2 is equivalent to the assumption that the conditional
distribution of $(Y_{it}(0),Y_{it}(1))$ given $X_{i}$ does not vary with $t$
. This is the key assumption that identifies treatment effects from time
variation in treatment. In this context $\delta
=E[Y_{it}(1)-Y_{it}(0)|D_{i}=1]$ is the ATE for individuals where both
treatment and nontreatment occurs during the observation period. Similarly, $
\delta _{\lambda }$ is the difference between the $\lambda $ quantile of the
distribution of $Y_{it}(1)$ and the $\lambda $ quantile for $Y_{it}(0)$
conditional on $D_{i}=1$. The ATE and QTE are not identified for those
individuals that either receive treatment in every time period or receive no
treatment in every time period.

In general the usual panel data within (linear fixed effects) estimator is
not a consistent estimator of $\delta .$ This inconsistency results because
the within estimator constrains the slope coefficient to be the same for
each $i$ when the slope is actually varying with $i$. For simplicity we
demonstrate this inconsistency in the binary $X_{it}$ example. The within
estimator $\hat{\delta}_{w}$ is given by
\begin{equation*}
\hat{\delta}_{w}=\frac{\sum_{i=1}^{n}\sum_{t=1}^{T}(X_{it}-\bar{X}_{i})Y_{it}
}{\sum_{i=1}^{n}\sum_{t=1}^{T}(X_{it}-\bar{X}_{i})^{2}},\bar{X}
_{i}=T^{-1} \sum_{t=1}^{T}X_{it}.
\end{equation*}
Let $\sigma _{i}^{2}=(T-1)^{-1}\sum_{t=1}^{T}(X_{it}-\bar{X}_{i})^{2}$ be the
sample variance over time of $X_{it}$.

\bigskip

\textsc{Theorem 1:} \textit{If Assumptions 1 and 2 are satisfied, } $X_{it} \in \{0,1\}$, $
E[Y_{it}^{2}]<\infty ,$\textit{\ }$(t=1,...,T)$\textit{, and }$E[D_{i}\sigma
_{i}^{2}]>0$, \textit{then }$\delta =E[D_{i}\{\bar{Y}_{i}(1)-\bar{Y}
_{i}(0)\}]/E[D_{i}]$ \textit{and}
\begin{equation}
\hat{\delta}_{w}\overset{p}{\longrightarrow }\delta _{w}=\frac{E[\sigma
_{i}^{2}D_{i}\{\bar{Y}_{i}(1)-\bar{Y}_{i}(0)\}]}{E[\sigma _{i}^{2}D_{i}]}.
\label{lpm}
\end{equation}

Note that the limit of the within estimator is a weighted average of
individual, least-squares estimates $\bar{Y}_{i}(1)-\bar{Y}_{i}(0)$ from
equation (\ref{ideff}). If $T\geq 4$ then the weights $\sigma _{i}^{2}$ vary
over the positive $\sigma _{i}^{2}$ and so the limit $\delta _{w}$ of $\hat{
\delta}_{w}$ is not the identified conditional ATE $\delta $.

Theorem 1 is different than Yitzhaki (1996) and Angrist (1998), who gave
weighted average interpretations of least squares in other, non-panel
settings. Theorem 1 is also different from Hahn (2001), who found that $\hat{
\delta}_{w}$ consistently estimates the ATE. Hahn (2001) considered $T=2$
and assumed $X_{i}=(0,1)^{\prime }$. As noted by Hahn (2001), those
conditions are quite special. Theorem 1 is also different from Wooldridge
(2005), who showed that if $b_{i}=E[Y_{it}(1)-Y_{it}(0)|\alpha _{i}]$ is
mean independent of $X_{it}-\bar{X}_{i}$ for each $t$ then linear fixed
effects is a consistent estimator of $\delta $. The problem is that the
mean-independence assumption is very strong when $X_{it}$ is discrete. For
instance, if $T=2$, $X_{i2}-\bar{X}_{i}$ takes on the values $0$ when $
X_{i}=(1,1)$ or $(0,0)$, $-1/2$ when $X_{i}=(1,0)\,,$ and $1/2$ when $
X_{i}=(0,1)$. Thus mean independence of $b_{i}$ and $X_{i2}-\bar{X}_{i}$
actually implies that
\begin{equation*}
E[b_{i}|X_{i}=(1,0)^{\prime }]=E[b_{i}|X_{i}=(0,1)^{\prime
}]=E[b_{i}|X_{i}\in \{(0,0)^{\prime },(1,1)^{\prime }\}].
\end{equation*}
This is quite close to independence of $b_{i}$ and $X_{i}$, which is not
very interesting if we want to allow the treatment effect to vary with $
X_{i} $.

The conditional ATE and QTE estimators can easily be modified to accommodate
the time effects of Assumption 4. The changes in $Y_{it}$ over time for
fixed $X_{it}$ can be used to identify and estimate the time effects that
can then be included in the estimation of the ATE and QTE. To describe this
approach, let $\hat{m}_{t}=\sum_{i=1}^{n}1(X_{it}=X_{i1})Y_{it}/
\sum_{i=1}^{n}1(X_{it}=X_{i1})$ and
\begin{equation*}
\hat{s}_{t}=\frac{\sum_{i=1}^{n}1(X_{it}=X_{i1})X_{i1}(Y_{it}-\hat{m}_{t})}{
\sum_{i=1}^{n}1(X_{it}=X_{i1})X_{i1}(Y_{i1}-\hat{m}_{1})},\hat{\tau}_{t}=
\hat{m}_{t}-\hat{s}_{t}\hat{m}_{1},t=2,...,T.
\end{equation*}
This $(\hat{\tau}_{t},\hat{s}_{t})^{\prime }$ is an instrumental variables
estimator where the residual is $Y_{it}-\tau _{t}-s_{t}Y_{i1}$, the
instruments are $(1,X_{i1})^{\prime }$, and the estimation is done on the
subsample where $X_{it}=X_{i1}$. These estimators will be consistent and
asymptotically normal as long as $Cov(X_{i1},Y_{i1}|X_{it}=X_{i1})\neq 0$
for each $t=2,...,T.$ One could also use other functions of $X_{i1}$ as
instrumental variables to improve efficiency. We focus on just $X_{i1}$ as
an instrument for simplicity. Graham and Powell (2011) use a similar
approach to identify time effects in a linear model with continuous
regressors.

The time effects are accounted for in ATE and QTE estimation by removing
time location and scale effects from all periods when estimating the first
period effect, and then putting the scale effects back for other periods.
Note first that under Assumption 4 $\delta $ is the conditional ATE for the
first time period. Let $\tilde{Y}_{it}=(Y_{it}-\hat{\mu}_{t})/\hat{s}_{t}$
be the $t^{th}$ period observation with estimated location and scale
removed. Replacing $Y_{it}$ by $\tilde{Y}_{it}$ in the formula for $\hat{
\delta}$ gives
\begin{equation*}
\tilde{\delta}=\frac{\sum_{i=1}^{n}D_{i}[\tilde{Y}_{i}(x^{a})-\tilde{Y}
_{i}(x^{b})]}{\sum_{i=1}^{n}D_{i}},\tilde{Y}_{i}(x)=\left\{
\begin{array}{c}
T_{i}(x)^{-1}\sum_{t=1}^{T}1(X_{it}=x)\tilde{Y}_{it},T_{i}(x)>0 \\
0,T_{i}(x)=0
\end{array}
\right. .
\end{equation*}
The conditional ATE for the $t^{th}$ time period is given by $s_{t}\delta $
and a time average by $(\sum_{t=1}^{T}s_{t}/T)\delta $ for $s_{1}=1$,
analogously to equation (\ref{tqte}). These can be estimated by $\hat{s}_{t}
\tilde{\delta}$ and $\bar{s}\tilde{\delta},$ respectively for $\bar{s}
=\sum_{t=1}^{T}\hat{s}_{t}/T$ and $\hat s_{1}=1.$ These estimators will be
consistent and asymptotically normal. Because of their multistage nature the
bootstrap may provide the easiest approach to carrying out inference on
these estimators, where one resamples from the empirical distribution of $
(Y_{i},X_{i}),$ $(i=1,...,n)$ to form confidence intervals for the true
parameter. For brevity we omit explicit results.

An analogous approach can be followed to account for time effects in the
QTE. The interpretation of $\delta _{\lambda }$ now becomes QTE for the
first time period conditional on $D_{i}=1$. An estimator of $
G(y,x|D_{i}=1)=\Pr (g_{0}(x,\alpha _{i},\varepsilon _{i1})\leq y|D_{i}=1)$
that adjusts for time, location and scale is given by
\begin{equation*}
\tilde{G}(y,x|D_{i}=1)=\frac{\sum_{i=1}^{n}D_{i}\tilde{G}_{i}(y,x)}{
\sum_{i=1}^{n}D_{i}},\tilde{G}_{i}(y,x)=\left\{
\begin{array}{c}
T_{i}(x)^{-1}\sum_{t=1}^{T}1(X_{it}=x)\Phi (\frac{y-\tilde{Y}_{it}}{h}
),T_{i}(x)>0 \\
0,T_{i}(x)=0
\end{array}
\right. .
\end{equation*}
Let $\tilde{q}_{\lambda }^{a}=\tilde{G}^{-1}(\lambda ,x^{a}|D_{i}=1)$ and $
\tilde{q}_{\lambda }^{b}=\tilde{G}^{-1}(\lambda ,x^{b}|D_{i}=1).$ Estimators
for the conditional QTE for the first period, other periods, and a time
average are given by $\tilde{\delta}_{\lambda }=\tilde{q}_{\lambda }^{a}-
\tilde{q}_{\lambda }^{b}$, $\hat{s}_{t}\tilde{\delta}_{\lambda
},(t=2,...,T), $ and $\bar{s}\tilde{\delta}_{\lambda }$, respectively. Here
again the bootstrap provides a convenient method for inference. One could
also use quantiles to estimate the time effects, but we avoid that for
simplicity.

\section{Nonparametric Bounds in the Static Model}

When $g_{0}(x,\alpha _{i},\varepsilon _{it})$ is bounded we can estimate
bounds for the ASF and corresponding bounds for the ATE. For the QSF and QTE
we can also estimate bounds without any restriction on $g_{0}$, using the
fact that there are known upper and lower bounds for the indicator function $
1(g_{0}(x,\alpha _{i},\varepsilon _{it})\leq y).$ The idea of the bounds is
an extension of the estimation of identified effects discussed in the
previous Section. Time homogeneity allows us to use time averages to
estimate the identified parts of the ASF or QSF when $x$ is an element of $
X_{i},$ i.e. $X_{it}=x$ for some $t$, and apply the lower or upper bounds
when $x$ does not appear in $X_{i}$.

We first describe bounds estimation for the ASF. These bounds depend on
bounds on $g_{0}$ imposed in the following condition:

\bigskip

\textsc{Assumption 6: }$B_{\ell }\leq g_{0}(x,\alpha _{i},\varepsilon
_{it})\leq B_{u}$ \textit{for constants }$B_{\ell }$\textit{\ and }$B_{u}$
\textit{and all }$x.$

\bigskip

For example, in the binary-choice model, where $Y_{it}\in \{0,1\}$, upper
and lower bounds are $B_{u}=1$ and $B_{\ell }=0$ respectively. We could
allow $B_{\ell }$\textit{\ }and $B_{u}$ to depend on $x$ and using that
information could tighten the ATE bounds given below. To avoid further
complication we do not allow this.

Let $T_{i}(x)$ and $\bar{Y}_{i}(x)$ be as in Section 3 and $\bar{P}
(x)=\sum_{i=1}^{n}1(T_{i}(x)=0)/n$ be the sample frequency of $x$ not
occurring in any time period. Estimated lower and upper bounds for $\mu (x)$
are
\begin{equation*}
\hat{\mu}_{\ell }(x)=n^{-1}\sum_{i=1}^{n}\bar{Y}_{i}(x)+\bar{P}(x)B_{\ell },
\hat{\mu}_{u}(x)=\hat{\mu}_{\ell }(x)+\bar{P}(x)(B_{u}-B_{\ell }).
\end{equation*}
Here $\bar{Y}_{i}(x)$ estimates the identified part of the ASF,
corresponding to $T_{i}(x)>0$, and the upper and lower bounds are applied
for observations where $T_{i}(x)=0\,.$ Corresponding estimated lower and
upper bounds for the ATE are $\hat{\Delta}_{\ell }=\hat{\mu}_{\ell }(x^{a})-
\hat{\mu}_{u}(x^{b})$ and $\hat{\Delta}_{u}=\hat{\mu}_{u}(x^{a})-\hat{\mu}
_{\ell }(x^{b}).$ The width of these estimated bounds is
\begin{equation*}
\hat{\Delta}_{u}-\hat{\Delta}_{\ell }=[\bar{P}(x^{a})+\bar{P}
(x^{b})](B_{u}-B_{\ell }).
\end{equation*}
For example, for binary choice with a binary regressor, where $B_{u}=1$ and $
B_{\ell }=0,$ the width of the estimated bounds for the ATE\ is $\bar{P}(0)+
\bar{P}(1),$ where $\bar{P}(0)$ and $\bar{P}(1)$ are the sample proportions
of $X_{i}$ with $X_{it}=1$ for all $t$ and $X_{it}=0$ for all $t$,
respectively

These estimators will be jointly asymptotically normal under i.i.d. $
(Y_{i},X_{i})$. The asymptotic variance can be estimated by $\hat{\Sigma}
=\sum_{i=1}^{n}\hat{\Psi}_{i}\hat{\Psi}_{i}^{\prime }/n$, where
\begin{equation*}
\hat{\Psi}_{i}=\left(
\begin{array}{c}
\bar{Y}_{i}(x^{a})-\bar{Y}_{i}(x^{b})+B_{\ell
}1(T_{i}(x^{a})=0)-B_{u}1(T_{i}(x^{b})=0)-\hat{\Delta}_{\ell } \\
\bar{Y}_{i}(x^{a})-\bar{Y}_{i}(x^{b})+B_{u}1(T_{i}(x^{a})=0)-B_{\ell
}1(T_{i}(x^{b})=0)-\hat{\Delta}_{u}
\end{array}
\right) .
\end{equation*}
Confidence intervals for the identified set can then be formed using results
of Chernozhukov, Hong, and Tamer (2007) or Beresteanu and Molinari (2008,
pp. 779-781) on estimators of intervals where the upper and lower endpoints
are jointly asymptotically normal.

Turning to the bounds for the QSF, lower and upper estimated bounds for the $
G(y,x)=\Pr (g_{0}(x,\alpha _{i},\varepsilon _{i1})\leq y)$ are $\hat{G}
_{\ell }(y,x)=\sum_{i=1}^{n}\bar{G}_{i}(y,x)/n$ and $\hat{G}_{u}(y,x)=\hat{G}
_{\ell }(y,x)+\bar{P}(x)$ respectively. The idea of these bounds is similar
to the ASF, with a known lower bound of $0$ and upper bound of $1$ for $
1(g_{0}(x,\alpha _{i},\varepsilon _{i1})\leq y)$. To obtain quantile bounds
we need to invert these functions of $y$. For a strictly
increasing function $G(y)$ with range contained in $[0,1]$ let
\begin{equation*}
Q(\lambda ,G(\cdot ))=\left\{
\begin{array}{l}
-\infty ,\text{ }\lambda \leq \inf_{y}G(y) \\
\multicolumn{1}{c}{G^{-1}(\lambda ),\text{ }\inf_{y}G(y)<\lambda
<\sup_{y}G(y)} \\
+\infty ,\text{ }\lambda \geq \sup_{y}G(y)
\end{array}
\right. .
\end{equation*}
This is a function with domain $[0,1]$ and range equal to the extended real
line that can be used to invert $\hat{G}_{u}(y,x)$ and $\hat{G}_{\ell
}(y,x). $ Estimators of lower and upper bounds on the QSF are given by
\begin{equation*}
\hat{q}_{\ell }(\lambda ,x)=Q(\lambda ,\hat{G}_{u}(\cdot ,x)),\hat{q}
_{u}(\lambda ,x)=Q(\lambda ,\hat{G}_{\ell }(\cdot ,x)).
\end{equation*}
Corresponding lower and upper bounds for the QTE are $\hat{\Delta}_{\lambda
\ell }=\hat{q}_{\ell }^{a}-\hat{q}_{u}^{b}$ and $\hat{\Delta}_{\lambda u}=
\hat{q}_{u}^{a}-\hat{q}_{\ell }^{b}$ where $\hat{q}_{\ell }^{a}=\hat{q}
_{\ell }(\lambda ,x^{a}),$ $\hat{q}_{u}^{a}=\hat{q}_{u}(\lambda ,x^{a}),$ $
\hat{q}_{\ell }^{b}=\hat{q}_{\ell }(\lambda ,x^{b}),$ and $\hat{q}_{u}^{b}=
\hat{q}_{u}(\lambda ,x^{b}).$ The width of these bounds depends on the shape
of the empirical distribution of $Y_{it}$ and on $\bar{P}(x).$ The width of
the bounds will be finite when
\begin{equation}
\max \{\bar{P}(x^{a}),\bar{P}(x^{b})\}<\lambda <\min \{1-\bar{P}(x^{a}),1-
\bar{P}(x^{b})\},  \label{quant ineq}
\end{equation}
and otherwise they are infinitely wide.

The bounds will be joint asymptotically normal under the following
regularity condition:

\bigskip

\textsc{Assumption 7: }$\Pr (g_{0}(x,\alpha _{i},\varepsilon _{i1})\leq
y|X_{i})$ \textit{is twice continuously differentiable in }$y$ \textit{with
uniformly bounded derivatives and }$G_{\ell }(y,x)=E[E[1(T_{i}(x)>0)|X_{i}]\Pr
(g_{0}(x,\alpha _{i},\varepsilon _{i1})\leq y|X_{i})]$ \textit{is strictly
increasing in }$y$\textit{\ on the interior of its range for all }$x$.
\textit{Also }$nh^{4}\longrightarrow 0$ and $nh^{2}\longrightarrow \infty $.

\bigskip

For $\lambda $ satisfying equation (\ref{quant ineq}) the asymptotic
variance can be estimated by $\hat{\Sigma}_{\lambda }=\sum_{i=1}^{n}\hat{\Psi
}_{\lambda i}\hat{\Psi}_{\lambda i}^{\prime }/n$, where
\begin{equation*}
\hat{\Psi}_{\lambda i}=\left(
\begin{array}{c}
\frac{\bar{G}_{i}(\hat{q}_{\ell }^{a},x^{a})+1(T_{i}(x^{a})=0)-\lambda }{
\hat{G}_{\ell }^{\prime }(\hat{q}_{\ell }^{a},x^{a})}-\frac{\bar{G}_{i}(\hat{
q}_{u}^{b},x^{b})-\lambda }{\hat{G}_{\ell }^{\prime }(\hat{q}_{u}^{b},x^{b})}
\\
\frac{\bar{G}_{i}(\hat{q}_{u}^{a},x^{a})-\lambda }{\hat{G}_{u}^{\prime }(
\hat{q}_{u}^{a},x^{a})}-\frac{\bar{G}_{i}(\hat{q}_{\ell
}^{b},x^{b})+1(T_{i}(x^{b})=0)-\lambda }{\hat{G}_{\ell }^{\prime }(\hat{q}
_{\ell }^{b},x^{b})}
\end{array}
\right) .
\end{equation*}
As in estimation of the conditional quantile effect, one might want to use
different bandwidths for numerators and denominators, or just bootstrap to
estimate the asymptotic variance.

Here is a result for both ATE and QTE bounds:

\bigskip

\textsc{Theorem 2:} \textit{Suppose that Assumptions 1, 2, and 5 are
satisfied.\ If Assumption 6 is satisfied} \textit{then there are }$\Delta
_{\ell },$\textit{\ }$\Delta _{u},$\textit{\ and }$\Sigma $ \textit{such
that } \textit{\ }
\begin{equation*}
\sqrt{n}[(\hat{\Delta}_{\ell },\hat{\Delta}_{u})^{\prime }-(\Delta _{\ell
},\Delta _{u})^{\prime }]\overset{d}{\longrightarrow }N(0,\Sigma ),\hat{
\Sigma}\overset{p}{\longrightarrow }\Sigma .
\end{equation*}
\textit{\ where }$\Delta _{\ell }\leq \Delta \leq \Delta _{u}$, \textit{and
these bounds are sharp. If Assumption 7 is satisfied then there are }$\Delta
_{\lambda \ell },$\textit{\ }$\Delta _{\lambda u},$\textit{\ and }$\Sigma
_{\lambda }$ \textit{such that}
\begin{equation*}
\sqrt{n}[(\hat{\Delta}_{\lambda \ell },\hat{\Delta}_{\lambda u})^{\prime
}-(\Delta _{\lambda \ell },\Delta _{\lambda u})^{\prime }]\overset{d}{
\longrightarrow }N(0,\Sigma _{\lambda }),\hat{\Sigma}_{\lambda }\overset{p}{
\longrightarrow }\Sigma _{\lambda }.
\end{equation*}
\textit{\ where }$\Delta _{\lambda \ell }\leq \Delta _{\lambda }\leq \Delta
_{\lambda u}$\textit{. If }$G_{\ell }(y,x)$ \textit{is also everywhere
strictly increasing in }$y$ \textit{then these bounds are sharp.}

\bigskip

The sharpness conclusion of Theorem 2 for the ATE depends on being able to
let $g_{0}(x,\alpha _{i},\varepsilon _{it})$ take any value between $B_{\ell
}$ and $B_{u}.$ That is not possible for binary choice, where the outcome is
restricted to zero or one. Nevertheless the bounds can still be shown to be
sharp.

Similarly to the treatment-effects literature, we may be interested in the
ATE or QTE, conditional on $X_{i}\in S$ for some set $S$. For example, if $
X_{it}\in \{0,1\}$ represents treatment then we might be interested in the
effect of treatment conditional on ever treated, i.e. conditional on $
X_{i}\neq (0,...,0)^{\prime }$. Tighter bounds for such effects can be
formed and in some cases the effects may be identified. These bounds can be
estimated by replacing $1(X_{it}=x)$ by $1(X_{i}\in S)1(X_{it}=x)$ in the
definition of $\bar{Y}_{i}(x)$ and $\bar{G}_{i}(y,x)$, $1(T_{i}(x)=0)$ by $
1(X_{i}\in S)1(T_{i}(x)=0)$ in the definition of $\bar{P}(x)$, and dividing
through by $\sum_{i=1}^{n}1(X_{i}\in S)/n.$ If $1(X_{i}\in S)\leq D_{i}$ for
$D_{i}$ from Section 3 the corresponding effects will be identified, and the
upper and lower estimated bounds will be identical.

Time effects can easily be allowed for in quantile-effect bounds by adapting
the approach used earlier. It is not clear that allowing for time effects in
that way makes sense for bounds on the ATE, e.g. for binary choice models
where the support of $Y_{it}$ is fixed. Therefore we focus just on time
effects in quantile bounds. For QTE bounds we can replace $Y_{it}$ by $
\tilde{Y}_{it}=(Y_{it}-\hat{\mu}_{t})/\hat{s}_{t}$ in the formula for $\hat{G
}_{\ell }(y,x)$ given above, and interpret $\hat{\Delta}_{\lambda \ell }$
and $\hat{\Delta}_{\lambda u}$ as estimators of the first period bounds.
Estimators of $t^{th}$ period lower and upper bounds for the QTE are then
given by $\hat{s}_{t}\hat{\Delta}_{\lambda \ell }$ and $\hat{s}_{t}\hat{
\Delta}_{\lambda u}$ respectively. Estimators of time average bounds are $
\bar{s}\hat{\Delta}_{\lambda \ell }$ and $\bar{s}\hat{\Delta}_{\lambda u},$
where $\bar{s}=\sum_{t=1}^{T}\hat{s}_{t}/T.$ These upper and lower bounds
will be joint asymptotically normal, and their asymptotic variance can be
estimated by the bootstrap.

\section{Nonparametric Bounds in the Dynamic Model}

Analysis of the dynamic model is more challenging than that of the static
one. In the dynamic model of Assumption 3 only the first-period regressor is
common to the conditioning sets for each time period. Consequently location
and scale time effects are not identified, because the conditioning set is
different for every time period. For this reason we do not consider time
effects in the nonparametric dynamic model. Also, the identification and
bounds analysis is limited to objects that are conditional on the first
period or are unconditional. For example, we cannot identify or bound the
ATE conditional on $X_{it}$ changing over time because that event involves
information about all time periods. We can bound unconditional objects and
ones that are conditional on just $X_{i1}.$ These bounds are simple and
novel, for example in providing partial-identification results for the
average effect of state dependence with heterogeneity in both location and
slope when $Y_{it}$ is binary and $X_{it}=Y_{it-1}$.

The model with a binary, lagged dependent variable has $
Y_{it}=g_{0}(Y_{i,t-1},\alpha _{i},\varepsilon _{it})$, and under Assumption
3,
\begin{eqnarray*}
\Pr (Y_{it} =1|X_{it},...,X_{i1},\alpha _{i})&=&\int g(Y_{i,t-1},\alpha
_{i},\varepsilon )dF(\varepsilon |\alpha _{i},Y_{i0}) \\
&=&\Pr (Y_{it}=1|Y_{i,t-1},\alpha _{i},Y_{i0}),
\end{eqnarray*}
where $F(\varepsilon |\alpha _{i},Y_{i0})$ denotes the conditional CDF of $
\varepsilon _{it}$ given $\alpha _{i}$ and $Y_{i0}.$ Here $\Pr
(Y_{it}=1|Y_{i,t-1},\alpha _{i},Y_{i0})$ does not vary with $t$, and the
model places no other restrictions on $\Pr (Y_{it}=1|Y_{i,t-1},\alpha
_{i},Y_{i0})$. Conditioning on $Y_{i0}$ is present to account correctly for
the initial condition, as in Honore and Tamer (2006) and Browning and Carro
(2007, 2009). The probabilities can be distributed across individuals in any
way at all through the individual effect $\alpha _{i}$. That is we can think
of the four conditional probabilities,
\begin{equation*}
\Pr (Y_{it}=1|1,\alpha _{i},1),\Pr (Y_{it}=1|0,\alpha _{i},1)\Pr
(Y_{it}=1|1,\alpha _{i},0),\Pr (Y_{it}=1|0,\alpha _{i},0),
\end{equation*}
as having an unrestricted distribution. Here the ATE is
\begin{equation*}
\Delta =\int [\Pr (Y_{it}=1|Y_{i,t-1}=1,\alpha ,Y_{0})-\Pr
(Y_{it}=1|Y_{i,t-1}=0,\alpha ,Y_{0})]dF(\alpha ,Y_{0}).
\end{equation*}
This object quantifies the effect of state dependence in the presence of
individual heterogeneity, an important problem posed by Feller (1943) and
Heckman (1981). The dynamic bounds here provide a simple, estimable,
identified set for this object. This model is considered by Browning and
Carro (2007, 2009), who derive properties of various estimators and
restrictions on $\alpha _{i}$ that lead to identification. We give
nonparametric bounds.

A partition of $X_{i}$ values that preserves the dynamic structure of
Assumption 3 is used to obtain bounds for the ASF and QSF. For each $x$ we
partition $X_{i}$ into realizations where the first occurrence of $x$ is at
time $t$ and the set where $x$ never occurs. This partition is given by $\{
\mathcal{\bar{X}}(x),\mathcal{X}_{1}(x),...,\mathcal{X}_{T}(x)\}$ where
\begin{equation*}
\mathcal{X}_{t}(x)=\{X:X_{t}=x,\ X_{s}\neq x\ \forall s<t\},t=1,...,T;
\mathcal{\bar{X}}(x)=\{X:X_{t}\neq x\ \forall t\}.
\end{equation*}
Define $\hat{Y}_{i}(x)=\sum_{t=1}^{T}1(X_{i}\in \mathcal{X}_{t}(x))Y_{it}$,
which picks out the $Y_{it}$ for the time period where $x$ first occurs.
Estimated lower and upper ASF bounds are
\begin{equation*}
\hat{\mu}_{\ell }(x)=n^{-1}\sum_{i=1}^{n}\hat{Y}_{i}(x)+\bar{P}(x)B_{\ell },
\hat{\mu}_{u}(x)=\hat{\mu}_{\ell }(x)+\bar{P}(x)(B_{u}-B_{\ell })\text{.}
\end{equation*}
Corresponding lower and upper bounds for $\Delta $ are $\hat{\Delta}_{\ell }=
\hat{\mu}_{\ell }(x^{a})-\hat{\mu}_{u}(x^{b})$ and $\hat{\Delta}_{u}=\hat{\mu
}_{u}(x^{a})-\hat{\mu}_{\ell }(x^{b}).$ A joint asymptotic-variance
estimator $\hat{\Sigma}$ can be constructed exactly as for the static case
with $\hat{Y}_{i}(x)$ replacing $\bar{Y}_{i}(x).$

It is interesting to note that the width $\bar{P}(x)(B_{u}-B_{\ell })$ of
the estimated ASF bounds is the same for the dynamic and static models.
Because the static model is a special case of the dynamic one we conjecture
that the bounds for the dynamic model are sharp like the bounds for the
static one, but have not yet been able to show this.

To construct estimated lower and upper bounds for the CDF of $g_{0}(x,\alpha
_{i},\varepsilon _{it})$ let $\hat{G}_{i}(y,x)=\sum_{t=1}^{T}1(X_{i}\in
\mathcal{X}_{t}(x))\Phi (\frac{y-Y_{it}}{h}).$ The estimated CDF bounds are
\begin{equation*}
\hat{G}_{\ell }(y,x)=\frac{1}{n}\sum_{i=1}^{n}\hat{G}_{i}(y,x),\hat{G}
_{u}(y,x)=\hat{G}_{\ell }(y,x)+\bar{P}(x).
\end{equation*}
Estimated lower and upper bounds for the QSF are then given by
\begin{equation*}
\hat{q}_{\ell }(\lambda ,x)=Q(\lambda ,\hat{G}_{u}(\cdot ,x)),\hat{q}
_{u}(\lambda ,x)=Q(\lambda ,\hat{G}_{\ell }(\cdot ,x)).
\end{equation*}
Corresponding lower and upper bounds for the QTE are $\hat{\Delta}_{\lambda
\ell }=\hat{q}_{\ell }(\lambda ,x^{a})-\hat{q}_{u}(\lambda ,x^{b})$ and $
\hat{\Delta}_{\lambda u}=\hat{q}_{u}(\lambda ,x^{a})-\hat{q}_{\ell }(\lambda
,x^{b}).$ A joint asymptotic variance estimator $\hat{\Sigma}_{\lambda }$
can be constructed just as for the static case with $\hat{G}_{i}(y,x)$
replacing $\bar{G}_{i}(y,x).$

\bigskip

\textsc{Theorem 3:} \textit{Suppose that Assumptions 1, 3, and 5 are
satisfied.\ If Assumption 6 is satisfied} \textit{then there are }$\Delta
_{\ell },$\textit{\ }$\Delta _{u},$\textit{\ and }$\Sigma $ \textit{such that
}
\begin{equation*}
\sqrt{n}[(\hat{\Delta}_{\ell },\hat{\Delta}_{u})^{\prime }-(\Delta _{\ell
},\Delta _{u})^{\prime }]\overset{d}{\longrightarrow }N(0,\Sigma ),\hat{
\Sigma}\overset{p}{\longrightarrow }\Sigma .
\end{equation*}
\textit{\ where }$\Delta _{\ell }\leq \Delta \leq \Delta _{u}$\textit{. Also
if Assumption 7 is satisfied with }$X_{i1}$ \textit{
replacing }$X_i$ \textit{then there are }$\Delta _{\lambda \ell },$
\textit{\ }$\Delta _{\lambda u},$\textit{\ and }$\Sigma _{\lambda }$ \textit{
such that}
\begin{equation*}
\sqrt{n}[(\hat{\Delta}_{\lambda \ell },\hat{\Delta}_{\lambda u})^{\prime
}-(\Delta _{\lambda \ell },\Delta _{\lambda u})^{\prime }]\overset{d}{
\longrightarrow }N(0,\Sigma _{\lambda }),\hat{\Sigma}_{\lambda }\overset{p}{
\longrightarrow }\Sigma _{\lambda }.
\end{equation*}
\textit{\ where }$\Delta _{\lambda \ell }\leq \Delta _{\lambda }\leq \Delta
_{\lambda u}$\textit{. }

\bigskip

Similarly to the static model we may be interested in effects conditional on
$X_{i1}\in S_{1}$ for some set $S_{1}$. For example, if $X_{it}\in \{0,1\}$
represents treatment then we might be interested in the effect of treatment
conditional on being treated in the first period, i.e. conditional on $
X_{i1}=1$. Tighter bounds for such effects can be estimated by replacing $
1(X_{i}\in \mathcal{X}_{t}(x))$ by $1(X_{i1}\in S_{1})1(X_{i}\in \mathcal{X}
_{t}(x))$ in the definition of $\hat{Y}_{i}(x)$ and $\hat{G}_{i}(y,x)$, $
1(T_{i}(x)=0)$ by $1(X_{i1}\in S_{1})1(T_{i}(x)=0)$ in the definition of $
\bar{P}(x)$, and dividing through by $\sum_{i=1}^{n}1(X_{i1}\in S_{1})/n.$

In the binary, lagged-dependent-variable example we have $B_{\ell }=0$ and $
B_{u}=1$, so the bounds on the ATE are
\begin{equation*}
\hat{\Delta}_{\ell }=\frac{1}{n}\sum_{i=1}^{n}[\hat{Y}_{i}(1)-\hat{Y}
_{i}(0)]-\bar{P}(0),\hat{\Delta}_{u}=\hat{\Delta}_{\ell }+\bar{P}(1)+\bar{P}
(0).
\end{equation*}
Here $\bar{P}(1)+\bar{P}(0)$ estimates the width of the bounds, providing a
very simple measure of the severity of the problem of identifying state
dependence in the presence of heterogeneity. The bounds will tend to be wide
in short panels but more informative in long ones.

Figure 1 shows the width of corresponding population bounds in a numerical
example based on a dynamic probit model where
\begin{equation*}
Y_{it}=1(\beta ^{\ast }Y_{i,t-1}+\alpha _{i}\geq \varepsilon
_{it}),\varepsilon _{it}\sim N(0,1),\alpha _{i}\sim N(0,1),\Pr (Y_{i0}=1)=.5.
\end{equation*}
We consider different DGPs indexed by $\beta ^{\ast }\in \lbrack -2,2]$ and
compute the width of the bounds for $T\in \{2,4,8,16,32,64\}$. The width is
asymmetric with respect to $\beta ^{\ast }=0$ because $\Pr
(X_{i}=(1,...,1)^{\prime })$ grows with $\beta ^{\ast }$, whereas $\Pr
(X_{i}=(0,...,0)^{\prime })$ does not depend on $\beta ^{\ast }$. The width
growing with $\beta ^{\ast }$ may therefore be explained by having fewer switches of $
Y_{it}$ between one and zero when $\beta ^{\ast }$ is larger. It is
presumably the changes that help identify the ATE. We find that the bounds
can be substantially wide for high values of $\beta ^{\ast }$ even for large
$T$, consistent with the width of the nonparametric bounds shrinking only at
rate $1/T,$ as shown in the next Section. Semiparametric bounds for this
model that impose the constancy of $\beta ^{\ast }$ across individuals, will
shrink much faster at $T$ grows, as shown in Section 7.

\section{The Impact of $T$}

Increasing $T$ improves identification, shrinking the estimated and
population-identified sets for the objects of interest. The rate at which
the identified set shrinks quantifies this improvement. Here we give rates
for the ASF and, for brevity, leave the quantile results to the
Supplementary Material.

The width of the population bounds for the ASF is $(B_{u}-B_{\ell })\mathcal{
\bar{P}}(x)$ where
\begin{equation*}
\mathcal{\bar{P}}(x)=\Pr (X_{i1}\neq x,...,X_{iT}\neq x).
\end{equation*}
Thus, the rate at which the identified set shrinks, that we will refer to as
the identification rate, is the same as the rate at which $\mathcal{\bar{P}}
(x)$ shrinks. Factors that determine this rate can be seen when $X_{it}$ is
i.i.d. conditional on $\alpha _{i}$. In that case
\begin{equation*}
\mathcal{\bar{P}}(x)=E[\Pr (X_{it}\neq x|\alpha _{i})^{T}].
\end{equation*}
The rate at which $\mathcal{\bar{P}}(x)$ goes to zero will be determined by how much
probability mass of $\Pr (X_{it}\neq x|\alpha _{i})$ is close to one. If $
\Pr (X_{it}\neq x|\alpha _{i})=1$ with positive probability then $\mathcal{
\bar{P}}(x)$ does not go to zero. This corresponds to nonidentification of
the ASF, where $x$ does not occur for some individuals as indexed by $\alpha
_{i}$ (see Theorem A11 of the Supplementary Material). On the other hand, if $
\ \Pr (X_{it}\neq x|\alpha _{i})$ is bounded away from one then the
identified set will shrink exponentially quickly, since $\Pr (X_{it}\neq
x|\alpha _{i})^{T}\leq (1-\varepsilon )^{T}$ for some $\varepsilon >0$. In
between the nonidentified and exponential rate cases there are a range of
rates depending on how much of the distribution of $\Pr (X_{it}\neq x|\alpha
_{i})$ is close to $1$. The following result shows the range of rates.

\bigskip

\textsc{Theorem 4:}\textit{\ Suppose that Assumptions 1, 3, 5, and 6 are
satisfied and }$(X_{i1},X_{i2},...)$\textit{\ is stationary and Markov of
order }$J$\textit{\ conditional on }$\alpha _{i}$\textit{. If for some }$
\varepsilon >0$\textit{,\ }$\Pr (X_{it}=x|X_{i,t-1},...,X_{i,t-J},\alpha
_{i})\geq \varepsilon $\textit{\ a.s. then }$\mu _{u}(x)-\mu _{\ell }(x)\leq
(B_{u}-B_{\ell })(1-\varepsilon )^{T-J}.$\textit{\ If }$X_{it}$\textit{\ is
i.i.d. conditional on }$\alpha _{i},$\textit{\ }$\Pr (X_{it}\neq x|\alpha
_{i})$\textit{\ is continuously distributed with pdf }$f_{P}(p),$\textit{\
and\ }
\begin{equation}
f_{P}(p)\leq Cp^{\gamma -1}(1-p)^{v-1},\gamma >0,v>0,\mathit{\ }
\end{equation}
\textit{then }$\mu _{u}(x)-\mu _{\ell }(x)=O(T^{-v}).$

\bigskip

The upper bound on the rate at which the pdf $f_{P}(p)$ of $\Pr (X_{it}\neq
x|\alpha _{i})$ grows or converges to zero as $p\longrightarrow 1$ provides
an upper bound on the rate at which the identified set shrinks. For example,
if $v=1$ so that $f_{P}(p)$ is bounded as $p\longrightarrow 1,$ then the
identified set shrinks at rate $1/T.$ All of the rates implied by this
result are slower than the exponential rate, reflecting how having $\Pr
(X_{it}\neq x|\alpha _{i})$ close to $1$ affects the rate. Also, $\gamma $
has no effect on the convergence rate because that rate is determined by
closeness of $\Pr (X_{it}\neq x|\alpha _{i})$ to $1$, and not to $0$.

The dynamic, binary-choice model is an example where more explicit
conditions can be given. Suppose $Y_{it}=1(\alpha _{i1}+(\alpha _{i2}-\alpha
_{i1})Y_{i,t-1}\geq \varepsilon _{it})$ and $\varepsilon _{it}$ is i.i.d.
and independent of $\alpha _{i}=(\alpha _{i1}, \alpha _{i2})$ with CDF $H(\varepsilon )$. Here $\Pr
(Y_{it}=1|Y_{i,t-1}=0,\alpha _{i})=H(\alpha _{i1})$ and $\Pr
(Y_{it}=1|Y_{i,t-1}=1,\alpha _{i})=H(\alpha _{i2}).$ Unbounded $\alpha _{i}$
and bounded $\varepsilon _{it}$ will correspond to the unidentified case.
Bounded $\alpha _{i}$ and unbounded $\varepsilon _{it}$ lead to an
exponential convergence rate. The following result covers the in-between
case. Let $f_{\varepsilon }(\varepsilon ),$ $f_{\alpha _{1}}(\alpha )$, and $
f_{\alpha _{2}}(\alpha )$ denote the pdfs of $\varepsilon _{it},$ $\alpha
_{i1},$ and $\alpha _{i2}$ respectively, all are assumed to be continuously
distributed.

\bigskip

\textsc{Theorem 5:}\textit{\ If }$Y_{it}=1(\alpha _{i1}+(\alpha _{i2}-\alpha
_{i1})Y_{i,t-1}\geq \varepsilon _{it}),$ \textit{where }$\varepsilon
_{it},(t=1,...,T)$ \textit{is i.i.d. and independent of }$(\alpha
_{i1},\alpha _{i2})$\textit{\ and there is }$v,C>0$ \textit{such that for
all }$\varepsilon $
\begin{equation}
\max_{j=1,2}f_{\alpha _{j}}(\varepsilon )\leq CH(\varepsilon
)^{v-1}[1-H(\varepsilon )]^{v-1}f_{\varepsilon }(\varepsilon ),
\label{density bound}
\end{equation}
\textit{then }$\Delta _{u}-\Delta _{\ell }=O(T^{-v}).$

\bigskip

Here we see that the identification rate in the nonparametric dynamic model
is related to the tail thickness of the distribution of $\alpha _{i1}$ and $
\alpha _{i2}$ relative to the distribution of $\varepsilon _{it}$. The
thinner the tail of $f_{\varepsilon }(\varepsilon )$ relative to the tails
of $f_{\alpha _{1}}(\alpha _{1})$ and $f_{\alpha _{2}}(\alpha _{2})$ the
smaller $v$ will need to be to satisfy the inequality in Theorem 5 and the
slower the identification rate will be. In this way the identification rate
is slower the less strong the signal provided by $\varepsilon _{it}$
relative to the individual effects. Here there is no $\gamma $ present
because both left and right tails matter, in order to bound the rate for the
ATE, and not just for the ASF at a particular $x$.

For a specific example consider $\alpha _{i1}$ and $\alpha _{i2}$ as $
N(0,\sigma _{\alpha }^{2})$ and $\varepsilon _{it}$ as $N(0,\sigma
_{\varepsilon }^{2})$ where $\sigma _{\varepsilon }^{2}\leq \sigma _{\alpha
}^{2}.$ Then for constants $C_{1},$ $C_{2},$ and $v=\sigma _{\varepsilon
}^{2}/\sigma _{\alpha }^{2}$ we have $f_{\alpha _{j}}(\varepsilon
)=C_{1}[f_{\varepsilon }(\varepsilon )]^{v}.$ Also, as is well known for the
Gaussian distribution, $f_{\varepsilon }(\varepsilon )\geq
C_{2}F_{\varepsilon }(\varepsilon )[1-F_{\varepsilon }(\varepsilon )],$
where $F_{\varepsilon }(\varepsilon )$ denotes the CDF of $\varepsilon $. It
follows by $v\leq 1$ that
\begin{equation*}
f_{\alpha _{j}}(\varepsilon )=C_{1}[f_{\varepsilon }(\varepsilon
)]^{v-1}f_{\varepsilon }(\varepsilon )\leq C_{1}C_{2}^{v-1}F_{\varepsilon
}(\varepsilon )^{v-1}[1-F_{\varepsilon }(\varepsilon )]^{v-1}f_{\varepsilon
}(\varepsilon ).
\end{equation*}
Thus equation (\ref{density bound}) is satisfied with $v=\sigma
_{\varepsilon }^{2}/\sigma _{\alpha }^{2}$ so that
\begin{equation*}
\Delta _{u}-\Delta _{\ell }=O(T^{-\sigma _{\varepsilon }^{2}/\sigma _{\alpha
}^{2}}).
\end{equation*}
Hence the width of the bounds shrinks at a rate no larger than $T^{-1}$ and
the rate is slower the smaller $\sigma _{\varepsilon }^{2}/\sigma _{\alpha
}^{2}$ is. It can also be shown that convergence is faster than $T^{-1}$
when $\sigma _{\varepsilon }^{2}>\sigma _{\alpha }^{2}$ and increases with $
\sigma _{\varepsilon }^{2}/\sigma _{\alpha }^{2}$. Thus we see that the
stronger the signal provided by $\varepsilon $ relative to that provided by $
\alpha ,$ in the sense that the higher $\sigma _{\varepsilon }^{2}$ is
relative to $\sigma _{\alpha }^{2}$, the faster will be the identification
rate.

One can obtain analogous results in a static model. If $X_{it}=1(\alpha
_{i}\geq \eta _{it})$ is a binary regressor where $\eta _{it}$ is i.i.d.
over time then the identification rate will be $T^{-v}$ when the inequality
in Theorem 5 is satisfied with the pdf $f_{\eta }(\eta )$ of $\eta _{it}$
replacing the pdf $f_{\varepsilon }(\varepsilon ).$ If $\alpha _{i}$ and $
\eta _{it}$ are distributed as $N(0,\sigma _{\alpha }^{2})$ and $\eta _{it}$
as $N(0,\sigma _{\eta }^{2})$ respectively with $\sigma _{\eta }^{2}\leq
\sigma _{\alpha }^{2}$, then the identified set shrinks at rate $T^{-\sigma
_{\eta }^{2}/\sigma _{\alpha }^{2}}$. For brevity we omit the details.

\section{Semiparametric Multinomial Choice Models}

The nonparametric bounds are informative but may be quite wide for small $T$
. They can be tightened by imposing additional structure on the model. One
way to do this is to specify a parametric model for the conditional
distribution of $Y_{i}$ given values for $(X_{i},\alpha _{i}).$ We focus
here on multinomial choice models. In those models $Y_{i}$ is one of a
finite number of outcomes, denoted here by $\{Y^{1},...,Y^{J}\}.$ The
parametric part of the model are the known conditional probabilities $
\mathcal{L}_{j}^{k}(\alpha ,\beta )$ of $Y_{i}=Y^{j}$\textit{\ }given $
\alpha _{i}$ and $X_{i}\in \mathcal{X}^{k},(k=1,...,K),$ where $\beta $ is a
parameter vector with true value $\beta ^{\ast }$, and $\mathcal{X}^{k}$ is
the set of $X_{i}$ values being conditioned on. Formulating the model in
this way allows for $X_{i}$ that are lagged dependent variables. The
nonparametric part of the model will be the unknown CDF's $F_{k}^{\ast
}(\alpha ),(k=1,...,K)$ of $\alpha _{i}$ conditional on $X_{i}$ in each $
\mathcal{X}^{k}.$ The model then satisfies

\bigskip

\textsc{Assumption 8:} $\Pr (Y_{i}=Y^{j}|X_{i}\in \mathcal{X}^{k})=\int
\mathcal{L}_{j}^{k}(\alpha ,\beta ^{\ast })dF_{k}^{\ast }(\alpha
),(j=1,...,J;k=1,...,K)$\textit{.}

\bigskip

Some examples may be helpful. An important example is a binary choice model
where $Y_{it}\in \{0,1\}$, $\alpha $ is a scalar location individual effect,
$\Pr (Y_{it}=1|X_{i},\alpha _{i},\beta ^{\ast })=H(X_{it}^{\prime }\beta
^{\ast }+\alpha _{i})$ for a CDF $H(\varepsilon ),$ and $Y_{i1},...,Y_{iT}$
are mutually independent conditional on $X_{i}$ and $\alpha _{i}$. In this
case we would let $\mathcal{X}^{k}$ be a singleton given by the $k^{th}$
value $X^{k}$ in the finite support of $X_{i}$ and
\begin{equation}
\mathcal{L}_{j}^{k}(\alpha ,\beta )=\prod_{t=1}^{T}H(X_{t}^{k\prime }\beta
+\alpha )^{Y_{t}^{j}}[1-H(X_{t}^{k\prime }\beta +\alpha )]^{1-Y_{t}^{j}}.
\label{semistat}
\end{equation}
Time effects can be included in this model by specifying that some
components of $X_{t}^{k}$ only depend on $t.$ This model can also be
generalized to allow for some slopes to vary across individuals by
specifying that
\begin{equation}
\mathcal{L}_{j}^{k}\left( \alpha ,\beta \right)
=\prod_{t=1}^{T}H(z_{t}^{\prime }\beta _{1}+X_{t1}^{k\prime }\beta
_{2}+X_{t2}^{k\prime }\alpha )^{Y_{t}^{j}}[1-H(z_{t}^{\prime }\beta
_{1}+X_{t1}^{k\prime }\beta _{2}+X_{t2}^{k\prime }\alpha )]^{1-Y_{t}^{j}}.
\label{semi}
\end{equation}
This model allows the coefficients of $X_{t2}^{k}$ to vary with individuals,
which will include a location effect when some element of $X_{t2}^{k}$ does
not vary with $t$ or $k.$

This set up also allows for dynamic models. For example, consider a binary
choice model with a lagged dependent variable where $\Pr
(Y_{it}=1|Y_{i,t-1},...,Y_{i0},\alpha _{i},\beta ^{\ast })=H(Y_{i,t-1}\beta
^{\ast }+\alpha _{i}).$ Here $X_{i}=(Y_{i,T-1},...,Y_{i0})$ and we take $
K=2, $ with $\mathcal{X}^{k}=\{X_{i}:X_{i1}=Y_{i0}=k-1\}.$ The parametric
part of the model is
\begin{eqnarray}
\mathcal{L}_{j}^{k}\left( \alpha ,\beta \right)
&=&\prod_{t=2}^{T}H(Y_{t-1}^{j}\beta +\alpha
)^{Y_{t}^{j}}[1-H(Y_{t-1}^{j}\beta +\alpha )]^{1-Y_{t}^{j}}  \label{semidyn}
\\
&&\times H((k-1)\beta +\alpha )^{Y_{1}^{j}}[1-H((k-1)\beta +\alpha
)]^{1-Y_{1}^{j}}.  \notag
\end{eqnarray}
This model could be generalized to allow individual specific coefficients
for the dynamic effect, time effects, and other covariates, including the
model of Browning and Carro (2009). For brevity we omit this generalization.

The ATE and its bounds can be decomposed into a weighted average of
conditional ATE and corresponding bounds, weighted by the identified $\Pr
(X_{i}\in \mathcal{X}^{k})$. The semiparametric model may restrict the
conditional bounds so we focus first on them. We will assume that a
conditional ATE takes the form
\begin{equation*}
\Delta ^{k}=\int \Delta (\alpha ,\beta ^{\ast })dF_{k}^{\ast }(\alpha ),
\end{equation*}
where $\Delta (\alpha ,\beta )$ denotes a treatment effect conditional on $
\alpha $. For example, in the model of equation (\ref{semistat}) we could
take $\Delta (\alpha ,\beta )=H(x^{a\prime }\beta +\alpha )-H(x^{b\prime
}\beta +\alpha ),$ in which case
\begin{equation*}
\Delta ^{k}=\int [H(x^{a\prime }\beta ^{\ast }+\alpha )-H(x^{b\prime }\beta
^{\ast }+\alpha )]dF_{k}^{\ast }(\alpha )
\end{equation*}
is the ATE conditional on $X_{i}=X^{k}$. One could also consider the ASF
conditional on $X_{i}=X^{k},$ that would be $\int H(x^{\prime }\beta ^{\ast
}+\alpha )dF_{k}^{\ast }(\alpha )$ in this example.

Neither $\Delta ^{k}$ nor $\beta ^{\ast }$ need be identified. Instead,
there may be sets of $\beta ^{\ast }$ and ATE values that are consistent
with the distribution of the data. To describe the identified sets let $
\mathcal{P}=(\mathcal{P}_{1}^{1},...,\mathcal{P}_{J}^{1},...,\mathcal{P}
_{J}^{K})^{\prime }$ denote the vector of population choice probabilities
with $\mathcal{P}_{j}^{k}=\Pr (Y_{i}=Y^{j}|X_{i}\in \mathcal{X}^{k})$ and
\begin{equation*}
\mathcal{F}_{k}(\beta ,\mathcal{P})=\{F_{k}:\mathcal{P}_{j}^{k}=\int
\mathcal{L}_{j}^{k}\left( \alpha ,\beta \right) dF_{k}(\alpha ),j=1,...,J\},
\end{equation*}
where $\mathcal{F}_{k}(\beta ,\mathcal{P})$ may be empty. The identified set
for $\beta ^{\ast }$ is
\begin{equation*}
B=\{\beta \text{ s.t. }\mathcal{F}_{k}(\beta ,\mathcal{P})\neq \varnothing
,\forall k=1,...,K\}.
\end{equation*}
That is, $B$ is the set where there exist individual effect distributions
such that integrals of model probabilities equal population choice
probabilities. Sharp upper and lower bounds $\Delta _{u}^{k}$ and $\Delta
_{\ell }^{k}$ for $\Delta ^{k}$ are given by
\begin{equation}
\Delta _{u}^{k}=\sup_{\beta \in B,F_{k}\in \mathcal{F}_{k}(\beta ,\mathcal{P}
)}\int \Delta (\alpha ,\beta )dF_{k}\left( \alpha \right) ,\text{ }\Delta
_{\ell }^{k}=\inf_{\beta \in B,F_{k}\in \mathcal{F}_{k}(\beta ,\mathcal{P}
)}\int \Delta (\alpha ,\beta )dF_{k}\left( \alpha \right) .  \label{sate}
\end{equation}
This characterization of bounds for the ATE extends that of Honore and Tamer
(2006) from a finite dimensional $F_{k}$, where $\alpha $ is restricted to a
known fixed grid, to infinite-dimensional $F_{k}$ where any distribution for
$\alpha $ is allowed.

For purposes of comparison with the nonparametric results we consider models
without trends, where the semiparametric models in equations (\ref{semistat}
) and (\ref{semidyn}) are nested in the nonparametric static or dynamic
model. In those models $\Delta ^{k}$ will be identified if it is also
identified in the nonparametric model. In the static case $\Delta ^{k}$ is
nonparametrically identified if $X_{t}^{k}$ takes on the values $x^{b}$ and $
x^{a}$ for some time periods$.$ This follows similarly to the identification
of the conditional effect $\delta $ in Section 3. Therefore, in static
models obtaining a smaller identified set by imposing the restrictions of a
semiparametric model is limited to those $\Delta ^{k}$ where at least one of
$x^{b}$ or $x^{a}$ does not appear in any time period. In what follows we
focus on these $\Delta ^{k}$.

When slopes vary across individuals the semiparametric bounds may be no
tighter than the nonparametric ones. To illustrate consider a binary-choice
model with a single binary regressor $X_{it},$ where $Y_{it}=1((\alpha
_{i2}-\alpha _{i1})X_{it}+\alpha _{i1}>\varepsilon _{it}),$ $\varepsilon
_{it}$ is independent of $(X_{i},\alpha _{i2},\alpha _{i1}),$ and $
\varepsilon _{it}$ has known CDF $H(\varepsilon )$ that is strictly
increasing on the entire real line. The joint distribution of $H(\alpha
_{i1})$ and $H(\alpha _{i2})$ conditional on $X_{i}=X^{k}$ is entirely
unrestricted. Therefore when $X^{k}=(0,...,0)^{\prime }$ the fact that $
E[H(\alpha _{i1})|X_{i}=X^{k}]=E[Y_{it}|X_{i}=X^{k}]$ for every every $t$,
and so is identified gives no information about $E[H(\alpha
_{i2})|X_{i}=X^{k}].$ Thus, $E[H(\alpha _{i2})|X_{i}=X^{k}]$ can be anything
in the unit interval. \ Therefore, the width of the bound for $\Delta
^{k}=E[H(\alpha _{i2})-H(\alpha _{i1})|X_{i}=X^{k}]$ will be equal to the
width in the nonparametric case, $\Delta _{u}^{k}-\Delta _{\ell }^{k}=1$.
More generally, in the panel binary choice model of equation (\ref{semi}),
when there are no time effects, every coefficient of $X_{it}$ varies across
individuals, and $X_{it}$ is fully saturated (e.g. is a complete set of
dummies, one for every possible value of $X_{it}$), the semiparametric
bounds will equal the nonparametric ones.

In the binary-regressor case the width of the overall bound on the ATE is
given by
\begin{equation}
\Delta _{u}-\Delta _{\ell }=\mathcal{\bar{P}}(0)(\Delta _{u}^{1}-\Delta
_{\ell }^{1})+\mathcal{\bar{P}}(1)(\Delta _{u}^{2}-\Delta _{\ell }^{2}).
\label{semi width}
\end{equation}
where we assume $X^{1}=(0,...,0)^{\prime }$ and\textit{\ }$
X^{2}=(1,...,1)^{\prime }$. The semiparametric bounds will be smaller than
the nonparametric bounds if and only if $\Delta _{u}^{1}-\Delta _{\ell }^{1}$
or $\Delta _{u}^{2}-\Delta _{\ell }^{2}$ are smaller than the nonparametric
values of $1.$ This decomposition also shows that the semiparametric
identification rate will be determined by the nonparametric rate, which
governs how fast $\mathcal{\bar{P}}(0)$ and $\mathcal{\bar{P}}(1)$ shrink,
and the rate that the conditional bounds converge. When the slope does not
vary across individuals it turns out that the conditional bounds can
converge very rapidly. The following result shows this in static and
dynamic, binary-choice logit models with binary regressors.

\bigskip

\textsc{Theorem 6: }\textit{Suppose that }$H(v)=e^{v}/(1+e^{v}),$ $\Delta
(\beta ,\alpha )=H(\beta +\alpha )-H(\alpha ),$\textit{\ and either equation
(\ref{semistat}) is satisfied with, }$X_{it}\in \{0,1\}$, \textit{and }$
X^{1}=(0,...,0)^{\prime }$ \textit{and }$X^{2}=(1,...,1)^{\prime },$ \textit{
or equation (\ref{semidyn}) is satisfied with }$k\in \{1,2\}$. \textit{Then
there are }$C>0$ \textit{and }$1>\varepsilon >0$\textit{\ such that}
\begin{equation*}
\Delta _{u}^{k}-\Delta _{\ell }^{k}\leq C(1-\varepsilon )^{T},k=1,2.
\end{equation*}

\bigskip

This fast rate occurs because $T$ conditional moments of a one-to-one
transformation of $\alpha _{i}$ are identified from probabilities of various
$Y$ values, and these moments lead to a fast approximation of the
conditional ATE. For example, $\Pr (Y_{i}=(1,...,1)^{\prime
}|X_{i}=X^{1})=E[H(\alpha _{i})^{T}|X_{i}=X^{1}]$, and other conditional
moments of $H(\alpha _{i})$ can be similarly identified. For the logit $
H(\alpha )$, identification of these moments leads to fast approximation of $
\Delta ^{1}=E[H(\beta ^{\ast }+\alpha _{i})-H(\alpha _{i})|X_{i}=X^{1}]$ and
hence to fast shrinkage of the conditional bound.

From equation (\ref{semi width}) we see that the semiparametric
identification rate in this example will be at least exponential, and may be
even faster, depending on the nonparametric rate. This result illustrates
how imposing a single, additive individual effect can speed up the
identification rate. We expect that this type of improvement will extend
beyond the logit model with binary regressors.

\section{Computation of Semiparametric Bounds}

In this section we discuss computation of population bounds, give examples,
and present theoretical results. A challenge for computation and for
estimation is the dimensionality of the unknown parameters and the nonlinearity
of the probabilities in those parameters. A useful feature of multinomial
panel models is that they are finite dimensional, in spite of the presence
of distributions. The following lemma shows that one only need consider
discrete distributions with $J$ unknown support points in the specification
of the likelihood and the bounds for the ATE. Let \textit{$\Upsilon $}
denote the set of possible values for the individual effect and $\mathbb{B}$
the set of parameters for $\beta .$

\bigskip

\textsc{Lemma 7: }\textit{If Assumptions 5 and 8 are satisfied and }$
\mathcal{L}_{j}^{k}\left( \alpha ,\beta \right) $ \textit{is a measurable
function of }$\alpha $ \textit{for each }$\beta \in \mathbb{B},$ \textit{
then for each }$\beta $\textit{\ and every CDF }$F_{k}$\textit{\ on $
\Upsilon $ there is a discrete distribution }$F_{k}^{J}$\textit{\ with no
more than }$J$\textit{\ support points such that }$\int \mathcal{L}
_{j}^{k}\left( \alpha ,\beta \right) dF_{k}^{J}(\alpha )=\int \mathcal{L}
_{j}^{k}(\alpha ,\beta )dF_{k}(\alpha )$\textit{\ }$(j=1,...,J).$ \textit{
If, in addition, }$\Delta (\alpha ,\beta )$ \textit{is bounded for each }$
\beta $ \textit{then }$\Delta _{u}^{k}$ \textit{and }$\Delta _{\ell }^{k}$
\textit{\ are not affected by restricting attention to }$F_{k}\in \mathcal{F}
_{k}(\beta )$\textit{\ that are discrete with no more than }$J$ \textit{
support points.}

\bigskip

Thus, no matter what the dimension of $\alpha $ is, the multinomial panel
model is finite dimensional, with the number of parameters given by $\dim
(\beta )+(2J-1)^{K}.$ Another implication of this result is that the
distribution of the individual effect is generally not identified in
multinomial models. For example, if the true distribution $F_{k}^{\ast }$
were continuous then Lemma 7 would imply that there is a discrete
distribution that gives exactly the same likelihood. The proof of this
result is similar to Lindsay's (1983) result that the maximum likelihood
estimator of a mixture model has a finite support. It is interesting that
the model takes a discrete mixture form, although the finite-dimensional
nature of the model is expected because the data have finite support.

Although the individual-effect distribution can be taken to be finite
dimensional, the dimension can be large, and the probabilities depend
nonlinearly on the support points for the individual effect. We overcome
this challenge by using an approximation with a fixed but large number of
support points for the individual effects. This approximation makes
approximate probabilities and the ATE linear in parameters, simplifying
computation. Honore and Tamer (2006) used a similar approach, but assumed
that the true distribution of individual effects had known support points.
We explicitly allow for approximation of unknown support points.

To describe how the approximation can be used to calculate the identified
set, let $M$ denote a number of support points for the individual effect and
\textit{$\Upsilon _{M}\mathcal{=}$}$(\bar{\alpha}_{1M},...,\bar{\alpha}
_{MM})^{\prime }$ be a grid of fixed values for the individual effect. Also
let $\pi =(\pi ^{1\prime },...,\pi ^{K\prime })^{\prime }$ denote a $
MK\times 1$ vector of possible probabilities, with each $\pi ^{k}$ an
element of the $M$ dimensional unit simplex $\mathcal{S}_{M}$. Approximate
model probabilities are
\begin{equation*}
P_{j}^{k}(\beta ,\pi ,M)=\sum_{m=1}^{M}\pi _{m}^{k}\mathcal{L}_{j}^{k}\left(
\bar{\alpha}_{mM},\beta \right) \text{.}
\end{equation*}
Consider the function
\begin{equation*}
T_{\lambda }(\beta ,\pi ,M)=\sum_{j,k}w_{j}^{k}\left[ \mathcal{P}
_{j}^{k}-P_{j}^{k}(\beta ,\pi ,M)\right] ^{2}+\lambda _{M}\pi ^{\prime }\pi ,
\end{equation*}
where $w_{j}^{k}$ are positive weights, such as the chi-square ones $
\mathcal{P}^{k}/\mathcal{P}_{j}^{k},$ for $\mathcal{P}^{k}=\Pr (X_{i}\in
\mathcal{X}^{k})$, and $\lambda _{M}>0$ is a penalty multiplier that
controls the impact of the penalty term $\lambda _{M}\pi ^{\prime }\pi $.
This term is present to help regularize the objective function and ensures a
nonsingular Hessian matrix. Let $\tilde{T}_{\lambda }(\beta ,M)=\min_{\pi
\in \mathcal{S}_{M}^{K}}T_{\lambda }(\beta ,\pi ,M)$ and let $\epsilon
_{M}>0 $ be a positive scalar. We approximate the identified set for $\beta $
by
\begin{equation*}
B(M)=\{\beta :\tilde{T}_{\lambda }(\beta ,M)\leq \epsilon _{M}\},\epsilon
_{M}>0.
\end{equation*}
The use of $\epsilon _{M}$ here in allowing a range of values of the
objective function is analogous to Manski and Tamer's (2002) estimation
method. A positive $\epsilon _{M}$ ensures that the set sequence $
(B(M))_{M=1}^{\infty }$ is lower hemi-continuous and that $B(M)$ need not be
smaller than the identified set, even though the individual effect
distributions are restricted by fixing their support points for each $M.$

We calculate the identified set by letting $M$ grow and $\lambda _{M}$ and $
\epsilon _{M}$ shrink until there is little change in $B(M)$. Calculation of
$\tilde{T}_{\lambda }(\beta ,M)$ is straightforward because it is the
minimum of a quadratic function. In practice we have found that $B(M)$
changes little as $M$ increases even when $M$ is quite small. As $M$ grows
and $\epsilon _{M}$ shrinks the set $B(M)$ will converge to the identified
set under conditions given below.

For the ATE bounds, note
\begin{equation*}
D^{k}(M)=\{\sum_{m=1}^{M}\pi _{m}^{k}\Delta (\bar{\alpha}_{mM},\beta
):T_{\lambda }(\beta ,\pi ,M)\leq \epsilon _{M}\}
\end{equation*}
is the set of possible conditional ATE (given $X \in \mathcal{X}^{k})$ that are consistent
with $\tilde{T}_{\lambda }(\beta ,M)\leq \epsilon _{M}$. Approximate lower
and upper bounds are
\begin{equation*}
\Delta _{\ell }^{k}(M)=\min D^{k}(M),\Delta _{u}^{k}(M)=\max D^{k}(M).
\end{equation*}
As $M$ grows and $\epsilon _{M}$ shrinks these bounds will converge to $
\Delta _{\ell }^{k}$ and $\Delta _{u}^{k}$ respectively, under conditions
given below.\

Computation of these ATE bounds is challenging because it requires searching
over a large dimensional set of possible $\pi $. In practice we start with a
smaller set of probabilities and then try others. Specifically, let $\tilde{
\pi}(\beta )\in \arg \min_{\pi \in \mathcal{S}_{M}^{K}}T_{\lambda }(\beta
,\pi ,M),$ $\tilde{S}^{k}(\beta )=\{\pi ^{k}:P_{j}^{k}(\beta ,\pi ,M)=$ $
P_{j}^{k}(\beta ,\tilde{\pi}(\beta ),M),$ $j=1,...,J\},$ and
\begin{equation*}
\tilde{\Delta}_{\ell }^{k}(M)=\min_{\beta \in B(M),\pi ^{k}\in \tilde{S}
^{k}(\beta )}\sum_{m=1}^{M}\pi _{m}^{k}\Delta (\bar{\alpha}_{mM},\beta ),
\text{ }\tilde{\Delta}_{u}^{k}(M)=\max_{\beta \in B(M),\pi ^{k}\in \tilde{S}
^{k}(\beta )}\sum_{m=1}^{M}\pi _{m}^{k}\Delta (\bar{\alpha}_{mM},\beta ).
\end{equation*}
For each $\beta $ these bounds are easy to calculate by linear programming.
We have done so and then checked to see if other values $\pi $ violate these
bounds. We have not found this to be so for values of $M$ that we use to
compute $\beta $. We conjecture that these bounds also converge to the
population bounds as $M\longrightarrow \infty $ although we have not yet
been able to prove this (because we have not been able to show that the ATE
bounds are continuous in the true probabilities).

We carry out some numerical calculations for the probit model where
\begin{equation*}
Y_{it}=1(\beta ^{\ast }X_{it}+\alpha _{i}\geq \varepsilon _{it}),\varepsilon
_{it}\sim N(0,1),X_{it}=1(\alpha _{i}\geq \eta _{it}),\eta _{it}\sim
N(0,1),\alpha _{i}\sim N(0,1).
\end{equation*}
We consider different DGPs indexed by $\beta ^{\ast }\in \lbrack -2,2]$ and $
T\in \{2,3\}$. Figures 2 and 3 show nonparametric bounds for ATEs and
semiparametric bounds for $\beta ^{\ast }$ and ATEs for $T=2$ and $T=3$,
respectively. The semiparametric bounds are obtained using the computational
algorithm described above with $M=100$ and $\lambda _{M}=1.3\times 10^{-8}$.
The elements of the fixed grid $\Upsilon _{M}$ are located at the
percentiles of the standard normal distribution. We find that $\beta ^{\ast }
$ is not identified for $T=2$, extending the result of Chamberlain (2010) to
this example without time dummy. This result also holds for $T=3$, although
it is difficult to appreciate in the figure because the identified set $B$
is very small. The nonparametric bounds for the ATEs (NP-bounds) can be very
wide, even when we impose monotonicity (NPM-bounds) as described in the
Supplementary Material. The semiparametric bounds for the ATEs (SP-bounds)
are tighter than the nonparametric bounds and shrink very fast with $T$. In
the Supplementary Material we report similar results for the logit,
including nonidentification of the ATEs, except that $\beta ^{\ast }$ is
identified, as is well known. Honore and Tamer (2006) also found tight
bounds for the coefficient of a dynamic model.

To show that the approximate sets converge to the identified set as $M$
grows we impose some conditions. Let $d(\alpha ,\tilde{\alpha})$ denote a
metric on the set \textit{$\Upsilon $ }of possible values for $\alpha $.

\bigskip

\textsc{Assumption 9:} \textit{(i) $\Upsilon $ is a compact metric space
with metric }$d(\alpha ,\tilde{\alpha})$\textit{; ii) }$\eta
(M)=\sup_{\alpha \in \Upsilon }\min_{\tilde{\alpha}\in \Upsilon
_{M}}d(\alpha ,\tilde{\alpha})$\textit{\ }$\longrightarrow 0$ as $
M\longrightarrow \infty ;$ \textit{(iii) }$\mathbb{B}$ \textit{is a compact
subset of }$\Re ^{b}$\textit{; (iv) there is }$C$ \textit{such that for all }
$(\alpha ,\beta ),(\tilde{\alpha},\tilde{\beta})\in $\textit{$\Upsilon $}$
\times \mathbb{B}$, $\left\vert \mathcal{L}_{j}^{k}\left( \tilde{\alpha},
\tilde{\beta}\right) -\mathcal{L}_{j}^{k}\left( \alpha ,\beta \right)
\right\vert \leq C[d(\tilde{\alpha},\alpha )+\left\Vert \tilde{\beta}-\beta
\right\Vert ];$ \textit{and v) }$\Delta (\alpha ,\beta )$ \textit{is
continuous on $\Upsilon $}$\times \mathbb{B}$\textit{.}

\bigskip

Although condition (i) seems restrictive, unbounded individual effects may
be allowed if $\Upsilon $\textit{\ }is chosen appropriately. For example, in
the binary-choice model of equation (\ref{semistat}) this condition will be
satisfied if \textit{$\Upsilon $} is taken to be a two-point
compactification of the real line and $d(\alpha ,\tilde{\alpha})$ is
specified appropriately, as shown in the following result.

\bigskip

\textsc{Lemma 8: }\textit{If Assumptions 5 and 8 and equation (\ref{semistat}
) are satisfied, where }$H(v)$\textit{\ is strictly monotonic on }$\Re $
\textit{with bounded continuous derivative, and }$\mathbb{B}$\textit{\ is a
compact subset of }$\Re ^{b},$\textit{\ then there is a metric }$d(\alpha ,
\tilde{\alpha})$\textit{\ and for each }$M$\textit{\ there is }$\Upsilon
_{M}=\{\bar{\alpha}_{1M},...,\bar{\alpha}_{MM}\}$\textit{\ such that
Assumption 9 is satisfied with }$\eta (M)=1/(M-1).$\textit{\ }

\bigskip

For the convergence results for the identified set we use the Hausdorff set
metric,
\begin{equation*}
d_{H}(A,B)=\max \{\sup_{a\in A}\inf_{b\in B}d(a,b),\sup_{b\in B}\inf_{a\in
A}d(a,b)\}.
\end{equation*}
\textit{\ }

\bigskip

\textsc{Theorem 9: } \textit{If Assumptions 5, 8, and 9 are satisfied, }$
\epsilon _{M}\longrightarrow 0$, \textit{and} $\left( \eta (M)+\lambda
_{M}\right) /\epsilon _{M}\longrightarrow 0$\textit{\ then as }$
M\longrightarrow \infty ,$
\begin{equation*}
d_{H}(B(M),B)\longrightarrow 0,\Delta _{\ell }^{k}(M)\longrightarrow \Delta
_{\ell }^{k},\Delta _{u}^{k}(M)\longrightarrow \Delta _{u}^{k}.
\end{equation*}

\section{Estimation and Inference}

Under Assumptions 5 and 8 the complete description of the data-generating
process is provided by the parameter vector $(${$P_{X}^{\prime }$}$,${$P$}$
^{\prime })^{\prime },$ where $P_{X}=(P^{k},k=1,...,K)^{\prime }$ and $
P=(P_{j}^{k},j=1,...,J,k=1,...,K)^{\prime }.$ The true value of the
parameter vector is $\Pi =({\mathcal{P}}_{X}^{\prime },\mathcal{P}^{\prime
})^{\prime }$, where $\mathcal{P}_{X}=(\mathcal{P}^{k},k=1,...,K)^{\prime }$
and $\mathcal{P}=(\mathcal{P}_{j}^{k},j=1,...,J,k=1,...,K)^{\prime },$ and
the empirical estimate is $\hat{\Pi}=({\hat{P}}_{X}^{\prime },\hat{P}
^{\prime })^{\prime }$, where $\hat{P}_{X}=(\hat{P}^{k},k=1,...,K)^{\prime }$
and $\hat{P}=(\hat{P}_{j}^{k},j=1,...,J,k=1,...,K)^{\prime }.$

The estimation method is like the computational one in using
linear-in-parameters approximations to the probabilities. Here we describe
the estimation method and give a consistency result, and in the
Supplementary Material we provide the implementation details. We follow
 the same steps as the computational one except that we use
estimated weights $\hat{w}_{j}^{k}$ and estimated probabilities $\hat{P}
_{j}^{k}$. Let $\hat{M}$ be a choice of $M$ that may depend on the data and
sample size, and
\begin{equation*}
\hat{T}_{\lambda }(\beta ,\pi )=\sum_{j,k}\hat{w}_{j}^{k}\left[ \hat{P}
_{j}^{k}-P_{j}^{k}(\beta ,\pi ,\hat{M})\right] ^{2}+\lambda _{n}\pi ^{\prime
}\pi .
\end{equation*}
Let $\hat{T}_{\lambda }(\beta )=\min_{\pi \in \mathcal{S}_{M}^{K}}\hat{T}
_{\lambda }(\beta ,\pi )$ and $\epsilon _{n}$ $>0$ be a positive scalar. We
estimate the identified set for $\beta $ by
\begin{equation*}
\hat{B}=\{\beta \in \mathbb{B}:\hat{T}_{\lambda }(\beta )\leq \epsilon
_{n}\},
\end{equation*}
where $\mathbb{B}$ is the parameter space and $\epsilon _{n}$ is a cut-off
parameter that shrinks to zero with the sample size, as in Manski and Tamer
(2002) and Chernozhukov, Hong, and Tamer (2007). The ATE bounds can be
estimated by
\begin{equation*}
\hat{\Delta}_{\ell }^{k}=\min \hat{D}^{k},\hat{\Delta}_{u}^{k}=\max \hat{D}
^{k},\hat{D}^{k}=\{\sum_{m=1}^{M}\pi _{m}^{k}\Delta (\bar{\alpha}_{mM},\beta
):\hat{T}_{\lambda }(\beta ,\pi )\leq \epsilon _{n}\}.
\end{equation*}

This approach to estimation (and computation) can be easily modified to
handle the case where the distribution of the individual effect is
restricted to be the same across some values of $k$. Such a modification
could be implemented by imposing equality of $\pi _{m}^{k}$ across those
values of $k.$ An example would be a model where the distribution of $\alpha
_{i}$ did not depend on some component of $X_{it}.$ That restriction could
be imposed setting $\pi _{m}^{k}$ to be equal across $k$ where the other
components of $X_{it}$ do not vary. Or in a case with a lagged dependent
variable we could restrict the distribution of $\alpha $ to only depend on
the initial condition by imposing equality of $\pi _{m}^{k}$ across all $k$
where $Y_{i0}$ takes on a particular value.

The following is a consistency result.

\bigskip

\textsc{Theorem 10: } \textit{If Assumptions 5, 8, and 9 are satisfied, }$
\hat{w}_{j}^{k}\overset{p}{\longrightarrow }w_{j}^{k}>0,$ $\hat{P}_{j}^{k}
\overset{p}{\longrightarrow }\mathcal{P}_{j}^{k}$, $\epsilon
_{n}\longrightarrow 0$, \textit{and }$\left( n^{-1}+\eta (\hat{M})+\lambda
_{n}\right) /\epsilon _{n}\overset{p}{\longrightarrow }0$,\textit{\ then} $
d_{H}(\hat{B},B)\overset{p}{\longrightarrow }0,\hat{\Delta}_{\ell }^{k}
\overset{p}{\longrightarrow }\Delta _{\ell }^{k},\hat{\Delta}_{u}^{k}\overset
{p}{\longrightarrow }\Delta _{u}^{k}.$

\bigskip

It is interesting to note that no upper limit is placed on $M$ in this
result or in Theorem 9. The reason for this is that the model is finite
dimensional, so there is no need for such a limit. Mathematically, a richer,
fixed grid simply corresponds to a bigger submodel of the finite-dimensional
model.

Turning now to the inference for the semiparametric models, we note that it
is rather challenging. The estimators of parameters and ATE are obtained by
nonlinear programming subject to data-dependent constraints that are
modified to respect the constraints of the model. The distributions of these
highly-complex estimators are not tractable, and are also non-regular in the
sense that the limit versions of these distributions do not vary with
perturbations of the DGP in a continuous\ fashion. This implies that the
usual bootstrap is not consistent. To overcome all of these difficulties we
will rely on a variation of the bootstrap, which we call the perturbed
bootstrap. We also give an alternative inference method based on a modified
projection in the Supplementary Material.

The usual bootstrap computes the critical value -- the $\alpha $-quantile of
the distribution of a test statistic -- given a consistently-estimated
data-generating process (DGP). If this critical value is not a continuous
function of the DGP, the usual bootstrap fails to consistently estimate the
critical value. We instead consider the perturbed bootstrap, where we
compute a set of critical values generated by suitable perturbations of the
estimated DGP and then take the most conservative critical value in the set.
If the perturbations cover at least one DGP that gives a more conservative
critical value than the true DGP does, then this approach yields a valid
inference procedure.

The approach outlined above is most closely related to the Monte-Carlo
inference approach of Dufour (2006); see also Romano and Wolf (2000) for a
finite-sample inference procedure for the mean that has a similar spirit. In
the set-identified context, this approach was first applied in the MIT
thesis work of Rytchkov (2007); see also Chernozhukov (2007).

We consider the problem of performing inference on a real parameter $\theta
^{\ast }$. For example, $\theta ^{\ast }$ can be an upper (or lower) bound
on the conditional ATE $\Delta ^{k}$ such as
\begin{equation*}
\theta ^{\ast }(P)=\max_{\beta \in B^{\ast }(P),F_{k}\in \mathcal{F}
_{k}(\beta ,P^{\ast }(P))}\int \Delta (\alpha ,\beta )dF_{k}\left( \alpha
\right) ,\
\end{equation*}
where $P^{\ast }$ denotes the projection of $P$ onto the model space $\Xi
=\{P:\exists \beta \in \mathbb{B}$ with $\mathcal{F}_{k}(\beta ,P)\neq
\varnothing ,\forall k=1,...,K\}$, i.e.
\begin{equation*}
P^{\ast }(P)=\arg \min_{\tilde{P}\in \Xi }W({\tilde{P}},P),\ \ W({\tilde{P}}
,P)=n\sum_{j,k}\hat{P}^{k}\frac{(P_{j}^{k}-{\tilde{P}}_{j}^{k})^{2}}{\tilde{P
}_{j}^{k}},
\end{equation*}
and $B^{\ast }(P)$ is the corresponding projection for the identified set of
the parameter, i.e.
\begin{equation*}
B^{\ast }(P)=\left\{ \beta \in \mathbb{B}:\exists \tilde{P}\in P^{\ast }(P)
\text{ with }\mathcal{F}_{k}(\beta ,\tilde{P})\neq \varnothing
,k=1,...,K\right\} .
\end{equation*}
Alternatively, $\theta ^{\ast }$ can be an upper (or lower) bound on a
scalar functional $c^{\prime }\beta ^{\ast }$ of the parameter $\beta ^{\ast
}$. Then we define
\begin{equation*}
\theta ^{\ast }(P)=\max_{\beta \in B^{\ast }(P)}c^{\prime }\beta .
\end{equation*}
In both cases we project $P$ onto the model space in order to address the
problem of infeasibility of constraints defining the parameters of interest
under misspecification or sampling error. Under misspecification, we
interpret our inference as targeting the parameters of interest in a best
approximating model; see the Supplementary Material on the modified
projection method for further details. Under correct specification, our inference targets
the parameters of interest in the true model.

In order to perform inference on the true value $\theta ^{\ast }=\theta
^{\ast }(\mathcal{P})$ of the parameter, we use the statistic
\begin{equation*}
S_{n}=\hat{\theta}-\theta ^{\ast },
\end{equation*}
where $\hat{\theta}=\theta ^{\ast }(\hat{P})$. Let $G_{n}(s,P)$ denote the
distribution function of $S_{n}(P)=\hat{\theta}-\theta ^{\ast }(P)$, when
the data follow the DGP $P$. The goal is to estimate the distribution of the
statistic $S_{n}$ under the true DGP $P=\mathcal{P}$, that is, to estimate $
G_{n}(s,\mathcal{P})$.

The method proceeds by constructing a confidence region $CR_{1-\gamma }(
\mathcal{P})$ that contains the true DGP $\mathcal{P}$ with probability $
1-\gamma $, close to one. For efficiency purposes, we also want the
confidence region to be an efficient estimator of $\mathcal{P}$, in the
sense that as $n\rightarrow \infty $, $d_{H}(CR_{1-\gamma }(\mathcal{P}),
\mathcal{P})=O_{p}(n^{-1/2}),$ where $d_{H}$ is the Hausdorff distance
between sets. Specifically, in our case we use
\begin{equation}
CR_{1-\gamma }(\mathcal{P})=\{P\in S_{J}^{K}:W(P,\hat{P})\leq c_{1-\gamma
}(\chi _{K(J-1)}^{2})\},  \label{coverage}
\end{equation}
where $c_{1-\gamma }(\chi _{K(J-1)}^{2})$ is the $(1-\gamma )$-quantile of
the $\chi _{K(J-1)}^{2}$ distribution and $W$ is the goodness-of-fit
statistic:
\begin{equation*}
W(P,\hat{P})=n\sum_{j,k}\hat{P}^{k}\frac{\left( \hat{P}_{j}^{k}-P_{j}^{k}
\right) ^{2}}{P_{j}^{k}}.
\end{equation*}
Then we define the estimates of the lower and upper bounds on the quantiles of $
G_{n}(s,\mathcal{P})$ as
\begin{equation}
\underline{G}_{n}^{-1}(\alpha ,\mathcal{P})/\overline{G}_{n}^{-1}(\alpha ,
\mathcal{P})=\inf /\sup_{P\in CR_{1-\gamma }(\mathcal{P})}G_{n}^{-1}(\alpha
,P),
\end{equation}
where $G_{n}^{-1}(\alpha ,P)=\inf \{s:G_{n}(s,P)\geq \alpha \}$ is the $
\alpha $-quantile of the distribution function $G_{n}(s,P)$. Then we
construct a $(1-\alpha -\gamma )\cdot 100\%$ confidence region for the
parameter of interest as
\begin{equation*}
CR_{1-\alpha -\gamma }(\theta ^{\ast })=\left[ \underline{\theta },\overline{
\theta }\right]
\end{equation*}
where, for $\alpha =\alpha _{1}+\alpha _{2}$,
\begin{equation*}
\underline{\theta }=\hat{\theta}-\overline{G}_{n}^{-1}(1-\alpha _{1},
\mathcal{P}),\ \overline{\theta }=\hat{\theta}-\underline{G}_{n}^{-1}(\alpha
_{2},\mathcal{P}).
\end{equation*}
This formulation allows for both one-sided intervals (either $\alpha _{1}=0$
or $\alpha _{2}=0$) or two-sided intervals ($\alpha _{1}=\alpha _{2}=\alpha
/2)$.

For the inference results we condition on the observed distribution of $X$
and thus set $P_{X}=\mathcal{P}_{X}=\hat{P}_{X}.$ We make the following
assumption about the data-generating process.

\bigskip

\textsc{Assumption 10: }\ $\Pi \in $\thinspace $\mathbb{P=}
\{(P_{X},P):P^{k}>\varepsilon ,P_{j}^{k}>\varepsilon ;j=1,...,J,k=1,...,K\}$
for some $\varepsilon >0$.

\bigskip

The following theorem shows that this method delivers (uniformly) valid
inference on the parameter of interest.

\bigskip

\textsc{Theorem 11:} \textit{If Assumptions 5, 8, and 9 are satisfied then
for any sequence of data-generating process }$\Pi = \Pi_n$ \textit{satisfying Assumption 10},
\begin{equation*}
\lim_{n\rightarrow \infty }\text{Pr}_{\Pi }(\theta ^{\ast }\in \left[
\underline{\theta },\overline{\theta }\right] )\geq 1-\alpha -\gamma .
\end{equation*}

\bigskip

In practice, we use the following simulation approach to compute the
confidence intervals.

\bigskip

\textsc{Algorithm: Perturbed Bootstrap}
\begin{enumerate}
\item \textit{Draw a potential DGP $P_{r}=(P_{r1}^{\prime
},...,P_{rK}^{\prime }),$ where $P_{rk}\sim \mathcal{M}(n\hat{P}^{k},(\hat{P}
_{1}^{k},...,\hat{P}_{J}^{k}))/(n\hat{P}^{k})$ and $\mathcal{M}$ denotes the
multinomial distribution. }

\item \textit{Keep $P_{r}$ if it passes the chi-square goodness-of-fit test
at the $\gamma $ level in equation (\ref{coverage}), using $K(J-1)$ degrees
of freedom, and proceed to the next step. Otherwise reject, and repeat step
1. }

\item \textit{Estimate the distribution $G_{n}(s,P_{r})$ of $S_{n}(P_{r})$
by simulation under the DGP $P_{r}$. }

\item \textit{Repeat steps 1 to 3 for $r=1,...,R$, obtaining $
\{G_{n}(s,P_{r})$, $r=1,...,R\}.$}

\item \textit{Let $\hat{\underline{G}}_{n}^{-1}(\alpha ,\mathcal{P})/\hat{
\overline{G}}_{n}^{-1}(\alpha ,\mathcal{P})=\min /\max \{G_{n}^{-1}(\alpha
,P_{1}),...,G_{n}^{-1}(\alpha ,P_{R})\},$ and construct a $1-\alpha -\gamma $
confidence region for the parameter of interest as $CR_{1-\alpha -\gamma
}(\theta ^{\ast })=\left[ \underline{\theta },\overline{\theta }\right] $,
where $\underline{\theta }=\hat{\theta}-\hat{\overline{G}}_{n}^{-1}(1-\alpha
_{1},\mathcal{P}),$ $\overline{\theta }=\hat{\theta}-\hat{\underline{G}}
_{n}^{-1}(\alpha _{2},\mathcal{P})$, and $\alpha _{1}+\alpha _{2}=\alpha .$ }
\end{enumerate}

\bigskip

\section{Empirical Examples}

We illustrate the estimation and inference results with two empirical
examples. One estimates identified effects and calculates bounds for the
effect of unions on earnings quantiles. The other compares nonparametric and
semiparametric bounds for the effect of fertility on women's labor force
participation.

\subsection{Union Premium}

We revisit the empirical question of how unions impact wage structure using
panel data. Our major contribution here is to estimate the effect without
imposing the assumption that unobserved heterogeneity is some additive term
that can be simply differenced out. In our model unobserved heterogeneity
can have an almost unrestricted impact on the structural/causal response
functions, with the time homogeneity serving as the only restriction.

Our analysis is motivated by previous empirical studies that find
differences in unobservables between union and nonunion workers. For
instance, in an influential study, Chamberlain (1982) finds strong evidence
of heterogeneity bias in the estimation of the union effect by comparing
estimates of cross-sectional models and panel data models with additive
heterogeneity. This finding demonstrates the important need of controlling
for unobserved heterogeneity. Also, Angrist and Newey (1991) reject the
hypothesis that the unobserved heterogeneity acts solely in an additive
fashion, motivating the need to control for more general unobserved
heterogeneity. Card (1996) found differences in the union and selection
effect across skill levels. Here we account fully for differences across
individuals in the union effect while allowing correlation of that effect
with union status, thus accounting for selection. Recently Frandsen (2011)
focused on quantile union effects using a regression-discontinuity design
that estimates union effects for those near a union election discontinuity
rather than for those whose union status changes. We find a flatter quantile
profile than he does, consistent with his theoretical results that suggest a
flatter profile away from the discontinuity.

We use data from the National Longitudinal Survey (Youth Sample). The sample
consists of full-time, young, working males, 20 to 29 years old in 1986,
followed over the period 1986 to 1993. We exclude individuals who failed to
provide sufficient information for each year, were in the active armed
forces or were students any year, or who reported too high (more than \$500
per hour) or too low (less than \$1 per hour) wages. The final sample
includes 2,065 men followed over 8 years. We use the union membership and
the log-hourly wage rate in 1980 dollars as the covariate and the outcome
variables. The union membership variable reflects whether or not the
individual had his wage set by a collective bargaining agreement. Vella and
Verbeek (1998) also used data from the NLSY for different years and found
evidence of important union effect heterogeneity with a random effects model.

We begin by imposing the stationarity condition that income with and without
union membership has the same distribution in each time period but also will
allow for location and scale time effects. It turns out that time effects
are not important in this data. Some covariates are also allowed for since
time-invariant covariates are absorbed in the individual effects.
Insensitivity to time effects also suggests that time-varying covariates may
not be important though a fuller exploration would be useful. For brevity we
focus on the case without covariates.

In our analysis, we focus on estimating the union quantile effect for the
subpopulations of workers that ever became unionized within the sample (47\%
of the sample) or that were unionized in the first year (20\% of the
sample). For these subpopulations, the union effect is not point-identified,
since there are 13\% of the ever-unionized workers that always stayed
unionized between 1986 and 1993, and there are 32\% of the workers unionized
in 1986 that remained unionized until 1993. However, we hope to construct
informative bounds on the union effect. We consider both a static model that
allows for the union membership decisions to be strictly exogenous with
respect to wage-setting decisions, and a dynamic model that allows for the
union-membership decisions to be only predetermined with respect to
wage-setting decisions. We shall also report the estimates of the union
effect for the subpopulation of workers who change their union status at
least once within the sample. For this subpopulation, the effect is
point-identified in the static model, that is, the bounds on the union
effect collapse to a point. We shall not estimate the union effect for the
entire population of workers, since the bounds are completely uninformative
in this case. This happens because more than half of the workers are never
unionized within the sample (see Table 1).

All the results are reported in Table 1 and Figure 4. Table 1 assesses the
plausibility of the time-homogeneity assumption by comparing moments and
quantiles of the cross-sectional distributions of log-wages across years for
workers that do not change union status. Under time homogeneity, these cross
sectional distributions should remain time invariant in the static model. In
the table we observe distributional changes across years, but most of the
variation can be captured by additive location effects for both
always-unionized and never-unionized workers.

Panels A and B of fig. 4 present the estimates of the union effect in the
static model for the subpopulation of workers who change their union status
at least once within the sample. In panel A we compare our panel data
estimates of quantile effects that control for individual heterogeneity with
pooled estimates that do not control for individual heterogeneity. In the
pooled estimates, we see that the quantile effect of union membership is
positive but declines sharply at the upper end of the distribution, which
agrees with previous cross-sectional findings (Chamberlain, 1994). A common
explanation for this phenomenon is that the high-skill workers at the lower
end of the earning distribution tend to join the union, whereas the
high-skill workers at the high end of the earning distribution tend not to
join the union. The estimated quantile effect in the cross-section therefore
captures this selection effect of unobserved skills. In the panel-data
estimates, which control for unobserved skills, we see that the quantile
effects of union membership become very flat across the quantile indices.
Thus, by controlling for individual heterogeneity, we have eliminated the
selection effect. Panel B shows that the results are not sensitive to the
inclusion of location and scale effects.

Panel C presents estimated bounds on the union effect for the subpopulation
of workers that ever became unionized within the sample using the static
model with time effects. The bounds are informative, and show that the
effect is positive for most of the quantile indices. The panel also shows
bounds obtained using the assumption of monotonic and positive union effect
on earnings described in the Supplementary Material. These bounds are also
informative, and in fact are substantially tighter than the bounds obtained
without the monotonicity assumption. Panel D presents similar bounds on the
union effect for the subpopulation of workers unionized in the first period
using the dynamic model. The bounds in this case are not informative, even
after imposing monotonicity.

All the panels include 90\% uniform confidence bands for the quantile union
effects constructed by bootstrap with 200 repetitions. These bands allow us
to make visual simultaneous inference on the entire quantile functions. For
example, we cannot reject that the identified union effect is constant and
positive for all the quantiles. For the ever unionized, the quantile union
effect is positive for a large range of quantiles.

\subsection{Female Labor Force Participation}

For an application of the semiparametric bounds we consider a binary choice
panel model of female labor force participation. We focus on the
relationship between participation and the presence of young children in the
household. Other studies that estimate similar models of participation in
panel data include Heckman and MaCurdy (1980, 1982), Chamberlain (1984),
Hyslop (1999), Chay and Hyslop (2000), Carrasco (2001), Carro (2007), and
Fern\'{a}ndez-Val (2009).

The empirical analysis is based on a sample of married women from the
National Longitudinal Survey of Youth 1979 (NLSY79). The sample consists of
1,587 married women. Only women continuously married, not students or in the
active forces, and with complete information on the relevant variables in
the entire sample period are selected from the survey. Descriptive
statistics for the sample are shown in Table 2. The labor force
participation variable ($LFP$) is an indicator that takes the value one if
the woman's employment status is \textquotedblleft in the labor
force\textquotedblright\ according to the CPS definition, and zero
otherwise. The fertility variable ($kids$) indicates whether the woman has
any children younger than 3 years. We focus on very young, preschool
children as most empirical studies find that their presences have the
strongest impact on the mother's participation decision. $LFP$ is stable
across the years considered, whereas $kids$ is decreasing. The proportion of
women that change fertility status grows steadily with the number of time
periods of the panel, but there are still $49\%$ of the women in the sample
for which the effect of fertility is not identified after 3 periods.

The empirical specification we use is similar to Chamberlain (1984). In
particular, we estimate the following equation
\begin{equation*}
LFP_{it}=\mathbf{1}\left\{ \beta ^{\ast }\cdot kids_{it}+\alpha
_{i} \geq \epsilon _{it} \right\} ,
\end{equation*}
where $\alpha _{i}$ is an individual-specific effect. The parameters of
interest are $\beta ^{\ast }$ and the ATE of fertility on participation. We
compute nonparametric and semiparametric probit and logit bounds for these
parameters. We also obtain linear and nonlinear fixed effects estimates,
together with large-$T$ analytical bias corrected estimates and conditional
fixed effects logit estimates.\footnote{
The analytical corrections use the estimators of the bias based on expected
quantities in Fern\'{a}ndez-Val (2009).} The nonparametric bounds impose
monotonicity on the effects. For the semiparametric bounds, we use the
method described in Section 9 with penalty $\lambda _{n}=1/(n\log n)$ and
iterate the quadratic program 3 times with initial weights $\hat{w}_{j}^{k}=
\hat{P}^{k}$. This iteration makes the estimates insensitive to the penalty
and weighting. We search over discrete distributions with $\hat{M}=23$
support points at $\{-\infty ,-4,-3.6,...,3.6,4,\infty \}$ for the parameter
$\beta ^{\ast }$, and with $\hat{M}=163$ support points at $\{-\infty
,-8,-7.9,...,7.9,8,\infty \}$ for the ATE. The estimates are based on panels
of 2 and 3 time periods, both of them starting in 1990.

Table 3 reports estimates and 95\% confidence regions for the parameters of
interest. The confidence regions for the nonparametric bounds are
constructed using the normal approximation $(95\%\ N)$ and nonparametric
bootstrap with 200 repetitions $(95\%\ B)$. The confidence regions for the
semiparametric bounds are obtained using the procedures described in Section
9 and the Supplementary Material. For the perturbed bootstrap method $
(95\%\ PB)$ we use $R=100$, $\gamma =.01$, $\alpha _{1}=\alpha _{2}=.02,$
and 200 simulations from each DGP to approximate the distribution of the
statistic. For the modified projection method $(95\%\ MP)$, the confidence
interval for $\mathcal{P}$ in the first stage is approximated by 5,000 DGPs
drawn from the empirical multinomial distributions that pass the
goodness-of-fit test. Together the modified projection and the perturbed
bootstrap took several days to compute on a personal computer. We also
include confidence intervals obtained by a canonical projection method $
(95\%\ CP)$ less robust to model misspecification than the modified
projection method, that intersects a nonparametric confidence interval for $
\mathcal{P}$ with the space of probabilities compatible with the
semiparametric model $\Xi $:
\begin{equation*}
CR_{1-\alpha }(\mathcal{P})=\left\{ P\in \Xi :W(P,\hat{P})\leq c_{1-\alpha
}(\chi _{K(J-1)}^{2})\right\} .
\end{equation*}
For the fixed-effects estimators, the confidence regions are based on the
asymptotic normal approximation. The semiparametric estimates are shown for $
\epsilon _{n}=0$, i.e., for the solution that gives the minimum value in the
quadratic problem.

Overall, we find that the nonparametric bound estimates and confidence
regions are too wide to provide informative evidence about the relationship
between participation and fertility. The semiparametric bounds offer a good
compromise between producing more informative results without adding too
much structure to the model. Thus, these estimates are always inside the
confidence regions of the nonparametric model and do not suffer important
efficiency losses relative to the fixed-effects estimates.
Another salient feature of the results is that the misspecification problem
of the canonical projection method clearly arises in this application. Thus,
this procedure gives empty confidence regions for the panel with 3 periods.
The perturbed bootstrap and modified projection methods produce similar
(non-empty) confidence regions for the model parameters and ATEs.

The semiparametric intervals for the ATE cover the -9.6\% estimate of
Chamberlain (1984) for the expected effect of having an additional young
child on the participation probability. He obtained this estimate from a
correlated, random-coefficient probit model, a richer specification that
includes education and fertility covariates, and a different sample from the
PSID.

\bigskip

\begin{thebibliography}{99}
\bibitem{} \textsc{Altonji, J., and R. Matzkin} (2005), \textquotedblleft
Cross Section and Panel Data Estimators for Nonseparable Models with
Endogenous Regressors,\textquotedblright\ \textit{Econometrica} 73,
1053-1102.

\bibitem{} \textsc{Alvarez, J., and M. Arellano} (2003), \textquotedblleft
The Time Series and Cross-Section Asymptotics of Dynamic Panel Data
Estimators,\textquotedblright \emph{Econometrica} 71, 1121-1159.

\bibitem{} \textsc{Angrist, J. D.} (1998), \textquotedblleft Estimating the
Labor Market Impact of Voluntary Military Service Using Social Security Data
on Military Applicants,\textquotedblright \emph{Econometrica} 66, 249--288.

\bibitem{} \textsc{Angrist, J. D. and W.K. Newey }(1991), \textquotedblleft
Over-Identification Tests in Earnings Functions with Fixed
Effects,\textquotedblright\ with J.A. Angrist, \textit{Journal of Business
and Economic Statistics} 9, 317-323.

\bibitem{} \textsc{Beresteanu, A., and Molinari, F.} (2008),
\textquotedblleft Asymptotic properties for a class of partially identified
models,\textquotedblright \emph{Econometrica} 76, 763--814.

\bibitem{} \textsc{Bester, A.C., and C. Hansen} (2008), \textquotedblleft
Flexible Correlated Random Effects Estimation in Panel Models with
Unobserved Heterogeneity,\textquotedblright\ working paper, GSB, University
of Chicago.

\bibitem{} \textsc{Bhargava A., and J.D. Sargan }(1983), "Estimating Dynamic
Random Effects Models from Panel Data Covering Short Time Periods," \emph{
Econometrica} 51, 1635---1660.

\bibitem{} \textsc{Blundell, R. and J.L. Powell} (2003), \textquotedblleft
Endogeneity in Nonparametric and Semiparametric Regression
Models,\textquotedblright\ in M. Dewatripont, L. P. Hansen and S. J.
Turnsovsky (eds.) \emph{Advances in Economics and Econometrics}, Cambridge:
Cambridge University Press.

\bibitem{} \textsc{Browning, M. and J. Carro} (2007), \textquotedblleft
Heterogeneity and Microeconometrics Modeling,\textquotedblright\ in
Blundell, R., W.K. Newey, T. Persson (eds.), \emph{Advances in Theory and
Econometrics, Vol. 3}, Cambridge: Cambridge University Press.

\bibitem{} \textsc{Browning, M. and J. Carro }(2009), "Dynamic Binary
Outcome Models with Maximal Heterogeneity," working paper, Oxford.

\bibitem{} \textsc{Card, D. }(1996), The Effect of Unions on the Structure
of Wages: A Longitudinal Analysis," \emph{Econometrica} 64, 957-979.

\bibitem{} \textsc{Carro, J. M.} (2007), \textquotedblleft Estimating
Dynamic Panel Data Discrete Choice Models with Fixed
Effects,\textquotedblright\ \emph{Journal of Econometrics} 140(2), 503-528.

\bibitem{} \textsc{Carrasco, R.} (2001), \textquotedblleft Binary Choice
With Binary Endogenous Regressors in Panel Data: Estimating the Effect of
Fertility on Female Labor Participation,\textquotedblright\ \emph{Journal of
Business and Economic Statistics} 19(4), 385-394.

\bibitem{} \textsc{Chamberlain, G. }(1980), \textquotedblleft Analysis of
Covariance with Qualitative Data,\textquotedblright\ \emph{Review of
Economic Studies}, 47, 225--238.

\bibitem{} \textsc{Chamberlain, G. }(1982), \textquotedblleft Multivariate
Regression Models for Panel Data,\textquotedblright\ \emph{Journal of
Econometrics}, 18, 5--46.

\bibitem{} \textsc{Chamberlain, G.} (1984), \textquotedblleft Panel
Data,\textquotedblright\ in Z. Griliches and M. Intriligator (eds), \emph{
Handbook of Econometrics}. Amsterdam: North-Holland.

\bibitem{} \textsc{Chamberlain, G. }(1987), \textquotedblleft Asymptotic
Efficiency in Estimation with Conditional Moment
Restrictions,\textquotedblright\ \textit{Journal of Econometrics} 34,
305-334.

\bibitem{} \textsc{Chamberlain, G. }(1994), "\textquotedblleft Quantile
Regression, Censoring, and the Structure of Wages,\textquotedblright\ in C.
Sims, ed., \textit{Advances in Econometrics: Sixth World Congress, Volume I}
, Cambridge: Cambridge University Press.

\bibitem{} \textsc{Chamberlain, G.} (2010), \textquotedblleft Binary
Response Models for Panel Data: Identification and
Information,\textquotedblright\ \emph{Econometrica }78, 159-168.

\bibitem{} \textsc{Chay, K. Y., and D. R. Hyslop} (2000), ``Identification
and Estimation of Dynamic Binary Response Panel Data Models: Empirical
Evidence using Alternative Approaches,'' unpublished manuscript, University
of California at Berkeley.

\bibitem{} \textsc{Chernozhukov, V.} (2007), ``Course Materials for 14.385
Nonlinear Econometric Analysis, Fall 2007,'' MIT OpenCourseWare
(http://ocw.mit.edu), MIT.

\bibitem{} \textsc{Chernozhukov, V., J.Hahn, and W.K.Newey }(2004),
\textquotedblleft Bound Analysis in Panel Models with Correlated Random
Effects,\textquotedblright\ \emph{unpublished manuscript},
http://econ-www.mit.edu/files/5239.

\bibitem{} \textsc{Chernozhukov, V., Fernandez-Val, I., Hahn, J., and W.K.Newey }(2007),
\textquotedblleft Identification and estimation of marginal effects in nonlinear panel models,\textquotedblright\ \emph{unpublished manuscript},
MIT.

\bibitem{} \textsc{Chernozhukov, V., Fernandez-Val, I., Hahn, J., and W.K.Newey }(2012),
\textquotedblleft Supplemental Material for Average and Quantile Effects in Nonseparable Panel Models,\textquotedblright\ \emph{unpublished manuscript},
MIT.

\bibitem{} \textsc{Chernozhukov, V., H. Hong, and E. Tamer} (2007),
\textquotedblleft Estimation and Confidence Regions for Parameter Sets in
Econometric Models,\textquotedblright\ \emph{Econometrica} 75(5), 1243--1284.

\bibitem{} \textsc{Dufour, J.-M.} (2006), \textquotedblleft Monte Carlo
Tests with Nuisance Parameters: A General Approach to Finite-Sample
Inference and Nonstandard Asymptotics,\textquotedblright\ \emph{Journal of
Econometrics} 133, 443--477.

\bibitem{} \textsc{Feller, W.} (1943), \textquotedblleft On a General Class
of Contagious Distributions,\textquotedblright\ \emph{Annals of Statistics, }
14, 389-400.

\bibitem{} \textsc{Fernandez-Val, I. }(2009), \textquotedblleft Fixed
Effects Estimation of Structural Parameters and Marginal Effects in Panel
Probit Models,\textquotedblright\ \textit{Journal of Econometrics} 150(1),
71-85.

\bibitem{} \textsc{Fernandez-Val, I. and J. Lee }(2010), "Panel Data Models
with Nonadditive Unobserved Heterogeneity: Estimation and Inference,"
working paper, Boston University.

\bibitem{} \textsc{Frandsen, B. }(2011), "Why Unions Still Matter: The
Effects of Unionization on the Distribution of Employee Earnings," working
paper, MIT.

\bibitem{} \textsc{Graham,B.W. J. Hahn, and J.L. Powell }(2009),
\textquotedblleft A quantile correlated random coefficient panel data
model\textquotedblright\ working paper, Berkeley.


\bibitem{} \textsc{Graham, B.W. and J.L. Powell }(2011), \textquotedblleft
Identification and Estimation of Average Partial Effects in `Irregular'
Correlated Random Coefficient Panel Data Models\textquotedblright , working
paper, Berkeley.

\bibitem{} \textsc{Hahn, J.} (2001), \textquotedblleft Comment: Binary
Regressors in Nonlinear Panel-Data Models with Fixed
Effects,\textquotedblright\ \emph{Journal of Business and Economic Statistics
} 19, 16-17.

\bibitem{} \textsc{Hahn, J., and G. Kuersteiner} (2002), \textquotedblleft
Asymptotically Unbiased Inference for a Dynamic Panel Model with Fixed
Effects when Both n and T Are Large,\textquotedblright\ \emph{Econometrica}
70, 1639-1657.

\bibitem{} \textsc{Hahn, J., and W. Newey} (2004), \textquotedblleft
Jackknife and Analytical Bias Reduction for Nonlinear Panel
Models,\textquotedblright\ \emph{Econometrica} 72, 1295-1319.

\bibitem{} \textsc{Heckman, J.J.} (1981), \textquotedblleft Statistical
Models for Discrete Panel Data,\textquotedblright\ in Manski, C.F. and D.
McFadden (eds.), \emph{Structural Analysis of Discrete Data with Econometric
Applications}, MIT Press, Cambridge, MA.

\bibitem{} \textsc{Heckman, J. J., and T. E. MaCurdy} (1980),
\textquotedblleft A Life Cycle Model of Female Labor
Supply,\textquotedblright\ \emph{Review of Economic Studies} 47, 47-74.

\bibitem{} \textsc{Heckman, J. J., and T. E. MaCurdy} (1982),
\textquotedblleft Corrigendum on: A Life Cycle Model of Female Labor
Supply,\textquotedblright\ \emph{Review of Economic Studies} 49, 659-660.

\bibitem{} \textsc{Hoderlein, S. and H. White }(2011), "Nonparametric
Identification in Nonseparable Panel Data Models with Generalized Fixed
Effects," working paper, Boston College.

\bibitem{} \textsc{Honore, B.E.} (1992), \textquotedblleft Trimmed Lad and
Least Squares Estimation of Truncated and Censored Regression Models with
Fixed Effects,\textquotedblright\ \textit{Econometrica} 60, 533-565.

\bibitem{} \textsc{Honore, B.E. and E. Tamer }(2003), \textquotedblleft
Bounds on Parameters in Dynamic Discrete Choice Models,\textquotedblright\
working paper.

\bibitem{} \textsc{Honore, B.E., and E. Tamer} (2006), \textquotedblleft
Bounds on Parameters in Dynamic Discrete Choice Models,\textquotedblright\
\emph{Econometrica} 74(3), 611-629.

\bibitem{} \textsc{Hyslop, D. R.} (1999), \textquotedblleft State
Dependence, Serial Correlation and Heterogeneity in Intertemporal Labor
Force Participation of Married Women,\textquotedblright\ \emph{Econometrica}
67(6), 1255-1294.

\bibitem{} \textsc{Imbens, G. and W.K. Newey }(2009), \ \textquotedblleft
Identification and Estimation of Triangular Simultaneous Equations Models
Without Additivity,\textquotedblright\ \textit{Econometrica} 77, 1481-1512.

\bibitem{} \textsc{Lehmann, E. L. }(1974), \emph{Nonparametrics: Statistical
Methods Based on Ranks}. San Francisco, CA: Holden-Day.

\bibitem{} \textsc{Lindsay, B.G. }(1983), \textquotedblleft The Geometry of
Mixture Likelihoods: A General Theory,\textquotedblright\ \emph{Annals of
Statistics }11, 86-94.

\bibitem{} \textsc{Manski, C. }(1987),\ \textquotedblleft Semiparametric
Analysis of Random Effects Linear Models From Binary Response
Data,\textquotedblright\ \textit{Econometrica }55, 357-362.

\bibitem{} \textsc{Manski, C.F., and E. Tamer} (2002), \textquotedblleft
Inference on Regressions with Interval Data on a Regressor or
Outcome,\textquotedblright\ \emph{Econometrica} 70, 519 - 546.

\bibitem{} \textsc{Romano, J. P., and M. Wolf} (2000), \textquotedblleft
Finite Sample Nonparametric Inference and Large Sample
Efficiency,\textquotedblright\ \emph{Annals of Statistics}, 28(3), 756--778.

\bibitem{} \textsc{Rytchkov, O. }(2007), \emph{Essays on Predictability of
Stock Returns}\textit{.} Doctoral Dissertation. MIT.

\bibitem{} \textsc{Vella, F. and M. Verbeek }(1998), \textquotedblleft Whose
Wages Do Unions Raise? A Dynamic Model of Unionism and Wage Rate
Determination for Young Men,\textquotedblright\ \textit{Journal of Applied
Econometrics,} 13, 163-183.

\bibitem{} \textsc{Wooldridge, J.M.} (2005), \textquotedblleft Fixed-Effects
and Related Estimators for Correlated Random-Coefficient and
Treatment-Effect Panel Data Models,\textquotedblright\ \emph{Review of
Economics and Statistics} 87, 385--390.

\bibitem{} \textsc{Woutersen, T.} (2002), \textquotedblleft Robustness
Against Incidental Parameters,\textquotedblright\ \emph{unpublished
manuscript}.

\bibitem{} \textsc{Yitzhaki, S. }(1996), \textquotedblleft On Using Linear
Regressions in Welfare Economics,\textquotedblright\ \textit{Journal of
Business \& Economic Statistics} 14, 478-486.

\bibitem{} \textsc{Yu, K. and M.C. Jones }(1998), \textquotedblleft Local
Linear Quantile Regression,\textquotedblright\ \textit{Journal of the
American Statistical Association} 93, 228-237.
\end{thebibliography}




\begin{figure*}

\begin{center}

\centering\epsfig{figure=table1.eps,width=8in,height=10in}


\end{center}

\end{figure*}



\begin{figure*}

\begin{center}

\centering\epsfig{figure=table2.eps,width=8in,height=10in}


\end{center}

\end{figure*}




\begin{figure*}

\begin{center}

\centering\epsfig{figure=table3.eps,width=7.5in,height=9in}


\end{center}

\end{figure*}




\begin{figure}

\begin{center}

\centering\epsfig{figure=width-bounds-dynamic-probit.eps,width=6in,height=8in}

\caption{\label{fig: width_bounds_dynamic_probit} Width of
nonparametric bounds for the ATE in dynamic binary choice probit
models with $ Y_{it} = 1(\beta^* Y_{i,t-1} + \alpha_i \geq
\varepsilon_{it})$, $\varepsilon_{it} \sim  N(0,1)$, $\alpha_i \sim
N(0,1)$, $\Pr(Y_{i0} = 1) = .5$, $\beta^* \in [-2, 2]$, and $T \in
\{2, 4, 8, 16, 32, 64 \}$.}

\end{center}

\end{figure}



\begin{figure}

\begin{center}

\centering\epsfig{figure=probit-QP-T2.eps,width=6.5in,height=6.5in}

\caption{\label{fig: probit-QP-T2} Identified set for parameter and
ATEs in binary choice probit models with $ Y_{it} = 1(\beta^{\ast}
X_{it} + \alpha_i \geq \varepsilon_{it})$, $\varepsilon_{it} \sim
N(0,1)$, $X_{it} = 1(\alpha_i \geq \eta_{it})$, $\eta_{it} \sim
N(0,1)$, $\alpha_i \sim N(0,1)$, $\beta^{\ast} \in [-2, 2]$, and $T
= 2$.}

\end{center}

\end{figure}



\begin{figure}

\begin{center}

\centering\epsfig{figure=probit-QP-T3.eps,width=6.5in,height=6.5in}

\caption{\label{fig: probit-QP-T3} Identified set for parameter and
ATEs in binary choice probit models with $ Y_{it} = 1(\beta^{\ast}
X_{it} + \alpha_i \geq \varepsilon_{it})$, $\varepsilon_{it} \sim
N(0,1)$, $X_{it} = 1(\alpha_i \geq \eta_{it})$, $\eta_{it} \sim
N(0,1)$, $\alpha_i \sim N(0,1)$, $\beta^{\ast} \in [-2, 2]$, and $T
= 3$.}

\end{center}

\end{figure}





\begin{figure}

\begin{center}

\centering\epsfig{figure=identified_effects.eps,width=.48\textwidth,height=.48\textwidth}
\epsfig{figure=identified_effects_te.eps,width=.48\textwidth,height=.48\textwidth}
\epsfig{figure=bounds_static_te.eps,width=.48\textwidth,height=.48\textwidth}
\epsfig{figure=bounds_dynamic.eps,width=.48\textwidth,height=.48\textwidth}


\caption{\label{Fig: union} Quantile union effects for male workers. Panel A displays point and interval estimates of the identified quantile union
effects in the static model with and without accounting for individual heterogeneity. Panel B displays point and interval estimates of the identified quantile union
effects in the static model with location and scale time effects, averaged across time periods with and without accounting for individual heterogeneity. Panel C displays point and interval estimates of the bounds for the quantile effect on the ever unionized in the static model with time effects, with and without imposing monotonicity. Panel D displays point and interval estimates of the bounds for the quantile effect on the unionized in the first period in the dynamic model, with and without imposing monotonicity. Estimates based on NLSY79 for the years 1986--1993. 90\% confidence intervals obtained by bootstrap with 200 repetitions.}

\end{center}

\end{figure}



\clearpage

\begin{center}
\Large{Supplemental Material for  Average and Quantile Effects in
Nonseparable Panel Models}

\bigskip

\normalsize{
Victor Chernozhukov,  Iv\'an Fern\'andez-Val,
Jinyong Hahn, and Whitney Newey}
\end{center}

\bigskip
\bigskip

\setcounter{section}{0}

\section{Introduction}

In this supplemental material we provide omitted discussions, results, and
proofs by Section in the same order they are referred to in the paper. Let
w.p.a.1 denote "with probability approaching one" and $C$ denote a generic
constant that may be different in different uses.