The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
141,374 characters
2530A Correlated Random Coefficient Panel Model with Time-Varying Endogeneity
\title{\fontsize{25}{30}\selectfont \vspace{-50pt} A Correlated Random Coefficient Panel Model \\ with Time-Varying Endogeneity \vspace{5pt}}
\author{\fontsize{15}{19}\selectfont Louise Laage\thanks{[email removed], Department of Economics, Georgetown University.
\newline
I am grateful to Donald W. K. Andrews, Xiaohong Chen and especially Yuichi Kitamura for their guidance and support. I thank Anna Bykhovskaya, Philip A. Haile, Yu Jung Hwang, John Eric Humphries, Ivana Komunjer, Rosa Matzkin, Patrick Moran, Peter C. B. Phillips, Alexandre Poirier, Pedro Sant'Anna, Masayuki Sawada, Francis Vella, Edward Vytlacil as well as participants at the Yale econometrics seminar for helpful conversations and comments on this project. I thank an associate editor and two anonymous referees for suggestions that greatly improved the paper. I thank the Toulouse School of Economics for hosting me during the academic year 2020-2021 and acknowledge financial support from the grant ERC POEMH 337665, from the ANR under grant ANR-17-EURE-0010 (Investissements d’Avenir program), and from the IAAE travel grant. All errors are mine.} \vspace{5pt}}
\date{This version : June 12, 2022}
\maketitle
\begin{abstract}
This paper studies a class of linear panel models with random coefficients. We do not restrict the joint distribution of the time-invariant unobserved heterogeneity and the covariates. We investigate identification of the average partial effect (APE) when fixed-effect techniques cannot be used to control for the correlation between the regressors and the time-varying disturbances. Relying on control variables, we develop a constructive two-step identification argument. The first step identifies nonparametrically the conditional expectation of the disturbances given the regressors and the control variables, and the second step uses ``between-group'' variations, correcting for endogeneity, to identify the APE. We propose a natural semiparametric estimator of the APE, show its $\sqrt{n}$ asymptotic normality and compute its asymptotic variance. The estimator is computationally easy to implement, and Monte Carlo simulations show favorable finite sample properties. Control variables arise in various economic and econometric models, and we propose applications of our argument in several models. As an empirical illustration, we estimate the average elasticity of intertemporal substitution in a labor supply model with random coefficients.
\end{abstract}
\newpage
\section{Introduction}\label{sec:intro}
Correlated Random Coefficient (CRC) models are linear models with random coefficients where the joint distribution of the random coefficients and the regressors is left unspecified. With panel data, panel CRC models relate an outcome variable $y_{it}$ to regressors $x_{it}$ through the following equation\footnote{We purposefully do not allow for the constant term to be included in $x_{it}$. More details are provided in Section \ref{sec:id_model}.}
\begin{align}\label{eq:model}
y_{it} = \, x_{it}' \, \mu_i + \alpha_i + \epsilon_{it}, \qquad \ i \leq n, \ t \leq T,
\end{align}
where $T$ is treated as fixed, $\epsilon_{it}$ is a time-varying residual and the unobserved random coefficients $(\mu_i', \alpha_i)$ can be arbitrarily related to the regressors $x_{it}$. An important and empirically relevant task is to recover properties of the random coefficients and the literature has focused on this question while imposing strict exogeneity of the regressors, that is, $\mathbb{E}(\epsilon_{it}|X_i)= 0$ where $X_i=(x_{i1}',...,x_{iT}')$ \citep{ch92,ab12}. This paper instead examines identification of $\mathbb{E}(\mu_i', \alpha_i)$ in situations where the regressors and the residuals are correlated.
Linear random coefficient models have many empirical applications, for instance to model heterogeneous returns to schooling \citep{c01}, heterogeneous production functions (\cite{mg88}, \cite{gu22}, or see \cite{s11} for an example of heterogeneous technology adoption), heterogeneous taxable income responses to tax changes \citep{kl20}, or in medical studies \citep{lw82}. With panel data, the absence of restrictions on the joint distribution of the random coefficients and the regressors corresponds to the \textit{fixed effect approach}. It depicts situations in which the researcher does not know which factors drive the heterogeneity in the impact of $x_{it}$, or does not observe them. Loosely speaking, the fixed-effect approach together with the strict exogeneity condition imply that the endogeneity of the model can be ``captured by a fixed effect''.
However strict exogeneity fails to hold in many cases, for instance whenever omitted variables correlated with the regressors $x_{it}$ vary over time, e.g., unobserved determinants of productivity in a production function. This leads to contemporaneous endogeneity.
Since strict exogeneity conditions on past and future values of the regressors, it also fails when the regressors are \textit{predetermined}, i.e., correlated with residuals of prior periods, a well-known phenomenon in the panel data literature.
To allow for such prevalent correlations, this paper focuses on situations in which there is \textit{time-varying endogeneity}. A rigorous definition of time-varying endogeneity here is that there is no time-invariant random vector $c_i$ such that the correlation between the regressors $x_{it}$ and the entire vector of unobserved heterogeneity $(\mu_i', \alpha_i,\epsilon_{it})'$ can be written as the correlation between $x_{it}$ and $c_i$.
We adopt a \textit{control function approach} (CFA) and assume that instruments $Z_i = (z_{i1}',...,z_{iT}')$ and potentially time-varying control variables are available such that once control variables are conditioned on, the residual is mean independent of the regressors.
The main contribution of this paper is to develop a two-step method which identifies $\mathbb{E}(\mu_i)$, the average partial effect\footnote{As defined in \cite{w05}, the APE is $\mathbb{E}_q[ \partial \mathbb{E}(y_t|x_t,q_t) / \partial x|_{x_t = \bar{x}}]$ where the outer expectation is over the vector of unobservables $q_t = (\mu, \alpha, \epsilon_t)$.} (APE), in (\ref{eq:model}) under time-varying endogeneity when the data satisfy such a control function approach assumption.
The identification argument proceeds as follows. The CFA implies that the conditional mean of $\epsilon_{it}$ given the collection of regressors and control variables at all periods is in fact a function of the control variables only.
Therefore the first step identifies nonparametrically the first differences of these control functions using the residuals of individual-specific linear regressions of the vector of differenced outcomes on differenced regressors.
An \textit{invertibility} assumption must be imposed to recover the control functions from these residuals. Once the control functions are known, individual-specific linear regressions are run and a cross-section average identifies the APE. Additional steps extend the argument to identify $\mathbb{E}(\alpha_i)$ as well as higher order moments of the random coefficients, using existing results in the literature.
Since this identification argument relies on several key assumptions, we study them in details and connect them to existing conditions in the literature. The most important and least standard ones are the control function approach and the invertibility assumptions. Under the CFA, control variables control for the correlation between the residual at period $t$ and the collection of regressors at all time periods: we
argue that it can be used to capture contemporaneous endogeneity as well as predeterminedness. Additionally, CFA does not restrict the joint distribution of the instruments and the random coefficients but only that of $(X_i,Z_i,(\epsilon_{i \, t})_{t \leq T})$: it remains comparable to exogeneity restrictions in models without random coefficients. The second assumption is a high-level assumption related to both support and rank conditions, as is common in the CFA literature \citep{in09}. However we show with examples that sufficient conditions include cases where the instrument has small support, as in \cite{fhmv08}, and even cases where the instrument is discretely distributed. We provide an example where a regressor following a Markov process is not strictly exogenous but predetermined and show that all assumptions hold.
The proof of identification of $\mathbb{E}(\mu_i)$ is constructive and suggests a natural multi-step estimator, taking sample analogs of each of the identification steps after estimating the control variables. The second contribution of this paper is the analysis of this estimator under a particular specification of the control variables. We show that it is consistent and asymptotically normal. The challenge in deriving the asymptotic properties of the estimator comes from the nonparametric regression estimators constructed using nonparametrically estimated regressors.
The results presented in this paper have various limitations.
First throughout the paper we maintain $T \geq d_x + 2$, where $d_x$ is the dimension of the regressor $x_{it}$ (without the constant term). To identify the APE when the regressor is scalar, the panel must therefore include 3 periods. As explained below, this is common in the analysis of CRC panel models \citep{ab12, ch92}, however an approach developed in \cite{gp12} may be extended so as to obtain identification when $T = d_x +1$. The associated estimation procedure being outside the scope of this paper, we leave it for future work. Furthermore, a crucial implication from a modeling perspective comes from the use of the control function approach. Indeed,
first-stage equations for which we can identify valid control variables typically have scalar unobserved heterogeneity when they model a scalar endogenous regressor. This highlights an underlying imbalance in the degree of heterogeneity between the outcome equation and the implicit selection (first-stage) equation. We further comment on this at the end of Section \ref{sec:CFAcomment}.
\medskip
\textit{\textbf{Related literature.}}
This paper directly contributes to the literature on CRC panel models such as (\ref{eq:model}) placing no restrictions on the distribution of $(X_i,\mu_i', \alpha_i)$. A seminal paper in this literature is \cite{ch92} studying a model similar to (\ref{eq:model}) in a fixed $T$ setting with an additional additively separable parametric term. Under the condition\footnote{The discussion on the number of periods is for (\ref{eq:model}) where we impose the inclusion of a constant regressor.} $ T \geq d_x + 2$, the author derives the semiparametric efficiency bound for $\mathbb{E}(\mu_i', \alpha_i)$ and provides an efficient estimator. Important recent papers studying comparable models include \cite{ab12} and \cite{gp12}. \cite{ab12} focus on identifying the conditional variance and distribution of $(\mu_i', \alpha_i)'$ under various restrictions on the serial dependence of $\epsilon_{it}$. \cite{gp12} relax a regularity assumption in \cite{ch92} which imposed sufficient variations of the regressors over time, and, as mentioned above, allow for $T=d_x+1$. More recent contributions are \cite{v20} and \cite{su21}. We differ from all these papers in that we allow for time-varying endogeneity.
A wide panel data literature studies models with time-invariant unobserved heterogeneity $(\mu_i', \alpha_i)$ and time-varying residuals $\epsilon_{it}$, allowing for time-varying endogeneity and adopting a fixed effect approach.
A well-known example is a linear regression model with an additive fixed effect, in which case the unobserved heterogeneity is scalar. One may use the fixed-effect instrumental variable estimator to consistently estimate the slope, see \cite{w05c}. An example using the control function approach when there is sample selection is \cite{w95}. Unlike these papers, we allow for the slope to be heterogeneous as well.
More recently, results have been developed for identification of nonseparable models in panel data with time-invariant unobserved heterogeneity and time-varying residuals, see, e.g., \cite{ab16}, \cite{abb17}. Using results from \cite{hs08}, these papers recover the distribution of the unobserved heterogeneity. The number of time periods required for identification is typically higher than we need: for instance \cite{ab16} need 5 time periods to identify the distribution of a bivariate fixed effect. Focusing on the particular functional specification of the model (\ref{eq:model}), we need fewer time periods and our identification argument suggests a computationally simpler estimator. See also \cite{fglv22}.
An existing alternative approach for panel models with time-invariant unobserved heterogeneity restricts its joint distribution with the regressors. This corresponds to the \textit{correlated random effect approach} (CRE).
It has been used in analyses of linear random coefficients models, see \cite{w05b}, but also of nonseparable panel data models. For instance \cite{am05} identify the local average response function assuming that the vector of both time-invariant unobserved heterogeneity and residual at time $t$ is independent of the regressor at time $t$ conditional on an instrument.
Some papers have developed a CRE approach in linear random coefficient models with time-varying endogeneity as well. For instance, \cite{mw08} study a linear panel random coefficient model and assume that the random coefficients are conditionally mean independent of the detrended instrument to show that the fixed-effect instrumental variables estimator is consistent to the average partial effect. See also \cite{mw16} and \cite{l21}.
Other recent examples of the CRE approach include \cite{bh09},
\cite{mkv11},
\cite{ghpp18}.
In contrast to these papers, we adopt a fixed-effect approach and do not restrict the joint distribution of the unobserved heterogeneity $(\mu_i', \alpha_i)$, the regressors and the instruments. The multiplicative specification in (\ref{eq:model}) allows us to ``difference out'' $(\mu_i', \alpha_i)$ entirely in one step and recover the term capturing the time-varying endogeneity.
An alternative approach to identification of the APE in (\ref{eq:model}) would be to consider each cross-section separately and to apply existing results on identification of cross-section random coefficient models with endogenous regressors. For instance, \cite{w97}, \cite{w03} and \cite{hv98} identify the average partial effect imposing an exclusion restriction on the random coefficients and a first-stage equation with homogeneous impact of the instruments on the regressors. More recently, \cite{mt16} use a nonseparable first-stage equation similar to \cite{in09}
and retrieve the conditional APE imposing a control function approach on the entire vector of random coefficients: in particular this vector is assumed independent of the instruments. A similar approach can be found in \cite{ns21} who study a set of regressions which encompasses a random coefficient model where the control function approach is applied on the random coefficients.
In these papers, the control function approach restricts the joint distribution of all random coefficients, the regressors and the instruments. In contrast, using panel data and imposing the random coefficients multiplying the regressors to be time-invariant allows us to obtain identification while imposing the control function approach on the residual $\epsilon_{i \, t}$ only.
\medskip
The structure of the paper is as follows. Section \ref{sec:id_model} defines the model, explains briefly the two-step argument proving identification of $\mathbb{E}(\mu_i)$ and introduces our assumptions. Section \ref{sec:assn} discusses at length these assumptions and their empirical content.
Section \ref{sec:identification}
provides the formal identification result and discusses other moments of the random coefficients.
In Section \ref{sec:estim}, we define our proposed estimator and provide its asymptotic properties.
Finally, Section \ref{sec:illus} turns to an empirical illustration of a labor supply model with random coefficients. We give conditions under which our main assumptions hold and with a dataset from \cite{z97} estimate the average elasticity of intertemporal substitution.
The proofs for all sections are in the Appendix, as well as some Monte Carlo simulations showing favorable finite sample properties of the estimator, and additional comments.
\section{Model \& Intuition}\label{sec:id_model}
We first introduce some notations used throughout the paper. Random variables are indexed with $i$ or $it$. We denote with $\mathcal{S}_{A_i}$ the support of a random variable $A_i$, $\mathcal{S}_{A_i|B_i=B}$ the support of $A_i$ given that the random variable $B_i$ is equal to $B$. $I$ is the identity matrix and its dimension is usually clear from the context.
For a sample of units indexed by $i$, for $i = 1,..,n$, the outcome variable $y_{it} \in \mathbb{R}$ in period $t=1,..,T$ is given by
\begin{equation}
y_{it} = \, x_{it}' \, \mu_i + \alpha_i + \epsilon_{it}, \quad \mathbb{E}(\epsilon_{it}) = 0, \tag{\ref{eq:model}}
\end{equation}
where $x_{it} \in \mathbb{R}^{d_x}$ is a vector of observed variables which does not include the constant regressor, $\epsilon_{it}$ is a time-varying disturbance, and $(\alpha_i,\mu_i')' \in \mathbb{R}^{d_x+1}$ is a time-invariant vector capturing individual unobserved heterogeneity. The reason why the constant term is not included in $x_{it}$ is discussed at the beginning of Section \ref{sec:FD}. We focus on short panels, where $T$ is fixed and $n$ large. The constraint $T \geq d_x + 2$ will be maintained throughout the paper. More details on this condition are given in Section \ref{sec:assn_CRC_matrix}. Denoting by $y_i = (y_{i1},..,y_{iT})'$ the vector of outcomes of unit $i$, $X_i = (x_{i1},..,x_{iT})'$ the matrix of regressors, and $\epsilon_i = (\epsilon_{i1},..,\epsilon_{iT})'$ the vector of error terms, we can rewrite (\ref{eq:model}) as
$$y_i = X_i \mu_i + \alpha_i 1_{T} + \epsilon_i,$$
where $1_{T}$ is the vector of size $T$ where each component is equal to $1$. The parameters of interest we focus on are the average effects $\mathbb{E}(\mu_i)$ and $\mathbb{E}(\alpha_i)$.
A standard assumption in the panel CRC literature is strict exogeneity of the regressors, that is, $\mathbb{E}(\epsilon_{i \, t}|X_i, \alpha_i, \mu_i) = 0$.
As pointed out in the introduction, this assumption does not allow for the presence of time-varying omitted variables or other time-varying sources of correlation between the regressors and the residual $\epsilon_{it}$. We seek to relax this condition.
Consider for instance an education production function with school class size as input. The impact of teacher's attention, thus of class size, on students' achievements is likely heterogeneous as it potentially depends on teacher's quality. For data observed at the class level across schools and cohorts, many omitted variables are correlated with class size, e.g.,
through school budget, such as students, parents and community characteristics. These omitted factors vary over cohort and school and thus cannot be captured by an additive fixed effect.
This framework also applies when $t$ captures alternative group structures. To study the impact of mother smoking habit on infant birth weight with a panel of mothers with multiple births (see, e.g., \cite{a06}), data indexed by $it$ describe the $t^{\text{th}}$ child of mother $i$. The impact of a mother's level of smoking is likely heterogeneous across mothers but moreover some omitted variables related to a mother's health behavior may impact the birthweight and vary across children with the smoking level.
In both examples, omitted variables are correlated with the regressor of interest yet can not be captured by additive fixed effects. To relax the strict exogeneity condition and allow for such correlations, we assume instead that there exist control variables $v_{it}\in\mathbb{R}^{d_v}$ which are known or identified functions of regressors and instrumental variables $z_{it} \in \mathbb{R}^{d_z}$, and satisfy a conditional mean independence condition in line with the control function approach. This condition is our first assumption, Assumption \ref{assn:idCFA}, and will be formally stated in Section \ref{sec:CFAcomment} together with discussions of the serial correlations it allows for and of how to construct control variables. This assumption implies that for $V_i = (v_{i1}',.. v_{iT}')'$, for all $0\leq t \leq T$ there exists a random variable $u_{it} $ such that
\begin{equation}\label{eq:cfa_bf_assn}
\epsilon_{it} = f_t(V_i) + u_{it}, \text{ for some function } f_t \text{ where } \mathbb{E}(u_{i \, t}|X_i, V_i) = 0 \text{ and } \mathbb{E}(f_t(V_i)) = 0.
\end{equation}
To identify the average effects $\mathbb{E}(\mu_i)$ and $\mathbb{E}(\alpha_i)$, we suggest a stepwise procedure which will isolate the unobserved heterogeneity from the regressor and the control variables. We describe this procedure below and define $f(V_i)\, = \, (f_1(V_i),.. \, , f_T(V_i))'$ and $Z_i = (z_{i1},...,z_{iT})$.
\subsection{First-Differencing}\label{sec:FD}
In (\ref{eq:model}), we deliberately treat the additive fixed effect $\alpha_i$ separately, that is, the constant regressor is not included in $x_{it}$.
Two alternative approaches could be used instead, or, specifically, the model could have been written in two other ways.
The first is to change (\ref{eq:model}) to write instead $y_{it} = \, \tilde{x}_{it}' \, \tilde{\mu}_i + \epsilon_{it}$ where $\tilde{\mu}_i = (\alpha_i,\mu_i')'$ and $\tilde{x}_{it} = (1,x_{it}')'$. However one of the main assumptions needed for identification, Assumption \ref{assn:id_M_GLn}, does not hold if the regressor is replaced with $\tilde{x}_{it}$. This assumption could be rewritten using $\tilde{x}_{it}$ so as to guarantee identification when it holds, but its formulation and analysis would be more complex. We thus prefer to treat the constant regressor separately from the other regressors.
Note that with our use of the control function approach, Model (\ref{eq:model}) can be written $y_{it} = x_{it}' \, \mu_i + \alpha_i + f_t(V_i) + u_{it}.$ Consequently, a second approach would be to define instead $\tilde{\alpha} = \mathbb{E}(\alpha_i)$, $\tilde{f}_t(V_i)= \mathbb{E}(\alpha_i|V_i) - \mathbb{E}(\alpha_i) + f_t(V_i)$ and $\tilde{u}_{it} = [\alpha_i - \mathbb{E}(\alpha_i|V_i)] + u_{it}$
, so as to obtain $y_{it} = x_{it}' \, \mu_i + \tilde{\alpha} + \tilde{f}_t(V_i) + \tilde{u}_{it}$. However, the control function approach mentioned below imposes the condition $\mathbb{E}(\alpha_{i}|X_i, V_i) = \mathbb{E}(\alpha_i|V_i) $, which is conflicting with the fixed effect approach we adopt. We thus treat $\alpha_i$ separately from the other unobserved heterogeneity terms $f_t(V_i)$ and $u_{it}$, i.e., from $\epsilon_{it}$.
It implies that the functions $f(V_i)$ and $\mathbb{E}(\alpha_i|V_i)$ are not separately identifiable. But the normalization $\mathbb{E}(f(V_i)) \, = \, 0$ can be leveraged to identify $\mathbb{E}(\alpha_i)$ once we identify first differences of $f$, as will be clear in Section \ref{sec:maindid}.
Exploiting the time invariance of $\alpha_i$, we take time differences to eliminate this term and focus on identification of $\mathbb{E}(\mu_i)$.
Thus we now focus on a first-differencing transformation of the model,
\begin{equation}
\dot{y}_{i t} = \, \dot{x}_{it}^{\, \prime} \, \mu_i + \, g_t(V_i) + \, \dot{u}_{it}, \label{eq:modelwTdiff}
\end{equation}
with
$\dot{y}_{i t} = y_{i \, t + 1} - y_{i t}$, $\dot{x}_{it} = x_{i \, t + 1} - x_{i t}$, $g_t(V_i) = f_{t+1}(V_i) - f_{t}(V_i)$, and $\dot{u}_{it} = u_{i \, t+1} - u_{it}$ for $ t \leq T-1$.
In vector form, define the $(T-1) \, \times d_x$ matrix $\dot{X}_i = ( \dot{x}_{i1},.. , \dot{x}_{i \, T-1})'$ and the $(T-1) \times 1$ vectors $\dot{y}_i = (\dot{y}_{i1},.. \, , \dot{y}_{i \, T-1})'$, $g(V_i)\, = \, (g_1(V_i),.. \, , g_{T-1}(V_i))'$ and $\dot{u}_i = (\dot{u}_{i1},.. \, , \dot{u}_{i \, T-1})'$. By assumption, $\mathbb{E}(\dot{u}_i | X_i , V_i) =0$ and
Equation (\ref{eq:modelwTdiff}) can then be rewritten
\begin{equation}\label{eq:modelwTdiffVector}
\dot{y}_i = \,\dot{X}_i \mu_i + \, g(V_i) \, + \, \dot{u}_i.
\end{equation}
\subsection{Two-Step Identification}\label{sec:intuition_id}
We now discuss identification of $\mathbb{E}(\mu_i)$ and introduce additional assumptions. Formal results are provided in Section \ref{sec:identification}. Since the random coefficients $\mu_i$ are heterogeneous, standard estimators such as TSLS are not consistent without further restrictions.
However an intuitive idea is to consider each unit $i$ separately: for each we observe several time periods and can run a unit-specific linear regression of differenced outcomes $\dot{y}_{it} = y_{i \, t + 1} - y_{i t}$ on differenced regressors $\dot x_{it} = x_{i \, t + 1} - x_{i t}$ over observations $t=1, .., T-1$. This produces a unit-specific estimator $\beta_{i}^{OLS}$.
first-differencing eliminates the additive fixed effects $\alpha_i$ and if the regressors are strictly exogenous, the average of $\beta_{i}^{OLS}$ across units is $\mathbb{E}(\mu_i)$. Note that to run such a regression, we need more observations than regressors, i.e., $T-1 \geq d_x$ as otherwise, if the matrix of observations $\dot X_i$ is of full rank, there is not a unique solution to the least squares optimization problem.
If there is time-varying endogeneity and (\ref{eq:cfa_bf_assn}) holds,
the argument described above does not identify $\mathbb{E}(\mu_i)$.
In fact, since the residual $\epsilon_{it}$ is composed of the two terms $u_{it}$ and $f_t(V_i)$, the unit-specific $\beta_{i}^{OLS}$ is the sum of three terms: $\mu_i$, the individual slope, $\beta_{i}^{OLS, \, g}$, the coefficient of the regression of $g_t(V_i)$ on $\dot x_{it}$ and $\beta_{i}^{OLS, \, u}$, the coefficient of the regression of $\dot u_{it}$ on $\dot x_{it}$. The regressors are strictly exogenous for the new residual $u_{it}$, thus $\beta_{i}^{OLS, \, u}$ averages to $0$ across units. However this does not hold for $\beta_{i}^{OLS, \, g}$.
Thus to recover $\mathbb{E}(\mu_i)$, we suggest a two-step identification argument. It is not uncommon to use two-step approaches to handle endogeneity, as in the construction of the TSLS estimator for a linear regression with endogenous regressors. Our two steps are more involved due to $\mu_i$ varying across units and the nonparametric take on the endogeneity. The first step will focus on recovering the functions $g_t = f_{t+1} - f_t$ and the second on recovering $\mathbb{E}(\mu_i)$. We now provide an intuitive explanation of what these steps are.
\textbf{\textit{Step 1:}} To isolate $g$, we use variations across units. Let us first fix a unit $i$ and compute the residuals of the unit-specific linear regression of $\dot y_{it} $ on $\dot x_{it}$. Denote the vector of these residuals with $R_i$, it is of size $T-1$ and obtained by multiplication of the vector $\dot{y}_i$ by the residual-maker matrix. There are two components in this residual: $R_{i}^u$, the vector of residuals of the regression of $\dot u_{it}$ on $\dot x_{it}$, and $R_{i}^g$, the vector of residuals of the regression of $g_t(V_i)$ on $\dot x_{it}$. We have
$$R_i = R_{i}^g+ R_{i}^u.$$
Note that $T-1 >d_x$ is now needed. Indeed, if $T-1=d_x$ and $\dot X_i$ is of full rank, the regression fit is perfect and $R_i = 0$ independently of $g$ and $u$. There is no residual variation left that we can exploit because any unit-specific vector is a linear combination of the regressors.
Both $g_t(V_i)$ and $\dot u_{it}$ vary over time, i.e., across observations. However $g_t$ is a function of $V_i$ while $\dot u_{it}$ is mean independent of $X_i$ given $V_i$.
Thus focusing on the units $i$ such that $V_i = V$ with $V$ a fixed value, by mean independence the residuals of the regression of $\dot u_{it}$, that is, $R_{i}^u$, average to $0$ in this sub-population.
On the other hand, the average over this same sub-population of the vector of residuals of the regression of $g_t(V_i)$ on $\dot x_{it}$, that is, of $R_{i}^g$, is an average of regression residuals over regressions which all have the same vector of outcomes $g(V)$, but have different values of regressors thus of the residual-maker matrices. Therefore, this vector of outcomes $g(V)$ can be isolated by multiplication of the average of $R_{i}^g$ in the subpopulation of interest, by the inverse of the average of the residual-maker matrices in the same population. Lastly, this is also true if we replace $R_{i}^g$ with the observed $R_{i}$ since $R_{i}^u$ averages to $0$.
\textbf{\textit{Step 2:}} Once the function $g$ is identified on the support of $V_i$, we can recover $\mathbb{E}(\mu_i)$ using the intuition described at the beginning of this section. Recall that the difference between $\beta_{i}^{OLS}$ and $\mu_i$ is $\beta_{i}^{OLS, \, g} + \beta_{i}^{OLS, \, u}$. Since $\beta_{i}^{OLS, \, g}$ is the coefficient of the regression of $g_t(V_i)$ on $\dot x_{it}$ for unit $i$ and $g$ is identified, it is known. Since $\beta_{i}^{OLS, \, u}$ averages to $0$ across units, we consider once again cross-section averages: $\mathbb{E}(\mu_i)$ is the average of the difference between $\beta_{i}^{OLS}$ and $\beta_{i}^{OLS, \, g}$ across units.
\medskip
We now formalize the steps in the argument described above. Crucial to this argument is the use of two matrices constructed with the regressors: the unit-specific regression matrix $Q_i= (\dot{X}_i' \dot{X}_i)^{-1} \dot{X}_i'$, and the unit-specific residual-maker matrix $M_i = I_{T-1} - \dot{X}_i (\dot{X}_i' \dot{X}_i)^{-1} \dot{X}_i'$ if $\dot{X}_i$ is of full rank or $ M_i = I - \dot{X}_i \dot{X}_i^{ +}$ if not, where $\dot{X}_i^{ +}$ is the Moore Penrose inverse.
The matrix $Q_i$ is defined only if $\dot{X}_i$ has full column rank and if so, is of size $d_x \times (T-1)$ while $M_i$ is of size $(T-1) \times (T-1)$.
These unit-specific matrices have been used in analysis of CRC panel models with short $T$, see e.g \cite{ab12}, \cite{gp12}, but also by a literature focusing more on the large $T$ case, see e.g \cite{s70,ps95,h14}. Note that $R_i = M_i \dot y_i$ and $\beta_{i}^{OLS} = Q_i \dot y_i$.
The identification argument relies on averages of quantities functions of these matrices: we assume that such moments exist and this corresponds to our second assumption, Assumption \ref{assn:id_finiteE}. See more details in Section \ref{sec:assn_CRC_matrix}.
\textbf{\textit{Step 1:}} Left multiplication of (\ref{eq:modelwTdiffVector}) with $M_i$ gives
\begin{align}
R_i = M_i \dot{y}_i \, = M_i g(V_i) + \, M_i \dot{u}_i , \qquad \mathbb{E}( M_i \dot{u}_i | X_i , V_i) =0, \label{eq:model_id_g}
\end{align}
where the conditional expectations are $0$ because $M_i$ is a function of $\dot{X}_i$. The discussion above suggests focusing on the subpopulation with the same value of $V_i$, which gives
\begin{equation}\label{eq:almost_closedform_g}
\mathbb{E}(M_i \dot{y}_i | V_i = V) \, = \, \mathbb{E}( \, M_i | V_i = V) \, g(V) \, = \, \mathcal{M}(V) g(V),
\end{equation}
where we write $\mathcal{M}(V) = \mathbb{E}( M_i | V_i = V)$.
If $\mathcal{M}(V)$ is invertible for a given value $V$ on the support of $V_i$, then (\ref{eq:almost_closedform_g}) gives the following closed-form expression, $g(V) = \mathcal{M}(V)^{-1} \mathbb{E}(M_i \dot{y}_i | V_i = V)$, that is, we retrieve $g(V) = \mathcal{M}(V)^{-1} \mathbb{E}(R_i | V_i = V)$. Thus our last assumption, Assumption \ref{assn:id_M_GLn} (see Section \ref{sec:IA}), will be invertibility of $\mathcal{M}(V_i)$ almost surely. Note that writing $\mathbb{E}(M_i \dot{y}_i | X_i , V_i) = M_i g(V_i)$ does not identify $g(V_i)$ because $M_i$ is a projection matrix and singular whenever $\dot X_i\neq 0$.
\textbf{\textit{Step 2:}} Left multiplication of (\ref{eq:modelwTdiffVector}) with $Q_i$ gives
\begin{align}
&\beta_{i}^{OLS} = Q_i \dot{y}_i = \mu_i + \, Q_i g(V_i) + Q_i \dot{u}_i, \text{ and }\mathbb{E}( Q_i \dot{u}_i | X_i , V_i) =0, \label{eq:model_id_mu} \\
& \Rightarrow \mathbb{E}(\mu_i) = \mathbb{E}\left(Q_i \dot{y}_i - Q_i g(V_i) \right), \label{eq:intuitionmu}
\end{align}
which identifies $\mathbb{E}(\mu_i)$. Note that (\ref{eq:intuitionmu}) can be rewritten $\mathbb{E}(\mu_i) = \mathbb{E}(\beta_{i}^{OLS} - \beta_{i}^{OLS, \, g})$.
This short exposition gave the main intuition of the identification strategy. We postpone the formal statement of the results to Section \ref{sec:identification} and first discuss the assumptions we introduced.
\section{Empirical Content of the Assumptions}\label{sec:assn}
This section discusses the main assumptions on which the identification results are based. As some of these assumptions are nonstandard, we provide intuition and introduce situations in which they might or might not hold. We start by discussing our use of the control function approach.
\subsection{Control Function Approach}\label{sec:CFAcomment}
To identify structural parameters of models with endogenous regressors, two well-known approaches are the instrumental variable approach and the control function approach. The assumption we use is in the tradition of the control function approach (CFA) literature. Under CFA assumptions, conditional on the control variables, the endogenous regressor is either independent or mean independent of the residual. In nonlinear models this has been leveraged to identify various parameters of interest,
see e.g \cite{npv99}, \cite{in09}, \cite{fhmv08}, \cite{df15}, \cite{t15}, etc.
The identification results of this paper similarly rely on control variables to control for the dependence between the regressors and the residual but the exact assumption that we maintain,
although it remains a conditional mean independence assumption,
differs from usual versions of CFA assumptions due to the panel aspect of the data.
It is formally stated below in Assumption \ref{assn:idCFA} (\ref{assn:idCFA_CFA}).
\begin{assumption}\label{assn:idCFA}
\
\begin{enumerate}[noitemsep,nolistsep]
\item \label{assn:idCFA_continuity}
$\left(X_i, Z_i, \epsilon_{i}, \mu_i, \alpha_i \right)$ is i.i.d.\ across $i$, and (\ref{eq:model}) holds with $\mathbb{E}(\epsilon_{it})=0$,
\item \label{assn:idCFA_CFA}
There exist scalar-valued functions $(f_t)_{t \leq T}$ and identified functions $(C_t)_{t \leq T}$ such that, defining $v_{
it} = C_t(X_i, Z_i) \in \mathbb{R}^{d_v}$,
\vspace{-14pt}
\begin{equation}\label{eq:assnCFA}
\mathbb{E}(\epsilon_{it} \, | \, x_{i1},.. x_{iT}, v_{i1},.. v_{iT}) = f_t(v_{i1},.. v_{iT})
\end{equation}
\end{enumerate}
\end{assumption}
Note that (\ref{eq:assnCFA}) can be rewritten $\mathbb{E}(\epsilon_{it}|X_i, V_i) = f_t(V_i)$. It can be compared in particular to an assumption maintained in \cite{npv99}, which studies a triangular simultaneous equation model in a cross-sectional setting where the regressors $x_{it}$ in the outcome equation include endogenous regressors $x_{it}^{en}$ and exogenous ones $x_{it}^{ex}$. Using our notations and applying their framework to a fixed period $t$ for the sake of the comparison, the control variable $v_{it}$ used by the authors is the residual of a nonparametric reduced form first-stage regression of $x_{it}^{en}$ on instruments $z_{it}$ and $x_{it}^{ex}$.
They assume that $\mathbb{E}(\epsilon_{it} | x_{it}^{en}, z_{it},v_{it})= f_t(v_{it})$, which we call CS-CFA. CS-CFA implies $\mathbb{E}(\epsilon_{it} | x_{it},v_{it})= f_t(v_{it})$. The main difference between this equality and (\ref{eq:assnCFA}) is that the latter conditional expectation conditions on the vector $(X_i,V_i)$ and not $(x_{it},v_{it})$: we discuss later how this brings additional flexibility and describe now how the practitioner can construct control variables.
In the CFA literature, control variables are typically provided by a first-step selection equation. Although we do not rely on a particular specification for the control variables, a natural approach is to construct, for each period, control variables studied in the cross-section CFA literature using cross-sectional data for this period so as to obtain $v_{it} = C_t(x_{it}, z_{it})$. For instance, following \cite{npv99}
with a linear first stage, we obtain the following two-equation model,
\vspace{-25pt}
\begin{align}
y_{it} & = \, x_{it}^{ex \ \prime} \, \mu_i^{ex} + x_{it}^{en \ \prime} \, \mu_i^{en} + \epsilon_{it}, \label{eq:modelX1X2} \\
x_{it}^{en} & = \pi_t x_{it}^{ex} + \gamma_t z_{it} + v_{it}, \qquad \mathbb{E}(v_{it}| x_{it}^{ex}, z_{it} )=0 \label{eq:modelX1X2lin1stage}
\end{align}
where $ x_{it}^{ex} \in \mathbb{R}^{d_1}$ is a vector of exogenous regressors and $x_{it}^{en} \in \mathbb{R}^{d_2}$ a vector of potentially endogenous regressors. Define $x_{it}=(x_{it}^{ex \ \prime},x_{it}^{en \ \prime})'$, $d_x = d_1+d_2$. If the instrument satisfies a natural extension of CS-CFA, i.e, $\mathbb{E}(\epsilon_{it} | Z_i, X_i^{ex}, V_i) = h_t(v_{it})$, then (\ref{eq:assnCFA}) holds with $C_t(x,z) = x^{en} - \mathbb{E}(x_{it}^{en}|x_{it}^{ex} = x^{ex},z_{it} = z)$ and $f_t(V_i)=h_t(v_{it})$.
To study the impact of class size on students' achievements in a panel of class level data, a well-known instrument is random variations in cohort size, see \cite{h00}, and control variables are the vector of residuals of a first-stage regression of class size or its logarithm on an instrument.
\cite{er99} use state cigarette taxes on birth data to instrument for mother level of smoking in an analysis of its impact on birth weight: the instrument can be used on a panel of mothers with multiple births to construct $v_{it}$ as the residual of the first-stage linear regression.
For nonseparable first stages $x_{it}^{en} = h(x_{it}^{ex}, z_{it},\eta_{it}) $ where $(\epsilon_{it},\eta_{it}) \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} z_{it}$ and $\eta_{it}$ is scalar-valued, another well-known control variable studied in \cite{bp03} and \cite{in09} is $v_{it} = F_{x^{en}|x^{ex},z}(x_{it}^{en}|x_{it}^{ex}, z_{it})$.
Aside from first-stage equations, the panel data literature has suggested various assumptions under which variables can be constructed without instrumental variables yet satisfy conditional independence conditions similar to (\ref{eq:assnCFA}). Exchangeability conditional on individual-level heterogeneity is one such condition, see \cite{am05}, or the use of sufficient statistics, as in \cite{ai19}. See also \cite{lps21} for a discussion of these conditions.
Assuming exchangeability in (\ref{eq:model}) would amount to assume that $\alpha_i,\epsilon_{it}|(x_{i1},...,x_{iT}) $ has the same distribution as $\alpha_i,\epsilon_{it}|(x_{i\tau(1)},...,x_{i\tau(T)})$ for any permutation $\tau$: this in general is contrary to our goal of addressing time-varying endogeneity where timing matters. However this motivates strengthening Assumption \ref{assn:idCFA} (\ref{assn:idCFA_CFA}) to $\mathbb{E}(\epsilon_{it}|X_i, V_i) = f_t(s(V_i))$ where $s(V_i)$ is scalar. For instance, one could have $s(V_i) = \frac{1}{T} \sum_{t=1}^T v_{it}$. An obvious benefit of this assumption is that it will alleviate the curse of dimensionality as all nonparametric estimations will be conditional on $s(V_i)$ instead of $V_i$, a vector of dimension $d_V \times T$.
One important difference between (\ref{eq:assnCFA}) and well-known applications of the CFA in nonseparable cross-sectional models is that we do not impose restrictions on the distribution of $(\mu_i, \alpha_i, (\epsilon_{it})_{t \leq T}, Z_i, X_i)$ but only on that of $((\epsilon_{it})_{t \leq T}, Z_i,X_i)$. The CFA typically includes the entire vector of unobserved heterogeneity, see e.g \cite{npv99}, \cite{in09}, \cite{mt16}, etc.
Under Assumption \ref{assn:idCFA} (\ref{assn:idCFA_CFA}), the instrument need not be independent of the impact of the regressor $\mu_i$, and of $\alpha_i$.
Despite the added dimensions of unobserved heterogeneity in (\ref{eq:model}) when compared to linear panel data model with constant slope, additive fixed effects and time-varying endogeneity, the requirement on the instrument is only on its correlation with $\epsilon_{it}$ thus remains comparable to exogeneity assumptions of the CFA type in the latter model. Thus the practitioner, when choosing instruments, need not worry about additional exogeneity requirements.
We now explore the serial correlations allowed in (\ref{eq:assnCFA}), noting that $\mathbb{E}(\epsilon_{it}|X_i,V_i)$ can depend on past and future values of $v$.
We consider the conceptually easier case where $v_{it} = C_t(x_{it}, z_{it})$. In the extension of CS-CFA mentioned above, i.e, $\mathbb{E}(\epsilon_{it} | Z_i, X_i^{ex}, V_i)= h_t(v_{it})$, the control variables are strictly exogenous with respect to the residual $\epsilon_{it} - h_t(v_{it})$. This approach is used for instance in \cite{w95}, which studies a linear panel data model with additive unobserved heterogeneity and sample selection issues.
However letting the conditional expectation $\mathbb{E}(\epsilon_{it}|X_i, V_i)$ take the exact form $f_t(V_i)$ allows for the error terms $\epsilon_{it}$ to be correlated with past and future values of the regressors as long as this dependence is captured by $V_i$. Of particular relevance in panel data is the correlation between contemporaneous error terms and future regressors, also called \textit{feedback}. To allow for a feedback effect, the assumption of strict exogeneity is often replaced with \textit{sequential exogeneity}. Sequentially exogenous or \textit{predetemined} regressors are uncorrelated with present and future values of the error term $\epsilon$ but may be correlated with past values of the error term, see e.g \cite{ah01}, \cite{a03}. In Section \ref{sec:relax_SE},
we show that we can formally leverage the flexibility of the functional form $f_t(V_i)$ to identify $\mathbb{E}(\mu_i)$ in a model with contemporaneously exogenous but predetermined regressors, which to the best of our knowledge is an open question in CRC panel models.
Moreover feedback can also occur when there is contemporaneous time-varying endogeneity. In the context of evaluating the impact of smoking on birthweight, a mother may change her smoking level after giving birth to a low birthweight infant ``$t$''. This can be captured by Assumption \ref{assn:idCFA} since in (\ref{eq:assnCFA}) $\epsilon_{it}$ is allowed to be correlated with the residuals of the first-stage regressions for the following births. Another example is in Section \ref{sec:illus}.
Note also that assuming $f_t(V_i) = h_t(v_{it})$ would not simplify the identification argument of Section \ref{sec:id_model} and cannot alleviate the curse of dimensionality as the left multiplication with $M_i$ in (\ref{eq:almost_closedform_g}) involves linear combinations of all periods.
A drawback of using the control function approach is that recovering valid control variables from first-stage equations is not always possible if there are additional unobserved sources of heterogeneity. Looking at a specific selection equation with two-dimensional unobserved heterogeneity, \cite{i07} shows that the conditional quantile of the regressor given the instrument is not a valid control variable anymore. There is consequently an apparent imbalance in the degree of heterogeneity between the implicit first stage and the main model equation (\ref{eq:model}).
Such imbalance can be found in other analyses of random coefficient models with endogenous regressors, see e.g \cite{hv98}, \citeauthor{w97} (\citeyear{w97}, \citeyear{w03}). As in \cite{mt16}, our approach allows for nonlinear first-stage equations which guarantee a certain degree of heterogeneity as partial effects vary across agents. It may be possible to use panel techniques to identify control variables in first-stage equations with more heterogeneity: for instance if a first stage is $x_{it} = m(z_{it})+ \xi_i + \eta_{it}$ where $\mathbb{E}(\eta_{it}|z_{it})=0$ then $v_{it} = \xi_i + \eta_{it}$ is identified if $T \geq 2$ and may be a valid control variable. However generalizing this approach to more complex models is outside the scope of this paper and is left for future research.
\subsection{Within and Between Matrices}\label{sec:assn_CRC_matrix}
To separately identify the function $g$ from $\mu_i$, the identification argument relies on the matrices
$ M_i = I_{T-1} - \dot{X}_i (\dot{X}_i' \dot{X}_i)^{-1} \dot{X}_i' \, $ and $Q_i = (\dot{X}_i' \dot{X}_i)^{-1} \dot{X}_i'$, called within- and between-matrices respectively, and considers expectations of vectors left-multiplied with these matrices. First, we require $\dot{X}_i$ to have full column rank with probability $1$. This section explains how the use of $M_i$ and $Q_i$ further imposes constraints on the data.
Note that if $T = d_x + 1$ then $M_i = 0$: as discussed in Section \ref{sec:intuition_id}, if $\dot{X}_i$ is of full rank then there is no variation left in the residual $R_i$ to identify the function $g$. This fact motivates the constraint $T \geq d_x + 2$, standard in the analysis of CRC panel models, see \cite{ch92,ab12}.
A second issue is that the matrix $Q_i$ and functions of it may not have finite moments. Indeed the norm of $ (\dot{X}'\dot{X})^{-1}\dot{X}$ increases without bound as $\det(\dot{X}'\dot{X})$ approaches $0$, that is, as the columns of $\dot{X}$ are ``close to'' being linearly dependent. Thus if the regressors are continuous and for instance the density of $\dot{X}$ is positive at values such that $\det(\dot{X}'\dot{X}) = 0$,
simple moments involving $Q_i$ such as $\mathbb{E}(||Q_i \dot{u}_{i}||)$ may not exist.
For these reasons, we assume that the needed moments do exist. The formal assumption is stated below. Note that this does not affect $M_i$: it is an orthogonal projection matrix and thus bounded.
\begin{assumption}\label{assn:id_finiteE}
\
\begin{enumerate}[noitemsep,nolistsep]
\item \label{assn:full_rk} $\dot{X}_i$ is of full column rank a.s,
\item $\mathbb{E}(||\dot{u}_{i}||) < \infty$, $\mathbb{E}(||Q_i \dot{u}_{i}||) < \infty$, $\mathbb{E}(||Q_i g(V_i)||) < \infty$, and
$\mathbb{E}(||(\mu_i', \alpha_i)||) < \infty$.
\end{enumerate}
\end{assumption}
Assumption \ref{assn:id_finiteE} implies that almost surely, units have strictly positive $\det(\dot{X}_i'\dot{X}_i)$, i.e., regressors variations must be linearly independent and thus each regressor must have variations over time of its first-differences. But Assumption \ref{assn:id_finiteE} also implies that the proportion of units with low values of $\det(\dot{X}_i'\dot{X}_i)$ must be sufficiently low. Such units with high persistence are called \textit{stayers}.
This issue also arises when there is no time-varying endogeneity and is discussed in details in \cite{gp12}.
The authors illustrate it with an extreme example where they consider the case $d_x = 1$ and $x_{it} = s_i w_{it}$, with $s_i \sim \mathcal{U}[a,b]$ and $w_{it} \sim \text{i.i.d }\mathcal{N}(0,1)$. They show that for fixed $b>0$ and independently of the number of periods, the moment $\mathbb{E}\left((X_i'X_i)^{-1}\right)$ is not finite if $a\leq 0$, i.e., if regressors with perfect persistence (when $s_i = 0$) have positive density despite the matrix of regressors being of full rank almost surely.
If Assumption \ref{assn:id_finiteE} does not hold and there are not sufficient regressor time variations, i.e., \textit{within}-variations of the regressors for all units in the sample, an alternative is to focus on units for which there is sufficient variation, i.e., such that $\det(\dot{X}_i' \dot{X}_i) > \delta$ for some threshold $\delta \geq 0$. We show in Section \ref{sec:maindid} that $\mathbb{E}(\mu_i | \det(\dot{X}_i' \dot{X}_i) > \delta)$ is also identified under standard assumptions and focus on an estimator of this parameter\footnote{This approach is pursued in \cite{ab12}.} in Section \ref{sec:estim} since samples with substantial shares of stayers are more likely. To identify the APE for the entire population in cases where Assumption \ref{assn:id_finiteE} does not hold or where $T = d_x + 1$, \cite{gp12} provide an alternative procedure: letting the threshold $\delta$ go to $0$.
Although it is beyond the scope of this paper, this approach can be used here.
\subsection{Invertibility Assumption}\label{sec:IA}
This last assumption is perhaps the least standard but is crucial for the derivation of the main results. For $\mathcal{M}(V) = \mathbb{E}( \, M_i \, | V_i = V) $, the assumption is stated below.
\begin{assumption}\label{assn:id_M_GLn}
The matrix $\mathcal{M}(V_i)$ is invertible, $\mathbb{P}_{V_i} \, \text{a.s}$.
\end{assumption}
We call Assumption \ref{assn:id_M_GLn} the invertibility assumption (IA). Imposing invertibility of a matrix for identification is not surprising: in the linear regression model in cross-section, $y_i = w_i'\beta+u_i$ with $w \in \mathbb{R}^{d_w}$, a necessary condition for identification of $\beta$ is nonsingularity of $\mathbb{E}(w w')$. It may be surprising however that the expectation of the singular matrices $M_i$ can be assumed to be nonsingular, but taking the same linear regression example, $w_iw_i'$ is also singular for each unit $i$.
Importantly, IA is an assumption on the conditional expectation of a function of $\dot{X}_i$ given $V_i$, which itself is a function of $X_i$ and $Z_i$. Thus IA is an assumption on the joint distribution of $(X_i,Z_i)$: it assesses relevance of the instrument.
However due to its unusual form, what IA imposes on $Z_i$ is not immediate and not easily comparable with standard conditions. Therefore we now have a closer look at its empirical content and obtain conditions on the distribution of $(Z_i,X_i,V_i)$ that are sufficient for IA to hold. Section \ref{sec:minsupport} provides a geometric intuition on why the instrument can have discrete support and derives the minimum number of points on its support.
Section \ref{sec:sufficientcondIA} specifies a generic first-stage equation and under various restrictions obtains sufficient conditions.
These conditions are not restrictive: if the instrument is continuous, we obtain a small support condition similar to \cite{fhmv08} and we show that IA holds for a binary instrument $Z_i$ when $x_{it}$ is scalar even if $Z_i$ impacts the regressor at only one period. We note here that our discussions of IA mostly assume that $\dot{X}_i$ is of full column rank $\mathbb{P}_{\dot{X}_i|V_i=V} \, \text{a.s}$, which relates with Assumption \ref{assn:id_finiteE}. We return to this point when discussing Result \ref{rslt:evdok_z}.
We start by showing that IA is necessary and sufficient to identify $g$.
\subsubsection{A Necessary Condition}
Under IA, if a function $\dot{h}: \mathcal{S}_{V_i} \mapsto \mathbb{R}^{T-1}$ is such that $\dot{h}(V_i) = \dot{X}_i \beta_i \text{ a.s}$, then $M_i \dot{h}(V_i) = 0 \Rightarrow \mathcal{M}(V_i) \dot{h}(V_i) = 0 \Rightarrow \dot{h}(V_i) = 0\text{ a.s}$.
IA precludes any nonzero function $\dot{h}$ of $V_i$ from being of the form $\dot{X}_i \beta_i$ or, to rephrase it, precludes any time-varying function $h: \mathcal{S}_{V_i} \mapsto \mathbb{R}^{T}$ of $V_i$ from being of the form $X_i \beta_i$. This separately identifies $g$ from $\dot{X}_i \mu_i $: we show that IA is a necessary and sufficient condition for identification of the function $g$.
\begin{result}\label{rslt:cns_g}
Let Assumptions \ref{assn:idCFA} and \ref{assn:id_finiteE} hold. If $V_i$ is continuously distributed, then $g$ is identified $\mathbb{P}_{V_i} \, \text{a.s}$ on $\mathcal{S}_{V_i}$ if and only if Assumption \ref{assn:id_M_GLn} holds.
\end{result}
\subsubsection{Minimum Number of Support Points for $Z$}\label{sec:minsupport}
Going back to the vocabulary used in Section \ref{sec:intuition_id}, identification of $g(V)$ focuses on
the subpopulation with fixed $V_i=V$.
For this subpopulation we obtain the average vector of residuals of the regression of $g_t(V_i)$ on $\dot x_{it}$, i.e., $\mathbb{E}(R_i^g|V_i=V)$.
For a unit $i$ such that $V_i=V$, $g(V_i)$ cannot be separated from $\dot X_i \mu_i$ if $g_t(V_i)$ is linear in the regressors $\dot x_{it}$, i.e., if $g(V_i)$ is a linear combination of the columns in $\dot X_i$, as in the standard ``no perfect multicollinearity assumption''.
Let $\operatorname*{Span} \dot{X}_i$ denote the set spanned by the columns of $\dot{X}_i$. The argument above implies that we must have $g(V) \notin \operatorname*{Span} \dot X_i$ for any $i$ such that $V_i=V$. Since $g(V)$ is nonparametrically specified and could take any value, we need to make sure that no nonzero vector can be in the intersection of all $\operatorname*{Span} \dot X_i$ for $i$ such that $V_i=V$. Result \ref{rslt:equiv_span} below follows from this intuition and its interpretation is that within the subpopulation with fixed $V_i=V$, the matrices of regressor variations must vary sufficiently.
\begin{result}\label{rslt:equiv_span}
Fix $V \in \mathcal{S}_{V_i}$, the matrix $\mathcal{M}(V)$ is invertible if and only if for all subsets $\tilde{\mathcal{S}}\subset \mathcal{S}_{X_i|V_i=V}$ such that $\mathbb{P}_{X_i|V_i=V}\left(\tilde{\mathcal{S}}\right)=1$, then $\ \bigcap_{X \in \tilde{\mathcal{S}}} \operatorname*{Span} \dot{X} = \{0\} $.
\end{result}
In the subpopulation with fixed $V_i=V$, cross-sectional variations of the matrix of regressor variations are generated by cross-sectional variations of the instrument. Thus Result \ref{rslt:equiv_span} provides some evidence that IA can hold when the instrument has discrete support. Consider the case where $d_x=1$ and the instrument is discrete. We maintain that $\dot{X}_i$ is of full column rank $\mathbb{P}_{\dot{X}_i|V_i=V} \, \text{a.s}$. Then $ \bigcap_{X \in \tilde{\mathcal{S}}} \operatorname*{Span}\dot{X} $ is an intersection of straight lines. Since variations of $\dot{X}_i$ holding $V_i$ fixed are generated by variations of $Z_i$ holding $V_i$ fixed, it is a finite intersection of lines. Such an intersection is always the trivial set $ \{0\} $ unless these straight lines are all equal: thus IA imposes that for some two different values of the vector of instruments, the corresponding vectors of regressors $\dot{X}$ are not collinear. This example gives a better understanding of the empirical content of IA: variations of $Z_i$ across individuals need to create time variations of the regressors that are sufficiently diverse in the population $V_i=V$.
Now if $d_x =2$ and $T=4$, then $\operatorname*{Span} \dot{X}$ is a plane in a 3-dimensional set. The intersection of two planes is either a plane or a line. But the intersection of three planes can be $\{0\}$: for $\mathcal{M}(V)$ to be nonsingular for a given $V$, the vector of instruments must have at least 3 points on its support. This points to a tension between the dimension of the random coefficients $d_x$, the number of time periods $T$ and the number of points in the support of $Z_i$, which is formalized in the following result.
\begin{result}\label{rslt:discrete_IV}
Fix $V \in \mathcal{S}_{V_i}$. Assume that $\dot{X}_i$ is of full column rank $\mathbb{P}_{\dot{X}_i|V_i=V} \, \text{a.s}$ and that $\mathcal{S}_{Z_i|V_i=V}$ has $k_V$ points. Then
\vspace{-5pt}
\begin{equation}\label{eq:minZdiscr}
\mathcal{M}(V) \text{is nonsingular } \Rightarrow k_V \geq \frac{T-1}{T-1 - d_x} .
\end{equation}
\end{result}
The minimum number of points on $\mathcal{S}_{Z_i|V_i=V}$ increases with $d_x$ and decreases with $T$. When $T$ is the minimum number of periods required for the framework of this paper to apply, i.e, $T=d_x+2$, then (\ref{eq:minZdiscr}) becomes $k_V \geq d_x+1 $, $d_x+1 $ being the dimension of the unobserved heterogeneity.
Recent contributions show that discrete instruments can be used with the control function approach in cross-sectional data. For instance, in a nonseparable model with scalar unobserved heterogeneity, \cite{df15} and \cite{t15} obtain nonparametric identification with a binary instrument. As the dimension of the unobserved heterogeneity increases, existing results document stricter requirements on the support of the instrument. To identify a nonseparable model with unobserved heterogeneity of unknown dimension, \cite{in09} requires the instrument to be continuous and have large support. On the other hand, \cite{ns18} studies an outcome equation polynomial in an endogenous variable with random coefficients and shows that the number of points needed on the support of the instrument conditional on the control variable is at least as large as the dimension of the unobserved heterogeneity.
See also \cite{ns21}, and \cite{mt16} for a similar result.
This is comparable to (\ref{eq:minZdiscr}) but for its dependence in $T$: higher $T$ translates into more observations for the same draw of $(\mu_i, \alpha_i)$, thus more information.
Note that the constraint (\ref{eq:minZdiscr}) is on the support of $Z_i$ and not $z_{it}$, and the cardinality of the support of $z_{it}$ can potentially be substantially lower.
\subsubsection{Sufficient Conditions for Triangular Systems}\label{sec:sufficientcondIA}
Although applications of the control function approach are often associated with large support conditions on the instruments, the previous discussion implies that IA does not require large support. We now look at specific first-stage equations and illustrate this point with various sufficient conditions on $Z_i$, when $\dot{X}_i$ is of full rank $\mathbb{P}_{\dot{X}_i|V_i=V} \, \text{a.s}$.
We impose the following general equation,
\begin{equation}\label{eq:1ststagenp}
x_{it} = m_t(Z_i,V_i) \in \mathbb{R}^{d_x}, \, \text{ or in vector form, } \, X_i = m(Z_i,V_i) \in \mathbb{R}^{(T-1) \times d_x},
\end{equation}
and its first-difference version $\dot{X}_i = \dot{m}(Z_i,V_i)$.
The discussion preceding Result \ref{rslt:equiv_span} states that to recover $g(V)$, we must ensure that it is not possible for any nonzero vector to be linearly generated by the matrix of regressor time variations $\dot X_i$, for all units $i$ such that $V_i = V$. Thus the regressors variations must be sufficiently different from each other among these units. Let us consider a simple example.
Let us write the components of $x_{it}$ as $x_{it,k}$ where $k=1,..,d_x$ and define similarly $\dot x_{it,k}$. If one regressor $x_{it,k}$ has constant variations among units $i$ such that $V_i = V$, i.e., if
$$ \text{for } t=1,..,T-1, \text{ there exists } \bar{x}_{t,V} \text{ such that } \dot x_{it,k} = \bar{x}_{t,V}, \, \mathbb{P}_{X_i|V_i=V}\text{ a.s.},$$
defining the vector $\bar x_V = (\bar{x}_{1,V},..,\bar{x}_{T-1,V})'$, then
$\bar x_V \in \operatorname*{Span} \dot X_i, \, \mathbb{P}_{X_i|V_i=V}\text{ a.s.}$ Result \ref{rslt:equiv_span} implies therefore that $\mathcal{M}(V)$ is singular if for some $t$, $\bar{x}_{t,V} \neq 0$. This counterexample shows that as $Z_i$ varies on $\mathcal{S}_{Z_i|V_i=V}$, it must generate variations of the regressor time variations in the cross-sectional dimension, that is, of the within-variations of the regressors.
But these cross-sectional variations need not be very rich.
To illustrate this, we consider (\ref{eq:1ststagenp}) in the simpler case $d_x = 1$, and
$Z_i | V_i = V$ has discrete distribution.
Result \ref{rslt:counterex_1cstt} shows that, conditional on $V_i=V$, even when there is so little cross-sectional variation that the instrument is binary and impacts the regressor at only one period $t_0$, $\mathcal{M}(V)$ is nonsingular as long as one of the regressors is not of the form $(b,..b,a,b,..b)'$.
\begin{result}\label{rslt:counterex_1cstt}
Let (\ref{eq:1ststagenp}) hold with $d_x =1$, and fix $V \in \mathcal{S}_{V_i}$, $t_0 \leq T$. Assume that $\dot{X}_i$ is of full column rank $\mathbb{P}_{\dot{X}_i|V_i=V} \, \text{a.s}$, $\mathcal{S}_{Z_i|V_i=V}$ is finite and
there exist $(Z^{(1)},Z^{(2)}) \in \mathcal{S}_{Z_i|V_i=V}$ such that $ m_{t_0}(Z^{(1)},V) \neq m_{t_0}(Z^{(2)},V)$ and $ m_{t}(Z^{(1)},V) = m_{t}(Z^{(2)},V)$ for $t \neq t_0$. Then $\mathcal{M}(V)$ is nonsingular if and only if
\vspace{-18pt}
$$ \exists (t_1,t_2), \ t_1 \neq t_2, \, t_1 \neq t_0, t_2 \neq t_0, \ \text{such that } m_{t_1}(Z^{(1)},V) \neq m_{t_2}(Z^{(1)},V). $$
\end{result}
This result combines variations in the population (at $t_0$), and uniform-over-$i$ variations over time (between $t_1$ and $t_2$).
Thus for $\mathcal{M}(V)$ to be nonsingular, if the instrument generates cross-sectional variations of the regressor at $t_0$ only as it varies on $\mathcal{S}_{Z_i|V_i=V}$, then the regressor must vary over time at least between two periods that are not $t_0$.
Another evidence for the need of both variations over time and across units comes from looking at models where a binary instrument does not vary over time.
Such an example with ``structural breaks'', i.e., time variations in the function $m_t$, is detailed in Appendix \ref{sec:app_struc_break}. It
illustrates that one structural break is in this case not enough, but IA can hold when there are two. Another example with a time-invariant instrument will be given in Section \ref{sec:relax_SE} below.
\vspace{8pt}
IA is a condition on the within-variations of $X_i$ and further insight can also be gained looking at units $i$ such that $x_{it+1}=x_{it}$. Let us call these units \textit{consecutive stayers}.
Fixing $V$, note that Results \ref{rslt:discrete_IV} and \ref{rslt:counterex_1cstt} as well as the discussions around these results examined $\mathcal{M}(V)$ while imposing for $\dot{X}_i$ to be of full column rank $\mathbb{P}_{\dot{X}_i|V_i=V} \, \text{a.s}$. This simplified some of the proofs but is also particularly relevant since we impose Assumption \ref{assn:id_finiteE} for identification of $\mathbb{E}(\mu_i)$. Recall that the second step of the identification argument runs unit-specific linear regressions.
Without this additional restriction, we can consider a dgp such that $\dot X_i = 0,$ $\mathbb{P}_{\dot{X}_i|V_i=V} \, \text{a.s}$. Then $M_i$ is the identity matrix $\mathbb{P}_{\dot{X}_i|V_i=V} \, \text{a.s}$ and $\mathcal{M}(V)$ is invertible.
It is easy to understand why $g(V)$ is identified with this dgp. Indeed $\dot X_i = 0$ implies that $ \forall t \leq T-1, \ x_{it+1}=x_{it}$: thus for a fixed $ t \leq T-1 $,
$$
\mathbb{E}(\dot y_{it} | V_i=V) = \mathbb{E}((x_{it+1} - x_{it})'\mu_i + g_t(V_i) + \dot{u}_{it} | V_i=V) = g_t(V),
$$
which identifies $g_t(V)$. We cannot not consider such a dgp but consecutive stayers can still be leveraged when regressor matrices are of full column rank and the instrument is discrete, as the following result states.
\begin{result}\label{rslt:evdok_z}
Fix $V \in \mathcal{S}_{V_i}$. Assume that $\dot{X}_i$ is of full column rank $\mathbb{P}_{\dot{X}_i|V_i=V} \, \text{a.s}$, $\mathcal{S}_{Z_i|V_i=V}$ is finite and for all $t \leq T-1$, $x_{it} = m_t(z_{it},v_{it})$.
Assume also that for all $t \leq T-1$, there exists $Z^{(t)} \in \mathcal{S}_{Z_i|V_i=V}$ such that $z_{t+1}^{(t)}=z_{t}^{(t)}$. Then $\mathcal{M}(V)$ is invertible.
\end{result}
The full rank condition implies that it is not possible that $z_{s}^{(t)}=z_{s'}^{(t)}$ for all $(s',s)$. There needs to be cross-sectional variation of the period at which consecutive stayers ``stay'', as $Z_i$ varies on $\mathcal{S}_{Z_i|V_i=V}$. This generates cross-sectional variation of the within-variations.
Leveraging these consecutive stayers is common in the panel data literature and important papers using this approach are for instance \cite{hk00,evdo10,hw12,gp12}.
\vspace{8pt}
Although existence of consecutive stayers guarantees invertibility, it is surely not a necessary condition. Indeed, if the instrument is continuous, we obtain a sufficient condition comparable to a well-known result in \cite{fhmv08}. This paper studies a cross-section outcome equation which is polynomial in a scalar endogenous regressor $w_i$ and with random coefficients. The authors impose a nonseparable first stage $w_i = m(z_i,v_i)$ with an instrument $z_i$ independent of the control variable $v_i$ and $m(z_i,v_i)$ assumed to be a continuous function of $v_i$. Then existence of an open interval contained in the support of $m(z_i,v_0)$ for all fixed $v_0$ is shown to be a sufficient condition for identifiability of some average effects. The condition we obtain is a panel version of this restriction, instead on the entire vector of regressors, and implies too that a large support is not needed.
\begin{result}\label{rslt:florenstype}
Let (\ref{eq:1ststagenp}) hold and fix $V \in \mathcal{S}_{V_i}$, assume that $\dot{X}_i$ is of full column rank $\mathbb{P}_{\dot{X}_i|V_i=V}\text{a.s}$ and that the support of $m(Z_i,V)$ contains an open ball, as $Z_i$ varies conditional on $V_i=V$.
Then $\mathcal{M}(V)$ is invertible.
\end{result}
\vspace{-5pt}
IA and the rank condition require the vector of instruments to generate variations of the regressors both over time and across individuals. Result \ref{rslt:florenstype} obtains these variations with a stronger than needed condition that the instrument has $X_i$ vary ``in all directions''.
Note that since we do not impose independence between $Z$ and $V$, the assumption is on the support of $Z$ conditional on a fixed value of $V$.
A direct application of this result is to a linear first stage where the impact of the instrument does not vary with time. Write
\vspace{-5pt}
\begin{equation}\label{eq:lin1ststage_csttA}
\forall \ t \leq T, \ x_{it} = A z_{it} + v_{it},
\end{equation}
\vspace{-7pt}
where $A$ is of size $d_x \times d_z$ and $\mathbb{E}(v_{it}|z_{it}) = 0$.
\begin{corollary}\label{rslt:lin1ststage_csttA}
Let (\ref{eq:lin1ststage_csttA}) hold, where $A$ is of full row rank and fix $V \in \mathcal{S}_{V_i}$. Assume that $\dot{X}_i$ is of full column rank $\mathbb{P}_{\dot{X}_i|V_i=V}\text{a.s}$ and that the support $\mathcal{S}_{Z_i|V_i=V} $ contains an open ball. Then $\mathcal{M}(V)$ is invertible.
\end{corollary}
If $Z \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} V$, one can replace the support condition with ``$\mathcal{S}_{Z_i} $ contains an open ball''.
The condition that $A$ is of full row rank implies that $d_z \geq d_x$.
\begin{note}\label{note:test}
The invertibility assumption is a condition on observable and estimable random variables which implies that it can be tested. A formal test would require testing for the value of the function $V \mapsto\det(\mathcal{M}(V))$ being nonzero for all values of $V$ on $\mathcal{S}_{V_i}$. One can consider for instance the minimum of this function over the support of $V_i$, that is, $\min_{V \in \mathcal{S}_{V_i}} \det(\mathcal{M}(V))$. A test can be constructed using the nonparametric estimator for $\mathcal{M}(V)$, however such a test is nonstandard and studying its test statistic is beyond the scope of this paper.
\end{note}
\subsection{Example: Sequential Exogeneity}\label{sec:relax_SE}
A particular failure of the strict exogeneity condition in (\ref{eq:model}) is sequential exogeneity, which we separate from contemporaneous endogeneity. Sequentially exogenous regressors satisfy the assumption $\mathbb{E}(\epsilon_{i \, t}|x_{i1},..,x_{it}) = 0$ for all $t$. To the best of our knowledge, there does not exist a proof of identification of $\mathbb{E}(\mu_i)$ when only this assumption holds.
In this section, we show that the two-step identification argument obtains identification under certain conditions. In particular, we assume that the regressors follow a Markov process and show that Assumption \ref{assn:idCFA} holds. The instrument is then the regressor at the first period: an interesting question is whether Assumption \ref{assn:id_M_GLn} holds when there are so few variations. It will be shown in a particular example that it is the case, underlying the minimal variations the invertibility assumption imposes on the regressor as the instrument varies.
Equation (\ref{eq:model}) holds and we look at the case where $x_{it}$ is a Markov process. We assume
\vspace{-5pt}
\begin{align}
\text{for }1 \leq t \leq T-1, \ x_{i \, t+1} \, = \, m_t(x_{it}) + \, \eta_{i \, t+1}, \text{ and } (\epsilon_{it}, \eta_{i\, t+1},.. \, , \eta_{iT}) \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} (x_{i1}, \eta_{i2}, \, .. \, , \eta_{it}), \label{eq:markovmodel}
\end{align}
\vspace{-5pt}
where $m_t$ is unknown. This implies that $(\epsilon_{it}, \eta_{i\, t+1},.. \, , \eta_{iT}) \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} (x_{i1}, \, .. \, , x_{it})$, thus sequential exogeneity holds but $x_{it+1}$ may be correlated with $\epsilon_{i \, t}$.
We define $V_i \, = \, (\eta_{i2},..\, , \eta_{iT})$, it is identified as the vector of residuals of nonparametric regressions of $x_{it+1}$ on $x_{it}$. Then for all $t \leq T$,
\vspace{-5pt}
\begin{align*}
\mathbb{E}(\epsilon_{it} \, | \, x_{i1},..\, , x_{iT}, \eta_{i2},..\, , \eta_{iT}) \, & = \, \mathbb{E}(\epsilon_{it}|x_{i1}, \eta_{i2},..\, , \eta_{iT}) \, = \, \mathbb{E}(\epsilon_{it}| \eta_{i\, t+1},..\, , \eta_{iT}) : = f_t(V_i).
\end{align*}
\vspace{-5pt}
Assumption \ref{assn:idCFA} holds\footnote{Note that
independence is stronger than needed. Conditional mean independence of $\mathbb{E}(\epsilon_{it}|x_{i1}, \eta_{i2},..\, , \eta_{iT})$ with respect to $x_{i1}$ is sufficient.} and $V_i$ is a valid vector of control variables.
Note that one could alternatively consider $x_{i \, t+1} = m_{t+1}(x_{it},\eta_{i \, t+1})$ with $\eta_{i \, t+1}$ scalar and $m_t$ strictly monotonic in $\eta_{i \, t+1}$ and use the control variable suggested in \cite{in09}. Defining $M_i \, = \, I - \dot{X}_i (\dot{X}_i' \dot{X}_i)^{-1} \dot{X}_i' \, $, $\mathcal{M}(V_i) \, = \, \mathbb{E}(M_i|V_i)$, $u_{it} = \epsilon_{it} - f_t(V_i)$ and $g_t(V_i) \, = \, f_{t+1}(V_i) - f_t(V_i)$,
\vspace{-5pt}
\begin{equation}\label{eq:idgwithSeqExog}
\mathbb{E}(M_i \dot{y}_{i} | V_i) \, = \mathcal{M}(V_i) g(V_i), \text{ and } \mathbb{E}(Q_i \dot{y}_i) = \mathbb{E}(\mu_i) + \mathbb{E}(Q_i g(V_i)).
\end{equation}
\vspace{-5pt}
It follows that the two-step identification of the main model applies here: a first step identifies the vector of functions $g$ and the second step identifies the average effect. We thus need to assume invertibility of $\mathcal{M}(V)$ on the support of $V$. This condition is nontrivial here because $V_i = (\eta_{i2},.. \eta_{iT})$ while $M_i$ is a function of the vector of variables $X_i$: this implies that the expectation of $M_i$ conditional on $V_i$ is an expectation over $x_{i1}$,
that is, the only source of variation once $V_i$ is fixed is $x_{i1}$. To show that Assumption \ref{assn:id_M_GLn} can hold with so little variation, let us look more closely at the case where
$x_{it}$ is a scalar AR(1) process.
\begin{assumption}\label{assn:relax_SE}
\
\begin{enumerate}[noitemsep,nolistsep]
\item (\ref{eq:model}) and (\ref{eq:markovmodel}) hold, and $m_t(x) = \rho x$, $\rho \notin \{0,1\}$
\item $(X_i, \mu_i, \alpha_i, \epsilon_i, \eta_i)$ is i.i.d across $i = 1..n$,
\item Either $x_{i1}$ has a discrete distribution with at least two support points, or $x_{i1}$ is continuously distributed.
Moreover, $(\eta_{i2},.. ,\eta_{iT})$ is continuously distributed on $\mathbb{R}^{T-1}$.
\end{enumerate}
\end{assumption}
In this case, the instrument $x_{i1}$ does not vary over time but its impact on each period $x_{it}$, i.e., $\rho^{t-1}$, creates sufficient time variation of the regressor to ensure that IA holds. We obtain the following result.
\begin{result}\label{rslt:relax_SE}
Under Assumptions \ref{assn:id_finiteE} and \ref{assn:relax_SE}, IA holds and $\mathbb{E}(\mu_i', \alpha_i)$ is identified.
\end{result}
We note that if there is contemporaneous endogeneity and if the control variables $v_{it}$ are identified by a cross-section regression of $x_{it}$ on $z_{it}$, the approach developed in this subsection can be applied to cases where $z_{it}$ is not strictly exogenous but sequentially exogenous, in particular if it follows a similar Markovian structure.
The instrument would then be $z_{i1}$ instead of $Z_i$. We use this approach in Section \ref{sec:illus} and in a structural model developed in Section \ref{sec:app_struct_model} of the Appendix.
\begin{note}
While the model exposed above allows for the regressor to be correlated with past disturbances, the assumptions are not compatible with a lagged dependent variable as regressor in the outcome equation (\ref{eq:model}). Indeed writing $x_{i \, t+1} \, = \, m_t(x_{it}) + \, \eta_{i \, t+1}$ implicitly imposes a homogeneous dependence on the past value that is not consistent with the specification $y_{i \, t+1} = \mu_i y_{it} + \epsilon_{it+1}$ and a fixed-effect approach.
\end{note}
\begin{note}
The control variables introduced in Assumption \ref{assn:idCFA} may come from various types of first-stage equations: this is leveraged here to obtain identification. It can be done in other panel models with random coefficients similar to Model (\ref{eq:model}), where other types of violations of the strict exogeneity condition occur. For example, in a previous version of this paper, we applied the idea to a panel CRC model with sample selection where the propensity score is used as the control variable. This will now appear in a separate paper.
\end{note}
\subsection{Sample selection}
Consider a panel model with random coefficients and sample selection. Loosely speaking, if the selection is correlated with the disturbance of the main equation, an endogeneity problem arises, since the regressors of the selected individuals will be correlated with the disturbance as well. \cite{dnv03} study a nonparametric model of sample selection in a cross sectional setting and address the endogeneity issue with a selection equation which provides them with a control variable.
The selection equation studied here is similar and some of the arguments closely follow theirs, but the outcome equation differs: as in (\ref{eq:model}) it is a panel random coefficients specification.
The selection model we consider is
\begin{align}
y_{it}^* \, & = \, x_{it}^{ \, \prime} \, \mu_i + \alpha_i + \epsilon_{it} , \nonumber \\
d_{it} \, & = \, \mathbbm{1} \left( \eta_{it} \leq C_t(x_{it},z_{it}) \right), \label{eq:modelwSelection} \\
y_{it} \, & = \, d_{it} \,y_{it}^*, \nonumber
\end{align}
where $z_{it}$ is an instrument.
Let $d_i \, = \, (d_{it})_{t \leq T}$, and write $d_i = 1$ to denote the event that $d_{it}=1$ for all $t\leq T$.
Also, let $p_{it} \, = \, \mathbb{E} \left( d_{it} | \, x_{it}^1,z_{it} \right) \, = \, \mathbb{P} \left( \eta_{it} \leq C_t(x_{it},z_{it}) \right)$, $P_i \, = \, (p_{it})_{t \leq T}$, and assume that for each $t$ there is a function $f_t$ such that for all $t \leq T$,
\begin{equation} \label{eq:CFAwSelection}
\mathbb{E} \left( \epsilon_{it} | d_i= 1, X_i, P_i \right) \, = \, f_t(P_i).
\end{equation}
Note that as pointed out in \cite{dnv03} in the cross-sectional case, this assumption is satisfied in particular if $ (\epsilon_{is}, \eta_{is})_{s \leq T} \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} (X_i,Z_i )$ and if the cdf of $\eta_{t}$, $F_t$, is strictly increasing. Indeed, in this case defining $\nu_{it} \, = \, F_{t}(\eta_{it})$, $\nu_i \, = \, (\nu_{it})_{t \leq T}$, $p_{it} \, = \, F_{t}(C_t(x_{it},z_{it}))$, then $d_{it} \, = \, \mathbbm{1} \left( \nu_{it} \leq p_{it} \right)$ and
\begin{equation*}
\mathbb{E} \left( \epsilon_{it} | d_i= 1, X_i, P_i \right) \, = \, \mathbb{E} \left( \mathbb{E} \left( \epsilon_{it} | \nu_i, X_i, Z_i \right) | d_i= 1, X_i, P_i \right) = \, \mathbb{E} \left( \mathbb{E} \left( \epsilon_{it} | \nu_i \right) | \nu_i \leq P_i \right) \, : = \, f_t(P_i),
\end{equation*}
as desired (where by an abuse of notation the inequality $\nu_i \leq P_i$ denotes the inequality component by component). Note that the joint distribution of $(\epsilon_{it}, \eta_)$ is unrestricted under these assumptions.
The conditional expectation has a form similar to the control function assumption we maintained in the identification section on the main model, where the control variable is now $p_{it}$ and is identified through a cross sectional regression of $d_{it}$, for each period $t$.
Identification can thus be obtained by a similar two-step argument. The important difference is that all the conditional expectations are evaluated for the subsample such that $d_i = 1$, that is, the subsample of individuals who are selected in all periods.
To be more precise, define $u_{it} \, = \, \epsilon_{it} - f_t(P_i)$, $\tilde{x}_{it} \, = \, d_{i\, t+1}x_{i\, t+1} - d_{it} x_{it}$, $\tilde{g}_t(P_i) \, = \, d_{i\, t+1}f_{ t+1}(P_i) - d_{it} f_{t}(P_i)$, and similarly, $\tilde{u}_{it}$ and the matrices and vectors $\tilde{X}_i$, $\tilde{u}_i$, $\tilde{g}(P_i)$, $\tilde{M}_i$ and $\tilde{Q}_i$. Note that for the subsample such that $d_i \, = \, 1$, we have $\tilde{X}_i = \dot{X}_i$. Hence,
\begin{align*}
\mathbb{E} \left(\tilde{M}_i \dot{y}_{i} \, | \, d_i \, = \, 1 , P_i \right) \, & = \, \mathbb{E}\left(\tilde{M}_i \tilde{X}_i \mu_i + \tilde{M}_i \tilde{g}(P_i) + \tilde{M}_i \tilde{u}_i \, | \, d_i \, = \, 1 , P_i \right) \\
& = \, \mathbb{E}\left(M_i \dot{X}_i \mu_i + M_i g(P_i) + M_i \dot{u}_i \, | \, d_i \, = \, 1 , P_i \right) \, = \, \mathcal{M}(P_i) g(P_i),
\end{align*}
where we define $\mathcal{M}(P_i) = \mathbb{E} \left(\tilde{M} \tilde{y}_i \, | \, (d_i \, = \, 1) , P_i \right)$. This first step equation identifies $g$ on the support of $P_i$ if $\mathcal{M}(P_i)$ is invertible a.s.
The second step equation will be given by
\begin{equation*}
\mathbb{E} \left( \tilde{Q}_i \tilde{y}_{i} - \, \tilde{Q}_i \tilde{g}(P_i) | d_i = 1 \right)\, = \, \mathbb{E} \left( \mu_i | d_i = 1 \right) .
\end{equation*}
\begin{assumption}\label{assn:selection_id}
$\mathbb{E} \left( \epsilon_{t} | d= 1, X, P \right) \, = \, f_t(P)$, and $\mathcal{M}(P)$ is invertible almost surely in $P$.
\end{assumption}
\begin{result}
Under Assumptions \ref{assn:finiteE} and \ref{assn:selection_id}, $\mathbb{E} \left( \mu | d = 1 \right)$ is identified.
\end{result}
The identified object is the average effect conditional on selection, $\mathbb{E}(\mu | d=1)$ which is in general different from $\mathbb{E}(\mu)$ unless $\mu \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} (\eta, X,Z)$. That additional restrictions are needed to identify $\mathbb{E}(\mu)$ is intuitive: when $d_{it} \neq 1$ the econometrician does not have any information on the unobserved heterogeneity.
The type of counterfactual that one can compute with this object would describe the effect of policies which impact the intensive margin, not the extensive margin, i.e, which do not affect whether individuals are selected. If $T > d_x + 2$, we recommend using the procedure described in Section \ref{sec:don_idea} and computing the average effect conditional on being selected in a subset of time periods. Averaging over all subsets identifies a conditional average effect under some additional conditions. This avoids using only the subsample of individuals for whom $d_{it}=1$ for all $t \leq T$, which can be quite small if $T$ is large (and still fixed).
Other cases studied in \cite{dnv03} can be handled here. For instance, the model allows for regressors $s_{it}$ to be suject to selection as well, that is, to be not observed for the population such that $d_{it} \, = \, 0$ (for example, the wage variable is not defined for unemployed individuals). As long as these regressors are not arguments of the function $C_{t}$ and the condition $\mathbb{E} \left( \epsilon_{t} | d= 1, X, S, P \right) \, = \, f_t(P)$ holds, the identification argument remains valid.
This sample selection model can also be adapted to the case where some of the regressors are endogenous. If these regressors are not multiplied by random coefficients, the argument of Section \ref{sec:mix_homog_heterog} can be applied. If they are accompanied by random coefficients, on the other hand, we suggest using the control function approach on the endogenous regressors, the control variables being for instance the residuals of the regression of the endogenous regressors on the exogenous regressors and instruments. The identification method developed above would then use a vector of control variables which include these residuals in addition to the propensity scores.
One last case worth mentioning, to which our two-step approach can be adapted, is when some regressors are endogenous and subject to selection.
We briefly explain how to construct the control variables. The model is
\begin{align}
y_{it} \, & = \, d_{it} \,y_{it}^*, \ \textrm{ with } \ y_{it}^* \, = \, x_{it}^{1 \ \prime} \, \mu_i^1 + x_{it}^{2 \ \prime} \, \mu_i^2 + \epsilon_{it}, \nonumber \\
x_{it}^{2} \, & = \, d_{it} \, x_{it}^{2 \, *}, \ \textrm{ with } \ x_{it}^{2 \, *} \, = \, \pi_t^2(x_{it}^1, z_{it}^1) + v_{it}, \label{eq:modelwSelection&Endog} \\
d_{it} \, & = \, \mathbbm{1} \left( \nu_{it} \leq p_t(x_{it}^1,z_{it}^1, z_{it}^2) \right) \, = \, \mathbbm{1} \left( \nu_{it} \leq p_{it} \right) , \nonumber
\end{align}
with $\nu_{it} \, \sim \mathcal{U}[0;1]$. Assume that
$(\epsilon_{is}, \nu_{is},v_{is})_{s \leq T} \raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} (x_{is}^1, z_{is}^1, z_{is}^2)_{s \leq T}$. In this model, $x^2$ is the endogenous regressor.
Identification of $p_{it}$ holds by $\mathbb{E} \left( d_{it} | \, x_{it}^1,z_{it}^1, z_{it}^2 \right) \, = \, p_{it}.$
Moreover
$$\mathbb{E} \left( v_{it} | d_{it}= 1, x_{it}^1, z_{it}^1, z_{it}^2 \right) \, = \, \mathbb{E} \left( \mathbb{E} \left( v_{it} | \nu_{it} \right) | (\nu_{it} \leq p_{it}), x_{it}^1, z_{it}^1, z_{it}^2 \right) \, : = \, \phi_t(p_{it}),$$ implying
$
\mathbb{E} \left( x_{it}^{2} | d_{it}=1 , x_{it}^1, z_{it}^1, z_{it}^2 \right) \, = \, \pi_t^2(x_{it}^1, z_{it}^1) + \phi_t(p_{it})
$.
This gives
\begin{equation}
x_{it}^{2} - \mathbb{E} \left( x_{it}^{2} | d_{it}=1 , x_{it}^1, z_{it}^1, z_{it}^2 \right) \, = \, v_{it} - \phi_t(p_{it}) \, := \, \bar{v}_{it},
\end{equation}
where $\bar{v}_{it}$ is identified. That is, the residuals for individuals selected in the sample are also control variables. The corresponding estimator will not need generated covariates.
We define again $\nu_i \, = \, (\nu_{is})_{s \leq T}$, and similarly $d_i$, $P_i$, $\bar{V}_i$, $X_i \, = \, (X_i^1, X_i^2)$ and $\phi(P_i) \, = \, (\phi_t(p_{it}))_{t \leq T}$. Note that the function $(V_i, P_i) \mapsto (\bar{V}_i,P_i)$ is one-to-one. Therefore,
\begin{align*}
\mathbb{E} \left( \epsilon_{it} | d_i=1 , X_i, \bar{V}_i, P_i \right) \, & = \, \mathbb{E} \left( \mathbb{E} \left( \epsilon_{it} | \nu_i, V_i, X_i^1, Z_i^1,Z_i^2 \right) | \, (\nu_i \leq P_i) , X_i, V_i, P_i \right) \\
& = \, \mathbb{E} \left( \mathbb{E} \left( \epsilon_{it} | \nu_i, V_i \right) | \, (\nu_i \leq P_i) , X_i, V_i, P_i \right) \, := \, h_t(V_i, P_i) , \\
& = \, h_t(\bar{V}_i + \phi(P_i), P_i) := f_t(\bar{V_i}, P_i).
\end{align*}
This conditional expectation is as in Assumption \ref{assn:idCFA}, where the control variables are $(\bar{V_i}, P_i)$.
This double use of the control function approach is already suggested in \cite{dnv03}. We presented here a slight modification such that the identification requires two steps instead of three. This for instance allows one to use the formula for the asymptotic variance given in Section \ref{sec:asymptoticnorm} (provided $\pi^2$ is not an object of interest) as it is known that increasing the number of steps typically changes the asymptotic variance matrix.
\fi
\section{Identification Results}\label{sec:identification}
\subsection{Main Result}\label{sec:maindid}
Under Assumptions \ref{assn:idCFA}, \ref{assn:id_finiteE} and \ref{assn:id_M_GLn}, we obtain
\begin{equation}\label{eq:closedform_g}
g(V_i) = \mathcal{M}(V_i)^{-1} \mathbb{E}(M_i \dot{y}_i | V_i), \ \mathbb{P}_{V_i} \ \text{a.s}.
\end{equation}
Under Assumption \ref{assn:id_finiteE}, the matrix $Q_i$ is well-defined with probability $1$.
Equation (\ref{eq:model_id_mu}) implies
\begin{align}
\mu_i \, & = \, Q_i \dot{y}_i \, - \, Q_i g(V_i) \, - \, \, Q_i \dot{u}_i, \label{eq:idMiou}, \\
\mathbb{E}( \mu_i ) \, & = \, \mathbb{E} (Q_i \dot{y}_i - \, Q_i g(V_i) ), \label{eq:Miou}
\end{align}
where the second line holds since by the law of iterated expectations and Assumption \ref{assn:id_finiteE}, $\mathbb{E}(Q_i \dot{u}_i)=0$. Equation (\ref{eq:Miou}) identifies $\mathbb{E}(\mu)$ since all elements on the right hand side are observed.
\begin{result}
Under Assumptions \ref{assn:idCFA}, \ref{assn:id_finiteE} and \ref{assn:id_M_GLn}, the average effect $\mathbb{E}(\mu_i)$ is identified and $\mathbb{E}( \mu_i ) = \mathbb{E} (Q_i \dot{y}_i - \, Q_i g(V_i) )$.
\end{result}
As mentioned in Section \ref{sec:assn_CRC_matrix}, the conditions $\mathbb{E}(||Q_i \dot{u}_{i}||) < \infty$ and $\mathbb{E}(||Q_i g(V_i)||) < \infty$ of Assumption \ref{assn:id_finiteE} may not hold. In this case, one can instead focus on $\mathbb{E}( \mu | \delta) := \mathbb{E}(\mu_i | \det(\dot{X}_i' \dot{X}_i) > \delta)$.
Define $\delta_i = \mathbbm{1}(\det(\dot{X}_i' \dot{X}_i) > \delta_0)$ and $Q_i^{\delta} = \delta_i Q_i $. Then
\begin{equation}\label{eq:Miou_delta}
\mathbb{E}( \mu | \delta) \, = \, \frac {\mathbb{E} (\delta_i Q_i \dot{y}_i - \, \delta_i Q_i g(V_i))}{\mathbb{P}(\det(\dot{X}_i' \dot{X}_i) > \delta_0)} \, = \, \frac {\mathbb{E} ( Q_i^\delta \dot{y}_i - \, Q_i^\delta g(V_i) )}{\mathbb{P}(\det(\dot{X}_i' \dot{X}_i) > \delta_0)},
\end{equation}
which identifies $\mathbb{E}( \mu | \delta) $ as all the terms on the right hand side of (\ref{eq:Miou_delta}) are identified. The considered expectations exist as long as $\mathbb{E}(||Q_i^\delta \dot{u}_{i}||) < \infty$ and $\mathbb{E}(||Q_i^\delta g(V_i)||) < \infty$, which holds under standard conditions: for instance if $X_i$ has bounded support, $\mathbb{E}(||\dot{u}_i||) < \infty$ and $\mathbb{E}(||g(V_i)||) < \infty$. Bounded support can be replaced with finite moments conditions.
It remains to identify $\mathbb{E}(\alpha_i)$, which we obtain using the variables in period $1$. We multiply (\ref{eq:idMiou}) with $x_{i1}$ and subtract $y_{i1}$, obtaining $y_{i1} - x_{i1}' \, \mu_i \, = \, y_{i1} \, - \, x_{i1}' \left[ Q_i \dot{y}_i \, - \, Q_i g(V_i) \, - \, Q_i \dot{u}_i \right].$
Since $y_{i1} - x_{i1}' \, \mu_i = \alpha_i + \epsilon_{it}$ where $\mathbb{E}(\epsilon_{it}) \, = \, 0$, we obtain the following identifying equation for $\mathbb{E}(\alpha_i)$,
\vspace{-5pt}
$$\mathbb{E}(\alpha_i) = \mathbb{E}(y_{i1} \, - \, x_{i1}' \left[ Q_i \dot{y}_i \, - \, Q_i g(V_i) \right] ).$$
\subsection{Average Effect of an Exogenous Intervention}\label{sec:exogIntervention}
Consider a policy intervention on $x_{it}$ for each unit $i$ in a given period $t$. The average effect of this exterior intervention is an object of interest and \cite{bp03} studies its identifiability in different models when the change in covariates is exogenous, i.e, independent of the unobservable error terms. The unobservables in the CRC model are $(\mu_i, \alpha_i, (\epsilon_{it})_{t \leq T})$, and if the exogenous shift is a variation $\Delta_{it}$ independent of $(\alpha_i, \mu_i, (\epsilon_{it})_{t \leq T})$, the average impact of the policy is $\mathbb{E}(\mu_i) \mathbb{E}(\Delta_{it})$ which is identified. However some policy interventions may impose a dependence between $\Delta_{it}$ is correlated with $x_{it}$, hence $\Delta_{it}$ will be correlated with $(\mu_i, \alpha_i)$ while being exogenous in the sense that it is independent of $(\epsilon_{it})_{t \leq T}$.
Consider an exogenous intervention that shifts $x_{it}$ to $l(x_{it})$, i.e, $\Delta_{it} = l(x_{it}) - x_{it}$. The average outcome after this intervention is $\mathbb{E}(l(x_{it})' \mu_i + \alpha_i + \epsilon_{it})$ and depends on the joint distribution of $(\mu_i,x_{it})$ where $\mu_i$ is unobservable and this joint distribution is unrestricted. Using Equation (\ref{eq:idMiou}) expressing $\mu_i$ as a function of the primitives,
the change in expected outcome is
\vspace{-7pt}
\begin{align*}
\mathbb{E}(l(x_{it})' \mu_i + \alpha_i + \epsilon_{it} - [x_{it}'\mu_i + \alpha_i + \epsilon_{it}]) &\, = \, \mathbb{E}([l(x_{it}) - x_{it}]'\mu_i ), \\
& \, = \, \mathbb{E} ( [l(x_{it}) - x_{it} ]' [Q_i y_i - \, Q_i g(V_i) - Q_i \dot{u}_i]),\\
& \, = \, \mathbb{E} ( [l(x_{it}) - x_{it} ]' [Q_i y_i - \, Q_i g(V_i)] ),
\end{align*}
where the second equality holds by exogeneity of the change in regressors. All elements in the last expectation are identified, thus so is the average change in outcome.
\subsection{Identifying Higher-Order Properties of $\mu_i$}\label{sec:mu_distrib}
The focus of this paper being on allowing for time-varying endogeneity in correlated random coefficient panel models, the parameter of interest is the simplest one, the average effect $\mathbb{E}(\mu_i)$. However more properties of the unobserved heterogeneity can be obtained. In particular, by $\mathbb{E}(\dot{u}_i |X_i)=0$ and (\ref{eq:idMiou}) we have $\mathbb{E}( \mu_i | X_i ) \, = \, \mathbb{E} (Q_i \dot{y}_i - \, Q_i g(V_i) |X_i)$.
In a model with strict exogeneity, \cite{ab12} extend the method in \cite{ch92} to identify the variance matrix and the distribution of $(\alpha_i,\mu_i)$ conditional on $X_i$ under various restrictions on the time-dependence of $\epsilon_{it}$ conditional on $X_i$ and on the joint distribution of $(\epsilon_i, \alpha_i, \mu_i)$ conditional on $X_i$. The argument first identifies the common parameters and subtracting the common part from the outcome variables, higher order moments of $\mu_i$ are separated from those of $\epsilon_{i}$ using the above-mentioned restrictions.
We note here that their argument can be combined with the assumptions made in the present paper so as to allow for endogeneity of the regressors. Indeed $g(V_i)$ being recovered using the method described in Section \ref{sec:maindid}, the analysis of \cite{ab12} can be conducted on $y_i - g(V_i) \, = \, X_i \mu_i + u_{i}$ which takes the same form as in their paper. We refer to the paper for more details on the procedure to recover these moments.
\section{Estimation}\label{sec:estim}
The proof of identification of $\mathbb{E}(\mu)$ is constructive as we obtained closed form expressions (\ref{eq:closedform_g}) and (\ref{eq:Miou}). The estimator we suggest follows the identification steps, replacing population moments with their sample analogs. Estimation is thus a multi-steps procedure.
First, for all $(i,t)$, $v_{it}$ is estimated as $\hat{v}_{it} = \hat{C}_t(X_i, Z_i)$ for $\hat{C}_t$ an estimator of the function $C_t$.
Second, the conditional expectation functions $\mathcal{M}(V) = \mathbb{E}( \, M_i | V_i = V)$ and $k(V) = \mathbb{E}(M_i \dot{y}_i | V_i = V)$ are estimated nonparametrically using the generated values $\hat{V}_i = (\hat{v}_{it})_{t \leq T}$ as regressors and the function $g = \mathcal{M}^{-1} k$ is estimated by plug in of the estimators $\hat{\mathcal{M}}(V)$ and $\hat{k}(V)$. Finally, the estimator for $\mathbb{E}(\mu | \delta)$ will be a sample analog of Equation (\ref{eq:Miou}), plugging in the estimator of $g$ and $V$.
The asymptotic properties of this estimator will depend on the definition of the control variables, that is, on $C_t$. We thus choose focus on the following model
\vspace{-5pt}
\begin{align}
& y_{it} = \, x_{it}^{ex \prime} \, \mu_i^{ex} + x_{it}^{en \, \prime} \, \mu_i^{en} + \alpha_i + \epsilon_{it}, \label{eq:model_for_estim}\\
& x_{it}^{en} = \, b_t(x_{it}^{ex}, z_{it}) + v_{it}, \qquad \mathbb{E}( v_{it} | x_{it}^{ex},z_{it}) = 0, \label{eq:modelCV_for_estim}
\end{align}
\vspace{-5pt}
where $x_{it}^{ex} \in \mathbb{R}^{d_1}$, $x_{it}^{en} \in \mathbb{R}^{d_2}$, $z_{it} \in \mathbb{R}^{d_z}$, and where Assumption \ref{assn:idCFA} holds. The regressors $x^{ex}$ are exogenous while $x^{en}$ can be endogenous. The control variables in this model are the residuals of the nonparametric regression of the endogenous regressors on the exogenous regressors and the instruments and will be estimated as residuals of the nonparametric regression estimation of $x_{it}^{en}$. All estimators of the nonparametric regressions will be series estimators. As highlighted in Section \ref{sec:assn_CRC_matrix}, the condition $\mathbb{E}(||Q_i||^2) < \infty$ may not hold thus out of caution we use (\ref{eq:Miou_delta}) to estimate $\mathbb{E}(\mu | \delta)$, where we fix $\delta_0$ and define $\delta_i = \mathbbm{1}(\det(\dot{X}_i' \dot{X}_i) > \delta_0)$. We proceed with an explicit definition of the estimators and a stepwise proof of asymptotic normality of $\hat{\mu}$. All proofs are in the Appendix, Section \ref{sec:appConsistency}.
\subsection{Definition of the estimators}\label{sec:def_estim}
We introduce some notations. For a vector $a \in \mathbb{R}^p$, $||a||$ is its Euclidean norm. We also denote by $||.||_F$ the Frobenius norm (the canonical norm) in the space of matrices $\mathcal{M}_p(\mathbb{R})$, and $||.||_2$ the matrix norm induced by $||.||$ on $\mathbb{R}^p$ (the spectral norm). By an abuse of notation, in Section \ref{sec:estim} we drop the index $i$ for random variables whenever it does not hinder clarity.
For $g$ a vector of functions of $x \in \mathcal{S}_x \subset R^{k}$, $||g||_{\infty}$ is $ \sup_{\mathcal{S}_x} ||g(.)||$. For $l = (l_1,\, .. \, , l_{k}) \in \mathbb{N}^{k}$, we define $|l| = \sum_{j = 1}^{k} l_j $, and the partial derivative $\partial^{l} g(x) = \partial^{|l|} g(x) / \partial^{l_1}x_1 ... \partial^{l_{k}}$. We will use the norm $|g|_d = \max_{|l| \leq d} \sup_{x \in \mathcal{S}_x} || \partial^{l} g(x) ||$ when $g$ is $d$ times differentiable.
We denote by $\partial g(x)$ the Jacobian matrix $( \partial g(x) / \partial x_1, ... , \partial g(x) / \partial x_k )$.
For a sequence $(c_n)_{n \in \mathbb{N}} \in \mathbb{R}^{\mathbb{N}}$, the notation $c_n \to 0$ should be understood as $c_n \to_{n \to \infty} 0$. For a random variable $x$, $f_x$ denotes its density.
By (\ref{eq:modelCV_for_estim}), we have $v_{it} = x_{it}^{en} - \mathbb{E}(x_{it}^{en} | x_{it}^{ex}, z_{it})$.
We write $\xi_{it} = (x_{it}^{ex},z_{it})$. Consider the $L \times 1$ vector of approximating functions $r^L(\xi_t)=(r_{1L}(\xi_t),.. \, ,r_{LL}(\xi_t))'$ and $r_{it} = r^L(\xi_{it})$. We define the series estimators of the regression function $\mathbb{E}(x_{it}^{en} | \xi_{it}=\xi_t) = b_t(\xi_t)$ to be $ \hat{\beta}_t' r^L(\xi_t)$ where $\hat{\beta}_t$ is $L \times d_2$, and
\begin{equation}\label{eq:estimV}
\hat{\beta}_t \, = \, (R_t R_t')^{-1} \sum_{i} r^L(\xi_{it}) x_{it}^{en \, \prime} \, = \, (R_t R_t')^{-1} R_t X_t^{en \, \prime},
\end{equation}
where $R_t \, = \, (r_{1t},.. \, ,r_{nt})$ is $L \times n$ and $X_t^{en} \, = \, (x_{1t}^{en},.. \, ,x_{nt}^{en})$ is $d_2 \times n$.
Write $r_{it} = r^L(\xi_{it})$ and $\hat{b}_{it} = \hat{\beta}_t' \, r_{it}$, define $b_{it} = b_t(\xi_{it})$, $\hat{b}_t = \hat{\beta}_t' r^L$ and the residuals $\tilde{v}_{it} \, = \, x_{it}^{en} - \hat{b}_{it} $ with $\tilde{V}_i = (\tilde{v}_{i1}',\, .. \, ,\tilde{v}_{iT}' )'$.
Later, the support $\mathcal{S}_V$ of $V$ will be assumed bounded. However, the values obtained using the estimated residuals might not be in $\mathcal{S}_V$: it will be convenient for the asymptotic analysis to introduce a transformation $\tau$ of the generated variables such that their transformed values lie in $\mathcal{S}_V$. Specifically, we assume that the support of $v_t$ is of the form $\bigtimes_{d=1}^{d_2} \, [\underline{v}_{td},\bar{v}_{td}]$ and that the support of $V$ is $\mathcal{S}_V = \bigtimes_{d \leq d_2, \, t \leq T} \, [\underline{v}_{td},\bar{v}_{td}]$.
We define $\tau$ such that for $V = (v_1',\, .. \, ,v_T')' \in \mathbb{R}^{Td_2}$, then $\tau(V)$ is the projection of $V$ onto $\mathcal{S}_V$.
The exact definition of $\tau$ is given in the Appendix \ref{ref:trimmingapp} by (\ref{eq:def_tau_consist}). Our estimator for $V_i$ will then be $\hat{V}_i = \tau(\tilde{V}_i)$. Note that for all draws of $V_i$, $\tau(V_i)=V_i$ and $|| \hat{V}_i - V_i || \leq ||\tilde{V}_i - V_i||$. We also assume that $\mathcal{S}_{\xi t}$ is of the form $\bigtimes_{d=1}^{d_\xi} \, [\underline{\xi}_{td};\bar{\xi}_{td}]$, where we use the notation $\xi_{td}$ for the d$^{th}$ component of $\xi_t$.
Let $p^K(V)=(p_{1K}(V),.. \, ,p_{KK}(V))'$ denote a $K \times 1$ vector of approximating functions, $p_i = p^K(V_i)$ and $\hat{p}_i = p^K(\hat{V}_i)$. An estimator of $h^W(V) = \mathbb{E}(w_i | V_i = V)$ for a generic scalar random variable $w_i$ using the generated $\hat{V}$ is $p^K(V)' \, \hat{\pi}^W$ where $\hat{\pi}^W$ is a vector of size $K$ given by
\begin{equation}\label{eq:def_2stepestim}
\hat{\pi}^W \, = \, (\hat{P} \hat{P}')^{-1} \sum_{i} p^K(\hat{V}_i) \, W_i' \, = \, (\hat{P} \hat{P}')^{-1} \hat{P} W,
\end{equation}
where $\hat{P} \, = \, ( \hat{p}_1,.. \, , \hat{p}_n)$ is $K \times n$ and $W \, = \, (w_1,.. \, ,w_n)'$ is a vector of size $ n$.
Using this general definition, we construct component by component estimators $\hat{\mathcal{M}}$ and $\hat{k}$ for the matrix and vector valued functions $\mathcal{M}$ and $k$. We obtain $p^K(V)' \, \hat{\pi}^{M,st}$ an estimator of the $(s,t)$ component of the matrix $\mathcal{M}$, taking $w_i$ to be $(M_{i})_{s,t}$. Similarly, an estimator of the $s$th component of $k$ will be $p^K(V)' \, \hat{\pi}^{k,s}$, taking $w_i = (M_i \dot{y}_i)_s$.
By (\ref{eq:closedform_g}) and (\ref{eq:Miou_delta}),
a plug-in estimator of $g$ is
$\hat{g}(V) \, = \, \hat{\mathcal{M}}(V)^{-1} \, \hat{k}(V)$
and a plug-in estimator of $\mathbb{E}(\mu|\delta)$ is
$$\hat{\mu} \, = \, \frac{\sum_{i=1}^n \delta_i Q_i \, [ \dot{y}_i - \, \hat{g}(\hat{V_i}) ]}{\sum_{i=1}^n \delta_i} \, = \, \frac{\sum_{i=1}^n Q_i^{\delta} \, [ \dot{y}_i - \, \hat{g}(\hat{V_i}) ]}{\sum_{i=1}^n \delta_i}.$$
The multi-step estimation procedure only uses closed form expressions: its ease of implementation comes with a layered asymptotic analysis as each step needs to be analyzed one by one to eventually obtain the asymptotic behavior of $\hat{\mu}$. This type of asymptotic analysis is the subject of a wide literature on nonparametric and semiparametric estimation with generated covariates. Before laying out the main results of our asymptotic analysis, we give here a brief overview of this literature.
Papers studying asymptotic normality of semiparametric estimators, such as \cite{n94}, \cite{clvk03}, \cite{ac03} and \cite{hl10} among many other references, have a level of generality which encompasses the case where the regressors are themselves estimated. However, the conditions given in these papers are ``high-level'' conditions and are not easily applied to the composition of nonparametrically estimated infinite-dimensional nuisance parameters.
Examples of asymptotic derivations in specific models with generated regressors are papers already cited such as \cite{npv99}, \cite{dnv03}, and \cite{in09}, as the use of a nonparametric control function approach naturally suggests an estimator with generated covariates. Others are, e.g, \cite{ap93}, \cite{bp04}, \cite{n09} and \cite{ejl16}.
Moreover, recent contributions have focused on obtaining general asymptotic results for such semiparametric estimators.
Among important recent contributions, \cite{hr13,hr19} derive in the spirit of \cite{n94} a general formula of the asymptotic variance of semiparametric estimators using generated regressors.
However they do not provide results on how to obtain asymptotic normality for particular classes of estimators. For estimators with generated regressors depending on a nonparametrically estimated function, this type of analysis can be found for instance in \cite{ejl14}, \cite{mrs16} and \cite{hlr18}.
\cite{ejl14} obtain a uniform expansion of a weighted sample average of residuals obtained from kernel-estimated nonparametric regressions with generated covariates, which can then be used to prove asymptotic normality of a class of semiparametric estimators.
\cite{mrs16} study the asymptotic normality of a general class of semiparametric GMM estimators depending on a nonparametric nuisance parameter, also constructed with generated covariates. Our estimator of the APE $\hat{\mu}$ belongs to this class of estimators, although of a simpler form since it has a closed-form expression. Moreover we use series to construct the nonparametric estimates while the infinite dimensional nuisance parameter in \cite{mrs16} is a conditional expectation estimated with local polynomial estimator and they do not specify an estimator for the generated covariates.
Estimators in \cite{hlr18} have a structure closer to that of $\hat{\mu}$: they study nonparametric two-step sieve M estimators, but focus on known functionals. They show asymptotic normality of their estimator when standardized by a finite sample variance and give a practical estimator of this variance. They do not however provide an explicit formula of the asymptotic variance.
The estimator we analyze in this section is instead an estimated functional of the two-step nonparametric estimators. Using a different type of proof techniques with lower level conditions on the primitives of a more specific class of models, we show asymptotic normality and obtain the asymptotic variance of a generic class of estimators to which ours belongs.
See, e.g, \cite{mrs16} for a literature review on semiparametric estimation with generated covariates and explanation on the specificity of this type of estimation.
\subsection{Consistency}\label{sec:consistency}
Consistency of $\hat{\mu}$ is obtained under the following assumptions.
\begin{assumption}\label{assn:NPV1stStep0Dsty}
There exists $\alpha_1 > 0$ and $\gamma_1 > 0$ such that $\sqrt{L/n} \ L^{\alpha_1 + 1} \xrightarrow[n \to \infty]{} 0$ and $\forall \, t \leq T$,
\begin{enumerate}[noitemsep,nolistsep]
\item $(x_{it}^{en},\xi_{it})$ is i.i.d over $i$, continuously distributed and $\operatorname*{Var}(x_t^{en}|\xi_t)$ is bounded,
\item $r^L(.)$ is the power series basis, and $\forall \, \xi_t \in \mathcal{S}_{\xi t}, \ f_{\xi t}(\xi_t) \, \geq \, \Pi_{d=1}^{d_\xi} (\xi_{dt} - \underline{\xi}_{dt})^{\alpha_1}(\bar{\xi}_{dt} - \xi_{dt})^{\alpha_1}$,
\item \label{assn:NPV1stStep0Dsty_approx} There exists $\beta_t^L$ such that $\sup_{\mathcal{S}_{\xi t}} || b_t(\xi_t) - \beta_t^{L \prime} \, r^L(\xi_t)|| \leq C L^{- \gamma_1}$.
\end{enumerate}
\end{assumption}
Assumption \ref{assn:NPV1stStep0Dsty} allows us to obtain a convergence rate for the sample mean-squared error of the generated covariates $\hat{V}_i$ following results in \cite{npv99}. Note that we use the power series basis so as to allow for the density of the regressors to be $0$ on the boundary of their support. A version of these results with arbitrary sieve basis but for a density bounded away from $0$ can be found in Appendix \ref{sec:appConsistency}.
If for instance $b_t$ is continuously differentiable up to order $p$, writing $d_{\xi} = d_1 + d_z$, then Assumption \ref{assn:NPV1stStep0Dsty} (\ref{assn:NPV1stStep0Dsty_approx}) holds with $\gamma_1 \, = \, p / d_{\xi}$ for different choices of sieve basis.
\begin{assumption}\label{assn:IN2stStep0Dsty_model}
\
\begin{enumerate}[noitemsep,nolistsep]
\item $\mathbb{E}(\dot{u} | X, Z)$ and $\operatorname*{Var}(|| \dot{u}|| \, | X, Z)$ are bounded on $\mathcal{S}_{X,Z}$,
\item \label{assn:IN2stStep0Dsty_model_lipschitz} $\mathcal{M}$ and $k$ are Lipschitz and $g$ is bounded on $\mathcal{S}_{V}$,
\item \label{assn:IN2stStep0Dsty_model_density} $p^K(.)$ is the power series basis, and $\forall V \in \mathcal{S}_{V}, \ f_{V}(V) \, \geq \, \Pi_{d \leq k_2, \, t \leq T} \, (v_{t,d} - \underline{v}_{td})^{\alpha_2}(\bar{v}_{td} - v_{t,d})^{\alpha_2}$,
\item \label{assn:IN2stStep0Dsty_model_approx} There exists $\gamma_2$ and $(\pi_{M,st}^{K})_{s,t}$ and $(\pi_{k,t}^{K})_{s,t}$ such that $\sup_{\mathcal{S}_{V}} | \mathcal{M}_{st}(V) - \, \pi_{M,st}^{K \prime} p^K(V) | \leq C K^{- \gamma_2}$ and $\sup_{\mathcal{S}_{V}} | k_t(V) - \, \pi_{k,t}^{K \prime} p^K(V) | \leq C K^{- \gamma_2}$ for all $s,t \leq T-1$,
\item $K^{\alpha_2 + 7/2} \Delta_n \to_{n \to \infty} 0$ and $ K^{\alpha_2 + 3/2} / \sqrt{n} \, \to_{n \to \infty} 0$.
\end{enumerate}
\end{assumption}
The second step of the proof of consistency is the derivation of a convergence rate for $\hat{\mathcal{M}}$ and $\hat{k}$ in sup norm and mean square norms, under Assumptions \ref{assn:IN2stStep0Dsty_model}. Convergence rates of nonparametric estimators of conditional expectations $\mathbb{E}(w_i| V_i=V) = h^{W}(V)$ using generated regressors are also derived in \cite{npv99}. However, they impose an orthogonality condition which by definition does not hold for our specific choices of $w_i$ and this rules out a direct application of their asymptotic results.
Specifically, writing $ e^W = w - h^W(V)$,
an additional assumption required to apply directly \cite{npv99} is $\mathbb{E}(e^W |X^{ex},V,Z)= 0$, that is, $e^W$ must be conditionally mean-independent of all variables involved in the first step, i.e, in the construction of the control variables.
This condition does not hold when $w_i$ is either a component of the matrix $ M= I - \dot{X} (\dot{X}' \dot{X})^{-1} \dot{X}' $ or of the vector $M \dot{y}$, because
$\mathbb{E}(M|X^{ex},X^{en}, Z)= M \neq \mathbb{E}(M|V)=\mathcal{M}(V)$ and $\mathbb{E}(M \dot{y}|X^{ex},X^{en}, Z) =M g(V) + M \mathbb{E}(\dot{u}|X^{ex},X^{en}, Z) \neq \mathbb{E}(M \dot{y}|V) = k(V) = \mathcal{M}(V) \, g(V)$.
To account for the difference $ \mathbb{E}(e^W|X, V) \neq \mathbb{E}(e^W|V)$, we write
\vspace{-5pt}
\begin{equation}\label{eq:reg_GalCase}
w = h^W(V) + \rho^W(X,Z) + e^{W*}, \ \mathbb{E}(e^{W*} | X,Z) = 0,
\end{equation}
\vspace{-5pt}
where $\rho^W(X,Z) = \mathbb{E}(e^W |X,Z) = \mathbb{E}(w | X,Z) - \mathbb{E}(w|V)$. For our choices of $w$, $\rho \neq 0$ and this will add an extra term to the convergence rate of the two-step estimator as well as, as will be clear in a later part of the paper, the asymptotic variance of a linear functional of this estimator.
This difference has been documented for instance in \cite{hr13} and \cite{mrs16}.
\begin{assumption}\label{assn:cv_g}
$\mathcal{M}$ and $g$ are continuous on $ \mathcal{S}_V$, $\mathcal{S}_V$ is a compact set and $\mathcal{M}(V)$ is invertible for all values $V \in \mathcal{S}_V$.
\end{assumption}
Recall that $\hat{g}(V) \, = \, \hat{\mathcal{M}}(V)^{-1} \, \hat{k}(V).$
We use Assumption \ref{assn:cv_g} to obtain the rate of convergence of $\hat{g}(.)$ with continuity arguments.
Note that while Assumption \ref{assn:id_M_GLn} requires the matrix to be invertible only $\mathbb{P}_V \, \text{a.s}$, under Assumption \ref{assn:cv_g} $\mathcal{M}(V)$ is invertible for all values of $V$ on the support $\mathcal{S}_V$.
\begin{assumption}\label{assn:finiteE}
Assume $\mathbb{E}(||Q^{\delta}||) < \infty$ and $\mathbb{E}(||Q^{\delta} \, \dot{y}||) < \infty$.
\end{assumption}
Under the above assumptions, and defining $\gamma_n = b_1(K) (K/n + K^{-2 \gamma_2} + \Delta_n^2 b_2(K)^2)^{1/2}$, we obtain the following consistency result.
\begin{result}\label{rslt:consistency}
Suppose Assumptions \ref{assn:NPV1stStep0Dsty}, \ref{assn:IN2stStep0Dsty_model}, \ref{assn:cv_g}, and \ref{assn:finiteE} hold. Assume also $\gamma_n \to 0$, $a_1(L) \Delta_n \to 0$ and that $g$ is continuously differentiable on $\mathcal{S}_V$. Then $ \hat{\mu}\, \to_{\mathbb{P}} \, \mathbb{E} (\mu | \delta).$
\end{result}
\subsection{Asymptotic normality}\label{sec:asymptoticnorm}
\subsubsection{Definitions}\label{sec:trimming}
Recall that $\hat{V}_i = \tau(\tilde{V}_i)$, where $\tilde{V}_i =(\tilde{v}_{it})_{t \leq T}$ is the vector of residuals from the sieve regression of $x_{it}^{en}$ on $\xi_{it} = (x_{it}^{ex}, z_{it})$ and where $\tau$ projects onto $\mathcal{S}_V = \bigtimes_{d \leq d_2, \, t \leq T} \, [\underline{v}_{td};\bar{v}_{td}]$.
The proof of asymptotic normality will require $\tau$ to be twice differentiable. Thus we change the definition of $\tau$ from (\ref{eq:def_tau_consist}) to (\ref{eq:def_tau_asnorm}), defined in Appendix \ref{ref:trimmingapp}.
The new trimming function is twice differentiable and projects onto a bounded superset of $\mathcal{S}_V$. For $\varsigma>0$, this superset is $\mathcal{S}_V^\varsigma := \bigtimes_{d \leq d_2, \, t \leq T} \, [\underline{v}_{td} - \varsigma ;\bar{v}_{td} + \varsigma]$. We will refer to $ \mathcal{S}_V^\varsigma $ as the ``extended support''.
On this extended support, we use extensions of the various regression functions.
To be precise, let $p \in \mathbb{N}$ and a twice-continuously differentiable function $m : \mathcal{S}_V \to \mathbb{R}^p$.
We define an extension of $m$ as $m^\varsigma : \mathcal{S}_V^\varsigma \to \mathbb{R}^p$ such that for all $V$ in $\mathcal{S}_V$, $m^\varsigma(V) = m(V)$, and $m^\varsigma$ is twice continuously differentiable on the extended support $ \mathcal{S}_V^\varsigma$.
We previously used, for $g$ a function of the variable $V$, the norm $|g|_d = \max_{|l| \leq d} \sup_{V \in \mathcal{S}_V} || \partial^{l} g(.) ||$. A corresponding norm for the extended functions will change the supremum to a supremum over the extended support, i.e., $|g|_d^\varsigma = \max_{|l| \leq d} \sup_{V \in \mathcal{S}_V^\varsigma} || \partial^{l} g(.) ||$.
We now use bases of functions $p^K(.)$ defined on the extended support.
We point out here that unlike in Section \ref{sec:consistency}, we will not allow for the density of the regressors to go to zero on the boundaries of their support. We are thus silent on the choice of the basis.
\subsubsection{Linearization}\label{sec:lineariz}
To study asymptotic normality of $\sqrt{n} (\hat{\mu} - \mathbb{E}(\mu))$, we define
$
\hat{\mu}^\delta = \sum_{i=1}^n Q_i^{\delta} \, [ \dot{y}_i - \, \hat{g}(\hat{V_i}) ] /n,
$
then $\hat{\mu} = \hat{\mu}^\delta / (\sum_{i=1}^n \delta_i/n)$. We thus focus first on $\hat{\mu}^\delta - \mathbb{E}(\mu \delta)$. Define $W_i = (X_i, V_i, Z_i, u_i, \mu_i, \alpha_i)$ the vector of primitive variables where we now write the variables as column vectors, e.g $V_i \, = \, (v_{i1}',... \, , v_{iT}')'$. Define also $\mathcal{G} \, = \, ((b_t)_{t \leq T}, k, \mathcal{M})$ a vector of generic functions with $b_t : \mathcal{S}_{\xi t} \mapsto \mathbb{R}^{d_2}$, $k : \mathcal{S}_{V}^\varsigma \mapsto \mathbb{R}^{T-1}$ and $\mathcal{M} : \mathcal{S}_{V}^\varsigma \mapsto \ \mathcal{M}_{T-1}(\mathbb{R})$. We now write $ \mathcal{G}_0 \, = \, ((b_{0t})_{t \leq T}, k_0, \mathcal{M}_0), $ for the true values of these functions, that is, for the nonparametric primitives of the model. Note that the functions we consider here are functions on the extended support. We dropped the exponent $\varsigma$ and will display it to avoid confusion whenever necessary.
We decompose
\begin{align}
\sqrt{n} & (\hat{\mu}^\delta - \mathbb{E}(\mu_i \delta_i)) \, = \ \frac{1}{\sqrt{n}} \sum_{i=1}^n Q_i^\delta \, [ \dot{y}_i - \ \hat{g}(\hat{V_i}) ] - \mathbb{E}(\mu_i \delta_i), \nonumber \\
= \, & \frac{1}{\sqrt{n}} \left[ \sum_{i=1}^n [ \delta_i \mu_i - \mathbb{E}(\mu_i \delta_i)] + \, \sum_{i=1}^n Q_i^\delta \dot{u}_i + \, \frac{1}{\sqrt{n}} \sum_{i=1}^n [ Q_i^\delta g_0(V_i) - \mathbb{E}(Q^\delta g_0(V))] - \sum_{i=1}^n [ Q_i^\delta \hat{g}(\hat{V_i}) - \mathbb{E}(Q^\delta g_0(V))] \right] , \nonumber \\
= \, & \frac{1}{\sqrt{n}} \sum_{i=1}^n [\delta_i \mu_i - \mathbb{E}(\mu_i \delta_i)] + \, \frac{1}{\sqrt{n}} \sum_{i=1}^n Q_i^\delta \dot{u}_i \, - \, \sqrt{n} \big[ \mathcal{X}_n(\hat{\mathcal{G}}) - \mathcal{X}_n(\mathcal{G}_0) \big], \label{eq:dvlptasNorm}
\end{align}
where we define
\begin{align*}
\chi(W_i, \mathcal{G}) & = Q_i^\delta \, \mathcal{M}\left( \tau \left[ (x_{it}^{en} - b_t(\xi_{it}))_{t \leq T} \right] \right)^{-1} \, k\left( \tau \left[ (x_{it}^{en} - b_t(\xi_{it}))_{t \leq T} \right] \right),\\
\mathcal{X}_n(\mathcal{G}) & = \frac{1}{n} \sum_{i=1}^n [\chi(W_i, \mathcal{G}) - \mathbb{E}(\chi(W_i, \mathcal{G}_0))] \text{ and } \mathcal{X}(\mathcal{G}) = \mathbb{E}(\chi(W_i, \mathcal{G})) - \mathbb{E}(\chi(W_i, \mathcal{G}_0)),
\end{align*}
where $\tau$
ensures that the argument of $\mathcal{M}$ and $k$ lies in $\mathcal{S}_V^\varsigma$.
Note that $\mathcal{X}(\mathcal{G}_0) = 0$.
The use of generated covariates
in place of the true value of the variables has a twofold impact on semiparametric estimators such as $\hat{\mu}$. First the nuisance parameters $\mathcal{M}$ and $k$ are estimated using the generated values. Second the estimators $\hat{\mathcal{M}}$ and $\hat{k}$ are evaluated at the generated values when plugged in in the sample average that defines $\hat{\mu}$. The dependence of $\mathcal{X}$ on $(b_t)_{t \leq T}$ highlights the latter aspect.
The two first terms in Equation (\ref{eq:dvlptasNorm}) are normalized sums of i.i.d random variables. Their asymptotic normality can be established by a standard CLT argument.
We focus on the last term, $\sqrt{n} \big[ \mathcal{X}_n(\hat{\mathcal{G}}) - \mathcal{X}_n(\mathcal{G}_0) \big]$, which includes a composition of estimated infinite dimensional nuisance parameters. Typically, under restrictions detailed in Appendix \ref{sec:lineq_app} its asymptotic distribution will be that of
$ \sqrt{n} \, \mathcal{X}_0^{(G)} [\hat{\mathcal{G}} - \mathcal{G}_0]$ where $\mathcal{X}_0^{(G)}$ is the pathwise derivative of $\mathcal{X}$ at $\mathcal{G}_0$ and it is evaluated at $\hat{\mathcal{G}} - \mathcal{G}_0$. It is a linear functional of a vector of nonparametric estimators.
The structure of our asymptotic analysis is thus as follows. First we derive the asymptotic distribution of $ \sqrt{n} \, \mathcal{X}_0^{(G)} [\hat{\mathcal{G}} - \mathcal{G}_0]$ and obtain its influence function, and second we show that $\sqrt{n} \big[ \mathcal{X}_n(\hat{\mathcal{G}}) - \mathcal{X}_n(\mathcal{G}_0) \big]$ is asymptotically equivalent to its linearization. Finally these results are used together with (\ref{eq:dvlptasNorm}) to obtain the influence function of $\hat{\mu}^\delta $ and thus asymptotic normality of $\hat{\mu}$.
Focusing on the linearized term, the pathwise derivative applied to the estimators can be decomposed as the sum of $T + 2$ partial pathwise derivatives applied to each nonparametrically estimated function. We denote with $v_t$ the t$^{th}$ component of $V$ and with $\frac{\partial g}{\partial v_t}(V_i)$ the Jacobian matrix of size $(T-1) \times d_2$.
We define $\mathcal{X}_0^{(k)}[\tilde{k}]$ the partial pathwise derivative of $\mathcal{X}$ with respect to $k$ at $\mathcal{G}_0$ and evaluated at $\tilde{k}$, and similarly $\mathcal{X}_0^{(M)}[\tilde{\mathcal{M}}]$ and $(\mathcal{X}_0^{(bt)}[\tilde{b}_t])_{t \leq T}$. Assuming one can interchange expectation and differentiation, we follow \cite{mrs16} and write these derivatives as
\begin{align}
\mathcal{X}_0^{(G)} [\mathcal{G} - \mathcal{G}_0] \, = & \, \mathcal{X}_0^{(k)} [k - k_0] + \mathcal{X}_0^{(M)}[\mathcal{M} - \mathcal{M}_0 ] + \sum_{t = 1}^T \mathcal{X}_0^{(bt)}[b_t - b_{0,t}], \nonumber \\
\, = &\, \int_V \left[ \lambda_{M}(v) \right.\, \left. \operatorname*{Vec} \left((\mathcal{M} - \mathcal{M}_0)(v)\right) \right] d F_{V}(v) + \int_V \left[ \lambda_{k}(v) \, (k - k_0)(v) \right] d F_{V}(v) \nonumber \\
& \ + \ \ \sum_{t = 1}^T \int_{\xi_t} \lambda_{bt}(\xi_t) \, (b_t - b_{0,t})(\xi_t) d F_{\xi_t}(\xi_t), \label{eq:linearX_int}
\end{align}
where the functions $\lambda_{.}$ are obtained by differentiation (see details in Appendix \ref{sec:lineq_app}),
\begin{align*}
\lambda_{M}(v) \, & = \, - \mathbb{E} \left( g(V_i)' \otimes (Q_i^\delta \mathcal{M}_0(V_i)^{-1}) \, | \, V_i = v \right) \, = \, - g(v)' \otimes [\mathbb{E}(Q_i^\delta | V_i = v) \mathcal{M}_0(v)^{-1}] ,\\
\lambda_{k}(v) \, & = \, \mathbb{E}( Q_i^\delta \mathcal{M}_0(V_i)^{-1} \, | \, V_i = v ) \, = \, \mathbb{E}(Q_i^\delta \, | \, V_i = v ) \mathcal{M}_0(v)^{-1}, \\
\lambda_{bt}(\xi_t) \, & = \, - \mathbb{E}\left(Q_i^\delta \frac{ \partial g(V_i)}{\partial v_t} \, | \, \xi_{it} = \xi_t \right) . \vspace{-2pt}
\end{align*}
\subsubsection{Asymptotics of the linear part}\label{sec:asympt_lin_model}
The term $\sqrt{n} \mathcal{X}_0^{(G)} [\mathcal{G} - \mathcal{G}_0]$ is a sum of linear functionals applied to the components of $\mathcal{G} -\mathcal{G}_0$, where $\mathcal{G} = ((b_{t})_{t \leq T}, k, \mathcal{M})$. To obtain its asymptotic distribution, we need to define the following random variables and functions,
\begin{align*}
& e_i^M = M_i - \mathbb{E}(M_i|V_i) = M_i - \mathcal{M}_0(V_i),\ e_i^{M*} = M_i - \mathbb{E}(M_i|X_i,Z_i) = 0, \\
& \rho^M(X_i, Z_i) = \mathbb{E}(M_i|X_i, Z_i) - \mathbb{E}(M_i|V_i) = M_i - \mathcal{M}_0(V_i),\\
& e_i^{k} = M_i \dot{y}_i - \mathbb{E}(M_i \dot{y}_i|V_i) = [M_i - \mathcal{M}_0(V_i)] g(V_i) + M_i \dot{u}_i,\\
& e_i^{k*} = M_i \dot{y}_i - \mathbb{E}(M_i \dot{y}_i|X_i, Z_i) = M_i[\dot{u}_i - \mathbb{E}(\dot{u}_i |X_i, Z_i)], \\
& \rho^{k}(X_i, Z_i) = \mathbb{E}(M_i \dot{y}_i|X_i, Z_i) - \mathbb{E}(M_i \dot{y}_i|V_i) = [M_i - \mathcal{M}_0(V_i)] g(V_i) + M_i \mathbb{E}(\dot{u}_i |X_i, Z_i).
\end{align*}
For a given $v$, $\lambda_M^j(v)$ is defined as the $j^{th}$ column of the matrix $\lambda_M(v)$ and
\begin{align*}
\Lambda^M = \int_v \left[ \, \lambda_M^1(v) p_1^K(v) , \ ... \ , \lambda_M^{1}(v) p_K^K(v) , \ ... \ , \lambda_M^{(T-1)^2}(v) p_1^K(v), \ ... \ , \lambda_M^{(T-1)^2}(v) p_K^K(v) \right]dF_V(v) ,
\end{align*}
of dimension $d_x \times K(T-1)^2$. Define similarly the matrix $\Lambda^{k}$. We will interchangeably index the columns of $\lambda_M$ as $\lambda_M^d$ with $d \leq (T-1)^2$ and as $\lambda_M^{st}$ with $1 \leq s,t \leq T-1$.
We also define
\begin{align*}
H_t^M \, & = \, \mathbb{E} \left[ \frac{\partial \operatorname*{Vec}( \mathcal{M}_0(V_i))}{\partial v_t} \otimes p_i \otimes r_{it}' \right], \ H_t^{k} \, = \, \mathbb{E} \left[\frac{\partial k_0(V_i)}{\partial v_t} \otimes p_i \otimes r_{it}' \right], \\
\text{d}P_t^M & = \mathbb{E} \left[ \operatorname*{Vec}(\rho_i^M) \otimes \frac{\partial p^K (V_i)}{\partial v_t} \otimes r_{it}' \right], \ \text{d}P_t^{k} = \mathbb{E} \left[ \rho_i^{k} \otimes \frac{\partial p^K (V_i)}{\partial v_t} \otimes r_{it}' \right].
\end{align*}
For a given $\xi_t$, $\lambda_{bt}^j(\xi_t)$ is defined as the $j^{th}$ column of the matrix $\lambda_{bt}(\xi_t)$ and for each $t$,
\begin{align*}
\Lambda^{bt} \, = \,
\int_{\xi_t} \left[ \, \lambda_{bt}^1(\xi_t) r_1^L(\xi_t) , \ ... \ , \lambda_{bt}^{1}(\xi_t) r_L^L(\xi_t) , \ ... \ , \lambda_{bt}^{T-1}(\xi_t) r_1^L(\xi_t) , \ ... \ , \lambda_{bt}^{T-1}(\xi_t) r_L^L(\xi_t) \right]dF_{\xi_t}(\xi_t).
\end{align*}
\begin{assumption}\label{assn:LAE_model}
\
\begin{enumerate}[noitemsep,nolistsep]
\item \label{assn:LAE_model_fctn} $\mathcal{M}_0, g_0$ are twice continuously differentiable with bounded first and second order derivatives,
\item \label{assn:LAE_model_mmt} $\mathbb{E}(||Q_i^\delta||) < \infty$, $\mathbb{E}(||\dot{u}|| \, | X, Z)$ is bounded on $\mathcal{S}_{X,Z}$, and $\mathcal{S}_{\xi t}$ for all $t \leq T$ and $\mathcal{S}_V$ are bounded,
\item \label{assn:LAE_model_approx} There exists $\gamma_1$ and $\beta_{t}^L$ such that for all $t \leq T$, $\sup_{\mathcal{S}_{\xi_t}} || b_{0t}(\xi_t) - \beta_{t}^{L \, \prime} r^L(\xi_t) || \leq C L^{- \gamma_1}$. There exists $\gamma_2$, $\pi_{M}^K$ and $\pi_{k}^K$ such that $\sup_{\mathcal{S}_{V}^\varsigma} || \mathcal{M}_0^\varsigma(v) - p^K(v)' \, \pi_{M}^{K} || \leq C K^{- \gamma_2}$ and $\sup_{\mathcal{S}_{V}^\varsigma} || k_0^\varsigma(v) - p^K(v)' \, \pi_{k}^{K} || \leq C K^{- \gamma_2}$,
\item \label{assn:LAE_model_sievebasis} For all $t \leq T$, there exists $\Gamma_{1t}$, a $L \times L$ nonsingular matrix such that for $R_t^L(\xi_t) = \Gamma_{1t} r_t^L(\xi_t)$, $\mathbb{E}(R_t^L(\xi_t) R_t^L(\xi_t)')$ has smallest eigenvalue bounded away from $0$ uniformly in $L$.
There exists $\Gamma_2$, a $K \times K$ nonsingular matrix such that for $P^K(V) = \Gamma_2 p^K(V)$, $\mathbb{E}(P^K(V) P^K(V)')$ has smallest eigenvalue bounded away from $0$ uniformly in $K$,
\item \label{assn:LAE_model_matrixbounded} $|| \Lambda^M||$, $|| \Lambda^{k}||$, and $|| \Lambda^{bt}||$ are bounded,
\item \label{assn:LAE_model_rate}For $|R_t^L(\xi_t)|_0 \leq a_1(L)$, $|P^K(V)|_0^\varsigma \leq b_1(K)$, $|P^K(V)|_1^\varsigma \leq b_2(K)$, $|P^K(V)|_2^\varsigma \leq b_3(K)$, we have
$\sqrt{n} K^{- \gamma_2} = o(1)$, $\max(\sqrt{K}, \sqrt{L} b_2(K)) \, a_1(L) \sqrt{L/n} = o(1)$,
$b_2(K) \sqrt{n} L^{-\gamma_1} \allowbreak = o(1)$, $b_2(K)[\sqrt{L/n} + L^{- \gamma_1}][K + L] = o(1)$, $ b_3(K) [L /\sqrt{n} + \sqrt{n} L^{- 2\gamma_1}] = o(1)$, $b_2(K)^2 \sqrt{K} [\sqrt{L/n} + L^{- \gamma_1}] = o(1)$,
\item \label{assn:LAE_model_varbounded} \label{assn:MSE_model_extrammt} $\mathbb{E}(\dot{u} |X, Z) = 0$, and $\mathbb{E}(||v_{t}||^2|\xi_{t})$ and $\mathbb{E}(||\dot{u}||^2|X,Z)$ are bounded on $\mathcal{S}_{\xi t}$ and $\mathcal{S}_{X,Z}$ respectively.
\end{enumerate}
\end{assumption}
\begin{result}\label{rslt:LAE_model}
Under Assumptions \ref{assn:idCFA}, \ref{assn:cv_g} and \ref{assn:LAE_model},
\begin{align}
\ \sqrt{n} & \mathcal{X}_0^{(G)} [\hat{\mathcal{G}} - \mathcal{G}_0] \, = \, o_{\mathbb{P}}(1) \nonumber \\
& + \ \frac{1}{\sqrt{n}} \sum_{i=1}^n \Lambda^M (I_{(T-1)^2} \otimes \Theta) \, \operatorname*{Vec}(e_i^M) \otimes p_i + \frac{1}{\sqrt{n}} \sum_{i=1}^n \Lambda^{k} (I_{T-1} \otimes \Theta) \, e_i^{k} \otimes p_i \label{eq:linearEq}\\
& + \ \frac{1}{\sqrt{n}} \sum_{i=1}^n \sum_{t = 1}^T \left[\Lambda^M (I_{(T-1)^2} \otimes \Theta) (H_t^M - \text{d}P_t^M) + \Lambda^{k} (I_{T-1} \otimes \Theta) (H_t^{k} - \text{d}P_t^{k}) + \Lambda^{bt}\right] (I_{k_2} \otimes \Theta_1) \, v_{it} \otimes r_{it}, \nonumber
\end{align}
where we define $\Theta = \mathbb{E}(p_i p_i^{\prime})$ and $\Theta_1 = \mathbb{E}(r_i r_i^{\prime})$.
\end{result}
Result \ref{rslt:LAE_model} is obtained using Appendix \ref{app:lin_galcase}, where we look at Model (\ref{eq:reg_GalCase}) and obtain the influence function in the general case of a linear functional of nonparametric two-step series estimators, with some comments comparing our results to the existing literature. Under Assumption \ref{assn:LAE_model}, these results can be applied to the functionals in (\ref{eq:linearX_int}), choosing $w_i$ to be either a component of $M_i \dot{y}_i$ or of $M_i$. The regression functions $b_{0t}$ are estimated nonparametrically and the asymptotic distribution of functionals of such objects is studied in \cite{n97}.
Note that Assumption \ref{assn:LAE_model} (\ref{assn:MSE_model_extrammt}) is imposed to simplify computations. It amounts to strengthening the control function assumption, that is, Assumption \ref{assn:idCFA} (\ref{assn:idCFA_CFA}).
Note also that in the influence function, the terms $\text{d}P_t^M$ and $\text{d}P_t^k$ depend on $\rho^M$ and $\rho^k$. As discussed in Appendix \ref{app:lin_galcase}, this extra term compared to other applications of the CFA is due to the failure of the orthogonality condition discussed above Equation (\ref{eq:reg_GalCase}).
Under Result \ref{rslt:LAE_model}, we can write $\frac{1}{\sqrt{n}} \sum_{i=1}^n [\delta_i \mu_i - \mathbb{E}(\mu \delta)] + \, \frac{1}{\sqrt{n}} \sum_{i=1}^n Q_i^\delta \dot{u}_i \, - \, \sqrt{n} \mathcal{X}_0^{(G)} [\hat{\mathcal{G}} - \mathcal{G}_0]$ as $\frac{1}{\sqrt{n}} \sum_{i=1}^n s_{i,n} + o_{\mathbb{P}}(1)$ where
\begin{align*}
s_{i,n} = & [\delta_i \mu_i - \mathbb{E}(\mu \delta)] + Q_i^\delta \dot{u}_i - \Lambda^M (I_{(T-1)^2} \otimes \Theta) \, \operatorname*{Vec}(e_i^M) \otimes p_i - \Lambda^{k} (I_{T-1} \otimes \Theta) \, e_i^{k} \otimes p_i \\
& + \sum_{t = 1}^T \left[\Lambda^M (I_{(T-1)^2} \otimes \Theta) (\text{d}P_t^M - H_t^M) + \Lambda^{k} (I_{T-1} \otimes \Theta) (\text{d}P_t^{k} - H_t^{k} ) - \Lambda^{bt}\right] (I_{k_2} \otimes \Theta_1) \, v_{it} \otimes r_{it}.
\end{align*}
We define $\Omega = \operatorname*{Var}(s_{i,n})$ and the following objects, $\tilde{Q}_i^\delta = Q_i^\delta - \, \mathbb{E}(Q_i^\delta | V_i) \mathcal{M}_0(V_i)^{-1} M_i$, and
$\Omega_0 = \operatorname*{Var} \Big( [\delta_i \mu_i - \mathbb{E}(\mu \delta)] + \tilde{Q}_i^\delta \dot{u}_i + \allowbreak \sum_{t = 1}^T \mathbb{E}\left( \tilde{Q}_i^\delta \frac{ \partial g(V_i)}{\partial v_t} | \, \xi_{it} \right) v_{it} \Big)$.
\begin{assumption}\label{assn:MSE_model}
\
\begin{enumerate}[noitemsep,nolistsep]
\item \label{assn:MSE_model_approx} For each function $\lambda_a(.)$, column of $\lambda_{M}(.)$ or $\lambda_{k}(.)$, $\lambda_a(.)$ is continuously differentiable and
there exists $\iota_{a}^K$ such that $|\lambda_a(V) - \iota_{a}^K p^K(V)|_1 =O(K^{-\gamma_3})$, and $(\iota_{aht}^L,\iota_{a\rho t}^L)$ such that as $L \to \infty$,
$\mathbb{E}\left(||\mathbb{E} \left[ \lambda_a(V) \frac{\partial h^W(V)}{\partial v_{td}} | \xi_t \right] - \iota_{aht}^L r^L(\xi_t)||^2 \right) \to 0$ and
$\mathbb{E}\left(||\mathbb{E} \left[ \rho_i^W \frac{\partial \lambda_a(V)}{\partial v_{td}} | \xi_t \right] - \iota_{a\rho t}^L r^L(\xi_t)||^2 \right) \to 0$.
For all $t \leq T$, $b_t$ is continuous and there exists $\iota_{2t}^L$ such that $\mathbb{E}\left(|| \lambda_{bt}(\xi_t) - \iota_{bt}^L r^L(\xi_t)||^2 \right) \to 0$ as $L \to \infty$,
\item \label{assn:MSE_model_rate} $b_2(K) K^{- \gamma_3} = o(1)$,
\item \label{assn:MSE_model_var} $\mathbb{E}(Q_i^\delta||^2) < \infty$, $\mathbb{E}(|| \mu_i||^2) < \infty$, and there exists $C>0$ such that $\Omega_0 \geq C I_{d_x}$.
\end{enumerate}
\end{assumption}
\begin{result}\label{rslt:as.var}
Under Assumptions \ref{assn:idCFA}, \ref{assn:cv_g}, \ref{assn:LAE_model}' and \ref{assn:MSE_model}, $\Omega^{-1/2} \to_{n \to \infty} \Omega_0^{-1/2} $. Moreover, $|| \Lambda^M||$, $|| \Lambda^{k}||$, $||\Lambda^{bt}||$, $|| \Lambda^M (I_{(T-1)^2} \otimes \Theta) (H_t^M - \text{d}P_t^M)||$, and $|| \Lambda^{k} (I_{T-1} \otimes \Theta) (H_t^{k} - \text{d}P_t^{k})||$ are bounded.
\end{result}
Result \ref{rslt:as.var} states that $\Omega_0$ is the asymptotic variance of $\sqrt{n} \mathcal{X}_0^{(G)} [\mathcal{G} - \mathcal{G}_0]$. It additionally guarantees that Assumption \ref{assn:LAE_model} (\ref{assn:LAE_model_matrixbounded}) holds. The boundedness of the two last matrices is added for later results on asymptotic normality. We thus write Assumption \ref{assn:LAE_model}' for Assumption \ref{assn:LAE_model} without its condition (\ref{assn:LAE_model_matrixbounded}). Thus, under Assumption \ref{assn:LAE_model}' and Assumption \ref{assn:MSE_model}, Equation (\ref{eq:linearEq}) on $\sqrt{n} \mathcal{X}_0^{(G)} [\hat{\mathcal{G}} - \mathcal{G}_0] $ holds. Note that the condition $\Omega_0 \geq C I_{d_x}$ holds if for instance $\operatorname*{Var}(\mu_i | X_i, Z_i, u_i, V_i) \geq C I_{d_x}$ for some $C > 0$, or if a similar condition holds on the conditional variance of $\dot{u_i}$, as is typically assumed.
\subsubsection{Asymptotic distribution of $\hat{\mu}$}\label{sec:as_miou_together}
Asymptotic normality of $\hat{\mu}^\delta$ is obtained from the decomposition in (\ref{eq:dvlptasNorm}) and asymptotic equivalence of $\sqrt{n} \big[ \mathcal{X}_n(\hat{\mathcal{G}}) - \mathcal{X}_n(\mathcal{G}_0) \big]$ to its linearization. The conditions under which the two terms are equivalent asymptotically are detailed in Appendix \ref{sec:lineq_app}, one main condition being that of stochastic equicontinuity. To guarantee that it holds, we follow \cite{clvk03} and impose smoothness conditions. In particular,
for $\mathcal{S}_W$ a bounded subset of $\mathbb{R}^k$, we define for a function $g : \mathcal{S}_W \mapsto \mathbb{R}$, and $\varrho > 0$, the norm $||g||_{\infty,\varrho} = |g|_{ \left \lfloor{\varrho}\right \rfloor } + \max_{ |r| = \left \lfloor{\varrho}\right \rfloor } \sup_{w \neq w'} \frac{|\partial^rg(w) - \partial^rg(w')|}{|| w - w' ||^{\varrho - \left \lfloor{\varrho}\right \rfloor }}$.
We define $\mathcal{C}_c^{\varrho}(\mathcal{S}_W)$ to be the set of continuous functions $g : \mathcal{S}_W \mapsto \mathbb{R}$ such that $||g||_{\infty,\varrho} \leq c$. The set $\mathcal{H}_{2t,c}^\varrho = \mathcal{C}_c^{\varrho}(\mathcal{S}_{\xi_t})^{k_2}$ will be the class of vector valued functions taking values in $\mathbb{R}^{k_2}$, each component of which lies in $\mathcal{C}_c^{\varrho}(\mathcal{S}_{\xi_t})$. Since $k$ and $\mathcal{M}$ are defined on the extended support $\mathcal{S}_V^\varsigma$, we define $\mathcal{H}_{\mathcal{M},c,c'}^\varrho = \mathcal{C}_c^{\varrho}(\mathcal{S}_{V}^\varsigma)^{(T-1)^2} \cap \{ g : \forall V \in \mathcal{S}_{V}^\varsigma , \, \lambda_{\min} (M_{g(V)})> c' \}$, where $M_{g(V)}$ is the matrix formed by the coefficients of $g(V)$, and $\mathcal{H}_{k,c}^\varrho = \mathcal{C}_c^{\varrho}(\mathcal{S}_{V}^\varsigma)^{T-1}$. Finally, for the entire vector of infinite dimensional parameters $\mathcal{G}$, we define the set $\mathcal{H}_{c,c'}^\varrho = \left( \times_{t \leq T}\mathcal{H}_{2t,c}^\varrho \right) \times \mathcal{H}_{\mathcal{M},c,c'}^\varrho \times \mathcal{H}_{k,c}^\varrho$.
\begin{assumption}\label{assn:SE}
$\mathcal{G}_0 \in \mathcal{H}_{c,c'}^\varrho$ for some $(c,c') \in (\mathbb{R}_+^*)^2$ and $\varrho > \max(Td_2,d_z + d_1)/2$.
\end{assumption}
Condition (\ref{assn:LAE_model_approx}) of Assumption \ref{assn:LAE_model}' needs to be modified to
\textit{``There exists $\gamma_1$ and $\beta_{t}^L$ such that for all $t \leq T$, $\sup_{\mathcal{S}_{\xi_t}} || g^{ent}(\xi_t) - \beta_{t}^{L \, \prime} r^L(\xi_t) || \leq C L^{- \gamma_1}$. There exists $\gamma_2$, $\pi_{M,st}^K$ and $\pi_{k,t}^K$ such that $ | \mathcal{M}_0^\varsigma(.)_{st} - p^K(.)' \, \pi_{M,st}^{K} |_1^\varsigma \leq C K^{- \gamma_2}$ and $| k_0^\varsigma(.)_t - p^K(.)' \, \pi_{k,t}^{K} |_1^\varsigma \leq C K^{- \gamma_2}$, for all $1 \leq s,t \leq T-1$''.}
This modification is a stronger assumption, changing the approximation rate to be over the $|.|_1$ norm instead of the sup norm. Assumption \ref{assn:LAE_model}'' is the modified version of Assumption \ref{assn:LAE_model}'. We can now state the main results of this section.
\begin{result}\label{rslt:linear_miou_num}
Under Assumptions \ref{assn:idCFA}, \ref{assn:cv_g}, \ref{assn:LAE_model}'', \ref{assn:MSE_model} and \ref{assn:SE}, assuming moreover that $a_1(L) \Delta_n = o(n^{-1/4})$ and $b_2(K) [K/n + K^{-2 \gamma_2} + \Delta_n^2 b_2(K)^2]^{1/2} = o(n^{-1/4})$, then
$$ [\hat{\mu}^\delta - \mathbb{E}(\mu_i \delta_i)] = \frac{1}{\sqrt{n}} \sum_{i=1}^n s_{i,n} + o_{\mathbb{P}}(1).$$
\end{result}
\begin{assumption}\label{assn:as.normlt}
$\mathbb{E}[||\mu_i - \mathbb{E}(\mu)||^4] < + \infty$, $\mathbb{E}[||Q_i^\delta||^4] < \infty$ and $\operatorname*{Var}(\delta_i) >0$. Also, $\mathbb{E}(|| v_{t}||^4|\xi_{t})$ and $\mathbb{E}(|| \dot{u}||^4|X,Z)$ are bounded on $\mathcal{S}_{\xi t}$ for all $t \leq T$ and on $\mathcal{S}_{X,Z}$ respectively.
\end{assumption}
\begin{result}\label{rslt:as.normlt}
Under Assumptions \ref{assn:idCFA}, \ref{assn:cv_g},
\ref{assn:LAE_model}'', \ref{assn:MSE_model}, \ref{assn:SE} and \ref{assn:as.normlt}, assuming moreover that $a_1(L) \Delta_n = o(n^{-1/4})$ and $b_2(K)[K/n + K^{-2 \gamma_2} + \Delta_n^2 b_2(K)^2]^{1/2} = o(n^{-1/4})$,
$$\sqrt{n} [\hat{\mu} - \mathbb{E}(\mu|\delta)] \to^d \mathcal{N}(0, \Phi^{-2} \Xi), $$
where $\Xi = \Omega_0 + \mathbb{E}( (\delta_i - \Phi) s_i) \mathbb{E}(\mu |\delta)' + \mathbb{E}(\mu | \delta) \mathbb{E}( (\delta_i - \Phi) s_i)' + (\Phi - \Phi^2) \mathbb{E}(\mu | \delta)\mathbb{E}(\mu | \delta)'$.
\end{result}
\vspace{-12pt}
\section{Illustration}\label{sec:illus}
As an empirical exercise, we apply our method to a model of labor supply. An important parameter of interest is the elasticity of intertemporal substitution (EIS), or Frisch elasticity, which quantifies how labor supply responds to anticipated wage changes over time.
To estimate the EIS, the literature derives a labor supply response equation from a life-cycle labor supply model. Agents face uncertainty in future wage rates and interest rates, and choose paths of consumption, hours worked and possibly an asset distribution to maximize the expected discounted present value of lifetime utility under a dynamic budget constraint. A linear relationship is often obtained between the log of labor supply and the log of wage, see e.g \cite{mc85,f04}.
Within this framework, to estimate the EIS \cite{z97} studies the following model
\begin{equation}\label{eq:model_empirics1}
\ln h_{it} = \alpha_i + \mu \ln \omega_{it} + b'\chi_{it} + \epsilon_{it}, \qquad i=1..n, \ t=1..T,
\end{equation}
where $h_{it}$ is the number of hours worked annually by agent $i$, $\omega_{it}$ is the hourly wage, $\mu$ is the EIS and $\chi_{it}$ is composed of age and additional demographics (taste shifters). The additive fixed effect $\alpha_i$ depends on the initial marginal utility of wealth.
The empirical exercise considers instead a version with heterogeneity and, using a reduced form approach, studies
\begin{equation}\label{eq:model_empirics}
\ln h_{it} = \alpha_i + \mu_i \ln \omega_{it} + b'\chi_{it} + \epsilon_{it}, \qquad i=1..n, \ t=1..T,
\end{equation}
where the parameter $\mu_i$ is now allowed to vary across agents and potentially covary with $\omega_{it}$ and $\chi_{it}$. In (\ref{eq:model_empirics1}) $\omega_{it}$ and $\epsilon_{it}$ are usually assumed correlated. The main reason is $\epsilon_{it}$ depends on unexpected shocks to the marginal utility of wealth that are realized at time $t$ and these shocks are correlated with realized wage at $t$. Another reason
is that wage is usually observed with error in survey data.
This thus justifies applying the method developed in this paper to (\ref{eq:model_empirics}), and to do so, we will use the data in \cite{z97}.
We now discuss how to obtain control variables and provide some support for Assumptions \ref{assn:idCFA} and \ref{assn:id_M_GLn}. We can rewrite (\ref{eq:model_empirics}) with first differences, $
\dot{\ln} \, h_{it} = \mu_i \, \dot{\ln} \, \omega_{it} + b'\dot{\chi}_{it} + \dot{\epsilon}_{it}$.
For a panel with periods $t=0,..,T$, we consider the following conditions,
\begin{equation}
\text{for all } t \leq T, \ \ln \omega_{it} = \gamma_0 + \gamma_1 \ln \omega_{it-1} + \eta_{it}, \text{ and } \{\dot{\epsilon}_{it},(\eta_{is})_{s=t}^{T} \} \,\raisebox{0.05em}{\rotatebox[origin=c]{90}{$\models$}} \, \{ \omega_{i0}, (\eta_{is})_{s=1}^{t-1}, (\chi_{is})_{s=1}^{t} \}. \label{eq:wageAR_argUtWealthdpdce}
\end{equation}
These conditions are a simplification and chosen for illustrative purposes: the purpose is to provide an intuitive framework in which Assumptions \ref{assn:idCFA} and \ref{assn:id_M_GLn} hold.
The assumption that $\ln \omega_{it}$ follows an autoregressive process can be found in, e.g, \cite{hnr88} : (\ref{eq:wageAR_argUtWealthdpdce}) uses a simple AR(1) model.
The independence condition is based on the fact that the first-difference $\dot{\epsilon}_{it} = \epsilon_{it+1} - \epsilon_{it}$ is composed of shocks unexpected by the agent prior to $t+1$. Thus
as a white noise can be assumed independent of prior values of $\omega_{is}$, $s \leq t-1$, and of past and present values of the exogenous demographics. It is potentially correlated with future values of these exogenous demographics as for instance, if $\chi_{it}$ includes an indicator for bad health at period $t$ then it will be positively impacted by a positive shock to hours of work at period $t-1$.
Define $v_{it} = \ln \omega_{it} - \mathbb{E}(\ln \omega_{it}| \ln \omega_{i0})$, $V_{i} = (v_{it})_{1 \leq t \leq T}$, $W_i = (\ln \omega_{i1},..,\ln \omega_{iT})'$. Note that $v_{it} = \sum_{k=0}^{t-1} \gamma_1^k [\eta_{it-k} + \gamma_0]$ and that this relationship can be inverted so as to write $\eta_{it}$ as a function of $v_{i1},..,v_{it}$. Since the impact of $\chi_{it}$, $b$, is not heterogeneous, we use the framework detailed in Appendix \ref{sec:mix_homog_heterog} and condition on additional instruments for $\chi_{it}$, $Z_i^{\chi} = (\chi_{i1}, \chi_{i0})$. Then under (\ref{eq:wageAR_argUtWealthdpdce}), we have
\vspace{-30pt}
\begin{align*}
\text{for all } t \leq T, \ \mathbb{E}(\dot{\epsilon}_{it} | V_i, W_i, Z_i^{\chi})& = \mathbb{E}\left(\mathbb{E} (\dot{\epsilon}_{it} | \eta_{i1},.., \eta_{iT}, \ln \omega_{i0}, Z_i^{\chi}) | V_i, W_i, Z_i^{\chi} \right)\\
&= \mathbb{E}\left( \mathbb{E} (\dot{\epsilon}_{it} | \eta_{it},.., \eta_{iT}) | V_i, \ln \omega_{i0}, Z_i^{\chi} \right) = \mathbb{E} (\dot{\epsilon}_{it} | \eta_{it},.., \eta_{iT})=: g_t(V_i)
\end{align*}
where the second-to-last equality holds because $( \eta_{it},.., \eta_{iT})$ is a deterministic function of $V_i$. Note that the control function approach is directly on the first-differenced residual which is slightly different from Assumption \ref{assn:idCFA} but does not change the two-step identification procedure.
As for Assumption \ref{assn:id_M_GLn}, Result \ref{rslt:IA_illus} states that it holds under very mild conditions on the support of $\ln \omega_{i0}$.
The dataset constructed in \cite{z97}
is described in Section 2.1 of the paper. It is a selected sample from the Survey Research Center subsample of the Panel Study of Income Dynamics composed of $532$ men aged $22$ to $55$, married and working at all periods of the panel. The demographics $\chi_{it}$ are number of children, age and a dummy variable for bad health. We take $T=3$ and since we use an initial period $t=0$ to construct the control variables through $\ln \omega_{i0}$, the total number of periods in the panel is $4$. We use a panel of years $1979$ to $1982$ where period $1$ is year $1980$, period $T$ is year $1982$.
The estimation steps are as follows.
First, we estimate the control variables $v_{it}$ using the following specification, slightly richer than described above
\vspace{-5pt}
$$\ln \omega_{it} = \lambda_{0} + \gamma_{1t} \ln \omega_{i0} + \gamma_{2t}' \chi_{i}^{\text{GC}} + v_{it},$$
\vspace{-5pt}
where $\chi_{i}^{\text{GC}}$ includes $\chi_{i1}$ and age$_{i1}^2$. We choose this simple linear specification adding only a quadratic in age to avoid the curse of dimensionality which potentially has a strong impact given our small sample size.
The subsequent steps are estimation of the vector $b$, the functions $g_t(.)$ for $t \leq T-1$ and the average partial effect $\mathbb{E}(\mu_i)$: more details are provided in the Online Appendix \ref{sec:app_illus} together with the estimation results.
\vspace{-6pt}
\section{Conclusion}
This paper proved that in a correlated random coefficient panel model, the average partial effect $\mathbb{E}(\mu_i)$ is identified when the strict exogeneity condition standard in the literature is relaxed to allow for time-varying endogeneity. The approach imposes the existence of a control function to control for the dependence between the regressors and the time-varying disturbance with first-stage variables function of instruments and regressors. Within-group and between-group operations are then combined to disentangle the heterogeneous impact of the regressors from the time-varying disturbances. An estimator of $\mathbb{E}(\mu_i | \det(\dot{X}_i' \dot{X}_i) > \delta_0)$ can be naturally constructed from the identification argument: we showed its asymptotic normality and computed its asymptotic variance.
We highlight two directions for future research, following the contribution of \cite{gp12}. First, it would be of interest to relax the condition $T > d_x +1$: long enough panels might not be available to identify average partial effects in models with multiple covariates and heterogeneous impacts.
A second direction for future work would address the dependence of the limit of the proposed estimator, $\mathbb{E}(\mu| \delta)$, on the constant $\delta_0$: $\delta_0$ is arbitrarily fixed and choosing a value when implementing the estimator can be an issue in practice. One could instead study the asymptotic properties of $\mathbb{E}(\mu_i | \det(\dot{X}_i' \dot{X}_i) > \delta_n)$ as $\delta_n \to 0$.
\bibliography{bibCRC}