EconBase
← Back to paper

Identification and Estimation of Partial Effects in Nonlinear Semiparametric Panel Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

78,540 characters · 18 sections · 78 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and Estimation of Partial Effects in Nonlinear Semiparametric Panel Models

\newsavebox{\tablebox} \newlength{\tableboxwidth}

\setstretch{1}

abstract\begin{spacing}{1} Average partial effects (APEs) are often not point identified in panel models with unrestricted unobserved individual heterogeneity, such as a binary response panel model with fixed effects and logistic errors as a special case. This lack of point identification occurs despite the identification of these models' common coefficients. We provide a unified framework to establish the point identification of various partial effects in a wide class of nonlinear semiparametric models under an index sufficiency assumption on the unobserved heterogeneity, even when the error distribution is unspecified and non-stationary. This assumption does not impose parametric restrictions on the unobserved heterogeneity and idiosyncratic errors. We also present partial identification results when the support condition fails. We then propose three-step semiparametric estimators for APEs, average structural functions, and average marginal effects, and show their consistency and asymptotic normality. Finally, we illustrate our approach in a study of determinants of married women's labor supply. \end{spacing}

Keywords: Average partial effects, panel data, nonlinear models, semiparametric estimation, unobserved individual heterogeneity, binary response models

JEL classification: C13, C14, C23, C25

\baselineskip19.5pt

Introduction

Nonlinear panel models with unobserved individual heterogeneity are commonly used in empirical research. This paper is concerned with panels where the outcome is generated from the nonlinear semiparametric model

align[align omitted — 80 chars of source]

for units $i=1 ,\ldots, N$ and time periods $t=1, \ldots, T$. Here, $X_{it}$ are covariates, $C_i$ are possibly multi-dimensional unobserved individual heterogeneity, and $U_{it}$ are possibly multi-dimensional unobserved idiosyncratic errors. The function $g_t$ is potentially unknown and may vary across time. We assume $N$ is large, but that $T$ is small and fixed, as is the case in many microeconomic datasets. This class of models includes fixed effects, random effects, and intermediate levels of structure on the conditional distribution of $C_i|\mathbf{X}_i$, where $\mathbf{X}_i = (X_{i1},\ldots,X_{iT})'$. A leading example of this class of models in a binary outcome panel model generated by

align[align omitted — 118 chars of source]

See Wooldridge2010 Chapter 15.8 for an exposition of such models. We will use this binary response model to illustrate some results from the more general model in (ref).

Identification results for the common parameters $\beta_0$ in (ref) are well-known and go back to the work of Rasch1960 in the case where $U_{it}$ is assumed to be logistic and Manski1987 when its distribution is unspecified. Identification results for $\beta_0$ in many other special cases of (ref) have been derived. However, in this paper we focus on features of the distribution of the potential outcome

align*[align* omitted — 92 chars of source]

where $\underline{x}_t$ is in the support of $X_t$. They include the average structural function (ASF), average partial effects (APE), and average marginal effects (AME). First, as studied in BlundellPowell2003, the ASF at potential value $\underline{x}_t$ is the unconditional expectation of potential outcome $Y_{it}(\underline{x}_t)$. In the binary response model, it is the conditional response probability $\mathbb{P}(Y_{it} = 1|X_{it} = \underline{x}_t, C_i = c)$ averaged over the marginal distribution of the unobserved heterogeneity, $F_C$. It is used to assess the average impact of interventions in which the value of $X_{it}$ is manipulated. Second, the APE is a derivative of the ASF with respect to one covariate, hence measuring the partial effect of this covariate on the conditional response probability averaged over the marginal distribution of $C_i$. Both the ASF and APE can be averaged over a distribution for $\underline{x}_t$ to evaluate their averages. For example, we can average the ASF or APE across all covariates except for one “treatment” variable. Third, the AME measures the impact on the average potential outcome of a marginal increase in a covariate across the entire population. These measures are commonly used to evaluate the causal impact of policies. See AbrevayaHsu2021 for a survey of various partial effects in panels.

When $C_i|\mathbf{X}_i$ is unrestricted, i.e., under fixed effects, the ASF, APE, and AME are often not point identified. This can be the case even when the error distribution is known and when $\beta_0$ is point identified: see DaveziesDHaultfoeuilleLaage2021 who show their partial identification in a binary panel logit model with fixed effects, a model in which $\beta_0$ is point identified.

The main contribution of this paper is to show the point identification of these features under an index sufficiency assumption on the unobserved heterogeneity. This assumption restricts this conditional distribution to depend on covariates only through $v(\mathbf{X}_i)$, a (multiple) index of $\mathbf{X}_i$. It is related to an assumption of AltonjiMatzkin2005 and BesterHansen2009 that they use to show the identification of the local average response (LAR). While the APE averages partial effects over the unconditional distribution of the unobservables, the LAR differs by conditioning on the covariates: see the discussion below and in Section (ref) for a comparison of these estimands. Here we assume that the index function is known, or known up to a finite-dimensional parameter. Under our assumption, $v(\mathbf{X}_i)$ acts as a control function that does not require the specification of a first stage or the existence of an instrument. As we discuss in Section (ref), whether $v(\mathbf{X}_i)$ is a suitable control variable with enough variation depends crucially on the panel structure of the data. We also allow for estimated indices of the form $v(\mathbf{X}_i)'\gamma_0$, as in multiple index models IchimuraLee1991.\footnote{In general, we can relax the linear index structure by replacing $X_{it}'\beta_0$ and $v(\mathbf{X}_i)'\gamma_0$ with $w(X_{it};\beta_0)$ and $v(\mathbf{X}_i;\gamma_0)$, where $w(\cdot;\cdot)$ and $v(\cdot;\cdot)$ are known functions with unknown finite-dimensional parameters $\beta_0$ and $\gamma_0$. The linear index structure is more commonly used in the literature, so we focus on it in this paper.} As in ImbensNewey2009, the support of this index variable plays an important role, which we study in detail.

Note that the identification results in this paper do not rely on parametric assumptions on the conditional distribution of $C_i|\mathbf{X}_i$, nor on the distribution of $U_{it}$. While the ASF, APE, and AME depend directly on the distribution of $C_i$, which is not specified or identified, these partial effects are identified despite this dependence. Our approach can be viewed as a unified framework for identifying various partial effects under this assumption in a broad class of nonlinear panel models.

Though the primary emphasis of this paper is on the point identification of the partial effects, we also derive the partial identification results when the support condition is not satisfied and establish sharp bounds for the ASF in Section (ref). The identified set is relatively narrow when deviations from the support condition are small. Moreover, these partial identification results also help shed light on the effects of the index sufficiency and support condition on the identified set.

We then propose three-step semiparametric estimators for the ASF, APE, and AME. For example, we show the ASF is the partial mean of the conditional expectation of $Y_{it}$ given $(X_{it}'\beta_0,v(\mathbf{X}_i))$, integrated over the marginal distribution of $v(\mathbf{X}_i)$. In the first step, we estimate $\beta_0$ using one of the many available estimators in the literature. In the second step, we estimate the above conditional expectation, replacing the unobserved $X_{it}'\beta_0$ by generated regressor $X_{it}'\widehat{\beta}$. We then use local polynomial regression to recover this conditional mean. In the third step, we average this estimated conditional mean over the empirical distribution of $v(\mathbf{X}_i)$. The APE estimator is analogous, replacing the conditional expectation estimate with an estimate of its derivative, which is obtained directly via the local polynomial regression. The AME is similarly estimated. The estimators are easy to implement, and their convergence rates are fast relative to other nonparametric estimators since, after integrating over the distribution of $v(\mathbf{X}_i)$, the ASF and APE are functions of one-dimensional $X_{it}'\beta_0$. We offer a full treatment of their asymptotic properties in Supplemental Appendix (ref).

In our empirical illustration, we study women's labor force participation using our semiparametric estimator and compare it with a random effects (RE) estimator and a correlated random effects (CRE) estimator. The RE and CRE estimators we consider are commonly used parametric estimators that assume $C_i|V_i$ is Gaussian and that $U_{it}$ is logistic (see the definitions in Section (ref)). Compared to the parametric RE/CRE, our semiparametric APE estimates are closer to zero for lower husband's incomes and more negative for higher ones, whereas the parametric RE/CRE estimates show less variation across husband's incomes. Additionally, the effects of the husband's income are no longer significant once we allow for flexibility in the distributions of the unobserved heterogeneity and idiosyncratic errors.

In addition, we discuss an extension of our identification results to cases with sequentially exogenous $U_{it}$ in Supplemental Appendix (ref). This includes models with lagged dependent variables, see, e.g., HonoreKyriazidou2000. The associated estimators are similar under strict or sequential exogeneity.

Finally, we also conduct Monte Carlo simulation experiments in Supplemental Appendix (ref). Our results show that the semiparametric estimator yields smaller biases but larger standard deviations, and the former channel tends to dominate when the true distributions of the unobserved heterogeneity and the idiosyncratic errors are non-Gaussian and non-logistic, respectively.

Related Literature

We now discuss the literature that is closely related to the main text. References to related literature on estimation and dynamic panels can be found in their respective sections in the Supplemental Appendix.

First, while we focus on functionals of $Y_{it}(\underline{x}_t)$, our work builds on an extensive literature on the identification and estimation of $\beta_0$ in model (ref). This literature can be subdivided based on its distributional assumptions on $C_i|\mathbf{X}_i$ and those on $U_{it}|C_i,\mathbf{X}_i$.

In the case where $g_t$ is known and both $C_i|\mathbf{X}_i$ and $U_{it}|C_i,\mathbf{X}_i$ are parametrized, the distribution of $Y_i|\mathbf{X}_i$ is fully parametrized and $\beta_0$ can be estimated via integrated maximum likelihood. The binary outcome case is studied in Chamberlain1980. This case includes random effects, where $C_i|\mathbf{X}_i \overset{d}{=} C_i$ and $C_i$ follow a parametric distribution. See Chapter 3 in Chamberlain1984 and Chapter 15.8 in Wooldridge2010 for a review of this approach.

Under fixed effects, the literature on the identification and estimation of $\beta_0$ in binary models with logistic errors goes back to the work of Rasch Rasch1960,Rasch1961. Also see Andersen1970 and Chamberlain1980. For general error distributions, Manski1987 shows the identification of $\beta_0$ in binary panels when $X_{it}$ contains a continuous regressor with support equal to $\mathbb{R}$, and when $U_{it}|C_i,\mathbf{X}_i$ is stationary. Abrevaya1999 considers the identification of $\beta_0$ in the model $Y_{it} = g(X_{it}'\beta_0 + C_i + U_{it})$ when $U_{it}$ is nonparametric. The identification argument generalizes the one from Manski1987 for binary panels. Also see Abrevaya2000 and BotosaruMurisPendakur2021 for identification results under weaker assumptions on the link function $g$. These papers mainly focused on the identification of $\beta_0$ and $g(\cdot)$, whereas BotosaruMuris2022 show a range of point and partial identification results for the distribution of $Y_{it}(\underline{x}_t)$ under the above assumptions. In contrast, we achieve point identification even when $Y_{it}$ is binary and without assuming stationarity, but we impose restrictions on the conditional distribution of the heterogeneity.

We consider an intermediate restriction on $F_{C_i|\mathbf{X}_i}$: we assume it depends on $\mathbf{X}_i$ only through a known potentially multivariate index $v(\mathbf{X}_i)$. We do not parametrize the distribution of $C_i|\mathbf{X}_i$ nor restrict how it depends on this index. As mentioned above, our primary focus is on aspects of the distribution of $Y_{it}(\underline{x}_t)$, such as the ASF, APE, and AME, rather than $\beta_0$.

Second, due to our conditional independence assumption between the heterogeneity and covariates conditional on an index, our work is also related to a large literature on control functions. \citet*{NeweyPowellVella1999} show the identification of structural functions in a triangular model, where a control variable $V_i$ is identified from a first stage. BlundellPowell2004 consider a binary response model with endogeneity and focus on the identification and estimation of the ASF. ImbensNewey2009 consider a nonseparable triangular model and, like us, focus on identifying functionals of the structural function, such as the ASF. The multiple index structure we obtain for the conditional expectation $\mathbb{E}[Y_{it}|\mathbf{X}_i]$ is related to the work of IchimuraLee1991 and EscancianoJacho-ChavezLewbel2016 among others. For results on the ASF in binary panels, MaurerKleinVella2011 use a semiparametric maximum likelihood approach and a control function assumption to identify and estimate the ASF. Also see Laage2020 for a panel data model with triangular endogeneity.

Third, although they focus on a different estimand, the work of AltonjiMatzkin2005 and BesterHansen2009 is closely related to ours. In AltonjiMatzkin2005, they consider an exchangeability assumption, where $F_{C|X_1,\ldots,X_T}$ is invariant to relabeling of the time indices on the regressors. They then assume that $C_i|X_{it},v(\mathbf{X}_i) \overset{d}{=} C_i|v(\mathbf{X}_i)$ where $v(\mathbf{X}_i)$ are known symmetric functions of $(X_{i1},\ldots,X_{iT})$. They consider a nonparametric outcome equation and show the identification of the LAR, which averages changes in the structural function over the conditional distribution of the heterogeneity. Note that the LAR differs from the APE since the former integrates over the conditional distribution of $C_i|X_{it}$ rather than its marginal distribution. We also achieve point identification of the LAR: see Theorem (ref). Because of the single-index structure of outcome equation $g_t(X_{it}'\beta_0,C_i,U_{it})$, this identification is achieved under weaker support conditions on $\mathbf{X}_i$ and the indices than in their model. We discuss in more detail in Section (ref) the difference in estimands, and related differences in assumptions on the support of the index are illustrated in the discussion after Theorem (ref). Moreover, our semiparametric structure allows for much faster rates of convergence for our APE when compared to the rates obtained for the LAR in their nonparametric outcome equation. In particular, their rate of convergence for their LAR estimator decreases with the dimension of $X_{it}$ while the rate of convergence of our APE estimator does not, since it depends on the dimension of $X_{it}'\beta_0$, which is fixed.

In BesterHansen2009 they also consider an index sufficiency assumption, where each index is a function of each covariate. Specifically, they have $v(\mathbf{X}_i) = (v_1(\mathbf{X}_{i}^{(1)}),\ldots,v_{d_X}(\mathbf{X}_{i}^{(d_X)}))$, where $d_X$ denotes the dimension of $X_{it}$ , and $\mathbf{X}_i^{(k)}$ denotes a $T\times 1$ vector with the $k$th components of $X_{it}$ for $t=1,\ldots,T$. Their indices $\{v_j(\cdot)\}_{j=1}^{d_X}$ are allowed to be unknown under the separability requirement, while our indices are assumed to be known, or known up to finite-dimensional parameters.

Other identification approaches in these models have also been proposed. GrahamPowell2012 consider a correlated random coefficients model and estimate averages of these coefficients, which correspond to the APEs. HoderleinWhite2012 consider the identification of the LAR for a subpopulation of stayers in a nonseparable model. BonhommeLamadonManresa2017 consider a discretization of the unobserved heterogeneity as an intermediate assumption between fixed and random effects.

Finally, we also consider partial identification when the support condition fails, and our proof builds on ImbensNewey2009. In fixed effects binary response models, DaveziesDHaultfoeuilleLaage2021 derive bounds for the AME and ASF when $U_{it}$ is assumed to be logistic. BotosaruMuris2022 develop partial identification results under weaker assumptions on the link function. ChernozhukovFernandez-ValHahnNewey2013 derive bounds on the ASF with nonparametric distributions of $C_i|\mathbf{X}_i$ and of $U_{it}$. Also see ChernozhukovFernandez-ValHoderleinHolzmannNewey2015.

The remainder of this paper is organized as follows. In Section (ref) we present the baseline model and provide our main identification results. Section (ref) briefly describes our proposed estimators for the ASF and APE. Section (ref) applies our ASF and APE estimators to an empirical illustration of female labor force participation. Finally, Section (ref) concludes. Appendix (ref) shows the identification of $\beta_0$ under the index condition, and Appendix (ref) contains the proofs for all propositions and theorems. A Supplemental Appendix is also available online with theoretical results on the asymptotic properties of our proposed estimators, an extension to a dynamic panel data model, a discussion of implementation details, Monte Carlo experiments, and additional tables and figures for the empirical illustration.

Model and Identification

In this section, we describe the panel model of interest and show the identification of the ASF, APE, and AME under an index sufficiency assumption on the conditional distribution of the heterogeneity.

Model and Estimands

Recall the baseline model in equation (ref)

align*[align* omitted — 63 chars of source]

where $i=1,\ldots,N$, $t=1,\ldots,T$. Here, $X_{it} \in \mathcal{X}_t \subseteq \mathbb{R}^{d_X}$ are covariates and $\beta_0 \in \mathcal{B} \subseteq \mathbb{R}^{d_X}$ are unknown parameters. Let $\mathbf{X}_i \in \mathcal{X} \subseteq \mathbb{R}^{T\times d_X}$ denote the observed covariate matrix which has $X_{it}'$ as its $t$th row, where $\mathbb{R}^{a \times b}$ denotes the set of matrices with $a$ rows, $b$ columns, and real entries. Let $C_i \in \mathcal{C} \subseteq \mathbb{R}^{d_C}$ denote the unobserved individual heterogeneity, and $U_{it} \in \mathcal{U}_t \subseteq \mathbb{R}^{d_U}$ are idiosyncratic errors. Let $Y_i = (Y_{i1},\ldots,Y_{iT})$ denote the vector of outcomes for unit $i$. Note that the outcome function $g_t$ can depend on the time period. For example, this allows for $g_t(X_{it}'\beta_0, C_i, U_{it}) = g(X_{it}'\beta_0 + \delta_t, C_i, U_{it})$, where $\delta_t$ is an additive time-effect, or for $g_t(X_{it}'\beta_0, C_i, U_{it}) = g(X_{it}'\beta_0 + \delta_t C_i, U_{it})$, an interactive effect.

The $i$ subscript is suppressed in the remainder of this section and when there is no confusion. We maintain the following assumptions on the baseline model. Let $\text{supp}(\cdot)$ denote the support of a random vector or variable.

\begin{IDsec_assump}[Model assumptions] For each $t = 1,\ldots,T$,

itemize$Y_t$ is generated according to equation (ref); • $U_t \perp \! \! \! \perp \mathbf{X}|C$; • $\mathbb{E}[|g_t(\underline{x}_t'\beta_0,C,U_t)|] < \infty$ for all $\underline{x}_t'\beta_0 \in \text{supp}(X_t'\beta_0)$.

\end{IDsec_assump}

Besides assuming model equation (ref) holds, A(ref).(ii) also imposes that unobserved variables $U_t$ are independent of covariates given the unobserved individual heterogeneity. This strict exogeneity assumption rules out the presence of lagged dependent variables in $\mathbf{X}$. We relax this assumption and consider models with lagged dependent variables in Supplemental Appendix (ref). The relationship between $C$ and $\mathbf{X}$ is unrestricted by A(ref), and this assumption allows for serial correlation and nonstationarity in $U_t$. Lastly, A(ref).(iii) ensures that the ASF is well defined.

Let $Y_t(\underline{x}_t) \equiv g_t(\underline{x}_t'\beta_0,C,U_t)$ denote the potential outcome at time $t$ evaluated at covariate value $\underline{x}_t \in \mathcal{X}_t$. We define the ASF at time $t$ evaluated at $\underline{x}_t$ by

align*[align* omitted — 142 chars of source]

It is the average outcome if $X_t$ were set to $\underline{x}_t$ in an exogenous manner. The ASF generally differs from the identified conditional expectation $\mathbb{E}[Y_t|X_t = \underline{x}_t]$ due to the dependence between $C$ and $X_t$, unless $C \perp \! \! \! \perp X_t$, a random effects assumption.

In the binary response model, it can alternatively be defined as a function of the conditional response probability, $\mathbb{P}(Y_t = 1 | X_t = \underline{x}_t, C = c)$. The ASF is then defined as the conditional response probability integrated over the marginal distribution of the unobserved effect $C$:

align*[align* omitted — 147 chars of source]

By Assumption A(ref).(ii), this definition coincides with ours.

Our second object of interest is the APE, which measures the partial effect of changing one covariate, averaged over the marginal distribution of $C$. If this covariate is continuously distributed, the APE is the derivative of the ASF with respect to this covariate, assuming that the derivative exists. Formally, define the APE of the $k$th element of $\underline{x}_t \in \mathcal{X}_t$, denoted by $\underline{x}_t^{(k)}$, as follows:

align*[align* omitted — 264 chars of source]

where $\beta_0^{(k)}$ is the $k$th element of $\beta_0$.

In the case where $X_t^{(k)}$ is discretely distributed, the APE is the difference between the ASF at two values, which can be interpreted as an average treatment effect. We let

align*[align* omitted — 213 chars of source]

where $\underline{x}_{t}^*$ is a vector that differs from $\underline{x}_t$ in its $k$th position.

Finally, assuming that the necessary derivatives exist, define the local average response as follows:\footnote{Note that we can also define an alternative LAR that conditions on covariate values at time $t$ only: $\partial \mathbb{E}[Y_t(\underline{x}_t)|X_t = x_t]/\partial \underline{x}_t^{(k)} |_{x_t = \underline{x}_t}$. This alternative LAR can be obtained from $\mathbb{E}[\text{LAR}_{k,t}(\mathbf{X})|X_t = \underline{x}_t]$.}

align*[align* omitted — 378 chars of source]

We define the average marginal effect as its average over the distribution of $\mathbf{X}$:

align*[align* omitted — 173 chars of source]

In contrast to the LAR, the APE averages this response over the entire population. Thus, the APE is analogous to average treatment effects (ATE) in the causal inference literature, which averages the difference between two potential outcomes over its unconditional distribution; meanwhile, the LAR is analogous to a local treatment effect, where the averaging occurs over the conditional distribution of the heterogeneity given $\mathbf{X} = \underline{\bold{x}}$.\footnote{Specifically, the local average response can be viewed as an average causal response on the treated (ACRT) since it conditions on the subpopulation with covariate values $\underline{\bold{x}}$, see CallawayGoodman-BaconSantAnna2021.} Neither estimand is more general since knowledge of the LAR for all $\underline{\bold{x}} \in \text{supp}(\mathbf{X})$ does not imply knowledge of APEs, and vice-versa. We later show that the LAR is identified under weaker index support assumptions than the APE.

The AME averages this local response over the distribution of covariates and essentially measures the impact of a small change in covariate $X_t^{(k)}$ for all units on average outcomes.

remark[Integrated estimands] It may also be of interest to consider averages of the ASF or APE over certain covariate values. For example, one can consider the APE's average over the marginal distribution of $X_t$, or over the distribution of all covariates except for one: \begin{align*} \widetilde{APE}_{k,t} &\equiv \int_{supp(X_t)} APE_{k,t}(x_t) \; dF_{X_t}(x_t)\\ \widetilde{APE}_{k,t}(\underline{x}_t^{(k)}) &\equiv \int_{\text{supp}(X_t^{(-k)})} \text{APE}_{k,t}(\underline{x}_t)\; dF_{X_t^{(-k)}}(\underline{x}_t^{(-k)}) \end{align*} where the $(-k)$ superscript denotes removal of the $k$th entry.

Identifying Assumptions

Without further assumptions, it is generally impossible to point identify these partial effects, even under parametric assumptions on $U_t$. Under fixed effects, the ASF, APE, and AME are generally partially identified.\footnote{ Point identification can be obtained if more structure is assumed, such as random effects: $C \perp \! \! \! \perp \mathbf{X}$. Another example is the set of assumptions in BotosaruMuris2022, which include $g_t$ being invertible in $X_t'\beta_0 + C - U_t$, $Y_t$ being continuous, and $(\beta_0,g_t)$ being identified.}

(Non-)Identification under Fixed Effects

To fix ideas, we illustrate this identification failure in the binary response model of equation (ref) with scalar individual effects, a special case of our general model.

Under fixed effects, $C_i|\mathbf{X}_i$ is unrestricted. Denote by $G_t(\underline{x}_t'\beta_0,\mathbf{x})$ the counterfactual conditional probability

align*[align* omitted — 256 chars of source]

Note that we observe the conditional probabilities $\mathbb{P}(Y_t = 1 | \mathbf{X} = \mathbf{x}) \equiv G_t(x_t'\beta_0,\mathbf{x})$ for all $\mathbf{x} \in \text{supp}(\mathbf{X})$ and $t \in \{1,\ldots,T\}$. By the law of total probability, the ASF for covariate value $\underline{x}_t$ is

align[align omitted — 746 chars of source]

We can see from equation (ref) that the ASF is an average over the distribution of $\mathbf{X}$ of conditional probability $G_t(\underline{x}_t'\beta_0,\mathbf{X})$. For the ASF to be point identified, we need $G_t(\underline{x}_t'\beta_0,\mathbf{x})$ to be identified for all $\mathbf{x} \in \text{supp}(\mathbf{X})$, but this generally fails since, given $X_t'\beta_0 = \underline{x}_t'\beta_0$, the support of $\mathbf{X}$ does not equal its marginal support. In equation (ref), the $G_t(\underline{x}_t'\beta_0,\mathbf{x})$ where $x_t'\beta_0 \neq \underline{x}_t'\beta_0$ are counterfactual probabilities that are not point identified from the data since they do not correspond to any conditional probability of $Y_t$ given $\mathbf{X} = \mathbf{x}$.

Unless restrictions are imposed on the distribution of $C|\mathbf{X}$ or other aspects of the model, this causes the ASF, and therefore the APE too, to be partially identified. In the logit case, \citet*{DaveziesDHaultfoeuilleLaage2021} provide partial identification results for the ASF and AME. If one further relaxes their assumption that the error distribution is known and logistic while retaining their other assumptions, the resulting bounds would become weakly wider, so point identification cannot be achieved in other binary response models either. In the nonparametric case, ChernozhukovFernandez-ValHahnNewey2013 obtain bounds on the ASF. In the context of the general nonlinear semiparametric model in equation (ref), we establish sharp bounds on the ASF in Corollary (ref) below.

Identification of Common Parameters

While $\beta_0$ is not the object of interest, its identification facilitates the identification of partial effects. In what follows, we also take the identification of $\beta_0$ as given. \begin{IDsec_assump}[Identification of coefficients] $\beta_0$ is point identified or point identified up to scale. \end{IDsec_assump}

This assumption can be justified by the fact that $\beta_0$'s identification can be established for many special cases of models (ref) satisfying Assumption A(ref). For example, when $Y_t$ is binary, $g_t(X_t'\beta_0,C,U_t) = \mathbbm{1}(X_t'\beta_0 + C - U_t \geq 0)$, and $U_t$ follows a standard logistic distribution, Rasch1960 showed that $\beta_0$ is point identified under minimal assumptions requiring variation in $X_t$ over time, allowing for all regressors to be discrete. Still in the binary outcome model, Manski1987 showed that $\beta_0$ is identified up to scale when $U_t|C,\mathbf{X}$ is stationary and when $X_t$ includes a continuous regressor with support equal to $\mathbb{R}$. Unlike the previous result, this does not require knowledge that $U_t$ follows a logistic distribution. Zhu2022 recently showed that this identification holds under weaker support assumptions on the regressors.

Manski's result is generalized to non-binary outcomes in Abrevaya1999 where he considers $g_t(X_t'\beta_0, C, U_t) = h(C + X_t'\beta_0 - U_t)$, where $h$ is weakly increasing. He also assumes the existence of a regressor with large support and shows $\beta_0$ is point identified up to scale. Abrevaya2000 and BotosaruMurisPendakur2021 generalize these results to nonseparable and time-varying models.

A variety of other identification approaches can also be used. These could include special regressors, as in HonoreLewbel2002, but see \citet*{ChenKhanTang2019} who point out that this may not be the case for certain dynamic binary choice panel data models. Lee1999 and \citet*{ChenSiZhangZhou2017} provide alternative assumptions that yield the identification of $\beta_0$. In Appendix (ref), we provide a new approach to point identify $\beta_0$ under the index sufficiency condition (see Assumption A(ref) below), where we consider both the baseline model with a known index function as well as the case where the index function is known up to finite-dimensional parameters.

As mentioned above, the point identification of $\beta_0$ does not imply the point identification of any partial effect as was shown in, for example, DaveziesDHaultfoeuilleLaage2021 for the binary panel logit model, and in Corollary (ref) below for our nonlinear semiparametric model.

An Index Assumption

To achieve point identification of partial effects, we consider an index sufficiency restriction that imposes additional structure on the conditional distribution of the heterogeneity. In the example of equation (ref), we assume that $G_t(\underline{x}_t'\beta_0,\mathbf{x})$ depends on $\mathbf{x}$ only through known index functions. \begin{IDsec_assump}[Index sufficiency] Given $V \equiv v(\mathbf{X})$, where $v:\mathbb{R}^{T \times d_X} \rightarrow \mathbb{R}^{d_V}$ is known, let $C \mid \mathbf{X} \overset{d}{=} C \mid V$. \end{IDsec_assump}

This assumption is a correlated random effects assumption that restricts the conditional distribution of $C|\mathbf{X}$ to depend solely on $v(\mathbf{X})$, which are indices of $\mathbf{X}$. The conditional distribution of $C|v(\mathbf{X})$ remains nonparametric though. This assumption is similar to Assumption 2.1 in AltonjiMatzkin2005. On its own, Assumption (ref) holds if $v$ is the identity function, i.e., under fixed effects. However, to identify various partial effects we will impose restrictions on the support of $V$ that rule out fixed effects. These support restrictions vary with the partial effects being considered, so we introduce and discuss them separately: see Theorems (ref) and (ref) in Section (ref) below.

Motivation for Assumption A(ref)

Such an index assumption is considered in AltonjiMatzkin2005 and BesterHansen2009, and can be motivated from several perspectives.

In AltonjiMatzkin2005, the exchangeability of $f_{C|\mathbf{X}}(c|x_1,\ldots,x_T)$ in $(x_1,\ldots,x_T)$ is assumed. They consider symmetric polynomials as candidates for the index function, e.g., $v(\mathbf{X}) = \left(\sum_{t=1}^T X_t, \sum_{1\leq t_1 < t_2 \leq T} X_{t_1}X_{t_2} \right)$ when the indices are the first two elementary symmetric functions and $X_t$ is scalar. Unlike us, BesterHansen2009 do not assume $v(\cdot)$ is known, but they do not allow for the indices to be arbitrary functions of $\mathbf{X}$: each component on the index may only depend on one component of $\mathbf{X}_t$. This requires that $T \geq 3$, that covariates are continuously distributed, and that the index $v$ satisfies their separability requirement. Note that if their assumptions are met, we could build on their results to relax A(ref) and assume unknown index functions. The focus of these two papers is also different from ours: they identify the LAR rather than the ASF or APE, and their identification of the LAR is shown for continuous covariates.

Covariate assignment models in panel data can also be used to find candidate indices. This is explored in ArkhangelskyImbens2023 where they assume the distribution of $\mathbf{X}|C$ is from an exponential family with a known sufficient statistic. For example, if $(X_1,\ldots,X_T)|C$ are assumed iid Gaussian, then $v(\mathbf{X}) = (\sum_{t=1}^T X_t, \sum_{t=1}^T X_t^2)$ forms a sufficient statistic for $\mathbf{X}$ by the Fisher-Neyman factorization theorem. Therefore, we have that $\mathbf{X}|C,v(\mathbf{X}) \overset{d}{=} \mathbf{X}|v(\mathbf{X})$, which implies A(ref) holds.

In a special case where the indices are time-averages, i.e., $v(\mathbf{X}) = \frac{1}{T}\sum_{t=1}^T X_t$, the index assumption is consistent with $C = \zeta\left(\left(\frac{1}{T}\sum_{t=1}^T X_t\right)'\gamma_0,\eta\right)$ where $\eta \perp \! \! \! \perp \mathbf{X}$ and $\zeta(\cdot,\cdot)$ is any function. This is a relaxation of the specification of the conditional distribution of $C$ given $\mathbf{X}$ in Mundlak1978 since we do not specify the distribution of $\eta$ nor restrict the functional form of $\zeta(\cdot,\cdot)$. In this specification for $C$, the one-dimensional index $v(\mathbf{X}) = \sum_{t=1}^T X_t'\gamma_0$ also satisfies A(ref), but is unknown due to its dependence on unknown $\gamma_0$. Proposition (ref) in the Appendix shows how the unknown index parameter can be identified building on the work of IchimuraLee1991 on the identification and estimation of multiple index models.

Identification of Partial Effects

We can now state our main identification results.

theoremLet $t \in \{1,\ldots,T\}$, $\underline{x}_t \in \text{supp}(X_t)$, and let Assumptions A(ref)--A(ref) hold. Then, \begin{enumerate} • $\text{ASF}_t(\underline{x}_t)$ is point identified from the distribution of $(Y,\mathbf{X})$ when $\text{supp}(V|X_t'\beta_0 = \underline{x}_t'\beta_0) = \text{supp}(V)$; • Let the partial derivative of $\text{ASF}_t(\underline{x}_t)$ with respect to $\underline{x}_t^{(k)}$ exist. Then, $\text{APE}_{k,t}(\underline{x}_t)$ is point identified from the distribution of $(Y,\mathbf{X})$ when $\text{supp}(V|X_t'\beta_0 = u) = \text{supp}(V)$ for all $u$ in a neighborhood of $\underline{x}_t'\beta_0$. \end{enumerate}

The APE part of this theorem assumes that $X_t^{(k)}$ is continuously distributed. For discretely distributed $X_t^{(k)}$, the APE is a difference between two ASFs, and its identification is achieved when the two corresponding ASFs are point identified. We omit this case for brevity.

For example, the point identification of the ASF occurs because we can write it as follows:

align[align omitted — 764 chars of source]

The second equality follows from iterated expectations and the third from the support assumption in the theorem's statement. The fourth follows from $(C,U_t) \perp \! \! \! \perp X_t'\beta_0|V$, which is implied by Assumptions A(ref).(ii) and A(ref). Equation (ref) depends only on $\left\{\mathbb{E}[Y_t|X_t'\beta_0 = \underline{x}_t'\beta_0, V = v] : v \in \text{supp}(V) \right\}$ and on the marginal distribution of $V$, which are both identified from the data. All identification results in this subsection still hold if $V$ is replaced by $V'\gamma_0$ and Assumption A(ref) in the Appendix holds.

remarkNote that Theorem (ref) can be used to point identify $\mathbb{E}[m(Y_t(\underline{x}_t))]$ for any known function $m$ without having to modify any assumptions. In particular, one can identify $F_{Y_t(\underline{x}_t)}(y) = \mathbb{E}[\mathbbm{1}(Y_t(\underline{x}_t) \leq y)]$, the potential outcomes cdf, by setting $m(a) = \mathbbm{1}(a \leq y)$ under the assumptions of Theorem (ref). Therefore, one can also identify the Quantile Structural Function (QSF),\footnote{See ImbensNewey2009 or ChernozhukovFernandez-ValHoderleinHolzmannNewey2015 for example.} $\text{QSF}_t(\tau;\underline{x}_t) \equiv F_{Y_t(\underline{x}_t)}^{-1}(\tau)$, whenever the ASF is identified.

To identify the APE of a continuous regressor, we note that the support assumption implies the ASF is point identified for values of $X_t$ near $\underline{x}_t$. Since the APE is a derivative of the ASF, we can identify the APE as a limit of finite differences between identified ASFs. Formally, we can write

align[align omitted — 208 chars of source]

All quantities in equation (ref) are identified, hence the APE is identified. Note that the identification of the ASF and APE bypasses the need to identify $F_C$, the distribution of the heterogeneity. As a side note, in the previous version of this paper liu2021identification, we show that $F_C$ is identified under stronger support assumptions on $(X_t'\beta_0,V)$ when outcomes are binary and $C$ is scalar.

We now contrast these with identification results for the LAR and AME.

theoremLet $t\in\{1,\ldots,T\}$, $\underline{\bold{x}} \in \text{supp}(\mathbf{X})$, and let Assumptions A(ref)--A(ref) hold. Then, \begin{enumerate} • $\text{LAR}_{k,t}(\underline{\bold{x}})$ is point identified from the distribution of $(Y,\mathbf{X})$ when $\mathbb{E}[Y_t(\underline{x}_t)|\mathbf{X} = \mathbf{x}]$ is differentiable in $\underline{x}_t^{(k)}$ for $\mathbf{x} = \underline{\bold{x}}$ and when $v(\underline{\bold{x}}) \in \text{supp}(V|X_t'\beta_0 = u)$ for all $u$ in a neighborhood of $\underline{x}_t'\beta_0$; • $\text{AME}_{k,t}$ is point identified from the distribution of $(Y,\mathbf{X})$ if the above condition holds for all $\underline{\bold{x}} \in \text{supp}(\mathbf{X})$ up to a $\mathbb{P}_\mathbf{X}$-measure zero set. \end{enumerate}

The following equations help explain how the LAR and AME are identified:

align[align omitted — 419 chars of source]

Note that the condition for the identification of the LAR is weaker than that for the APE. The APE requires that $\text{supp}(V) = \text{supp}(V|X_t'\beta_0 = u)$ for $u$ in a neighborhood of $\underline{x}_t'\beta_0$, while the LAR requires that $v(\underline{\bold{x}}) \in \text{supp}(V|X_t'\beta_0 = u)$ for $u$ in a neighborhood of $\underline{x}_t'\beta_0$. This is weaker because $v(\underline{\bold{x}}) \in \text{supp}(V)$ by construction.

These support conditions indirectly but critically rely on the panel structure of the data and on the index structure of $X_t'\beta_0$. We discuss these support conditions below.

Discussion of the Support Conditions in Theorems (ref)--(ref)

The support condition for the APE in Theorem (ref) requires $X_t'\beta_0$ to be continuously distributed in a neighborhood of $\underline{x}_t'\beta_0$. This allows for some components of $X_t$ to be discretely distributed.

We do not require $X_t'\beta_0$ to be supported on the entire real line, but for the APE, we do require that the support of the sufficient statistic is independent of the value of $X_t'\beta_0$ in a neighborhood of $\underline{x}_t'\beta_0$. This support assumption is related to the common support assumption of ImbensNewey2009, although we only restrict the support of $V|X_t'\beta_0$ rather than the support of $V|X_t$. This is a key benefit of the semiparametric setup where the outcome equation depends on index $X_t'\beta_0$, as the support of $V|X_t'\beta_0$ is by construction a superset of the support of $V|X_t$. As opposed to ImbensNewey2009, we do not posit the existence of a first stage or exogenous excluded variables since our indices are functions of $\mathbf{X}$ only. Also see Remark (ref) below.

AltonjiMatzkin2005 do not consider the identification of the ASF/APE and instead focus on the LAR. To understand the difference in identifying assumptions, in their nonparametric setting, identification of the ASF or APE would require $\text{supp}(v(\mathbf{X})|X_t = \underline{x}_t) = \text{supp}(v(\mathbf{X}))$. This is significantly stronger than our condition whenever more than one covariate is present. To see this, assume $d_X = T = 2$, $X_t^{(1)}$ is continuously distributed on $\mathbb{R}$, and that $X_{t}^{(2)} \in \{0,1\}$ is binary. Let $v(\mathbf{X}) = \sum_{t=1}^2 X_t = (\sum_{t=1}^2 X_t^{(1)}, \sum_{t=1}^2 X_t^{(2)})$. Then, under minimal assumptions, $\text{supp}(v(\mathbf{X})) = \mathbb{R} \times \{0,1,2\}$ but $\text{supp}(v(\mathbf{X})|X_1 = \underline{x}_1) = \mathbb{R} \times \{\underline{x}_1^{(2)}, \underline{x}_1^{(2)} + 1\} \neq \text{supp}(v(\mathbf{X}))$. On the other hand, the conditional support of $v(\mathbf{X})$ given $\{X_t'\beta_0 = x_t'\beta_0\}$ equals $\text{supp}(v(\mathbf{X}))$ when $\beta_0^{(1)} \neq 0$. Therefore, in this example, the ASF/APE will be identified under our assumptions in the semiparametric model, but not in their nonparametric model.

This important condition also has implications on the dimension of $v(\mathbf{X})$. For example, this condition is violated when $v(\mathbf{X}) = \mathbf{X}$, i.e., no index restrictions are imposed and, equivalently, we have fixed effects. This is because the support of $\mathbf{X}$ does not equal its conditional support given $X_t'\beta_0$: $\text{supp}(\mathbf{X}|X_t'\beta_0 = \underline{x}_t'\beta_0) \neq \text{supp}(\mathbf{X})$. On the other hand, if $v(\mathbf{X}) = \sum_{t=1}^T X_t \in \mathbb{R}^{d_X}$, this condition is written as $\text{supp}(\sum_{t=1}^T X_t | X_t'\beta_0 = u) = \text{supp}(\sum_{t=1}^T X_t)$. For example, we can see that this holds in the simple case where $(X_1,\ldots,X_T)$ are jointly normally distributed.

Although not studied here, we note that the support conditions are potentially testable since they only depend on the observed variables $\mathbf{X}$ and identified parameter $\beta_0$.

Finally, in Section (ref) below, we show that while the support condition may not always be warranted, the ASF and APE are partially identified when it fails.

remark[Excluded control variable] A more general version of A(ref) is that we can identify a variable $V$ such that $C|\mathbf{X},V \overset{d}{=} C|V$. This is a control variable assumption, where $V$ may be an unobservable that is not functionally related to $\mathbf{X}$. For example, it could be the residual in a first-stage equation relating $\mathbf{X}$ to some excluded instruments, whose existence we do not assume in this paper. See, for example, ImbensNewey2009 in the nonseparable cross-sectional case or Laage2020 for a panel model with triangular endogeneity and control functions. If this control variable satisfies the support conditions in either Theorem (ref) or (ref), the corresponding identification result still applies provided that $U_t \perp \! \! \! \perp \mathbf{X}|(C,V)$, a modification of A(ref).(ii). As for estimation, if $V$ is identified from a first-stage equation, we should substitute $\widehat{V}_i$ for $V_i$, where $\widehat{V}_i$ is a suitable estimator for the control variable. This additional generated regressor's impact on the limiting distribution would then have to be taken into account. In this paper, we focus on the case where no such $V$ is observed or identified from a first-stage model, and instead where $V$ is an index of $\mathbf{X}$.

Relaxing Support Assumptions and Partial Identification

The validity of the support assumptions in Theorems (ref) and (ref) depends intricately on the support of $\mathbf{X}$, and on the considered value $\underline{x}_t$. For a given $\underline{x}_t$, these assumptions may fail. When they do, we can show the ASF, APE, LAR, and AME are partially identified instead. These partial identification results could help us better understand how the index sufficiency and support condition affect the identified set. In this section we focus on the ASF and the related APE.

Let $\mathcal{V}_t(u) = \text{supp}(V|X_t'\beta_0 = u)$, and assume that $\mathcal{V}_t(\underline{x}_t'\beta_0) \subsetneq \mathcal{V} \equiv \text{supp}(V)$, where $\subsetneq$ is used to denote a proper subset. Then, the conditional expectation $\mathbb{E}[Y_t|X_t'\beta_0 = \underline{x}_t'\beta_0, V = v]$ is identified for all $v \in \mathcal{V}_t(\underline{x}_t'\beta_0)$. Therefore, building on Theorem 4 in ImbensNewey2009, we obtain the following sharp bounds on the ASF.

theoremLet $t \in \{1,\ldots,T\}$, $\underline{x}_t \in \text{supp}(X_t)$, and let Assumptions A(ref)--A(ref) hold. Also, let $g_t(\underline{x}_t'\beta_0,C,U_t) \in [\underline{g},\overline{g}]$ with probability 1.\footnote{The point identification of $\beta_0$ could require additional assumptions, which in turn may further sharpen the bounds. We do not incorporate these potential additional assumptions as they are model-specific and beyond the current general setup of the paper. Also, we allow for unrestricted $g_t$ with only boundedness required. Imposing additional conditions on $g_t$, such as monotonicity and parametric distribution, could lead to narrower bounds.} Then, the identified set for $\text{ASF}_t(\underline{x}_t)$ is \begin{align} \mathcal{A}_t(x_t) &= \left[\int_{\mathcal{V}_t(x_t'\beta_0)} \mathbb{E}[Y_t|X_t'\beta_0 = x_t'\beta_0, V = v] \, dF_{V}(v) + g \cdot \mathbb{P}(V \in \mathcal{V} \setminus \mathcal{V}_t(x_t'\beta_0)),\right.\notag\\ &\left. \qquad \qquad \int_{\mathcal{V}_t(x_t'\beta_0)} \mathbb{E}[Y_t|X_t'\beta_0 = \underline{x}_t'\beta_0, V = v] \, dF_{V}(v) + \overline{g}\cdot\mathbb{P}(V \in \mathcal{V} \setminus \mathcal{V}_t(\underline{x}_t'\beta_0))\right]. \end{align}

In the case where $Y_t$ is binary, $[\underline{g},\overline{g}] = [0,1]$ and the bounds take on a simpler form. The width of these bounds depends only on $\overline{g} - \underline{g}$ and the probability that $V$ falls outside of $\mathcal{V}_t(\underline{x}_t'\beta_0)$, which is small if $V$ is continuously distributed and the measure of $\mathcal{V} \setminus \mathcal{V}_t(\underline{x}_t'\beta_0)$ is close to zero. Hence, small violations of the support condition yield a narrow identified set.

For fixed effects as in Section (ref), $C|\mathbf{X}$ is unrestricted, and equivalently, $V=\mathbf{X}$. The support condition fails because $\mathcal{V}=\text{supp}(\mathbf{X})\neq\text{supp}(\mathbf{X}|X_t'\beta_0 = \underline{x}_t'\beta_0)=\mathcal{V}_t(\underline{x}_t'\beta_0),$ and thus the ASF cannot be point identified in general. The following corollary characterizes the sharp bounds on the ASF under fixed effects.

corollaryLet the conditions in Theorem (ref) hold. Under fixed effects, i.e., \(V=\mathbf{X} \), the identified set for $\text{ASF}_t(\underline{x}_t)$ is \begin{align} & \left[\int_{supp(\mathbf{X}|X_t'\beta_0 = x_t'\beta_0)} \mathbb{E}[Y_t|\mathbf{X} = \mathbf{x}] \, dF_{\mathbf{X}}(\mathbf{x}) + g \cdot \mathbb{P}(X_t'\beta_0 \neq x_t'\beta_0),\right.\notag\\ &\left. \qquad \qquad \int_{supp(\mathbf{X}|X_t'\beta_0 = \underline{x}_t'\beta_0)} \mathbb{E}[Y_t|\mathbf{X} = \mathbf{x}] \, dF_{\mathbf{X}}(\mathbf{x}) + \overline{g}\cdot\mathbb{P}(X_t'\beta_0 \neq \underline{x}_t'\beta_0)\right]. \end{align}

Comparing the bounds in Theorem (ref) and Corollary (ref), we see that with index sufficiency but no support conditions, the bounds are as specified in equation (ref); when we further drop the index sufficiency condition, these bounds become even wider as in equation (ref). For example, the identified set in (ref) equals the trivial set $[\underline{g},\overline{g}]$ whenever $X_t'\beta_0$ is continuously distributed: these bounds are informative only when $\mathbb{P}(X_t'\beta_0 = \underline{x}_t'\beta_0) > 0$. This implies that ASF bounds under fixed effects can be uninformative even if $\beta_0$ is point identified. In the end, the difference between the point identification and the identified set in (ref) shows the importance of the support condition, and the difference between the identified sets in (ref) and (ref) highlights the role of the index sufficiency condition.

The ASF bounds can be used to construct sharp bounds on the APE with discrete covariates. Let $\underline{\text{ASF}}_t(\underline{x}_t)$ and $\overline{\text{ASF}}_t(\underline{x}_t)$ denote the lower and upper bounds of $\mathcal{A}_t(\underline{x}_t)$, we have

align*[align* omitted — 278 chars of source]

To obtain bounds on the APE for a continuous covariate, bounds on $\frac{\partial}{\partial a}\mathbb{E}[g_t(a,C,U_t)|C=c]$ are needed. In the case where $\frac{\partial}{\partial a}\mathbb{E}[g_t(a,C,U_t)|C=c] \in [\underline{g}',\overline{g}']$ and $\beta_0^{(k)}>0$, the APE bounds are given by\footnote{We leave a formal proof of the APE bounds' sharpness for future work.}

align*[align* omitted — 795 chars of source]

The bounds are reversed if $\beta_0^{(k)}<0$. With binary outcomes and under the assumption that $U_t \perp \! \! \! \perp C$, the partial derivative $\frac{\partial}{\partial a}\mathbb{E}[g_t(a,C,U_t)|C=c]$ is the density $f_{U_t}$. Its lower bound is trivially 0, but it is harder to postulate an upper bound when $f_{U_t}$ is unrestricted. In the specific case of logit models, however, this density attains a maximal value of 1/4 at the origin. Therefore, substituting $[\underline{g}',\overline{g}'] = [0,1/4]$ yields bounds on the APE in binary panel logit when common support fails.

Bounds for the LAR and AME can be similarly obtained. The estimation methods we provide below for the point identified ASF, APE, or AME can be adapted to estimate these bounds under support condition violations.

Estimation

In this section we briefly propose estimators for the ASF, APE, LAR, and AME. Detailed asymptotic results can be found in Supplemental Appendix (ref). The estimators we construct are sample analogs of (ref) for the ASF, and of (ref) for the APE. We assume throughout that we have access to a random sample $\{(Y_i,\mathbf{X}_i)\}_{i=1}^N$.

All the partial effects estimators are obtained in three steps. We first obtain a consistent estimator $\widehat{\beta}$ of the common parameters using one of the many approaches proposed in the related literature. In the second step, we nonparametrically estimate the conditional expectation $h(u,v) \equiv \mathbb{E}[Y_t|X_t'\beta_0 = u, V = v]$ using a local polynomial regression of $Y_t$ on generated regressor $X_t'\widehat{\beta}$ and $V$. Local polynomial regression naturally yields estimates of the derivatives of $h$, which are used in estimating the APE, LAR, and AME. In the final step, we average features of the local polynomial regression estimates over the empirical distribution of $V_i$, i.e., a partial mean structure.

For example, the ASF estimator is constructed analogously to equation (ref) above. We average this conditional mean over the empirical marginal distribution of $V_i$ to obtain the ASF estimator:

align*[align* omitted — 168 chars of source]

where $\hat\pi_{it}$ is a trimming function that regularizes the behavior of the estimator. See Supplemental Appendix (ref) for more details on this function. The APE is obtained similarly, replacing the estimated $h$ function by the estimate of its partial derivative with respect to its first component, denoted as $h_u$:

align*[align* omitted — 202 chars of source]

Finally, analogously to equations (ref) and (ref), we can define estimators for the LAR and AME:

align*[align* omitted — 444 chars of source]

In Supplemental Appendix (ref), we provide conditions on the convergence rate of $\widehat{\beta}$ and the bandwidth, as well as on the order of the local polynomial regression, that allow us to establish the consistency and asymptotic normality of the ASF and APE estimators. We leave a complete asymptotic analysis of the AME estimator for future work. It is worth noting that these estimators exhibit fast convergence rates compared to other nonparametric estimators, which can be attributed to the fact that upon integrating over the distribution of $v(\mathbf{X}_i)$, the ASF and APE are essentially functions of one-dimensional $X_{it}'\beta_0$. Their convergence rates do not depend on the dimension of $\mathbf{X}_i$, which is $T \times d_X$. For example, when the index is one-dimensional, the ASF's and APE's rates of convergence are similar to the standard rates of convergence of univariate nonparametric kernel regression estimators, which are fast within the class of nonparametric estimators. In particular, we show that the ASF can converge at a rate faster than $N^{2/5}$ when the index is one-dimensional.

Empirical Illustration

In this section we compare the performance of our proposed estimators with that of commonly used alternative estimators.

Alternative Estimators

The alternative estimators we consider are a parametric random effects (RE) and a correlated random effects (CRE) estimator. See, for example, Wooldridge2010. Both assume a standard logistic distribution for the error term $U_t$, but differ in their specifications for the distribution of individual effects $C$. For the RE, \[ C \sim \mathcal{N}(\mu_c,\sigma_c^2) \] and is independent of $V$. For the CRE, \[ C|V\sim \mathcal{N}(\mu_{c0}+\mu_{c1}'V,\sigma_c^2). \] Then, the CRE is equivalent to an augmented RE with $V$ being additional regressors. As is standard in the literature, we use the Maximum Likelihood Estimator (MLE) to jointly estimate $\beta_0$ and the distribution parameters $(\mu_c,\sigma_c^2)$ or $(\mu_{c0},\mu_{c1},\sigma_c^2)$.

In the same spirit as the semiparametric estimator in Section (ref), we allow the marginal distribution of $V$ to be unrestricted. The conditional expectation of the binary outcome and its derivative are calculated based on the MLE estimates, and the ASF and APE are obtained by averaging out $V$.

Background and Specification

In this empirical illustration, we examine women's participation in the labor market using our semiparametric approach. See the handbook chapter by killingsworth1986female for an extensive review of the literature on female labor supply. For illustrative purposes, our analysis is based on the static setup of Fernandez-Val2009, where covariates $X_{t}$ include numbers of children in three age categories, log husband's income, a quadratic function of age, as well as time dummies.\footnote{Fernandez-Val2009 proposes bias-corrected estimators of marginal effects when $T$ is large and when $U_{t}$ follows a normal distribution. CharlierMelenbergSoest1995 and ChenSiZhangZhou2017, among others, also considered female labor force participation in their empirical applications. They used similar model specifications, but most of these papers focused on the estimation of common parameters $\beta_0$ instead of the ASF or APE.}

The sample consists of $N=1461$ married women observed for $T=9$ years from the PSID between 1980--1988. We use the dataset kindly made available on Iv{\'a}n Fern{\'a}ndez-Val's website, originally sourced from Jes{\'u}s Carro. In the Supplemental Appendix, we plot the distributions of the covariates in Figure (ref), and summarize the corresponding descriptive statistics in Table (ref). Roughly 45% of the women in the sample always participated in the labor market, less than 10% never participated, and around 45% changed their status during the sample period. Movers tended to be younger and have more children in all children's age categories. Never participants were relatively uniformly distributed between ages 30 and 50, whereas the women in other subgroups were generally younger. All subgroups exhibited heavy tails in log husband's income.

The unobserved individual effects $C$ could be interpreted as an individual's willingness to work. In the benchmark specification, we construct indices $V$ based on the initial values of the covariates \(X_{i1}\). Women's ages and numbers of children are discrete variables, and we consider a cell-by-cell analysis.\footnote{For a more comprehensive empirical analysis, one could handle discrete index variables using a discrete kernel as suggested in LiRacine2004, which would be outside the scope of the current empirical illustration.} These covariates generate over 1000 cells in this sample, and some cells do not contain sufficient observations to use a semiparametric estimator within them. Therefore, we collapse the discrete index variables as follows. First, we sum the number of children under 18 in the initial period, then categorize this number into a low, medium, or high group based on the 33rd and 67th percentiles. Similarly, we collapse the initial age into a binary indicator based on the median. This coarsening scheme results in 6 cells, with a range of 156 to 314 observations per cell. Thus, we have three index variables: a trinary fertility variable, a binary age variable, and a continuously distributed average log husband's income. The number of continuous index variables is $d_V=1$.

Various robustness checks regarding alternative choices of \(V_i\) (e.g., constructed from \(X_{i1}\) or \(\overline X_i=\frac 1 T \sum_t X_{it}\)), alternative estimators (e.g., multiple indices and local logit), and alternative coarsening schemes are explored in Supplemental Appendix (ref) as well as the previous version of this paper liu2021identification. The semiparametric estimator is generally robust with respect to these variations.

Results

Table (ref) reports the estimated common coefficients on key covariates. Figure (ref) in the Supplemental Appendix also plots the estimated coefficients on time dummies, which capture the time-variation in aggregate participation rates. We see that women are more inclined to withdraw from the labor force when they have more children, especially younger ones, and when their husbands earn a higher income. Compared to the RE and CRE, the flexible smoothed maximum score estimator provides slightly larger (in magnitude) estimates with larger standard errors.

table[table omitted — 2,520 chars of source]

In our empirical example, we focus on the effects of the husband's income, which may affect the wife's reservation wage. We select evaluation points $\underline{x}_t$ such that the log husband's income ranges from its 20th to 80th quantiles, and other variables are equal to their medians. These choices correspond to a hypothetical woman who is 35 years old, has 0 children between 0 and 2, 0 children between 3 and 5, 1 child between 6 and 17, and whose husband's income ranges from \$21K to \$55K. All time dummies are set to zero in this counterfactual.

figure[figure omitted — 680 chars of source]

Figure (ref) shows estimates of the ASF and APE across $\underline{x}_t$ together with the 90% bootstrap confidence intervals based on 500 bootstrap samples.\footnote{For the ASF, all bootstrap estimates are between 0 and 1, and so is the symmetric percentile-$t$ confidence band based on bootstrap standard deviations. For the APE, the smoothed maximum score in the first step requires monotonicity, i.e., $\frac{d}{d u}\mathbb{P}(Y_t = 1|X_t'\beta_0 = u)|_{u = \underline{x}_t'\beta_0} \ge 0$. In the bootstrap, this constraint occasionally binds, so we censor it at zero and employ the percentile bootstrap to account for the possible non-standard distribution due to censoring. Note that in principle, the bootstrap band for the APE could still contain positive values since the estimated coefficient for the log husband's income could be positive in some bootstrap samples. However, this incidence is very rare in our empirical example.} For the ASF, all point estimates are downward sloping with respect to the husband's income. The semiparametric estimator yields slightly higher participation probabilities compared to the RE and CRE.

For the APE, the semiparametric estimates are closer to zero for lower husband's incomes and more negative for higher ones, while their RE and CRE counterparts are rather flat. Note that for continuous $\underline{x}_t^{(k)}$, \[\text{APE}_{k,t}(\underline{x}_t)=\beta_0^{(k)}\cdot f_{U_t-C}\left(\underline{x}_t'\beta_0\right)=\beta_0^{(k)} \cdot \int_{\mathcal{C}} f_{U_t}(\underline{x}_t'\beta_0 + c) \, dF_C(c),\] where $f_{U_t-C}$ denotes the pdf of $U_{t}-C$, i.e., a convolution of $-C$ and $U_{t}$. Thus, the slope of the APE with respect to $\underline{x}_t^{(k)}$ reflects the shapes of $f_C$ and $f_{U_t}$ as well as the magnitude of $\beta_0^{(k)}$. In this sense, the flatter APE profile with respect to the husband's incomes in the RE and CRE could be due to the following three sources: (i) The RE and CRE feature a Gaussian $f_{C|V}$ with estimated mean and variance. The estimated Gaussian variance could be fairly large to accommodate potential non-Gaussian heterogeneity in $C|V$, and the resulting $\widehat f_{U_t-C}$ could be flatter (around the peak) than the true distribution. (ii) The RE and CRE assume a logistic $f_{U_t}$, which may deviate from the true data generating process. (iii) The smaller magnitudes of $\widehat\beta_0^{(k)}$ for RE and CRE could be due to misspecification of the distributions of $U_t$ and $C$ and, in turn, further lead to a milder slope of the APE profile, though the discrepancy in the coefficients alone does not fully account for the differences in the slopes of the APE profiles. In contrast, the semiparametric estimator does not require the parametrization of $f_{C|V}$ or $f_{U_t}$, thus reducing potential biases due to misspecification.

Moreover, our flexible semiparametric estimator, which does not impose parametric assumptions on the distributions of $C$ or $U_t$, yields statistically insignificant APEs with respect to husband's incomes. This finding raises concerns about the potential bias in the highly significant APEs by RE and CRE estimators due to their parametric restrictions. The insignificant APEs are consistent with the empirical observation that married women's labor supply choices became less sensitive to their husbands' income around 1980 when baby boomers started constituting a larger portion of the labor force, and both partners contribute to housework and earnings more equally. Hence fewer married women were at the margin of labor force participation that could be nudged by temporary fluctuations in husbands' income.

Conclusion

The distributions of the unobserved individual heterogeneity and the idiosyncratic errors play a crucial role in identifying partial effects in nonlinear panel models. In this paper, we first show the identification of the ASF, APE, and AME in a nonlinear semiparametric panel model with potentially unspecified distributions of the unobserved heterogeneity and idiosyncratic errors. To achieve point identification, we assume that units with the same value of the index $V$ have correspondingly similar distributions of their unobserved heterogeneity $C$. We also establish partial identification results when the support condition fails. We then develop three-step semiparametric estimators for the ASF and APE, and show their consistency and asymptotic normality in the Supplemental Appendix. Finally, we illustrate our semiparametric estimator in a study of determinants of women's labor supply.

In Supplemental Appendix (ref) we provided an identification result that applies to dynamic panel models. Extending our results to a broader class of dynamic models would be an interesting direction for further research.

singlespace