Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
63,877 characters · 17 sections · 80 citation commands
Identification of time-varying counterfactual parameters in nonlinear panel models
We study counterfactual or policy parameters for nonlinear panel models with structural equation
where $i$ indexes individual units, $t$ indexes time, $h_{t}$ is a weakly-monotone transformation function that can vary over $t$ in an unrestricted way, $\beta\in\mathbb{R}^{k}$ is a vector of regression coefficients, $\alpha_{i}$ is an individual-specific effect, and $U_{it}$ is a stochastic error term.\footnote{There is a large literature on the identification and estimation of structural parameters in the cross-sectional version of this model, e.g., see the work on single index models by Han1987, PowellStockStoker, Ichimura, Ahn2018, and references therein.} For each unit $i$, we observe covariates $X_{it}$, $t=1,\dots,T$. The dependence between $\alpha_{i}$ and $X_{i}=\left(X_{i1},\dots,X_{iT}\right)$ is left unrestricted, so that $\alpha_{i}$ is a fixed effect, c.f., e.g., GrahamPowell2012. The class of models with outcome equation as in ((ref)) includes the binary choice model with two-way fixed effects, the ordered choice model with fixed effects and time-varying cut-offs, the censored regression model with time-varying censoring, and various transformation models for continuous dependent variables.\footnote{For example, letting $v\equiv\alpha_{i}+x\beta-U_{it}$, the structural equation for the binary choice model with two-way fixed effects is $h_{t}\left(v\right)=1\left\{ v\geq\lambda_{t}\right\} $, for the ordered choice model with time-varying cutoffs it is $h_{t}\left(v\right)=\sum_{j=1}^{J}1\left\{ v\geq\lambda_{jt}\right\} $, $J\in\mathbb{N}$, and for censored regression with time-varying censoring it is $h_{t}\left(v\right)=\max\left\{ \lambda_{t},v\right\} $.}
For a subpopulation of individuals defined by their sequence of regressor values $X_{i}$, our parameter of interest is the counterfactual survival probability:\footnote{The counterfactual survival probability answers the question “For a subpopulation defined by their sequence of regressor values $X_{i}$, what is the ceteris paribus probability that their period-$t$ outcome $Y_{it}$ exceeds $y$ if their period-$t$ regressor values were exogenously set to $x$?”}
where $Y_{it}\left(x\right)$ is given by ((ref)), $x$ is a fixed counterfactual value of the period-$t$ regressors, and $y\in\underline{\mathcal{Y}}\equiv\mathcal{Y}\setminus\text{inf}\mathcal{Y}$ is a fixed cut-off value, where $\mathcal{Y}\subseteq\mathbb{R}$ denotes the support of the observed $Y_{it}=Y_{it}\left(X_{it}\right)$.\footnote{Conditioning on the sequence $X_{i}$ allows us to identify the same parameter across different exogeneity and time-stationarity assumptions.} The parameter in ((ref)) is a building block for most partial and marginal effects of interest in applied practice that are based on the average structural function (ASF) as defined in the pioneering work of BlundellPowell2003,BlundellPowell2004. For example, when $Y_{it}\left(x\right)$ is non-negative , the ASF at time $t$ can be obtained as:\footnote{See, for example, SongWang for the integrated tail probability expectation formula that uses a survival function as defined in ((ref)).}
The ASF can then be used to define partial effects based on partial derivatives or marginal effects based on discrete differences, see, e.g., LinWooldridge2015.
The challenge is to identify ((ref)) in nonlinear panel models with structural equation as in ((ref)) when $\alpha_{i}$ are fixed effects and $T<\infty$. To see that this is challenging, consider the special case of the binary choice model with two-way fixed effects. A recent literature has made progress in identifying certain counterfactual parameters for this model provided that the error terms follow a standard logistic distribution, see e.g. aguirregabiriaidentification, davezies2021identification, dobronyi2021identification.\footnote{Earlier work by honore2006boundson provides partial identification of marginal effects under a more general structure with dynamics and arbitrary but known error term distributions. See also PakelWeidner for an approach that applies to parametric models covered by the results in bonhomme2012functional. Finally, see honore_censored_marginal for results on marginal effects for the censored regression model.}
A separate literature provides identification results under time-homogeneity assumptions that do not allow for arbitrary time-effects, see the benchmark results in hoderlein2012nonparametric, chernozhukov2013average, and chernozhukov2015nonparametric.\footnote{There, the authors consider a nonseparable structural function and impose no parametric assumptions on the error terms. When $h_{t}=h$, the class we models we study here is nested in their analysis.} To the best of our knowledge, nothing is known about the identification of the ASF for the binary choice model with two way fixed effects without logistic errors.\footnote{We do not consider here the case of correlated random effects or the case of large-$T$. Progress on counterfactual parameters for the former case has been made by, e.g., ArellanoCarrasco2003, altonji2005crosssection, bester2009identification, ChenKhanTang2019, liu2021identification, while for the latter by, e.g., FernandezVal2009, FernandezValWeidner2018, and bartolucci_partial_effect2023.}
We derive (partial) identification results for ((ref)) without parametric restrictions on the distribution of $U_{it}$ for panel models with outcome equation as in ((ref)). Our results are for short-$T$. Relevant examples of models to which our results apply are (i) binary choice with two-way fixed effects and nonlogistic errors, (ii) ordered choice with time-varying thresholds, (iii) censored regression with time-varying censoring. Additionally, since ((ref)) varies over time whenever $h_{t}$ is time-varying, policy parameters that are functionals of ((ref)) are also time-varying.
For nonseparable panel models with fixed effects, the results in benchmark work such as hoderlein2012nonparametric, chernozhukov2013average, and chernozhukov2015nonparametric establish limitations on what can be learned from panel data in terms of counterfactual parameters. By imposing additional structure on the latent outcome, such as additivity in a linear index and the fixed effects, we show that (partial) identification of counterfactual parameters can be obtained without time-homogeneity assumptions on the outcome equation and no parametric assumptions on the distribution of the error terms. Because we make no time-homogeneity assumptions on $h_{t}$, the counterfactual parameters can vary over time in an arbitrary way. To the best of our knowledge, it is the combination of time-varyingness of the outcome equation (hence, of the counterfactual parameters) and no parametric distributional assumptions on the error terms that constitutes our contribution relative to the literature on partial effects in nonlinear panel models with fixed effects.
For our results, we treat $\beta$ and $h_{t}$ as given, i.e. either known or previously point- or partially-identified.\footnote{Sufficient conditions for the identification of $\beta$ and time-varying $h_{t}$ for models with structural equation as in ((ref)) are provided in BotosaruMuris2017 and BotosaruMurisPendakur2021 under strict exogeneity and weak monotonicity of $h_{t}$, and BotosaruMurisSokullu2022 under endogeneity and strict invertibility of $h_{t}$. With a parametric structure on both $h_{t}$ and the distribution of the stochastic errors, one may use the results in bonhomme2012functional. Consistent estimators for $\beta$ or/and time-invariant transformation $h_{t}=h$ in nonlinear panel models without parametric assumptions on the error terms have been derived by, e.g., abrevaya1999leapfrog, Chen2010, ChenWang2018, WangChen2020, ChenLuWang2022 and references therein. For specific panel models, such as binary choice, ordered choice, linear models, duration models, censored regression, see, e.g., manski1987semiparametric, Honore1992, Honore1993, HorowitzLee2004, ChenDahlKahn, Lee2008, muris2017estimation. Recent work on sharp identification regions for structural parameters includes KhanPonomarevaTamer2011, KhanPonomarevaTamer2016, HonoreHu2020, KhanPonomarevaTamer2021, and aristodemou2021semiparametric. See also Ghanem2017 on testing identifying assumptions in the class of models we consider here.} Our key insight is that the linear index structure allows us to classify each observed probability at time $s$,
as either an upper bound on $\tau_{t,x,y}\left(X_{i}\right)$, a lower bound on $\tau_{t,x,y}\left(X_{i}\right)$, or both. We show that ((ref)) is an upper bound only if the observed index at time $s$, $X_{is}\beta-h_{s}^{-}\left(y\right)$, is at least the counterfactual index at time $t$, $x\beta-h_{t}^{-}\left(y\right)$; and a lower bound only if the observed index at time $s$ is at most the counterfactual index at time $t$. Without any additional time-stationarity or exogeneity assumptions, bounds on the period-$t$ counterfactual probability $\tau_{t,x,y}\left(X_{i}\right)$ can be constructed from the period-$t$ observed probabilities $P\left(\left.Y_{it}\geq y^{\prime}\right|X_{i}\right)$, $y^{\prime}\in\underline{\mathcal{Y}}$. Under a conditional time-stationarity assumption on the errors, outcomes from all periods are informative for the period-$t$ counterfactual probability. The bounds under conditional time-stationarity are tighter than those without exogeneity assumptions. Point identification is obtained when the transformation function $h_{t}$ is invertible or when the counterfactual index equals one of the observed indices.
The remainder of this paper is organized as follows. Section (ref) presents our main results: Section (ref) constructs bounds without any additional time-stationarity or exogeneity assumptions, while Section (ref) constructs bounds under a conditional time-stationarity assumption on the error terms. Section (ref) applies our results to a few examples: binary choice model, ordered choice model, and censored regression. We present a numerical experiment for the binary choice model in Section (ref) and one for the ordered choice model in Section (ref). All proofs and an additional numerical experiment can be found in the Appendix.
We provide two results on the identification of $\tau_{t,x,y}$ defined in ((ref)). Our first result in Theorem (ref) provides bounds on $\tau_{t,x,y}$ without imposing any exogeneity assumptions or time-stationarity assumptions on $U_{it},\,X_{i}$. Our second result in Theorem (ref) uses a conditional time-stationarity assumption on $\left.U_{it}\right|\alpha_{i},X_{i}$ to tighten those bounds. Conditioning on the sequence $X_{i}$ allows us to identify the same parameter, $\tau_{t,x,y}$, across different exogeneity and time-stationarity assumptions.
The following two assumptions are maintained throughout the paper:
Footnote (ref) lists work that provides sufficient assumptions for either the point- or the partial-identification of $\beta$ and $h_{t}$.
Assumption (ref) does not restrict the way that $h_{t}$ can vary over $t$. Our results apply to the case of time-invariant $h_{t}=h$, in which case parameters such as ((ref)) and ((ref)) are also time-invariant. This assumption allows $h_{t}$ to have flat parts and jumps, or to be continuous. Hence, our setting accommodates both discrete and continuous outcomes.
Given Assumption (ref), we define the generalized inverse of $h_{t}$ as:\footnote{Existence of $h_{t}^{-}$ is ensured by Assumption (ref).}
i.e. it is the smallest value of the latent variable that yields a value of the observed outcome $Y_{t}\geq y$.
For what follows we fix $t,x$ and we fix $y\in\underline{\mathcal{Y}}$.
For our first set of results, we define the following sets:
The set $\mathcal{Y}_{L}$ collects the values $y^{\prime}\in\underline{\mathcal{Y}}$ for which the counterfactual index at time $t$, $x\beta-h_{t}^{-}\left(y\right)$, is greater than the observed index at time $t$, $X_{it}\beta-h_{t}^{-}\left(y^{\prime}\right)$, while $\mathcal{Y}_{U}$ collects the values $y^{\prime}\in\underline{\mathcal{Y}}$ for which the counterfactual index at time $t$ is smaller than the observed index at time $t$. Note that $\mathcal{Y}_{L}\cup\mathcal{Y}_{U}=\underline{\mathcal{Y}}$.
Theorem (ref) uses the linear-index structure of ((ref)) and knowledge of $\beta$ and $h_{t}$ to classify $P\left(\left.Y_{it}\geq y^{\prime}\right|X_{i}\right)$ as a lower (upper) bound on $\tau_{t,x,y}\left(X_{i}\right)$ for any value of $y^{\prime}\in\underline{\mathcal{Y}}$. By varying $y^{\prime}\in\underline{\mathcal{Y}},$ we obtain the sets $\mathcal{Y}_{L}$ ($\mathcal{Y}_{U})$ of values $y^{\prime}$ that provide lower (upper) bounds. Equation ((ref)) intersects these bounds. Since $\mathcal{Y}_{L}\cup\mathcal{Y}_{U}=\underline{\mathcal{Y}}$, every value of $y^{\prime}$ provides an upper bound, a lower bound, or both. Note that since either $\mathcal{Y}_{L}$ or $\mathcal{Y}_{U}$ can be empty, the trivial lower (upper) bound is selected by the intersection with $\left[0,1\right]$.
Section (ref) applies Theorem (ref) to binary choice, ordered choice, and censored regression. In the case of binary choice, the bounds of Theorem (ref) simplify significantly. In that case, $P\left(\left.Y_{it}\geq1\right|X_{i}\right)$ provides an upper bound for $P\left(\left.Y_{it}\left(x\right)\geq1\right|X_{i}\right)$ if $X_{it}\beta>x\beta$, and a lower bound if $X_{it}\beta<x\beta$. When $X_{it}\beta=x\beta$, $P\left(\left.Y_{it}\geq1\right|X_{i}\right)$ point-identifies $P\left(\left.Y_{it}\left(x\right)\geq1\right|X_{i}\right)$.
The bounds in Theorem (ref) can be wide, for example when the dependent variable has few points of support (see Remark (ref)). The following assumption obtains tighter bounds by using information across all time periods rather than information from period $t$ only.
Conditional time-stationarity requires that the conditional distribution of the error terms conditional on $\alpha_{i},X_{i}$ be the same in each time period. In chernozhukov2013average,chernozhukov2015nonparametric and hoderlein2012nonparametric\footnote{See, e.g., manski1987semiparametric, Honore1992, Abrevaya2000, hoderlein2012nonparametric, GrahamPowell2012, chernozhukov2013average, chernozhukov2015nonparametric, KhanPonomarevaTamer2016, ChenKhanTang2019, KhanPonomarevaTamer2021.} the authors refer to this assumption as both “strict exogeneity” and time-homogeneity of the error terms.\footnote{For a discussion of strict exogeneity, as well as other notions of exogeneity, in the context of linear models, see Chamberlain1984, ArellanoHonore2001, ArellanoBonhomme2011.} Note that the error terms are required to have a time-stationary (“time-homogeneous”) distribution conditional on $\alpha_{i}$ and the entire sequence $X_{i1},\dots,X_{iT}$. The assumption allows for serial correlation in the errors $U_{it}$ and in some components of $X_{i}$, and it leaves the distribution of $\alpha_{i}$ conditional on $X_{i}$ unrestricted. Note that, in our set-up with time-varying structural equation, the conditional distribution of $Y_{it}$ given $X_{i}$ can still vary over time.
Fix $t,x,$$y\in\underline{\mathcal{Y}}$ and define the sets:
As in the case of in Theorem (ref), the result above applies directly to point-identified $\beta$ and $h_{t}$. If $\left(\beta,h_{t}\right)$ are partially identified, bounds on $\tau_{t,x,y}\left(X_{i}\right)$ can be obtained by taking the worst-case bounds across parameter values in the identified set.
According to Theorem (ref), we can use information from any period $s$ to construct bounds on the counterfactual survival distribution in period $t$. The resulting bounds can be much more informative than those in Theorem (ref) without time-stationarity and exogeneity assumptions; this can be seen from the expressions for binary and ordered choice in Section (ref), and from the numerical experiments in Section (ref). The gains can be substantial, especially if the number of time periods is large, if there is variation in the values of the sequence $X_{i}$, and if there is a large degree of variation over time in $h_{t}$.\footnote{Although, we do not prove sharpness of the bounds in Theorem (ref), we were unable to construct a situation where they were not. As we show in the numerical exercises, the bounds are (very) informative.}
In this section, we apply our results to a number of empirically relevant choices for $h_{t}$. These examples are helpful in understanding how informative our bounds are, and help relate them to existing bounds in the literature derived under related but different conditions.
Let $v\equiv\alpha_{i}+X_{it}\beta-U_{it}$. We start by applying our results to the fixed effects linear transformation model with invertible $h_{t}$. We then study binary choice models with two-way fixed effects $h_{t}\left(v\right)=1\left\{ v\geq\lambda_{t}\right\} $, ordered choice models with time-varying thresholds $h_{t}\left(v\right)=\sum_{j=1}^{J}1\left\{ v\geq\lambda_{jt}\right\} $, $J\in\mathbb{N},$ and censored regression with $h_{t}\left(v\right)=\max\left\{ \lambda_{t},v\right\} $.
Let $h_{t}$ be invertible, so that $h_{t}^{-}=h_{t}^{-1}$. Examples include linear regression with two-way fixed effects, i.e. $h_{t}\left(v\right)=v+\lambda_{t}$; transformation models used in duration analysis; and the Box-Cox transformation model.\footnote{Such models have been studied extensively in the cross-sectional setting, see, e.g., amemiyaComparisonBoxCoxMaximum1981,barnettNonparametricSemiparametricMethods1991,powellRescaledMethodsofmomentsEstimation1996.}
Let $h_{t}$ be defined on $\mathbb{R},$or $v\in\mathbb{R}.$ There always exists a $y^{\prime}$ such that
Theorem (ref) applies and the counterfactual survival probability ((ref)) is point-identified. To see this, consider the argument below:\footnote{See also Remark 2 in Botosaru and Muris (2017), and Botosaru et al. (2021).}
where invertibility of $h_{t}$ was used in the second equality.
We are not aware of results for $\tau_{t,x,y}$ -- or derived quantities such as partial/marginal effects -- for the case considered here, other than those in BotosaruMuris2017,BotosaruMurisPendakur2021,BotosaruMurisSokullu2022.
The link function \[ h_{t}\left(v\right)=1\left\{ v-\lambda_{t}\geq0\right\} \] obtains the panel binary choice model with two-way fixed effects with structural function
Our results yield bounds on ((ref)), hence on partial effects, without parametric assumptions on the distribution of the error terms and for a variety of exogeneity conditions. The only existing results without parametric assumptions on the error term that we are aware of are those in chernozhukov2013average, but those require that regressors be discrete and that $\lambda_{t}=0$ for all $t$.
In what follows, we fix the time period for the counterfactual to $t=1$ and use the abbreviated notation
For this model, $\underline{\mathcal{Y}}=\left\{ 1\right\} $ and $h_{t}^{-}\left(1\right)=\lambda_{t}$.
For $y^{\prime}=1=y,$ Theorem (ref) leads to three different cases depending on the sign of $X_{i1}\beta-x\beta$. These cases are:
The results of Theorem (ref) can be summarized as follows: \[ \tau_{x}\left(X_{i}\right)\in
\]
To see that these bounds are valid, we adapt the key derivation underlying Theorem (ref) to this specific case:
where the third line denotes that the direction of the inequality depends on the sign of the difference between the observed index $X_{i1}\beta$ and the counterfactual index $x\beta$.
Under Assumption (ref), Theorem (ref) implies that any period can be used to construct counterfactuals for period $1$. Instead of stating the bounds as implied by Theorem (ref), we show how to construct them from first principles.
Suppose that there exists a time period $s$ such that
Then ((ref)) is point identified with
where the second equality follows by Assumption ((ref)) and the third equality follows by ((ref)).\footnote{As a special case, in a model without time dummies and with a binary treatment indicator, e.g., $X_{it}\in\left\{ 0,1\right\} $, we can point-identify the distribution under treatment $x=1$ for any subpopulation that is treated at some point, $\exists t:X_{it}=1$.}
If there does not exist a time period such that ((ref)) holds, Theorem (ref) can be operationalized as follows. Fix $y^{\prime}=1$ and, for each period $s\in\left\{ 1,\dots,T\right\} $, compare $X_{is}\beta-\lambda_{s}$ to $x\beta-\lambda_{1}$, and group the time periods according to whether they provide an upper bound $\left(s\in\mathcal{T}_{U}\right)$ or a lower bound $\left(s\in\mathcal{T}_{L}\right)$:
These sets correspond to $\mathcal{U}$ and $\mathcal{L}$ in Theorem (ref) with $y^{\prime}=1=y$, counterfactual period $t=1$, and $h_{t}^{-}\left(1\right)=\lambda_{t}.$
The best lower (upper) bound on ((ref)) is constructed using periods in $\mathcal{T}_{L}$ ($\mathcal{T}_{U}$):
or \[ \tau_{x}\left(X_{i}\right)\in
\]
These bounds have a few interesting properties. First, they can be informative for stayers, i.e. even when $X_{it}=x$ for all $t$, nontrivial bounds can be derived as long as there are time effects $\lambda_{t}\neq\lambda_{1}$ for some $t$. Second, under Assumption (ref), the bounds tighten as compared to the bounds without this assumption. In Section (ref), we investigate these and other properties through a numerical experiment.
The link function \[ h_{t}\left(v\right)=\sum_{j=1}^{J}1\left\{ v\geq\lambda_{jt}\right\} ,J\in\mathbb{N}, \] obtains an ordered choice model with fixed effects and time-varying thresholds:\footnote{dasPanelDataModel1999,johnsonphdthesis,baetschmannIdentificationEstimationThresholds2012,baetschmann2012identification,muris2017estimation,BotosaruMurisPendakur2021 discuss identification of the $\left(\beta,\lambda_{tj}\right)$ under various conditions. With logistic errors, the parameters can be estimated via composite conditional maximum likelihood estimation. Without logistic errors, the parameters can be estimated using maximum score methods. In the logistic case, the results in davezies2021identification can then be used to bound average marginal effects.}
For this model, $\mathcal{\underline{Y}}=\left\{ 2,\dots,J\right\} $ and $h_{t}^{-}\left(y\right)=\lambda_{ty},\,y\in\underline{\mathcal{Y}}.$
We fix the time period for the counterfactual to $t=1$. The parameter of interest is then
Fix $\left(y,x,X_{i}\right)$. To bound $\tau_{1,x,y}\left(X_{i}\right)$ in ((ref)), for each $y^{\prime}\in\mathcal{\underline{Y}}$ we compare the observed index $X_{i1}\beta-\lambda_{1y^{\prime}}$ to the counterfactual index $x\beta-\lambda_{1y}$, and construct the sets
These sets can be used to form lower/upper bounds on $\tau_{1,x,y}\left(X_{i}\right)$ according to Theorem (ref).
Note that for the ordered choice model, for a fixed $y$, there may be multiple $y^{'}$s yielding nontrivial bounds. This is different from the binary choice case where $\underline{\mathcal{Y}}=\left\{ 1\right\} $ has one element, that provides either a non-trivial lower bound or a non-trivial upper bound. For the ordered choice model, $\underline{\mathcal{Y}}$ has at least two elements, so there may be nontrivial upper and lower bounds. To see this, let $y^{\prime}=y$ (so that $\lambda_{1y}=\lambda_{1y^{\prime}}$) and assume $X_{i1}\beta\geq x\beta$. Then $y^{\prime}\in\mathcal{Y}_{U}$, so that the observed probability $P\left(\left.Y_{i1}\geq y^{\prime}\right|X_{i}\right)$ provides a nontrivial upper bound: \[ P\left(\left.Y_{i1}\geq y^{\prime}\right|X_{i}\right)=P\left(\left.Y_{i1}\geq y\right|X_{i}\right)\geq\tau_{1,x,y}\left(X_{i}\right). \] However, there may be other values in $\underline{\mathcal{Y}}$, call them $y^{\prime},$ for which $y^{\prime}\in\mathcal{Y}_{L}$. This would happen if, e.g., $\lambda_{1y}$ and $\lambda_{1y^{\prime}}$ are such that $X_{i1}\beta-\lambda_{1y^{\prime}}\leq x\beta-\lambda_{1y}$. In this case, the observed probability $P\left(\left.Y_{i1}\geq y^{\prime}\right|X_{i}\right)$ is a nontrivial lower bound: \[ P\left(\left.Y_{i1}\geq y^{\prime}\right|X_{i}\right)\leq\tau_{1,x,y}\left(X_{i}\right). \]
The bounds given by Theorem (ref) are the best bounds across $\mathcal{Y}_{U}$ and $\mathcal{Y}_{L}$, i.e.
setting $\min\emptyset=-\infty$ and $\max\emptyset=+\infty$ to deal with the case when all of $\mathcal{Y}$ provides an upper (lower) bound.
Fix $s\in\left\{ 1,\dots,T\right\} $. For each $y^{\prime}\in\mathcal{\underline{Y}}$, compare $X_{is}\beta-\lambda_{sy^{\prime}}$ to $x\beta-\lambda_{1y}$, and compute the sets:
The bounds under time-stationary errors are then the intersection of the bounds in ((ref)) across all time periods:
The structural function \[ Y_{it}=\max\left\{ 0,\alpha_{i}+X_{it}\beta-U_{it}\right\} \] corresponds to a censored regression model.\footnote{For ease of exposition, we do not consider here extensions covered by our setup such as $Y_{it}=\max\left\{ \lambda_{t},g_{t}\left(\alpha_{i}+X_{it}\beta-U_{it}\right)\right\} $ with time-varying censoring cutoff, and unknown time-varying $g$ function. For the identification and estimation of the common parameters in censored regression models, see e.g. Honore1992,honorePairwiseDifferenceEstimators1994,honoreEstimationTobittypeModels2000a,charlierEstimationCensoredRegression2000,abrevayaIntervalCensoredRegression2020. For the cross-sectional case, see e.g. powellLeastAbsoluteDeviations1984,honorePairwiseDifferenceEstimators1994.} For this model, honore_censored_marginal shows that one can point-identify a meaningful marginal effect using knowledge of $\beta$. Because this model is encompassed by our framework, with $\mathcal{Y}=\left[0,\infty\right)$ and \[ h_{t}^{-}\left(y\right)=h^{-}\left(y\right)=
\] we can use our Theorems 1 and 2 to generate additional point and partial identification results for partial effects in this model.
The object of interest is \[ \tau_{1,x,y}\left(X_{i}\right)=P\left(\left.Y_{i1}\left(x\right)\geq y\right|X_{i}\right). \]
The case $y=0$ is not informative because $P\left(\left.Y_{i1}\left(x\right)\geq0\right|X_{i}\right)=1$. Thus, we restrict attention to $y>0$ and use that $h_{1}^{-}\left(y\right)=y$. If there exists a $y^{\prime}>0$ such that \[ X_{i1}\beta-y^{\prime}=x\beta-y, \] i.e. if \[ y^{\prime}\equiv y+\left(X_{i1}-x\right)\beta>0, \] then
so that Theorem (ref) implies point identification
If $y+\left(X_{i1}-x\right)\beta\leq0$ then \[ X_{i1}\beta-h_{1}^{-}\left(0\right)\geq x\beta-h_{1}^{-}\left(y\right) \] because $h_{1}^{-}\left(0\right)=-\infty$, which yields the trivial upper bound \[ P\left(\left.Y_{i1}\left(x\right)\geq y\right|X_{i}\right)\leq1. \] Each $y^{\prime}>0$ provides a lower bound, since
Then \[ X_{i1}\beta-h_{1}^{-}\left(y^{\prime}\right)=X_{i1}\beta-y^{\prime}\leq x\beta-y=x\beta-h_{1}^{-}\left(y\right) \] so that \[ P\left(\left.Y_{i1}\geq y^{\prime}\right|X_{i}\right)\leq P\left(\left.Y_{i1}\left(x\right)\geq y\right|X_{i}\right). \] Hence, Theorem (ref) obtains \[ \sup_{y^{\prime}>0}P\left(\left.Y_{i1}\geq y^{\prime}\right|X_{i}\right)\leq P\left(\left.Y_{i1}\left(x\right)\geq y\right|X_{i}\right)\leq1. \]
With time-stationary errors, point identification occurs if there exists a time period $t$ and a $y^{\prime}>0$ such that \[ X_{it}\beta-y^{\prime}=x\beta-y, \] because then
This only requires the existence of one time period for which we can find such a $y^{\prime}>0$.
Partial identification thus only results if, for each time period, \[ y+\left(X_{it}-x\right)\beta\leq0, \] in which case the resulting bound is \[ \max_{t}\sup_{y^{\prime}>0}P\left(\left.Y_{it}\geq y^{\prime}\right|X_{i}\right)\leq P\left(\left.Y_{i1}\left(x\right)\geq y\right|X_{i}\right)\leq1. \] This partial identification result only applies when the subpopulation $X_{i}$ is such that for each $t$, $y+\left(X_{it}-x\right)\beta\leq0$.
We report on two numerical experiments that explore the bounds in Theorem 1 and Theorem (ref) for the two-way binary choice and ordered choice models. We show that our bounds are informative without exogeneity or time-homogeneity assumptions on the error terms. The bounds tighten as $T$ grows, and they are more informative for ordered choice models than for binary choice since the bounds tighten as the cardinality of $\underline{\mathcal{Y}}$ grows. Section (ref) presents results for a two-way binary choice probit model with continuous regressors. Section (ref) presents results for a staggered adoption design with both binary and ordered outcomes.
Consider the following data generating process for a binary choice model with two-way fixed effects:
with $\sigma_{x}^{2}=1$ and $\rho=\frac{1}{2}$. The former parameter controls the cross-sectional heterogeneity in $X_{i}$, while the latter controls the degree of variation in the sequence $X_{i}$. The smaller each parameter is, the tighter the bounds are expected to be.
Figure (ref) plots the bounds on the sequence $\mathbb{E}\left[Y_{it}\left(0\right)\right]=\mathbb{E}_{X}\left[\tau_{t,0,1}\left(X_{i}\right)\right]$, $t=1,2,\dots,T$. The bounds are computed according to Theorem 2. We find that the bounds get tighter as the number of periods increases. For example, the width of the interval is $0.26$ when computed with up to 5 periods, $0.07$ when using all $T=20$ periods. In a separate experiment with $T=100$ (not reported), the width shrinks to $0.01$ when using all periods. The identified region may not collapse to a point as $T\rightarrow\infty$, since its width depends on the distribution of $X_{i}$, the stationary distribution of $\left.\alpha_{i}-U_{it}\right|X_{i}$, and the shape of $h_{t}$.
Figure (ref) presents results for the sequence $\mathbb{E}\left[Y_{it}\left(0\right)\right],t\geq1$ for some variations on the model above. The first row, left column shows results for $T=8$, keeping the other design parameters unchanged. Note that, because $\lambda_{t}=\left(-1+2\frac{\left(t+1\right)}{T}\right)^{2}$, this changes the evolution of the expectation. All subpanels of Figure (ref) display results for a deviation from the top left panel. All results in Figure (ref) are for the sequence $\mathbb{E}\left[Y_{it}\left(0\right)\right],t\geq1$, except for the top right panel, which shows the bounds for the sequence $\mathbb{E}\left[Y_{it}\left(1\right)\right],t\geq1$, under the same model as that for the top left panel. The bounds are wider because $X_{it}$ has less mass around $1$ than around $0$. The second row reports bounds on the sequence $\mathbb{E}\left[Y_{it}\left(0\right)\right],t\geq1$, when there are no time-effects (left column) and when there is no persistence in $X_{it}$, i.e. $\rho=0$ (right column). The bounds are slightly wider when the transformation function is time-invariant, and they are tighter when there is no persistence in $X_{i}$. The third row shows the bounds when $\sigma_{x}^{2}$ is $1/4$ (left column; compare to $\sigma_{x}^{2}=1$ in the benchmark case) and $\sigma_{x}^{2}=4$ (right column). The bounds when there is smaller cross-sectional variation in $X_{it}$ at a given time period $t$ are tighter than when there is greater cross-sectional variation.
For the specification consider in this section -- binary probit with two way fixed-effects and short $T$ -- there are no other results in the literature. In Appendix (ref), we present another numerical exercise for binary probit with discrete regressors. That DGP is the same as the one in Section 8 of chernozhukov2013average when there are no time-effects, i.e. $\lambda_{t}=0$ for all $t$. In that case, we recover the bounds in chernozhukov2013average.
In this section, we consider a staggered adoption design with both binary and $J$ ordered outcomes.
The population consists of $G$ groups and individuals in each group $g\in\left\{ 1,\cdots,G\right\} $ are observed over $T$ periods. Individuals are untreated up to and including period $g$; they are treated at period $g+1$, and then stay treated for $t>g+1$, i.e. \[ X_{it}=
\] where $g\left(i\right)$ is individual $i$'s group. Here, $G=19$ and $T\in\left\{ 5,\dots,20\right\} $.
The latent outcome is given by
and $U_{it}$ is standard logistic. The observed outcome is generated as \[ Y_{it}\geq j\Leftrightarrow Y_{it}^{*}\geq\lambda_{jt}, \] where $j\in J\in\left\{ 2,4,6,8\right\} $\footnote{For $J=2$, the outcome is binary, while for $J>2$, the outcome is ordered.}, and the threshold $\lambda_{jt}$ is generated as follows. Set $\overline{J}\equiv\frac{J}{2}+1\in\left\{ 1,3,4,5\right\} $, so that $\left\{ 1,\cdots,\overline{J}-1\right\} $ are the lowest $J/2$ outcomes, and $\left\{ \overline{J},\cdots,J\right\} $ are the top $J/2$ outcomes. Set $\lambda_{\overline{J}1}=0,\,\lambda_{\overline{J}t}\sim\mathcal{U}\left[-1,1\right]$, then draw $\eta_{+},\eta_{-}\sim\mathcal{U}\left[0,1\right]$ and construct \[ \lambda_{jt}=
\] We draw one set of $\left(\left(\lambda_{\overline{J}t}\right)_{t=2}^{T},\eta_{+},\eta_{-}\right)$ and condition our results on them. This allows us to compare the same parameter across different values of $J$ and $T$.
The parameter of interest is \[ \tau\equiv\mathbb{E}\left[\tau_{1,1,\overline{J}}\left(X_{i}\right)\right]=\mathbb{E}\left[P\left(\left.Y_{i1}\left(1\right)\geq\overline{J}\right|X_{i}\right)\right]=P\left(Y_{i1}\left(1\right)\geq\overline{J}\right),\,\overline{J}\in\left\{ 1,3,4,5\right\} . \] This parameter gives the probability of being in the upper half of possible outcomes.\footnote{A two-period version of this setup resembles the nonlinear difference-in-difference setup in athey2006identification. The results in this section differ from theirs because we allow for the combination of fixed effects and discrete outcomes, which is not covered by their results.}
Our results for this design are presented in Figure (ref). The solid black line is $\tau=P\left(Y_{i1}\left(1\right)\geq\overline{J}\right)$. Bounds for different values of $J$ are in colored, dashed lines. Binary choice is in red. At $T=5$, the width of the bounds for $\tau$ are approximately $0.35$, and become as narrow as $0.05$ when $T=20$. As $J$ increases, the bounds narrow substantially. This is especially evident when $T$ is small. For example, at $T=5$, the width of the bounds for $J=6$ is $0.07$, and for $J=8$ it is $0.02$.
Figure (ref) provides further insight, and further clarifies the construction of the bounds in Theorem (ref). Consider $J=4$, $G=6$ and $T=8$. Each panel corresponds to a group $g\left(i\right)$, so that panel “3” corresponds to the population of individuals treated in periods $4-8$. The vertical error bars, at each $t$, correspond to the best bounds that can be constructed for the period-$1$ counterfactual using period-$t$ data, see Remark (ref): \[ \left[\sup_{y^{\prime}\in\mathcal{Y}_{tL}}P\left(\left.Y_{is}\geq y^{\prime}\right|X_{i}\right),\inf_{y^{\prime}\in\mathcal{Y}_{tU}}P\left(\left.Y_{is}\geq y^{\prime}\right|X_{i}\right)\right]\cap\left[0,1\right], \] with
From Figure (ref), it is clear that $J=4$ is not sufficient to provide non-trivial bounds in each period, see for example periods 1-3. The bounds in period 4 are also trivial, but in a way that complements the period 1-3 bounds. Furthermore, the thresholds in periods 6-8 are such that these periods supply non-trivial bounds. By using Assumption (ref), and taking the best bounds in each panel over the time periods as in Theorem (ref), relatively narrow bounds are obtained on $\tau$, see Figure (ref).
This paper discusses identification of counterfactual parameters in a class of nonlinear semiparametric panel models with fixed effects and arbitrary time effects. We derive bounds on the counterfactual survival probability and related functionals that depend on either outcomes from the same time period as the counterfactual (without exogeneity assumptions) or outcomes from across all available time periods (under a “strict exogeneity” or conditional time-stationarity assumption on the errors). The bounds tighten as the cardinality of the support of the dependent variable increases, and, under a time-stationarity assumption on the errors, as $T$ increases. The bounds need not collapse to a point as $T$ grows, rather they collapse to a point under particular assumptions on the time effects and the observed regressors, i.e. when the counterfactual index equals one of the observed indexes. Our bounds are valid for continuous and discrete outcomes and covariates, for both movers and stayers.
Although our focus is on identification, we describe here a potential way of addressing the issue of inference. Note that the bounds in Theorem (ref) can be written as a set of conditional moment inequalities: for fixed $t,x,X_{i},y\in\underline{\mathcal{Y}}$ and for all $\left(s,y^{\prime}\in\underline{\mathcal{Y}}\right)$: \[ \mathbb{E}\left[\left.\left(X_{is}\beta-h_{s}^{-}\left(y^{\prime}\right)-\left(x\beta-h_{t}^{-}\left(y\right)\right)\right)\times\left(1\left\{ Y_{is}\geq y^{\prime}\right\} -\tau_{t,x,y}\left(X_{i}\right)\right)\right|X_{i}\right]\geq0. \] When $\left(\beta,h_{t}\right)$ are partially identified via a set of moment inequalities, these moment inequalities can be added to the program. For either pointwise or uniform in $x$ inference, one could implement the full vector approach of CoxShi if the regressors have finite support, or AndrewsShi if the regressors are continuously distributed.
The bounds of Theorem (ref) are valid for dynamic panel models. However, the bounds may be wider than those obtained by studying specific models, where the dynamic structure is known, e.g., binary choice with lagged outcomes as covariates. In ongoing work, we are studying how to extend our approach to such models. It is an open question as to whether our approach could be extended to nonseparable models.