EconBase
← Back to paper

Identification and Estimation of Average Causal Effects in Fixed Effects Logit Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

94,365 characters · 0 sections · 33 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and Estimation of Average Causal Effects in Fixed Effects Logit Models

abstractThis paper studies identification and estimation of average causal effects, such as average marginal or treatment effects, in fixed effects logit models with short panels. Relating the identified set of these effects to an extremal moment problem, we first show how to obtain sharp bounds on such effects simply, without any optimization. We also consider even simpler outer bounds, which, contrary to the sharp bounds, do not require any first-step nonparametric estimators. We build confidence intervals based on these two approaches and show their asymptotic validity. Monte Carlo simulations suggest that both approaches work well in practice, the second being typically competitive in terms of interval length. Finally, we show that our method is also useful to measure treatment effect heterogeneity. Keywords: Fixed effects logit models, panel data, partial identification. \\ \noindentJEL Codes: C14, C23, C25.

\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Introduction}

In this paper, we consider the identification and estimation of average causal effects in the fixed effects (FE) static binary logit model, with short panels. Estimation of the slope parameters dates back to rasch1961 and andersen1970 chamberlain1980analysis but up to now, there has been no study of the identification and estimation of average causal effects such as average marginal effects (AME) or average treatment effects (ATE) in this model. These parameters are yet of more direct interest than the slope parameter, which only provides information on relative marginal effects. For this reason, and following the influential work of angrist2001estimation angrist2008mostly, many applied economists have turned to using FE linear probability models, as we illustrate in Section (ref).

However, this approach can be misleading, for at least two reasons. Firstly, the linearity of the conditional expectation is often implausible with binary outcome. In a simple setup where covariates only include time dummies and an additional binary variable, the “treatment” of interest, this means that the so-called parallel trend assumption would typically be violated. In Appendix (ref), we give an example where, because of this, the FE linear model estimand is negative, even though the true ATE and average treatment effect on the treated (ATT) are positive. Secondly, even if the parallel trend assumption holds, the two-way fixed effect estimator underlying the FE linear probability model with time dummies is in general inconsistent for the ATT if treatment effects are heterogeneous dCDH2020. And treatment effects are unlikely to be homogenous with binary outcomes, again because probabilities are constrained to be in $[0,1]$.

Unlike the FE linear model, and even if it does impose other restrictions, the FE logit model allows for heterogeneity of treatment effects and does not rely on the parallel trend restriction discussed above. Moreover, we demonstrate in this paper that estimation and inference on average causal effects in this model can be performed simply, without any optimization once the slope parameter is estimated.

We first study in Section (ref) the identification of a class of average causal parameters, including the AME and the ATE, in the FE logit model. Such parameters are generally not point identified, but sharp bounds can be obtained by solving an extremal moment problem, that is, maximizing a moment over probability distributions given the knowledge of some other moments. Using existing results on such problems, we show that the bounds are simple functions of some matrices of moments. One drawback of this approach, yet, is that it depends on functions that need to be estimated nonparametrically in a first step. We then consider even simpler outer bounds avoiding this issue. Importantly, these outer bounds are still informative: the length of the corresponding interval quickly decreases to $0$ as $T$, the number of periods, tends to infinity. We also study identification of heterogeneity measures using the same apparatus.

Next, we consider in Section (ref) estimators of the sharp and outer bounds and corresponding confidence intervals on the average causal effects. We establish the root-$n$ consistency of the estimators of the sharp bounds under regularity conditions. The estimators of the bounds are asymptotically normal except if a function of the slope coefficient is zero. We build a confidence interval of the true effect that is asymptotically valid in both cases. We also show asymptotic validity of confidence intervals based on the outer bounds.

We study in Section (ref) the finite sample properties of our two estimation and inference methods. In line with the theory, they show that the estimated bounds are very informative in practice. Also, the two confidence intervals have coverage close to their nominal level already for moderate sample sizes. Interestingly, we also find in our simulations that the second inference method leads to confidence intervals of similar size as those obtained with the first method. This may seem surprising given that they rely on outer bounds but as it turns out, for typical sample sizes and number of periods, the difference between the outer and the sharp bounds is small compared to the standard errors of their estimators.

Finally, we show in Online Appendix that our method also applies to the (static) ordered and dynamic (binary) FE logit models. Also, we developed with Christophe Gaillac and Ma\"el Laoufi the R package MarginalFElogit and the Stata command mfelogit, which perform inference on the AME and ATE (depending on whether $X$ is continuous or binary) with the two methods considered here, and accommodates the case of an individual-specific number of observations.\footnote{The R package can be found \href{https://github.com/cgaillac/MarginalFElogit}{here}. The more preliminary Stata command is available on the SSC repository.}

\paragraph{Related literature}

Our work is related to the literature on the identification and estimation of average causal effects with panel data. Bias-correcting approaches have been developed for panels with large $T$ for both the logit and probit models fernandez2009fixed,fernandez2016individual. Few papers have studied average causal effects with fixed $T$ in parametric models with FE, as we do here. In the dynamic FE logit model, aguirregabiria2021identification show point identification of some average effects of lagged values of the outcome, and dobronyi2021 study the partial identification of average causal effects that can be obtained from the knowledge of the slope parameter. We complement these papers by considering other parameters, including average causal effects (AME or ATE) of exogenous covariates in the same dynamic model, though our main focus is on the static, binary FE logit model.

Several papers have studied identification of causal effects in a nonparametric context, see in particular altonji2005cross,Hoderlein2012,chernozhukov2013average,chernozhukov2015nonparametric,chernozhukov2019,botosaru2017binarization. Compared to these papers, we consider a more constrained model, with the aim of providing a simple characterization and estimation of the bounds in this set-up. Note that chernozhukov2013average describe in their Section 8 a generic method for computing sharp bounds on average causal effects of semiparametric models, which also applies to FE logit models honore2006bounds. It is based on the idea that the distribution of individual effects can be approximated arbitrarily well by a distribution with fixed and finite support, if the number of support points is large enough. Then, bounds can be computed by solving as many linear programming problems as, typically, twice the number of units. In addition to being computationally more costly than our method for the FE logit model, this approach relies on an approximation, and it remains unclear how to quantify the error of this approximation. Thus, even if our approach is much more specialized than theirs, it offers important advantages for logit models.

Finally, our work relies on results about moment problems, which have been extensively studied since Chebyshev and Markov. We refer to schmudgen2017moment for a recent mathematical exposition and to dette1997 for applications to various statistical problems. DR17 use similar results on moment problems to obtain bounds on segregation measures with small units. dobronyi2021 use other results on moment problems to characterize the identified set of slope parameters in dynamic FE logit models, generalizing the work of honore2022moment. Even if the primary goal of dobronyi2021 differs from ours, in both papers we use similar ideas to rewrite the initial identification problem as a more standard moment problem.

\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Literature review on applications with FE and binary outcomes}

To assess how often and in which manner researchers estimate models with FE and binary outcomes, we conducted a literature review of all papers published in 2022 in the five following economic journals: the American Economic Review (AER), Econometrica (ECTA), the Journal of Political Economy (JPE), the Quarterly Journal of Economics (QJE) and the Review of Economic Studies (RES). That year, these five journals published 402 papers, excluding comments and corrigenda. 74% of them have at least some empirical content; note that here, we even include mostly theoretical papers. Among those (but excluding their supplementary material), we looked for linear or binary choice models, such as logit and probit models, that include FEs. We considered “fixed effect” as being either individual effects in a standard panel data, or dummies corresponding to a discrete variable whose number of support points is at least 10% of the number of observations, so that if used in a nonlinear model, the incidental parameter problem becomes potentially important. For papers matching these criteria, we checked which model the authors estimated.

table[table omitted — 1,064 chars of source]

The details of our findings are displayed in Table (ref). We identified 26 papers with binary outcomes and FE, representing 9% of the set of all papers with some empirical content. In all but two of these papers, only linear probability models (LPM) were used. These findings support that (i) models with FE and binary outcomes are quite common in practice; (ii) in such cases, the common practice is to estimate linear probability models; (iii) even if the FE logit model is probably one of the most well-known nonlinear model with FE, it is seldom used today. One explanation for (iii) could be that inference methods for the AME and the ATE, which are of primary interest in empirical work, were not available. By developing computationally simple methods for them below, we thus hope that more applied researchers will turn to FE logit models.

\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Identification}

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{The set-up and identification of $\beta_0$}

We consider a panel with $T$ periods and observe binary outcomes $Y_1,...,Y_T$ and for each period $t$, a vector of covariates $X_t:=(X_{t1},...,X_{tp})' \in \mathbb R^p$. We let $Y:=(Y_1,...,Y_T)'$, $X:=(X'_1,...,X'_T)'$ and assume in this section that the joint distribution of $(X,Y)$ is identified. We also rely on the following notation hereafter. For any random variables $A$ and $B$, we let $F_A$ and $F_{A|B=b}$ denote the cumulative distribution function (cdf) of $A$ and its cdf conditional on $B=b$, respectively. We also let $\text{Supp}(A)$ and $\text{Supp}(A|B=b)$ denote the support of $A$ and its support conditional on $B=b$. We also make the following assumption.

hypWe have $Y_t=\mathds{1}\left\{X'_t\beta_0 + \alpha + \varepsilon_t\geq 0\right\}$, where the $(\varepsilon_t)_{t=1,...,T}$ are i.i.d., independent of $(\alpha,X)$ and follow a logistic distribution.

Importantly, the individual effect $\alpha$ is allowed to be correlated in an unspecified way with $X$. In this model, $S:=\sum_{t=1}^T Y_{t}$ is a sufficient statistic for $\alpha$ andersen1970. As a result, identification of $\beta_0$ can be achieved by maximizing the expected conditional log-likelihood, conditioning not only on $X$ but also on $S$. For any $y=(y_1,...y_T)\in\{0,1\}^T$, let us define

align*[align* omitted — 231 chars of source]

The second term $\ell_c(y|x;\beta)$ is the conditional log-likelihood. To ensure that $\beta_0$ is identified as the unique maximizer of the expected conditional log-likelihood, we impose the following.

hyp$E[ \sum_{t,t'} (X_t-X_{t'})(X_t-X_{t'})']$ is nonsingular.

Assumption (ref) is necessary and sufficient for the identification of the slope parameter in fixed effects linear models. The following proposition ensures that this is also the case in FE logit models. It must be well-known, but we have not been able to find it in the literature. Its proof and the proofs of other identification results are presented in Appendix (ref).

propSuppose that the distribution of $(X,Y)$ is identified, Assumption (ref) holds and for all $t\neq t'$ and $k\in\{1,...,p\}$, $E[(X_{tk} - X_{t'k})^2]<\infty$. Then $\beta_0$ is point identified if and only if Assumption (ref) holds. In this case, $\beta_0 = \arg\max_\beta E\left(\ell_c(Y|X,\beta)\right)$ and $\mathcal{I}_0=-E\left(\partial {}^2\ell_c/\partial \beta\partial\beta'(Y|X;\beta_0)\right)$ is nonsingular.

The second part of Proposition (ref) shows that $\beta_0$ can be identified as the unique maximizer of the average expected log-likelihood. Under mild regularity conditions, $\mathcal{I}^{-1}_0$ is the asymptotic variance of the conditional maximum likelihood estimator (CMLE) but also the semiparametric efficiency bound for $\beta_0$ Hahn1997.

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Partial identification of average causal effects}

We now turn to the partial identification of average effects. Specifically, we consider parameters of the form $$\delta_0 := E\left[g(X,\alpha,\beta_0)\right],$$ where the function $g$ is known. We will impose an additional restriction on $g$ below but our analysis will cover the following three key parameters:

example[Average marginal effect, AME] Suppose that the distribution of $X_{tk}$ is continuous. Then, the average marginal effect of $X_{tk}$ (considering here the $t$-th period) is the infinitesimal change of $X_{tk}$ on the probability that $Y_t=1$: $$\delta_0 := E\left[\frac{\partial P(Y_t=1|X,\alpha)}{\partial X_{tk}}\right] = \beta_{0k} E\left[\Lambda'(X_t'\beta_0 + \alpha)\right],$$ where $\Lambda(x):=1/(1+\exp(-x))$ denotes the cdf of the logistic distribution. Because $\Lambda'=\Lambda\times (1-\Lambda)$, we also have \begin{equation} \delta_0 = \beta_{0k} E\left[\Lambda(X_t'\beta_0 + \alpha)(1-\Lambda(X_t'\beta_0 + \alpha)) \right]. \end{equation}
example[Average treatment effect, ATE] Suppose now that $\text{Supp}(X_{tk})=\{0,1\}$. Then, the average treatment effect of $X_{tk}$ is the average effect on $Y_t$ of moving $X_{tk}$ from $0$ to $1$ for all units: $$\delta_0 := E\left[\Lambda\left(X_t^{(1)}{}'\beta_0+\alpha\right)- \Lambda\left(X_t^{(0)}{}'\beta_0+\alpha\right)\right].$$ There, $X^{(d)}_t$ is such that its $k$-th component is equal to $d\in\{0,1\}$ and any of its other $j$-th component ($j\ne k$) is equal to $X_{tj}$. We can rewrite the ATE as $$\delta_0 = E\left[(2X_{tk}-1)\left(Y_t - \Lambda\left(X_t^{(1-X_{tk})}{}'\beta_0+\alpha\right)\right)\right],$$ where the term $E[(2X_{tk}-1)Y_t]$ is identified. For convenience, we focus in this section on the remaining term, namely \begin{equation} E\left[-(2X_{tk}-1)\Lambda\left(X_t^{(1-X_{tk})}'\beta_0+\alpha\right)\right]. \end{equation}
example[Average structural function, ASF] This function, evaluated at $\widetilde{x}_t\in \text{Supp}(X_t)$, corresponds to the counterfactual probability that $Y_t=1$ if everyone had $X_t$ equal to $\widetilde{x}_t$: \begin{equation} \delta_0 := E\left[\Lambda(\widetilde{x}_t'\beta_0+\alpha)\right]. \end{equation}

The model does not impose any restriction between $F_{\alpha|X=x}$ and $F_{\alpha|X=x'}$. As a result, to obtain the sharp identified set $\Delta$ for $\delta_0$, it suffices to obtain it on $$\delta_0(x):= E\left[g(X,\alpha,\beta_0)|X=x\right]$$ and then integrate the corresponding sets over $x$ to obtain the sharp identified set of $\delta_0$ (note that $\delta_0(x)$ could also be of interest by itself). The only restrictions on $F_{\alpha|X=x}$ come from the data, namely from the distribution of $Y|X=x$. Actually, because $S$ is a sufficient statistic for $\alpha$, its (conditional) distribution exhausts all the information available on $\alpha$. Then, the restrictions on $F_{\alpha|X=x}$ reduce to

equation[equation omitted — 174 chars of source]

In order to obtain the identified set of $\delta_0(x)$, we thus have to find all possible values of $E\left[g(X,\alpha,\beta_0)|X=x\right]$ under the $T+1$ constraints on $F_{\alpha|X=x}$ given by (ref).

To derive simple formulas for the sharp and an outer identified set of $\delta_0(x)$ and $\delta_0$, we use two ideas. The first is to reparameterize the problem, so as to replace, under conditions on $g$, the moment function $a\mapsto g(x,a,\beta_0)$ by $a\mapsto a^{T+1}$, and accordingly simplify the moment restrictions (ref). This leads us to consider the identified set of $\theta:=\int u^{T+1} d\mu(u)$ when the unknown positive measure $\mu$ is constrained by moments of the kind $\int u^k d\mu(u)=m_k$ $(k=0,...,T)$. The second idea, then, is to rely on existing results, in particular the theory of moment problems schmudgen2017moment, to obtain simple expressions for the sharp and an outer identified set for $\theta$. This, in turn, allows us to obtain the sharp and an outer identified set on $\delta_0(x)$ and $\delta_0$.

\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont}{Reparametrization of the problem}

Our reparameterization relies on the fact that for any $v(x,\beta)\in\mathbb R$, if we let $U:=\Lambda(v(X,\beta_0)+\alpha)$, the constraints (ref) satisfy, for $k=0,...,T$,

align[align omitted — 272 chars of source]

with $\Omega_{x,\beta}(u):=\prod_{t=1}^T [u(\exp(x_t'\beta -v(x,\beta)) - 1)+1]$. We then assume that $\delta_0(x)$ takes a related form $E[Q(U)/\Omega_{x,\beta_0}(U)|X=x]$, for some polynomial $Q$ with degree at most $T+1$:

hypFor all $(x,\beta)\in\text{Supp}(X)\times \mathbb R^p$, there exists $v(x,\beta)\in\mathbb R$ such that letting $\Omega_{x,\beta}(u):=\prod_{t=1}^T [u(\exp(x_t'\beta -v(x,\beta)) - 1)+1]$, $u\mapsto g(x,\Lambda^{-1}(u) - v(x,\beta),\beta)\times \Omega_{x,\beta}(u)$ defined on $(0,1)$ is a polynomial of degree at most $T+1$. We let $\lambda_t(x,\beta)$ denote the coefficient of $u^t$ of this polynomial.

Under Assumption (ref), $\delta_0(x)=E[h_{x,\beta_0}(U)|X=x]$ with $U:=\Lambda(v(X,\beta_0)+\alpha)$ and $h_{x,\beta}(u):=g(x,\Lambda^{-1}(u) - v(x,\beta),\beta)$. Then, we have

align[align omitted — 242 chars of source]

Importantly, Assumption (ref) holds for the three previous average causal parameters:

\setcounter{example}{0}

example[AME] If we let $v(x,\beta)=x_t'\beta$, we obtain, in view of (ref), $h_{x,\beta}(u) = \beta_k u (1-u)$. Then, remark that $\Omega_{x,\beta}$ is of degree at most $T-1$, since its $t$-th term in the product simplifies to one. As a result, $h_{x,\beta}\times \Omega_{x,\beta}$ is a polynomial of degree at most $T+1$, and Assumption (ref) holds.
example[ATE] Recall that $\delta_0$ is defined by (ref), as the remaining part of the ATE is identified. By letting $v(x,\beta)=x^{(1-x_{tk})}_t{}'\beta$, we obtain $h_{x,\beta}(u) = -(2x_{tk}-1) u$. Since $\Omega_{x,\beta}$ is of degree at most $T$, $h_{x,\beta}\times \Omega_{x,\beta}$ is a polynomial of degree at most $T+1$.
example[ASF] If we let $v(x,\beta)=\widetilde{x}_t'\beta$, we obtain, in view of (ref), $h_{x,\beta}(u) = u$. Since $\Omega_{x,\beta}$ is of degree at most $T$, $h_{x,\beta}\times \Omega_{x,\beta}$ is a polynomial of degree at most $T+1$.

We showed in (ref) that the probabilities $(P(S=k|X=x))_{k=0,...,T}$ are linear combinations of the expectations $(E[U^t/\Omega_{x,\beta_0}(U)|X=x])_{t=0,...,T}$. Actually, Lemma (ref) below establishes that there is a one-to-one relationship between the two, so that these moments are identifiable. Specifically, we show that

align[align omitted — 92 chars of source]

with $$Z_t(x,s,\beta) := \binom{T-t}{s-t} \frac{\exp(sv(x,\beta))}{C_s(x;\beta)}, \quad Z_t := Z_t(X,S,\beta_0),$$ and where we use the convention $\binom{T-t}{s-t}=0$ if $s<t$. Hence, the first $T+1$ terms of the sum in (ref) are identified. The last term of the sum, on the other hand, is not identified in general. To write this term in a more convenient way and complete our reparameterization, let us define the following probability measure on $[0,1]$ (remark that $\Omega_{x,\beta_0}(U)>0$ almost surely):

equation[equation omitted — 171 chars of source]

for any Borel set $A\subseteq [0,1]$. Then, in view of (ref)-(ref), we obtain

equation[equation omitted — 158 chars of source]

Moreover, the first $T$ moments of $\mu_x$ are identified:

equation[equation omitted — 113 chars of source]

In other words, $\delta_0(x)$ satisfies (ref) with $\mu_x$ a probability measure whose vector of first moments is equal to $m(x):=(m_0(x),...,m_T(x))$. Lemma (ref) below shows that the converse holds as well: for any probability measure $\mu$ on $[0,1]$ with a vector of first moments equal to $m(x)$ and such that $\mu(\{0,1\})=0$, we prove that $$\sum_{t=0}^T \lambda_t(x,\beta_0)E[Z_t|X=x] + \lambda_{T+1}(x,\beta_0) E[Z_0|X=x] \int_0^1 u^{T+1}d\mu(u).$$ is in the identified set of $\delta_0(x)$. To state Lemma (ref), we denote by $\mathcal{D}$ the set of positive measures on $[0,1]$ and for any $m=(m_0,...,m_T)\in [0,1]^{T+1}$, we let $$\mathcal{D}(m)=\left\{\mu\in\mathcal{D}: \int u^k d\mu(u)=m_k, \; k=0,...,T\right\}.$$ In words, $\mathcal{D}(m)$ is the subset of positive measures on $[0,1]$ whose vector of first $T+1$ raw moments (including the moment of order 0) is equal to $m$.

lemSuppose that the distribution of $(Y,X)$ is identified and Assumptions (ref)-(ref) hold. Then, for all $x\in\text{Supp}(X)$, the identified set of $\delta_0(x)$ is \begin{align} &\bigg\{\sum_{t=0}^T \lambda_t(x,\beta_0) E[Z_t|X=x] + \lambda_{T+1}(x,\beta_0) E[Z_0|X=x]\int_0^1 u^{T+1}d\mu(u): \; \notag \\ & \quad \mu\in \mathcal{D}(m(x)), \; \mu(\{0,1\})=0\bigg\}. \end{align}

We prove this result by, basically, showing that there is a one-to-one mapping between $F_{\alpha|X=x} $ and $\mu_x$ defined by (ref). This implies that we can express the problem using $\mu_x$ and the $(m_t(x))_{t=0,...,T}$ instead of $F_{\alpha|X=x}$ and the $(P(S=s|X=x))_{s=0,...,T}$, respectively.

\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont}{Sharp bounds}

Lemma (ref) ensures that finding the sharp identification region of $\delta_0(x)$ reduces to finding the set of values of $\int_0^1 u^{T+1}d\mu_x(u)$ for $\mu_x\in \mathcal{D}(m(x))$ satisfying $\mu_x(\{0,1\})=0$. Without this last constraint, the problem is known as the truncated Hausdorff problem. The solution, accounting for the constraint, is given in Proposition (ref) below. Let us first introduce additional notation. For any $t\ge 1$, $s\geq t$ and $m=(m_0,...,m_s)\in\mathbb R^{s+1}$, we define the Hankel matrices $\underline{\mathbb{H}}_t(m)$ and $\overline{\mathbb{H}}_t(m)$ as $$

array[array omitted — 396 chars of source]

$$ Then, let $H_t(m)=\det \left(\mathbb{H}_t(m)\right)$ and $\overline{H}_t(m)=\det(\overline{\mathbb{H}}_t(m))$. Now, for $m\in \mathbb{R}^{T+1}$ and $q\in\mathbb R$, consider $H_{T+1}(m,q)$. By expanding this determinant along its last column, we see that $q\mapsto H_{T+1}(m,q)$ is linear. Then, let $a_{T+1}(m)$ and $b_{T+1}(m)$ denote the corresponding intercept and slope parameters, so that $\underline{H}_{T+1}(m,q) = \underline{a}_{T+1}(m)+\underline{b}_{T+1}(m)q$. We define similarly $ \overline{a}_{T+1}(m)$ and $\overline{b}_{T+1}(m)$. Finally, let $|A|$ denote the cardinal of set $A$ and $\lfloor x\rfloor$ denote the integer part of $x\in \mathbb R$.

propLet $\theta_0=\int_0^1 u^{T+1} d\mu(u)$ for some unknown measure $\mu\in \mathcal{D}(m)$ satisfying $\mu(\{0,1\})=0$ and some identified $m\in[0,1]^{T+1}$. Then, $\underline{H}_T(m)\ge 0$, $\overline{H}_T(m)\ge 0$ and $\Theta$, the closure of the identified set of $\theta_0$, satisfies $\Theta=[\underline{q}_T(m),\, \overline{q}_T(m)]$, where the functions $\underline{q}_T$ and $\overline{q}_T$ are such that: \begin{enumerate} • If $\underline{H}_T(m) \times \overline{H}_T(m)> 0$, then $\underline{q}_T(m)<\overline{q}_T(m)$ with $\underline{q}_T(m)=-\underline{a}_{T+1}(m)/$ $\underline{b}_{T+1}(m)$ and $\overline{q}_T(m)=-\overline{a}_{T+1}(m)/\overline{b}_{T+1}(m)$. Moreover, $\left|\text{Supp}(\mu)\right|> \lfloor T/2 \rfloor$. • If $\underline{H}_T(m) \times \overline{H}_T(m)= 0$, then $\mathcal{D}(m)=\{\mu\}$, $\underline{q}_T(m)=\overline{q}_T(m)=\theta_0$ and $\Theta=\{\theta_0\}$. Moreover, letting $T'=\min\{t\leq T: \underline{H}_t(m) \times \overline{H}_t(m)= 0\}$, $\underline{q}_T(m)$ is equal to \begin{equation} \begin{array}{rcl} -a_{T'}(m_{T-T'+1},...,m_T)/b_{T'}(m_{T-T'+1},...,m_T) & if H_{T'}(m)=0, \\ -\overline{a}_{T'}(m_{T-T'+1},...,m_T)/\overline{b}_{T'}(m_{T-T'+1},...,m_T) & if \overline{H}_{T'}(m)=0. \end{array} \end{equation} Finally, $\left|\text{Supp}(\mu)\right|\leq \lfloor T'/2 \rfloor$. \end{enumerate}

Proposition (ref) shows that the sharp bounds of $\theta_0$ are very simple to obtain as rational functions of determinants of some identified matrices. Point 1 follows from classical results in moment theory, see in particular Theorem 10.8 and Proposition 10.15 in schmudgen2017moment. The first part of the Point 2 is also well-known. On the other hand, to the best of our knowledge, its second part is new. To understand why we get respectively partial identification and point identification in Point 1 and Point 2, let us assume $T=2$. Since the support of $\mu$ is included in $[0,1]$, we have $m_1^2 \le m_2 \le m_1$. This implies that $\underline{H}_2(m)\ge 0$ and $\overline{H}_2(m)\ge 0$. Moreover, if $m_1^2 = m_2$, so that $\overline{H}_2(m)=0$, the variance of the distribution is 0. Then, $\mathcal{D}(m)=\{\mu\}$ where $\mu$ is the Dirac distribution at $m_1$.\footnote{The case $m_1 = m_2$ corresponds to a Bernoulli distribution with parameter $m_1$, which is excluded here because $\mu(\{0,1\})=0$.} If, on the other hand, $m_1 > m_2 > m_1^2$, so that $\underline{H}_2(m) \times \overline{H}_2(m)>0$, the variance is strictly positive and there are infinitely many distributions with first two moments equal to $(m_1, m_2)$. More intuition on Proposition (ref) is given at the beginning of its proof.

By combining Lemma (ref) with Propositions (ref) and (ref), we obtain the following result on the closure of the sharp identified sets of $\delta_0(x)$ and $\delta_0$, denoted respectively by $\Delta(x)$ and $\Delta$.

thmIf the distribution of $(X,Y)$ is identified and Assumptions (ref)-(ref) hold, $\Delta(x)=[\underline{\delta}(x),\, \overline{\delta}(x)]$, with \begin{equation} \begin{array}{rcl} \delta(x) & := & \sum_{t=0}^T \lambda_t(x,\beta_0) E[Z_t|X=x] + \lambda_{T+1}(x,\beta_0) E[Z_0|X=x] \big( q_T(m(x)) \\[2mm] & & \mathds{1}\left\{\lambda_{T+1}(x,\beta_0)\geq 0\right\} + \overline{q}_T(m(x)) \mathds{1}\left\{\lambda_{T+1}(x,\beta_0)< 0\right\} \big), \\[2mm] \overline{\delta}(x) & := & \sum_{t=0}^T \lambda_t(x,\beta_0) E[Z_t|X=x] + \lambda_{T+1}(x,\beta_0) E[Z_0|X=x] \big( \overline{q}_T(m(x)) \\[2mm] & & \mathds{1}\left\{\lambda_{T+1}(x,\beta_0)\geq 0\right\} + q_T(m(x)) \mathds{1}\left\{\lambda_{T+1}(x,\beta_0)< 0\right\} \big). \end{array} \end{equation} Moreover, $\Delta =[\underline{\delta},\, \overline{\delta}]$, with $\underline{\delta}=E(\underline{\delta}(X))$, $\overline{\delta}=E(\overline{\delta}(X))$. $\delta_0$ is point identified if and only if $$P\left(\big\{\lambda_{T+1}(X,\beta_0)=0\big\} \cup \big\{|\text{Supp}(\alpha|X)|\leq \lfloor T/2\rfloor \big\}\right)=1.$$

From (ref), point identification of $\delta_0(x)$ holds if and only if $\lambda_{T+1}(x_t,\beta_0)=0$ or $\underline{q}_T(m(x))=\overline{q}_T(m(x))$. The first case occurs for the AME and ATT if $\beta_{0k}=0$. It also holds for the AME at $t$ if $\min_{s\ne t} |(x_s-x_t)'\beta_0|=0$. This could be expected, as Hoderlein2012 show that average marginal effects are nonparametrically identified on “stayers”, namely individuals for whom $X_{it}$ remains constant between two periods. Finally, point identification can also be achieved if $|\text{Supp}(\alpha|X)|\le \left\lfloor T/2\right\rfloor$.\footnote{Such identification is achieved using the logit structure: as a complement to Hoderlein2012, chernozhukov2019 show that average marginal effects are not identified nonparametrically for non-stayers.} This corresponds to the second case described in Proposition (ref). Intuitively, if $\alpha|X$ has few points of support, its full distribution is characterized by its first moments. Then, given these moments, the higher moments are fully determined. As an illustration, assume that $T=2$ and $\alpha|X$ is degenerate and equal to $\alpha_0$. Then, some algebra shows that $m(X)=(1,\Lambda(\alpha_0+X_T'\beta_0),\Lambda(\alpha_0+X_T'\beta_0)^2)'$ which implies that the variance of any distribution in $\mathcal{D}(m(X))$ is zero. Thus, $\mathcal{D}(m(X))$ reduces to the Dirac distribution at $\Lambda(\alpha_0+X_T'\beta_0)$, implying $\underline{q}_2(m(X))=\overline{q}_2(m(X))= \Lambda(\alpha_0+X_T'\beta_0)^3$.

\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont}{Outer bounds}

A drawback of the previous approach, in terms of estimation, is that even if we are interested in $\delta_0$ rather than $\delta_0(x)$, we need to estimate nonparametrically the functions $m_k(x)$ for all $x$ observed in the data, because they enter nonlinearly into the bounds. This requires regularity conditions and the choice of tuning parameters. Also, we need to account for the constraint that $m(x)$ is a vector of moments. An estimator $\widehat{m}(x)$ need not be a vector of moments and, for instance, need not satisfy the variance constraint $\widehat{m}_2(x)\ge \widehat{m}^2_1(x)$. If not, the set $\mathcal{D}(\widehat{m}(x))$ is empty and the bounds $(\underline{q}_T(\widehat{m}(x)),\overline{q}_T(\widehat{m}(x)))$ are not defined anymore. We suggest in Section (ref) below a constrained estimator which is a vector of moments but this complicates the estimation of the bounds.

As an alternative, we consider simple outer bounds that, when focusing on the parameter $\delta_0$, do not require nonparametric estimation of $m_k(x)$. The idea is to find good approximations, in a sup-norm sense on $[0,1]$, of $u\mapsto u^{T+1}$ by a polynomial of degree $T$. Let us define, as in Proposition (ref), $\theta_0:=\int_0^1 u^{T+1}d\mu(u)$ for some $\mu\in\mathcal{D}(m)$ and suppose that there exists $K>0$ and $(b_0,...,b_T)$ such that $$\sup_{u\in[0,1]} \left|u^{T+1}-\sum_{k=0}^T b_k u^k\right|\le K.$$ By Jensen's inequality, this implies that $|\theta_0 - \sum_{k=0}^T b_k m_k| \le K$. As a result, we obtain that $[\sum_{k=0}^T b_k m_k - K,\; \sum_{k=0}^T b_k m_k +K]$ is an outer set for $\theta_0$. Moreover, an optimal solution to the uniform approximation problem is very simple, and the corresponding $K$ is small, as the following proposition shows. Let $\mathbb{T}^c_{T+1}$ denote the Chebyshev polynomial of degree $T+1$, renormalized so that its leading coefficient is equal to 1.\footnote{Recall that the unnormalized Chebyshev polynomials are defined by $\mathbb{T}^u_0(x)=1$, $\mathbb{T}^u_1(x)=x$ and $\mathbb{T}^u_{k+1}(x)=2x\mathbb{T}^u_{k}(x)-\mathbb{T}^u_{k-1}(x)$ for any $k\ge 1$.} Then, let $\mathbb{T}_{T+1}(u):= 2^{-T-1}\mathbb{T}^c_{T+1}(2u-1)$ and let $-b^*_{k,T}$ denote the coefficient of degree $k$ of $\mathbb{T}_{T+1}$. Finally, recall that $\Theta$ denotes the closure of the identified set of $\theta_0$.

propWe have $(b^*_{0,T},...,b^*_{T,T})=\arg\min_{(b_0,...,b_T)\in\mathbb R^{T+1}} \sup_{u\in[0,1]} \left|u^{T+1}-\sum_{k=0}^T b_K u^k\right|$. Moreover, under the conditions of Proposition (ref), we have $\Theta \subseteq \Theta^o$, with \begin{align*} \Theta^o & := \left[\sum_{k=0}^T b^*_{k,T}m_k - \frac{1}{2\times 4^T} ,\,\sum_{k=0}^T b^*_{k,T}m_k + \frac{1}{2\times 4^T}\right]. \end{align*} Finally, we may have $\Theta=\Theta^o$.

Figure (ref) displays $u\mapsto u^{T+1}$ and its best uniform approximation $P^*_T(u):=\sum_{k=0}^T b^*_{k,T} u^k$ for $T=2, 3$ and 4. As we can see, the approximation is already good for $T=2$, and the two functions become almost indistinguishable for $T=4$. This could be expected as the bound $1/(2\times 4^T)$ decreases very quickly with $T$.

figure[figure omitted — 208 chars of source]

By combining Lemma (ref) with Propositions (ref) and (ref), we obtain the following outer identified sets for $\delta_0(x)$ and $\delta_0$. Recall that $\Delta(x)$ and $\Delta$ denote the closure of their identified sets.

thmIf the distribution of $(X,Y)$ is identified and Assumptions (ref)-(ref) hold, we have $\Delta(x)\subseteq \Delta^o(x) :=[\tilde{\delta}(x)\pm\overline{b}(x)]$ and $\Delta \subseteq \Delta^o:=[\tilde{\delta}\pm \overline{b}]$, with \begin{align*} \tilde{\delta}(x) & := \sum_{t=0}^T \left(\lambda_t(x,\beta_0) + b^*_{t,T}\lambda_{T+1}(x,\beta_0)\right) E(Z_t|X=x), \\ \overline{b}(x) & := \frac{1}{2\times 4^T} |\lambda_{T+1}(x,\beta_0)| E\left[Z_0|X=x\right],\\ \tilde{\delta} & := E\bigg[\sum_{t=0}^T \left(\lambda_t(X,\beta_0) + b^*_{t,T}\lambda_{T+1}(X,\beta_0)\right)Z_t \bigg],\\ \overline{b} & := \frac{1}{2\times 4^T} E\left[|\lambda_{T+1}(X,\beta_0)|\times Z_0\right]. \end{align*} Moreover, we may have $\Delta^o(x)= \Delta(x)$ and $\Delta^o= \Delta$.

As mentioned above, the outer bounds on $\delta_0$ are very simple and do not require nonparametric estimation of $m(X)$. Moreover, because $[\underline{\delta},\overline{\delta}] \subseteq [\tilde{\delta}-\overline{b},\tilde{\delta}+\overline{b}]$, we obtain the following bound for the length of the identified set of $\delta_0$:

equation[equation omitted — 152 chars of source]

For some distributions of $X$, this inequality yields an upper bound on the rate of decrease of the size of the identified set as $T$ increases. Specifically, assume that for all $t\le T$, $P[|X_t'\beta_0-v(X,\beta_0)|\leq \ln(2)]=1$. Then, additional algebra shows that $E\left[|\lambda_{T+1}(X,\beta_0)|\times Z_0\right]\leq 1$, which in turn implies $$\overline{\delta} - \underline{\delta} \leq \frac{1}{4^T}.$$ Similarly if for all $t\le T$, $P(|(X_t'\beta_0-v(X,\beta_0)|\leq c)=1$ for some $c\in [\ln(2), \ln(5))$ then $\overline{\delta} - \underline{\delta} \leq (e^c-1)^{T-1}/4^T \leq K^{T}$ for $K=(e^c-1)/4< 1$. This discussion may convey the impression that for $X$ with larger support, possibly unbounded, the identified set could be large. However, recall that the right-hand side of (ref) is only an upper bound on the true length $\overline{\delta} - \underline{\delta}$. In Appendix (ref), we show how the sharp bounds evolve when varying the support of $X$, for a specific distribution of $\alpha|X$. We find that even if $X$ has large support, the bounds remain extremely informative and their length quickly decreases with $T$.

chernozhukov2013average also obtain, in their Theorem 4, an exponential rate of decrease on the length of the identified set of the ASF. Actually, their result imposes substantially weaker conditions on the distribution of $Y_t|X_t,\alpha$. On the other hand, it only holds for finitely supported $X$ and imposes additional restrictions on the distribution of $(X,\alpha)$.

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Measuring the heterogeneity of treatment effects}

Whereas in the linear probability model, $(x,\alpha) \mapsto g(x,\alpha,\beta_0)$ is constant when considering the AME or ATE, this function varies in the FE logit model. Understanding this heterogeneity is important for, e.g., policy design manski2004statistical. We have shown above how to get sharp and outer bounds for $\delta_0(x)=E[g(X,\alpha,\beta_0)|X=x]$. However, $X$ has dimension $pT$, so even if the $\alpha_i$ were observed, $\delta_0(x)$ would typically be inaccurately estimated due to the curse of dimensionality. On the other hand, if $\widetilde{X}$ is a low-dimensional subvector of $X$ (e.g., $\widetilde{X}=X_{tj}$), we can measure how $g(X,\alpha,\beta_0)$ varies with $\widetilde{X}$ by partially identifying $$\delta_{\widetilde{X}}(\widetilde{x}):=E[g(X,\alpha,\beta_0)|\widetilde{X}=\widetilde{x}]=E[\delta_0(X)|\widetilde{X}=\widetilde{x}].$$ One may also be interested by heterogeneity with respect to factors not included in $X$, e.g. time-invariant features. For instance, if one considers participation to the labor market as a function of the number of children, we may be interested by heterogeneity with respect to education. We then consider the parameter $\delta_W(w):= E[g(X,\alpha,\beta_0)|W=w]$. This parameter has a similar interpretation as $\delta_{\widetilde{X}}(\widetilde{x})$, even if $W$ is not a function of $X$, provided that the following conditional independence holds:

equation[equation omitted — 108 chars of source]

Then, if we consider for instance the AME, we have $g(x,\alpha,\beta_0)=\partial P(Y_T=1|X=x,W=w,\alpha)/\partial x_{tk}$, and $\delta_W(w)$ is the average marginal effect of $X_k$ on the subpopulation $W=w$. When $W$ is multivariate, $W=(W_1,...,W_p)$, one may also want to learn about the main drivers of heterogeneity, by considering the coefficient $\varsigma_j$ of $W_j \ (j\in\{1,...,p\})$ in the best linear prediction of $g(X,\alpha,\beta_0)$ by $W$.

We obtain sharp bounds on $\delta_W(w)$ by using $\delta_W(w)= E[\delta_0(W,X)|W=w]$, with $\delta_0(w,x)=E[g(X,\alpha,\beta_0)|W=w, X=x]$, and remarking that Lemma (ref) holds when conditioning on $(X,W)$ instead of $X$ only. Then, the same reasoning as for obtaining the sharp bounds on $\delta_0(x)$ applies to $\delta_0(w,x)$. For instance, its sharp lower bound satisfies

align*[align* omitted — 337 chars of source]

where $m_t(w,x):=E[Z_t|W=w, X=x]/E[Z_0|W=w, X=x]$. Then the sharp lower bound on $\delta_W(w)$ is simply $E[\underline{\delta}(W,X)|W=w]$, and similarly for its upper bound. Also, following bontemps2012set, the sharp lower bound on $\varsigma_j$ is $$\underline{\varsigma}_j = \frac{E\left[W^r_j\left(\underline{\delta}(W,X)\mathds{1}\left\{W^r_j>0\right\}+ \overline{\delta}(W,X)\mathds{1}\left\{W^r_j<0\right\}\right)\right]}{E[W^{r2}_j]},$$ where $\overline{\delta}(W,X)$ is the sharp upper bound on $\delta_0(W,X)$ and $W^r_j$ is the residual of the theoretical regression of $W_j$ on 1 and the $(W_k)_{k\ne j}$.

The sharp bounds above rely on the nonparametric functions $E[Z_t|W=w, X=x]$, with $(W,X)$ possibly of high dimension. The following proposition shows that we can avoid this by considering, again, outer bounds. Its proof is identical to that of Theorem (ref) and therefore omitted.

propIf the distribution of $(W,X,Y)$ is identified, Assumptions (ref)-(ref) and (ref) holds, an outer identified set on $\delta_W(w)$ (resp. on $\varsigma_j$) is $[\tilde{\delta}_W(w)\pm \overline{b}_W(w)]$ (resp. $[\tilde{\varsigma}_j\pm \overline{b}_{\varsigma_j}]$), with \begin{align*} \tilde{\delta}_W(w) & = E\left[\sum_{t=0}^T (\lambda_t(X,\beta_0) + b^*_{t,T} \lambda_{T+1}(X,\beta_0)) Z_t \big| W=w\right], \\ \overline{b}_W(w) & = \frac{1}{2\times 4^T} E\left[|\lambda_{T+1}(X,\beta_0)| Z_0 \big| W=w \right],\\ \tilde{\varsigma}_j & = \frac{E\left[W^r_j\sum_{t=0}^T \left(\lambda_t(X,\beta_0) + b^*_{t,T} \lambda_{T+1}(X,\beta_0)\right)Z_t\right]}{E[W^{r2}_j]}, \\ \overline{b}_{\varsigma_j} & = \frac{E\left[\left|W^r_j\lambda_{T+1}(X,\beta_0) \right|\times Z_0 \right]}{E[W^{r2}_j]}. \end{align*}

\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Estimation and inference}

We turn to the estimation of bounds on $\delta_0$, and inference on this parameter, using a sample $(Y_i,X_i)_{i=1,...,n}$. We first consider sharp bounds, before considering outer bounds. Though we omit details to save space, the methodology developed below can be applied to estimate bounds on the heterogeneity parameters $\delta_W(w)$ and $\theta_j$.

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Estimation and inference based on the sharp bounds}

\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont}{Definition of the estimators}

By Theorem (ref) and the law of iterated expectations, we have $\underline{\delta} = E[\underline{h}(X,S)]$, with

align*[align* omitted — 270 chars of source]

with $r(x,s,\beta) := \sum_{t=0}^T \lambda_t(x,\beta) Z_t(x,s,\beta)$, so that $r(X,S,\beta_0) = \sum_{t=0}^T \lambda_t(X,\beta_0)Z_t$.\footnote{Recall that for the ATE, we actually focused so far on a part of it, defined by (ref). To estimate the ATE itself, we thus need to replace $r(x,s,\beta)$ by $(2x_{kt}-1)y_t+ \sum_{t=0}^T Z_t(x,s,\beta) \lambda_t(x,\beta)$.} Also, $\overline{\delta}= E[\overline{h}(X,S)]$, with $\overline{h}(X,S)$ defined similarly as $\underline{h}(X,S)$. We estimate $\underline{h}$ and $\overline{h}$, and in turn $\underline{\delta}$ and $\overline{\delta}$ by plug-in, in four steps:

enumerate• Estimation of $\beta_0$ by the conditional maximum likelihood estimator: $$\widehat{\beta}:=\arg\max_{b\in B}\sum_{i=1}^n\ell_c(Y_i|X_i,b).$$ • Initial estimation of $m=(m_0,...,m_T)$:\\ Let $\gamma_{0j}(x):=P \left(S=j|X=x\right)$ for $j=0,...,T$. By definition, $m_t(x)=c_t(x)/c_0(x)$, with \begin{equation} c_t(x):=E(Z_t|X=x)=\sum_{j=t}^T\binom{T-t}{j-t}\frac{\gamma_{0j}(x)\exp(j v(x,\beta_0))}{C_{j}(x,\beta_0)}. \end{equation} We first estimate the $(\gamma_{0j})_{j=0,...,T}$ nonparametrically. We consider a local polynomial estimator $\widehat{\gamma}_j$ of order $\ell$, with kernel function $K$ and bandwidth $h_n$. Then, let $\widehat{c}_t(x)$ be the plug-in estimator of $c_t(x)$ based on (ref), and let $\widetilde{m}_t(x)=\widehat{c}_t(x)/\widehat{c}_0(x)$ for $t=0,...,T$. Remark that by construction, $\widetilde{m}_0(x)=1$. • Constrained estimation of $m$: \\ $\widetilde{m}$ has an important drawback: because of sampling uncertainty, it is not necessarily a vector of moments. Formally, if we let $\mathcal{M}_T:=\left\{m\in [0,1]^{T+1}: \mathcal{D}(m)\ne \emptyset\right\}$, we may have $\widetilde{m}\not\in\mathcal{M}_T$, in which case $\underline{q}_T(\widetilde{m}(X_i))$ and $\overline{q}_T(\widetilde{m}(X_i))$ are not defined.\footnote{In our simulations, this already occurs with $T=3$ and $n=1,000$, even if $m(x)$ is in the interior of $\mathcal{M}_T$.} We thus construct an estimator $\widehat{m}$ such that for all $i=1,...,n$, $\widehat{m}(X_i)\in\mathcal{M}_T$, by exploiting Proposition (ref). For any $(m_t)_{t\geq 0}$ and $t\in\{0,...,T\}$, let $m_{\rightarrow t}:=(m_0,...,m_t)$. The idea of the estimator is to use the first elements of $\widetilde{m}(x)$, until $\widetilde{m}_t(x) \not\in [\underline{q}_{t-1}(\widetilde{m}_{\rightarrow t-1}(x)), \overline{q}_{t-1}(\widetilde{m}_{\rightarrow t-1}(x))]$, or (for technical reasons) $\widetilde{m}_t(x)$ is close to $\underline{q}_{t-1}(\widetilde{m}_{\rightarrow t-1}(x))$ or $\overline{q}_{t-1}(\widetilde{m}_{\rightarrow t-1}(x))$. In such a case, we simply replace $\widetilde{m}_t(x)$ by $\underline{q}_{t-1}(\widetilde{m}_{\rightarrow t-1}(x))$ or $\overline{q}_{t-1}(\widetilde{m}_{\rightarrow t-1}(x))$. We finally complete the vector using the second part of Proposition (ref). Specifically, let $c_n$ be a sequence tending to 0 at a rate specified later and define \begin{equation} \widehat{I}(x):= \max\left\{t\in \{1,...,T\}: \;\forall s\le t,\; H_s(\widetilde{m}_{\rightarrow s}(x))\times \overline{H}_s(\widetilde{m}_{\rightarrow s}(x)) > c_n\right\}, \end{equation} with the convention that $\max \emptyset = 0$. We then let $\widehat{m}_{\rightarrow \widehat{I}(x)}(x):= \widetilde{m}_{\rightarrow \widehat{I}(x)}(x)$. If $\widehat{I}(x)=T$, $\widehat{m}(x)$ is fully defined. Otherwise, we complete $\widehat{m}(x)$ by first letting $$\widehat{m}_{\widehat{I}(x)+1}(x):= \left|\begin{array}{cl} \underline{q}_{\widehat{I}(x)}(\widetilde{m}_{\rightarrow \widehat{I}(x)}(x)) & \text{ if } \underline{H}_{\widehat{I}(x)+1}(\widetilde{m}_{\rightarrow \widehat{I}(x)+1}(x)) < c_n^{1/2}, \\ \overline{q}_{\widehat{I}(x)}(\widetilde{m}_{\rightarrow \widehat{I}(x)}(x)) & \text{ otherwise.} \end{array}\right.$$ Next, if $\widehat{I}(x)+1<T$, we have, by construction, $$\underline{H}_{\widehat{I}(x)+1}(\widehat{m}_{\rightarrow \widehat{I}(x)+1}(x))\times \overline{H}_{\widehat{I}(x)+1}(\widehat{m}_{\rightarrow \widehat{I}(x)+1}(x)) = 0. $$ Then, Part 2 of Proposition (ref) shows that there are unique moments $\widehat{m}_{\widehat{I}(x)+2},...,\widehat{m}_{T}$ that are compatible with $\widehat{m}_{\rightarrow \widehat{I}(x)+1}(x)$. We construct them by induction using this proposition. In the end, the corresponding vector $\widehat{m}(x)$ satisfies $\widehat{m}(x)\in\mathcal{M}_T$. • Plug-in estimation of $\underline{h}$, $\overline{h}$ and the sharp bounds:\\ We compute the estimator $\widehat{\underline{\delta}} = \frac{1}{n}\sum_{i=1}^n \widehat{\underline{h}}(X_i,S_i)$, with \begin{align} \widehat{h}(x,s) & = r(x,s,\widehat{\beta}) + \widehat{c}_0(x) \lambda_{T+1}(x, \widehat{\beta}) \bigg[\overline{q}_T(\widehat{m}(x)) \mathds{1}\left\{\lambda_{T+1}(x,\widehat{\beta})\geq 0\right\} \nonumber \\ & + q_T(\widehat{m}(x)) \mathds{1}\left\{\lambda_{T+1}(x,\widehat{\beta}) < 0\right\} \bigg]. \end{align} We define $\widehat{\overline{\delta}}$ and $\widehat{\overline{h}}(x,s)$ similarly.

Two remarks on this estimation method are in order. First, we produce estimators of the bounds that never cross, even if the model is misspecified. In particular, misspecification can induce $m(x)\not\in\mathcal{M}_T$ for some $x$, but still $\widehat{m}(x)\in \mathcal{M}_T$ in this case.\footnote{ In theory, one could exploit the constraint that $m(X)\in\mathcal{M}_T$ almost surely (a.s.) to derive a specification test for the FE logit model. However, we would recommend using only the conditional moment restrictions $E[\partial \ell_c/\partial \beta(Y|X,\beta_0)|X]=0$. This leads to a much simpler test and it is unclear whether adding the former constraint would lead to large power gains.} Second, we can define in an exact similar way estimators of the sharp bounds of parameters related to the heterogeneity of treatment effects described in Subsection (ref).

\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont}{Asymptotic properties of the estimated bounds}

We first derive the asymptotic distribution on the estimated bounds $(\widehat{\underline{\delta}}, \widehat{\overline{\delta}})$. To this end, we rely on Assumptions (ref)-(ref) and the assumptions below. We refer to Theorem (ref) in the Online Appendix for a consistency result on $(\widehat{\underline{\delta}}, \widehat{\overline{\delta}})$ under weaker conditions than those below. To state some regularity conditions below, we let (rearranging terms here) $X=(X^{c \, \prime},X^{d \, \prime})'$ with $X^c$ a vector of $p_cT$ continuous regressors and $X^d$ a vector of $(p-p_c)T$ discrete regressors.

hyp\ \begin{enumerate} • The variables $(X_i, \alpha_i, \varepsilon_{i1},...,\varepsilon_{iT})$ are i.i.d across $i$. • $\text{Supp}(X)$ is a compact set and $\beta_0$ belongs to the interior of $B$, a convex compact set. \end{enumerate}
defin[Regularity Condition $Reg(j)$]\ A function $f$ from $\text{Supp}(X)\times \mathbb{R}^p$ to $\mathbb{R}$ is $Reg(j)$ if $\beta\in \mathbb{R}^{p}\mapsto f(x,\beta)$ admits continuous derivatives of order $j$ for any $x\in \text{Supp}(X)$ and $x\in \text{Supp}(X)\mapsto \partial^{|k|} f(x,\beta)/(\partial^{k_1} \beta_{1}... \partial^{k_p} \beta_{p})$ is continuous for any $\beta\in \mathbb{R}^{p}$ and any $(k_1,...,k_p)\in \mathbb{N}^p$ such that $|k|:=\sum_{r=1}^p k_r\leq j$.

\newcounter{counterhypprod} \setcounter{counterhypprod}{\value{hyp}}

hypThe functions $v$, $\lambda_{0},...,$ and $\lambda_{T}$ are $Reg(2)$. Moreover, there exist $k$, $a$ and $\rho$ such that $\lambda_{T+1}(x,\beta)=a(\beta_k)\rho(x,\beta)$, $\rho$ is $Reg(2)$, $a$ is twice continuously differentiable and either $a(\beta_{0k}) = 0$ or $P\left(\rho(X,\beta_0)=0\right)=0$.
hyp\begin{enumerate} • The support of $X^d$ is finite and $X^c|X^d=x^d$ admits a density $f_{X^c|X^d=x^d}$ with respect to the Lebesgue measure on $\mathbb R^{p_cT}$. $f_{X^c|X^d=x^d}$ is $C^{\ell+2}$ and bounded away from $0$ on $\text{Supp}(X^c|X^d=x^d)$, which is convex. Also, $\ell \geq p_cT/2$.\footnote{If $p_c=p$, this condition and Assumption (ref).(ref) should be understood replacing $X^d$ by a constant variable. If $p_c=0$, all the conditions on $X^c|X^d=x^d$ should be removed.} • $x^c \mapsto \gamma_0(x^c,x^d)$ is $C^{\ell+1}$ on $\text{Supp}(X^c|X^d=x^d)$ for all $x_d$. • $K$ is a density on $\mathbb R^{pT}$ with compact support bounded away from 0 in a neighborhood of 0. $K$ is $C^{\ell+2}$ on $\mathbb R^{pT}$. • $h_n=C_h n^{-\xi_h}, c_n=C_c n^{-\xi_c}$ with $0<\xi_c<\xi_h(\ell+1)$, $\frac{1}{4(\ell+1)}<\xi_h<\frac{1}{p_cT+2(\ell+1)}$ and $C_c,C_h>0$. \end{enumerate}
hypEither $|\text{Supp}(\alpha|X=x)|>\left\lfloor T/2\right\rfloor$ for all $x \in \text{Supp}(X)$, or the function $x\mapsto |\text{Supp}(\alpha|X=x)|$ is constant on $\text{Supp}(X)$.

Note that the decomposition of $\lambda_{T+1}(x,\beta)$ in Assumption (ref) may not be unique. Also, this assumption holds for the three average parameters introduced in Section (ref), under regularity conditions on $X$:

\setcounter{example}{0}

example[AME] In view of (ref), $\lambda_{T+1}(x,\beta)=-\beta_k \prod_{s\neq t}\left(\exp[(x_s-x_t)'\beta]-1\right)$. Let $a(\beta_{k}) := \beta_{k}$ and $\rho(x,\beta) =-\prod_{s \neq t}\left(\exp\left[(x_s-x_t)'\beta\right]-1\right)$. Assumption (ref) then holds if either $\beta_{0k}=0$ or $(X_s - X_t)'\beta_0 \neq 0$ a.s. for all $s\ne t$.
example[ATE] Recall that $\delta_0$ is defined by (ref), as the remaining part of the ATE is identified and $\lambda_{T+1}(x,\beta)=-(2x_{tk} - 1) \prod_{s=1}^T\left(\exp[(x_s-x_t^{(1-x_{tk})})'\beta]-1\right)$. Let $a(\beta_k)=1-\exp(\beta_k)$ and $\rho(x,\beta)=\exp\left((x_{tk}-1)\beta_k\right)\prod_{s \neq t} \left(\exp[(x_s -x_t^{(1 - x_{tk})})'\beta]-1 \right)$. Assumption (ref) then holds if either $\beta_{0k}=0$ or $(X_s-X_{t}^{(1-X_{tk})})'\beta_0 \neq 0$ a.s. for all $s\ne t$.
example[ASF] In view of (ref), $\lambda_{T+1}(x,\beta)=\prod_{s=1}^T\left(\exp[(x_s-\tilde{x}_t)'\beta]-1\right)$. Let $a(\beta_k):=1$ and $\rho(x,\beta)=\prod_{s=1}^T\left(\exp[(x_s-\tilde{x}_t)'\beta]-1\right)$. Assumption (ref) then holds if the distribution of $X_s'\beta_0$ is continuous for all $s\in\{1,...,T\}$.

Assumption (ref) is a standard regularity condition when estimating, as here, parameters involving first-step nonparametric estimators. It ensures that $\widehat{\gamma}$ converges uniformly to $\gamma_0$ over the support of $X$ at a rate of at least $\eta_n:= (\ln n/(n h_n^{p_cT}))^{1/2} + h_n^{\ell+1}=o(n^{-1/4})$. We also impose restrictions on the $c_n$ used in our constrained estimation of $m(x)$. In particular, the restrictions on $\xi_h$ and $\xi_c$, when combined with Assumption (ref), ensure that $\widehat{I}(x)$ defined in (ref) converges to $I(x):=\max\left\{t\in \{1,...,T\}: \;\forall s\le t,\; \underline{H}_s(m_{\rightarrow s}(x))\times \overline{H}_s(m_{\rightarrow s}(x)) > 0\right\}$. The restrictions on $\xi_h$ and $\xi_c$ also ensure that the derivatives of $\widehat{\gamma}(x)$ remain bounded in probability.

Assumption (ref) is less standard than the other conditions. It restricts the number of support points of $\alpha$ given $X=x$. This number should either be large enough, in which case it may vary with $x$, or remain constant as a function of $x$. We impose this condition because $\underline{q}_T$ and $\overline{q}_T$ are not regular everywhere for $T\geq 3$: whereas they are infinitely differentiable on the relative interior of the moment space $\mathcal{M}_T$, they may not even be directionally differentiable on the boundary of $\mathcal{M}_T$.\footnote{See DR17 for a proof of the first statement. Regarding the second, one can show that, e.g., $m_1\mapsto \underline{q}_3(m_0,m_1,m_2,m_3)$ is not differentiable at $m=(1,m_1, m_1^2,m_1^3)$.}

Before presenting the asymptotic distribution of $\left(\widehat{\underline{\delta}}, \widehat{\overline{\delta}}\right)$, we introduce additional notation. Let $\phi=\mathcal{I}_0^{-1}\partial \ell_c(Y|X,\beta_0)/\partial \beta$ denote the influence function of $\widehat{\beta}$ and $\Gamma = (\mathds{1}\left\{S=0\right\},...,$ $\mathds{1}\left\{S=T\right\})'$. In general, the estimated bounds are asymptotically normal. Their asymptotic variance is $\Sigma:=V[\underline{\psi}, \overline{\psi}]$, with

align*[align* omitted — 300 chars of source]

where the vectors $\underline{v}_\beta$ and $\underline{v}_\gamma(x,s)$ (and similarly $\overline{v}_\beta$ and $\overline{v}_\gamma(x,s)$) are, respectively, the gradient of $E(\underline{h}(X, S))$ with respect to $\beta$ and the gradient of $\underline{h}(x, s)$ with respect to $\gamma(x)$. The exact expressions of these four vectors are given in Online Appendix (ref). The term $\underline{v}_\beta' \phi$ captures the effect of estimating $\beta_0$, whereas $\underline{v}_\gamma(X,S)' [\Gamma - \gamma_0(X)]$ captures the effect of estimating the nonparametric function $\gamma_0$.

In some cases, the estimated bounds are not asymptotically normal. Then, their asymptotic distribution depends on $\Omega$, the variance matrix of $(r(X,S,\beta_0),\phi)$ and on the functions $\underline{K}(u):=E[\underline{k}(X,S,u)]$ and $\overline{K}(u):=E[\overline{k}(X,S,u)]$, with $u\in \mathbb{R}^p$ and where

align*[align* omitted — 299 chars of source]

$\overline{k}$ is defined similarly, with the minimum replaced by the maximum.

thmSuppose that Assumptions (ref)-(ref) hold and let $(\underline{N},\overline{N})\sim \mathcal{N}(0,\Sigma)$ and $(N_1,N_2)\sim \mathcal{N}(0, \Omega)$. Then: $$\sqrt{n} \left(\widehat{\underline{\delta}} - \underline{\delta}, \widehat{\overline{\delta}} - \overline{\delta}\right) \stackrel{d}{\longrightarrow} \left|\begin{array}{ll} (\underline{N},\overline{N}) & \text{if } a(\beta_{0k})\ne 0, \\ \left(N_1+\underline{K}(N_2),N_1+\overline{K}(N_2)\right) & \text{otherwise.} \end{array} \right. $$

The proof of Theorem (ref), as the proofs of other estimation results, is in Section (ref) of the Online Appendix. The main difficulty in this proof is to deal with the nonlinear terms $\underline{q}_T(\widehat{m}_T)$ and $\overline{q}_T(\widehat{m}_T)$, given that $\underline{q}_T$ and $\overline{q}_T$ may not be differentiable. The construction of $\widehat{m}_T$ above and Assumption (ref) are key for dealing with this issue.

If $a(\beta_{0k})\ne 0$ (equivalently, $\beta_{0k}\ne 0$ for the AME and the ATE), the limit distribution is normal. But this is generally not the case when $a(\beta_{0k})=0$, as the functions $\underline{K}$ and $\overline{K}$ are not linear. We still get asymptotic normality of the AME and the ATE, however, if the whole vector $\beta_0$ is equal to 0: then, the functions $\underline{K}$ and $\overline{K}$ are equal and linear. Note that because $a(\beta_{0k})=1$ for the ASF, the estimated bounds are always asymptotically normal.

\@startsection{subsubsection}{3}{0mm}{-0.8\baselineskip}{0.4\baselineskip}{\normalfont}{Inference on $\delta_0$}

With Theorem (ref) at hand, we can construct confidence intervals on $\delta_0$ that are asymptotically valid, whether or not the estimated bounds are asymptotically normal. For simplicity, we focus hereafter on the AME, the ATE and the ASF. First, we estimate $\Sigma$ by $\widehat{\Sigma}=\frac{1}{n}\sum_{i=1}^n(\widehat{\underline{\psi}}_i,\widehat{\overline{\psi}}_i)' (\widehat{\underline{\psi}}_i,\widehat{\overline{\psi}}_i)$, where $$\widehat{\underline{\psi}}_i := \widehat{\underline{h}}(X_i, S_i) - \widehat{\underline{\delta}} + \widehat{\underline{v}}_\beta' \widehat{\phi}_i + \widehat{\underline{v}}_\gamma' [\Gamma_i - \widehat{\gamma}(X_i)],$$ with $\widehat{\phi}_i:=-\left[\frac{1}{n}\sum_{j=1}^n \partial {}^2 \ell_c/\partial \beta \partial \beta'(Y_j|X_j;\widehat{\beta})\right]^{-1}\partial \ell_c/\partial \beta(Y_i|X_i;\widehat{\beta})$ and $\widehat{\underline{v}}_\beta$ and $\widehat{\underline{v}}_\gamma$ are estimators of $\underline{v}_\beta$ and $\underline{v}_\gamma$ defined in Section (ref) of the Online Appendix. $\widehat{\overline{\psi}}_i$ is defined similarly.

Next, let $\varphi_\alpha=1$ for the ASF and otherwise, let $\varphi_\alpha=1$ denote a consistent test of asymptotic level $\alpha$ of $\beta_{0k}=0$. For instance we can use $\varphi_\alpha=1$ if a Wald test rejects $\beta_{0k}=0$, $\varphi_\alpha=0$ otherwise. Following imbens2004confidence, let $c_\alpha$ denote the unique solution to

equation[equation omitted — 253 chars of source]

with $\Phi$ the cdf of a standard normal distribution and $\Sigma_{ij}$ the $(i,j)$ term of $\Sigma$. Then, we define $\text{CI}^{\, 1}_{1-\alpha}$ as $$CI^{\, 1}_{1-\alpha}:=\left|

array[array omitted — 463 chars of source]

\right.$$ The following proposition shows that $CI^{\, 1}_{1-\alpha}$ is pointwise valid as $n\to\infty$.

propSuppose that $\delta_0$ is the AME, the ATE or the ASF, Assumptions (ref)-(ref) hold and $\min(\Sigma_{11},\Sigma_{22})>0$. Then $\liminf_n \inf_{\delta_0\in[\underline{\delta},\overline{\delta}]} P(\delta_0\in\text{CI}^{\, 1}_{1-\alpha})\geq 1-\alpha$, with equality if $a(\beta_{0k})\neq 0$.

Intuitively, $\text{CI}^{\, 1}_{1-\alpha}$ asymptotically reaches its nominal level when $a(\beta_{0k})\neq 0$ because it includes $[\widehat{\underline{\delta}} - c_\alpha (\widehat{\Sigma}_{11}/n)^{1/2}, \; \widehat{\overline{\delta}} + c_\alpha (\widehat{\Sigma}_{22}/n)^{1/2}]$, and the latter interval has asymptotic coverage $1-\alpha$, by Theorem (ref). When $a(\beta_{0k})=0$, which applies only to the AME and the ATE, the asymptotic coverage of $\text{CI}^{\, 1}_{1-\alpha}$ is also at least $1-\alpha$ because $\delta_0=0\in\text{CI}^{\, 1}_{1-\alpha}$ as soon as $\varphi_\alpha=0$. Finally, for simplicity we did not consider other parameters, for which $\delta_0$ may be unknown when $a(\beta_{0k})=0$, but we can still adapt $\text{CI}^{\, 1}_{1-\alpha}$ to such cases. We simply have to replace (i) $c_\alpha$ and $\varphi_\alpha$ by $c_{\alpha_1}$ and $\varphi_{\alpha_1}$ for some $\alpha_1\in (0,\alpha)$; (ii) $0$ in $\text{CI}^{\, 1}_{1-\alpha}$ by the lower and upper bounds of confidence intervals with nominal level $\alpha-\alpha_1$ (for some $\alpha_1\in (0,\alpha)$) on $E[r(X,S,\beta_0)]$, since $\delta_0=E[r(X,S,\beta_0)]$ when $a(\beta_{0k})=0$.

The interval $\text{CI}^{\, 1}_{1-\alpha}$ may have a uniform coverage over an appropriate set of data generating processes (DGPs). Establishing this formally would however require to establish the uniform convergence in distribution of $(\widehat{\underline{\delta}}, \widehat{\overline{\delta}})$, a multistep estimator with a nonparametric first step. We leave this issue for future research.

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Estimation and inference based on the outer bounds}

We now turn to the estimation of the outer bounds, and inference on $\delta_0$ based on these outer bounds. First, by Theorem (ref) and the law of iterated expectations, these outer bounds are $\tilde{\delta} \pm \overline{b}$, with $\tilde{\delta} = E[p(X, S,\beta_0)]$, and\footnote{Again, in the case of the ATE, we must add $(2x_{kt}-1)y_t$ to $p(x,s,\beta)$.}

align*[align* omitted — 318 chars of source]

We estimate $\tilde{\delta}$ and $\overline{b}$ by plug-in: $\widehat{\tilde{\delta}} = \sum_{i=1}^n p(X_i,S_i,\widehat{\beta})/n$ and $$\widehat{\overline{b}} = \frac{1}{2\times 4^T} \frac{1}{n}\sum_{i=1}^n \left|\lambda_{T+1}(X_i,\widehat{\beta})\right| \binom{T}{S_i}\frac{\exp(S_i v(X_{i},\widehat{\beta}))}{C_{S_i}(X_i,\widehat{\beta})}.$$ Lemma (ref) in the Online Appendix implies that $\widehat{\tilde{\delta}} - \widehat{\overline{b}}$ (resp. $\widehat{\tilde{\delta}} + \widehat{\overline{b}}$) is a consistent estimator of the outer bound $\tilde{\delta} - \overline{b}$ (resp. $\tilde{\delta} + \overline{b}$).

To define confidence intervals on $\delta_0$ based on $\widehat{\tilde{\delta}}$ and $\widehat{\overline{b}}$, we introduce additional notation. Let

align*[align* omitted — 372 chars of source]

Then, we define $\sigma^2=V(\psi)$ and $\widehat{\sigma}^2 = \sum_{i=1}^n \widehat{\psi}_i^2/n$. The first confidence interval we consider is $$\text{CI}^{\, 2}_{1-\alpha} = \left[\widehat{\tilde{\delta}} \pm q_\alpha\left(\frac{n^{1/2}\widehat{\overline{b}}}{\widehat{\sigma}}\right) \frac{\widehat{\sigma}}{n^{1/2}}\right],$$ where $q_\alpha(b)$ denotes the quantile of order $1-\alpha$ of a $|\mathcal{N}(b,1)|$. The intuition for the asymptotic validity of this confidence interval is as follows. We prove in Lemma (ref) in the Online Appendix that

equation[equation omitted — 155 chars of source]

and $\widehat{\sigma}\stackrel{P}{\longrightarrow} \sigma$. For simplicity, let us assume that $\widehat{\sigma}=\sigma$, $\widehat{\overline{b}}=\overline{b}$ and the asymptotic approximation (ref) is exact. Then

equation[equation omitted — 196 chars of source]

It is not difficult to show that $b \mapsto q_\alpha(b)$ is symmetric and increasing on $[0,\infty)$. Then, because $|\tilde{\delta} - \delta_0| \le \overline{b}$, we have

equation[equation omitted — 230 chars of source]

Theorem (ref) below shows that (ref) holds asymptotically even if $\widehat{\sigma}\ne \sigma$, $\widehat{\overline{b}}\ne \overline{b}$ and (ref) is not exact, as soon as $|\delta_0 - \tilde{\delta}| < \overline{b}$. The only difference between $\text{CI}^{\, 2}_{1-\alpha}$ and a standard confidence interval is that because of the possible bias, we consider $q_\alpha\left(n^{1/2}\widehat{\overline{b}}/\widehat{\sigma}\right)$ instead of the usual normal quantile $q_\alpha\left(0\right)$. This difference is important however: it implies that $\text{CI}^{\, 2}_{1-\alpha}$ converges to the outer set $[\tilde{\delta} \pm \overline{b}]$, rather than to $\{\tilde{\delta}\}$.

thmSuppose that Assumptions (ref)-(ref) hold, $v$, $\lambda_{0}, \, ...,\lambda_{T+1}$ are $Reg(1)$, $\sigma^2>0$ and either $\overline{b}=0$ or $\left\vert\delta_0-\tilde{\delta}\right\vert\neq\overline{b}$. Then: $$\liminf_{n\to\infty} P\left(\delta_0\in \text{CI}^{\, 2}_{1-\alpha}\right) \geq 1-\alpha.$$

Hence, $\text{CI}^{\, 2}_{1-\alpha}$ is pointwise asymptotically conservative under the condition that $\overline{b}=0$ or $|\delta_0-\tilde{\delta}|<\overline{b}$. This condition is very weak, as the following lemma shows:

lemSuppose that Assumptions (ref)-(ref) hold. Then, $|\delta_0-\tilde{\delta}|=\overline{b}>0$ implies \ \begin{align} & P(\lambda_{T+1}(X,\beta_0)\ne 0)>0 and either P\left(Supp(U|X)\subseteq \mathcal{R}_{T,X} | \lambda_{T+1}(X,\beta_0)\ne 0\right)=1 \notag \\ & or P\left(Supp(U|X)\subseteq \mathcal{R}'_{T,X} | \lambda_{T+1}(X,\beta_0)\ne 0\right)=1, \end{align} where $\mathcal{R}_{T,x}$ is defined as \begin{equation} \mathcal{R}_{T,x}=\left|\begin{array}{ccl} \arg\max_{u\in[0,1]} \mathbb{T}_{T+1}(u) & if & \lambda_{T+1}(x,\beta_0)>0 \\ \arg\min_{u\in[0,1]} \mathbb{T}_{T+1}(u) & if & \lambda_{T+1}(x,\beta_0)<0 \end{array} \right. \end{equation} and $\mathcal{R}'_{T,X}$ is defined similarly, simply switching the min and the max in (ref).

Condition (ref) strongly restricts the support of $\alpha|X$. First, because $\mathbb{T}_{T+1}$ has at most $\left\lfloor T/2\right\rfloor+1$ maxima and minima, one must have $|\text{Supp}(\alpha|X)|\le \left\lfloor T/2\right\rfloor+1$. Second, (ref) imposes that the support points of $\alpha$ change discontinuously around any $x_0$ for which $\lambda_{T+1}(x_0,\beta_0)=0$ and $\partial \lambda_{T+1}(x_0,\beta_0)/\partial x\neq 0$. In any case, our simulations below suggest that $\text{CI}^{\, 2}_{1-\alpha}$ still has good coverage when (ref) holds. Another potential issue with $\text{CI}^{\, 2}_{1-\alpha}$ is that because it does not account for the variability of $\widehat{\overline{b}}$, it may not be asymptotically uniformly valid for the AME and ATE for sequences of DGPs such that $\beta_{0k}$ tends to 0 at the rate $n^{1/2}$. We consider in Appendix (ref) another confidence interval that is uniformly valid on a set of DGPs allowing for such sequences of $\beta_{0k}$.

\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Monte Carlo simulations}

We now study the estimators of the bounds on the AME, and inference based on them, through simulations.\footnote{For an application of our methodology to a real dataset with both continuous and discrete regressors, see the \href{https://github.com/cgaillac/MarginalFElogit/blob/master/MarginalFElogit.pdf}{documentation} of our R package MarginalFELogit.} We first compare our two methods. Then, we compare our approach with FE linear probability models. Finally, we study measures of heterogeneous effects.

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Comparison of the two inference methods}

We first compare the finite sample performances of our two methods. We consider $T\in\{2,3\}$ and $n\in\{250; 500;1,000\}$ and four DGPs. In all of them, we assume that $(X_1,...,X_T)$ are i.i.d., with $X_t\in \mathbb R$, uniformly distributed on $[-1/2,1/2]$ and $\beta_0=1$. We also suppose that $\alpha = - X_T'\beta_0 +\eta$. Then, we consider different distributions for $\eta|X$. In DGP1, we let $\eta=0$. By Theorem (ref), $\delta_0$ is point identified for all $T\ge 2$ in this case. In DGP2, we let $\eta|X\sim \mathcal{N}(0,1)$. Again by Theorem (ref), $\delta_0$ is partially identified in this case for all $T\ge 2$. We also consider DGP$3_T$ where $\eta|X$ is uniformly distributed over $\Lambda^{-1}(\mathcal{R}_{T,X})$, where $\mathcal{R}_{T,X}$ is defined in (ref). Note that we index DGP3 by $T$ because contrary to DGP1 and DGP2, this DGP actually varies with $T$. Table (ref) shows the true parameter, the sharp bounds and the outer bounds $\underline{\delta}^o := \tilde{\delta}-\overline{b}$ and $\overline{\delta}^o:= \tilde{\delta}+\overline{b}$ for $T\in\{2,3\}$. In the partially identified case of DGP2, the sharp bounds are very informative, even with $T=2$. The outer bounds are also very informative in all DGPs.

table[table omitted — 1,146 chars of source]

For each of the DGPs above, we consider and perform 3,000 simulations for each such $(T,n)$. We then compute the estimators of the sharp bounds, $\widehat{\underline{\delta}}$ and $\widehat{\overline{\delta}}$, those of the outer bounds, $\widehat{\underline{\delta}^o} :=\widehat{\tilde{\delta}} - \widehat{\overline{b}}$ and $\widehat{\overline{\delta}^o} :=\widehat{\tilde{\delta}} + \widehat{\overline{b}}$, and $\text{CI}^1_{0.95}$ and $\text{CI}^2_{0.95}$. To estimate nonparametrically $\gamma_0$, we use local linear estimators with a Gaussian product kernel. We use data-driven bandwidths $h_n$ and thresholds $c_n$, on which further details are given in Section (ref) of the Online Appendix.

Table (ref) displays the properties of the estimators underlying the two methods. The estimators of the bounds appear to have a small bias in all cases, except perhaps with DGP3$_2$. With this DGP, the distribution of $\eta|X=x$ does not vary in a smooth way with $x$: $\eta=\Lambda^{-1}(1/4)$ when $x_1 \le x_2$ while $\eta=\Lambda^{-1}(3/4)$ otherwise. As a result, the regularity condition we impose on $\gamma_0$ (see Assumption (ref).(ref)) is actually violated, which could explain the larger bias.

table[table omitted — 2,022 chars of source]

On the other hand, the estimated outer bounds exhibit very little bias. Also, in DGP3$_2$, the estimated outer bounds are more precise than the estimated sharp bounds. This already suggests that the corresponding inference may be more precise than that based on the sharp bounds.

Table (ref) presents the coverage rate and length of both confidence intervals. The second confidence interval shows a very good coverage, always greater than 94%. This is the case even with DGP3$_T$, for which Theorem (ref) does not provide any guarantee. Hence, neglecting the variability of $\widehat{\overline{b}}$ does not seem to lead to undercoverage here. The first method also leads to coverage very close to the 95% nominal level.

table[table omitted — 1,177 chars of source]

In terms of length of the confidence intervals, the first method delivers slightly shorter intervals for DGP1 and DGP2, but the difference is very small, for all $T$ and $n$. This was not obvious: when $n\to\infty$, the length of $\text{CI}^1_{0.95}$ becomes smaller than that of $\text{CI}^2_{0.95}$ because $\overline{\delta}- \underline{\delta}<\overline{\delta}^o- \underline{\delta}^o$. Hence, at least with our four DGPs, the difference in the length of the sharp and outer identified sets is small compared to the standard errors of $\widehat{\underline{\delta}}$, $\widehat{\overline{\delta}}$ and $\widehat{\tilde{\delta}}$, even with $n=1,000$.\footnote{Additional simulations show that, as expected, $\text{CI}^1_{0.95}$ becomes shorter than $\text{CI}^2_{0.95}$ as $n$ increases. For $n=10^4$ for instance, we observe a gain of around 10%-15% for DGP1 and DGP2 with $T=2$. Nonetheless, the gain remains negligible for this sample size with $T=3$.} The second method actually leads to shorter intervals with DGP3$_2$ and DGP3$_3$, especially for small $n$.

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Comparison with the linear probability model estimator}

Next, we compare our confidence interval $\text{CI}^2_{0.95}$ with $\text{CI}^{\text{LPM}}_{0.95}$, obtained using the linear probability model (LPM) estimator and the usual standard error accounting for clustering at the individual level. We consider DGP1 but also two incorrectly specified models.\footnote{ Remark that the estimated outer bounds never cross, even if the model is misspecified.} In DGP4, the $(\varepsilon_t)_{t=1,...,T}$ still marginally follow a logistic distribution, so that $\delta_0$ is the same as in DGP1, but these variables are autocorrrelated: their copula is gaussian with a variance matrix $(\Sigma_{s,t})_{1\le s,t\le T}$ satisfying $\Sigma_{s,t}=1/2^{|s-t|}$. In DGP5, we assume instead that the $\varepsilon_t$ are independent over time but $\varepsilon_t\sim \mathcal{N}(0,8/\pi)$. We chose this variance so that again, the AME is the same as in DGP1. We consider $T\in\{2,3,4\}$, $n=1,000$ and two possible values of $\beta_0$, namely $\beta_0=1$ and $\beta_0=2$.

Table (ref) displays the results. It first shows that in the two misspecified DGPs we consider, our confidence interval still performs very well, with a coverage very close to or above 95%. Inference based on the linear probability model estimator also works well when $\beta_0=1$, with a coverage always larger than 93%. However, when $\beta_0=2$, its performance deteriorates, especially for larger $T$. Note that this sensitivity on $\beta_0$ may be more exacerbated with fixed effects. To see this, we consider the same DGP as DGP1 but with $\alpha=0$. Then, the coverage rate of $\text{CI}^{\text{LPM}}_{0.95}$ only decreases from 95% when $\beta_0=1$ to 89% when $\beta_0=2$ with $T=4$, as opposed to the decrease from 94% to 52% that we observe with DGP1.

table[table omitted — 1,365 chars of source]

\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Measuring heterogeneous effects}

Finally, we study the estimator of $\delta_W(w)$, one of the parameter considered in Section (ref) above. We focus on DGP2, defining $W$ as $W=\mathds{1}\left\{|\eta+\nu|>1\right\}$, with $\nu|\eta,X,\varepsilon \sim\mathcal{N}(0,0.3^2)$. In this setup, we have $\delta_W(0)\simeq 0.2317$ and $\delta_W(1)\simeq 0.1575$. We consider below the coverage of the confidence intervals on $\delta_W(0)$ and $\delta_W(1)$ and the length of these confidence intervals. We also consider the test of $\delta_W(1)=\delta_W(0)$, by considering the critical region $\{|T_h|>q^d_\alpha\}$, where the test statistic for homogeneity $T_h$ satisfies $$T_h= n^{1/2} \frac{\widehat{\tilde{\delta}}_W(1)-\widehat{\tilde{\delta}}_W(0)}{\left(\widehat{\sigma}_0^2 +\widehat{\sigma}_1^2\right)^{1/2}}, \quad q^d_\alpha=q_\alpha\left(n^{1/2}\frac{\widehat{\overline{b}}_0+\widehat{\overline{b}}_1}{\left(\widehat{\sigma}_0^2+\widehat{\sigma}_1^2\right)^{1/2}}\right),$$ where $\widehat{\sigma}_d^2$ is a consistent estimator of the asymptotic variance of $\widehat{\tilde{\delta}}_W(d)$. Reasoning as in Section (ref), we obtain, under the null hypothesis, that $\lim_{n\to\infty} P(|T_h|>q^d_\alpha)\le \alpha$, so this test is valid albeit asymptotically conservative.

The results are displayed in Table (ref). As one could expect, larger sample sizes are necessary to detect heterogeneous treatment effects. But already with $n=2,000$, our test of $\delta_W(1)=\delta_W(0)$ has good power when $T=3$.

table[table omitted — 899 chars of source]

\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Conclusion}

We have shown that in the FE logit model, we can partially identify a class of average causal effects including the AME, the ATE and ASF in a simple way, without requiring any optimizaton. Inference based on the outer bounds, in particular, is computationally cheap, does not require any tuning parameter and is shown to fare very well compared to inference based on the sharp bounds. For these reasons, we recommend this approach in practice.

The theory is simple here because only raw moments are involved; but similar results hold with other moments, provided that the corresponding functions form a so-called Chebyshev system krein_1977. Results on these systems have already been applied to the optimal design of experiments dette1997 and the measure of segregation with small units DR17. By drawing further attention on these tools, we hope that this paper will contribute to their use in econometrics.

\linespread{1}\selectfont \linespread{1.3}\selectfont