Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
43,734 characters · 0 sections · 35 citation commands
Fixed Effects Binary Choice Models with Three or More Periods
\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Introduction}
In this paper, we revisit the classical binary choice model with fixed effects. Specifically, let $T$ denote the number of periods and let us suppose to observe, for individual $i$, $(Y_{it}, X_{it})_{t=1, \dotsc, T}$ with
where $\beta_0\in \mathbb R^K$ is unknown and $\varepsilon_{it} \in \mathbb R$ is an idiosyncratic shock. The nonlinear nature of the model and the absence of restriction on the distribution of $\gamma_i$ conditional on $X_i:=(X_{i1}',\dotsc,X_{iT}')'$ renders the identification of $\beta_0$ difficult. rasch1960 shows that if the $(\varepsilon_{it})_{t=1, \dotsc, T}$ are i.i.d. with a logistic distribution, a conditional maximum likelihood can be used to identify and estimate $\beta_0$. chamberlain2010 establishes a striking converse of Rasch's result: if the $(\varepsilon_{it})_{t=1, \dotsc, T}$ are i.i.d. with distribution $F$ and the support of $X_i$ is bounded, $\beta_0$ is point identified only if $F$ is logistic. Other papers have circumvented such an impossibility result by either considering large support regressors manski1987, HonoreLewbel2002 or allowing for dependence between the shocks Magnac2004.
It turns out, however, that chamberlain2010 only proves his result for $T=2$. And in fact, we show that his result does not generalize to $T\geq 3$. Specifically, we consider distributions $F$ satisfying
with $T\geq \tau+1$, $(w_1,...,w_\tau) \in (0,\infty)^\tau$ and $1=\lambda_{1} < \dotsc < \lambda_{\tau}$. We study the identification of $\beta_0$, assuming that $\lambda:=(\lambda_{1},\dotsc,\lambda_{\tau})$ is known. The weights $w_1,\dotsc,w_{\tau}$ remain unknown, thus allowing for much more flexibility on the distribution of $\varepsilon_{it}$ than in the logit case. In particular, it may either be left- or right-skewed, platykurtic or leptokurtic. Our main insight is that for any $F$ satisfying (ref), a conditional moment restriction holds. We also obtain some results on the corresponding identified set $B$. For instance, if, roughly speaking, $X_i$ is continuous, we show that $B$ includes at most $T!-1$ points (2 if $T=3$) and relative marginal effects are point identified. Note that johnson_2004 considers the same family with $\tau=2$ and $T=3$. However, he does not study the general case and does not show any formal identification result based on the corresponding moment conditions.
Obviously, the conditional moment condition can be used to construct GMM estimators. This means, in particular, that $\sqrt{n}$-consistent estimation is possible beyond the logit case when $T>2$, overturning again the impossibility results of chamberlain2010 and Magnac2004. Further, we show that if $T=3$ and mild additional restrictions hold, the optimal GMM estimator based on our conditional moment conditions reaches the semiparametric efficiency bound of the model. Hence, at least when $T=3$, these moment conditions contain all the information of the model.
Finally, we showcase the empirical relevance of our approach by studying whether budget deficits and economic growth affect reelection, revisiting BrenderDrazen2008. The authors investigate this issue using simple and fixed effects logit models. However, the assumption of logistic errors is not warranted, so we consider whether the results are robust to this assumption on the unobserved terms. Our results suggest that the relative effects of budget deficits and economic growth or other variables are fairly robust to the logistic assumption.
Our paper is related to the seminal work of Bonhomme_2012, who develops a unified approach for models where the conditional distribution of $(Y_1,...,Y_T)$ given $(X_i,\gamma_i)$ is parametrized by $\beta_0$, but no restriction on the distribution of $\gamma_i|X_i$ is imposed. In such set-ups, he shows that the identification and estimation of $\beta_0$ depends on the existence of functions $m\neq 0$ satisfying $$\mathbb{E}(m(Y,X,\beta_0)|X,\gamma)=0.$$ This approach has been fruitfully applied to the dynamic logit model by Kitazawa2022 and honore2020. Our paper may be seen as yet another application of this approach, focusing on static models but dropping the logistic assumption.
The remainder of the paper is organized as follows. Section (ref) describes the moment condition we use for identification of $\beta_0$ and establishes some properties of the identified set based on these moments. Section (ref) discusses GMM estimation of $\beta_0$, links it with the semiparametric efficiency bound of the model and discusses the case of unbalanced panel data. Section (ref) is devoted to the application. Section (ref) concludes. All the proofs are collected in the appendix.
\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Identification}
\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{The model and moment conditions}
We drop the subscript $i$ in the absence of ambiguity and let $Y=(Y_{1}',\dotsc, Y_{T}')'$, $X=(X_{1}',\dotsc, X_{T}')'$, $X_t=(X_{1,t},\dotsc ,X_{K,t})'$ , $X_{k,\cdot}=(X_{k,1}, \ldots, X_{k,T})'$, $X_{-k}= (X_{k',t})_{k'\neq k, t=1, \dotsc, T}$, $X_{k,-t}=(X_{k,s})_{s\neq t}$, and $X_{-k,t}=(X_{k',t})_{k'\neq k}$. $\text{Supp}(X)\subset \mathbb R^{KT}$ denotes the support of the random variable $X$. For any set $A\subset \mathbb R^p$ (for any $p\geq 1$), we let $A^*:=A\backslash\{0\}$ and denote by $|A|$ the cardinal of $A$. Hereafter, we maintain the following conditions.
The first condition is also considered in chamberlain2010. The second condition is a standard moment restriction on the covariates. Finally, we exclude in the third condition the case $\beta_0=0$ here. This case can be treated separately, as the following proposition shows.
Condition (ref) can be tested by a specification test on the nonparametric regression of $D=Y_t(1-Y_{t'})$ on $(X_t,X_{t'})$, conditional on the event $Y_t+Y_{t'}=1$. See, e.g., bierens1990consistent or hong1995consistent.
Turning to identification on $\mathbb R^{K*}$, we first recall the impossibility result of chamberlain2010. We say below that $F$ is logistic if $G(u):=F(u)/(1-F(u))=w\exp(\lambda u)$ for some $(w,\lambda)\in\mathbb R^{+*2}$.
This result implies in particular that when $T=2$ and $F$ is not logistic, relative effects $\beta_{0j}/\beta_{0k}$, for $k$ such that $\beta_{0k}\neq 0$, may not be identified. Such relative effects are important as they are equal to relative marginal effects if both $X_{j,t}$ and $X_{k,t}$ are continuous. If only $X_{k,t}$ is continuous (say), $-\beta_{0j}/\beta_{0k}$ still corresponds to a compensating variation.\footnote{To see the first point, note that under Assumptions (ref)-(ref), $$\mu_{k,t}(x):=\frac{\partial \mathbb P(Y_t=1|X_{k,t} = x_{k,t}, X_{k,-t} = x_{k,-t} , X_{-k}=x_{-k})}{\partial x_{k,t}} = \beta_{0k}\mathbb{E}[F'(x_t'\beta_0 + \gamma)|X=x]$$ and thus $\mu_{j,t}(x)/\mu_{k,t}(x) = \beta_{0j}/\beta_{0k}$. Also, $-\beta_{0j}/\beta_{0k}$ corresponds to the change in $X_{k,t}$ necessary to keep $\mathbb P(Y_t=1|X_t,\alpha)$ constant when $X_{j,t}$ increases by one unit.}
The key step in Chamberlain's proof is that if $\beta_0$ is identified for all data generating process satisfying the restrictions of the theorem, the conditional probabilities (conditional on $X$ and $\gamma$) of the four possible trajectories for $(Y_1,Y_2)$ are necessarily affinely dependent. Moreover, by letting $|\gamma|$ tend to infinity, the stable trajectories $(0,0)$ and $(1,1)$ disappear from this relationship. This leads to the following functional equation for $G$:
for all $u\in \mathbb R$, $\alpha$ in an open subset of $\mathbb R$ and some functions $\psi_1(\cdot),\psi_2(\cdot)$ such that for all $\alpha$, $(\psi_1(\alpha),\psi_2(\alpha))\neq (0,0)$. The result follows by noting that the solutions necessarily have the form $u\mapsto w\exp(\lambda u)$.
Equation (ref) relies on the time dummy variable $\mathds{1}\left\{t=2\right\}$. However, the proof of Theorem 2 of chamberlain2010 shows that even without such a dummy variable, (ref) is necessary for the semiparametric efficiency bound not to be zero, or, equivalently, for the existence of regular, root-n consistent estimators of $\beta_0$. In this case, $\alpha$ corresponds to $(x_2-x_1)'\beta_0$, for $(x_1,x_2)$ in a set of positive measure.
In any case, the same reasoning with $T=3$ leads to the following equation for $G$:
for all $u\in \mathbb R$, $\bm{\alpha}:=(\alpha_1,\alpha_2)$ in an open subset of $\mathbb R^2$ and some functions $\psi_k(\cdot)$, $k=1,...,6$, such that for for all $\bm{\alpha}$, $(\psi_1(\bm{\alpha}),...,\psi_6(\bm{\alpha}))\neq (0,...,0)$. We now have $6=2^3 - 2$ terms instead of just $2=2^2-2$, and thus we can expect to have other solutions than just $u\mapsto w\exp(\lambda u)$. And indeed, one can check that if $G$ has the form $u\mapsto w_1\exp(\lambda_1 u)+w_2\exp(\lambda_2 u)$, we can construct $(\psi_1(\bm{\alpha}),\psi_2(\bm{\alpha}),\psi_3(\bm{\alpha}))\neq (0,0,0)$ such that (ref) holds, with $\psi_4(\bm{\alpha})=\psi_5(\bm{\alpha})=\psi_6(\bm{\alpha})=0$. Similarly, if $1/G$ has the form $u\mapsto w_1\exp(\lambda_1 u)+w_2\exp(\lambda_2 u)$, we can construct $(\psi_4(\bm{\alpha}),\psi_5(\bm{\alpha}),\psi_6(\bm{\alpha}))\neq (0,0,0)$ such that (ref) holds, with $\psi_1(\bm{\alpha})=\psi_2(\bm{\alpha})=\psi_3(\bm{\alpha})=0$. Note that there may still be other solutions to (ref) that are increasing and have a limit of $\infty$ (resp. $0$) at $\infty$ (resp. at $-\infty$). The question of identifying all such solutions is left for future research.
Generalizing this reasoning to any $T>2$, we see that combinations of at most $T-1$ exponential functions satisfy the functional restrictions tantamount to (ref) and which render identification of $\beta_0$ possible. This suggests that identification may be achieved for the corresponding family of distribution, which we now formally introduce. Hereafter, $\Lambda_{\tau}$ denotes a subset of $\{(\lambda_1,\dotsc, \lambda_{\tau})\in\mathbb R^{\tau}: 1=\lambda_1<\dotsc< \lambda_{\tau}\}$.
Fixing $\min\{\lambda_{1},\dotsc, \lambda_{\tau}\}$ to 1 is without loss of generality, as we can always multiply $\beta_0$, $\gamma_i$ and $\varepsilon_{it}$ by this factor. If $F$ is of the second type, then one can show that the cdf of $-\varepsilon_{it}$ is of the first type. Thus, up to changing $(Y_{t},X_{t})$ into $(1-Y_{t},-X_{t})$, we can assume without loss of generality, as we do afterwards, that $F$ is of the first type. We shall see that $\tau+1$ periods are sufficient to achieve identification. Hence, we assume, again without loss of generality, that $T=\tau+1$: if $T>\tau+1$, we can always focus on $\tau+1$ periods.
Before describing our identification strategy of $\beta_0$ when $F$ is a generalized logistic distribution, two remarks are in order. First, we obtain our results below irrespective of the vector $w$.\footnote{ We do impose however that all the components of $w$ are non-zero, for normalization purposes. Otherwise, the model with $w=(w_1,0)$ and $\beta_0$, for instance, would be equivalent to the model with $w=(0,w_1)$ and $\beta_0/\lambda_2$. A similar issue arises with, e.g., $w=(w_1,w_2,0)$ if $\lambda_3/\lambda_2=\lambda_2/\lambda_1$.} Hence, in contradistinction with the fixed effect logistic model, we do not fix the distribution of $\varepsilon$, but simply impose that it belongs to a family of distributions indexed by two parameters. Members of this family differ in particular by their skewness and kurtosis. In linear regressions, the residuals are often found to have a skewed distribution with either positive or negative excess kurtosis. Then, there is no reason why the latent variables corresponding to $Y_{it}$ would not exhibit a similar pattern. On the other hand, we do fix $\lambda$. Identification of $\lambda$ could also be of interest but is not addressed in this paper.
Now, the idea behind the identification of $\beta_0$ is to construct a function $m\neq 0$ such that $\mathbb{E}(m(Y,X,\beta_0)|X,\gamma)=0$ almost surely. Thus, as mentioned in the introduction, we apply Bonhomme_2012's general idea of functional differencing. The function $m$ is related to the functions $\psi_k$ in (ref) when $T=3$, and the generalization of (ref) when $T>3$. For any $x=(x'_1,...,x'_T)'\in\mathbb R^{KT}$, let $x_s^{-t}{}=x_s$ if $s<t$, $x_s^{-t}{}=x_{s+1}$ else. We let $$M_t(x;\beta)= (-1)^{t+1} \det
.$$ Then define, for any $(y,x,\beta)\in\{0,1\}^T\times \text{Supp}(X)\times \mathbb R^{K*}$, $$m(y, x; \beta):=\sum_{t=1}^T \mathds{1}\{y_t=1, y_{t'}=0 \;\forall t'\neq t\} M_t(x;\beta).$$ Our first result shows that $m$, indeed, satisfies a conditional moment restriction:
Theorem (ref) shows there exists a known moment condition which potentially identifies $\beta_0$ in a more general model than the logistic one. Also, as the number of periods $T$ increases, the class of distributions $F$ for which $\beta_0$ can be point identified increases. This is consistent with the idea that if $T=\infty$, $\beta_0$ is point identified for any $F$, by using variations in $X_t$ of a single individual. Note however that the class of generalized logistic distribution is not dense for the set of all cdf's: any cdf $F$ belonging to the closure of this class should be such that either $F/(1-F)$ or $(1-F)/F$ is convex. Theorem (ref) also complements the results of chernozhukov2013average showing that bounds on $\beta_0$ for general $F$ shrink quickly as $T$ increases.
Theorem (ref) holds with $T=\tau+1=2$. In such a case, the conditional moment condition can be written $$\mathbb{E}\left[\mathds{1}\{Y_1>Y_2\} \exp(X'_2\beta_0)- \mathds{1}\{Y_2>Y_1\} \exp(X'_1\beta_0)|X\right]=0.$$ This conditional moment generates the first-order conditions of the maximization of the theoretical conditional likelihood, since these the first-order conditions are equivalent to $$\mathbb{E}\left[\frac{(X_1-X_2)}{\exp(X_1'\beta_0) + \exp(X_2'\beta_0)}\left(\mathds{1}\{Y_1>Y_2\} \exp(X'_2\beta_0)- \mathds{1}\{Y_2>Y_1\} \exp(X'_1\beta_0)\right)\right]=0.$$
\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Necessary and sufficient conditions for identification}
The discussion above implies that with $T=\tau+1=2$, $\beta_0$ is identified by (ref) as soon as $\mathbb{E}\left[(X_1-X_2)(X_1-X_2)'\right]$ is nonsingular. We now turn to the more difficult case where $T-1=\tau>1$. Let $B$ denote the identified set of $\beta_0$ obtained with our conditional moment conditions, namely $$B := \left\{b \in \mathbb R^{K*} : \mathbb{E}[m(Y,X;b)|X] = 0 \; \text{a.s.}\right\}.$$ We also denote by $B_k:=\{b_k:\exists b=(b_1,...,b_k,...,b_K)\in B\}$ ($k=1,...,K$) the identified set of $\beta_{0k}$. Our first result shows that $B$ is included in a set depending on the distribution of $X$ only. To define this set, let us introduce
and, for all $b\in\mathbb R^{K*}$, let
Because $D_j(x;\beta_0)=0$ for all $x \in \text{Supp}(X)$, we have $\mathbb P(X \in \mathcal{D}(\beta_0))=0$. The following lemma shows that $B$ is actually included in the set of $b$'s satisfying this property.
This result follows because the moment condition can be written as a weighted sum of the $D_j(x;b)$'s, with positive weights. It shows that $\beta_0$ is identified if for all nonzero $b \neq \beta_0$, we can find some $x\in\text{Supp}(X)$ such that all nonzero $D_j(x;b)$ have the same sign, and the set of such nonzero determinants is not empty.
The set $\widetilde{B}$ is convenient in that it does not depend on the unknown distribution of $\gamma|X$; but it is hard to characterize in general. Nevertheless, we are able to obtain results under either of the conditions below.
The first assumption corresponds to a case where all components of $X$ are discrete. It imposes that for all $k$ and $t$, the support of $X_{k,t}$ includes 0 and at least $T-1$ additional elements. Because we can always replace $X_{k,\cdot}$ by $X_{k,\cdot}-c_k$ for any $c_k\in\mathbb R^T$, the condition $0\in\text{Supp}(X_{k,t})$ for all $k, t$ holds as long as $\cap_{t=1}^T \text{Supp}(X_{k,t})$ is not empty (for all $k$). The second condition imposes that all components of $X_t$ are continuous. It also imposes that for at least two periods $s$ and $t$, $\text{Supp}(X_s)\cap \text{Supp}(X_t)$ is not empty. This last condition holds for instance if $(X_t)_{t\geq 1}$ is strictly stationary.
Whether Assumption (ref) or (ref) holds, Theorem (ref) shows that under-identification is at most finite, namely $|B|<\infty$. This implies that $\beta_0$ is locally identified in the sense that there exists a neighborhood of $\beta_0$ in which the unique solution to the equation $\mathbb{E}[m(Y,X;b) | X]=0$ is $b=\beta_0$. Further, the first result of Theorem (ref) shows that with discrete regressors satisfying Assumption (ref), the “length” of the identified set on $\beta_{0k}$, defined as $$\max_{(b_{1k},b_{2k})\in B_k^2} |b_{1k}-b_{2k}|,$$ cannot exceed $\beta_{0k}(\lambda_{T-1}-1/\lambda_{T-1})$ if $0\not\in B_k$. Note that under Assumption (ref), we can actually identify whether or not $\beta_{0k}=0$ without relying on our conditional moments, since the sign of $\beta_{0k}$ is equal to that of $\mathbb{E}[Y_t - Y_s | X_{-k,s}=X_{-k,t}, X_{k,t}>X_{k,s}]$. The second result on continuous regressors is stronger. It shows that if Assumption (ref) holds, $\beta_0$ is identified up to a scale $c$, with $c$ belonging at most to $(1/\lambda_{T-1}, \lambda_{T-1})$. This directly implies point identification of relative marginal effects. The second result also states that $B$ includes at most $T!-1$ points, and even only 2 points when $T=3$. Importantly, all these result hold for any possible distribution of $\gamma|X$. Thus, point identification may actually hold for many distributions of $\gamma|X$, a point we shall come back to below.
The proof of Theorem (ref) relies on the following ideas. In the first case, when $b_k\not\in \{c\beta_{0k}: c \in \{0\}\cup (1/\lambda_{T-1},\lambda_{T-1})\}$, we construct a subset of $\text{Supp}(X)$ of positive probability such that all nonzero $D_j(x;b)$ have the same sign. The result then follows by Lemma (ref). We use a similar reasoning to prove (ref). To establish the upper bounds on $|B|$, we exploit the fact that the family of exponential functions $(v\mapsto \exp(\zeta_k v))_{k=1,\dotsc, K}$ with distinct coefficients $\zeta_k$ forms a Chebyshev system krein_1977. This property implies that some key determinants do not vanish, and any non-zero “exponential polynomial” $v\mapsto \sum_{k=1}^K \alpha_k \exp(\zeta_k v)$ does not have more than $K-1$ zeros.
We now turn to necessary conditions for identification. The following result is a partial converse of Lemma (ref) and Theorem (ref) above.
The first result shows that for our conditional moments to have any identifying power, there must exist trajectories of $X=(X_1,\dotsc, X_T)$ with distinct values at all periods. Since we focus here on $T\geq 3$, this excludes in particular the case where $X_t$ is binary. More generally, if all components of $X_t$ are binary, one must have $K> \log(T)/\log(2)$ for our moment conditions to have some identifying power. The second result shows that when $T=3$, one cannot improve (ref), at least in a uniform sense over conditional distributions of $\gamma$. Specifically, for any $b\in R$, there exists a data generating process satisfying Assumptions (ref)-(ref) and for which $b\in B$. Note however that failure of point identification at $b$ implies strong restrictions on the distribution of $\gamma|X$. If $b\in B$ with $b\neq \beta_0$, then, for almost all $x$,
where $a_i(\gamma,x)$ is defined in (ref). Namely, the distribution of $\gamma|X$ should satisfy a conditional moment restriction (note that (ref) trivially holds when $b$ is replaced by $\beta_0$, because $D_1(x,\beta_0)=D_2(x,\beta_0)=0$). A violation of (ref) on a set of $x$ of positive measure is sufficient to discard $b$ from $B$.\footnote{Related to this, we establish point identification of $\beta_0$ under some restrictions on the conditional distribution of $\gamma|X$ in a \href{https://arxiv.org/abs/2009.08108v1}{previous version} of the paper.}
\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{GMM estimation}
\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Efficiency bounds}
We now suppose point identification based on (ref) (namely, $B=\{\beta_0\}$) and discuss estimation of $\beta_0$. Let $R(X) = \mathbb{E}[\nabla_{\beta}m(Y,X;\beta_0) | X]$, $\Omega(X) = \mathbb{V}[m(Y,X;\beta_0)|X]$ (so that $\Omega(X) \in \mathbb R$) and define, provided that it exists, $$V_0:=\mathbb{E} \left[ \Omega(X)^{-1}R(X)R(X)'\right]^{-1}.$$ As shown by chamberlain_1987, asymptotically optimal estimators of $\beta_0$ based on (ref) have an asymptotic variance equal to $V_0$. The standard way to construct such estimators consists in two steps: first, one uses the unconditional moment $g(X) m(Y,X;\beta)$ for some $g(\cdot)$ and second, one estimates the optimal instruments $g^\star(X):= R(X)/\Omega(X)$. Such estimators, however, are not consistent if $$\mathbb{E}[g(X) m(Y,X;\beta)]=0 \; \text{ or }\; \mathbb{E}[g^\star(X) m(Y,X;\beta)]=0$$ for $\beta\neq \beta_0$; see dominguez_lobato_2004. Instead, we can use an efficient GMM estimator exploiting the continuum of moment conditions associated with (ref). We refer in particular to Sections 4 in HSU201187 and Section 2.5 in lavergne2013 for the construction of such estimators.
These GMM estimators are optimal among those based on (ref). However, it is not obvious that (ref) actually exhausts all the possible restrictions induced by the model, and therefore that $V_0$ is the semiparametric efficiency bound of $\beta_0$. Theorem (ref) below shows that this is the case for $T=\tau+1=3$ under the following conditions.
The first condition is a mild restriction on $X$. The second condition is a local identifiability condition, which is neither weaker nor stronger than $B=\{\beta_0\}$. The third condition is weaker than that imposed by chamberlain2010, namely $\text{Supp}(\gamma|X)=\mathbb R$. Intuitively, if $\gamma|X$ has few points of support, moments of $\gamma|X$ are restricted, and we may exploit this to produce additional restrictions that would improve an estimation of $\beta_0$ based solely on (ref).
Intuitively, this result states that all the information content of the model is included in the conditional moment restriction $\mathbb{E}[m(Y,X;\beta_0)|X]=0$. It complements, for $T=\tau+1=3$, the result of Hahn1997, which states that the conditional maximum likelihood estimator is the efficient estimator of $\beta_0$ if $F$ is logistic. Note however that we cannot compare his bound with ours in the logistic case: for this distribution, $w_2=0$, and for identification reasons, this case is excluded from our family of generalized logistic distributions with $\tau=2$. We refer to Footnote (ref) above for more detials about this.
\@startsection{subsection}{2}{0mm}{-1.2\baselineskip}{1\baselineskip}{\normalfont}{Unbalanced panel}
In many applications, as that considered below, panel data are unbalanced. To handle this case, we can simply consider, for each individual, all possible subsets of periods of size $\tau+1$ and form the corresponding moment conditions. Specifically, suppose that the set of periods available for individual $i$ is $\mathcal{T}_i \subset \{1,...,T\}$. Thus, we observe the sample $((Y_{it},X_{it})_{t\in \mathcal{T}_i})_{i=1,...,n}$. Let us assume that the selection of periods is (conditionally) exogenous, namely
Then, we basically get back to the case $T=\tau+1$ by considering the moment vector
for some function $g(X_{it_1},...,X_{it_{\tau+1}}) \in \mathbb R^L$, with $L\ge K$. Condition (ref) ensures that $\mathbb{E}\left[\psi(Y_i,X_i,\mathcal T_i,\beta_0)\right]=0$. Then, we can consider the GMM estimator
for some symmetric positive definite $\widehat{W}$. This idea also applies to balanced panel data for which $T>\tau+1$. In such a case, $\mathcal{T}_i=\{1,...,\tau+1\}$ and (ref) automatically holds.
\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Application to BrenderDrazen2008}
BrenderDrazen2008 study how budget deficits and economic growth affect reelection. To this end, they gather data from multiple sources on 74 countries, over the period 1960-2003. They use two definitions for their binary outcome variable REELECT, one where reelection is defined in a “narrow” sense and another where it is “expanded”, following here their terminology. This also leads to two different samples, as REELECT may be missing in the narrow sense but equal to 0 in the expanded sense. The covariates related to budget deficits are BALCH_term and BALCH_ey. BALCH_term corresponds to the change in ratio of the central government's balance to GDP over the term in office. BALCH_ey is the change in the balance/GDP ratio between the year preceding the election and the election year. The variable GDPPC_gr is the average annual growth rate of real GDP per capita between two election years. The authors also include in their models two controls, namely a dummy for a new democracy and a dummy of having a majoritarian electoral system. We refer to BrenderDrazen2008 for more details about the data.
In their main specification, BrenderDrazen2008 consider a simple logit model, see Table 2 therein. Then, as a robustness check (see their Table 3), they estimate a fixed effect logit model. They show that their main results are robust to including fixed effects. However, the assumption of logistic errors is not warranted, so we investigate whether the results are robust to this assumption, by considering instead the family of generalized logistic distribution, with $\tau=2$. We focus on the sample of developed countries as the sample of less developped countries is very small, and thus leads to noisy estimates. Note that the data are not balanced at all: some countries are only observed for $4$ periods in the narrow sample (resp. $5$ in the expanded sample), while others are observed over $13$ (resp. $14$) periods. We thus apply the procedure mentioned in Section (ref). The vector of instruments $g(X_{it_1},X_{it_2},X_{it_3})$ is simply the list of the corresponding 15 variables (as $X_{it}\in\mathbb R^5$), demeaned over these three periods. We consider $\lambda_2 = 1.2, 1.4, 1.6$ and $1.8$. We do not consider larger values of $\lambda_2$ as they seem to lead to numerical instabilities.\footnote{This may be because $|M_t(x;\beta)|$ increases quickly with $\lambda_2$, due to the exponential function.} Finally, as the GMM objective function may have local optima, we consider 200 random initial points and pick the vector of parameters minimizing the corresponding final objective function.
The results are presented in Table (ref). Because the coefficients themselves are not comparable, we focus on the sign of BALCH_ey and on the relative effects with respect to BALCH_ey; note that we were able to recover the exact same estimates as BrenderDrazen2008 in their Tables 2 and 3. We choose BALCH_ey as the reference variable for relative effects because its coefficient should not be 0, and it has the largest t-test on the logit and fixed effect logit model. For the three methods, the $t$-statistics of relative effects under the null hypothesis are obtained using the estimated asymptotic variance of $\widehat{\beta}$.
Overall, at least two important results seem robut to the distributional assumption on the unobserved terms. First, the sign of BALCH_ey is always positive. Second, the relative effect of BALCH_term and BALCH_ey remain quite stable when considering our FE generalized logistic model, with fluctuations between 0.27 and 0.52 depending on the sample and value of $\lambda_2$ that we consider. At the 10% level, we cannot reject that the effect of BALCH_term is actually 0, except in the narrow sample with $\lambda_2=1.8$. But the test was already close to not being rejected with the FE logit model on the narrow sample (p-value=0.097), and not rejected with the simple and FE logit models based on the expanded sample (p-values=0.124 and 0.204 respectively). So the most important results seem overall robust to the change of specification we consider. Other results fluctuate slightly more: the fact of being a new democracy had a positive and borderline significant effect with the expanded sample (p-value=0.099). It is not significant anymore with our model, the coefficient being sometimes even negative.
\@startsection{section}{2}{0mm}{-1.5\baselineskip}{1\baselineskip}{\normalfont}{Conclusion}
This paper studies the identification and root-n estimation of the common slope parameter in a static panel binary model with exogenous and bounded regressors. We first show that when $T\geq 3$ and the unobserved terms belong to a family of generalized logistic distribution, a conditional moment restriction holds. Then, we study the identified set corresponding to these restrictions. In particular, under a restriction on the distribution of covariates only, relative effects are point identified, no matter the distribution of the individual effect. Our identification results lead to a GMM estimator that reaches the semiparametric efficiency bound when $T=3$. Estimating this model may serve as a robustness check for the fixed effect logit model, something we illustrate in the application.
Our paper also leaves a few questions unanswered. A first one is whether the family of $F$ considered here is the only one for which point identification can be achieved. Another one is whether the GMM estimator still reaches the semiparametric efficiency bound when $T>3$. Both questions raise difficult issues and deserve future investigation.