EconBase
← Back to paper

New possibilities in identification of binary choice models with fixed effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

80,047 characters · 8 sections · 67 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

New possibilities in identification of binary choice models with fixed effects

\global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long \global\long

abstractWe study the identification of binary choice models with fixed effects. We propose a condition called sign saturation and show that this condition is sufficient for identifying the model. In particular, this condition can guarantee identification even when all the regressors are bounded, including multiple discrete regressors. We also establish that without this condition, the model is not identified unless the error distribution belongs to a special class. Moreover, we show that sign saturation is also essential for identifying the sign of treatment effects. Finally, we introduce a measure for sign saturation and develop tools for its estimation and inference.

Key words: identification, panel model, binary choice, fixed effects

Introduction

This paper considers panel models with binary outcomes in the presence of fixed effects. These models are convenient in economic analysis as they allow for fairly general unobserved individual heterogeneity. The nonlinear nature of the binary choice models makes it difficult to eliminate the fixed effects by differencing. Since we view the fixed effects or their conditional distribution as a nuisance parameter, the identification of the parameter of interest becomes tricky. In this paper, we provide new results and insights on what drives the identification of the coefficients on the regressors and what this means for applied work.

Consider independent and identically distributed (i.i.d) observations $\{(Y_{i},X_{i})\}_{i=1}^{n}$ with $Y_{i}=(Y_{i,0},Y_{i,1})$ and $X_{i}=(X_{i,0},X_{i,1})$ from the following model \[ Y_{i,t}=\mathbf{1}\{X_{i,t}'\beta+\alpha_{i}\geq u_{i,t}\}\qquad t\in\{0,1\}, \] where the fixed effects are represented by the scalar variable $\alpha_{i}$ and $\beta$ is a non-random vector of coefficients. This is a semiparametric model as the distribution of $\alpha_{i}$ given $X_{i}$ is unrestricted. For notational simplicity, we drop the $i$ subscript in the rest of the paper. Therefore, we write $Y=(Y_{0},Y_{1})$ and $X=(X_{0},X_{1})\in\mathcal{X}_{0}\times\mathcal{X}_{1}$ with

equation[equation omitted — 103 chars of source]

In this paper, we mainly focus on the identification of $\beta$ but we will also discuss quantities related to treatment effects. There are roughly two approaches to identifying $\beta$, depending on whether or not we impose a parametric model on the distribution of $u_{t}$ given $(X,\alpha)$. Methods without parametric assumptions on the error distribution are typically based on manski1987semiparametric and maximum-score-type estimation. This approach only assumes that the distribution $u_{t}\mid(X,\alpha)$ does not depend on $t$. In contrast, the more parametric approach relies on additional assumptions on the functional form of the distribution $u_{t}\mid(X,\alpha)$. For example, perhaps the most popular parametric assumption is that $u_{t}$'s are i.i.d logistic errors across $t$ and are independent of $(X,\alpha)$. This approach typically adopts an estimation scheme based on the (conditional) likelihood.

Despite the strong restrictions on the functional form, the approach relying on parametric restrictions has received considerable attention arguably due to identification reasons. The literature has pointed out the widespread identification failure, such as arellano2011nonlinear. In particular, identification seems infeasible outside the logistic case unless the support of $X$ is unbounded. For example, Assumption 2 in manski1987semiparametric requires the unboundedness of at least one component of $W=X_{1}-X_{0}$: for $W=(W_{1},...,W_{K})'$ and $\beta=(\beta_{1},...,\beta_{K})'$, there exists $k$ such that $\beta_{k}\neq0$ and the conditional distribution $W_{k}\mid(W_{1},...,W_{k-1},W_{k+1},...,W_{K})$ has support equal to $\mathbb{R}$ almost surely. Theorem 1 of chamberlain2010binary goes even further: for bounded $X$, the identification fails in certain regions of the parameter space if no further restrictions are imposed on the distribution of the error term $u_{t}$.

commentThe literature has pointed out the widespread identification failure, such as arellano2011nonlinear.

In this paper, we provide a more precise picture than chamberlain2010binary by showing that identification without parametric assumptions on the error distribution is quite possible even for bounded $X$. The main motivation of this paper is to explore new possibilities without functional-form assumptions. This raises concerns about stability. If identification crucially hinges on the imposed functional form, the reliability of the analysis could be a concern. After all, if the model parameter is unidentified under every distribution function other than the logistic one, how much should we trust this model? Hence, it is helpful to explore a more robust setting. In this direction, this paper considers the identification issue and leaves the estimation problem to future research. We make the following contributions.

First, we provide tight and simple identification conditions without parametric restrictions on the error distribution. We show that a sufficient condition for identification is what we refer to as sign saturation (Assumption (ref)), which states that $P\left(E(Y_{1}-Y_{0}\mid X)>0\right)$ and $P\left(E(Y_{1}-Y_{0}\mid X)<0\right)$ are both strictly positive. This condition can hold with bounded regressors $X$ and allow for multiple discrete regressors and interaction terms. Moreover, we show that this is also a necessary condition for identification unless the distribution of $u_{t}$ is in a special class. Therefore, although we know (from chamberlain2010binary) that for bounded regressors the identification fails at some values of $\beta$, our results pinpoint these values: the identification fails exactly at points where the sign saturation fails. Therefore, whether the regressors are bounded is not what really drives the identification. The identifiability is more closely related to the sign saturation condition, which is simple and intuitive.

A key advantage of the proposed sign saturation condition is that it is stated in terms of observed variables only and does not involve the parameter that we are trying to identify. Hence, it can be used to answer the question of whether identification holds under the distribution of observed data, thereby making identification directly testable. This is different from the manski1987semiparametric-type condition. For example, suppose that $K=2$, $W_{1}$ is bounded and the conditional distribution $W_{2}\mid W_{1}$ has support $\mathbb{R}$ almost surely. The identification condition in manski1987semiparametric becomes $\beta_{2}\neq0$. However, how can we check $\beta_{2}\neq0$ when the identifiability of $\beta$ is not yet established? It does not seem obvious how to do so with the data. In contrast, under the sign saturation condition, we check statements only on observed data: $P\left(E(Y_{1}-Y_{0}\mid X)>0\right)$ and $P\left(E(Y_{1}-Y_{0}\mid X)<0\right)$.

Second, we show that the sign saturation condition is also sufficient and necessary for learning the sign of the marginal effects. Even though the magnitude of different notions of the marginal effects is usually not point identified, the sign of the effects is typically the same as the sign of a component of $\beta$. It turns out that outside a special class of distributions, if the sign saturation condition fails, we cannot guarantee to distinguish between zero effects and strictly positive effects. In many empirical studies, one important task is to check whether the identified set for the effects includes zero. Therefore, it is worthwhile to check the sign saturation condition before exploring additional assumptions that give bounds to treatment effects.

Third, we introduce a tool that can be used to assess the sign saturation condition empirically. We show that $E\max\{P(Y_{1}-Y_{0}\mid X),0\}$ and $E\min\{P(Y_{1}-Y_{0}\mid X),0\}$ can be directly estimated from the data and confidence intervals can be constructed by a simple bootstrap without any tuning parameters such as bandwidth or number of basis functions. For bounded regressors, the existing literature highlights the risk of identification failure: the identification possibly (not definitely) fails. This generic warning may have discouraged the application of fixed effect models in practice. The proposed tool serves as a more accurate diagnosis so that the identifiability can be assessed from the data.

commentIt turns out that the computation is already familiar to applied researchers in this area as it only involves computing the usual maximum score estimator and its bootstrapping. We are not aware of any existing tools for checking identification in this model.
commentOne important objective of our work is to provide robust tools for empirical research. Although conditional maximum likelihood or any other approach based on logistic errors is popular in applied work, we argue that this assumption can be problematic in applications. If so, it might be prudent to leave the error distribution nonparametric. Since estimation and inference based on maximum score estimates are already available (e.g., kim1990cube, seo2018local and cattaneo2020bootstrap), we only need to make sure that we can reasonably assume the identification. This paper fills this very need: tight and testable conditions for identification.

Our work is closely related to the literature of binary response models with logistic errors. rasch1960probabilistic, andersen1970asymptotic and chamberlain1980analysis consider the estimation and inference in this case. The logistic link function is also singled out in chamberlain2010binary as the “nice” link function: the identification does not require unbounded regressors. Recently, the work of Mugnier2009.08108 points out that if there are more than 2 time periods, then the nice link function can be extended to a generalized version of the logistic distribution. This suggests that imposing a special class of parametric structures is a useful strategy of achieving identification and this special class seems to be the logistic functions. We provide new insight on this. We show that this special class also includes those such that $\dot{G}(\cdot)$ is periodic, where $\dot{G}(\cdot)$ is the derivative of $\ln\frac{F(\cdot)}{1-F(\cdot)}$ and $F(\cdot)$ is the distribution function of $u_{t}$. If $F(\cdot)$ is logistic, then $\dot{G}(\cdot)$ is a constant function, which is clearly periodic. However, we find that identification is possible for any periodic $\dot{G}(\cdot)$. Although chamberlain2010binary tells us that for non-logistic distributions, identification fails in an open neighborhood, we show that for any periodic $\dot{G}(\cdot)$, identification must hold in another open neighborhood. Therefore, it seems to us that the truly problematic distributions are those with non-periodic $\dot{G}(\cdot)$. Indeed, outside this special class of periodic $\dot{G}(\cdot)$, we show that identification is impossible at every point if the sign saturation condition is not satisfied. Thus, for generic distribution functions, sign saturation is a necessary condition for identification. We summarize the relation between identifiability and error distribution in Table (ref).

commentLogistic errors are also assumed in the study of identification in dynamic models, including recent works of Honore2005.05942, khan2020identification among many others.

Our work is also closely related to works that do not assume logistic errors. First, some results in the literature allow for bounded regressors but require all regressors to be continuous. For example, Assumption 3.3 of shi2018estimating gives a sufficient condition for identification under bounded support. As commented in the paper, their assumption essentially requires all regressors to be continuous. Continuous distribution on all the regressors are also required by Assumption 6 of toth2017 and Assumption 4' of gao2023logical. Examples on nonseparable models include hoderlein2012nonparametric, chernozhukov2015nonparametric and chernozhukov2019nonseparable; their arguments rely on the derivatives with respect to $X$ and thus $X$ needs to be continuous. Here, we allow for one or multiple discrete regressors, which are important in empirical studies as many treatment variables are binary. Second, Corollary 4.1 of horowitz2009semiparametric considers a related model in the cross-sectional setting. It is possible to translate this to the panel data in terms of conditional densities, but these are still not the most general conditions; for example, some strictly increasing continuous distribution functions have zero density almost everywhere, e.g., salem1943some and takacs1978increasing. Most importantly, none of the aforementioned works show whether their conditions are necessary. We contribute to the literature by finding what must be assumed and developing empirical tools to assess it.

commentFor this model, proposes a condition for the identification of $\beta$. Under the normalization of $\beta=(1,\beta_{-1}')'$ and the corresponding partition $X=(X_{1},X_{-1}')'$, the key requirement is that there exist some small $\delta>0$ and some set $N_{\delta}$ with $P(X_{-1}\in N_{\delta})>0$ such that for almost every $x_{-1}\in N_{\delta}$, the density of $X_{1}+X_{-1}'\beta_{-1}$ conditional on $X_{-1}=x_{-1}$ is everywhere positive on $[-\delta,\delta]$. One might wonder if a similar condition can be stated for the panel setting. Like manski1987semiparametric, this condition involves the unknown parameter $\beta$: to determine the identifiability of $\beta$ from the observed data, we need to check the density of $X'\beta$ condition on $X_{k}$ for some $k$. As discussed above, the sign saturation condition is stated only in terms of observed data and is thus directly testable. The proposed test also has no tricky tuning parameters or curse of dimensionality that plague many nonparametric estimators, especially in multiple dimensions. Another technical issue is that statements of densities might not give the most general conditions; for example, some strictly increasing distribution functions have zero density almost everywhere, e.g., takacs1978increasing.
comment\begin{rem} We note that this condition does not cover all the identified cases. For example, consider $X=(1,Z)'$ and $\beta=(0,3)'$, i.e., $X'\beta=1+3Z$. If we set $X_{1}=1$ and $X_{-1}=Z$, then $X'\beta$ conditional on $X_{-1}$ has a point mass; if we set $X_{1}=Z$ and $X_{-1}=1$, then Due to the normalization, we can only condition on $Z$ because the coefficient on $1$ is zero; however, $X'\beta=1+3Z$ conditional on $Z$ has no density. Moreover, \end{rem}
commentNotice that this condition rules out a discrete $X_{1}$; if $X_{1}$ is discrete, then the density of $X_{1}+X_{-1}'\beta_{-1}$ conditional on $X_{-1}=x_{-1}$ does not exist. This can be viewed as a conditional sign saturation, namely, conditional on $X_{-1}$, $X'\beta$ has sign saturation. Our sign saturation is an unconditional statement. The conditional sign saturation is stronger than necessary even for continuous $X$; in Lemma (ref) in the appendix, we give an example with $X=(X_{1},X_{2})'$ in which almost surely, one of $P(X'\beta>0\mid X_{1})$ and $P(X'\beta<0\mid X_{1})$ is exactly zero and one of $P(X'\beta>0\mid X_{2})$ and $P(X'\beta<0\mid X_{2})$ is exactly zero.
comment\begin{lem} Let $X=(X_{1},X_{2})'$ with $X_{2}=X_{1}\xi$, where ${\rm supp}(X_{1})=[-1,1]$, ${\rm supp}(\xi)=[1/2,1]$ and $X_{1}$ and $\xi$ are independent. Assume that $\beta=(\beta_{1},\beta_{2})'$ with $\beta_{1},\beta_{2}>0$. Then \\ (1) for any $x_{1}\in(-1,1)$, one of $P(X'\beta>0\mid X_{1}=x_{1})$ and $P(X'\beta<0\mid X_{1}=x_{1})$ is exactly zero. \\ (2) for any $x_{2}\in(-1,1)$, one of $P(X'\beta>0\mid X_{2}=x_{2})$ and $P(X'\beta<0\mid X_{2}=x_{2})$ is exactly zero. \end{lem} \begin{proof} We first show the result for $x_{1}$. Notice that $X'\beta=X_{1}(\beta_{1}+\beta_{2}\xi)$. If $x_{1}>0$, then $P(X'\beta<0\mid X_{1}=x_{1})=P(\beta_{1}+\beta_{2}\xi<0\mid X_{1}=x_{1})=0$ since ${\rm supp}(\xi)=[1/2,1]$ and $\beta_{1},\beta_{2}>0$. If $x_{1}<0$, then $P(X'\beta>0\mid X_{1}=x_{1})=P(\beta_{1}+\beta_{2}\xi<0\mid X_{1}=x_{1})=0$. If $x_{1}=0$, then both $P(X'\beta>0\mid X_{1}=x_{1})$ and $P(X'\beta<0\mid X_{1}=x_{1})$ are zero. To see the result for $x_{2}$, notice that $X_{1}=X_{2}/\xi$. If $x_{2}>0$, then \[ P(X'\beta<0\mid X_{2}=x_{2})=P(\beta_{1}X_{2}/\xi+\beta_{2}X_{2}<0\mid X_{2}=x_{2})=P(\beta_{1}/\xi+\beta_{2}<0\mid X_{2}=x_{2})=0. \] If $x_{2}<0$, then \[ P(X'\beta>0\mid X_{2}=x_{2})=P(\beta_{1}X_{2}/\xi+\beta_{2}X_{2}>0\mid X_{2}=x_{2})=P(\beta_{1}/\xi+\beta_{2}<0\mid X_{2}=x_{2})=0. \] If $x_{2}=0$, then both $P(X'\beta>0\mid X_{1}=x_{1})$ and $P(X'\beta<0\mid X_{1}=x_{1})$ are zero. \end{proof}
commentMoreover, checking this condition on density in practice seems to require (1) finding a correct normalization (which component of $\beta$ is positive or at least non-zero) and (2) multiple-dimensional nonparametric estimation, which might involve tricky tuning parameters. In contrast, the sign saturation condition proposed here is stated in terms of observed variables without any unknown parameters and can be tested in a straight-forward way.

Our necessity results complement those in pakes2024moment. By Proposition 3 therein, $\arg\max_{\beta}E(Y_{1}-Y_{0})\cdot\mathbf{1}\{W'\beta\geq0\}$, the set of maximizers of the maximum score criterion function, is the sharp identification set when we only assume the stationarity of the distribution $u_{t}\mid(X,\alpha)$. Our results show that this sharp identified reduces to a singleton up to scale under the sign saturation condition. A natural question is the necessity of this condition, especially if we are willing to impose additional assumptions, namely i.i.d errors that are independent of $(X,\alpha)$. By classical results, we know that under logistic errors, $\arg\max_{\beta}E(Y_{1}-Y_{0})\cdot\mathbf{1}\{W'\beta\geq0\}$ is not the sharp identified set. Our results in Section (ref) further complete the puzzle by characterizing point identification as the periodicity of $\dot{G}(\cdot)$.

commentHowever, the proof of their Theorem 2 shows that every point in the identified set can be rationalized by a stationary error $u_{t}\mid X$ without any fixed effects $P(\alpha=0)=1$. Therefore, the class of models with arbitrary stationary errors is so large that the fixed effects are not important any more. In this sense, our sign saturation condition is necessary for point identification up to scale. \footnote{One can show that the maximum score criterion function is maximized at a unique point (up to scale) under the sign saturation condition if the support of $Z$ is convex with non-empty interior.}

Identification of coefficients

In the setting of ((ref)), manski1987semiparametric identifies $\beta$ with the distributional stationarity condition:

assumptionThe distribution of $u_{t}\mid(X,\alpha)$ does not depend on $t\in\{0,1\}$. Let the cumulative distribution function (c.d.f) of this common distribution be $F(\cdot|x,\alpha)$. Assume that for any $(x,\alpha)$, $F(\cdot\mid x,\alpha)$ is a strictly increasing function on $\mathbb{R}$.

Assumption (ref) rules out dynamic models. In this paper, we focus on static models with two time periods. Without any restriction on the magnitude of $\alpha$ and $u_{t}$ in ((ref)), it is impossible to identify the magnitude of $\beta$ but the following result by manski1987semiparametric gives the identification of the “direction” of $\beta$; in other words, $\beta$ is identified up to scaling.

lem[manski1987semiparametric] Let Assumption (ref) hold. Then with probability one, \begin{equation} {\rm sgn}\left(E(Y_{1}-Y_{0}\mid X)\right)={\rm sgn}\left(W'\beta\right), \end{equation} where $W=X_{1}-X_{0}$ and ${\rm sgn}(\cdot)$ is defined as ${\rm sgn}(t)=1$ for $t>0$, ${\rm sgn}(t)=-1$ for $t<0$ and ${\rm sgn}(0)=0$.

The key idea of our result is based on the condition in ((ref)). Clearly, the conditional mean function $E(Y_{1}-Y_{0}\mid X)$ is identified. For simplicity, suppose that $P(E(Y_{1}-Y_{0}\mid X)=0)=0$. Then ((ref)) implies that there is a hyperplane of $W$ that perfectly classifies ${\rm sgn}(E(Y_{1}-Y_{0}\mid X))$. This is an ideal support vector machine (SVM) as illustrated in Figure (ref), where the red and blue dots represent $-1$ and $1$, respectively for ${\rm sgn}(E(Y_{1}-Y_{0}\mid X))$.\footnote{In komarova2013binary, the geometry based on SVM is also to study binary response models with a median restriction; the model there does not have a panel structure. We thank Christopher Walker for this reference. } If we want to identify the classification boundary in the SVM, we obviously need to have both classes (red and blue) in the data. Since these two classes correspond to $E(Y_{1}-Y_{0}\mid X)>0$ and $E(Y_{1}-Y_{0}\mid X)<0$, this simple requirement is formalized as the following sign saturation condition.

figure[figure omitted — 357 chars of source]
assumptionBoth $P(E(Y_{1}-Y_{0}\mid X)>0)$ and $P(E(Y_{1}-Y_{0}\mid X)<0)$ are strictly positive.

From Figure (ref), it is also clear that in addition to having both classes, we also need the data points to be “dense” around the hyperplane. For simplicity, we will require that the support of $W$ is convex and has no-empty interior so the support of $W$ is dense. This is all we need for uniquely determining the hyperplane and there is nothing about unbounded supports. We will formalize this intuition in in the next subsection and extend the results to the case of multiple discrete regressors in Section (ref).

remPhrasing the problem in terms of support vector machines makes it easy to see that when all the regressors are discrete, we cannot in general expect to point identify $\beta$ up to scale. To see this, notice that by pakes2024moment, the sharp identified set for $\beta$ is $\arg\max_{b}E(Y_{1}-Y_{0}){\rm sgn}(W'b)$, which can be written as \[ \arg\max_{b}E\left[\left|E(Y_{1}-Y_{0}\mid X)\right|\cdot{\rm sgn}(W'\beta)\cdot{\rm sgn}(W'b)\right]. \] Therefore, if ${\rm sgn}(W'\tilde{\beta})={\rm sgn}(W'\beta)$ with probability one, then $\tilde{\beta}$ is in the identified set. When the regressors are discrete, we can in general move $\beta$ a bit and retain the same ${\rm sgn}(W'\beta)$. As illustrated in Figure (ref), the black and gray lines both perfectly classify the red and blue dots and thus represent two points in the identified set that are not scalar multiple of each other.
figure[figure omitted — 382 chars of source]

The sign saturation condition is sufficient for identification

Following chamberlain2010binary, we assume that one of the regressors is binary in that it is equal to zero at $t=0$ and is equal to one at $t=1$; in other words, one component of $W=X_{1}-X_{0}$ is one. Thus, without loss of generality, we can partition $W=(Z',1)'\in\mathbb{R}^{K}$. If all the regressors are continuous, then the result can still be applied because the coefficient for the binary regressor is allowed to be zero, see Corollary (ref). The first main result of this paper is to show that the sign saturation condition (together with additional weak assumptions) is enough to guarantee identification up to scaling. To state the formal result, we first recall the definition of the support of a random variable (or its corresponding probability measure): the support of probability measure $\lambda$ is the smallest closed set $A$ such that $\lambda(A)=1$, e.g., page 227 of dudley_2002.

thmLet Assumption (ref) hold. Partition $W=(Z',1)'\in\mathbb{R}^{K}$. Suppose that the support of $Z$ is convex and has non-empty interior. Let Assumption (ref) hold. Then $b=\mu\beta$ for some $\mu>0$ if and only if $R(b)=0$, where \begin{equation} R(b)=P\left({\rm sgn}\left(E(Y_{1}-Y_{0}\mid X)\right)\neq{\rm sgn}(W'b)\right). \end{equation} Therefore, $\beta$ is identified up to scale.

The requirement on the distribution of $Z$ is mild. Let $\mathcal{Z}$ be the support of $Z$. First, we do not require $\mathcal{Z}$ to be an unbounded set. Second, the distribution of $Z$ does not have to admit a density and “atoms” (or point masses) are allowed. Theorem (ref) imposes this via the convexity and non-empty interior of $\mathcal{Z}$. When $\mathcal{Z}$ is not convex, the result is still useful: as long as $\mathcal{Z}$ contains a convex subset with non-empty interior, we can apply the result on this subset by restricting the sample, i.e., the sign saturation condition holds on the restricted sample. If there are multiple discrete variables, then we set $Z$ to be the difference in continuous variables, see Section (ref).

By Theorem (ref), to check whether $b$ is a rescaled version of the true $\beta$, we only need to check $R(b)$ from the distribution of the observed data. Notice that the identification does not assume that $u_{t}$'s are independent across $t$; this is because Assumption (ref) only requires them to have the same marginal distribution given $(X,\alpha)$. To give some intuition on the identifying power of sign saturation under bounded $\mathcal{Z}$, we use the following simple example to illustrate why the criterion function $R(\cdot)$ fails to deliver identification when Assumption (ref) is not satisfied.

exampleSuppose that $W=(Z,1)$, where $Z$ is a continuous random variable with support $[-1,1]$. For simplicity, we consider $\beta=(\beta_{1},\beta_{2})'$ and $b=(b_{1},b_{2})'$ with $\beta_{1},b_{1}>0$. The question is whether we can derive $b=\mu\beta$ for some $\mu>0$ from $R(b)=0$. Here, identifying $\beta$ up to scale is equivalent to identifying $\beta_{2}/\beta_{1}$. After straight-forward calculations, we can see that \[ R(b)=P\left(Z\in\left(-\frac{\beta_{2}}{\beta_{1}},-\frac{b_{2}}{b_{1}}\right)\bigcup\left(-\frac{b_{2}}{b_{1}},-\frac{\beta_{2}}{\beta_{1}}\right)\right). \] Suppose that the sign saturation condition fails, say $P(W'\beta>0)=0$, which implies that $-\beta_{2}/\beta_{1}\geq1$. In this case, the set $\left(-\frac{\beta_{2}}{\beta_{1}},-\frac{b_{2}}{b_{1}}\right)\bigcup\left(-\frac{b_{2}}{b_{1}},-\frac{\beta_{2}}{\beta_{1}}\right)$ has empty intersection with $[-1,1]$ whenever $-b_{2}/b_{1}>1$. In other words, if $P(W'\beta>0)=0$, then $R(b)=0$ whenever $-b_{2}/b_{1}>1$; as a result, we cannot always verify $b_{2}/b_{1}=\beta_{2}/\beta_{1}$ through $R(b)$. On the other hand, if the sign saturation condition holds, then we have $-1<\beta_{2}/\beta_{1}<1$. In this case, the intersection of $\left(-\frac{\beta_{2}}{\beta_{1}},-\frac{b_{2}}{b_{1}}\right)\bigcup\left(-\frac{b_{2}}{b_{1}},-\frac{\beta_{2}}{\beta_{1}}\right)$ and $[-1,1]$ always has non-empty interior whenever $b_{2}/b_{1}\neq\beta_{2}/\beta_{1}$. Consequently, $R(b)=0$ if and only $b_{2}/b_{1}=\beta_{2}/\beta_{1}$. \qed

The following example demonstrates how Theorem 1 gives identification conditions for $\beta$ with interaction terms.

example[Interaction terms] Let $D_{t}=\mathbf{1}\{t=1\}$ and consider a continuous variable $H_{t}$. Suppose that we are interested in the treatment effect of $D_{t}$ controlling for $H_{t}$. For a more flexible specification, we include the interaction $D_{t}H_{t}$. Then $X_{t}=(D_{t},H_{t},D_{t}H_{t})'\in\mathbb{R}^{3}$ with the corresponding partition $\beta=(\beta_{1},\beta_{2},\beta_{3})'\in\mathbb{R}^{3}$. For simplicity, suppose that the support for $(H_{0},H_{1})'$ is $[-1,1]\times[-1,1]$. Then $Z=(H_{1},D_{1}H_{1})'-(H_{0},D_{0}H_{0})'=(H_{1}-H_{0},H_{1})'$. One can easily verify that $\mathcal{Z}$ is convex and has non-empty interior.\footnote{To see the convexity, simply notice that $Z=\begin{pmatrix}-1 & 1\\ 0 & 1 \end{pmatrix}\begin{pmatrix}H_{0}\\ H_{1} \end{pmatrix}$ and $(H_{0},H_{1})'$ has a convex support.} By Lemma (ref), straight-forward calculations show that the sign saturation condition becomes $|\beta_{2}+\beta_{3}|+|\beta_{3}|>|\beta_{1}|$. For example, this condition holds when $\beta_{1}=\beta_{3}=0$ and $\beta_{2}\neq0$. Therefore, to test the null hypothesis of “zero effects whatsoever”, it suffices to have a relevant control variable. \qed

We now discuss an important special case with only continuous covariates.

corLet Assumption (ref) hold. If the interior of the support of $X_{1}-X_{0}$ contains $\boldsymbol{0}_{\dim(X_{t})}$ (zero in $\mathbb{R}^{\dim(X_{t})})$ and $\beta\neq\boldsymbol{0}_{\dim(X_{t})}$, then $\beta$ is identified up to scale.

A comparison between Example (ref) and Corollary (ref) highlights the role of discrete variables in identification. If all the covariates are continuous and the change in covariates contains zero in the support, then we can generally expect identification (up to scale) of $\beta$, regardless of whether the covariates are bounded. In contrast, if the covariates include discrete variables (such as the time fixed effect in chamberlain2010binary), we no longer have such a generic guarantee of identification.

\textcolor{blue}Extension to multiple discrete regressors

In this subsection, we consider the case of multiple discrete regressors.\footnote{I thank Xavier D'Haultf{\oe}uille for a discussion of this case.} We partition $X_{t}=(X_{(1),t}',X_{(2),t}')'$, where $X_{(1),t}\in\mathbb{R}^{K_{1}}$ are discrete regressors and $X_{(2),t}\in\mathbb{R}^{K_{2}}$ are continuous regressors. We can partition the corresponding $W=X_{1}-X_{0}=(D',Z')'$ with $D=X_{(1),1}-X_{(1),0}$ and $Z=X_{(2),1}-X_{(2),0}$. Let $\mathcal{D}$ denote the support of $D$. We show that identification still holds under a conditional version of sign saturation.

thmLet Assumption (ref) hold. Suppose that there exist linearly independent $d_{1},...,d_{K_{1}}\in\mathcal{D}$ with such that for $j\in\{1,...,K_{1}\}$,\\ (1) both $P(E(Y_{1}-Y_{0}\mid X,D=d_{j})>0)$ and $P(E(Y_{1}-Y_{0}\mid X,D=d_{j})<0)$ are strictly positive\\ (2) the support of $Z$ conditional on $D=d_{j}$ is convex and has non-empty interior. Then $\beta$ is identified up to scale.

We note that $|\mathcal{D}|$ can be much larger than $K_{1}$. For example, if $X_{(1),t}$ represents $K_{1}$ binary variables, then the support of $X_{(1),t}$ is $\{0,1\}^{K_{1}}$ and $\mathcal{D}=\{-1,0,1\}^{K_{1}}$, which means $|\mathcal{D}|=3^{K_{1}}$. Theorem (ref) is convenient in that it does not require the sign saturation to hold conditional on $D=d$ for every $d\in\mathcal{D}$. It suffices to require this for $K_{1}$ linearly independent elements of $\mathcal{D}$. If $K_{1}=1$ ($X_{(1),t}$ is a time dummy), then this requirement reduces to Assumption (ref). Theorem (ref) also allows for more complicated cases as demonstrated in the following example.

example[Time dummy and categorical variables] Suppose that $R_{t}$ is a categorical variable taking values in $\{1,...,K_{1}\}$. Consider $X_{(1),t}=(\mathbf{1}\{t=1\},\mathbf{1}\{R_{t}=1\},...,\mathbf{1}\{R_{t}=K_{1}-1\})'\in\mathbb{R}^{K_{1}}$, i.e., a time dummy and $K_{1}-1$ dummy variables for $R_{t}$. Then we can write the support of $X_{(1),t}\in\mathbb{R}^{K_{1}}$ as $\{\mathbf{1}\{t=1\}\}\times\{e_{1},...,e_{K_{1}-1},\boldsymbol{0}_{K_{1}-1}\}$, where $e_{j}$ is the $j$-th column of the $(K_{1}-1)\times(K_{1}-1)$ identity matrix and $\boldsymbol{0}_{K_{1}-1}=(0,...,0)'\in\mathbb{R}^{K_{1}-1}$. Let us assume that the support of $(R_{0},R_{1})'$ contains $(K_{1},j)$ for any $j\in\{1,...,K_{1}\}$. Then $\mathcal{D}$ contains $\{1\}\times\{e_{1},...,e_{K_{1}-1},\boldsymbol{0}_{K_{1}-1}\}$. In other words, $\mathcal{D}$ contains all the columns of the following matrix: \[ \begin{pmatrix}1 & 1 & 1 & \cdots & 1\\ 0 & 1 & 0 & \cdots & 0\\ 0 & 0 & 1 & \cdots & 0\\ \vdots & \vdots & \vdots & \ddots & \vdots\\ 0 & 0 & 0 & \cdots & 1 \end{pmatrix}\in\mathbb{R}^{K_{1}\times K_{1}}. \] We can easily see that the determinant of this upper triangular matrix is equal to one. Therefore, the above matrix has full rank, which means that $\mathcal{D}$ contains $K_{1}$ linearly independent elements. Then $\beta$ is identified up to scale if for every $j\in\{1,...,K_{1}\}$, $P(E(Y_{1}-Y_{0}\mid Z,R_{0}=K_{1},R_{1}=j)>0)$ and $P(E(Y_{1}-Y_{0}\mid Z,R_{0}=K_{1},R_{1}=j)<0)$ are strictly positive and the support of $Z$ conditional on $(R_{0},R_{1})=(K_{1},j)$ is convex and has non-empty interior.\qed

Is the sign saturation condition is also necessary?

It turns out that in the many situations, sign saturation characterizes identifiability. To illustrate this point, we consider the setting studied by chamberlain2010binary and assume that $u_{t}$ is independent of $(X,\alpha)$ and is from a known distribution.\footnote{To establish necessity, it is enough to consider the case in which $u_{t}$ is independent of $(X,\alpha)$. This is similar to showing minimax lower bounds. If identification cannot be guaranteed even with the additional assumption of independence between $u$ and $(X,\alpha)$, then it is definitely not guaranteed without this assumption.} We show that without sign saturation, identification fails at every point unless the distribution of $u_{t}$ is from a special class.

Suppose that in time period $t\in\{0,1\}$, we observe $(Y_{t},X_{t})$, where $X_{t}=(X_{t,1}',X_{t,2})'$ with $X_{t,1}\in\mathbb{R}^{K-1}$ and $X_{t,2}=\mathbf{1}\{t=1\}$. Suppose that $u_{t}$ is independent of $(X,\alpha)$ and has c.d.f $F(\cdot)$. Then from the data, the distribution of $Y$ given $(X,\alpha)$ is determined by the vector

align[align omitted — 403 chars of source]

where $\beta=(\beta_{1}',\beta_{2})$ is partitioned as $\beta_{1}\in\mathbb{R}^{K-1}$ and $\beta_{2}\in\mathbb{R}$, and $W:=X_{1}-X_{0}=(Z',1)'$ with $Z=X_{1,1}-X_{0,1}$. As pointed out in chamberlain2010binary, here we should aim for identification of $\beta$, not just up to scaling, because “our scale normalization is built in to the given specification for the $u_{t}$ distribution”. Let $\mathcal{Z}$ be the support of $Z$, i.e., again the smallest closed set with the full probability mass. Since $W=(Z',1)'$, the support of $W$ is $\mathcal{W}=\mathcal{Z}\times\{1\}$.

The special class of functions is best described in terms of a transformation of $F(\cdot)$. Define the function $G(\cdot)=\ln\frac{F(\cdot)}{1-F(\cdot)}$. Let $\dot{G}(\cdot)$ denote the derivative of $G(\cdot)$. If $F(\cdot)$ is the logistic function, then $G(\cdot)$ is an affine function and $\dot{G}(\cdot)$ is a constant function. It turns out that if we are willing to rule out cases with $\dot{G}(\cdot)$ being a periodic function,\footnote{A function $h(\cdot)$ on $\mathbb{R}$ is a periodic function if there exists $c\neq0$ such that $h(c+t)=h(t)$ for any $t\in\mathbb{R}$. In this case, $c$ is said to be a period of $h(\cdot)$. Notice that we assume that a period has to be non-zero; otherwise, every function would be a periodic function with a period being zero.} then sign saturation is a necessary condition for identification. This special class includes the logistic function, whose $\dot{G}(\cdot)$ is a constant function and is clearly periodic. (Any real number is a period of a constant function.)

We now introduce the notations for describing identification. Let $\Pi$ denote the set of probability measures on $\mathbb{R}$. The distribution of $Y\mid X=x$ is determined by $\int L(x;\beta,\alpha)d\pi_{x}(\alpha)$ for some $\pi_{x}\in\Pi$. Here, $\pi_{x}$ denotes the distribution $\alpha\mid X=x$. We now recall the notation of identification failure discussed in chamberlain2010binary.

defn[Identification failure at $\beta$] We say that the identification fails at $\beta$ if there exists $b\neq\beta$ such that for any $x$ in the support, there exist $\pi_{x},\tilde{\pi}_{x}\in\Pi$ depending on $x$ such that $\int L(x;\beta,\alpha)d\pi_{x}(\alpha)=\int L(x;b,\alpha)d\tilde{\pi}_{x}(\alpha)$.

This definition says that there exist two “mixing distributions” $\pi_{x}$ and $\tilde{\pi}_{x}$ in $\mathcal{G}$ such that the two mixtures (representing the distribution $Y\mid X=x$) $\int L(x;\beta,\alpha)d\pi_{x}(\alpha)$ and $\int L(x;b,\alpha)d\tilde{\pi}_{x}(\alpha)$ are identical. Following chamberlain2010binary, we can state identification in terms of convex hulls of $L(x;\beta,\alpha)$. The identification fails at $\beta$ if there exists $b\neq\beta$ such that ${\rm conv}\{L(x;\beta,\alpha):\alpha\in\mathbb{R}\}\bigcap{\rm conv}\{L(x;b,\alpha):\alpha\in\mathbb{R}\}\neq\emptyset$ for any $x$, where ${\rm conv}$ denotes the convex hull.

Notice that $X_{0,1}'\beta_{1}$ can be absorbed by $\alpha$ since the distribution of $\alpha$ is allowed to have arbitrary dependence on $x$. Without loss of generality, we view the entire $X_{0,1}'\beta_{1}+\alpha$ as the fixed effects to simplify notations. Then we can restate Definition (ref) as follows.

defnWe say that the identification fails at $\beta$ if there exists $b\neq\beta$ such that $\mathcal{A}(w'\beta)\bigcap\mathcal{A}(w'b)\neq\emptyset$ for any $w\in\mathcal{W}$, where for any $t\in\mathbb{R}$, $\mathcal{A}(t)={\rm conv}\{p(t,\alpha):\alpha\in\mathbb{R}\}$ and \[ p(t,\alpha):=\begin{pmatrix}F(\alpha)\\ F(t+\alpha)\\ F(\alpha)\cdot F(t+\alpha) \end{pmatrix}. \]

Perhaps the simplest way to see the equivalence between the two definitions is to notice that ${\rm conv}\{L(x;\beta,\alpha):\alpha\in\mathbb{R}\}=\mathcal{A}(w'\beta)$, where $w=x_{1}-x_{0}$. We now define the parameter space in which the sign saturation fails: $\mathcal{B}_{+}=\{\beta\in\mathbb{R}^{K}:\ w'\beta>0\ \forall w\in\mathcal{W}\}$. The main result for the necessity is the following.

thmSuppose that $\mathcal{Z}$ is bounded. Suppose that $F(\cdot)$ is continuously differentiable with support on $\mathbb{R}$. If $\dot{G}(\cdot)$ is not a periodic function, then the identification fails at every point in $\mathcal{B}_{+}$.

We compare this with Theorem 1 of chamberlain2010binary, which states that if $F(\cdot)$ is outside a special class, identification fails in an open neighborhood. Theorem (ref) complements this result by showing that if $F(\cdot)$ is outside a special class, identification fails at every point, not just in an open neighborhood. This is also why the proof of Theorem (ref) follows a very different strategy from the proof of Theorem 1 of chamberlain2010binary; the latter essentially only needs to find one point at which the identification fails whereas we need to show that identification fails universally in $\mathcal{B}_{+}$.

The special class in chamberlain2010binary only includes logistic functions, but here our special class is larger and includes all functions with periodic $\dot{G}(\cdot)$. This enlargement of the special class cannot be avoided because we show that there are instances of non-logistic $F(\cdot)$ for which identification holds, e.g., $F(t)=[1+\exp(-2t-\sin(t))]^{-1}$. We can push our arguments further and show that the identification under periodic $\dot{G}(\cdot)$ is “robust” and is based on a set with strictly positive probability mass.

thmSuppose that $\dot{G}(\cdot)$ is a continuous periodic function that is non-constant. Let $\beta=(\beta_{1}',\beta_{2})'\in\mathcal{B}_{+}$. Then\\ (1) there exists a minimal positive period $\eta>0$ for $\dot{G}(\cdot)$, i.e., the set $\{a>0:\ \dot{G}(a+t)=\dot{G}(t)\ \forall t\in\mathbb{R}\}$ has a smallest element. \\ (2) If $(z'\beta_{1}+\beta_{2})/\eta$ is an integer for some $z$ in the interior of $\mathcal{Z}$, then $\beta$ is identified; moreover, for any $b\in\mathcal{B}_{+}$ with $b\neq\beta$, $P\left(\mathcal{A}(W'\beta)\bigcap\mathcal{A}(W'b)=\emptyset\right)>0$, where $W=(Z',1)'$. \\ (3) If $\mathcal{Z}$ is compact and has non-empty interior, then there exists an open ball $D\subset\mathcal{B}_{+}$ such that every point in $D$ is identified.

By Theorem (ref), as long as $\dot{G}(\cdot)$ is periodic and $z'\beta_{1}+\beta_{2}$ is in $\eta\cdot\mathbb{N}$ for some interior point $z$ ($\mathbb{N}$ denotes the set of positive integers), we still have identification. This complements Theorem 1 of chamberlain2010binary in an interesting way. For non-logistic $F(\cdot)$, it is true that the identification fails at every point in an open set. On the other hand, we show that the identification also holds at every point in another open set for periodic $\dot{G}(\cdot)$. It is worth noting that in Theorem (ref), the identification of $\beta$ is robust in the sense that it is not based on a small number of “unimportant points” in $\mathcal{W}$; the points in $\mathcal{W}$ that allow us to identify $\beta$ have strictly positive probability mass. Therefore, the truly hopeless cases for identification are those with non-periodic $\dot{G}(\cdot)$. Finally, for the case of compact and convex $\mathcal{Z}$ with non-empty interior, we summarize our results and the existing literature in Table (ref).

table[table omitted — 1,450 chars of source]

Exploiting the periodicity of $\dot{G}(\cdot)$ can be seen as a generalization of the conditional maximum likelihood estimator for logistic distributions. To see this, we notice that \[ \frac{P(Y_{1}=1,Y_{0}=0\mid X,\alpha)}{P(Y_{1}=0,Y_{1}=1\mid X,\alpha)}=\exp\left[G(W'\beta+\alpha)-G(\alpha)\right]. \] Under logistic errors, $\dot{G}(\cdot)$ is a constant and thus $G(W'\beta+\alpha)-G(\alpha)$ does not depend on $\alpha$ for any $W$. If $\dot{G}(\cdot)$ is a periodic function with the smallest positive period $\eta>0$, then $G(W'\beta+\alpha)-G(\alpha)$ also does not depend on $\alpha$ whenever $W'\beta/\eta$ is an integer. Since $W=(Z',1)'$, we only need to have an interior point $z$ such that $(z',1)'\beta/\eta$ is an integer.

remAs we have seen, the sign saturation condition guarantees the identification of $\beta$ up to scale. Does it also guarantee the parametric rate for estimation? The answer is no because Theorem 2 of chamberlain2010binary shows that the information bound is always zero outside the logistic case, regardless of the identification status. This highlights a key difference between identification and the parametric rate in estimation. The question of identification depends on the sign saturation condition, whereas the root-$n$ estimability depends on the semiparametric efficiency calculation or the functional projection argument developed by bonhomme2012functional. \begin{comment} The above analysis shows that distributions in the special class of periodic $\dot{G}(\cdot)$ can have sufficient identification power without the sign saturation condition. One natural question is whether this special class of periodic $\dot{G}(\cdot)$ also has special properties for estimation. One thing we can say is that the parametric rate is still not possible outside the logistic case. This is because Theorem 2 of chamberlain2010binary still applies. Of course, periodicity of $\dot{G}(\cdot)$ might provide other benefits in terms of estimation even though the root-$n$ rate is impossible. We will leave the exploration of this direction to future research. \end{comment}

Implications on marginal effects

The identification of $\beta$ is directly related to the identification of marginal effects in panel models with binary responses. The bounds on the marginal effects often assume point identification of $\beta$, see e.g., \citet*{DaveziesID2021}, liu2105.12891 and Theorem 6 of \citet*{chernozhukov2013average}.

Here, we explore an important link between $\beta$ and the treatment effects through the sign of components of $\beta$. Let us consider the setting of Section (ref): (1) the errors $u_{t}$'s are i.i.d across $t$ with distribution $F(\cdot)$ and are independent of $(X,\alpha)$ and (2) $X_{2,t}$ is binary with $X_{2,t}=\mathbf{1}\{t=1\}$. We can interpret $X_{2,t}$ as a binary treatment; in period $t=0$, no one is treated and in period $t=1$, everyone is treated. For $\beta=(\beta_{1}',\beta_{2})'$, the sign of $\beta_{2}$ corresponds to the sign of common measures of treatment effects. For example, the average partial effect effect at $X_{1,t}=z$ is \[ \Delta_{APE}(z)=P(z'\beta_{1}+\beta_{2}+\alpha\geq u_{t})-P(z'\beta_{1}+\alpha\geq u_{t}), \] see e.g., chernozhukov2013average.

Although the magnitude of average partial effect is often not point identified (see e.g., honore2006bounds and chernozhukov2013average), there is hope that the sign of the effect is point identified, which means that the identified set contains only positive numbers, only negative numbers or only zero. Under Assumption (ref), the support of $u_{t}$ is $\mathbb{R}$, which means that ${\rm sgn}(\Delta_{APE}(z))={\rm sgn}(\beta_{2})$ for any $z$. Hence, identifying ${\rm sgn}(\Delta_{APE}(z))$ is equivalent to identifying ${\rm sgn}(\beta_{2})$. In the literature, there are also terms such as average treatment effects or ceteris paribus effects, e.g., hoderlein2012nonparametric and chernozhukov2015nonparametric. Under the assumptions of a linear index, the sign of these other measures of treatment effects is also the sign of $\beta_{2}$.

remThere are alternative ways of defining treatment effects but due to the single-index nature of the model, the sign of other notions of treatment effects is still related to ${\rm sgn}(\beta_{2})$. For example, consider the individual effect $\mathbf{1}\{X_{1,t}\beta_{1}+\beta_{2}+\alpha\geq u_{t}\}-\mathbf{1}\{X_{1,t}\beta_{1}+\alpha\geq u_{t}\}$. Clearly, this effect is negative with zero probability if and only if $\beta_{2}>0$.

We now show that the sign saturation condition is necessary for this purpose. To formally state the result, we rephrase Definition (ref).

defnWe say that $\beta$ and $b$ are observationally equivalent if for any $x$ in the support, there exist $\pi_{x},\tilde{\pi}_{x}\in\Pi$ depending on $x$ such that $\int L(x;\beta,\alpha)d\pi_{x}(\alpha)=\int L(x;b,\alpha)d\tilde{\pi}_{x}(\alpha)$, where $L$ is defined in ((ref)).

Recall that $\mathcal{B}_{+}=\{\beta\in\mathbb{R}^{K}:\ w'\beta>0\ \forall w\in\mathcal{W}\}$ is the set of parameter values that do not satisfy the sign saturation condition ($E(Y_{1}-Y_{0}\mid X)$ always positive). Consider two subsets $\mathcal{B}_{+,0}=\{\beta=(\beta_{1}',\beta_{2})'\in\mathcal{B}_{+}:\ \beta_{2}=0\}$ and $\mathcal{B}_{+,+}=\{\beta=(\beta_{1}',\beta_{2})'\in\mathcal{B}_{+}:\ \beta_{2}>0\}$, which denote the set of parameter values that imply zero effects and positive effects of $X_{2,t}$, respectively. The following result states the identification failure of ${\rm sgn}(\beta_{2})$.

thmSuppose that $\mathcal{Z}$ is bounded. Suppose that $F(\cdot)$ is continuously differentiable such that $\dot{G}(\cdot)$ is not a periodic function. Then every point in $\mathcal{B}_{+,0}$ is observationally equivalent to a point in $\mathcal{B}_{+,+}$.

By Theorem (ref), the sign saturation condition guarantees point identification of $\beta$ up to scale and thus point identification of ${\rm sgn}(\beta_{2})$. By Theorem (ref), when the sign saturation condition fails, we cannot distinguish between $\Delta_{APE}(z)=0$ and $\Delta_{APE}(z)>0$ unless $\dot{G}(\cdot)$ is periodic. Therefore, unless $\dot{G}(\cdot)$ is periodic, the sign saturation condition is sufficient and necessary to guarantee point identification of the sign of the marginal effects of $X_{2,t}$.

commentBy a continuity argument, one can also show that . For any $\varepsilon>0$, we define $\mathcal{B}_{+}^{(\varepsilon)}:=\{\beta\in\mathbb{R}^{K}:\ w'\beta>\varepsilon\ \forall w\in\mathcal{W}\}$. Clearly, $\mathcal{B}_{+}=\mathcal{B}_{+}^{(0)}$. For any $\varepsilon>0$, $\mathcal{B}_{+,+,\varepsilon}=\{\beta=(\beta_{1}',\beta_{2})'\in\mathcal{B}_{+}:\ 0<\beta_{2}<\varepsilon\}$ and $\mathcal{B}_{+,+,-\varepsilon}=\{\beta=(\beta_{1}',\beta_{2})'\in\mathcal{B}_{+}:\ -\varepsilon<\beta_{2}<0\}$. We show that there exists $\varepsilon>0$ such that every point in $\mathcal{B}_{+,+,-\varepsilon}$ is observationally equivalent to a point in $\mathcal{B}_{+,+}$. \begin{proof} We now show that there exists $\delta>0$ such that for any $z\in\mathcal{Z}$, $\mathcal{A}(w'\beta)\bigcap\mathcal{A}(w'b)\neq\emptyset$, where $w=(z',1)'$. We proceed by contradiction. Suppose that for any $\delta>0$, there exists $z_{\delta}\in\mathcal{Z}$ such that $\mathcal{A}(w_{\delta}'b)\bigcap\mathcal{A}(w_{\delta}'\beta)=\emptyset$, where $w_{\delta}=(z_{\delta}',1)'$. By Lemma (ref), there exists $v_{\delta}\neq0$ such that $\sup_{\alpha\in\mathbb{R}}v_{\delta}'p(w_{\delta}'b,\alpha)\leq\inf_{\alpha\in\mathbb{R}}v_{\delta}'p(w_{\delta}'\beta,\alpha)$. By Lemma (ref) and $w_{\delta}'b-w_{\delta}'\beta=\delta>0$, we have that $\sup_{\alpha\in\mathbb{R}}v_{\delta}'p(w_{\delta}'b,\alpha)\geq\inf_{\alpha\in\mathbb{R}}v_{\delta}'p(w_{\delta}'\beta,\alpha)$. Hence, \[ \sup_{\alpha\in\mathbb{R}}v_{\delta}'p(w_{\delta}'\beta+\delta,\alpha)=\sup_{\alpha\in\mathbb{R}}v_{\delta}'p(w_{\delta}'b,\alpha)=\inf_{\alpha\in\mathbb{R}}v_{\delta}'p(w_{\delta}'\beta,\alpha). \] By Lemma (ref) (with $t=w_{\delta}'\beta$), \[ \inf_{\alpha\in\mathbb{R}}[G(w_{\delta}'\beta+\delta+\alpha)-G(\alpha)]>\sup_{\xi\in\mathbb{R}}[G(w_{\delta}'\beta+\xi)-G(\xi)]. \] Notice that $w_{\delta}'\beta\in T$. Therefore, we have shown that for any $\delta>0$, there exists $t_{\delta}\in T$ such that \[ \inf_{\alpha\in\mathbb{R}}[G(t_{\delta}+\delta+\alpha)-G(\alpha)]>\sup_{\xi\in\mathbb{R}}[G(t_{\delta}+\xi)-G(\xi)]. \] By Lemma (ref), $\dot{G}$ is a periodic function. However, this contradicts the assumption that $\dot{G}$ is not a periodic function. \end{proof} Define $\mathcal{B}_{+,-}=\{\beta=(\beta_{1}',\beta_{2})'\in\mathcal{B}_{+}:\ \beta_{2}<0\}$. We fix an arbitrary $\beta_{*}=(\beta_{*,1}',\beta_{,*2})\in\mathcal{B}_{+,-}$. Clearly, $(\beta_{*,1}',0)'\in\mathcal{B}_{+,0}$ since $\beta_{*,2}<0$. In the proof of Theorem (ref), we proved that $(\beta_{*,1}',0)'\in\mathcal{B}_{+,0}$ is observationally equivalent to $(\beta_{*,1}',\delta)'$ for some $\delta_{*}>0$. Notice that we can choose $\delta_{*}$ to depend only on $\beta_{1,*}$. We now consider $\mathcal{B}_{+,+,-\varepsilon}$ with $\varepsilon=\delta_{*}/2$. We also know that $z'\beta_{*,1}>-\beta_{*,2}$ for any $z$. We fix an arbitrary $\beta=(\beta_{1}',\beta_{2})'\in\mathcal{B}_{+,+,-\varepsilon}$. We define Now we consider any $\beta=(\beta_{1}',\beta_{2})\in\mathcal{B}_{+}$ with $\beta_{2}<0$. Now we pick any $\beta\in\mathcal{B}_{+,0}$. Then $b=(\beta_{1}',\beta_{2}+\delta)'=(\beta_{1}',\delta)\in\mathcal{B}_{+,+}$. The desired result follows. Define $A=\{z'\beta_{1}:z\in\mathcal{Z}\}$. Then $T=A+\beta_{2}$ is contained by $\overline{T}=[]$.

Assessing the sign saturation condition

We have seen that the sign saturation condition is sufficient and necessary for the identification. The conditional mean function $\phi(X)=E(Y_{1}-Y_{0}\mid X)$ is identified. Therefore, ideally the sign saturation condition can be checked in data. Here, we provide a simple check that does not involve nonparametric estimation of $\phi$.

By Lemma (ref), $P(\phi(X)\leq0)=1$ if and only if $P(W'\beta\leq0)=1$. Hence, we define $\rho(q)=E\mathbf{1}\{W'q\geq0\}(Y_{1}-Y_{0})$ for $q\in\mathbb{R}^{K}$. Here, we notice that we can replace $\mathbb{R}^{K}$ with $[-1,1]^{K}$ or any set that contains an open neighborhood of zero. It turns out that the sign saturation condition can be written in terms of $\rho(\cdot)$ once we observe \[ \rho(\beta)=E\left(\mathbf{1}\{\phi(X)\geq0\}\cdot\phi(X)\right)\quad{\rm and}\quad\rho(-\beta)=E\left(\mathbf{1}\{\phi(X)\leq0\}\cdot\phi(X)\right). \]

We give the formal statement below.

lemLet Assumption (ref) hold. Then \[ \sup_{q\in\mathbb{R}^{K}}\rho(q)=E\max\left\{ \phi(X),0\right\} \ {\rm and}\ \inf_{q\in\mathbb{R}^{K}}\rho(q)=E\min\left\{ \phi(X),0\right\} . \]

Define $\tau_{*}=\min\{\tau_{1},-\tau_{2}\}$ with $\tau_{1}=\sup_{q\in\mathbb{R}^{K}}\rho(q)$ and $\tau_{2}=\inf_{q\in\mathbb{R}^{K}}\rho(q)$. By Lemma (ref), measuring $\tau_{*}$ can give us some indication of whether (or how “well”) the sign saturation condition holds. In particular, the sign saturation fails if and only if $\tau_{*}=0$. From the data, we can construct a one-sided confidence interval for $\tau_{*}$. The natural estimate for $\tau_{*}$ is $\hat{\tau}_{*}=\min\{\hat{\tau}_{1},-\hat{\tau}_{2}\}$, where $\hat{\tau}_{1}=\sup_{q\in\mathbb{R}^{K}}\hat{\rho}_{n}(q)$, $\hat{\tau}_{2}=\inf_{q\in\mathbb{R}^{K}}\hat{\rho}_{n}(q)$ and \[ \hat{\rho}_{n}(q)=n^{-1}\sum_{i=1}^{n}\mathbf{1}\{W_{i}'q\geq0\}(Y_{i,1}-Y_{i,0}). \]

In the proof of Theorem (ref) below, we show that \[ \sqrt{n}\left(\hat{\tau}_{*}-\tau_{*}\right)\leq\left(\sup_{q\in\mathbb{R}^{K}}S_{n}(q)\right)\cdot\mathbf{1}\{\tau_{1}\leq-\tau_{2}\}+\left(\sup_{q\in\mathbb{R}^{K}}(-S_{n}(q))\right)\cdot\mathbf{1}\{\tau_{1}>-\tau_{2}\}, \] where $S_{n}(q)=\sqrt{n}(\hat{\rho}_{n}(q)-\rho(q))$. By the standard arguments of empirical processes, $S_{n}$ converges to a mean-zero Gaussian process, which means that $\sup_{q\in\mathbb{R}^{K}}S_{n}(q)$ and $\sup_{q\in\mathbb{R}^{K}}(-S_{n}(q))$ have the same distribution. Thus, a simple $(1-\alpha)$-confidence interval for $\tau_{*}$ can be obtained once we approximate the $(1-\alpha)$ quantile of $\sup_{q\in\mathbb{R}^{K}}S_{n}(q)$. This motivates the following bootstrapping algorithm.

lyxalgorithmFor a $(1-\alpha)$-confidence interval for $\tau_{*}$: \begin{enumerate} • Collect data $\{(W_{i},Y_{i,1},Y_{i,0})\}_{i=1}^{n}$. • Compute the estimate $\hat{\tau}_{*}=\min\{\hat{\tau}_{1},-\hat{\tau}_{2}\}$. • Draw a random sample $\{(\tilde{W}_{i},\tilde{Y}_{i,1},\tilde{Y}_{i,0})\}_{i=1}^{n}$ with replacement from the data and compute $\sup_{q\in\mathbb{R}^{K}}\tilde{S}_{n}(q)$, where $\tilde{S}_{n}(q)=\sqrt{n}\left(n^{-1}\sum_{i=1}^{n}\mathbf{1}\{\tilde{W}_{i}'q>0\}(\tilde{Y}_{i,1}-\tilde{Y}_{i,0})-\hat{\rho}_{n}(q)\right)$. • Repeat the previous step many times and compute $c_{1-\alpha}$, the $(1-\alpha)$ quantile of $\sup_{q\in\mathbb{R}^{K}}\tilde{S}_{n}(q)$.\footnote{In practice, optimization over $q\in\mathbb{R}^{K}$ can be done by a grid search over $N$ points on the unit sphere $\{v\in\mathbb{R}^{K}:\ \|v\|_{2}=1\}$. The computation seems reasonably fast for a normal sample size. For example, in the setting of Section (ref) with $\beta_{2}=1$, $n=5000$ and grid size $N=2000$, computing $\sup_{q}\tilde{S}_{n}(q)$ with 500 bootstrap samples take less than 17 seconds in Matlab on a 2022 MacBook Pro laptop. Further speedups are possible depending on $n$, $N$ and available memory.} • The $(1-\alpha)$-confidence interval for $\tau_{*}$ is $\left[\max(\hat{\tau}_{*}-c_{1-\alpha}n^{-1/2},0),\ 1\right]$. \end{enumerate}

Notice that computing $\hat{\tau}_{1}$ is equivalent to computing the maximum score estimator\footnote{Notice that maximizing $\hat{\rho}_{n}(q)$ is equivalent to maximizing $n^{-1}\sum_{i=1}^{n}{\rm sgn}(W_{i}'q)(Y_{i,1}-Y_{i,0})$ because ${\rm sgn}(W_{i}'q)=-1+2\cdot\mathbf{1}\{W_{i}'q\geq0\}$.} and all the existing computational algorithms and software for the maximum score estimator can be used. Finding $\hat{\tau}_{2}$ also reduces to computing the maximum score estimator once we swap $Y_{1}$ and $Y_{0}$. The above algorithm can be justified by the following result.

thmLet Assumption (ref) hold. Assume that $P(Y_{1}=Y_{0})<1$. Then \[ \limsup_{n\rightarrow\infty}P\left(\sqrt{n}(\hat{\tau}_{*}-\tau_{*})>c_{1-\alpha}\right)\leq\alpha. \]

Note that the cube-root asymptotics typically associated with the maximum score estimator (see e.g., kim1990cube and seo2018local) does not arise here. The reason is that $\hat{\tau}_{1}$ is the maximum of $\hat{\rho}_{n}(\cdot)$, rather than the argmax. The cube-root asymptotics of the maximum score estimator is due to the non-standard rate of some terms in the expansion for analyzing the argmax. Fortunately, we do not have to deal with such terms for our purpose.

It is worth pointing out that Algorithm (ref) is asymptotically exact in the sense that there exists a data-generating process under which the asymptotic coverage probability is exactly $\alpha$. For example, if $\rho(q)=0$ for any $q$, then $\sqrt{n}(\hat{\tau}_{*}-\tau_{*})=\sup_{q\in\mathbb{R}^{K}}S_{n}(q)$ and $P\left(\sqrt{n}(\hat{\tau}_{*}-\tau_{*})>c_{1-\alpha}\right)\rightarrow\alpha$.

commentIf the support of $Z_{i}$ (with $W_{i}=(1,Z_{i}')'$) is not convex but contains a convex open neighborhood of zero, we can restrict $Z_{i}$ to this neighborhood.
commentWe now revisit the empirical example in Section (ref). Since we have seen strong evidence against the logistic assumption, we would like to check whether $\beta$ can be identified without imposing parametric restrictions on the error terms. To accommodate multiple time periods, we modify the test statistic by using \[ \hat{\rho}_{n}(q)=n^{-1}\sum_{i=1}^{n}\sum_{s>t}\mathbf{1}\{(X_{i,s}-X_{i,t})'q\geq0\}(Y_{i,s}-Y_{i,t}). \] The bootstrapped version is adjusted accordingly. We now report the p-value based on 2000 bootstrap samples. \begin{center} \begin{tabular}{cc} Null hypotheses & p-value\tabularnewline \hline $H_{0}:\sup_{q\in\mathbb{R}^{K}}\rho(q)\leq0$ & 0.000\tabularnewline $H_{0}:\inf_{q\in\mathbb{R}^{K}}\rho(q)\geq0$ & 0.027\tabularnewline \hline \hline & \tabularnewline \end{tabular} \end{center} We see that at least at 5% significance level, we reject the null hypothesis of
comment\begin{rem} An alternative test is to construct a one-sided confidence interval for $\tau_{1}$ and then for $\tau_{2}$. However, doing inference simultaneously on both $\tau_{1}$ and $\tau_{2}$ is not straight-forward and might invoke Bonferroni-type methods that sacrifice power. Instead, Algorithm (ref) circumvents this problem and gives an asymptotically exact test. \end{rem} testing $E(Y_{1}-Y_{0}\mid X)\leq0$ almost surely can be done by testing $\sup_{q\in\mathbb{R}^{K}}\rho(q)\leq0$. This test amounts to constructing a one-sided confidence interval for $\sup_{q\in\mathbb{R}^{K}}\rho(q)$. A similar strategy can be done for testing $E(Y_{1}-Y_{0}\mid X)\geq0$ almost surely. Since the sign saturation condition is split into two hypotheses here, one could consider a Bonferroni-type correction. However, Bonferroni-type procedures can have low power. We now construct one statistic that directly tests the sign saturation condition. By Lemma (ref), $P\left(E(Y_{1}-Y_{0}\mid X)>0\right)=0$, i.e., $P\left(E(Y_{1}-Y_{0}\mid X)\leq0\right)=1$, is equivalent to $\tau_{1}\leq0$, where $\tau_{1}=\sup_{q\in\mathbb{R}^{K}}\rho(q)$. Similarly, $P\left(E(Y_{1}-Y_{0}\mid X)<0\right)=0$, i.e., $P\left(E(Y_{1}-Y_{0}\mid X)\geq0\right)=1$, is equivalent to $\tau_{2}\geq0$, where $\tau_{2}=\inf_{q\in\mathbb{R}^{K}}\rho(q)$. Therefore, the testing problem of \[ H_{0}:\ P\left(E(Y_{1}-Y_{0}\mid X)>0\right)=0\ \text{or}\ P\left(E(Y_{1}-Y_{0}\mid X)<0\right)=0 \] versus \[ H_{1}:\ P\left(E(Y_{1}-Y_{0}\mid X)>0\right),P\left(E(Y_{1}-Y_{0}\mid X)<0\right)>0 \] is equivalent to the testing problem of \begin{equation} H_{0}:\ \min\{\tau_{1},-\tau_{2}\}\leq0\quad{\rm versus}\quad H_{1}:\ \min\{\tau_{1},-\tau_{2}\}>0. \end{equation} To construct a test statistic, we consider the natural estimate $\min\{\hat{\tau}_{1},-\hat{\tau}_{2}\}$, where $\hat{\tau}_{1}=\sup_{q\in\mathbb{R}^{K}}\hat{\rho}(q)$, $\hat{\tau}_{2}=\inf_{q\in\mathbb{R}^{K}}\hat{\rho}(q)$ and \[ \hat{\rho}_{n}(q)=n^{-1}\sum_{i=1}^{n}\mathbf{1}\{W_{i}'q\geq0\}(Y_{i,1}-Y_{i,0}). \] The next question is how to construct the critical value. It turns out that a one-sided confidence interval for $\min\{\tau_{1},-\tau_{2}\}$ is $[\min\{\hat{\tau}_{1},-\hat{\tau}_{2}\}-c_{1-\alpha}^{*},1]$, where $c_{1-\alpha}^{*}$ satisfies $P(\sup_{q\in\mathbb{R}^{K}}S_{n}(q)>c_{1-\alpha}^{*})=\alpha$ with $S_{n}(q)=\sqrt{n}(\hat{\rho}_{n}(q)-\rho(q))$. Since $S_{n}(\cdot)$ is a simple empirical process, we can use a nonparametric bootstrap to obtain $c_{1-\alpha}^{*}$. We summarize the test below.

Our results characterize what drives the identification and introduce $\tau_{*}$ as a measure for the identification strength. This theoretical insight motivates a preliminary diagnostic step in applied work. Reporting the confidence interval of $\tau_{*}$ in Algorithm (ref) can offer important insight on the identification strength. For example, if this confidence interval contains zero, then one should be concerned about identification failure.

However, testing $\tau_{*}=0$ might not detect all the problematic cases. In practice, the identification might not be clear-cut and additional caution is warranted if we worry about the “weak” identification scenario, which can complicate econometric analysis due to the generic issue of post-selection inference, see e.g., Leeb2005.

Fortunately, since $\tau_{*}$ can be learned from the data, one option is to proceed with the identification-based inference (i.e., cube-root asymptotics) only when the identification is strong enough. For instance, one might adopt a robust default inference method and switch to the cube-root approach only when a test fails to reject $H_{0}:\ \tau_{*}\geq\tau_{0}$, where $\tau_{0}>0$ is a pre-specified fixed threshold. This is uniformly valid asymptotically as it rules out $\tau_{*}$ being “local” to zero.

This strategy is conceptually analogous to a common treatment of weak instruments in the linear instrumental variable (IV) models: low correlation between the instrument and the exogenous variable suggests potential identification failure and one solution is to proceed with the classical asymptotics only when this correlation is large enough, e.g., measured by the first-stage $F$-statistic stock2002testing.

This raises a natural question: what should this default identification-robust method be? In the IV literature, approaches such as the Anderson-Rubin test have been developed. In our model, one could consider inverting the maximum score criterion function, in a manner similar to the Anderson--Rubin approach. While more refined methods may be possible, they would require substantial additional analysis on asymptotic theory, which lies beyond the current focus on characterizing identification and are left for future research.

commentMore broadly, the results here still offer applied researchers a practical diagnostic tool to rule out some (but not all) problematic situations. If the confidence interval for $\tau_{*}$ in Algorithm (ref) contains zero, then we should definitely worry about identification and avoid cube-root asymptotics. includes zero, this would signal insufficient identification strength, suggesting that cube-root asymptotics should be avoided. On the other hand, the results here can still provide applied researchers with helpful diagnostic. If the identification does not seem Here, $\tau_{*}=0$ corresponds to the failure of the sign saturation condition and thus implies lack of point-identification; similarly, zero correlation between the IV and the exogenous variable implies the identification failure in an IV model. To ensure uniform validity of inference, one can choose to apply the classical asymptotic theory only when the correlation between the IV and the exogenous variable is above a fixed positive threshold. Of course, this approach does not provide a concrete answer to how to proceed when the identification is not strong enough. In the IV literature, various identification-robust solutions, such as the Anderson-Rubin test, have been developed. In our model, one could consider inverting the maximum score criterion function, in a manner analogous to the Anderson--Rubin approach. While more refined methods may be possible, they would require substantial technical analysis on inference that deviates from the current focus on characterizing identification---a direction we leave to future research.
commentFor the weak IV problem, one solution is to check whether the IV is strong enough and proceed accordingly, such as Stock and Yogo (2005)??. In both problems, the identification strength can be assessed from the data. In our model, we can learn $\tau_{*}$; in the IV model, we can learn the correlation between the IV and the exogenous variable. We recommend the following approach for inference: start with an identification-robust method for inference, check the identification strength and only switch to the cubic-root asymptotics when we are satisfied with strong identification.
commentADD weak identification discussion. We should set the threshold to be high enough.
commentthis approach does not provide a complete answer. What should we do when the identification is not strong enough?
commentFor example, suppose that we have an identification-robust inference method as a default method. We can set a threshold $\tau_{0}$ and only switch from the default method to a method based on cube-root asymptotics when we fail to reject

Numerical illustration

We have seen that $\tau_{*}$ is a measure of the sign saturation condition and thus serves a gauge for identification. Using Monte Carlo simulations, we now illustrate how $\tau_{*}$ relates to the identification strength and the accuracy of the cube-root asymptotics.

Consider the model in ((ref)) with $X_{t}=(\mathbf{1}\{t=0\},X_{2,t})'$ for $t\in\{0,1\}$, where $X_{2,t}$ is from the uniform distribution on $[-1,1]$, $u_{t}$ is from the standard normal distribution and $\alpha_{t}=(X_{2,0}+X_{2,1})/2$. Here, $X_{2,0}$, $X_{2,1}$, $u_{0}$ and $u_{1}$ are mutually independent. We set $\beta=(1,\beta_{2})'$. Therefore, $W'\beta=1+\beta_{2}(X_{2,1}-X_{2,0})$. Clearly, the support of $W'\beta$ is $[-2|\beta_{2}|+1,2|\beta_{2}|+1]$, which means that the sign saturation holds if and only if $|\beta_{2}|>0.5$. In Figure (ref), we plot $\tau_{*}$ (computed using $\hat{\tau}_{*}$ with $n=10^{8}$) as a function of $\beta_{2}$. It confirms the intuition that $\beta_{2}=0.5$ corresponds to $\tau_{*}=0$, which means that the sign saturation condition fails. Higher value of $\beta_{2}$ corresponds to a higher degree to which the sign saturation condition holds.

figure[figure omitted — 135 chars of source]

When the identification holds (together with other regularity conditions), one can expect the cube-root asymptotics. Given a particular value of $\tau_{*}$, we compare the finite-sample distribution of the maximum score estimator with its asymptotic distribution. Explicitly deriving the cube-root asymptotic distribution requires quantities that are difficult to compute. To circumvent this problem, we treat the distribution of the estimator with a sample size of $n=1,000,000$ as the asymptotic distribution.\footnote{Here, the goal is to assess the importance of identification by examining how well the asymptotic theory applies. In practice, there is another issue of approximating the asymptotic distribution either with sub-sampling or with a modified bootstrap.} The finite-sample distribution is computed using $n=1,000$. To make these two distributions comparable, we compare the maximum score estimator $\hat{\beta}_{2}$ and consider the distribution of $n^{1/3}(\hat{\beta}_{2}-\beta_{2})$, which should converge to a limiting distribution under the cube-root asymptotics. We present the comparison in Figure (ref) for $\beta_{2}\in\{0.6,\ 1\}$ based on 5,000 repetitions. We see that when the identification is stronger ($\beta_{2}=1$) and the finite-sample distribution of the maximum score estimator is better approximated by the asymptotic distribution.

comment\section{Discussions} This paper establishes the sign saturation condition as the key identification condition and provides a simple way of checking this condition in the data. Although the estimation and inference has been studied assuming identification, we are not aware of any previous works that prescribe a formal procedure for checking the identification. In practice, the identification might not be a clear-cut issue and extra caution should be exercised if we worry about the “weak” identification scenario due to the post-selection inference problem by ???. We draw a parallel to the identification problem with weak instrumental variables (IV's). Notice that $\tau_{*}=0$ (defined in Section (ref)) denotes the failure of the sign saturation condition and thus of point-identification. Similarly, zero correlation between the IV and the exogenous variable implies the identification failure in an IV model. For the weak IV problem, one solution is to check whether the IV is strong enough and proceed accordingly, such as Stock and Yogo (2005)??. In both problems, the identification strength can be assessed from the data. In our model, we can learn $\tau_{*}$; in the IV model, we can learn the correlation between the IV and the exogenous variable. We recommend the following approach for inference: start with an identification-robust method for inference, check the identification strength and only switch to the cubic-root asymptotics when we are satisfied with strong identification. We now outline the details of this approach and establish its validity. First, an identification-robust inference approach can be constructed in a way similar to the Anderson-Rubin test in the linear IV model. By pakes2024moment, the identified set for $\beta$ is $\arg\max_{q}\rho(q)$. \begin{thm} Let Assumption (ref) hold. Assume that $P(Y_{1}=Y_{0})<1$. Consider $c_{1-\alpha}$ defined in Algorithm (ref). Define \begin{equation} CS(1-\alpha)=\left\{ b:\ \hat{\rho}(b)\geq\sup_{q}\hat{\rho}(q)-n^{-1/2}c_{1-\alpha}\right\} . \end{equation} Then \[ \limsup_{n\rightarrow\infty}P\left(\beta\in CS(1-\alpha)\right)\geq1-\alpha. \] \end{thm} Therefore, Algorithm (ref) gives us the identification-robust confidence set in ((ref)) as well as a confidence interval for $\tau_{*}$, which is $[\hat{\tau}_{*}-n^{-1/2}c_{1-\alpha},1]$. Suppose that we give a Alternatively, the literature has recommended an identification-robust approach using methods such as inverting the Anderson-Rubin test???. Our recommendation is to start with an identification-robust approach here. If we have further evidence of strong idenfication, we can adopt results in this paper suggest the following steps in practice. We can start with blindly computing the maximum score estimator without worrying about whether the identification fails. Then we check the identification by verifying the sign saturation condition using Algorithm (ref); in essence, the test statistic and the critical values are essentially the objective function values of the maximum score estimator and its bootstrapped version. Once we are confident that the model is identified, the estimate we already computed is econometrically justified and the bootstrap procedure in Algorithm (ref) can also be used for inference via a test inversion. In this paper, we do not consider non-standard asymptotics under which $P(E(Y_{1}-Y_{0}\mid X)>0)$ and $P(E(Y_{1}-Y_{0}\mid X)<0)$ are local to zero in some sense. A similar issue arises in pre-tests for the strength of instruments in instrumental variables regressions. Perhaps a simple idea is to employ confidence intervals. For example, the one-sided confidence interval with nominal coverage $1-\alpha$ for $\min\{\tau_{1},-\tau_{2}\}$ is simply $[\min\{\hat{\tau}_{1},-\hat{\tau}_{2}\}-n^{-1/2}c_{1-\alpha},1]$. Even if we fail to reject $H_{0}$ in ((ref)), we would still like to see a relatively high value of $\min\{\hat{\tau}_{1},-\hat{\tau}_{2}\}-n^{-1/2}c_{1-\alpha}$ to be comfortable with the identification strength. We will leave a more complete solution to this issue as future research.
figure[figure omitted — 262 chars of source]