EconBase
← Back to paper

Robust Tests of Model Incompleteness in the Presence of Nuisance Parameters

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

101,700 characters · 17 sections · 69 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Robust Tests of Model Incompleteness in the Presence of Nuisance Parameters

abstractEconomic models may exhibit incompleteness depending on whether or not they admit certain policy-relevant features such as strategic interaction, self-selection, or state dependence. We develop a novel test of model incompleteness and analyze its asymptotic properties. A key observation is that one can identify the least-favorable parametric model that represents the most challenging scenario for detecting local alternatives without knowledge of the selection mechanism. We build a robust test of incompleteness on a score function constructed from such a model. The proposed procedure remains computationally tractable even with nuisance parameters because it suffices to estimate them only under the null hypothesis of model completeness. We illustrate the test by applying it to a market entry model and a triangular model with a set-valued control function. Keywords: Incomplete models, Discrete choice models, Strategic interaction, Score tests

Introduction

Discrete choice models are used widely. A common empirical strategy is to combine a theory of choice (e.g., utility maximization) that predicts a unique outcome value with distributional assumptions on latent variables mcfadden1981econometric. This approach allows the researcher to derive the conditional distribution of the outcome given covariates and apply likelihood-based inference methods. However, recent economic applications often involve models that permit multiple outcome values, which we call an incomplete prediction. Such an incomplete prediction occurs when the researcher is willing to work only with weak assumptions or has limited knowledge of the data-generating process.

This paper considers a form of incompleteness summarized as follows. An observable discrete outcome $Y\in \mathcal Y$ satisfies

align[align omitted — 64 chars of source]

where $G$ collects all outcome values compatible with the model given observed and unobserved variables $(X,U)$ and a structural parameter $\theta$. This structure arises in a variety of contexts. For example, multiple outcomes are predicted in single-agent discrete choice models when the agent's unobservable choice set is consistent with a wide range of choice set formation processes barseghyan2021heterogeneous. Multiple equilibria may exist in discrete games such as firms' market entry or household labor supply decisions, but one may not know how an equilibrium outcome gets selected BRESNAHANJoE91,CilibertoTamer2009. Recent empirical studies have applied inference methods for incomplete models in different areas; they include English auctions haile2003inference, strategic voting kawai2013inferring, product offerings eizenberg2014upstream,wollmann2018trucks, network formation depaulaEtAl2018,sheng2020, school choice fack2019beyond, and major choice henry2020revealing.

This paper focuses on developing tests to determine if a structural model is incomplete. The completeness of a model is important when considering its policy implications. For example, in a canonical market-entry model, multiple equilibria exist only if firms strategically interact. Testing for strategic interaction effects and inferring their signs can provide valuable information for policymakers depaulatang2012. In a triangular model with a discrete endogenous variable, the control function approach yields an incomplete model, but it remains complete if the assignments are exogenous. Detecting self-selection can aid practitioners in selecting a proper strategy for evaluating treatment effects. Dynamic discrete choice models can make incomplete predictions when they admit state dependence, but they are complete without it. The presence of state dependence can significantly impact program evaluation and counterfactual analyses card2005estimating,handel2013adverse.

In many of these examples, one can state the null hypothesis of model completeness as a restriction on the parameter's subvector. We develop a novel score test for such a restriction. Advantages of this approach are (i) the score statistic only requires estimation of nuisance parameters in the complete restricted model; (ii) one can use package software to estimate the nuisance parameters; and (iii) simulating the score statistic's limiting distribution is straightforward.

Our focus is on models that are complete under the null hypothesis. This structure makes score-based tests appealing. First, one can obtain a unique likelihood function $q_{\theta_0}$ for any null parameter value $\theta_0$, thanks to the model's completeness. For each alternative parameter value $\theta_0+h$, the model implies multiple (typically infinitely-many) likelihood functions. However, the recent work of kz demonstrates that one can identify a "least favorable" density $q_{\theta_0+h}$ that is most difficult to distinguish from $q_{\theta_0}$. We may then consider the family $\{q_{\theta_0+h}\}_{h\in\mathbb R^d}$ as a "least favorable parametric model". Second, we can use the least favorable parametric family to calculate a score function. We then construct a score-based statistic that maximizes a measure of local discrimination. The resulting test is robust to incompleteness because it detects any local deviation from the null hypothesis regardless of how $Y$ is selected from the predicted set.

The score test has another appealing feature wherein one can estimate nuisance parameters within the complete restricted model. Typically, the null hypothesis does not restrict some components of $\theta$. By exploiting the model completeness, we demonstrate that one can construct a restricted maximum likelihood estimator (RMLE) of these nuisance components that is $\sqrt n$-consistent. This estimator is usually simple to calculate using package software. Next, we insert the estimator into the score formula to compute a test statistic. The suggested procedure is computationally tractable since it avoids evaluating the test statistic over a grid of nuisance parameters. Finally, we derive the score-based statistic's limiting distribution and demonstrate that it can be easily simulated. Although there are existing methods to test restrictions on subvectors of $\theta$, they can be computationally expensive since they are not designed to utilize the model completeness under the null hypothesis. This paper presents a procedure that makes use of the model structure to simplify its implementation. Through a Monte Carlo experiment, we also show that the proposed test has significantly higher power than a general subvector test that does not utilize the structure.

Relation to the Literature

Our paper belongs to the literature on inference in incomplete models pioneered by Wald1950 in the context of simultaneous equations models and by jovanovich89 in the context of models with multiple equilibria. The seminal work of tamer shows an incomplete model induces multiple distributions and implies partially identifying restrictions on parameters. Developments in the literature galichon2011set,bmm,chesher2017generalized provided tools to systematically derive so-called sharp identifying restrictions, which convert all model information into a set of equality and inequality restrictions on the conditional moments of observable variables. Inference methods based on the sample analogs of moment restrictions are extensively studied canay/shaikh:2017. Our approach builds on recent developments in likelihood-based inference methods for incomplete models Chen2018,kz. In particular, we combine the sharp identifying restrictions with the framework of kz to derive the least favorable parametric model and its score. To our knowledge, this approach is new. One can view our procedure as an analog of deriving a score function from a parametric family in a complete model.

Hypothesis testing in incomplete models is studied extensively. As discussed earlier, many of them are based on the sample analogs of conditional or unconditional moment restrictions. A challenge in making inferences is the high computational cost of implementing existing methods, as noted in molinari2020microeconometrics. There are attempts to improve the computational tractability of moment-based inference methods within specific classes of models or testing problems, such as those made by ARP and Cox2020, who assume that moment inequality restrictions are linear conditional on observable variables. This paper focuses on another class in which the model is complete under the null hypothesis. This structure makes our test computationally tractable by combining (i) the score function associated with the least favorable parametric model and (ii) a point estimator of the nuisance components.

Practitioners can use this paper's framework to test various hypotheses. For example, one can examine the exsitence of strategic interaction effects and multiple equilibria in static complete information games. Related problems are studied in other models. For incomplete information games, depaulatang2012 introduced a semiparametric inference procedure on the signs of strategic interaction effects. For finite-state Markov games, OtsuPT2016 developed techniques to test whether the conditional choice probabilities, state transition, and other features of games are homogeneous across cross-sectional units. Rejecting their null hypothesis may indicate the presence of multiple equilibria. pelican2020optimal studied a testing problem that involves determining whether agents' preferences are interdependent in a network formation model. The problem includes nuisance parameters that account for degree heterogeneity and homophily. Using the logit structure, they constructed a sufficient statistic for the nuisance parameters and developed a conditional score test. This paper and pelican2020optimal consider testing in settings with a complete model under the null and an incomplete model under the alternative. The two papers take different approaches by exploiting the structures of the respective models. Specifically, we use the least favorable parametric model and estimate nuisance parameters using the restricted MLE. In contrast, pelican2020optimal utilized a sufficient statistic for the nuisance parameters.

One can apply our framework to triangular systems involving a binary outcome and a discrete endogenous variable. We show that taking a control function approach in such a setting leads to a model with an incomplete prediction.\footnote{A nonparametric identification analysis based on set-valued control functions is undertaken in another paper.} We provide a test of endogenous treatment assignments under weak assumptions. To our knowledge, this test is new and provides an alternative to the existing proposal by wooldridge2014JoE, who makes additional high-level assumptions.

Set-up

Let $Y$ be a discrete outcome taking values in a finite set $\mathcal Y$. Let $X\in\mathcal X\subseteq\mathbb R^{d_X}$ be a vector of observable covariates and let $U\in \mathcal U\subseteq\mathbb R^{d_U}$ be a vector of unobservable variables. We equip $\mathcal Y,\mathcal X$, and $\mathcal U$ with their Borel $\sigma$-algebra. In what follows, we use upper case letters for random elements (e.g., $X$) and lower case letters (e.g., $x$) for the values they can take. Let $\theta\in\Theta\subset\mathbb R^{d_\theta}$ be a finite-dimensional parameter, where $\Theta$ is a convex parameter space with a nonempty interior.

A set-valued map $G:\mathcal U\times\mathcal X\times\Theta\leadsto \mathcal Y$ summarizes the prediction of a structural model. We assume $G(\cdot|\cdot;\theta)$ is weakly measurable for every $\theta\in\Theta$ and $Y$ takes one of the values in $G(U|X;\theta)$ with probability 1.\footnote{A set-valued map is weakly measurable if its weak inverse image $G_{-1}(A)\equiv\{s\in\mathcal S :G(s)\cap A\ne\emptyset\}$ is measurable for any open set $A$.} A random element $Y$ with this property is said to be a measurable selection of the random closed set $G(U|X;\theta)$ molchanov2005theory. The map $G$ describes how observable and unobservable characteristics translate into a set of possible outcome values. It reflects restrictions imposed by theory, such as the functional form of utility/profit functions, forms of strategic interaction, and any equilibrium or optimality concepts. Importantly, $G(u|x,\theta)$ can contain multiple values. This feature allows the researchers to encode their lack of understanding of parts of the structural model.

The formulation above also nests the standard setting in which the model is characterized by a reduced form equation:

align[align omitted — 33 chars of source]

for a function $g:\mathcal U\times \mathcal X\times\Theta\to \mathcal Y$. In this case, $G$ is almost surely singleton-valued, i.e., $G(U|X;\theta)=\{g(U|X;\theta)\}$, and we say the model makes a complete prediction.

Throughout, we assume $U$'s law belongs to a parametric family $F=\{F_\theta,\theta\in\Theta\}$, where, for each $\theta$, $F_\theta$ is a probability distribution on $\mathcal U$. To keep notation concise, we use the same $\theta$ for parameters that enter $G$ and that index $F_\theta$. Also, we focus on settings in which $U$ is independent of $X$. However, the framework can be easily extended to settings where $U$ is correlated with $X$, and the researcher specifies its conditional distribution $F_\theta(u|x)$. Furthermore, our framework accommodates settings in which some of the observable covariates are endogenous, and one can construct a set-valued control function (See Example (ref) below).

Motivating examples

We illustrate the objects introduced above with examples. Our first example is a discrete game of complete information BRESNAHANJoE91,CilibertoTamer2009.

example[Discrete Games of Strategic Substitution]\rm There are two players (e.g., firms). Each player may either choose $y^{(j)}=0$ or $y^{(j)}=1$. The payoff of player $j$ is \begin{align} \pi^{(j)}=y^{(j)}\big(x^{(j)}{'}\delta^{(j)}+\beta^{(j)}y^{(-j)}+u^{(j)}\big), \end{align} where $y^{(-j)}\in\{0,1\}$ is the opponent's action, $x^{(j)}$ is player $j$'s observable characteristics, and $u^{(j)}$ is an unobservable payoff shifter. The payoff is summarized in the table below and is assumed to belong to the players' common knowledge. \begin{table}[H] {2pt} \begin{tabular}{cc|c|c|} & \multicolumn{1}{c} & \multicolumn{2}{c}{Player $2$}\\ & \multicolumn{1}{c} & \multicolumn{1}{c}{$y^{(2)}=0$} & \multicolumn{1}{c}{$y^{(2)}=1$} \\\cline{3-4} \multirow{2}*{Player $1$} & $y^{(1)}=0$ & $0, 0$ & $0, x^{(2)}{}{'}\delta^{(2)}+u^{(2)}$ \\\cline{3-4} & $y^{(1)}=1$ & $x^{(1)}{}{'}\delta^{(1)}+u^{(1)}, 0$ & $x^{(1)}{}{'}\delta^{(1)}+\beta^{(1)}+u^{(1)}$, $x^{(2)}{}{'}\delta^{(2)}+\beta^{(2)}+u^{(2)}$ \\\cline{3-4} \end{tabular} \end{table} The key parameter is the strategic interaction effect $\beta^{(j)}$ which captures the impact of the opponent's taking $y^{(-j)}=1$ on player $j$'s payoff. Suppose that $\beta^{(j)}\le 0$ for both players.\footnote{This restriction is often used in models of market entry. Games with strategic complementarity (i.e., $\beta^{(j)}\ge 0$) can be analyzed similarly.} Suppose that the players play a pure strategy Nash equilibrium (PSNE). Then, one can summarize the set of PSNEs by the following correspondence: \begin{align} G(u|x;\theta)= \begin{cases} \{(0, 0)\} & u^{(1)}<-x^{(1)}'\delta^{(1)}, u^{(2)}<-x^{(2)}'\delta^{(2)},\\ \{(0, 1)\} & u\in S_{\theta,(0,1)},\\ \{(1, 0)\} & u\in S_{\theta,(1,0)},\\ \{(1, 1)\} & u^{(1)}>-x^{(1)}'\delta^{(1)}-\beta^{(1)}, u^{(2)}>-x^{(2)}'\delta^{(2)}-\beta^{(2)},\\ \{(1, 0), (0, 1)\} & -x^{(j)}'\delta^{(j)}<u^{(j)}<-x^{(j)}'\delta^{(j)}-\beta^{(j)}, \quad j=1, 2, \end{cases} \end{align} where, $\theta = (\beta', \delta')'$, $S_{\theta,(1,0)}=\{u^{(1)}>-x^{(1)}{}'\delta^{(1)}-\beta^{(1)}, u^{(2)}<-x^{(2)}{}'\delta^{(2)}-\beta^{(2)}\}\cup\{-x^{(1)}{}'\delta^{(1)}<u^{(1)}<-x^{(1)}{}'\delta^{(1)}-\beta^{(1)}, u^{(2)}<-x^{(2)}{}'\delta^{(2)}\}$ and $S_{\theta,(0,1)}=\{u^{(1)}<-x^{(1)}{}'\delta^{(1)}, u^{(2)}>-x^{(2)}{}'\delta^{(2)}\}\cup\{-x^{(1)}{}'\delta^{(1)}<u^{(1)}<-x^{(1)}{}'\delta^{(1)}-\beta^{(1)}, u^{(2)}>x^{(2)}{}'\delta^{(2)}-\beta^{(2)}\}$. Figure (ref) shows the $U$-level sets of $G$ for a given $(x,\theta)$. When $\beta^{(j)}=0$ for either of the players, the model predicts a unique equilibrium for any value of $u=(u^{(1)},u^{(2})'$ (left panel of Figure (ref)). When $\beta^{(j)}<0$ for both players, the model predicts multiple equilibria $\{(0,1),(1,0)\}$ when each $u^{(j)}$ is between the two thresholds $x^{(j)}{}'\delta^{(j)}$ and $x^{(j)}{}'\delta^{(j)}-\beta^{(j)}$ (the blue region in Figure (ref)).
figure[figure omitted — 2,504 chars of source]

The next example is a parametric version of triangular nonseparable equations with binary outcome and treatment chesher2003,shaikh_vytlacil2011. We consider a control function approach applied to the triangular system.

example[Triangular Models with a Set-valued Control Function]\rm Consider a triangular model, in which a binary outcome $Y_{i}$ is determined by a binary treatment $D_i$, a vector $W_i$ of exogenous covariates, and an unobserved variable $\epsilon_i$. The binary treatment $D_i$ is determined by a vector of instrumental variables $Z_i$ and an unobserved variable $V_i$. \begin{align} Y_{i}&=1\{\alpha D_i+W_i'\eta+ \epsilon_i\ge 0 \},\\ D_i&=1\{Z_i'\gamma+V_i\ge0\}. \end{align} The unobserved variables $(\epsilon_i,V_i)$ may be dependent rendering $D_i$ potentially endogenous. We assume that $(W_i,Z_i)$ is independent of $(\epsilon_i,V_i)$. If one could recover $V_i$ from the observables (which would be possible with a continuous $D_i$), conditioning on $V_i$ would make $\epsilon_i$ independent of $D_i$, i.e. $\epsilon_i|D_i,V_i\sim \epsilon_i|V_i$. This control function approach would allow us to recover structural parameters Blundell:2004td,imbens2009identification,wooldridgeJHR. With a binary endogenous variable, we cannot uniquely recover $V_i$.\footnote{Alternatively, wooldridge2014JoE uses the generalized residual $r_i=d_i\lambda(z_i'\gamma)-(1-d_i)\lambda(-z_i'\gamma)$ from the first stage MLE, where $\lambda$ is the inverse Mills ratio. He makes additional high-level assumptions so that $r_i$ is a sufficient statistic for capturing the endogeneity of $d_i$ and proposes an estimator of the average structural function. Instead of taking this approach, we explore what can be learned from the set-valued control function.} However, the model restricts $V_i$ to the following set-valued control function: \begin{align} \mathbf V(D_i,Z_i;\gamma) &\equiv\begin{cases} [-Z_i'\gamma,\infty)& if D_i=1\\ (-\infty,-Z_i'\gamma]& if D_i=0. \end{cases} \end{align} This set contains the actual control function $V_i$ as its measurable selection. Suppose that the conditional distribution $\epsilon_i|V_i$ belongs to a location family, in which the location parameter is $\beta V_i$. Then, one may write $\epsilon_i=\beta V_i+U_i$ for some $U_i$ independent of $(D_i,V_i)$. Substituting this expression into (ref), we can summarize the model's prediction by \begin{align} G(u_i|x_i;\theta)=\Big\{y_i\in\{0,1\}:y_{i}=1\{\alpha d_i+w_i'\eta+ \beta v_i+u_i\ge 0 \}, for some v_i\in \mathbf V(d_i,z_i;\gamma)\Big\}, \end{align} where $x_i=(d_i,w_i',z_i')'$ and $\theta=(\beta,\delta')'$ with $\delta=(\alpha,\eta',\gamma')'$. When $\beta$ is positive, one can simplify $G$ further: \begin{align} G(u_i|1,w_i,z_i;\theta)=\begin{cases} \{1\} & u_i\ge -\alpha d_i-w_i'\eta+\beta z_i'\gamma\\ \{0,1\} & u_i< -\alpha d_i-w_i'\eta+\beta z_i'\gamma. \end{cases} \end{align} As we show below, the control function approach allows one to test the endogeneity of $D_i$ even if the control function is set-valued.\footnote{We take a control function approach that conditions on $v_i$, which only requires specification of the conditional distribution of $\epsilon_i$ given $v_i$. Alternatively, one could take the vector $(y_i,d_i)$ as endogenous variables and specify the joint distribution of $(\epsilon_i,v_i)$. This alternative approach with a stronger assumption would imply a complete model lewbel2007ier.}

The next example is a panel dynamic discrete choice model heckman78,chamberlain_1985,hyslop99.

example[Panel Dynamic Discrete Choice Models]\rm An individual makes binary decisions across multiple periods according to \begin{align} Y_{it}=1\{X_{it}'\eta+Y_{it-1}\beta+\alpha_i+\epsilon_{it}\ge 0\}, i=1,\dots, n, t=1,\dots,T, \end{align} where $Y_{it}$ is a binary outcome for individual $i$ in period $t$, $X_{it}$ is a vector of observable covariates, $\alpha_i$ is an unobservable individual specific effect, and $\epsilon_{it}$ is an unobserved idiosyncratic error. We use $a_i\in\mathbb R$ and $e_{it}\in\mathbb R$ to denote realizations of $\alpha_i$ and $\epsilon_{it}$ respectively. If $\beta$ is nonzero, the individual's choice in period $t$ depends on her past choice, rendering the decision state dependent. Suppose the researcher observes $(Y_{it},X_{it})$ for $i=1,\dots,n$ and $t=1,\dots,T$. Without any knowledge of $Y_{i0}$, the dynamic restrictions (ref) alone do not fully determine the value of $Y_{i1},\dots,Y_{iT}$ heckman78,heckman1987incidental,honore2006bounds.\footnote{As an alternative, one could work with the likelihood function conditional on the initial observation. However, this approach can be problematic if one wants to be internally consistent across different periods honore2006bounds,wooldridge2005simple.} Consider $T=2$. Suppose for the moment $y_{i0}=0$. For a given $(x_i,a_i, e_{i1},e_{i2})$ and $(\beta,\eta)$, the outcome $y_i=(y_{i1},y_{i2})$ must satisfy \begin{align} y_{i1}&=1\{x_{i1}'\eta+a_i+e_{i1}\ge 0\}\\ y_{i2}&=1\{x_{i2}'\eta+y_{i1}\beta+a_i+e_{i2}\ge 0\}. \end{align} Similarly, if $y_{i0}=1$, the outcome must satisfy \begin{align} y_{i1}&=1\{x_{i1}'\eta+\beta+a_i+e_{i1}\ge 0\}\\ y_{i2}&=1\{x_{i2}'\eta+y_{i1}\beta+a_i+e_{i2}\ge 0\}. \end{align} Without further assumptions, the model permits both possibilities. Let $U_i=(U_{i1},U_{i2})'$ with $U_{it}=\alpha_i+\epsilon_{it}$. One can summarize the model prediction by \begin{align} G(u_i|x_i;\theta)=\Big\{y_i=(y_{i1},y_{i2})\in \{0,1\}^2: y_i satisfies either (ref)-(ref) or (ref)-(ref)\Big\}. \end{align} If $\beta\ge 0$, one can express this correspondence as follows:\footnote{Appendix (ref) provides details and a graphical illustration of $G$. } \begin{align} G(u_i|x_i;\theta)=\begin{cases} \{(0,0)\} & u_{i1}<-x_{i1}'\eta-\beta, u_{i2}<-x_{i2}'\eta,\\ \{(0,1)\} & u_{i1}<-x_{i1}'\eta-\beta, u_{i2}\ge-x_{i2}'\eta,\\ \{(1,0)\} & u_{i1}\ge-x_{i1}'\eta, u_{i2}<-x_{i2}'\eta-\beta,\\ \{(1,1)\} & u_{i1}\ge-x_{i1}'\eta, u_{i2}\ge-x_{i2}'\eta-\beta,\\ \{(0,0),(1,0)\} & -x_{i1}'\eta-\beta\le u_{i1}< -x_{i1}'\eta, u_{i2}\le -x_{i2}'\eta-\beta,\\ \{(0,0),(1,1)\} & -x_{i1}'\eta-\beta\le u_{i1}< -x_{i1}'\eta, -x_{i2}'\eta-\beta\le u_{i2}<-x_{i2}'\eta,\\ \{(0,1),(1,1)\} & -x_{i1}'\eta-\beta\le u_{i1}< -x_{i1}'\eta, u_{i2}\ge-x_{i2}'\eta. \end{cases} \end{align} Figure (ref) summarizes $G$. The model makes a complete prediction when there is no state dependence, i.e., $\beta=0$ (left panel). \begin{figure}[h] \caption{Level sets of $u\mapsto G(u|x;\theta)$ when $\beta\ge 0$} \begin{tikzpicture} [scale=0.8, domain=-3:3,>=latex] \draw[->] (-3,-3) -- (3,-3) node[right] {$u_1$}; \draw[->] (-3,-3) -- (-3,3) node[above] {$u_2$}; \draw[thick, dashed] (-3,1.1) -- (-1,1.1); \draw[thick, dashed] (1.2,1.1) -- (3,1.1); \draw[thick, dashed] (1.2,-3) -- (1.2,1.1); \draw[thick, dashed] (-1,1.1) -- (1.2,1.1); \draw[thick, dashed] (1.2,1.1) -- (1.2,3); \filldraw (1.2,1.1) circle (2pt); \draw (1.25,1.4) node[right]{$B$}; \draw (-1,-2) node[left] { $\{(0,0)\}$}; \draw (1.3,2) node [right]{ $\{(1,1)\}$}; \draw (-1,2) node[left] { $\{(0,1)\}$}; \draw (1.3,-2) node[right] { $\{(1,0)\}$}; \end{tikzpicture} \begin{tikzpicture} [scale=0.8, domain=-3:3,>=latex] \draw[->] (-3,-3) -- (3,-3) node[right] {$u_1$}; \draw[->] (-3,-3) -- (-3,3) node[above] {$u_2$}; \draw[thick, dashed] (-3,1.1) -- (-1,1.1); \draw[thick, dashed] (-1,-1) -- (-1,3); \draw[thick, dashed] (-1,-1) -- (3,-1); \filldraw (-1,-1) circle (2pt); \draw (-1,-0.5) node[right]{$A$} ; \draw[thick, dashed] (1.2,-3) -- (1.2,1.1); \draw[thick, dashed] (-1,1.1) -- (1.2,1.1); \draw[thick, dashed] (1.2,1.1) -- (1.2,3); \draw[thick, dashed] (-1,-1) -- (-1,-3); \filldraw (1.2,1.1) circle (2pt); \draw (1.25,1.4) node[right]{$B$}; \fill[fill=blue,opacity=0.4] (-1,-1) -- (1.2,-1) -- (1.2,1.1) -- (-1,1.1); \fill[fill=red,opacity=0.4] (-1,1.1) -- (1.2,1.1) -- (1.2,3) -- (-1,3); \fill[fill=green,opacity=0.4] (-1,-1) -- (1.2,-1) -- (1.2,-3) -- (-1,-3); \draw (-1,-2) node[left] { $\{(0,0)\}$}; \draw (1.3,2) node [right]{ $\{(1,1)\}$}; \draw (-1,2) node[left] { $\{(0,1)\}$}; \draw (1.3,-2) node[right] { $\{(1,0)\}$}; \draw (-0.8,0.2) node [right]{ $\{(0,0),$}; \draw (-0.2,-0.5) node [right]{ $(1,1)\}$}; \draw (-0.8,2.4) node [right]{ $\{(0,1),$}; \draw (-0.2,1.7) node [right]{ $(1,1)\}$}; \draw (-0.8,-1.7) node [right]{ $\{(0,0),$}; \draw (-0.2,-2.4) node [right]{ $(1,0)\}$}; \end{tikzpicture} \begin{minipage}{0.7\textwidth} { Note: The level sets of $G$ when $\beta=0$ (left) and $\beta>0$ (right). $A=(-x_{i1}{}{'}\eta-\beta, -x_{i2}{}{'}\eta-\beta)$; $B=(-x_{i1}{}{'}\eta, -x_{i2}{}{'}\eta)$. Multiple outcome values are predicted in the red, blue, and green regions. } \end{minipage} \end{figure}

Testing Hypotheses

Let $\beta\in\Theta_\beta\subset\mathbb R^{d_\beta}$ denote the subvector of $\theta$ whose value determines whether the model is complete or not. Let $\delta\in\Theta_\delta\subset\mathbb R^{d_\delta}$ collect the remaining components of $\theta$. Given a sample of data $(Y_i,X_i),i=1,\dots, n$, consider testing

align[align omitted — 84 chars of source]

where $B_1\subset\Theta_\beta$ is a set not containing $\beta_0$. For instance, in entry games (i.e. Example (ref)), the presence of strategic substitution effects can be tested by letting $\beta_0=0$ and $B_1=\{\beta:\beta^{(j)}<0,j=1,2\}$.\footnote{Our framework also nests settings in which the researcher tests one of the interaction effects, e.g., $\beta^{(1)}_0=0$ and $B_1=\{\beta:\beta^{(1)}<0\}$. In this case, the model is complete under both hypotheses. Our test then reduces to a conventional score test.} Similarly, we may test the potential endogeneity of treatment assignments (Example (ref)) and the presence of state dependence (Example (ref)) by setting $\beta_0=0$ and choosing suitable alternative hypotheses. In what follows, we let $\Theta_0=\{\beta_0\}\times \Theta_\delta$ and $\Theta_1=B_1\times\Theta_\delta$.

Let $\Delta_{Y|X}$ denote the set of conditional distributions of $Y$ given $X$. For each $\theta=(\beta',\delta')'$, an incomplete model admits the following set of conditional distributions:

multline[multline omitted — 228 chars of source]

The conditional distribution $\eta(\cdot|x,u)$ represents the unknown selection mechanism according to which an outcome gets selected from $G(u|x;\theta)$. Reflecting the lack of understanding of the selection, we allow any law supported on $G(u|x;\theta)$. Consequently, the model can admit (infinitely) many likelihood functions for a given $\theta$. Let $\mu$ be the counting measure on $\mathcal Y.$ For each $\theta$, define

align[align omitted — 93 chars of source]

This set collects all (conditional) densities compatible with a given $\theta$. In the case of discrete games (Examples (ref)), this set contains all densities of equilibrium outcomes that are compatible with the game's description. Similarly, in the context of panel discrete choice (Example (ref)), this set collects all densities of individual choices consistent with arbitrary specifications of the initial condition. Observe that $\mathfrak q_\theta$ reduces to a singleton set $\{q_\theta\}$ if the model is complete, in which case $q_\theta=dQ_\theta/d\mu$ with $Q_\theta(A|x)=\int 1\{g(u|x;\theta)\in A\}dF_\theta(u)$.

While the multiplicity of likelihood functions may appear challenging, $\mathfrak q_\theta$ can be simplified, and this property also simplifies our tests. By Artstein's inequality, we may rewrite $\mathfrak q_\theta$ as follows galichon2011set,molinari2020microeconometrics:

align[align omitted — 149 chars of source]

where

align[align omitted — 98 chars of source]

is the conditional containment functional (or belief function) associated with the random set $G(u|x;\theta)$. This function gives the sharp lower bound for the conditional probability $Q(A|x)$ across all $Q$'s belonging to $\mathcal Q_\theta$.\footnote{The upper bound for $Q(A|x)$ is given by the capacity functional $\nu^*(A|X)=F_\theta(G(u|x;\theta)\cap A\ne\emptyset|x)$. It is sufficient to use either of the lower or upper bounds in (ref) because the bounds are related to each other through the conjugate relationship $\nu(A|x)=1-\nu^*(A^c|x)$.} Theoretical properties of the containment functional and numerical approximation methods are well studied.\footnote{See molchanov2005theory for a general treatment. For numerical approximations, see CilibertoTamer2009,galichon2011set. We briefly review them in Appendix (ref). } For us, it is important that the linear inequalities in (ref) characterize $\mathfrak q_\theta$. Together with an extended Neyman-Pearson lemma reviewed below, this characterization makes the score computation feasible. In the following subsection, we briefly review the existing results we will rely on.

Least Favorable Parametric Model

Let $p_0(y|x)$ denote the true conditional density of $Y$ given $X$. Consider distinguishing a parameter value $\theta_0$ from another value $\theta_1$ in a parametric model $\{p_\theta,\theta\in\Theta\}$ of conditional densities. This amounts to testing a simple null hypothesis $p_0=p_{\theta_0}$ against a simple alternative hypothesis $p_0=p_{\theta_1}$. By the Neyman-Pearson lemma, the most powerful test is the likelihood-ratio (LR) test. In incomplete models, corresponding null and alternative hypotheses would be $p_0\in\mathfrak q_{\theta_0}$ and $p_0\in\mathfrak q_{\theta_1}$ rendering both hypotheses composite. kz showed that it was possible to extend the Neyman-Pearson lemma to such settings, building on a general result by HS. We briefly summarize their results below.

Let $\phi:\mathcal Y\times\mathcal X\to[0,1]$ be a test and $E_q[\phi(Y,X)]$ be its rejection probability, under conditional density $q$ and the marginal distribution of $X$.\footnote{The rejection probability can be written as $E_q[\phi(Y,X)]=E[E_q[\phi(Y,X)|X]]$. Only the conditional expectation depends on $q$.} For a given $\theta$, the power guarantee of $\phi$ is $\pi_{\theta}(\phi)\equiv\inf_{q\in\mathfrak q_{\theta}}E_q[\phi(Y,X)]$, which is the power value certain to be obtained regardless of the unknown selection mechanism. kz seeked for a level-$\alpha$ minimax test lehmann2005testing such that

align[align omitted — 97 chars of source]

and

align[align omitted — 130 chars of source]

The minimax test is a procedure that maximizes the power guarantee among tests that meet the uniform size control requirement.

Results from HS imply that, when $\mathfrak q_{\theta_0}\cap\mathfrak q_{\theta_1}=\emptyset$, the rejection region of a minimax test is of the form $\{(y,x):\Lambda(y,x)>t\}$ for a measurable function $\Lambda:\mathcal Y\times\mathcal X\to\mathbb R$. Furthermore, there is a least favorable pair (LFP) of densities $(q_{\theta_0},q_{\theta_1})\in \mathfrak q_{\theta_0}\times\mathfrak q_{\theta_1}$ such that for all $t\ge 0$,

align[align omitted — 116 chars of source]

and

align[align omitted — 118 chars of source]

where $\Lambda(y,x)=q_{\theta_1}(y|x)/q_{\theta_0}(y|x)$. That is, $q_{\theta_0}\in \mathfrak q_{\theta_0}$ is a density consistent with $\theta_0$ and least favorable for controlling the size of a test among all elements of $\mathfrak q_{\theta_0}$, and $q_{\theta_1}\in\mathfrak q_{\theta_1}$ is a density consistent with $\theta_1$ and least favorable for maximizing a measure of power among all elements of $\mathfrak q_{\theta_1}$. kz exploited these features to show that a level-$\alpha$ minimax test is an LR test based on the LFP:

align*[align* omitted — 130 chars of source]

where $(C,\gamma)$ solves $E_{q_{\theta_0}}[\phi(Y,X)]=\alpha.$

One can also characterize the LFP $(q_{\theta_0},q_{\theta_1})$ as a solution to the following convex program.

align[align omitted — 389 chars of source]

The constraints in (ref) and (ref) are the sharp identifying restrictions.\footnote{A common way to use them for identification analysis is to define the sharp identified set as $\Theta_I=\{\theta:P(A|x)\ge \nu_\theta(A|x),a.s.\}.$ That is, given the conditional probability $P(\cdot|x)$ identified from data, one collects all values of $\theta$ satisfying the sharp identifying restrictions. For hypothesis testing, we instead fix $\theta$ and ask what would be a distribution among all distributions satisfying the sharp identifying restrictions, which is least favorable for controlling the size or maximizing the power.} In view of (ref), they are equivalent to imposing the restrictions that $q_0$ belongs to $\mathfrak q_{\theta_0}$ and $q_1$ belongs to $\mathfrak q_{\theta_1}$ respectively. These restrictions are useful for computing the LFP because they are linear in $(q_0,q_1)$. Also, the objective function is strictly convex in $(q_0,q_1)$. One can solve the convex program above numerically in general. For the examples we discussed earlier, it is also possible to compute the LFP analytically (see Appendix (ref)).

To illustrate, let us consider Example (ref). Suppose that the latent payoff shifters $(U^{(1)},U^{(2)})$ follow a bivariate standard normal distribution. We may then compute $\nu_\theta(A|x)$ for each event. Let us take $A=\{(1,0)\}$ as an example. Using (ref) and (ref), we obtain

multline[multline omitted — 210 chars of source]

The last expression corresponds to the probability assigned to the green region in Figure (ref) (right panel) and is the sharp lower bound for the probability of $A=\{(1,0)\}$.

Now consider two parameter values $\theta_0=(0'_2,\delta')'$ and $\theta_1=(\beta',\delta')'$, where $\beta=(\beta^{(1)},\beta^{(2)})'$ with $\beta^{(j)}<0$ for $j=1,2$. As discussed in more detail below, the model is complete when $\beta=0.$ One can show that (ref) reduces to the following equality restrictions:

align[align omitted — 411 chars of source]

They uniquely determine the least-favorable null density $q_{\theta_0}$ as follows

multline[multline omitted — 211 chars of source]

where, to ease notation, we use $\Phi_{1}$ and $\Phi_{2}$ to denote $\Phi(x^{(1)}{}'\delta^{(1)})$ and $\Phi(x^{(2)}{}'\delta^{(2)})$.

When $\beta^{(j)}<0,j=1,2$, there are multiple densities satisfying (ref). The least favorable alternative density $q_{\theta_1}$ can be found by minimizing (ref) with respect to $q_1$ subject to (ref). The solution can be expressed analytically. For example, when player 1's strategic interaction effect on player 2 is relatively high, it is given by the following form:\footnote{Appendix (ref) gives a full characterization of the LFP for Example (ref). }

multline[multline omitted — 407 chars of source]

Comparing (ref) and (ref), one can see that $q_{\theta_1}$ tends to $q_{\theta_0}$ as $\beta$ approaches its null value (i.e. 0). Hence, one may view $\theta\mapsto q_{\theta}$ as a “parametric” model. For each $\theta_1$, the density $q_{\theta_1}$ corresponds to the data generating process that is least favorable in detecting $\beta$'s deviation from its null value among all densities compatible with $\theta_1$. By varying $\theta_1$, we may trace out a family of such densities and form a parametric model. We define a least favorable (LF) parametric model as follows.

definition[Least favorable parametric model] Let $\tilde\Theta\subseteq \Theta$ be an open set containing $\Theta_0$. A family of densities $\{q_\theta:\theta\in\tilde\Theta\}$ is the least favorable (LF) parametric model indexed by $\theta\in\tilde\Theta$ if (a) $q_\theta$ is the unique element of $\mathfrak q_\theta$ for any $\theta\in\Theta_0$; and (b) $q_\theta$ is the density of the least-favorable alternative distribution $Q_1$ if $\theta\in \tilde\Theta\setminus \Theta_0$, i.e., $Q_1$ constitutes a LFP $(Q_0,Q_1)\in \mathcal P_{\theta_0}\times\mathcal P_{\theta}$ with $\theta_0=(\beta_0',\delta')'$ and $\theta=(\beta',\delta')'$.

Part (a) of the definition above is natural because the model is complete under the null hypothesis. Part (b) deserves discussion. We associate each alternative parameter value $\theta\not\in\Theta_0$ with a density that corresponds to the least favorable density for testing between $\theta_0=(\beta_0,\delta)$ and $\theta=(\beta,\delta)$. In other words, for each $\delta$, we focus on testing $\beta_0$ against $\beta$ using the least-favorable distribution for maximizing power. We construct $q_\theta$ this way because we aim at detecting local deviations in terms of $\beta$. As we see below, this approach allows us to capture the sensitivity of $q_\theta$ with respect to a change in the parameter of interest through a score function.\footnote{An alternative choice of $\theta_0\in\Theta_0$ is also possible, which we do not seek here because it does not seem to lead to a tractable score test.} We provide a sufficient condition for the existence of an LF parametric model in the next section.

Let us also note the following points. First, having a complete null model is not sufficient for obtaining a score function. It also requires us to define parametric densities at local alternatives. Our approach is to select the density associated with the least favorable selection mechanism for power maximization. Second, we do not need to know the precise form of the selection mechanism that induces $q_{\theta}$ (for $\theta\in \Theta_1$). Solving the convex program, we “profile out” the selection mechanism and directly obtain the induced density $q_{\theta}$. This is why $q_{\theta}$ is a function of $\theta$ only and does not involve any selection mechanism.

Coming back to equation (ref), the formula suggests we may pretend as if data were generated by a parametric discrete choice model with the given density. Thanks to this feature, most of our analysis below will resemble that of standard discrete choice models.

Model Completeness under the Null Hypothesis

In this section, we start with an assumption to ensure that the LF parametric model is well-defined over a parameter set that contains the null parameter space $\Theta_0$ and its neighborhood. For this, let $\theta_{h}=(\beta_0'+h',\delta')'$, and let $\mathbb C_\epsilon$ denote an open cube centered at the origin with edges of length $2\epsilon$.

assumption(i) Under any null parameter value $\theta_0=(\beta_0',\delta')'$ with $\delta\in\Theta_\delta$, the model makes a complete prediction so that $\mathfrak q_{\theta_0}=\{q_{\theta_0}\}$ is a singleton set; (ii) There exists $\epsilon>0$ such that the two sets $\mathfrak q_{\theta_0}$ and $\mathfrak q_{\theta_{h}}$ are disjoint for any $h\in (B_1-\beta_0)\cap \mathbb C_\epsilon.$

Assumption (ref) (i) holds whenever the model makes a complete prediction under the null hypothesis and is satisfied in the examples discussed in Section (ref). The model can be incomplete under the alternative hypothesis. Assumption (ref) (ii) requires $\mathfrak q_{\theta_{h}}$ does not share any element with $\mathfrak q_{\theta_0}$. Under this condition, it is possible to detect local deviations from the null hypothesis regardless of the unknown selection mechanism. Such alternatives are robustly testable in the sense that there exists a test that has nontrivial power against any distribution in $\mathfrak q_{\theta_{h}}$ kz. For this condition, it suffices to have an event $A\subset\mathcal Y$ such that $q_{\theta_0}(A|x)<\nu_{\theta_{h}}(A|x)$ (or $q_{\theta_0}(A|x)>\nu_{\theta_{h}}^*(A|x)$) for all $\tau>0$ for some $x\in\mathcal X$. All of the examples in Section (ref) satisfy Assumption (ref) (ii), and we demonstrate how to show this condition in Appendix (ref). We construct score tests that have power against robustly testable local alternatives.\footnote{kz extend the notion of local alternatives and analyze a more general setting that does not require Assumption (ref) (ii). We conjecture that we may extend our framework similarly. Since all of our examples satisfy Assumption (ref) (ii), we leave this extension elsewhere.}

Let us revisit the examples. \setcounter{example}{0}

example[Discrete Games of Strategic Substitution]\rm Consider testing the presence of strategic substitution effects by testing $H_0:\beta^{(1)}=\beta^{(2)}=0$ against $H_0:\beta^{(1)}<0,\beta^{(2)}<0$. Under the null hypothesis, there is no strategic interaction between the players, which leads to the following complete prediction: \begin{align} G(u|x;\theta_0)= \begin{cases} \{(0, 0)\} & u^{(1)}<-x^{(1)}'\delta^{(1)}, u^{(2)}<-x^{(2)}'\delta^{(2)},\\ \{(1, 1)\} & u^{(1)}>-x^{(1)}'\delta^{(1)}, u^{(2)}>-x^{(2)}'\delta^{(2)},\\ \{(1, 0)\} & u^{(1)}>-x^{(1)}'\delta^{(1)}, u^{(2)}\le -x^{(2)}'\delta^{(2)},\\ \{(0, 1)\} & u^{(1)}\le -x^{(1)}'\delta^{(1)}, u^{(2)}>-x^{(2)}'\delta^{(2)}. \end{cases} \end{align} Hence, for any value of the observed and unobserved variables, $G(u|x;\theta_0)$ contains a unique equilibrium outcome (left panel of Figure (ref)). Combining the complete prediction with a parametric assumption on $U$ ensures Assumption (ref) (i) We use this example to discuss Assumption (ref) (ii). Under any $\theta_1$ with $\beta^{(j)}<0,j=1,2$, the probability allocated to $A=\{(1,1)\}$ is lower than that under the null as (right panel in Figure (ref)), provided $U$ is continuously distributed over $\mathbb R^2$. Hence, $\nu_{\theta_0}^*(\{(1,1)\}|x)<\nu_{\theta_1}(\{(1,1)\}|x)$, which ensures Assumption 1 (ii).

We also note that Examples 2-3 reduce to complete models under the null hypothesis.\footnote{To save space, we show Assumption (ref) (ii) for these examples in Appendix (ref).}

example[Triangular Models with a Set-valued Control Function]\rm Consider testing the endogeneity of the treatment by testing the hypothesis that the coefficient $\beta$ on the control function $v$ is 0. When the null hypothesis is true, the model's prediction reduces to \begin{align} y_{i}=1\{\alpha d_i+w_i'\eta+u_i\ge 0 \}, i=1,\dots,n. \end{align} This is a standard binary choice model with exogenous covariates. For a given $(x_i,u_i)$, $y_i$ is uniquely determined. There is no need to control for $V_i$ because $U_i$ is independent of $(D_i,W_i)$.
example[Panel Dynamic Discrete Choice Models]\rm Consider testing the presence of state dependence by testing whether the coefficient $\beta$ on the lagged dependent variable is 0. When $\beta=0$ in (ref), the model reduces to a static panel binary choice model: \begin{align} Y_{it}=1\{X_{it}'\eta+\alpha_i+\epsilon_{it}\ge 0\}, i=1,\dots, n, t=1,\dots,T, \end{align} which makes (ref)-(ref) and (ref)-(ref) equivalent. Under $H_0$, $G(u_i|x_i;\theta_0)$ contains the unique outcome value satisfying (ref).

We conclude this subsection with the following proposition.

propositionSuppose Assumption (ref) holds. Then, a least-favorable parametric model $\{q_\theta:\theta\in\tilde\Theta\}$ exists for $\tilde\Theta=\{\theta=(\beta_0+h,\delta):h\in (B_1-\beta_0)\cap \mathbb C_\epsilon,\delta\in\Theta_\delta\}$.

Score Tests

Score-based tests such as Rao's score (or Lagrange multiplier) test and Neyman's $C(\alpha)$ test are widely used. They require the estimation of the restricted model only, which is particularly attractive in our setting. The restricted model is complete and typically admits point estimation of nuisance parameters under reasonably weak conditions. We take advantage of this property to carry out a score-based test. Below, we briefly review the core ideas behind the classic score tests and discuss extensions to handle potential model incompleteness under the alternative. For expositional purposes, we assume $q_\theta$ is differentiable with respect to $\theta$ for now and will weaken this assumption later.

Consider testing the null parameter value $\theta_0=(\beta_0',\delta')'$ against a local alternative hypothesis $\theta_h=(\beta_0'+h',\delta')'$, where $h\in\mathbb R^{d_\beta}$. As discussed earlier, the optimal test in terms of guaranteed power is the likelihood-ratio test based on the LFP. The test is also robust in that, under Assumption (ref) (ii), the log-likelihood ratio can detect any deviation from the null hypothesis with non-trivial power regardless of the selection mechanism.

One can locally approximate the log-likelihood ratio by $\sum_{i=1}^n h's_{\beta}(Y_i|X_i;\beta_0,\delta)$, where $s_{\beta}(y|x;\beta,\delta)=\frac{\partial}{\partial \beta }\ln q_\theta(y|x)|_{\theta=(\beta,\delta)}$ is the score function. Let $\Sigma_{\beta_0}=\text{Var}(\sum_{i=1}^ns_{\beta}(Y_i|X_i;\beta_0,\delta))$. For i.i.d. data, $\Sigma_{\beta_0}=n I_{\beta_0}$ where $I_{\beta_0}=E[s_{\beta}(Y_i|X_i;\beta_0,\delta)s_{\beta}(Y_i|X_i;\beta_0,\delta)']$. For a fixed $h$, the normalized quantity

align[align omitted — 118 chars of source]

serves as a robust measure of discrimination between $\beta_0$ and $\beta_0+h$. It locally approximates the log-likelihood ratio $\ln (q_{\theta_h}/q_{\theta_0})$. The direction $h$ that maximizes (ref) is $h^*=I_{\beta_0}^{-1}\frac{1}{\sqrt n}\sum_{i=1}^n s_{\beta}(Y_i|X_i;\beta_0,\delta)$, which motivates Rao's score statistic:\footnote{See Bera:2001va for a more detailed argument for complete models. The same argument can be applied to incomplete models by replacing the standard likelihood function with the LF density $q_\theta$.}

align[align omitted — 288 chars of source]

Suppose that the nuisance parameter $\delta$ can be estimated by a point estimator $\hat\delta_n$. Evaluating the sample mean of the score at $\delta=\hat\delta_n$ and imposing the null hypothesis yields

align[align omitted — 94 chars of source]

A feasible version of (ref) is

align[align omitted — 82 chars of source]

where $\hat V_n$ is an estimator of the asymptotic variance $V_0\equiv I_{\beta_0}$. For example, one can use the sample analog $\hat V_n= n^{-1}\sum_{i=1}^ns_\beta(Y_i|X_i;\beta_0,\hat\delta_n)s_\beta(Y_i|X_i;\beta_0,\hat\delta_n)'$ or its regularized version.\footnote{For example, the following estimator proposed by andrews_barwick12 ensures that it is always nonsingular and is equivariant to scale changes

align[align omitted — 97 chars of source]

where $\hat\Sigma_n=n^{-1}\sum_{i=1}^ns_\beta(Y_i|X_i;\beta_0,\hat\delta_n)s_\beta(Y_i|X_i;\beta_0,\hat\delta_n)'$, $\hat D_n=\text{diag}(\hat\Sigma_n)$, and $\hat\Omega_n=\hat D_n^{-1/2}\hat\Sigma_n\hat D_n^{-1/2}$. } Under regularity conditions, $\hat T_n$ converges in distribution to a $\chi^2$-distribution with $d_\beta$ degrees of freedom under the null hypothesis.

The analysis so far presumed that $q_\theta$ was differentiable, and $h\in\mathbb R^{d_\beta}$ was unrestricted. These assumptions may be restrictive in our context. For example, in discrete games of complete information, the least favorable parametric model $h\mapsto q_{\theta_h}$ and its score can take different functional forms depending on whether the alternative hypothesis admits strategic substitution (i.e. $h<0$ as in Example (ref)) or strategic complementarity (i.e. $h>0$). It is then natural to analyze these two cases separately. Below, we weaken differentiability requirements to accommodate these features and allow the alternative hypothesis to be restricted (e.g., one-sided).

Recall that a set $\Gamma\subseteq\mathbb R^d$ is said to be locally equal to set $\Upsilon\subseteq\mathbb R^d$ if $\Gamma\cap \mathbb C_\epsilon=\Upsilon\cap \mathbb C_\epsilon$ for some $\epsilon>0$ Andrews:1999aa.

assumption[$L^2$-directional differentiability] (i) $B_1-\beta_0$ is locally equal to a convex cone $\mathcal V_1$; (ii) For any $\zeta\in \mathcal V_1\times \mathbb R^{d_\delta}$, there exists a square integrable function $s_{\theta}=(s_\beta',s_\delta')':\mathcal Y\times\mathcal X\to\mathbb R^d$ such that \begin{align} \Big\|q_{\theta_0+\tau \zeta}^{1/2}-q_{\theta_0}^{1/2}(1+\frac{1}{2}\tau \zeta's_{\theta}(\cdot|\cdot;\beta_0,\delta))\Big\|_{L^2_\mu}=o(\tau), \end{align} as $\tau\downarrow0$.

Assumption (ref) (i) requires the set of deviations (from $\beta_0$) can be locally approximated by a convex cone. In Example (ref), consider testing $H_0:\beta=(0,0)'$ against $H_1:\beta^{(1)}<0,\beta^{(2)}<0$. Then, $B_1-\beta_0$ is locally equal to

align[align omitted — 70 chars of source]

Assumption (ref) (ii) uses the notion of differentiability in quadratic mean Van-der-Vaart:2000aa, but it only requires that a unique score, in the sense of the $L^2$-derivative of the square-root density, exists for the set $\mathcal V_1$ of local deviations from the null hypothesis. This weaker assumption is appropriate for incomplete models, and $s_\theta$ can be derived from the least favorable parametric model similar to the standard parametric models (see Appendix (ref)).

To accommodate the one-sided nature of the alternative hypothesis, we define a test statistic by

align[align omitted — 153 chars of source]

This test statistic is a modification of (ref) and follows the construction in Silvapulle:1995tm. It requires the same functions of data as $T_n$, but it is designed to direct power against the local alternatives in $\mathcal V_1$. If the alternative hypothesis is locally unrestricted, i.e., $\mathcal V_1=\mathbb R^{d_\beta}\setminus 0$, the test statistic reduces to $T_n$.

The asymptotic distribution of $\hat S_n$ is no longer a $\chi^2$-distribution. However, its critical value is easy to compute using simulations. Let

align[align omitted — 91 chars of source]

where

align[align omitted — 109 chars of source]

which can be simulated by drawing $Z$ repeatedly from a zero mean multivariate normal distribution with estimated variance $\hat V_n.$

Restricted Maximum Likelihood Estimator

Let $q_{\beta_0,\delta}$ be the conditional density of $Y_i$ given $X_i$. By Assumption (ref) (i), this density is unique. A natural estimator of $\delta$ is the restricted maximum likelihood estimator (RMLE) $\hat\delta_n$, which maximizes the log-likelihood function

align[align omitted — 114 chars of source]

The complete model (under $H_0$) is often a standard discrete choice problem. Hence, one can use package software (e.g., R, Stata) to compute the RMLE. Let us revisit the examples.

\setcounter{example}{0}

example[Discrete Games of Strategic Substitution]\rm Under $H_0:\beta^{(1)}=\beta^{(2)}=0$, the model has a unique likelihood function as discussed in Section (ref). The RMLE $\hat\delta_n$ maximizes \begin{align*} \mathbb M_n(\delta)&=\sum_{i=1}^n \Big( 1\{Y_i=(0,0)\} \ln[(1-\Phi_{1,i})(1-\Phi_{2,i})] +1\{Y_i=(0,1)\}\ln [(1-\Phi_{1,i})\Phi_{2,i}]\\ &\qquad+1\{Y_i=(1,0)\}\ln[\Phi_{1,i}(1-\Phi_{2,i})]+1\{Y_i=(1,1)\}\ln[\Phi_{1,i}\Phi_{2,i}]\Big), \end{align*} where $\Phi_{j,i}=\Phi(X_i^{(j)}{}'\delta^{(j)}),j=1,2$.
example[Triangular Models with a Set-valued Control Function]\rm When $\beta=0$, there is no correlation between the errors in the outcome and selection equations. The outcome equation reduces to a binary choice model with exogenous covariates. If we assume $U_i\sim N(0,1)$, we obtain a probit model with \begin{align} P(Y_i=1|D_i=d_i,W_i=w_i)= q_{\beta_0,\delta}(1|d_i,w_i,z_i)=\Phi(\alpha d_i+w_i'\eta). \end{align} One can compute the RMLE of $\delta=(\alpha,\eta')'$ using package software. Similarly, the selection equation is another binary choice model. One can estimate the coefficients on the instruments $Z$ similarly.
example[Panel Dynamic Discrete Choice Models]\rm A random effects probit model assumes $\alpha_i$ is independent of $X_i$ and follows $N(0,\gamma^2)$, and $\epsilon_{i1},\dots,\epsilon_{iT}$ are independent standard normal random variables. This specification yields the following conditional density function: \begin{align} q_{\beta_0,\delta}(y_i|x_i)=\int \prod_{t=1}^T\Phi\big[(2y_{it}-1)(x_{it}'\eta+\gamma a)\big]\phi(a)da. \end{align} One can construct a simulated likelihood function based on (ref) to obtain a restricted MLE of $\delta=(\eta',\gamma)'$ train_2009.

Asymptotic Properties

This section collects results on the asymptotic properties of the score test. The proofs of all theoretical results are in Appendix (ref). Throughout, we assume that $U^n=(U_1,\dots,U_n)$ is an independent and identically distributed (i.i.d.) sample drawn from $F_\theta$, and $X^n=(X_1,\dots,X_n)$ is also an i.i.d. sample following $q_X^n$. The joint distribution of the outcome sequence $Y^n=(Y_1,\dots,Y_n)\in\mathcal Y^n$ conditional on $X^n=x^n$ is not uniquely determined due to the potential incompleteness of the model. For $\theta\in\Theta$, the distribution belongs to the following set:

multline[multline omitted — 250 chars of source]

where $F^n_\theta$ denotes the joint law of $U^n$, and $G^n(u^n|x^n;\theta)=\prod_{i=1}^n G(u_i|x_i;\theta)$ is the Cartesian product of the set-valued predictions. We let $\mathcal P_{\theta}^n$ collect joint laws of $(Y^n,X^n)$; each element $P^n$ of $\mathcal P_{\theta}^n$ is such that the conditional law of $Y^n$ given $X^n$ belongs to $\mathcal Q_{\theta}^n$, and the law of $X^n$ is $q^n_X$. Assuming $U^n$ and $X^n$ are i.i.d. does not imply $Y^n$ is i.i.d. The set $\mathcal Q^n_\theta$, in general, contains dependent and heterogeneous laws because the behavior of the selection mechanism across experiments is unrestricted eks. This feature does not create an issue for the size properties of our test because $\mathcal Q_\theta^n$ reduces to a single i.i.d. law under the null hypothesis.

For the asymptotic properties of the RMLE, we also allow $\beta$ to be in a local neighborhood of $\beta_0$. For such settings, we provide conditions under which $\hat\delta_n$ is $\sqrt n$-consistent. Let $h\in \mathcal V_1$ and $\delta_0\in\Theta_\delta$. We assume data are generated from $P^n \in \mathcal Q^n_{\beta_0+h/\sqrt n,\delta_0}$. The null hypothesis corresponds to the setting with $h=0$.

Fixing $\beta=\beta_0$, one can view $q_{\beta_0,\delta}$ as the conditional density of $Y$ in a regular parametric model, in which $\delta$ is the only unknown parameter. For each $\delta\in\Theta_\delta$, let $\mathbb M(\delta)\equiv E[\ln q_{\beta_0,\delta}(Y_i|X_i)]$, where expectation is taken with respect to the conditional density $q_{\beta_0,\delta_0}$ and the distribution of $X.$ Let $\mathbb{M}_n(\delta)$ be the sample counterpart of $\mathbb M$ defined in (ref).

assumption(i-a) There is a continuous function $M:\Theta_\delta\to\mathbb R_+$ such that $$\sup_{(y,x)\in\mathcal Y\times\mathcal X}|\ln q_{\beta_0,\delta}(y|x)|\le M(\delta),~\text{ for all }\delta\in\Theta_\delta;$$ (i-b) The map $\delta\mapsto \lnq_{\beta_0,\delta}(y|x)$ is Lipschitz continuous uniformly in $(y,x)$. That is, \begin{align} \sup_{(y,x)\in\mathcal Y\times\mathcal X}\big|\ln q_{\beta_0,\delta}(y|x)-\ln q_{\beta_0,\delta'}(y|x)\big|\lesssim \|\delta-\delta'\| \forall \delta,\delta'\in\Theta_\delta. \end{align} (i-c) $\delta\ne\delta_0\Rightarrow q_{\beta_0,\delta}(y|x)\ne q_{\beta_0,\delta_0}(y|x)$ with positive probability; (ii) $\Theta_\delta$ is a nonempty compact subset of Euclidean space; (iii) The restricted MLE $\hat\delta_n$ satisfies $\mathbb{M}_n(\hat\delta_n)\ge \inf_{\delta\in\Theta_\delta}\mathbb{M}_n(\delta)+r_n$, for some sequence $\{r_n\}$. For any $\epsilon>0$ and $\theta$ in a neighborhood of $\theta_0$, $\sup_{P^n\in\mathcal P^n_{\theta}}P^n(|r_n|>\epsilon)\to 0$.

Assumption (ref) imposes sufficient conditions for identification and uniform law of large numbers standard in the literature. We also assume the density of $F_\theta$ depends on $\theta$ smoothly. For this, for any integrable function $f$ defined on a measure space $(A,\mathfrak F,\zeta)$, let $\|f\|_{L^1_\zeta}$ be the $L^1$-norm of $f$.

assumptionFor each $\theta\in\Theta$, $F_\theta$ is absolutely continuous with respect to a $\sigma$-finite measure $\zeta$ on $U$. The Radon-Nikodym density $f_\theta=dF_\theta/d\zeta$ satisfies \begin{align} \|f_\theta-f_{\theta'}\|_{L^1_\zeta}\le C\|\theta-\theta'\|, \forall \theta,\theta'\in \Theta, \end{align} for some $C>0.$

Finally, the following condition ensures the population objective function is locally well behaved so that its value is informative about $\delta_0$.

assumption(i) $\mathbb M$ is twice continuously differentiable at $\delta_0$; (ii) The Hessian matrix \begin{align} H(\delta_0)=\frac{\partial^2}{\partial\delta\partial\delta'}\mathbb M(\delta)\Big|_{\delta=\delta_0} \end{align} is negative definite.

Under these assumptions, the restricted MLE $\hat\delta_n$ is $\sqrt n$-consistent.

propositionSuppose Assumptions (ref)-(ref) hold. Then, \begin{align} \sqrt n\|\hat\delta_n-\delta_0\|=O_{P^n}(1), \end{align} uniformly in $P^n\in \mathcal P^n_{\theta_0+h/\sqrt n}.$

Below, let $P^n_0\in \mathcal P^n_{\theta_0}$ be the joint law of $(Y^n,X^n)$ under the null hypothesis. Let $s_{\theta,j}$ be the $j$-th component of $s_\theta$. Let

align[align omitted — 121 chars of source]

and let $\Xi=\big\{f:\mathcal Y\times\mathcal X\to\mathbb R|f(y,x)=\xi_{j,k}(y,x;\delta),~1\le j,k\le d,\delta\in\Theta_{\delta}\big\}.$ Next, we add a condition for the asymptotic distribution of $\hat S_n$ and consistent estimation of the asymptotic variance $V_0$.

assumption(i) $\frac{\partial}{\partial \delta}E[s_\beta(Y|X;\beta_0,\delta)]$ exists on a neighborhood of $\delta_0$; (ii) $\sup_{f\in\Xi}\big|\frac{1}{n}\sum_{i=1}^n f(Y_i,X_i)-E_{P_0}[f(Y_i,X_i)]\big|=o_{P_0^n}(1),$ and $\delta\mapsto E[\xi_{j,k}(Y,X;\delta)]$ is continuous for any $1\le j,k\le d$. (iii) $V_0$ is nonsingular.

Here we assume the expected score can be linearized and the elements of $\Xi$ obey a uniform law of large numbers. Suppose $\hat S_n$ as defined in (ref). The following theorem shows that the test controls its asymptotic size.

theoremSuppose Assumptions (ref)-(ref) hold. Let $c_\alpha$ be defined as in (ref). Then, for any $\alpha\in (0,1)$, \begin{align} \lim_{n\to\infty}P^n_0(\hat S_n>c_\alpha)=\alpha. \end{align}

Inference on Parameters

In some applications, the ultimate goal may be to make inference on the underlying parameter, for example, to construct confidence intervals for components of $\theta$. While we defer a formal analysis to future work, we suggest a hybrid procedure that aims at controlling the potential distortion of the model selection step, borrowing insights from the moment selection literature andrews2010inference,romano2014practical.

Consider constructing confidence intervals for a component or linear combination $\gamma_0=p'\delta_0$ of $\delta_0$.\footnote{Since the null hypothesis pins $\beta$'s value down, it is natural to consider inference on the parameters that are estimated under both null and alternative hypotheses.} Due to Proposition (ref), $\hat\gamma_n=p'\hat\delta_n$ is a $\sqrt n$-consistent estimator of $\gamma_0$ as long as the true value of $\beta$ is in a neighborhood of $\beta_0$ whose radius is of order $n^{-1/2}$. It would be natural to use such an estimator to construct a confidence interval for $\delta_0$ if the complete model is selected. A well-known challenge for such post-model selection inference is that a naive asymptotic approximation that disregards the model selection step may not be valid uniformly over a large class of data generating processes leeb2005model,andrews2009hybrid. Given this, we consider the following hybrid method.

Step 1: Compute $\hat S_n$ and $c_{n}=(\kappa_n\wedge 1)c_\alpha$, where $\kappa_n$ is a sequence of shrinkage factors that tends to 0 slowly, e.g. $\kappa_n=(\ln n)^{-1/2}$;

Step 2:

itemize• Reject $H_0:\beta=\beta_0$ if $\hat S_n>c_n$. Construct a robust confidence interval for $\gamma_0$ using methods such as BCS,KMS; • Do not reject $H_0:\beta=\beta_0$ if $\hat S_n\le c_n$. Construct the Wald confidence interval $[\hat\gamma_n-z_{\alpha/2} SE(\hat\gamma_n),~\hat\gamma_n+z_{\alpha/2} SE(\hat\gamma_n)]$, where $SE(\cdot)$ is the (estimated) standard error of its argument, and $z_{\alpha}$ is the $1-\alpha$ quantile of the standard normal distribution.

The heuristic behind this procedure is as follows. First, we compare $\hat S_n$ to a critical value $c_n$ that tends to 0 slowly. For DGPs whose $\beta$ is outside local neighborhoods of $\beta_0$, we cannot ensure the asymptotic validity of the Wald confidence interval. In such settings, the procedure above uses a robust confidence interval asymptotically, which controls the asymptotic coverage probability. Since the critical value tends to 0, we use the Wald confidence interval only if $\beta$ is in a local neighborhood of $\beta_0$. The shrinkage factor $\kappa_n$, therefore, introduces a conservative distortion, which is expected to make the resulting confidence interval's coverage probability above its nominal over a wide range of $\beta$ values.

Empirical Illustrations

We illustrate the score test through two empirical applications.

Testing Strategic Interaction Effects

The first application revisits the analysis of the airline industry by klinetamer2016bayesian. We test the presence of strategic interaction effects between two types of firms: low-cost carriers (LCC) and other airlines (OA). Below, we briefly summarize the setup and refer to klinetamer2016bayesian for details. A market is defined as trips between airports regardless of intermediate stops. The two types of firms, LCC and OA, decide whether or not to serve each market. The binary variable $y^{(\ell)}_i$ takes value 1 if airline $\ell\in\{\text{LCC, OA}\}$ serves market $i$. Airline $\ell$'s payoff in market $i$ equals $$y_{i}^{(\ell)}(\delta^{cons}_{\ell}+\delta^{size}_{\ell}X_{i, size}+\delta^{pres}_{\ell}X^{(\ell)}_{i, pres}+\beta_{\ell}y^{(-\ell)}_{i}+u^{(\ell)}_{i}),$$ where $\beta_{\ell}$ captures the impact of the competitor's entry decision, $y^{(-\ell)}_{i}$. The airline-specific intercepts and observable covariates determine each firm's payoff. The covariates include the market size $X_{i, size}$ and the market presence $X^{(\ell)}_{i,pres}$. The market size $X_{i, size}$ is defined as the population at the endpoints of each trip. The latter variable $X^{(\ell)}_{i,pres}$ measures the presence of firm $\ell$ in market $i$ (see klinetamer2016bayesian p.356 for its definition). This airline-and-market-specific variable shows up only in firm $\ell$'s payoff. The data come from the second quarter of the 2010 Airline Origin and Destination Survey (DB1B) and contain 7882 markets.\footnote{The data are available on Brendan Kline's \href{www.brendankline.com}{website}.}

Our hypothesis of interest is whether the LCCs and OAs compete strategically, which can be formulated as a one-sided test. The null hypothesis is $H_0:\beta_{LCC} = \beta_{OA} = 0$, and the alternative hypothesis is $H_1:\beta_{\ell}<0,\ell \in \{\text{LCC,OA}\}$. Finally, the vector of coefficients $\delta=(\delta^{cons}_{LCC},\delta^{size}_{LCC},\delta^{pres}_{LCC},\delta^{cons}_{OA},\delta^{size}_{OA},\delta^{pres}_{OA})$ is the nuisance parameter in this model. We estimate $\delta$ by the restricted MLE under the null hypothesis.

The value of the test statistic is 24.668. The 5% critical value is 5.050. Hence, we reject the null hypothesis at the 5% level. This result is consistent with the finding of klinetamer2016bayesian whose credible sets for the strategic interaction effects $\beta_\ell,\ell\in\{\text{LCC,OA}\}$ do not contain the origin. Table (ref) reports the RMLE of the index coefficients. The estimates suggest that the effect of market presence is larger for LCCs than other airlines, and the monopoly profits (captured by the constant terms) in a market with below-median size and below-median market presence are smaller for the LCCs. These observations are also consistent with klinetamer2016bayesian's findings, although we note that these estimates are obtained by imposing the restriction rejected by the score test.

table[table omitted — 783 chars of source]

Testing the Endogeneity of Catholic School Attendance

The second application concerns the causal effect of Catholic school attendance on academic achievements studied by AltonjiElderTaber2005. Whether Catholic schools provide a better education than public ones is important for education policies, but the analysis is complicated by the concern that selection into Catholic schools is nonrandom. Using the framework in Example (ref), we examine the endogeneity of Catholic school attendance by testing if the coefficient on the control function is zero.

The data source is a subset of the National Educational Longitudinal Survey of 1988 (NELS:88). We use a version of the data available from WooldridgeText. We refer to AltonjiElderTaber2005 for a detailed discussion of the data. The dependent variable $y_{i}$ is a binary variable indicating whether the student graduated from high school by the year 1994. The binary treatment $d_{i}$ indicates whether the student attended a Catholic high school. The vector of exogenous control variables $w_{i}$ includes each parent's years of education and log family income. The instrument variable is a dummy variable indicating whether a parent was reported to be Catholic. The sample size $n$ is 5970 after we remove missing observations on $y_{i}$.

Table (ref) reports the point estimates of nuisance parameters $\delta$ under the null hypothesis $H_0:\beta=0$. The value of the test statistic is 154.848. The 5% critical value is 2.755. We, therefore, reject the null hypothesis at the 5% level. Our test provides strong evidence supporting the students' selection into Catholic schools based on their unobservable characteristics. This result is in line with the concern expressed in AltonjiElderTaber2005.

table[table omitted — 935 chars of source]

Monte Carlo Experiments

Size and Power of the Score Test

We examine the size and power properties of the score test through simulations. The data generating process is based on Example (ref) and is motivated by the empirical illustration in the previous section. There are player-specific covariates $X_i=(X_i^{(1)},X_i^{(2)})'$, each of which is generated as an independent Rademacher random variable taking values on $\{-1,1\}$. We then generate $U_i = (U^{(1)}_i, U^{(2)}_i)$ from the bivariate standard normal distribution. For each $u_i$ and $x_i$, we determine the predicted set of outcomes $G(u_i|x_i;\theta)$ based on the payoff functions with $\delta_0=(\delta_0^{(1)},\delta_0^{(2)})=(2,1.5)'$. We then test

align[align omitted — 89 chars of source]

As discussed earlier, the model is complete under $H_0$. We estimate $\delta_0$ using the restricted MLE. The sample size is set to 2500, 5000, or 7500. This choice is motivated by the sample size used in the empirical application.

The size of the score test is reported in Table (ref). The size of the test is controlled properly across all sample sizes, while it tends to be slightly conservative when $n$ is small.

table[table omitted — 271 chars of source]

Under alternative hypotheses, multiple equilibria may be predicted. If this is the case, we select an outcome according to one of the following selection mechanisms. The first design uses a selection mechanism, which selects $(1, 0)$ out of $G(u_i|x_i;\theta)=\{(1, 0), (0, 1)\}$ if an i.i.d. Bernoulli random variable $\nu_{i}$ takes 1. In the second design, we generate data from the least favorable distribution, which draws an independent outcome sequence from the least favorable distribution $Q_{\theta_1}\in \mathcal Q_{\theta_1}$.

The power of the score test is calculated against local alternatives with $\beta^{(j)}_1=-h/\sqrt n,h>0$ for $j=1,2.$ For this exercise, we introduce a grid of values for $h$ and generate the data described above. We then compare the rejection frequency of our test to that of the moment-based testing procedure by BCS. Their test checks if a hypothesized value $(\beta^{(1)},\beta^{(2)})'=(0,0)'$ is compatible with a set of moment restrictions. Their statistic and bootstrap critical value are calculated using a sample analog of the following moment inequality and equality restrictions

align*[align* omitted — 365 chars of source]

which are the sharp identifying restrictions that characterize $\mathfrak q_\theta$ in (ref).\footnote{Since the example resembles the specification used in their Monte Carlo experiments, we added minimal changes to their replication code posted on the repository of Quantitative Economics to implement their procedure.}

Figures (ref)-(ref) show the rejection frequencies of the score and moment-based tests. The results are similar across the two designs. In each design, the score test outperforms the moment-based test in terms of power by a significant margin. This difference in performance may potentially be due to the proposed score test's exploitation of completeness under the null to estimate the nuisance parameter and direct power against $\mathcal V_1$. The moment-based test is designed for general subvector inference and does not necessarily exploit the model completeness.\footnote{The procedure by BCS tests if the null parameter value is consistent with the model restrictions and deals with nuisance parameters by a profiling method combined with regularization to ensure its uniform validity.} The simulation results suggest that taking advantage of the model structure may provide considerable benefits in terms of power.

Concluding Remarks

Economic models exhibit incompleteness for various reasons. They are, for example, consequences of strategic interaction, state dependence, or self-selection. This paper shows that one can test these important features using a score statistic even if the model is incomplete under the alternative hypothesis. The proposed test exploits the model completeness under the null hypothesis to simplify its computation. An avenue for future research includes a theory for the uniform validity of inference for post-model selection procedures based on the score test.

thebibliography{71} \expandafter\ifx\csname natexlab\endcsname\relax\def\natexlab#1{#1}\fi \bibitem[\citeauthoryear{Aliprantis and Border}{Aliprantis and Border}{2006}]{aliprantisborder} Aliprantis, C. D. and K. C. Border (2006): Infinite Dimensional Analysis: A Hitchhiker's Guide, Springer. \bibitem[\citeauthoryear{Altonji, Elder, and Taber}{Altonji et al.}{2005}]{AltonjiElderTaber2005} Altonji, J. G., T. E. Elder, and C. R. Taber (2005): “An Evaluation of Instrumental Variable Strategies for Estimating the Effects of Catholic Schooling,” Journal of Human Resources, 40, 791--821. \bibitem[\citeauthoryear{Andrews}{Andrews}{1999}]{Andrews:1999aa} Andrews, D. W. K. (1999): “Estimation When a Parameter is on a Boundary,” Econometrica, 67, 1341--1383. \bibitem[\citeauthoryear{Andrews and Barwick}{Andrews and Barwick}{2012}]{andrews_barwick12} \textsc{Andrews, D. W. K. and P. J. Barwick} (2012): “Inference for Parameters Defined by Moment Inequalities: A Recommended Moment Selection Procedure,” \emph{Econometrica}, 80, 2805--2826. \bibitem[\citeauthoryear{Andrews and Guggenberger}{Andrews and Guggenberger}{2009}]{andrews2009hybrid} \textsc{Andrews, D. W. K. and P. Guggenberger} (2009): “Hybrid and size-corrected subsampling methods,” \emph{Econometrica}, 77, 721--762. \bibitem[\citeauthoryear{Andrews and Soares}{Andrews and Soares}{2010}]{andrews2010inference} \textsc{Andrews, D. W. K. and G. Soares} (2010): “Inference for Parameters Defined by Moment Inequalities Using Generalized Moment Selection,” \emph{Econometrica}, 78, 119--157. \bibitem[\citeauthoryear{Andrews, Roth, and Pakes}{Andrews et al.}{2019}]{ARP} \textsc{Andrews, I., J. Roth, and A. Pakes} (2019): “Inference for Linear Conditional Moment Inequalities,” Discussion Paper, Harvard University. \bibitem[\citeauthoryear{Barseghyan, Coughlin, Molinari, and Teitelbaum}{Barseghyan et al.}{2021}]{barseghyan2021heterogeneous} \textsc{Barseghyan, L., M. Coughlin, F. Molinari, and J. C. Teitelbaum} (2021): “Heterogeneous Choice Sets and Preferences,” \emph{Econometrica}, 89, 2015--2048. \bibitem[\citeauthoryear{Bera and Bilias}{Bera and Bilias}{2001}]{Bera:2001va} \textsc{Bera, A. K. and Y. Bilias} (2001): “Rao's Score, Neyman's $C(\alpha)$ and Silvey's LM tests: An Essay on Historical Developments and Some New Results,” \emph{Journal of Statistical Planning and Inference}, 97, 9--44. \bibitem[\citeauthoryear{Beresteanu, Molchanov, and Molinari}{Beresteanu et al.}{2011}]{bmm} \textsc{Beresteanu, A., I. Molchanov, and F. Molinari} (2011): “Sharp Identification Regions in Models with Convex Moment Predictions,” \emph{Econometrica}, 79, 1785--1821. \bibitem[\citeauthoryear{Blundell and Powell}{Blundell and Powell}{2004}]{Blundell:2004td} \textsc{Blundell, R. W. and J. L. Powell} (2004): “Endogeneity in Semiparametric Binary Response Models,” \emph{The Review of Economic Studies}, 71, 655--679. \bibitem[\citeauthoryear{Bresnahan and Reiss}{Bresnahan and Reiss}{1991}]{BRESNAHANJoE91} \textsc{Bresnahan, T. F. and P. C. Reiss} (1991): “Empirical models of discrete games,” \emph{Journal of Econometrics}, 48, 57--81. \bibitem[\citeauthoryear{Bugni, Canay, and Shi}{Bugni et al.}{2017}]{BCS} \textsc{Bugni, F., I. Canay, and X. Shi} (2017): “Inference for Subvectors and Other Functions of Partially Identified Parameters in Moment Inequality Models,” \emph{Quantitative Economics}, 8, 1--38. \bibitem[\citeauthoryear{Canay and Shaikh}{Canay and Shaikh}{2017}]{canay/shaikh:2017} \textsc{Canay, I. A. and A. M. Shaikh} (2017): “Practical and Theoretical Advances in Inference for Partially Identified Models,” in \emph{Advances in Economics and Econometrics: Eleventh World Congress}, ed. by B. Honor{\'e}, A. Pakes, M. Piazzesi, and L. Samuelson, Cambridge University Press, vol. 2 of \emph{Econometric Society Monographs}, 271--306. \bibitem[\citeauthoryear{Card and Hyslop}{Card and Hyslop}{2005}]{card2005estimating} \textsc{Card, D. and D. R. Hyslop} (2005): “Estimating the effects of a time-limited earnings subsidy for welfare-leavers,” \emph{Econometrica}, 73, 1723--1770. \bibitem[\citeauthoryear{Chamberlain}{Chamberlain}{1985}]{chamberlain_1985} \textsc{Chamberlain, G.} (1985): \emph{Heterogeneity, omitted variable bias, and duration dependence}, Cambridge University Press, 3–38, Econometric Society Monographs. \bibitem[\citeauthoryear{Chen, Christensen, and Tamer}{Chen et al.}{2018}]{Chen2018} \textsc{Chen, X., T. M. Christensen, and E. Tamer} (2018): “Monte Carlo confidence sets for identified sets,” \emph{Econometrica}, 86, 1965--2018. \bibitem[\citeauthoryear{Chesher}{Chesher}{2003}]{chesher2003} \textsc{Chesher, A.} (2003): “Identification in Nonseparable Models,” \emph{Econometrica}, 71, 1405--1441. \bibitem[\citeauthoryear{Chesher and Rosen}{Chesher and Rosen}{2017}]{chesher2017generalized} \textsc{Chesher, A. and A. M. Rosen} (2017): “Generalized Instrumental Variable Models,” \emph{Econometrica}, 85, 959--989. \bibitem[\citeauthoryear{Ciliberto and Tamer}{Ciliberto and Tamer}{2009}]{CilibertoTamer2009} \textsc{Ciliberto, F. and E. Tamer} (2009): “Market Structure and Multiple Equilibria in Airline Markets,” \emph{Econometrica}, 77, 1791--1828. \bibitem[\citeauthoryear{Cox and Shi}{Cox and Shi}{2020}]{Cox2020} \textsc{Cox, G. and X. Shi} (2020): “Simple Adaptive Size-Exact Testing for Full-Vector and Subvector Inference in Moment Inequality Models,” Working Paper. \bibitem[\citeauthoryear{de Paula, Richards-Shubik, and Tamer}{de Paula et al.}{2018}]{depaulaEtAl2018} \textsc{de Paula, {\'A}., S. Richards-Shubik, and E. Tamer} (2018): “Identifying Preferences in Networks With Bounded Degree,” \emph{Econometrica}, 86, 263--288. \bibitem[\citeauthoryear{de Paula and Tang}{de Paula and Tang}{2012}]{depaulatang2012} \textsc{de Paula, {\'A}. and X. Tang} (2012): “Inference of Signs of Interaction Effects in Simultaneous Games With Incomplete Information,” \emph{Econometrica}, 80, 143--172. \bibitem[\citeauthoryear{Dempster}{Dempster}{1967}]{dempster67} \textsc{Dempster, A.} (1967): “{Upper and Lower Probabilities Induced by a Multivalued Mapping},” \emph{The Annals of Mathematical Statistics}, 38, 325--339. \bibitem[\citeauthoryear{Eizenberg}{Eizenberg}{2014}]{eizenberg2014upstream} \textsc{Eizenberg, A.} (2014): “Upstream Innovation and Product Variety in the U.S. Home PC Market,” \emph{The Review of Economic Studies}, 81, 1003--1045. \bibitem[\citeauthoryear{Epstein, Kaido, and Seo}{Epstein et al.}{2016}]{eks} \textsc{Epstein, L., H. Kaido, and K. Seo} (2016): “Robust Confidence Regions for Incomplete Models,” \emph{Econometrica}, 84, 1799--1838. \bibitem[\citeauthoryear{Fack, Grenet, and He}{Fack et al.}{2019}]{fack2019beyond} \textsc{Fack, G., J. Grenet, and Y. He} (2019): “Beyond Truth-Telling: Preference Estimation with Centralized School Choice and College Admissions,” \emph{American Economic Review}, 109, 1486--1529. \bibitem[\citeauthoryear{Galichon and Henry}{Galichon and Henry}{2011}]{galichon2011set} \textsc{Galichon, A. and M. Henry} (2011): “Set identification in models with multiple equilibria,” \emph{The Review of Economic Studies}, 78, 1264--1298. \bibitem[\citeauthoryear{Gilboa and Schmeidler}{Gilboa and Schmeidler}{1989}]{GILBOA1989141} \textsc{Gilboa, I. and D. Schmeidler} (1989): “Maxmin Expected Utility with Non-Unique Prior,” \emph{Journal of Mathematical Economics}, 18, 141--153. \bibitem[\citeauthoryear{Haile and Tamer}{Haile and Tamer}{2003}]{haile2003inference} \textsc{Haile, P. A. and E. Tamer} (2003): “Inference with an Incomplete Model of English Auctions,” \emph{Journal of Political Economy}, 111, 1--51. \bibitem[\citeauthoryear{Handel}{Handel}{2013}]{handel2013adverse} \textsc{Handel, B. R.} (2013): “Adverse selection and inertia in health insurance markets: When nudging hurts,” \emph{American Economic Review}, 103, 2643--82. \bibitem[\citeauthoryear{Heckman}{Heckman}{1978}]{heckman78} \textsc{Heckman, J. J.} (1978): “Simple Statistical Models for Discrete Panel Data Developed and Applied to Test the Hypothesis of True State Dependence Against The Hypothesis of Spurious State Dependence,” \emph{Annales de INSEE}, 227--269. \bibitem[\citeauthoryear{Heckman}{Heckman}{1981}]{heckman1987incidental} --------- (1981): “The Incidental Parameters Problem and the Problem of Initial Conditions in Estimating a Discrete Time-Discrete Data Stochastic Process,” in \emph{Structural Analysis of Discrete Data With Econometric Applications}, ed. by C. Manski and D. McFadden, MIT Press. \bibitem[\citeauthoryear{Henry, Meango, and Mourifi{\'e}}{Henry et al.}{2020}]{henry2020revealing} \textsc{Henry, M., R. Meango, and I. Mourifi{\'e}} (2020): “Revealing Gender-Specific Costs of STEM in an Extended Roy Model of Major Choice,” Working Paper. \bibitem[\citeauthoryear{Honor{\'e} and Tamer}{Honor{\'e} and Tamer}{2006}]{honore2006bounds} \textsc{Honor{\'e}, B. E. and E. Tamer} (2006): “Bounds on Parameters in Panel Dynamic Discrete Choice Models,” \emph{Econometrica}, 74, 611--629. \bibitem[\citeauthoryear{Huber and Strassen}{Huber and Strassen}{1973}]{HS} \textsc{Huber, P. and V. Strassen} (1973): “Minimax Tests and Neyman--Pearson Lemma for Capacities,” \emph{The Annals of Statistics}, 1, 251--263. \bibitem[\citeauthoryear{Hyslop}{Hyslop}{1999}]{hyslop99} \textsc{Hyslop, D. R.} (1999): “State Dependence, Serial Correlation and Heterogeneity in Intertemporal Labor Force Participation of Married Women,” \emph{Econometrica}, 67, 1255--1294. \bibitem[\citeauthoryear{Imbens and Newey}{Imbens and Newey}{2009}]{imbens2009identification} \textsc{Imbens, G. W. and W. K. Newey} (2009): “Identification and Estimation of Triangular Simultaneous Equations Models Without Additivity,” \emph{Econometrica}, 77, 1481--1512. \bibitem[\citeauthoryear{Jovanovic}{Jovanovic}{1989}]{jovanovich89} \textsc{Jovanovic, B.} (1989): “Observable Implications of Models with Multiple Equilibria,” \emph{Econometrica}, 57, 1431--1437. \bibitem[\citeauthoryear{Kaido, Molinari, and Stoye}{Kaido et al.}{2019}]{KMS} \textsc{Kaido, H., F. Molinari, and J. Stoye} (2019): “Confidence Intervals for Projections of Partially Identified Parameters,” \emph{Econometrica}, 87, 1397--1432. \bibitem[\citeauthoryear{Kaido and Zhang}{Kaido and Zhang}{2019}]{kz} \textsc{Kaido, H. and Y. Zhang} (2019): “Robust Likelihood-Ratio Tests for Incomplete Economic Models,” Working Paper. \bibitem[\citeauthoryear{Kawai and Watanabe}{Kawai and Watanabe}{2013}]{kawai2013inferring} \textsc{Kawai, K. and Y. Watanabe} (2013): “Inferring Strategic Voting,” \emph{American Economic Review}, 103, 624--62. \bibitem[\citeauthoryear{Kline and Tamer}{Kline and Tamer}{2016}]{klinetamer2016bayesian} \textsc{Kline, B. and E. Tamer} (2016): “Bayesian inference in a class of partially identified models,” \emph{Quantitative Economics}, 7, 329--366. \bibitem[\citeauthoryear{Leeb and P{\"o}tscher}{Leeb and P{\"o}tscher}{2005}]{leeb2005model} \textsc{Leeb, H. and B. M. P{\"o}tscher} (2005): “Model selection and inference: Facts and fiction,” \emph{Econometric Theory}, 21, 21--59. \bibitem[\citeauthoryear{Lehmann and Romano}{Lehmann and Romano}{2005}]{lehmann2005testing} \textsc{Lehmann, E. L. and J. P. Romano} (2005): \emph{Testing statistical hypotheses}, vol. 3, Springer. \bibitem[\citeauthoryear{Lewbel}{Lewbel}{2007}]{lewbel2007ier} \textsc{Lewbel, A.} (2007): “Coherency and Completeness of Structural Models Containing a Dummy Endogenous Variable,” \emph{International Economic Review}, 48, 1379--1392. \bibitem[\citeauthoryear{McFadden}{McFadden}{1981}]{mcfadden1981econometric} \textsc{McFadden, D.} (1981): “Econometric Models of Probabilistic Choice,” in \emph{Structural Analysis of Discrete Data With Econometric Applications}, ed. by C. Manski and D. McFadden, MIT Press. \bibitem[\citeauthoryear{Molchanov}{Molchanov}{2005}]{molchanov2005theory} \textsc{Molchanov, I. S.} (2005): \emph{Theory of random sets}, vol. 19, Springer. \bibitem[\citeauthoryear{Molinari}{Molinari}{2020}]{molinari2020microeconometrics} \textsc{Molinari, F.} (2020): “Microeconometrics with Partial Identification,” in \emph{Handbook of Econometrics}, ed. by S. N. Durlauf, L. P. Hansen, J. J. Heckman, and R. L. Matzkin, Elsevier, vol. 7, 355--486. \bibitem[\citeauthoryear{Newey}{Newey}{1994}]{newey1994asymptotic} \textsc{Newey, W. K.} (1994): “The asymptotic variance of semiparametric estimators,” \emph{Econometrica: Journal of the Econometric Society}, 1349--1382. \bibitem[\citeauthoryear{Newey and McFadden}{Newey and McFadden}{1994}]{newey1994large} \textsc{Newey, W. K. and D. McFadden} (1994): “Large Sample Estimation and Hypothesis Testing,” in \emph{Handbook of Econometrics}, ed. by J. J. Heckman and E. Leamer, Elsevier, vol. 4, chap. 36, 2111--2245. \bibitem[\citeauthoryear{Otsu, Pesendorfer, and Takahashi}{Otsu et al.}{2016}]{OtsuPT2016} \textsc{Otsu, T., M. Pesendorfer, and Y. Takahashi} (2016): “Pooling Data Across Markets in Dynamic Markov Games,” \emph{Quantitative Economics}, 7, 523--559. \bibitem[\citeauthoryear{Pelican and Graham}{Pelican and Graham}{2021}]{pelican2020optimal} \textsc{Pelican, A. and B. S. Graham} (2021): “An Optimal Test for Strategic Interaction in Social and Economic Network Formation Between Heterogeneous Agents,” Working Paper. \bibitem[\citeauthoryear{Philippe, Debs, and Jaffray}{Philippe et al.}{1999}]{philippe1999decision} \textsc{Philippe, F., G. Debs, and J.-Y. Jaffray} (1999): “Decision Making with Monotone Lower Probabilities of Infinite Order,” \emph{Mathematics of Operations Research}, 24, 767--784. \bibitem[\citeauthoryear{Romano, Shaikh, and Wolf}{Romano et al.}{2014}]{romano2014practical} \textsc{Romano, J. P., A. M. Shaikh, and M. Wolf} (2014): “A Practical Two-Step Method for Testing Moment Inequalities,” \emph{Econometrica}, 82, 1979--2002. \bibitem[\citeauthoryear{Shafer}{Shafer}{1976}]{shafer1976mathematical} \textsc{Shafer, G.} (1976): \emph{A Mathematical Theory of Evidence}, Princeton University Press. \bibitem[\citeauthoryear{Shaikh and Vytlacil}{Shaikh and Vytlacil}{2011}]{shaikh_vytlacil2011} \textsc{Shaikh, A. M. and E. J. Vytlacil} (2011): “Partial Identification in Triangular Systems of Equations With Binary Dependent Variables,” \emph{Econometrica}, 79, 949--955. \bibitem[\citeauthoryear{Sheng}{Sheng}{2020}]{sheng2020} \textsc{Sheng, S.} (2020): “A Structural Econometric Analysis of Network Formation Games Through Subnetworks,” \emph{Econometrica}, 88, 1829--1858. \bibitem[\citeauthoryear{Silvapulle and Silvapulle}{Silvapulle and Silvapulle}{1995}]{Silvapulle:1995tm} \textsc{Silvapulle, M. J. and P. Silvapulle} (1995): “A Score Test Against One-Sided Alternatives,” \emph{Journal of the American Statistical Association}, 90, 342--349. \bibitem[\citeauthoryear{Talagrand}{Talagrand}{1994}]{Talagrand:1994aa} \textsc{Talagrand, M.} (1994): “Sharper Bounds for Gaussian and Empirical Processes,” \emph{The Annals of Probability}, 22, 28--76. \bibitem[\citeauthoryear{Tamer}{Tamer}{2003}]{tamer} \textsc{Tamer, E.} (2003): “Incomplete Simultaneous Discrete Response Model with Multiple Equilibria,” \emph{The Review of Economic Studies}, 70, 147--165. \bibitem[\citeauthoryear{Train}{Train}{2009}]{train_2009} \textsc{Train, K.} (2009): \emph{Discrete Choice Methods with Simulation}, Cambridge University Press, 2 ed. \bibitem[\citeauthoryear{van der Vaart}{van der Vaart}{2000}]{Van-der-Vaart:2000aa} \textsc{van der Vaart, A.} (2000): \emph{Asymptotic Statistics}, Cambridge University Press. \bibitem[\citeauthoryear{van der Vaart and Wellner}{van der Vaart and Wellner}{1996}]{VanderVaarta} \textsc{van der Vaart, A. and J. A. Wellner} (1996): \emph{Weak Convergence and Empirical Processes: With Applications to Statistics.}, Springer, New York. \bibitem[\citeauthoryear{Wald}{Wald}{1950}]{Wald1950} \textsc{Wald, A.} (1950): “Remarks on the estimation of unknown parameters in incomplete systems of equations,” \emph{Statistical Inference in Dynamic Economic Models, ed. TC Koopmans. John Wiley, New York}, 305--310. \bibitem[\citeauthoryear{Wasserman}{Wasserman}{1990}]{wasserman1990belief} \textsc{Wasserman, L. A.} (1990): “Belief Functions and Statistical Inference,” \emph{Canadian Journal of Statistics}, 18, 183--196. \bibitem[\citeauthoryear{Wollmann}{Wollmann}{2018}]{wollmann2018trucks} \textsc{Wollmann, T. G.} (2018): “Trucks Without Bailouts: Equilibrium Product Characteristics for Commercial Vehicles,” \emph{American Economic Review}, 108, 1364--1406. \bibitem[\citeauthoryear{Wooldridge}{Wooldridge}{2005}]{wooldridge2005simple} \textsc{Wooldridge, J. M.} (2005): “Simple Solutions to the Initial Conditions Problem in Dynamic, Nonlinear Panel Data Models with Unobserved Heterogeneity,” \emph{Journal of Applied Econometrics}, 20, 39--54. \bibitem[\citeauthoryear{Wooldridge}{Wooldridge}{2014}]{wooldridge2014JoE} --------- (2014): “{Quasi-Maximum Likelihood Estimation and Testing for Nonlinear Models with Endogenous Explanatory Variables},” \emph{Journal of Econometrics}, 182, 226--234. \bibitem[\citeauthoryear{Wooldridge}{Wooldridge}{2015}]{wooldridgeJHR} --------- (2015): “Control Function Methods in Applied Econometrics,” \emph{Journal of Human Resources}, 50, 420--445. \bibitem[\citeauthoryear{Wooldridge}{Wooldridge}{2019}]{WooldridgeText} --------- (2019): \emph{Introductory Econometrics: A Modern Approach}, Cengage Learning, 7 ed.