Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
101,700 characters · 17 sections · 69 citation commands
Robust Tests of Model Incompleteness in the Presence of Nuisance Parameters
Discrete choice models are used widely. A common empirical strategy is to combine a theory of choice (e.g., utility maximization) that predicts a unique outcome value with distributional assumptions on latent variables mcfadden1981econometric. This approach allows the researcher to derive the conditional distribution of the outcome given covariates and apply likelihood-based inference methods. However, recent economic applications often involve models that permit multiple outcome values, which we call an incomplete prediction. Such an incomplete prediction occurs when the researcher is willing to work only with weak assumptions or has limited knowledge of the data-generating process.
This paper considers a form of incompleteness summarized as follows. An observable discrete outcome $Y\in \mathcal Y$ satisfies
where $G$ collects all outcome values compatible with the model given observed and unobserved variables $(X,U)$ and a structural parameter $\theta$. This structure arises in a variety of contexts. For example, multiple outcomes are predicted in single-agent discrete choice models when the agent's unobservable choice set is consistent with a wide range of choice set formation processes barseghyan2021heterogeneous. Multiple equilibria may exist in discrete games such as firms' market entry or household labor supply decisions, but one may not know how an equilibrium outcome gets selected BRESNAHANJoE91,CilibertoTamer2009. Recent empirical studies have applied inference methods for incomplete models in different areas; they include English auctions haile2003inference, strategic voting kawai2013inferring, product offerings eizenberg2014upstream,wollmann2018trucks, network formation depaulaEtAl2018,sheng2020, school choice fack2019beyond, and major choice henry2020revealing.
This paper focuses on developing tests to determine if a structural model is incomplete. The completeness of a model is important when considering its policy implications. For example, in a canonical market-entry model, multiple equilibria exist only if firms strategically interact. Testing for strategic interaction effects and inferring their signs can provide valuable information for policymakers depaulatang2012. In a triangular model with a discrete endogenous variable, the control function approach yields an incomplete model, but it remains complete if the assignments are exogenous. Detecting self-selection can aid practitioners in selecting a proper strategy for evaluating treatment effects. Dynamic discrete choice models can make incomplete predictions when they admit state dependence, but they are complete without it. The presence of state dependence can significantly impact program evaluation and counterfactual analyses card2005estimating,handel2013adverse.
In many of these examples, one can state the null hypothesis of model completeness as a restriction on the parameter's subvector. We develop a novel score test for such a restriction. Advantages of this approach are (i) the score statistic only requires estimation of nuisance parameters in the complete restricted model; (ii) one can use package software to estimate the nuisance parameters; and (iii) simulating the score statistic's limiting distribution is straightforward.
Our focus is on models that are complete under the null hypothesis. This structure makes score-based tests appealing. First, one can obtain a unique likelihood function $q_{\theta_0}$ for any null parameter value $\theta_0$, thanks to the model's completeness. For each alternative parameter value $\theta_0+h$, the model implies multiple (typically infinitely-many) likelihood functions. However, the recent work of kz demonstrates that one can identify a "least favorable" density $q_{\theta_0+h}$ that is most difficult to distinguish from $q_{\theta_0}$. We may then consider the family $\{q_{\theta_0+h}\}_{h\in\mathbb R^d}$ as a "least favorable parametric model". Second, we can use the least favorable parametric family to calculate a score function. We then construct a score-based statistic that maximizes a measure of local discrimination. The resulting test is robust to incompleteness because it detects any local deviation from the null hypothesis regardless of how $Y$ is selected from the predicted set.
The score test has another appealing feature wherein one can estimate nuisance parameters within the complete restricted model. Typically, the null hypothesis does not restrict some components of $\theta$. By exploiting the model completeness, we demonstrate that one can construct a restricted maximum likelihood estimator (RMLE) of these nuisance components that is $\sqrt n$-consistent. This estimator is usually simple to calculate using package software. Next, we insert the estimator into the score formula to compute a test statistic. The suggested procedure is computationally tractable since it avoids evaluating the test statistic over a grid of nuisance parameters. Finally, we derive the score-based statistic's limiting distribution and demonstrate that it can be easily simulated. Although there are existing methods to test restrictions on subvectors of $\theta$, they can be computationally expensive since they are not designed to utilize the model completeness under the null hypothesis. This paper presents a procedure that makes use of the model structure to simplify its implementation. Through a Monte Carlo experiment, we also show that the proposed test has significantly higher power than a general subvector test that does not utilize the structure.
Our paper belongs to the literature on inference in incomplete models pioneered by Wald1950 in the context of simultaneous equations models and by jovanovich89 in the context of models with multiple equilibria. The seminal work of tamer shows an incomplete model induces multiple distributions and implies partially identifying restrictions on parameters. Developments in the literature galichon2011set,bmm,chesher2017generalized provided tools to systematically derive so-called sharp identifying restrictions, which convert all model information into a set of equality and inequality restrictions on the conditional moments of observable variables. Inference methods based on the sample analogs of moment restrictions are extensively studied canay/shaikh:2017. Our approach builds on recent developments in likelihood-based inference methods for incomplete models Chen2018,kz. In particular, we combine the sharp identifying restrictions with the framework of kz to derive the least favorable parametric model and its score. To our knowledge, this approach is new. One can view our procedure as an analog of deriving a score function from a parametric family in a complete model.
Hypothesis testing in incomplete models is studied extensively. As discussed earlier, many of them are based on the sample analogs of conditional or unconditional moment restrictions. A challenge in making inferences is the high computational cost of implementing existing methods, as noted in molinari2020microeconometrics. There are attempts to improve the computational tractability of moment-based inference methods within specific classes of models or testing problems, such as those made by ARP and Cox2020, who assume that moment inequality restrictions are linear conditional on observable variables. This paper focuses on another class in which the model is complete under the null hypothesis. This structure makes our test computationally tractable by combining (i) the score function associated with the least favorable parametric model and (ii) a point estimator of the nuisance components.
Practitioners can use this paper's framework to test various hypotheses. For example, one can examine the exsitence of strategic interaction effects and multiple equilibria in static complete information games. Related problems are studied in other models. For incomplete information games, depaulatang2012 introduced a semiparametric inference procedure on the signs of strategic interaction effects. For finite-state Markov games, OtsuPT2016 developed techniques to test whether the conditional choice probabilities, state transition, and other features of games are homogeneous across cross-sectional units. Rejecting their null hypothesis may indicate the presence of multiple equilibria. pelican2020optimal studied a testing problem that involves determining whether agents' preferences are interdependent in a network formation model. The problem includes nuisance parameters that account for degree heterogeneity and homophily. Using the logit structure, they constructed a sufficient statistic for the nuisance parameters and developed a conditional score test. This paper and pelican2020optimal consider testing in settings with a complete model under the null and an incomplete model under the alternative. The two papers take different approaches by exploiting the structures of the respective models. Specifically, we use the least favorable parametric model and estimate nuisance parameters using the restricted MLE. In contrast, pelican2020optimal utilized a sufficient statistic for the nuisance parameters.
One can apply our framework to triangular systems involving a binary outcome and a discrete endogenous variable. We show that taking a control function approach in such a setting leads to a model with an incomplete prediction.\footnote{A nonparametric identification analysis based on set-valued control functions is undertaken in another paper.} We provide a test of endogenous treatment assignments under weak assumptions. To our knowledge, this test is new and provides an alternative to the existing proposal by wooldridge2014JoE, who makes additional high-level assumptions.
Let $Y$ be a discrete outcome taking values in a finite set $\mathcal Y$. Let $X\in\mathcal X\subseteq\mathbb R^{d_X}$ be a vector of observable covariates and let $U\in \mathcal U\subseteq\mathbb R^{d_U}$ be a vector of unobservable variables. We equip $\mathcal Y,\mathcal X$, and $\mathcal U$ with their Borel $\sigma$-algebra. In what follows, we use upper case letters for random elements (e.g., $X$) and lower case letters (e.g., $x$) for the values they can take. Let $\theta\in\Theta\subset\mathbb R^{d_\theta}$ be a finite-dimensional parameter, where $\Theta$ is a convex parameter space with a nonempty interior.
A set-valued map $G:\mathcal U\times\mathcal X\times\Theta\leadsto \mathcal Y$ summarizes the prediction of a structural model. We assume $G(\cdot|\cdot;\theta)$ is weakly measurable for every $\theta\in\Theta$ and $Y$ takes one of the values in $G(U|X;\theta)$ with probability 1.\footnote{A set-valued map is weakly measurable if its weak inverse image $G_{-1}(A)\equiv\{s\in\mathcal S :G(s)\cap A\ne\emptyset\}$ is measurable for any open set $A$.} A random element $Y$ with this property is said to be a measurable selection of the random closed set $G(U|X;\theta)$ molchanov2005theory. The map $G$ describes how observable and unobservable characteristics translate into a set of possible outcome values. It reflects restrictions imposed by theory, such as the functional form of utility/profit functions, forms of strategic interaction, and any equilibrium or optimality concepts. Importantly, $G(u|x,\theta)$ can contain multiple values. This feature allows the researchers to encode their lack of understanding of parts of the structural model.
The formulation above also nests the standard setting in which the model is characterized by a reduced form equation:
for a function $g:\mathcal U\times \mathcal X\times\Theta\to \mathcal Y$. In this case, $G$ is almost surely singleton-valued, i.e., $G(U|X;\theta)=\{g(U|X;\theta)\}$, and we say the model makes a complete prediction.
Throughout, we assume $U$'s law belongs to a parametric family $F=\{F_\theta,\theta\in\Theta\}$, where, for each $\theta$, $F_\theta$ is a probability distribution on $\mathcal U$. To keep notation concise, we use the same $\theta$ for parameters that enter $G$ and that index $F_\theta$. Also, we focus on settings in which $U$ is independent of $X$. However, the framework can be easily extended to settings where $U$ is correlated with $X$, and the researcher specifies its conditional distribution $F_\theta(u|x)$. Furthermore, our framework accommodates settings in which some of the observable covariates are endogenous, and one can construct a set-valued control function (See Example (ref) below).
We illustrate the objects introduced above with examples. Our first example is a discrete game of complete information BRESNAHANJoE91,CilibertoTamer2009.
The next example is a parametric version of triangular nonseparable equations with binary outcome and treatment chesher2003,shaikh_vytlacil2011. We consider a control function approach applied to the triangular system.
The next example is a panel dynamic discrete choice model heckman78,chamberlain_1985,hyslop99.
Let $\beta\in\Theta_\beta\subset\mathbb R^{d_\beta}$ denote the subvector of $\theta$ whose value determines whether the model is complete or not. Let $\delta\in\Theta_\delta\subset\mathbb R^{d_\delta}$ collect the remaining components of $\theta$. Given a sample of data $(Y_i,X_i),i=1,\dots, n$, consider testing
where $B_1\subset\Theta_\beta$ is a set not containing $\beta_0$. For instance, in entry games (i.e. Example (ref)), the presence of strategic substitution effects can be tested by letting $\beta_0=0$ and $B_1=\{\beta:\beta^{(j)}<0,j=1,2\}$.\footnote{Our framework also nests settings in which the researcher tests one of the interaction effects, e.g., $\beta^{(1)}_0=0$ and $B_1=\{\beta:\beta^{(1)}<0\}$. In this case, the model is complete under both hypotheses. Our test then reduces to a conventional score test.} Similarly, we may test the potential endogeneity of treatment assignments (Example (ref)) and the presence of state dependence (Example (ref)) by setting $\beta_0=0$ and choosing suitable alternative hypotheses. In what follows, we let $\Theta_0=\{\beta_0\}\times \Theta_\delta$ and $\Theta_1=B_1\times\Theta_\delta$.
Let $\Delta_{Y|X}$ denote the set of conditional distributions of $Y$ given $X$. For each $\theta=(\beta',\delta')'$, an incomplete model admits the following set of conditional distributions:
The conditional distribution $\eta(\cdot|x,u)$ represents the unknown selection mechanism according to which an outcome gets selected from $G(u|x;\theta)$. Reflecting the lack of understanding of the selection, we allow any law supported on $G(u|x;\theta)$. Consequently, the model can admit (infinitely) many likelihood functions for a given $\theta$. Let $\mu$ be the counting measure on $\mathcal Y.$ For each $\theta$, define
This set collects all (conditional) densities compatible with a given $\theta$. In the case of discrete games (Examples (ref)), this set contains all densities of equilibrium outcomes that are compatible with the game's description. Similarly, in the context of panel discrete choice (Example (ref)), this set collects all densities of individual choices consistent with arbitrary specifications of the initial condition. Observe that $\mathfrak q_\theta$ reduces to a singleton set $\{q_\theta\}$ if the model is complete, in which case $q_\theta=dQ_\theta/d\mu$ with $Q_\theta(A|x)=\int 1\{g(u|x;\theta)\in A\}dF_\theta(u)$.
While the multiplicity of likelihood functions may appear challenging, $\mathfrak q_\theta$ can be simplified, and this property also simplifies our tests. By Artstein's inequality, we may rewrite $\mathfrak q_\theta$ as follows galichon2011set,molinari2020microeconometrics:
where
is the conditional containment functional (or belief function) associated with the random set $G(u|x;\theta)$. This function gives the sharp lower bound for the conditional probability $Q(A|x)$ across all $Q$'s belonging to $\mathcal Q_\theta$.\footnote{The upper bound for $Q(A|x)$ is given by the capacity functional $\nu^*(A|X)=F_\theta(G(u|x;\theta)\cap A\ne\emptyset|x)$. It is sufficient to use either of the lower or upper bounds in (ref) because the bounds are related to each other through the conjugate relationship $\nu(A|x)=1-\nu^*(A^c|x)$.} Theoretical properties of the containment functional and numerical approximation methods are well studied.\footnote{See molchanov2005theory for a general treatment. For numerical approximations, see CilibertoTamer2009,galichon2011set. We briefly review them in Appendix (ref). } For us, it is important that the linear inequalities in (ref) characterize $\mathfrak q_\theta$. Together with an extended Neyman-Pearson lemma reviewed below, this characterization makes the score computation feasible. In the following subsection, we briefly review the existing results we will rely on.
Let $p_0(y|x)$ denote the true conditional density of $Y$ given $X$. Consider distinguishing a parameter value $\theta_0$ from another value $\theta_1$ in a parametric model $\{p_\theta,\theta\in\Theta\}$ of conditional densities. This amounts to testing a simple null hypothesis $p_0=p_{\theta_0}$ against a simple alternative hypothesis $p_0=p_{\theta_1}$. By the Neyman-Pearson lemma, the most powerful test is the likelihood-ratio (LR) test. In incomplete models, corresponding null and alternative hypotheses would be $p_0\in\mathfrak q_{\theta_0}$ and $p_0\in\mathfrak q_{\theta_1}$ rendering both hypotheses composite. kz showed that it was possible to extend the Neyman-Pearson lemma to such settings, building on a general result by HS. We briefly summarize their results below.
Let $\phi:\mathcal Y\times\mathcal X\to[0,1]$ be a test and $E_q[\phi(Y,X)]$ be its rejection probability, under conditional density $q$ and the marginal distribution of $X$.\footnote{The rejection probability can be written as $E_q[\phi(Y,X)]=E[E_q[\phi(Y,X)|X]]$. Only the conditional expectation depends on $q$.} For a given $\theta$, the power guarantee of $\phi$ is $\pi_{\theta}(\phi)\equiv\inf_{q\in\mathfrak q_{\theta}}E_q[\phi(Y,X)]$, which is the power value certain to be obtained regardless of the unknown selection mechanism. kz seeked for a level-$\alpha$ minimax test lehmann2005testing such that
and
The minimax test is a procedure that maximizes the power guarantee among tests that meet the uniform size control requirement.
Results from HS imply that, when $\mathfrak q_{\theta_0}\cap\mathfrak q_{\theta_1}=\emptyset$, the rejection region of a minimax test is of the form $\{(y,x):\Lambda(y,x)>t\}$ for a measurable function $\Lambda:\mathcal Y\times\mathcal X\to\mathbb R$. Furthermore, there is a least favorable pair (LFP) of densities $(q_{\theta_0},q_{\theta_1})\in \mathfrak q_{\theta_0}\times\mathfrak q_{\theta_1}$ such that for all $t\ge 0$,
and
where $\Lambda(y,x)=q_{\theta_1}(y|x)/q_{\theta_0}(y|x)$. That is, $q_{\theta_0}\in \mathfrak q_{\theta_0}$ is a density consistent with $\theta_0$ and least favorable for controlling the size of a test among all elements of $\mathfrak q_{\theta_0}$, and $q_{\theta_1}\in\mathfrak q_{\theta_1}$ is a density consistent with $\theta_1$ and least favorable for maximizing a measure of power among all elements of $\mathfrak q_{\theta_1}$. kz exploited these features to show that a level-$\alpha$ minimax test is an LR test based on the LFP:
where $(C,\gamma)$ solves $E_{q_{\theta_0}}[\phi(Y,X)]=\alpha.$
One can also characterize the LFP $(q_{\theta_0},q_{\theta_1})$ as a solution to the following convex program.
The constraints in (ref) and (ref) are the sharp identifying restrictions.\footnote{A common way to use them for identification analysis is to define the sharp identified set as $\Theta_I=\{\theta:P(A|x)\ge \nu_\theta(A|x),a.s.\}.$ That is, given the conditional probability $P(\cdot|x)$ identified from data, one collects all values of $\theta$ satisfying the sharp identifying restrictions. For hypothesis testing, we instead fix $\theta$ and ask what would be a distribution among all distributions satisfying the sharp identifying restrictions, which is least favorable for controlling the size or maximizing the power.} In view of (ref), they are equivalent to imposing the restrictions that $q_0$ belongs to $\mathfrak q_{\theta_0}$ and $q_1$ belongs to $\mathfrak q_{\theta_1}$ respectively. These restrictions are useful for computing the LFP because they are linear in $(q_0,q_1)$. Also, the objective function is strictly convex in $(q_0,q_1)$. One can solve the convex program above numerically in general. For the examples we discussed earlier, it is also possible to compute the LFP analytically (see Appendix (ref)).
To illustrate, let us consider Example (ref). Suppose that the latent payoff shifters $(U^{(1)},U^{(2)})$ follow a bivariate standard normal distribution. We may then compute $\nu_\theta(A|x)$ for each event. Let us take $A=\{(1,0)\}$ as an example. Using (ref) and (ref), we obtain
The last expression corresponds to the probability assigned to the green region in Figure (ref) (right panel) and is the sharp lower bound for the probability of $A=\{(1,0)\}$.
Now consider two parameter values $\theta_0=(0'_2,\delta')'$ and $\theta_1=(\beta',\delta')'$, where $\beta=(\beta^{(1)},\beta^{(2)})'$ with $\beta^{(j)}<0$ for $j=1,2$. As discussed in more detail below, the model is complete when $\beta=0.$ One can show that (ref) reduces to the following equality restrictions:
They uniquely determine the least-favorable null density $q_{\theta_0}$ as follows
where, to ease notation, we use $\Phi_{1}$ and $\Phi_{2}$ to denote $\Phi(x^{(1)}{}'\delta^{(1)})$ and $\Phi(x^{(2)}{}'\delta^{(2)})$.
When $\beta^{(j)}<0,j=1,2$, there are multiple densities satisfying (ref). The least favorable alternative density $q_{\theta_1}$ can be found by minimizing (ref) with respect to $q_1$ subject to (ref). The solution can be expressed analytically. For example, when player 1's strategic interaction effect on player 2 is relatively high, it is given by the following form:\footnote{Appendix (ref) gives a full characterization of the LFP for Example (ref). }
Comparing (ref) and (ref), one can see that $q_{\theta_1}$ tends to $q_{\theta_0}$ as $\beta$ approaches its null value (i.e. 0). Hence, one may view $\theta\mapsto q_{\theta}$ as a “parametric” model. For each $\theta_1$, the density $q_{\theta_1}$ corresponds to the data generating process that is least favorable in detecting $\beta$'s deviation from its null value among all densities compatible with $\theta_1$. By varying $\theta_1$, we may trace out a family of such densities and form a parametric model. We define a least favorable (LF) parametric model as follows.
Part (a) of the definition above is natural because the model is complete under the null hypothesis. Part (b) deserves discussion. We associate each alternative parameter value $\theta\not\in\Theta_0$ with a density that corresponds to the least favorable density for testing between $\theta_0=(\beta_0,\delta)$ and $\theta=(\beta,\delta)$. In other words, for each $\delta$, we focus on testing $\beta_0$ against $\beta$ using the least-favorable distribution for maximizing power. We construct $q_\theta$ this way because we aim at detecting local deviations in terms of $\beta$. As we see below, this approach allows us to capture the sensitivity of $q_\theta$ with respect to a change in the parameter of interest through a score function.\footnote{An alternative choice of $\theta_0\in\Theta_0$ is also possible, which we do not seek here because it does not seem to lead to a tractable score test.} We provide a sufficient condition for the existence of an LF parametric model in the next section.
Let us also note the following points. First, having a complete null model is not sufficient for obtaining a score function. It also requires us to define parametric densities at local alternatives. Our approach is to select the density associated with the least favorable selection mechanism for power maximization. Second, we do not need to know the precise form of the selection mechanism that induces $q_{\theta}$ (for $\theta\in \Theta_1$). Solving the convex program, we “profile out” the selection mechanism and directly obtain the induced density $q_{\theta}$. This is why $q_{\theta}$ is a function of $\theta$ only and does not involve any selection mechanism.
Coming back to equation (ref), the formula suggests we may pretend as if data were generated by a parametric discrete choice model with the given density. Thanks to this feature, most of our analysis below will resemble that of standard discrete choice models.
In this section, we start with an assumption to ensure that the LF parametric model is well-defined over a parameter set that contains the null parameter space $\Theta_0$ and its neighborhood. For this, let $\theta_{h}=(\beta_0'+h',\delta')'$, and let $\mathbb C_\epsilon$ denote an open cube centered at the origin with edges of length $2\epsilon$.
Assumption (ref) (i) holds whenever the model makes a complete prediction under the null hypothesis and is satisfied in the examples discussed in Section (ref). The model can be incomplete under the alternative hypothesis. Assumption (ref) (ii) requires $\mathfrak q_{\theta_{h}}$ does not share any element with $\mathfrak q_{\theta_0}$. Under this condition, it is possible to detect local deviations from the null hypothesis regardless of the unknown selection mechanism. Such alternatives are robustly testable in the sense that there exists a test that has nontrivial power against any distribution in $\mathfrak q_{\theta_{h}}$ kz. For this condition, it suffices to have an event $A\subset\mathcal Y$ such that $q_{\theta_0}(A|x)<\nu_{\theta_{h}}(A|x)$ (or $q_{\theta_0}(A|x)>\nu_{\theta_{h}}^*(A|x)$) for all $\tau>0$ for some $x\in\mathcal X$. All of the examples in Section (ref) satisfy Assumption (ref) (ii), and we demonstrate how to show this condition in Appendix (ref). We construct score tests that have power against robustly testable local alternatives.\footnote{kz extend the notion of local alternatives and analyze a more general setting that does not require Assumption (ref) (ii). We conjecture that we may extend our framework similarly. Since all of our examples satisfy Assumption (ref) (ii), we leave this extension elsewhere.}
Let us revisit the examples. \setcounter{example}{0}
We also note that Examples 2-3 reduce to complete models under the null hypothesis.\footnote{To save space, we show Assumption (ref) (ii) for these examples in Appendix (ref).}
We conclude this subsection with the following proposition.
Score-based tests such as Rao's score (or Lagrange multiplier) test and Neyman's $C(\alpha)$ test are widely used. They require the estimation of the restricted model only, which is particularly attractive in our setting. The restricted model is complete and typically admits point estimation of nuisance parameters under reasonably weak conditions. We take advantage of this property to carry out a score-based test. Below, we briefly review the core ideas behind the classic score tests and discuss extensions to handle potential model incompleteness under the alternative. For expositional purposes, we assume $q_\theta$ is differentiable with respect to $\theta$ for now and will weaken this assumption later.
Consider testing the null parameter value $\theta_0=(\beta_0',\delta')'$ against a local alternative hypothesis $\theta_h=(\beta_0'+h',\delta')'$, where $h\in\mathbb R^{d_\beta}$. As discussed earlier, the optimal test in terms of guaranteed power is the likelihood-ratio test based on the LFP. The test is also robust in that, under Assumption (ref) (ii), the log-likelihood ratio can detect any deviation from the null hypothesis with non-trivial power regardless of the selection mechanism.
One can locally approximate the log-likelihood ratio by $\sum_{i=1}^n h's_{\beta}(Y_i|X_i;\beta_0,\delta)$, where $s_{\beta}(y|x;\beta,\delta)=\frac{\partial}{\partial \beta }\ln q_\theta(y|x)|_{\theta=(\beta,\delta)}$ is the score function. Let $\Sigma_{\beta_0}=\text{Var}(\sum_{i=1}^ns_{\beta}(Y_i|X_i;\beta_0,\delta))$. For i.i.d. data, $\Sigma_{\beta_0}=n I_{\beta_0}$ where $I_{\beta_0}=E[s_{\beta}(Y_i|X_i;\beta_0,\delta)s_{\beta}(Y_i|X_i;\beta_0,\delta)']$. For a fixed $h$, the normalized quantity
serves as a robust measure of discrimination between $\beta_0$ and $\beta_0+h$. It locally approximates the log-likelihood ratio $\ln (q_{\theta_h}/q_{\theta_0})$. The direction $h$ that maximizes (ref) is $h^*=I_{\beta_0}^{-1}\frac{1}{\sqrt n}\sum_{i=1}^n s_{\beta}(Y_i|X_i;\beta_0,\delta)$, which motivates Rao's score statistic:\footnote{See Bera:2001va for a more detailed argument for complete models. The same argument can be applied to incomplete models by replacing the standard likelihood function with the LF density $q_\theta$.}
Suppose that the nuisance parameter $\delta$ can be estimated by a point estimator $\hat\delta_n$. Evaluating the sample mean of the score at $\delta=\hat\delta_n$ and imposing the null hypothesis yields
A feasible version of (ref) is
where $\hat V_n$ is an estimator of the asymptotic variance $V_0\equiv I_{\beta_0}$. For example, one can use the sample analog $\hat V_n= n^{-1}\sum_{i=1}^ns_\beta(Y_i|X_i;\beta_0,\hat\delta_n)s_\beta(Y_i|X_i;\beta_0,\hat\delta_n)'$ or its regularized version.\footnote{For example, the following estimator proposed by andrews_barwick12 ensures that it is always nonsingular and is equivariant to scale changes
where $\hat\Sigma_n=n^{-1}\sum_{i=1}^ns_\beta(Y_i|X_i;\beta_0,\hat\delta_n)s_\beta(Y_i|X_i;\beta_0,\hat\delta_n)'$, $\hat D_n=\text{diag}(\hat\Sigma_n)$, and $\hat\Omega_n=\hat D_n^{-1/2}\hat\Sigma_n\hat D_n^{-1/2}$. } Under regularity conditions, $\hat T_n$ converges in distribution to a $\chi^2$-distribution with $d_\beta$ degrees of freedom under the null hypothesis.
The analysis so far presumed that $q_\theta$ was differentiable, and $h\in\mathbb R^{d_\beta}$ was unrestricted. These assumptions may be restrictive in our context. For example, in discrete games of complete information, the least favorable parametric model $h\mapsto q_{\theta_h}$ and its score can take different functional forms depending on whether the alternative hypothesis admits strategic substitution (i.e. $h<0$ as in Example (ref)) or strategic complementarity (i.e. $h>0$). It is then natural to analyze these two cases separately. Below, we weaken differentiability requirements to accommodate these features and allow the alternative hypothesis to be restricted (e.g., one-sided).
Recall that a set $\Gamma\subseteq\mathbb R^d$ is said to be locally equal to set $\Upsilon\subseteq\mathbb R^d$ if $\Gamma\cap \mathbb C_\epsilon=\Upsilon\cap \mathbb C_\epsilon$ for some $\epsilon>0$ Andrews:1999aa.
Assumption (ref) (i) requires the set of deviations (from $\beta_0$) can be locally approximated by a convex cone. In Example (ref), consider testing $H_0:\beta=(0,0)'$ against $H_1:\beta^{(1)}<0,\beta^{(2)}<0$. Then, $B_1-\beta_0$ is locally equal to
Assumption (ref) (ii) uses the notion of differentiability in quadratic mean Van-der-Vaart:2000aa, but it only requires that a unique score, in the sense of the $L^2$-derivative of the square-root density, exists for the set $\mathcal V_1$ of local deviations from the null hypothesis. This weaker assumption is appropriate for incomplete models, and $s_\theta$ can be derived from the least favorable parametric model similar to the standard parametric models (see Appendix (ref)).
To accommodate the one-sided nature of the alternative hypothesis, we define a test statistic by
This test statistic is a modification of (ref) and follows the construction in Silvapulle:1995tm. It requires the same functions of data as $T_n$, but it is designed to direct power against the local alternatives in $\mathcal V_1$. If the alternative hypothesis is locally unrestricted, i.e., $\mathcal V_1=\mathbb R^{d_\beta}\setminus 0$, the test statistic reduces to $T_n$.
The asymptotic distribution of $\hat S_n$ is no longer a $\chi^2$-distribution. However, its critical value is easy to compute using simulations. Let
where
which can be simulated by drawing $Z$ repeatedly from a zero mean multivariate normal distribution with estimated variance $\hat V_n.$
Let $q_{\beta_0,\delta}$ be the conditional density of $Y_i$ given $X_i$. By Assumption (ref) (i), this density is unique. A natural estimator of $\delta$ is the restricted maximum likelihood estimator (RMLE) $\hat\delta_n$, which maximizes the log-likelihood function
The complete model (under $H_0$) is often a standard discrete choice problem. Hence, one can use package software (e.g., R, Stata) to compute the RMLE. Let us revisit the examples.
\setcounter{example}{0}
This section collects results on the asymptotic properties of the score test. The proofs of all theoretical results are in Appendix (ref). Throughout, we assume that $U^n=(U_1,\dots,U_n)$ is an independent and identically distributed (i.i.d.) sample drawn from $F_\theta$, and $X^n=(X_1,\dots,X_n)$ is also an i.i.d. sample following $q_X^n$. The joint distribution of the outcome sequence $Y^n=(Y_1,\dots,Y_n)\in\mathcal Y^n$ conditional on $X^n=x^n$ is not uniquely determined due to the potential incompleteness of the model. For $\theta\in\Theta$, the distribution belongs to the following set:
where $F^n_\theta$ denotes the joint law of $U^n$, and $G^n(u^n|x^n;\theta)=\prod_{i=1}^n G(u_i|x_i;\theta)$ is the Cartesian product of the set-valued predictions. We let $\mathcal P_{\theta}^n$ collect joint laws of $(Y^n,X^n)$; each element $P^n$ of $\mathcal P_{\theta}^n$ is such that the conditional law of $Y^n$ given $X^n$ belongs to $\mathcal Q_{\theta}^n$, and the law of $X^n$ is $q^n_X$. Assuming $U^n$ and $X^n$ are i.i.d. does not imply $Y^n$ is i.i.d. The set $\mathcal Q^n_\theta$, in general, contains dependent and heterogeneous laws because the behavior of the selection mechanism across experiments is unrestricted eks. This feature does not create an issue for the size properties of our test because $\mathcal Q_\theta^n$ reduces to a single i.i.d. law under the null hypothesis.
For the asymptotic properties of the RMLE, we also allow $\beta$ to be in a local neighborhood of $\beta_0$. For such settings, we provide conditions under which $\hat\delta_n$ is $\sqrt n$-consistent. Let $h\in \mathcal V_1$ and $\delta_0\in\Theta_\delta$. We assume data are generated from $P^n \in \mathcal Q^n_{\beta_0+h/\sqrt n,\delta_0}$. The null hypothesis corresponds to the setting with $h=0$.
Fixing $\beta=\beta_0$, one can view $q_{\beta_0,\delta}$ as the conditional density of $Y$ in a regular parametric model, in which $\delta$ is the only unknown parameter. For each $\delta\in\Theta_\delta$, let $\mathbb M(\delta)\equiv E[\ln q_{\beta_0,\delta}(Y_i|X_i)]$, where expectation is taken with respect to the conditional density $q_{\beta_0,\delta_0}$ and the distribution of $X.$ Let $\mathbb{M}_n(\delta)$ be the sample counterpart of $\mathbb M$ defined in (ref).
Assumption (ref) imposes sufficient conditions for identification and uniform law of large numbers standard in the literature. We also assume the density of $F_\theta$ depends on $\theta$ smoothly. For this, for any integrable function $f$ defined on a measure space $(A,\mathfrak F,\zeta)$, let $\|f\|_{L^1_\zeta}$ be the $L^1$-norm of $f$.
Finally, the following condition ensures the population objective function is locally well behaved so that its value is informative about $\delta_0$.
Under these assumptions, the restricted MLE $\hat\delta_n$ is $\sqrt n$-consistent.
Below, let $P^n_0\in \mathcal P^n_{\theta_0}$ be the joint law of $(Y^n,X^n)$ under the null hypothesis. Let $s_{\theta,j}$ be the $j$-th component of $s_\theta$. Let
and let $\Xi=\big\{f:\mathcal Y\times\mathcal X\to\mathbb R|f(y,x)=\xi_{j,k}(y,x;\delta),~1\le j,k\le d,\delta\in\Theta_{\delta}\big\}.$ Next, we add a condition for the asymptotic distribution of $\hat S_n$ and consistent estimation of the asymptotic variance $V_0$.
Here we assume the expected score can be linearized and the elements of $\Xi$ obey a uniform law of large numbers. Suppose $\hat S_n$ as defined in (ref). The following theorem shows that the test controls its asymptotic size.
In some applications, the ultimate goal may be to make inference on the underlying parameter, for example, to construct confidence intervals for components of $\theta$. While we defer a formal analysis to future work, we suggest a hybrid procedure that aims at controlling the potential distortion of the model selection step, borrowing insights from the moment selection literature andrews2010inference,romano2014practical.
Consider constructing confidence intervals for a component or linear combination $\gamma_0=p'\delta_0$ of $\delta_0$.\footnote{Since the null hypothesis pins $\beta$'s value down, it is natural to consider inference on the parameters that are estimated under both null and alternative hypotheses.} Due to Proposition (ref), $\hat\gamma_n=p'\hat\delta_n$ is a $\sqrt n$-consistent estimator of $\gamma_0$ as long as the true value of $\beta$ is in a neighborhood of $\beta_0$ whose radius is of order $n^{-1/2}$. It would be natural to use such an estimator to construct a confidence interval for $\delta_0$ if the complete model is selected. A well-known challenge for such post-model selection inference is that a naive asymptotic approximation that disregards the model selection step may not be valid uniformly over a large class of data generating processes leeb2005model,andrews2009hybrid. Given this, we consider the following hybrid method.
Step 1: Compute $\hat S_n$ and $c_{n}=(\kappa_n\wedge 1)c_\alpha$, where $\kappa_n$ is a sequence of shrinkage factors that tends to 0 slowly, e.g. $\kappa_n=(\ln n)^{-1/2}$;
Step 2:
The heuristic behind this procedure is as follows. First, we compare $\hat S_n$ to a critical value $c_n$ that tends to 0 slowly. For DGPs whose $\beta$ is outside local neighborhoods of $\beta_0$, we cannot ensure the asymptotic validity of the Wald confidence interval. In such settings, the procedure above uses a robust confidence interval asymptotically, which controls the asymptotic coverage probability. Since the critical value tends to 0, we use the Wald confidence interval only if $\beta$ is in a local neighborhood of $\beta_0$. The shrinkage factor $\kappa_n$, therefore, introduces a conservative distortion, which is expected to make the resulting confidence interval's coverage probability above its nominal over a wide range of $\beta$ values.
We illustrate the score test through two empirical applications.
The first application revisits the analysis of the airline industry by klinetamer2016bayesian. We test the presence of strategic interaction effects between two types of firms: low-cost carriers (LCC) and other airlines (OA). Below, we briefly summarize the setup and refer to klinetamer2016bayesian for details. A market is defined as trips between airports regardless of intermediate stops. The two types of firms, LCC and OA, decide whether or not to serve each market. The binary variable $y^{(\ell)}_i$ takes value 1 if airline $\ell\in\{\text{LCC, OA}\}$ serves market $i$. Airline $\ell$'s payoff in market $i$ equals $$y_{i}^{(\ell)}(\delta^{cons}_{\ell}+\delta^{size}_{\ell}X_{i, size}+\delta^{pres}_{\ell}X^{(\ell)}_{i, pres}+\beta_{\ell}y^{(-\ell)}_{i}+u^{(\ell)}_{i}),$$ where $\beta_{\ell}$ captures the impact of the competitor's entry decision, $y^{(-\ell)}_{i}$. The airline-specific intercepts and observable covariates determine each firm's payoff. The covariates include the market size $X_{i, size}$ and the market presence $X^{(\ell)}_{i,pres}$. The market size $X_{i, size}$ is defined as the population at the endpoints of each trip. The latter variable $X^{(\ell)}_{i,pres}$ measures the presence of firm $\ell$ in market $i$ (see klinetamer2016bayesian p.356 for its definition). This airline-and-market-specific variable shows up only in firm $\ell$'s payoff. The data come from the second quarter of the 2010 Airline Origin and Destination Survey (DB1B) and contain 7882 markets.\footnote{The data are available on Brendan Kline's \href{www.brendankline.com}{website}.}
Our hypothesis of interest is whether the LCCs and OAs compete strategically, which can be formulated as a one-sided test. The null hypothesis is $H_0:\beta_{LCC} = \beta_{OA} = 0$, and the alternative hypothesis is $H_1:\beta_{\ell}<0,\ell \in \{\text{LCC,OA}\}$. Finally, the vector of coefficients $\delta=(\delta^{cons}_{LCC},\delta^{size}_{LCC},\delta^{pres}_{LCC},\delta^{cons}_{OA},\delta^{size}_{OA},\delta^{pres}_{OA})$ is the nuisance parameter in this model. We estimate $\delta$ by the restricted MLE under the null hypothesis.
The value of the test statistic is 24.668. The 5% critical value is 5.050. Hence, we reject the null hypothesis at the 5% level. This result is consistent with the finding of klinetamer2016bayesian whose credible sets for the strategic interaction effects $\beta_\ell,\ell\in\{\text{LCC,OA}\}$ do not contain the origin. Table (ref) reports the RMLE of the index coefficients. The estimates suggest that the effect of market presence is larger for LCCs than other airlines, and the monopoly profits (captured by the constant terms) in a market with below-median size and below-median market presence are smaller for the LCCs. These observations are also consistent with klinetamer2016bayesian's findings, although we note that these estimates are obtained by imposing the restriction rejected by the score test.
The second application concerns the causal effect of Catholic school attendance on academic achievements studied by AltonjiElderTaber2005. Whether Catholic schools provide a better education than public ones is important for education policies, but the analysis is complicated by the concern that selection into Catholic schools is nonrandom. Using the framework in Example (ref), we examine the endogeneity of Catholic school attendance by testing if the coefficient on the control function is zero.
The data source is a subset of the National Educational Longitudinal Survey of 1988 (NELS:88). We use a version of the data available from WooldridgeText. We refer to AltonjiElderTaber2005 for a detailed discussion of the data. The dependent variable $y_{i}$ is a binary variable indicating whether the student graduated from high school by the year 1994. The binary treatment $d_{i}$ indicates whether the student attended a Catholic high school. The vector of exogenous control variables $w_{i}$ includes each parent's years of education and log family income. The instrument variable is a dummy variable indicating whether a parent was reported to be Catholic. The sample size $n$ is 5970 after we remove missing observations on $y_{i}$.
Table (ref) reports the point estimates of nuisance parameters $\delta$ under the null hypothesis $H_0:\beta=0$. The value of the test statistic is 154.848. The 5% critical value is 2.755. We, therefore, reject the null hypothesis at the 5% level. Our test provides strong evidence supporting the students' selection into Catholic schools based on their unobservable characteristics. This result is in line with the concern expressed in AltonjiElderTaber2005.
We examine the size and power properties of the score test through simulations. The data generating process is based on Example (ref) and is motivated by the empirical illustration in the previous section. There are player-specific covariates $X_i=(X_i^{(1)},X_i^{(2)})'$, each of which is generated as an independent Rademacher random variable taking values on $\{-1,1\}$. We then generate $U_i = (U^{(1)}_i, U^{(2)}_i)$ from the bivariate standard normal distribution. For each $u_i$ and $x_i$, we determine the predicted set of outcomes $G(u_i|x_i;\theta)$ based on the payoff functions with $\delta_0=(\delta_0^{(1)},\delta_0^{(2)})=(2,1.5)'$. We then test
As discussed earlier, the model is complete under $H_0$. We estimate $\delta_0$ using the restricted MLE. The sample size is set to 2500, 5000, or 7500. This choice is motivated by the sample size used in the empirical application.
The size of the score test is reported in Table (ref). The size of the test is controlled properly across all sample sizes, while it tends to be slightly conservative when $n$ is small.
Under alternative hypotheses, multiple equilibria may be predicted. If this is the case, we select an outcome according to one of the following selection mechanisms. The first design uses a selection mechanism, which selects $(1, 0)$ out of $G(u_i|x_i;\theta)=\{(1, 0), (0, 1)\}$ if an i.i.d. Bernoulli random variable $\nu_{i}$ takes 1. In the second design, we generate data from the least favorable distribution, which draws an independent outcome sequence from the least favorable distribution $Q_{\theta_1}\in \mathcal Q_{\theta_1}$.
The power of the score test is calculated against local alternatives with $\beta^{(j)}_1=-h/\sqrt n,h>0$ for $j=1,2.$ For this exercise, we introduce a grid of values for $h$ and generate the data described above. We then compare the rejection frequency of our test to that of the moment-based testing procedure by BCS. Their test checks if a hypothesized value $(\beta^{(1)},\beta^{(2)})'=(0,0)'$ is compatible with a set of moment restrictions. Their statistic and bootstrap critical value are calculated using a sample analog of the following moment inequality and equality restrictions
which are the sharp identifying restrictions that characterize $\mathfrak q_\theta$ in (ref).\footnote{Since the example resembles the specification used in their Monte Carlo experiments, we added minimal changes to their replication code posted on the repository of Quantitative Economics to implement their procedure.}
Figures (ref)-(ref) show the rejection frequencies of the score and moment-based tests. The results are similar across the two designs. In each design, the score test outperforms the moment-based test in terms of power by a significant margin. This difference in performance may potentially be due to the proposed score test's exploitation of completeness under the null to estimate the nuisance parameter and direct power against $\mathcal V_1$. The moment-based test is designed for general subvector inference and does not necessarily exploit the model completeness.\footnote{The procedure by BCS tests if the null parameter value is consistent with the model restrictions and deals with nuisance parameters by a profiling method combined with regularization to ensure its uniform validity.} The simulation results suggest that taking advantage of the model structure may provide considerable benefits in terms of power.
Economic models exhibit incompleteness for various reasons. They are, for example, consequences of strategic interaction, state dependence, or self-selection. This paper shows that one can test these important features using a score statistic even if the model is incomplete under the alternative hypothesis. The proposed test exploits the model completeness under the null hypothesis to simplify its computation. An avenue for future research includes a theory for the uniform validity of inference for post-model selection procedures based on the score test.