Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
65,171 characters · 13 sections · 83 citation commands
Universal Inference for Incomplete Discrete Choice Models
We are grateful to \'{A}ureo de Paula, Hidehiko Ichimura, Toru Kitagawa, Seojeong Lee, Simon Lee, Kirill Ponomarev, Azeem Shaikh, and Alex Torgovitsky for their helpful comments. We thank, for their comments, the seminar, lecture, and conference participants at the following places: Chicago, CyberAgent (AI lab), KER International Conference, Kyoto Summer Workshop in Applied Economics, and EEA-ESEM. We thank Undral Byambadalai, Yan Liu, Junwen Lu, Jiahui Guo and Jie Huang for their excellent research assistance. Financial support from NSF grants SES-1824344 and SES-2018498 is gratefully acknowledged. Part of the research was conducted at Microsoft Research (New England) and the University of Tokyo while Kaido was visiting them. Financial support from the Young Scientists Fund of the National Natural Science Foundation of China (NSFC) grant No.72203077 is gratefully acknowledged.}}
\onehalfspacing Discrete choice models are widely used to study economic agents' decisions. Such models may make set-valued predictions when the researcher is willing to relax restrictive assumptions. A commonly used structure is that given observable and unobservable exogenous variables $(X,U)$, multiple values $G(U|X;\theta)$ of an outcome variable $Y$ are predicted. However, how $Y$ is selected from the predicted set is unspecified. A growing number of empirical models exhibit such predictions. Examples include but are not limited to discrete games bresnahan1990entry,Ciliberto:2009aa, dynamic panel data models Honore:2006aa,chesher2024robust, discrete choice models with heterogeneous choice sets bar:cou:mol:tei21, models with instrumental variables che:ros17,Berry:2022aa, auctions Haile:2003to, network formation Miyauchi:2016aa,Sheng2020, product offerings EIZENBERG:2014aa, exporter's decisions Dickstein:2018aa and school choices FackGrenetHe17. This class nests classic discrete choice models such as binary, multinomial, and ordered choice models, in which $G$ contains a single value of $Y$. Despite the recent developments in the econometrics literature, the existing methods for estimating such incomplete models often face challenges in applications. In particular, the limiting distributions of the existing statistics for full-vector and sub-vector inference are often non-standard and require complex regularity conditions or tuning parameters.
We propose a novel inference method for conducting hypothesis tests and constructing confidence sets (or intervals) in incomplete models. Specifically, we consider testing the composite hypotheses:
for subsets $\Theta_0$ and $\Theta_1$ of a parameter space $\Theta$. Our proposal is to compare a tailor-made likelihood-ratio statistic to a fixed critical value, which builds on the universal inference method of Wasserman:2020aa. We show that the test is universally valid, meaning the proposed test has a correct size in any finite sample without complex regularity conditions.
Since the critical value is fixed, the proposed method remains computationally tractable in models of practical relevance. For example, we allow models to involve a number of discrete or continuous covariates and their associated parameters. The proposed test also avoids the use of moment selection or other regularization tuning parameters. Inverting the test yields confidence sets (or intervals) for the structural parameter, their subcomponents, and counterfactual objects.
This paper aims to develop a versatile inference procedure that is widely applicable. We demonstrate that any incomplete discrete choice model possesses a structure that allows us to conduct the proposed test. The main idea is as follows: When trying to distinguish between the null parameter set $\Theta_0$ and any (unrestricted) parameter value $\theta_1\in \Theta$, a classic idea is to use a likelihood ratio (LR) to compare the model fit at each $\theta\in\Theta_0$ against $\theta_1$. While the model incompleteness makes it difficult to apply this idea to the current setting, we show that a special pair of densities called least-favorable pair exists for testing $\theta\in\Theta_0$ against $\theta_1$. This insight allows us to construct a parametric model $\{q_\theta\}$ indexed only by $\theta$, which is useful for constructing a robust test.\footnote{This parametric model can be derived analytically in many existing models kaido2019robust. We also provide a command to implement the procedure as part of a general Python library for incomplete discrete choice models (\url{https://github.com/hkaido0718/IncompleteDiscreteChoice}).} Using this construction, we obtain a novel LR statistic for universal inference.
Our framework allows the researcher to leave parts of individuals' decision processes unspecified and heterogeneous. For example, bar:cou:mol:tei21 left the unknown choice-set formation process flexible. The choice of an insurance plan may involve some buyers forming their choice set using an online platform, while others may form their choice sets based on recommendations by their insurance agents. These potentially heterogeneous choice set formation processes are unobservable to the econometrician and may not be understood well enough to be modeled.\footnote{If the researcher understands how the process operates, they can formulate it using a probabilistic model. Some studies take this approach bjorn_vuong_1984,Bajari:2010aa.} In an extreme case, each decision maker may have their own choice-set formation process, rendering them incidental parameters. While some empirical studies impose regularity conditions that implicitly limit the heterogeneity of decision processes, our proposed procedure remains valid without such assumptions.\footnote{Existing studies often impose some assumption on the heterogeneity of outcome distributions to apply limit theorems. A leading example is to assume identically distributed outcomes, which holds under identically distributed selection mechanisms. A weaker assumption is that samples are stationary and strongly mixing chernozhukov2007estimation,andrews2010inference, which imposes implicit restrictions on the selection mechanisms. An exception is Epstein:2016qv, which allowed arbitrary heterogeneity and dependence.}
The proposed procedure has two novel features compared to standard treatments of LR tests. Integrating them into the standard inference toolkit is straightforward. First, it utilizes a tailor-made likelihood function that can be derived by solving a convex program. We provide an algorithm and demonstrate that leading examples have closed-form likelihoods. Second, the test employs cross-fitting. Specifically, we obtain a parameter estimate from one subsample, use the other subsample to evaluate the likelihood ratio, and aggregate the statistics after swapping the roles of the two subsamples. This approach enables us to apply a straightforward conditioning argument and a Chernoff-style bound developed in Wasserman:2020aa, yielding a simple yet non-trivial critical value.
The method proposed here offers a tractable way for practitioners to perform inference on various models while avoiding ad-hoc assumptions. It is particularly effective in settings where the outcome is relatively low-dimensional, involves discrete and continuously distributed covariates, and $\theta$ may contain nuisance parameters. Such examples are ubiquitous. For the theoretical front, this paper provides a new likelihood-based inference method with finite sample validity with a focus on inference for subvectors and counterfactual objects. To keep a tight focus, we limit our attention to the finite-sample validity of the test and evaluate its power only through simulations. The proposed procedure is not designed to be optimal due to the use of sample-splitting and a Chernoff-style bound. In a companion paper kaido2019robust, we build a theoretical framework for analyzing the power of likelihood-ratio tests and their optimality.
The study of incomplete systems has a long history dating back to the work of Wald1950. Economic models with multiple equilibria are well-known examples of incomplete models. jovanovic1989observable developed a theoretical framework to examine the empirical content of such models. The class of incomplete models considered here also covers a wide range of empirical models beyond them, as discussed earlier. Systematic ways to derive partially identifying restrictions for such models have been developed tamer2003incomplete,galichon2011set,beresteanu2011sharp,che:ros17. Building on this line of work, we use the sharp identifying restrictions to incorporate all information in the original structural model to construct a likelihood function.
We contribute to the literature on inference by providing a novel test that has the following properties: (i) it has finite sample validity without complex regularity conditions (universal validity); (ii) it is applicable to models with mixed data types (e.g., continuous and discrete covariates); (iii) one can construct confidence regions for the entire parameter or its functions; (iv) it is robust to the presence of incidental parameters (selections); (v) it can accommodate nonparametric parameter components.\footnote{Each of the properties has been studied somewhat separately. For (i), other than Li:2022aa, Horowitz:2023aa develop a method with finite sample validity for models represented by optimization problems; (ii) For discrete covariates, one can use unconditional moment inequalities Canay_Shaikh_2017. For continuous covariates, one needs to work with a continuum (or increasing number) of moment inequalities Andrews:2013aa,chernozhukov_lee_rosen13. See also Kaido:2023aa for this point who develop a method related to this paper's approach; (iii) Subvector inference is studied, for example, by BCS,KMS,Cox:2022aa,Andrews:2023aa; For (iv), Hahn:2010aa tackles the problem using panel data. Epstein:2015aa,Epstein:2016qv develop robust Bayesian and frequentist inference methods; For (v), there is a vast literature on semiparametric models POWELL19942443.} To our knowledge, Li:2022aa is the only existing work that develops a finite-sample valid method for incomplete models using the idea of Monte Carlo tests. Our proposal differs from theirs mainly in two respects. First, we focus on composite hypothesis tests with inference on subvectors and counterfactual objects in mind, while they focus on full-vector inference. Second, we use a fixed critical value, which avoids simulated draws of latent variables used in their method to calculate critical values. Our test uses a likelihood-ratio statistic. Likelihood-based methods are also considered by Chen:2011aa,Chen_2018 and Kaido:2023aa who show their methods' asymptotic uniform validity. In contrast, this paper aims to achieve finite-sample validity.
Likelihood-based inference is commonly used. However, as Wasserman:2020aa notes “The (limiting) null distribution of the classical likelihood-ratio statistic is often intractable when used to test composite null hypotheses in irregular statistical models.” They develop inference methods that do not require complex regularity conditions using split-sample and cross-fit versions of LR statistics. This approach is also appealing for incomplete models, which are typically irregular. Nonetheless, Wasserman:2020aa's Wasserman:2020aa framework is not directly applicable to the current setting because their method uses a unique likelihood function as a key input and requires the knowledge of the sampling process.\footnote{As a baseline, they assume random sampling. To be precise, they assume the knowledge of the conditional likelihood based on a subsample given another subsample, which is unknown in incomplete models. See Section (ref).} Neither of them is readily available in incomplete models. This is because, for each parameter, the model implies multiple (typically infinitely many) likelihoods. Furthermore, the unspecified selection mechanism (together with observable and unobservable variables) can vary across experiments. We address these issues by carefully constructing a likelihood function using the model structure, hence the name “tailor-made” likelihood.
The universal validity is related to the asymptotic uniform validity requirements shown for many of the existing proposals based on moment inequality restrictions Canay_Shaikh_2017. Both are the notions of the validity of an inference procedure over a wide class of data-generating processes. The former requires validity in any finite samples, whereas the latter requires it in large samples. Wasserman:2020aa also emphasize that they avoid regularity conditions in defining the set of data-generating processes, which we also follow here. Some of the recently developed moment-based inference methods study subvector inference and attain computational tractability by focusing on specific classes of models or testing problems Cox:2022aa,Andrews:2023aa. This paper focuses on the class of incomplete discrete choice models, which is not nested by (nor nests) the class of models considered in these proposals.
Let $Y\in \mathcal{Y}\subseteq\mathbb R^{d_Y}$ and $X\in \mathcal{X}\subseteq\mathbb R^{d_X}$ denote, respectively, observable endogenous and exogenous variables, and $U\in \mathcal{U}\subseteq\mathbb R^{d_U}$ denote latent variables. We assume $\mathcal{Y}$ is a finite set. Let $\mathcal{Z}=\mathcal{Y}\times\mathcal{X}$. We equip $\mathcal{Z}$ with the Borel $\sigma$-algebra $\Sigma_{\mathcal{Z}}$ and let $\Delta(\mathcal{Z})$ denote the set of all Borel probability measures on $(\mathcal{Z},\Sigma_\mathcal{Z})$. Let $\Theta$ be a parameter space. We do not restrict $\Theta$. Hence, both parametric and nonparametric components are accommodated. Let $F_\theta(\cdot|x)$ denote the conditional distributions of $U$ given $X=x$. We let $F=\{F_\theta,\theta\in\Theta\}$ be the collection of the conditional laws.
For each $\theta\in\Theta$, let $G(\cdot|\cdot;\theta):\mathcal{U}\times\mathcal{X}\twoheadrightarrow \mathcal{Y}$ be a weakly measurable correspondence, which collects permissible outcome values for each $(x,u)$. The observable outcome $Y$ is a random vector satisfying
Such a random vector is called a measurable selection of $G(U|X;\theta)$. The model does not impose any restrictions on how $Y$ is selected. This structure nests models with a complete prediction, characterized by a function $g(\cdot|\cdot;\theta):\mathcal{X}\times\mathcal{U}\to\mathcal{Y}$ such that $Y=g(U|X;\theta),~a.s.$
Let $\mathcal{C}$ be the collection of all subsets of $\mathcal{Y}$. Define the containtment functional of $G$ by
This functional characterizes all conditional distributions of the measurable selections of $G(U|x;\theta)$ by the following set, known as the core of $\nu_\theta$ art83:
where $\mathcal{M}(\Sigma_Y,\mathcal{X})$ is the collection of laws of random variables supported on $\mathcal{Y}$ conditional on $X$. The inequalities $Q(\cdot|x)\ge \nu_\theta(\cdot|x)$ characterizing $\mathcal P_{\theta,x}$ are known as the sharp identifying restrictions as they can partially identify $\theta$ without losing information galichon2011set,mol:mol18. Our method applies to any model characterized by such restrictions.
Any (conditional) distribution in $\mathcal P_{\theta,x}$ can also be expressed as follows philippe1999decision:
The unknown function $\eta(\cdot|u,x)$ is the conditional distribution of $Y$ over the prediction, representing an unknown selection mechanism. It may represent objects with different interpretations, such as equilibrium selection mechanisms (Example (ref)), choice-set formation processes (Example (ref)), fixed effects, and the initial conditions (Example (ref)).\footnote{Other examples include the unobserved true value of variables in a set-valued covariate or control function manskitamer02,han2024setvalued and details of agents' behavior beyond minimal rationality or stability restrictions Haile:2003to,FackGrenetHe17.} The characterization in (ref) is tractable as it allows us to restrict the conditional distribution of $Y$ only through inequality restrictions without introducing $\eta(\cdot|u,x)$ explicitly.
Let $\mu$ be the counting measure on $\mathcal{Y}.$ For each $x\in\mathcal{X}$, let
collect the set of conditional densities compatible with $\theta$ at $x$. We then let $\mathfrak{q}_{\theta}\equiv\{\mathfrak{q}_{\theta,x},x\in \mathcal{X}\}$.
We illustrate $G$ with well-known examples.
The next example concerns a single-agent discrete choice problem with unobserved choice sets.
The next example is a panel dynamic discrete choice model with short panel data Honore:2006aa,KHAN2023105515,chesher2024robust.
This section overviews the proposed procedure. Let $(Y_i,X_i),i=1,\dots, n$ be a sample of outcome and covariates. Consider testing
for subsets $\Theta_0,\Theta_1$ of the parameter space. The proposed procedure takes the following steps.
Step 1: Split samples into $D_0$ and $D_1$. Let $\hat\theta_1$ be any estimator of $\theta$ computed from sample $D_1$.
Step 2: Using $D_0$, construct a tailor-made likelihood function
where $q_\theta(y|x)$ is the LFP-based density we introduce below. Let $\hat\theta_0$ be the restricted maximum likelihood estimator (RMLE) based on $D_0$:
Step 3: Compute the split-sample likelihood-ratio (LR) statistic:
The cross-fit LR statistic is
where $T_n^{\text{swap}}$ is calculated in the same way as $T_n$ after swapping the roles of $D_0$ and $D_1$.
Step 4: Reject $H_0$ if $S_n> 1/\alpha$. Do not reject $H_0$ otherwise.
One can construct confidence regions for $\theta$ and its functions by defining $\Theta_0$ properly and inverting the test (see Section (ref)).
We focus on key properties of the proposed test for now and defer discussion on how to construct $\hat\theta_1$ and $\mathcal L_0$ to the next sections. The cross-fit LR test controls its size in any finite sample. That is, for any $n$ and the class $\mathcal{P}^n_0$ of data-generating processes (DGPs) compatible with the null hypothesis, the following statement holds
Our test builds on Wasserman:2020aa, who introduced the split-sample and cross-fit LR statistics for probabilistic models. Following them, we call the procedure above universal inference (or universal hypothesis test). The idea is that the inference applies universally to any model described by (ref) without further regularity conditions.
In the next few sections, we provide details on how to construct the LR statistic. The key is to construct $\mathcal L_0$ from a least favorable parametric model $\{q_\theta,\theta\in\Theta_0\}$ against a density $q_{\hat\theta_1}$ in the unrestricted model.
Any estimator of $\theta$ based on sample $D_1$ can be used as an unrestricted estimator $\hat\theta_1$. The main role of this estimator is to find a “representative” parameter value that fits $D_1$ well. Hence, we do not require $\hat\theta_1$ to be a consistent estimator. Nevertheless, we recommend choosing $\hat\theta_1$ to maximize a measure of fit to $D_1$ for power consideration.
A natural choice is an extremum estimator that minimizes some sample criterion function $\theta\mapsto\hat {\mathsf Q}_1(\theta)$.\footnote{We use the subscript “1” to indicate that the sample criterion function onlyuses $D_1$.} An example of $\hat {\mathsf Q}_1$ is
where $\hat P_1(A_j|x)$ is an estimator of the conditional probability $P(A_j|x)$ using $D_1$, and $\hat{s}_{\theta,1}$ is an estimator of the standard error of $\hat P_1$.\footnote{These criterion functions are commonly used in practice chernozhukov2007estimation,chernozhukov_lee_rosen13. The supremum operation over $j$ (or $x$) can be replaced with sum (or integral). See Andrews:2013aa,chernozhukov_lee_rosen13.} Another possibility is to use the (negative) log-likelihood function:
It uses a log-likelihood $p_\theta(\cdot|\cdot;\hat p_n)$, which is the Kullback-Leibler divergence projection of a nonparametric estimator of the conditional choice probability to $\mathfrak{q}_{\theta,x}$ Kaido:2023aa. Our simulation results suggest both estimators perform similarly.
Next, we find a positive density $p(\cdot|x)$ compatible with the unrestricted estimator, by solving the following linear feasibility problem:
Any solution can be used as long as $p(y|x)>0$ for all $y\in\mathcal{Y}$. We then set $q_{\hat\theta_1}=p$. One may view $q_{\hat\theta_1}$ as a representative density in the unrestricted model.
Below, we condition on $D_1$, assume $q_{\hat\theta_1}$ is already computed, and treat it as fixed. Given $q_{\hat\theta_1}$, consider testing whether the data in $D_0$ are compatible with the null hypothesis or $q_{\hat\theta_1}$. If there is a parametric model $\{q_\theta,\theta\in\Theta_0\}$ such that $q_\theta\in\mathfrak{q}_{\theta,x}, \theta\in\Theta_0$, it is natural to base our test on:
However, naively selecting a parametric model can be problematic. For each $\theta$, some density in $\mathfrak{q}_{\theta,x}$ may be easier to distinguish from $q_{\hat\theta_1}$ than others. Using such densities to compute the denominator of (ref) could cause over-rejection if the true density is close to $q_{\hat\theta_1}$.
To address this issue, we define the least favorable pair (LFP)-based parametric model. For this, define the Kullback-Leibler (KL) divergence by
where $S_x=\{y\in\mathcal{Y}:f(y|x)>0\}$.
We call $q_\theta$ the LFP-based density. The theory of minimax tests ensures that $q_\theta(\cdot|x)$ is the least-favorable density for distinguishing $\mathfrak{q}_{\theta,x}$ from $q_{\hat\theta_1}(\cdot|x)$ (see Proposition (ref)), suitable for constructing an LR statistic and achieving size control. A practical aspect of this construction is that one can compute $q_\theta(\cdot|x)$ by solving the following convex program:
The program (ref) imposes the sharp identifying restrictions as constraints. As such, they retain all restrictions from the model. We show that $q_\theta$ can be derived analytically in leading examples. A companion Python program can also be used to compute $q_\theta$ numerically.
Now we define our tailor-made likelihood function and the restricted maximum likelihood estimator as follows:\footnote{To compute $T_n$, we only need $q_\theta$ to be defined over $\Theta_0\cup \{\hat\theta_1\}$, which is a non-random domain conditional on $D_1$. The LFP-based parametric model is defined similarly for $T_n^{\text{swap}}.$}
We make the following assumption on sampling.
We assume $(X_i,U_i)$ are independently and identically distributed across $i$.\footnote{Allowing heterogeneity of $F_{U|X}$ across $i$ is straightforward but requires additional notation.} We also assume $Y_i$ is independently distributed but allow its law to be heterogeneous across $i$. Allowing $Y_i$'s (cross-sectional) dependence is also possible at the expense of modifying $\hat\theta_1$. We discuss this extension in Appendix (ref).
For each $\theta\in\Theta$, let
Let $\mathcal{P}^n_0\equiv\{P^n\in\mathcal{P}^n_\theta:\theta\in\Theta_0\}$ be the set of data generating processes (DGPs) compatible with $H_0$. The following theorem establishes the universal validity of the proposed test and its robustness.
The theorem ensures that the test's finite-sample size is at most $\alpha$ under any distribution $P^n$ compatible with the null hypothesis regardless of how selections operate across $i$. The result does not require any regularity conditions beyond Assumption (ref). The proof is in Appendix (ref). It uses the properties of the least-favorable pair and Markov's inequality interpreted as a Chernoff-style bound for log likelihood-ratio as in Wasserman:2020aa. For our result, the construction of the null model $\{q_\theta,\theta\in\Theta_0\}$ is the key, which is possible because $\mathfrak{q}_{\theta,x}$ is characterized by a containment functional. We elaborate on this theoretical background in Section (ref).
Let us compare our test with the standard LR test. They are based on the idea of comparing the likelihood values with and without restrictions imposed by $H_0$. The standard LR test's rejection rule is $2\ln\big(\mathcal L_{full}(\hat\theta_1)/\mathcal L_{full}(\hat\theta_0)\big)> c_{df,\alpha}$, where $\mathcal L_{full}$ is the full-sample likelihood, and the critical value $c_{df,\alpha}$ is the $1-\alpha$ quantile of a $\chi^2$-distribution, which increases as the degrees of freedom ($df$) increases. In contrast, a split-sample version of the proposed test uses $2\ln\left(\frac{\mathcal L_0(\hat\theta_1)}{\mathcal L_0(\hat\theta_0)}\right)> c_\alpha$ with $c_\alpha=2\ln(1/\alpha)$. This critical value does not increase with the degrees of freedom. This difference arises because we use the out-of-sample likelihood $\mathcal L_0(\hat\theta_1)$ whose behavior differs from the in-sample likelihood $\mathcal L_{full}(\hat\theta_1)$.\footnote{See Wasserman:2020aa for further discussion on this point.}
Confidence regions for functions of $\theta$ can be constructed in a simple manner. Let $\varphi:\Theta\to \mathbb R^{d_\varphi}$. Examples of $\varphi(\theta)$ are subvectors of $\theta$ and counterfactual objects as we discuss below using examples.\footnote{One can also construct a confidence region for the entire parameter by letting $\varphi$ be the identity map.} For each $\varphi^*\in\mathbb R^{d_\varphi}$, let $\Theta_0(\varphi^*)\equiv\{\theta\in\Theta:\varphi(\theta)=\varphi^*\}$ and $\Theta_1(\varphi^*)\equiv\{\theta\in\Theta:\varphi(\theta)\ne\varphi^*\}$. Let $T_n(\varphi^*)$ be defined as in (ref) with $\Theta_0=\Theta_0(\varphi^*)$, and let $S_n(\varphi^*)$ be defined similarly. Define
For each $P^n$, let $\mathcal H_{P^n}[\varphi]\equiv\{\varphi^*:\varphi(\theta)=\varphi^*, P^n\in\mathcal{P}^n_\theta,\text{ for some }\theta\in\Theta\}$ be the sharp identification region for $\varphi(\theta)$. Let $\mathcal F^n\equiv \big\{(\varphi^*,P^n):\varphi(\theta)=\varphi^*,~P^n\in\mathcal{P}^n_\theta,\text{ for some }\theta\in\Theta\big\}$ be the collection of data generating processes compatible with the model restrictions. Then, the following result holds.
The coverage statement can also be written as $\inf_{P^n\in \mathcal P^n_\theta,\theta\in\Theta}\inf_{\varphi^*\in \mathcal H_{P^n}[\varphi]}P^n(\varphi^*\in CS_n)\ge 1-\alpha$. Therefore, the corollary ensures that $CS_n$ covers elements of the sharp identification region uniformly across $P^n$.
We illustrate the main result by revisiting Example (ref).
\setcounter{example}{0}
One can use Proposition (ref) to test various hypotheses on $\theta$. The presence of strategic interaction effects can be examined by testing
For this hypothesis, we let $\Theta_0=\{\theta=(\beta',\delta')':\beta=0\}.$\footnote{Further simplification of Step 2 (in Section (ref)) is possible for this null hypothesis because $\eta_2(\theta;x)=\eta_3(\theta;x)$ if $\beta=0$ due to the completeness of the model under $H_0$. Hence, the LFP-based likelihood satisfies $q_{\theta}(y|x)=F_\theta(S_{\{y\}|x;\theta}|x)$ for all $y$.} One can also construct confidence intervals for counterfactual probabilities. Let
be the potential entry decision by player $j$ when the covariates and the other player's action are set to $(x^{(j)},y^{(-j)})$. The counterfactual entry probability is
where $F_\theta$ is the marginal distribution of $U$. Letting $\Theta_0(\varphi^*)=\{\theta\in\Theta:\varphi(\theta)=\varphi^*\}$, one can construct a confidence interval as in (ref). Corollary (ref) ensures the finite sample coverage of this interval.\footnote{Alternatively, one may be interested in a counterfactual equilibrium outcome $Y(x)$, which would realize when we set $X$ to $x$. Due to the incompleteness, the counterfactual equilibrium probability $P(Y(x)=y)$ is not unique, but its bounds are functions of $\theta$. For example, the sharp upper bound on $P(Y(x)=y)$ is $\varphi(\theta) =F_\theta(G(U|x;\theta)\cap \{(1,0)\}\ne\emptyset)=F_\theta(S_{\{(1,0)\}|x;\theta})+F_\theta(S_{\{(0,1),(1,0)\}|x;\theta}).$} Various other functionals can also be handled. In related work, han2024setvalued apply this paper's procedure to construct confidence intervals for average structural functions, average treatment effects, and the counterfactual probability of an individual's switching behavior.
This section outlines the machinery behind Theorem (ref) and how it is related to existing results.
Wasserman:2020aa consider a probabilistic model $\{P_\theta,\theta\in\Theta\}$, where each $P_\theta$ is a probability distribution over a measurable space $\mathcal Z$. Given a sample $Z_i\in \mathcal Z,i=1,\dots, n$, they construct split and cross-fit LR statistics with $\mathcal L_0(\theta)=\prod_{i\in D_0}p_{\theta}(Z_i)$, where $p_\theta=dP_\theta/d\mu$ and show the their universal validity.\footnote{To be precise, they apply Markov's inequality to $T_n$, which can be viewed as bounding the tail probability using the moment generating function (MGF) of the log-likelihood. The key is to use a simple yet nontrivial result $E_{P_\theta}[e^{t\ln T_n}]\le 1$ for $t=1$. They call this approach “poor man's Chernoff bound”. See their discussion on page 16882.} In their setting, the likelihood $\mathcal L_0(\cdot)$ represents the conditional distribution of the observations $\{Z_i,i\in D_0\}$ over subsample $D_0$, given $D_1$. In our setting, the form of this distribution is unknown. Hence, instead of applying their argument directly, we construct $\mathcal L_0$ using the theory of minimax tests.\footnote{See Lehmann:2006aa (Ch. 8) for a general treatment of the topic.}
Let $P$ be the joint distribution of $Z=(Y,X)\in\mathcal{Z}$, $P_{Y|X}$ be the conditional distribution of $Y|X$, and $P_X$ be the marginal distribution of $X$. For each $\theta\in\Theta$, the set of permissible distributions of $Z$ is $\mathcal{P}_{\theta}=\{P\in \Delta(\mathcal{Z}):P_{Y|X}(\cdot|x)\in \mathcal{P}_{\theta,x},x\in \mathcal{X}\}$.
For $\theta_0, \theta_1 \in \Theta$ such that $\mathcal P_{\theta_0}$ and $\mathcal P_{\theta_1}$ are disjoint, consider testing $H_0: \theta=\theta_0$, against $H_1:\theta=\theta_1$. For any test $\phi: \mathcal{Z} \mapsto [0,1]$, its rejection probability under $P$ is $E_{P}[\phi(Z)]=\int \phi(z)dP.$ Let $\pi_{\theta}(\phi)\equiv\inf_{P \in \mathcal{P}_{\theta}}E_P [\phi(Z)]$ be the power guarantee of $\phi$ at $\theta$. This is the power value certain to be obtained regardless of the unknown selection mechanism. A test $\phi$ is a level-$\alpha$ minimax test if it satisfies the following conditions:
and
That is, the test maximizes the power guarantee subject to the uniform size control constraint (over $\mathcal P_{\theta_0}$).
Define the lower envelope of $\mathcal{P}_\theta$ by
Define the correspondence $\Gamma(x,u;\theta)\equiv\{(y,x)\in\mathcal{Z}: y\in G(u|x;\theta)\}$. By Choquet's theorem Choquet1954,philippe1999decision,Molchanov:2006aa, we may represent $\nu_\theta$ using the distribution of the random set $\Gamma(X,U;\theta)$:
The set function $\nu_\theta$ belongs to a class of two-monotone capacities whose properties have proven powerful for conducting robust inference Huber:1981aa.\footnote{See Appendix (ref). $\nu_\theta$ is called a belief function or a totally monotone capacity. The total monotonicity of $\nu_\theta$ follows from philippe1999decision (Theorem 3). The foundations of belief functions are given by dempster1967upper and shafer1982belief. See Gul:2014aa and epstein2015exchangeable for the axiomatic foundations of the use of belief functions in incomplete models.} For this class, the rejection region of a minimax test takes the form $\{z:\Lambda(z)>t\}$ for a measurable function $\Lambda$. Furthermore, we may obtain the following result Huber:1973aa,kaido2019robust.
Heuristically, $Q_0$ is the probability distribution consistent with the null parameter value, which makes size control most difficult. $Q_1$ is the distribution consistent with the alternative parameter value, which is least favorable for maximizing power. The LR-test based on this pair is minimax optimal.\footnote{Proposition (ref) is a special case of Theorem 3.1 in kaido2019robust. They provide explicit forms of the minimax tests for several examples.}
We use Proposition (ref) in the proof of Theorem (ref) (see Appendix (ref)). The following is a sketch of the proof arguments. Let $\theta_1\in \Theta$ and let $Q_{\theta_1}\in\mathcal{P}_{\theta_1}$. For each $\theta\in\Theta_0$, consider distinguishing $\mathcal{P}_{\theta_1}$ against a singleton set $\{Q_{\theta_1}\}$ and form the LFP $(Q_{\theta},Q_{\theta_1})$. Repeating this exercise across $\theta\in\Theta_0$, we form a null parametric model $\{q_{\theta},\theta\in\Theta_0\}$ with $q_{\theta}=dQ_{\theta}/d\mu$. Let $T_n(\theta_1)=\mathcal L_0(\theta_1)/\mathcal L_0(\hat\theta_0)$. Suppose $P^n\in \mathcal{P}^n_{\theta}$ for some $\theta\in\Theta_0$. Then,
The first inequality is due to Markov's inequality, and the second inequality is due to $\mathcal L_0(\hat\theta_0)\ge \mathcal L_0(\theta)$ by $\hat\theta_0$ being the RMLE. We show that the supremum on the right-hand side of (ref) is attained by the product of the least-favoarable distribution $Q_\theta$ at $\theta$ and use this result to show $ \sup_{\tilde P^n\in \mathcal P^n_{\theta}}E_{\tilde P^n}\Big[\frac{\mathcal L_0(\theta')}{\mathcal L_0(\theta)}\Big]\le 1$. This ensures that the size of the test is at most $\alpha$. The argument above fixed $\theta_1$, but an estimator $\hat\theta_1$ is used in practice. In the proof of Theorem (ref), we use a conditioning argument to handle the randomness of $\hat\theta_1$.
We examine the performance of the proposed test through simulations. First, we use the two-player entry game example with the following payoff:
We then test $H_0:\theta^{(j)}=0,j=1,2$ against $H_1:\theta^{(j)}<0$ for some $j$. As discussed in Example (ref), the model is complete under the null hypothesis, which determines $\mathcal L_0$. Hence, one only needs to determine $\hat\theta_1$. We consider two options. Both are extreme estimators. The first one is a minimizer of a sample criterion function based on the sample analog of the sharp identifying restrictions as in (ref). We call this estimator a moment-based estimator. The second one maximizes the information-based objective function as in (ref). We call this estimator an MLE.
We set the sample sizes to $n\in\{50,100,200\}$ and calculate the rejection probability of the tests at the alternatives with $\theta^{(j)}=-h$ with $h\ge 0$. Under each alternative, the outcome $Y_i=(1,0)$ is selected with probability 0.5 whenever the model admits multiple equilibria. Table (ref) reports the size and power of the two cross-fit tests. The two tests have similar power profiles, suggesting the choice of the initial estimator $\hat\theta_1$ does not seem to matter, at least for this example. Overall, the proposed tests' rejection probabilities under $H_0$ are nearly 0. However, they exhibit considerable power even in small samples as $h$ deviates from 0. They could detect mild strategic interaction effects (e.g., $h=0.5$, which is a half standard-deviation unit of $u^{(j)}$) with high rejection probabilities, even in small samples.
In the second experiment, we use the specification of payoff functions in (ref). For each $j$, $X^{(j)}$ is a covariate that takes $K=5$ discrete values $\{-2,-1,0,1,2\}$, and the two covariates are generated independently. When multiple equilibria are predicted, one of them gets selected with a probability of 0.5. We test $H_0:\delta^{(j)}=0,j=1,2$ against $H_1:\delta^{(j)}\ne 0$ for some $j$. A notable difference from the previous specification is that the model is incomplete under the null hypothesis, and $\theta$ contains nuisance components (i.e., $\beta^{(j)},j=1,2$ for this design). We evaluate the power of the test against alternatives with $\delta^{(j)}=h,h>0$ for samples of size $n\in\{50,100,200,300\}$. For this experiment, we use the moment-based estimator as $\hat\theta_1$. Table (ref) summarizes the result. As in the previous design, the test's size is nearly 0, but it has meaningful power against alternatives, even in small samples. The test also exhibits monotonically increasing power curves. Hence, it again shows promise in detecting violations of the restrictions.
In the third experiment, we compare our test to an existing test of BCS using large samples ($n=5000,7500$).\footnote{With covariates taking 25 different values with equal probabilities, the average sample size in each bin would become extremely small if we use a small sample (e.g., $n=100$). This feature was not an issue for the cross-fit LR test. However, it caused computational issues for implementing BCS's BCS test. Instead of adding modifications not considered in their original paper, we work in an environment in which their procedure is reliable. We found that their test was computationally reliable for sample sizes $n=5000,7500$.} We call their procedure a moment-based test. Figure (ref) reports the power curves of the two tests. Both tests show monotonically increasing power curves, and neither of them is uniformly dominant. This result shows that the cross-fit LR test can compete with a well-established test in terms of power. Table (ref) reports the computation time required to implement the tests.\footnote{They were computed using Boston University's Shared Computing Cluster (SCC) nodes equipped with 2.6 GHz Intel Xeon processor (E5-2650v2) and 128GB memory (per node). } BCS's BCS test uses a bootstrap critical value. To mimic a realistic scenario, we parallelized their bootstrap replications ($B=500$) across multiple processors. The table shows that the cross-fit LR test takes about 14 seconds (without any parallelization), which is significantly below the computation time required for the moment-based test with parallelization.
In sum, the simulation results show that the cross-fit LR test has considerable power, even in small samples. In large samples for which existing tests are applicable, the proposed test has power properties comparable to a well-established test.
This paper develops a novel likelihood-based test and confidence sets for incomplete models. They apply to a wide range of discrete choice models involving set-valued predictions. Yet, they are simple to implement. To retain simplicity, this paper uses simple two-fold cross-fitting. An avenue for further research is to examine whether alternative sample-splitting schemes can improve the statistical properties and replicability of the proposed method. For the latter, Ritzwoller:2024aa recently proposed a way to control an error rate due to sample splitting by sequentially aggregating statistics, which is a fruitful direction. Our simulation results suggest that the proposed test's power is comparable to that of a well-established moment-based test in large samples. It is worthwhile to study the proposed test's asymptotic power properties further.