EconBase
← Back to paper

Information Based Inference in Models with Set-Valued Predictions and Misspecification

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

108,628 characters · 13 sections · 103 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Information Based Inference in Models with Set-Valued Predictions and Misspecification

frontmatter\runtitle{Information Based Inference with Misspecification} \begin{aug} \address[id=add1]{ \orgdiv{Department of Economics}, \orgname{Boston University}} \address[id=add2]{ \orgdiv{Department of Economics}, \orgname{Cornell University}} \end{aug} \support{We thank anonymous referees, Xiaoxia Shi, and audiences at Bonn/Mannheim, Bristol/Warwick, BU, CEMFI, Chicago, Columbia, Cornell, FGV-EESP, JHU, Michigan, Nebraska, NYU, Queen's Mary, Tokyo, Toulouse, UCD, UCL, UCLA, UCSB, UPF, USC, Yale, Wisconsin, ESAM21, AMES23, ESWC25, Workshop on Econometrics and Models of Strategic Interaction, and Chamberlain Seminar, for comments. U. Byambadalai, S. Chen, Q. Han, L. Hoderlein, Yan Liu, Yiqi Liu, P. Power, R. Strong, Y. Sun provided excellent research assistance. We gratefully acknowledge financial support from NSF grants SES-2018498 (Kaido) and 1824375 (Molinari).} \begin{abstract} This paper proposes an information-based inference method for partially identified parameters in incomplete models that is valid both when the model is correctly specified and when it is misspecified. Key features of the method are: (i) it is based on minimizing a suitably defined Kullback-Leibler information criterion that accounts for incompleteness of the model and delivers a non-empty pseudo-true set; (ii) it is computationally tractable; (iii) its implementation is the same for both correctly and incorrectly specified models; (iv) it exploits all information provided by variation in discrete and continuous covariates; (v) it relies on Rao's score statistic, which is shown to be asymptotically pivotal. \end{abstract} \begin{keyword} \kwd{Misspecification} \kwd{Partial Identification} \kwd{Rao's score statistic} \end{keyword}

Introduction

Applied research is rarely based on the empirical evidence alone: exogeneity assumptions, behavioral restrictions, distributional and functional form specifications, etc., are routinely imposed to approximate features of complex social and economic phenomena. Early on, koo:rei50 highlighted the importance of imposing restrictions based on prior knowledge of the phenomenon under analysis and some criteria of simplicity, but argued against choosing these restrictions primarily for the purpose of identifiability of the structure the researcher happens to be interested in. Yet, even after embracing koo:rei50's perspective, concerns for misspecification remain.

To answer these concerns, we propose a novel information theoretic inference method in the spirit of White1982 that is asymptotically valid both when the model is correctly or incorrectly specified and point or partially identified. The latter case often results when, based on prior knowledge, a researcher is willing to impose modeling restrictions on some features of the model, but remains agnostic about others. Once the researcher has taken a stand on the model, we provide inference methods for its features that are robust to misspecification. Our method is easy to implement and addresses three challenges that arise in misspecified partially identified models: (i) the set of observationally equivalent parameters may be spuriously tight or empty; (ii) confidence sets constructed assuming correct model specification may (severely) undercover; (iii) their tightness may be misinterpreted as highly informative data.\footnote{Counterpart challenges arise in point identified models, with various solutions put forward since at least White1982 han:lee21.} The method applies to a specific but wide class of partially identified models that predict a set of values for the endogeneous variables ($Y$) given the exogenous observed and unobserved ones ($X$ and $U$, respectively), yielding a set of conditional distributions for $Y|X$. Many examples belong to this class, including: games with multiple equilibria; discrete choice models with either interval data on covariates, counterfactual choice sets, endogeneous explanatory variables, or unobserved heterogeneity in choice sets; dynamic discrete choice models; network formation models; and auctions and school choice models under weak assumptions on behavior MolinariHOE.

We adapt the textbook method for (point identified) models that predict a singleton conditional distribution for $Y|X$, to models that predict a set of distributions for $Y|X$, through three steps that jointly yield our main innovations. In step one, leveraging a result in art83, we characterize the exact set of model predicted distributions consistent with all maintained assumptions, and show that with discrete $Y$ it is a finite dimensional convex polytope.\footnote{Using only a subset of model implications may yield misleading conclusions ked:li:mou21,MolinariHOE,ber:mol:mol11; nonetheless, even in this case our method remains applicable.} Doing so greatly simplifies our next steps as it allows us to bypass the need to model all possible distribution of $Y|X$ through the use of probability mixtures with infinite dimensional nuisance mixing functions. In step two, we define a never-empty pseudo-true set, denoted $\Theta^*$, for the parameter vector $\theta$ characterizing the model. This is the collection of minimizers of a Kullback-Leibler (KL) information criterion measuring the divergence of the set of model-predicted distributions from the distribution of the observed data. The set $\Theta^*$ shrinks to the pseudo-true parameter vector in White1982 if the modeling assumptions are augmented so that the model predicts a single distribution. When the model predicts multiple distributions, $\Theta^*$ collects the parameter values that minimize the researcher's ignorance about the true structure Akaike1973,White1982, recognizing that this ignorance extends to the selection mechanism that picks an element from the set of model predictions. As in the point identified case, one may wonder to what extent a pseudo-true set is of substantive interest. In our view, models are only approximations to the true data generating process (DGP) and hence pseudo-true sets are often what researchers estimate in practice; hence, inference methods robust to misspecification are needed.

To this end, in the third step we obtain a profiled likelihood function by projecting, with respect to the KL divergence measure, the distribution of the observed data on the set of model implied distributions. A key advantage, yielded by steps one and two, is that this projection is carried out through a computationally simple convex program, which in our leading examples with discrete outcomes features a strictly convex objective and linear constraints. As in the textbook case, the pseudo-true set equals the collection of maximizers of the profiled likelihood function. We next derive a novel score representation for this function, based on $d_\theta$ estimating (score) equations, with $d_\theta$ the number of model's parameters. These equations depend on the conditional distribution of $Y|X$, which is unknown and needs to be estimated nonparametrically. We leverage classic results in the semiparametric inference literature, specifically Newey:1994aa, to establish that an orthogonality property holds. Provided the convergence rate of the nonparametric conditional density estimator is sufficiently fast ($o_p(n^{-1/4})$), this implies that the limit distribution of the averaged score function is insensitive to estimation of the distribution of the data. We use this result to construct a Rao's score statistic with asymptotically pivotal limit distribution $\chi^2_{d_\theta}$, which we use to test the hypothesis that a candidate parameter vector belongs to $\Theta^*$. We invert the test to construct a confidence set and show that it is robust to misspecification: it covers each element of $\Theta^*$ with asymptotic probability at least equal to the nominal level $1-\alpha$, uniformly over a large class of DGPs. While the size of $\Theta^*$ may or may not be impacted by the extent to which the model is misspecified (see Section (ref)), relative to $\Theta^*$'s size the volume of the confidence set depends only on the sampling variability of the score statistic.

\noindentRelated Literature. Chen:2011aa,Chen_2018 put forward inference methods that asymptotically cover the $\arg\max$ of a profiled likelihood function, which by textbook arguments is related to the $\arg\min$ of the KL divergence that we focus on. However, their method relies on infinite dimensional nuisance functions to represent each possible distribution of $Y|X$, and its validity is established for correctly specified models only. Our method completely bypasses the use of infinite dimensional nuisance functions, yielding computational advantages and allowing us to consider broader classes of incomplete models.\footnote{See the discussion following (ref) for further comparisons. For a dynamic model where misspecification results when subjective beliefs deviate from rational expectations, che:han:han21 show that a confidence set built using the results in Chen_2018 is asymptotically valid.}

Only a few recent papers put forth tools for construction of confidence sets that are valid in the presence of misspecification, covering each element of a pseudo-true identified set with an asymptotic probability at least as large as a prespecified nominal level. and:kwo22 show that model misspecification can lead to spuriously tight confidence sets while statistical tests have low power at detecting misspecification. They propose a notion of pseudo true set, a specification test, and an asymptotically uniformly valid inference method for partially identified models defined by a finite number of unconditional moment inequalities. Their inference method aggregates violations of the sample moment conditions relaxed by the minimum amount that guarantees that at least one parameter vector in the parameter space satisfies them. Stoye20 studies interval identified scalar parameters with asymptotic normality of the estimators of the endpoints. He obtains a valid and never-empty confidence interval that is free of tuning parameters and simple to compute. In contrast, our method applies to models for conditional density functions of outcome variables given discrete and continuous covariates. Allowing for the latter is important: they are commonplace in practice and may yield substantial identifying information. While a fine discretization or the use of instrument functions and:shi13 may transform conditional moment inequalities in unconditional ones, doing so may incur computational costs associated with an increase in the number of moment inequalities proportional to the cardinality of the discretized support or the number of instrument functions. We bypass this problem by evaluating directly the contribution of each observation to the score function. In our simulations (Section (ref)), doing so reduces computational time by 228 to 3,564 times relative to and:shi13's method, depending on how fine a discretization one uses. Our pseudo-true set and inference method are insensitive to which inequalities one uses to characterize the sharp collection of model implied distributions for $Y|X$. In comparison, much of the related literature often requires moment selection either for computational tractability or as part of the inference procedure, which may substantially impact the population region that the researcher targets ked:li:mou21 and the properties of confidence sets and:shi13,bug:can:shi17,kai:mol:sto19.

\noindentOutline. Section (ref) introduces the class of models we study. Section (ref) provides the notion of pseudo-true set and derives the misspecification robust inference method. Section (ref) discusses computational aspects of the method. Section (ref) provides an empirical illustration revisiting the analysis in kli:tam16 and Section (ref) Monte Carlo evidence on the size and power of the test. Section (ref) concludes. Appendix (ref) provides proofs of our main results. The Online Appendix includes auxiliary Lemmas and additional examples.

Notation and Motivating Example

Let $Y\in \mathcal{Y}\subseteq\mathbb{R}^{d_Y}$, $X\in \mathcal{X}\subseteq\mathbb{R}^{d_X}$ and $U\in \mathcal{U}\subseteq\mathbb{R}^{d_U}$ denote, respectively, observable endogenous and exogenous variables, and unobservable variables, with realizations $y,x,u$. Let $P_0\in\mathcal{P}(\mathcal{Y}\times\mathcal{X})$ denote the distribution of $(Y,X)$.\footnote{For a space $\mathcal{S}$ with Borel $\sigma$-algebra $\Sigma_\mathcal{S}$, $\mathcal{P}(\mathcal{S})$ denotes the set of all Borel probability measures on $(\mathcal{S},\Sigma_\mathcal{S})$.} Assume the conditional law $P_0(\cdot|x)$ is absolutely continuous with respect to a $\sigma$-finite measure $\mu$ on $\mathcal{Y}$. Let $p_{0,y|x}$ be the Radon-Nikodym derivative of $P_0(\cdot|x)$ with respect to $\mu$ and $p_{0}\equiv\{p_{0,y|x},x\in\mathcal{X}\}$. Throughout, for some $\underline{c}>0$ and all $x\in\mathcal{X},y\in\mathcal{Y}$, let $p_{0,y|x}(y|x)\ge\underline{c}$. We consider a framework where, based on prior knowledge, a researcher is willing to maintain some parametric structural restrictions on the joint behavior of $(Y,X,U)$, but is agnostic about other features of the model or whether they continue to hold after a policy intervention. We denote $\theta\in\Theta\subset\mathbb{R}^{d_\theta}$ the parameters characterizing the model. We let the structure associate with each $(u,x,\theta)$ a set of predicted outcomes through a closed-valued measurable correspondence $G:\mathcal{U}\times\mathcal{X}\times\Theta\mapsto\mathcal{Y}$.\footnote{Given a probability space $(\Omega,\mathfrak{F},\mathbf{P})$ and $\mathcal{C}$ the family of closed sets in $\mathbb{R}^d$, a correspondence $G:\Omega\mapsto\mathcal{C}$ is measurable if, for every compact set $K$ in $\mathbb{R}^d$, $G^{-1}(K) =\{ \omega \in \Omega :G(\omega)\cap K\neq \emptyset \} \in \mathfrak{F}$.} This nests as special case the textbook model with singleton predictions with $Y=g(U|X;\theta) ~\text{a.s.}$ for $g:\mathcal{U}\times\mathcal{X}\times\Theta\mapsto\mathcal{Y}$ a measurable function. We assume the family of distributions for the latent variables $U$ is known up to finite dimensional parameter vector that is part of $\theta$, and, omitting specific notation for subvectors of $\theta$, we denote $\{F_\theta:\theta\in\Theta\}$ the family of distributions for $U$ and assume that $F_\theta$ is independent of $X$.\footnote{This can easily be relaxed if the researcher is willing to specify the conditional distribution of $U|X$.} While assuming a parametric distribution for the entire vector $U$ simplifies our presentation, in some applications the researcher might be unwilling to take a stand on the distribution of some of its components. Our method is applicable in these cases too, though it requires a parametric distribution for a subvector of $U$, as we show in Online Appendix Example (ref) for the case of panel dynamic discrete choice models, where the researcher may lack prior knowledge of the distribution of the initial condition hon:tam06.

The next example, used throughout the paper to illustrate results, clarifies notation. More examples are provided in the Online Appendix and in MolinariHOE.

example[Static entry game] Consider a two player entry game as in tam03, with each player $i=1,2$ choosing to enter ($Y_i=1$) or stay out of the market ($Y_i=0$). Let $(X_1,X_2)$ and $(U_1,U_2)\simF_\theta$ be, respectively, observable and unobservable payoff shifters and player's payoffs be $ \pi_j=Y_j(X_j\beta_j+\delta_jY_{(3-j)}+U_j), j=1,2, $ with $\delta_1\le 0,\delta_2\le 0$ the interaction effects and $(\beta_1,\beta_2,\delta_1,\delta_2)$ part of $\theta$. Let each player enter the market if and only if $\pi_j\geq 0$. Given $\theta\in\Theta$ and $x\in\mathcal{X}$, the model has multiple pure strategy Nash equilibria (PSNE), depicted in Figure (ref) as a function of $(u_1,u_2)$. In our notation, the set of PSNE is the measurable correspondence $G(\cdot|x;\theta)$ ber:mol:mol11, with: \begingroup \allowdisplaybreaks { \begin{align} G(U|x;\theta)&=\{(0,0)\} if U\in S_{\{(0,0)\}|x;\theta} \equiv \{u:u_j<-x_j\beta_j, j=1,2\},\\ G(U|x;\theta)&=\{(1,1)\} if U\in S_{\{(1,1)\}|x;\theta} \equiv \{u:u_j\ge-x_j\beta_j-\delta_j, j=1,2\},\\ G(U|x;\theta)&=\{(1,0)\} if U\in S_{\{(1,0)\}|x;\theta} \equiv \{u: u_1\ge -x_1\beta_1 -\delta_1, u_2<-x_2\beta_2-\delta_2)\}\notag\\ &\bigcup\{u: -x_1\beta_1\le u_1< -x_1\beta_1-\delta_1, u_2<-x_2\beta_2\},\\ G(U|x;\theta)&=\{(0,1)\} if U\in S_{\{(0,1)\}|x;\theta} \equiv \{u:u_1< -x_1\beta_1, u_2\ge -x_2\beta_2\} \notag\\ &\bigcup \{u:-x_1\beta_1\le u_1< -x_1\beta_1-\delta_1, u_2\ge -x_2\beta_2-\delta_2\},\\ G(U|x;\theta)&=\{(1,0),(0,1)\} if U\in M_{x;\theta}\equiv \{u: -x_j\beta_j\le u_j<-x_j\beta_j-\delta_j, j=1,2\}. \end{align}} \endgroup \begin{figure} \begin{center} \begin{tikzpicture}[thick,scale=0.525, every node/.style={transform shape}] \draw[->,dashed,gray!30] (-4,0) -- (13,0) node[below] {$\color{gray}u_1$}; \draw[->,dashed,gray!30] (2,-2) -- (2,4) node[left] {$\color{gray}u_2$}; \draw[white] (2,2.5) -- (2,3.7); \draw[white] (8,0) -- (12.7,0); \draw (2,3.25) node {$G(U|x;\theta)=\{(0,1)\}$}; \draw (8.5,-1.5) node {$G(U|x;\theta)=\{(1,0)\}$}; \draw (5,1.25) node {$G(U|x;\theta)=\{(1,0),(0,1)\}$}; \draw[dashed,gray!30] (13,2.5) -- (2,2.5) node[left] {$\color{gray}-x_2\beta_2-\delta_2$}; \draw[dashed,gray!30] (8,4) -- (8,0) node[below] {$\color{gray}-x_1\beta_1-\delta_1$}; \draw[dashed,gray!30] (11,2.5) -- (2,2.5) node[left] {$\color{gray}-x_2\beta_2-\delta_2$}; \draw[gray!30] (2,0) -- (2,2.5); \draw[gray!30] (8,0) -- (8,2.5); \draw[gray!30] (2,0) -- (8,0); \draw[gray!30] (2,2.5) -- (8,2.5); \draw (10.5,3.25) node {$G(U|x;\theta)=\{(1,1)\}$}; \draw (-1,-1.5) node {$G(U|x;\theta)=\{(0,0)\}$}; \draw[gray!30] (2,0.2) node[left] {$\color{gray}-x_2\beta_2$}; \draw[gray!30] (2.2,0) node[below] {$\color{gray}-x_1\beta_1$}; \end{tikzpicture} \caption{{Stylized depiction of $G(\cdot|x;\theta)$ in Example (ref) with $\delta_1< 0,\delta_2< 0$. }} \end{center} \end{figure} If one assumes $\delta_1\cdot\delta_2=0$ (a “principal assumption” in the econometrics literature on simultaneous equation models with dummy endogeneous variables predating tam03, tam03), the region $M_{x;\theta}$ occurs with probability zero, and $G(U|x;\theta)$ reduces to a measurable function. Doing so, however, removes the simultaneity in player's actions that many applications aim to capture. One may consider completing the model through a selection mechanism that picks an outcome in the region of multiplicity, $M_{x;\theta}$. Doing so is also often undesirable, as various plausible selection mechanisms may lead to notably different conclusions about the nature of firms' competition ber92, and which firms enters when the market can sustain only one profitable entrant ($U\in M_{x;\theta}$) may be impacted by counterfactual interventions. The specification of $F_\theta$ can be made flexible through the use of mixtures, although our approach requires the mixture to be finite. $\square$

Information-Based Inference Robust to Misspecification

The Set of Model-Implied Density Functions

One can view $G(U|x;\theta)$ as the collection of its measurable selections mol:mol18, i.e., all random vectors $\tilde{Y}$ such that $\tildeY\inG(U|x;\theta)$ a.s. Each selection $\tildeY$ is a model predicted outcome. In order to obtain a set-valued analog of a likelihood model, one needs to be able to characterize the distribution of each of these predicted outcomes. To do so in a computationally feasible manner, we denote by $\mathcal{C}$ the collection of closed subsets of $\mathcal{Y}$ and define the law of $G(U|x;\theta)$ induced by the model's structure:

align[align omitted — 154 chars of source]

The containment functional $\nu_\theta(A|x)$ uniquely determines the distribution of $G(U|x;\theta)$ when it is evaluated at all $A\in\mathcal{C}$ mol:mol18.

Given $\theta\in\Theta$, $x\in\mathcal{X}$, and $\nu_\theta(\cdot|x)$, by art83 it is possible to characterize all distributions of measurable selections of $G(U|x;\theta)$ as the set

align[align omitted — 176 chars of source]

where $\mathcal{M}(\Sigma_Y,\mathcal{X})$ is the collection of laws of random variables supported on $\mathcal{Y}$ conditional on $X$. The characterization in (ref) is sharp, in the sense that, up to an ordered coupling mol:mol18, given $\tildeY\sim Q(\cdot|x)$, $\tildeY\inG(U|x;\theta)$ a.s. if and only if $Q(A|x)\ge\nu_\theta(A|x),$ for all $A\in\mathcal{C}$, $x$-a.s.

\noindentExample 1 (Continued). Let $Y_{x;\theta}(u)$ be the unique element of $G(u|x;\theta)$ if $u\notin M_{x;\theta}$ and (arbitrary) $Y_{x;\theta}(u)\equiv(0,0)$ if $u\in M_{x;\theta}$. mol:mol18 show that all measurable selections of $G(\cdot|x;\theta)$ in Example (ref) can be represented as

align[align omitted — 165 chars of source]

for a random variable $R\in\{0,1\}$ with any distribution in $\mathcal{P}(\{0,1\})$ and unrestricted dependence on $U$ given $X$. Each of these distributions is a selection mechanism cil:tam09 that assigns to $(1,0)$ and $(0,1)$ the probability that each is played given $X=x$ and $U\in M_{x;\theta}$. ber:mol:mol11 show that the distribution of each selection in (ref) belongs to $\mathop{\rm core}(\nu_\theta(\cdot|x))$, and that only those distributions do. The containment functional of $G(\cdot|x;\theta)$ satisfies: $\nu_\theta(\{(0,0)\}|x)=F_\theta(S_{\{(0,0)\}|x;\theta})$, $\nu_\theta(\{(0,1),(1,0)\}|x)=1-F_\theta(S_{\{(0,0)\}|x;\theta})-F_\theta(S_{\{(1,1)\}|x;\theta})$, and similarly for all $A\subseteq\mathcal{Y}=\{(0,0),(1,0),(0,1),(1,1)\}$, where for a given set $B\subset\mathcal{U}$, $F_\theta(B)=\int_\mathcal{U} \mathbf{1}(u\in B)dF_\theta(u)$ and the sets $S_{\{y\}|x;\theta},y\in\mathcal{Y}$ are defined in (ref)-(ref).$\square$

Assume that there are $\sigma$-finite measures $\mu$ on $(\mathcal{Y},\Sigma_Y)$ and $\xi$ on $(\mathcal{X},\Sigma_X)$, a product measure $\zeta\equiv\mu\times\xi$ on $(\mathcal{Y}\times\mathcal{X},\Sigma_Y\times\Sigma_X)$, and for all $\theta\in\Theta$, $x\in\mathcal{X}$, and $Q\in\mathop{\rm core}(\nu_\theta(\cdot|x))$, $Q\ll\mu$ White1982. Let the set of conditional densities associated with $\mathop{\rm core}(\nu_\theta(\cdot|x))$ be

align[align omitted — 261 chars of source]

\noindentExample 1 (Continued). Given $\theta\in\Theta$ and denoting $\Delta$ the unit simplex in $\mathbb{R}^4$, the set of all model predicted probability mass functions corresponding to selections of $G(\cdot|x;\theta)$ is

multline[multline omitted — 334 chars of source]

with $S_{\{(0,0)\}|x;\theta}$, $S_{\{(1,1)\}|x;\theta}$, $S_{\{(1,0)\}|x;\theta}$, $M_{x;\theta}$ defined in (ref), (ref), (ref), (ref). $\square$

Correct Specification, Misspecification, and Pseudo-True Set

Define a model $\mathfrak{Q}\equiv\left\{\mathfrak{q}_{\theta}:\theta\in\Theta\right\}$ as the collection of sets $\mathfrak{q}_{\theta}$ across $\theta\in\Theta$. We propose a generalization of the standard definition of correct specification for models with singleton predictions White1996 to models with set-valued predictions.

definition[Correctly Specified Model & Misspecified Model] A model is correctly specified if $p_{0}\in\mathfrak{q}_{\theta}$ for some $\mathfrak{q}_{\theta}\in\mathfrak{Q}\equiv\left\{\mathfrak{q}_{\vartheta}:\vartheta\in\Theta\right\}$, and misspecified otherwise.
remarkIn models that yield a singleton prediction $Y=g(U|X;\theta)$ a.s., with $g:\mathcal{U}\times\mathcal{X}\times\Theta\mapsto\mathcal{Y}$ a measurable function, there is a unique implied law for $g|X=x$: $Q_{\theta}(A|x)=\int_\mathcal{U} \mathbf{1}(g(u|x;\theta)\in A)dF_\theta(u),~\forall A\in\mathcal{C} $, with associated conditional density function $q_{\theta,y|x}=dQ_{\theta}(\cdot|x)/d\mu$ (compare with (ref) and (ref)). The model is defined as the collection of (singleton) $q_{\theta,y|x}$ across $\theta\in\Theta$ and $x\in\mathcal{X}$, $\mathcal{Q}=\left\{[q_{\theta,y|x},~x\in\mathcal{X}]:~\theta\in\Theta\right\}$. The model is correctly specified if $p_{0}=q_{\theta}$ for some $q_{\theta}\in\mathcal{Q}$, and misspecified otherwise.

Given density functions $f$ and $f'$, it has been common to measure their similarity through the Kullback-Leibler Information Criterion (KLIC) White1996. Here we extend the standard definition of KLIC to measure divergence from $f$ of a set of density functions $\mathfrak{f}$.

definition[KLIC for set of density functions] Let $(\Omega,\mathfrak{F},\zeta)$ be a measure space. Let $f:\Omega\mapsto\mathbb{R}_+$ be a measurable function satisfying $\int f d\zeta<\infty$ and $\int_S f\ln f d\zeta<\infty$ where $S=\{\omega\in\Omega:f(\omega)>0\}$. Let $\mathfrak{f}$ denote a set of measurable functions $f':\Omega\mapsto\mathbb{R}_+$ satisfying $\int_S f\ln f' d\zeta<\infty$ and $I(f||f')\equiv\int_{S}f\ln\tfrac{f}{f'}d\zeta$. The Kullback-Leibler divergence measure from $f$ of a set $\mathfrak{f}$ is $I(f||\mathfrak{f})\equiv\inf_{f'\in\mathfrak{f}}I(f||f').$

It follows from White1996 that when $\inf_{f'\in\mathfrak{f}}\int_S(f-f')d\zeta\ge 0$, $I(f||\mathfrak{f})=0$ if $f\in\mathfrak{f}$, and $I(f||\mathfrak{f})>0$ otherwise. As our approach is based on measuring divergence between conditional density functions, denoted $f(y|x)$ and $f'(y|x)$, in $I(f||\mathfrak{f})$ we replace the unconditional KLIC with the conditional one, $I(f||f')\equiv\int_{\mathcal{Y}\times\mathcal{X}}f(y,x)\ln\tfrac{f(y|x)}{f'(y|x)}d\zeta(y,x)$. Following White1982's treatment of point-identified models, we let the pseudo true set, $\Theta^*(p_{0})$, be the set of minimizers of the researcher's ignorance about the true structure. But here one is also ignorant about which selection from the model predicted set is closest to the data. Hence, minimization occurs with respect to both $\vartheta\in\Theta$, as in the textbook case, and $q_{}\in\mathfrak{q}_{\vartheta}$. If the model is correctly specified, $\Theta^*(p_{0})$ equals the sharp identification region, as in point-identified models where the pseudo-true value matches the data-generating one.

definitionThe pseudo-true identified set is given by \begin{align} \Theta^*(p_{0})&\equiv \Big\{\theta\in\Theta:I(p_{0}||\mathfrak{q}_{\theta})=\inf_{\vartheta\in\Theta}I(p_{0}||\mathfrak{q}_{\vartheta}) \Big\}. \end{align}

To understand the effect of minimizing KLIC with respect to $q_{}\in\mathfrak{q}_{\vartheta}$, note that

align[align omitted — 366 chars of source]

with $\mathfrak{q}_{\vartheta,x}$ defined in (ref). No unknown selection mechanisms are used in (ref) to formalize all ways in which a measurable selection could be picked from $G(\cdot|x;\theta)$ and all associated likelihoods obtained. Rather, (ref) relies on a convex program, with strictly convex objective and convex constraints (a finite dimensional convex program with linear constraints if $\mathcal{Y}$ is finite). It delivers the density function in $\mathfrak{q}_{\vartheta,x}$ closest with respect to KLIC to $p_{0,y|x}$,

align[align omitted — 175 chars of source]

which can be calculated analytically or numerically.\footnote{Existence and uniqueness of $q_{\vartheta,y|x}^*$ is guaranteed under mild conditions (see Online Appendix Lemma (ref)).} It can be interpreted as a profiled (quasi)-likelihood where a convex optimization program profiles out the selection mechanism, which is left completely unspecified and may arbitrarily depend on $(X,U,\vartheta)$. The support of $X$ is also unrestricted. In contrast, related likelihood-based inference methods rely on an infinite-dimensional parameter space to represent the selection mechanism that picks measurable selections from $G(\cdot|x;\theta)$ (as in (ref)) and profile it out via non-convex optimization programs with increasing number of (sieve) coefficients Chen:2011aa; or restrict the class of selection mechanisms by assuming that they do not depend on $U$ after conditioning on $X$ and that $X$ has finite support Chen_2018.\footnote{Chen:2011aa's inference method assumes correct model specification. Chen_2018 suggest that their method may remain valid for some misspecified separable models with discrete covariates.} Doing so may substantially increase computational burden or narrow the class of models allowed for. For example, in discrete choice models with unobserved heterogeneity in choice sets bar:cou:mol:tei21 it would rule out choice set formation based on sequential search or rational inattention (see Online Appendix Example (ref)).

Putting together (ref) and (ref) we obtain

align*[align* omitted — 158 chars of source]

Hence, the pseudo-true set $\Theta^*(p_{0})$ in (ref) is equal to the set of maximizers of $L(\vartheta)$, with

align[align omitted — 189 chars of source]

\noindentExample 1 (Continued). Given $\theta\in\Theta,x\in\mathcal{X}$, and $S_{\{(0,0)\}|x;\theta}$, $S_{\{(1,1)\}|x;\theta}$, $S_{\{(1,0)\}|x;\theta}$, $M_{x;\theta}$ as in (ref), (ref), (ref), (ref), let

align[align omitted — 291 chars of source]

In words, $\eta_1(\theta;x)$ is the probability allocated by the model to either $(1,0)$ or $(0,1)$ occurring as outcome of the game; $\eta_2(\theta;x)$ [$\eta_3(\theta;x)$] is the upper [lower] bound implied by the model on the probability that $(1,0)$ is the outcome of the game. Define the parameter sets:

align[align omitted — 562 chars of source]

Then the profiled likelihood is given by (see Proposition (ref) in the Online Appendix):

align[align omitted — 659 chars of source]

Intuitively, when $\theta\in\Theta_1(x,p_{0})$, $(1,0)$ can be assigned a share of $\eta_1(\theta;x)$ equal to the share of $p_{0,y|x}(\{(0,1),(1,0)\}|x)$ that $(1,0)$ has in the data. When $\theta\in\Theta_2(x,p_{0})$, that allocation yields a probability for $(1,0)$ larger than the model's upper bound $\eta_2(\theta;x)$, and the KL divergence is minimized setting $q_{\theta,y|x}^*((1,0)|x)=\eta_2(\theta;x)$. Similarly for $\theta\in\Theta_3(x,p_{0})$. In Section (ref) below we discuss the general topological properties of $\Theta^*$. $\square$

The Score Function

Here we characterize the score function associated with the singleton-valued likelihood function in (ref). Let $\mathcal{A}^{(*e)}(\theta;x)\subseteq \mathcal{C}$ denote a smallest core determining class in the sense of luo:pon:wan25, i.e., a subset of $\mathcal{C}$ of minimal cardinality such that $\mathop{\rm core}(\nu_\theta(\cdot|x))=\big\{Q\in\mathcal{M}(\Sigma_Y,\mathcal{X}):Q(A|x)\ge\nu_\theta(A|x),\forall A\in\mathcal{A}^{(*e)}(\theta;x)\big\}$.

assumption(a) $\mathcal{Y}$ is finite and $\mathcal{X}$ is compact. (b) $\nu_\theta(A|x)$ is continuously differentiable with respect to $\theta$ and $\nabla_\theta\nu_\theta(A|x)$ is square integrable for all $A\subset \mathcal{Y}$, $P_0-a.s.$ (c) $\Theta^*(p_{0})\subset \operatorname{int}\Theta$. (d) There exists a constant $c>0$ such that for all $\theta\in\Theta$ and $y\in\mathcal{Y}$, $q_{\theta,y|x}^*(y|x)>c$, $P_0-a.s.$ (e) There is a collection $\mathcal{A}_G(x)\subset 2^\mathcal{Y}$, that does not depend on $\theta$, such that $\operatorname{supp}(G(\cdot|x;\theta))\equiv \{A\subseteq \mathcal{Y}:F_\theta(G(U|x;\theta)=A)>0\}=\mathcal{A}_G(x)$ for all $\theta\in\Theta$, $P_0-a.s.$ (f) The set of (in)equalities in $\mathcal{A}^{(*e)}(\theta;x)$ active at $q_{\theta,y|x}^*$ are irredundant, in the sense that removing any of them strictly enlarges the active face of $\mathfrak{q}_{\theta,x}$ at $q_{\theta,y|x}^*$.

Assumption (ref)-(a) restricts attention to models with discrete outcomes and compact support for $X$. Part (b) is easily verified when $F_\theta$ is differentiable in $\theta$. Part (c) implies that first order conditions hold at all $\theta^*\in\Theta^*(p_{0})$. Part (d) bounds $q_{\theta,y|x}^*$ away from zero. Assumption (ref)-(e) requires the support of $G(\cdot|x,\theta)$ not to vary with $\theta\in\Theta$, though it can vary with $x\in\mathcal{X}$. It holds in several applications, including discrete choice models with unobserved heterogeneity in choice sets bar:cou:mol:tei21 and dynamic monopoly entry models ber:com22. Moreover, gu:rus:str25 show that with finite $\mathcal{Y}$, $\Theta$ can be finitely partitioned in $x$-dependent partitions, so that within each partition $\mathcal{A}_G(x)$ does not depend on $\theta$. luo:pon:wan25 show that under this condition, $\mathcal{A}^{(*e)}(\theta;x)=\mathcal{A}^{(*e)}(x)$ for all $\theta$. Lemma (ref) shows that when $q_{\theta,y|x}^*$ belongs to the relative interior of $\mathfrak{q}_{\theta,x}$, parts (a)-(e) of Assumption (ref) imply that the Lagrange multiplier vector for the convex program in (ref) is unique and hence $L(\theta|x)$ is differentiable. The lemma also shows that Assumption (ref)-(f), which requires that removing a binding inequality strictly enlarges the face of $\mathfrak{q}_{\theta,x}$ that $q_{\theta,y|x}^*$ belongs to (see Definition (ref)), ensures the same conclusion when $q_{\theta,y|x}^*$ belongs to the relative boundary of $\mathfrak{q}_{\theta,x}$. Conditions that imply uniqueness of the Lagrange multipliers can replace parts (e)(f). These two conditions are restrictive. For example, as stated they rule out that $\Theta$ includes both values of $\theta$ at which $G(\cdot|x;\theta)$ collapses to a function and values at which it is a non-singleton correspondence che:kai21. In Online Appendix (ref) we show that we can dispense with these conditions and work with the subdifferential of $L(\theta|x)$, but doing so makes our procedure less tractable. We verify all conditions for Example (ref) below and other examples in Online Appendix (ref).

theoremUnder Assumption (ref), (i) $L(\theta|x)$ is differentiable with respect to $\theta$ on $\mathrm{int}(\Theta)$, $P_0-a.s.$ (ii) There exists a function $s:\Theta\times \mathcal{Y}\times\mathcal{X}\times\Delta\to \mathbb{R}^{d_\theta}$, with $\Delta$ the unit-simplex in $\mathbb{R}^{|\mathcal{Y}|-1}$, such that $\mathbb{E}[\|s_{\theta}(Y|X;p_{0,y|x})\|^2]<\infty$, and \begin{align} \tfrac{\partial}{\partial\theta}L(\theta|x)&=\mathbb{E}[s_{\theta}(Y|X;p_{0,y|x})|X=x], \\ \mathbb{E}[s_{\theta}(Y|X;p_{0,y|x})]&=0, for all \theta\in\Theta^*(p_{0}). \end{align}

The score function depends on $p_{0,y|x}$. When $\theta\mapstoL(\theta|x)$ is concave, $\Theta^*(p_{0})$ equals the set of $\theta\in\Theta$ for which (ref) holds. When concavity does not hold, this set includes $\Theta^*(p_{0})$. As one of our goals is to avoid spuriously tight confidence sets, we view the benefit of an easy-to-implement method to outweigh the cost of a sometimes wider confidence set which asymptotically uniformly covers the set of $\theta$'s satisfying (ref). Our Monte Carlo results in Section (ref) show that, for the examples analyzed there, our procedure performs well relative to existing methods. The proof of Theorem (ref) is in Appendix (ref) and leverages results in Gauvin:1990uz to establish differentiability with respect to $\theta$ of

align[align omitted — 186 chars of source]

where $\mathfrak{q}_{\theta,x}$ is defined in (ref). The proof uses results in luo:pon:wan25 by which under Assumption (ref)-(e), the smallest collection of inequalities among the ones in (ref) that suffice to sharply characterize $\mathfrak{q}_{\theta,x}$ does not depend on $\theta$. As further discussed in Section (ref), Lemma (ref)-(ii) in the Online Appendix shows that $q_{\theta,y|x}^*$ and the value of $L(\theta|x)$, and hence its differentiability and the score function $s_{\theta}(y|x;p_{0,y|x})$, are insensitive to inclusion of additional inequalities from (ref) in the maximization problem in (ref). Hence, so are the pseudo-true set $\Theta^*(p_{})$ and our inference procedure. This is in contrast with much related literature, where moment selection is often required for computational tractability or for the inference procedure, but may have substantial implications both on the population region that the researcher targets ked:li:mou21 and on the properties of the inference procedure and:shi13,bug:can:shi17,kai:mol:sto19.

\noindentExample 1 (Continued). Assumption (ref)-(a) holds. Part (b) holds as long as $F_\theta(S_{\{y\}|x;\theta})$, $y\in\mathcal{Y}$, and $F_\theta(M_{x;\theta})$ are differentiable with respect to $\theta$ (e.g., for $(U_1,U_2)$ bivariate normal). Part (d) holds by compactness of $\mathcal{X}$ and $\Theta$, as one can find $c>0$ such that $\eta_j(\theta;x)\ge c$ and $\eta_1(\theta;x)-\eta_j(\theta;x)\ge c,j=2,3$, $F_\theta(S_{\{y\}|x;\theta})>c,~y\in\{(0,0),(1,1)\}$, for all $x\in\mathcal{X}$ and $\theta\in\Theta$. Part (e) holds if, e.g., $\delta_1,\delta_2<0$, whence $\mathcal{A}_G=\{\{(0,0)\},\{(0,1)\},\{(1,0)\},\{(1,1)\},$ $\{(1,0),(0,1)\}\}$ for all $\theta\in\Theta$. Part (f) holds when $\delta_1\cdot\delta_2\neq 0$: at most one inequality binds at $q_{\theta,y|x}^*$ and is irredundant. If one allows for $\delta_1\cdot\delta_2=0$, at those values the model is complete, the support of $G(\cdot|x,\theta)$ changes to $\{\{(0,0)\},\{(0,1)\},\{(1,0)\},\{(1,1)\}\}$, and two inequalities with identical gradient bind at $q_{\theta,y|x}^*$, violating Assumption (ref)-(e)(f); in that case, one can extend our approach using subdifferentials, as we show in Online Appendix (ref). In Proposition (ref) in the Online Appendix we show that:

align[align omitted — 1,071 chars of source]

In Section (ref) below we discuss how to accurately and rapidly compute the score numerically when analytic representations are not available.

Geometry of the Pseudo-True Set

We next discuss the topological properties of the pseudo-true set. By its definition in (ref), $\Theta^*$ can be viewed as the $\mathop{\rm arg\,min}$-set of an optimization problem indexed by $p_{}$:

align[align omitted — 287 chars of source]
theoremSuppose $\Theta$ is compact, Assumption (ref) holds, and there exists a neighborhood $V$ of $p_{0}$ and $M>0$ such that for all $p\in V$, $\sup_{\theta\in\Theta}\mathbb{E}[\|s_{\theta}(Y|X;p_{y|x})\|^2]\le M$. Then the mapping $p_{}\mapsto \Theta^*(p_{})$ is nonempty, compact-valued, and upper hemicontinuous.

We prove this theorem by showing that under its assumptions, $\phi(\cdot,\cdot)$ is jointly continuous in $(\theta,p_{})$ (Lemma (ref)), which along with compactness of $\Theta$ guarantees applicability of Berge's maximum theorem. The result establishes non-emptyness of $\Theta^*(p_{})$ and yields a first step towards characterizing how the geometry of the pseudo-true set changes as the extent of model misspecification changes. Given an observed density $p_{0}$, let $\{p_{\gamma}\}\in\Delta$, $\gamma\in\Gamma\subset\mathbb{R}$, be a net with corresponding pseudo-true values $\{\theta_\gamma\}$ such that: (i) $\theta_\gamma\in \Theta^*(p_{\gamma})$ for all $\gamma\in\Gamma$, (ii) $p_{\gamma}\to p_{0}$, and (iii) $\theta_\gamma\to\theta^*$ for some $\theta^*\in\Theta$. Then Theorem (ref) yields that $\theta^*\in \Theta^*(p_{0})$: the limits of such nets form elements of the pseudo true set at $p_{0}$. This property, which holds under weak conditions, rules out that $\Theta^*(p_{\gamma})$ is persistently larger than $\Theta^*(p_{0})$, as $p_{\gamma}\to p_{0}$. However, it does not preclude the possibility that $\Theta^*(p_{\gamma})$ shrinks from a set at $p_{0}$ (e.g., an arc in the entry game example) to a smaller set or a singleton at $p_{\gamma}$, no matter how close $p_{\gamma}$ is to $p_{0}$.

This possibility is instead precluded when $p_{\gamma}\mapsto\Theta^*(p_{\gamma})$ is lower hemicontinuous, yielding the second step characterizing how the geometry of $\Theta^*(p_{\gamma})$ depends on the extent of misspecification. It is well known that lower hemicontinuity at some $\gamma_0\in\Gamma$ is guaranteed under stronger and more challenging to verify conditions than upper hemicontinuity Rockafellar_Wets2005aBK. Equivalent statements include that for all $\theta\in\Theta$, the mapping $\gamma\mapsto\mathrm{dist}(\theta,\Theta^*(p_{\gamma}))$ is upper semicontinuous at $\gamma_0$; or that for every $\rho>0$ and $\epsilon>0$, there is a neighborhood $V$ of $\gamma_0$ such that $\Theta^*(p_{\gamma_0})\cap\rho\mathbb{B}\subset\Theta^*(p_{\gamma})+\epsilon\mathbb{B}$ for all $\gamma\in V\cap\Gamma$, with $\mathbb{B}$ the unit ball in $\mathbb{R}^{d_\theta}$ Rockafellar_Wets2005aBK. Sufficient conditions for lower hemicontinuity are given, e.g., in Rockafellar_Wets2005aBK and Aubin:1990aa.

Next, we provide an entry game example where we can transparently show that the correspondence $\gamma\mapsto\Theta^*(p_{\gamma})$ is both lower and upper hemicontinuous, guaranteeing that the geometric properties under correct specification of the sharp identification region are preserved as the model becomes misspecified. In particular, $\Theta^*(\gamma)$ does not abruptly become a singleton as $p_{\gamma}$ turns incompatible with the model (see Definition (ref)).

\noindentExample 1 (Specialized). Let $(U_1,U_2)\sim \mathrm{Uniform}[0,1]^2$, $\mathcal{Y}=\{(1,1),(0,1),(1,0)\}$. For $X^*\in\{0,1\}$ and $P_0(X^*=1)=1/2$, set $\pi_j=Y_j(\delta_{j,0}(X^*)Y_{3-j}+U_j)$, $\delta_{j,0}(X^*)=-0.5 + \theta_{j,0} X^*$, $\Theta=[-0.45,0]^2$, and $\theta_{j,0}<0$ for $j=1,2$. Let $y=(0,1)$ be always selected in the region of multiplicity. Misspecification occurs because $X^*$ is replaced by a binary proxy $X$ with $P_0(X^*=1|X=x)=\kappa(x,\gamma)\equiv(1-\gamma)x+\gamma(1-x)$. Expressing the observed density of $Y|X$, indexed by $\gamma$ and denoted $p_{\gamma}(Y|X)$, as a mixture of the density $p_{0}(Y|X^*=x)$, for $x=0,1$, we obtain $p_{\gamma}(y|X=x)=p_{0}(y|X^*=1)\kappa(x,\gamma)+0.25(1-\kappa(x,\gamma))$, for $y=(1,1),(1,0)$, and $p_{\gamma}((0,1)|X=x)=1- p_{\gamma}((1,1)|X=x)-p_{\gamma}((1,0)|X=x)$. At $\gamma=0$, $p_{0}(Y|X=x)=p_{0}(Y|X^*=x)$ for all $x\in\{0,1\}$ and the model is correctly specified. Regardless of the value of $\gamma$, denoting $\delta_{j,\theta}(x)=-0.5 + \theta_j x$ for some $\theta_j\in\Theta$, we have

multline*[multline* omitted — 224 chars of source]

Ignoring possible misspecification, one would state $\theta$'s sharp identification region as $\Theta_I(p_{\gamma})=\{\theta\in\Theta:p_{\gamma}(y|x)\in\mathfrak{q}_{\theta,y|x},x=0,1\}=\big\{(\vartheta-[0.5~0.5]^\top)\in\Theta:\vartheta_1\vartheta_2=p_{\gamma}((1,1)|1)$; $\vartheta_1\le 1-p_{\gamma}((0,1)|1);~\vartheta_2\le 1-p_{\gamma}((1,0)|1);~p_{\gamma}((1,1)|0)=0.25;~p_{\gamma}((1,0)|0)\in[0.25,0.5]\big\}$. For $\gamma=0$, $\Theta_I(p_{0})$ is an arc defined by $p_{0}(y|1)\in\mathfrak{q}_{\theta,y|1}$, with $p_{0}(y|0)\in\mathfrak{q}_{\theta,y|0}$ restricting $p_{0}(y|0)$ but not $\theta$. But for any $\gamma>0$, no matter how close to $0$, $\Theta_I(p_{\gamma})=\emptyset$, as one can verify that $p_{\gamma}((1,1)|0)<(1+\delta_{1,\theta}(0))(1+\delta_{2,\theta}(0))=0.25$ for all $\theta\in\Theta$.

In contrast, our pseudo-true set $\Theta^*(p_{\gamma})$ remains robust to misspecification and preserves its arc shape. After plugging the uniform distribution for $F_\theta$ into (ref)-(ref) to define $\eta_j(\theta;x)$ and the sets $\Theta_j(x,p_{\gamma})$, and into (ref)-(ref) to obtain $q_{\theta,y|x}^*$ and from that $s_{\theta}(y|x;p_{\gamma,y|x})$ for $y\in\mathcal{Y}$, one can verify that

multline[multline omitted — 561 chars of source]

Even when $\gamma>0$, the equality restriction in (ref) is analogous to $p_{0}((1,1)|1)=(1+\delta_{1,0}(1))(1+\delta_{2,0}(1))$. Along with the other two inequalities, it yields a region with the same shape as $\Theta_I(p_{\gamma})$ for $\gamma=0$. In Proposition (ref) in the Online Appendix we derive (ref) and prove that the mapping $\gamma\mapsto\Theta^*(p_{\gamma})$ is both upper and lower hemicontinuous. $\square$

Asymptotic Distribution of the Average Score and Rao's Test Statistic

Any $\theta^*\in\Theta^*(p_{0})$ satisfies the population first order condition in (ref). By the sample analog principle, we propose to estimate $\mathbb{E}[s_{\theta^*}(Y|X;p_{0,y|x})]$ through $\bar{s}_{\theta^*,n}(\hat p_{n,y|x})$, with

align[align omitted — 147 chars of source]

and $\hat p_{n,y|x}$ a nonparametric estimator of $p_{0,y|x}$, e.g., a cell mean estimator when $X$ has a discrete distribution or a sieve estimator when $X$ has a continuous distribution. Our core result consists of showing that $\sqrt{n}\bar{s}_{\theta^*}(\hat p_{n,y|x})$ has an asymptotically normal distribution, which is insensitive to estimation of $p_{0,y|x}$. We do so leveraging the literature on semiparametric estimation, in particular Newey:1994aa, to prove that $\mathbb{E}[s_{\theta^*}(Y,X;p_{y|x})]$ has an orthogonality property with respect to $p_{0,y|x}$. Here we provide high-level conditions under which our results attain. In Online Appendix (ref) we verify these conditions for the entry game example, both with discrete and continuous covariates. To state these conditions, for any $\theta\in\Theta$, let $m_\theta(x;p_{y|x})\equiv \mathbb{E}[s_{\theta}(Y|X;p_{y|x})|X=x]$. Let $\mathcal{H}$ be a parameter space to which $p_{0,y|x}$ belongs, with $\dim(\mathcal{H})=d_Y\times d_X<\infty$ if $X$ is finitely supported, and $\mathcal{H}$ infinite dimensional otherwise. Let $\|p-p'\|_{\mathcal{H}}$ be a pseudo-metric on $\mathcal{H}$ (e.g., the sup-norm $\|p\|_{\mathcal{H}}=\sup_{x\in\mathcal{X}}\sup_{y\in\mathcal{Y}}|p(y|x)|$). For any $p\in\mathcal{H}$ and $\theta\in\Theta$, let $\mathbb G_{n,\theta}(p)\equiv\sqrt{n}(\bar{s}_{\theta,n}(p)-\mathbb{E}[s_{\theta}(Y_i|X_i;p)])$.

assumptionFor each $\theta^*\in\Theta^*(p_{0})$, the pathwise derivative \begin{align*} D(\theta^*,p_{0,y|x})[p_{y|x}-p_{0,y|x}]=\lim_{\tau\to 0}\tfrac{\mathbb{E}[m_{\theta^*}(X,p_{0,y|x}+\tau(p_{y|x}-p_{0,y|x}))-m_{\theta^*}(X,p_{0,y|x})]}{\tau} \end{align*} exists in all directions $(p_{y|x}-p_{0,y|x})\in \mathcal{H}$. For any $\delta_n=o(1)$ and all $\|p_{y|x}-p_{0,y|x}\|_{\mathcal{H}}\le\delta_n$, \begin{align*} \big\|\mathbb{E}[m_{\theta^*}(X;p_{y|x})]-\mathbb{E}[m_{\theta^*}(X,p_{0,y|x})]-D(\theta^*,p_{0,y|x})[p_{y|x}-p_{0,y|x}]\big\|\le c\|p_{y|x}-p_{0,y|x}\|^2_{\mathcal H}. \end{align*}
assumption\begin{enumerate}[label=(\roman*)] • The data is a random sample $(Y_i,X_i)_{i=1}^n$ drawn from $P_0$. • $\hat p_{n,y|x}\in\mathcal H$ with probability approaching 1 and $\|\hat p_{n,y|x}-p_{0,y|x}\|_{\mathcal H}=o_P(n^{-1/4})$. • For each $\theta^*\in\Theta^*(p_{0})$, $\mathbb G_{n,\theta^*}(p_{0,y|x})\stackrel{d}{\to}N(0,\Sigma_{\theta^*})$, with $\Sigma_{\theta^*}\equiv\mathbb{E}[s_{\theta^*}(Y_i|X_i;p_{0})s_{\theta^*}(Y_i|X_i;p_{0})^\top]$ the population variance-covariance matrix of the score function. • For each $\theta^*\in\Theta^*(p_{0})$ and all sequences of positive numbers $\{\delta_n\}$ with $\delta_n=o(1)$, \begin{align*} \sup_{\|p_{y|x}-p_{0,y|x}\|_{\mathcal H}\le \delta_n}\big\|\mathbb G_{n,\theta^*}(p_{y|x})-\mathbb G_{n,\theta^*}(p_{0,y|x})\big\|=o_P(1). \end{align*} \end{enumerate}

Assumptions (ref) and (ref), in their use of $\|\cdot\|_{\mathcal H}$, refer to the same norm. In Assumption (ref), we follow Chen2003 and impose a smoothness condition with respect to $p_{y|x}$ on $\mathbb{E}[s_{\theta^*}(Y,X;p_{y|x})]$. Assumption (ref) (ref) is a standard random sampling condition (see eps:kai:seo16 for a discussion of inference under different assumptions). Assumption (ref) (ref) requires that the estimation error of the nuisance parameter $p_{0,y|x}$ vanishes fast enough. Assumption (ref) (ref) follows from the central limit theorem. Assumption (ref) (ref) is a stochastic equicontinuity condition with well known primitive conditions Vaart:1996wk. In Propositions (ref)-(ref) in the Online Appendix, we verify all these assumptions in the two players entry game Example (ref), both with discrete and continuous covariates. Under these assumptions, we obtain:

theoremSuppose Assumptions (ref), (ref), and (ref) hold. Then, for each $\theta^*\in\Theta^*(p_{0})$, \begin{align} \tfrac{1}{\sqrt n}\sum_{i=1}^n s_{\theta^*}(Y_i|X_i;\hat p_{n,y|x})=\tfrac{1}{\sqrt n}\sum_{i=1}^ns_{\theta^*}(Y_i|X_i;p_{0,y|x})+o_p(1)\stackrel{d}{\to}N(0,\Sigma_{\theta^*}). \end{align}

Armed with the result in Theorem (ref), we propose to use a Rao's score statistic to test at prespecified asymptotic level $\alpha\in(0,1)$ hypotheses of the form

align[align omitted — 129 chars of source]

and to obtain confidence sets by test inversion. The Rao-type test statistic takes the form

align[align omitted — 239 chars of source]

Given the score function, the test statistic in (ref) is easy to compute even when the covariates have a continuous distribution. The weight matrix $\tilde\Sigma_{n,\theta}$ is a consistent estimator of $\Sigma_{\theta}$ when $\Sigma_{\theta}$ is not nearly singular, and assures an asymptotically valid test when $\Sigma_{\theta}$ is nearly singular, as shown below. Let $\hat\Sigma_{n,\theta}=\tfrac{1}{n}\sum_{i=1}^n (s_{\theta}(Y_i|X_i;\hat p_{n,y|x})-\bar{s}_{\theta}(\hat p_{n,y|x}))(s_{\theta}(Y_i|X_i;\hat p_{n,y|x})-\bar{s}_{\theta}(\hat p_{n,y|x}))^\top$ be the sample analog estimator of $\Sigma_\theta$; $\hat\Xi_{n,\theta}=\hat\Psi_{n,\theta}^{-1/2}\hat\Sigma_{n,\theta}\hat\Psi_{n,\theta}^{-1/2}$ the correlation matrix associated with $\hat\Sigma_{n,\theta}$; $\hat\Psi_{n,\theta}=\rm{diag}(\hat\Sigma_{n,\theta})$; and $\varepsilon>0$ a regularization constant. We recommend using the estimator proposed in and:bar12, which introduces an adjustment insuring that the weight matrix is always nonsingular and equivariant to scale changes in the score function:

align[align omitted — 168 chars of source]
corollaryLet Assumptions (ref), (ref), and (ref) hold. Then, under $\mathbb{H}_0$ in (ref), for any $\theta^*\in\Theta^*(p_{0})$ such that $\min_{j=1,\dots,d_\theta}\{\rm{diag}(\Sigma_{\theta^*})\}_j>0$ and \begin{align} \hat\Sigma_{n,\theta^*}\stackrel{p}{\to}\Sigma_{\theta^*}, \end{align} (a) If $\Sigma_{\theta^*}$ is nonsingular, $T_n(\theta^*)\stackrel{d}{\to}\chi^2_{d_\theta}$; (b) Both for singular and nonsingular $\Sigma_{\theta^*}$, $\lim\sup_{n\to\infty}P(T_n(\theta^*)>c_{d_\theta,\alpha})\le \alpha$, with $c_{d_\theta,\alpha}$ the $1-\alpha$ quantile of the $\chi^2_{d_\theta}$ distribution.

Example 1 (Continued). The score function for Example 1 on p. (ref) implies that $\Theta^*(p_{\gamma})$ is characterized by a single equality restriction, and hence the rank of $\Sigma_{\theta^*}$ in this case is lower than $d_\theta$ and the critical value in Corollary (ref) is asymptotically conservative.$\square$

Corollary (ref) requires, in (ref), that the population covariance matrix can be consistently estimated. In the semiparametric literature with point identification, this is a standard requirement;\footnote{In the semiparametric literature, (ref) is imposed for $\hat\Sigma_{n,\hat{\theta}_n}$, with $\hat{\theta}_n$ a consistent estimator of a singleton $\theta^*$; in our use of it, $\hat\Sigma_{n,\theta^*},\Sigma_{\theta^*}$ are both evaluated at the same $\theta^*$, but the requirement applies to all $\theta^*\in\Theta^*(p_{0})$.} in the moment inequalities literature with partial identification, it is also common to assume that the covariance matrix of the moment functions can be consistently estimated for all $\theta$ in the identified set. The result in Corollary (ref) is valuable because it implies that no simulations are needed to compute the quantiles of the limiting distribution, and that the critical values used to test the hypothesis in (ref) and to construct the confidence set via test inversion are constant across candidates $\theta\in\Theta$. This is in contrast with much of the related literature, where the asymptotic distribution of the test statistic is nonpivotal and the critical values need to be recomputed for each $\theta$.\footnote{Nonpivotal asymptotic distributions appear, e.g., in and:kwo22,and:shi13,kai:mol:sto19, and bug:can:shi17, while Chen_2018's test statistic converges to $\chi^2$ distribution.} One can construct a confidence region that covers each point in $\Theta^*(p_{0})$ with asymptotic probability $1-\alpha$ as

align[align omitted — 95 chars of source]

In practice, $CS_n$ is computed by specifying a grid of values $\Theta_n$ through which to explore the parameter space, and letting $CS_n=\{\theta\in\Theta_n:T_n(\theta)\le c_{d_\theta,\alpha}\}$.

We next show that $CS_n$ is an asymptotically uniformly valid confidence set. We posit that $P_0$, the distribution of the observed data, belongs to a class of distributions denoted by $\mathcal{P}$, where the conditional law $P(\cdot|x)$ for each $P\in\mathcal{P}$ is absolutely continuous with respect to $\mu$ on $\mathcal{Y}$. We let $p_{y|x}$ denote the Radon-Nykodim derivative of $P(\cdot|x)$. We write stochastic order relations that hold uniformly over $P \in \mathcal{P}$ using the notations $o_{\mathcal P}$ and $O_{\mathcal P}$.

theoremFor constants $c>0$ and all $P\in\mathcal{P}$, let Assumptions (ref), (ref), and (ref) hold, with the following conditions replacing the corresponding ones in the original assumptions:\footnote{The constants $c$ may differ across appearances but do not depend on $P$; $\mathbb{N}$ denotes the natural numbers; and $\mathbb{B}_c(\theta)$ denotes a ball of radius $c$ centered at $\theta$.} \begin{enumerate} • $\Theta^*(p_{})\subset \operatorname{int} \Theta^{-c}\equiv\{\theta \in \Theta:\mathbb{B}_c(\theta)\subset\Theta\}$. • $\mathcal{A}_G(x)=\operatorname{supp}(G(\cdot|x;\theta))\equiv \{A\subseteq \mathcal{Y}:F_\theta(G(U|x;\theta)=A)>c\}$ for all $\theta\in\Theta$, $P-a.s.$ • The constant $c$ is the same for all $P\in\mathcal{P}$. • For all $\epsilon>0$ there exists $N\in\mathbb{N}$, with $\epsilon$ and $N$ not dependent on $P\in\mathcal{P}$, such that $P(\hat p_{n,y|x}\in\mathcal H)\ge1-\epsilon,~\forall n\ge N$, and $\|\hat p_{n,y|x}-p_{0,y|x}\|_{\mathcal H}=o_\mathcal{P}(n^{-1/4})$. • For all sequences of positive numbers $\{\delta_n\}$ with $\delta_n=o(1)$, \begin{align*} \sup_{\theta^*\in\Theta^*(p_{0})}\sup_{\|p_{y|x}-p_{0,y|x}\|_{\mathcal H}\le \delta_n}\Big\|\mathbb G_{n,\theta^*}(p_{y|x})-\mathbb G_{n,\theta^*}(p_{0,y|x})\Big\|=o_\mathcal{P}(1). \end{align*} \end{enumerate} Suppose that for all $P\in\mathcal{P}$ and $\theta^*\in\Theta^*(p_{0})$, $\min_{j=1,\dots,d_\theta}\{\rm{diag}(\Sigma_{\theta^*})\}_j>0$, and $\Vert\hat\Sigma_{n,\theta^*}-\Sigma_{\theta^*}\Vert=o_\mathcal{P}(1)$. Then, for $CS_n$ in (ref), we have $$\liminf_{n\to\infty}\inf_{P\in\mathcal{P}}\inf_{\theta^*\in\Theta^*(p_{})}P(\theta^*\in CS_n)\ge 1-\alpha.$$

Under the assumptions of Theorem (ref), Corollary (ref) also applies uniformly over $P\in\mathcal{P}$.

remarkOur confidence set can be interpreted as the union over $\theta^*\in\Theta^*$ of ellipsoids centered at $\theta^*$ whose volume is determined by $\tilde{\Sigma}_{n,\theta^*}$ only, and not by the extent of misspecification. The lever through which the latter might impact the volume of $CS_n$, is through $\Theta^*$'s volume, which may or may not be impacted by misspecification (see Section (ref)). Relative to $\Theta^*$'s volume, $CS_n$'s volume depends only on the sampling variability of the score statistic, and as $\Theta^*$ is always non-empty, $CS_n$ is always non-empty as well.

Role of Sharp Identifying Restrictions

Our analysis is able to exploit, through (ref), all identifying restrictions associated with the economic model. Doing so is in part motivated by our desire to avoid the risk of potentially discordant conclusions driven by model misspecification and different choices of subsets of moment inequalities on which to base inference ked:li:mou21. And in part because it allows us to provide a general proof, under Assumption (ref)-(e)(f), of Lemma (ref). This lemma establishes the existence of a representation of $L(\theta|x)$, for all $\theta\in\Theta$, characterized by $q_{\theta,y|x}^*$ and unique Lagrange multipliers. Importantly, one does not need to calculate that representation, because $L(\theta|x)$ coincides with it pointwise (see Lemma (ref)-(ii)). Yet, $L(\theta|x)$ in (ref) inherits its differentiability properties.\footnote{The presence of redundant (in)equality constraints in the original formulation of $L(\theta|x)$ does not affect differentiability, as there exists an equivalent formulation without such redundancies with unique Lagrange multipliers.}

In practice, when the number of inequalities in (ref) is relatively small, users of our method can derive the score in closed form. When it is moderate (a few hundred), the numerical score can be computed directly based on (ref) or many redundant inequalities can be quickly eliminated without resorting to specialized methods.\footnote{See the \href{https://github.com/hkaido0718/IncompleteDiscreteChoice}{Python library} created by Hiroaki Kaido to carry out this task for discrete choice models, and the computational simplifications in bon:kum20 for a certain class of multi-players entry models.} When the number of inequalities in (ref) is substantially larger, one can use Algorithm 3 in luo:pon:wan25 to eliminate all redundant ones. This algorithm remains practically feasible, with reported computational times of seconds for a Julia implementation that removes all redundant inequalities in entry game examples where (ref) includes $10^{14}$ inequalities, and dynamic discrete choice examples with $10^{154}$ inequalities luo:pon:wan25.

Yet, when the number of inequalities is prohibitive, one may need to carry out inference based on a subset of inequalities.\footnote{Inference methods that rely on discretization of $\mathcal{X}$ and aim to use all information in the model may face an additional computational bottleneck, as those methods have a final number of inequalities given at least by the number of nonredundant inequalities in (ref) multiplied by the cardinality of the discretization of $\mathcal{X}$.} Then, our procedure remains valid provided the conclusions of Lemma (ref) continue to hold. Hence, among the criteria guiding the selection of inequalities, one might include ensuring this is the case. We note that a rich literature studies how to remove redundant constraints in linear programs. For a given chosen subset of inequalities in (ref), which by construction defines a set linear in $q$ that includes $\mathfrak{q}_{\theta}$, as long as the identity of the non-redundant inequalities does not change with $\theta$ and those binding at $q_{\theta,y|x}^*$ are irredundant in the sense of Definition (ref), the conclusions of Lemma (ref) continue to hold (for the new linear program with enlarged constrained set) and our method remains valid. In this case, one obtains a different pseudo-true set than the one yielded by sharp restrictions, with weakly lower minimal KL divergence, and our confidence set then covers the elements of this different pseudo-true set with a prespecified asymptotic probability. Without the conclusions of Lemma (ref), existence of the score is no longer guaranteed, rendering a score-based approach inapplicable. A likelihood-based approach may nonetheless be possible, but is beyond the scope of this paper.

\noindentExample 1 (Non-sharp). Consider the two player entry game on p. (ref) with $\mathfrak{q}_{\theta}$ in (ref) and $\delta_1<0,\delta_2<0$. Instead of using sharp identifying restrictions, replace $\mathfrak{q}_{\theta}$ with $\big\{q_{y|x}\in\Delta:~q_{y|x}(y|x)\geF_\theta(S_{\{y\}|x;\theta}),~y\in\mathcal{Y},x\in\mathcal{X}\big\}$. Then $L(\theta|x)$ in (ref) can be represented using a set of constraints of the form $Aq_{y|x}\ge b(\theta)$, with $A$ a $4\times4$ matrix with three rows equal to standard basis vectors and one a vector of $1$. Hence, Lemma (ref) holds. $\square$

Computation of the Score Function

Sometimes it is possible to obtain a closed-form expression for $s_{\theta}(y|x;p_{0,y|x})$ as gradient of $\ln q_{\theta,y|x}^*$ with respect to $\theta$, as in Example (ref) (p. (ref)). If $q_{\theta,y|x}^*$ does not have a closed form expression, one needs to compute the score numerically. Here we describe how to do so, adapting the method in for23. We omit the dependence of $s_{\theta}(y|x)$ on $p_{0,y|x}$ or its estimator. We presume that one can compute $q_{\theta,y|x}^*$ relatively easily (e.g., using cvxpy).

Consider a smoothed version $f_\varsigma$ of $f(\theta)\equiv\ln q_{\theta,y|x}^*$, defined by the convolution:

align*[align* omitted — 70 chars of source]

where $\phi$ is a smooth kernel decaying to 0 in the tails, such as the Gaussian density function. The derivative of $f_\varsigma$ exists. If it admits integration by parts, one has: {

algorithm[algorithm omitted — 1,338 chars of source]

}

align*[align* omitted — 270 chars of source]

with the last expectation taken with respect to $Z\sim N(0, I_d)$. One can then approximate the derivative of $f$ by that of $f_\varsigma$. Letting $Z_r,r=1,\dots,R$ be i.i.d. draws from $N(0,I_d)$, and noting that $\big(\tfrac{\partial}{\partial z} \phi(Z)\big)/\phi(Z)=\nabla \ln \phi(z)=-z$, an unbiased estimator for $\tfrac{\partial}{\partial\theta}f_\varsigma(\theta)$ is

align[align omitted — 126 chars of source]

Replacing $f$ with $f_\varsigma$ introduces a bias proportional to $\varsigma$; when $f(\theta)$ is evaluated with noise, the variance of $[f(\theta+\varsigma Z_r)-f(\theta)]Z_r/\varsigma$ grows with $1/\varsigma^2$. In practice one needs to take a stand on this bias-variance trade off. We analyze how to do so in research-in-progress (available from the authors upon request), where numerical experiments lead us to recommend setting $\varsigma = c(nR)^{-1/4}$ for $c \in [0.03, 0.12]$. The Monte Carlo approximation in (ref) inflates the variance by a factor of $(1+\tfrac{\bar{c}}{R})$, for some constant $\bar{c}>0$, as in the method of simulated moments. This factor can easily be incorporated in the estimator of the asymptotic variance of the score. Letting $f(\theta;Y_i,X_i)=\ln q_{\theta,y|x}^*(Y_i,X_i)$, one can obtain the estimator in (ref) for each value of $(Y_i,X_i)$. The average score can then be approximated by

align[align omitted — 217 chars of source]

Algorithm (ref) presents pseudo-code with the steps and tuning parameters required to build the confidence set in (ref). When a closed form expression for $s_{\theta}(Y_i|X_i;\hat p_{n,y|x})$ is available, one can set $\varsigma=0$ and plug the observed values $\{Y_i,X_i\}_{i=1}^n$ in (ref); when it is not available, one needs to choose $\varsigma>0$ and plug $\{Y_i,X_i\}_{i=1}^n$ in (ref).

Empirical Illustration

We illustrate the usefulness of our method by applying it to answer the question addressed in kli:tam16: “what explains the decision of an airline to provide service between two airports.” kli:tam16 analyze data for the second quarter of the year 2010, documenting the entry decisions of two types of airline companies: Low Cost Carriers ($LCC$) versus Other Airlines ($OA$).\footnote{We use their data, downloading it from Quantitative Economics' \href{https://www.econometricsociety.org/publications/quantitative-economics/2016/07/01/Bayesian-inference-in-a-class-of-partially-identified-models}{online repository}.} They define a market as a trip between two airports, irrespective of intermediate stops. They record the entry decision $Y_{j,m}$ of player $j\in\{LCC,OA\}$ in market $m$ as a $1$ if a firm of type $j$ serves market $m$, and $0$ otherwise. They posit that player $j$'s decision to serve a market depends not only on their opponent's entry decision, but also on observable payoff shifters $X_{j,m}$ and unobservable payoff shifters $U_{j,m}$. The observable payoff shifters $X_{j,m}$ include the constant and two continuously distributed variables, $X_{j,m}^{pres}$ and $X_m^{size}$. The first variable is firm-and-market-specific: it measures the market presence of firms of type $j$ in market $m$ kli:tam16. Market presence of the LCC airline, $X_{LCC,m}^{pres}$ (respectively, $X_{OA,m}^{pres}$), is excluded from the payoff of firm $OA$ (respectively, $LCC$). The second variable, market size, enters the payoff of firms of both types; it measures population size at the two endpoints of the trip and is market-specific. The unobservables $U_{j,m}$, $j\in\{LCC,OA\}$, are assumed to have a bivariate normal distribution with $\mathbb{E}(U_{j,m})=0$, $Var(U_{j,m})=1$, $Corr(U_{LCC,m},U_{OA,m})=r$, and to be i.i.d. across $m$.\footnote{We assume $r\in [-0.9,0.9]$ and estimate it as part of the vector $\theta$. We ensure that the strategic interaction parameters $\delta_{LCC}$ and $\delta_{OA}$ are less than a constant $c<0$ and that $q_{\theta^*,y|x}>c$ for another constant $c>0$.}\newline Both kli:tam16 and we assume that players enter the market if doing so yields non-negative payoffs. However, we posit different payoff functions. They posit:

multline[multline omitted — 250 chars of source]

In words, kli:tam16 transform each of market size and of the two market presence variables into binary variables, based on whether each of these variables realizes above or below their respective median. Doing so yields a finite number of unconditional moment inequalities, which they need for their inference procedure, at the cost of foregoing the information provided in the variation in $X_{j,m}$ past whether each variable is above or below its median, and of using an arguably more restrictive payoff function.

Leveraging our new method, we are able to avoid discretizing the continuously distributed covariates, thereby exploiting all identifying power in their variation and allowing $X_{j,m}$ to impact payoffs proportionally to their value. We assume that payoffs take the form:

align[align omitted — 146 chars of source]

We study how the decision of an $LCC$ airline to enter the market is affected by whether an $OA$ airline is in the market and by the extent of $LCC$ airlines market presence. To do so, we define the potential entry decision of an $LCC$ player as $Y_{LCC}(d)=\mathbf{1}(X_{LCC,m}^\top\beta_{LCC}+\delta_{LCC} d+U_{LCC,m}\ge 0)$, with $\beta_j=(\beta_j^0,\beta_j^{pres},\beta_j^{size})$, $j\in\{LCC,OA\}$. This is the entry outcome of an $LCC$ airline when we fix the $OA$’s entry to take value $d\in\{0,1\}$. Based on our model, the entry probability of the $LCC$ airline is $P(Y_{LCC}(d)=1|X_{LCC,m})=\Phi(X_{LCC,m}^\top\beta_{LCC}+\delta_{LCC} d)$. We obtain a confidence interval for this parameter for each of $d=0,1$ and for specific values of $X_{LCC,m}$. To ease the reporting of results, we set $X_m^{size}$ equal to the median of its distribution and compute the $\tau$-quantile of the distribution of $X_{LCC,m}^{pres}$ for $\tau\in\mathcal{T}\equiv\{0.125,0.250,0.375,0.5,0.625,0.750,0.875\}$; we evaluate our parameter of interest for $X_{LCC,m}^{pres}$ set equal to each of these values, but our confidence set construction uses all information in $X_{j,m}$. Letting $\theta=(\beta_{LCC},\delta_{LCC},\beta_{OA},\delta_{OA},r)$, we report confidence intervals $CI_n(x,d)=\big[\min_{\theta:~T_n(\theta)\le c_{d_\theta,\alpha}} \Phi(x_{LCC,m}^\top\beta_{LCC}+\delta_{LCC} d)$, $\max_{\theta:~T_n(\theta)\le c_{d_\theta,\alpha}} \Phi(x_{LCC,m}^\top\beta_{LCC}+\delta_{LCC} d)\big]$ for the values of $x$ corresponding to $\tau\in\mathcal{T}$ and for $d\in\{0,1\}$.\footnote{As in the Monte Carlo experiments in Section (ref), we estimate $p_{0,y|x}$ using a series estimator with $J$-th order (tensor product) B-spline basis functions and set $\varepsilon= 0.05$ in (ref) to compute $\tilde\Sigma_{n,\theta}$.} Under the conditions of Theorem (ref), by standard arguments $$\liminf_{n\to\infty}\inf_{P\in\mathcal{P}}\inf_{\theta^*\in\Theta^*(p_{})}P(\Phi(x_{LCC,m}^\top\beta^*_{LCC}+\delta^*_{LCC} d)\in CI_n)\ge 1-\alpha.$$

figure[figure omitted — 522 chars of source]

Figure (ref)-Panel (a) reports our results, displaying on the horizontal axis the value of $\tau$ and on the vertical axis the candidate value for $\Phi(x_{LCC,m}^\top\beta^*_{LCC}+\delta^*_{LCC} d)$. The results show potentially sizable heterogeneity in the treatment effects of interest and reject the hypothesis that they are constant across $\tau$. While the confidence intervals for $d=0,1$ are not disjoint for several values of $\tau$, when $OA$ opponents are not in the market (blue segments in Figure (ref) for $d=0$) the entry probability can be much larger than when they are present (orange segments for $d=1$) across all values of $\tau$, with the effect largest for $\tau= 0.750$.

For each fixed value of the entry decision of $OA$, as the market presence of the $LCC$ airlines increases, so does the probability that $LCC$ firms enter a market. When OA firms are in the market, the difference is statistically significant for all $\tau$, although the impact of $x_{LCC,m}^{pres}$ on the entry probability is low until market presence reaches its 0.625 quantile, at which point the slope increases rapidly. This suggests that in order to overcome the presence of $OA$ opponents and enter the market, $LCC$ firms need large market presence. On the other hand, when $OA$ firms are not in the market, the impact of $x_{LCC,m}^{pres}$ on the entry probability is sizable starting with $\tau=0.375$ (and further increases with $\tau$).

We compare our results to what one would obtain using the likelihood based inference method in Chen_2018, which is designed for correctly specified models with discrete covariates. We note that Chen_2018 assume the payoff function in (ref), whereas we use the specification in (ref), and therefore the coefficient estimates on $X_{j,m}$ and $Y_{-j}$ are not directly comparable to each other. Nonetheless, we believe it to be instructive to compare the counterfactual model-implied entry probabilities across the two approaches.\footnote{We use the replication package provided by Chen_2018 at Econometrica's \href{https://www.econometricsociety.org/publications/econometrica/browse/2018/11/01/monte-carlo-confidence-sets-identified-sets}{online repository}, where the payoffs are specified as $\tilde\pi_{j,m}=Y_{j,m}(\tilde\beta_j^0+\tilde\beta_j^{pres}\mathbf{1}(X_{j,m}^{pres}> Med(X_j^{pres}))+\tilde\beta_j^{size}\mathbf{1}(X_m^{size}> Med(X^{size}))+\tilde\delta_j Y_{-j,m}+U_{j,m})$ (compare with (ref)). We compute the confidence intervals using the payoffs in their code and obtain a confidence interval on $P(Y_{LCC}(d)=1|X_{LCC,m})$ through the projection method that they propose.} Figure (ref)-Panel (b) reports confidence intervals based on Chen_2018's projection method for $d=0,1$ and for $X_{LCC,m}^{pres}$ below the median and above the median. The figure shows that aggregating the value of $X_{LCC,m}^{pres}$ at this coarse level hides interesting patterns in the results. Bundling “above the median” and “below the median” as single values for the covariates does not allow one to learn the extent of the heterogeneity in the effect of market presence on the probability of entry. Moreover, using Chen_2018's inference method yields very large treatment effects for the presence of an OA opponent, which instead we do not find. Of note, our confidence sets are shorter than those based on Chen_2018 for all values of $\tau$, often substantially, and hence this difference is not driven by a reduction in precision but likely by a different model and our use of all information in the covariates. While one could use a finer discretization of $(X_{LCC,m}^{pres},X_{OA,m}^{pres})$ in (ref) combined with Chen_2018's method, doing so would result in a substantially harder computational problem. Indeed, in Chen_2018's approach the selection probabilities, which are allowed to depend on $X$ (but not $U$) and have cardinality at least equal to the cardinality of $\mathcal{X}$, are part of the parameters to be estimated. In contrast, the computational complexity of our procedure does not change with the cardinality of $\mathcal{X}$.

Monte Carlo Experiments

We carry out an empirical Monte Carlo exercise where the data-generating process is calibrated to the data we use for the application in Section (ref); the notation is as in that section and the payoffs are as in (ref). We normalize each covariate to the unit interval (and continue to denote them $\{X_{LCC}^{pres},X_{OA}^{pres},X^{size}\}$) and assume $(U_{LCC,m},U_{OA,m})$ is distributed i.i.d. bivariate standard normal. We let $\theta=(\beta_{LCC},\delta_{LCC},\beta_{OA},\delta_{OA})$.

To calibrate a DGP value for the parameter vector $\theta$ from kli:tam16's data, we introduce a parameter $\kappa\in[0,1]$, independent of $(X,U)$, representing the probability of selecting outcome $Y=(1,0)$ when multiple equilibria are present. Define

align[align omitted — 458 chars of source]

where the functions $\eta_1(\cdot;x),\eta_2(\cdot;x),\eta_3(\cdot;x)$ are defined in (ref), (ref), (ref). We report in Table (ref) the value for $(\theta,\kappa)$ estimated by maximizing the likelihood based on (ref)-(ref). We use this estimate as baseline DGP value in our simulations.\footnote{To ensure that the DGP induces conditional choice probabilities that are bounded away from 0 and 1, we estimate $(\theta,\kappa)$ using observations such that the estimated conditional choice probabilities are in the interval $(\epsilon,1-\epsilon)$ with $\epsilon=1e-3$. This gives us a sample of size 7,017.}

table[table omitted — 409 chars of source]

We consider two designs for our DGPs. Design 1 uses only the two player-specific covariates, $\mathbf{X}^{D1}=(X^{pres}_{LCC},X^{pres}_{OA})$, and sets $(\theta_0,\kappa_0) = (\beta^0_{LCC}, \beta^{pres}_{LCC}, \delta_{LCC}, \beta^0_{OA}, \beta^{pres}_{OA}, \delta_{OA}, \kappa)$ to the corresponding maximum likelihood estimates in Table (ref). Design 2 incorporates the full set of covariates, $\mathbf{X}^{D2}=(X^{pres}_{LCC},X^{pres}_{OA},X^{size})$ and sets $(\theta_0,\kappa_0)$ to the full MLE vector in Table (ref). As we further explain below, we implement Design 1 because of computational difficulties with Design 2 for the comparator inference method of and:shi13.

Within each design $k=1,2$, we resample from the original kli:tam16's dataset covariates $\mathbf{X}^{Dk}_m,m=1,\dots,n$. We evaluate the performance of our inference method and:shi13 when the model is correctly specified, by generating $(Y_{LCC,m},Y_{OA,m})$, $m=1,\dots,n$ through inverse CDF transformation, using as DGP the probability mass functions in (ref)-(ref) in which we plug $\mathbf{X}^{Dk}_m$ and, for $(\theta,\kappa)$, the MLE values in Table (ref) for that design.

To evaluate performance under model misspecification, we simulate data from a two-player entry game where the true payoff for player $j$ in Design $k$ is given by:

align*[align* omitted — 135 chars of source]

with $X^*$ a binary variable omitted from the model, $X_{j}^{D1}=[1~~X^{pres}_{j}]^\top$, and $X_j^{D2}=[1~~X^{pres}_{j}~~X^{size}]^\top$. Given $(X^{pres}_{LCC},X^{pres}_{OA})=(\tilde x_{LCC},\tilde x_{OA})$, $X^*=1$ with probability $\Phi(\tfrac{\tilde x_{LCC}-\mu_{LCC}}{\sigma_{LCC}}+\tfrac{\tilde x_{OA}-\mu_{OA}}{\sigma_{OA}})$, with $\mu_j$ and $\sigma^2_j$ the mean and variance of $X^{pres}_{j,m}$. The value of $\gamma$ determines the extent of misspecification. We report results for $\gamma\in\{-.1,-.2,-.3,-.4\}$.

figure[figure omitted — 318 chars of source]

Figure (ref) shows the projections of $\Theta^*(p_{0})$ onto the space of $(\delta_{LCC},\delta_{OA})$ under Design 2.\footnote{A similar figure for Design 1 is omitted to conserve space but available from the authors upon request.} When the model is correctly specified (left panel), $\Theta^*(p_{0})$ coincides with the sharp identification region of $\theta$. With misspecification ($\gamma=-0.4$), the optimal value of the KL divergence measure in (ref) is strictly positive (see Table (ref) for the value of $I(p_{0}||q_{\theta}^*)$), indicating that the sharp identification region is empty. In contrast, the pseudo-true set $\Theta^*(p_{0})$ remains nonempty, as shown in the right panel of Figure (ref), and the shape of $\Theta^*(p_{0})$ remains similar both in the correctly specified and in the misspecified case. Nonetheless, under misspecification it shifts to the (lower) left. This is because when $X^*=1$, the true DGP allocates a large mass to either $(1,0)$ or $(0,1)$ (whereas with $X^*=0$ that mass is allocated to $(1,1)$). This is not captured by the model in (ref). A reduction in the values of $(\delta_1,\delta_2)$ (an increase in absolute value) enlarges the region of multiplicity, thereby allowing for a larger mass to be allocated to $(1,0)$ and $(0,1)$ than the model in (ref) allows for.

To implement our test, we estimate $p_{0,y|x}$ by a series estimator with $J$-th order (tensor-product) B-spline basis functions.\footnote{We compute $\tilde\Sigma$ in (ref) setting $\varepsilon= 0.05$ as done in and:shi13; additional simulations (available from the authors) with $\varepsilon= 0.025$ and $0.1$ indicate that increasing $\varepsilon$ makes our procedure more conservative.} For Design 1, we compare the performance of our procedure to that of and:shi13.\footnote{We implement the method in and:shi13 using their $S_3$ test statistic, with moment functions

align*[align* omitted — 376 chars of source]

where $m(z,\theta)=(m_{\le}(z,\theta)^\top,m_{=}(z,\theta)^\top)^\top$, and with $\bar m_n(\theta)=\tfrac{1}{n}\sum_{i=1}^n m(Z_i,\theta)$. The number of inequalities is two times the number of hyper-cubes used, and similarly for the equalities.} We report a power comparison where the parameter value $\theta^*$ lies on the boundary of $\Theta^*(p_0)$. Specifically, we examine the rejection probabilities for local alternatives of the form $\theta_{0,h} = (\beta^*_{LCC}, \delta^*_{LCC}+\tfrac{h}{\sqrt n},\beta^*_{OA}, \delta^*_{OA}+\tfrac{h}{\sqrt n})$, with $h > 0$. This corresponds to drifting the strategic interaction effects toward $(0,0)$ while holding the other components fixed. and:shi13's test transforms the conditional moment inequalities into unconditional ones using as instruments indicator functions of whether each covariate belongs to specified hypercubes. The side length of each hypercube is $1/(2r)$, for $r = 1, \dots, r_{1n}$, where a larger value of $r_{1n}$ corresponds to finer conditioning information. Following and:shi13, we set $r_{1n} = 3$, and also report results for $r_{1n} = 2$ and $4$. For Design 2, we report results only for the score-based test, as the moment-based test was computationally infeasible in this setting.

Table (ref) reports the results of this exercise for $500$ Monte Carlo repetitions. Panel (A) documents the size and power of our test as well as the moment inequality-based tests for the case that the model is correctly specified. The test of and:shi13 over-rejects slightly in this correctly specified DGP, while our Rao's score-based test has valid size but under-rejects. Nonetheless, the power curve of our test quickly dominates that of the moment inequality based test.

table[table omitted — 7,835 chars of source]
table[table omitted — 449 chars of source]

In the misspecified case (Panels (B)-(F)), as expected the moment inequality-based test is oversized. The extent of the size distortion grows with the extent to which the model is misspecified. To quantify the latter across DGPs, we compute the rejection probability of an infeasible Information Matrix test White1982 that uses knowledge of the fact that the selection mechanism $R$ in (ref) is distributed $\mathrm{Bernoulli}(\kappa_0)$. Enriched with this information, the model yields a unique prediction $q_{(\theta,\kappa_0),y|x}$ and a well defined likelihood function, and hence we can obtain the (point identified) maximum likelihood estimator $\hat{\theta}^{MLE}$ that a researcher would obtain if they knew the selection mechanism. We compute the Hessian and the outer product forms for the covariance matrix, and evaluate them at $\hat{\theta}^{MLE}$ to carry out the Information Matrix test. We report the rejection probability of this infeasible test in Table (ref), labeling it “IM rej.” As can be seen from the table, for levels of $\gamma\in\{-.1,-.2\}$ the rejection probability is low, reaching at most $7.6\%$ under Design 1; yet, and:shi13's test already shows non-trivial size distortions ($16\%$ and $37\%$, respectively). For $\gamma=-0.3,-.4$, the power is higher, reaching $91.6\%$ under Design 1. For such settings, the size of and:shi13's test is substantially distorted ($75\%$ and $96\%$, respectively, against a $5\%$ nominal level). In contrast, our test has correct size throughout all simulations, and maintains a power curve that is very similar to the one it displays in the case of correct model specification.

Table (ref) reports average computational time in seconds to calculate test statistics and critical values in Design 1 for 500 Monte Carlo replications on Boston University's computing cluster (with Intel Xeon Gold 6132 Processors and 192GB RAM). Our test statistic is 35-294 times faster to compute than and:shi13's and:shi13, but the most substantial gain comes from calculation of the critical value: 0.0002 seconds for us, against 48-818 for and:shi13, for a total computation time reduction of 228-3,564 times.

Conclusions

This paper is concerned with statistical inference in incomplete models with set valued predictions. Such models are typically partially identified, and can be misspecified. Misspecification can make the identification region of the model's parameters spuriously tight or even empty, raising a challenge for interpreting identification results, and can cause existing testing procedures to severely overreject. We propose to resolve these problems through an information-based method. Our method delivers a non-empty pseudo true set which can be interpreted as the set of minimizers of the researcher's ignorance about the true structure, as in White1982. For any given parameter value, our inference method solves a convex program to find the density function that is closest to the data generating process with respect to the Kullback-Leibler information criterion. It then obtains the score of the likelihood function associated with this density and a Rao score test statistic. We show that the test statistic has an asymptotically pivotal distribution, is easy to compute, and does not require moment selection. The associated test has uniformly valid asymptotic size, is applicable to both correctly specified and misspecified models, and allows for discrete and continuous covariates. Monte Carlo simulations confirm the good computational and statistical properties of our proposed inference method.