EconBase
← Back to paper

Robust Likelihood Ratio Tests for Incomplete Economic Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

250,324 characters · 41 sections · 129 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Robust Likelihood Ratio Tests for Incomplete Economic Models

abstractThis study develops a framework for testing hypotheses on structural parameters in incomplete models. Such models make set-valued predictions and hence do not generally yield a unique likelihood function. The model structure, however, allows us to construct tests based on the least favorable pairs of likelihoods using the theory of Huber and Strassen (1973). We develop tests robust to model incompleteness that possess certain optimality properties. We also show that sharp identifying restrictions play a role in constructing such tests in a computationally tractable manner. A framework for analyzing the local asymptotic power of the tests is developed by embedding the least favorable pairs into a model that allows local approximations under the limits of experiments argument. Examples of the hypotheses we consider include those on the presence of strategic interaction effects in discrete games of complete information. Monte Carlo experiments demonstrate the robust performance of the proposed tests.

\onehalfspacing

Keywords: Incomplete models, Robust inference, Likelihood ratio tests, Limits of experiments

Introduction

Incomplete structures arise in a wide class of economic models when the researcher's theory does not fully describe how a particular outcome occurs given the primitives of the model. In this study, we consider a class of models in which, given structural parameter $\theta\in \Theta$ and latent variable $u\in U$, the model predicts the set $G(u|\theta)$ of values for discrete outcome $s$.\footnote{Introducing covariates does not fundamentally change the structure. We therefore treat this case in Section (ref) as an extension.} The researcher observes $s$, but his/her theory is silent about the mechanism that determines how $s$ is selected from the predicted set. This class encompasses various models studied in the empirical literature. Examples include models of market entry bresnahan1990entry,bresnahan1991empirical,berry1992estimation,Ciliberto:2009aa where the theory does not specify how a pure strategy Nash equilibrium is selected, models of self-selection Heckman:1990aa,Mourifie:2018aa where an individual's choice of the sector of activity interacts with unobserved skills, and models of English auction Haile:2003to,Aradillas-Lopez:2008ab where the researcher wants to allow solutions that satisfy weak rationality restrictions.\footnote{Other examples include models of voting Kawai:2013aa, choice of product variety EIZENBERG:2014aa, and network formation models Miyauchi:2016aa.} In many of these settings, empirical questions can be investigated by testing hypotheses on the structural parameter. Given the incompleteness of the theory, it is desirable to conduct tests without adding assumptions on how selections operate. Despite the need for such robustness, the theoretical study of robust testing procedures and their properties has been limited.\footnote{Studies of tests of the moment inequality model are related, but their model differs from the one we consider. See Section (ref).} This study develops a framework for hypotheses testing in incomplete models, shows how to construct robust and optimal tests, and provides asymptotic tools to evaluate their performance.

Each of the hypotheses we consider can be written as

align[align omitted — 109 chars of source]

for some function $\varphi:\Theta\to \mathbb R^k$ and mutually exclusive sets $K_0,K_1\subset\mathbb R^k$. Such hypotheses naturally arise in applications of incomplete models. For example, in an entry game, a key parameter is the strategic interaction effect, which measures the effect of an opponent firm's entry on a firm's profit. An important empirical question is whether the presence of such interaction effects can be supported by the observed data dePaula2012inference. One way to address the question is to formally test the null hypothesis that the strategic interaction does not exist, namely $H_0:\varphi(\theta)=0$, against an alternative hypothesis that negative externalities exist, $H_1:\varphi(\theta)< 0$, by choosing a suitable functional $\varphi$.

Such a hypothesis testing problem, however, faces several challenges. First, without further assumptions, the model permits multiple distributions of the observables even if each hypothesis fully specifies the value of $\theta$. To see this, consider a simplified problem in which $H_0:\theta=\theta_0~v.s. ~H_1:\theta=\theta_1$. This problem may appear as testing a simple null hypothesis against a simple alternative. However, under each hypothesis, multiple distributions (of $s$) may be compatible with the theory because any distribution $P$ of the outcome is consistent with $\theta$ as long as one can augment the model by finding a suitable selection mechanism that induces $P$. Therefore, even under the simplest setting, both null and alternative hypotheses can be composite (in terms of permitted distributions).\footnote{Furthermore, the hypotheses in (ref) allow the presence of additional nuisance parameters such as sub-components of $\theta$. We address this issue separately as an extension of the base framework in Section (ref).} The problem becomes even more challenging when data are obtained from a sequence of experiments. If one stays agnostic about the selection, the unknown selection mechanism is allowed to be arbitrary across experiments. For example, across experiments, the true selection mechanism may vary with and be correlated through a specific variable; however, the researcher does not even know the identity of this variable. From the researcher's viewpoint, the resulting outcome sequence is then heterogeneous and dependent in an unknown way, which in turn makes it hard to characterize the large-sample distribution of test statistics and apply standard asymptotic tools to analyze the power of the tests.

We develop tests that overcome these challenges. For this, we exploit the fact that the sampling uncertainty and lack of understanding of the selection can be represented by a {\em belief function\/}, a capacity (or non-additive probability), which belongs to a broader class of two-monotone capacities. Capacities in this class are known to have properties useful for conducting robust statistical inference Huber:1973aa.\footnote{See also Huber:1974aa for corrigendum and Huber:1981aa for the broader area of robust statistics.} We start by demonstrating that for testing between simple hypotheses, robust tests can be constructed for any finite sample. The proposed test, which takes the form of a likelihood ratio (LR) test, controls the size in finite samples regardless of the unknown selection mechanism and maximizes a measure of power, which we call {\em lower power\/}. One may wonder how such an LR test can be constructed because incomplete models generally admit infinitely many likelihoods. A key observation is that HS's theory ensures that there exists the least favorable pair (LFP) of likelihoods: one compatible with the null that is the least favorable for size control and the other compatible with the alternative that is the least favorable for power maximization.\footnote{They also show that such a pair is unique up to its Radon-Nikodym derivative.} Distinguishing two such extreme distributions turns out to be the best way to test one parameter value against another while staying agnostic about the selection.

We then develop LR tests for repeated experiments. Our first main contribution is to show that despite the potential heterogeneity and dependence of the data, the LFP consists of product measures as long as the latent variables are independent across experiments. Heuristically, this means that under the least favorable distribution for size control (or power maximization), the observables can be viewed as independent across experiments, while the true data-generating process (DGP) may not satisfy such regularity. This leads to a number of desirable results. In particular, it allows us to construct robust LR tests that are optimal in the minimax sense, provide a simple critical value based on a large-sample Gaussian approximation, and develop an asymptotic framework for evaluating the power of the tests.

Our second contribution is on the practical side. While HS's theory ensures the existence of the LFP, in practice, one needs to find a way to compute it. We show that in the class of models we consider, the LFP can be computed by solving a finite-dimensional convex program in which the constraints of the program are the {\em sharp identifying restrictions\/} studied in the identification literature beresteanu2011sharp,galichon2011set,Chesher:2014aa. These restrictions simplify the constraints by making them linear in the control variable, and they therefore play a crucial role in computing the LFP and implementing the robust optimal tests. While the restrictions are useful for characterizing sharp identified sets, little is known about whether they lead to statistically optimal inference. Our result shows they are indeed crucial for likelihood-based inference that has a certain optimality property. Our theoretical result on the LFP also has a practical implication. In particular, under mild conditions, the distributions forming the LFP are independently and identically distributed (i.i.d.) laws, and hence the researcher only needs to find the LFP in a “single” experiment rather than finding it from the entire sequence of experiments. This result also contributes to a significant reduction in the computational cost of our tests.

Our third main contribution is to provide a framework for analyzing the asymptotic power of the tests by embedding the product LFPs into a model that admits local approximations. Specifically, we show that under regularity conditions, a sequence of experiments characterized by the ratio of the LFPs, obtained from a null parameter value and a local alternative, converge to a limit in the sense of Le-Cam:1972aa,LeCam:1986aa. We use this property to characterize the upper bound of the asymptotic lower power of the tests for one-sided hypotheses. Our approach uses the limits of experiments argument and can potentially be used in other statistical decision problems in incomplete models. The main advantage of this approach is that once the LFPs are embedded into a probabilistic model whose limit is tractable, most of the power analysis can be performed using standard tools.

Our framework, however, also incorporates some non-standard features. First, the underlying model in which we embed the LFPs may not satisfy the well-known differentiability in quadratic mean condition, which is sufficient for the local asymptotic normality (LAN) of the experiments over the entire local parameter space. Instead, the model is typically directionally differentiable (in the $L^2$ sense) and satisfies the LAN property separately on a collection of convex cones that partition the local parameter space. Second, perhaps more importantly, incomplete models may yield alternatives that are not robustly testable. Such an alternative admits a selection mechanism that makes the lower power of any level-$\alpha$ test weakly below the nominal level. We clarify the notion of robust testability and relate it to the observational equivalence concepts studied in the identification literature Chesher:2014aa. To conduct a meaningful power analysis, we then introduce an extended notion of alternatives, which we call {\em shifted local alternatives\/}. The asymptotic power envelope is shown to be non-trivial against such alternatives.

We further extend our analysis to a model that permits the presence of nuisance components of the parameter vector. Setting up a statistical decision problem, we construct a robust LR test that minimizes a certain risk function. We call this test a Bayes--Dempster--Shafer (BDS) test as it minimizes a risk that treats parameter uncertainty in a Bayesian way and incorporates ambiguity due to incompleteness through a belief function. Finally, we establish a minimax theorem for this setting, which suggests that a level-$\alpha$ test that maximizes a weighted average of lower power can be approximated using a sequence of BDS tests.

Relation to the Literature

Our study is most closely related to Epstein:2016qv who developed a theoretical framework for modeling repeated experiments with incompleteness.\footnote{epstein2015exchangeable provided axiomatic foundations for robust subjective inference and decision making in such a setting.} We adopt their framework and use the (product) belief function to characterize the set of joint distributions of outcomes across experiments. This allows us to study the robustness of tests even in settings where selections are heterogeneous and dependent in an unknown way. This study then takes a step further and develop ways to examine the optimality of tests in such settings.

Our study is also related to earlier work on incomplete models.\footnote{The analysis of an incomplete system of equations dates back to the early work of Wald1950. Here, we focus on reviewing more recent developments in models with multiple equilibria.} In particular, our framework for the single experiment builds on that of jovanovic1989observable, who pointed out that models with multiple equilibria lead to incomplete structures and face potential difficulty in identifying structural parameters. tamer2003incomplete studied identifying restrictions in an incomplete simultaneous discrete response model with multiple equilibria. Since his seminal work, it has become common to use partially identifying inequality restrictions to bound parameters of interest. galichon2011set, beresteanu2011sharp, and Chesher:2014aa characterized sharp identifying restrictions for a wide range of incomplete models using the theory of random sets. We also use, as a central tool, the capacities associated with random sets. As discussed above, the sharp identifying restrictions play an important role in the construction of tests that achieve robustness and statistical optimality.

Commonly used identifying restrictions take the form of moment inequalities. As such, inference methods developed for moment inequality models (chernozhukov2007estimation, andrews2010inference, bugni2010bootstrap, andrews2012inference) have been commonly used. Some of them galichon2006inference,galichon2009test,galichon2013dilation use test statistics based on capacities to construct confidence regions. These methods combine the implications of incomplete models on moments with an additional assumption on the sampling process (e.g., i.i.d. sampling). By contrast, our approach uses the model's implications on certain likelihoods and does not restrict the sampling process. Chen:2011aa considered a sieve MLE-based inference, which can be applied to incomplete models. Their approach profiles out a non-parametric nuisance parameter (selection) from the likelihood function using a sieve. Our approach, which picks out the LFP, can also be interpreted as a way to average out the nuisance parameter, in which the weights are the least favorable priors, and averaging is carried out without explicitly introducing a functional space.

The results on the optimality of the tests in related settings are somewhat limited. Within a moment inequality framework, CANAY2010408 found that a test based on the empirical LR statistic is optimal with respect to the large deviations criterion. In a more specialized setting in which moment restrictions are convex in the parameter, Kaido:2014aa showed that a test based on a semiparametrically efficient estimator of the identified set achieves the asymptotic power envelope against some local alternatives. In models characterized by conditional moment inequalities, Armstrong:2014aa,Armstrong:2018aa compared the relative power of the testing procedures based on Cramer--von Mises and weighted Kolmogorov--Smirnov statistics. These studies deal with testing problems in models characterized by moment inequalities, which differ from ours in terms of (i) the hypotheses they test and (ii) how they extend a single experiment to repeated experiments. For the former, these studies consider testing whether $\theta$ is in the identified set, while our focus is on testing hypotheses of the form in (ref), which does not involve identified sets. For the latter, they assume that an i.i.d. sample is available, and hence the robustness issue against heterogeneity and the dependence of selection does not arise.\footnote{See Epstein:2016qv for this distinction as well as MR3753715 (Section 5.3).}

Finally, our framework for inference is related to others that use limit theorems based on the thoery of random sets. As mentioned earlier, we use a Gaussian approximation to compute the critical value for the LR statistic, which is similar in spirit to the central limit theorem (CLT) in Epstein:2016qv, whereas a different tool is used to obtain this result because of the non-trivial difference between the LR statistic we use here and their Kolmogorov--Smirnov-type statistic. In a different class of models, in which observations are set-valued, beresteanu2008asymptotic applied a central limit theorem for random sets to make their inference.

Throughout, for any metric space $A$, we let $\Sigma_A$ denote its Borel $\sigma$-algebra. We then denote the set of Borel probability measures on $A$ by $\Delta(A)$ and equip it with the topology of weak convergence. Let $N(\mu,V)$ denote the law of a normal random vector with mean $\mu\in\mathbb R^k$ and variance-covariance matrix $V\in\mathbb R^{k\times k}$. For any integrable random vector $X$, we let $E_P[X]$ denote its expectation with respect to probability measure $P$.

The remainder of the paper is organized as follows. Section (ref) introduces the model and provides illustrative examples. In Section (ref), after reviewing the theory of Huber:1973aa (Section (ref)), we present our main result on minimax tests in repeated experiments (Section (ref)). We also discuss the robust testability of the hypotheses and computational aspects (Sections (ref)--(ref)). Section (ref) provides results on the local asymptotic power of the tests. Section (ref) then provides simulation evidence. Section (ref) contains extensions of the baseline framework and Section (ref) concludes. Appendices (ref) and (ref) collect the proofs of the theoretical results.

Setup

Let $S$ be a finite set of observable outcomes and let $u\in U$ denote a variable unobservable to the researcher, where $U$ is assumed to be a Polish space. Let $\Theta$ denote the parameter space. We let $m=\{m_\theta,\theta\in\Theta\}$ denote a family of Borel probability measures on $U$. For each $\theta\in\Theta$, let $G(\cdot|\theta):U\twoheadrightarrow S$ be a weakly measurable correspondence. This map shows how latent variable $u$ is mapped to a set of permissible outcomes. Observable outcome $s$ is then a measurable selection of random set $G(u|\theta)$. As such, the model does not impose any restrictions on how $s$ is selected. One may also introduce observable covariates to this model. As the core analysis remains unaffected, we defer the analysis of this case to Section (ref).

The incomplete structure above is summarized by tuple $(S,U,m,\Theta;G)$. Such structures arise in various economic models. To fix the ideas, we present several examples based on simplifications of well-known models. The first example is a binary response game, which is commonly used to analyze environments such as firms' entry into markets and households' joint labor supply decisions bresnahan1990entry,bresnahan1991empirical,berry1992estimation,Ciliberto:2009aa.

example[Binary response game]\rm Consider a two-player binary response game with the following payoff: \begin{center} \begin{tabular} { r|c|c| } \multicolumn{1}{r} & \multicolumn{1}{c}{out} & \multicolumn{1}{c}{in} \\ \cline{2-3} out & $0,0$ & $0,u^{(2)}$\\ \cline{2-3} in & $u^{(1)},0$ & $u^{(1)}+\theta^{(1)}, u^{(2)}+\theta^{(2)}$ \\ \cline{2-3} \end{tabular} \end{center} The effect of the other player's action (e.g., entry) on player $k$'s payoff is represented by $\theta^{(k)}$. Throughout, we call $\theta=(\theta^{(1)},\theta^{(2)})'\in\Theta\subset \mathbb R^2$ the players' {\em strategic interaction effects\/}. Let $U=\mathbb{R}^2$. The latent payoff shifter $u=(u^{(1)},u^{(2)})'$ follows a continuous distribution $m_\theta$. Consider pure strategy Nash equilibria in this game when $\theta^{(1)}\le 0$ and $\theta^{(2)}\le 0$.\footnote{For simplicity, we focus on games with strategic substitutes throughout. Games with strategic complements, in which $\theta^{(1)}>0,\theta^{(2)}>0$, can be analyzed similarly.} There are four possible equilibrium outcomes: $S=\{(0,0),(1,1),(1,0),(0,1)\}$. How $u$ and $\theta$ are mapped to the equilibrium outcomes is summarized by the following correspondence: \begin{equation} G(u|\theta) = \left\{ \begin{array}{cl} \{(0,0)\} & u^{(1)}<0, u^{(2)}<0\\ \{(1,1)\} & u^{(1)}\ge-\theta^{(1)}, u^{(2)}\ge-\theta^{(2)}\\ \{(1,0)\} & u\in U_1,\\ \{(0,1)\} & u\in U_2,\\ \{(1,0),(0,1)\} & 0\le u^{(1)}<-\theta^{(1)}, 0\le u^{(2)}<-\theta^{(2)}, \end{array} \right. \end{equation} where $U_1=\{u:u^{(1)}\ge -\theta^{(1)}, u^{(2)}<-\theta^{(2)})\cup\{u:0\le u^{(1)}<-\theta^{(1)}, u^{(2)}<0\}$ and $U_2=\{u:0\le u^{(1)}<-\theta^{(1)}, u^{(2)}\ge -\theta^{(2)}\}\cup \{u:u^{(1)}<0, u^{(2)}\ge 0\}$. The model predicts multiple equilibria when each player's latent payoff shifter is between the two thresholds ($0$ and $-\theta^{(k)}$, $k=1,2$).

The second example is the (binary) Roy model studied in Mourifie:2018aa.

example[Roy model]\rm Consider an individual who chooses a sector of activity $D\in \{0,1\}$ and whether to work $Y\in \{0,1\}$ in the sector. The binary outcome is given by $Y=Y_1D+Y_0(1-D)$, where selection indicator $D$ is determined by binary potential outcomes $(Y_0,Y_1)$ through the following structure: \begin{align} D=\begin{cases} 1&Y_1>Y_0\\ 0 or 1& Y_1=Y_0\\ 0&Y_1<Y_0. \end{cases} \end{align} Binary potential outcome $Y_d$ represents whether one has good economic prospects in sector $d\in\{0,1\}$.\footnote{We focus on the case in which the potential outcomes are binary. Mourifie:2018aa extended their analysis to more general settings in which $Y_d$ is discrete or continuous (or both). As discussed in Section 2 of their paper, one could also think of the binary Roy model as a consequence of a two-step decision process in which $D$ is determined first by potential wage $Y^*_d$ in sector $d$, and whether to work in section $d$ is determined by whether $Y^*_d$ crosses a threshold.} The sector choice is not uniquely determined if $Y_1=Y_0$. This model can be mapped to the present framework by letting $s=(y,d)\in S=\{(0,0),(0,1),(1,0),(1,1)\}$ be observable outcomes and $u=(Y_0,Y_1)\in U\equiv \{(0,0),(0,1),(1,0),(1,1)\}$ be latent variables. Since $u$ is discrete, we take the probability mass function of $u$ as a parameter vector. For this, let $\theta=(\theta^{(0,0)},\theta^{(0,1)},\theta^{(1,0)})'\in\Theta$, where $\theta^{(0,0)}=m_\theta((Y_0,Y_1)=(0,0))$, for instance, and $\Theta=\{\theta\in [0,1]^3:\theta^{(0,0)}+\theta^{(0,1)}+\theta^{(1,0)}\le 1\}$. The Roy selection in (ref) then yields the following correspondence: \begin{align} G(u)=\begin{cases} \{(0,0),(0,1)\}&u=(0,0)\\ \{(1,1)\}&u=(0,1)\\ \{(1,0)\}&u=(1,0)\\ \{(1,0),(1,1)\}&u=(1,1). \end{cases} \end{align} The model implies a unique outcome only if the potential outcomes are ordered (e.g., an individual works in sector 1 when $Y_0=0$ and $Y_1=1$). Otherwise, it predicts multiple outcome values.

The third example is an incomplete model of an English auction Haile:2003to.

example[English auction]\rm For each auction, there are $k=1,\cdots,\bar N$ potential bidders whose valuations $u^{(k)},k=1,\cdots,\bar N$ are drawn independently from common distribution $F_\theta$ with support $[\underline u,\overline u]\subset\mathbb R,$ which is indexed by parameter $\theta\in\Theta.$ There is reserve price $r$ and minimum bid increment $\bar\Delta>0$. Each bidder's set of actions is $\{r,r+\bar\Delta,r+2\bar\Delta,\cdots, r+ K\bar\Delta\}$, where $K\in\mathbb N$ is such that $r+K\bar\Delta>\overline u$. Bidders with valuations above the reserve price bid in the auction. Let $N\le \bar N$ be the number of such bidders. Haile:2003to assumed the following weak restrictions on observed bids $s=(s^{(1)},\cdots,s^{(N)})$: (i) bidders do not bid more than their valuations, implying $s^{(k)}\le u^{(k)},k=1,\cdots,N$, and (ii) bidders do not allow an opponent to win at a price they can beat, which implies $u^{(N-1,N)}\le s^{(N,N)}+\bar\Delta$, where $x^{(k,N)}$ denotes the $k$-th (ascending) order statistic within a sample $(x^{(1)},\cdots,x^{(N)})$. Let $S=(\emptyset\cup \{r,r+\bar\Delta,r+2\bar\Delta,\cdots, r+ K\bar\Delta\})^{\bar N}$ be the set of bids. Let $U=[\underline u, \overline u]^{\bar N}$ be the set of valuations and $m_\theta=F_{\theta}^N$ be the ($N$-fold) product measure on $U$, which represents the joint distribution of private valuations. The prediction of the model is then given by \begin{align} G(u)=\big\{s\in S:s^{(k)}\le u^{(k)}, u^{(N-1,N)}\le s^{(N,N)}+\bar\Delta, k=1,\cdots,N \big\}. \end{align}

Set of Permitted Distributions and Robustness

To develop tests for incomplete models, we start by defining the family of probability distributions compatible with the model structure. For each $\theta\in\Theta$, define

align[align omitted — 165 chars of source]

where $P_u$ is a conditional law of $s$ (supported on $G(u|\theta)$), which represents the unknown {\em selection mechanism\/}. This set collects probability distributions $P$, for which one can find a suitable selection mechanism and make it consistent with a given parameter value $\theta$ and the model structure. Economic theory rarely provides any guidance on selection. The researcher therefore views any distribution in $\mathcal P_\theta$ as consistent with $\theta$.

Within this model, consider testing parameter value $\theta_0$ against another value $\theta_1$ on the basis of observed outcome $s\in S.$ This is equivalent to testing the null hypothesis, $P\in\mathcal{P}_{\theta_0}$, against the alternative hypothesis, $P\in\mathcal{P}_{\theta_1}$. Note that $\mathcal P_{\theta_0}$ and $\mathcal P_{\theta_1}$ may contain multiple (typically infinitely many) elements because the selection is left unspecified. Therefore, even for testing a single value of $\theta$ against another value, the hypotheses are composite in terms of the permitted distributions.\footnote{The composite nature of the hypotheses arises because the unknown selection is a nuisance parameter. It is possible to allow some components of structural parameter $\theta$ to be additional nuisance parameters. We analyze this extension in Section (ref).} Given this challenge, we pursue a robust approach to inference. That is, we construct tests that (i) control the size uniformly across distributions permitted under the null and (ii) maximize certain measures of power under the alternative.

Robust Tests for Incomplete Models

We provide the main theoretical results below. For this, we start with preliminaries including an introduction of the key technical tools and the important extension of the Neyman--Pearson lemma presented by Huber:1973aa. We then discuss the computational aspects of our LR tests, novel results on minimax tests, and LFPs in repeated experiments as well as a local asymptotic power analysis, which builds on our main theorem (Theorem (ref)).

Preliminaries

Belief Functions

For any $\mathcal P\subseteq\Delta(S)$, define the upper and lower probabilities of $\mathcal P$ pointwise by $\nu^{*}(A)\equiv \sup_{P \in \mathcal{P} } P(A)$ and $\nu(A)\equiv\inf_{P \in \mathcal{P}}P(A),A\subset S$, respectively. These functions are conjugate to each other in the sense that $\nu^{*}(A)=1-\nu(A^{c})$ for any $A\subset S$. Under mild restrictions on $\mathcal P$, they define set functions called {\em capacities\/}.\footnote{Appendix A provides the details. Some authors distinguish a capacity from its conjugate (co-capacity). For simplicity, we call both of these “capacities” throughout.}

For each $\theta\in\Theta$ and $A\subset S$, define $\nu_\theta$ and $\nu_\theta^*$ as the lower and upper probabilities of $\mathcal P_\theta$ defined in (ref):

align[align omitted — 192 chars of source]

The key factor for our analysis is that the lower probability $\nu_\theta$ of $\mathcal P_\theta$ is a {\em belief function\/} (or infinitely monotone capacity).\footnote{The infinite monotonicity of $\nu_\theta$ follows from philippe1999decision (Theorem 3). The foundations of belief functions are given by dempster1967upper and shafer1982belief. See Gul:2014aa and epstein2015exchangeable for the axiomatic foundations of the use of belief functions in incomplete models.} From Choquet's theorem Choquet1954,philippe1999decision,Molchanov:2006aa, it is related to the probability distribution of random set $G(u|\theta)$ as follows:

align[align omitted — 109 chars of source]

This representation allows us to obtain $\nu_\theta$ without explicitly solving the minimization (or maximization) in (ref) by computing the right-hand side of (ref) directly. Another key property of the belief function is that $P\in\mathcal P_\theta$ is equivalent to the following statement:

align[align omitted — 61 chars of source]

galichon2011set used the restrictions above to characterize the smallest possible (or “sharp”) identification region of the parameters.\footnote{galichon2011set used the conjugate of $\nu_\theta$, which yields equivalent identifying restrictions.} Following the literature, we call these the {\em sharp identifying restrictions\/} beresteanu2011sharp,Chesher:2014aa.\footnote{While the restrictions play a role in constructing robust tests, the sharp identified set does not play a role as the latter is an object of interest when the sampling processes reveals the unique data generating process in the limit, which is not guaranteed in our setting. See Epstein:2016qv for a discussion.}

Theory of Huber:1973aa

Our starting point is an analog of the Neyman--Pearson framework, which builds upon HS. For $\theta_0, \theta_1 \in \Theta$ such that $\mathcal P_{\theta_0}$ and $\mathcal P_{\theta_1}$ are disjoint, consider testing a simple null hypothesis, $H_0: \theta=\theta_0$, against a simple alternative hypothesis, $H_1:\theta=\theta_1$. In {\em complete\/} models, in which $G$ is singleton-valued, a well-defined reduced form induces a unique likelihood function tamer2003incomplete. In such settings, an optimal test is an LR test, as is well known from the Neyman--Pearson lemma. In incomplete models, however, the model generally admits a (non-singleton) set $\mathcal P_\theta$ of likelihoods, which prevents us from directly applying the Neyman--Pearson lemma.

In this setting, it is useful to consider {\em minimax tests\/} Lehmann:2006aa. Let $\phi: S \mapsto [0,1]$ denote a (possibly randomized) test. For each $P$ on $(S,\Sigma_S)$, the rejection probability of $\phi$ is

equation[equation omitted — 47 chars of source]

Let $\pi_{\theta_1}(\phi)\equiv\inf_{P_1 \in \mathcal{P}_{\theta_1}}E_P [\phi(s)]$ be the {\em lower power\/} of $\phi$ under $\theta_1$. This is the power value certain to be obtained regardless of the unknown selection mechanism. We then call test $\phi$ a {\em level-$\alpha$ minimax test\/} if it satisfies the following conditions:

equation[equation omitted — 92 chars of source]

and

equation[equation omitted — 141 chars of source]

The condition in (ref) imposes a uniform size control requirement. In (ref), tests are ranked in terms of their lower power. This reflects the researcher's preference for tests that exhibit robust power performance across selections.

A belief function (and its conjugate) is a special case of two-monotone (and two-alternating) capacities whose properties have proven powerful for conducting robust inference Huber:1981aa.\footnote{Capacity $\nu$ is said to be {\em monotone of order $k$\/} or, for short, {\em k-monotone\/} if for any $A_i\subset S,i=1\cdots,k$,

align[align omitted — 137 chars of source]

Conjugate $\nu^*(A)=1-\nu(A^c)$ is then called a {\em $k$-alternating\/} capacity.} For a class of models whose lower probabilities are two-monotone, HS showed that the rejection region of a minimax test takes the form $\{s:\Lambda(s)>t\}$ for a measurable function, $\Lambda:S\to\mathbb R$, which they called the Radon--Nikodym derivative of $\nu^*_{\theta_1}$ with respect to $\nu^*_{\theta_0}$. Further, they showed that there exists an LFP of distributions $(Q_0,Q_1)\in\mathcal{P}_{\theta_0}\times\mathcal{P}_{\theta_1}$ such that for all $t\in\mathbb{R}_+$,

equation[equation omitted — 76 chars of source]

and

equation[equation omitted — 71 chars of source]

where $\Lambda$ can be taken to be a version of the Radon--Nikodym derivative:

align[align omitted — 118 chars of source]

where $\upsilon$ is a measure that dominates $Q_j,j=0,1$. Below, we take $\upsilon$ to be the counting measure.

Heuristically, this means that $Q_0$ is the probability distribution consistent with the null parameter value, under which the size of the test is maximal. Similarly, $Q_1$ is the distribution consistent with the alternative parameter value, which is the least favorable for power maximization. The following extension of the classic Neyman--Pearson lemma (tailored to our setting) then follows from HS.

lemmaLet $\mathcal{P}_{\theta_0}$ and $\mathcal{P}_{\theta_1}$ be defined as in (ref) with $\theta=\theta_0$ and $\theta=\theta_1$, respectively. Then, there is a level-$\alpha$ minimax test $\phi$: $S \rightarrow [0,1]$ such that \begin{align} \phi(s) = \left\{ \begin{array}{ll} 1 & if \quad \Lambda(s)>C\\ \gamma & if \quad \Lambda(s)=C\\ 0 & if \quad \Lambda(s)<C,\\ \end{array} \right. \end{align} where $\Lambda\in dQ_1/dQ_0$ is a version of the Radon--Nikodym derivative of the LFP $(Q_0,Q_1)\in\mathcal P_{\theta_0}\times\mathcal P_{\theta_1}$, and $(C,\gamma)$ solves $E_{Q_0}[\phi(s)]=\alpha$.

Lemma (ref) characterizes a level-$\alpha$ minimax test as an LR test in which the ratio is formed by the LFP. Recall that $Q_1$ is the least favorable for maximizing the test's power, while $Q_0$ is the least favorable for controlling the size. Heuristically, a large value of their ratio can then be taken as evidence against the null hypothesis. Lemma (ref) states that it is indeed optimal in the minimax sense to reject $H_0$ when this ratio is sufficiently high.\footnote{In addition, the binary experiment $(S,\Sigma_S,P\in\{Q_0,Q_1\})$ in which one tests $Q_0$ against $Q_1$ is the hardest (or least informative) in terms of Bayes risk among all binary experiments such that $(S,\Sigma_S,P\in\{P_0,P_1\})$ with $P_j\in\mathcal P_{\theta_j},j=0,1$ Bednarski1982.}

Lemma (ref) is an existence and characterization result useful for obtaining the more general results below. To implement LR tests in practice, one needs to compute the LFPs (typically in a single experiment). We discuss the computational aspects in Section (ref).

Testability of Hypotheses

Before proceeding further, we comment on the testability of the hypotheses. The theory of HS requires that $\mathcal P_{\theta_0}$ and $\mathcal P_{\theta_1}$ are disjoint. Otherwise, any test is vacuous from the minimax viewpoint because probability distribution $P\in \mathcal P_{\theta_0}\cap \mathcal P_{\theta_1}$ is consistent with both hypotheses. If this is the case, we say $\theta_1$ is not {\em robustly testable\/} relative to $\theta_0$ because the lower power of any level-$\alpha$ test cannot exceed $\alpha$. This issue does not arise in complete models as long as the likelihood function satisfies $f(s;\theta_0)\ne f(s;\theta_1),a.s.$ for any $\theta_0\ne\theta_1$. One should therefore expect non-trivial lower power only if an alternative hypothesis induces set $\mathcal P_{\theta_1}$ that does not intersect with $\mathcal P_{\theta_0}$. For this, there needs to be an event $\bar A\subset S$ such that $\nu^*_{\theta_0}(\bar A)<\nu_{\theta_1}(\bar A)$ (or $\nu^*_{\theta_1}(\bar A)<\nu_{\theta_0}(\bar A)$).\footnote{In Example (ref), $\bar A=\{(1,1)\}$ (or $\bar A=\{(1,0),(0,1)\}$) constitutes such an event for testing $H_0:\theta=0$ against $H_1:\theta=\theta_1$ with $\theta_1<0$ when $u$ is continuously distributed over $\mathbb R^2$.}

The lack of robust testability is also related to the notion of {\em observational equivalence\/} Chesher:2014aa. Let $s$ follow distribution $P$ and suppose $P$ is known. Consider parameter values $\theta,\theta'\in\Theta$ such that $\theta\ne\theta'$, $P\in \mathcal P_{\theta}$, and $P\in \mathcal P_{\theta'}$. In other words, the true distribution can be justified by structure $\theta$ augmented with some selection or by another structure $\theta'$ (again augmented with some selection). When this holds, $\theta$ and $\theta'$ are said to be observationally equivalent with respect to $P$. In incomplete models, $P$ is not in general identifiable, as the sampling process does not necessarily reveal it even asymptotically Maccheroni:2005aa,Epstein:2016qv. Following Chesher:2014aa, we say that $\theta$ and $\theta'$ are {\em potentially observationally equivalent\/} if two structures are observationally equivalent for some $P$. Clearly, any pair of potentially observationally equivalent parameter values are not robustly testable, as $\mathcal P_{\theta}$ and $\mathcal P_{\theta'}$ share a distribution in common. This feature of the model raises a challenge for analyzing the local power of the tests because some local alternatives may not be robustly testable. Evaluating the power of the tests under such alternatives does not lead to a meaningful comparison. We therefore introduce a suitably modified notion of local alternatives if such an issue arises (see Section (ref)).

Computing LFPs

A key step toward implementing our tests is the computation of the LFPs, in which the sharp identifying restrictions play a role. Let $H:[0,1]\to\mathbb R$ be a twice-continuously differentiable convex function. Our proposal is to find the LFP through the following characterization:

align[align omitted — 259 chars of source]

where the constraints on $(P_0,P_1)$ are the sharp identifying restrictions.\footnote{An alternative approach would be to use the sharp identifying restrictions of beresteanu2011sharp, which also yield a finite number of linear restrictions. While we do not pursue that here, the insights presented in this paper may be useful for constructing optimal tests in models with endogeneity. Such models are studied by Chesher:2014aa, who obtained sharp identifying restrictions using generalized instrumental variables.} The number of restrictions can be reduced further by restricting the class of events to the {\em core determining class\/} galichon2011set,Luo:2017aa. This is a convex program with a convex objective function and linear constraints.\footnote{The convexity of the objective function follows from the convexity of the perspective $g(x,t)= tH(x/t)$ on its domain Boyd:2004jk.}

Our proposal builds on Theorem 6.1 in HS, which characterizes the LFP as a solution to a more general and abstract optimization problem in which $(P_0,P_1)$ is constrained to $\mathcal P_{\theta_0}\times\mathcal P_{\theta_1}$. However, because of the presence of an unknown selection in the definition of $\mathcal P_\theta$ (see (ref)), directly imposing such constraints does not lead to a tractable program. Restating the constraints using the sharp identifying restrictions, we may reduce the problem to a convex one with linear constraints, which can then be solved using efficient algorithms Boyd:2004jk. Computing $\nu_{\theta_0}$ and $\nu_{\theta_1}$ in the constraints is often straightforward (see Example (ref) and the supplementary material of Epstein:2016qv).

Emphasizing the role of the sharp identifying restrictions is worthwhile. Instead of using them to characterize the set of identifiable parameter values, we use them to obtain the LFP. To the best of our knowledge, this way of using the sharp identifying restrictions is new. Further, imposing only a subset of them in (ref) does not generally yield an LFP. In this sense, these restrictions are crucial for robust and optimal inference.

remark\rm Since $S$ is finite, the program in (ref) can be simplified further. Let $p_0$ denote the probability mass function of $P_0\in \mathcal P_{\theta_0}$ and $p_1$ be defined similarly. For simplicity, suppose $p_0(s)>0$ for all $s\in S$ and let $H(x)=-\ln x$. Then, one may solve \begin{align} (q_0,q_1)=\operatorname*{arg\,min}_{(p_0,p_1)\in \Delta(S)^2}& \sum_{s\in S}\ln\Big(\frac{p_0(s)+p_1(s)}{p_0(s)}\Big)(p_0(s)+p_1(s)) \\ s.t. & \nu_{\theta_0}(A)\le \sum_{s\in A}p_0(s), A\subset S \notag\\ & \nu_{\theta_1}(A)\le \sum_{s\in A}p_1(s), A\subset S \notag. \end{align} In this finite-dimensional convex program, one minimizes Kullback--Leibler divergence $D_{KL}(p_0+p_1\|p_0)$ subject to linear constraints on $(p_0,p_1)$. One may then use efficient numerical solvers (e.g., \verb1CVX1) to obtain the LFP.

We illustrate the computation of an LFP and minimax test using Example (ref).

\setcounter{example}{0}

example[Binary response game (continued)]\rm Let $0<\alpha<1/2$. Consider testing $H_0:\theta=0$ against $H_1: \theta=\theta_1$, where $\theta_1^{(k)}<0,k=1,2$. Suppose that $u$ follows the standard bivariate normal distribution $N(0,I_2)$. It is straightforward to calculate $\nu_\theta(A)$. As discussed earlier, a key feature of the belief function is that it is related to a probability distribution of a random set in (ref). This allows us to compute $\nu_\theta(A)$ analytically. For example, let $A=\{(1,0),(1,1)\}$. From (ref) and (ref), the lower probability of $A$ is then \begin{multline} \nu_\theta\big(\{(1,0),(1,1)\}\big)=m_\theta\big(G(u|\theta)\subseteq \{(1,0),(1,1)\}\big) \\ =m_\theta\big(G(u|\theta)=\{(1,0)\}\big)+m_\theta\big(G(u|\theta)=\{(1,1)\}\big)+m_\theta\big(G(u|\theta)=\{(1,0),(1,1)\}\big)\\ =\frac{1}{4}+\frac{\Phi(\theta^{(1)})}{2}, \end{multline} where $\Phi$ is the CDF of a standard normal random variable (see Table (ref) in Appendix (ref) for $\nu_\theta(A)$ for other events). In more complex models, simulation-based methods can be used galichon2011set,Ciliberto:2009aa,Epstein:2016qv. Suppose that $\Phi(\theta^{(k)}_1)(1-\Phi(\theta^{(-k)}_1)) \le \frac{1}{4}$ for $k=1,2$.\footnote{The form of the minimax test depends on the relative magnitude of $\theta^{(1)}_1$ and $\theta^{(2)}_1$. This assumption is made for analyzing one of the subcases. See Section (ref) for the full description of the minimax test in Example (ref).} Solving (ref), we obtain the following probability mass functions of the LFP $(Q_0,Q_1)$: \begin{align} (q_0(0,0),q_0(1,1),q_0(1,0),q_0(0,1))&=\Big(\frac{1}{4},\frac{1}{4}, \frac{1}{4}, \frac{1}{4}\Big)\\ (q_1(0,0),q_1(1,1),q_1(1,0),q_1(0,1))&=\Big(\frac{1}{4},\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)}), \frac{3-4\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})}{8},\frac{3-4\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})}{8}\Big). \end{align} The LR statistic $\Lambda$ is then given by \begin{align} \Lambda(s)&= \begin{cases} 1& s=(0,0)\\ 4\Phi(\theta^{(1)})\Phi(\theta^{(2)})& s=(1,1)\\ \frac{3-4\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})}{2}&s=(1,0)\\ \frac{3-4\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})}{2}&s=(0,1). \end{cases} \end{align} An LR test based on $\Lambda$ is level-$\alpha$ when $C=\frac{3-4\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})}{2}$ and $\gamma=2\alpha$. Hence, the test can be simplified as \begin{align} \phi(s)= \begin{cases} 2\alpha & s=(1,0) or (0,1)\\ 0& otherwise. \end{cases} \end{align} This test rejects the null hypothesis with probability $\gamma=2\alpha$ when either $s=(1,0)$ or $(0,1)$ is observed. Otherwise, the null hypothesis is retained. The intuition behind this test is as follows. When $H_0$ is true ($\theta^{(j)}=0$ for both players), the model is indeed complete. The four possible outcomes occur with equal probabilities because of $u\sim N(0,I_2)$ (Figure (ref), left). When $H_1$ is true, there exists a region of incompleteness: the set of values of $u$ for which multiple equilibria $\{(1,0),(0,1)\}$ are predicted. While the model is silent about the exact allocation of the probabilities across equilibria, it predicts a higher probability of $s\in \{(1,0),(0,1)\}$ under $H_1$ than $H_0$ (Figure (ref), right). The robust LR test then interprets $s=(1,0)$ or $s=(0,1)$ as evidence of the presence of strategic interaction and rejects the null hypothesis with a positive probability. This mechanism does not rely on any knowledge of the selection. \begin{remark} \rm Consider a special case of the example above in which the alternative hypothesis is symmetric: $\theta_1^{(1)}=\theta_1^{(2)}=\theta$. The minimax test in (ref) does not depend on the value of $\theta$ under the alternative. Hence, it can be interpreted as a “uniformly most powerful” test in terms of the lower power for testing $H_0:\theta=0$ against $H_1:\theta<0$.\footnote{If $\theta_1^{(1)}=\theta_1^{(2)}$ is not imposed, the form of the minimax test depends on the relative magnitude of the interaction effects. See Table (ref).} \end{remark} \begin{figure} [htbp] \tmpsmall\sc \begin{center} \caption{Level sets of $G$ under $H_0$ (left) and $H_1$ (right)} \begin{tikzpicture} [scale=0.9,domain=-3:3,>=latex] \fill[fill=red,opacity=0.4] (0,0) -- (3,0) -- (3,3) -- (0,3) ; \fill[fill=Green,opacity=0.4] (0,0) -- (3,0) -- (3,-3) -- (0,-3) ; \fill[fill=Green,opacity=0.4] (0,0) -- (-3,0) -- (-3,3) -- (0,3) ; \draw[->] (-3,0) -- (3,0) node[right] {$u_1$}; \draw[->] (0,-3) -- (0,3) node[above] {$u_2$}; \draw (0,-0.275)node[right]{$\theta=(0,0)$} ; \draw (1,1.2)node[right]{$\{(1,1)\}$} ; \draw (1,-1.2)node[right]{$\{(1,0)\}$} ; \draw (-1,1.2)node[left]{$\{(0,1)\}$} ; \draw (-1,-1.2)node[left]{$\{(0,0)\}$} ; \draw (0,0) circle [radius = 0.035]; \end{tikzpicture} \begin{tikzpicture} [scale=0.9,domain=-3:3,>=latex] \fill[fill=red,opacity=0.4] (2,1.5) -- (3,1.5) -- (3,3) -- (2,3) ; \fill[fill=Green,opacity=0.4] (3,-3) -- (0,-3) -- (0,0) -- (-3,0) -- (-3,3) -- (0,3) -- (2,3) -- (2,0) -- (2,1.5) -- (3,1.5) ; \draw[-] (-3,0) -- (2,0) ; \draw[->] (2,0) -- (2,3) node[above] {$u_2$}; \draw[->] (0,1.5) -- (3,1.5) node[right] {$u_1$}; \draw[-] (0,-3) -- (0,1.5) ; \draw (1.9,2.1) node [right]{$\{(1,1)\}$} ; \draw (1,-1.2) node [right]{$\{(1,0)\}$} ; \draw (-1,1.2) node [left]{$\{(0,1)\}$} ; \draw (-1,-1.2) node [left]{$\{(0,0)\}$} ; \draw (0.3,0.9) node [right]{$\{(0,1),$}; \draw (0.5,0.5) node [right]{$(1,0)\}$}; \draw (2,1.5) circle [radius = 0.035]; \draw (2,1.8)node[left]{$(-\theta^{(1)},-\theta^{(2)})$} ; \end{tikzpicture} \end{center} Note: The area in red represents the values of $u$ under which $\{(1,1)\}$ is predicted. The area in green represents the values of $u$ under which $\{(0,1)\}$, $\{(1,0)\}$, or their union is predicted. \end{figure}

Minimax Tests in Repeated Experiments

The sequence of outcomes $s^n=(s_1,s_2,\cdots,s_n)$ is commonly generated from repeated experiments. In this section, we present a set of theoretical results that characterize a minimax test in such a setting and provide an asymptotic Gaussian approximation to its (upper) rejection probability.

For any set $B$, let $B^n$ denote the $n$-fold Cartesian product of $B$. For each $n\in\mathbb N$, let $S^n$ and $U^n$ be the sets of outcome sequences $s^n=(s_1,s_2,\cdots,s_n)$ and latent variable sequences $u^n=(u_1,\cdots,u_n)$, respectively. Below, we use the abbreviation $m^n$ to denote the family $\{m^n_{\theta}\}_{\theta \in \Theta}$ of $u^n$'s joint laws permitted by the model. We then make the following assumption on each member of this family.

assumptionFor each $\theta\in\Theta$, $m^n_\theta\in\Delta(U^n)$ is a product measure.

This assumption requires that $u_i$'s are distributed independently across experiments. A leading case is that $(u_1,\dots,u_n)$ are i.i.d. This can also accommodate heteroskedasticity and other types of heterogeneity across cross-sectional units or clusters of them (e.g., group-specific effects).

Without further assumptions, $s^n$ takes values in the Cartesian product of the sets of permissible outcome values:

align[align omitted — 74 chars of source]

where $G(\cdot|\theta)$ is given as in (ref).\footnote{It is possible to allow the functional form of $G(u_i|\theta)$ to vary across $i$ as well. For notational simplicity, we do not explicitly consider this extension here. We introduce the heterogeneity of $G$ due to covariates in Section (ref).} This set collects outcome sequences that are compatible with the model and $\theta$. We represent the repeated experiments by the tuple $(S^n,U^n,\Theta, G^n;m^n).$ Although $u^n$ is assumed to be independent, the outcome sequence $s^n$ can be dependent because the model does not restrict the selection mechanism. Similarly, even if one makes the stronger assumption that the $u_i$'s are i.i.d., the distribution of outcome $s_i$ may be heterogeneous because of the potential heterogeneity of selection across experiments. This feature arises because the joint selection mechanism (across all experiments) is left unspecified and the joint distribution of the outcome sequence depends on this incidental parameter.

For each $\theta\in\Theta$ and $n\in \mathbb N$, the set of distributions compatible with the model is

align[align omitted — 162 chars of source]

This set collects all the distributions of $s^n$ consistent with $\theta$. $P_u$ is unrestricted in the sense that the selection may be heterogeneous and dependent across experiments. Hence, $\mathcal P^n_\theta$ contains a broad range of distributions that can exhibit arbitrary dependence and heterogeneity. In particular, $\mathcal P^n_\theta$ allows measures under which the distributions of sample moments are not well approximated by classical limit theorems---even in large samples (see Epstein:2016qv).

Finding the LFP in such a rich set of distributions may be challenging. However, under Assumption (ref) and with the correspondence in (ref), the model has a tractable “product” structure, which significantly simplifies the characterization of the LFPs.

Let $\nu_\theta^{n,*}$ and $\nu_\theta^n$ denote the upper and lower probabilities of $\mathcal P_\theta^n$. For each $i\in\{1,\dots,n\}$ and $\theta\in\Theta$, let

align[align omitted — 171 chars of source]

where $m_{\theta,i}$ is the $i$-th marginal distribution of $m_\theta^n.$ The following theorem shows that the minimax test in repeated experiments is an LR test and that the LFPs are product measures.

theoremSuppose Assumption (ref) holds. Then, (i) an LFP $(Q_{0}^n,Q_{1}^n)\in\mathcal P_{\theta_0}^n\times\mathcal P_{\theta_1}^n$ exists such that for all $t\in\mathbb R_+$, \begin{align} \nu^{*,n}_{\theta_0}(\Lambda_n>t)&=Q_0^n(\Lambda_n>t)\\ \nu^{n}_{\theta_1}(\Lambda_n>t)&=Q_1^n(\Lambda_n>t), \end{align} where $\Lambda_n$ is a version of the Radon--Nikodym derivative $dQ_1^n/dQ_0^n$. The LFP consists of the product measures: \begin{align} Q^n_0=\bigotimes_{i=1}^n Q_{0,i}, and Q^n_1=\bigotimes_{i=1}^n Q_{1,i}, \end{align} where, for each $i\in\mathbb N$, $(Q_{0,i},Q_{1,i})\in \mathcal P_{\theta_0,i}\times\mathcal P_{\theta_1,i}$ is the LFP in the $i$-th experiment: (ii) A minimax test $\phi_n$: $S^n \rightarrow [0,1]$ can be constructed as \begin{align} \phi_n(s^n) = \left\{ \begin{array}{ll} 1 & if \quad \Lambda_n(s^n)>C_n\\ \gamma_n & if \quad \Lambda_n(s^n)=C_n\\ 0 & if \quad \Lambda_n(s^n)<C_n,\\ \end{array} \right. with \Lambda_n(s^n)=\prod_{i=1}^{n} \Lambda_i, \end{align} where $\Lambda_i\in dQ_{1,i}/dQ_{0_i}$ for all $i$, and $C_n$ and $\gamma_n$ are chosen so that $E_{Q_0^n}[\phi_n(s^n)]=\alpha$.

The LFP consists of the product measures.\footnote{This result does not follow from Corollary 4.2 of HS, who assumed that a sample is independently distributed (p. 258).} Heuristically, this means that either for controlling size or maximizing power, the least favorable distribution in $\mathcal P^n_{\theta_0}$ (or $\mathcal P^n_{\theta_1}$) is a law that multiplies up the least favorable distributions in the individual experiments. When the $u_i$'s are i.i.d., this characterization has a particularly useful implication for the implementation. To construct a minimax test, it suffices to find the LFP $(Q_0,Q_1)\in\mathcal P_{\theta_0}\times\mathcal P_{\theta_1}$ in a {\em single\/} experiment $(S,U,\Theta, G,m)$. One may then obtain the LR statistic by taking the product of their ratios across experiments. We state this result as a corollary.

corollarySuppose $(u_1,\dots,u_n)$ are identically and independently distributed. Then, a minimax test $\phi_n$: $S^n \rightarrow [0,1]$ can be constructed as \begin{align} \phi_n(s^n) = \left\{ \begin{array}{ll} 1 & if \quad \Lambda_n(s^n)>C_n\\ \gamma_n & if \quad \Lambda_n(s^n)=C_n\\ 0 & if \quad \Lambda_n(s^n)<C_n,\\ \end{array} \right. with \Lambda_n(s^n)=\prod_{i=1}^{n}\Lambda_i, \end{align} where $\Lambda_i\in dQ_{1}/dQ_{0}$, $(Q_0,Q_1)\in\mathcal P_{\theta_0}\times \mathcal P_{\theta_1}$ is the LFP in $(S,U,\Theta, G,m)$, and $C_n$ and $\gamma_n$ are chosen so that $E_{Q_0^n}[\phi_n(s^n)]=\alpha$.

The LFP consisting of the product measures in Theorem (ref) provides an important link through which we may connect the incomplete model to standard frameworks. Below, we demonstrate this by studying the large-sample approximations and asymptotic local power of the tests. While both issues can be analyzed without assuming identically distributed latent variables, we maintain this assumption to simplify the notation and our analysis.\footnote{If $u_i$ is not identically distributed, one should invoke a central limit theorem for independent and not identically distributed (i.n.i.d.) sequences under $Q_0^n$ White:2001aa to obtian a Gaussian approximation. Similarly, for local approximations of experiments with an i.n.i.d. sequence, Rieder_1994 (Section 2.3) provides a general framework, which can be applied to the sequence $\{\mathsf Q^n_{\theta_{n,\xi,h}}\}$ defined below.}

Gaussian Approximation

One of the consequences of the product structure is that the upper rejection probability of $\phi_n$ admits a Gaussian approximation in large samples. For ease of exposition, we assume that the $u_i$'s are i.i.d., which in turn implies that $Q_0^n$ is an i.i.d. law from Corollary (ref). Hence, properly normalized sample averages follow classical limit theorems under this law. We use this insight to obtain an asymptotically valid critical value. Since $Q_0^n$ is the least favorable, the asymptotic size is controlled under any distribution under the null hypothesis.

Let $z_\alpha$ be the $1-\alpha$ quantile of the standard normal distribution and let $\Lambda\in dQ_1/dQ_0$. For each $n$, let

align[align omitted — 99 chars of source]

where $\mu_{Q_0}\equiv E_{Q_0}[\ln \Lambda(s)],$ and $\sigma^2_{Q_0}\equiv\text{Var}_{Q_0}(\ln\Lambda(s))$. Observe that $\mu_{Q_0}$ and $\sigma_{Q_0}$ depend only on the LFP but not on the unknown DGP. Once the LFP is found, computing $\mu_{Q_0}$ and $\sigma_{Q_0}$ is straightforward because $Q_0$ is a discrete distribution and $\Lambda$ is known. The critical value in (ref) is constructed in such a way that the following convergence holds:

align[align omitted — 168 chars of source]

where $Z$ is a standard normal random variable. This critical value is computed without any resampling or simulation and therefore can be done so easily. Despite its simplicity, it has the advantage of being asymptotically valid even if the true DGP is highly heterogeneous and dependent.

Let $\phi^*_n$ be a test that rejects the null hypothesis if and only if $\Lambda_n>C^*_n$. The following proposition then follows.

propositionSuppose Assumption (ref) holds and that $0\le\sigma^2_{Q_0}<\infty$. Then, the test controls the asymptotic size: \begin{align} \limsup_{n\to\infty}\sup_{P\in\mathcal P^n_{\theta_0}}E_{P}[\phi^*_n(s^n)]\le \alpha. \end{align} Furthermore, (ref) holds with equality when $\sigma^2_{Q_0}>0$.

Asymptotic Local Power

Building on Theorem (ref), we analyze the asymptotic local power of the tests. In what follows, suppose that $\Theta$ is a subset of Euclidean space $\mathbb R^d$ and let $\varphi:\Theta\to\mathbb R$ be a continuously differentiable function with gradient $\dot\varphi_{\theta}:\mathbb R^d\to\mathbb R$. Consider the following hypotheses:

align[align omitted — 98 chars of source]

Various hypotheses of empirical interest can be formulated in this way. Our goal here is to characterize the upper envelope of the lower power of the tests for (ref) in an asymptotic framework. Theorem (ref) serves as a building block for this purpose, as it allows us to embed our problem into a more standard one. In this section, we assume that $u_i,i=1,\dots,n$ are i.i.d. throughout.

We consider localized experiments. Let $\theta_0\in\Theta$ be a parameter such that $\varphi(\theta_0)=0$ and let $\{\theta_{n,\xi,h}\},(\xi,h)\in\mathbb R^d\times \mathbb R^d$ be a sequence of alternative parameter values, which we specify below. We call $\xi$ a {\em fixed shift\/} and $h$ a {\em local parameter\/}. The sequence of parameters induces a sequence of belief functions $\nu^n_{\theta_{n,\xi,h}}$. Suppose that for each $n$ and $(\xi,h)$, the conditions of Theorem 3.1 hold for $\nu^n_{\theta_0}$ and $\nu^n_{\theta_{n,\xi,h}}$. Then, there exists an LFP $(Q^n_0,Q^n_1)\in \mathcal P^n_{\theta_0}\times \mathcal P^n_{\theta_{n,\xi,h}}$. If there exists model $\theta\mapsto \mathsf Q_\theta$ indexed by $\theta$ defined on a neighborhood of $\theta_0$ such that $Q_0^n=\mathsf Q^n_{\theta_0}$ and $Q_1^n=\mathsf Q^n_{\theta_{n,\xi,h}}$ for all $n$, one may consider the following sequence of experiments:

align[align omitted — 129 chars of source]

In this hypothetical environment, observations are generated using a sequence of probability laws that are the least favorable for testing $\theta_0$ against $\theta_{n,\xi,h}$. The fact that the experiments are characterized by probabilities instead of capacities allows us to employ standard asymptotic tools. In particular, we employ the {\em limits of experiments\/} argument in the style of Le-Cam:1972aa,LeCam:1986aa. Heuristically, if one wants to obtain an asymptotic upper envelope of $\pi_{\theta_{n,\xi,h}}$, one may consider the least favorable sequence of DGPs $\{\mathsf Q^n_{\theta_{n,\xi,h}}\}$ for power maximization. It turns out that the lower power of any test can be matched by a power function of a limit experiment, which is often more straightforward to analyze. The limit experiment can then be used to derive an asymptotic power envelope.

While the argument above suggests that we may use the standard limits of experiments framework, a few non-standard features arise. First, the underlying model may not satisfy the differentiability in quadratic mean condition, which is sufficient for the LAN of the experiments. Instead, the model is typically directionally differentiable (in the $L^2$ sense) and satisfies the LAN property separately on a finite number of convex cones that partition the local parameter space, which requires us to consider sub-experiments of (ref). Second, some alternatives are not robustly testable. Hence, to conduct a meaningful power analysis, one needs to construct local alternatives with care.\footnote{We modify the definition of the local alternative so that the sequences of LFPs $\{Q^n_{\theta_0}\}$ and $\{Q^n_{\theta_{n,\xi,h}}\}$ are contiguous.}

Asymptotic Power against Robustly Testable Local Alternatives

Recall that in Example (ref), setting the strategic interaction effects to $\theta_1<0$ made $\mathcal P_{\theta_0}$ and $\mathcal P_{\theta_{1}}$ disjoint no matter how small the deviation from the null $\theta_0=0$ was. Below, we start with a relatively simple setting in which $H_0:\theta=\theta_0$ and $H_1:\theta=\theta_0+h/\sqrt n$ induce sets of distributions that are disjoint for any $n$. In this setting, it suffices to index local alternatives using $h$ only, and hence we use $\theta_{n,h}$ instead of $\theta_{n,\xi,h}$.\footnote{We elaborate on the role of $\xi$ in the next section.}

We define the notion of differentiability below. For this, let $L^2_{P}$ denote the set of square integrable functions on $S$ with respect to measure $P$. For each $f,g\in L^2_P$, we then let $\langle f,g\rangle_{L^2_P}$ and $\|f\|_{L^2_P}$ denote the $L^2$ inner product and $L^2$ norm, respectively. Consider a parametric family of distributions $\{P_\theta,\theta\in V\}$, where $V$ is an open subset of $\Theta.$

definitionLet $\theta\mapsto P_\theta$ be a model such that $P_\theta$ is absolutely continuous with respect to a $\sigma$-finite measure $\mu$ on $S$. The model is said to be {\em $L^2$ differentiable\/} at $\theta\in V$ tangentially to set $\mathcal T\subset\mathbb R^d$ if there exists a square integrable (w.r.t. $P_\theta$) function $\dot\ell_{\theta}:S\to\mathbb R^d$ such that, for every $h\in \mathcal T$, \begin{align} \Big\|p_{\theta+\tau h}^{1/2}-p_\theta^{1/2}(1+\frac{1}{2}\tau h'\dot\ell_\theta)\Big\|_{L^2_\mu}=o(\tau), \end{align} as $\tau\downarrow0$, where $p_{\theta}=dP_{\theta}/d\mu$ for all $\theta$.

This is the $L^2$ differentiability of the square root density commonly used in the literature. If one could embed the LFPs into an $L^2$ differentiable model $\theta\mapsto\mathsf Q_\theta$ with $\mathcal T$ being a suitable limit of the local parameter space $\{h\in \sqrt n(\Theta-\theta_0):\varphi(\theta_0+h/\sqrt n)>0\}$, asymptotic local power can be analyzed in a standard way. However, as we show below, the model is often only directionally differentiable. That is, the form of the derivative, $\dot\ell_\theta$, varies across the subsets of $\mathcal T$ (see the discussions below). The following high-level assumption states this formally.\footnote{For the results that follow, it suffices that a model satisfies Assumption (ref) for an LFP. The LFPs are unique up to the Radon--Nikodym derivative (HS, 1973), and thus they all lead to the same quadratic expansion of the log-LR process. While one could alternatively take (ref) as a high-level condition, we do not do so because Assumption (ref) is often easier to check. A similar comment applies to Assumption (ref).} For this, let $\stackrel{P^n}{\leadsto}$ denote weak convergence under the sequence $\{P^n\}$ of distributions. Let $\mathbb C(0,\epsilon)$ denote an open cube centered on the origin with edges of length $2\epsilon.$ A set $\Gamma\subset\mathbb R^d$ is said to be locally equal to set $\Upsilon\subset\mathbb R^d$ if $\Gamma\cap \mathbb C(0,\epsilon)=\Upsilon\cap \mathbb C(0,\epsilon)$ for some $\epsilon>0$ Andrews:1999aa.

assumption[Local parameter cones and $L^2$ directional differentiability] (i) Set $\{\xi\in \Theta-\theta_0:\varphi(\theta_0+\xi)>0\}$ is locally equal to convex cone $\mathcal T(\theta_0)$; (ii) There exists set $\mathbb J$ and a collection of convex cones (containing 0) $\{\mathcal T_{j}(\theta_0),j\in \mathbb J\}$ such that $\mathcal T_{j}(\theta_0)\cap\mathcal T_{j'}(\theta_0)=\{0\},\forall j\ne j'$ and $\bigcup_j\mathcal T_{j}(\theta_0)=\mathcal T(\theta_0)$; (iii) For each $j\in\mathbb J$, there exists model $\theta\mapsto \mathsf Q_{j,\theta}$ defined on a neighborhood of $\theta_0$ such that an LFP $(Q_{0,j,\tau h},Q_{1,j,\tau h})\in \mathcal P_{\theta_0}\times\mathcal P_{\theta_0+\tau h}$ satisfies \begin{align} Q_{0,j,\tau h}=\mathsf Q_{j,\theta_0}, and Q_{1,j,\tau h}=\mathsf Q_{j,\theta_0+\tau h}, \end{align} for all $\tau \in(0,\bar \tau]$ for some $\bar \tau>0$, and $\theta\mapsto \mathsf Q_{j,\theta}$ is $L^2$ differentiable at $\theta_0$ tangentially to $\mathcal T_{j}(\theta_0).$

The assumption above imposes a regularity condition on the LFP for testing $H_0:\theta=\theta_0$ against $H_1:\theta=\theta_{0}+\tau h$ for $\tau>0$. While $\theta_0$ is fixed, both the least favorable distributions (under $H_0$ and $H_1$) may depend on deviation $\tau h$. Hence, we index the LFP using $\tau h$. The restrictions in (ref) require that the least favorable distribution $Q_{0,j,\tau h}$ under $H_0$ remains the same for all sufficiently small $\tau$, and this is given by point $\mathsf Q_{j,\theta_0}$ in the $L^2$ differentiable model $\theta\mapsto \mathsf Q_{j,\theta}$. As we demonstrate below through examples, Assumption (ref) (and similarly Assumption (ref)) can be checked by analyzing the LFP. In all our examples, the cardinality of $\mathbb J$ is finite. Appendix (ref) also provides the primitive conditions that ensure the key condition ($L^2$ differentiability) (see Assumption (ref), Proposition (ref), and Corollary (ref)).

Below, we let $\dot \ell_{j,\theta_0}$ denote the $L^2$ derivative defined for $h\in\mathcal T_{j}(\theta_0)$. Under Assumption (ref), one may expand the log-LR for every $h\in \mathcal T_{j}(\theta_0)$ as follows:

align[align omitted — 172 chars of source]

where $\Delta_{j,n}\stackrel{\mathsf Q^n_{j,\theta_0}}{\leadsto} \Delta_j\sim N(0,C_{j})$ and $C_j=E_{\mathsf Q_{j,\theta_0}}[\dot\ell_{j,\theta_0}\dot\ell_{j,\theta_0}'].$ In what follows, we call $\Delta_{j,n}$ the central sequence (or normalized score) and $C_{j}$ the information matrix. Consider the following sub-experiments:

align[align omitted — 156 chars of source]

When Assumption (ref) holds, for each $j$, the limit of the experiments (as $n\to\infty$) is

align[align omitted — 130 chars of source]

which is also equivalent to $(\mathbb R^d,\Sigma_{\mathbb R^d},N( h,C_{j}^{-1}): h\in \mathcal T_{j}(\theta_0))$ if $C_j$ is non-singular Van-der-Vaart:2000aa. In other words, the experiment is equivalent to the one in which the researcher observes a single normal random vector whose mean and variance are $h\in \mathcal T_{j}(\theta_0)$ and $C_j^{-1}$, respectively. The asymptotic local (lower) power of a test is then bounded from above by the corresponding power in the limit experiment. The power envelope can be derived by considering the highest possible power for testing $H_0:\dot\varphi_{\theta_0} h\le 0$ against $H_1:\dot\varphi_{\theta_0} h>0$ at level-$\alpha.$

As is well known, these limit experiments are Gaussian shift experiments defined on suitable subsets of $\mathbb R^d$. If Assumption (ref) holds with $\mathbb J=\{ \textup{\uppercase\expandafter{\romannumeral1}} \}$ and $\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta_0}=\mathbb R^d$, we obtain the LAN LeCam:1986aa. Assumption (ref) slightly extends the LAN, and this allows us to consider experiments defined separately on the subsets (cones) of the local parameter space. This extension is motivated by the fact that central sequence $\Delta_{j,n}$ and information matrix $C_{j}$ may differ across the local parameter cones. This is because as $h$ varies (and hence $\nu_{\theta_{n,h}}$ varies), the LFP defined through the convex program in (ref) may change in a non-differentiable (but directionally differentiable) way.\footnote{See Shapiro:1988aa, Dempe_Directional_1993, and the references therein for the directional differentiability of solutions to parametric convex programs. Our primitive conditions (see Appendix (ref)) for Assumptions (ref) and (ref) are based on Shapiro:1988aa.} This may lead to distinct central sequence and information matrix pairs across the cones. Owing to this non-standard feature, our results below concern asymptotically optimal statistical decisions when the underlying model may only be directionally differentiable (in the $L^2$ sense). This complements the recent developments on statistical inference and decisions in non-standard models in which the parameters of interest are only directionally differentiable, whereas the underlying model is regular Hirano_Impossibility_2012,Song2014JMA,Fang:2014aa,Fang2018,HongLi2018JoE.

To characterize the power envelope, we define the tangent cone of the score functions and efficient influence function of $\varphi$. For each $j\in\mathbb J$, let

align[align omitted — 141 chars of source]

We call the set above the {\em tangent cone\/} of the model. The influence curve $\varrho_j\in L^2_{\mathsf Q_{j,\theta_0}}(S)$ of $\varphi$ is such that, for any $g\in\mathcal G_{j,\theta_0},$

align[align omitted — 158 chars of source]

as $\tau\downarrow 0.$ The {\em efficient influence function\/} (or canonical gradient) $\tilde\varrho_j$ of $\varphi$ is then defined as the projection of $\varrho_j$ on the closure of $\mathcal G_{j,\theta_0}$ (which is often called the {\em tangent set\/}).

The following theorem characterizes the asymptotic upper bound of the lower power and provides a test that achieves the bound (for a given cone).

theoremSuppose Assumption (ref) holds. Let $j\in\mathbb J.$ Suppose that $\varphi$ is such that $\varphi(\theta_0)=0$ and has influence curve $\varrho_j$. Let $\phi_n$ be a level-$\alpha$ test for $H_0:\varphi(\theta)\le 0$ against $H_1:\varphi(\theta)>0$ and $\pi_{n,\theta}(\phi_n)$ be its lower power under $\nu_{\theta}^n$. Then, for any $h\in\mathcal T_{j}(\theta_0)$, \begin{align} \limsup_{n\to\infty}\pi_{n,\theta_0+h/\sqrt n}(\phi_n)\le 1-\Phi\Bigg(z_\alpha-\frac{\langle \varrho_j,h'\dot\ell_{j,\theta_0}\rangle_{L^2_{\mathsf Q_{j,\theta_0}}}}{\|\tilde{\varrho}_j\|_{L^2_{\mathsf Q_{j,\theta_0}}}}\Bigg). \end{align} Let $T_{j,n}$ be a statistic such that \begin{align} T_{j,n}=\frac{n^{-1/2}\sum_{i=1}^n\tilde\varrho_j(s_i)}{\|\tilde\varrho_j\|_{L^2_{\mathsf Q_{j,\theta_0}}}}+o_{\mathsf Q^n_{j,\theta_0}}(1). \end{align} Let $\phi^*_{j,n}$ be a test that rejects the null hypothesis iff $T_{j,n}\ge z_\alpha.$ Then, the test is of asymptotically level-$\alpha$ and, for any $h\in \mathcal T_{j}(\theta_0)$, \begin{align} \lim_{n\to\infty}\pi_{n,\theta_0+h/\sqrt n}(\phi^*_{j,n})=1-\Phi\Bigg(z_\alpha-\frac{\langle \varrho_j,h'\dot\ell_{j,\theta_0}\rangle_{L^2_{\mathsf Q_{j,\theta_0}}}}{\|\tilde{\varrho}_j\|_{L^2_{\mathsf Q_{j,\theta_0}}}}\Bigg). \end{align}

The asymptotic power envelope in (ref) coincides with that for the one-sided test $\dot\varphi_{\theta_0} h\le 0$ against $\dot\varphi_{\theta_0} h>0$ in the Gaussian shift experiment. The theorem also implies that the test based on the (rescaled) efficient influence function achieves the power envelope toward alternatives with $h\in\mathcal T_{j}(\theta_0)$.\footnote{The power envelope can be expressed using $\dot\varphi h$ and $C_j$ as in Theorem 15.4 in Van-der-Vaart:2000aa if $C_j$ is non-singular and $\mathcal G_{j,\theta_0}$ is a linear subspace. The description above is slightly more general to handle cases in which $C_j$ may be singular and $\mathcal G_{j,\theta_0}$ is a convex cone Rieder2014.} Therefore, the key factor is to find the efficient influence function.

We next revisit Example (ref). \setcounter{example}{0}

example[Binary response game (continued)]\rm Let $\Theta=\{\theta\in\mathbb R^2:\theta^{(1)}\le 0, \theta^{(2)}\le 0\}$. Consider testing the hypothesis as in (ref) with $\varphi(\theta)=p'\theta=-\theta^{(1)}-\theta^{(2)}$, where $p=(-1,-1)'$. To localize the experiment at $\theta_0=(0,0)'$, we start by describing the LFPs (and minimax tests) for all the parameter values under consideration. Let $\Theta_1\equiv\{\theta\in \Theta:\varphi(\theta)>0\}$ be the set of parameter values under the alternative. Let $\mathbb J=\{ \textup{\uppercase\expandafter{\romannumeral1}} , \textup{\uppercase\expandafter{\romannumeral2}} , \textup{\uppercase\expandafter{\romannumeral3}} \}$. Then, define the following parameter sets: \begin{align} \Theta_{ \uppercase\expandafter{\romannumeral1} }&\equiv\big\{\theta_1\in\Theta_1:\Phi(\theta_1^{(1)})(1-\Phi(\theta_1^{(2)})) \le \frac{1}{4}, \Phi(\theta_1^{(2)})(1-\Phi(\theta_1^{(1)})) \le \frac{1}{4}\big\}\\ \Theta_{ \uppercase\expandafter{\romannumeral2} }&\equiv\big\{\theta_1\in\Theta_1:\Phi(\theta_1^{(1)})(1-\Phi(\theta_1^{(2)})) > \frac{1}{4} \big\}\\ \Theta_{ \uppercase\expandafter{\romannumeral3} }&\equiv\big\{\theta_1\in\Theta_1: \Phi(\theta_1^{(2)})(1-\Phi(\theta_1^{(1)})) > \frac{1}{4} \big\}. \end{align} Figure (ref) (left) shows these sets. In addition to the case in which $\theta_1^{(1)}$ and $\theta_1^{(2)}$ are both strictly negative and comparable in magnitude (as discussed in Section (ref)), we consider two other cases here. Across all subcases, the density of $Q_0$ is $(q_0(0,0),q_0(1,1),q_0(1,0),q_0(0,1))=(\frac{1}{4},\frac{1}{4}, \frac{1}{4}, \frac{1}{4})$. However, $Q_1$ and the minimax test vary (Table (ref)).\footnote{Proposition (ref) in Appendix (ref) provide these results formally.} If $\theta_1^{(1)}$ is substantially smaller than $\theta_1^{(2)}$ (i.e., $\theta_1\in\Theta_{ \textup{\uppercase\expandafter{\romannumeral2}} }$), a larger mass moves to the region over which $s=(1,0)$ is predicted than to the region over which $s=(0,1)$ is predicted (recall Figure (ref)). The minimax test then rejects $H_0$ with a positive probability ($\gamma=4\alpha$) only when $s=(1,0)$ is observed. A similar comment applies to the case in which $\theta\in\Theta_{ \textup{\uppercase\expandafter{\romannumeral3}} }$. \begin{table}[htbp] \begin{center} \caption{$Q_1$ and minimax tests} \resizebox{6.75in}{!}{ \begin{tabular}{cccccl} \hline\hline & \multicolumn{4}{c}{$Q_1$} & \\ \cline{2-5} Sets & $q_1(0,0)$ & $q_1(1,1)$ & $q_1(1,0)$ & $q_1(0,1)$ & \multicolumn{1}{c}{Minimax test} \\ \hline $\Theta_{ \textup{\uppercase\expandafter{\romannumeral1}} }$ & $\frac{1}{4}$ & $\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})$ & $\frac{3}{8}-\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})/2$ & $\frac{3}{8}-\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})/2$ & $ \phi(s)= \begin{cases} 2\alpha & s\in\{(1,0),(0,1)\}\\ 0& \text{otherwise} \end{cases}$ \\ $\Theta_{ \textup{\uppercase\expandafter{\romannumeral2}} }$ & $\frac{1}{4}$ & $\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})$ & $\frac{1}{4}-\Phi(\theta_1^{(1)}) (\Phi(\theta_1^{(2)})-\frac{1}{2})$ & $\frac{1}{2}(1-\Phi(\theta_1^{(1)}))$ & $\phi(s)= \begin{cases} 4\alpha & s=(1,0)\\ 0& \text{otherwise} \end{cases}$ \\ $\Theta_{ \textup{\uppercase\expandafter{\romannumeral3}} }$ & $\frac{1}{4}$ & $\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})$ & $\frac{1}{2}(1-\Phi(\theta_1^{(2)}))$ & $\frac{1}{4}-\Phi(\theta_1^{(2)}) (\Phi(\theta_1^{(1)})-\frac{1}{2})$ & $\phi(s)= \begin{cases} 4\alpha & s=(0,1)\\ 0& \text{otherwise} \end{cases}$\\ \hline\hline \end{tabular}} \end{center} \end{table} \begin{figure} [htbp] \begin{center} \tmpsmall\sc \caption{Subcases in Proposition 1 and local parameter cones} \tmpsmall\sc Left panel: Parameter space and subcases; Right panel: Local parameter cones at $\theta_0=(0,0)'$: $\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta_0}$ (solid line), $\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral2}} }(\theta_0)$ (red), and $\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral3}} ,\theta_0}$ (blue). \end{center} \end{figure} Consider a local alternative $\theta_{n,h}=\theta_0+h/\sqrt n$ such that $\varphi(\theta_{n,h})>0$. Suppose that $\theta_{n,h}\in \Theta_j$ for some $j\in\mathbb J$ for all sufficiently large $n$. Then, the local parameter must belong to one of the following cones: \begin{align} \mathcal T_{ \uppercase\expandafter{\romannumeral1} }(\theta_0)&=\{h\in\mathbb R^2:h=(\bar h,\bar h)',\bar h\in(-\infty,0] \}\\ \mathcal T_{ \uppercase\expandafter{\romannumeral2} }(\theta_0)&=\{h\in\mathbb R^2:h=(h^{(1)},h^{(2)})',-\infty<h^{(2)}<h^{(1)}\le 0\}\\ \mathcal T_{ \uppercase\expandafter{\romannumeral3} }(\theta_0)&=\{h\in\mathbb R^2:h=(h^{(1)},h^{(2)})',-\infty<h^{(1)}<h^{(2)}\le 0\}. \end{align} These cones are localized versions of the parameter subsets ($\Theta_{ \textup{\uppercase\expandafter{\romannumeral1}} }$-$\Theta_{ \textup{\uppercase\expandafter{\romannumeral3}} }$), as shown in Figure (ref) (right). Below, as an example, we consider the case $\theta_{n,h}\in\Theta_{ \textup{\uppercase\expandafter{\romannumeral2}} }$ for all sufficiently large $n$. Since $Q_1$ is as shown in Table (ref), one may embed the LFP into model $\theta\mapsto \mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\theta}$ whose density is \begin{multline} (\mathsf q_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\theta}(0,0),\mathsf q_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\theta}(1,1),\mathsf q_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\theta}(1,0),\mathsf q_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\theta}(0,1))\\ =\Big(\frac{1}{4},\Phi(\theta^{(1)})\Phi(\theta^{(2)}),\frac{1}{4}-\Phi(\theta^{(1)}) (\Phi(\theta^{(2)})-\frac{1}{2}),\frac{1}{2}(1-\Phi(\theta^{(1)}))\Big). \end{multline} Then, the $L^2$ derivative of the model for $h\in\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral2}} }(\theta_0)$ is \begin{align} \dot \ell_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\theta_0}(s) &=1\{s=(1,1)\}\begin{pmatrix} \frac{2}{\sqrt{2\pi}}\\ \frac{2}{\sqrt{2\pi}} \end{pmatrix} +1\{s=(1,0)\}\begin{pmatrix} 0\\ \frac{-2}{\sqrt{2\pi}} \end{pmatrix}+1\{s=(0,1)\} \begin{pmatrix} \frac{-2}{\sqrt{2\pi}}\\ 0 \end{pmatrix}. \end{align} The log-likelihood function can be expanded as in (ref) with $\Delta_{ \textup{\uppercase\expandafter{\romannumeral2}} ,n}=\frac{1}{\sqrt n}\sum_{i=1}^n\dot \ell_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\theta_0}$ and the information matrix: $C_{ \textup{\uppercase\expandafter{\romannumeral2}} } =\big( \begin{smallmatrix} \frac{1}{\pi}&\frac{1}{2\pi}\\ \frac{1}{2\pi}&\frac{1}{\pi} \end{smallmatrix}\big). $ Therefore, the limit experiment is \begin{align} \mathcal E_{ \textup{\uppercase\expandafter{\romannumeral2}} }=\Big(\mathbb R^2,\Sigma_{\mathbb R^2},N(h,C_{ \textup{\uppercase\expandafter{\romannumeral2}} }^{-1}):h\in \mathcal T_{ \textup{\uppercase\expandafter{\romannumeral2}} }(\theta_0)\Big), \end{align} in which one observes a single random vector $Z\sim N(h,C_{ \textup{\uppercase\expandafter{\romannumeral2}} }^{-1})$. From Theorem (ref), it then suffices to consider the power for testing $H_0:p'h=0$ against $p'h>0$ in this simple experiment. With $p=(-1,-1)'$, it can be shown that the power envelope is \begin{align} \limsup_{n\to\infty}\pi_{n,\theta_0+h/\sqrt n}(\phi_n)\le 1-\Phi\Big(z_\alpha-\frac{-h^{(1)}-h^{(2)}}{\sqrt{4\pi/3}}\Big). \end{align} This bound can be achieved using a test that rejects the null when the following statistic exceeds $z_\alpha$: \begin{align} T_{ \textup{\uppercase\expandafter{\romannumeral2}} ,n}&=\frac{1}{\sqrt n}\sum_{i=1}^n\Big[-\frac{4}{3}\sqrt{\frac{3}{2}}1\{s_i=(1,1)\}+\frac{2}{3}\sqrt{\frac{3}{2}}(1\{s_i=(1,0)\}+1\{s_i=(0,1)\})\Big]. \end{align} Heuristically, this means that the test rejects $H_0$ when one observes $(1,0)$ or $(0,1)$ frequently relative to $(1,1)$. In general, the power envelope and optimal test depend on local cone $\mathcal T_{j}(\theta_0)$ and the functional of interest. Appendix (ref) describes the efficient influence functions for this example.
remark\rm The optimal test above compares the relative frequencies of two events $\{(1,1)\}$ and $\{(1,0),(0,1)\}$. It therefore only uses the information on the {\em number of entrants\/} (i.e., duopoly v.s. monopoly) in each market, which is the feature (or transformation) of the outcome $s$ used in bresnahan1990entry,bresnahan1991empirical and berry1992estimation. For testing competing hypotheses on $\varphi(\theta)=-\theta^{(1)}-\theta^{(2)}$, using such a transformed outcome indeed leads to an optimal test. However, our theory suggests that the choice of transformation depends, in general, on the functional of interest ($\varphi$) and the direction of alternatives ($\mathcal T_j(\theta_0)$). For example, with $\varphi(\theta)=-\theta^{(1)}-2\theta^{(2)}$ and $h\in\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral2}} }(\theta_0)$, the optimal test compares the relative frequencies of $\{(1,1)\}$ and $\{(1,0)\}$.
remark\rm Theorem (ref) applies to each cone $\mathcal T_{j}(\theta_0)$ and hence can be used to obtain a test that asymptotically achieves the power envelope against the alternatives in $\mathcal T_{j}(\theta_0)$. The optimal test, however, may differ by cone.\footnote{When a functional satisfies $\dot\varphi_{\theta_0}=c\times (-1,-1)$ for some $c>0$ in Example (ref), the optimal test is common across the cones because the projection of the influence function onto $cl(\mathcal G_{j,\theta_0})$ is common across $j$. However, this does not hold with the other functionals.}

Asymptotic Power against Shifted Local Alternatives

As discussed earlier, not all alternatives are robustly testable. Consider the following example. \setcounter{example}{1}

example[Roy model (continued)]\rm Suppose that the researcher wants to know if the share of individuals who have higher economic prospects in sector 0 is above a certain percentage. The parameter of interest is $\theta^{(1,0)}=m_\theta((Y_0,Y_1)=(1,0))$, and the null and alternative hypotheses can be expressed as $H_0:p'\theta=\theta^{(1,0)}\le c$ and $H_1:p'\theta>c$ for some $c\in [0,1]$ with $p=(0,0,1).$ Mourifie:2018aa showed that the sharp identifying restrictions are \begin{align} \theta^{(1,0)}&\le P(\{(1,0)\})\\ \theta^{(0,1)}&\le P(\{(1,1)\})\\ \theta^{(0,0)}&= P(\{(0,0)\})+P(\{(0,1)\}). \end{align} For simplicity, suppose that $\theta^{(0,0)}$ is known to be $1/6$. This implies $P(\{(1,0)\})+P(\{(1,1)\})=5/6$, which in turn simplifies the restrictions to \begin{align} \theta^{(1,0)}&\le P(\{(1,0)\})\le \frac{5}{6}-\theta^{(0,1)}\\ \frac{1}{6}&= P(\{(0,0)\})+P(\{(0,1)\}). \end{align} Now, consider testing $\theta_0$ against the alternative $\theta_1=\theta_0+\xi$, where $\xi=(0,\xi^{(0,1)},\xi^{(1,0)})'$ with $\xi^{(1,0)}>0$. As shown in Figure (ref) (Alt. 1), the interval $[\theta^{(1,0)}_0+\xi^{(1,0)}, \frac{5}{6}-\theta^{(0,1)}-\xi^{(0,1)}]$ to which $P(\{(1,0)\})$ belongs under this alternative has a non-empty intersection with the interval under the null until $\xi^{(1,0)}$ becomes sufficiently large. Indeed, $\mathcal P_{\theta_0+\xi}$ becomes disjoint from $\mathcal P_{\theta_0}$ only when $\xi^{(1,0)}> \frac{5}{6}-\theta^{(0,1)}_0-\theta^{(1,0)}_0$ (Alt. 2).\footnote{Another way to make $\mathcal P_{\theta_0}$ and $\mathcal P_{\theta_0+\xi}$ disjoint is to shift the interval to the left in Figure (ref) by a sufficiently large amount. For this, one needs $5/6-\theta_0^{(0,1)}-\xi^{(0,1)}<\theta_0^{(1,0)}$. However, $\theta_0+\xi$ does not satisfy $\varphi(\theta_0+\xi)>0$ in this case.} This means that against any local alternative of the form $\theta_0+h/\sqrt n$, the lower power of any level-$\alpha$ test is eventually dominated (weakly) by $\alpha$, which does not lead to useful comparisons. \begin{figure}[h] \begin{center} \caption{Bounds on $P(\{(1,0)\})$} \begin{tikzpicture} [scale=11,domain=0:1,range=-1:1,>=latex] \draw (0,0.15) -- (1,0.15); \draw (0,0.13) -- (0,0.17); \draw (1,0.13) -- (1,0.17); \draw (0,0.13) node[below] {0}; \draw (1,0.13) node[below] {1}; \filldraw[black] (0,0.15) circle [radius = 0] node[left] {Null:}; \draw (0,0) -- (1,0); \draw (0,-0.02) -- (0,0.02); \draw (1,-0.02) -- (1,0.02); \draw (0,-0.02) node[below] {0}; \draw (1,-0.02) node[below] {1}; \filldraw[black] (0,0) circle [radius = 0.] node[left] {Alt.1:}; \draw (0,-0.15) --(1,-0.15); \draw (0,-0.17) -- (0,-0.13); \draw (1,-0.17) -- (1,-0.13); \draw (0,-0.17) node[below] {0}; \draw (1,-0.17) node[below] {1}; \filldraw[black] (0,-0.15) circle [radius = 0] node[left] {Alt.2:}; \draw (1/6,0.15) node {[}; \draw (1/6+0.02,0.17) node[above] {$\theta_0^{(1,0)}$}; \draw (1/3,0.15) node {]}; \draw (1/3+0.02,0.17) node[above] {$5/6-\theta_0^{(0,1)}$}; \draw (1/3.5,0) node {[}; \draw (1/3.5,0.02) node[above] {$\theta_0^{(1,0)}+\xi^{(1,0)}$}; \draw (1/4+1/6,0) node {]}; \draw (1/4+1/6+0.02,-0.02) node[below] {$5/6-\theta_0^{(0,1)}-\xi^{(0,1)}$}; \fill[fill=Green,opacity=0.4] (1/3.5,0) -- (1/3.5,0.01) -- (1/3,0.01) -- (1/3,0); \draw (1/2,-0.15) node {[}; \draw (1/2,-0.13) node[above] {$\theta_0^{(1,0)}+\tilde\xi^{(1,0)}$}; \draw (1/2+1/4+1/6-1/3.5,-0.15) node {]}; \draw (1/2+1/4+1/6-1/3.5+0.02,-0.17) node[below] {$5/6-\theta_0^{(0,1)}-\tilde\xi^{(0,1)}$}; \end{tikzpicture} \end{center} {\tmpsmall\sc Note: The interval under Alt.1 has a non-empty intersection (in green) with the interval under the null. When $\tilde\xi^{(1,0)}> \frac{5}{6}-\theta^{(0,1)}_0-\theta^{(1,0)}_0$ (Alt.2), the two intervals are disjoint, i.e., $\mathcal P_{\theta_0+\tilde \xi}\cap \mathcal P_{\theta_0}=\emptyset.$} \end{figure}

Given the challenge above, we conduct a local power analysis as follows. First, we {\em shift\/} $\theta_0$ by vector $\xi$, which does not depend on $n$. Hence, at $\theta_0+\xi$, certain local deviations can be robustly detectable. We then analyze the limit of experiments constructed from a sequence of LFPs induced by such alternatives. Specifically, the {\em shifted local alternative\/} is

align[align omitted — 54 chars of source]

where $\theta_0\in\Theta$ is such that $\varphi(\theta_0)= 0$, and $\xi$ and $h$ take values in the following sets:

align[align omitted — 292 chars of source]

In other words, $\Xi_{\theta_0}$ collects the deviations for which $\theta_0+\xi$ are not robustly testable. Set $\mathcal T_n(\theta_0,\xi)$ then collects local deviations that make $\theta_{n,\xi,h}$ robustly testable. Among the points in $\Xi_{\theta_0}$, we focus on those for which $\mathcal T_n(\theta_0,\xi)$ is non-empty.

Under this construction, we may apply Theorem (ref) for any $\tau>0$ to obtain the LFP $(Q_{0,j,\tau h},Q_{1,j,\tau h})\in \mathcal P_{\theta_0}\times\mathcal P_{\theta_0+\xi+\tau h}$. Suppose that the following assumption holds for $\theta_0\in\Theta$ and $\xi\in\Xi_{\theta_0}.$

assumption[Local parameter cones and $L^2$ directional differentiability] (i) Set $\{h\in \Theta-\theta_0-\xi:\mathcal P_{\theta_0}\cap \mathcal P_{\theta_0+\xi+h}=\emptyset\}$ is locally equal to convex cone $\mathcal T(\theta_0,\xi)$; (ii) There exists set $\mathbb J$ and a collection of convex cones (containing 0) $\{\mathcal T_{j}(\theta_0,\xi),j\in \mathbb J\}$ such that $\mathcal T_{j}(\theta_0,\xi)\cap\mathcal T_{j'}(\theta_0,\xi)=\{0\},\forall j\ne j'$ and $\bigcup_j\mathcal T_{j}(\theta_0,\xi)=\mathcal T(\theta_0,\xi)$; (iii) For each $j\in\mathbb J$, there exists model $\vartheta\mapsto \mathsf Q_{j,\vartheta}$ defined on a neighborhood of $\vartheta=0$ such that the LFP $(Q_{0,j,\tau h},Q_{1,j,\tau h})\in \mathcal P_{\theta_0}\times\mathcal P_{\theta_0+\xi+\tau h}$ satisfies \begin{align} Q_{0,j,\tau h}=\mathsf Q_{j,0}, and Q_{1,j,\tau h}=\mathsf Q_{j,\tau h}, \end{align} for all $\tau \in(0,\bar \tau]$ for some $\bar \tau>0$, and $\vartheta\mapsto \mathsf Q_{j,\vartheta}$ is $L^2$ differentiable at $0$ tangentially to $\mathcal T_{j}(\theta_0,\xi).$

In what follows, we let $\dot \ell_{j,0},j\in\mathbb J$ denote the $L^2$ derivatives and focus on testing the linear hypotheses, namely $\varphi(\theta)=p'\theta-c$. Define the influence curve $\varrho_j$ of $\varphi(\theta)=p'\theta-c$ as a square integrable function $\varrho_j$ that satisfies, for every $h\in\mathcal T_{j}(\theta_0,\xi)$, $p'h=\langle \varrho_j,g\rangle_{L^2_{\mathsf Q_{j,0}}}$, where $g=h'\dot \ell_{j,0}$. Let $\mathcal G_{j,\theta_0}=\{g\in L^2_{\mathsf Q_{j,0}}:g=h'\dot \ell_{j,0},h\in\mathcal T_j(\theta_0,\xi)\}$. Let $\tilde\varrho_j$ be the efficient influence function, which is the projection of $\varrho_j$ to the closure of $\mathcal G_{j,0}.$ We then obtain the following result.

theoremSuppose Assumption (ref) holds. Let $j\in\mathbb J.$ Let $\theta_0\in\Theta$ such that $\varphi(\theta_0)=0$. Let $\phi_n$ be a level-$\alpha$ test for $H_0:\varphi(\theta)\le 0$ against $H_1:\varphi(\theta)>0$ and $\pi_{n,\theta}(\phi_n)$ be its lower power under $\nu_{\theta}^n$. Then, for any $h\in\mathcal T_{j}(\theta_0,\xi)$, \begin{align} \limsup_{n\to\infty}\pi_{n,\theta_{n,\xi,h}}(\phi_n)\le 1-\Phi\Bigg(z_\alpha-\frac{\langle \varrho_j,h'\dot\ell_{j,0}\rangle_{L^2_{\mathsf Q_{j,0}}}}{\|\tilde{\varrho}_j\|_{L^2_{\mathsf Q_{j,0}}}}\Bigg). \end{align} Let $T_{j,n}$ be a statistic such that \begin{align} T_{j,n}=\frac{n^{-1/2}\sum_{i=1}^n\tilde\varrho_j(s_i)}{\|\tilde\varrho_j\|_{L^2_{\mathsf Q_{j,0}}}}+o_{\mathsf Q^n_{j,0}}(1). \end{align} Let $\phi^*_{j,n}$ be a test that rejects the null hypothesis iff $T_{j,n}\ge z_\alpha.$ Then, the test is of level-$\alpha$ and, for any $h\in \mathcal T_{j,\theta_0}$, \begin{align} \liminf_{n\to\infty}\pi_{n,\theta_{n,\xi,h}}(\phi^*_{j,n})=1-\Phi\Bigg(z_\alpha-\frac{\langle \varrho_j,h'\dot\ell_{j,0}\rangle_{L^2_{\mathsf Q_{j,0}}}}{\|\tilde{\varrho}_j\|_{L^2_{\mathsf Q_{j,0}}}}\Bigg). \end{align}

Below, we again use Example (ref) to illustrate the local power analysis. \setcounter{example}{1}

example[Roy model (continued)]\rm Figure (ref) shows $\Xi_{\theta_0}$ and the cones in Assumption (ref). Since $\theta^{(0,0)}=1/6$ is known, we can plot the parameters in a two-dimensional simplex $\{(\theta^{(1,0)},\theta^{(0,1)})\in[0,1]^2:\theta^{(0,1)}+\theta^{(1,0)}\le 5/6\}.$ Here, we take $\theta^{(1,0)}_0,=1/6,\theta^{(0,1)}=1/2,$ and test $H_0:\theta^{(1,0)}\le 1/6$ against $H_1:\theta^{(1,0)}>1/6.$ This configuration implies that the alternative $\theta_0+\xi$ is not robustly testable unless $\xi^{(1,0)}>1/6.$ The green region in Figure (ref) shows the set of not robustly testable alternatives.\footnote{Once translated by $-\theta_0$, this region represents the set $\Xi_{\theta_0}$ of fixed shifts.} One can see that the local parameter space $\mathcal T(\theta_0,\xi)$ is non-empty when $\xi$ is a boundary point of $\Xi_{\theta_0}$ toward $p$. For example, at $\theta_0+\xi_A$, the local parameter space $\mathcal T(\theta_0,\xi_A)$ is given by a half space $\{h=(h^{(0,0)},h^{(0,1)},h^{(1,0)}):h^{(0,0)}=0,h^{(1,0)}>0\}$. At $\theta_0+\xi_B$, the local parameter space is $\{h=(h^{(0,0)},h^{(0,1)},h^{(1,0)}):h^{(0,0)}=0,h^{(1,0)}>0,h^{(0,1)}+h^{(1,0)}=0\}$. At each boundary point, one can then conduct a local power analysis. \begin{figure} [htbp] \begin{center} \caption{Set of not robustly testable alternatives and local parameter spaces} \begin{tikzpicture} [scale=5.5,domain=0:1,range=-0.2:1.2,>=latex] \draw[->] (0,0) -- (1,0) node[right] {$\theta^{(1,0)}$}; \draw[->] (0,0) -- (0,1) node[above] {$\theta^{(0,1)}$}; \draw[-] (5/6,0) -- (0,5/6); \draw[dashed] (1/6,-0.2) -- (1/6,1); \draw (1/6,-0.2) node[right] {$p'\theta>0$}; \draw (1/6,-0.2) node[left] {$p'\theta\le 0$}; \draw[->] (1/6,-0.1) -- (1/6-0.05,-0.1); \draw[->] (1/6,-0.1) -- (1/6+0.05,-0.1); \draw (0,0) node[left] {$0$}; \filldraw[black] (1/6,1/2) circle [radius = 0.005] node[left] {$\theta_0$}; \fill[fill=Green,opacity=0.4] (1/6,0) -- (1/6,2/3) -- (1/3,1/2) -- (1/3,0); \draw (0.25,0.5) node[above]{$\xi_{B}$}; \draw[dashed,->] (1/6,1/2) -- (1/3,1/2); \draw (0.25,0.4) node[below]{$\xi_{A}$}; \draw[dashed,->] (1/6,1/2) -- (1/3,1/3); \filldraw[black] (1/3,1/2) circle [radius=0.005]; \filldraw[black] (1/3,1/3) circle [radius=0.005]; \filldraw[gray,opacity=0.2] (1/3,1/2) circle [radius =0.03]; \filldraw[gray,opacity=0.2] (1/3,1/3) circle [radius =0.03]; \end{tikzpicture} \end{center} \begin{center} \tmpsmall\sc \begin{tikzpicture} [scale=1.5,domain=0:1,range=0:1,>=latex] \filldraw[black] (0,0) circle [radius=0.025]; \draw[->] (-1,0) -- (1,0) node[right] {$h^{(1,0)}$}; \draw[dashed,->] (0,-1) -- (0,1) node[above] {$h^{(0,1)}$}; \fill[fill=blue,opacity=0.2] (0,1) -- (1,1) -- (1,-1) -- (0,-1); \draw (0.5,-1) node[below] {$\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} }(\theta_0,\xi_{A})$}; \end{tikzpicture} \begin{tikzpicture} [scale=1.5,domain=0:1,range=0:1,>=latex] \filldraw[black] (0,0) circle [radius=0.025]; \draw[->] (-1,0) -- (1,0) node[right] {$h^{(1,0)}$}; \draw[dashed,->] (0,-1) -- (0,1) node[above] {$h^{(0,1)}$}; \fill[fill=blue,opacity=0.2] (0,0) -- (1,-1) -- (0,-1); \draw (0.5,-1) node[below] {$\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} }(\theta_0,\xi_{B})$}; \draw (1,-1) node[right] {$\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral2}} }(\theta_0,\xi_{B})$}; \draw[color=black,line width=0.25mm] (0,0) -- (1,-1); \end{tikzpicture} \end{center} Note: $\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} }(\theta_0,\xi_{A})$ (half space) coincides with the local parameter space at $\theta_0+\xi_A$. $\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} }(\theta_0,\xi_{B})$ (shaded area not including the solid line) and $\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral2}} }(\theta_0,\xi_{B})$ (solid line) are tangent cones at $\theta_0+\xi_B$. \end{figure} As an illustration, take $\xi=\xi_A\equiv (0,1/6,-1/6)'$. It turns out that in this setting, it suffices to consider a single convex cone $\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} }(\theta_0,\xi_A)\equiv\mathcal T(\theta_0,\xi_A)$. Let $\theta_{n,\xi_A,h}=\theta_0+\xi_A+h/\sqrt n$ with $h\in\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} }(\theta_0,\xi_A)$ be our shifted local alternative.\footnote{At $\theta_0+\xi_A$, one needs to consider a single cone that coincides with the local parameter space.} Hence, the LFP for $\mathcal P_{\theta_0}$ and $\mathcal P_{\theta_{n,\xi_A,h}}$ has the densities \begin{align} (q_0(0,0),q_0(0,1),q_0(1,0),q_0(1,1))&=(\frac{1}{12},\frac{1}{12},\frac{1}{3},\frac{1}{2})\\ (q_1(0,0),q_1(1,1),q_1(1,0),q_1(1,1))&=(\frac{1}{12},\frac{1}{12},\frac{1}{3}+\frac{h^{(1,0)}}{\sqrt n},\frac{1}{2}-\frac{h^{(1,0)}}{\sqrt n}). \end{align} These LFPs can be embedded into model $\vartheta\mapsto \mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\vartheta}$ whose density $\mathsf q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\vartheta}$ is given by \begin{align} (\mathsf q_{ \uppercase\expandafter{\romannumeral1} ,\vartheta}(0,0),\mathsf q_{ \uppercase\expandafter{\romannumeral1} ,\vartheta}(1,1),\mathsf q_{ \uppercase\expandafter{\romannumeral1} ,\vartheta}(1,0),\mathsf q_{ \uppercase\expandafter{\romannumeral1} ,\vartheta}(1,1))=(\frac{1}{12},\frac{1}{12},\frac{1}{3}+\vartheta^{(1,0)},\frac{1}{2}-\vartheta^{(1,0)}). \end{align} This model is $L^2$ directionally differentiable (at $0$) tangentially to $\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} }(\theta_0,\xi_A)$ with the following directional derivative: \begin{align} \dot\ell_{ \uppercase\expandafter{\romannumeral1} ,0}(s) &=1\{s_i=(1,0)\}\begin{pmatrix} 0\\ 0\\ 3 \end{pmatrix}+ 1\{s_i=(1,1)\}\begin{pmatrix} 0\\ 0\\ -2 \end{pmatrix}. \end{align} The log-likelihood function can then be expanded as in (ref) with $\Delta_{ \textup{\uppercase\expandafter{\romannumeral1}} ,n}=\frac{1}{\sqrt n}\sum_{i=1}^n\dot \ell_{ \textup{\uppercase\expandafter{\romannumeral1}} ,0}$ and information matrix $C_{ \textup{\uppercase\expandafter{\romannumeral1}} }= \big(\begin{smallmatrix} \mathbf{0}_2&0\\ 0&5 \end{smallmatrix}\big), $ which is singular, where $\mathbf 0_2$ is a two-by-two matrix of zeros. The limit experiment is \begin{align} \mathcal E_{ \uppercase\expandafter{\romannumeral1} }=\Big(\mathbb R^3,\Sigma_{\mathbb R^3},N(C_{ \textup{\uppercase\expandafter{\romannumeral1}} }h,C_{ \textup{\uppercase\expandafter{\romannumeral1}} }):h\in \mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} }(\theta_0,\xi_A)\Big), \end{align} in which one observes a single random vector $Z\sim N(C_{ \textup{\uppercase\expandafter{\romannumeral1}} }h,C_{ \textup{\uppercase\expandafter{\romannumeral1}} })$. Here, the information matrix is not full rank. This experiment essentially involves a single normal random variable with mean $5h^{(1,0)}$ and variance $5$. With $p=(0,0,1)'$, the efficient influence function is \begin{align} \tilde\varrho_{ \textup{\uppercase\expandafter{\romannumeral1}} }(s)=\frac{3}{5}1\{s=(1,0)\}-\frac{2}{5}1\{s=(1,1)\}. \end{align} Theorem (ref) then implies, for any level-$\alpha$ test $\phi_n$, \begin{align} \limsup_{n\to\infty}\pi_{n,\theta_{n,\xi_A,h}}(\phi_n)\le 1-\Phi\Big(z_\alpha-\sqrt 5h^{(1,0)}\Big). \end{align} This bound can be achieved using a test that rejects the null when the following statistic exceeds $z_\alpha$: \begin{align} T_{ \textup{\uppercase\expandafter{\romannumeral1}} ,n}&=\frac{1}{\sqrt n}\sum_{i=1}^n\Big[\frac{3}{\sqrt 5}1\{s_i=(1,0)\}-\frac{2}{\sqrt 5}1\{s_i=(1,1)\}\Big]. \end{align} Heuristically, the test rejects $H_0$ when one observes $(1,0)$ sufficiently more frequently than $(1,1)$, which is the robust prediction of the model when $\theta^{(1,0)}$ is sufficiently large.\footnote{The case with $\xi_B=(0,1/6,0)'$ can be analyzed similarly. For $\theta_{n,\xi_B,h}=\theta_0+\xi_B+h/\sqrt n$, one needs to consider two tangent cones, $\mathcal T_j(\theta_0,\xi_B),j\in\{ \textup{\uppercase\expandafter{\romannumeral1}} , \textup{\uppercase\expandafter{\romannumeral2}} \}$, because the form of the LFP changes depending on the cone $h$ to which belongs (see Appendix (ref)). Interestingly, the optimal test statistic is still given by (ref) in both cases.}

Monte Carlo Experiments

We conduct Monte Carlo experiments to examine the performance of our tests.

Tests of Strategic Interaction Effects

The first set of experiments evaluates the size and power of the tests on the strategic interaction effects. The design of the experiment is based on Example (ref), in which we generate $u_{i}\stackrel{i.i.d.}{\sim}N(0,I_2)$ for $i=1,\cdots,n.$ Whenever multiple equilibria exist, we select an outcome using one of the three selection mechanisms below. The first one is an i.i.d. selection mechanism, which selects $(1,0)$ out of $G(u|\theta)=\{(1,0),(0,1)\}$ if an i.i.d. Bernoulli random variable $v_i$ takes 1. Otherwise, $(0,1)$ is selected. The second mechanism selects $(1,0)$ when another Bernoulli random variable $\tilde v_i$ takes 1, where $\{\tilde v_i\}$ is a i.n.i.d. sequence. Let $N^*_k$ be an increasing sequence of integers.\footnote{In our simulations, we set $N^*_k=2^{2^k}$.} For each $i$, let $h(i)=N^*_k$, where $N^*_{k-1}<i\le N^*_k$. We define

align[align omitted — 445 chars of source]

where $\Psi_n^G(u^{\infty})=\frac{\sum_{i=1}^nI[G(u_i|\theta)=\{(1,0)\}]}{\sum_{i=1}^nI[G(u_i|\theta)=\{(1,0)\} or \{(0,1)\}]}$. One may view observations for which $N^*_{k-1}<i\le N^*_k$ as members of a cluster. Under this selection mechanism, the outcomes are dependent on each cluster, which in turn makes the outcome sequence heterogeneous (and non-ergodic). The third mechanism generates data from the LFP, which draws an outcome sequence from $Q_0^n$ when $\theta=\theta_0$ and from $Q^n_1$ when $\theta=\theta_1$.

We evaluate the size and power of the test in Section (ref) on $H_0:p'\theta=0$ against $H_1:p'\theta>0$ with $p=(-1,-1)'.$ The test based on the statistic in (ref) achieves the power envelope (see Appendix (ref)). We therefore evaluate the size and power of this test. To make a comparison, we consider another test, namely the Wald--Wolfowitz runs test, which non-parametrically tests the i.i.d.-ness of the outcome sequence Wald:1940aa,Cho:2011aa. Since the model is complete under $H_0$ and $u_i$ is i.i.d., the resulting outcome sequence is i.i.d. under the null hypothesis. The test therefore should have power only when the selection introduces heterogeneity and/or dependence.

Figures (ref) and (ref) show the power of the optimal test and runs test, respectively. The power of our test changes little across the selection mechanisms. This can be explained as follows. Our test statistic in (ref) treats the two outcomes, $(0,1)$ and $(1,0)$, symmetrically. While the selection mechanism affects the relative frequencies of the two outcomes, what matters for the statistic is the frequency of $\{(0,1),(1,0)\}$ (relative to $(1,1)$), and hence its power curve is insensitive to the selection mechanism.\footnote{This insensitivity is not a generic feature of the optimal test. See the next example.} The Wald--Wolfowitz test has non-zero power when the selection is i.n.i.d.; it becomes noticeable only when $h$ is very large and its lower power is significantly below that of the optimal test. As expected, it does not have any power when the selection mechanism is i.i.d. or the LFP.

Tests on the Distribution of Potential Outcomes

The second set of experiments is based on Example (ref), which we use to evaluate the performance of the tests when the model is incomplete under $H_0$. The model may be complete under certain alternatives. As described in Section (ref), we consider testing $H_0:p'\theta\le c$ against $H_1:p'\theta>c$, where $p=(0,0,1)'$ and $c=1/6$. As before, we assume that $\theta^{(0,0)}=1/6$ is known and localize the experiment at $\theta_0=(\theta_0^{(0,0)},\theta_0^{(0,1)},\theta_0^{(1,0)})'=(1/6,1/2,1/6)'.$

The set of selection mechanisms is the same as the one in Section (ref). One difference is that the model predicts multiple outcomes when $u=(0,0)$ or $u=(1,1)$. In both cases, an i.i.d. selection mechanism selects one of the outcomes ($s=(0,0)$ when $G(u)=\{(0,0),(0,1)\}$ and $s=(1,0)$ when $G(u)=\{(1,0),(1,1)\}$) when an i.i.d. Bernoulli random variable $v_i$ is 1. Similarly, an i.n.i.d. selection mechanism selects one of the predicted outcomes when $\tilde v_i$ in (ref) is 1. The LFP selection mechanism is defined in the same way as before.

We consider the following shifted local alternatives:

align[align omitted — 297 chars of source]

Under Alternative A, we consider a sequence of the parameters that tends to $\theta_0+\xi_A$, where $h\in\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} }(\theta_0,\xi_A)$. Under Alternatives B-I and B-II, we consider the parameters that tend to $\theta_0+\xi_B$, where $h\in\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} }(\theta_0,\xi_B)$ and $h\in\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral2}} }(\theta_0,\xi_B)$, respectively.

Figures (ref) and (ref) report the results with $n=1000$ and $S=2000$, respectively. The right panel of Figure (ref) shows the power curves of the optimal test against Alternative A. Because of the model incompleteness, its performance varies significantly across the selection mechanisms. As predicted by the theory, the power of the test under the LFP selection mechanism essentially coincides with the power envelope. Both the i.i.d. and the i.n.i.d. selection mechanisms are considerably more favorable; that is, the actual power of the test under these mechanisms is much higher than under the LFP selection. In particular, the power of the test under the i.i.d. selection mechanism is already 1 even when $\bar h=0$ because the test can detect deviations from the null under this selection mechanism even for $\theta=\theta_0+\xi$ that are not robustly testable. The left panel of Figure (ref) shows the power curves against alternatives of the form $\theta_0+w\xi_A$ with $w\in[0,1]$. The figure shows that the test may have non-trivial power even for such alternatives when the selection mechanisms are i.i.d. or i.n.i.d.\footnote{Under the LFP selection, the data are drawn from $Q_{\theta_0}\in\mathcal P_{\theta_0+w\xi_A}$, and hence the power of the test does not exceed the nominal level.}

Figure (ref) shows the power curves against Alternative B-I. When $\bar h=0$, the model is indeed complete.\footnote{The model remains incomplete regarding selection from $G(u)=\{(0,0),(0,1)\}$ when $u=(0,0)$, however, these outcomes are not used by the test and hence do not affect the power.} However, as $\bar h$ increases, the region of incompleteness enlarges, leading to differences in the local power across the selection mechanisms. Under Alternative B-II, the model stays complete under the local alternatives. Therefore, the power of the test is essentially the same across the selection mechanisms.

Extensions

Tests in the Presence of Nuisance Components

We now consider testing the hypotheses on subcomponents of $\theta$. Let $\theta=(\beta',\delta')'\in \Theta_\beta\times\Theta_\delta$, where $\beta$ is a $k\times 1$-sub-vector of interest and $\delta$ is a $(d-k)\times 1$ vector of nuisance parameters. Consider the following hypotheses:

align[align omitted — 140 chars of source]

This problem can be recast as a special case of (ref) with $\varphi(\theta)=\beta$, $K_0=\{\beta_0\}$, and $K_1=\{\beta\in\mathbb R^k:\beta\ne\beta_0\}.$\footnote{Sub-vector inference has been actively studied in the context of partially identified models, particularly moment inequality models romano2008inference,Shicanaybugni2015inference,Kaido:2016aa. Here, we focus on hypothesis tests on the sub-vectors of the structural parameters.}

In this general setting, both hypotheses are composite in terms of the structural parameters. Therefore, Lemma (ref) is not directly applicable. However, the result is still useful for constructing tests that have desirable properties. To this end, we partition the parameter space into two sets, namely $\Theta_0$ and $\Theta_1$, where $\Theta_0=\{\beta_0\}\times \Theta_\delta$ and $\Theta_1=\{\beta:\beta\ne\beta_0\}\times\Theta_\delta$. We focus on this setting, whereas the theory below applies more generally to the hypotheses of the form in (ref) by taking $\Theta_0=\{\theta:\varphi(\theta)\in K_0\}$ and $\Theta_1=\{\theta:\varphi(\theta)\in K_1\}$.

Throughout, the researcher's action is binary, that is $a=1$ (reject) or $a=0$ (accept). For each $\theta\in\Theta$ and action $a\in\{0,1\}$, define a loss function $\mathsf L:\Theta\times \{0,1\}\to\mathbb R_+$ by

align*[align* omitted — 95 chars of source]

where $\zeta>0$. The loss from the Type-I error is normalized to 1. The trade-off between the Type-I and Type-II errors is determined by parameter $\zeta$.

For each test $\phi$ and $\theta\in\Theta$, define the {\em upper risk\/} by

align[align omitted — 269 chars of source]

where the integrals in (ref) are Choquet integrals (see Appendix (ref)). The upper risk determines the trade-off between the size ($R_0(\theta,\phi)\equiv\sup_{P\in\mathcal P_\theta}\int \phi dP=\int\phi d\nu^*_\theta$ for $\theta\in\Theta_0$) and the lower power ($\inf_{P\in\mathcal P_\theta}\int \phi dP= \int\phi d\nu_\theta$ for $\theta\in\Theta_1$). What remains is to incorporate parameter uncertainty. For this, let $\mu$ be a (prior) probability distribution over $\Theta$. We write $\mu$ as $\mu=\tau\mu_0+(1-\tau)\mu_1,$ where $\tau\in (0,1)$ and $\mu_0,\mu_1$ are suitable probability measures supported on $\Theta_0$ and $\Theta_1$, respectively. Define

align[align omitted — 173 chars of source]

where $\kappa_0^*=\int_{\Theta_0}\nu_\theta^*d\mu_0$ and $\kappa_1=\int_{\Theta_1}\nu_\theta d\mu_1$.\footnote{The second equality in (ref) is established in the proof of Theorem (ref).} This risk function uses the prior probability to reflect parameter uncertainty, while it uses the belief function (and its conjugate) to incorporate the decision maker's willingness to be robust against incompleteness. In what follows, we call $r$ the {\em Bayes--Dempster--Shafer (BDS) risk\/}. We then call $\phi$ a {\em BDS test\/} if it minimizes the BDS risk.\footnote{The axiomatic foundations for this type of preference (when $S$ is the payoff-relevant state space) is given in Gul:2014aa and epstein2015exchangeable (in the context of repeated experiments).} One of the components of the BDS risk is $\pi_{\kappa_1}(\phi)=\int \phi d\kappa_1$. We call this object the {\em weighted average lower power\/} (WALP). The interpretation of $\pi_{\kappa_1}(\phi)$ is similar to that of the standard weighted average power Andrews:1994aa,Andrews:1995aa. Choosing a suitable $\mu_1$, one may direct the power of a test toward certain alternatives. However, $\pi_{\kappa_1}$ takes the average of the guaranteed power value instead of the actual unknown power.

The following theorem characterizes the BDS test. For this, let $core(\kappa)\equiv \{P\in\Delta(S):P(A)\ge \kappa(A),\forall A\subset S\}$. In what follows, we assume that $core(\kappa_0)\cap core(\kappa_1)=\emptyset$.\footnote{To ensure this condition, it is sufficient to have at least one $\bar A\subset S$ such that $P_0(\bar A)<P_1(\bar A)$ (or $P_0(\bar A)>P_1(\bar A)$) for all $(P_0,P_1)\in \mathcal P_{\theta_0}\times\mathcal P_{\theta_1}$ and $(\theta_0,\theta_1)\in \Theta_0\times\Theta_1$. In Example (ref), one may take $\bar A=\{(1,1)\}$.}

lemmaLet the BDS risk be defined as in (ref). Then, there exists a BDS test such that, for any $\zeta>0$, \begin{align} \phi(s) = \left\{ \begin{array}{ll} 1 & if \quad \Lambda(s)>C\\ \gamma & if \quad \Lambda(s)=C\\ 0 & if \quad \Lambda(s)<C,\\ \end{array} \right. \end{align} where $C=\tau/\zeta(1-\tau)$, and $\Lambda$ is a version of $dQ_1/dQ_0$ for the LFP $(Q_0,Q_1)\in core(\kappa_0)\times core(\kappa_1)$ such that, for all $t\in\mathbb R_+$, \begin{align} Q_0(\Lambda>t)=\kappa_0^{*}(\Lambda>t), and Q_1(\Lambda>t)=\kappa_1(\Lambda>t). \end{align}

One may view this as an analog of Lemma (ref). A key difference is that the LFP belongs to the product of the cores of the capacities $\kappa_0$ and $\kappa_1$. Hence, $\kappa_0$ and $\kappa_1$ are both belief functions, which in turn allows us to compute the LFP in a tractable way.

The analysis above is useful for constructing optimal tests for minimizing risk. However, the BDS tests are not designed to control size uniformly over $\Theta_0$. Therefore, we consider a test that controls size and maximizes the WALP. For this, we fix $\mu_1$ throughout and follow the developments on the tests in the presence of nuisance parameters Chamberlain:2000aa,Elliott:2015aa,Moreira:2013aa.

The following minimax theorem characterizes the minimax test as a BDS test for the least favorable prior (if it exists). For this, let $\mathcal M(\mu_1)\equiv\{\mu:\mu=\tau\mu_0+(1-\tau)\mu_1,\mu_0\in\Delta(\Theta_0),\tau\in[0,1]\}$, where $\mu_1$ is fixed. In what follows, we drop $\mu_1$ from the argument of $\mathcal M$, but its dependence should be understood. We then let $\mathbf\Phi$ be the set of randomized tests.

theoremLet the upper risk $R$ be defined as in (ref). Suppose that $\Theta$ is compact. Then, \begin{align} \sup_{\mu\in\mathcal M}\inf_{\phi\in\mathbf\Phi}\int_\Theta R(\theta,\phi)d\mu(\theta)=\inf_{\phi\in\mathbf\Phi}(\sup_{\theta\in\Theta_0}R_{0}(\theta,\phi)\vee R_1(\phi)), \end{align} where $R_{0}(\theta,\phi)=\int \phi(s)d\nu^*_\theta(s)I_{\Theta_0}(\theta)$ and $R_1(\phi)=\zeta(1-\pi_{\kappa_1}(\phi))$. Furthermore, there exists $\phi^\dagger$ that achieves equality in (ref).
remark\rm Suppose that $\zeta$ is chosen so that the maximum BDS risk (left-hand side of (ref)) equals $\alpha$. Then, $\phi^\dagger$ is a level-$\alpha$ test that maximizes the WALP. The theorem suggests that such a test can be approximated (in terms of risk) by the sequence of tests $\{\phi_\ell,\ell=1,2,\dots\}$ such that $\phi_\ell$ is a BDS test for some prior $\mu_\ell$, and $\int_\Theta R(\theta,\phi_\ell)d\mu_\ell(\theta)\to \sup_{\mu\in\mathcal M}\int_\Theta R(\theta,\phi^\dagger)d\mu(\theta)$.

Covariates

This section extends the base framework to incorporate observable covariates. Each individual experiment is described by $(S,X,U,G,\Theta;\upsilon,m)$, where $S,U,G,\Theta$ are defined as before. We let $X$ denote the finite set of covariate values and $\{\upsilon_\theta,\theta\in\Theta\}$ be a family of distributions on $X$. Throughout, we assume that each $\upsilon_\theta\in\Delta(X)$ has full support on $X$. Measure $m_\theta(\cdot|x)$ then determines the conditional law of $u$ given $x$. The prediction of the model is then summarized by a weakly measurable correspondence $(u,x)\mapsto F(u,x|\theta)\equiv \{(s,x):s\in G(u|\theta,x) \}\subset S\times X$ for each $\theta\in\Theta$. As before, this correspondence induces a belief function on $S\times X$

align[align omitted — 115 chars of source]

If $A$ is a rectangle $A=A_s\times A_x$ for some $A_s\subset S$ and $A_x\subset X$, one may write it as

align[align omitted — 141 chars of source]

which can be viewed as the mean of the conditional belief function $\nu_\theta(\cdot|x)$. The subsequent analysis, starting with the Neyman--Pearson lemma, is then essentially the same as before.

A simplification occurs when $\upsilon$ does not depend on $\theta$. Consider $\theta_0,\theta_1\in\Theta$ with $\theta_0\ne\theta_1$. Because of the additivity of $\upsilon$, it suffices to consider sets of the form $B\times \{x\}$, where $B\subset S$ and $x\in X.$ Then, the program that determines the LFP is

align[align omitted — 304 chars of source]

Observe that the constraints simplify to

align[align omitted — 119 chars of source]

Taking $B=S$, one obtains $\upsilon(x)\le \upsilon_j(x)$ for all $x\in X$ with $j=0,1.$ Since $\upsilon$ is a measure, this implies $\upsilon_0=\upsilon_1=\upsilon,$ and the resulting LR statistic does not depend on $\upsilon$. The LR statistic $dQ_1/dQ_0=q_1(s|x)/q_0(s|x)$ can then be calculated by solving, for each $x$,

align[align omitted — 304 chars of source]

Concluding Remarks

This study explored robust likelihood-based inference methods for incomplete economic models. A key result is the existence of an LFP consisting of product measures, through which we may connect incomplete models to standard frameworks and thus obtain asymptotic approximations and analyze optimality properties while remaining agnostic about the selection. However, some problems need further work. They include methods of inference on the subcomponents of $\theta$, especially when the dimension of the parameter is moderately high and a framework that can handle solution concepts involving some form of mixing (e.g., mixed Nash, Bayes correlated equilibria) which are undertaken in ongoing work. Another important avenue for future research is an extension of the current results to statistical inference or decision problems outside hypothesis testing.

\singlespacing \ifx\undefined\leavevmode\rule[.5ex]{3em}{.5pt}\ \fi \ifx\undefined\textsc \let\tmpsmall\tmpsmall\sc \fi

thebibliography\harvarditem[Aliprantis and Border]{Aliprantis and Border}{2006}{aliprantis2006infinite} {\sc Aliprantis, C. D., {\tmpsmall\sc and} K. Border} (2006): {\em Infinite dimensional analysis: a hitchhiker's guide\/}. Springer Science & Business Media. \harvarditem[Andrews and Barwick]{Andrews and Barwick}{2012}{andrews2012inference} {\sc Andrews, D. W., {\tmpsmall\sc and} P. J. Barwick} (2012): “Inference for parameters defined by moment inequalities: A recommended moment selection procedure,” {\em Econometrica\/}, 80(6), 2805--2826. \harvarditem[Andrews and Ploberger]{Andrews and Ploberger}{1994}{Andrews:1994aa} {\sc Andrews, D. W., {\tmpsmall\sc and} W. Ploberger} (1994): “Optimal tests when a nuisance parameter is present only under the alternative,” {\em Econometrica: Journal of the Econometric Society\/}, pp. 1383--1414. \harvarditem[Andrews and Ploberger]{Andrews and Ploberger}{1995}{Andrews:1995aa} {\sc \leavevmode\rule[.5ex]{3em}{.5pt}\ } (1995): “Admissibility of the likelihood ratio test when a nuisance parameter is present only under the alternative,” {\em The Annals of Statistics\/}, pp. 1609--1629. \harvarditem[Andrews and Soares]{Andrews and Soares}{2010}{andrews2010inference} {\sc Andrews, D. W., {\tmpsmall\sc and} G. Soares} (2010): “Inference for parameters defined by moment inequalities using generalized moment selection,” {\em Econometrica\/}, 78(1), 119--157. \harvarditem[Andrews]{Andrews}{1999}{Andrews:1999aa} {\sc Andrews, D. W. K.} (1999): “Estimation When a Parameter is on a Boundary,” {\em Econometrica\/}, 67(6), 1341--1383. \harvarditem[Aradillas-Lopez and Tamer]{Aradillas-Lopez and Tamer}{2008}{Aradillas-Lopez:2008ab} {\sc Aradillas-Lopez, A., {\tmpsmall\sc and} E. Tamer} (2008): “The Identification Power of Equilibrium in Simple Games,” {\em Journal of Business & Economic Statistics\/}, 26(3), 261--283. \harvarditem[Armstrong]{Armstrong}{2014}{Armstrong:2014aa} {\sc Armstrong, T. B.} (2014): “Weighted KS statistics for inference on conditional moment inequalities,” {\em Journal of Econometrics\/}, 181(2), 92 -- 116. \harvarditem[Armstrong]{Armstrong}{2018}{Armstrong:2018aa} {\sc \leavevmode\rule[.5ex]{3em}{.5pt}\ } (2018): “On the choice of test statistic for conditional moment inequalities,” {\em Journal of Econometrics\/}, 203(2), 241 -- 255. \harvarditem[Bednarski]{Bednarski}{1982}{Bednarski1982} {\sc Bednarski, T.} (1982): “Binary experiments, minimax tests and 2-alternating capacities,” {\em The Annals of Statistics\/}, pp. 226--232. \harvarditem[Beresteanu, Molchanov, and Molinari]{Beresteanu, Molchanov, and Molinari}{2011}{beresteanu2011sharp} {\sc Beresteanu, A., I. Molchanov, {\tmpsmall\sc and} F. Molinari} (2011): “Sharp identification regions in models with convex moment predictions,” {\em Econometrica\/}, 79(6), 1785--1821. \harvarditem[Beresteanu and Molinari]{Beresteanu and Molinari}{2008}{beresteanu2008asymptotic} {\sc Beresteanu, A., {\tmpsmall\sc and} F. Molinari} (2008): “Asymptotic properties for a class of partially identified models,” {\em Econometrica\/}, 76(4), 763--814. \harvarditem[Berry]{Berry}{1992}{berry1992estimation} {\sc Berry, S. T.} (1992): “Estimation of a Model of Entry in the Airline Industry,” {\em Econometrica: Journal of the Econometric Society\/}, pp. 889--917. \harvarditem[Boyd and Vandenberghe]{Boyd and Vandenberghe}{2004}{Boyd:2004jk} {\sc Boyd, S., {\tmpsmall\sc and} L. Vandenberghe} (2004): {\em Convex optimization\/}. Cambridge university press. \harvarditem[Bresnahan and Reiss]{Bresnahan and Reiss}{1990}{bresnahan1990entry} {\sc Bresnahan, T. F., {\tmpsmall\sc and} P. C. Reiss} (1990): “Entry in monopoly market,” {\em The Review of Economic Studies\/}, 57(4), 531--553. \harvarditem[Bresnahan and Reiss]{Bresnahan and Reiss}{1991}{bresnahan1991empirical} {\sc \leavevmode\rule[.5ex]{3em}{.5pt}\ } (1991): “Empirical models of discrete games,” {\em Journal of Econometrics\/}, 48(1), 57--81. \harvarditem[Bugni, Canay, and Shi]{Bugni, Canay, and Shi}{2017}{Shicanaybugni2015inference} {\sc Bugni, F., I. Canay, {\tmpsmall\sc and} X. Shi} (2017): “Inference for functions of partially identified parameters in moment inequality models,” {\em Quantitative Economics\/}, 8, 1--38. \harvarditem[Bugni]{Bugni}{2010}{bugni2010bootstrap} {\sc Bugni, F. A.} (2010): “Bootstrap inference in partially identified models defined by moment inequalities: Coverage of the identified set,” {\em Econometrica\/}, 78(2), 735--753. \harvarditem[Canay]{Canay}{2010}{CANAY2010408} {\sc Canay, I. A.} (2010): “EL inference for partially identified models: Large deviations optimality and bootstrap validity,” {\em Journal of Econometrics\/}, 156(2), 408 -- 425. \harvarditem[Chamberlain]{Chamberlain}{2000}{Chamberlain:2000aa} {\sc Chamberlain, G.} (2000): “Econometric applications of maxmin expected utility,” {\em Journal of Applied Econometrics\/}, pp. 625--644. \harvarditem[Chen, Tamer, and Torgovitsky]{Chen, Tamer, and Torgovitsky}{2011}{Chen:2011aa} {\sc Chen, X., E. T. Tamer, {\tmpsmall\sc and} A. Torgovitsky} (2011): “Sensitivity Analysis in Semiparametric Likelihood Models,” {\em Cowles Foundation Discussion Paper No. 1836\/}. \harvarditem[Chernozhukov, Hong, and Tamer]{Chernozhukov, Hong, and Tamer}{2007}{chernozhukov2007estimation} {\sc Chernozhukov, V., H. Hong, {\tmpsmall\sc and} E. Tamer} (2007): “Estimation and confidence regions for parameter sets in econometric models1,” {\em Econometrica\/}, 75(5), 1243--1284. \harvarditem[Chesher and Rosen]{Chesher and Rosen}{2017}{Chesher:2014aa} {\sc Chesher, A., {\tmpsmall\sc and} A. Rosen} (2017): “Generalized instrumental variable models,” {\em Econometrica\/}, 85(3), 959--989. \harvarditem[Cho and White]{Cho and White}{2011}{Cho:2011aa} {\sc Cho, J. S., {\tmpsmall\sc and} H. White} (2011): “Generalized runs tests for the IID hypothesis,” {\em Journal of econometrics\/}, 162(2), 326--344. \harvarditem[Choquet]{Choquet}{1954}{Choquet1954} {\sc Choquet, G.} (1954): “Theory of capacities,” in {\em Annales de l'institut Fourier\/}, vol. 5, pp. 131--295. \harvarditem[Ciliberto and Tamer]{Ciliberto and Tamer}{2009}{Ciliberto:2009aa} {\sc Ciliberto, F., {\tmpsmall\sc and} E. Tamer} (2009): “Market Structure and Multiple Equilibria in Airline Markets,” {\em Econometrica\/}, 77(6), 1791--1828. \harvarditem[de Paula and Tang]{de Paula and Tang}{2012}{dePaula2012inference} {\sc de Paula, {\'A}., {\tmpsmall\sc and} X. Tang} (2012): “Inference of signs of interaction effects in simultaneous games with incomplete information,” {\em Econometrica\/}, pp. 143--172. \harvarditem[Dempe]{Dempe}{1993}{Dempe_Directional_1993} {\sc Dempe, S.} (1993): “Directional differentiability of optimal solutions under Slater's condition,” {\em Math Program\/}, 59(1-3), 49--69. \harvarditem[Dempster]{Dempster}{1967}{dempster1967upper} {\sc Dempster, A. P.} (1967): “Upper and lower probabilities induced by a multivalued mapping,” {\em The annals of mathematical statistics\/}, pp. 325--339. \harvarditem[Eizenberg]{Eizenberg}{2014}{EIZENBERG:2014aa} {\sc Eizenberg, A.} (2014): “Upstream Innovation and Product Variety in the U.S. Home PC Market,” {\em The Review of Economic Studies\/}, 81(3 (288)), 1003--1045. \harvarditem[Elliott, M{\"u}ller, and Watson]{Elliott, M{\"u}ller, and Watson}{2015}{Elliott:2015aa} {\sc Elliott, G., U. K. M{\"u}ller, {\tmpsmall\sc and} M. W. Watson} (2015): “Nearly optimal tests when a nuisance parameter is present under the null hypothesis,” {\em Econometrica\/}, 83(2), 771--811. \harvarditem[Epstein, Kaido, and Seo]{Epstein, Kaido, and Seo}{2016}{Epstein:2016qv} {\sc Epstein, L. G., H. Kaido, {\tmpsmall\sc and} K. Seo} (2016): “Robust Confidence Regions for Incomplete Models,” {\em Econometrica\/}, 84(5), 1799--1838. \harvarditem[Epstein and Seo]{Epstein and Seo}{2015}{epstein2015exchangeable} {\sc Epstein, L. G., {\tmpsmall\sc and} K. Seo} (2015): “Exchangeable capacities, parameters and incomplete theories,” {\em Journal of Economic Theory\/}, 157, 879--917. \harvarditem[Fang]{Fang}{2014}{Fang:2014aa} {\sc Fang, Z.} (2014): “Optimal plug-in estimators of directionally differentiable functionals,” {\em Working Paper, Texas A&M\/}. \harvarditem[Fang and Santos]{Fang and Santos}{2018}{Fang2018} {\sc Fang, Z., {\tmpsmall\sc and} A. Santos} (2018): “Inference on directionally differentiable functions,” {\em The Review of Economic Studies\/}, 86(1), 377--412. \harvarditem[Galichon and Henry]{Galichon and Henry}{2006}{galichon2006inference} {\sc Galichon, A., {\tmpsmall\sc and} M. Henry} (2006): “Inference in incomplete models,” {\em Available at SSRN 886907\/}. \harvarditem[Galichon and Henry]{Galichon and Henry}{2009}{galichon2009test} {\sc \leavevmode\rule[.5ex]{3em}{.5pt}\ } (2009): “A test of non-identifying restrictions and confidence regions for partially identified parameters,” {\em Journal of Econometrics\/}, 152(2), 186--196. \harvarditem[Galichon and Henry]{Galichon and Henry}{2011}{galichon2011set} {\sc \leavevmode\rule[.5ex]{3em}{.5pt}\ } (2011): “Set identification in models with multiple equilibria,” {\em The Review of Economic Studies\/}, p. rdr008. \harvarditem[Galichon and Henry]{Galichon and Henry}{2013}{galichon2013dilation} {\sc \leavevmode\rule[.5ex]{3em}{.5pt}\ } (2013): “Dilation bootstrap,” {\em Journal of Econometrics\/}, 177(1), 109--115. \harvarditem[Gauvin]{Gauvin}{1977}{Gauvin:1977aa} {\sc Gauvin, J.} (1977): “A necessary and sufficient regularity condition to have bounded multipliers in nonconvex programming,” {\em Mathematical Programming\/}, 12(1), 136--138. \harvarditem[Gul and Pesendorfer]{Gul and Pesendorfer}{2014}{Gul:2014aa} {\sc Gul, F., {\tmpsmall\sc and} W. Pesendorfer} (2014): “Expected uncertain utility theory,” {\em Econometrica\/}, 82(1), 1--39. \harvarditem[Haile and Tamer]{Haile and Tamer}{2003}{Haile:2003to} {\sc Haile, P. A., {\tmpsmall\sc and} E. Tamer} (2003): “Inference with an Incomplete Model of English Auctions,” {\em Journal of Political Economy\/}, 111(1). \harvarditem[H{\"a}usler and Luschgy]{H{\"a}usler and Luschgy}{2015}{Hausler:2015aa} {\sc H{\"a}usler, E., {\tmpsmall\sc and} H. Luschgy} (2015): {\em Stable Convergence and Stable Limit Theorems\/}, vol. 74. Springer. \harvarditem[Heckman and Honor{\'e}]{Heckman and Honor{\'e}}{1990}{Heckman:1990aa} {\sc Heckman, J. J., {\tmpsmall\sc and} B. E. Honor{\'e}} (1990): “The Empirical Content of the Roy Model,” {\em Econometrica\/}, 58(5), 1121--1149. \harvarditem[Hirano and Porter]{Hirano and Porter}{2012}{Hirano_Impossibility_2012} {\sc Hirano, K., {\tmpsmall\sc and} J. R. Porter} (2012): “Impossibility Results for Nondifferentiable Functionals,” {\em Econometrica\/}, 80(4), 1769--1790. \harvarditem[Hong and Li]{Hong and Li}{2018}{HongLi2018JoE} {\sc Hong, H., {\tmpsmall\sc and} J. Li} (2018): “The numerical delta method,” {\em Journal of Econometrics\/}, 206(2), 379--394. \harvarditem[Huber]{Huber}{1981}{Huber:1981aa} {\sc Huber, P. J.} (1981): {\em Robust statistics.\/} Wiley, New York. \harvarditem[Huber and Strassen]{Huber and Strassen}{1973}{Huber:1973aa} {\sc Huber, P. J., {\tmpsmall\sc and} V. Strassen} (1973): “Minimax tests and the Neyman-Pearson lemma for capacities,” {\em The Annals of Statistics\/}, pp. 251--263. \harvarditem[Huber and Strassen]{Huber and Strassen}{1974}{Huber:1974aa} {\sc \leavevmode\rule[.5ex]{3em}{.5pt}\ } (1974): “Note: Correction to minimax tests and the Neyman-Pearson lemma for capacities,” {\em The Annals of Statistics\/}, 2(1), 223--224. \harvarditem[Jovanovic]{Jovanovic}{1989}{jovanovic1989observable} {\sc Jovanovic, B.} (1989): “Observable implications of models with multiple equilibria,” {\em Econometrica: Journal of the Econometric Society\/}, pp. 1431--1437. \harvarditem[Kaido, Molinari, and Stoye]{Kaido, Molinari, and Stoye}{2019}{Kaido:2016aa} {\sc Kaido, H., F. Molinari, {\tmpsmall\sc and} J. Stoye} (2019): “Confidence intervals for projections of partially identified parameters,” {\em Econometrica\/}, 87(4), 1397--1432. \harvarditem[Kaido and Santos]{Kaido and Santos}{2014}{Kaido:2014aa} {\sc Kaido, H., {\tmpsmall\sc and} A. Santos} (2014): “Asymptotically Efficient Estimation of Models Defined by Convex Moment Inequalities,” {\em Econometrica\/}, 82(1), 387--413. \harvarditem[Kawai and Watanabe]{Kawai and Watanabe}{2013}{Kawai:2013aa} {\sc Kawai, K., {\tmpsmall\sc and} Y. Watanabe} (2013): “Inferring Strategic Voting,” {\em The American Economic Review\/}, 103(2), 624--662. \harvarditem[Le Cam]{Le Cam}{1972}{Le-Cam:1972aa} {\sc Le Cam, L.} (1972): “Limits of experiments,” in {\em Proceedings of the Sixth Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Theory of Statistics\/}, pp. 245--261, Berkeley, Calif. University of California Press. \harvarditem[Le Cam]{Le Cam}{1986}{LeCam:1986aa} {\sc Le Cam, L.} (1986): “Asymptotic Methods in Statistical Decision Theory,” {\em Springer Series in Statistics\/}. \harvarditem[Lehmann and Romano]{Lehmann and Romano}{2006}{Lehmann:2006aa} {\sc Lehmann, E. L., {\tmpsmall\sc and} J. P. Romano} (2006): {\em Testing statistical hypotheses\/}. Springer Science & Business Media. \harvarditem[Luo and Wang]{Luo and Wang}{2017a}{Luo:2017aa} {\sc Luo, Y., {\tmpsmall\sc and} H. Wang} (2017a): “Core determining class and inequality selection,” {\em American Economic Review, Papers & Proceedings\/}. \harvarditem[Luo and Wang]{Luo and Wang}{2017b}{Luo:2017ab} {\sc Luo, Y., {\tmpsmall\sc and} H. Wang} (2017b): “Core Determining Class: Construction, Approximation and Inference,” Working Paper, University of Florida. \harvarditem[Maccheroni and Marinacci]{Maccheroni and Marinacci}{2005}{Maccheroni:2005aa} {\sc Maccheroni, F., {\tmpsmall\sc and} M. Marinacci} (2005): “A strong law of large numbers for capacities,” {\em The Annals of Probability\/}, 33(3), 1171--1178. \harvarditem[Miyauchi]{Miyauchi}{2016}{Miyauchi:2016aa} {\sc Miyauchi, Y.} (2016): “Structural estimation of pairwise stable networks with nonnegative externality,” {\em Journal of Econometrics\/}, 195(2), 224 -- 235. \harvarditem[Molchanov]{Molchanov}{2006}{Molchanov:2006aa} {\sc Molchanov, I.} (2006): {\em Theory of random sets\/}. Springer Science & Business Media. \harvarditem[Molchanov and Molinari]{Molchanov and Molinari}{2018}{MR3753715} {\sc Molchanov, I., {\tmpsmall\sc and} F. Molinari} (2018): {\em Random sets in econometrics\/}, vol. 60 of {\em Econometric Society Monographs\/}. Cambridge University Press, Cambridge. \harvarditem[Moreira and Moreira]{Moreira and Moreira}{2013}{Moreira:2013aa} {\sc Moreira, H. A., {\tmpsmall\sc and} M. J. Moreira} (2013): “Contributions to the theory of optimal tests,” Working Paper. \harvarditem[Mourifie, Henry, and Meango]{Mourifie, Henry, and Meango}{2018}{Mourifie:2018aa} {\sc Mourifie, I., M. Henry, {\tmpsmall\sc and} R. Meango} (2018): “Sharp Bounds and Testability of a Roy Model of STEM Major Choices,” SSRN Electronic Journal. \harvarditem[Philippe, Debs, and Jaffray]{Philippe, Debs, and Jaffray}{1999}{philippe1999decision} {\sc Philippe, F., G. Debs, {\tmpsmall\sc and} J.-Y. Jaffray} (1999): “Decision making with monotone lower probabilities of infinite order,” {\em Mathematics of Operations Research\/}, 24(3), 767--784. \harvarditem[Rieder]{Rieder}{1994}{Rieder_1994} {\sc Rieder, H.} (1994): “Robust Asymptotic Statistics,” {\em Springer Series in Statistics\/}. \harvarditem[Rieder]{Rieder}{2014}{Rieder2014} {\sc \leavevmode\rule[.5ex]{3em}{.5pt}\ } (2014): “One-sided confidence about functionals over tangent cones,” {\em arXiv preprint arXiv:1412.1701\/}. \harvarditem[Rockafellar]{Rockafellar}{1972}{Rockafellar:1972aa} {\sc Rockafellar, R. T.} (1972): {\em Convex analysis\/}. Princeton university press. \harvarditem[Romano and Shaikh]{Romano and Shaikh}{2008}{romano2008inference} {\sc Romano, J. P., {\tmpsmall\sc and} A. M. Shaikh} (2008): “Inference for identifiable parameters in partially identified econometric models,” {\em Journal of Statistical Planning and Inference\/}, 138(9), 2786--2807. \harvarditem[Shafer]{Shafer}{1982}{shafer1982belief} {\sc Shafer, G.} (1982): “Belief functions and parametric models,” {\em Journal of the Royal Statistical Society. Series B (Methodological)\/}, pp. 322--352. \harvarditem[Shapiro]{Shapiro}{1988}{Shapiro:1988aa} {\sc Shapiro, A.} (1988): “Sensitivity Analysis of Nonlinear Programs and Differentiability Properties of Metric Projections,” {\em SIAM Journal on Control and Optimization\/}, 26(3), 628--645. \harvarditem[Song]{Song}{2014}{Song2014JMA} {\sc Song, K.} (2014): “Local asymptotic minimax estimation of nonregular parameters with translation-scale equivariant maps,” {\em Journal of Multivariate Analysis\/}, 125, 136 -- 158. \harvarditem[Strasser]{Strasser}{1985}{Strasser:1985aa} {\sc Strasser, H.} (1985): {\em Mathematical theory of statistics: statistical experiments and asymptotic decision theory\/}, vol. 7. Walter de Gruyter. \harvarditem[Tamer]{Tamer}{2003}{tamer2003incomplete} {\sc Tamer, E.} (2003): “Incomplete simultaneous discrete response model with multiple equilibria,” {\em The Review of Economic Studies\/}, 70(1), 147--165. \harvarditem[Van der Vaart]{Van der Vaart}{2000}{Van-der-Vaart:2000aa} {\sc Van der Vaart, A. W.} (2000): {\em Asymptotic statistics\/}. Cambridge university press. \harvarditem[van der Vaart and Wellner]{van der Vaart and Wellner}{1996}{vandervaart:wellner:1996} {\sc van der Vaart, A. W., {\tmpsmall\sc and} J. A. Wellner} (1996): {\em Weak Convergence and Empirical Processes: with Applications to Statistics\/}. Springer-Verlag, New York. \harvarditem[Wald]{Wald}{1950}{Wald1950} {\sc Wald, A.} (1950): “Remarks on the estimation of unknown parameters in incomplete systems of equations,” in {\em Statistical Inference in Dynamic Economic Models,\/}, ed. by T. Koopmans, pp. 305--310. John Wiley, New York. \harvarditem[Wald and Wolfowitz]{Wald and Wolfowitz}{1940}{Wald:1940aa} {\sc Wald, A., {\tmpsmall\sc and} J. Wolfowitz} (1940): “On a test whether two samples are from the same population,” {\em The Annals of Mathematical Statistics\/}, 11(2), 147--162. \harvarditem[White]{White}{2001}{White:2001aa} {\sc White, H.} (2001): {\em Asymptotic theory for econometricians\/}. Academic Press, San Diego, rev. ed edn.

Figures

figure[figure omitted — 538 chars of source]
figure[figure omitted — 596 chars of source]
figure[figure omitted — 347 chars of source]

\pagenumbering{arabic}\origappendix \onehalfspacing

center[center omitted — 124 chars of source]

\tmpsmall\sc

Capacities

Let $\Omega$ be a compact metric space and let $\Sigma_\Omega$ denote its Borel $\sigma$-algebra. Let $\mathcal K(\Omega)$ be the set of compact subsets of $\Omega$ endowed with the Hausdorff metric. Let $\mathcal C(\Omega)$ be the set of continuous functions on $\Omega$. Let $\Delta(\Omega)$ be the set of Borel probability measures on $\Omega$ endowed with the weak topology.

A set function $\nu^{*}$ is said to be a {\em capacity\/} if $\nu^{*}$ satisfies the following conditions:

enumerate[label=(\roman*)] • {$\nu^{*}(\emptyset)=0, \nu^{*}(\Omega)=1$,} • {$A\subset B \Rightarrow \nu^{*}(A) \leq \nu^{*}(B)$, for all $A,B\in \Sigma_\Omega$.} • {$A_n \uparrow A \Rightarrow \nu^{*}(A_n) \uparrow \nu^{*}(A)$, for all $\{A_n,n\ge 1\}\subset \Sigma_\Omega$ and $A\in\Sigma_\Omega$.} • {$F_n \downarrow F, F_n$ closed $\Rightarrow \nu^{*}(F_n) \downarrow \nu^{*}(F)$.}

One may define integral operations with respect to capacities as follows. Let $f:\Omega\to\mathbb R$ be a measurable function. The {\em Choquet integral\/} of $f$ with respect to $\nu$ is defined by

align[align omitted — 140 chars of source]

where the integrals on the right hand side are Riemann integrals.

The following result due to Choquet follows from Theorems 1-3 in philippe1999decision.

lemmaLet $\Omega$ be a Polish space. Let $M$ be a probability measure on $\mathcal K(\Omega)$. Let $\mathcal P=\{P\in\Delta(\Omega):P= \int P_KdM(K),P_K\in \Delta(K)\}$. Then, $\nu(\cdot)=\inf_{P\in\mathcal P}P(\cdot)$ is a belief function and satisfies \begin{align} \nu(A)=M(\{K\subset A\}). \end{align}

In each experiment characterized by the tuple $(S,U,\Theta,G;m)$, one may apply the lemma above with $\mathcal P=\mathcal P_{\theta}$, $\nu=\nu_\theta$, $K=G(u|\theta)$, and $M$ is the law of $K$ induced by correspondence $G$ and $m_\theta$.

Proofs

Proof of Lemma (ref)

proof[\rm Proof of Lemma (ref)] It is straightforward to show $\nu^*_\theta$ is a capacity satisfying conditions (i)-(iv) in Appendix (ref). Since $G(\cdot|\theta)$ is weakly measurable, the map $u\mapsto G(u|\theta)$ defines a measurable map from $U$ to $\mathcal K(S)$. Let $\tilde m$ be the induced measure of $m_\theta$ on $\mathcal K(S)$ by this map. Then, by Lemma (ref), $P\in \mathcal P_\theta$ is equivalent to $P\in core(\nu)$ for an infinitely monotone capacity $\nu$ such that $\nu(A)=\tilde m(K\subset A)$ for all $A\in \Sigma_\Omega.$ By Lemma 2.5 in HS, \begin{align} \nu(A)=\inf_{P\in\mathcal P_\theta}P(A)=\nu_\theta(A), for all A\in \Sigma_S, \end{align} and hence $\nu_\theta$ is infinitely monotone. By the previous step, $\nu_{\theta_1}^{*}$ and $\nu_{\theta_0}^{*}$ are also 2-alternating and 2-monotone capacities respectively. Let $\Lambda$ be the Radon-Nikodym derivative of $\nu_{\theta_1}^{*}$ and $\nu_{\theta_0}^{*}$ in the sense of HS (Section 3). Then, by their Theorem 4.1, the conclusion of the lemma follows.

Proof of Theorem (ref), Corollary (ref) and Auxiliary Lemmas

We use Theorem 8.1.1 in Lehmann:2006aa to show Theorem (ref). For ease of reference, we copy their theorem below (with a slight change of notation to avoid conflicts). For this, let $E,E'$ be measurable spaces and let $\mathcal P=\{P_\eta\in \Delta(S):\eta\in E\cup E'\}$ be families of probability distributions on $S$ with densities $p_\eta=dP_\eta/d\upsilon$ parameterized by $\eta\in E\cup E'$. Throughout, we assume that the map $(s,\eta)\mapsto p_\eta(s)$ is jointly measurable.

theorem[Theorem 8.1.1. of Lehmann:2006aa] For any distributions $\mu,\mu'$ over $\Sigma_{E}$ and $\Sigma_{E'}$, let $\phi_{\mu,\mu'}$ be the most powerful test for testing \begin{align} f(s)=\int_E p_\eta(s)d\mu(\eta) \end{align} at level $\alpha$ against \begin{align} f'(s)=\int_{E'} p_\eta(s)d\mu(\eta) \end{align} and let $\beta_{\mu,\mu'}$ be its power against the alternative $f'.$ If there exist $\mu$ and $\mu'$ such that \begin{align} \sup_{\eta\in E}E_{P_\eta}[\phi_{\mu,\mu'}(s)]&\le \alpha\\ \inf_{\eta\in E'}E_{P_\eta'}[\phi_{\mu,\mu'}(s)]&=\beta_{\mu,\mu'}, \end{align} then: \begin{itemize} • $\phi_{\mu,\mu'}$ maximizes $\inf_{\eta\in E'}E_{P_\eta'}[\phi_{\mu,\mu'}(s)]$ among all level-$\alpha$ tests of the hypothesis $H:\eta\in E$ and is the unique test with this property if it is the unique most powerful level-$\alpha$ test for testing $f$ against $f'$. • The pair of distributions $\mu,\mu'$ is least favorable in the sense that for any other pair $\tilde \mu,\tilde\mu'$ we have \begin{align} \beta_{\mu,\mu'}\le\beta_{\tilde\mu,\tilde\mu'}. \end{align} \end{itemize}
lemmaLet $\nu_{\theta}$ be defined as in ((ref)), and let $\nu_{\theta}^{*}$ be its conjugate. Let $f: S\rightarrow \mathbb{R}$ be a measurable function. Similarly, for each $i\in \mathbb N$, let $f_i: S\rightarrow \mathbb{R}$ be a measurable function. Then, \begin{itemize} • There exists a minimizing measure $Q \in \Delta(S)$ and a maximizing measure $Q^*\in\Delta(S)$ such that, for any $t \in \mathbb{R}$, \begin{equation} {\nu}_{\theta}(f(s) > t)=Q(f(s) > t), \end{equation} and \begin{equation} {\nu}_{\theta}^{*}(f(s) > t)=Q^{*}(f(s) > t). \end{equation} • If, for each $i\in\mathbb N$ and each measurable function $f_i$, $Q_i,Q_i^*$ are the minimizing and maximizing measures in the sense of (ref)-(ref), it follows that \begin{equation} \nu_\theta^n\Big(\sum_{i=1}^nf_i(s_i) > t\Big)=Q^n\Big(\sum_{i=1}^nf_i(s_i) > t\Big), \end{equation} and \begin{equation} \nu_\theta^{*,n}\Big(\sum_{i=1}^nf_i(s_i) > t\Big)=Q^{*n}\Big(\sum_{i=1}^nf_i(s_i) > t\Big), \end{equation} for all $t\in\mathbb R$, where $Q^{n}=\bigotimes_{i=1}^nQ_i$ and $Q^{* n}=\bigotimes_{i=1}^nQ_i^* \in \Delta(S^{n})$. \end{itemize}
proof(i) As shown in the proof of Lemma (ref), $\nu_{\theta}^{*}$ is a $2$-alternating capacity. Since $S$ is finite, any function on $S$ is upper semi-continuous by the continuity of $f$. By Lemma 2.4 in HS, there exists a probability measure $p^*\in \Delta(S)$ such that for all $t \in \mathbb{R}$, ${\nu}_{\theta}^{*}(f(s) > t)={p}^*(f(s) > t)$. This ensures (ref). Similarly, let $g=-f$ and note that $g$ is again upper semicontinuous. Applying Lemma 2.4 in HS to the event $\{g\ge -t\}$, there exists $p\in\Delta(s)$ such that, for any $t\in \mathbb R,$ \begin{align} \nu^*_\theta(g\ge -t)=p(g\ge t)& \Leftrightarrow 1-\nu^*_\theta(g< -t)=1-p(g<t)\notag\\ & \Leftrightarrow \nu_\theta(f>t)=p(f>t). \end{align} This therefore establishes (ref). (ii) For each $i$, let $Y_i\equiv \min_{s_i \in G(u_i|\theta)} f_i(s_i)$ and $Z_i\equiv f_i(s_i)$. Note that $Y_i$ is a function of $u_i$, and hence we use $m_{\theta,i}$ for the law of $Y_i$ induced by $u_i$. For each experiment, we have \begin{equation} G(u|\theta) \subseteq \{s\in S: f_i(s)>t \} \Leftrightarrow \min_{s \in G(u|\theta)} f_i(s) > t. \end{equation} Therefore, by Lemma (ref), \begin{equation} \begin{split} \nu_{\theta,i}(f_i(s_i) > t) & = m_{\theta,i}(\min_{s_i \in G(u_i|\theta)} f_i(s_i) > t)= m_{\theta,i}(Y_i> t), \forall t\in\mathbb R. \end{split} \end{equation} By (i), there is $Q_i\in\Delta (S)$ such that \begin{equation} \nu_{\theta,i}(f_i(s_i) > t)=Q_i(Z_i > t), \forall t\in\mathbb R. \end{equation} Hence, by (ref)-(ref), $Y_i \buildrel d \over = Z_i$ for all $i$. Let $\mathcal{P}_{\theta}^{n}$ be defined as in ((ref)) and let $\nu^{n}_{\theta}$,$\nu_{\theta}^{* n}$ be the lower and upper probabilities of $\mathcal{P}_{\theta}^{n}$ respectively. By Lemma (ref), $\nu_{\theta}^{n}$ is a belief function and $\nu_{\theta}^{*n}$ is its conjugate. Therefore, \begin{equation} \nu_{\theta}^{n}\Big(\sum_{i=1}^nf_i({s}_i) > t \Big) =m_{\theta}^{n}\Big(u^{n} \in U^{n}:G^{n}(u^{n}|\theta) \subseteq \{\sum_{i=1}^nf_i({s}_i) > t\} \} \Big). \end{equation} Since $G^{n} (u^{n}|\theta)=\prod_{i=1}^{n}G({u}_i|\theta)$, inside the parenthesis we have: \begin{align} \prod_{i=1}^{n}G({u}_i|\theta) &\subseteq \{{s}^{n}: \sum_{i=1}^nf_i(s_i) > t\}\notag\\ &\Leftrightarrow \min_{s^{n} \in \prod_{i=1}^{n}G(u_i|\theta)} \sum_{i=1}^nf_i(s_i) > t \Leftrightarrow \sum_{i=1}^{n} \min_{s_i \in G(u_i|\theta)} f_i(s_i) > t. \end{align} By (ref)-(ref) and recalling that $Y_i=\min_{s_i\in G({u}_i|\theta)} f_i(s_i)$, we have \begin{align} \nu_{\theta}^{n}\Big(\sum_{i=1}^nf_i({s}_i) > t \Big)=m^{n}_{\theta}\Big(\sum_{i=1}^{n} \min_{s_i\in G({u}_i|\theta)} f_i(s_i) > t\Big)=m_{\theta}^{n} \Big(\sum_{i=1}^{n} Y_i >t \Big). \end{align} Let $\{Y_1,Y_2, \dotsi, Y_i, \dotsi \}$ be independently distributed according to $m_{\theta}^n$, and let $\{Z_1,Z_2, \dotsi, Z_i \dotsi \}$ be independently distributed according to $Q^n$. Then, $\sum_{i=1}^{n} Y_i \buildrel d \over = \sum_{i=1}^{n}Z_i$ because $(Y_1,\dots,Y_n) \buildrel d \over = (Z_1,\dots,Z_n)$. Therefore, for all $t\in\mathbb R$, \begin{align} m^{n}_{\theta}\Big(\sum_{i=1}^{n} Y_i > t\Big)=Q^{n}(\sum_{i=1}^{n} Z_i > t). \end{align} By (ref)-(ref) and $\nu^n_\theta$ being the lower probability of $\mathcal P_\theta^n$, we have \begin{align} \min_{P \in \mathcal{P}_{\theta}^{n}}P(\sum_{i=1}^nf_i(s_i) > t)=\nu_{\theta}^{n} \Big(\sum_{i=1}^{n}f_i(s_i) > t\Big)=Q^{n}(\sum_{i=1}^nf_i(s_i) > t). \end{align} This establishes (ref). One may show (ref) by a similar argument.
proof[\rm Proof of Theorem (ref)] Recall that $Q^n_0=\otimes_{i=1}^n Q_{0,i}$, $Q^n_1=\otimes_{i=1}^n Q_{1,i}$, and $\Lambda_n$ is a version of the Radon-Nikodym derivative of them. We follow Section 8.3 in Lehmann:2006aa and show the following statements: \begin{itemize} • When $s^n$ is distributed according to a distribution in $\mathcal P^n_{\theta_0}$, the probability of the event $\{s^n:\Lambda_n>t\}$ is largest (for any $t$), i.e. $\Lambda_n$ is stochastically largest, when the distribution of $s^n$ is $Q^n_0=\otimes_{i=1}^n Q_{0,i}$. • When $s^n$ is distributed according to a distribution in $\mathcal P^n_{\theta_1}$, the probability of the event $\{s^n:\Lambda_n>t\}$ is smallest (for any $t$), i.e. $\Lambda_n$ is stochastically smallest, when the distribution of $s^n$ is $Q^n_1=\otimes_{i=1}^n Q_{1,i}$. • $\Lambda_n$ is stochastically larger when the distribution of $s$ is $Q_1^n$ than when it is $Q_0^n$. \end{itemize} These statements are summarized by \begin{align} Q^{n,\prime}_0(\Lambda_n>t)\stackrel{(a)}{\le} Q_0^n(\Lambda_n>t)\stackrel{(c)}{\le} Q_1^n(\Lambda_n>t)\stackrel{(b)}{\le} Q^{n,\prime}_1(\Lambda_n>t), \end{align} for all $t$, $Q^{n,\prime}_0\in\mathcal P^n_{\theta_0}$, and $Q^{n,\prime}_1\in\mathcal P^n_{\theta_1}.$ Below, we invoke Lemma (ref). For this, let $f_i(\cdot)=\ln \Lambda_i(\cdot)$, where $\Lambda_i\in dQ_{1,i}/dQ_{0,i}$. Let $(Q^{*n},Q^n)$ be the product measures in Lemma (ref) with $f_i=\ln \Lambda_i$ for $i=1,\dots,n$. Note that $\Lambda_n>t$ is equivalent to $\sum_{i=1}^nf_i(s_i)>\ln t$. By Lemma (ref) with $t'=\ln t$, it then follows that \begin{align} \nu^{*n}_{\theta_0}(\Lambda_n>t)=\nu_{\theta_0}^{*n}\big(\sum_{i=1}^n f_i(s_i)>t'\big)=Q^{*n}\big(\sum_{i=1}^n f_i(s_i)>t'\big)=Q^{*n}(\Lambda_n>t), \end{align} where $Q^{*n}=Q_0^n$. Recall that $\nu^{*n}_{\theta_0}(\Lambda_n>t)=\sup_{Q_0^{n,\prime}\in \mathcal P_{\theta_0}^n}Q_0^{n,\prime}(\Lambda_n>t)$. This therefore means $Q^{n}_0$ makes $\Lambda_n$ stochastically largest among all distributions in $\mathcal P^n_{\theta_0}$ and hence ensures inequality (a) in (ref). Similarly, again by Lemma (ref), \begin{align} \nu^n_{\theta}(\Lambda_n>t)=\nu_{\theta}^n\big(\sum_{i=1}^n f_i(s_i)>t'\big)=Q^n\big(\sum_{i=1}^n f_i(s_i)>t'\big)=Q^n(\Lambda_n>t),. \end{align} where $Q^n=Q^{n}_1$. Therefore, $Q^n$ makes $\Lambda_n$ stochastically smallest and hence ensures inequality (c) in (ref). The middle inequality in (ref) follows from Corollary 3.2.1 in Lehmann:2006aa and the Neyman-Pearson lemma. Let $E=\mathcal P^n_{\theta_0},E'=\mathcal P^n_{\theta_1}$. Let $\mu,\mu'\in \Delta(E)$ be distributions, each assigning probability 1 to a single distribution, $\mu$ to $Q_0^n\in\mathcal P^n_{\theta_0}$ and $\mu'$ to $Q_1^n\in\mathcal P^n_{\theta_0}$. Let $(C_n,\gamma_n)$ be chosen so that $E_{Q_0^n}[\phi_n(s^n)]=\alpha$, where $\phi_n$ is the likelihood-ratio test defined as in (ref). The argument above shows that $\mu,\mu'$ satisfy (ref)-(ref). The conclusion of the theorem then follows from applying Theorem (ref) to the present setting.
proof[\rm Proof of Corollary (ref)] By Theorem (ref), the LFP $(Q_0^n,Q_i^n)$ exists, and they are product measures. Note that, in the application of Lemma (ref), $Q_i$ (and $Q^*_i$) is identical across $i$ because $\nu_\theta$ (and $\nu^*_\theta$) is identical across $i$. The conclusion of the Corollary then follows by arguing as in the proof of Theorem (ref).

Proof of Proposition (ref) and Theorems (ref)-(ref)

proof[\rm Proof of Proposition (ref)] Note that \begin{align} \sup_{P\in\mathcal P^n_{\theta_0}}E_{P}[\phi_n^*(s^n)]&= \sup_{P\in\mathcal P^n_{\theta_0}}P(\Lambda_n(s^n)> C_n^*)\notag\\ &=\nu^{*n}_{\theta_0}(\Lambda_n(s^n)> C_n^*)\notag\\ &=Q_0^n\big(\Lambda_n(s^n)> C_n^*\big)\notag\\ &=Q_0^n\Big(\frac{1}{\sqrt n}\sum_{i=1}^n\ln\frac{dQ_1}{dQ_0}(s_i)-E_{Q_0}[\ln \frac{dQ_1}{dQ_0}(s_i)]> \sigma_{Q_0}z_{\alpha}\Big), \end{align} where the second equality follows from $\nu^{*n}$ being the upper probability of $\mathcal P^n_{\theta_0}$, and the third equality follows from $Q_0^n$ being the least favorable null distribution by Theorem (ref). For each $i$, let $Z_i\equiv \ln \frac{dQ_1}{dQ_0}(s_i)$. Under $Q_0^n$, $\{Z_i\}_{i=1}^n$ is an i.i.d. sequence with a finite variance due to $\sigma^2_{Q_0}<\infty$. Hence, if $\sigma_{Q_0}>0$, by the CLT for i.i.d. random variables, one obtains \begin{align} \lim_{n\to\infty} Q_0^n\Big(\frac{1}{\sqrt n}\sum_{i=1}^n\frac{\ln\frac{dQ_1}{dQ_0}(s_i)-E_{Q_0}[\ln \frac{dQ_1}{dQ_0}(s_i)]}{\sigma_{Q_0}}>z_{\alpha}\Big)=Pr\big(Z> z_{\alpha}\big)= \alpha, \end{align} where $Z\sim N(0,1)$. If $\sigma_{Q_0}=0$, the summand in (ref) is identically 0 and hence the probability of the event is zero and hence $\limsup_{n\to\infty}\sup_{P\in\mathcal P^n_{\theta_0}}E_{P}[\phi_n^*(s^n)]\le \alpha.$
proof[\rm Proof of Theorem (ref)] Let $\phi_n$ be a level-$\alpha$ test for $H_0:\varphi(\theta)\le 0$ against $H_1:\varphi(\theta)>0$. Since $\varphi(\theta_0)\le 0$, $\phi_n$ is necessarily a level-$\alpha$ test for testing $\theta=\theta_0$ against $\theta_1=\theta_0+h/\sqrt n$. For any $n$, the lower power of $\phi_n$ is then bounded from above by that of the minimax test in Theorem (ref), which we denote by $\phi^*_n$ below. Thus, \begin{align} \pi_{n,\theta_{n,h}}(\phi_n)=\inf_{P^n\in\mathcal P^n_{\theta_{n,h}}}E_{P^n}[\phi_n]\le \inf_{P^n\in\mathcal P^n_{\theta_{n,h}}}E_{P^n}[\phi^*_n]=\pi_{\theta_{n,h}}(\phi^*_n). \end{align} Let $j\in\mathbb J$ and let $h\in\mathcal T_j(\theta_0).$ By Assumption (ref), the LFP $(Q_{0,j,\tau h},Q_{1,j,\tau h})\in \mathcal P_{\theta_0}\times\mathcal P_{\theta_0+\tau h}$ satisfies \begin{align} Q_{0,j,\tau h}=\mathsf Q_{j,\theta_0}, and Q_{1,j,\tau h}=\mathsf Q_{j,\theta_0+\tau h}, for all 0<\tau\le\bar\tau. \end{align} This and Theorem (ref) imply, for each $j\in\mathbb J$ and $h\in\mathcal T_j(\theta_0)$, \begin{align} \pi_{n,\theta_0+h/\sqrt n}(\phi^*_n)=\int \phi^*_nd\nu^n_{\theta_0+h/\sqrt n}=\int \phi^*_n d\mathsf Q^n_{j,\theta_0+h/\sqrt n}, \end{align} for $n$ sufficiently large. Hence, it suffices to analyze the asymptotic power under $\{\mathsf Q^n_{j,\theta_0+h/\sqrt n}\}.$ The underlying model $\theta\mapsto \mathsf Q_{j,\theta}$ is $L^2$ differentiable tangentially to $\mathcal T_j(\theta_0)$. By Lemma 25.14 in Van-der-Vaart:2000aa, the log-likelihood ratio of the LFP can be expanded as \begin{align} L_n=\ln \frac{d\mathsf Q^n_{j,\theta_0+h/\sqrt n}}{d\mathsf Q^n_{j,\theta_0}}=h' \Delta_{j,n}-\frac{1}{2} h'C_{j} h+o_{\mathsf Q^n_{j,\theta_0}}(1), \end{align} where $\Delta_{j,n}=n^{-1/2}\sum_{i=1}^n\dot\ell_{j,\theta_0}(s_i)$ and hence $L_n\stackrel{\mathsf Q^n_{j,\theta_0}}{\leadsto}N(-\frac{\sigma^2}{2},\sigma^2)$ with $\sigma^2=h'C_jh.$ By Theorem 9.4 of Van-der-Vaart:2000aa, the sequence $\mathcal E_{j,n}$ of localized experiments in (ref) then converges to the Gaussian limit experiment $\mathcal E_j$ in (ref). This ensures that, for any $j\in\mathbb J$ and $h\in\mathcal T_{j}(\theta_0)$, there is a subsequence along which $\pi_{n,\theta_0+h/\sqrt n}(\phi^*_n)\to \pi_h,$ where $\pi_h$ is a power function in the Gaussian limit experiment Van-der-Vaart:2000aa. For any $h\in \mathcal T_{j}(\theta_0)$ with $\dot\varphi_{\theta_0}h<0$, we have $\varphi(\theta_0+h/\sqrt n)<0$ for all $n$ sufficiently large. Hence, by $\theta_0+h/\sqrt n$ satisfying the null for all $n$ sufficiently large and $\phi^*_n$ being level-$\alpha$, \begin{align} \pi_h\le \limsup_{n\to\infty}\pi_{n,\theta_0+h/\sqrt n}(\phi^*_n)\le \alpha. \end{align} By continuity, $\pi_h\le \alpha$ for all $h$ such that $\dot\varphi_{\theta_0}h\le 0$. This, in turn, implies that $\pi_h$ is a power function of a level-$\alpha$ test for testing $H_0:\dot\varphi_{\theta_0}h\le 0$ against $H_1:\dot\varphi_{\theta_0}h> 0$ in $\mathcal E_j$, and hence it is bounded by the power of the uniformly most powerful test. The rest of the proof parallels the proof of Theorem 15.4 in Van-der-Vaart:2000aa if $C_j$ is non-singular and the tangent set is a linear subspace. The first claim of the theorem then holds with $\tilde\varrho_j=\dot\varphi_{\theta_0}C_j^{-1}\dot \ell_{j,\theta_0}.$ In case $C_j$ is singular or the tangent set is a convex cone (not a linear subspace), we follow the argument in the proof of Theorem 2.1 in Rieder2014. Let $h\in\mathcal T_{j}(\theta_0)$ be a vector such that $\dot\varphi_{\theta_0}h = c>0$. We rewrite it as $h=\tau a$, where $\tau>0$ and $a$ is a unit vector. We then let $g=a'\dot\ell_{j,\theta_0}\in\mathcal G_{j,\theta_0}$. Note that, by the definition of $\varrho_j$, \begin{align} \dot\varphi_{\theta_0}a=\langle\varrho_j,g\rangle_{L^2_{\mathsf Q_{j,\theta_0}}}, \end{align} and hence $\tau=c/\langle\varrho_j,g\rangle_{L^2_{\mathsf Q_{j,\theta_0}}}$. One may now rewrite (ref) as \begin{align} L_n=\ln \frac{d\mathsf Q^n_{j,\theta_0+h/\sqrt n}}{d\mathsf Q^n_{j,\theta_0}}=\tau\frac{1}{\sqrt n}\sum_{i=1}^ng(s_i)-\frac{\tau^2}{2} \|g\|_{L^2_{\mathsf Q_{j,\theta_0}}}+o_{\mathsf Q^n_{j,\theta_0}}(1). \end{align} By Corollary 3.4.2 in Rieder_1994, the asymptotic power of any test satisfying (ref) is then dominated by $ 1-\Phi(z_\alpha-\tau\|g\|_{L^2_{\mathsf Q_{j,\theta_0}}}).$ Now let $g\to \tilde\varrho_j$ in $L^2_{\mathsf Q_{j,\theta_0}}(S)$. Then, \begin{align} \tau\|g\|_{L^2_{\mathsf Q_{j,\theta_0}}}=\frac{c \|g\|_{L^2_{\mathsf Q_{j,\theta_0}}}}{\langle\varrho_j,g\rangle_{L^2_{\mathsf Q_{j,\theta_0}}}}\to \frac{c\|\tilde\varrho_j\|_{L^2_{\mathsf Q_{j,\theta_0}}}}{\langle\varrho_j,\tilde\varrho_j\rangle_{L^2_{\mathsf Q_{j,\theta_0}}}}=\frac{c}{\|\tilde\varrho_j\|_{L^2_{\mathsf Q_{j,\theta_0}}}}, \end{align} where the last equality follows from $\langle\varrho_j,\tilde\varrho_j\rangle_{L^2_{\mathsf Q_{j,\theta_0}}}=\|\tilde\varrho_j\|_{L^2_{\mathsf Q_{j,\theta_0}}}^2$ due to $\tilde\varrho_j$ being the projection of $\varrho_j$ to $cl(\mathcal G_{j,\theta_0})$. Therefore, the power bound is obtained as the following limit \begin{align} \lim_{g\to\tilde\varrho_j}1-\Phi\Big(z_\alpha-\tau\|g\|_{L^2_{\mathsf Q_{j,\theta_0}}}\Big)=1-\Phi\Big(z_\alpha-\frac{c}{\|\tilde\varrho_j\|_{L^2_{\mathsf Q_{j,\theta_0}}}}\Big). \end{align} The first claim of the theorem then follows from noting that $c=\dot\varphi_{\theta_0}h=\langle\varrho_j,h'\dot\ell_{j,\theta_0}\rangle_{L^2_{\mathsf Q_{j,\theta_0}}}.$ For the second claim, let $h\in\mathcal T_{j}(\theta_0)$. Then, by Le Cam's third lemma, \begin{align} \frac{1}{\sqrt n}\sum_{i=1}^n\tilde\varrho_j(s_i)\stackrel{\mathsf Q^n_{j,\theta_{n,h}}}{\leadsto}N\big(\langle \tilde\varrho_j,h'\dot\ell_{j,\theta_0}\rangle_{L^2_{\mathsf Q_{j,\theta_0}}},\|\tilde\varrho_j\|_{L^2_{\mathsf Q_{j,\theta_0}}}^2\big). \end{align} Therefore, \begin{align} \lim_{n\to\infty}\pi_{n,\theta_0+h/\sqrt n}(\phi^*_{j,n})&=1-\Phi\Bigg(z_\alpha-\frac{\langle \tilde\varrho_j,h'\dot\ell_{\theta_0}\rangle_{L^2_{\mathsf Q_{j,\theta_0}}}}{\|\tilde{\varrho}_j\|_{L^2_{\mathsf Q_{j,\theta_0}}}}\Bigg) \ge 1-\Phi\Bigg(z_\alpha-\frac{\langle \varrho_j,h'\dot\ell_{\theta_0}\rangle_{L^2_{\mathsf Q_{j,\theta_0}}}}{\|\tilde{\varrho}_j\|_{L^2_{\mathsf Q_{j,\theta_0}}}}\Bigg), \end{align} where the inequality follows from $\langle \tilde\varrho_j,h'\dot\ell_{\theta_0}\rangle_{L^2_{\mathsf Q_{j,\theta_0}}}\ge \langle \varrho_j,h'\dot\ell_{\theta_0}\rangle_{L^2_{\mathsf Q_{j,\theta_0}}}$ by $\langle\tilde\varrho_j,g\rangle_{L^2_{\mathsf Q_{j,\theta_0}}}\ge \langle\varrho_j,g\rangle_{L^2_{\mathsf Q_{j,\theta_0}}}$ for any $g\in\mathcal G_{j,\theta_0}$ due to $\tilde\varrho_j$ being the projection of $\varrho_j$ to $cl(\mathcal G_{j,\theta_0}).$ This establishes the claim of the theorem.
proof[\rm Proof of Theorem (ref)] Let $j\in\mathbb J$. Consider $h\in\mathcal T_j(\theta_0,\xi).$ By Assumption (ref), for any $\tau\in (0,\bar \tau]$, the LFP $(Q_{0,j,\tau h},Q_{1,j,\tau h})\in \mathcal P_{\theta_0}\times\mathcal P_{\theta_0+\xi+\tau h}$ satisfies \begin{align} Q_{0,j,\tau h}=\mathsf Q_{j,0}, and Q_{1,j,\tau h}=\mathsf Q_{j,\tau h}, \end{align} where $\vartheta\mapsto \mathsf Q_{j,\vartheta}$ is $L^2$ differentiable tangentially to $\mathcal T_j(\theta_0,\xi)$. The rest of the argument then parallels that of Proof of Theorem (ref).

Proof of Theorems in Section (ref)

In what follows, we repeatedly use the fact that, for any nonnegative measurable function $g$ on $S$, belief function $\nu$ and its conjugate $\nu^*$, one has

align[align omitted — 115 chars of source]

where $M_\nu$ is the probability measure on $\mathcal K(S)$ associated with $\nu$ (see Lemma (ref)).

proof[\rm Proof of Lemma (ref)] We first start with showing (ref) and (ref). For this, observe that \begin{align} R(\theta,\phi)&=\max_{P\in \mathcal P_\theta}\int \phi(s)I_{\Theta_0}(\theta)+\zeta(1-\phi(s)) I_{\Theta_1}(\theta)dP(s)\notag\\ &=\max_{P\in \mathcal P_\theta}( I_{\Theta_0}(\theta)-\zeta I_{\Theta_1}(\theta))\int\phi(s)dP(s)+\zeta I_{\Theta_1}(\theta)\notag\\ &= I_{\Theta_0}(\theta)\max_{P\in \mathcal P_\theta}\int\phi(s)dP(s) -\zeta I_{\Theta_1}(\theta)\min_{P\in \mathcal P_\theta}\int\phi(s)dP(s)+\zeta I_{\Theta_1}(\theta)\notag\\ &=\int \phi(s)d\nu^*_\theta(s)I_{\Theta_0}(\theta)+\zeta(1-\int\phi(s)d\nu_\theta(s))I_{\Theta_1}(\theta), \end{align} where the third equality follows from the fact that $ I_{\Theta_0}(\theta)- \zeta I_{\Theta_1}(\theta)>0$ if and only if $I_{\Theta_0}(\theta)=1$ (and $ I_{\Theta_0}(\theta)-\zeta I_{\Theta_1}(\theta)\le 0$ if and only if $I_{\Theta_1}(\theta)=1$). The last equality follows from $core(\nu_\theta)=\mathcal P_\theta$ by Theorem 3 in philippe1999decision and the fact that, for any nonnegative bounded function $g$ on $S$, $\int g d\nu\le \int gdP\le \int gd\nu^*$ for all $P\in core(\nu)$. Using this, write \begin{align} r(\mu,\phi)&\equiv \int_\Theta R(\theta,\phi)d\mu(\theta)\notag\\ &=\tau\int\int \phi(s)I_{\Theta_0}(\theta)d\nu^*_\theta(s)d\mu_0(\theta)+\zeta(1-\tau)(1-\int\int\phi(s)I_{\Theta_1}(\theta)d\nu_\theta(s)d\mu_1(\theta)). \end{align} By Lemma (ref), for each $\nu_\theta$, there is a unique Borel probability measure $M_\theta$ on $\mathcal K(S)$ such that \begin{align} \nu_\theta(A)=M_\theta(K\subset A), \forall A\subset S. \end{align} Let $q$ be a $\sigma$-finite measure on $\mathcal K(S)$ such that $M_\theta\ll q$. Let $M_{\kappa_1}$ be a Borel probability measure on $\mathcal K(S)$ such that $dM_{\kappa_1}/dq=\int_{\Theta_1}\frac{dM_\theta}{dq}d\mu_1(\theta)$. For any $A\subset S$, it then follows that \begin{align} \int_{\Theta_1}\nu_\theta(A)d\mu_1(\theta)&=\int_{\Theta_1}M_\theta(K\subset A)d\mu_1(\theta)\notag\\ &= \int_{\Theta_1}\int_{\mathcal K(S)} 1\{K\subset A\}dM_\theta d\mu_1(\theta)\notag\\ &=\int_{\mathcal K(S)} 1\{K\subset A\}\int_{\Theta_1}\frac{dM_\theta}{dq}d\mu_1(\theta)dq(K)\notag\\ &=\int_{\mathcal K(S)}1\{K\subset A\}dM_{\kappa_1}(K)\notag\\ &=\int_{S}1\{s\in A\}d\kappa_1(s)\notag\\ &=\kappa_1(A), \end{align} where the third equality follows from Fubini's theorem. The existence and uniqueness of $\kappa_1$ follows again from Choquet's theorem (Lemma (ref)). Note that by $\phi\ge 0$ and the definition of the Choquet integral, one can show by the same argument \begin{align} \int\int \phi(s)I_{\Theta_1}(\theta)d\nu_\theta(s)d\mu_1(\theta)&=\int_{\Theta_1}\int \inf_{s\in K}\phi(s) dM_{\theta}(K)d\mu_1(\theta)\notag\\ &=\int \inf_{s\in K}\phi(s)\int_{\Theta_1}\frac{dM_\theta}{dq} d\mu_1(\theta)dq(K)\notag\\ &=\int \phi(s)d\kappa_1(s), \end{align} where the second equality follows from Fubini's theorem. Similarly, it follows that \begin{align} \int\int \phi(s)I_{\Theta_0}(\theta)d\nu^*_\theta(s)d\mu_0(\theta)=\int\phi(s)d\kappa_0^*(s). \end{align} By (ref), (ref), and (ref), we have \begin{align} r(\mu,\phi)=\tau\int\phi(s)d\kappa_0^*(s)+\zeta(1-\tau)(1-\int\phi(s)d\kappa_1(s)). \end{align} Therefore, (ref) holds. Minimizing the BDS risk is then equivalent to minimizing \begin{align} \tilde r_t(\mu,\phi)=t\int\phi(s)d\kappa_0^*(s)-\int\phi(s)d\kappa_1(s), \end{align} where $t=\tau/\zeta(1-\tau)>0$. Let $A\equiv\{s:\phi(s)>0\}.$ Minimizing the risk function above with respect to $\phi$ is then equivalent to minimizing the 2-alternating function $w_t(A)\equiv t\kappa_0^*(A)-\kappa_1(A)$ with respect to $A\subset S$. By Lemmas 3.1 and 3.2 in HS, for each $t\in[0,\infty]$, there exists a set $A_t\subset S$ such that \begin{align} w_t(A_t)=\inf_{A\subset S}w_t(A), \end{align} and $\{A_t,t\ge 0\}$ forms an increasing family of sets. Now define $\Lambda(s)\equiv\inf\{t|s\in A_t\}.$ By Theorem 4.1 in HS, the conclusion of the theorem then follows.
proof[\rm Proof of Theorem (ref)] Note that $\mu_1$ is fixed throughout and $\mathcal M=\{\mu:\mu=\tau\mu_0+(1-\tau)\mu_1,\mu_0\in\Delta(\Theta_0),\tau\in[0,1]\}$. In what follows, we therefore redefine $R$ in (ref) as \begin{align} R(\theta,\phi)&=\int \phi(s)d\nu^*_\theta(s)I_{\Theta_0}(\theta)+\zeta(1-\int\phi(s)d\kappa_1(s))I_{\Theta_1}(\theta)\notag\\ &=R_0(\theta,\phi)I_{\Theta_0}(\theta)+R_1(\phi)I_{\Theta_1}(\theta), \end{align} where $\kappa_1=\int_{\Theta_1}\nu_\theta d\mu_1(\theta).$ First, we show $\sup_{\mu\in \mathcal{M}}\inf_{\phi\in\mathbf\Phi}\int R(\theta,\phi)d\mu\le \inf_{\phi\in\mathbf\Phi}\sup_{\theta\in\Theta}R(\theta,\phi)$. This follows because for any $(\theta,\phi)$, one has $\inf_{\phi'}R(\theta,\phi') \le R(\theta,\phi)\le \sup_{\theta'}R(\theta',\phi),$ and hence \begin{align*} \sup_{\theta}\inf_{\phi'}R(\theta,\phi') \le \inf_{\phi} \sup_{\theta'}R(\theta',\phi). \end{align*} Note that $\sup_{\theta}\inf_{\phi'}R(\theta,\phi')\ge \inf_{\phi\in\mathbf\Phi}\int R(\theta,\phi)d\mu$ for any $\mu$, and hence, the first claim follows. The other direction follows from Lemma (ref). To see this, let \begin{align} \beta\equiv\sup_{\mu\in \mathcal{M}}\inf_{\phi\in\mathbf\Phi} \int R(\theta,\phi)d\mu. \end{align} If $\beta=\infty$, the result is trivial. If $\beta<\infty$, set $f(\theta)=\beta$ for all $\theta\in\Theta.$ By construction, $f(\theta)\ge \inf_{\phi\in\mathbf\Phi} \int R(\theta,\phi)d\mu$ for every $\mu.$ By Lemma (ref), this is equivalent to the existence of $\phi^\dagger\in\mathbf\Phi$ such that $\beta\ge R(\theta,\phi^\dagger),~\forall\theta\in\Theta.$ This implies \begin{align} \beta \ge \sup_{\theta\in\Theta}R(\theta,\phi^\dagger) \ge \inf_{\phi\in\mathbf\Phi}\sup_{\theta\in\Theta}R(\theta,\phi). \end{align} Finally, observe that by (ref), $\sup_{\theta\in\Theta}R(\theta,\phi)=\sup_{\theta\in\Theta_0}R_0(\theta,\phi)\vee R_1(\phi).$ This completes the proof.

Auxiliary Lemmas

Below, we identify each decision function (randomized test) $\phi$ with a Markov kernel $\phi:S\times \mathcal B_{\{0,1\}}\to [0,1]$ and let $\mathbf\Phi$ be the set of all decision functions. We then equip $\mathbf\Phi$ with the weak topology Hausler:2015aa. The following lemma is an analog of Lemma 46.1 in Strasser:1985aa for the BDS risk. We state it as a lemma because the BDS risk $R$ is defined through Choquet integrals with respect to capacities (instead of measures) and hence Lemma 46.1 in Strasser:1985aa is not directly applicable.\footnote{They also show their results to generalized decision functions, which we do not pursue here.}

lemmaSuppose $S$ is finite and $\Theta$ is compact. Let $R$ be defined as in (ref). For every $f:\Theta\to\mathbb R$, the following assertions are equivalent. (i) There exists $\phi\in\mathbf\Phi$ such that $f(\theta)\ge R(\theta,\phi)$ for every $\theta\in\Theta$. (ii) $\int fd\mu\ge \inf_{\phi\in \mathbf\Phi}\int R(\theta,\phi)d\mu(\theta)$ for every $\mu\in \Delta(\Theta).$
proofThe implication $(i)\Rightarrow (ii)$ is obvious. We therefore prove the other implication. Consider the following sets of functions: \begin{align*} M_1=\{f\}, M_2=\{h\in\mathcal C(\Theta):h(\cdot)=R(\cdot,\phi),\phi\in \mathbf\Phi\}. \end{align*} We mimic the proof of Lemma 46.1 in Strasser:1985aa while replacing $M_2$ with the set above. For this, let $M\subseteq\mathcal C(\Theta)$ be an arbitrary set. For any $m\in \Delta(\Theta)$, the lower envelope of $M$ is defined as \begin{align} \psi_M(m)\equiv\inf\{\int fdm,f\in M\}. \end{align} Also define $\alpha(M)\equiv \bigcup_{f\in M}\{g\in\mathcal C(\Theta):f\le g\}.$ This is the set of continuous functions that dominate some function in $M$. Since $M_2$ is compact by Lemma (ref), $\alpha(M_2)$ is closed and hence coincides with its closure $\overline{\alpha (M_2)}$ Strasser:1985aa. By (ii), $\psi_{M_2}(m)\le \psi_{M_1}(m)$ for all $m\in \Delta(\Theta)$. By Lemma (ref), $M_2$ is subconvex. Then, by Theorem 45.6 in Strasser:1985aa, for every $f\in M_1$, there is $g\in\alpha(M_2)=\overline{\alpha (M_2)}$ such that $g\le f$. By the construction of $\alpha(M_2)$, this means there exists $\phi\in \mathbf\Phi$ such that \begin{align} R(\cdot,\phi)\le g(\cdot)\le f(\cdot). \end{align} This completes the proof.

Below, a set $M$ is said to be {\em subconvex\/}, if for any $\alpha\in (0,1)$ and $h_1,h_2\in M,$ there exists $h_3\in M$ such that $h_3\le \alpha h_1+(1-\alpha)h_2$.

lemma$M_2$ is subconvex.
proofLet $h_1,h_2\in M_2$. Then, there exist $\phi_1,\phi_2\in\mathbf\Phi$ such that $h_j(\cdot)=R(\cdot,\phi_j),j=1,2$. Therefore, for any $\alpha\in(0,1)$, \begin{align} \alpha h_1(\theta)+(1-\alpha)h_2(\theta) =&\alpha R(\theta,\phi_1)+(1-\alpha)R(\theta,\phi_2)\notag\\ =&\zeta\big(\alpha\int \phi_1(s)d\nu_{\theta}^*(s)+(1-\alpha)\int \phi_2(s)d\nu_{\theta}^*(s)\big)I_{\Theta_0}(\theta)\notag\\ &+\big(\alpha(1-\int\phi_1(s)d\nu_{\theta}(s))+(1-\alpha)(1-\int\phi_2(s)d\nu_{\theta}(s))\big)I_{\Theta_1}(\theta). \end{align} Note that $\nu_{\theta}^*$ is 2-alternating. Therefore, the Choquet integral with respect to $\nu_{\theta}^*$ is subadditive. The Choquet integral is also positively homogeneous. Therefore, \begin{align} \alpha\int\phi_1(s)d\nu_{\theta}^*(s)+(1-\alpha)\int \phi_2(s)d\nu_{\theta}^*(s)&=\int \alpha\phi_1(s)d\nu_{\theta}^*(s)+\int (1-\alpha)\phi_2(s)d\nu_{\theta}^*(s)\notag\\ &\ge \int \alpha\phi_1(s)+(1-\alpha)\phi_2(s)d\nu_{\theta}^*(s). \end{align} Similarly, by the 2-monotonicity of $\nu_{\theta}$, the Choquet integral with respect to it is superadditive and positively homogeneous. Therefore, \begin{align} \alpha(1-\int\phi_1(s)d\nu_{\theta}(s))+(1-\alpha)(1-\int\phi_2(s)d\nu_{\theta}(s))\ge 1-\int \alpha\phi_1(s)+(1-\alpha)\phi_2(s)d\nu_{\theta}(s). \end{align} Combining (ref)-(ref), we obtain $\alpha h_1(\cdot)+(1-\alpha)h_2(\cdot)\ge h_3(\cdot),$ where $h_3(\cdot)=R(\cdot, \phi_3)$ with $\phi_3=\alpha\phi_1+(1-\alpha)\phi_2$. Hence, $h_3\in M_2$. Conclude that $M_2$ is subconvex.
lemmaSuppose that $S$ is finite and $\Theta$ is compact. Then, $M_2$ is weakly compact.
proofEquip $M_2$ with the weak topology. Let $\varrho$ be the counting measure. We let $\mathbf\Phi=K^1(\varrho)$ denote the set of Markov kernels equipped with the weak topology. It is the coarsest topology that makes any functional of the following form continuous: \begin{align} T(\phi)= \int_S\int_{\{0,1\}}h(a)\phi(s,da)f(s)d\varrho(s), f\in L^1_\varrho(S),h\in\mathcal C_b(\{0,1\}). \end{align} By Theorem 2.7 in Hausler:2015aa, $\mathbf\Phi$ is compact if the set $\varrho\mathbf\Phi\equiv\{ \varrho\phi:\varrho\phi(\cdot)=\int_S \phi(s,\cdot)d\varrho(s),\phi\in \mathbf\Phi\}$ is relatively compact in $\Delta(\{0,1\}).$ Note that $\{0,1\}$ is compact. Hence, $\varrho\mathbf\Phi$ is uniformly tight, which implies that $\varrho\mathbf\Phi$ is relatively compact by Prohorov's theorem vandervaart:wellner:1996. This ensures the compactness of $\mathbf\Phi$. Below, let $\mathcal K(S)$ the set of all nonempty (and necessarily closed) subsets of $S$. By Choquet's theorem, a belief function $\nu_\theta$ can be expressed by its canonical representation $(\mathcal K(S),K,\hat m_\theta)$, where $K$ is a random set following a probability measure $\hat m_\theta$ on $\mathcal K(S)$ such that $\nu_\theta(A)=\hat m_\theta(K\subset A)$ for all $A\in\mathcal K(S)$. Below, we adopt this canonical representation and also denote the measure on $\mathcal K(S)$ by $m_\theta$ rather than $\hat m_\theta.$ We also note that we write $\phi(s)=\int_{\{0,1\}}\phi(s,da)$ in what follows. Define $g:\mathbf\Phi\to\mathcal C(\Theta)$ pointwise by $\phi\mapsto R(\cdot,\phi)$. We argue that this map is continuous, where we equip $\mathbf\Phi$ and $\mathcal C(\Theta)$ with weak topologies. Let $\phi_n,n=1,2,\cdots$ be a sequence such that $\phi_n\to\phi\in \mathbf\Phi$ weakly. \begin{align} \int_S\int_{\{0,1\}}\phi_n(s,da)d\nu^*_\theta(s)= \int_S \phi_n(s)d\nu^*_\theta(s)=\int_S \max_{s\in K}\phi_n(s)dm_\theta(K) \end{align} Since $S$ is finite and $1\{s=s'\}\in L^1_\varrho(S)$ for any $s'\in S$, $\phi_n$ converging weakly implies \begin{align} \phi_n(s')=\sum_{s\in S}\phi_n(s)1\{s=s'\}\to \sum_{s\in S}\phi(s)1\{s=s'\}=\phi(s'), \forall s'\in S. \end{align} Therefore, $\phi_n$ converges pointwise to $\phi$. Fix $K\in\mathcal K(S)$. Note that $(s,n)\mapsto \phi_n(s)$ is continuous with respect to the discrete topology. By Berge's maximum theorem, $ \max_{s\in K}\phi_n(s)\to \max_{s\in K}\phi(s)$. Hence, \begin{align} \lim_{n\to\infty}\int_S\int_{\{0,1\}}\phi_n(s,da)d\nu^*_\theta(s)&=\lim_{n\to\infty}\int_S \max_{s\in K}\phi_n(s)dm_\theta(K)\notag\\ &=\int_S\lim_{n\to\infty}\max_{s\in K}\phi_n(s)dm_\theta(K)\notag\\ &=\int_S\max_{s\in K}\phi(s)dm_\theta(K)\notag\\ &=\int_S\phi (s)d\nu^*_\theta(s)\notag\\ &=\int_S\int_{\{0,1\}}\phi (s,a)d\nu^*_\theta(s). \end{align} where the second equality follows from the convergence of $\max_{s\in K}\phi_n(s)$ and the dominated convergence theorem. Consider any finite Borel measure $\mu$ on $\Theta$. The result above and the dominated convergence theorem imply \begin{align} \lim_{n\to\infty}\int_\Theta \int_S\int_{\{0,1\}}\phi_n(s,da)d\nu^*_\theta(s)I_{\Theta_0}(\theta)d\mu(\theta)&=\int_\Theta\lim_{n\to\infty}\int_S\int_{\{0,1\}}\phi_n(s,da)d\nu^*_\theta(s)I_{\Theta_0}(\theta)d\mu(\theta)\\ &=\int_\Theta\int_S\int_{\{0,1\}}\phi(s,da)d\nu^*_\theta I_{\Theta_0}(\theta)d\mu(\theta). \end{align} By a similar argument, one can also show \begin{align} \lim_{n\to\infty}\int_\Theta \int_S\int_{\{0,1\}}\phi_n(s,da)d\nu_\theta(s)I_{\Theta_1}(\theta)d\mu(\theta)&=\int_\Theta\lim_{n\to\infty}\int_S\int_{\{0,1\}}\phi_n(s,da)d\nu_\theta(s)I_{\Theta_1}(\theta)d\mu(\theta)\\ &=\int_\Theta\int_S\int_{\{0,1\}}\phi(s,da)d\nu_\theta I_{\Theta_1}(\theta)d\mu(\theta). \end{align} Note that $\Theta$ is a compact set in a metric space. Corollary 14.15 in aliprantis2006infinite then ensures that the dual space of $\mathcal C(\Theta)$ is the set of finite Borel measures on $\Theta$. Combining these results and noting that $R(\theta,\phi)=\zeta\int_S\int_{\{0,1\}}\phi(s,da)d\nu^*_\theta I_{\Theta_0}(\theta)+(1-\int_S\int_{\{0,1\}}\phi(s,da)d\nu_\theta)I_{\Theta_1}(\theta)$, it follows that $R(\cdot,\phi_n)\to R(\cdot,\phi)$ in $\mathcal C(\Theta)$ with respect to the weak topology. This establishes that $g$ is continuous. Hence, $M_2$ is the continuous image of a compact set. Conclude that $M_2$ is weakly compact.

Examples

Example 1: Binary response game

In this section, we provide details on Example (ref) including the computation of the belief function, least favorable pair, and minimax test. Recall that $S=\{(0,0),(1,1),(1,0),(0,1)\}$. There exist 14 subsets to be considered (without considering the empty set and $S$). One can then compute the the lower and upper bounds of the probability of each event by mimicking the calculation in (ref). The results are summarized in Table (ref).

table[table omitted — 2,279 chars of source]

LFP

The following proposition characterizes the LFP and minimax tests.

propositionLet $(S,U,\Theta,G)$ be as defined in Example (ref). Suppose that $u$ follows the bivariate standard normal distribution. Let $\alpha\in (0,1/4)$, $\theta_0=(0,0)'$, and $\theta_1<0$. Then, for any $j\in\mathbb J=\{ \textup{\uppercase\expandafter{\romannumeral1}} , \textup{\uppercase\expandafter{\romannumeral2}} , \textup{\uppercase\expandafter{\romannumeral3}} \}$, the density of $Q_0$ is $(q_0(0,0),q_0(1,1),q_0(1,0),q_0(0,1))=(\frac{1}{4},\frac{1}{4}, \frac{1}{4}, \frac{1}{4})$. The densities of $Q_1$ for $\theta_1\in \Theta_{j},j\in\mathbb J$ and minimax tests are as in Table (ref).
proof[\rm Proof of Proposition (ref)] First, observe that the upper and lower probabilities of $A_1,\cdots, A_4$ fully characterize the constraints in the convex program. To see this, observe that, for example, \begin{multline} \nu_\theta(A_5)=m_\theta(G(u|\theta)\subset \{(0,0),(1,1)\})\\ =m_\theta(G(u|\theta)= \{(0,0)\})+m_\theta(G(u|\theta)= \{(1,1)\})=\nu_\theta(A_1)+\nu_\theta(A_2)=\frac{1}{4}+\Phi(\theta^{(1)}) \Phi(\theta^{(2)}), \end{multline} where we note that the additivity of $\nu_\theta$ for this event is due to the form of the correspondence in (ref) and does not hold in general. One can compute $\nu_\theta(A_6)$ and $\nu_\theta(A_7)$ similarly. The upper bounds $\nu^*_\theta(A_j)$ for $j=8,\cdots,14$ can then be computed using the conjugacy of $\nu_\theta$ and $\nu^*_\theta.$ Similarly, the upper bounds $\nu^*_\theta(A_j)$ for $j=1,\cdots,7$ imply the lower bounds $\nu_\theta(A_j)$ for $j=8,\cdots,14$. In sum, it suffices to impose the constraints that arise from the upper and lower probabilities of $A_1,\cdots,A_4$. Further, $q_0(0,0)=q_1(0,0)=1/4$ regardless of the parameter value. This allows to simplify the convex program as \begin{align} \min_{(q_0,q_1)}-\ln&(\frac{q_0(1,1)}{q_0(1,1)+q_1(1,1)})(q_0(1,1)+q_1(1,1)) -\ln(\frac{q_0(1,0)}{q_0(1,0)+q_1(1,0)})(q_0(1,0)+q_1(1,0))\notag\\ &\qquad-\ln(\frac{q_0(0,1)}{q_0(0,1)+q_1(0,1)})(q_0(0,1)+q_1(0,1))\notag\\ s.t. & \frac{1}{4}-\Phi(\theta_1^{ (1)}) \Phi(\theta_1^{(2)})+\frac{\Phi(\theta_1^{(2)})}{2} \leq q_1(0,1) \leq \frac{1}{2}(1-\Phi(\theta_1^{(1)}))\\ & \frac{1}{4}-\Phi(\theta_1^{(1)}) \Phi(\theta_1^{(2)})+\frac{\Phi(\theta_1^{(1)})}{2} \leq q_1(1,0) \leq \frac{1}{2}(1-\Phi(\theta_1^{(2)}))\\ & q_1(1,1)=\Phi(\theta_1^{(1)}) \Phi(\theta_1^{(2)})\\ & q_1(1,1)+q_1(1,0)+q_1(0,1)=\frac{3}{4}\\ & q_0(1,1)=q_0(1,0)=q_0(0,1)=\frac{1}{4}. \end{align} Note that (ref)-(ref) imply that the values of $q_0(1,1),q_0(1,0),q_0(0,1)$, and $q_1(1,1)$ are determined uniquely. Hence, it remains to optimize the problem with respect to $q_1(1,0)$ and $q_1(0,1)$. For this, let $y=q_1(1,0)$. Then, one may write $q_1(0,1)=3/4-\Phi(\theta_1^{(1)}) \Phi(\theta_1^{(2)})-y$ due to (ref) and (ref). Hence, the problem reduces to an optimization problem with a single control variable. Using this, define the Lagrangian by \begin{multline} \mathcal L(y,\lambda)\equiv-\ln \bigg(\frac{1/4}{1/4+y}\bigg)(\frac{1}{4}+y)-\ln \bigg(\frac{1/4}{\frac{1}{4}+\frac{3}{4}-\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})-y}\bigg)(\frac{1}{4}+\frac{3}{4}-\Phi(\theta_1^{(1)}) \Phi(\theta_1^{(2)})-y)\\- \lambda_1 \Big(\frac{1}{2}-\frac{\Phi(\theta_1^{(2)})}{2}-y\Big)- \lambda_2 \Big(y-\frac{1}{4}-\frac{\Phi(\theta_1^{(1)})}{2}+\Phi(\theta_1^{(1)}) \Phi(\theta_1^{(2)})\Big). \end{multline} By Theorem 28.3 in Rockafellar:1972aa, the saddle point of the Lagrangian characterizes the optimal solution of the original problem. The Karush-Kuhn-Tucker (KKT) conditions are as follows: \begin{align} &1-\ln\bigg(\frac{1/4}{1/4+y} \bigg) -1+\ln \bigg(\frac{1/4}{1-\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})-y} \bigg)+\lambda_1-\lambda_2=0 \\ &\lambda_1 \Big(\frac{1}{2}-\frac{\Phi(\theta_1^{(2)})}{2}-y\Big)=0\\ &\lambda_2 \Big(y-\frac{1}{4}-\frac{\Phi(\theta_1^{(1)})}{2}+\Phi(\theta_1^{(1)}) \Phi(\theta_1^{(2)})\Big)=0\\ &\lambda_1,\lambda_2\ge 0. \end{align} Below, we consider three cases based on the value of the Lagrange multipliers. Case 1 ($\lambda_1=\lambda_2=0$): Suppose that $\lambda_1=0$ and $\lambda_2=0$. Then, the solution from (ref) is \begin{equation} y=\frac{3}{8}-\frac{\Phi(\theta_1^{(1)}) \Phi(\theta_1^{(2)})}{2}. \end{equation} Substituting this into the complementary slackness conditions (ref) and (ref) yields \begin{align} \frac{1}{2}-\frac{\Phi(\theta_1^{(2)})}{2}-\frac{3}{8}+\frac{\Phi(\theta_1^{(1)}) \Phi(\theta_1^{(2)})}{2} \ge0, \frac{3}{8}-\frac{\Phi(\theta_1^{(1)}) \Phi(\theta_1^{(2)})}{2}-\frac{1}{4}-\frac{\Phi(\theta_1^{(1)})}{2}+\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})\ge 0, \end{align} which can be simplified as \begin{equation} \Phi(\theta_1^{(2)})(1-\Phi(\theta_1^{(1)})) \le \frac{1}{4},\qquad \Phi(\theta_1^{(1)})(1-\Phi(\theta_1^{(2)})) \le \frac{1}{4}. \end{equation} By $y=q_1(1,0)$, $q_1(0,1)=3/4-\Phi(\theta_1^{(1)}) \Phi(\theta_1^{(2)})-y$, and (ref), the LFP is \begin{align} (q_0(0,0),q_0(1,1),q_0(1,0),q_0(0,1))&=\Big(\frac{1}{4},\frac{1}{4}, \frac{1}{4}, \frac{1}{4}\Big)\\ (q_1(0,0),q_1(1,1),q_1(1,0),q_1(0,1))&=\Big(\frac{1}{4},\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)}), \frac{3-4\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})}{8},\frac{3-4\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})}{8}\Big). \end{align} Hence, the likelihood-ratio statistic is given by \begin{align} \Lambda(s)&= \begin{cases} 1& s=(0,0)\\ 4\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})& s=(1,1)\\ \frac{3-4\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})}{2}&s=(1,0)\\ \frac{3-4\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})}{2}&s=(0,1). \end{cases} \end{align} Under $Q_0$, $\Lambda(s)$ is supported on $\{4\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)}),1,(3-4\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)}))/2\}$ with probabilities $(1/4,1/4,1/2)$. For $\theta^{(j)}<0,j=1,2$, one has $4\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})<1<(3-4\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)}))/2.$ The largest value of the support is therefore $(3-4\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)}))/2$. Setting $C$ to this value and solving \begin{align} \alpha=E_{Q_0}[\phi(s)]=\gamma Q_0(\pi(s)\ge C)=\gamma Q_0\big(s=(1,0) \cup s=(0,1)\big)=\frac{\gamma}{2}, \end{align} one obtains $\gamma=2\alpha.$ This gives the level-$\alpha$ minimax test. Case 2 ($\lambda_1=0$ and $\lambda_2>0$): Suppose that $\lambda_1=0$ and $\lambda_2>0$. By the complementary slackness condition (ref), the solution is obtained at the lower bound $y=\frac{1}{4}-\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})+\frac{\Phi(\theta_1^{(1)})}{2}$. Substituting this into (ref) and noting that $\lambda_1=0$, we have \begin{equation} \lambda_2=\ln \bigg(\frac{\frac{1}{2}-\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})+\frac{\Phi(\theta_1^{(1)})}{2}}{\frac{3}{4}-\frac{\Phi(\theta_1^{(1)})}{2}} \bigg). \end{equation} The difference between the numerator and denominator in the logarithm above is \begin{equation} \frac{1}{2}-\Phi(\theta_1^{(1)}) \Phi(\theta_1^{(2)})+\frac{\Phi(\theta_1^{(1)})}{2}-(\frac{3}{4}-\frac{\Phi(\theta_1^{(1)})}{2})=\Phi(\theta_1^{(1)})(1-\Phi(\theta_1^{(2)})-\frac{1}{4}. \end{equation} Therefore, $\lambda_2 > 0$ if and only if $\Phi(\theta_1^{(1)})(1-\Phi(\theta_1^{(2)}) > 1/4.$ Similarly, the complementary slackness condition (ref) is satisfied with $\lambda_1=0$ and $y=\frac{1}{4}-\Phi(\theta_1^{(1)}) \Phi(\theta_1^{(2)})+\frac{\Phi(\theta_1^{(1)})}{2}$. Note that the constraint $\frac{1}{2}-\frac{\Phi(\theta_1^{(2)})}{2}-y\ge 0$ is trivially satisfied because \begin{align} \frac{1}{2}-\frac{\Phi(\theta_1^{(2)})}{2}-y&= \frac{1}{2}-\frac{\Phi(\theta_1^{(2)})}{2}-\bigg(\frac{1}{4}-\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)})+\frac{\Phi(\theta_1^{(1)})}{2} \bigg) \\ &=(\frac{1}{2}-\Phi(\theta_1^{(1)}))(\frac{1}{2}-\Phi(\theta_1^{(2)}))\ge0, \end{align} where the last inequality follows from $\theta_1^{(j)}\le 0$ for $j=1,2.$ Hence, if $\Phi(\theta_1^{(1)})(1-\Phi(\theta_1^{(2)})) > \frac{1}{4}$, the least favorable pair is \begin{align} (q_0(0,0),q_0(1,1),q_0(1,0),q_0(0,1))&=\Big(\frac{1}{4},\frac{1}{4}, \frac{1}{4}, \frac{1}{4}\Big)\\ (q_1(0,0),q_1(1,1),q_1(1,0),q_1(0,1))&=\Big(\frac{1}{4},\Phi(\theta_1^{(1)})\Phi(\theta_1^{(2)}), \frac{1}{4}-\Phi(\theta_1^{(1)}) (\Phi(\theta_1^{(2)})-\frac{1}{2}),\frac{1}{2}(1-\Phi(\theta_1^{(1)}))\Big), \end{align} where we used $y=q_1(1,0)$, $q_1(0,1)=3/4-\Phi(\theta_1^{(1)}) \Phi(\theta_1^{(2)})-y$, and $y=\frac{1}{4}-\Phi(\theta_1^{(1)}) \Phi(\theta_1^{(2)})+\frac{\Phi(\theta_1^{(1)})}{2}$. The rest of the analysis is similar to Case 1. Case 3 ($\lambda_1>0$ and $\lambda_2=0$): The argument is similar to the one for Case 2. Hence, we omit the proof.

Efficient influence function and optimal tests

propositionSuppose that the conditions of Proposition (ref) hold. Let $p=(p^{(1)},p^{(2)})'$ with $p^{(j)}< 0$ for $j=1,2$ and let $\varphi(\theta)=p'\theta$. Then, Assumption (ref) holds.
proof[\rm Proof of Proposition (ref)] (i) Observe that $\{\xi\in \Theta-\theta_0:\varphi(\theta_0+\xi)>0\}=\{\xi\in (-\infty,0]^2:\xi^{(1)}<0,\text{ or }\xi^{(2)}<0\}$, which is indeed a convex cone; (ii) Let $\mathbb J=\{ \textup{\uppercase\expandafter{\romannumeral1}} , \textup{\uppercase\expandafter{\romannumeral2}} , \textup{\uppercase\expandafter{\romannumeral3}} \}$ and let $\{\mathcal T_{j}(\theta_0),j\in \mathbb J\}$ be defined as in (ref)-(ref). These cones satisfy the requirements in Assumption (ref) (ii); (iii) Based on Table (ref), let $\mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta}$ be a model with the following density \begin{multline} (\mathsf q_{ \uppercase\expandafter{\romannumeral1} ,\theta}(0,0),\mathsf q_{ \uppercase\expandafter{\romannumeral1} ,\theta}(1,1),\mathsf q_{ \uppercase\expandafter{\romannumeral1} ,\theta}(1,0),\mathsf q_{ \uppercase\expandafter{\romannumeral1} ,\theta}(0,1))\\ =\Big(\frac{1}{4},\Phi(\theta^{(1)})\Phi(\theta^{(2)}), \frac{3-4\Phi(\theta^{(1)})\Phi(\theta^{(2)})}{8},\frac{3-4\Phi(\theta^{(1)})\Phi(\theta^{(2)})}{8}\Big). \end{multline} Similarly, for $j= \textup{\uppercase\expandafter{\romannumeral2}} $ and $ \textup{\uppercase\expandafter{\romannumeral3}} $, let $\mathsf q_{j}$ be defined accordingly based on Table (ref). By Proposition (ref), for each $j\in\mathbb J$, the LFP $(Q_{0,j,\tau h},Q_{1,j,\tau h})\in \mathcal P_{\theta_0}\times\mathcal P_{\theta_0+\tau h}$ satisfies \begin{align} Q_{0,j,\tau h}=\mathsf Q_{j,\theta_0}, and Q_{1,j,\tau h}=\mathsf Q_{j,\theta_0+\tau h}, \end{align} for all $\tau \in(0,\bar \tau]$ for some $\bar \tau>0$. Furthermore, $\theta\mapsto \mathsf Q_{j,\theta}$ is $L^2$ differentiable at $\theta_0$ tangentially to $\mathcal T_{j}(\theta_0)$ by Proposition (ref).
propositionSuppose that the conditions of Proposition (ref) hold. Let $p=(p^{(1)},p^{(2)})'$ with $p^{(j)}< 0$ for $j=1,2$ and let $\varphi(\theta)=p'\theta$. Then, for each $j\in\{ \textup{\uppercase\expandafter{\romannumeral1}} , \textup{\uppercase\expandafter{\romannumeral2}} , \textup{\uppercase\expandafter{\romannumeral3}} \}$, the model $\theta\mapsto \mathsf Q_{j,\theta}$ is $L^2$ differentiable at $\theta_0$ tangentially to $\mathcal T_{j}(\theta_0)$. Furthermore, the efficient influence functions are \begin{align} \tilde\varrho_{ \uppercase\expandafter{\romannumeral1} }(s)&= \frac{2\sqrt{2\pi}}{3}(p^{(1)}+p^{(2)})1\{s=(1,1)\}-\frac{\sqrt{2\pi}}{3}(p^{(1)}+p^{(2)})(1\{s=(1,0)\}+1\{s=(0,1)\}\\ \tilde\varrho_{ \uppercase\expandafter{\romannumeral2} }(s)&=(b^{(1)}_{ \uppercase\expandafter{\romannumeral2} }(p)+b^{(2)}_{ \uppercase\expandafter{\romannumeral2} }(p))1\{s=(1,1)\}-b^{(2)}_{ \uppercase\expandafter{\romannumeral2} }(p)1\{s=(1,0)\}-b^{(1)}_{ \uppercase\expandafter{\romannumeral2} }(p)1\{s=(0,1)\}\\ \tilde\varrho_{ \textup{\uppercase\expandafter{\romannumeral3}} }(s)&=(b^{(1)}_{ \textup{\uppercase\expandafter{\romannumeral3}} }(p)+b^{(2)}_{ \textup{\uppercase\expandafter{\romannumeral3}} }(p))1\{s=(1,1)\}-b^{(1)}_{ \textup{\uppercase\expandafter{\romannumeral3}} }(p)1\{s=(1,0)\}-b^{(2)}_{ \textup{\uppercase\expandafter{\romannumeral3}} }(p)1\{s=(0,1)\}, \end{align} where \begin{align} b_{ \textup{\uppercase\expandafter{\romannumeral2}} }(p)&=\operatorname*{arg\,min}_{b^{(2)}\le b^{(1)}\le 0}\frac{1}{4}[(\sqrt{2\pi}p-b)'1]^2+\frac{1}{4}(\sqrt{2\pi}p^{(2)}-b^{(2)})^2+\frac{1}{4}(\sqrt{2\pi}p^{(1)}-b^{(1)})^2\\ b_{ \textup{\uppercase\expandafter{\romannumeral3}} }(p)&=\operatorname*{arg\,min}_{b^{(1)}\le b^{(2)}\le 0}\frac{1}{4}[(\sqrt{2\pi}p-b)'1]^2+\frac{1}{4}(\sqrt{2\pi}p^{(1)}-b^{(1)})^2+\frac{1}{4}(\sqrt{2\pi}p^{(2)}-b^{(2)})^2. \end{align}
proof[\rm Proof of Proposition (ref)] Case $ \textup{\uppercase\expandafter{\romannumeral1}} $: First, we consider the case in which $h\in\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} }(\theta_0).$ Let \begin{align} \dot\ell_{ \uppercase\expandafter{\romannumeral1} ,\theta_0}&\equiv 1\{s_i=(1,1)\}\begin{pmatrix} \frac{2}{\sqrt{2\pi}}\\ \frac{2}{\sqrt{2\pi}} \end{pmatrix} +(1\{s_i=(1,0)\}+1\{s_i=(0,1)\}) \begin{pmatrix} \frac{-1}{\sqrt{2\pi}}\\ \frac{-1}{\sqrt{2\pi}} \end{pmatrix}. \end{align} Let $\mathsf q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta}$ denote the density of $\mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta}.$ Then, by Proposition (ref) (Case I), for $h=(\bar h,\bar h)'\in \mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} }(\theta_0),$ \begin{align} &\mathsf q_{ \uppercase\expandafter{\romannumeral1} ,\theta_0+\tau h}^{1/2}-\mathsf q_{ \uppercase\expandafter{\romannumeral1} ,\theta_0}^{1/2}\notag\\ &\quad =\Big(\frac{1}{2}1\{s=(0,0)\}+\Phi(\tau\bar h)1\{s=(1,1)\}+\Big(\frac{3-4\Phi(\tau\bar h)^2}{8}\Big)^{\frac{1}{2}}(1\{s=(1,0)\}+1\{s=(0,1)\})\Big)\notag\\ &\quad\quad -\frac{1}{2}\Big(1\{s=(0,0)\}+1\{s=(1,1)\}+1\{s=(1,0)\}+1\{s=(0,1)\}\Big)\notag\\ &\quad = (\Phi(\tau\bar h)-1/2)1\{s=(1,1)\}+\Big(\Big(\frac{3-4\Phi(\tau\bar h)^2}{8}\Big)^{\frac{1}{2}}-\frac{1}{2}\Big)\Big(1\{s=(1,0)\}+1\{s=(0,1)\}\Big)\notag\\ &\quad=(\Phi(0)+\Phi'(0)\tau \bar h+o(\tau)-1/2)1\{s=(1,1)\}\notag\&\quad\quad+\Big(\Big(\frac{3-4\Phi(0)^2}{8}\Big)^{\frac{1}{2}} -\frac{1}{2}\Big(\frac{3-4\Phi(0)^2}{8}\Big)^{-\frac{1}{2}}\Phi(0)\Phi'(0)\tau\bar h+o(\tau)-\frac{1}{2}\Big)\Big(1\{s=(1,0)\}+1\{s=(0,1)\}\Big)\notag\\ &\quad=(\frac{1}{\sqrt{2\pi}}\tau\bar h+o(\tau))1\{s=(1,1)\}-(\frac{1}{2\sqrt{2\pi}}\tau\bar h+o(\tau))\Big(1\{s=(1,0)\}+1\{s=(0,1)\}\Big), \end{align} where the third equality follows from taking a Taylor expansion of $\mathsf q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta_0+\tau h}^{1/2}$ with respect to $\tau$ at 0. Note that, by (ref), $h=(\bar h,\bar h)'$, and $\mathsf q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta_0}^{1/2}=\frac{1}{2}(1\{s=(0,0)\}+1\{s=(1,1)\}+1\{s=(1,0)\}+1\{s=(0,1)\})$ by Proposition (ref), it follows that \begin{align} \frac{1}{2}\tau h'\dot\ell_{ \uppercase\expandafter{\romannumeral1} ,\theta_0} \mathsf q_{ \uppercase\expandafter{\romannumeral1} ,\theta_0}^{1/2}=\frac{1}{\sqrt{2\pi}}\tau\bar h1\{s=(1,1)\}-\frac{1}{2\sqrt{2\pi}}\tau\bar h\Big(1\{s=(1,0)\}+1\{s=(0,1)\}\Big) \end{align} By (ref), (ref), and the triangle and Cauchy-Schwarz inequalities, for any given $h\in\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} }(\theta_0)$, \begin{align} \Big\| \mathsf q_{ \uppercase\expandafter{\romannumeral1} ,\theta_0+\tau h}^{1/2}-\mathsf q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta_0}^{1/2}(1+\frac{1}{2}\tau h'\dot\ell_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta_0})\Big\|_{L^2_\mu} =o(\tau). \end{align} This establishes that $\dot\ell_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta_0}$ in (ref) is the $L^2$ derivative for $h\in\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} }(\theta_0)$. Therefore, the tangent cone is \begin{align} \mathcal G_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta_0}&=\Big\{g\in L^2_{\mathsf Q_{\theta_0}}:g=\frac{4\bar h}{\sqrt{2\pi}}1\{s=(1,1)\}-\frac{2\bar h}{\sqrt{2\pi}}(1\{s=(1,0)\}+1\{s=(0,1)\}),\bar h\le 0\Big\}\notag\\ &=\Big\{g\in L^2_{\mathsf Q_{\theta_0}}:g=2 b 1\{s=(1,1)\}-b(1\{s=(1,0)\}+1\{s=(0,1)\}),b\le 0\Big\}. \end{align} Let $p=(p^{(1)},p^{(2)})'$ with $p^{(j)}< 0$ for $j=1,2$. The influence curve $\varrho_{ \textup{\uppercase\expandafter{\romannumeral1}} }$ must satisfy \begin{align} p'h=E_{\mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta_0}}\big[\varrho_{ \textup{\uppercase\expandafter{\romannumeral1}} }h'\ell_{\theta_0}\big]. \end{align} Let $b<0$ and let $\varrho_{ \textup{\uppercase\expandafter{\romannumeral1}} }=2 b 1\{s=(1,1)\}-b(1\{s=(1,0)\}+1\{s=(0,1)\}.$ Then, \begin{align} E_{\mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta_0}}\big[\varrho_{ \textup{\uppercase\expandafter{\romannumeral1}} }h'\ell_{\theta_0}\big]&=\frac{8b\bar h}{\sqrt{2\pi}}E_{\mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta_0}}[1\{s=(1,1)\}]+\frac{2b\bar h}{\sqrt{2\pi}}(E_{\mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta_0}}[1\{s=(1,0)\}]+E_{\mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta_0}}[1\{s=(0,1)\}])\\ &=\frac{2b\bar h}{\sqrt{2\pi}}+\frac{b\bar h}{\sqrt{2\pi}}=\frac{3b\bar h}{\sqrt{2\pi}}. \end{align} Note that $p'h=(p^{(1)}+p^{(2)})\bar h$, and hence setting $b=\frac{\sqrt{2\pi}}{3}(p^{(1)}+p^{(2)})$ gives the following efficient influence function as a projection of the influence curve on the closure of $\mathcal G_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta_0}:$ \begin{align*} \tilde\varrho_{ \textup{\uppercase\expandafter{\romannumeral1}} }=\frac{2\sqrt{2\pi}}{3}(p^{(1)}+p^{(2)})1\{s=(1,1)\}-\frac{\sqrt{2\pi}}{3}(p^{(1)}+p^{(2)})(1\{s=(1,0)\}+1\{s=(0,1)\}. \end{align*} Case $ \textup{\uppercase\expandafter{\romannumeral2}} $: Suppose $h\in\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral2}} }(\theta_0).$ Arguing as in Case $ \textup{\uppercase\expandafter{\romannumeral1}} $, it can be shown that \begin{align} \dot\ell_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\theta_0}(s)=1\{s=(1,1)\}\begin{pmatrix} \frac{2}{\sqrt{2\pi}}\\ \frac{2}{\sqrt{2\pi}} \end{pmatrix} +1\{s=(1,0)\}\begin{pmatrix} 0\\ \frac{-2}{\sqrt{2\pi}} \end{pmatrix}+1\{s=(0,1)\} \begin{pmatrix} \frac{-2}{\sqrt{2\pi}}\\ 0 \end{pmatrix} \end{align} is the $L^2$ derivative when $h\in\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral2}} }(\theta_0).$ Therefore, the tangent cone can be written \begin{multline} \mathcal G_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\theta_0} =\Big\{g\in L^2_{\mathsf Q_{\theta_0}}:g=(b^{(1)}+b^{(2)})1\{s=(1,1)\}-b^{(2)}1\{s=(1,0)\}-b^{(1)}1\{s=(0,1)\},\\-\infty< b^{(2)}<b^{(1)}\le 0\Big\}. \end{multline} Let $\varrho_{ \textup{\uppercase\expandafter{\romannumeral2}} }(s)\equiv(\sqrt{2\pi}p^{(1)}+\sqrt{2\pi}p^{(2)})1\{s=(1,1)\}-\sqrt{2\pi}p^{(2)}1\{s=(1,0)\}-\sqrt{2\pi}p^{(1)}1\{s=(0,1)\}$. It is straightforward to show $\varrho_{ \textup{\uppercase\expandafter{\romannumeral2}} }$ is an influence curve for $\varphi$. The efficient influence function is then given by the projection of $\varrho_{ \textup{\uppercase\expandafter{\romannumeral2}} }$ onto $cl(\mathcal G_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\theta_0})$, which is $\tilde\varrho_{ \textup{\uppercase\expandafter{\romannumeral2}} }=\operatorname*{arg\,min}_{g\in cl(\mathcal G_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\theta_0})}\|\varrho_{ \textup{\uppercase\expandafter{\romannumeral2}} }-g\|^2_{L^2_{\mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\theta_0}}}.$ Note that $b\mapsto b'\dot\ell_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\theta_0}$ is a continuous map from $\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral2}} }(\theta_0)$ to $\mathcal G_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\theta_0}$ by $\dot\ell_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\theta_0}$ being square integrable. Hence, the efficient influence function $\tilde\varrho_{ \textup{\uppercase\expandafter{\romannumeral2}} }$ is as given in (ref) with \begin{align*} b_{ \textup{\uppercase\expandafter{\romannumeral2}} }(p)&=\operatorname*{arg\,min}_{b\in cl(\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral2}} }(\theta_0))}E_{\mathsf Q_{\theta_0}}\big[(\varrho_{ \textup{\uppercase\expandafter{\romannumeral2}} }-((b^{(1)}+b^{(2)})1\{s=(1,1)\}-b^{(2)}1\{s=(1,0)\}-b^{(1)}1\{s=(0,1)\}))^2\big]\\ &=\operatorname*{arg\,min}_{b^{(2)}\le b^{(1)}\le 0}\frac{1}{4}[(\sqrt{2\pi}p-b)'1]^2+\frac{1}{4}(\sqrt{2\pi}p^{(2)}-b^{(2)})^2+\frac{1}{4}(\sqrt{2\pi}p^{(1)}-b^{(1)})^2. \end{align*} The analysis of Case $ \textup{\uppercase\expandafter{\romannumeral3}} $ ($h\in\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral3}} }(\theta_0)$) is similar to the one above. Therefore, we omit its proof.

Example 2: Roy model

In this section, we provide details on Example (ref). Recall that $S=\{(0,0),(0,1),(1,0),(1,1)\}$, and the sharp identifying restrictions are given as (ref)-(ref).

LFP

We start with the following characterization of the LFP for the hypotheses considered in the text. For this, let $\bar c>0$ be a known constant.

propositionLet $(S,U,G,\Theta)$ be defined as in Example (ref). Let $m_\theta$ be a discrete distribution on $S$ whose distribution is uniquely determined by $\theta=(\theta^{(0,0)},\theta^{(0,1)},\theta^{(1,0)})'.$ Suppose (i) $\theta^{(0,0)}_0=\theta^{(0,0)}_1=\bar c>0$ and $\theta^{(1,0)}_0<1-\bar c-\theta^{(0,1)}_0$; and (ii) $\theta^{(1,0)}_1>1-\bar c-\theta^{(0,1)}_0$. Case 1: If $\theta_1^{(1,0)}< 1-\bar c-\theta^{(0,1)}_1$, the densities of the LFP $(Q_0,Q_1)\in\mathcal P_{\theta_0}\times\mathcal P_{\theta_1}$ are: \begin{align} (q_0(0,0),q_0(0,1),q_0(1,0),q_0(1,1))&=(\frac{\bar c}{2},\frac{\bar c}{2},1-\bar c-\theta^{(0,1)}_0,\theta^{(0,1)}_0)\\ (q_1(0,0),q_1(0,1),q_1(1,0),q_1(1,1))&=(\frac{\bar c}{2},\frac{\bar c}{2},\theta^{(1,0)}_1,1-\bar c-\theta^{(1,0)}_1). \end{align} Case 2: If $\theta_1^{(1,0)}= 1-\bar c-\theta^{(0,1)}_1$, the densities of the LFP $(Q_0,Q_1)\in\mathcal P_{\theta_0}\times\mathcal P_{\theta_1}$ are: \begin{align} (q_0(0,0),q_0(0,1),q_0(1,0),q_0(1,1))&=(\frac{\bar c}{2},\frac{\bar c}{2},1-\bar c-\theta^{(0,1)}_0,\theta^{(0,1)}_0)\\ (q_1(0,0),q_1(0,1),q_1(1,0),q_1(1,1))&=(\frac{\bar c}{2},\frac{\bar c}{2},\theta^{(1,0)}_1,\theta^{(0,1)}_1). \end{align}

Before giving the proof of the claim above, a few remarks are in order. To simplify the analysis, we assume that $\theta^{(0,0)}_0=\theta^{(0,0)}_1$ are known. Recall that the sharp identifying restrictions in (ref)-(ref) imply

align[align omitted — 78 chars of source]

The assumption $\theta^{(1,0)}_0<1-\bar c-\theta^{(0,1)}_0$ therefore makes the model incomplete at $\theta_0$. Finally, the assumption $\theta^{(1,0)}_1>1-\bar c-\theta^{(0,1)}_0$ is equivalent to $\nu_{\theta_1}(\{(1,0)\})>\nu_{\theta_0}(\{(1,0)\})$, which ensures that $\mathcal P_{\theta_0}\cap \mathcal P_{\theta_1}\ne\emptyset.$

proof[\rm Proof of Proposition (ref)] The following convex program characterizes the LFP: \begin{align} \min_{(p_0,p_1)\in \Delta^3\times\Delta^3}&\sum_{s\in\{(0,0),(0,1),(1,0),(1,1)\}}-\ln \big(\frac{p_0(s)}{p_0(s)+p_1(s)}\big)(p_0(s)+p_1(s))\\ s.t. & \theta^{(1,0)}_j\le p_j(1,0), j=0,1,\\ &\theta^{(0,1)}_j\le p_j(1,1), j=0,1,\\ &\bar c= p_j(0,0)+p_j(0,1), j=0,1. \end{align} First, we concentrate out $p_0(0,0),p_1(0,0),p_0(0,1),p_1(0,1)$ from the problem. The subset of the KKT conditions that involves these components are, for $s\in\{(0,0),(0,1)\}$ \begin{align} &-\frac{p_0(s)+p_1(s)}{p_0(s)}\frac{p_0(s)+p_1(s)-p_0(s)}{(p_0(s)+p_1(s))^2}(p_0(s)+p_1(s))-\ln \big(\frac{p_0(s)}{p_0(s)+p_1(s)}\big)-\lambda_1=0\\ &-\frac{p_0(s)+p_1(s)}{p_0(s)}\frac{-p_0(s)}{(p_0(s)+p_1(s))^2}(p_0(s)+p_1(s))-\ln \big(\frac{p_0(s)}{p_0(s)+p_1(s)}\big)-\lambda_2=0\\ & p_j(0,0)+p_j(0,1)=\bar c, j=0,1. \end{align} for some $\lambda_1\ne 0$ and $\lambda_2\ne 0.$ The first two conditions can be simplified as \begin{align} \frac{p_1(s)}{p_0(s)}=1+\lambda_2-\lambda_1, s\in\{(0,0),(0,1)\}, \end{align} which implies $\frac{p_1(0,0)}{p_0(0,0)}=\frac{p_1(0,1)}{p_0(0,1)},$ and hence, if $p_0(0,0)\in (0,\bar c),$ $\frac{p_1(0,0)}{p_1(0,1)}=\frac{p_0(0,0)}{p_0(0,1)}=\beta$ for some $0< \beta<\infty$. This, together with (ref) yields \begin{align} (p_j(0,0),p_j(0,1))=(\frac{\beta \bar c}{1+\beta},\frac{\bar c}{1+\beta}), j=0,1. \end{align} For example, one can take the following as a solution $(q_j(0,0),q_j(0,1))=(\frac{\bar c}{2},\frac{\bar c}{2})$. After concentrating out $p_0(0,0),p_1(0,0),p_0(0,1),p_1(0,1)$, the problem becomes maximizing \begin{align} \min_{(p_0,p_1)\in \Delta^3\times\Delta^3}&\sum_{s\in\{(1,0),(1,1)\}}-\ln \big(\frac{p_0(s)}{p_0(s)+p_1(s)}\big)(p_0(s)+p_1(s)) \end{align} subject to the constraints in (ref)-(ref). The KKT conditions are \begin{align} &-\frac{p_1(1,0)}{p_0(1,0)}-\ln \big(\frac{p_0(1,0)}{p_0(1,0)+p_1(1,0)}\big)-\chi_1-\chi_5=0 \\ &-\frac{p_1(1,1)}{p_0(1,1)}-\ln \big(\frac{p_0(1,1)}{p_0(1,1)+p_1(1,1)}\big)-\chi_3-\chi_5=0 \\ & 1-\ln \big(\frac{p_0(1,0)}{p_0(1,0)+p_1(1,0)}\big)-\chi_2-\chi_6=0 \\ & 1-\ln \big(\frac{p_0(1,1)}{p_0(1,1)+p_1(1,1)}\big)-\chi_4-\chi_6=0 \\ & \chi_1(p_0(1,0) - \theta^{(1,0)}_0)=0 \\ & \chi_2(p_1(1,0) - \theta^{(1,0)}_1)=0 \\ & \chi_3(p_0(1,1) - \theta^{(0,1)}_0)=0 \\ & \chi_4(p_1(1,1) - \theta^{(0,1)}_1)=0. \end{align} where $\chi_j\ge 0$ for $j=1,\dots,4$, and the original inequality and equality constraints are also imposed. Case 1: ($\chi_1=0,\chi_2>0,\chi_3>0,\chi_4=0$) Suppose $\chi_3>0$. Then, $p_0(1,1) =\theta^{(0,1)}_0$ by the complementary slackness condition (ref). It also implies $p_0(1,0)=1-\bar c-\theta^{(0,1)}_0$ by $ p_0(1,0)+ p_0(1,1)=1-\bar c$. Further, $\chi_1=0$ because the constraints associated with $\chi_1$ and $\chi_3$ cannot bind simultaneously due to the assumption that $\theta_0^{(1,0)}<1-\bar c-\theta_0^{(0,1)}.$ Suppose further that $\chi_2>0$. Then, $p_1(1,0)=\theta^{(1,0)}_1$ and $p_1(1,1)=1-\bar c-\theta^{(1,0)}_1.$ We also assume that $\chi_4=0.$ Now note that (ref)-(ref) reduce to \begin{align} &-\frac{\theta^{(1,0)}_1}{1-\bar c-\theta^{(0,1)}_0}-\ln \big(\frac{1-\bar c-\theta^{(0,1)}_0}{1-\bar c-\theta^{(0,1)}_0+\theta^{(1,0)}_1}\big)-\chi_5=0 \\ &-\frac{1-\bar c-\theta^{(1,0)}_1}{\theta^{(0,1)}_0}-\ln \big(\frac{\theta^{(0,1)}_0}{\theta^{(0,1)}_0+1-\bar c-\theta^{(1,0)}_1}\big)-\chi_3-\chi_5=0 \\ & 1-\ln \big(\frac{1-\bar c-\theta^{(0,1)}_0}{1-\bar c-\theta^{(0,1)}_0+\theta^{(1,0)}_1}\big)-\chi_2-\chi_6=0 \\ & 1-\ln \big(\frac{\theta^{(0,1)}_0}{\theta^{(0,1)}_0+1-\bar c-\theta^{(1,0)}_1}\big)-\chi_6=0 \end{align} It can be shown that, when $\theta_1^{(1,0)}>1-\bar c-\theta_0^{(0,1)}$, the system can be solved for $(\chi_2,\chi_3,\chi_5,\chi_6)$ that satisfies $\chi_j>0$ for $j=2,3,$ and hence the solution (the remaining components of LFP) is \begin{align} (q_0(1,0),q_0(1,1))&=(1-\bar c-\theta^{(0,1)}_0,\theta^{(0,1)}_0)\\ (q_1(1,0),q_1(1,1))&=(\theta^{(1,0)}_1,1-\bar c-\theta^{(1,0)}_1). \end{align} Case 2: ($\chi_1=0,\chi_2>0,\chi_3>0,\chi_4>0$) The difference from Case 1 is that $\chi_4>0$ is assumed. By the complementary slackness condition, this implies $p_1(1,1) = \theta^{(0,1)}_1$, which also equals $1-\bar c-\theta^{(1,0)}_1$ due to $\chi_2>0$ and $ p_1(1,0)+ p_1(1,1)=1-\bar c$. Therefore, the solutions are \begin{align} (q_0(1,0),q_0(1,1))&=(1-\bar c-\theta^{(0,1)}_0,\theta^{(0,1)}_0)\\ (q_1(1,0),q_1(1,1))&=(\theta^{(1,0)}_1,\theta^{(0,1)}_1). \end{align} Similar to Case 1, one may show that there exists $(\chi_2,\dots,\chi_6)$ that solves the KKT conditions with $\chi_j>0$ for $j=2,3,4$.

Efficient influence function, power envelope, and optimal tests

As in Section (ref), we consider testing $H_0:p'\theta=\theta^{(1,0)}\le c$ against $H_1:p'\theta>c$ for some $c\in (0,1)$ with $p=(0,0,1)'.$ Consider the following three configurations analyzed in the text (see also Figure (ref)) with $\bar c=\frac{1}{6}$. In each configuration, the null and alternative parameter values are

align[align omitted — 303 chars of source]

where $\xi$ and $h$ are taken from one of the following specifications:

description• : $\xi_{A}=(0,\xi^{(0,1)},\xi^{(1,0)})=(0,\frac{-1}{6},\frac{1}{6})$, $\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} }(\theta_0,\xi_A)=\{h=(h^{(0,0)},h^{(0,1)},h^{(1,0)}):h^{(0,0)}=0,h^{(1,0)}>0\}$; • : $\xi_{B}=(0,\xi^{(0,1)},\xi^{(1,0)})=(0,0,\frac{1}{6})$, $\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral1}} }(\theta_0,\xi_B)=\{h=(h^{(0,0)},h^{(0,1)},h^{(1,0)}):h^{(0,0)}=0,h^{(1,0)}>0,h^{(0,1)}+h^{(1,0)}<0\}$; • : $\xi_{B}=(0,\xi^{(0,1)},\xi^{(1,0)})=(0,0,\frac{1}{6})$, $\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral2}} }(\theta_0,\xi_B)=\{h=(h^{(0,0)},h^{(0,1)},h^{(1,0)}):h^{(0,0)}=0,h^{(1,0)}>0,h^{(0,1)}+h^{(1,0)}=0\}$.
propositionSuppose that the conditions of Proposition (ref) hold. Let $p=(0,0,1)'$ and let $\varphi(\theta)=p'\theta$. Then, Assumption (ref) holds for each of the cases above. Furthermore, for all cases, the efficient influence function is \begin{align} \tilde\varrho(s)=\frac{3}{5}1\{s=(1,0)\}-\frac{2}{5}\{s=(1,1)\}), \end{align} and the asymptotic power envelope $1-\Phi\Big(z_\alpha-\sqrt 5h^{(1,0)}\Big)$ is achieved by a test that rejects $H_0$ when the following statistic exceeds $z_\alpha$: \begin{align} T_{ \uppercase\expandafter{\romannumeral1} ,n}&=\frac{1}{\sqrt n}\sum_{i=1}^n\Big[\frac{3}{\sqrt 5}1\{s_i=(1,0)\}-\frac{2}{\sqrt 5}1\{s_i=(1,1)\}\Big]. \end{align}
proof[\rm Proof of Proposition (ref)] We analyze the three cases separately. Case A-$ \textup{\uppercase\expandafter{\romannumeral1}} $: The alternative parameter configuration satisfies the assumptions of Case 1 in Proposition (ref). Therefore, the LFPs are \begin{align} (q_0(0,0),q_0(0,1),q_0(1,0),q_0(1,1))&=(\frac{\bar c}{2},\frac{\bar c}{2},1-\bar c-\theta^{(0,1)}_0,\theta^{(0,1)}_0)=(\frac{1}{12},\frac{1}{12},\frac{1}{3},\frac{1}{2})\\ (q_1(0,0),q_1(1,1),q_1(1,0),q_1(1,1))&=(\frac{\bar c}{2},\frac{\bar c}{2},\theta^{(1,0)}_1,1-\bar c-\theta^{(1,0)}_1)=(\frac{1}{12},\frac{1}{12},\frac{1}{3}+\frac{h^{(1,0)}}{\sqrt n},\frac{1}{2}-\frac{h^{(1,0)}}{\sqrt n}). \end{align} Let $\vartheta\mapsto \mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\vartheta}$ be a model defined on a neighborhood of $\vartheta=0$ for which the density $\mathsf q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\vartheta}$ is given by \begin{align} (\mathsf q_{ \uppercase\expandafter{\romannumeral1} ,\vartheta}(0,0),\mathsf q_{ \uppercase\expandafter{\romannumeral1} ,\vartheta}(1,1),\mathsf q_{ \uppercase\expandafter{\romannumeral1} ,\vartheta}(1,0),\mathsf q_{ \uppercase\expandafter{\romannumeral1} ,\vartheta}(1,1))=(\frac{1}{12},\frac{1}{12},\frac{1}{3}+\vartheta^{(1,0)},\frac{1}{2}-\vartheta^{(1,0)}). \end{align} Due to the form of the LFP, Assumption (ref) is satisfied with the underlying model $\vartheta\mapsto \mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\vartheta}$. Calculations similar to the ones in (ref) can ensure that the $L^2$-derivative is \begin{align} \dot\ell_{ \uppercase\expandafter{\romannumeral1} ,0}(s)= 1\{s=(1,0)\} \begin{pmatrix} 0\\ 0\\ 3 \end{pmatrix} - 1\{s=(1,1)\} \begin{pmatrix} 0\\ 0\\ 2 \end{pmatrix}. \end{align} Hence, the tangent cone of the model is \begin{align} \mathcal G_{ \textup{\uppercase\expandafter{\romannumeral1}} ,0}=\big\{g\in L^2_{Q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,0}}:g=3h^{(1,0)} 1\{s=(1,0)\}-2h^{(1,0)}1\{s=(1,1)\}),h^{(1,0)}>0\big\}. \end{align} The influence curve of $\varphi$ must satisfy $p'h=\langle \varrho_{ \textup{\uppercase\expandafter{\romannumeral1}} },g\rangle_{L^2_{\mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,0}}}$ for $g\in \mathcal G_{ \textup{\uppercase\expandafter{\romannumeral1}} ,\theta_0+\xi}$. Since $p=(0,0,1)',$ $\varrho_{ \textup{\uppercase\expandafter{\romannumeral1}} }=\frac{3}{5}1\{s=(1,0)\}-\frac{2}{5}\{s=(1,1)\})$ satisfies the requirement. Observe also that $\varrho_{ \textup{\uppercase\expandafter{\romannumeral1}} }$ is in $\mathcal G_{ \textup{\uppercase\expandafter{\romannumeral1}} ,0}$ and hence in its closure. Therefore, it is the efficient influence function, i.e. $\varrho_{ \textup{\uppercase\expandafter{\romannumeral1}} }=\tilde\varrho_{ \textup{\uppercase\expandafter{\romannumeral1}} },$ which in turn implies $\|\tilde{\varrho}_{ \textup{\uppercase\expandafter{\romannumeral1}} }\|_{L^2_{\mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral1}} },0}}=1/\sqrt 5.$ By Theorem (ref), for any level-$\alpha$ test $\phi_n$, \begin{align} \limsup_{n\to\infty}\pi_{n,\theta_{n,\xi_A,h}}(\phi_n)\le 1-\Phi\Big(z_\alpha-\sqrt 5h^{(1,0)}\Big). \end{align} Again, by Theorem (ref), this bound can be achieved by a test that rejects the null when the statistic in (ref) exceeds $z_\alpha$. \textbf{Case B-$ \textup{\uppercase\expandafter{\romannumeral1}} $}: The alternative parameter configuration again satisfies the assumptions of Case 1 in Proposition (ref). Therefore, the LFPs are given as in (ref)-(ref). The rest of the analysis parallels Case A-$ \textup{\uppercase\expandafter{\romannumeral1}} $ and is omitted. \textbf{Case B-$ \textup{\uppercase\expandafter{\romannumeral2}} $}: The alternative parameter configuration satisfies the assumptions of Case 2 in Proposition (ref). Hence, the LFP is \begin{align} (q_0(0,0),q_0(0,1),q_0(1,0),q_0(1,1))&=(\frac{\bar c}{2},\frac{\bar c}{2},1-\bar c-\theta^{(0,1)}_0,\theta^{(0,1)}_0)=(\frac{1}{12},\frac{1}{12},\frac{1}{3},\frac{1}{2}),\\ (q_1(0,0),q_1(0,1),q_1(1,0),q_1(1,1))&=(\frac{\bar c}{2},\frac{\bar c}{2},\theta^{(1,0)}_1,\theta^{(0,1)}_1)=(\frac{1}{12},\frac{1}{12},\frac{1}{3}+\frac{h^{(1,0)}}{\sqrt n},\frac{1}{2}+\frac{h^{(0,1)}}{\sqrt n}) \end{align} Let $\vartheta\mapsto \mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\vartheta}$ be a model defined on a neighborhood of $\vartheta=0$ for which the density $\mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\vartheta}$ is given by \begin{align} (\mathsf q_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\vartheta}(0,0),\mathsf q_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\vartheta}(1,1),\mathsf q_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\vartheta}(1,0),\mathsf q_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\vartheta}(1,1))=(\frac{1}{12},\frac{1}{12},\frac{1}{3}+\vartheta^{(1,0)},\frac{1}{2}+\vartheta^{(0,1)}). \end{align} Assumption (ref) is satisfied with the underlying model $\vartheta\mapsto \mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral2}} ,\vartheta}$. Calculations similar to the ones in (ref) can ensure that the $L^2$-derivative is \begin{align} \dot\ell_{ \textup{\uppercase\expandafter{\romannumeral2}} ,0}(s)= 1\{s=(1,0)\} \begin{pmatrix} 0\\ 0\\ 3 \end{pmatrix} + 1\{s=(1,1)\} \begin{pmatrix} 0\\ 2\\ 0 \end{pmatrix}. \end{align} Recall that the local parameter space is $\mathcal T_{ \textup{\uppercase\expandafter{\romannumeral2}} }(\theta_0,\xi_B)=\{h=(h^{(0,0)},h^{(0,1)},h^{(1,0)}):h^{(0,0)}=0,h^{(1,0)}>0,h^{(0,1)}+h^{(1,0)}=0\}$. Hence, the tangent cone of the model is \begin{align} \mathcal G_{ \textup{\uppercase\expandafter{\romannumeral2}} ,0}=\big\{g\in L^2_{\mathsf Q_{ \textup{\uppercase\expandafter{\romannumeral1}} ,0}}:g=3h^{(1,0)} 1\{s=(1,0)\}-2h^{(1,0)}1\{s=(1,1)\}),h^{(1,0)}>0\big\}, \end{align} where we used $h^{(0,1)}=-h^{(1,0)}$. Observe that the tangent cone coincides with the one in (ref). The rest of the analysis parallels Case A-$ \textup{\uppercase\expandafter{\romannumeral1}} $ and is omitted.

Computing the $L^2$-derivative

In practice, the $L^2$ derivative, a key object for constructing the optimal tests, can be computed by an analytical method or a quadratic-programming method.\footnote{Another possibility is to use numerical differentiation, which is studied by HongLi2018JoE in a related context.} In what follows, let $\theta_1=\theta_0+\xi+\tau h$ with $\tau>0$ and assume that $\mathcal P_{\theta_0}\cap\mathcal P_{\theta_1}=\emptyset.$ A robustly testable alternative is a special case, and the argument below can be applied with $\xi=0.$

Analytical method

Among the two, the analytical method is more straightforward and is recommended whenever possible. Let $h$ be given and assume that Assumption (ref) or (ref) holds. For calculating the $L^2$ derivative, one can take the following steps.

Step 1: For $\tau>0$, derive the LFP $(Q_{0,\tau h},Q_{1,\tau h})$ in closed form by solving (ref). Find a collection of models $\vartheta\mapsto \mathsf Q_{j,\vartheta},j\in\mathbb J$ such that (ref) (or (ref)) holds.

Step 2: For each $j$, calculate $\dot\ell_{j,0}$ using the analytical form of $\mathsf Q_{j,\vartheta}.$

This is the approach taken in the analysis of Examples (ref) and (ref) (see Appendix (ref) and (ref)). Also, in Step 1, one should check if (ref) (or (ref)) is satisfied for $Q_{0,\tau h}$. In Example (ref), it is satisfied because $\mathcal P_{\theta_0}=\{Q_0\}$ is a singleton and by an inspection of the underlying model. In Example 2, the complementary slackness condition determines $Q_{0,\tau h}$ uniquely for all $\tau$ sufficiently small, and again the condition can be checked analytically.

Quadratic-Programming method

Under certain conditions, the $L^2$-derivative of the model can also be calculated as a solution to an auxiliary quadratic programming (QP) problem. Below, we discuss this alternative approach and provide primitive conditions for this method to work.

This approach is based on the stability analysis of a solution to a parametric convex program. Toward this end, we relate the problem in (ref) to a parametric optimization problem studied in Shapiro:1988aa. Let $L$ be the cardinality of $S$ and let $ K$ be the cardinality of $2^S$, the power set of $S$. Let $p_0,p_1\in[0,1]^L$ denote vectors of probability mass functions on $S$. Let $x=(p_0',p_1')'\in [0,1]^{2L}$ denote a vector of control variables, and let $y=\theta_1$ denote a parameter vector in the optimization problem (defined below).

Define

align[align omitted — 86 chars of source]

and note that it does not depend on $y$. We impose constraints $\tilde g_\ell(x,y)\le 0$ for $\ell=1,\cdots,2(K+L)$ and $\tilde g_\ell(x,y)= 0$ for $\ell=2(K+L)+1,2(K+L)+2$ with

align[align omitted — 519 chars of source]

These restrictions impose the sharp identifying restrictions together with the constraints that restrict $p_j, j=0,1$ in the probability simplex. Before proceeding, one can reduce the constraints by the following procedures.\footnote{This step is not strictly required, but it is recommended as the reduced system is more likely to satisfy Assumption (ref).}

First, reduce the first $2K$ inequality constraints to the ones that are generated by a {\em core determining class\/} $\mathcal C\subset 2^S$ galichon2006inference,galichon2011set.\footnote{galichon2011set provide a tractable characterization of the core determining class for incomplete models with $G$ possessing a certain monotonicity property. Luo:2017aa,Luo:2017ab provide algorithms to construct the core determining class for general incomplete models.} Second, define

align[align omitted — 185 chars of source]

which collects events that remain complete both under the null and local alternatives. For any $A\in \mathcal C$, combine the two inequality constraints, $\nu_{\theta_j}(A)-p_j(A)\le 0$ and $\nu_{\theta_j}(A^c)-p_j(A^c)\le 0$, and impose them as an equality constraint $\nu_{\theta_j}(A)=p_j(A).$ Finally, let $\mathcal B\equiv \mathcal C\setminus \mathcal A$, which collects the events associated with the remaining inequality restrictions. We note that some of the inequality constraints associated with events in this collection may still bind when $\tau=0.$ Let $M$ be the number constraints after reducing the restrictions and let the first $M_1\le M$ of the restrictions collect equality constraints.

The resulting convex program can then be written as

align[align omitted — 150 chars of source]

for some $g_\ell:\mathbb R^{2L}\times \Theta\to\mathbb R,\ell=1,\dots, M$. This is the setting analyzed in Shapiro:1988aa.

Let $x_\tau=(q_{0,\tau}',q_{1,\tau}')'$ be a solution to the problem above when $y$ is set to $y_\tau=\theta_0+\xi+\tau h$. Let $x_0=(q_{0,0}',q_{1,0}')'$ denote the solution when $\tau=0$. We then construct two sets $\mathbf A$ and $\mathbf B$, which collect gradient vectors of the equality and binding inequality constraints at $x_0$ as follows.

enumerate• Gradients associated with equality constraints ($\mathbf A$): For each $A\in\mathcal A$ and $j\in \{0,1\}$, write the equality restriction $\nu_{\theta_j}(A)-\sum_{s\in A}q_{j,0}(s)= 0$ as $\nu_{\theta_j}(A)-a'x_0=0$ for a vector $a\in \{0,1\}^{2L}$. Let $\mathbf A$ collect such vectors together with two additional vectors $(1'_L,0'_L)'$ and $(0'_L,1'_L)'$ that are associated with the equality constraints $\sum_{s\in S}q_{j,0}(s)-1=0$ for $j=0,1$. • Gradients associated with inequality constraints ($\mathbf B$): Let $\mathcal B_0\equiv\{A\in\mathcal B:\nu_{\theta_0}(A)=\sum_{s\in A}q_{0,0}(s)\text{ or }\nu_{\theta_1}(A)=\sum_{s\in A}q_{1,0}(s)\}$ be the collection of events for which an inequality constraint binds at $\tau=0.$ For each $A\in\mathcal B_0$ and $j\in \{0,1\}$, write the active inequality restriction $\nu_{\theta_j}(A)-\sum_{s\in A}q_{j,0}(s)= 0$ as $\nu_{\theta_j}(A)-b'x_0=0$ for a vector $b\in \{0,1\}^{2L}$. Let $\mathbf B$ collect all such vectors.

Recall that $H$ is a twice continuously differentiable convex function. The following condition is sufficient for the directional differentiability of the model.

assumption(i) $H$ is such that $H(z)>0$, $H'(z)>0$ (or $H'(z)<0$), and $H''(z)>0$ for all $z\in [0,1]$; (ii) The solution $x_\tau=(q_{0,\tau}',q_{1,\tau}')'$ is unique at $\tau=0$; (iii) The vectors in $\mathbf A$ are linearly independent, and there is a vector $u\in\mathbb R^{2L}$ such that $u'a=0$ for all $a\in\mathbf A$ and $u'b<0$ for all $b\in\mathbf B.$

Assumption (ref) (i) is a regularity condition on the objective function, which ensures that the solution of the problem satisfies a weak second-order condition. One can choose $H$ that satisfies these conditions. Assumption (ref) (ii) requires that the solution, hence the LFP (not only its ratio), is unique when $\tau=0$. This condition is satisfied in general when $\mathcal P_{\theta_0}\cap \mathcal P_{\theta_0+\xi}$ is a singleton. A special case is that the local alternatives are robustly testable and the model is complete at $\tau=0$, in which case $\mathcal P_{\theta_0}$ and $\mathcal P_{\theta_0+\tau h}$ coincide at $\tau=0$ and is a singleton. This occurs, for example, in Example (ref). Assumption (ref) (iii) imposes the Mangasarian-Fromovitz constraint qualification (MFCQ). This condition ensures that the set of Lagrange multipliers is bounded for all $\tau\in (0,\bar\tau]$ Gauvin:1977aa,Shapiro:1988aa, which is the key for the stability of the solution to local perturbations. Note that the constraints are linear and the gradient vectors in $\mathbf A\subset \{0,1\}^{2L}$ and $\mathbf B\subset \{0,1\}^{2L}$ do not need to be estimated. Hence, checking this condition is relatively straightforward.

Write the solution to (ref)-(ref) as $x(y)$ and denote its directional derivative with direction $h$ by

align*[align* omitted — 81 chars of source]

Let $\mathcal X_\tau\subset\mathbb R^M$ be the set of Lagrange multipliers under $\tau$ in a neighborhood of 0, which is, under MFCQ, bounded and the convex hull of a finite set $E_\tau$ of extreme points Shapiro:1988aa. For each $\chi\in\mathbb R^M$, let $J_+(\chi)=\{\ell:\chi^{(\ell)}>0,\ell=M_1+1,\dots,M\}$, $J_0=\{\ell:\chi^{(\ell)}=0,\ell=M_1+1,\dots,M\}$, and $J(\chi)=\{1,\dots,M_1\}\cup J_+(\chi)$. Each element $\chi$ in $E_0$ is such that the gradient vectors $\{\nabla_x g_\ell(x_0,y_0),\ell\in J(\chi)\}$ are linearly independent. For each $\ell=1,\dots, M$, define

align*[align* omitted — 84 chars of source]

and note that $\alpha_\ell(u,h)=u'a+h'\nabla_{y}\nu_{y_0}(A)$ for some $a\in \{0,1\}^{2L}$ if $g_\ell$ originated from one of the constraints in (ref) and $\alpha_\ell(u,h)=u'a$ otherwise because the other constraints (in (ref) and (ref)-(ref)) do not involve $\theta_1$. We then have the following directional differentiability result.

propositionSuppose Assumption (ref) holds. Then, $x(y)$ is directionally differentiable at $y=\theta_0+\xi$ with the directional derivative \begin{align} Dx(\theta_0+\xi)[h]=\operatorname*{arg\,min}_{u\in\bar\Sigma(h)}\bar\zeta(u), \end{align} where \begin{align} \bar\zeta(u)&=\sum_{\ell=1}^LH”(\frac{q_0(s^{(\ell)})}{q_0(s^{(\ell)})+q_1(s^{(\ell)})})\frac{(u_0^{(\ell)}q_1(s^{(\ell)})-u_1^{(\ell)}q_0(s^{(\ell)}))^2}{(q_0(s^{(\ell)})+q_1(s^{(\ell)}))^3}\\ \bar\Sigma(h)&=\Big\{u:\alpha_\ell(u,h)=0,\ell\in J(\chi),\alpha_\ell(u,h)\le 0,\ell\in J_0(\chi), \chi\in E_0\Big\}. \end{align}

Note that each $\alpha_\ell$ is linear in $u$ and the number of constraints is finite because $E_0$ is finite. Hence, $Dx(\theta_0+\xi)[h]$ is a solution to a finite-dimensional (convex) QP.

To satisfy Assumption (ref) (or Assumption (ref)), we also need to ensure (ref) (or (ref)). For this, among the inequality conditions in (ref), let $J_{H_0}\subset \{M_1+1,\dots,M\}$ collect the indices associated with the constraints that restrict $q_{0,\tau}$, i.e. those of the form: $\nu_{\theta_0}(A_\ell)-\sum_{s\in A_\ell}p_{0}(s)\le 0$. We then let $J_{+,H_0}(\chi)\equiv\{\chi^{(\ell)}>0,\ell\in J_{H_0}\}$. The following condition is sufficient for (ref) (or (ref)).

assumptionThere is a path $\tau\mapsto\chi_\tau$ and $\bar\tau>0$ such that (i) $\chi_\tau\in\mathcal X_\tau,\forall \tau\in (0,\bar\tau]$; (ii) $J_{+,H_0}(\chi_\tau)=J_{+,H_0}(\chi_0)$ for all $\tau\in (0,\bar\tau]$; and (iii) the solution $q_{0,0}\in \Delta(S)$ is uniquely determined by \begin{align} \nu_{\theta_0}(A_\ell)-\sum_{s\in A_\ell}q_{0,0}(s)=0, \ell\in J_{+,H_0}(\chi_0). \end{align}

In other words, there is a configuration of the Lagrange multipliers, under which the binding constraints under $H_0$ uniquely determines $Q_0$, which remains constant across $\tau\in (0,\bar\tau]$. As $\tau$ varies, the remaining components of the solution, $q_{1,\tau},\chi_\tau$, are properly adjusted to satisfy the first order conditions (see (ref)).

For each $j\in\mathbb J$ and $h\in\mathcal T_{j}(\xi,\theta_0)$, $Dx(\theta_0+\xi)[h]$ is a $2L$-dimensional vector. Let us write it as $(u_0^{*,(1)},\dots,u_0^{*,(L)},u_1^{*,(1)},\dots,u_1^{*,(L)})'$. The following corollary gives a formula to calculate the $L^2$-derivative for $h\in\mathcal T_j(\xi,\theta_0)$.

corollarySuppose the conditions of Proposition (ref) hold. Suppose Assumption (ref) holds. Then, the $L^2$-derivative $\dot\ell_{j,0}$ satisfies \begin{align} h'\dot\ell_{j,0}(s)=\sum_{\ell=1}^L\frac{u_1^{*,(\ell)}1\{s=s^{(\ell)}\}}{\mathsf q_{j,0}(s)}. \end{align}

The proof of the results above are collected at the end of this appendix.

Proof of Proposition (ref) and Corollary (ref)

proof[\rm Proof of Proposition (ref)] The proof proceeds by showing the conditions of Theorem 5.1 in Shapiro:1988aa. First, for any $y\in\Theta$, the set of feasible solutions $\Omega(y)=\{x=(p_0',p_1')':g_\ell(x,y)= 0,\ell=1,\dots,M_1,g_\ell(x,y)\le 0,\ell=M_1+1,\dots,M\}$ is a subset of a compact set in $\mathbb R^{2L}$ because $p_j$ is in the probability simplex for $j=0,1.$ This ensures Assumption 1 of Shapiro:1988aa. Assumption (ref) then ensures Assumption 2 in Shapiro:1988aa. Similarly, Assumption (ref) (iii) ensures Assumption 3 of Shapiro:1988aa. Finally, Assumption 4 in Shapiro:1988aa is a second-order condition on the Lagrangian, which we show below. Define \begin{align} \mathcal L(x,y,\chi)=\sum_{s\in S}H(\frac{p_0(s)}{p_0(s)+p_1(s)})[p_0(s)+p_1(s)]+\sum_{k=1}^M\chi^{(k)} g_k(x,y) \end{align} Note that $g_k(x,y)$ is affine in $x$ for all $k$ so that one may write $a_k'x-b_k(y)$ for some $a_k\in\{0,1\}^{2L}$ and $b_k:\Theta\to \mathbb R$. Therefore, for $\ell=1,\dots,L$, \begin{align} \frac{\partial}{\partial x^{(\ell)}}\mathcal L(x,y,\chi) &=H'(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})\frac{p_1(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})}+H(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})+\sum_{k=1}^M\chi^{(k)}a_{k}^{(\ell)},\end{align} and for $\ell=L+1,\dots,2L,$ \begin{align} \frac{\partial}{\partial x^{(\ell)}}\mathcal L(x,y,\chi) &=H'(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})\frac{-p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})}+H(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})+\sum_{k=1}^M\chi^{(k)}a_{k}^{(\ell)}. \end{align} Below, we derive the Hessian matrix of the Lagrangian. First, suppose $\ell\in \{1,\dots,L\}$ and $m\in\{1,\dots,L\}$ and $\ell=m$. Then, \begin{align*} \frac{\partial^2}{\partial x^{(\ell)}\partial x^{(m)}}\mathcal L(x,y,\chi)&=H”(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})\frac{p_1(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})}\times\frac{p_0(s^{(\ell)})+p_1(s^{(\ell)})-p_0(s^{(\ell)})}{(p_0(s^{(\ell)})+p_1(s^{(\ell)}))^2}\\ &\quad+H'(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})\frac{-p_1(s^{(\ell)})}{(p_0(s^{(\ell)})+p_1(s^{(\ell)}))^2}\\ &\quad+H'(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})\frac{p_0(s^{(\ell)})+p_1(s^{(\ell)})-p_0(s^{(\ell)})}{(p_0(s^{(\ell)})+p_1(s^{(\ell)}))^2}\\ &=H”(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})\frac{p_1(s^{(\ell)})^2}{(p_0(s^{(\ell)})+p_1(s^{(\ell)}))^3} \end{align*} Next, suppose $\ell\in \{1,\dots,L\}$ and $m\in \{L+1,\dots, 2L\}$ and $m=\ell+L$. Then, a similar calculation yields \begin{align*} \frac{\partial^2}{\partial x^{(\ell)}\partial x^{(m)}}\mathcal L(x,y,\chi)&=H”(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})\frac{-p_0(s^{(\ell)})p_1(s^{(\ell)})}{(p_0(s^{(\ell)})+p_1(s^{(\ell)}))^3} \end{align*} Finally, suppose $\ell\in \{L+1,\dots,2L\}$ and $m\in \{L+1,\dots, 2L\}$ and $\ell=m$. Then, \begin{align*} \frac{\partial^2}{\partial x^{(\ell)}\partial x^{(m)}}\mathcal L(x,y,\chi) &=H”(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})\frac{p_0(s^{(\ell)})^2}{(p_0(s^{(\ell)})+p_1(s^{(\ell)}))^3}. \end{align*} All other cases (with $\ell\ne m$) lead to 0 second order derivatives. Therefore, \begin{align} u'\nabla^2_{xx}\mathcal L(x_0,y_0,\chi)u &=\sum_{\ell=1}^LH”(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})\frac{(u_0^{(\ell)}p_1(s^{(\ell)})-u_1^{(\ell)}p_0(s^{(\ell)}))^2}{(p_0(s^{(\ell)})+p_1(s^{(\ell)}))^3}\ge 0. \end{align} where $u=(u_0',u_1')'.$ The object above is non-negative in general. For the second order condition in Shapiro:1988aa, we need it to be strictly positive for any non-zero vector $u$ in the following critical cone $C$ \begin{align} C=\{u:u'\nabla_xg_\ell(x_0,y_0)=0,\ell=1,\dots,M_1; u'\nabla_xg_\ell(x_0,y_0)\le 0, \ell=M_1+1,\dots,M; u'\nabla f_x(x_0,y_0)\le 0\}. \end{align} The restriction, $u'\nabla f_x(x_0,y_0)\le 0$, can also be written as \begin{multline} \sum_{\ell=1}^LH'(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})\frac{u_0^{(\ell)}p_1(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})}+u_0^{(\ell)}H(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})\\ +\sum_{\ell=1}^LH'(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})\frac{-u_1^{(\ell)}p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})}+u_1^{(\ell)}H(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})\le 0, \end{multline} and hence \begin{align} \sum_{\ell=1}^LH'(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})\frac{u_0^{(\ell)}p_1(s^{(\ell)})-u_1^{(\ell)}p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})}\le -\sum_{\ell=1}^L(u_0^{(\ell)}+u_1^{(\ell)})H(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})}). \end{align} Now note that, by taking singleton events $A=\{s^{(\ell)}\},\ell=1,\dots,L$, part of the restrictions $u'\nabla_xg_\ell(x_0,y_0)\le 0,~\ell=M_1+1,\dots,M$ must include gradients of the form $\nabla_x g_\ell(x_0,y_0)=(-e_\ell',0_L')',\ell=1,\dots,L$ where $e_\ell$ is a vector of zeros whose $\ell$-th component is 1. This therefore requires $u_0^{(\ell)}\ge 0$ for all $\ell$. Similarly, $u_1^{(\ell)}\ge 0$ for all $\ell$. Further, Shapiro' second order condition requires $u\ne 0$, and hence, $u_j^{(\ell)}>0$ for some $\ell$ and $j\in \{0,1\}$. This implies that the RHS of (ref) is strictly negative if we choose $H$ to be strictly positive on $[0,1]$ as assumed in Assumption (ref). This in turn implies $u$ must satisfy \begin{align} \sum_{\ell=1}^LH'(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})\frac{u_0^{(\ell)}p_1(s^{(\ell)})-u_1^{(\ell)}p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})}<0. \end{align} Since $H'$ is strictly positive (or strictly negative) for all $z\in [0,1]$, for at least one $\ell$, one must have $u_0^{(\ell)}p_1(s^{(\ell)})-u_1^{(\ell)}p_0(s^{(\ell)})<0$ (or $>0$) to satisfy the inequality above. Therefore, for any $u\in C$ and $u\ne 0$, one must have \begin{align} u'\nabla^2_{xx}\mathcal L(x_0,y_0,\chi)u>0 \end{align} This ensures Assumption 4 of Shapiro:1988aa. Note that $\bar\zeta$ in Proposition (ref) corresponds to $\zeta_{v,w}$ in Eq. (4.10) in Shapiro:1988aa, where $w=0$ in our setting. To see this, observe that one of the terms in this function is $\xi_\lambda$ defined in Eq. (3.3) in Shapiro:1988aa, which in our setting equals $u'\nabla^2_{xx}\mathcal L(x_0,y_0,\chi)u$ because the components of $\nabla^2_{xy}\mathcal L(x_0,y_0,\chi)$ and $\nabla^2_{yy}\mathcal L(x_0,y_0,\chi)$ are 0 due to $\nabla_y \mathcal L(x,y,\chi)$ being a constant as $g_k$ is separable in $(x,y)$. Therefore, by (ref) and this function being independent of $\chi$, we obtain \begin{align} \bar\zeta(u)=\sum_{\ell=1}^LH”(\frac{p_0(s^{(\ell)})}{p_0(s^{(\ell)})+p_1(s^{(\ell)})})\frac{(u_0^{(\ell)}p_1(s^{(\ell)})-u_1^{(\ell)}p_0(s^{(\ell)}))^2}{(p_0(s^{(\ell)})+p_1(s^{(\ell)}))^3}. \end{align} Also observe that $\bar \Sigma(h)$ in (ref) corresponds to $\bar\Sigma(v)$ in Eq. (2.12) in Shapiro:1988aa. The claim of the proposition now follows from Theorem 5.1 of Shapiro:1988aa.
proof[\rm Proof of Corollary (ref)] Let $(q_{0,j,\tau h},q_{1,j,\tau h})$ be the densities of the LFP between $\theta_0$ and $\theta_0+\xi+\tau h$ and recall that they are the solutions of a parametric convex program. By Assumption (ref), there is a model $\vartheta\mapsto \mathsf Q_{j,\vartheta}$ such that $q_{0,j,\tau h}=\mathsf q_{j,0}$ and $q_{1,j,\tau h}=\mathsf q_{j,\tau h}$ for all $\tau\in (0,\bar\tau]$ for some $\bar \tau>0$, where $\mathsf q_{j,\vartheta}$ is the density of $\mathsf Q_{j,\vartheta}$. By Proposition (ref), the LFP density $\vartheta\mapsto \mathsf q_{j,\vartheta}$ is directionally differentiable with the directional derivative $u_1^{*,(\ell)}=\lim_{\tau\downarrow 0}\frac{\mathsf q_{j,\tau h}(s^{(\ell)})-\mathsf q_{j,0}(s^{(\ell)})}{\tau}$. By the chain rule, this implies \begin{align} \lim_{\tau\downarrow 0}\frac{\mathsf q^{1/2}_{j,\tau h}(s^{(\ell)})-\mathsf q^{1/2}_{j,0}(s^{(\ell)})}{\tau}=\frac{u_1^{*,(\ell)}}{\mathsf q_{j,0}(s^{(\ell)})}. \end{align} Noting that $u_1^{*,(\ell)}$ is finite, and interchanging limits and sum, we obtain \begin{align} \lim_{\tau\downarrow 0}\int_S (\mathsf q^{1/2}_{j,\tau h}(s)-\mathsf q^{1/2}_{j,0}(s)-&\tau\sum_{\ell=1}^L\frac{u_1^{*,(\ell)}1\{s=s^{(\ell)}\}}{\mathsf q_{j,0}(s)})^2\mathsf q_{j,0}d\mu(s)\notag\\ &=\lim_{\tau\downarrow 0}\sum_{\ell=1}^L(\mathsf q^{1/2}_{j,\tau h}(s^{(\ell)})-\mathsf q^{1/2}_{j,0}(s^{(\ell)})-\tau\frac{u_1^{*,(\ell)}}{\mathsf q_{j,0}(s^{(\ell)})})^2\mathsf q_{j,0}(s^{(\ell)})\notag\\ &=\sum_{\ell=1}^L\big(\lim_{\tau\downarrow 0}\big[\mathsf q^{1/2}_{j,\tau h}(s^{(\ell)})-\mathsf q^{1/2}_{j,0}(s^{(\ell)})-\tau\frac{u_1^{*,(\ell)}}{\mathsf q_{j,0}(s^{(\ell)})}\big]\big)^2\mathsf q_{j,0}(s^{(\ell)})=0. \end{align} This ensures (ref).