EconBase
← Back to paper

Monte Carlo Confidence Sets for Identified Sets

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

158,429 characters · 28 sections · 62 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Monte Carlo Confidence Sets for Identified Sets

abstract\singlespacing In complicated/nonlinear parametric models, it is generally hard to know whether the model parameters are point identified. We provide computationally attractive procedures to construct confidence sets (CSs) for identified sets of full parameters and of subvectors in models defined through a likelihood or a vector of moment equalities or inequalities. These CSs are based on level sets of optimal sample criterion functions (such as likelihood or optimally-weighted or continuously-updated GMM criterions). The level sets are constructed using cutoffs that are computed via Monte Carlo (MC) simulations directly from the quasi-posterior distributions of the criterions. We establish new Bernstein-von Mises (or Bayesian Wilks) type theorems for the quasi-posterior distributions of the quasi-likelihood ratio (QLR) and profile QLR in partially-identified regular models and some non-regular models. These results imply that our MC CSs have exact asymptotic frequentist coverage for identified sets of full parameters and of subvectors in partially-identified regular models, and have valid but potentially conservative coverage in models with reduced-form parameters on the boundary. Our MC CSs for identified sets of subvectors are shown to have exact asymptotic coverage in models with singularities. We also provide results on uniform validity of our CSs over classes of DGPs that include point and partially identified models. We demonstrate good finite-sample coverage properties of our procedures in two simulation experiments. Finally, our procedures are applied to two non-trivial empirical examples: an airline entry game and a model of trade flows.

Introduction

It is often difficult to verify whether parameters in complicated nonlinear structural models are globally point identified. This is especially the case when conducting a sensitivity analysis to examine the impact of various model assumptions on the estimates of parameters of interest, where relaxing some suspect assumptions may lead to loss of point identification. This difficulty of verifying point identification naturally calls for inference procedures that are valid whether or not the parameters of interest are point identified. Our goal is to contribute to this sensitivity literature by proposing relatively simple inference procedures that allow for partial identification in models defined through a likelihood or a vector of moment equalities or inequalities.

To that extent, we provide computationally attractive and asymptotically valid confidence set (CS) constructions for the identified set $\Theta_I$ of the full vector of parameters $\theta \equiv (\mu,\eta)\in \Theta$,\footnote{Following the literature, the identified set $\Theta_I$ is the argmax of a population criterion over the whole parameter space $\Theta$. A model is point identified if $\Theta_I$ is a singleton, say $\{\theta_0\}$, and partially identified if $\{\theta_0\} \subsetneq \Theta_I \subsetneq \Theta$.} and for the identified sets $M_I$ of subvectors $\mu$. As a sensitivity check in an empirical study, a researcher could report conventional CSs based on inverting a $t$ or Wald statistic, which are valid under point identification only, alongside our new CSs that are asymptotically optimal under point identification and robust to failure of point identification.

Our CS constructions are criterion-function based, as in CHT (CHT) and the subsequent literature on CSs for identified sets. That is, contour sets of the sample criterion function are used as CSs for $\Theta_I$ and contour sets of the sample profile criterion are used as CSs for $M_I$. However, our CSs are constructed using critical values that are calculated differently from those in the existing literature. In two of our proposed CS constructions, we estimate critical values using quantiles of the sample criterion function (or profile criterion) that are simulated from a quasi-posterior distribution, which is formed by combining the sample criterion function with a prior over the model parameter space $\Theta$.\footnote{In correctly-specified likelihood models the quasi-posterior is a true posterior distribution over $\Theta$. We refer to the distribution as a quasi-posterior because we accommodate non-likelihood based models, such as moment-based models with GMM criterions.}

We propose three procedures for constructing various CSs. To construct a CS for the identified set $\Theta_I$, our Procedure 1 draws a sample $\{\theta^1,...,\theta^B\}$ from the quasi-posterior, computes the $\alpha$-quantile of the sample criterion evaluated at the draws, and then defines our CS $\wh\Theta_\alpha$ for $\Theta_I$ as the contour set at said $\alpha$-quantile. The computational complexity here is simply as hard as the problem of taking draws from the quasi-posterior, a well-researched and understood area in the literature on Monte Carlo (MC) algorithms in Bayesian computation (see, e.g., liu, CRobert). Many MC samplers (including the popular Markov Chain Monte Carlo (MCMC) algorithms) could, in principle, be used for this purpose. In our simulations and empirical applications, we use an adaptive sequential Monte Carlo (SMC) algorithm that is well-suited to drawing from irregular, multi-modal (quasi-)posteriors and is also easily parallelizable for fast computation (see, e.g., HS2014, DDJ2012, DG2014). Our Procedure 2 produces a CS $\wh M_\alpha$ for $M_I$ of a general subvector using the same draws from the quasi-posterior as in Procedure 1. Here an added computation step is needed to obtain critical values that guarantee the exact asymptotic coverage for $M_I$. Finally, our Procedure 3 CS for $M_I$ of a scalar subvector is simply the contour set of the profiled quasi-likelihood ratio (QLR) with its critical value being the $\alpha$ quantile of a chi-square distribution with one degree of freedom. Our Procedure 3 CS is simple to compute but is valid only for scalar subvectors.

Our CS constructions are valid for “optimal” criterions, which include (but are not limited to) correctly-specified likelihood models, GMM models with optimally-weighted or continuously-updated or GEL criterions,\footnote{Moment inequality-based models are special cases of moment equality-based models as one can add nuisance parameters to transform moment inequalities into moment equalities. Although moment inequality models are allowed, our criterion differs from the popular GMS criterion for moment inequalities in AndrewsSoares and others; see Subsections (ref), (ref) and (ref).} or sandwich quasi-likelihoods. For point- or partially-identified regular models, optimal criterions correspond to criterions that satisfy a generalized information equality. But our optimal criterions also allow for some correctly-specified non-regular (or non-standard) models such as models with parameter-dependent support, an important feature of set identified models (see Appendix (ref)). Our Procedure 1 and 2 CSs, $\wh\Theta_\alpha$ and $\wh M_\alpha$, are shown to have exact asymptotic coverage for $\Theta_I$ and $M_I$ in potentially partially identified regular models, and are valid but possibly conservative in potentially partially identified models with reduced-form parameters on the boundary (in which the local tangent space is a convex cone). Our Procedure 1 and 2 CSs are also shown to be uniformly valid over DGPs that include both point- and partially identified models (see Appendix (ref)). Moreover, our Procedure 2 CS is shown to have exact asymptotic coverage for $M_I$ in models with singularities, which are particularly relevant in applications when parameters are close to point-identified or point-identified. Our Procedure 3 CS has exact asymptotic coverage in regular models that are point-identified\footnote{In fact, all three of our procedures are efficient in point-identified regular models.}. Although theoretically slightly conservative in partially identified models, our Procedure 3 CS performs well in our simulations and empirical examples.

Our Procedure 1 and 2 CSs are Monte Carlo (MC) based. To establish their theoretical validity, we derive new Bernstein-von Mises (or Bayesian Wilks) type theorems for the (quasi-)posterior distributions of the QLR and profile QLR in partially identified models, allowing for regular models and some important non-regular cases (e.g. models in which the local tangent space is a convex cone, models with singularities, and models with parameter-dependent support). These theorems establish that the (quasi-)posterior distributions of the QLR and profile QLR converge to their frequentist counterparts in regular models; see Section (ref) and Appendix (ref) for similar results in some non-regular cases. As an illustration we briefly mention some results for Procedure 1 here: Section (ref) presents conditions under which the sample QLR statistic and the (quasi-)posterior distribution of the QLR both converge to a chi-square distribution with unknown degree of freedom in regular models.\footnote{In point-identified models, Wilks-type asymptotics imply the degree of freedom is equal to the dimension of $\theta$ for QLR statistics. In partially identified models, the degree of freedom is some $d^*$, typically less than or equal to $\dim(\theta)$. The correct $d^*$ may not be easy to infer from the context, which is why we refer to it as “unknown”.} Appendix (ref) shows that the QLR and the (quasi-)posterior of the QLR both converge to a gamma distribution with scale parameter of 2 and unknown shape parameter in more general partially-identified models. These results ensure that the quantiles of the QLR evaluated at the MC draws from its quasi-posterior consistently estimate the correct critical values needed for Procedure 1 CS to have exact asymptotic coverage for $\Theta_I$. See Section (ref) for similar results for the profile QLR and Procedure 2 CSs for $M_I$ for subvectors.

We demonstrate the computational feasibility and good finite-sample coverage of our proposed methods in two simulation experiments: a missing data example and a complete information entry game with correlated payoff shocks. We use the missing data example to illustrate the conceptual difficulties in a transparent way, studying both numerically and theoretically the behaviors of our CSs when this model is partially-identified, close to point-identified, and point-identified. Although the length of a confidence interval for the identified set $M_I$ of a scalar $\mu$ is by definition no shorter than that for $\mu$ itself, our simulations demonstrate that the differences in length between our Procedures 2 and 3 CSs for $M_I$ and the GMS CSs of AndrewsSoares for $\mu$ are negligible. Finally, our CS constructions are applied to two real data examples: an airline entry game with correlated payoff shocks and an empirical trade flow model. The airline entry game example has 17 partially-identified structural parameters. Our empirical findings using Procedures 2 and 3 CSs show that the data are informative about some equilibrium selection probabilities. The trade example has 46 structural parameters. Here, point-identification may be difficult to verify, especially when conducting a sensitivity analysis of restrictive model assumptions.

\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Literature Review.} Several papers have recently proposed Bayesian (or pseudo Bayesian) methods for constructing CSs for $\Theta_I$ that have correct frequentist coverage properties. See section 3.3 in 2009 NBER working paper version of MoonSchorfheide, Kitagawa, NoretsTang, KlineTamer, LiaoSimoni and the references therein. All these papers consider separable regular models and use various renderings of a similar intuition. First, there exists a finite-dimensional reduced-form parameter, say $\phi$, that is (globally) point-identified and $\sqrt n$-consistently and asymptotically normal estimable from the data, and is linked to the model structural parameter $\theta$ via a known global mapping. Second, a prior is placed on the reduced-form parameter $\phi$, and third, a classical Bernstein-von Mises theorem stating the asymptotic normality of the posterior distribution for $\phi$ is assumed to hold. Finally, the known global mapping between the reduced-form and the structural parameters is inverted, which, by step 3, guarantees correct coverage for $\Theta_I$ in large samples. In addition to this literature's focus on separable models, it is not clear whether the results there remain valid in various non-regular models we study.

Our approach is valid regardless of whether the model is separable or not. We show that for general separable or non-separable partially identified likelihood or moment-based models, a {\it local reduced-form reparameterization} exists (see Section (ref)). We use this local reparameterization as a proof device to show that the (quasi-)posterior distributions of the QLR and the profile QLR statistics have a frequentist interpretation in large samples. Importantly, since our Procedures 1 and 2 impose priors on the model parameter $\theta$ only, there is no need for obtaining a global reduced-form reparameterization or deriving its dimension to implement our procedures. This is in contrast with the above-mentioned existing Bayesian methods for partially identified separable models, for which researchers need to impose priors on global reduced-form parameters $\phi$ to ensure that its posterior lies on $\{\phi(\theta) : \theta \in \Theta\}$ (i.e. the set of reduced-form parameters consistent with the structural model), which could be difficult even in some empirically relevant separable models; see the airline entry game application in Section (ref). Moreover, our new Bernstein-von Mises type theorems for the (quasi-)posterior distributions of the QLR and profile QLR allow for several important non-regular cases in which the local reduced-form parameter is typically not $\sqrt n$-consistent and asymptotically normally estimable.

When specialized to point- or partially-identified likelihood models, our Procedure 1 CS for $\Theta_I$ is equivalent to Bayesian credible set for $\theta$ based on inverting a LR statistic. With flat priors, these CSs are also the highest posterior density (HPD) credible sets. Our general theoretical results imply that HPD credible sets give correct frequentist coverage in partially identified regular models and conservative coverage in some non-standard circumstances. These findings complement those of MoonSchorfheide who showed that HPD credible sets can under-cover (in a frequentist sense) in separable partially identified regular models under their conditions.\footnote{Note that this is not a contradiction since our priors are imposed on structural parameters $\theta$ only, violating Assumption 2 in MoonSchorfheide.} In point-identified regular models satisfying a generalized information equality with $\sqrt n$-consistent and asymptotically normally estimable parameters $\theta$, CH (CH hereafter) propose constructing CSs for scalar subvectors $\mu$ by taking the upper and lower quantiles of the MCMC draws $\{\mu^1,\ldots,\mu^B\}$ where $(\mu^b,\eta^b) \equiv \theta^b$. Our CS constructions for scalar subvectors are asymptotically equivalent to CH's CSs in such models, but they differ otherwise. Our CS constructions, which are based on quantiles of the {\it criterion} evaluated at the MC draws $\{\theta^1,\ldots,\theta^B\}$ rather than of the raw parameter draws themselves, are valid irrespective of whether the model is point- or partially-identified. Intuitively, this is because the population criterion is always point-identified irrespective of whether $\theta$ is point- or partially-identified.

There are several published works on frequentist CS constructions for $\Theta_I$: see, e.g., CHT and RomanoShaikh where subsampling based methods are used for general partially identified models, Bugni and Armstrong14 where bootstrap methods are used for moment inequality models, and BM where random set methods are used when $\Theta_I$ is strictly convex. For inference on identified sets of subvectors, both the subsampling-based papers of CHT and RomanoShaikh deliver valid tests with a judicious choice of the subsample size for a profiled criterion function. The subsampling-based CS construction allows for general criterion functions, but is computationally demanding and sensitive to choice of subsample size in realistic empirical structural models.\footnote{There is a large literature on frequentist approach for {\it inference on the true parameter} $\theta \in \Theta_I$ or $\mu \in M_I$ (e.g., IM, Rosen, AndrewsGuggenberger, Stoye, AndrewsSoares, andrews2012inference, canay, RSW, bugni/canay/shi:16 and KMS among many others), which generally uses discontinuous-in-parameters asymptotic (repeated sampling) approximations to test statistics. These existing frequentist methods are difficult to implement in realistic empirical models.} Our methods are computationally attractive and typically have asymptotically correct coverage, but require “optimal” criterion functions.

The rest of the paper is organized as follows. Section (ref) describes our new procedures for CSs for identified sets $\Theta_I$ and $M_I$. Section (ref) presents simulations and real data applications. Section (ref) first establishes new BvM (or Bayesian Wilks) results for the QLR and profile QLR in partially identified models. It then derives the frequentist validity of our CSs. Section (ref) provides some sufficient conditions to the key regularity conditions for the general theory in Section (ref). Section (ref) briefly concludes. Appendix (ref) describes the implementation details for the simulations and real data applications in Section (ref). Appendix (ref) shows that our CSs for $\Theta_I$ and $M_I$ are valid uniformly over a class of DGPs. Appendix (ref) verifies the main regularity conditions for uniform validity in the missing data and a moment inequality examples. Appendix (ref) presents results on local power. Appendix (ref) establishes a new BvM (or Bayesian Wilks) result which shows that the limiting (quasi-)posterior distribution of the QLR in a partially identified model is a gamma distribution with unknown shape parameter and scale parameter of 2. There, results on models with parameter-dependent support are given. Appendix (ref) contains all the proofs and additional lemmas.

Description of our Procedures

In this section we first describe our method for constructing CSs for $\Theta_I$. We then describe methods for constructing CSs for $M_I$ of any subvector. We finally present an extremely simple method for constructing CSs for $M_I$ of a scalar subvector in certain situations.

Let $\mf X_n = (X_1,\ldots,X_n)$ denote a sample of i.i.d. or strictly stationary and ergodic data of size $n$. Consider a population objective function $L: \Theta \to \mb R$, such as a log-likelihood function for correctly specified likelihood models, an optimally-weighted or continuously-updated GMM objective function, or a sandwich quasi-likelihood function. The function $L$ is assumed to be an upper semicontinuous function of $\theta$ with $\sup_{\theta \in \Theta} L(\theta) < \infty$. The population objective $L$ may not be maximized uniquely over $\Theta$, but rather its maximizers, the {\it identified set}, may be a nontrivial set of parameters:

equation[equation omitted — 147 chars of source]

The set $ \Theta_I$ is our first object of interest. In many applications, it may be of interest to provide a CS for a {\it subvector} of interest. Write $\theta \equiv (\mu,\eta)$ where $\mu$ is the subvector of interest and $\eta$ is a nuisance parameter. Our second object of interest is the identified set for the subvector $\mu$:

equation[equation omitted — 97 chars of source]

Given the data $\mf X_n$, we seek to construct computationally attractive CSs that cover $\Theta_I$ or $M_I$ with a pre-specified probability (in repeated samples) as sample size $n$ gets large.

To describe our approach, let $L_n$ denote an (upper semicontinuous) sample criterion function that is a jointly measurable function of the data $\mf X_n$ and $\theta$. This objective function $L_n$ can be a natural sample analogue of $L$. We give a few examples of objective functions that we consider.

{\bf Parametric likelihood:} Given a parametric model: $\{ P_\theta : \theta \in \Theta\},$ with a corresponding density $p_\theta(.)$ (with respect to some dominating measure), the identified set is $\Theta_I = \{ \theta \in \Theta: P_0= P_\theta\}$ where $P_0$ is the true data distribution. We take $L_n$ to be the average log-likelihood function:

equation[equation omitted — 90 chars of source]

{\bf GMM models:} Consider a set of {\it moment equalities} $E[\rho_\theta(X_i)] = 0$ such that the solution to this vector of equalities may not be unique. The identified set is $\Theta_I = \{ \theta \in \Theta: \, E[\rho_\theta(X_i)] = 0\}$. The sample objective function $L_n$ can be the continuously-updated GMM objective function:

equation[equation omitted — 101 chars of source]

where $\rho_n(\theta) = \frac{1}{n} \sum_{i=1}^n \rho_\theta(X_i)$ and $W_n(\theta) = \left( \frac{1}{n} \sum_{i=1}^n \rho_\theta(X_i)\rho_\theta(X_i)' - \rho_n(\theta) \rho_n(\theta)' \right)^{-}$ (the superscript $^-$ denotes generalized inverse) for iid data or other suitable choices. Given an optimal weighting matrix $\wh W_n$, we could also use an optimally-weighted GMM objective function:

equation[equation omitted — 99 chars of source]

Generalized empirical likelihood objective functions could also be used with our procedures.

Our main CS constructions (Procedures 1 and 2 below) are based on Monte Carlo (MC) simulation methods from a quasi-posterior. Given $L_n$ and a prior $\Pi$ over $\Theta$, the quasi-posterior distribution $\Pi_n$ for $\theta$ given $\mf X_n$ is defined as

equation[equation omitted — 158 chars of source]

Our procedures 1 and 2 require drawing a sample $\{\theta^1,\ldots,\theta^B\}$ from the quasi-posterior $\Pi_n$. In practice we use an adaptive sequential Monte Carlo (SMC) algorithm which is known to be well suited to drawing from irregular, multi-modal distributions, but any MC sampler could, in principle, be used. The SMC algorithm is described in detail in Appendix (ref).

Confidence sets for the identified set $\Theta_I$

Here we seek a 100$\alpha$% CS $\wh \Theta_{\alpha}$ for $\Theta_I$ using $L_n(\theta)$ that has asymptotically exact coverage, i.e.: \[ \lim_{n \to \infty} \mb P(\Theta_I \subseteq \wh \Theta_{\alpha}) = \alpha\,. \]

\centerline{\it \sc [Procedure 1: Confidence sets for the identified set]}

enumerate• Draw a sample $\{\theta^1,\ldots,\theta^B \}$ from the quasi-posterior distribution $\Pi_n$ in ((ref)). • Calculate the $(1-\alpha)$ quantile of $\{L_n(\theta^1),\ldots,L_n(\theta^B)\}$; call it $\zeta_{n,\alpha}^{mc}$. • Our 100$\alpha$% confidence set for $\Theta_I$ is then: \begin{equation} \wh\Theta_\alpha = \{ \theta \in \Theta : L_n(\theta) \geq \zeta_{n,\alpha}^{mc}\}\,. \end{equation}

Notice that no optimization of $L_n$ itself is required in order to construct $\wh \Theta_\alpha$. Further, an exhaustive grid search over the full parameter space $\Theta$ is not required as the MC draws $\{\theta^1,\ldots,\theta^B\}$ will concentrate around $\Theta_I$ and thereby indicate the regions in $\Theta$ over which to search.

CHT considered inference on the set of minimizers of a {\it nonnegative population criterion function} $Q: \, \Theta \to \mb R_+$ using a sample analogue $Q_n$ of $Q$. Let $\xi_{n,\alpha}$ denote a consistent estimator of the $\alpha$ quantile of $\sup_{\theta \in \Theta_I} Q_n(\theta)$. The 100$\alpha$% CS for $\Theta_I$ at level $\alpha \in (0,1)$ proposed is $\wh \Theta_\alpha^{CHT} = \{ \theta \in \Theta : Q_n(\theta) \leq \xi_{n,\alpha}\}$. In the existing literature, subsampling or bootstrap based methods have been used to compute $\xi_{n, \alpha}$ which can be tedious to implement. Instead, our procedure replaces $\xi_{n,\alpha}$ with a cut off based on Monte Carlo simulations. The next remark provides an equivalent approach to Procedure 1 but that is constructed in terms of $Q_n$, which is the quasi likelihood ratio statistic associated with $L_n$.

remarkLet $\hat \theta \in \Theta$ denote an approximate maximizer of $L_n$, i.e.: \[ L_n(\hat \theta) = \sup_{\theta \in \Theta}L_n(\theta) + o_\mb P(n^{-1})\,. \] and define the quasi-likelihood ratio (QLR) (at a point $\theta \in \Theta$) as: \begin{equation} Q_n(\theta) = 2n [L_n(\hat \theta) - L_n (\theta)]\,. \end{equation} Let $\xi_{n,\alpha}^{mc}$ denote the $\alpha$ quantile of $\{Q_n(\theta^1),\ldots,Q_n(\theta^B)\}$. The confidence set: \[ \wh\Theta_\alpha' = \{ \theta \in \Theta : Q_n(\theta) \leq \xi_{n,\alpha}^{mc}\} \] is equivalent to $\wh \Theta_\alpha$ defined in ((ref)) because $L_n(\theta) \geq \zeta_{n,\alpha}^{mc}$ if and only if $Q_n(\theta) \leq \xi_{n,\alpha}^{mc}$.

In Procedure 1 and Remark (ref) above, the posterior-like quantity involves the use of a prior distribution $\Pi$ over $\Theta$. This prior is user chosen and typically would be the uniform prior but other choices are possible. In our simulations, various choices of prior did not matter much, unless they assigned extremely small mass near the true parameter (which is avoided by using a uniform prior whenever $\Theta$ is compact).

The next lemma presents high-level conditions under which any 100$\alpha$% criterion-based CS for $\Theta_I$ has asymptotically correct (frequentist) coverage. Similar statements appear in CHT. Let $F_W (c):= \Pr(W \leq c )$ denote the (probability) distribution function of a random variable $W$ and $w_{\alpha}:=\inf \{c \in \mb R: F_W (c) \geq \alpha \}$ be the $\alpha$ quantile of $F_W$.

lemmaLet (i) $\sup_{\theta \in \Theta_I} Q_n(\theta) \rightsquigarrow W$ where $W$ is a random variable for which $F_W$ is continuous at $w_\alpha$, and (ii) $(w_{n,\alpha})_{n \in \mb N}$ be a sequence of random variables such that $w_{n,\alpha} \geq w_\alpha + o_\mb P(1)$. Define: \[ \wh \Theta_\alpha = \{ \theta \in \Theta : Q_n(\theta) \leq w_{n,\alpha}\}\,. \] Then: $\liminf_{n \to \infty} \mb P(\Theta_I \subseteq \wh \Theta_\alpha) \geq \alpha$. Moreover, if condition (ii) is replaced by the condition $w_{n,\alpha} = w_\alpha + o_\mb P(1)$, then: $\lim_{n \to \infty} \mb P(\Theta_I \subseteq \wh \Theta_\alpha) = \alpha$.

Our MC CSs for $\Theta_I$ are shown to be valid by verifying parts (i) and (ii) with $w_{n,\alpha} = \xi_{n,\alpha}^{mc}$. To verify part (ii), we shall establish a new Bernstein-von Mises (BvM) (or a new Bayesian Wilks) type result for the quasi-posterior distribution of the QLR under loss of identifiability.

Confidence sets for the identified set $M_I$ of subvectors

We seek a CS $\widehat M_{\alpha}$ for $M_I$ such that: \[ \lim_{n \to \infty} \mb P(M_I \subseteq \wh M_{\alpha}) = \alpha\,. \] A well-known method to construct a CS for $M_I$ is based on projection, which maps a CS $\wh \Theta_\alpha $ for $\Theta_I$ into one for $M_I$. The projection CS:

equation[equation omitted — 132 chars of source]

is a valid $100\alpha$% CS for $M_I$ whenever $\wh \Theta_\alpha$ is a valid $100\alpha$% CS for $\Theta_I$. As is well documented, $\wh M_\alpha^{proj}$ is typically conservative, and especially so when the dimension of $\mu$ is small relative to the dimension of $\theta$. Indeed, our simulations below indicate that $\wh M_\alpha^{proj}$ is very conservative even in reasonably low-dimensional parametric models.

We propose CSs for $M_I$ based on a profile criterion for $M_I$. Let $M = \{\mu: (\mu,\eta) \in \Theta \mbox{ for some } \eta\}$ and $H_\mu = \{ \eta : (\mu,\eta) \in \Theta\}$. The profile criterion for a point $\mu \in M$ is $\sup_{\eta \in H_\mu} L_n(\mu,\eta)$, and the profile criterion for $M_I$ is

equation[equation omitted — 102 chars of source]

Let $\Delta(\theta^b)$ be an equivalence set for $\theta^b$. In likelihood models we define $\Delta(\theta^b) = \{\theta \in \Theta : p_\theta = p_{\theta^b}\}$ and in moment-based models we define $\Delta(\theta^b) = \{ \theta \in \Theta : E[\rho(X_i,\theta)] = E[\rho(X_i,\theta^b)] \}$. Let $M(\theta^b) = \{ \mu : (\mu,\eta) \in \Delta(\theta^b) \mbox{ for some } \eta\}$, and the profile criterion for $M(\theta^b)$ is

equation[equation omitted — 122 chars of source]

\centerline{\it \sc [Procedure 2: Confidence sets for subvectors]}

enumerate• Draw a sample $\{\theta^1,\ldots,\theta^B\}$ from the quasi-posterior distribution $\Pi_n$ in ((ref)). • Calculate the $(1-\alpha)$ quantile of $\big\{ PL_n(M(\theta^b)) : b = 1,\ldots,B \big\}$; call it $\zeta_{n,\alpha}^{mc,p}$. • Our 100$\alpha$% confidence set for $M_I$ is then: \begin{equation} \wh M_\alpha = \Big\{ \mu \in M : \sup_{\eta \in H_\mu} L_n(\mu,\eta) \geq \zeta_{n,\alpha}^{mc,p} \Big\} \,. \end{equation}

By forming $\wh M_\alpha$ in terms of the profile criterion we avoid having to do an exhaustive grid search over $\Theta$. An additional computational advantage is that the subvectors of the draws, say $\{\mu^1,\ldots,\mu^B\}$, concentrate around $M_I$, thereby indicating the region in $M$ over which to search.

remarkRecall the definition of the QLR $Q_n$ in ((ref)), we define the profile QLR for the set $M(\theta^b)$ analogously as \begin{equation} PQ_n(M(\theta^b)) \equiv 2n [L_n(\hat \theta) - PL_n(M(\theta^b))] \;=\; \sup_{\mu \in M(\theta^b)} \inf_{\eta \in H_\mu} Q_n (\mu, \eta)\,. \end{equation} Let $\xi_{n,\alpha}^{mc,p}$ denote the $\alpha$ quantile of the profile QLR draws $\big\{PQ_n(M(\theta^b)) : b=1,\ldots,B\big\}$. The confidence set: \[ \wh M_\alpha' = \Big\{ \mu \in M : \inf_{\eta \in H_\mu} Q_n(\mu,\eta) \leq \xi_{n,\alpha}^{mc,p}\Big\} \] is equivalent to $\wh M_\alpha$ because $\sup_{\eta \in H_\mu} L_n(\mu,\eta) \geq \zeta_{n,\alpha}^{mc,p}$ if and only if $\inf_{\eta \in H_\mu} Q_n(\mu,\eta) \leq \xi_{n,\alpha}^{mc,p}$.

Our Procedure 2 and Remark (ref) above are different from taking quantiles of the MC parameter draws. A popular percentile CS (denoted as $\widehat{M}_{\alpha}^{perc}$) for a scalar subvector $\mu$ is computed by taking the upper and lower $100(1-\alpha)/2$ percentiles of $\{\mu^1,\ldots,\mu^B\}$. For point-identified regular models with $\sqrt n$-consistent and asymptotically normal parameters $\theta$, this approach is known to be valid for correctly-specified likelihood models in the standard Bayesian literature and its validity for criterion-based models satisfying a generalized information equality has been established by CH. However, in partially identified models this approach is no longer valid and under-covers, as evidenced in the simulation results below.

The following result presents high-level conditions under which any 100$\alpha$% criterion-based CS for $M_I$ is asymptotically valid. A similar statement appears in RomanoShaikh.

lemmaLet (i) $\sup_{\mu \in M_I} \inf_{\eta \in H_\mu} Q_n(\mu,\eta) \rightsquigarrow W$ where $W$ is a random variable for which $F_W$ is continuous at $w_\alpha$, and (ii) $(w_{n,\alpha})_{n \in \mb N}$ be a sequence of random variables such that $w_{n,\alpha} \geq w_\alpha + o_\mb P(1)$. Define: \[ \wh M_\alpha =\Big\{ \mu \in M : \inf_{\eta \in H_\mu} Q_n(\mu,\eta) \leq w_{n,\alpha}\Big\}\,. \] Then: $\liminf_{n \to \infty} \mb P(M_I \subseteq \wh M_\alpha) \geq \alpha$. Moreover, if condition (ii) is replaced by the condition $w_{n,\alpha} = w_\alpha + o_\mb P(1)$, then: $\lim_{n \to \infty} \mb P(M_I \subseteq \wh M_\alpha) = \alpha$.

Our MC CSs for $M_I$ are shown to be valid by verifying parts (i) and (ii) with $w_{n,\alpha} = \xi_{n,\alpha}^{mc,p}$. To verify part (ii), we shall derive a new BvM type result for the quasi-posterior of the profile QLR under loss of identifiability.

A simple but slightly conservative CS for $M_I$ of scalar subvectors

For a class of partially identified models with one-dimensional subvectors of interest, we now propose another CS $\wh M_\alpha^\chi$ which is extremely simple to construct. This new CS for $M_I$ is slightly conservative (whereas $\wh M_\alpha$ could be asymptotically exact), but its coverage is much less conservative than that of the projection-based CS $\widehat{M}_{\alpha}^{proj}$.

\centerline{\it \sc [Procedure 3: Simple conservative CSs for scalar subvectors]}

enumerate• Calculate a maximizer $\hat \theta$ for which $L_n(\hat \theta) \geq \sup_{\theta \in \Theta} L_n(\theta) + o_\mb P(n^{-1})$. • Our 100$\alpha$% confidence set for $M_I \subset \mb R$ is then: \begin{equation} \wh M_\alpha^\chi = \Big\{ \mu \in M : \inf_{\eta \in H_\mu} Q_n(\mu,\eta) \leq \chi^2_{1,\alpha} \Big\} \end{equation} where $Q_n$ is the QLR in ((ref)) and $\chi^2_{1,\alpha}$ denotes the $\alpha$ quantile of the $\chi^2_1$ distribution.

Procedure 3 above is justified when the limit distribution of the profile QLR for $M_I$ is stochastically dominated by the $\chi^2_1$ distribution (i.e., $F_W (z)\geq F_{\chi^2_1} (z)$ for all $z \geq 0 $ in Lemma (ref)). This allows for computationally simple construction using repeated evaluations on a scalar grid. Unlike $\wh M_\alpha$, the CS $\wh M_\alpha^\chi$ for $M_I$ is typically asymptotically conservative and is only valid for scalar functions of $\Theta_I$ (see Section (ref)). Nevertheless, the CS $\wh M_\alpha^\chi$ is asymptotically exact when $M_I$ happens to be a singleton belonging to the interior of $M$, and, for confidence levels of $\alpha \geq 0.85$, its degree of conservativeness for the set $M_I$ is negligible (see Section (ref)). It is extremely simple to implement and performs very favorably in simulations. As a sensitivity check in empirical estimation of a complicated structural model, one could report the conventional CS based on a $t$-statistic (that is valid under point identification only) as well as our CS $\wh M_\alpha^\chi$ (that remains valid under partial identification); see Section (ref).

Simulation Evidence and Empirical Applications

This section presents simulation evidence and empirical applications to demonstrate the good performances of our new procedures for general possibly partially identified models. See Appendix (ref) for implementation details.

Simulation evidence

In this subsection we investigate the finite-sample behavior of our proposed CSs in two leading examples of partially identified models: missing data and entry game with correlated payoff shocks. Both have been studied in the existing literature as leading examples of partially-identified moment inequality models; we instead use them as examples of likelihood and moment {\it equality} models.

We use samples of size $n = 100$, $250$, $500$, and $1000$. For each sample, we calculate the posterior quantile of the QLR or profile QLR statistic using $B=10000$ draws from an adaptive SMC algorithm (see Appendix (ref) for a description of the algorithm).

Example 1: missing data

We first consider the simple but insightful missing data example. Suppose we observe a random sample $\{(D_i,Y_iD_i)\}_{i=1}^n$ where both the outcome variable $Y_i$ and the selection variable $D_i$ take values in $\{0,1\}$. The parameter of interest is the true mean $\mu_0 = \mb E[Y_i]$. Without further assumptions, $\mu_0$ is not point identified when $\Pr(D_i = 0) > 0$ as we only observe $Y_i$ when $D_i=1$.

Denote the true probabilities of observing $(D_i,Y_iD_i) = (1,1)$, $(0,0)$ and $(1,0)$ by $\tilde \gamma_{11}$, $\tilde \gamma_{00}$, and $\tilde \gamma_{10} = 1-\tilde \gamma_{11} - \tilde \gamma_{00}$ respectively. We view $\tilde \gamma_{00}$ and $\tilde \gamma_{11}$ as true {\it reduced-form parameters} that are consistently estimable. The reduced-form parameters are functions of the structural parameter $\theta = (\mu, \eta_1, \eta_2)$ where $\mu = \mb E[Y_i]$, $\eta_1 = \Pr(Y_i = 1|D_i =0)$, and $\eta_2 = \Pr(D_i = 1)$. Under this model parameterization, $\theta$ is related to the reduced form parameters via $\tilde \gamma_{00} (\theta ) = 1-\eta_2$ and $\tilde \gamma_{11} (\theta ) = \mu - \eta_1 (1- \eta_2)$. The parameter space $\Theta$ for $\theta$ is defined as:

equation[equation omitted — 137 chars of source]

The identified set for $\theta$ is:

equation[equation omitted — 161 chars of source]

Here, $\eta_2$ is point-identified but only an affine combination of $\mu$ and $\eta_1$ are identified. The identified set for $\mu = E[Y_i]$ is: \[ M_I = [\tilde \gamma_{11},\tilde \gamma_{11}+\tilde \gamma_{00}] \] and the identified set for the nuisance parameter $\eta_1$ is $[0,1]$.

We set the true values of the parameters to be $\mu = 0.5$, $\eta_1 = 0.5$, and take $\eta_2 = 1-c/\sqrt n$ for $c = 0,1,2$ to cover both partially-identified but “drifting-to-point-identification” ($c = 1,2$) and point-identified $(c = 0)$ cases. We first implement the procedures using a likelihood criterion and a flat prior on $\Theta$. The likelihood function of $(D_i,Y_iD_i)=(d,yd)$ is

align*[align* omitted — 173 chars of source]

In Appendix (ref) we present and discuss additional results for a likelihood criterion with a curved prior and a continuously-updated GMM criterion based on the moments $E [ 1\!\mathrm{l}\{D_i = 0 \} - \tilde \gamma_{00} (\theta)] = 0$ and $E [ 1\!\mathrm{l}\{(D_i,Y_iD_i) = (1,1) \} - \tilde \gamma_{11} (\theta) ] = 0$ with a flat prior (this GMM case may be interpreted as a moment inequality model with $\eta_1(1-\eta_2)$ playing the role of a slackness parameter).

We implement the SMC algorithm as described in Appendix (ref). To illustrate sampling via the SMC algorithm and the resulting posterior of the QLR, Figure (ref) displays histograms of the draws for $\mu$, $\eta_1$ and $\eta_2$ for one run of the adaptive SMC algorithm for a sample of size $1000$ with $\eta_2 = 0.8$. Here $\mu$ is partially identified with $M_I = [0.4,0.6]$. The histograms in Figure (ref) show that the draws for $\mu$ and $\eta_1$ are both approximately flat across their identified sets. In contrast, the draws for $\eta_2$, which is point identified, are approximately normally distributed and centered at the MLE. The Q-Q plot in Figure (ref) shows that the quantiles of $Q_n(\theta)$ computed from the draws are very close to the quantiles of a $\chi^2_2$ distribution, as predicted by our theoretical results below.

figure[figure omitted — 936 chars of source]
sidewaystable[p]{\begin{center} \begin{tabular}{|c|cccccc|cccccc|cccccc|} \hline & \multicolumn{6}{c|}{$\eta_2 = 1-\frac{2}{\sqrt n}$} & \multicolumn{6}{c|}{$\eta_2 = 1-\frac{1}{\sqrt n}$}& \multicolumn{6}{c|}{$\eta_2 = 1$ (Point ID)} \\ & \multicolumn{2}{c}{0.90} & \multicolumn{2}{c}{0.95} & \multicolumn{2}{c|}{0.99} & \multicolumn{2}{c}{0.90} & \multicolumn{2}{c}{0.95} & \multicolumn{2}{c|}{0.99} & \multicolumn{2}{c}{0.90} & \multicolumn{2}{c}{0.95} & \multicolumn{2}{c|}{0.99} \\ \hline & & \multicolumn{16}{c}{$\widehat{\Theta}_{\alpha}$ (Procedure 1)} & \\ 100 & .910 & --- & .957 & --- & .994 & --- & .903 & --- & .953 & --- & .993 & --- & .989 & --- & .997 & --- & 1.000 & --- \\ 250 & .901 & --- & .947 & --- & .991 & --- & .912 & --- & .955 & --- & .992 & --- & .992 & --- & .997 & --- & 1.000 & --- \\ 500 & .913 & --- & .956 & --- & .991 & --- & .908 & --- & .957 & --- & .991 & --- & .995 & --- & .997 & --- & \phantom{0}.999 & --- \\ 1000 & .910 & --- & .958 & --- & .992 & --- & .911 & --- & .958 & --- & .994 & --- & .997 & --- & .999 & --- & 1.000 & --- \\ & & \multicolumn{16}{c}{$\widehat{M}_{\alpha}$ (Procedure 2)} & \\ 100 & .920 & [$ .32 ,\! .68 $] & .969 & [$ .30 ,\! .70 $] & .997 & [$ .27 ,\! .73 $] & .918 & [$ .37 ,\! .63 $] & .964 & [$ .35 ,\! .65 $] & .994 & [$ .32 ,\! .68 $] & .911 & [$ .42 ,\! .59 $] & .958 & [$ .40 ,\! .60 $] & \phantom{0}.990 & [$ .37 ,\! .63 $] \\ 250 & .917 & [$ .39 ,\! .61 $] & .961 & [$ .38 ,\! .62 $] & .992 & [$ .36 ,\! .64 $] & .920 & [$ .42 ,\! .58 $] & .963 & [$ .41 ,\! .59 $] & .991 & [$ .39 ,\! .61 $] & .915 & [$ .45 ,\! .55 $] & .959 & [$ .44 ,\! .56 $] & \phantom{0}.991 & [$ .42 ,\! .58 $] \\ 500 & .914 & [$ .42 ,\! .58 $] & .961 & [$ .41 ,\! .59 $] & .993 & [$ .40 ,\! .60 $] & .914 & [$ .44 ,\! .56 $] & .958 & [$ .43 ,\! .57 $] & .992 & [$ .42 ,\! .58 $] & .916 & [$ .46 ,\! .54 $] & .959 & [$ .46 ,\! .54 $] & \phantom{0}.990 & [$ .44 ,\! .56 $] \\ 1000 & .917 & [$ .44 ,\! .56 $] & .956 & [$ .44 ,\! .56 $] & .993 & [$ .43 ,\! .57 $] & .914 & [$ .46 ,\! .54 $] & .955 & [$ .45 ,\! .55 $] & .993 & [$ .44 ,\! .56 $] & .916 & [$ .47 ,\! .53 $] & .959 & [$ .47 ,\! .53 $] & \phantom{0}.992 & [$ .46 ,\! .54 $] \\ & & \multicolumn{16}{c}{$\widehat{M}^{\chi}_{\alpha}$ (Procedure 3)} & \\ 100 & .920 & [$ .32 ,\! .68 $] & .952 & [$ .31 ,\! .69 $] & .990 & [$ .28 ,\! .72 $] & .916 & [$ .37 ,\! .63 $] & .946 & [$ .36 ,\! .64 $] & .989 & [$ .33 ,\! .67 $] & .902 & [$ .42 ,\! .58 $] & .937 & [$ .41 ,\! .60 $] & \phantom{0}.986 & [$ .38 ,\! .63 $] \\ 250 & .915 & [$ .39 ,\! .61 $] & .952 & [$ .38 ,\! .62 $] & .990 & [$ .36 ,\! .64 $] & .914 & [$ .42 ,\! .58 $] & .954 & [$ .41 ,\! .59 $] & .990 & [$ .39 ,\! .61 $] & .883 & [$ .45 ,\! .55 $] & .949 & [$ .44 ,\! .56 $] & \phantom{0}.991 & [$ .42 ,\! .58 $] \\ 500 & .894 & [$ .42 ,\! .58 $] & .954 & [$ .41 ,\! .59 $] & .989 & [$ .40 ,\! .60 $] & .906 & [$ .44 ,\! .56 $] & .949 & [$ .44 ,\! .56 $] & .990 & [$ .42 ,\! .58 $] & .899 & [$ .46 ,\! .54 $] & .945 & [$ .46 ,\! .54 $] & \phantom{0}.988 & [$ .44 ,\! .56 $] \\ 1000 & .909 & [$ .44 ,\! .56 $] & .950 & [$ .44 ,\! .56 $] & .993 & [$ .43 ,\! .57 $] & .904 & [$ .46 ,\! .54 $] & .954 & [$ .45 ,\! .55 $] & .989 & [$ .45 ,\! .55 $] & .906 & [$ .48 ,\! .52 $] & .946 & [$ .47 ,\! .53 $] & \phantom{0}.991 & [$ .46 ,\! .54 $] \\ & & \multicolumn{16}{c}{$\widehat{M}^{proj}_{\alpha}$ (Projection)} & \\ 100 & .972 & [$ .30 ,\! .70 $] & .990 & [$ .28 ,\! .71 $] & .999 & [$ .25 ,\! .75 $] & .969 & [$ .35 ,\! .65 $] & .989 & [$ .33 ,\! .67 $] & .998 & [$ .30 ,\! .70 $] & .989 & [$ .37 ,\! .63 $] & .997 & [$ .36 ,\! .64 $] & 1.000 & [$ .33 ,\! .67 $] \\ 250 & .971 & [$ .37 ,\! .63 $] & .986 & [$ .36 ,\! .64 $] & .998 & [$ .34 ,\! .66 $] & .976 & [$ .40 ,\! .60 $] & .988 & [$ .39 ,\! .61 $] & .998 & [$ .37 ,\! .63 $] & .992 & [$ .42 ,\! .58 $] & .997 & [$ .41 ,\! .59 $] & 1.000 & [$ .39 ,\! .61 $] \\ 500 & .972 & [$ .41 ,\! .59 $] & .985 & [$ .40 ,\! .60 $] & .999 & [$ .39 ,\! .61 $] & .972 & [$ .43 ,\! .57 $] & .989 & [$ .42 ,\! .58 $] & .999 & [$ .41 ,\! .59 $] & .995 & [$ .44 ,\! .56 $] & .997 & [$ .43 ,\! .57 $] & \phantom{0}.999 & [$ .42 ,\! .58 $] \\ 1000 & .973 & [$ .44 ,\! .56 $] & .990 & [$ .43 ,\! .57 $] & .999 & [$ .42 ,\! .58 $] & .973 & [$ .45 ,\! .55 $] & .988 & [$ .45 ,\! .55 $] & .999 & [$ .44 ,\! .56 $] & .997 & [$ .45 ,\! .55 $] & .999 & [$ .45 ,\! .55 $] & 1.000 & [$ .44 ,\! .56 $] \\ & & \multicolumn{16}{c}{$\widehat{M}^{perc}_{\alpha}$ (Percentile)} & \\ 100 & .416 & [$ .38 ,\! .62 $] & .676 & [$ .36 ,\! .64 $] & .945 & [$ .32 ,\! .68 $] & .661 & [$ .40 ,\! .59 $] & .822 & [$ .39 ,\! .61 $] & .963 & [$ .35 ,\! .65 $] & .896 & [$ .42 ,\! .58 $] & .946 & [$ .40 ,\! .60 $] & \phantom{0}.989 & [$ .38 ,\! .63 $] \\ 250 & .402 & [$ .42 ,\! .58 $] & .669 & [$ .41 ,\! .59 $] & .917 & [$ .38 ,\! .62 $] & .662 & [$ .44 ,\! .56 $] & .822 & [$ .43 ,\! .57 $] & .960 & [$ .41 ,\! .59 $] & .899 & [$ .45 ,\! .55 $] & .950 & [$ .44 ,\! .56 $] & \phantom{0}.990 & [$ .42 ,\! .58 $] \\ 500 & .400 & [$ .44 ,\! .56 $] & .652 & [$ .43 ,\! .57 $] & .914 & [$ .42 ,\! .58 $] & .652 & [$ .46 ,\! .54 $] & .812 & [$ .45 ,\! .55 $] & .955 & [$ .43 ,\! .57 $] & .903 & [$ .46 ,\! .54 $] & .953 & [$ .46 ,\! .54 $] & \phantom{0}.988 & [$ .44 ,\! .56 $] \\ 1000 & .405 & [$ .46 ,\! .54 $] & .671 & [$ .45 ,\! .55 $] & .917 & [$ .44 ,\! .56 $] & .662 & [$ .47 ,\! .53 $] & .819 & [$ .46 ,\! .54 $] & .953 & [$ .45 ,\! .55 $] & .905 & [$ .47 ,\! .53 $] & .953 & [$ .47 ,\! .53 $] & \phantom{0}.990 & [$ .46 ,\! .54 $] \\ & & \multicolumn{16}{c}{Comparison with GMS CSs for $\mu$ via moment inequalities} & \\ 100 & .815 & [$ .34 ,\! .66 $] & .908 & [$ .32 ,\! .68 $] & .981 & [$ .29 ,\! .71 $] & .803 & [$ .39 ,\! .61 $] & .904 & [$ .37 ,\! .63 $] & .980 & [$ .34 ,\! .66 $] & .889 & [$ .42 ,\! .58 $] & .938 & [$ .40 ,\! .60 $] & \phantom{0}.973 & [$ .39 ,\! .62 $] \\ 250 & .798 & [$ .40 ,\! .60 $] & .899 & [$ .39 ,\! .61 $] & .979 & [$ .36 ,\! .63 $] & .811 & [$ .43 ,\! .57 $] & .897 & [$ .42 ,\! .58 $] & .980 & [$ .40 ,\! .60 $] & .896 & [$ .45 ,\! .55 $] & .944 & [$ .44 ,\! .56 $] & \phantom{0}.981 & [$ .42 ,\! .57 $] \\ 500 & .794 & [$ .43 ,\! .57 $] & .898 & [$ .42 ,\! .58 $] & .976 & [$ .40 ,\! .60 $] & .789 & [$ .45 ,\! .55 $] & .892 & [$ .44 ,\! .56 $] & .975 & [$ .43 ,\! .57 $] & .897 & [$ .46 ,\! .54 $] & .948 & [$ .46 ,\! .54 $] & \phantom{0}.986 & [$ .45 ,\! .55 $] \\ 1000 & .802 & [$ .45 ,\! .55 $] & .900 & [$ .44 ,\! .56 $] & .978 & [$ .43 ,\! .57 $] & .812 & [$ .46 ,\! .54 $] & .900 & [$ .46 ,\! .54 $] & .978 & [$ .45 ,\! .55 $] & .898 & [$ .47 ,\! .53 $] & .949 & [$ .47 ,\! .53 $] & \phantom{0}.990 & [$ .46 ,\! .54 $] \\ \hline \end{tabular} \parbox{14cm}{\caption{ Missing data example: average coverage probabilities for $\Theta_I$ and $M_I$ and average lower and upper bounds of CSs for $M_I$ across 5000 MC replications. Procedures 1--3, Projection and Percentile are implemented using a likelihood criterion and flat prior.}} \end{center} }

\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Confidence sets for $\Theta_I$:}

The top panel of Table (ref) displays MC coverage probabilities of $\wh \Theta_\alpha$ for 5000 replications. The MC coverage probability should be equal to its nominal value in large samples when $\eta_2 < 1$ (see Theorem (ref)). It is perhaps surprising that the nominal and MC coverage probabilities are close even in samples as small as $n = 100$. When $\eta_2 = 1$ the CSs for $\Theta_I$ are conservative, as predicted by our theoretical results (see Theorem (ref)).

\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Confidence sets for $M_I$:}

We now consider various CSs for the identified set $M_I$ for $\mu$. We first compute the projection CS $\wh M_\alpha^{proj}$, as defined in ((ref)), for $M_I$. As we can see from Table (ref), this results in conservative CSs for $M_I$. For example, when $\alpha = 0.90$ the projection CSs cover $M_I$ in around 97% of repeated samples. As the models with $c = 1,2$ are close to point-identified, one might be tempted to report simple percentile CSs $\widehat{M}_{\alpha}^{perc}$ for $M_I$ using CH procedure, which is valid under point identification, and taking the upper and lower $100(1-\alpha)/2$ quantiles from of the draws for $\mu$.\footnote{Note that we use exactly the same draws for implementing the percentile CS and procedures 1 and 2. As the SMC algorithm uses a particle approximation to the posterior, in practice we compute posterior quantiles for $\mu$ using the particle weights in a manner similar to ((ref)).} The results in Table (ref) show that $\wh M_\alpha^{perc}$ has correct coverage when $\mu$ is point identified (i.e. $\eta_2 =1$) but it under-covers when $\mu$ is not point identified. For instance, the coverage probabilities of 90% CSs for $M_I$ are about 66% with $c = 1$.

In contrast, our criterion-based procedures 2 and 3 remain valid under partial identification. We show below (see Theorem (ref)) that the coverage probabilities of our procedure 2 CS $\wh M_\alpha$ (for $M_I$) should be equal to their nominal values $\alpha$ when $n$ is large irrespective of whether the model is partially identified with (i.e. $\eta_2 < 1$) or point identified (i.e. $\eta_2 = 1$). The results in Table (ref) show that this is indeed the case, and that the coverage probabilities for procedure 2 are close to their nominal level even for small values of $n$, irrespective of whether the model is point- or partially-identified. In Section (ref), we show that the asymptotic distribution of the profile QLR for $M_I$ is stochastically dominated by the $\chi^2_1$ distribution. Table (ref) also presents results for procedure 3 using $\wh M_\alpha^\chi$ as in ((ref)). As we can see from these tables, the coverage results look remarkably close to their nominal values even for small sample sizes and for all values of $\eta_2$.

Finally, we compare our the length of CSs for $M_I$ using procedures 2 and 3 with the length of CSs for the parameter $\mu$ constructed using the generalized moment selection (GMS) procedure of AndrewsSoares. We implement their procedure using the inequalities

align[align omitted — 97 chars of source]

with their smoothing parameter $\kappa_n = (\log n)^{1/2}$, their GMS function $\varphi_j^{(1)}$, and with critical values computed via a multiplier bootstrap. Of course, GMS CSs are for the parameter $\mu$ rather than the set $M_I$, which is why the coverage for $M_I$ reported in Table (ref) appears lower than nominal under partial identification (GMS CSs are known to be asymptotically valid CSs for $\mu$). Importantly, the average lower and upper bounds of our CSs for $M_I$ constructed using Procedures 2 and 3 are very close to those using GMS, whereas projection-based CSs are, in turn, larger. On the other hand, CSs computed using percentiles of the draws for $\mu$ are narrower.

Example 2: entry game with correlated payoff shocks

We now consider the complete information entry game example described in Table (ref). We assume that $(\epsilon_1, \epsilon_2)$, observed by the players, are jointly normally distributed with variance 1 and correlation $\rho$, an important parameter of interest. We also assume that $\Delta_1$ and $\Delta_2$ are both negative and that players play a pure strategy Nash equilibrium. When $-\beta_j \leq \epsilon_j \leq -\beta_j -\Delta_j$, $j=1,2$, the game has two equilibria: for given values of the epsilons in this region, the model predicts $(1,0)$ {\it and} $(0,1)$. Let $D_{a_1 a_2}$ denote a binary random variable taking the value $1$ if and only if player 1 takes action $a_1$ and player 2 takes action $a_2$. We observe a random sample of $\{(D_{00,i},D_{10,i},D_{01,i},D_{11,i})\}_{i=1}^n$. So the data provides information of four choice probabilities $(P(0,0), P(1,0), P(0,1), P(1,1))$, but there are six parameters that need to be estimated: $\theta = (\beta_1, \beta_2, \Delta_1, \Delta_1, \rho, s)$ where $s \in [0,1]$ is the equilibrium selection probability. The model parameter is partially identified as we have 3 non-redundant choice probabilities from which we need to learn about 6 parameters.

table[table omitted — 748 chars of source]

We can link the choice probabilities (reduced-form parameters) to $\theta$ via:

align*[align* omitted — 555 chars of source]

and $\tilde \gamma_{01}(\theta) = 1-\tilde \gamma_{00}(\theta) - \tilde \gamma_{11}(\theta) - \tilde \gamma_{10}(\theta)$, where $Q_\rho$ denotes the joint probability distribution of $(\epsilon_1,\epsilon_2)$ indexed by the correlation parameter $\rho$. Let $(\tilde \gamma_{00}, \tilde \gamma_{10}, \tilde \gamma_{01}, \tilde \gamma_{11})$ denote the true choice probabilities $(P(0,0), P(1,0), P(0,1), P(1,1))$. This naturally suggests a likelihood approach, where the likelihood of $(D_{00,i},D_{10,i},D_{11,i},D_{01,i})=(d_{00},d_{10},d_{11},1-d_{00}-d_{10}-d_{11})$ is: \[ p_\theta(d_{00},d_{10},d_{11}) = [\tilde \gamma_{00}(\theta)]^{d_{00}}[\tilde \gamma_{10}(\theta)]^{d_{10}}[\tilde \gamma_{11}(\theta)]^{d_{11}}[1-\tilde \gamma_{00}(\theta)-\tilde \gamma_{10}(\theta)-\tilde \gamma_{11}(\theta)]^{1-d_{00}-d_{10}-d_{11}} \,. \] In the simulations, we use a likelihood criterion with parameter space: \[ \Theta = \{ (\beta_1, \beta_2, \Delta_1, \Delta_2, \rho,s) \in \mb R^6 : (\beta_1,\beta_2) \in [-1,2]^2, \; (\Delta_1,\Delta_2) \in [-2,0]^2, \; (\rho,s) \in [0,1]^2 \}\,. \] We simulate the data using $\beta_1 = \beta_2 = 0.2$, $\Delta_1 = \Delta_2 = -0.5$, $\rho = 0.5$ and $s = 0.5$.

figure[figure omitted — 1,171 chars of source]

We put a flat prior on $\Theta$ and implement the SMC algorithm as described in Appendix (ref). Figure (ref) displays histograms of the marginal draws for $\Delta_1$, $\beta_1$, $s$ and $\rho$ for one run of the SMC algorithm with a sample of size $n = 1000$. The plots for $\Delta_2$ and $\beta_2$ are very similar to those for $\Delta_1$ and $\beta_1$ (which is to be expected as the parameters are symmetric) and are therefore omitted. The draws for $\Delta_1$ and $\beta_1$ are supported on and around their respective identified sets, which are approximately $[-1.42,0]$ and $[-0.05,0.66]$ (the identified sets for $\rho$ and $s$ are $[0,1]$). Note that here the draws for $\Delta_1$, $\beta_1$ and $\rho$ are not flat over their identified sets, in contrast with the draws for $\mu$ and $\eta_1$ in Figure (ref). To see why, consider the marginal identified set for $(\Delta_1,\rho)$, plotted as the shaded region in Figure (ref). This plot shows that when $\Delta_1$ is close to the upper bound of its identified set, $(\Delta_1,\rho)$ is in the shaded region for any $\rho \in [0,1]$. However, when $\Delta_1$ is close to the lower bound of its identified set, $(\Delta_1,\rho)$ is only in the shaded region for very large values of $\rho$. This structure of $\Theta_I$, together with the flat prior on $\Theta$, means that the marginal posterior for $\Delta_1$ assigns relatively more mass towards the upper limit of the identified set for $\Delta_1$. Similar logic applies for $\beta_1$ and $\rho$. Figure (ref) also shows that the quantiles of $Q_n(\theta)$ computed from the draws are very close to the $\chi^2_3$ quantiles, as predicted by our theoretical results below.

Table (ref) reports average coverage probabilities and CS limits for the various procedures across 1000 replications. We form CSs for $\Theta_I$ using procedure 1, as well as CSs for the identified sets of scalar subvectors $\Delta_1$ and $\beta_1$ using procedures 2 and 3.\footnote{As the parameterization is symmetric, the identified sets for $\Delta_2$ and $\beta_2$ are the same as for $\Delta_1$ and $\beta_1$ so we omit them. We also omit CSs for $\rho$ and $s$, whose identified sets are both $[0,1]$.} We also compare our CS for identified sets for $\Delta_1$ and $\beta_1$ with projection-based and percentile-based CSs. Appendix (ref) provides additional details on computation of $M(\theta)$ for implementation of procedure 2. We do not use the reduced-form reparameterization in terms of choice probabilities to compute $M(\theta)$. Coverage of $\wh \Theta_\alpha$ for $\Theta_I$ is extremely good, even with the small sample size $n = 100$. Coverage of procedures 2 and 3 for the identified sets for $\Delta_1$ and $\beta_1$ is slightly conservative for the small sample size $n$, but close to nominal for $n =1000$. As expected, projection CSs are valid but very conservative (the coverage probabilities of 90% CSs are all at least 98%) whereas percentile-based CSs undercover.

table[table omitted — 5,379 chars of source]

Empirical applications

This subsection implements our procedures in two non-trivial empirical applications. The first application estimates an entry game with correlated payoff shocks using data from the US airline industry. Here there are 17 model parameters to be estimated. The second application estimates a model of trade flows initially examined in HMR (HMR henceforth). We use a version of the empirical model in HMR with 46 parameters to be estimated.

Although the entry game model is separable, we do not make use of separability in implementing our procedures. In fact, the existing Bayesian approaches that impose priors on the globally-identified reduced-form parameters will be problematic in this example. This separable model has 24 non-redundant choice probabilities (global reduced-form parameters, i.e., $\dim (\phi ) =24$) and 17 model structural parameters (i.e., $\dim (\theta ) =17$), and there is no explicit closed form expression for the identified set. Both MoonSchorfheide and KlineTamer would sample from the posterior for the reduced-form parameter $\phi$. But, unless the posterior for $\phi$ is constrained to lie on $\{\phi(\theta) : \theta \in \Theta\}$ (i.e. the set of reduced-form probabilities consistent with the model, rather than the full 24-dimensional space), certain values of $\phi$ drawn from their posteriors for $\phi$ will not be consistent with the model.

The empirical trade example is a {\it nonseparable} likelihood model that cannot be handled by either (a) existing Bayesian approaches that rely on a point-identified, $\sqrt n$-estimable and asymptotically normal reduced-form parameter, or (b) inference procedures based on moment inequalities.

In both applications, our approach only puts a prior on the model structural parameter $\theta$ so it does not matter whether the model is separable or not. Both applications illustrate how our procedures may be used to examine the robustness of estimates to various ad hoc modeling assumptions in a theoretically valid and computationally feasible way.

Bivariate Entry Game with US Airline Data

This section estimates a version of the entry game that we study in Subsection (ref) above. We use data from the second quarter of $2010$'s Airline Origin and Destination Survey (DB1B) to estimate a binary game where the payoff for firm $i$ from entering market $m$ is \[ \beta_i + \beta_i^xx_{im} + \Delta_iy_{3-i} + \epsilon_{im} \quad i=1,2 \] where the $\Delta_i$ are assumed to be negative (as usually the case in entry models). The data contain 7882 markets which are formally defined as trips between two airports irrespective of stopping and we examine the entry behavior of two kinds of firms: LC (low cost) firms,\footnote{The low cost carriers are: JetBLue, Frontier, Air Tran, Allegiant Air, Spirit, Sun Country, USA3000, Virgin America, Midwest Air, and Southwest.} and OA (other airlines) which includes all the other firms. The unconditional choice probabilities are $(.16, .61, .07, .15)$ which are respectively the probabilities that OA and LC serve a market, that OA and not LC serve a market, that LC and not OA serve a market, and finally whether no airline serve the market.

The regressors are {\it market presence} and {\it market size}. Market presence is a market- and airline-specific variable defined as follows: from a given airport, we compute the ratio of markets a given carrier (we take the maximum within the category OA or LC, as appropriate) serves divided by the total number of markets served from that given airport. The market presence variable $MP$ is the average of the ratios from the two endpoints and it provides a proxy for an airline's presence in a given airport (See berry for more on this variable). This variable acts as an excluded regressor: the market presence for OA only enters OA's payoffs, so $MP$ is both market- and airline-specific. The second regressor we use is {\it market size} $MS$ which is defined as the population at the endpoints, so this variable is market-specific. We discretize both $MP$ and $MS$ into binary variables that take the value of one if the variable is higher than its median (in the data) value and zero otherwise. The choice probabilities are $P(y_{OA}, y_{LC}|MS, \, MP_{OA},\, MP_{LC})$ are conditional on the three-dimensional vector $(MS, MP_{OA}, MP_{LC})$. We therefore have 4 choice probabilities for every value of the conditioning variables (and there are 8 values for these).\footnote{With binary values, the conditioning set takes the following eight values: (1,1,1), (1,1,0), (1,0,1), (1,0,0), (0,1,1), (0,1,0), (0,0,1), (0,0,0).} To use notation similar to that in Subsection (ref), let OA be player $1$ and firm LC be player $2$. Denote $\beta_1 (x_{mOA}):=\beta^0_{OA} + \beta_{OA}'x_{mOA}$ and $\beta_2 (x_{mLC}):= \beta^0_{LC} + \beta_{LC}'x_{mLC}$ with $x_{mOA} = (MS_m, MP_{mOA})'$ and $x_{mLC} = (MS_m, MP_{mLC})'$. The likelihood for market $m$ depends on the choice probabilities:

\vskip -20pt

{

align*[align* omitted — 737 chars of source]

}

\vskip -20pt

Here $s(x_m)$ is a nuisance parameter which corresponds to the various {\it aggregate} equilibrium selection probabilities. Here $s(\cdot)$ is defined on the support of $x_m$, so in the model this function takes $2^3=8$ values each belonging to $[0,1]$. In the {\it full model} we make no assumptions on the equilibrium selection mechanism. Therefore, the full model has 17 parameters: 4 parameters per profit function (namely $\Delta_i$, $\beta_i^0$, $\beta_i^{MS}$, and $\beta_i^{MP}$), the correlation $\rho$ between $\epsilon_{i1}$ and $\epsilon_{i2}$, and the 8 parameters in the aggregate equilibrium choice probabilities $s(\cdot)$. We also estimate a restricted version of the model called {\it fixed $s$} in which we restrict the aggregate selection probabilities to be the same across markets, for a total of 10 parameters. Note that these are just one version of the econometric model for a game; a less parsimonious version would allow, for example, for the parameters to change with regressor values, or allow for the regressors' support to be richer (rather than binary). We analyze this case precisely to highlight the fact that our CSs provide coverage guarantees regardless of whether the parameter vector is point identified.

sidewaystable[p] \begin{center}{ \begin{tabular}{|c|cccc|cccc|} \hline & \multicolumn{4}{c|}{Full model} & \multicolumn{4}{c|}{Fixed-$s$ model}\\ \hline & Procedure 2 & Procedure 3 & Projection & Percentile & Procedure 2 & Procedure 3 & Projection & Percentile \\ $\Delta_{OA}$ & [$ -1.599 ,\! -1.178 $] & [$ -1.539 ,\! -1.303 $] & [$ -1.707 ,\! -0.701 $] & [$ -1.515 ,\! -1.117 $] & [$ -1.563 ,\! -1.335 $] & [$ -1.543 ,\! -1.363 $] & [$ -1.655 ,\! -1.194 $] & [$ -1.536 ,\! -1.326 $] \\ $\Delta_{LC}$ & [$ -1.527 ,\! -1.218 $] & [$ -1.503 ,\! -1.246 $] & [$ -1.719 ,\! -1.018 $] & [$ -1.489 ,\! -1.225 $] & [$ -1.567 ,\! -1.343 $] & [$ -1.547 ,\! -1.367 $] & [$ -1.671 ,\! -1.222 $] & [$ -1.548 ,\! -1.339 $] \\ $\beta_{OA}^0$ & [$ 0.443 ,\! 0.581 $] & [$ 0.455 ,\! 0.575 $] & [$ 0.341 ,\! 0.695 $] & [$ 0.447 ,\! 0.578 $] & [$ 0.431 ,\! 0.551 $] & [$ 0.437 ,\! 0.539 $] & [$ 0.365 ,\! 0.611 $] & [$ 0.427 ,\! 0.540 $] \\ $\beta_{OA}^{MS}$ & [$ 0.365 ,\! 0.539 $] & [$ 0.383 ,\! 0.521 $] & [$ 0.238 ,\! 0.665 $] & [$ 0.389 ,\! 0.544 $] & [$ 0.347 ,\! 0.479 $] & [$ 0.353 ,\! 0.467 $] & [$ 0.275 ,\! 0.551 $] & [$ 0.348 ,\! 0.477 $] \\ $\beta_{OA}^{MP}$ & [$ 0.413 ,\! 0.581 $] & [$ 0.425 ,\! 0.569 $] & [$ 0.275 ,\! 0.713 $] & [$ 0.424 ,\! 0.579 $] & [$ 0.479 ,\! 0.641 $] & [$ 0.497 ,\! 0.623 $] & [$ 0.389 ,\! 0.719 $] & [$ 0.504 ,\! 0.648 $] \\ $\beta_{LC}^0$ & [$ -1.000 ,\! -0.729 $] & [$ -1.000 ,\! -0.723 $] & [$ -1.000 ,\! -0.453 $] & [$ -0.993 ,\! -0.751 $] & [$ -0.910 ,\! -0.627 $] & [$ -0.874 ,\! -0.657 $] & [$ -1.000 ,\! -0.507 $] & [$ -0.917 ,\! -0.655 $] \\ $\beta_{LC}^{MS}$ & [$ 0.226 ,\! 0.431 $] & [$ 0.238 ,\! 0.419 $] & [$ 0.064 ,\! 0.599 $] & [$ 0.220 ,\! 0.405 $] & [$ 0.299 ,\! 0.443 $] & [$ 0.305 ,\! 0.431 $] & [$ 0.220 ,\! 0.527 $] & [$ 0.303 ,\! 0.442 $] \\ $\beta_{LC}^{MP}$ & [$ 1.591 ,\! 1.868 $] & [$ 1.633 ,\! 1.832 $] & [$ 1.423 ,\! 1.988 $] & [$ 1.615 ,\! 1.821 $] & [$ 1.573 ,\! 1.790 $] & [$ 1.597 ,\! 1.760 $] & [$ 1.489 ,\! 1.880 $] & [$ 1.590 ,\! 1.776 $] \\ $\rho$ & [$ 0.874 ,\! 0.986 $] & [$ 0.910 ,\! 0.978 $] & [$ 0.713 ,\! 0.998 $] & [$ 0.867 ,\! 0.977 $] & [$ 0.938 ,\! 0.990 $] & [$ 0.948 ,\! 0.986 $] & [$ 0.886 ,\! 0.998 $] & [$ 0.935 ,\! 0.986 $] \\ $s$ & --- & --- & --- & --- & [$ 0.926 ,\! 0.980 $] & [$ 0.932 ,\! 0.976 $] & [$ 0.888 ,\! 0.992 $] & [$ 0.927 ,\! 0.977 $] \\ $s_{000}$ & [$ 0.587 ,\! 0.964 $] & [$ 0.679 ,\! 0.950 $] & [$ 0.000 ,\! 1.000 $] & [$ 0.572 ,\! 0.934 $] & --- & --- & --- & --- \\ $s_{001}$ & [$ 0.812 ,\! 1.000 $] & [$ 0.854 ,\! 1.000 $] & [$ 0.439 ,\! 1.000 $] & [$ 0.797 ,\! 0.995 $] & --- & --- & --- & --- \\ $s_{010}$ & [$ 0.000 ,\! 1.000 $] & [$ 0.000 ,\! 0.906 $] & [$ 0.000 ,\! 1.000 $] & [$ 0.018 ,\! 0.828 $] & --- & --- & --- & --- \\ $s_{100}$ & [$ 0.637 ,\! 0.998 $] & [$ 0.794 ,\! 0.998 $] & [$ 0.000 ,\! 1.000 $] & [$ 0.612 ,\! 0.990 $] & --- & --- & --- & --- \\ $s_{011}$ & [$ 0.916 ,\! 1.000 $] & [$ 0.930 ,\! 1.000 $] & [$ 0.804 ,\! 1.000 $] & [$ 0.915 ,\! 0.999 $] & --- & --- & --- & --- \\ $s_{101}$ & [$ 0.491 ,\! 0.920 $] & [$ 0.607 ,\! 0.842 $] & [$ 0.000 ,\! 1.000 $] & [$ 0.449 ,\! 0.799 $] & --- & --- & --- & --- \\ $s_{110}$ & [$ 0.000 ,\! 1.000 $] & [$ 0.000 ,\! 1.000 $] & [$ 0.000 ,\! 1.000 $] & [$ 0.042 ,\! 0.986 $] & --- & --- & --- & --- \\ $s_{111}$ & [$ 0.942 ,\! 1.000 $] & [$ 0.966 ,\! 1.000 $] & [$ 0.856 ,\! 1.000 $] & [$ 0.941 ,\! 0.999 $] & --- & --- & --- & --- \\ \hline \end{tabular}} \vskip 4pt \parbox{14cm}{\caption{ Entry game application: 95% CSs for structural parameters computed via our Procedures 2 and 3 as well as via Projection and Percentile methods. The full model contains a general specification for equilibrium selection while the fixed-$s$ model restricts the equilibrium selection probability to be the same across markets with different regressor values.}} \end{center}

We again take a flat prior on $\Theta$ and implement the procedures using a likelihood criterion. We restrict the support of $\Delta_i$ to $[-2,0]$, $\beta_i$ to $[-1,2]^3$, $\rho$ to $[0,1]$ and the selection probabilities to $[0,1]$. We implement the procedure using the adaptive SMC algorithm as described in Appendix (ref) with $B = 10000$ draws. Histograms of the SMC draws for the selection probabilities are presented in Figure (ref); histograms of draws for the profit function parameters and $\rho$ are presented in Figures (ref) and (ref) in Appendix (ref). We construct CSs for each of the parameters using procedure 2 and procedure 3 and compare these to projection-based CSs (projecting $\wh \Theta_\alpha$ using our procedure 1) and percentile CSs. The empirical findings are presented in Table (ref) below. Appendix (ref) contains further details on computation of $M(\theta)$ for implementation of procedure 2. As with the game simulation, we do not explicitly use the reduced-form reparameterization $\theta \mapsto \tilde \gamma(\theta)$ when computing $M(\theta)$.

The results in Table (ref) show that CSs computed via procedures 2 and 3 are generally similar (though there are some differences, with CSs via procedure 2, which is valid under weaker conditions than procedure 3, appearing wider for some of the selection probabilities in the full model). On the other hand, projection CSs are very wide, especially in the full model. For instance, the projection CS for $s_{101}$ is $[0,1]$ whereas CSs via procedures 2 and 3 are $[0.49,0.92]$ and $[0.61,0.84]$ respectively. As expected, percentile CSs are narrower than procedure 2 and 3 CSs, reflecting the fact that percentile CSs under-cover in partially identified models.

Starting with the {full model} results, we see that the estimates are meaningful economically and are inline with recent estimates obtained in the literature. For example, fixed costs (the intercepts) are positive and significant for the large airlines (OA) but are negative for the LC carriers. Typically, the presence of higher fixed costs can signal various barriers to entry prevent LCs from entering: the higher these fixed costs the less likely it is for LCs to enter. On the other hand, higher fixed costs of large airlines are associated with a bigger presence (such as a hub) and so OAs are more likely to enter. As expected, both market presence and market size are associated with a positive probability of entry for both OA and LC. Results for the fixed-$s$ model are in agreement with the corresponding ones for the full model and tell a consistent story. Note also the very high correlation in the errors, which could indicate missing profitability variables whereby firms enter a particularly profitable markets regardless of competition.

One interesting observation are the CSs for the selection probabilities (also see Figure (ref)). Consider $s_{010}$ and $s_{110}$: these are the aggregate selection probabilities which, according to the results, are not identified. This is likely due to the rather small number of markets with small size, large presence for OA but small presence for LC (for $s_{010}$) and the small number of markets with large market size, large presence for OA but small presence for LC (for $s_{110}$). The strength of our approach is its {\it adaptivity} to lack of identification in a particular data set: for example, 95% CSs for the identified set for $s_{010}$ are $[0,1]$ (via procedure 2), indicating that the model (and data) has no information about this parameter, while the corresponding CS for the identified set for $s_{111}$ is the narrow and informative interval $[0.94,1.00]$.

figure[figure omitted — 1,214 chars of source]

An empirical model of trade flows

In an influential paper, HMR examine the extensive margin of trade using a structural model estimated with current trade data. The following is a brief description of their empirical framework. Let $M_{ij}$ denote the value of country $i$'s imports from country $j$. This is only observed if country $j$ exports to country $i$. If a random draw for productivity from country $j$ to $i$ is sufficiently high then $j$ will export to $i$. To model this, HMR introduce a latent variable $z_{ij}^*$ which measures trade volume between $i$ and $j$. Here $z_{ij}^*$ takes the value zero if $j$ does not export to $i$ and is strictly positive otherwise. We adapt slightly their empirical model to obtain a selection model of the form:

align*[align* omitted — 320 chars of source]

in which $\lambda_j$, $\chi_i$, $\lambda_j^*$ and $\chi_i^*$ are exporting and importing continent fixed effects, $f_{ij}$ is a vector of observable trade frictions between $i$ and $j$, and $u_{ij}$ and $\eta_{ij}^*$ are error terms described below. Exclusion restrictions can be imposed by setting at least one of the elements of $\nu$ equal to zero.

There are three differences between our empirical model and that of HMR. First, we let $z_{ij}^*$ enter the outcome equation linearly instead of nonlinearly.\footnote{Their nonlinear specification is known to be problematic (see, e.g., HMRcritical).} Second, we use continent fixed effect instead of country fixed effects. This reduces the number of parameters from over 400 to 46. Third, we allow for heteroskedasticity in the selection equation, which is known to be a problem in trade data. This illustrates the robustness approach we advocate which relaxes parametric assumptions on part of the model that is suspect (homoskedasticity) without worrying about loss of point identification.

sidewaystable[p] { \begin{center} \begin{tabular}{|c|cc|ccccc|} \multicolumn{1}{c} & \multicolumn{2}{c}{Homoskedastic} & \multicolumn{5}{c}{Heteroskedastic} \\ \hline Variable & MLE & $t$-stat CI & MLE & $t$-stat CI & Procedure 2 & Procedure 3 & Percentile \\ \hline Distance & 2.352 & [1.154,3.549] & 0.314 & [0.273,0.355] & [0.216,0.749] & [0.242,0.509] & [0.207,0.397] \\ Border & -5.191 & [-7.077,-3.304] & -2.265 & [-2.452,-2.077] & [-2.651,-1.859] & [-2.611,-1.898] & [-2.618,-1.816] \\ Island & -1.302 & [-1.913,-0.691] & -0.728 & [-0.868,-0.589] & [-1.060,-0.308] & [-1.060,-0.308] & [-0.983,-0.397] \\ Landlock & -7.275 & [-10.769,-3.780] & -1.369 & [-1.602,-1.137] & [-2.194,-0.914] & [-2.194,-0.890] & [-1.801,-0.954] \\ Legal & 0.358 & [0.002,0.715] & -0.122 & [-0.183,-0.061] & [-0.254,0.004] & [-0.242,-0.009] & [-0.248,0.011] \\ Language & -4.098 & [-6.430,-1.766] & -0.095 & [-0.168,-0.021] & [-0.868,0.049] & [-0.868,0.026] & [-0.237,0.067] \\ Colonial & -17.378 & [-26.002,-8.755] & -2.822 & [-3.029,-2.615] & [-4.980,-2.373] & [-4.980,-2.461] & [-3.231,-2.298] \\ Currency & -1.550 & [-2.780,-0.320] & -0.631 & [-0.946,-0.315] & [-1.315,0.020] & [-1.282,-0.013] & [-1.274,0.062] \\ FTA & -19.540 & [-29.783,-9.298] & -2.151 & [-2.410,-1.892] & [-2.686,-1.589] & [-2.631,-1.616] & [-2.680,-1.577] \\ \hline \end{tabular} \vskip 4pt \parbox{14cm}{\caption{ Maximum likelihood estimates of the coefficients $\nu$ of the trade friction variables in the outcome equation (MLE) together with their 95% confidence sets based on inverting $t$-statistics ($t$-stat CI). Also shown are 95% CSs computed by our Procedures 2 and 3 as well as via Percentile methods.}} \end{center} }

To allow for heteroskedasticity, we suppose that the distribution of $(u_{ij},\eta_{ij}^*)$ conditional on observables is Normal with mean zero and covariance: \[ \Sigma(X_{ij}) = \left[

array[array omitted — 122 chars of source]

\right] \] where $X_{ij}$ denotes $f_{ij}$, the exporter's continent, and the importer's continent and where \[ \sigma_z(X_{ij}) = \exp(\varpi_1 \log(\mathrm{distance}_{ij}) + \varpi_2 \log(\mathrm{distance}_{ij})^2) \,. \]

We estimate the model from data on 24,649 country pairs in the selection equation and 11,146 country pairs in the outcome equation using the same data from 1986 as in HMR. We also impose the exclusion restriction that the coefficient in $\nu$ corresponding to religion is equal to zero, else there is an exact linear relationship between the coefficients in the outcome and selection equation. This leaves a total of 46 parameters to be estimated. We only report estimates for the trade friction coefficients $\nu$ in the outcome equation as these are the most important. We estimate the model first by maximum likelihood under homoskedasticity and report conventional ML estimates for $\nu$ together with 95% CSs based on inverting $t$-statistics. We then re-estimate the model under heteroskedasticity and report conventional ML estimates together with confidence sets based on inverting $t$-statistics, percentile CSs (i.e. the CH procedure under point identification), and our procedures 2 and 3. To implement our Procedure 2 and percentile CSs, we use the adaptive SMC algorithm as described in Appendix (ref) with $B = 10000$ draws.

The results are presented in Table (ref).\footnote{Note that the friction variables enter negatively in the outcome equation. The coefficient of distance is positive meaning that distance negatively affects trade flows; the remaining variables are dummy variables, so a negative coefficient of border means that sharing a border positively affects trade flows, and so forth.} Overall, though the model is sensitive to the presence of heteroskedasticity, the results for the heteroskedastic specification show that the confidence sets seem reasonably insensitive to the type of procedure used, which suggests that partial identification may not be an issue even allowing for heteroskedasticity. We also notice some difference in results relative to HMR. For instance, they document strong positive effects of common legal systems and currency unions on trade flows, whereas we find much weaker evidence for this. We also find a positive effect of landlocked status on trade flows whereas they document a negative effect. Under heteroskedasticity, the magnitudes of coefficients of the trade friction variables are generally smaller than under homoskedasticity but of the same sign. The exception is the legal variable, whose coefficient is positive under homoskedasticity but negative under heteroskedasticity. A remaining question is whether the estimates are also sensitive to the normality assumption on the errors. This question can be examined within the context of our results by, for example, using a flexible form for the joint distribution of the errors.

Large Sample Properties

This section provides regularity conditions under which $\wh \Theta_\alpha$ (Procedure 1), $\wh M_\alpha$ (Procedure 2) and $\wh M_\alpha^\chi$ (Procedure 3) are asymptotically valid confidence sets for $\Theta_I$ and $M_I$. The main new theoretical contributions are the derivations of the large-sample (quasi)-posterior distributions of the QLR for $\Theta_I$ and of the profile QLR for $M_I$ under loss of identifiability.

Coverage properties of $\wh \Theta_\alpha$ for $\Theta_I$

We first state some high-level regularity conditions. A discussion of these assumptions follows.

assumption(Posterior contraction) \\ (i) $L_n(\hat \theta) = \sup_{\theta \in \Theta_{osn}} L_n(\theta) + o_\mb P(n^{-1})$, with $(\Theta_{osn})_{n \in \mb N}$ a sequence of local neighborhoods of $\Theta_I$; \\ (ii) $\Pi_n(\Theta_{osn}^c |\,\mf X_n) = o_\mb P(1)$, where $\Theta_{osn}^c = \Theta \!\setminus \!\Theta_{osn}$.

We presume the existence of a fixed neighborhood $\Theta_I^N$ of $\Theta_I^{\phantom *}$ (with $\Theta_{osn} \subset \Theta_I^N$ for all $n$ sufficiently large) upon which there exists a local reduced-form reparameterization $\theta \mapsto \gamma(\theta)$ from $\Theta_I^N$ into $\Gamma \subseteq \mb R^{d^*}$ for a possibly unknown dimension $d^* \in [1,\infty)$, with $\gamma(\theta) =\gamma_0 \equiv 0$ if and only if $\theta \in \Theta_I$. Here $\gamma (\cdot) $ is merely a proof device and is only required to exist for $\theta$ in a fixed neighborhood of $\Theta_I$. To accommodate situations in which the true reduced-form parameter value $\gamma_0 =0$ may be “on the boundary” of $\Gamma$, we assume that the sets $T_{osn} \equiv \{ \sqrt n \gamma(\theta) : \theta \in \Theta_{osn}\}$ cover\footnote{We say that a sequence of sets $A_n \subseteq \mb R^{d^*}$ covers a set $A \subseteq \mb R^{d^*}$ if (i) $\sup_{b : \|b\| \leq M} |\inf_{a \in A_n} \| a- b\|^2 - \inf_{a \in A} \| a - b\|^2 | = o_\mb P(1)$ for each $M$, and (ii) there is a sequence of closed balls $B_{k_n}$ of radius $k_n \to \infty$ centered at the origin with each $C_n := A_n \cap B_{k_n}$ convex, $C_n \subseteq C_{n'}$ for each $n' \geq n$, and $A = \overline{ \cup_{n \geq 1} C_n }$ (almost surely).} a closed convex cone $T\subseteq \mb R^{d^*}$. We note that this is trivially satisfied with $T = \mb R^{d^*}$ whenever each $T_{osn}$ contains a ball of radius $k_n \to \infty$ centered at the origin. A similar approach is taken for point-identified models by Chernoff1954, Geyer1994, and Andrews1999. Let $\|\gamma\|^2:=\gamma'\gamma$ and for any $v \in \mb R^{d^*}$, let $\mf T v= \mathrm{arg}\min_{t \in T} \|v -t\|^2$ denote the orthogonal (or metric) projection of $v$ onto $T$.

assumption(Local quadratic approximation) \\ There exist sequences of random variables $\ell_n$ and $\mb R^{d^*}$-valued random vectors $\hat \gamma_n$ (both measurable in $\mf X_n$) such that as $n \to \infty$: \begin{equation} \sup_{\theta \in \Theta_{osn}} \left| n L_n(\theta) - \left(\ell_n +\frac{1}{2} \| \sqrt n \hat \gamma_n \|^2 - \frac{1}{2} \|\sqrt n (\hat \gamma_n - \gamma(\theta))\|^2 \right) \right| = o_\mb P(1) \end{equation} with $\sup_{\theta \in \Theta_{osn}} \|\gamma(\theta)\| \to 0$ and $\sqrt n \hat \gamma_n = \mf T \mb V_n$ where $\mb V_n \rightsquigarrow N(0,\Sigma)$.

Let $\Pi_\Gamma$ denote the image measure (under the map $\theta \mapsto \gamma(\theta)$) of the prior $\Pi$ on $\Theta_I^N$, namely $\Pi_\Gamma (A) = \Pi(\{\theta \in \Theta_I^N : \gamma(\theta) \in A\})$. Let $B_\delta \subset \mb R^{d^*}$ be a ball of radius $\delta$ centered at the origin.

assumption(Prior) \\ (i) $\int_\Theta e^{nL_n(\theta)} \,\mr d \Pi(\theta) < \infty $ almost surely; \\ (ii) $\Pi_\Gamma$ has a continuous, strictly positive density $\pi_\Gamma$ on $B_\delta \cap \Gamma$ for some $\delta > 0$.

\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Discussion of Assumptions:} Assumption (ref)(i) is a standard condition on any approximate extremum estimator, and Assumption (ref)(ii) is a standard posterior contraction condition. The choice of $\Theta_{osn}$ is deliberately general and will depend on the particular model under consideration. See Section (ref) for verification of Assumption (ref). Assumption (ref) is a standard local quadratic expansion condition imposed on the local reduced form parameter around $\gamma =0$. It is readily verified for likelihood and GMM models (see Section (ref)) with $\hat \gamma_n =\gamma (\hat \theta )$ and $\mb V_n$ typically a normalized score function of the data. For these models with i.i.d. data the vector $\mb V_n$ is typically of the form: $\mb V_n = n^{-1/2} \sum_{i=1}^n v(X_i)+ o_\mb P(1)$ with $\mb E[v(X_i) ] = 0$ and $\mathrm{Var}[v(X_i)]=\Sigma$. In fact, Appendix (ref) shows that this quadratic expansion assumption is satisfied uniformly over a large class of DGPs in models with discrete data. Assumption (ref)(i) requires the quasi-posterior to be proper. Assumption (ref)(ii) is a standard prior mass and smoothness condition used to establish BvM theorems for identified parametric models (see, e.g., Section 10.2 of vdV) but applied to $\Pi_\Gamma$. Under a flat prior on $\Theta$ and a continuous local mapping $\gamma :\Theta_I^N \mapsto \Gamma$, this assumption is easily satisfied (see its verification in examples of Section (ref)).

Assumptions (ref)(i) and (ref) imply that the QLR statistic for $\Theta_I$ satisfies

equation[equation omitted — 102 chars of source]

(see Lemma (ref)). Therefore, under the {\it generalized information equality} $\Sigma = I_{d^*}$, which holds for a correctly-specified likelihood, an optimally-weighted or continuously-updated GMM, or various (generalized) empirical-likelihood criterions, the asymptotic distribution of $\sup_{\theta \in \Theta_I} Q_n(\theta)$ becomes $F_T$, which is defined as

equation[equation omitted — 74 chars of source]

where $\mb P_Z$ denotes the distribution of a $N(0,I_{d^*})$ random vector $Z$. This recovers the known asymptotic distribution result for QLR statistics under point identification. If $T = \mb R^{d^*}$ then $F_T$ reduces to $ F_{\chi^2_{d^*}}$, the cdf of $\chi^2_{d^*}$ (a chi-square random variable with $d^*$ degree of freedom). If $T$ is polyhedral then $F_T$ is the distribution of a chi-bar-squared random variable (i.e. a mixture of chi-squared distributions with different degrees of freedom where the mixture weights depend on $T$).

Let $\mb P_{Z|\mf X_n}$ denote the distribution of a $N(0,I_{d^*})$ random vector $Z$ (conditional on the data), and $T -v$ denote the convex cone $T$ translated to have vertex at $-v$. The next lemma establishes the large sample behavior of the posterior distribution of the QLR statistic.

lemmaLet Assumptions (ref), (ref) and (ref) hold. Then: \begin{equation} \sup_z \left| \Pi_n \big(\{ \theta : Q_n(\theta) \leq z \}\big|\,\mf X_n\big) - \mb P_{Z | \mf X_n} \Big( \|Z\|^2 \leq z \Big| Z \in T - \sqrt n \hat \gamma_n \Big) \right| = o_\mb P(1)\,. \end{equation} And hence we have:\\ (i) If $T \subsetneq \mb R^{d^*}$ then: $\Pi_n \big(\{ \theta : Q_n(\theta) \leq z \}\big|\,\mf X_n\big) \leq F_T (z)$ for all $z \geq 0$.\\ (ii) If $T = \mb R^{d^*}$ then: $\sup_z \left| \Pi_n \big(\{ \theta : Q_n(\theta) \leq z\} \,\big|\,\mf X_n\big) - F_{\chi^2_{d^*}}\!\!( z ) \right| = o_\mb P(1)$.

This result shows that the posterior distribution of the QLR statistic is asymptotically $\chi^2_{d^*}$ when $T = \mb R^{d^*}$, which may be viewed as a Bayesian Wilks theorem for partially identified models, and asymptotically (first-order) stochastically dominates $F_T$ when $T$ is a closed convex cone. Note that Lemma (ref) does not require the generalized information equality $\Sigma=I_{d^*}$ to hold. This lemma extends known BvM results for possibly misspecified likelihood models with point-identified $\sqrt n$-consistent and asymptotically normally estimable parameters (see KleijnvdV and the references therein) to allow for other models with failure of $\Sigma=I_{d^*}$, with partially-identified parameters and/or parameters on a boundary.

Let $\xi_{n,\alpha}^{post}$ denote the $\alpha$ quantile of $Q_n(\theta)$ under the posterior distribution $\Pi_n$, and let $\xi_{n,\alpha}^{mc}$ be as stated in Remark (ref).

assumption(MC convergence)\\ $\xi_{n,\alpha}^{mc} = \xi_{n,\alpha}^{post} + o_\mb P(1)$.

Lemma (ref) and Assumption (ref) together imply that our Procedure 1 CS $\wh \Theta_\alpha$ is always a well-defined (quasi-)Bayesian credible set (BCS) regardless of whether $\Sigma=I_{d^*}$ holds or not. Further, together with Equation ((ref)), they imply the following result.

theoremLet Assumptions (ref), (ref), (ref), and (ref) hold with $\Sigma = I_{d^*}$. Then for any $\alpha$ such that $F_T (\cdot )$ is continuous at its $\alpha$ quantile, we have: \newline (i) $\liminf_{n \to \infty} \mb P(\Theta_I \subseteq \wh \Theta_{\alpha}) \geq \alpha $; \newline (ii) If $T = \mb R^{d^*}$ then: $\lim_{n \to \infty} \mb P(\Theta_I \subseteq \wh \Theta_{\alpha}) = \alpha $.

Theorem (ref) shows that we need the generalized information equality $\Sigma=I_{d^*}$ to hold so that our Procedure 1 CS $\wh \Theta_\alpha$ has valid frequentist coverage for $\Theta_I$ in large samples.\footnote{This is consistent with the fact that percentile CSs also need $\Sigma=I_{d^*}$ in order to have a correct coverage for a point-identified scalar parameter (see, e.g., CH and CRobert).} This is because the asymptotic distribution of $\sup_{\theta \in \Theta_I} Q_n(\theta)$ is $F_T$ only under $\Sigma=I_{d^*}$. It follows that, with a criterion satisfying $\Sigma = I_{d^*}$, our CS $\wh \Theta_\alpha$ will be asymptotically exact (for $\Theta_I$) when $T = \mb R^{d^*}$, and asymptotically valid but possibly conservative when $T$ is a convex cone.

remarkTheorem (ref) is still applicable to misspecified, separable partially-identified likelihood models. We can write the density in such models as $p_\theta(\cdot) = q_{\tilde \gamma(\theta)}(\cdot)$ where $\tilde \gamma(\theta)$ is an identifiable reduced-form parameter (see Section (ref) below). Under misspecification the identified set is $\Theta_I = \{ \theta : \tilde \gamma(\theta) = \tilde \gamma^*\}$ where $\tilde \gamma^*$ is the unique maximizer of $E[\log q_{\tilde \gamma}(X_i)]$ over $\wt \Gamma = \{\tilde \gamma(\theta) : \theta \in \Theta\}$. Following the insight of Mueller, we could base our inference on the sandwich log-likelihood function: \[ L_n(\theta) = -\frac{1}{2} (\check \gamma - \tilde \gamma(\theta)) ' (\wh \Sigma_S)^{-1} (\check \gamma - \tilde \gamma(\theta)) \] where $\check \gamma$ approximately maximizes $\frac{1}{n} \sum_{i=1}^n \log q_\gamma(X_i)$ over $\wt \Gamma$ and $\wh \Sigma_S$ is the sandwich covariance matrix estimator for $\check \gamma$. If $\sqrt n (\check \gamma - \tilde \gamma^*) \rightsquigarrow N(0,\Sigma_S)$ and $\hat \Sigma_S \to_p \Sigma_S$ with $\Sigma_S$ positive definite, then Assumption (ref) will hold with $\hat \gamma_n = \Sigma^{-1/2}_S(\check \gamma - \tilde \gamma^*)$ where $\sqrt n \hat \gamma_n \to_d N(0,I_{d^*})$ and $\gamma(\theta) = \Sigma^{-1/2}_S(\tilde \gamma(\theta) - \gamma^*)$.
remarkIn correctly specified likelihood models with flat priors, one may interpret $\wh \Theta_\alpha$ as a HPD 100$\alpha$% BCS for $\theta$. MoonSchorfheide (MS hereafter) show that BCSs for a partially identified parameter $\theta$ (or subvectors) can under-cover asymptotically. As CSs for $\Theta_I$ should be larger than CSs for $\theta$, MS's result might appear to suggest that our Procedure 1 CS $\wh \Theta_\alpha$ would under-cover for $\Theta_I$. The “apparent contradiction” is because a key regularity condition in MS's under-coverage result (their Assumption 2 on p. 767) is violated in our setting. For partially identified separable models, MS put a prior on the globally identified reduced-form parameter $\gamma$, say $\Pi(\gamma)$, and then a conditional prior, say $\Pi(\theta|\gamma)$, on the structural parameter $\theta$ given $\gamma$. The conditional prior $\Pi(\cdot|\gamma)$ needs to be supported on what would be the identified set for $\theta$ if $\gamma$ were the true reduced form parameter. Their Assumption 2 requires that $\Pi(\cdot|\gamma)$ is (locally) Lipschitz in $\gamma$, which is violated in our setting. We only put a prior on $\theta$. This prior on $\theta$ induces a prior on $\gamma = \gamma(\theta)$ and a conditional prior $\Pi(\theta|\gamma)$ that is supported on $\{\theta \in \Theta : \gamma(\theta) = \gamma\}$. Since $\gamma(\theta) = 0$ if and only if $\theta \in \Theta_I$, for any $\bar \gamma \neq 0$ our induced conditional prior satisfies \[ \Pi(\{\theta \in \Theta: \gamma(\theta) = 0\} | \gamma = 0) - \Pi(\{\theta \in \Theta : \gamma(\theta) = 0\} | \gamma = \bar \gamma ) = 1-0 = 1, \] thereby violating MS's Lipschitz condition (their Assumption 2). See Remark 3 in MS for additional discussion of violation of their Lipschitz condition.

Models with singularities

In this subsection we consider (possibly) partially identified models with singularities.\footnote{Such models are also referred to as non-regular models or models with non-regular parameters.} In identifiable parametric models $\{ P_\theta : \theta \in \Theta\}$, the standard notion of differentiability in quadratic mean requires that the mass of the part of $P_\theta$ that is singular with respect to the true distribution $P_0 = P_{\theta_0}$ vanishes faster than $\|\theta - \theta_0\|^2$ as $\theta \to \theta_0$ LeCamYang. If this condition fails then the log-likelihood will not be locally quadratic at $\theta_0$. By analogy with the identifiable case, we say a non-identifiable model has a singularity if it does not admit a local quadratic approximation (in the reduced-form reparameterization) like that in Assumption (ref). One example is the missing data model under identification (see Subsection (ref) below).

To allow for partially identified models with singularities, we first generalize the notion of the local reduced-form reparameterization to be of the form $\theta \mapsto (\gamma(\theta),\gamma_\bot(\theta))$ from $\Theta_I^N$ into $\Gamma \times \Gamma_\bot$ where $\Gamma \subseteq \mb R^{d^*}$ and $\Gamma_\bot \subseteq \mb R^{\dim(\gamma_\bot)}$ with $(\gamma(\theta),\gamma_\bot(\theta)) = 0$ if and only if $\theta \in \Theta_I$. The following regularity conditions generalize Assumptions (ref) and (ref) to allow for singularity.

\setcounter{aprime}{1}

aprime(Local quadratic approximation with singularity) \\ (i) There exist sequences of random variables $\ell_n$ and $\mb R^{d^*}$-valued random vectors $\hat \gamma_n$ (both measurable in $\mf X_n$), and a sequence of functions $f_{n,\bot}: \Gamma_\bot \to \mb R_{+}$ (measurable in $\mf X_n$) with $f_{n,\bot}(0)=0$ (almost surely), such that as $n \to \infty$: \begin{equation} \sup_{\theta \in \Theta_{osn}} \left| n L_n(\theta) - \left( \ell_n + \frac{1}{2} \|\sqrt n \hat \gamma_n\|^2 - \frac{1}{2} \| \sqrt n(\hat \gamma_n - \gamma(\theta))\|^2 - f_{n,\perp}(\gamma_\perp(\theta)) \right) \right| = o_\mb P(1) \end{equation} with $\sup_{\theta \in \Theta_{osn}} \|(\gamma(\theta),\gamma_\bot(\theta))\| \to 0$ and $\sqrt n \hat \gamma_n = \mf T \mb V_n$ where $\mb V_n \rightsquigarrow N(0,\Sigma)$; \\ (ii) $\{(\gamma(\theta),\gamma_{\bot}(\theta)) : \theta \in \Theta_{osn}\} = \{\gamma(\theta) : \theta \in \Theta_{osn}\} \times \{\gamma_{\bot}(\theta) : \theta \in \Theta_{osn}\}$.

Let $\Pi_{\Gamma^*}$ denote the image of the measure $\Pi$ under the map $\Theta_I^N \ni \theta \mapsto (\gamma(\theta),\gamma_\bot(\theta))$. Let $B_r^* \subset \mb R^{d^*+\dim(\gamma_\bot)}$ denote a ball of radius $r$ centered at the origin.

aprime(Prior with singularity) \\ (i) $\int_\Theta e^{nL_n(\theta)} \,\mr d \Pi(\theta) < \infty $ almost surely\\ (ii) $\Pi_{\Gamma^*}$ has a continuous, strictly positive density $\pi_{\Gamma^*}$ on $B_\delta^* \cap (\Gamma \times \Gamma_\bot)$ for some $\delta > 0$.

\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Discussion of Assumptions:} Assumption (ref)' is generalizes of Assumption (ref) to the singular case. Assumption (ref)' implies that the peak of the likelihood does not concentrate on sets of the form $\{ \theta : f_{n,\perp}(\gamma_\perp(\theta)) > \epsilon >0\}$. Recently, BochkinaGreen established a BvM result for identifiable parametric likelihood models with singularities. They assume the likelihood is locally quadratic in some parameters and locally linear in others (similar to Assumption (ref)'(i)) and that the local parameter space satisfies conditions similar to our Assumption (ref)'(ii). Assumption (ref)' generalizes Assumption (ref) to the singular case. We impose no further restrictions on the set $\{\gamma_\bot(\theta) : \theta \in \Theta_I^N \}$.

The next lemma shows that the posterior distribution of the QLR asymptotically (first-order) stochastically dominates $F_T$ in partially identified models with singularity.

lemmaLet Assumptions (ref), (ref)' and (ref)' hold. Then: \begin{equation} \sup_z \left( \Pi_n \big(\{ \theta : Q_n(\theta) \leq z \}\big|\,\mf X_n\big) - \mb P_{Z | \mf X_n} \Big( \|Z\|^2 \leq z \Big| Z \in T - \sqrt n \hat \gamma_n \Big) \right) \leq o_\mb P(1) \,. \end{equation} Hence: $\sup_z \left( \Pi_n \big(\{ \theta : Q_n(\theta) \leq z \}\big|\,\mf X_n\big) - F_T(z) \right) \leq o_\mb P(1)$.

Lemma (ref) immediately implies the following result.

theoremLet Assumptions (ref), (ref)', (ref)', and (ref) hold with $\Sigma=I_{d^*}$. Then for any $\alpha$ such that $F_T (\cdot )$ is continuous at its $\alpha$ quantile, we have: $\liminf_{n \to \infty} \mb P(\Theta_I \subseteq \wh \Theta_{\alpha}) \geq \alpha$.

For non-singular models, Theorem (ref) establishes that $\wh \Theta_\alpha$ is asymptotically valid for $\Theta_I$, with asymptotically exact coverage when $T$ is linear and can be conservative when $T$ is a closed convex cone. For singular models, Theorem (ref) shows that $\wh \Theta_\alpha$ is still asymptotically valid for $\Theta_I$ but can be conservative even when $T$ is linear.\footnote{It might be possible to establish asymptotically exact coverage of $\wh \Theta_\alpha$ for $\Theta_I$ in singular models where the singular part $f_{n,\perp}(\gamma_\perp(\theta))$ in Assumption (ref)' possesses some extra structure.} When applied to the missing data example, Theorems (ref) and (ref) imply that $\wh \Theta_\alpha$ for $\Theta_I$ is asymptotically exact under partial identification but conservative under point identification. This is consistent with simulation results reported in Table (ref); see Section (ref) below for details.

Coverage properties of $\wh M_\alpha$ for $M_I$

Here we present conditions under which $\wh M_\alpha$ has correct coverage for the identified set $M_I$ of subvectors $\mu$. Recall the definition of $M(\theta) \equiv \{ \mu : (\mu,\eta) \in \Delta(\theta) \mbox{ for some } \eta\}$ from Section (ref). The profile criterion $PL_n(M(\theta))$ for $M(\theta)$ and the profile QLR $PQ_n(M(\theta))$ for $M(\theta)$ are defined the same way as those in ((ref)) and ((ref)) respectively: \[ PL_n(M(\theta)) \equiv \inf_{\mu \in M(\theta)} \sup_{\eta \in H_\mu} L_n(\mu,\eta) \quad \mbox{and} \quad PQ_n(M(\theta))\equiv 2n [L_n(\hat \theta) - PL_n(M(\theta))]. \]

assumption(Profile QL) \\ There exists a measurable $f : \mb R^{d^*} \to \mb R_{+}$ such that: \begin{align*} & \sup_{\theta \in \Theta_{osn}} \left| nP L_n(M(\theta)) - \left( \ell_n + \frac{1}{2} \|\sqrt n \hat \gamma_n \|^2 - \frac{1}{2} f \left( \sqrt n (\hat \gamma_n - \gamma(\theta)) \right) \right) \right| = o_\mb P(1) \end{align*} with $\hat \gamma_n$ and $\gamma(\cdot)$ from Assumption (ref) or (ref)'.

\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Discussion of Assumption (ref):} By definition of $M_I$ (in display ((ref))) we have: $M_I = \{\mu : \gamma(\mu,\eta)=0 \mbox{ for some } \eta \in H_\mu\}$ and also $M_I = M(\theta)$ for any $\theta \in \Theta_I$. Thus \[ PQ_n(M_I) = \sup_{\mu \in M_I} \inf_{\eta \in H_\mu} Q_n (\mu, \eta) =PQ_n(M(\theta))\quad \mbox{for all $\theta \in \Theta_I$}~. \] Assumption (ref) imposes some structure on the profile QLR statistic for $M_I$ over the local neighborhood $\Theta_{osn}$. It implies that the profile QLR for $M_I$ is of the form:

equation[equation omitted — 85 chars of source]

When $\Sigma = I_{d^*}$, the asymptotic distribution of $\sup_{\theta \in \Theta_I} PQ_n(M(\theta))=PQ_n(M_I)$ becomes $G_T$: \[ G_T (z) := \mb P_Z (f( \mf T Z ) \leq z) \,~~~\mbox{where}~~Z \sim N(0,I_{d^*})~. \] The functional form of $f$ depends on the local reparameterization $\gamma$ and the geometry of $M_I$. When $M_I$ is a singleton and $T = \mb R^{d^*}$ then equation ((ref)) is typically satisfied with $f(v) = \inf_{t \in T_1} \|v - t\|^2$ where $T_1= \mb R^{d_1^*}$ with $d_1^* < d^*$ and the QLR statistic is asymptotically $\chi^2_{d^*-d_1^*}$. For a non-singleton set $M_I$, $f$ will typically be more complex. In the missing data example, we show in Section (ref) that $f(v) = \max_{\mu \in \{ \ul \mu,\ol \mu\}} \inf_{t \in T_\mu} \|v - t\|^2$ where $T_{\ul \mu}$ and $T_{\ol \mu}$ are halfspaces and $T = \mb R^{d^*}$. Here the resulting profile QLR statistic for $M_I$ is asymptotically the maximum of two mixtures of $\chi^2$ random variables. Note that the existence of $f$ is merely a proof device, and one does not need to know its precise expression in the implementation of our Procedure 2 CS $\wh M_{\alpha}$ for $M_I$.

The next lemma is a new BvM-type result for the posterior distribution of the profile QLR for $M_I$. Note that this result also allows for singular models.

lemmaLet Assumptions (ref), (ref), (ref), and (ref) or (ref), (ref)', (ref)', and (ref) hold. Then for any interval $I$ such that $\mb P_{Z} ( f(Z) \leq z )$ is continuous on $I$, we have: \begin{align*} \sup_{z \in I} \left| { \Pi_n \big( \{\theta:PQ_n( M(\theta)) \leq z \} \,\big|\, \mf X_n \big) } - \mb P_{Z | \mf X_n} \Big( f(Z) \leq z \Big| Z \in \sqrt n \hat \gamma_n - T \Big) \right| = o_\mb P(1)\,. \end{align*} And hence we have:\\ (i) If $T \subsetneq \mb R^{d^*}$ and $f$ is subconvex,\footnote{We say that $f : \mb R^{d^*} \to \mb R_+$ is quasiconvex if $f^{-1}(z) := \{ v : f(v) \leq z\}$ is convex for each $z \geq 0$ and subconvex if, in addition, $f(v) = f(-v)$ for all $v \in \mb R^{d^*}$. The conclusion of Lemma (ref)(i) remains valid under the weaker condition that (i) $f$ is quasiconvex and (ii) $\mb P_Z( Z \in ( f^{-1}(\xi_\alpha)- T^o)) \leq \mb P_Z( f(\mf T Z) \leq \xi_\alpha) $ holds, where $\xi_\alpha$ is the $\alpha$ quantile of $G_T$ and $T^o$ is the polar cone of $T$.} then: $\Pi_n \big(\{ \theta : Q_n(\theta) \leq z \}\big|\,\mf X_n\big) \leq G_T (z)$ for all $z \geq 0$.\\ (ii) If $T= \mb R^{d^*}$ then: $\sup_z \left| \Pi_n \big(\{ \theta : Q_n(\theta) \leq z\} \,\big|\,\mf X_n\big) - \mb P_{Z} \big( f ( Z ) \leq z\big) \right| = o_\mb P(1)$.

Let $\xi_{n,\alpha}^{post,p}$ denote the $\alpha$ quantile of the profile QLR $PQ_n(M(\theta))$ under the posterior distribution $\Pi_n$, and $\xi_{n,\alpha}^{mc,p}$ be given in Remark (ref).

assumption(MC convergence) \\ $\xi_{n,\alpha}^{mc,p} = \xi_{n,\alpha}^{post,p} + o_\mb P(1)$.

The next theorem is an important consequence of Lemma (ref).

theoremLet Assumptions (ref), (ref), (ref), (ref), and (ref) or (ref), (ref)', (ref)', (ref), and (ref) hold with $\Sigma=I_{d^*}$ and suppose that $G_T (\cdot )$ is continuous at its $\alpha$ quantile. \newline (i) If $T \subsetneq \mb R^{d^*}$ and $f$ is subconvex,\footnote{The conclusion of Theorem (ref)(i) remains valid under the weaker condition stated in footnote for Lemma (ref)(i).} then: $\liminf_{n \to \infty} \mb P(M_I \subseteq \wh M_{\alpha}) \geq \alpha $ ; \newline (ii) If $T = \mb R^{d^*}$ then: $\lim_{n \to \infty} \mb P(M_I \subseteq \wh M_{\alpha}) = \alpha $.

Theorem (ref)(ii) shows that our Procedure 2 CSs $\wh M_{\alpha}$ for $M_I$ can have asymptotically exact coverage if $T = \mb R^{d^*}$ even if the model is singular. In the missing data example, Theorem (ref)(ii) implies that $\wh M_\alpha$ for $M_I$ is asymptotically exact irrespective of whether the model is point-identified or not (see Subsection (ref) below). Theorem (ref)(i) shows that the CSs $\wh M_{\alpha}$ for $M_I$ can have conservative coverage when $T$ is a convex cone (see Appendix (ref) for a moment inequality example).

Coverage properties of $\wh M_\alpha^\chi$ for $M_I$ of scalar subvectors

This section presents one sufficient condition for validity of our Procedure 3 CS $\wh M_\alpha^\chi$ for $M_I \subset \mb R$. We say a half-space is regular if it is of the form $\{v \in \mb R^{d^*} : a'v \leq 0\}$ for some $a \in \mb R^{d^*}$.

assumption(Profile QLR, $\chi^2$ bound) \\ $PQ_n(M(\theta)) \rightsquigarrow W \leq \max_{i \in \{1,2\}} \inf_{t \in T_i} \| Z - t\|^2$ for all $\theta \in \Theta_I$, where $Z \sim N(0,I_{d^*})$ for some $d^* \geq 1$ and $T_1$ and $T_2$ are regular half-spaces in $\mb R^{d^*}$.
theoremLet Assumption (ref) hold and let the distribution of $W$ be continuous at its $\alpha$ quantile. Then: $\liminf_{n \to \infty} \mb P( M_I \subseteq \wh M_\alpha^\chi ) \geq \alpha$.

The following proposition presents a set of sufficient conditions for Assumption (ref).

propositionLet the following hold:\\ (i) Assumptions (ref)(i), (ref) or (ref)' hold with $\Sigma = I_{d^*}$ and $T = \mb R^{d^*}$; \\ (ii) $\inf_{\mu \in M_I} \sup_{\eta \in H_\mu} L_n(\mu,\eta) = \min_{\mu \in \{ \ul \mu, \ol \mu\}} \sup_{\eta \in H_\mu} L_n(\mu,\eta) + o_\mb P(n^{-1})$; \\ (iii) for each $\mu \in \{ \ul \mu, \ol \mu\}$ there exists a sequence of sets $(\Gamma_{\mu,osn})_{n \in \mb N}$ with $\Gamma_{\mu,osn} \subseteq \Gamma$ for each $n$ and a halfspace $T_\mu$ in $\mb R^{d^*}$ such that: \[ \sup_{\eta \in H_\mu} n L_n(\mu,\eta) = \sup_{\gamma \in \Gamma_{\mu,osn}} \left( \ell_n + \frac{1}{2} \|\mb V_n\|^2 - \frac{1}{2} \| \sqrt n \gamma - \mb V_n \|^2 \right) + o_\mb P(1) \] and $\inf_{\gamma \in \Gamma_{\mu,osn}} \| \sqrt n \gamma - \mb V_n \|^2 = \inf_{t \in T_\mu} \|t - \mb V_n\|^2 + o_\mb P(1)$. \\ Then: Assumption (ref) holds with $W=\max_{i \in \{ \ul \mu, \ol \mu\}} \inf_{t \in T_i} \| Z - t\|^2$.

Suppose $M_I = [\ul \mu, \ol \mu]\subsetneq \mb R$ (which is true when $\Theta_I$ is connected and bounded). If $\sup_{\eta \in H_\mu} L_n(\mu,\eta)$ is strictly concave in $\mu$ then condition (ii) of Proposition (ref) holds. The other conditions of Proposition (ref) are easy to verify as in the missing data example (see Subsection (ref) below).

The exact distribution of $\max_{i \in \{1,2\}} \inf_{t \in T_i} \| Z - t\|^2$ depends on the geometry of $T_1$ and $T_2$. For the missing data example, the polar cones of $T_1$ and $T_2$ are at least $90^o$ apart. The worst-case coverage (i.e., the case in which asymptotic coverage of $\wh M_\alpha^\chi$ will be most conservative) will occur when the polar cones of $T_1$ and $T_2$ are orthogonal, in which case $\max_{i \in \{1,2\}} \inf_{t \in T_i} \| Z - t\|^2$ has the mixture distribution $W^* :=\frac{1}{4} \delta_0 + \frac{1}{2}\chi^2_1 + \frac{1}{4} (\chi^2_1 \cdot \chi^2_1) $ where $\delta_0$ is a point mass at zero and $\chi^2_1 \cdot \chi^2_1$ is the distribution of the product of two independent $\chi^2_1$ random variables. The quantiles of the distribution of $\max_{i \in \{1,2\}} \inf_{t \in T_i} \| Z - t\|^2$ are continuous in $\alpha$ for all $\alpha > \frac{1}{4}$. For all configurations of $T_1$ and $T_2$ in this example, the distribution of $\max_{i \in \{1,2\}} \inf_{t \in T_i} \| Z - t\|^2$ (first-order) stochastically dominates $F_{W^*}$ and is (first-order) stochastically dominated by $F_{\chi^2_1}$ (i.e., $F_{W^*} (w) \geq F_{W} (w) \geq F_{\chi^2_1} (w)$). Notice that this is different from the usual chi-bar-squared case encountered when testing whether a parameter belongs to the identified set on the basis of finitely many moment inequalities Rosen.

To get an idea of the degree of conservativeness of $\wh M_\alpha^\chi$, consider the class of models satisfying conditions for Proposition (ref). Figure (ref) plots the asymptotic coverage of $\wh M_\alpha$ and $\wh M_\alpha^\chi$ against nominal coverage for models in this class where $\wh M_\alpha^\chi$ is most conservative for the missing data example (i.e., the worst-case coverage). For each model in this class, the asymptotic coverage of $\wh M_\alpha$ and $\wh M_\alpha^\chi$ is between the nominal coverage and the worst-case coverage. As can be seen, the coverage of $\wh M_\alpha$ is exact at all levels $\alpha \in (0,1)$ for which the distribution of the profile QLR is continuous at its $\alpha$ quantile, as shown in Theorem (ref)(ii). On the other hand, $\wh M_\alpha^\chi$ is asymptotically conservative, but the level of conservativeness decreases as $\alpha$ increases towards one. Indeed, for levels of $\alpha$ in excess of $0.85$ the level of conservativeness is negligible.

figure[figure omitted — 447 chars of source]

Since empirical papers typically report CSs for scalar parameters, Theorem (ref) will be very useful in applied work. One could generalize $\wh M_\alpha^\chi$ to deal with vector-valued subvectors by allowing $\chi_d^2$ quantiles with higher degrees of freedom $d \in (1,\dim(\theta))$, but it might be difficult to provide sufficient conditions as those in Proposition (ref) to establish results like Theorem (ref).

Sufficient Conditions and Examples

This section provides sufficient conditions for the key regularity condition, Assumption (ref), in possibly partially identified likelihood and moment-based models with i.i.d. data. See Appendix (ref) for low-level conditions to ensure that Assumption (ref) holds uniformly over a large class of DGPs in discrete models. We also verify Assumptions (ref), (ref) (or (ref)'), (ref) and (ref)) in examples.

We use standard empirical process notation: $P_0 g$ denotes the expectation of $g(X_i)$ under the true probability measure $P_0$, $\mb P_n g = n^{-1} \sum_{i=1}^n g(X_i)$ denotes expectation of $g(X_i)$ under the empirical measure, and $\mb G_n g = \sqrt n (\mb P_n - P_0)g$ denotes the empirical process.

Partially identified likelihood models

Consider a parametric likelihood model $\mc P = \{p_\theta : \theta \in \Theta\}$ where each $p_\theta(\cdot)$ is a probability density with respect to a common $\sigma$-finite dominating measure $\lambda$. Let $p_0 \in \mc P$ be the true density under the data-generating probability measure, $D_{KL}(p \| q)$ denote the Kullback-Leibler divergence, and $h(p,q)^2 = \int (\sqrt p - \sqrt q)^2 \, \mr d \lambda$ denote the squared Hellinger distance between densities $p$ and $q$. The identified set is $\Theta_I = \{\theta \in \Theta : D_{KL}(p_0 \| p_\theta) = 0\}= \{\theta \in \Theta : h(p_0, p_\theta) = 0\}$.

Separable likelihood models

For a large class of partially identified parametric likelihood models $\mc P= \{p_\theta : \theta \in \Theta\}$, there exists a function $\tilde \gamma : \Theta \to \wt \Gamma \subset \mb R^{d^*}$ for some possibly unknown $d^* \in [1, + \infty)$, such that $p_\theta(\cdot) = q_{\tilde \gamma(\theta)}(\cdot)$ for each $\theta \in \Theta$ and some densities $\{q_{\tilde \gamma(\theta)}(\cdot) : \tilde \gamma \in \wt \Gamma\}$. In this case we say that the model $\mc P$ is separable and admits a (global) reduced-form reparameterization. The reparameterization is assumed to be identifiable, i.e. $D_{KL}(q_{\tilde \gamma_0} \| q_{\tilde \gamma}) > 0$ for any $\tilde \gamma \neq \tilde \gamma_0$. The identified set is $\Theta_I = \{\theta \in \Theta : \tilde \gamma(\theta) = \tilde \gamma_0\}$ where $\tilde \gamma_0$ is the true parameter, i.e. $p_0 = q_{\tilde \gamma_0}$. Models with discrete choice probabilities (such as the missing data and entry game designs we used in simulations) fall into this framework, where the vector $\tilde \gamma$ maps the structural parameters $\theta$ to the model-implied probabilities of discrete outcomes and the true probabilities $\tilde \gamma_0 \in \wt \Gamma$ of discrete outcomes are point-identified.

The following result presents one set of sufficient conditions for Assumptions (ref)(ii) and (ref) under conventional smoothness assumptions.

Let $\ell_{\tilde \gamma}(\cdot) := \log q_{\tilde \gamma}(\cdot)$, let $\dot \ell_{\tilde \gamma}$ and $\ddot \ell_{\tilde \gamma}$ denote the score and Hessian, let $\mb I_{0} := -P_0( \ddot \ell_{\tilde \gamma_0}^{\phantom\prime} )$ and let $\gamma(\theta) = \mb I_{0}^{1/2} (\tilde \gamma(\theta) - \tilde \gamma_0)$ and $\Gamma = \{\mb I_{0}^{1/2} (\tilde \gamma\ - \tilde \gamma_0) : \tilde \gamma \in \wt \Gamma\}$.

propositionSuppose that $\{q_{\tilde \gamma} : \tilde \gamma \in \wt \Gamma\}$ satisfies the following regularity conditions:\\ (a) $X_1,\ldots,X_n$ is an i.i.d. sample from $q_{\tilde \gamma_0}$ with $\tilde \gamma_0$ identifiable and on the interior of $\wt \Gamma$; \\ (b) $\tilde \gamma \mapsto P_0 \ell_{\tilde \gamma}$ is continuous and there is a neighborhood $U$ of $\tilde \gamma_0$ on which $\ell_{\tilde \gamma}(x)$ is twice continuously differentiable for each $x$, with $\dot \ell_{\tilde \gamma_0} \in L^2(P_0)$ and $\sup_{\tilde \gamma \in U} \|\ddot \ell_{\tilde \gamma}(x) \| \leq \bar \ell(x)$ for some $\bar \ell \in L^2(P_0)$; \\ (c) $P_0 \dot \ell_{\tilde \gamma} = 0$, $\mb I_0$ is non-singular, and $\mb I_0 = P_0( \dot \ell_{\tilde \gamma_0}^{\phantom \prime}\dot \ell_{\tilde \gamma_0}^{\prime})$; \\ (d) $\wt \Gamma$ is compact and $\pi_\Gamma$ is strictly positive and continuous on $U$. \\ Then: there exists a sequence $(r_n)_{n \in \mb N}$ with $r_n \to \infty$ and $r_n = o(n^{1/4})$ such that Assumptions (ref)(ii) and (ref) hold for the average log-likelihood ((ref)) over $\Theta_{osn}:= \{ \theta \in \Theta : \|\gamma(\theta)\| \leq r_n/\sqrt n\}$ with $\ell_n = n\mb P_n \log p_0$, $\sqrt n \hat \gamma_n = \mb V_n = \mb I_{0}^{-1/2} \mb G_n (\dot \ell_{\tilde \gamma_0})$, $\Sigma = I_{d^*}$ and $T = \mb R^{d^*}$.

General non-identifiable likelihood models

It is possible to define a local reduced-form reparameterization for non-identifiable likelihood models, even when $\mc P= \{p_\theta : \theta \in \Theta\}$ does not admit an explicit (global) reduced-form reparameterization. Let $\mc D \subset L^2(P_0)$ denote the set of all limit points of: \[ \mc D_\epsilon := \left\{ \frac{\sqrt{p/p_0}-1}{h(p,p_0)} : p \in \mc P, 0 < h(p,p_0) \leq \epsilon\right\} \] as $\epsilon \to 0$ and let $\overline{\mc D}_{\epsilon} = \mc D_\epsilon \cup \mc D$. The set $\mc D$ is the set of generalized Hellinger scores,\footnote{It is possible to define sets of generalized scores via other measures of distance between densities. See LiuShao and AGM. Our results can easily be adapted to these other cases.} which consists of functions of $X_i$ with mean zero and unit variance. The cone $\mc T = \{ \tau d : \tau \geq 0, d \in \mc D\}$ is the tangent cone of the model $\mc P$ at $p_0$. We say that $\mc P$ is differentiable in quadratic mean (DQM) if each $p \in \mc P$ is absolutely continuous with respect to $p_0$ and for each $p \in \mc P$ there are elements $g_p \in \mc T$ and remainders $R_p \in L^2(\lambda)$ such that: \[ \sqrt{p_{\phantom{.}}} - \sqrt{p_0} = g_p \sqrt{p_0} + h(p,p_0) R_p \] with $\sup\{ \| R_p \|_{L^2(\lambda)} : h(p,p_0) \leq \varepsilon \} \to 0$ as $\varepsilon \to 0$. If the linear hull $\mathrm{Span}(\mc T)$ of $\mc T$ has finite dimension $d^* \geq 1$, then we can write each $g \in \mc T$ as $g = c(g)' \psi$ where $c(g) \in \mb R^{d^*}$ and the elements of $\psi = (\psi_1,\ldots,\psi_{d^*})$ form an orthonormal basis for $\mathrm{Span}(\mc T)$ in $L^2(P_0)$. Let $\mb T$ denote the orthogonal projection\footnote{If $\mc T \subseteq L^2(P_0)$ is a closed convex cone, the projection $\mb T f$ of any $f \in L^2(P_0)$ is defined as the unique element of $\mc T$ such that $\|f - \mb T f\|_{L^2(P_0)} = \inf_{t \in \mc T} \|f - t\|_{L^2(P_0)}$.} onto $\mc T$ and let $\gamma(\theta)$ be given by

align[align omitted — 90 chars of source]
propositionSuppose that $\mc P$ satisfies the following regularity conditions:\\ (a) $\{\log p : p \in \mc P\}$ is $P_0$-Glivenko Cantelli; \\ (b) $\mc P$ is DQM, $\mc T$ is closed and convex and $\mathrm{Span}(\mc T)$ has finite dimension $d^* \geq 1$;\\ (c) there exists $\varepsilon > 0$ such that $\overline{\mc D}_\varepsilon$ is Donsker and has envelope $D \in L^2(P_0)$.\\ Then: there exists a sequence $(r_n)_{n \in \mb N}$ with $r_n \to \infty$ and $r_n = o(n^{1/4})$, such that Assumption (ref) holds for the average log-likelihood ((ref)) over $\Theta_{osn}:= \{ \theta : h(p_\theta,p_0) \leq r_n/\sqrt n\}$ with $\ell_n = n\mb P_n \log p_0$, $\sqrt n \hat \gamma_n = \mb V_n = \mb G_n(\psi)$, $\Sigma = I_{d^*}$ and $\gamma(\theta)$ defined in ((ref)).

Proposition (ref) is a set of sufficient conditions for i.i.d. data; see Lemma (ref) in Appendix (ref) for a more general result. Assumption (ref)(ii) can be verified under additional mild conditions (see, e.g., Theorem 5.1 of GGV2000).

GMM models

Consider the GMM model $\{ \rho_\theta : \theta \in \Theta\}$ with $\rho : \mcr X \times \Theta \to \mb R^{d_\rho}$. Let $g(\theta) = E[\rho_\theta(X_i)]$ and the identified set be $\Theta_I = \{ \theta \in \Theta : g(\theta) = 0\}$ (we assume throughout this subsection that $\Theta_I$ is non-empty). When $\rho$ is of higher dimension than $\theta$, the set $\mc G = \{ g(\theta) : \theta \in \Theta\}$ will not contain a neighborhood of the origin. But, if the map $\theta \mapsto g(\theta)$ is smooth (e.g. $\mc G$ is a smooth manifold) then $\mc G$ can typically be locally approximated at the origin by a closed convex cone $\mc T \subset \mb R^{d_\rho}$.

To simplify notation, we assume that for any $v \in \mathrm{Span}(\mc T)$ we may partition $\Omega^{-1/2} v$ so that its upper $d^*$ elements $[\Omega^{-1/2} v]_1$ are (possibly) non-zero and the remaining $d_\rho -d^*$ elements $[\Omega^{-1/2} v]_2 = 0$ (this can always be achieved by multiplying the moment functions by a suitable rotation matrix).\footnote{See our July 2016 working paper version for details.} If $\mc G$ contains a neighborhood of the origin then we simply take $\mc T = \mb R^{d_\rho}$ and $[\Omega^{-1/2} v]_1 = \Omega^{-1/2} v$. Let $\mb T g(\theta)$ denote the projection of $g(\theta)$ onto $\mc T \subset \mb R^{d_\rho}$ and note that $[\Omega^{-1/2} \mb T g(\theta)]_2=0$. Finally, define $\Theta_I^\varepsilon = \{ \theta \in \Theta : \| g(\theta)\| \leq \varepsilon\}$.

propositionSuppose that $\{ \rho_\theta : \theta \in \Theta\}$ satisfies the following regularity conditions: \\ (a) there exists $\varepsilon_0 > 0$ such that $\{\rho_\theta : \theta \in \Theta_I^{\varepsilon_0}\}$ is Donsker; \\ (b) $E[\rho_\theta(X_i) \rho_\theta(X_i)'] = \Omega$ for each $\theta \in \Theta_I$ and $\Omega$ is positive definite; \\ (c) there exists $\theta^* \in \Theta_I$ such that $\sup_{\theta \in \Theta_I^\varepsilon} E[\|\rho_\theta(X_i) - \rho_{\theta^*}(X_i)\|^2] = o(1)$ as $\varepsilon \to 0$;\\ (d) there exists $\delta > 0$ such that $\sup_{\theta \in \Theta_I^\varepsilon} \| g(\theta) - \mb T g(\theta)\| = o(\varepsilon^{1+\delta})$ as $\varepsilon \to 0$.\\ Then: there exists a sequence $(r_n)_{n \in \mb N}$ with $r_n \to \infty$ and $r_n = o(n^{1/4})$ such that Assumption (ref) holds for the CU-GMM criterion ((ref)) over $\Theta_{osn} = \{ \theta \in \Theta : \|g(\theta)\| \leq r_n/\sqrt n\}$, where $\ell_n = -\frac{1}{2}Z_n'\Omega^{-1} Z_n$, $Z_n=\mb G_n (\rho_{\theta^*})$, $\gamma(\theta) = [\Omega^{-1/2} \mb Tg(\theta)]_1$, and $\sqrt n \hat \gamma_n = \mb V_n = -[\Omega^{-1/2}Z_n]_1$ and $\Sigma = I_{d^*}$.\\ If $\mc G$ contains a neighborhood of the origin then $\gamma(\theta) = \Omega^{-1/2} g(\theta)$ and $\sqrt n \hat \gamma_n = \mb V_n = - \Omega^{-1/2} Z_n$.
propositionLet all the conditions of Proposition (ref) hold and let: (e) $\|\wh W - \Omega^{-1}\| = o_\mb P(1)$. \\ Then: the conclusions of Proposition (ref) hold for the optimally-weighted GMM criterion ((ref)).

Moment inequality models

Consider the moment inequality model $\{ \tilde \rho(X_i,\mu) : \mu \in M\}$ where $\tilde \rho$ is a $d_\rho$ vector of moments and the space is $M \subseteq \mb R^{d_\mu}$. The identified set for $\mu$ is $M_I = \{ \mu \in M : E[\tilde \rho(X_i,\mu)] \leq 0\}$ (the inequality is understood to hold element-wise). We may reformulate the moment inequality model as a moment equality model by augmenting the parameter vector with a vector of slackness parameters $\eta \in H = \mb R^{d_\rho}_+$. Thus we re-parameterize the model by $\theta = (\mu,\eta) \in \Theta = M \times H$ and write the inequality model as a GMM model with

equation[equation omitted — 137 chars of source]

where the identified set for $\theta$ is $\Theta_I = \{ \theta \in \Theta : E[\rho_\theta(X_i)] = 0\}$ and $M_I$ is the projection of $\Theta_I$ onto $M$. Here the objective function would be as in display ((ref)) or ((ref)) using $\rho_\theta(X_i) = \tilde \rho(X_i,\mu) + \eta$. We may then apply Propositions (ref) or (ref) to the reparameterized GMM model ((ref)).

As the parameter of interest is $\mu$, one could use our Procedures 2 or 3 for inference on $M_I$. These procedures involve the profile criterion $\sup_{\eta \in H} L_n(\mu,\eta)$ which is simple to compute because the GMM objective function is quadratic in $\eta$ for given $\mu$ (since the optimal weighting or continuous updating weighting matrix will typically not depend on $\eta$). See Example 3 in Subsection (ref).

Examples

Example 1: missing data model in Subsection (ref)

We revisit the missing data example in Subsection (ref), where the parameter space $\Theta$ for $\theta = (\mu, \eta_1, \eta_2)$ is given in ((ref)), the identified set for $\theta$ is $\Theta_I$ given in ((ref)), and the identified set for $\mu$ is $M_I = [\tilde \gamma_{11},\tilde \gamma_{11}+\tilde \gamma_{00}]$.

\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Inference under partial identification:}

Consider the case in which the model is partially identified (i.e. $0 < \eta_2 < 1$). The likelihood of the $i$-th observation $(D_i,Y_iD_i)=(d,yd)$ is

align*[align* omitted — 204 chars of source]

where: \[ \tilde \gamma(\theta) = \left(

array[array omitted — 114 chars of source]

\right) \] with $\wt \Gamma = \{\tilde \gamma(\theta) : \theta \in \Theta\} = \{(g_{11}-\tilde \gamma_{11},g_{00}-\tilde \gamma_{00}) : (g_{11},g_{00}) \in[0,1]^2 , 0 \leq g_{11} \leq 1-g_{00}\}$. Conditions (a)-(b) of Proposition (ref) hold and Assumption (ref) is satisfied with $\gamma(\theta) = \mb I_{0}^{1/2} \tilde \gamma(\theta)$,

align*[align* omitted — 632 chars of source]

$\Sigma = I_2$ and $T = \mb R^2$. A flat prior on $\Theta$ in ((ref)) induces a flat prior on $\Gamma$, which verifies Condition (c) of Proposition (ref) and Assumption (ref). Therefore, Theorem (ref)(ii) implies that our CSs $\wh \Theta_\alpha$ for $\Theta_I$ has asymptotically exact coverage.

Now consider CSs for $M_I = [\tilde \gamma_{11},\tilde \gamma_{11}+\tilde \gamma_{00}]$. Here $H_\mu = \{ (\eta_1,\eta_2) \in [0,1]^2 : 0 \leq \mu - \eta_1(1-\eta_2) \leq \eta_2\}$. By concavity in $\mu$, the profile log-likelihood for $M_I$ is:

align*[align* omitted — 115 chars of source]

where $\ul \mu= \tilde \gamma_{11}$ and $\ol \mu = \tilde \gamma_{11}+\tilde \gamma_{00}$. The inner maximization problem is: \[ \sup_{\eta \in H_\mu} \mb P_n \log p_{(\mu,\eta)} = \sup_{ \substack{ 0 \leq g_{11} \leq \mu \\ \mu \leq g_{11}+g_{00} \leq 1 } } \!\!\!\!\! \mb P_n \Big( yd \log g_{11} + (d-yd) \log (1-g_{11}-g_{00}) + (1- d) \log g_{00} \Big) .\label{e:md:lower} \] Let $g = (g_{11},g_{00})'$ and $\tilde \gamma = (\tilde \gamma_{11},\tilde \gamma_{00})'$ and let:

align*[align* omitted — 203 chars of source]

where $r_n$ is from Proposition (ref). It follows that:

align*[align* omitted — 269 chars of source]

Equation ((ref)) and Assumption (ref) therefore hold with $f(v) = \max_{\mu \in \{\ul \mu,\ol \mu\}} \inf_{t \in T_\mu} \|v - t\|^2$ where $T_{\ul \mu}$ and $T_{\ol \mu}$ are regular halfspaces in $\mb R^2$. Theorem (ref) implies that the CS $\wh M_\alpha^\chi$ is asymptotically valid (but conservative) for $M_I$.

To verify Assumption (ref), take $n$ sufficiently large that $\gamma(\theta) \in \mathrm{int}(\Gamma)$ for all $\theta \in \Theta_{osn}$. Then:

align[align omitted — 213 chars of source]

This is geometrically the same as the profile QLR for $M_I$ up to a translation of the local parameter space from $(\tilde\gamma_{11},\tilde\gamma_{00})'$ to $(\tilde\gamma_{11}(\theta),\tilde\gamma_{00}(\theta))'$. The local parameter spaces are approximated by $T_{\ul \mu}(\theta) = T_{\ul \mu} + \sqrt n \gamma(\theta)$ and $T_{\ol \mu} (\theta) = T_{\ol \mu} + \sqrt n \gamma(\theta)$. It follows that uniformly in $\theta \in \Theta_{osn}$,

align*[align* omitted — 150 chars of source]

verifying Assumption (ref). Theorem (ref)(ii) implies that $\wh M_\alpha$ has asymptotically exact coverage.

\@startsection{paragraph}{4}{\z@} {0pt \@plus1ex \@minus.2ex} {-1em} {\normalfont}{Inference under identification:}

Now consider the case in which the model is identified (i.e. $\eta_2 = 1$ and $\tilde \gamma_{00} = 0$) and $M_I = \{\mu_0\}$. Here each $D_i = 1$ so the likelihood of the $i$-th observation $(D_i,Y_iD_i)=(1,y)$ is

align*[align* omitted — 160 chars of source]

Lemma (ref) in Appendix (ref) shows that with $\Theta$ as in ((ref)) and a flat prior, the posterior $\Pi_n$ concentrates on the local neighborhood $\Theta_{osn} = \{ \theta : |\tilde\gamma_{11}(\theta)-\tilde\gamma_{11}| \leq r_n/\sqrt n, \tilde\gamma_{00}(\theta) \leq r_n/n\}$ for any positive sequence $(r_n)_{n \in \mb N}$ with $r_n \to \infty$, $r_n/\sqrt n = o(1)$.

In this case, the reduced-form parameter is $\tilde\gamma_{11}(\theta)$ and the singular part is $\gamma_\perp(\theta) = \tilde \gamma_{00}(\theta) \geq 0$. Uniformly over $\Theta_{osn}$ we obtain: \[ n L_n(\theta) = \ell_n - \frac{1}{2} \frac{(\sqrt n (\tilde\gamma_{11}(\theta)-\tilde\gamma_{11}))^2}{\tilde\gamma_{11}(1-\tilde\gamma_{11})} + \frac{\sqrt n( \tilde \gamma_{11}(\theta)-\tilde\gamma_{11})}{{\tilde\gamma_{11}(1-\tilde\gamma_{11})}} \mb G_n(y) - n \tilde\gamma_{00}(\theta)+ o_\mb P(1) \] which verifies Assumption (ref)'(i) with

align*[align* omitted — 295 chars of source]

and $T = \mb R$. The remaining parts of Assumption (ref)' are easily shown to be satisfied. Therefore, Theorem (ref) implies that $\wh \Theta_\alpha$ for $\Theta_I$ will be asymptotically valid but conservative.

For inference on $M_I = \{\mu_0\}$, the profile LR statistic is asymptotically $\chi^2_1$ and equation ((ref)) holds with $f(v) = v^2$ and $T = \mb R$. To verify Assumption (ref), for each $\theta \in \Theta_{osn}$ we need to solve

align*[align* omitted — 229 chars of source]

at $\mu = \tilde\gamma_{11}(\theta)$ and $\mu = \tilde\gamma_{11}(\theta)+\tilde\gamma_{00}(\theta)$. The maximum is achieved when $g_{00}$ is as small as possible, i.e., when $g_{00} = \mu - g_{11}$. Substituting in and maximizing with respect to $g_{11}$: \[ \sup_{\eta \in H_\mu} \mb P_n \log p_{(\mu,\eta)} = \mb P_n \big(y \log \mu + (1-y) \log(1-\mu) \big)\,. \] Therefore, we obtain the following expansion uniformly for $\theta \in \Theta_{osn}$:

align*[align* omitted — 354 chars of source]

where the last equality holds because $\sup_{\theta \in \Theta_{osn}} \tilde\gamma_{00}(\theta) \leq r_n/n = o(n^{-1/2})$. This verifies that Assumption (ref) holds with $f(v) = v^2$. Thus Theorem (ref)(ii) implies that $\wh M_\alpha$ has asymptotically exact coverage for $M_I$, even though $\wh \Theta_\alpha$ is conservative for $\Theta_I$ in this case.

Example 2: entry game with correlated shocks in Subsection (ref)

Consider the bivariate discrete game with payoffs described in Subsection (ref). Here we consider a slightly more general setting, in which $Q_\rho$ denotes a general joint distribution (not just bivariate Gaussian) for $(\epsilon_1,\epsilon_2)$ indexed by a parameter $\rho$. This model falls into the class of models dealt with in Proposition (ref). Conditions (a)-(b) and (d) of Proposition (ref) hold with $\tilde \gamma(\theta) = (\tilde \gamma_{00}(\theta),\tilde \gamma_{10}(\theta),\tilde \gamma_{11}(\theta))'$ and $\wt \Gamma = \{\tilde \gamma(\theta) : \theta \in \Theta\}$ under very mild conditions on the parameterization $\theta \mapsto \tilde \gamma(\theta)$ (which, in turn, is determined by the specification of $Q_\rho$). Assumption (ref) is therefore satisfied with: \[ \mb I_{0} = \left[

array[array omitted — 134 chars of source]

\right] + \frac{1}{1-\tilde \gamma_{00}-\tilde \gamma_{10}-\tilde \gamma_{11}} \mf 1_{3 \times 3} \] where $\mf 1_{3 \times 3}$ denotes a $3 \times 3$ matrix of ones, \[ \sqrt n \hat \gamma_n = \mb V_n = \mb I_{0}^{-1/2} \mb G_n \left(

array[array omitted — 409 chars of source]

\right) \rightsquigarrow N(0,I_3) \] and $T = \mb R^3$. Condition (c) of Proposition (ref) and Assumption (ref) can be verified under mild conditions on the map $\theta \mapsto \tilde \gamma(\theta)$ and the prior $\Pi$. For instance, consider the parameterization $\theta = (\Delta_1,\Delta_2,\beta_1,\beta_2,\rho,s)$ where the joint distribution of $(\epsilon_{1},\epsilon_{2})$ is a bivariate Normal with mean zero, standard deviations one and positive correlation $\rho \in [0,1]$. The parameter space is \[ \Theta = \{ (\Delta_1, \Delta_2, \beta_1, \beta_2, \rho,s) \in \mb R^6 : \ul \Delta \leq \Delta_1, \Delta_2 \leq \ol \Delta , \ul \beta \leq \beta_1, \beta_2 \leq \ol \beta, 0 \leq \rho,s \leq 1\}\,. \] where $-\infty < \ul \Delta < \ol \Delta < 0$ and $-\infty < \ul \beta < \ol \beta < \infty$. The image measure $\Pi_\Gamma$ of a flat prior on $\Theta$ is positive and continuous on a neighborhood of the origin, which verifies Condition (c) of Proposition (ref) and Assumption (ref). Therefore, Theorem (ref)(ii) implies that our MC CSs for $\Theta_I$ will have asymptotically exact coverage.

Example 3: a moment inequality model

As a simple illustration, suppose that $\mu \in M = \mb R_+$ is identified by the inequality $\mb E[\mu - X_i] \leq 0$ where $X_1,\ldots,X_n$ are i.i.d. with unknown mean $\mu^* \in \mb R_+$ and unit variance. The identified set for $\mu$ is $M_I = [0,\mu^*]$, which is the argmax of the population criterion function $L(\mu) = -\frac{1}{2} (( \mu - \mu^*) \vee 0)^2$ (see Figure (ref)). The sample criterion $-\frac{1}{2 }((\mu - \bar X_n ) \vee 0)^2$ is typically used in the moment inequality literature but violates our Assumption (ref). However, we can rewrite the model as the moment equality model: $\mb E[ \mu + \eta - X_i] = 0$ where $\eta \in H = \mb R_+$ is a slackness parameter. The parameter space for $\theta = (\mu,\eta)$ is $\Theta = \mb R_+^2$. The identified set for $\theta$ is $\Theta_I = \{(\mu,\eta) \in \Theta : \mu + \eta = \mu^*\}$ and the identified set for $\mu$ is $M_I$ (see Figure (ref)). The GMM objective function is then: \[ L_n(\mu,\eta) = -\frac{1}{2 } (\mu + \eta - \bar X_n)^2\,. \] It is straightforward to show that $2nL_n(\hat \mu,\hat \eta) = -((\mb V_n + \sqrt n \mu^* ) \wedge 0)^2$ where $\mb V_n = \sqrt n (\bar X_n - \mu^*)$. Moreover, $\sup_{\eta \in H_\mu} 2nL_n (\mu,\eta) = -((\mb V_n + \sqrt n (\mu^* - \mu)) \wedge 0)^2$ and so the profile QLR for $M_I$ is $PQ_n(M_I) = (\mb V_n \wedge 0)^2 - ((\mb V_n + \sqrt n \mu^* ) \wedge 0)^2$.

figure[figure omitted — 2,347 chars of source]

For the posterior of the profile QLR, we also have $\Delta(\theta^b) = \{ \theta \in \Theta : \mu + \eta = \mu^b + \eta^b\}$ and $M(\theta^b) = [0,\mu^b + \eta^b]$. The profile QLR for $M(\theta^b)$ is \[ PQ_n(M(\theta^b)) = ((\mb V_n - \sqrt n (\mu^b + \eta^b - \mu^*)) \wedge 0)^2 - ((\mb V_n + \sqrt n \mu^* ) \wedge 0)^2 \] This maps into our framework with the local reduced-form parameter $\gamma(\theta) = \mu + \eta - \mu^*$. Consider the case $\mu^* \in (cn^{\alpha-1/2},\infty)$ where $c > 0$ and $\alpha \in (0,\frac{1}{2}]$ are positive constants (we consider this case for the moment just to illustrate verification of our conditions). Here $T = \mb R$ and a positive continuous prior on $\mu$ and $\eta$ induces a prior on $\gamma$ that is positive and continuous at the origin. Moreover, Assumption (ref) holds with $f(\kappa) = (\kappa \wedge 0)^2$. The regularity conditions of Theorem (ref) hold, and hence $\wh M_\alpha$ has asymptotically exact coverage for $M_I$.

More generally, Appendix (ref) shows that under very mild conditions our CS $\wh M_\alpha$ is uniformly valid over a class of DGPs $\mf P$, i.e.: \[ \liminf_{n \to \infty} \inf_{\mb P \in \mf P} \mb P( \mb M_I(\mb P) \subseteq \wh M_\alpha ) \geq \alpha \] where $M_I(\mb P) = [0,\mu^*(\mb P)]$ and the set $\mf P$ allows for any mean $\mu^*(\mb P) \in \mb R_+$ (encompassing, in particular, point-identified, partially identified, and drifting-to-point identified cases). In contrast, we construct sequences of DGPs $(\mr P_n)_{n \in \mb N} \subset \mf P$ along which bootstrap-based CSs $\wh M_\alpha^{boot}$ fail to cover with the prescribed coverage probability, i.e.: \[ \limsup_{n \to \infty} \mr P_n( \mb M_I(\mr P_n) \subseteq \wh M_\alpha^{boot} ) < \alpha\,. \] This reinforces the fact that our MC CSs for $M_I$ have very different asymptotic properties from bootstrap-based CSs for $M_I$.

Conclusion

We propose new methods for constructing CSs for identified sets in partially-identified econometric models. Our CSs are relatively simple to compute and have asymptotically valid frequentist coverage uniformly over a class of DGPs, including partially- and point- identified parametric likelihood and moment based models. We show that under a set of sufficient conditions, and in broad classes of models, our set coverage is asymptotically exact. We also show that in models with singularities (such as the missing data example), our MC CSs for $\Theta_I$ may be slightly conservative, but our MC CSs for identified sets of subvectors could still be asymptotically exact. Simulation experiments demonstrate the good finite-sample coverage properties of our proposed CS constructions in standard difficult situations. We also illustrate our proposed CSs in two realistic empirical examples.

There are numerous extensions we plan to address in the future. The first natural extension is to allow for semiparametric likelihood or moment based models involving unknown and possibly partially-identified nuisance functions. We think this paper's MC approach could be extended to the partially-identified sieve MLE based inference in CTT. A related, important extension is to allow for nonlinear structural models with latent state variables. Finally, we plan to study possibly misspecified and partially identified models.

{ \singlespacing }