EconBase
← Back to paper

Inference under partial identification with minimax test statistics

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

93,030 characters · 6 sections · 86 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference under partial identification with minimax test statistics

\address{Department of Economics and Finance, UNC Wilmington, Wilmington North Carolina 28403} \email{[email removed]} \subjclass[2000]{Primary 62G10; Secondary 62G20}

abstractWe provide a means of computing and estimating the asymptotic distributions of statistics based on an outer minimization of an inner maximization. Such test statistics, which arise frequently in moment models, are of special interest in providing hypothesis tests under partial identification. Under general conditions, we provide an asymptotic characterization of such test statistics using the minimax theorem, and a means of computing critical values using the bootstrap. Making some light regularity assumptions, our results augment several asymptotic approximations that have been provided for partially identified hypothesis tests, and extend them by mitigating their dependence on local linear approximations of the parameter space. These asymptotic results are generally simple to state and straightforward to compute (esp.\ adversarially).

Introduction

This paper is concerned with the computation and estimation of the asymptotic distribution of statistics that are based on minimax values. Such statistics are often of interest when one is interested in testing for the existence of a nonempty identified set, i.e.\ when hypothesis testing is conducted under the assumption of partial identification. When an identified set is characterized as the solution to a system of moment equations, the existence of a nonempty identified set can be consistently tested with adequate knowledge of how the moment functions will asymptotically behave thereon. We illustrate that test statistics designed to capture the behavior of these identifying relations often have a minimax formulation whose limiting distribution can be systematically computed, or at least bounded above, to form a hypothesis test.

Our results occur in the setting where one is concerned with minimizing a real-valued criterion function $\ell(\theta)$ over a parameter space $\Theta$. If $\ell$ can be estimated, consideration of the asymptotic distribution of the optimal value of the estimate in finite samples has been of interest since, at least, the $J$-test was proposed for overidentified GMM (Sargan1958, H1982). Lately, there has been extensive interest in extending this methodology to the partially identified setting, and fruitful work in both describing hypothesis tests (see S2012, CNS2023) and confidence regions (Tao2015, Zhu2020, Fan2023) therein. In practice, in order to test a given hypothesis, one often takes $\Theta$ to be a subset containing points of a larger, fixed parameter space which conform to that hypothesis. Our results can flexibly accommodate a range of parameter spaces, and can thereby facilitate tests of a range of analogous hypotheses. As an auxiliary exercise, we also analytically characterize the distance of certain random vectors from convex sets, such as those representing shape restrictions enforced on the random element (c.f.\ Fang2021).

As a point of departure, we note that $\ell$ can oftentimes be written as a supremum of a class of test functions $v$ indexed by $t \in \mathcal{T}$: \[ \ell(\theta) = \sup_{t \in \mathcal{T}} v(\theta, t). \] For instance, when $\ell$ is a norm or seminorm of a moment function of parameter $\theta$ over some Banach space, it always admits such a representation with linear $t$. This is true more generally if $\ell$ is any convex function of a moment function of $\theta$. Helpfully, such a decomposition will generally extend to empirical analogues of $\ell$, which will follow the corresponding form: \[ \ell_n(\theta) = \sup_{t \in \mathcal{T}} v_n(\theta, t), \] where $v_n(\theta, t)$ is an empirical analogue of test function $v$.

In order to compute the asymptotic distribution of a test statistic formed around the minimum value of $\ell_n(\theta)$ over $\Theta$, we take advantage of the fact that such a statistic will necessarily adopt a minimax formulation---the outer minimum being taken over parameter space $\Theta$, and the inner maximization being over its dual space $\mathcal{T}$. Such a representation is especially amenable to treatment by familiar tools arising from convex analysis. In particular, we show that an application of Sion's celebrated minimax theorem (Sion1958) leads to a streamlined way of bounding the distribution of such minimax test statistics from above, with asymptotic equality under certain rate or convexity conditions. Under some reasonable regularity conditions, these minimax bounds coincide with asymptotic approximations arising from S2012, hong2017, and CNS2023, and are closely tied to convex-analytical approaches to hypothesis testing (e.g.\ the test of shape restrictions in Fang2021). They provide a systematic basis for regarding such results which connects them back to the seminal work of Sargan1958 and H1982.

The intuition for our results in a simple, finite-dimensional setting can be stated as follows. If, say, $\sqrt{n}(v_n(\theta, t) -v(\theta, t))$ is an empirical process $\mathbb{G}_n(\theta, t)$ over $\Theta \times \mathcal{T}$ that converges to a tight limit $\mathbb{G}$, one can show by Taylor's theorem that the minimized value of $\ell$ over $\Theta$ may be rewritten as:

align[align omitted — 334 chars of source]

where we have employed in the last line the fact that a typical solution to the optimization problem will converge in probability towards $\Theta_0$. In the previous display, $(c_n)$ is a sequence diverging to $\infty$.

Equipped with (ref), previous work has provided means for estimating the image of the Jacobian $\frac{\partial v}{\partial \theta}$, which may be substituted with a Fr\'{e}chet derivative in infinite dimensional settings (S2012, hong2017, Fan2023, CNS2023). As the limiting distribution of $\mathbb{G}_n(\theta, t)$ can usually be consistently estimated with, say, the bootstrap, this provides a convenient means of estimating the asymptotic distribution of a minimax test statistic. However, this approach is disadvantaged by its reliance on the local linearity of $v$ in neighborhoods of the identified set. For instance, S2012 studies a setting in which $v$ is linear in the choice of parameter $\theta$. This assumption is relaxed in hong2017 and CNS2023, but finite sample analysis of the derivatives of $v$ is still e.g.\ generally dependent upon the use of linear sieve approximations of $\Theta$. Thus, the flexible employment of neural networks and associated nonlinear estimated schemes, which is highly desirable in nonparametric settings, is precluded.

This paper observes that there is a preponderance of $\ell$ and $v$ in the literature which satisfy the necessary convexity properties to rewrite (ref) as

align[align omitted — 324 chars of source]

As $c_n$ is a diverging sequence, asymptotically, the parameter $t$ in the supremum of (ref) will be taken to satisfy $\frac{\partial v(\theta, t)}{\partial \theta} =0$. As a consequence, we show that under general conditions, (ref) may be rewritten as

align[align omitted — 202 chars of source]

(ref) compares favorably to (ref) in several ways. In many popular Hilbert space settings, it shows that a test statistic based on (ref) has a projection interpretation which extends the $\chi^2$-limiting distribution of the $J$-test. Moreover, we show that it is quite straightforward to obtain estimates for critical values using (ref) without any direct computation of the derivatives of $\frac{\partial v}{\partial \theta}$ at all, even in settings with potentially nonlinear sieves. The minimax approach also allows for approximation methods devised for seminorm-based criterion functions to be extended quite broadly and systematically to, say, criterion functions involving general convex functions. Finally, our inference results accommodate the adversarial employment of nonconvex sieve spaces of test functions, such as neural networks. Thus, they blend a growing literature on adversarial econometric estimation (see K2023 and references therin) with inference in the partially identified setting.

The structure of the paper is as follows. Section (ref) formalizes our model and provides a method for obtaining the asymptotic distribution, or upper bounds thereof, of minimax test statistics. Section (ref) shows that, under general conditions, critical values can be obtained for those distributions using the bootstrap. The main focus of the paper is on single hypothesis tests, but Section (ref) of the appendix extends our distributional results uniformly over a class of parameter spaces and underlying probability distributions. Proofs are relegated to Section (ref) in the appendix.

Model and Asymptotic Distribution

Given a parameter space $\Theta$, our basic model is one where an identified set $\Theta_0 \subset \Theta$ is characterized as the set of all $\theta \in \Theta$ satisfying

align[align omitted — 47 chars of source]

where $\ell: \Theta \rightarrow \mathbb{R}$ is some function. We impose the additional restriction that $\ell$ takes the form

align[align omitted — 80 chars of source]

where $\{v(\cdot, t) : t \in \mathcal{T}\}$ is a set of test functions, and $\mathcal{T}$ is some set dual to $\Theta$, motivated by the idea that one often wants to investigate sets identified by criterion functions of this form:

exm\begin{enumerate} • With GMM, one investigates an identified set which can be expressed as the set of all $\theta$ which solve \begin{align*} \ell(\theta) \equiv \norm{\mathrm{E}\left[ g(Y, \theta) \right]}_W = 0, \end{align*} where $\rho$ takes on values in $\mathbb{R}^m$, $W$ is a weighting matrix, and $\norm{v}_W = v' W v$. In this case, one can write \begin{align*} \ell(\theta) = \sup_{t \in \mathcal{T}} v(\theta, t) \end{align*} by setting $v(\theta, t) = \left\langle \mathrm{E}\left[ g(Y, \theta), t \right] \right\rangle$ and $\mathcal{T} =\{t: \norm{t}_W \le 1\}$. • In a general formulation of GMM, one can consider sets of $\theta$ identified by the relation \begin{align*} \ell(\theta) = q(m(\theta)) = 0, \end{align*} where $m$ takes values in a vector space $\mathfrak{X}$ over $\mathbb{R}$ (e.g.\ a Banach space) and $q: \mathfrak{X} \rightarrow \mathbb{R}$ is a sublinear functional, such as a norm or a seminorm (see C1994). In this case, the Hahn-Banach theorem implies that we may write \begin{align*} \ell(\theta) = \sup_{t \in \mathcal{T}} v(\theta, t), \end{align*} where $v(\theta, t): (t, \theta) \mapsto t(m(\theta))$ and $\mathcal{T}$ is the set of linear functionals over $\mathfrak{X}$ which are continuous and bounded above by $q$, i.e.\ \begin{align*} \mathcal{T} = \{t: \mathfrak{X} \rightarrow \mathbb{R}: \, t is linear, t(x) \le q(x) for all x\in \mathfrak{X}\} \end{align*} • A range of models identify parameters $\theta \in \Theta_0$ by virtue of the fact that they satisfy a conditional moment equality \begin{align} \mathrm{E}\left[ g(Y, \theta)|Z \right] = 0. \end{align} for e.g.\ an instrumental variable $Z$. For instance, in a standard nonparametric instrumental variables model (NPIV, see e.g.\ NP2003, S2012), points $\theta$ in $\Theta_0$ are precisely which satisfy \begin{align*} \mathrm{E}\left[ Y_2 - \theta(Y_1) | Z \right] = 0, \end{align*} where $Y_2$ is an outcome variable, $Y_1$ is a regressor, and $Z$ is an instrumental variable complete for $Y_1$. To verify the identifying relation (ref), one can let $\mathcal{T}$ be a set of test functions in variable $Z$ (S2012 takes these to be a compact set of complex exponential functions, and CNS2023 allows more general choices) and set \begin{align} v(\theta, t) = \mathrm{E}\left[ g(Y, \theta) t(Z) \right]. \end{align} Note that, in the NPIV case, the function $v(\theta, t)$ is linear in both $\theta$ and $t$. Let $\mathrm{P}$ be the law of variable $Z$. Let $m(\theta)$ be the function $z \mapsto \mathrm{E}\left[ g(Y, \theta)|Z \right]$, and suppose that this function is in the Banach space $L^p(\mathrm{P})$, $p \in [1, \infty]$. If $q$ is conjugate to $p$ in the sense that $\frac{1}{p} + \frac{1}{q} = 1$, and $t$ (with abuse of notation) signifies the linear mapping $t(\psi)= \mathrm{E}\left[ \psi(Z) t(Z) \right]$ for $\psi \in L^p(\mathrm{P})$, then the relation \begin{align*} v(\theta, t) = t(m(\theta)) \end{align*} is true and is consistent with the following examples. Our standing model encompasses generalizations of (ref) employing moment inequalities, as we show next. • If $m$ maps $\Theta$ to a subset of a Hausdorff and locally convex vector space $\mathfrak{X}$ over $\mathbb{R}$ and $\ell = \gamma \circ m$ where $\gamma: \mathfrak{X} \rightarrow \mathbb{R}$ is any proper convex and lower semi-continuous function, then the Fenchel-Moreau theorem (Convex2005, Theorem 4.2.1) implies that we may write \begin{align*} \ell(\theta) = \sup_{t \in \mathcal{T}} v(\theta,t) \end{align*} where $v: (t,\theta) \mapsto t(m(\theta))$, and $\mathcal{T}$ is the set of affine functions from $\Theta$ to $\mathbb{R}$ that are dominated by $\gamma$, i.e.\ \begin{align*} \mathcal{T} = \{ t: \mathcal{X} \rightarrow \mathbb{R}: \, t affine, t(x) \le \gamma(x) for all x \in \mathfrak{X}\}. \end{align*} Lemma (ref) below shows that this formulation of our inference problem is consistent with the main assumptions of the paper. \begin{itemize} • For instance, let $\mathfrak{X}$ be a separable real Hilbert space and $C\subset \mathfrak{X}$ a convex set. Then the distance \begin{align*} \ell(\theta) = d(m(\theta), C) = \inf_{c \in C} \norm{m(\theta) - c} \end{align*} can be expressed as \begin{align*} \ell(\theta) = \sup_{\norm{t} \le 1} (\left\langle m(\theta), t \right\rangle - \sup_{c \in C} \left\langle c, t \right\rangle). \end{align*} (e.g.\ Deutsch2012, \S 7.). One may equivalently express the inner supremum in the previous display as being over the convex set of affine functions which are nonpositive over $C$, omitting the $\sup_{c \in C}$ term. When $C$ is additionally a convex cone, we have the relation \begin{align} \ell(\theta) = \sup_{\substack{\norm{t} \le 1 \\ \sup_{c \in C} \left\langle c,t \right\rangle \le 0}} \left\langle m(\theta), t \right\rangle. \end{align} Fang2021 propose a method for inference on the distance from a noisily observed parameter to a convex cone and also apply (ref), albeit under point identification. Section (ref) shortly motivates our main results in this setting. • Suppose $\mathfrak{X}$ is a vector space of functions over a domain $\mathcal{Z}$ and $C$ is the subset of $\mathfrak{X}$ consisting of (a.s.) positive functions. For instance, $\mathfrak{X}$ may be the space $L^p(\mathrm{P})$ of $L^p$-integrable functions in variable $Z$, containing functions of the form $m(\theta) = \mathrm{E}\left[ g(Y, \theta)|Z = z \right]$, as in (ref). In this case, one may set $\ell(\theta) = d(m(\theta), C)$, and $\Theta_0$ is the set of $\theta$ which satisfy the identifying relation $\mathrm{E}\left[ g(Y, \theta)|Z \right] \overset{\mathrm{a.s.}}{\ge} 0$, which extends (ref). The problem of conducting inference in models identified by conditional moment inequalities has been studied by andrews2013, csr2013, armstrong2016, and chet2018, among others. Our model allows for such conditional inequalities to be generalized to arbitrary convex restrictions. \end{itemize} • Suppose that $g$ maps $\Theta$ to a random distribution $m(\theta)$ supported on measurable space $(\mathfrak{X}, \mathcal{F})$, which admits some reference measure $\mu$. Then, one can metrize a host of notions of convergence for probability measures using (ref). \begin{itemize} • If one takes $\mathcal{T} = \{f: \mathfrak{X} \rightarrow [-1,1]\}$ and \begin{align} v(\theta, t) = t(m(\theta)) - t(\mu), \end{align} then $\ell$ is the total variation distance between $m(\theta)$ and $\mu$ \begin{align*} \ell(\theta) = \sup_{t \in \mathcal{T}} \int t \, \mathrm{d}m(\theta) - \int t \, \mathrm{d}\mu = \norm{m(\theta) - \mu}_{\mathrm{TV}}. \end{align*} • If $\mathfrak{X}$ is the Euclidean space $\mathbb{R}^m$, then letting $\mathcal{T}$ be a collection of maps $t: \nu \mapsto \int e^{isx} \, \mathrm{d} \nu$ and $v$ be as in (ref), one can check for the distance between $m(\theta)$ and $\mu$ in a weak topology (see S2012, D2010 \S 3.3.2) using L\'{e}vy's continuity theorem. • If $(\mathfrak{X},d)$ is a metric space and $\mu$ is a probability measure, it is also possible to metrize the Wasserstein-$1$ distance between $m(\theta)$ and $\mu$ using (ref) and the Kantorovich-Rubinstein identity (Ar2017). To do this one, lets $\mathcal{T}$ be the set of $1$-Lipschitz maps $t: \mathfrak{X} \rightarrow \mathbb{R}$ and again applies (ref). \end{itemize} \end{enumerate}

Note that, in every case of example (ref), all of the functions $v$ are linear in the parameter $t$.

Suppose that $\ell$ cannot be directly observed, but a researcher has access to a noisy estimate $\ell_n$ of $\ell$ which satisfies

align*[align* omitted — 71 chars of source]

where $v_n$ is a noisy estimate of $v$. Our approach to inference will consist of finding estimates for the asymptotic distribution of test statistic

align[align omitted — 156 chars of source]

for an appropriate normalizing sequence $r_n$ (most often $\sqrt{n}$), and then determining critical values for this distribution by bootstrapping.

Motivating example: distances to convex sets in Banach spaces

Our minimax approach allows for the straightforward calculation of asymptotic distributions of distances to convex sets. Fang2021 characterize and estimate such asymptotic distributions when the sets are convex cones in Hilbert spaces with the intent of testing shape restrictions. The distributions which arise are continuous functions of (usually) tight Borel measures which may be directly estimated with the bootstrap.

In a Banach space $\mathfrak{X}$, the function that assigns to every point its distance to some fixed convex set $C$ is convex, and Example (ref) suggests that it is compatible with the setting of the paper. Indeed, let $m(\theta)$ be an element of a Banach space $\mathfrak{X}$ for every $\theta \in \Theta$ and $C \subset \mathfrak{X}$ a convex set. Then, letting $\ell(\theta)$ denote the distance of $m(\theta)$ to $C$ and $\mathcal{T}$ the unit ball of the dual $\mathfrak{X}^*$ of $\mathfrak{X}$, equipped with the weak-* topology, one can write

align*[align* omitted — 170 chars of source]

Suppose that $m_n(\theta)$ is a random estimate of $m(\theta)$ satisfying that $r_n(m_n(\theta) - m(\theta)) \rightsquigarrow \mathbb{G}(\theta)$ for a sequence $(r_n)$, where `$\rightsquigarrow$' signifies weak convergence to a random element $\mathbb{G}(\theta)$. Then, an empirical analogue of $\ell(\theta)$ is

align[align omitted — 192 chars of source]

Fang2021 principally provide means of computing the limiting distribution of $r_n \ell_n(\theta)$ for a fixed $\theta$. To motivate the results of the paper, we show that we can provide such computations in a very broad range of cases.

The easiest case to consider is the one where $C$ is a convex cone, closed under dilations of its elements by positive multiples. Let $\delta^*(\cdot|C):t \mapsto \sup_{c \in C} t(c) $ denote the support function of $C$ (see Rock1970, \S13). For a convex cone $C$ in a real Banach space $\mathfrak{X}$, define the polar $C^\circ \subset \mathfrak{X}^*$ following fabian2011 as

align*[align* omitted — 153 chars of source]

(as opposed to the absolute polar, e.g.\ C1994). Then, using weak-* continuity of the maps $t \in \mathcal{T}$ and weak-* compactness of $\mathfrak{X}^*$, the Sion minimax theorem (Sion1958, see also the proof of Lemma (ref) below) allows us to rewrite (ref) as

align[align omitted — 505 chars of source]

Note that, if $m(\theta)$ is actually an element of $C$, the last line of (ref) is bounded above by

align[align omitted — 265 chars of source]

(see Lemma (ref) in the appendix; throughout, we will find it convenient to use e.g.\ $d_{\mathfrak{X}}$ as the Hausdorff distance in $\mathfrak{X}$ between a point and a set). (ref) and (ref) hold irrespective of the underlying probability distribution $P$ of $m_n$ or of $\theta$. Therefore, from these bounds arises our first motivating result:

propLet $\mathcal{P}$ index a family of probability distributions $P$, $C \subset \mathfrak{X}$ be some convex cone, and $\Theta$ some parameter space. Suppose that $\mathfrak{X}$-valued random elements $\mathbb{G}_{n,P}(\theta) \equiv r_n(m_{n,P}(\theta) - m_P(\theta))$ converge uniformly in $P$ to limiting distributions $\mathbb{G}_P(\theta)$ in the sense\footnote{This is uniform convergence in the bounded Lipschitz metric; see VW1996 \S2.8 and also Section (ref) below.} that $\limsup_{n \rightarrow \infty} \sup_{\substack{ P \in \mathcal{P} \\ h \in \mathrm{BL}_1}} \mathrm{E}_{P}\left[ h( \mathbb{G}_{n,P}) \right] - \mathrm{E}\left[ h( \mathbb{G}_{P}) \right] = 0$, where $\mathrm{BL}_1$ is the set of maps bounded by $1$ that are Lipschitz from the set of maps from $\Theta$ to $\mathfrak{X}$, equipped with the uniform norm, to $\mathbb{R}$. Suppose that $m_P(\theta) \in C$ for all $P$ and $\theta$. Then, \begin{align} \inf_{\theta \in \Theta} r_n d_{\mathfrak{X}}(m_{n,P}(\theta), C) &= \inf_{\theta \in \Theta} \sup_{\substack{t \in C^\circ \\ \norm{t}_{\mathfrak{X}^*} \le 1}} (\left\langle \mathbb{G}_{n,P}(\theta), t \right\rangle + r_n\left\langle m_P(\theta), t \right\rangle) \le \inf_{\theta \in \Theta} d_{\mathfrak{X}}(\mathbb{G}_{n,P}(\theta), C ) \nonumber \\ & \rightsquigarrow \inf_{\theta \in \Theta} d_{\mathfrak{X}}(\mathbb{G}_{P}(\theta), C ), \end{align} where `$\rightsquigarrow$' signifies uniform convergence over $P$ in the bounded Lipschitz sense.
proofThe result is a direct consequence of (ref) and (ref), and the observation that $G \mapsto \inf_{\theta \in \Theta} \sup_{\substack{t \in C^\circ \\ \norm{t}_{\mathfrak{X}^*} \le 1}} \left\langle G(\theta), t \right\rangle$ is bounded and Lipschitz from the space of maps from $\Theta$ to $\mathfrak{X}$ (equipped with the uniform norm) to $\mathbb{R}$ (of course, any other Lipschitz continuous map in $\theta$ may be employed in place of $\inf_{\theta \in \Theta}$).

Proposition (ref) complements and extends Theorem 3.1 of Fang2021 in a Banach space setting featuring arbitrary parameter spaces. In particular, note that the inequality in the first line of (ref) becomes an equality precisely in the `least favorable' case discussed therein. This paper further develops the minimax rearrangement displayed in (ref) and provides asymptotic characterizations of a much broader class of test statistics, as well as means of estimating those limiting distributions. For instance, Theorem (ref) below allows (ref) to be further decomposed using the structure of $\Theta$ as it relates to $C$.

The next proposition provides a similar characterization when $C$ is more generally held to be a convex set in $\mathfrak{X}$. For such a $C$ and a point $c \in C$, define the tangent cone $T_cC$ of $C$ at $c$ to be the subset of $\mathfrak{X}$ consisting of all points taking the form $\alpha (c' - c)$, $c' \in C$, $\alpha > 0$ (c.f.\ Rock1970, Theorem I.2.6.3). Similarly, define the normal cone $N_cC$ of $C$ at $c$ to be the subset $(T_cC)^\circ$ of $\mathfrak{X}^*$. For conciseness, we say that a class of random sequences $(X(\theta, P))$ indexed by $\theta \in \Theta, P \in \mathcal{P}$ is uniformly pre-tight in $\mathfrak{X}$ if, for every $\varepsilon > 0$, there is some totally bounded and measurable subset $S$ of $\mathfrak{X}$ which satisfies $\inf_{P \in \mathcal{P}} \mathrm{P}_{P}\left( X(\theta, P) \in S, \, \forall \theta \in \Theta \right) > 1 - \varepsilon$. Sufficient conditions for uniform pre-tightness are given in VW1996 (e.g.\ Problem 1.12.1 therein).

propMake the Assumptions of Proposition (ref) with $C$ a convex set. Suppose that the collection $\{\mathbb{G}_{P}(\theta): \theta \in \Theta, P \in \mathcal{P}\}$ is uniformly pre-tight in $\mathfrak{X}$. Impose that there is some topology $\mathcal{W}$ on $\mathcal{P} \times \Theta$ such that $(\mathcal{P}\times \Theta, \mathcal{W})$ is compact, $(P, \theta) \mapsto \mathcal{T} \cap N_{m_P(\theta)}C $ is $\mathcal{W}$-lower hemicontinuous, and $(P, \theta, t) \mapsto \left\langle m_P(\theta), t \right\rangle$ is $\mathcal{W} \times \mathcal{U}$-upper semicontinuous. Then, uniformly for $P \in \mathcal{P}$, \begin{align} r_n \inf_{\theta \in \Theta} d_{\mathfrak{X}} (m_{n,P}(\theta), C) \rightsquigarrow \inf_{\theta \in \Theta} \sup_{\substack{t \in N_{m_P(\theta)} C \\ \norm{t}_{\mathfrak{X}^*} \le 1}} \left\langle \mathbb{G}_P(\theta), t \right\rangle = \inf_{\theta \in \Theta} d_{\mathfrak{X}} (\mathbb{G}_P(\theta), T_{m_p(\theta)}C). \end{align}

In the statement of Proposition (ref), the map $\inf_{\theta \in \Theta}$ may be replaced by any other function that is Lipschitz continuous with respect to the uniform norm. When $\Theta$ is a singleton and the conditions of Proposition (ref) apply, (ref) forms a tighter lower bound than the right hand side of (ref), because $N_{m_P}C$ is the subset of $C^\circ$ consisting of $t$ for which $\left\langle m_P, t \right\rangle = 0$ when $C$ is a cone.

The lower hemicontinuity condition imposed by Proposition (ref) is important for a portion of our bootstrap consistency arguments and is further discussed in Section (ref). Lemma (ref) therein shows that the lower hemicontinuity condition is fulfilled if the correspondence sending pairs $(P,\theta)$ to the tangent cone $T_{m_P(\theta)}C \subset \mathfrak{X}$ is merely upper hemicontinuous with respect to the weak topology on $\mathfrak{X}$. On the other hand, by boundedness of the maps $t \in \mathcal{T}$, continuity of the map $(P, \theta, t) \mapsto \left\langle m_P(\theta), t \right\rangle$ may be ascertained by verifying norm-continuity of the map $(P,\theta, t) \mapsto m_P(\theta)$, or by confining the points $m_P(\theta)$ to lie in a norm-compact set and requiring $(P,\theta, t) \mapsto m_P(\theta)$ to be merely weakly continuous.\footnote{Sufficiency of the latter conditions can be verified using the Arzel\`{a} Ascoli theorem; see the discussion following Assumption (ref).}

Asymptotic distribution

To discuss inference around $T_n$, it is necessary to impose some structure on the spaces $\Theta$ and $\mathcal{T}$. We continue to let $\Theta_0 \subset \Theta$ be the set of $\theta$ satisfying (ref), and say that $\Theta$ is locally convex in a neighborhood of $\Theta_0$ if it is a subset of a real vector space and, for all $\theta \in \Theta_0$, there is some $\varepsilon(\theta) > 0$ such that the intersection of $\Theta$ with an $\varepsilon(\theta)$-ball around $\theta$ is convex. All Banach spaces concerned in this paper are presumed to be over the real numbers, although it would be straightforward to extend our results to vector spaces over $\mathbb{C}$.

asm$\Theta$ is a compact subset of a real Banach space $\mathfrak{B}$ with norm $\norm{\cdot}_{\mathfrak{B}}$, and is locally convex in a neighborhood of $\Theta_0$. $\mathcal{T}$ is a convex subset of a linear space. The function $\ell$ follows (ref), maps $\Theta \rightarrow \mathbb{R}_+$, and is lower semicontinuous.

Assumption (ref) imposes some regularity on $\Theta$ and $\mathcal{T}$ by requiring that they are convex subsets of linear topological spaces (at least in a neighborhood of the identified set). We have remarked that, when the asymptotic distribution of a test statistic following (ref) is of interest, the set $\Theta$ is often chosen relative to a hypothesis that is imposed on a larger parameter space. Therefore, it is essential that Assumption (ref) should be made as accommodating in its choice of $\Theta$ as possible. Firstly, we note that what is actually required is that $\Theta$ is star-shaped in a neighborhood of all points in the identified set, so that derivatives can be defined and considered on those neighborhoods (but convexity is hardly more onerous an assumption). The requirement for local convexity can also be achieved by choosing an appropriate embedding of $\Theta$ into a Banach space, as the following example illustrates:

exmSuppose that $\Theta$ is a compact subset of a Banach manifold $M$ (Abraham2012). By definition, there is a collection of charts $(U_i, \phi_i)$, indexed by $i \in I$ for which one can write $\Theta = \bigcup_{i \in I} U_i$. Here, each $U_i$ is open and $\phi_i: U_i \rightarrow E_i$ is a homeomorphism from $\Theta$ to Banach space $E_i$ for each $i$. By compactness, there is a finite subset $I_0 \subset I$ of charts satisfying that $\Theta$ is contained in the union of $U_i, i \in I_0$. We may also assume that $\phi(U_i)$ is bounded for every $i$, and therefore that $\phi_i(U_i)$ does not contain the origin for all $i \in I_0$. Let $P_i$ be the projection map from $E_i$ to the Banach space $\mathfrak{B} = \sum_{j \in I_0} E_j$ which sends $x$ to the point in having $x$ in its $i^\text{th}$ coordinate and $0$ elsewhere.\footnote{Here, $\sum_{j \in I_0} E_j$ signifies the direct sum of the Banach spaces $E_j$, which is a Banach space under e.g.\ the $L^1$ norm given by the sum of the individual $E_i$-norms.} We claim that a sufficient condition for the local convexity condition to hold for a suitable analogue of $\Theta$ in a Banach space is that the charts $\phi_i$ can be chosen in such a way that $\phi_i (\Theta \cap U_i)$ is locally convex at every point of $\phi_i(\Theta_0 \cap U_i)$ for all $i \in I_0$. Indeed, if this is the case, we may identify $\Theta$ with the set $\tilde{\Theta} = \bigsqcup_{i \in I_0} P_i \phi_i(\Theta \cap U_i) \subset \mathfrak{B}$. If we let $\rho: \tilde{\Theta} \rightarrow \Theta$ be a map which satisfies that $\rho \circ \phi_i = \mathrm{Id}_{U_i}$ for all $i$, then (ref) can be rewritten \begin{align*} T_n = r_n \inf_{\theta \in \tilde{\Theta}} \ell_n(\rho(\theta)). \end{align*} Moreover, the set identified by $\ell(\rho(\theta)) = 0$ is precisely $\bigsqcup_{i \in I_0} \phi_i (\Theta_0 \cap U_i)$, every point of which is contained in a ball on which $\phi_i(\Theta \cap U_i)$ is convex.

We also require that our parameter space $\Theta$ is a compact subset of a Banach space, over which we can eventually define derivatives. See FM2019 for a discussion of compactness in Banach spaces. Note that, if $\Theta$ is not necessarily convex, the closed convex hull of $\Theta$ is still compact (A1999, Theorem 5.35), so it may suffice to consider the convex closure of $\Theta$. It is also often the case that the map $v(\theta, \cdot)$ is at least quasi-concave in $t$ (in fact, every case of Example (ref) featured $v$ that were linear over parameter $t$), in which case it is equivalent to work with the convex hull of $\mathcal{T}$ in (ref). Note that the Donsker properties we require in Assumption (ref) are preserved under such set operations (VW1996 \S2.10).

It would likely be possible to extend the results of this paper to noncompact settings using bounded entropy conditions and finite sample concentration inequalities coming from empirical process theory (e.g.\ CNS2023, Assumption 3.3 and references therein). However, imposing compactness greatly streamlines the exposition and proofs of this paper, and is congruous with assumptions that are commonly imposed for the purposes of estimation and inference.

Throughout, we will let $d_{\mathfrak{B}}$ denote the metric induced by $\norm{\cdot}_{\mathfrak{B}}$ on $\mathfrak{B}$. The requirements that $\ell$ is lower semicontinuous and nonnegative are not very restrictive. In many examples, $\ell$ can be expected to be continuous in parameter $\theta$. For instance, this is the case when $\ell(\theta)$ is some seminorm of a continuous function $m$ of $\theta$.

In order to estimate the identified set $\Theta_0$, one must introduce an estimator $\ell_n$ of the criterion function $\ell$. Suppose that this is accomplished by introducing an estimator $v_n$ of the functions $v$, which obeys a central limit theorem in the sense of empirical process theory (see VW1996, Section 1.5):

asmFor some sequence $(r_n) \rightarrow \infty$, the empirical process $\mathbb{G}_n(\theta, t) \equiv r_n(v_n(\theta, t) - v(\theta, t))$ converges weakly to a tight Borel measurable element $\mathbb{G}$ in $L^\infty(\Theta, \mathcal{T})$, i.e.\ \begin{align*} \mathbb{G}_n(\theta, t) \rightsquigarrow \mathbb{G}(\theta, t), \end{align*} Moreover, $\mathbb{G}_n$ is asymptotically equicontinuous with respect to a pseudometric $\rho: (\Theta \times \mathcal{T})^2 \rightarrow \mathbb{R}_+$.

Most often, one takes $r_n = \sqrt{n}$, and $\mathbb{G}$ is a Gaussian process. There are a number of ways of guaranteeing the convergence of Assumption (ref). For instance, if one has $v(\theta, t) = \mathrm{E}\left[ g(Y,\theta,t) \right]$ for some function $g$ and $v_n(\theta, t)$ is an empirical analogue of $v$, then Assumption (ref) holds if the class of maps $g$ is $P$-Donsker over $\Theta \times \mathcal{T}$ (VW1996, \S2). The delta-method for empirical processes extends such results to more general settings where $v$ is a nonlinear function of population moments, and $v_n$ its empirical analogue. In this setting, $\rho$ can be taken to be the pseudometric induced by seminorm $(\theta, t) \mapsto \mathrm{E}\left[ (g(Y, \theta, t) - \mathrm{E}\left[ g(Y, \theta, t) \right])^p \right]theta, t) \right])^p}^{1/p}$, for $p \ge 1$ (VW1996 Sections 1.5, 2.1).

We now wish to approximate the functions $v(\cdot, t)$ with linear approximations taken at points in the identified $\Theta_0$. A standard tool for this purpose is the Fr\'{e}chet derivative $D$, although we allow for a slight relaxation of Fr\'{e}chet differentiability in the following assumption. We say that a function $D$ defined on a real vector space $\mathfrak{B}$ is positive homogeneous if $D(\lambda h) = \lambda D(h)$, for all $h \in \mathfrak{B}$ and $\lambda \ge 0$.

asmFor all $(\theta, t) \in \Theta_0 \times \mathcal{T}$ there is a lower semicontinuous, convex, and positive homogeneous functional $D_{\theta; t}v: \mathfrak{B} \rightarrow \mathbb{R} $ and function $f_v(\delta) = o(\delta)$ which satisfy \begin{enumerate} • \begin{align} \sup_{\substack{\theta \in \Theta_0, \theta' \in \Theta\\ t \in \mathcal{T} \\ \norm{ \theta' - \theta}_{\mathfrak{B}} \le \delta}} |v(\theta',t) - v(\theta, t) - D_{\theta; t}v (\theta' - \theta)| = O(f_v(\delta)), \end{align} • For all $n \in \mathbb{N}$, $\theta \in \Theta_0$, and $\theta' \in \Theta$, $t \mapsto v_n(\theta, t) + D_{\theta; t}v (\theta')$ is almost surely quasi-concave \end{enumerate}

The first part of Assumption (ref) requires that the functions $v(\cdot, t)$ can be approximated locally by positive homogeneous functionals $D_{\theta; t}$ in neighborhoods of points $\theta$ in the identified set. When $D_{\theta; t}$ is required to be bounded and linear, Assumption (ref) is principally the requirement that function $v$ is Fr\'{e}chet differentiable for all $\theta \in \Theta_0$ and $t \in \mathcal{T}$, with a uniform bound on the linear approximation provided by the derivative. Note that, if the left hand side of (ref) is $o(\delta)$, then (ref) trivially holds by defining $f_v(\delta)$ to be equal to its left hand side.

By Taylor's theorem, if $\Theta$ and $\mathcal{T}$ are subsets of Euclidean space and $v$ is smooth, Assumption (ref).1 may be established by uniformly bounding the second derivatives of $v(\cdot, t)$ over $\Theta_0$ and $\mathcal{T}$. In this case, one may take $f_v(\delta) = \delta^2$. Note that, if $\Theta_0$ is a singleton $\{\theta_0\}$ (point identification holds), Assumption (ref).1 is implied by Fr\'{e}chet differentiability at $\theta_0$ of the map $\theta \mapsto v(\theta, \cdot)$ from $\Theta$ to the set of bounded maps from $\mathcal{T}$ to $\mathbb{R}$ equipped with uniform norm.

The second part of Assumption (ref) appears onerous, but we stress that in all of the situations discussed in Example (ref), the function $v$ is linear over the parameter $t$ for fixed $\theta$. For $t, t' \in \mathcal{T}$ and $\alpha \in [0,1]$, passing to limits with Assumption (ref).1 and linearity of $v(\theta, \cdot)$ in hand allows one to write

align*[align* omitted — 104 chars of source]

Hence, it may be reasoned that, in many settings, $D_{\theta; t}v (h)$ is actually linear in parameter $t$, whence concave. The same applies to $v_n(\theta,t)$, which is often itself linear in $t$ (especially when $v(\theta, t)$ is linear in $t$). In particular, if we are in one of the situations delineated by Example (ref):

align*[align* omitted — 41 chars of source]

for some function $m$, then we may take $v_n(\theta, t) = t(m_n(\theta))$, where $m_n$ converges to $m$ as an empirical process. This makes $v_n$, and the whole expression $v_n(\theta, t) + D_{\theta; t} v(\theta')$, linear (whence, concave) in $t$.

Our final main assumption is placed on the derivative invoked in Assumption (ref). We are interested in the magnitude of the derivative, measured as how much it dilates points close to $0$. One may impose the usual dual norm on the functional $D_{\theta; t}v$ by letting $\norm{D_{\theta; t}v} = \sup_{\norm{h}_{\mathfrak{B}} \le 1} D_{\theta; t}v (h)$, although this norm might be stronger than necessary. Instead, we are concerned with the behavior of $D_{\theta; t}v$ only on the space tangent to $\Theta$ at $\theta$, and only in the negative direction. Formally, for each $\theta$ and $t$, we define

align[align omitted — 220 chars of source]

If $\Theta$ is convex in a neighborhood of $\theta$, it is straightforward to show that the term inside the limit is nonincreasing as $\delta \rightarrow 0$, so the limit must exist. Indeed, under the preceding assumptions it is immediate (see the proof of Theorem (ref)) that $\gamma(\theta, t)$ is the decreasing limit of $\inf_{\substack{\theta' \in \Theta \\ \norm{\theta' - \theta}_{\mathfrak{B}} \le \delta }} \delta^{-1} D_{\theta; t}v (\theta' - \theta)$ for $\delta < \underline{\delta}$, where $\underline{\delta} > 0$ can be chosen uniformly for $\theta$ and $t$, so that the $\liminf$ in (ref) can be replaced with a $\lim$.

Evidently, one has $\gamma \le 0$ and $0 \le |\gamma(\theta, t)| \le \norm{D_{\theta; t}v}_{\mathfrak{B}^*}$.\footnote{We impose the convention $\frac{0}{0} = 0$.} Also, when $\Theta$ is such that the ball $B(\theta, \delta) \subset \mathfrak{B}$ is contained in $\Theta$ for some $\delta > 0$, one immediately has by definition of the dual norm that $\gamma(\theta, t) = - \norm{D_{\theta; t}v}_{\mathfrak{B}^*}$.

In the statement of the next assumption, we write that pseudometric $\rho$ is continuous with respect to a topology $\mathcal{V}$ on $\mathcal{X}$ if, whenever $x_\alpha \rightarrow x$ is a convergent net in $\mathcal{X}$, one has $\rho(x_\alpha, x) \rightarrow 0$.

asm$\mathcal{T}$ is equipped with a topology $\mathcal{U}$ which makes $\rho: \Theta \times \mathcal{T} \text{ (product topology)} \rightarrow \mathbb{R}$ continuous, $\mathcal{T}$ compact, and the maps \begin{align} &t \mapsto D_{\theta; t}v(\theta' - \theta)\nonumber \\ &t \mapsto v_n(\theta, t) + D_{\theta; t }v(\theta' - \theta) \end{align} almost surely $\mathcal{U}$-upper semicontinuous for all $\theta \in \Theta_0, \theta' \in \Theta$.

Assumption (ref) requires the existence of a topology on $\mathcal{T}$ that is strong enough that $t \mapsto v_n(\theta, t) + D_{\theta; t }v(\theta' - \theta)$ becomes upper semicontinuous for every $\theta$, but also weak enough such that $\mathcal{T}$ is compact. In the proof of Theorem (ref), it is shown that an essential implication of Assumption (ref) is that the map $t \mapsto \gamma(\theta, \cdot)$ also is $\mathcal{U}$-continuous.

We now follow Example (ref) in outlining some instances in which Assumptions (ref) is met:

enumerate• Suppose that \begin{align} &v(\theta, t) = t(m(\theta)) and v_n(\theta, t) = t(m_n(\theta)), where \nonumber \\ &m(\theta) = \mathrm{E}\left[ g(Y, \theta) \right] and m_n(\theta) = \mathrm{E}_{n}\left[ g(Y, \theta) \right] \end{align} for some moment function $g: \Theta \mapsto \mathfrak{X}$, $\mathfrak{X}$ some Banach space. Make the additional assumption that $\mathcal{T}$ is a convex set of affine functions mapping $\mathfrak{X}$ to $\mathbb{R}$ which are uniformly bounded on some neighborhood of $0$ in $\mathfrak{X}$.\footnote{Note that, if $\mathcal{T}$ is regarded as a subset of a topological vector space with induced topology $\mathcal{U}$, then an affine $t$ is continuous if and only if there exists a neighborhood of $0$ in $\mathfrak{X}$ on which $t - t(0)$ is bounded (baggett1991, Theorem 3.5)} In addition, impose that the Fr\'{e}chet derivative $\nabla m(\cdot)$ exists as a mapping from $\mathfrak{B}$ to $\mathfrak{X}$. This setting is a general formulation of cases 1, 2, and 4 of Example (ref). If $t$ is an element of $\mathcal{T}$, then $t(\cdot) - t(0)$ is linear. By continuity and linearity, the derivative invoked in Assumption (ref) satisfies \begin{align} D_{\theta; t} v(\cdot) = t(\nabla m(\theta)(\cdot)) - t(0), \end{align} where $\nabla$ signifies a Fr\'{e}chet derivative taken with respect to $\theta$. The main task of Assumption (ref) is to provide a topology $\mathcal{U}$ on $\mathcal{T}$ with respect to which both lines of (ref) are upper continuous and $\mathcal{T}$ is compact. We deal with the stronger notion of $\mathcal{U}$-continuity. A straightforward choice of such a $\mathcal{U}$ for this task is the weak topology (the topology of pointwise convergence, see\ fabian2011 \S3). As long as $\mathcal{T}$ consists of only bounded affine maps, this choice of topology automatically makes $\mathcal{T}$ precompact (note that a set of affine maps over $\mathfrak{X}$ can be viewed as a subset of the dual of $\mathbb{R} \times \mathfrak{X}$, whose closed unit ball is compact by the Banach-Alaoglu theorem). By (ref) and (ref), the topology of pointwise convergence makes the maps in (ref) continuous. Of course, some more restrictive choices of test functions $\mathcal{T}$ would be compact under stronger choices of $L^p$ or Sobolev topologies (e.g.\ FM2019). The remainder of Assumption (ref) rests on continuity of $\rho$ with respect to the product of the norm topology on $\Theta$ with $\mathcal{U}$. Under (ref), a very typical seminorm $\rho$ which verifies Assumption (ref) is the one which has \begin{align} \rho ((\theta, t), (\theta', t')) &= \mathrm{E}\left[ ( t(g(Y, \theta)) - t'(g(Y, \theta')) - \mathrm{E}\left[ t(g(Y, \theta)) - t'(g(Y, \theta') \right] ) )^2 \right]ta') \right] ) )^2 }^{1/2} \nonumber \\ & \le \mathrm{E}\left[ ( t(g(Y, \theta)) - t'(g(Y, \theta') ) )^2 \right]^{1/2} \nonumber \\ & \le \mathrm{E}\left[ ( (t - t')(g(Y, \theta) ) )^2 \right]^{1/2} + \mathrm{E}\left[ ( t'(g(Y, \theta)- g(Y, \theta') ) - t'(0) )^2 \right]^{1/2} \end{align} (e.g.\ VW1996 \S2.1). Continuity of $\rho$ can be deduced from (ref) in a number of ways. Under uniform boundedness of maps in $\mathcal{T}$, continuity of the second summand in (ref) is a consequence of continuity of the pseudometric $(\theta, \theta') \mapsto \mathrm{E}\left[ \norm{ g(Y, \theta) - g(Y, \theta')}_\mathfrak{X}^2 \right]^{1/2}$ with respect to the norm topology on $\mathfrak{B}$, so we focus on the first summand. To establish continuity of the first summand with respect to the weak topology on $\mathcal{T}$ at $(\theta, t)$, we claim that it is sufficient to impose the relatively innocuous assumptions that $Y$ is a tight random variable (VW1996, \S1.3), the map $y \mapsto g(y, \theta)$ is continuous from $\mathrm{supp}\left( Y \right)$ to $\mathfrak{X}$, and $\mathrm{E}\left[ \norm{g(Y, \theta)}_{\mathfrak{X}}^2 \right]$ exists.\footnote{Actually, Lemma 1.3.2 of VW1996 implies that the first two assumptions can be replaced with mere separability of the distribution of $\mathfrak{X}$-valued random variable $g(Y, \theta)$, which is satisfied whenever e.g.\ $\mathfrak{X}$ is itself separable.} Under these restrictions, there exists some sequence $(K_n)$ of nested compact subsets of $\mathrm{supp}\left( Y \right)$ increasing to the whole set for which $\{g(y, \theta): y \in K_n\}$ is compact for all $n$. Let $t_\alpha \rightarrow t$ be a convergent net in the weak topology $\mathcal{U}$. Then, by uniform equicontinuity of the functions in $\mathcal{T}$, one has $\mathrm{E}\left[ ((t - t_\alpha)(g(Y, \theta)))^2 \mathbf{1}_{Y \in K_n} \right] \rightarrow 0$ for all $n$, and by consequence that \begin{align*} \limsup_\alpha \mathrm{E}\left[ ((t - t_\alpha)(g(Y, \theta)))^2 \right] \le \mathrm{E}\left[ ((t - t_\alpha)(g(Y, \theta)))^2 \mathbf{1}_{K_n^c} \right]. \end{align*} By dominated convergence, $K_n$ can be chosen to make the $\limsup$ above arbitrarily small, so it must actually equal $0$. This establishes the claim. • We now turn to a general case of the conditional moment restriction model described in case 3 of Example (ref) in which one has \begin{align} &v(\theta, t) = \mathrm{E}\left[ g(Y, \theta) t(Z) \right] and v_n(\theta, t) = \mathrm{E}_{n}\left[ g(Y, \theta)t(Z) \right] \end{align} for $t$ in some class of test functions $\mathcal{T}$, and \begin{align} \rho ((\theta, t), (\theta', t'))& = \mathrm{E}\left[ (g(Y, \theta) t(Z) - g(Y, \theta')t'(Z) - \mathrm{E}\left[ g(Y, \theta) t(Z) - g(Y, \theta')t'(Z) \right])^2 \right]a')t'(Z) \right])^2}^{1/2} \nonumber \\ & \le \mathrm{E}\left[ (g(Y, \theta)(t - t')(Z))^2 \right]^{1/2} + \mathrm{E}\left[ ((g(Y, \theta) - g(Y, \theta'))t'(Z))^2 \right]^{1/2}. \end{align} With sufficient regularity, one has $D_{\theta; t}v (\theta' - \theta) = \mathrm{E}\left[ \nabla g(Y, \theta)(\theta' - \theta) t(Z) \right]$. A topology satisfies the continuity requirements of (ref) if this map is continuous with respect to the choice of $t$ for all $\theta'$ and $v_n(\theta, t)$ is also continuous with respect to $t$. A minimal choice of topology for which is feasibly true is the topology of pointwise convergence which makes all of the maps $t \mapsto t(z)$, $z \in \mathrm{supp}\left( Z \right)$ continuous. By the Tychonoff theorem, $\mathcal{T}$ is automatically precompact in this topology if its constituent functions are uniformly bounded over $\mathrm{supp}\left( Z \right)$. Even this weak choice of topology $\mathcal{U}$ can be employed to establish the continuity properties of the expectations involved in Assumption (ref). Suppose that the distribution of $Z$ is pre-tight on some metric space and that the functions $t \in \mathcal{T}$ are uniformly bounded and uniformly equicontinuous on bounded subsets of $\mathrm{supp}\left( Z \right)$.\footnote{Again, \S1.3 of VW1996 implies that it is enough that $Z$ is Borel measurable and takes on values in a separable metric space.} For instance, S2012 uses a class of complex exponential functions $t: z \mapsto \exp(i \zeta_t z)$ as test functions $t$, which satisfy the equicontinuity property over bounded sets of $z$. Then, under existence of the integral $\mathrm{E}\left[ |g(Y, \theta)|^2 \right]$, an argument exactly like the one that proved continuity of the first summand in (ref) implies the continuity of the first summand of (ref) with respect to the topology of pointwise convergence. • We have already exploited the fact that the topology of pointwise convergence is as strong as the topology of uniform convergence for uniformly equicontinuous sets of functions over compact domains. This leads to the observation that, when the domain of the functions $t$ is compact, one can often take $\mathcal{U}$ to be a pseudometric (or, indeed, metric) topology on $\mathcal{T}$ generated by a pseudometric $d_{\mathcal{T}}$. For instance, if one has \begin{align} v(\theta, t) = \mathrm{E}\left[ g(Y, \theta, t) \right] and v_n(\theta, t) = \mathrm{E}_{n}\left[ g(Y, \theta, t) \right] \end{align} for some moment function $g(\cdot, \theta, t)$ (note that this nests models (ref) and (ref)), $\mathrm{supp}\left( Y \right) \times \Theta$ is compact, and the set of functions $\{(y,\theta) \mapsto g(y, \theta, t): t \in \mathcal{T}\}$ is uniformly equicontinuous, then the Arzel\`{a}-Ascoli theorem implies that $\mathcal{T}$ is at least precompact under pseudometric $d_{\mathcal{T}}$: \begin{align} d_{\mathcal{T}}(t,t') = \sup_{(y,\theta) \in \mathrm{supp}\left( Y \right) \times \Theta} |g(y, \theta, t) - g(y, \theta, t')|. \end{align} Letting $\mathcal{U}$ be the associated pseudometric topology on $\mathcal{T}$ has the helpful property of forcing $(\theta, t) \mapsto g(y, \theta, t)$ to be product topology-continuous for each $y$, which is useful for proving continuity of the pseudometric $\rho$ as in the arguments treating (ref) and (ref). Of course, it is simple to extend this principle to larger classes of functions. A requirement of (ref) is $\mathcal{U}$-continuity of the derivatives of $v$ with respect to $\theta$. If the maps $\{\theta \mapsto \nabla \mathrm{E}\left[ g(Y, \theta, t) \right]: t \in \mathcal{T}\}$ are also uniformly equicontinuous with respect to (say) the operator norm topology, this can be accomplished by adding a term $\sup_{\theta \in \Theta} \norm{\nabla \mathrm{E}\left[ g(Y, \theta, t) \right] - \nabla \mathrm{E}\left[ g(Y, \theta, t') \right]}_{\mathrm{op}}$ to the definition of $d_{\mathcal{T}}$ in (ref). The weak operator topology may be employed with similar effect when it is metrizable. \begin{exm} Consider a version of (ref) wherein one has \begin{align*} v_n(\theta, t) = t(m_n(\theta)) - t(0) = \mathrm{E}_{n}\left[ t(g(Y, \theta)) - t(0) \right], \mathcal{t \in \mathcal{T}} \end{align*} where the collection of functions $\mathcal{T}$ is uniformly equicontinuous and bounded over the domain $g(\mathrm{supp}\left( Y \right) \times \Theta)$. For instance, when $g$ is continuous and $\mathrm{supp}\left( Y \right) \times \Theta$ is compact, these two regularity properties are satisfied in general cases of parts 1, 2, and 4 of example (ref) (Lemma (ref) below handles part 4). Then, the Arzel\`{a}-Ascoli theorem implies that the the collection of maps $(y, \theta) \mapsto t(g(y, \theta))$ indexed by $t \in \mathcal{T}$ is precompact with respect to the uniform norm metric. As it is generally no inconvenience to substitute in (ref) between $\mathcal{T}$ and its uniform norm closure, we may typically also assume that $\mathcal{T}$ is compact in the uniform norm sense. When $\mathcal{U}$ is taken to be the uniform norm topology imposed on $t$ over $g(\mathrm{supp}\left( Y \right) \times \Theta)$ and $g$ is continuous, the pseudometric $((\theta, t) , (\theta', t')) \mapsto \sup_{y \in \mathrm{supp}\left( Y \right)} |t(g(y, \theta)) - t'(g(y, \theta'))|$) becomes continuous with respect to the product topology on $\Theta \times \mathcal{T}$. As this pseudometric is an upper bound for $\rho$ (as it is defined in (ref)), the first continuity requirement of Assumption (ref) is met. The subsequent continuity requirements may be addressed with a similar treatment of the derivatives of $g$ with respect to parameter $\theta$. When the supports of the measures addressed in part 5 of Example (ref) are compact and the space $\mathcal{T}$ is a uniformly equicontinuous space of test functions, then one can reason similarly for parameters identified by certain relations between their induced measures. \end{exm}

The preceding assumptions are sufficient to provide an upper bound for the asymptotic distribution of $T_n$. We now state a condition that turns the upper bound into an equality, which references the function $f$ introduced in Assumption (ref).

asmThe following both hold: \begin{enumerate} • $v(\theta, \cdot )$ is $\mathcal{U}$-upper semicontinuous for all $\theta \in \Theta_0$$f_v$ is right-continuous, there exist random variables $\widehat{\theta}_n$ which satisfy $f_v(d_{\mathfrak{B}}(\widehat{\theta}_n , \Theta_0))= o_p(r_n^{-1})$, and $\sup_{t \in \mathcal{T}} v_n(\widehat{\theta}_n, t) \le \inf_{\theta \in \Theta} \sup_{t \in \mathcal{T}} v_n(\theta, t)+ o_p(r_n^{-1})$ or $v(\cdot, t)$ is convex in $\theta$ in some neighborhood of $\Theta_0 \times \mathcal{T}$.\footnote{The proof of Theorem (ref) shows that it is possible to weaken this assumption slightly to convexity on a neighborhood of $\bigsqcup_{\theta \in \Theta_0}\{\theta \} \times \tilde{K}(\theta) \subset \Theta_0 \times \mathcal{T}$.} \end{enumerate}

In some cases, $v(\theta, \cdot)$ vanishes for $\theta \in \Theta_0$, and Assumption (ref).1 is vacuously true. For instance, consider the first part of Example (ref), where $v(\theta, t) = t(m(\theta))$ for some linear functional $t$, and $\ell(\theta)$ is some seminorm of $m(\theta)$. Suppose that $\ell(\theta)$ is a norm, and vanishes if and only if $m(\theta)$ does, as is the case with GMM estimation. Then $\theta$ is in $\Theta_0$ if and only if $m(\theta) = 0$, which is true if and only if $v(\theta, \cdot) = t(m(\theta)) = 0$ for all linear functionals $t$. The discussion following Assumption (ref) provides a number of other situations in which Assumption (ref).1 is satisfied.

Assumption (ref).2 may be fulfilled in two ways. The first is by supplying a convergence rate for an approximate minimizer $\widehat{\theta}_n$ of the criterion function $\sup_{t \in \mathcal{T}} v_n(\widehat{\theta}_n, t)$ to the identified set $\Theta_0$. If, for instance, $r_n = \sqrt{n}$ and the function $f$ bounding the approximation error given in Assumption (ref) is the map $f_v(\delta) = \delta^2$, the first case of Assumption (ref).2 is fulfilled if one has $d_{\mathfrak{B}} (\widehat{\theta}_n, \Theta_0) = o_p(n^{-1/4})$. If $\Theta_0 = \{\theta_0\}$ is a singleton and $\widehat{\theta}_n$ signifies an $M$-estimator of $\theta_0$, then $\sqrt{n}$-consistency of $\widehat{\theta}_n$ (or consistency at any rate faster than $n^{-1/4}$) is sufficient to meet this first case. S2012, hong2017, and CNS2023 use similar rate conditions to ensure that the parameter space local to $\Theta_0$ can be adequately approximated with linear expansions around points in the identified set.

The second case of Assumption (ref).2 is most simply fulfilled when $v(\theta,t)$ is convex in $\theta$ for all $t \in \mathcal{T}$. In some settings, $v(\theta, t)$ is linear in both $\theta$ and $t$---take for instance the NPIV model and related formulations discussed in Example (ref). Note that, when $v$ is bilinear, its Fr\'{e}chet derivative $D$ satisfies $D_{\theta; t}v(\theta' - \theta) = v(\theta' , t) - v(\theta, t)$, so one can take $f = 0$ in Assumption (ref). This automatically satisfies the rate condition of the first case of Assumption (ref).2.

We now state a result which bounds the asymptotic distribution of $T_n = r_n\inf_{\theta \in \Theta} \sup_{t \in \mathcal{T}} v_n(\theta, t)$, and gives an exact limiting distribution under the additional imposition of Assumption (ref)

theoremLet Assumptions (ref), (ref), (ref), and (ref) hold. For $\theta \in \Theta$, let $K(\theta)$ and $\tilde{K}(\theta)$ be defined as follows: \begin{align*} K(\theta) & = \{t \in \mathcal{T}: \gamma(\theta, t) = 0\}\\ \tilde{K}(\theta) & = \{t \in \mathcal{T}: \gamma(\theta,t) = v(\theta, t) = 0\} \end{align*} Then, if $\Theta_0$ is nonempty, $K(\theta)$ is nonempty for all $\theta \in \Theta_0$ and \begin{align} T_n = r_n\inf_{\theta \in \Theta} \sup_{t \in \mathcal{T}} v_n(\theta, t) &\le \inf_{\theta \in \Theta_0} \sup_{t \in K(\theta)} \mathbb{G}_n(\theta,t) + o_p(1) \nonumber\\ & \rightsquigarrow \inf_{\theta \in \Theta_0} \sup_{t \in K(\theta)} \mathbb{G}(\theta, t). \end{align} If $\Theta_0$ is empty, then $r_n = O_p(T_n)$. If Assumption (ref).1 also holds, then the preceding also holds with $K(\theta)$ replaced by $\tilde{K}(\theta)$. If all of Assumption (ref) also holds, then the preceding is an equality with $K(\theta)$ replaced by $\tilde{K}(\theta)$.

Examples (ref) and (ref) have discussed several instances in which one has $v(\theta, t) = t(m(\theta))$ for $m$ a mapping from $\Theta$ to a Banach space $\mathfrak{X}$, and $t$ a continuous affine function from $\mathfrak{X}$ to $\mathbb{R}$. When $t$ is continuous and affine, the map $x \mapsto t(x) - t(0)$ is linear (Rock1970, \S 1). Suppose $m$ is Fr\'{e}chet differentiable, and let $\nabla m$ denote the Fr\'{e}chet derivative of $m$ at $\theta$. Then, a straightforward calculation with (ref) shows that

align[align omitted — 77 chars of source]

for all $\theta \in \Theta_0$ and $t \in \mathcal{T}$. Therefore, we may rewrite the defining equation for $\gamma$ and infer that $K(\theta)$ is the set of $t$ which satisfy

align[align omitted — 73 chars of source]

for all $\theta'$ which are close enough to $\theta$ in the $d_\mathfrak{B}$-metric (the notion of “close enough" may be defined uniformly for $\theta \in \Theta_0$---see the proof of Theorem (ref)). (ref) has applications to several hypothesis testing scenarios of interest. An interesting case arises when $\mathfrak{X}$ is a Hilbert space, and the moment functions $m_n(\theta)$ converges weakly to a tight $L^\infty(\Theta)$-valued element $\mathbb{W}_n(\theta)$:

align[align omitted — 118 chars of source]
corSuppose that $v(\theta, t) = t(m(\theta))$ and $v_n(\theta, t) = t(m_n(\theta))$, where $m$ is Fr\'{e}chet differentiable and maps $\Theta$ to a Hilbert space $\mathfrak{X}$, and $\mathcal{T}$ is the unit ball in $\mathfrak{X}^*$ (so that $\ell(\theta) = \norm{m(\theta)}_{\mathfrak{X}}$). Make Assumptions (ref), (ref), (ref), and (ref). Then, if $\Theta_0$ is nonempty and contained in the interior of $\Theta$, one has \[ K(\theta) = \{t \in \mathfrak{X}^*: t(\nabla m ) = 0\} \] and, by consequence, \begin{align*} T_n &\le \inf_{\theta \in \Theta_0} \norm{M_{ \nabla m} r_n(m_n (\theta ) - m(\theta)) }_{\mathfrak{X}} + o_p(1) \\ & = \inf_{\theta \in \Theta_0} \norm{M_{\nabla m} \mathbb{W}_n(\theta)}_{\mathfrak{X}} + o_p(1), \end{align*} where $M_{\nabla m}$ is the projection to the orthogonal complement of the closure of the image of $\nabla m: \mathfrak{B} \rightarrow \mathfrak{X}$.
proofBy the preceding discussion, $t$ is in $K(\theta)$ if and only if $t(\nabla m(\theta' - \theta))$ for $\theta'\in \Theta$ which are in a $\underline{\delta}$-ball around $\theta$. Because $\Theta_0$ is contained in the interior of $\Theta$, this is true if and only if $t$ vanishes on the image of $\nabla m$, when the latter is viewed as a map from $\mathfrak{B}$ to $\mathfrak{X}$. By continuity of $t$, this is true if and only if $t$ vanishes on the closure $\overline{\mathrm{im}(\nabla m)}$. Let $P_{\nabla m}$ denote the projection onto this closed subspace (C1994), let $M_{\nabla m}$ denote its complement $I - P_{\nabla m}$, and with some abuse of notation let $t$ denote its own Riesz representor. Then, for all $x \in \mathfrak{X}$, \begin{align*} \sup_{t \in K(\theta)} t(x) & = \sup_{\substack{\norm{t}_{\mathfrak{X}} \le 1 \\ P_{\nabla m } t = 0}} M_{\nabla m} t(x) = \sup_{\norm{M_{\nabla m} t}_{\mathfrak{X}} \le 1} \left\langle M_{\nabla m} t, M_{\nabla m} x \right\rangle \\ & = \norm{M_{\nabla m}x }_{\mathfrak{X}}. \end{align*} Theorem (ref) then concludes.

Corollary (ref) complements tests of overidentifying restrictions provided under the assumption of point identification with method of moments estimators (e.g.\ Sargan1958 and H1982 for a generalization). In particular, if $\mathfrak{X}$ is some Euclidean space $\mathbb{R}^k$, and $r_n(m_n(\theta) - m(\theta))$ has a limiting $N(0, I_{k \times k})$ distribution (as is the case with efficiently weighted GMM), then the upper bound implied by Corollary (ref) has an asymptotic distribution characterized by writing

align[align omitted — 315 chars of source]

where $\mathbf{Z}$ is a standard normal random vector. Under point identification and a full-rank condition on $\nabla m$, (ref) states the limiting $\chi^2$ distribution of the $J$-statistic of overidentifying relations. We note that, under the usual assumptions of GMM, one typically verifies that Assumption (ref) holds, so that the asymptotic distribution of $T_n$ is precisely (ref).

It is also fruitful to examine the result of Corollary (ref) when $\mathfrak{X}$ is more generally held to be a Banach space and $\ell$ the composition of a moment function with some convex function over the Banach space (like a norm). Continue to suppose that $v(\theta, t) = t(m(\theta))$ for $t$ in the dual $\mathfrak{X}^*$ of some Banach space $\mathfrak{X}$. S2012, hong2017, CNS2023, and Fan2023, among other papers, provide Banach space bounds for the distribution of $T_n$ which have in common the structure:

align[align omitted — 160 chars of source]

where $S_\theta$ is some subset of the tangent space $T_\theta \Theta$ of $\Theta$ at $\theta$ (or its closure), which for our purposes may best be defined as the convex cone:

align*[align* omitted — 169 chars of source]

(convexity follows under local convexity of $\Theta$ and Rock1970, Theorem 2.5; note that $T_\theta \Theta$ may not be a subspace if, for instance, $\theta$ is contained in the boundary of $\Theta$). In the aforementioned results, one typically defines $S_\theta$ as the closed limit of linear subspaces formed by linear sieve approximations to $\Theta$.

Let $h \in \overline{T_\theta \Theta}$ be a nonzero element of the closure of the tangent space of $\Theta$ at $\theta$. Then, there exists a sequence of $\theta'_k \in \Theta$ converging to $\theta$, and a diverging sequence of $\lambda_k$, satisfying $\lambda_k (\theta'_k - \theta) \rightarrow h$ in $\mathfrak{B}$. If one lets $t$ be an affine function over $\mathfrak{X}$ and an element of $K(\theta)$, then by (ref) and (ref),

align[align omitted — 603 chars of source]

From (ref), one can readily show that the upper bound in (ref) is a lower bound for bounds of the type (ref), whence less conservative, especially under the assumptions of Theorem (ref):

lemmaLet $v(\theta, t) = t(m(\theta)) -t(0)$ and $v_n(\theta, t) = t(m_n(\theta))$, where $m$ is Fr\'{e}chet differentiable and maps $\Theta$ to a Banach space $\mathfrak{X}$, $\mathcal{T}$ is some set of continuous affine functions mapping $\mathfrak{X} \rightarrow \mathbb{R}$, and for all $\theta \in \Theta_0$, $S_\theta$ is some subset of the closure $\overline{T_\theta \Theta}$ of the tangent space of $\Theta$ at $\theta$. Then, under Assumption (ref), one has the bound \begin{align} \inf_{\theta \in \Theta_0} \sup_{t \in K(\theta)} \mathbb{G}_n(\theta, t) = \inf_{\theta \in \Theta_0} \sup_{t \in K(\theta)} t(\mathbb{W}_n(\theta)) \le \inf_{\theta \in \Theta_0} \inf_{h \in S_\theta} \sup_{t \in \mathcal{T}} t(\mathbb{W}_n(\theta) + \nabla m(h)). \end{align} In particular, if $\mathcal{T}$ is the unit ball in $\mathfrak{X}^*$, then one has \begin{align} \inf_{\theta \in \Theta_0} \sup_{t \in K(\theta)} \mathbb{G}_n(\theta, t) = \inf_{\theta \in \Theta_0} \sup_{t \in K(\theta)} t(\mathbb{W}_n(\theta)) \le U_n, \end{align} with equality if $S_\theta \supset T_\theta \Theta, \, \forall \theta \in \Theta_0$.
remUsing the definition of $\gamma$, it is straightforward to show that, in the setting of Lemma (ref) and under Assumption (ref), (ref) is actually a defining property for $K(\theta)$. More precisely, an element $t$ is in $K(\theta)$ if and only if $t(\nabla m(h)) - t(0) \ge 0$ for all $h \in T_\theta \Theta$.

Critical values

By Theorem (ref), critical values for the distribution of \[ \inf_{\theta \in \Theta_0} \sup_{t \in K(\theta)} \mathbb{G}(\theta, t) \] can be used to describe valid, if conservative, critical values for the test statistic $T_n$. The chief challenge in determining these critical values is in estimating the sets $\Theta_0$ and $K(\theta)$ for each point in $\Theta_0$, but this is not onerous if uniformly consistent estimators for the criterion function $\ell$ and the derivative-norm map $\gamma$ are available. The former can be estimated using $\ell_n$ under our assumptions, whereas an estimator for the latter can be obtained if one has access to an estimate $\nabla v_n$ of $v$.

Our recommended approach for inference in this partially identified setting (see especially S2012, hong2017, Zhu2020) is to apply the bootstrap. For a bootstrapped statistic $Z_n^*$ and a limit distribution $Z$, we follow VW1996 in writing that $Z_n^* \overset{\mathrm{P}}{\rightsquigarrow} Z$ if the convergence $d_{\mathrm{BL}_1}(Z_n^*, Z) = \sup_{h \in \mathrm{BL}_1} \mathrm{E}^*[h(Z_n^*)] - \mathrm{E}\left[ h(Z) \right]\overset{p}{\rightarrow} 0$ (in outer probability) holds in the bounded lipschitz metric $d_{\mathrm{BL}_1}$, where $\mathrm{BL}_1$ is the set of 1-Lipschitz functions bounded in magnitude by $1$, and $\mathrm{E}^*$ is an expectation conditional upon the sample. We write $Z_n^* \overset{\mathrm{a.s.}}{\rightsquigarrow} Z$ if the same convergence holds almost surely with regard to outer probability. With some abuse of notation, we let expectations and probabilities refer to outer measure when quantities are asymptotically measurable.

The following assumptions and theorem demonstrate a strategy for approximating the asymptotic distribution of $T_n$ using the bootstrap. Hemicontinuity in our assumptions is relative to topology $\mathcal{U}$ on $\mathcal{T}$ and the $d_{\mathfrak{B}}$ metric on $\Theta$ (see A1999).

asm[Bootstrap Consistency] The following hold: \begin{enumerate} • There exists an asymptotically measurable bootstrap analogue $\mathbb{G}_n^*$ of $\mathbb{G}_n$ which satisfies \begin{align*} \mathbb{G}_n^*(\theta, t) \overset{\mathrm{P}}{\rightsquigarrow} \mathbb{G}(\theta, t) \end{align*} • $|\gamma|$ is bounded over $\Theta \times \mathcal{T}$ by a constant $C_\gamma < \infty$. For all $\theta \in \Theta$, there exist a sequence of nonpositive and $d_\mathcal{T}$-upper semicontinuous functions $(\psi_n(\theta, \cdot) )$ decreasing pointwise to $\gamma(\theta, \cdot)$ and estimators $(\widehat{\psi}_n(\theta, \cdot))$ which satisfy \begin{align*} \sup_{\substack{\theta \in \Theta \\ t \in \mathcal{T}}} |\widehat{\psi}_n(\theta, t) - \psi_n(\theta, t)| = O_p(q_n) \end{align*} where $(q_n)\rightarrow 0$ is a convergent sequence of constants. • There is some topology $\mathcal{V}$ at least as strong as the metric topology on $\Theta$ such that $\theta \mapsto K(\theta)$ is $\mathcal{V}$-lower hemicontinuous at every point in $\Theta_0$ and every $\mathcal{V}$-open neighborhood of $\Theta_0$ contains a $d_{\mathfrak{B}}$-open neighborhood of $\Theta_0$, or there is a known nondecreasing $f_\ell: \mathbb{R}_+ \rightarrow \mathbb{R}_+$ such that one has $\sup_{\theta \in \Theta} \frac{ d_{\mathfrak{B}}(\theta, \Theta_0)}{f_\ell(\ell(\theta))} < \infty$ and $\lim_{x \downarrow 0} f_\ell(x) = 0$, and there is a $d_{\mathfrak{B}}$-open neighborhood $U$ of $\Theta_0$ on which $\gamma$ satisfies the Lipschitz condition: \begin{align*} \inf_{\substack{\theta \in \Theta_0, \theta' \in U \\ t \in K(\theta)}} \frac{\gamma(\theta' , t) }{d_\mathfrak{B} (\theta , \theta') } > - \infty. \end{align*} \end{enumerate}

The first two parts of Assumption (ref) are fairly light bootstrap consistency and regularity conditions. The functions $\psi_n$ can be taken to be $\gamma$ if the latter can be directly estimated (e.g.\ by computing the Fr\'{e}chet derivatives of $v$). In other cases, it may be more appropriate to let

align[align omitted — 194 chars of source]

where $(\delta_n)$ is an sequence of constants converging slowly enough to $0$ so that $\psi_n$ can be adequately estimated. Such a choice of $\psi_n$ satisfies the properties indicated in Assumption (ref) and is further discussed below, along with the corresponding sequence $(q_n)$.

The third part of Assumption (ref) consists of two possible cases. The first is satisfied under a hemicontinuity condition on the correspondence ${K}(\theta)$ with respect to a topology $\mathcal{V}$, which can be either the metric topology associated with $d_{\mathfrak{B}}$ or a stronger topology. Suppose that $\Theta$ and $\mathcal{T}$ are finite dimensional spaces, and assume an appropriate degree of smoothness for the map $(\theta, t) \mapsto D_{\theta; t} v$ (the right hand side is the Jacobian of $v(\cdot, t)$ evaluated at $\theta$). We have remarked that, as long as $\Theta_0$ is contained in the interior of $\Theta$, $K(\theta)$ is defined as the zero set of the vector-valued map $(\theta, t) \mapsto D_{\theta; t}v$. Its lower hemicontinuity may thus be established as a consequence of the implicit function theorem if the Jacobian of this map, regarded as a function in $t$, has full rank at all points $(\theta, t)$, $\theta \in \Theta_0, t \in K(\theta)$ (e.g.\ A1999). In more general settings, lower hemicontinuity of the correspondence can be ascertained using infinite-dimensional versions of the implicit function theorem (e.g.\ Loomis1990). Section (ref) below gives equivalent characterizations for the lower hemicontinuity of the correspondence $K$ in the strong (norm) topology on any Banach space.

The second case of Assumption (ref).3 is essentially a requirement that the function $x \mapsto \sup_{\theta: \ell(\theta) \le x} d_{\mathfrak{B}}(\theta, \Theta_0)$ can be bounded above up to some constant factor by a nondecreasing function $f_\ell$ (note that one can take $f_\ell$ to be exactly this increasing function, so that $f_\ell$ always exists). This requirement of local identification in a neighborhood of $\Theta_0$ resembles similar conditions imposed under partial identification, such as Assumption 3.4 in CNS2023 and Assumption B.3 in Zhu2020.

Assumption (ref) is sufficient for the estimation of an upper bound of the distribution of $T_n$ using $K(\theta)$. Theorem (ref) demonstrates that a tighter bound may potentially be obtained using the smaller set $\tilde{K}(\theta)$, so we also state a supplementary assumption which enables such an estimate to be made. We have observed that, in many settings of interest, one has $v(\theta, \cdot) = 0$ whenever $\theta \in \Theta_0$. This forces the relation $\tilde{K}(\theta) = K(\theta)$ and renders the following assumption unnecessary.

asm[Bootstrap Consistency II] The correspondence $\theta \mapsto \tilde{K}(\theta)$ is $\mathcal{V}$-lower hemicontinuous at every point in $\Theta_0$, or there is a $d_{\mathfrak{B}}$-open neighborhood $U$ of $\Theta_0$ on which $\gamma$ and $v$ satisfy the Lipschitz conditions: \begin{align*} \inf_{\substack{\theta \in \Theta_0, \theta' \in U \\ t \in \tilde{K}(\theta)}} \frac{\gamma(\theta' , t) }{d_\mathfrak{B} (\theta , \theta') }, \inf_{\substack{\theta \in \Theta_0, \theta' \in U \\ t \in \tilde{K}(\theta)}} \frac{ v(\theta', t)}{d_\mathfrak{B} (\theta , \theta') } > - \infty. \end{align*}
theoremMake Assumptions (ref), (ref), (ref), (ref), (ref).1, and (ref). Let $(\lambda_n) \rightarrow \infty$ and $(\mu_n), (\tilde{\mu}_n) \rightarrow \infty$ be divergent sequences chosen so that $\lambda_n q_n, \mu_n r_n^{-1},\tilde{\mu}_nr_n^{-1} = o(1) $ and $\lambda_n = o(\mu_n)$. If the second case of Assumption (ref).3 is presumed to hold, suppose also that $\lambda_n f_\ell(c_\gamma\mu_n^{-1} \lambda_n )$ and $\tilde{\mu}_n f_\ell(c_\gamma\mu_n^{-1} \lambda_n )$ are $o(1)$ for some constant $c_\gamma > C_\gamma$. Then, one has \begin{align} \inf_{\theta \in \Theta} (\mu_n \ell_n(\theta) + (\sup_{t \in \mathcal{T}} \mathbb{G}_n^*(\theta, t) + \lambda_n \widehat{\psi}_n(\theta,t))) \overset{\mathrm{P}}{\rightsquigarrow} \inf_{\theta \in \Theta_0} \sup_{t \in K(\theta)} \mathbb{G}(\theta, t). \end{align} If Assumption (ref) also holds, then \begin{align} &\inf_{\theta \in \Theta} (\mu_n \ell_n(\theta) + (\sup_{t \in \mathcal{T}} \mathbb{G}_n^*(\theta, t) + \lambda_n \widehat{\psi}_n(\theta,t) + \tilde{\mu}_n v_n(\theta, t) ) ) \overset{\mathrm{P}}{\rightsquigarrow} \inf_{\theta \in \Theta_0} \sup_{t \in \tilde{K}(\theta)} \mathbb{G}(\theta, t). \end{align}

Theorem (ref) includes as a corollary a simpler way of obtaining a critical values for the asymptotic distribution of $T_n$. The simplification omits the outer minimization over $\Theta$ in the left hand sides of (ref) and (ref) as long as an appropriate value of $\theta$ is used. When point identification holds so that $\Theta_0$ is a singleton, and some additional regularity on the functions $v, \psi$, and $\psi_n$ is imposed, it is straightforward to show that such a bootstrap procedure is asymptotically equivalent to the more computationally intensive approach indicated in Theorem (ref). The additional regularity comes in the form of convergence requirements upon the nonstochastic functions $\psi_n(\theta, t)$ to $\gamma$ and lower semicontinuity of $\gamma$:

asm$\Theta = \{\theta_0\}$ is a singleton, and there is a neighborhood $\mathcal{N}$ of $\{\theta_0\} \times K(\theta_0) \subset \Theta \times \mathcal{T}$ on which \begin{enumerate} • $\gamma$ is upper semicontinuous • The functions $\psi_n$ converge uniformly to $\gamma$. \end{enumerate}
corLet $\widehat{\theta}_n$ satisfy $\ell_n(\widehat{\theta}_n) \le \inf_{\theta \in \Theta} \ell_n(\theta) + O_p(r_n^{-1})$. Then, the assumptions implying (ref) also imply \begin{align} \sup_{t \in \mathcal{T}} \mathbb{G}_n^*(\widehat{\theta}_n, t) + \lambda_n \widehat{\psi}_n(\widehat{\theta}_n, t) - Z_n^* \overset{\mathrm{P}}{\rightsquigarrow} \inf_{\theta \in \Theta_0} \sup_{t \in K(\theta)} \mathbb{G}(\theta, t), \end{align} and the assumptions implying (ref) also imply \begin{align} \sup_{t \in \mathcal{T}} \mathbb{G}_n^*(\widehat{\theta}_n, t) + \lambda_n \widehat{\psi}_n(\widehat{\theta}_n, t) + \tilde{\mu}_n v_n(\widehat{\theta}_n , t) - Z_n^* \overset{\mathrm{P}}{\rightsquigarrow} \inf_{\theta \in \Theta_0} \sup_{t \in K(\theta)} \mathbb{G}(\theta, t), \end{align} where $Z_n^* \ge 0$. Under Assumption (ref), one can take $Z_n^* = 0$ in (ref). Under Assumption (ref) and upper semicontinuity of $v$ on $\mathcal{N}$, one can take $Z_n^* = 0$ in (ref).

It remains to show that one can construct a sequence of estimators $\widehat{\psi}_n$ of functions $\psi_n$ satisfying the second condition of Assumption (ref). By local convexity of $\Theta \subset \mathfrak{B}$ and positive homogeneity of $D_{\theta; t}$ (Assumption (ref)), we have remarked that the maps $\psi_{\delta} :(\theta, t) \mapsto \frac{D_{\theta; t} v(\theta' - \theta)}{\delta}$ are pointwise decreasing in $\delta$ for $\delta$ small enough. Moreover, each map $\psi_\delta$ may be estimated with a sample analogue using the estimate provided in Assumption (ref).1 with $v_n$ approximating $v$. Such a strategy provides a consistent sequence of estimators $\widehat{\psi}_n$ for the sequence $(\psi_n)$ defined in (ref) under the assumptions we have already introduced:

lemmaLet $(\delta_n) \rightarrow 0$ be a convergent sequence and $\psi_n$ be as in (ref). Let $\widehat{\psi}_n(\theta, t) = \inf_{\substack{\theta' \in \Theta \\ \norm{\theta' - \theta}_{\mathfrak{B}} \le \delta_n}} \frac{v_n(\theta',t) - v_n(\theta, t)}{\delta_n}$. Make Assumptions (ref) and (ref), and suppose that Assumption (ref).1 holds for all $\theta \in \Theta$: \begin{align*} \sup_{\substack{\theta, \theta' \in \Theta \\ t \in \mathcal{T} \\ \norm{\theta' - \theta}_{\mathfrak{B}} < \delta}} | v(\theta' t) - v(\theta, t) - D_{\theta; t} v(\theta' - \theta)| = O(f_v(\delta)). \end{align*} Then, \begin{align*} \sup_{\substack{\theta \in \Theta \\t \in \mathcal{T}}} | \widehat{\psi}_n(\theta, t) - \psi_n(\theta, t)| = O_p( \frac{f_v(\delta_n) + r_n^{-1}}{\delta_n}). \end{align*} In particular, if $f_v(\delta) = o(\delta)$, then $\widehat{\psi}_n$ is consistent for $\psi_n$ whenever $r_n = o(\delta_n)$.

Lemma (ref) suggests that a tractable way of estimating e.g.\ (ref) is to solve the triple optimization problem:

align*[align* omitted — 268 chars of source]

where $\delta_n$ is such that $q_n \equiv \frac{f_v(\delta_n) + r_n^{-1}}{\delta_n} = o(1)$ and $\mu_n$ and $\lambda_n$ are as in Theorem (ref). In a remark following the proof of Lemma (ref), we demonstrate that, under convexity of $\Theta$ and Lipschitz continuity of $v$, one can introduce a Lagrangian form of $\widehat{\psi}_n$ with negligible error: \[ \tilde{\psi}_n(\theta, t) = \inf_{\substack{\theta' \in \Theta \\ \norm{\theta - \theta'}_{\mathfrak{B}}} } \frac{v_n(\theta', t) - v_n(\theta, t)}{\delta_n} + \nu_n \frac{(\norm{\theta' - \theta}_{\mathfrak{B}} - \delta_n)_+}{ \delta_n }. \] The imposition of lower hemicontinuity on $K$ and $\tilde{K}$ in Assumptions (ref) and (ref) bears significance for our estimation results, inasmuch as it greatly simplifies the choice of tuning parameters for the calculation of critical values. Therefore, it is important to provide conditions under which lower hemicontinuity can be expected to hold.

Sufficient conditions for lower hemicontinuity

We turn again to the situation first described by Example (ref).2, and also in Lemma (ref), wherein $\ell = \gamma \circ m$ is the composition of a convex function $\gamma: \mathfrak{X} \rightarrow \mathbb{R}$ with a moment function $m: \mathfrak{B} \rightarrow \mathfrak{X}$ ($\mathfrak{X}$ a Banach space). By the Fenchel-Moreau theorem, the corresponding choice of $\mathcal{T}$ is a set of continuous affine maps $t : \mathfrak{X} \rightarrow \mathbb{R}$ majorized by $\gamma$. One also has $D_{\theta; t} v (h) = t(\nabla m(\theta)(h)) - t(0)$, and $t \in K(\theta)$ if and only if this quantity is nonnegative for all $h \in T_\theta \Theta$ (Remark (ref)). In particular, $t \in K(\theta)$ if and only if $t(x) - t(0) \ge 0$ for all $x \in \nabla m(\theta)(T_\theta \Theta)$, where the latter is a convex cone in $\mathfrak{X}$. In Lemma (ref) below, we characterize when exactly the set of linear functionals defined on $\mathfrak{X}$ satisfying this inequality can be expected to be lower hemicontinuous.

According to the discussion above, $-K(\theta) \subset \mathfrak{X}^*$ can be identified with the polar $(\nabla m(\theta)(T_\theta \Theta))^\circ$. Clearly, lower hemicontinuity of $-K(\theta)$ guarantees the lower hemicontinuity of $K(\theta)$.

Conveniently, one can provide conditions under which the polar of a correspondence sending points $\theta$ to cones $C(\theta) \subset \mathfrak{X}$ is lower hemicontinuous. These conditions depend on the upper hemicontinuous of the correspondence $\theta \mapsto C(\theta)$. In fact, we can show that a weaker version of upper hemicontinuity of this correspondence, which maps a topological space into a Banach space, is both necessary and sufficient for lower hemicontinuity of the polar correspondence in the strong topology on $\mathfrak{X}^*$. The weak version of hemicontinuity is relative to the weak topology on $\mathfrak{X}$. For the purposes of this paper, we will say that a set $A$ of a topological vector space $\mathfrak{X}$ is strongly contained in another set $V$ if there exists some open neighborhood $D$ of $0$ such that $A + D \subset V$. For instance, if $A$ is compact, strong containment of $A$ in $V$ is equivalent to containment if $V$ is open (C1994, Lemma IV.3.8).

If $\mathfrak{X}$ is a Banach space, we will refer to a correspondence $H$ mapping a topological space $\mathcal{A}$ to $\mathfrak{X}$ as being weakly upper-hemicontinuous at point $a$ if, whenever $H(a)$ is strongly contained in a weakly open $V$, there exists a neighborhood $U$ of $a$ such that $a' \in U$ implies that $C(a')$ is a subset of $V$. Such a weak notion of upper hemicontinuity in the correspondence $\theta \mapsto \nabla m(\theta)(T_\theta \Theta)$ is precisely what is needed to ensure (norm-topology) lower hemicontinuity of its polar. If $H$ maps $\mathcal{A}$ to a dual space $\mathfrak{X}^*$, say that it is weak-* upper hemicontinuous if it conforms to the usual definition of upper hemicontinuity with respect to the weak-* topology on $\mathfrak{X}^*$.

lemmaLet $\mathcal{A}$ be a topological space and $\mathfrak{X}$ a Banach space with closed unit ball $B_1$, whose dual $\mathfrak{X}^*$ has closed unit ball $B_1^*$. Let $C: \mathcal{A} \rightarrow \mathfrak{X}$ be a correspondence sending points $a \in \mathcal{A}$ to convex cones $C(a) \subset \mathfrak{X}$. Let $C^\circ(a): \mathcal{A} \rightarrow \mathfrak{X}^*$ be the correspondence mapping $a \mapsto \{y \in \mathfrak{X}^*: y(x) \le 0, \text{ for all } x \in C(a)\}$. Then, the following are equivalent: \begin{enumerate} • The correspondence $a \mapsto C(a) \cap B_1$ is weakly upper-hemicontinuous at $a$$a \mapsto C^\circ(a)$ is lower hemicontinuous with respect to the strong (norm) topology on $\mathfrak{X}^*$ at $a$$a \mapsto C^\circ(a) \cap B_1^*$ is lower hemicontinuous with respect to the strong topology at $a$ \end{enumerate} as are \begin{enumerate} • $a \mapsto C(a) \cap B_1$ is lower hemicontinuous with respect to the norm topology on $\mathfrak{X}$ at $a$$a \mapsto C^\circ(a) \cap B_1^*$ is weak-* upper hemicontinuous at $a$ \end{enumerate}

Lemma (ref) allows us to show that many choices of convex function $\gamma: \mathfrak{X} \rightarrow \mathbb{R}$ actually can be written in a way which makes the correspondence $\theta \mapsto K(\theta)$ lower hemicontinuous. For instance, in the following lemma, all that is required of $\gamma$ is that it is Lipschitz continuous in a neighborhood of $m(\Theta)$. For $\mathfrak{X}$ a Banach space and $\mathcal{A}$ the vector space of continuous and affine maps $t: \mathfrak{X} \rightarrow \mathbb{R}$, we define the weak topology on $\mathcal{A}$ to be the topology of pointwise convergence, and the norm topology to be the one generated by $\norm{t}_{\mathcal{A}} = \sup_{\norm{x}_{\mathfrak{X}} \le 1} |t(x)|$. For a convex function $\gamma$ defined on $\mathfrak{X}$, we also define the subgradient $\partial \gamma(x)$ of $\gamma$ at $x$ to be the set of $y \in \mathfrak{X}^*$ satisfying that $\gamma(x') \ge \gamma(x) + \left\langle x' - x, y \right\rangle$ for all $x' \in \mathfrak{X}$.

lemmaSuppose that $\ell(\theta) = \gamma(m(\theta))$, where $m$ is a Fr\'{e}chet differentiable mapping from $\Theta$ to a Banach space $\mathfrak{X}$, and $\gamma: \mathfrak{X} \rightarrow \mathbb{R}$ is Lipschitz continuous on some closed, convex, and bounded domain $D$ containing $m(\Theta)$. Then, there is a convex and weakly compact set of affine functions $\mathcal{T} \subset \mathcal{A}$ such that $\gamma(x) = \sup_{t \in \mathcal{T}} t(x) = \max_{t \in \mathcal{T}} t(x)$ for all $x \in D$. Under this choice of $\mathcal{T}$, if $m(\Theta_0)$ is contained in the interior of $D$ (e.g.\ if $\gamma$ is Lipschitz continuous) and $v(\theta, t) = t(m(\theta))$, then for $\theta \in \Theta_0$ one has \begin{align} \tilde{K}(\theta) = \{y - y(m(\theta)): y \in -(\nabla m(\theta)(T_\theta(\Theta))^\circ \cap \partial \gamma (m(\theta))\}. \end{align} If the correspondence $\theta \mapsto \nabla m(\theta)(T_\theta \Theta) \cap B_1$ is also weakly upper-hemicontinuous at every point in $\Theta_0$, then the correspondence $\theta \mapsto K(\theta)$ is lower hemicontinuous with respect to the norm topology on $\mathcal{A}$ at every point in $\Theta_0$. If $\theta \mapsto \partial \gamma (m(\theta))$ is also lower hemicontinuous at every point in $\Theta_0$, then $\theta \mapsto \tilde{K}(\theta)$ is likewise hemicontinuous.

Lemma (ref) furnishes a host of examples of $\mathcal{T}$ corresponding to arbitrary convex functions defined on Banach spaces (Example (ref) presents several convex functions that are of particular interest). These $\mathcal{T}$ are at least compact in the topology of pointwise convergence, which the discussion following Assumption (ref) shows is generally sufficient to apply Theorem (ref). The lemma also characterizes the correspondence $\theta \mapsto \tilde{K}(\theta) \subset \mathcal{T}$ in terms of the subgradient of $\gamma$. The topology of weak convergence automatically makes $t \mapsto v(\theta, t) = t(m(\theta))$ a continuous map satisfying Assumption (ref).1, so that $\tilde{K}(\theta)$ is the correspondence most relevant to Theorem (ref).