Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
93,030 characters · 6 sections · 86 citation commands
Inference under partial identification with minimax test statistics
\address{Department of Economics and Finance, UNC Wilmington, Wilmington North Carolina 28403} \email{[email removed]} \subjclass[2000]{Primary 62G10; Secondary 62G20}
This paper is concerned with the computation and estimation of the asymptotic distribution of statistics that are based on minimax values. Such statistics are often of interest when one is interested in testing for the existence of a nonempty identified set, i.e.\ when hypothesis testing is conducted under the assumption of partial identification. When an identified set is characterized as the solution to a system of moment equations, the existence of a nonempty identified set can be consistently tested with adequate knowledge of how the moment functions will asymptotically behave thereon. We illustrate that test statistics designed to capture the behavior of these identifying relations often have a minimax formulation whose limiting distribution can be systematically computed, or at least bounded above, to form a hypothesis test.
Our results occur in the setting where one is concerned with minimizing a real-valued criterion function $\ell(\theta)$ over a parameter space $\Theta$. If $\ell$ can be estimated, consideration of the asymptotic distribution of the optimal value of the estimate in finite samples has been of interest since, at least, the $J$-test was proposed for overidentified GMM (Sargan1958, H1982). Lately, there has been extensive interest in extending this methodology to the partially identified setting, and fruitful work in both describing hypothesis tests (see S2012, CNS2023) and confidence regions (Tao2015, Zhu2020, Fan2023) therein. In practice, in order to test a given hypothesis, one often takes $\Theta$ to be a subset containing points of a larger, fixed parameter space which conform to that hypothesis. Our results can flexibly accommodate a range of parameter spaces, and can thereby facilitate tests of a range of analogous hypotheses. As an auxiliary exercise, we also analytically characterize the distance of certain random vectors from convex sets, such as those representing shape restrictions enforced on the random element (c.f.\ Fang2021).
As a point of departure, we note that $\ell$ can oftentimes be written as a supremum of a class of test functions $v$ indexed by $t \in \mathcal{T}$: \[ \ell(\theta) = \sup_{t \in \mathcal{T}} v(\theta, t). \] For instance, when $\ell$ is a norm or seminorm of a moment function of parameter $\theta$ over some Banach space, it always admits such a representation with linear $t$. This is true more generally if $\ell$ is any convex function of a moment function of $\theta$. Helpfully, such a decomposition will generally extend to empirical analogues of $\ell$, which will follow the corresponding form: \[ \ell_n(\theta) = \sup_{t \in \mathcal{T}} v_n(\theta, t), \] where $v_n(\theta, t)$ is an empirical analogue of test function $v$.
In order to compute the asymptotic distribution of a test statistic formed around the minimum value of $\ell_n(\theta)$ over $\Theta$, we take advantage of the fact that such a statistic will necessarily adopt a minimax formulation---the outer minimum being taken over parameter space $\Theta$, and the inner maximization being over its dual space $\mathcal{T}$. Such a representation is especially amenable to treatment by familiar tools arising from convex analysis. In particular, we show that an application of Sion's celebrated minimax theorem (Sion1958) leads to a streamlined way of bounding the distribution of such minimax test statistics from above, with asymptotic equality under certain rate or convexity conditions. Under some reasonable regularity conditions, these minimax bounds coincide with asymptotic approximations arising from S2012, hong2017, and CNS2023, and are closely tied to convex-analytical approaches to hypothesis testing (e.g.\ the test of shape restrictions in Fang2021). They provide a systematic basis for regarding such results which connects them back to the seminal work of Sargan1958 and H1982.
The intuition for our results in a simple, finite-dimensional setting can be stated as follows. If, say, $\sqrt{n}(v_n(\theta, t) -v(\theta, t))$ is an empirical process $\mathbb{G}_n(\theta, t)$ over $\Theta \times \mathcal{T}$ that converges to a tight limit $\mathbb{G}$, one can show by Taylor's theorem that the minimized value of $\ell$ over $\Theta$ may be rewritten as:
where we have employed in the last line the fact that a typical solution to the optimization problem will converge in probability towards $\Theta_0$. In the previous display, $(c_n)$ is a sequence diverging to $\infty$.
Equipped with (ref), previous work has provided means for estimating the image of the Jacobian $\frac{\partial v}{\partial \theta}$, which may be substituted with a Fr\'{e}chet derivative in infinite dimensional settings (S2012, hong2017, Fan2023, CNS2023). As the limiting distribution of $\mathbb{G}_n(\theta, t)$ can usually be consistently estimated with, say, the bootstrap, this provides a convenient means of estimating the asymptotic distribution of a minimax test statistic. However, this approach is disadvantaged by its reliance on the local linearity of $v$ in neighborhoods of the identified set. For instance, S2012 studies a setting in which $v$ is linear in the choice of parameter $\theta$. This assumption is relaxed in hong2017 and CNS2023, but finite sample analysis of the derivatives of $v$ is still e.g.\ generally dependent upon the use of linear sieve approximations of $\Theta$. Thus, the flexible employment of neural networks and associated nonlinear estimated schemes, which is highly desirable in nonparametric settings, is precluded.
This paper observes that there is a preponderance of $\ell$ and $v$ in the literature which satisfy the necessary convexity properties to rewrite (ref) as
As $c_n$ is a diverging sequence, asymptotically, the parameter $t$ in the supremum of (ref) will be taken to satisfy $\frac{\partial v(\theta, t)}{\partial \theta} =0$. As a consequence, we show that under general conditions, (ref) may be rewritten as
(ref) compares favorably to (ref) in several ways. In many popular Hilbert space settings, it shows that a test statistic based on (ref) has a projection interpretation which extends the $\chi^2$-limiting distribution of the $J$-test. Moreover, we show that it is quite straightforward to obtain estimates for critical values using (ref) without any direct computation of the derivatives of $\frac{\partial v}{\partial \theta}$ at all, even in settings with potentially nonlinear sieves. The minimax approach also allows for approximation methods devised for seminorm-based criterion functions to be extended quite broadly and systematically to, say, criterion functions involving general convex functions. Finally, our inference results accommodate the adversarial employment of nonconvex sieve spaces of test functions, such as neural networks. Thus, they blend a growing literature on adversarial econometric estimation (see K2023 and references therin) with inference in the partially identified setting.
The structure of the paper is as follows. Section (ref) formalizes our model and provides a method for obtaining the asymptotic distribution, or upper bounds thereof, of minimax test statistics. Section (ref) shows that, under general conditions, critical values can be obtained for those distributions using the bootstrap. The main focus of the paper is on single hypothesis tests, but Section (ref) of the appendix extends our distributional results uniformly over a class of parameter spaces and underlying probability distributions. Proofs are relegated to Section (ref) in the appendix.
Given a parameter space $\Theta$, our basic model is one where an identified set $\Theta_0 \subset \Theta$ is characterized as the set of all $\theta \in \Theta$ satisfying
where $\ell: \Theta \rightarrow \mathbb{R}$ is some function. We impose the additional restriction that $\ell$ takes the form
where $\{v(\cdot, t) : t \in \mathcal{T}\}$ is a set of test functions, and $\mathcal{T}$ is some set dual to $\Theta$, motivated by the idea that one often wants to investigate sets identified by criterion functions of this form:
Note that, in every case of example (ref), all of the functions $v$ are linear in the parameter $t$.
Suppose that $\ell$ cannot be directly observed, but a researcher has access to a noisy estimate $\ell_n$ of $\ell$ which satisfies
where $v_n$ is a noisy estimate of $v$. Our approach to inference will consist of finding estimates for the asymptotic distribution of test statistic
for an appropriate normalizing sequence $r_n$ (most often $\sqrt{n}$), and then determining critical values for this distribution by bootstrapping.
Our minimax approach allows for the straightforward calculation of asymptotic distributions of distances to convex sets. Fang2021 characterize and estimate such asymptotic distributions when the sets are convex cones in Hilbert spaces with the intent of testing shape restrictions. The distributions which arise are continuous functions of (usually) tight Borel measures which may be directly estimated with the bootstrap.
In a Banach space $\mathfrak{X}$, the function that assigns to every point its distance to some fixed convex set $C$ is convex, and Example (ref) suggests that it is compatible with the setting of the paper. Indeed, let $m(\theta)$ be an element of a Banach space $\mathfrak{X}$ for every $\theta \in \Theta$ and $C \subset \mathfrak{X}$ a convex set. Then, letting $\ell(\theta)$ denote the distance of $m(\theta)$ to $C$ and $\mathcal{T}$ the unit ball of the dual $\mathfrak{X}^*$ of $\mathfrak{X}$, equipped with the weak-* topology, one can write
Suppose that $m_n(\theta)$ is a random estimate of $m(\theta)$ satisfying that $r_n(m_n(\theta) - m(\theta)) \rightsquigarrow \mathbb{G}(\theta)$ for a sequence $(r_n)$, where `$\rightsquigarrow$' signifies weak convergence to a random element $\mathbb{G}(\theta)$. Then, an empirical analogue of $\ell(\theta)$ is
Fang2021 principally provide means of computing the limiting distribution of $r_n \ell_n(\theta)$ for a fixed $\theta$. To motivate the results of the paper, we show that we can provide such computations in a very broad range of cases.
The easiest case to consider is the one where $C$ is a convex cone, closed under dilations of its elements by positive multiples. Let $\delta^*(\cdot|C):t \mapsto \sup_{c \in C} t(c) $ denote the support function of $C$ (see Rock1970, \S13). For a convex cone $C$ in a real Banach space $\mathfrak{X}$, define the polar $C^\circ \subset \mathfrak{X}^*$ following fabian2011 as
(as opposed to the absolute polar, e.g.\ C1994). Then, using weak-* continuity of the maps $t \in \mathcal{T}$ and weak-* compactness of $\mathfrak{X}^*$, the Sion minimax theorem (Sion1958, see also the proof of Lemma (ref) below) allows us to rewrite (ref) as
Note that, if $m(\theta)$ is actually an element of $C$, the last line of (ref) is bounded above by
(see Lemma (ref) in the appendix; throughout, we will find it convenient to use e.g.\ $d_{\mathfrak{X}}$ as the Hausdorff distance in $\mathfrak{X}$ between a point and a set). (ref) and (ref) hold irrespective of the underlying probability distribution $P$ of $m_n$ or of $\theta$. Therefore, from these bounds arises our first motivating result:
Proposition (ref) complements and extends Theorem 3.1 of Fang2021 in a Banach space setting featuring arbitrary parameter spaces. In particular, note that the inequality in the first line of (ref) becomes an equality precisely in the `least favorable' case discussed therein. This paper further develops the minimax rearrangement displayed in (ref) and provides asymptotic characterizations of a much broader class of test statistics, as well as means of estimating those limiting distributions. For instance, Theorem (ref) below allows (ref) to be further decomposed using the structure of $\Theta$ as it relates to $C$.
The next proposition provides a similar characterization when $C$ is more generally held to be a convex set in $\mathfrak{X}$. For such a $C$ and a point $c \in C$, define the tangent cone $T_cC$ of $C$ at $c$ to be the subset of $\mathfrak{X}$ consisting of all points taking the form $\alpha (c' - c)$, $c' \in C$, $\alpha > 0$ (c.f.\ Rock1970, Theorem I.2.6.3). Similarly, define the normal cone $N_cC$ of $C$ at $c$ to be the subset $(T_cC)^\circ$ of $\mathfrak{X}^*$. For conciseness, we say that a class of random sequences $(X(\theta, P))$ indexed by $\theta \in \Theta, P \in \mathcal{P}$ is uniformly pre-tight in $\mathfrak{X}$ if, for every $\varepsilon > 0$, there is some totally bounded and measurable subset $S$ of $\mathfrak{X}$ which satisfies $\inf_{P \in \mathcal{P}} \mathrm{P}_{P}\left( X(\theta, P) \in S, \, \forall \theta \in \Theta \right) > 1 - \varepsilon$. Sufficient conditions for uniform pre-tightness are given in VW1996 (e.g.\ Problem 1.12.1 therein).
In the statement of Proposition (ref), the map $\inf_{\theta \in \Theta}$ may be replaced by any other function that is Lipschitz continuous with respect to the uniform norm. When $\Theta$ is a singleton and the conditions of Proposition (ref) apply, (ref) forms a tighter lower bound than the right hand side of (ref), because $N_{m_P}C$ is the subset of $C^\circ$ consisting of $t$ for which $\left\langle m_P, t \right\rangle = 0$ when $C$ is a cone.
The lower hemicontinuity condition imposed by Proposition (ref) is important for a portion of our bootstrap consistency arguments and is further discussed in Section (ref). Lemma (ref) therein shows that the lower hemicontinuity condition is fulfilled if the correspondence sending pairs $(P,\theta)$ to the tangent cone $T_{m_P(\theta)}C \subset \mathfrak{X}$ is merely upper hemicontinuous with respect to the weak topology on $\mathfrak{X}$. On the other hand, by boundedness of the maps $t \in \mathcal{T}$, continuity of the map $(P, \theta, t) \mapsto \left\langle m_P(\theta), t \right\rangle$ may be ascertained by verifying norm-continuity of the map $(P,\theta, t) \mapsto m_P(\theta)$, or by confining the points $m_P(\theta)$ to lie in a norm-compact set and requiring $(P,\theta, t) \mapsto m_P(\theta)$ to be merely weakly continuous.\footnote{Sufficiency of the latter conditions can be verified using the Arzel\`{a} Ascoli theorem; see the discussion following Assumption (ref).}
To discuss inference around $T_n$, it is necessary to impose some structure on the spaces $\Theta$ and $\mathcal{T}$. We continue to let $\Theta_0 \subset \Theta$ be the set of $\theta$ satisfying (ref), and say that $\Theta$ is locally convex in a neighborhood of $\Theta_0$ if it is a subset of a real vector space and, for all $\theta \in \Theta_0$, there is some $\varepsilon(\theta) > 0$ such that the intersection of $\Theta$ with an $\varepsilon(\theta)$-ball around $\theta$ is convex. All Banach spaces concerned in this paper are presumed to be over the real numbers, although it would be straightforward to extend our results to vector spaces over $\mathbb{C}$.
Assumption (ref) imposes some regularity on $\Theta$ and $\mathcal{T}$ by requiring that they are convex subsets of linear topological spaces (at least in a neighborhood of the identified set). We have remarked that, when the asymptotic distribution of a test statistic following (ref) is of interest, the set $\Theta$ is often chosen relative to a hypothesis that is imposed on a larger parameter space. Therefore, it is essential that Assumption (ref) should be made as accommodating in its choice of $\Theta$ as possible. Firstly, we note that what is actually required is that $\Theta$ is star-shaped in a neighborhood of all points in the identified set, so that derivatives can be defined and considered on those neighborhoods (but convexity is hardly more onerous an assumption). The requirement for local convexity can also be achieved by choosing an appropriate embedding of $\Theta$ into a Banach space, as the following example illustrates:
We also require that our parameter space $\Theta$ is a compact subset of a Banach space, over which we can eventually define derivatives. See FM2019 for a discussion of compactness in Banach spaces. Note that, if $\Theta$ is not necessarily convex, the closed convex hull of $\Theta$ is still compact (A1999, Theorem 5.35), so it may suffice to consider the convex closure of $\Theta$. It is also often the case that the map $v(\theta, \cdot)$ is at least quasi-concave in $t$ (in fact, every case of Example (ref) featured $v$ that were linear over parameter $t$), in which case it is equivalent to work with the convex hull of $\mathcal{T}$ in (ref). Note that the Donsker properties we require in Assumption (ref) are preserved under such set operations (VW1996 \S2.10).
It would likely be possible to extend the results of this paper to noncompact settings using bounded entropy conditions and finite sample concentration inequalities coming from empirical process theory (e.g.\ CNS2023, Assumption 3.3 and references therein). However, imposing compactness greatly streamlines the exposition and proofs of this paper, and is congruous with assumptions that are commonly imposed for the purposes of estimation and inference.
Throughout, we will let $d_{\mathfrak{B}}$ denote the metric induced by $\norm{\cdot}_{\mathfrak{B}}$ on $\mathfrak{B}$. The requirements that $\ell$ is lower semicontinuous and nonnegative are not very restrictive. In many examples, $\ell$ can be expected to be continuous in parameter $\theta$. For instance, this is the case when $\ell(\theta)$ is some seminorm of a continuous function $m$ of $\theta$.
In order to estimate the identified set $\Theta_0$, one must introduce an estimator $\ell_n$ of the criterion function $\ell$. Suppose that this is accomplished by introducing an estimator $v_n$ of the functions $v$, which obeys a central limit theorem in the sense of empirical process theory (see VW1996, Section 1.5):
Most often, one takes $r_n = \sqrt{n}$, and $\mathbb{G}$ is a Gaussian process. There are a number of ways of guaranteeing the convergence of Assumption (ref). For instance, if one has $v(\theta, t) = \mathrm{E}\left[ g(Y,\theta,t) \right]$ for some function $g$ and $v_n(\theta, t)$ is an empirical analogue of $v$, then Assumption (ref) holds if the class of maps $g$ is $P$-Donsker over $\Theta \times \mathcal{T}$ (VW1996, \S2). The delta-method for empirical processes extends such results to more general settings where $v$ is a nonlinear function of population moments, and $v_n$ its empirical analogue. In this setting, $\rho$ can be taken to be the pseudometric induced by seminorm $(\theta, t) \mapsto \mathrm{E}\left[ (g(Y, \theta, t) - \mathrm{E}\left[ g(Y, \theta, t) \right])^p \right]theta, t) \right])^p}^{1/p}$, for $p \ge 1$ (VW1996 Sections 1.5, 2.1).
We now wish to approximate the functions $v(\cdot, t)$ with linear approximations taken at points in the identified $\Theta_0$. A standard tool for this purpose is the Fr\'{e}chet derivative $D$, although we allow for a slight relaxation of Fr\'{e}chet differentiability in the following assumption. We say that a function $D$ defined on a real vector space $\mathfrak{B}$ is positive homogeneous if $D(\lambda h) = \lambda D(h)$, for all $h \in \mathfrak{B}$ and $\lambda \ge 0$.
The first part of Assumption (ref) requires that the functions $v(\cdot, t)$ can be approximated locally by positive homogeneous functionals $D_{\theta; t}$ in neighborhoods of points $\theta$ in the identified set. When $D_{\theta; t}$ is required to be bounded and linear, Assumption (ref) is principally the requirement that function $v$ is Fr\'{e}chet differentiable for all $\theta \in \Theta_0$ and $t \in \mathcal{T}$, with a uniform bound on the linear approximation provided by the derivative. Note that, if the left hand side of (ref) is $o(\delta)$, then (ref) trivially holds by defining $f_v(\delta)$ to be equal to its left hand side.
By Taylor's theorem, if $\Theta$ and $\mathcal{T}$ are subsets of Euclidean space and $v$ is smooth, Assumption (ref).1 may be established by uniformly bounding the second derivatives of $v(\cdot, t)$ over $\Theta_0$ and $\mathcal{T}$. In this case, one may take $f_v(\delta) = \delta^2$. Note that, if $\Theta_0$ is a singleton $\{\theta_0\}$ (point identification holds), Assumption (ref).1 is implied by Fr\'{e}chet differentiability at $\theta_0$ of the map $\theta \mapsto v(\theta, \cdot)$ from $\Theta$ to the set of bounded maps from $\mathcal{T}$ to $\mathbb{R}$ equipped with uniform norm.
The second part of Assumption (ref) appears onerous, but we stress that in all of the situations discussed in Example (ref), the function $v$ is linear over the parameter $t$ for fixed $\theta$. For $t, t' \in \mathcal{T}$ and $\alpha \in [0,1]$, passing to limits with Assumption (ref).1 and linearity of $v(\theta, \cdot)$ in hand allows one to write
Hence, it may be reasoned that, in many settings, $D_{\theta; t}v (h)$ is actually linear in parameter $t$, whence concave. The same applies to $v_n(\theta,t)$, which is often itself linear in $t$ (especially when $v(\theta, t)$ is linear in $t$). In particular, if we are in one of the situations delineated by Example (ref):
for some function $m$, then we may take $v_n(\theta, t) = t(m_n(\theta))$, where $m_n$ converges to $m$ as an empirical process. This makes $v_n$, and the whole expression $v_n(\theta, t) + D_{\theta; t} v(\theta')$, linear (whence, concave) in $t$.
Our final main assumption is placed on the derivative invoked in Assumption (ref). We are interested in the magnitude of the derivative, measured as how much it dilates points close to $0$. One may impose the usual dual norm on the functional $D_{\theta; t}v$ by letting $\norm{D_{\theta; t}v} = \sup_{\norm{h}_{\mathfrak{B}} \le 1} D_{\theta; t}v (h)$, although this norm might be stronger than necessary. Instead, we are concerned with the behavior of $D_{\theta; t}v$ only on the space tangent to $\Theta$ at $\theta$, and only in the negative direction. Formally, for each $\theta$ and $t$, we define
If $\Theta$ is convex in a neighborhood of $\theta$, it is straightforward to show that the term inside the limit is nonincreasing as $\delta \rightarrow 0$, so the limit must exist. Indeed, under the preceding assumptions it is immediate (see the proof of Theorem (ref)) that $\gamma(\theta, t)$ is the decreasing limit of $\inf_{\substack{\theta' \in \Theta \\ \norm{\theta' - \theta}_{\mathfrak{B}} \le \delta }} \delta^{-1} D_{\theta; t}v (\theta' - \theta)$ for $\delta < \underline{\delta}$, where $\underline{\delta} > 0$ can be chosen uniformly for $\theta$ and $t$, so that the $\liminf$ in (ref) can be replaced with a $\lim$.
Evidently, one has $\gamma \le 0$ and $0 \le |\gamma(\theta, t)| \le \norm{D_{\theta; t}v}_{\mathfrak{B}^*}$.\footnote{We impose the convention $\frac{0}{0} = 0$.} Also, when $\Theta$ is such that the ball $B(\theta, \delta) \subset \mathfrak{B}$ is contained in $\Theta$ for some $\delta > 0$, one immediately has by definition of the dual norm that $\gamma(\theta, t) = - \norm{D_{\theta; t}v}_{\mathfrak{B}^*}$.
In the statement of the next assumption, we write that pseudometric $\rho$ is continuous with respect to a topology $\mathcal{V}$ on $\mathcal{X}$ if, whenever $x_\alpha \rightarrow x$ is a convergent net in $\mathcal{X}$, one has $\rho(x_\alpha, x) \rightarrow 0$.
Assumption (ref) requires the existence of a topology on $\mathcal{T}$ that is strong enough that $t \mapsto v_n(\theta, t) + D_{\theta; t }v(\theta' - \theta)$ becomes upper semicontinuous for every $\theta$, but also weak enough such that $\mathcal{T}$ is compact. In the proof of Theorem (ref), it is shown that an essential implication of Assumption (ref) is that the map $t \mapsto \gamma(\theta, \cdot)$ also is $\mathcal{U}$-continuous.
We now follow Example (ref) in outlining some instances in which Assumptions (ref) is met:
The preceding assumptions are sufficient to provide an upper bound for the asymptotic distribution of $T_n$. We now state a condition that turns the upper bound into an equality, which references the function $f$ introduced in Assumption (ref).
In some cases, $v(\theta, \cdot)$ vanishes for $\theta \in \Theta_0$, and Assumption (ref).1 is vacuously true. For instance, consider the first part of Example (ref), where $v(\theta, t) = t(m(\theta))$ for some linear functional $t$, and $\ell(\theta)$ is some seminorm of $m(\theta)$. Suppose that $\ell(\theta)$ is a norm, and vanishes if and only if $m(\theta)$ does, as is the case with GMM estimation. Then $\theta$ is in $\Theta_0$ if and only if $m(\theta) = 0$, which is true if and only if $v(\theta, \cdot) = t(m(\theta)) = 0$ for all linear functionals $t$. The discussion following Assumption (ref) provides a number of other situations in which Assumption (ref).1 is satisfied.
Assumption (ref).2 may be fulfilled in two ways. The first is by supplying a convergence rate for an approximate minimizer $\widehat{\theta}_n$ of the criterion function $\sup_{t \in \mathcal{T}} v_n(\widehat{\theta}_n, t)$ to the identified set $\Theta_0$. If, for instance, $r_n = \sqrt{n}$ and the function $f$ bounding the approximation error given in Assumption (ref) is the map $f_v(\delta) = \delta^2$, the first case of Assumption (ref).2 is fulfilled if one has $d_{\mathfrak{B}} (\widehat{\theta}_n, \Theta_0) = o_p(n^{-1/4})$. If $\Theta_0 = \{\theta_0\}$ is a singleton and $\widehat{\theta}_n$ signifies an $M$-estimator of $\theta_0$, then $\sqrt{n}$-consistency of $\widehat{\theta}_n$ (or consistency at any rate faster than $n^{-1/4}$) is sufficient to meet this first case. S2012, hong2017, and CNS2023 use similar rate conditions to ensure that the parameter space local to $\Theta_0$ can be adequately approximated with linear expansions around points in the identified set.
The second case of Assumption (ref).2 is most simply fulfilled when $v(\theta,t)$ is convex in $\theta$ for all $t \in \mathcal{T}$. In some settings, $v(\theta, t)$ is linear in both $\theta$ and $t$---take for instance the NPIV model and related formulations discussed in Example (ref). Note that, when $v$ is bilinear, its Fr\'{e}chet derivative $D$ satisfies $D_{\theta; t}v(\theta' - \theta) = v(\theta' , t) - v(\theta, t)$, so one can take $f = 0$ in Assumption (ref). This automatically satisfies the rate condition of the first case of Assumption (ref).2.
We now state a result which bounds the asymptotic distribution of $T_n = r_n\inf_{\theta \in \Theta} \sup_{t \in \mathcal{T}} v_n(\theta, t)$, and gives an exact limiting distribution under the additional imposition of Assumption (ref)
Examples (ref) and (ref) have discussed several instances in which one has $v(\theta, t) = t(m(\theta))$ for $m$ a mapping from $\Theta$ to a Banach space $\mathfrak{X}$, and $t$ a continuous affine function from $\mathfrak{X}$ to $\mathbb{R}$. When $t$ is continuous and affine, the map $x \mapsto t(x) - t(0)$ is linear (Rock1970, \S 1). Suppose $m$ is Fr\'{e}chet differentiable, and let $\nabla m$ denote the Fr\'{e}chet derivative of $m$ at $\theta$. Then, a straightforward calculation with (ref) shows that
for all $\theta \in \Theta_0$ and $t \in \mathcal{T}$. Therefore, we may rewrite the defining equation for $\gamma$ and infer that $K(\theta)$ is the set of $t$ which satisfy
for all $\theta'$ which are close enough to $\theta$ in the $d_\mathfrak{B}$-metric (the notion of “close enough" may be defined uniformly for $\theta \in \Theta_0$---see the proof of Theorem (ref)). (ref) has applications to several hypothesis testing scenarios of interest. An interesting case arises when $\mathfrak{X}$ is a Hilbert space, and the moment functions $m_n(\theta)$ converges weakly to a tight $L^\infty(\Theta)$-valued element $\mathbb{W}_n(\theta)$:
Corollary (ref) complements tests of overidentifying restrictions provided under the assumption of point identification with method of moments estimators (e.g.\ Sargan1958 and H1982 for a generalization). In particular, if $\mathfrak{X}$ is some Euclidean space $\mathbb{R}^k$, and $r_n(m_n(\theta) - m(\theta))$ has a limiting $N(0, I_{k \times k})$ distribution (as is the case with efficiently weighted GMM), then the upper bound implied by Corollary (ref) has an asymptotic distribution characterized by writing
where $\mathbf{Z}$ is a standard normal random vector. Under point identification and a full-rank condition on $\nabla m$, (ref) states the limiting $\chi^2$ distribution of the $J$-statistic of overidentifying relations. We note that, under the usual assumptions of GMM, one typically verifies that Assumption (ref) holds, so that the asymptotic distribution of $T_n$ is precisely (ref).
It is also fruitful to examine the result of Corollary (ref) when $\mathfrak{X}$ is more generally held to be a Banach space and $\ell$ the composition of a moment function with some convex function over the Banach space (like a norm). Continue to suppose that $v(\theta, t) = t(m(\theta))$ for $t$ in the dual $\mathfrak{X}^*$ of some Banach space $\mathfrak{X}$. S2012, hong2017, CNS2023, and Fan2023, among other papers, provide Banach space bounds for the distribution of $T_n$ which have in common the structure:
where $S_\theta$ is some subset of the tangent space $T_\theta \Theta$ of $\Theta$ at $\theta$ (or its closure), which for our purposes may best be defined as the convex cone:
(convexity follows under local convexity of $\Theta$ and Rock1970, Theorem 2.5; note that $T_\theta \Theta$ may not be a subspace if, for instance, $\theta$ is contained in the boundary of $\Theta$). In the aforementioned results, one typically defines $S_\theta$ as the closed limit of linear subspaces formed by linear sieve approximations to $\Theta$.
Let $h \in \overline{T_\theta \Theta}$ be a nonzero element of the closure of the tangent space of $\Theta$ at $\theta$. Then, there exists a sequence of $\theta'_k \in \Theta$ converging to $\theta$, and a diverging sequence of $\lambda_k$, satisfying $\lambda_k (\theta'_k - \theta) \rightarrow h$ in $\mathfrak{B}$. If one lets $t$ be an affine function over $\mathfrak{X}$ and an element of $K(\theta)$, then by (ref) and (ref),
From (ref), one can readily show that the upper bound in (ref) is a lower bound for bounds of the type (ref), whence less conservative, especially under the assumptions of Theorem (ref):
By Theorem (ref), critical values for the distribution of \[ \inf_{\theta \in \Theta_0} \sup_{t \in K(\theta)} \mathbb{G}(\theta, t) \] can be used to describe valid, if conservative, critical values for the test statistic $T_n$. The chief challenge in determining these critical values is in estimating the sets $\Theta_0$ and $K(\theta)$ for each point in $\Theta_0$, but this is not onerous if uniformly consistent estimators for the criterion function $\ell$ and the derivative-norm map $\gamma$ are available. The former can be estimated using $\ell_n$ under our assumptions, whereas an estimator for the latter can be obtained if one has access to an estimate $\nabla v_n$ of $v$.
Our recommended approach for inference in this partially identified setting (see especially S2012, hong2017, Zhu2020) is to apply the bootstrap. For a bootstrapped statistic $Z_n^*$ and a limit distribution $Z$, we follow VW1996 in writing that $Z_n^* \overset{\mathrm{P}}{\rightsquigarrow} Z$ if the convergence $d_{\mathrm{BL}_1}(Z_n^*, Z) = \sup_{h \in \mathrm{BL}_1} \mathrm{E}^*[h(Z_n^*)] - \mathrm{E}\left[ h(Z) \right]\overset{p}{\rightarrow} 0$ (in outer probability) holds in the bounded lipschitz metric $d_{\mathrm{BL}_1}$, where $\mathrm{BL}_1$ is the set of 1-Lipschitz functions bounded in magnitude by $1$, and $\mathrm{E}^*$ is an expectation conditional upon the sample. We write $Z_n^* \overset{\mathrm{a.s.}}{\rightsquigarrow} Z$ if the same convergence holds almost surely with regard to outer probability. With some abuse of notation, we let expectations and probabilities refer to outer measure when quantities are asymptotically measurable.
The following assumptions and theorem demonstrate a strategy for approximating the asymptotic distribution of $T_n$ using the bootstrap. Hemicontinuity in our assumptions is relative to topology $\mathcal{U}$ on $\mathcal{T}$ and the $d_{\mathfrak{B}}$ metric on $\Theta$ (see A1999).
The first two parts of Assumption (ref) are fairly light bootstrap consistency and regularity conditions. The functions $\psi_n$ can be taken to be $\gamma$ if the latter can be directly estimated (e.g.\ by computing the Fr\'{e}chet derivatives of $v$). In other cases, it may be more appropriate to let
where $(\delta_n)$ is an sequence of constants converging slowly enough to $0$ so that $\psi_n$ can be adequately estimated. Such a choice of $\psi_n$ satisfies the properties indicated in Assumption (ref) and is further discussed below, along with the corresponding sequence $(q_n)$.
The third part of Assumption (ref) consists of two possible cases. The first is satisfied under a hemicontinuity condition on the correspondence ${K}(\theta)$ with respect to a topology $\mathcal{V}$, which can be either the metric topology associated with $d_{\mathfrak{B}}$ or a stronger topology. Suppose that $\Theta$ and $\mathcal{T}$ are finite dimensional spaces, and assume an appropriate degree of smoothness for the map $(\theta, t) \mapsto D_{\theta; t} v$ (the right hand side is the Jacobian of $v(\cdot, t)$ evaluated at $\theta$). We have remarked that, as long as $\Theta_0$ is contained in the interior of $\Theta$, $K(\theta)$ is defined as the zero set of the vector-valued map $(\theta, t) \mapsto D_{\theta; t}v$. Its lower hemicontinuity may thus be established as a consequence of the implicit function theorem if the Jacobian of this map, regarded as a function in $t$, has full rank at all points $(\theta, t)$, $\theta \in \Theta_0, t \in K(\theta)$ (e.g.\ A1999). In more general settings, lower hemicontinuity of the correspondence can be ascertained using infinite-dimensional versions of the implicit function theorem (e.g.\ Loomis1990). Section (ref) below gives equivalent characterizations for the lower hemicontinuity of the correspondence $K$ in the strong (norm) topology on any Banach space.
The second case of Assumption (ref).3 is essentially a requirement that the function $x \mapsto \sup_{\theta: \ell(\theta) \le x} d_{\mathfrak{B}}(\theta, \Theta_0)$ can be bounded above up to some constant factor by a nondecreasing function $f_\ell$ (note that one can take $f_\ell$ to be exactly this increasing function, so that $f_\ell$ always exists). This requirement of local identification in a neighborhood of $\Theta_0$ resembles similar conditions imposed under partial identification, such as Assumption 3.4 in CNS2023 and Assumption B.3 in Zhu2020.
Assumption (ref) is sufficient for the estimation of an upper bound of the distribution of $T_n$ using $K(\theta)$. Theorem (ref) demonstrates that a tighter bound may potentially be obtained using the smaller set $\tilde{K}(\theta)$, so we also state a supplementary assumption which enables such an estimate to be made. We have observed that, in many settings of interest, one has $v(\theta, \cdot) = 0$ whenever $\theta \in \Theta_0$. This forces the relation $\tilde{K}(\theta) = K(\theta)$ and renders the following assumption unnecessary.
Theorem (ref) includes as a corollary a simpler way of obtaining a critical values for the asymptotic distribution of $T_n$. The simplification omits the outer minimization over $\Theta$ in the left hand sides of (ref) and (ref) as long as an appropriate value of $\theta$ is used. When point identification holds so that $\Theta_0$ is a singleton, and some additional regularity on the functions $v, \psi$, and $\psi_n$ is imposed, it is straightforward to show that such a bootstrap procedure is asymptotically equivalent to the more computationally intensive approach indicated in Theorem (ref). The additional regularity comes in the form of convergence requirements upon the nonstochastic functions $\psi_n(\theta, t)$ to $\gamma$ and lower semicontinuity of $\gamma$:
It remains to show that one can construct a sequence of estimators $\widehat{\psi}_n$ of functions $\psi_n$ satisfying the second condition of Assumption (ref). By local convexity of $\Theta \subset \mathfrak{B}$ and positive homogeneity of $D_{\theta; t}$ (Assumption (ref)), we have remarked that the maps $\psi_{\delta} :(\theta, t) \mapsto \frac{D_{\theta; t} v(\theta' - \theta)}{\delta}$ are pointwise decreasing in $\delta$ for $\delta$ small enough. Moreover, each map $\psi_\delta$ may be estimated with a sample analogue using the estimate provided in Assumption (ref).1 with $v_n$ approximating $v$. Such a strategy provides a consistent sequence of estimators $\widehat{\psi}_n$ for the sequence $(\psi_n)$ defined in (ref) under the assumptions we have already introduced:
Lemma (ref) suggests that a tractable way of estimating e.g.\ (ref) is to solve the triple optimization problem:
where $\delta_n$ is such that $q_n \equiv \frac{f_v(\delta_n) + r_n^{-1}}{\delta_n} = o(1)$ and $\mu_n$ and $\lambda_n$ are as in Theorem (ref). In a remark following the proof of Lemma (ref), we demonstrate that, under convexity of $\Theta$ and Lipschitz continuity of $v$, one can introduce a Lagrangian form of $\widehat{\psi}_n$ with negligible error: \[ \tilde{\psi}_n(\theta, t) = \inf_{\substack{\theta' \in \Theta \\ \norm{\theta - \theta'}_{\mathfrak{B}}} } \frac{v_n(\theta', t) - v_n(\theta, t)}{\delta_n} + \nu_n \frac{(\norm{\theta' - \theta}_{\mathfrak{B}} - \delta_n)_+}{ \delta_n }. \] The imposition of lower hemicontinuity on $K$ and $\tilde{K}$ in Assumptions (ref) and (ref) bears significance for our estimation results, inasmuch as it greatly simplifies the choice of tuning parameters for the calculation of critical values. Therefore, it is important to provide conditions under which lower hemicontinuity can be expected to hold.
We turn again to the situation first described by Example (ref).2, and also in Lemma (ref), wherein $\ell = \gamma \circ m$ is the composition of a convex function $\gamma: \mathfrak{X} \rightarrow \mathbb{R}$ with a moment function $m: \mathfrak{B} \rightarrow \mathfrak{X}$ ($\mathfrak{X}$ a Banach space). By the Fenchel-Moreau theorem, the corresponding choice of $\mathcal{T}$ is a set of continuous affine maps $t : \mathfrak{X} \rightarrow \mathbb{R}$ majorized by $\gamma$. One also has $D_{\theta; t} v (h) = t(\nabla m(\theta)(h)) - t(0)$, and $t \in K(\theta)$ if and only if this quantity is nonnegative for all $h \in T_\theta \Theta$ (Remark (ref)). In particular, $t \in K(\theta)$ if and only if $t(x) - t(0) \ge 0$ for all $x \in \nabla m(\theta)(T_\theta \Theta)$, where the latter is a convex cone in $\mathfrak{X}$. In Lemma (ref) below, we characterize when exactly the set of linear functionals defined on $\mathfrak{X}$ satisfying this inequality can be expected to be lower hemicontinuous.
According to the discussion above, $-K(\theta) \subset \mathfrak{X}^*$ can be identified with the polar $(\nabla m(\theta)(T_\theta \Theta))^\circ$. Clearly, lower hemicontinuity of $-K(\theta)$ guarantees the lower hemicontinuity of $K(\theta)$.
Conveniently, one can provide conditions under which the polar of a correspondence sending points $\theta$ to cones $C(\theta) \subset \mathfrak{X}$ is lower hemicontinuous. These conditions depend on the upper hemicontinuous of the correspondence $\theta \mapsto C(\theta)$. In fact, we can show that a weaker version of upper hemicontinuity of this correspondence, which maps a topological space into a Banach space, is both necessary and sufficient for lower hemicontinuity of the polar correspondence in the strong topology on $\mathfrak{X}^*$. The weak version of hemicontinuity is relative to the weak topology on $\mathfrak{X}$. For the purposes of this paper, we will say that a set $A$ of a topological vector space $\mathfrak{X}$ is strongly contained in another set $V$ if there exists some open neighborhood $D$ of $0$ such that $A + D \subset V$. For instance, if $A$ is compact, strong containment of $A$ in $V$ is equivalent to containment if $V$ is open (C1994, Lemma IV.3.8).
If $\mathfrak{X}$ is a Banach space, we will refer to a correspondence $H$ mapping a topological space $\mathcal{A}$ to $\mathfrak{X}$ as being weakly upper-hemicontinuous at point $a$ if, whenever $H(a)$ is strongly contained in a weakly open $V$, there exists a neighborhood $U$ of $a$ such that $a' \in U$ implies that $C(a')$ is a subset of $V$. Such a weak notion of upper hemicontinuity in the correspondence $\theta \mapsto \nabla m(\theta)(T_\theta \Theta)$ is precisely what is needed to ensure (norm-topology) lower hemicontinuity of its polar. If $H$ maps $\mathcal{A}$ to a dual space $\mathfrak{X}^*$, say that it is weak-* upper hemicontinuous if it conforms to the usual definition of upper hemicontinuity with respect to the weak-* topology on $\mathfrak{X}^*$.
Lemma (ref) allows us to show that many choices of convex function $\gamma: \mathfrak{X} \rightarrow \mathbb{R}$ actually can be written in a way which makes the correspondence $\theta \mapsto K(\theta)$ lower hemicontinuous. For instance, in the following lemma, all that is required of $\gamma$ is that it is Lipschitz continuous in a neighborhood of $m(\Theta)$. For $\mathfrak{X}$ a Banach space and $\mathcal{A}$ the vector space of continuous and affine maps $t: \mathfrak{X} \rightarrow \mathbb{R}$, we define the weak topology on $\mathcal{A}$ to be the topology of pointwise convergence, and the norm topology to be the one generated by $\norm{t}_{\mathcal{A}} = \sup_{\norm{x}_{\mathfrak{X}} \le 1} |t(x)|$. For a convex function $\gamma$ defined on $\mathfrak{X}$, we also define the subgradient $\partial \gamma(x)$ of $\gamma$ at $x$ to be the set of $y \in \mathfrak{X}^*$ satisfying that $\gamma(x') \ge \gamma(x) + \left\langle x' - x, y \right\rangle$ for all $x' \in \mathfrak{X}$.
Lemma (ref) furnishes a host of examples of $\mathcal{T}$ corresponding to arbitrary convex functions defined on Banach spaces (Example (ref) presents several convex functions that are of particular interest). These $\mathcal{T}$ are at least compact in the topology of pointwise convergence, which the discussion following Assumption (ref) shows is generally sufficient to apply Theorem (ref). The lemma also characterizes the correspondence $\theta \mapsto \tilde{K}(\theta) \subset \mathcal{T}$ in terms of the subgradient of $\gamma$. The topology of weak convergence automatically makes $t \mapsto v(\theta, t) = t(m(\theta))$ a continuous map satisfying Assumption (ref).1, so that $\tilde{K}(\theta)$ is the correspondence most relevant to Theorem (ref).