EconBase
← Back to paper

A Generalized Argmax Theorem with Applications

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

53,686 characters · 8 sections · 71 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Generalized Argmax Theorem with Applications

abstractThe argmax theorem is a useful result for deriving the limiting distribution of estimators in many applications. The conclusion of the argmax theorem states that the argmax of a sequence of stochastic processes converges in distribution to the argmax of a limiting stochastic process. This paper generalizes the argmax theorem to allow the maximization to take place over a sequence of subsets of the domain. If the sequence of subsets converges to a limiting subset, then the conclusion of the argmax theorem continues to hold. We demonstrate the usefulness of this generalization in three applications: estimating a structural break, estimating a parameter on the boundary of the parameter space, and estimating a weakly identified parameter. The generalized argmax theorem simplifies the proofs for existing results and can be used to prove new results in these literatures.

\\ \\

{\bf Keywords:} Argmax Theorem, M-Estimator, Painlev\'{e}-Kuratowski Convergence, Structural Break, Change-Point Estimation, Parameter on the Boundary, Weak Identification.

Introduction

The argmax theorem is a useful result for deriving the limiting distribution of estimators in many applications. The argmax theorem starts with a sequence of stochastic processes, $\mathbb{M}_n(h)$, indexed by a metric space $H$, that converges, in some sense, to a limiting stochastic process, $\mathbb{M}(h)$. The usual statement of the argmax theorem imposes conditions so that $\hat h_n\in\text{argmax}_{h\in H} \mathbb{M}_n(h)$ converges in distribution to $\hat h\in\text{argmax}_{h\in H}\mathbb{M}(h)$. See KimPollard1990, VaartWellner1996, and Kosorok2008 for statements and applications of the argmax theorem.

We are concerned with generalizing the domain of maximization from $H$ to a sequence of subsets, $\Lambda_n$, converging, in some sense, to a limit subset, $\Lambda$. This generalization is natural because $h$ is often a local parameter that comes from rescaling the parameter space; see Section 3.1 in VaartWellner1996. Theorem (ref), below, shows that we only need $\Lambda_n$ to converge to $\Lambda$ in the Painlev\'{e}-Kuratowski (PK) sense. PK convergence is a weak type of convergence for a sequence of sets. Assuming only a weak type of setwise convergence is important for being able to verify it in a variety of applications.

We demonstrate the usefulness of this generalization in three applications. See Section 3 for the literature related to each application.

(1) In time series, estimating a structural break is estimating the date at which a change in the distribution of the sample occurred. The estimator maximizes over a discrete set of dates that are observed in the sample. The generalization of the argmax theorem given in Theorem (ref) is useful as the rescaled set of dates becomes denser and converges to an interval in the PK sense. The generalization is also useful when a trimming parameter is used to narrow the possible break dates to a middle fraction of the sample. Theorem (ref) implies new results when the trimming interval is misspecified to be too narrow.

(2) When estimating a parameter that is on or near the boundary of the parameter space, the rescaled parameter space is locally approximated by a tangent cone. Existing derivations of the limiting distribution require an extra step showing that this approximation is negligible. The generalization of the argmax theorem given in Theorem (ref) eliminates this extra step, simplifying the proof. The generalization also covers new results in non-regular cases, when the rescaled parameter space may converge to any closed set.

(3) When a parameter is weakly identified, the estimator converges at a slower rate than strongly identified parameters. To derive the limiting distribution, the rescaled parameter space needs to be scaled differently for strongly identified and weakly identified parameters. The generalization of the argmax theorem given in Theorem (ref) is useful for handling this scaling of the parameter space. The generalization also covers new results when inequalities are available that provide information about the true value of a weakly identified parameter.

Other related papers include Ferger2004 and SeijoSen2011, which generalize the argmax theorem to a non-unique argmax in the limit. More widely, there is a broadly related literature on the convergence of the value of a sequence of optimization problems following Berge1963. There is also a broadly related literature on the stability of nonlinear programming problems to perturbations of the objective function and constraints following EvansGould1970. We focus on the statement of the argmax theorem given in VaartWellner1996 because of its usefulness in the applications of interest.

Section 2 states Theorem (ref). Section 3 discusses the applications. Section 4 proves Theorem (ref). Section 5 concludes. An appendix contains additional proofs.

For clarity, we briefly introduce some notation used throughout the paper. We write $(a,b)$ for $(a',b')'$ for stacking two column vectors, $a$ and $b$. Given a function, $f$, we write $f(a,b)$ for $f((a,b))$ for simplicity. We write $d_a$ or $d_f$ to denote the dimension of a vector, $a$, or the dimension of the range of a function, $f$.

A Generalized Argmax Theorem

This section states a generalization of the argmax theorem. We allow the maximization to take place over a sequence of subsets, $\Lambda_n$, that converges to a limiting set, $\Lambda$, in the Painlev\'{e}-Kuratowski sense. We give the definition of PK convergence in a metric space $H$ equipped with metric $d$. For any $h\in H$ and $\Lambda\subset H$, let $d(h,\Lambda)=\inf_{g\in \Lambda}d(h,g)$.

definitionLet $\Lambda_n, \Lambda$ be subsets of a metric space, $H$, equipped with metric $d$. Define \begin{align*} \limsup_{n\rightarrow\infty}\Lambda_n&=\{h\in H: \liminf_{n\rightarrow\infty} d(h,\Lambda_n)=0\}\\ \liminf_{n\rightarrow\infty}\Lambda_n&=\{h\in H: \lim_{n\rightarrow\infty} d(h,\Lambda_n)=0\}. \end{align*} We say that $\Lambda_n$ Painlev\'{e}-Kuratowski (PK) converges to $\Lambda$, denoted by $\Lambda_n\rightarrow\Lambda$, if \[ \limsup_{n\rightarrow\infty}\Lambda_n=\liminf_{n\rightarrow\infty}\Lambda_n=\Lambda. \]

Remarks. (1) Definition (ref) is a standard definition given, for example, in AubinFrankowska1990. It follows from the definition that PK limits are always closed. Note that there is no ambiguity in using the “$\rightarrow$” notation for PK convergence because it applies only to sequences of subsets of $H$, while convergence in $H$ applies only to elements of $H$. Even then, PK convergence is equivalent to convergence in $H$ when applied to sets of singletons.

(2) PK convergence is known to be a weak type of setwise convergence. It implies other types of setwise convergence, such as Hausdorff convergence. In fact, PK convergence is so weak that it has a sequential compactness property: any sequence of subsets of a separable metric space has a PK limit along a subsequence; see Theorem 1.1.7 in AubinFrankowska1990. This property is used in the applications to verify PK convergence. \qed

We next state the generalized argmax theorem. Let $H$ be a metric space and let “$\rightsquigarrow$” denote weak convergence of random elements in $H$ using the Hoffmann-J\orgensen definition of weak convergence. For any compact $K\subset H$, let $\ell^\infty(K)$ denote the space of real-valued bounded functions on $K$ equipped with the supremum norm. We also let “$\rightsquigarrow$” denote weak convergence in $\ell^\infty(K)$. Let $o_P(1)$ denote a sequence of scalar random variables that converges in outer probability to zero.

theoremLet $\mathbb{M}_n,\mathbb{M}$ be stochastic processes indexed by a separable metric space $H$ such that $\mathbb{M}_n\rightsquigarrow \mathbb{M}$ in $\ell^\infty(K)$ for every compact $K\subset H$. Let $\Lambda_n,\Lambda\subset H$ be such that $\Lambda_n\rightarrow\Lambda$. Suppose almost all sample paths $h\mapsto\mathbb{M}(h)$ are continuous and possess a unique maximum over $h\in \Lambda$ at a random point $\hat h$, which as a random map in $H$ is tight. If the sequence $\hat h_n\in \Lambda_n$ is uniformly tight and satisfies $\mathbb{M}_n(\hat h_n)\ge \sup_{h\in \Lambda_n}\mathbb{M}_n(h)-o_P(1)$, then $\hat h_n\rightsquigarrow \hat h$ in $H$.

Remarks: (1) The statement of Theorem (ref) should be compared to the statement of Theorem 3.2.2 in VaartWellner1996. The primary difference is that the maximization is taken over $\Lambda_n$, which is assumed to PK converge to $\Lambda$. This is more general because one can take $\Lambda_n=\Lambda=H$ for all $n$ to return to the original statement. The trade-off for this generality is the assumption that $H$ is separable and $\mathbb{M}(h)$ is almost surely continuous. This seems to be a small price to pay considering the usefulness of Theorem (ref) in the applications.

(2) The proof of Theorem (ref) does not follow directly from the argument used to prove Theorem 3.2.2 in VaartWellner1996. Section 4 gives the proof and details the modifications required.

(3) Lemma 9.10 in AndrewsCheng2012 is a similar generalization of the argmax theorem. The given proof contains a mistake. Specifically, AndrewsCheng2012 invoke the extended continuous mapping theorem (Theorem 1.11.1 in VaartWellner1996) to get

equation[equation omitted — 112 chars of source]

for an arbitrary closed set, $F$. The argument used to verify the condition of the extended continuous mapping theorem presumes $\Lambda_n\cap F\rightarrow \Lambda\cap F$, but this does not follow. For example, if $\Lambda_n=\{1/n, 1-1/n\}$, $\Lambda=\{0,1\}$, and $F=[0,1/2]\cup\{1\}$, then $\Lambda_n\cap F=\{1/n\}\rightarrow\{0\}\neq\{0,1\}=\Lambda\cap F$. Fortunately, Lemma 9.10 in AndrewsCheng2012 is still true, as long as $H$ is separable, because the proof of Theorem (ref) given in Section 4 provides a solution. \qed

Applications

This section reviews three applications of Theorem (ref) and shows how the generalization is useful.

Structural Break Estimation

Researchers using time series data often suspect the value of a parameter changed at some date in the sample. Structural break estimation is an attempt to estimate the date at which the change occurred. Structural break estimation is a commonly used statistical technique with a large statistical literature; see CasiniPerron2019 for a recent survey. Here, following Bai1997, we consider a linear regression model with a break in the value of the coefficients. We derive the limiting distribution of the estimator of the break date using Theorem (ref).

We start with a sample, $(y_t,x_t)$ for $t=1,...,T$, where $y_t$ is scalar and $x_t$ is a $p$-vector. The model is defined by a linear regression of $y_t$ on $x_t$ with a break in the value of the coefficients. Suppose

align[align omitted — 165 chars of source]

where $\beta$ and $\delta$ are unknown $p$-vectors of coefficients and $k_0$ is the unknown break date. For any fixed value of $k$, we can estimate $\beta$ and $\delta$ by defining $z_t=x_t\mathds{1}\{t>k\}$ and running a regression of $y_t$ on $x_t$ and $z_t$. Let $\hat \delta_k$ denote the estimator of $\delta$ given the value of $k$.

For any fixed value of $k$, let $V_T(k)=\hat\delta'_k(Z'_kMZ_k)\hat\delta_k$, where $X=(x_1,...,x_T)'$, $M=I-X(X'X)^{-1}X'$, and $Z_k=(0,...,0,x_{k+1},...,x_T)'$. We define $\hat k$ to maximize $V_T(k)$ over a set of possible $k$ values. This is equivalent to minimizing the sum of squared residuals; see equation 5 in Bai1997. The maximization set is given by $\Lambda_T=\{[\lambda_1 T], [\lambda_1 T]+1, ..., [\lambda_2 T]\}$, where $[\cdot]$ is the greatest integer function and where $\lambda_1\in(0,1)$ and $\lambda_2\in(\lambda_1,1)$ are trimming parameters. Trimming parameters are commonly used to narrow the possible values of the break date to a middle $\lambda_2-\lambda_1$ fraction of the sample. This can avoid some technical complications that arise near the beginning/end of the sample.

We next discuss the break date. Let $v_T$ be a sequence of positive constants such that $v_T\rightarrow 0$ and $T^{1/2}v_T\rightarrow\infty$. We consider two possible assumptions for the timing of the break date.

assumptionThe break date, $k_0$, depends implicitly on $T$ and satisfies one of the following. \begin{enumerate}[label=(\alph*)] • $k_0/T\rightarrow \tau$ for some $\tau\in(\lambda_1,\lambda_2)$. • $v_T^2(\lambda_2T-k_0)\rightarrow a$ for some $a\in\mathbb{R}$. \end{enumerate}

Remarks. (1) Assumption (ref)(a) requires the true break date to be strictly within the interval defined by the trimming parameters. In this case, the derivation of the limiting distribution of $\hat k$ is closely related to the proof of Proposition 3 in Bai1997. We show how Theorem (ref) can be used to simplify the argument.

(2) Assumption (ref)(b) considers the case that the true break date is near the end of the interval defined by the trimming parameters. When $a$ is positive, the true break date belongs to the interval but is close to the end. When $a$ is negative, the true break date lies outside the interval. One can view this as a local misspecification of Assumption (ref)(a) where the trimming interval has been chosen to be too narrow. \qed

The goal is to derive the limiting distribution of $\hat k$. We use $\delta_T$ to denote the true value of $\delta$ in ((ref)). We consider asymptotics where $\delta_T$ depends on the sample size. For $i\in\{1,2\}$, let $B_i(r)$ be a $p$-dimensional multivariate Gaussian processes on $[0,\infty)$ with mean zero and covariance $\mathbb{E}B_i(u)B_i(v)'=min(u,v)\Omega_i$, where $\Omega_i$ is positive definite and $B_i(r)$ is independent across $i$. We use the following assumptions.

assumption\begin{enumerate}[label=(\alph*)] • $\delta_T=\delta_0 v_T$ for some $\delta_0\neq 0$. • $\hat k=k_0+O_P(v_T^{-2})$. • For every $J\in\{p,...,T-p\}$, $\sum_{t=1}^Jx_tx'_t$ and $\sum_{t=J+1}^Tx_tx'_t$ are positive definite with probability 1. • For every $J_T\le k_0$ such that $J_T\rightarrow\infty$, $J_T^{-1}\sum_{t=k_0-J_T+1}^{k_0}x_tx'_t\rightarrow Q_1$ and for every $J_T\le T-k_0$ such that $J_T\rightarrow\infty$, $J_T^{-1}\sum_{t=k_0+1}^{k_0+J_T} x_tx'_t\rightarrow Q_2$, where $Q_1$ and $Q_2$ are positive definite. • $k_0^{-1/2}\sum_{t=1}^{k_0}x_t\epsilon_t=O_P(1)$ and $(T-k_0)^{-1/2}\sum_{t=k_0+1}^{T}x_t\epsilon_t=O_P(1)$. • For every $C<\infty$, \begin{align} v_T\sum_{t=[k_0-rv_T^{-2}]+1}^{k_0}x_t\epsilon_t&\rightsquigarrow B_1(r)\nonumber\\ v_T\sum_{t=k_0+1}^{[k_0+rv_T^{-2}]}x_t\epsilon_t&\rightsquigarrow B_2(r)\nonumber \end{align} in $\ell^\infty([0,C])$ jointly, for $r\in[0,C]$. \end{enumerate}

Remarks. (1) Assumption (ref)(a) is a restriction on the break size. It requires the break to be small, in the sense that the break size, $\|\delta_T\|$, converges to zero. It also requires the break size to not be too small, in the sense that $T^{1/2}\|\delta_T\|\rightarrow\infty$. This assumption is commonly used in the structural break estimation literature to simplify the limiting distribution. For simplicity, we proceed with Assumption (ref)(a) in order to demonstrate the usefulness of Theorem (ref) in this case. Theorem (ref) may also be useful when the break size is fixed, as in Hinkley1970, or when it is very small, as in ElliottMueller2007.

(2) Assumption (ref)(b) assumes the estimator converges at a particular rate. The rate depends on the break size, given by $v_T$. We note that the conditions in Bai1997 are sufficient for Assumption (ref)(b). Since our focus is on the application of the argmax theorem, we postpone the statement of this result to Lemma (ref) in the appendix.

(3) Assumptions (ref)(c)-(f) are high-level conditions used to derive the limit of the localized objective function. Assumption (ref)(c) is a no perfect multicollinearity condition. It makes sure $\hat\delta_k$ is well defined for all $k\in\{p,...,T-p\}$. Assumption (ref)(d) ensures the partial sums of $x_tx'_t$ before and after the break converge to positive definite matrices. Assumption (ref)(e) ensures standardized sums of $x_t\epsilon_t$ have the usual rate of convergence. Assumption (ref)(f) ensures the partial sums of $x_t\epsilon_t$ before and after the break converge weakly to multivariate Gaussian processes. Note that Assumptions (ref)(d) and (f) also allow a change in the distribution of $(x_t,\epsilon_t)$ at the break date. It is easy to state low-level conditions for Assumptions (ref)(c)-(f) by limiting the dependence in the distribution of $(x_t,\epsilon_t)$ and requiring certain finite moments; for example see Assumptions A2-A6 in Bai1997. We abstract from stating low-level conditions in order to focus on the role of the optimization set in this application of Theorem (ref). \qed

Next, we localize the objective function around $k_0$ at the $v_T^{-2}$ rate. For every $s\in \mathbb{R}$, let

equation[equation omitted — 62 chars of source]

Note that $\mathbb{M}_T(s)$ is eventually defined over every bounded set of $s$ values. By definition, $v_T^2(\hat k-k_0)$ maximizes $\mathbb{M}_T(s)$ over $s\in v_T^2(\Lambda_T-k_0)$, a discrete set of points. The following lemma shows that this set PK converges to $\mathbb{R}$ under Assumption (ref)(a) and PK converges to $(-\infty,a]$ under Assumption (ref)(b).

lemma\begin{enumerate}[label=(\alph*)] • Under Assumption (ref)(a), $v_T^2(\Lambda_T-k_0)\rightarrow \mathbb{R}$. • Under Assumption (ref)(b), $v_T^2(\Lambda_T-k_0)\rightarrow (-\infty,a]$. \end{enumerate}

Remark. The proof of Lemma (ref) uses the fact that every sequence of sets has a subsequence that PK converges. This is a very useful property of PK convergence because it means that, in order to verify PK convergence, all we need to do is show that the limiting set does not depend on which subsequence was taken. \qed

The limit of $\mathbb{M}_T(s)$ is defined as

equation[equation omitted — 178 chars of source]

The following lemma gives the convergence result that we need to verify the conditions of Theorem (ref). It is closely related to Lemma A.5 in Bai1997.

lemmaUnder Assumption (ref)(a) or (b) and Assumption (ref), for any $C>0$, \[ \mathbb{M}_T(s)\rightsquigarrow \mathbb{M}(s) \] in $\ell^\infty([-C,C])$.

We can now state the following result, based on Theorem (ref).

corollary\begin{enumerate}[label=(\alph*)] • Under Assumptions (ref)(a) and (ref), \[ v_T^2(\hat k - k_0)=\text{argmax}_{s\in v_T^2(\Lambda_T-k_0)}\mathbb{M}_T(s)\rightsquigarrow \text{argmax}_{s\in\mathbb{R}} \mathbb{M}(s). \] • Under Assumptions (ref)(b) and (ref), \[ v_T^2(\hat k-k_0)=\text{argmax}_{s\in v_T^2(\Lambda_T-k_0)}\mathbb{M}_T(s)\rightsquigarrow \text{argmax}_{s\in(-\infty,a]} \mathbb{M}(s). \] \end{enumerate}

Remarks. (1) Corollary (ref) is stated as a corollary to Theorem (ref) because the proof consists of verifying the conditions of Theorem (ref). In fact, all the conditions of Theorem (ref) are easily satisfied by Lemmas (ref) and (ref) in this section and Lemma (ref) in the appendix. Lemma (ref) in the appendix verifies that the argmax of the limit exists and is unique almost surely.

(2) Part (a) gives the limiting distribution for $\hat k$ when the true break date is strictly inside the trimming interval. The limiting distribution coincides with the limit without any trimming, as stated in Proposition 3 in Bai1997. Part (b) gives the limiting distribution for $\hat k$ when the true break date is near the right endpoint of the trimming interval. Instead of maximizing over all of $s\in\mathbb{R}$, the maximization is taken only over $s\in(-\infty,a]$. Thus, the presence of the trimming has an effect on the limit, by preventing the maximum from being achieved at any value larger than $a$.

(3) The limiting distribution from part (b) is new to the literature. It can be used to analyze the consequences of local misspecification of the trimming interval. When $a\ge 0$, the trimming parameter is helpful because it prevents $\hat k$ from being too much larger than $k_0$. However, when $a<0$, the trimming parameter bounds the limiting distribution of $\hat k$ away from zero, introducing a significant distortion. This distortion clarifies the reason why $\lambda_2$ should be chosen large enough that $k_0\le [\lambda_2T]$.

(4) The limit in part (b) with $a>0$ is similar to the limits obtained by JiangWangYu2018 and CasiniPerron2021 using in-fill asymptotics. These limits restrict the maximum by both an upper and a lower bound, while the limit in part (b) has only an upper bound. Both JiangWangYu2018 and CasiniPerron2021 argue that their limit theory provides a better approximation to the finite sample distribution of $\hat k$, especially replicating the multi-modality of the distribution found in simulations. For the same reason, the limit in part (b) may improve the approximation to the finite sample distribution over the limit in part (a). (The limit in part (a) is unimodal while the limit in part (b) with $a>0$ is bimodal.)

(5) Analogous results could be stated that allow the break date to be near the beginning of the sample, or that allow only some of the regressors to have coefficients that change. In addition, related results could be explored in other contexts: models with two or more breaks with a minimum distance between them, as in QuPerron2007; change-point estimation in nonparametric regression models, as in Muller1992; or threshold regression models, as in Hansen2000 or HidalgoLeeSeo2019. \qed

Parameter on the Boundary

Suppose we are estimating a parameter, $\theta$, that belongs to a parameter space, $\Theta$. Often the definition of $\Theta$ includes inequalities that provide a priori information on the possible values of $\theta$. That is, $\Theta=\{\theta\in\mathbb{R}^{d_\theta}: g(\theta)\le 0\}$ for a vector of inequalities given by $g(\theta)$. These inequalities can come from nonnegativity of variance parameters (Andrews1999, Andrews2002 and Ketz2018), known signs of some coefficients in regression models (AndrewsGuggenberger2010, AndrewsArmstrong2017, and KetzMcCloskey2021), or parameter restrictions in GARCH models (Andrews1997, Andrews2001 and FrazierRenault2020). These inequalities can also come from moment inequalities treated as overidentifying restrictions, as in MoonSchorfheide2009 or CCT2018. When $\theta$ satisfies one or more of the inequalities with equality, then this prior information is useful for estimating $\theta$ and influences the limiting distribution of the estimator. We consider the case that the true value of $\theta$ is on or near the boundary of $\Theta$ and use Theorem (ref) to derive the limiting distribution of the estimator.

The parameter-on-the-boundary problem has a long history in the statistics literature; see SilvapulleSen2005 for an overview. Most of this literature considers the case that $\Theta$ is Chernoff regular. That is, $\Theta$ can be approximated by a tangent cone, $\Lambda$, near a fixed point, $\theta_0\in\Theta$; see Chernoff1954 or Geyer1994 for a formal definition. The literature then derives the limiting distribution of the estimator to be the projection of a normal random vector onto $\Lambda$. This gives rise to a mixture of normal distributions for the estimator and a chi-bar-squared distribution for the likelihood ratio statistic.

Beginning with Feder1968 and Moran1971, the true value of $\theta$ need not be exactly on the boundary of $\Theta$, but it may be close to the boundary. Let $\theta_n\in\Theta$ denote the true value of $\theta$, indexed by $n$ because it takes a sequence of values that converges to $\theta_0$ as the sample size increases. If $\theta_n$ is close to the boundary in the sense that $\sqrt{n}g(\theta_n)\rightarrow b$, and if $g(\theta)$ is continuously differentiable in a neighborhood of $\theta_0$ with derivative $G(\theta_0)$, then the limiting set is given by $\Lambda=\{\lambda\in\mathbb{R}^{d_\theta}: b+G(\theta_0)\lambda\le 0\}$. Note that $\Lambda$ is a polyhedral set defined by a collection of affine inequalities. This result is formally stated in the following lemma. Let $e_j$ denote the $j$th unit basis vector in $\mathbb{R}^{d_g}$.

lemmaSuppose $g(\theta)$ is continuously differentiable in a neighborhood of $\theta_0$ with matrix of derivatives denoted by $G(\theta)$. Also suppose $\sqrt{n}g(\theta_n)\rightarrow b$. If there exists $\tilde \lambda\in\Lambda=\{\lambda\in\mathbb{R}^{d_\theta}: b+G(\theta_0)\lambda\le 0\}$ such that $e'_jb+e'_jG(\theta_0)\tilde \lambda<0$ for all $j\in\{1,...,d_g\}$, then \[ \sqrt{n}(\Theta-\theta_n)\rightarrow \Lambda. \]

Remarks. (1) Lemma (ref) states a fairly standard result that converts nonlinear inequalities to linear ones locally around a point. It is similar to Proposition 4.7.3 in SilvapulleSen2005, except with PK convergence to a polyhedral set instead of local approximation by a cone. PK convergence is more general than approximation by a cone simply because it allows the limit to be any closed set.

(2) The assumption on $\tilde\lambda\in\Lambda$ is a type of Mangasarian-Fromowitz constraint qualification. See KaidoMolinariStoye2022 for a discussion of constraint qualifications in partially identified models.

(3) If some of the components of $b$ are $-\infty$, then the convergence $\sqrt{n}g(\theta_n)\rightarrow b$ should be interpreted elementwise. The corresponding inequalities drop out from the definition of $\Lambda$ so they have no influence on the limit. \qed

Let $\hat\theta_n$ be an estimator of $\theta$ that maximizes a random objective function, $V_n(\theta)$ over $\theta\in\Theta$. We next state a corollary of Theorem (ref) that gives the limiting distribution of $\hat\theta_n$ when $\theta_n$ is local to the boundary. We use the following high-level assumptions.

assumption\begin{enumerate}[label=(\alph*)] • There exists a set, $\Lambda$, such that $\sqrt{n}(\Theta-\theta_n)\rightarrow \Lambda$. • There exists an almost surely continuous random objective function, $\mathbb{M}(h)$, such that for every compact $K\subset\mathbb{R}^{d_\theta}$, $V_n(\theta_n+n^{-1/2}h)\rightsquigarrow \mathbb{M}(h)$ in $\ell^\infty(K)$. • $\sqrt{n}(\hat\theta_n-\theta_n)=O_P(1).$ • The argmax of $\mathbb{M}(h)$ over $\Lambda$ is nonempty and unique almost surely. \end{enumerate}

Remarks. (1) When $\Theta$ is Chernoff regular and $\theta_n=\theta_0$, then Assumption (ref)(a) is satisfied with a cone $\Lambda$. When $\theta_n\rightarrow\theta_0$, then Assumption (ref)(a) can be verified using Lemma (ref), above. Even parameter spaces that are not Chernoff regular will satisfy Assumption (ref)(a) for some set $\Lambda$, along a subsequence.

(2) Assumption (ref)(b) requires the objective function, localized at the $n^{-1/2}$ rate, to converge to a limiting objective function. In standard models, the limit is quadratic in $h$. Assumption (ref)(c) requires the estimator to converge at the $\sqrt{n}$ rate. There are many approaches to verifying Assumptions (ref)(b) and (c) in various models. Fairly general sufficient conditions are available in VaartWellner1996. These arguments are typically unaffected by $\theta_n$ being close to the boundary of $\Theta$. Here we abstract from stating low-level conditions in order to focus on the role of the rescaled parameter space in the argmax theorem.

(3) Assumption (ref)(d) is satisfied easily when $\mathbb{M}(h)$ is strictly concave and $\Lambda$ is convex. This holds in standard cases when $\mathbb{M}(h)$ is quadratic with negative definite form and $\Lambda$ is a polyhedral set. In non-concave/convex cases, Assumption (ref)(d) can be verified using Cox2020. \qed

corollaryUnder Assumption (ref), \[ \sqrt{n}(\hat\theta_n-\theta_n)=\text{argmax}_{h\in\sqrt{n}(\Theta-\theta_n)}V_n(\theta_n+n^{-1/2}h)\rightsquigarrow \text{argmax}_{h\in\Lambda}\mathbb{M}(h). \]

Remarks. (1) Corollary (ref) follows directly from Theorem (ref). The focus is on the convergence of the rescaled parameter space. In the parameter-on-the-boundary setting, PK convergence of the rescaled parameter space follows naturally.

(2) When $\Lambda$ is a polyhedral cone and $\mathbb{M}(h)$ is quadratic in $h$ with normally distributed coefficients on the linear term, then Corollary (ref) coincides with typical results available in the parameter-on-the-boundary literature: the limiting distribution is the projection of a normal random vector onto a cone. Still, the proof is arguably simpler than what is available in the literature. For example, Lemma 2 in Andrews1999 and Proposition 4.7.4 in SilvapulleSen2005 are challenging results to prove that are needed to show that one can replace $\sqrt{n}(\Theta-\theta_n)$ by its approximating cone. More generally, if $\mathbb{M}(h)$ is not quadratic in $h$ or $\Theta$ is not Chernoff regular, then Corollary (ref) is new to the literature. \qed

Weak Identification

A vector of parameters, $\beta$, is weakly identified if the objective function used to estimate it is flat, or nearly flat, in a neighborhood of the maximum. Let $V_n(\beta,\pi)$ denote a random objective function that depends on a random sample of size $n$ and an additional vector of identified parameters, $\pi$. Suppose for some value of $\pi$, say $\pi_0$, $V_n(\beta,\pi_0)$ does not depend on the value of $\beta$. This designates $\pi_0$ as a problematic point in the parameter space where $\beta$ is not identified. Weak identification of $\beta$ arises when the true value of $\pi$, say $\pi_n$, is allowed to be a sequence that depends on the sample size and converges to $\pi_0$; see StockWright2000, AndrewsCheng2012, or Cox_weak_id_w_bounds. Sometimes, a careful reparameterization is needed to fit a given model into this setup; see HanMcCloskey2019.

The identified set for $\beta$ is the set of possible values that are observationally equivalent, or that cannot be distinguished from each other, based on the distribution of the sample. When $\pi=\pi_0$, no two values of $\beta$ can be distinguished by $V_n(\beta,\pi_0)$, and thus the identified set for $\beta$ is delineated by the boundary of the parameter space. Let $\Theta$ denote the parameter space for $\theta=(\beta,\pi)$. As in Section 3.2, $\Theta$ is defined by a collection of inequalities that give a priori restrictions on the possible values of the parameters. Suppose $\Theta=\{\theta\in\mathbb{R}^{d_\theta}: g(\theta)\le 0\}$, for some function $g(\theta)$.

In weak identification, these inequalities are especially important because they can provide information about the value of $\beta$. As a simple example, suppose $\beta$ and $\pi$ are both scalars. If the inequalities defined by $g(\theta)$ are $\beta\ge 0$ and $\beta\le\pi$, then the identified set for $\beta$ is given by the interval $[0,\pi]$. This simple example just demonstrates that the identified set for $\beta$ can be more or less informative, measured by the length of the interval, depending on the value of $\pi$. Most papers on weak identification exclude the possibility of informative inequalities by assuming $\Theta$ is a product space between the $\beta$ parameters and the $\pi$ parameters; see page 1060 in StockWright2000 for an example. Theorem (ref) can be used to derive the limiting distribution of an estimator of a weakly identified parameter without the product space assumption, and thus covers weak identification with informative inequalities.

Denote the estimators by $(\hat\beta_n,\hat\pi_n)$, which maximize $V_n(\beta,\pi)$ over $(\beta,\pi)\in\Theta$. The weak identification literature has derived two types of limiting distributions distinguished by the rate of convergence of $\hat\beta_n$ to the true value of $\beta$. Let $\beta_n$ denote the true value of $\beta$, which is allowed to be a sequence that depends on the sample size and converges to a limit, $\beta_0$. “Weak identification” arises when $\hat\beta_n$ is inconsistent, while “semi-strong identification” arises when $\hat\beta_n$ is consistent for the true value of $\beta$ at a rate given by $a_n(\hat\beta_n-\beta_n)=O_P(1)$ for a sequence of positive constants, $a_n$, satisfying $a_n\rightarrow\infty$ and $n^{-1/2}a_n\rightarrow 0$.

Under weak identification, the relevant scaling of the local parameter space is given by $\Lambda^{\hspace{-0.48mm}\text{W}}_n=\{(\beta,\sqrt{n}(\pi-\pi_n)): (\beta,\pi)\in\Theta\}$. Under semi-strong identification, the relevant scaling of the local parameter space is given by $\Lambda^{\text{SS}}_n=\{(a_n(\beta-\beta_n),\sqrt{n}(\pi-\pi_n)): (\beta,\pi)\in\Theta\}$. Note that for $\Lambda^{\hspace{-0.48mm}\text{W}}_n$, $\beta$ is not scaled, while for $\Lambda^{\text{SS}}_n$, $\beta$ is scaled at the $a_n$ rate, corresponding to the rate of convergence of $\hat\beta_n$. The following lemma gives sufficient conditions for these local parameter spaces to PK converge. Let $cl(A)$ and $int(A)$ denote the closure and interior of a set, $A$, respectively.

lemma\begin{enumerate}[label=(\alph*)] • Let $\mathcal{B}^{\text{W}}=\{\beta\in\mathbb{R}^{d_\beta}: (\beta,\pi_0)\in\Theta\}$. Suppose $\Theta$ is closed and that $\mathcal{B}^{\text{W}}=cl(\{\beta\in\mathbb{R}^{d_\beta}: (\beta,\pi_0)\in int(\Theta)\})$. Then $\Lambda^{\hspace{-0.48mm}\text{W}}_n\rightarrow\mathcal{B}^{\text{W}}\times \mathbb{R}^{d_\pi}$. • Suppose $g(\beta,\pi)$ is continuously differentiable in a neighborhood of $(\beta_0,\pi_0)$ with derivative $G(\beta,\pi)=[G_\beta(\beta,\pi),G_\pi(\beta,\pi)]$. Also suppose $a_ng(\beta_n,\pi_n)\rightarrow b\in[-\infty,0]^{d_g}$. Let $\mathcal{B}^{\text{SS}}=\{\lambda\in\mathbb{R}^{d_\beta}: b+G_\beta(\beta_0,\pi_0)\lambda\le 0\}$ and suppose there exists a $\tilde\lambda\in\mathcal{B}^{\text{SS}}$ such that $e'_jb+e'_jG_\beta(\beta_0,\pi_0)\tilde\lambda<0$ for every $j\in\{1,...,d_g\}$. Then $\Lambda^{\text{SS}}_n\rightarrow \mathcal{B}^{\text{SS}}\times \mathbb{R}^{d_\pi}$. \end{enumerate}

Remarks. (1) Part (a) gives the limit of the rescaled parameter space under weak identification. Part (b) gives the limit under semi-strong identification. In both cases, the limits are given by Cartesian products between sets in the $\beta$ direction and $\mathbb{R}^{d_\pi}$ in the $\pi$ direction. This arises because the $\pi$ directions are scaled at a faster rate, effectively turning the inequalities so that they only enforce in the $\beta$ directions.

(2) The condition in part (a) says that $\mathcal{B}^{\text{W}}$ is the closure of the interior of $\Theta$ intersected with the set satisfying $\pi=\pi_0$. It mostly rules out isolated points in $\mathcal{B}^{\text{W}}$ and atypical shapes for $\Theta$. The condition in part (b) is a type of Mangasarian-Fromowitz constraint qualification.

(3) The limit in part (b) is very similar to the limit of the local parameter space when the parameter is near the boundary from Lemma (ref) in Section 3.2. Both limits are polyhedral sets, defined by a collection of affine inequalities. The difference is that the limit in part (b) only uses the derivative with respect to $\beta$. This difference comes from the different rates of convergence of $\hat\beta_n$ and $\hat\pi_n$. This implies a distinct change in the limit theory between strong and semi-strong identification when the parameter is near the boundary. \qed

The limiting distribution of the estimator depends on the properties of $V_n(\beta,\pi)$ in a neighborhood of the identified set. We use the following high-level assumptions to derive the limiting distribution using Theorem (ref).

namedassumption[4(a)] Under weak identification, the following hold. \begin{enumerate}[label=(\roman*)] • There exists a set, $\Lambda^{\hspace{-0.48mm}\text{W}}$, such that $\Lambda_n^{\hspace{-0.48mm}\text{W}}\rightarrow \Lambda^{\hspace{-0.48mm}\text{W}}$. • $\sqrt{n}(\hat\pi_n-\pi_n)=O_P(1)$ and $\hat\beta_n=O_P(1)$. • There exists an almost surely continuous random objective function, $\mathbb{M}^{\text{W}}(\beta,h_\pi)$ such that for all compact sets, $K$, $V_n(\beta,\pi_n+n^{-1/2}h_\pi)\rightsquigarrow \mathbb{M}^{\text{W}}(\beta,h_\pi)$ in $\ell^\infty(K)$ for $(\beta,h_\pi)\in K$. • The argmax of $\mathbb{M}^{\text{W}}(\beta,h_\pi)$ over $(\beta,h_\pi)\in\Lambda^{\hspace{-0.48mm}\text{W}}$ is nonempty and unique almost surely. \end{enumerate}
namedassumption[4(b)] Under semi-strong identification, the following hold. \begin{enumerate}[label=(\roman*)] • There exists a set, $\Lambda^{\text{SS}}$, such that $\Lambda_n^{\text{SS}}\rightarrow \Lambda^{\text{SS}}$. • $\sqrt{n}(\hat\pi_n-\pi_n)=O_P(1)$ and $a_n(\hat\beta_n-\beta_n)=O_P(1)$. • There exists an almost surely continuous random objective function, $\mathbb{M}^{\text{SS}}(h_\beta,h_\pi)$ such that for all compact sets, $K$, $V_n(\beta_n+a_n^{-1}h_\beta,\pi_n+n^{-1/2}h_\pi)\rightsquigarrow \mathbb{M}^{\text{SS}}(h_\beta,h_\pi)$ in $\ell^\infty(K)$ for $(h_\beta,h_\pi)\in K$. • The argmax of $\mathbb{M}^{\text{SS}}(h_\beta,h_\pi)$ over $(h_\beta,h_\pi)\in\Lambda^{\text{SS}}$ is nonempty and unique almost surely. \end{enumerate}

Remarks. (1) There are two different versions of Assumption 4, one for weak identification and one for semi-strong identification. The only difference between the two is the rescaling of the parameter space, corresponding to the rate of convergence of $\hat\beta_n$.

(2) Part (i) of both versions assumes the rescaled parameter space PK converges. This holds using Lemma (ref) if the conditions are satisfied. Otherwise, it holds along a subsequence.

(3) Parts (ii) and (iii) of both versions are high-level conditions for the rate of convergence and the limit of the standardized objective functions. Existing results in the weak identification literature give sufficient conditions for various objective functions. StockWright2000 and AndrewsCheng2014 consider a generalized method of moments objective function, AndrewsCheng2013 considers a maximum likelihood objective function, and Cox_weak_id_w_bounds considers a minimum distance objective function. AndrewsCheng2012 give intermediate-level sufficient conditions for a general objective function. We abstract from stating the sufficient conditions here in order to focus on the role of PK convergence of the rescaled parameter space.

For all cases considered in the above literature, $\mathbb{M}^{\text{SS}}(h_\beta,h_\pi)$ is quadratic in $(h_\beta,h_\pi)$ with normally distributed coefficients on the linear term. Also, $\mathbb{M}^{\text{W}}(\beta,h_\pi)$ is quadratic in $h_\pi$ but may be non-quadratic in $\beta$. For brevity, we abstain from stating explicit formulas for $\mathbb{M}^{\text{SS}}(h_\beta,h_\pi)$ and $\mathbb{M}^{\text{W}}(\beta,h_\pi)$.

(4) Part (iv) for both versions is a uniqueness condition on the limit of the localized objective function maximized over the limit of the rescaled parameter space. When $\mathbb{M}^{\text{SS}}(h_\beta,h_\pi)$ is a negative definite quadratic form in $(h_\beta,h_\pi)$, this follows trivially from convexity of $\Lambda^{\text{SS}}$. Cox2020 gives sufficient conditions for verifying uniqueness of the argmax of $\mathbb{M}^{\text{W}}(\beta,h_\pi)$. \qed

The above assumptions allow us to derive the limiting distribution of the estimator using Theorem (ref).

corollary(a) Under Assumption 4(a), \begin{align*} \left(\hat\beta_n,\sqrt{n}(\hat\pi_n-\pi_n)\right)&=argmax_{(\beta,h_\pi)\in\Lambda_n^{W}}V_n(\beta,\pi_n+n^{-1/2}h_\pi)\\ &\rightsquigarrow argmax_{(\beta,h_\pi)\in\Lambda^{W}}\mathbb{M}^{W}(\beta,h_\pi). \end{align*} \textup{(b)} Under Assumption 4(b), \begin{align*} \left(a_n(\hat\beta_n-\beta_n),\sqrt{n}(\hat\pi_n-\pi_n)\right)&=\text{argmax}_{(h_\beta,h_\pi)\in\Lambda_n^{\text{SS}}}V_n(\beta_n+a_n^{-1}h_\beta,\pi_n+n^{-1/2}h_\pi)\\ &\rightsquigarrow \text{argmax}_{(h_\beta,h_\pi)\in\Lambda^{\text{SS}}}\mathbb{M}^{\text{SS}}(h_\beta,h_\pi). \end{align*}

Remark. Corollary (ref) states an important result for the weak identification literature. By incorporating inequalities into the weak identification limiting distribution, hypothesis tests that are identification-robust can rely on information available in inequalities when identification is weak. Cox_weak_id_w_bounds proposes such a test based on Theorem (ref) and Lemma (ref), below. Other identification-robust tests available in the literature are either invalid in the presence of inequalities or do not have power against hypotheses that violate the inequalities; see the simulations in Cox_weak_id_w_bounds. \qed

Proof of Theorem (ref)

The proof of Theorem (ref) does not follow directly from the argument used to prove Theorem 3.2.2 in VaartWellner1996. A problem comes from the fact that, for an arbitrary compact set, $K$, $\Lambda_n\cap K$ need not PK converge to $\Lambda\cap K$. Otherwise, one could show using the extended continuous mapping theorem that $\sup_{h\in \Lambda_n\cap K}\mathbb{M}_n(h)\rightsquigarrow\sup_{h\in\Lambda\cap K}\mathbb{M}(h)$. The proof of Theorem (ref) would then follow from the argument used to prove Theorem 3.2.2 in VaartWellner1996. The following lemma gives a solution based on bounding $K$ inside and outside by compact sets such that PK convergence holds along a subsequence.

lemmaLet $H$ be a separable metric space equipped with metric $d$. Suppose $\Lambda_n\rightarrow\Lambda$ for a sequence of sets $\Lambda_n$ and $\Lambda$. If $K$ is a compact subset of $H$, then there exist compact sets, $K_1\subset K\subset K_2\subset K_3$, and there exists a subsequence, $n_q$, such that $\Lambda_{n_q}\cap K\rightarrow \Lambda\cap K_1$ and $\Lambda_{n_q}\cap K_3\rightarrow \Lambda\cap K_2$.

Remark. The proof of Lemma (ref) is simple in finite dimensional spaces. In an arbitrary separable metric space, it is less simple. In particular, the set $K_3$ must be defined carefully. The proof is given in the appendix. \qed

The proof of Theorem (ref) also relies on convergence of the value function. The following lemma gives the statement that we need.

lemmaLet $\mathbb{M}_n,\mathbb{M}$ be stochastic processes indexed by a separable metric space $H$ such that $\mathbb{M}_n\rightsquigarrow \mathbb{M}$ in $\ell^\infty(K)$ for every compact $K\subset H$. Let $\Lambda_n,\Lambda\subset H$ be such that $\Lambda_n\rightarrow\Lambda$. Suppose that almost all sample paths $h\mapsto\mathbb{M}(h)$ are continuous and possess a maximum over $h\in \Lambda$ at a random point $\hat h$, which as a random map in $H$ is tight. If the sequence $\hat h_n\in \Lambda_n$ is uniformly tight and satisfies $\mathbb{M}_n(\hat h_n)\ge \sup_{h\in \Lambda_n}\mathbb{M}_n(h)-o_P(1)$, then $\sup_{h\in\Lambda_n}\mathbb{M}_n(h)\rightsquigarrow \sup_{h\in\Lambda}\mathbb{M}(h)$. Furthermore, if $\mathcal{D}$ is a metric space containing random elements $\mathbb{D}_n$ and $\mathbb{D}$, where $\mathbb{D}$ is Borel measurable with separable range, and if for every compact $K\subset H$, $\mathbb{M}_n\rightsquigarrow\mathbb{M}$ in $\ell^{\infty}(K)$ holds jointly with $\mathbb{D}_n\rightsquigarrow\mathbb{D}$, then $\sup_{h\in\Lambda_n}\mathbb{M}_n(h)\rightsquigarrow \sup_{h\in\Lambda}\mathbb{M}(h)$ holds jointly with $\mathbb{D}_n\rightsquigarrow\mathbb{D}$.

Remarks. (1) The statement of Lemma (ref) is the same as Theorem (ref), except that the conclusion is convergence of the value function and uniqueness of the maximizer in the limit is not needed. For completeness, the furthermore tells us that this convergence holds jointly with other weak convergence conditions.

(2) Convergence of the value functions has independent interest. Cox_weak_id_w_bounds uses it to derive the limit theory for likelihood ratio statistics when testing hypotheses on weakly identified parameters.

(3) The proof of Lemma (ref) is also affected by the fact that, for an arbitrary compact set, $K$, $\Lambda_n\cap K$ need not PK converge to $\Lambda\cap K$. The solution is again to use Lemma (ref), combined with a subsequencing argument. The proof is given in the appendix. \qed

We next give the proof of Theorem (ref). Note how Lemmas (ref) and (ref) are used, together with a subsequencing argument. (The $K_2$ and $K_3$ from Lemma (ref) are not used in the proof of Theorem (ref), but they are used in the proof of Lemma (ref).) Let $P^\ast$ denote the outer probability measure.

namedproof[of Theorem (ref)] We first note that for any compact set, $K$, for any closed set $F$, and for any subsequence, $n_m$, Lemma (ref) implies that there exists a further subsequence, $n_q$, and a compact set $K_1\subset K\cap F$ such that $\Lambda_{n_q}\cap K\cap F\rightarrow \Lambda\cap K_1$. We can apply the extended continuous mapping theorem to show \begin{equation} \sup_{h\in\Lambda_{n_q}\cap K\cap F}\mathbb{M}_{n_q}(h)\rightsquigarrow \sup_{h\in \Lambda\cap K_1}\mathbb{M}(h). \end{equation} The condition of the extended continuous mapping theorem is satisfied by Lemma (ref), in the appendix. We also note that by Lemma (ref), $\sup_{h\in\Lambda_n}\mathbb{M}_n(h)\rightsquigarrow \sup_{h\in\Lambda}\mathbb{M}(h)$, and this holds jointly with ((ref)) along the subsequence $n_q$. We also note that, almost surely, \begin{equation} \mathbb{M}(\hat h)>\sup_{h\in \Lambda\cap K\cap F} \mathbb{M}(h), \end{equation} for every compact set $K$ and for every closed set $F$ such that $\hat h\notin K\cap F$. If this were not true, then there would exist a sequence $\{h_m\}_{m=1}^{\infty}\subset \Lambda\cap K\cap F$ with $\mathbb{M}(h_m)\rightarrow \mathbb{M}(\hat h)$. Since $\Lambda\cap K\cap F$ is compact, the sequence may be chosen to be convergent; by continuity the value $\mathbb{M}(h)$ at the limit would be $\mathbb{M}(\hat h)$. This contradicts the fact that $\hat h$ is unique, because $h$ is contained in the closed set $\Lambda\cap K\cap F$ and hence cannot equal $\hat h$. Thus, for an arbitrary closed set $F$, \begin{align} &\limsup_{n\rightarrow\infty}P^*\left(\hat h_n\in F\right)\\ \le&\limsup_{n\rightarrow\infty}P^*\left(\hat h_n\in K\cap F\right)+\limsup_{n\rightarrow\infty}P^*\left(\hat h_n\notin K\right)\nonumber\\ \le&\limsup_{n\rightarrow\infty}P^*\left(\sup_{h\in \Lambda_n\cap K\cap F}\mathbb{M}_n(h)\ge \sup_{h\in \Lambda_n}\mathbb{M}_n(h)-o_P(1)\right)+\limsup_{n\rightarrow\infty}P^*\left(\hat h_n\notin K\right)\nonumber\\ \le &P\left(\sup_{h\in \Lambda\cap K\cap F}\mathbb{M}(h)\ge \sup_{h\in \Lambda}\mathbb{M}(h)\right)+\limsup_{n\rightarrow\infty}P^*\left(\hat h_n\notin K\right)\nonumber\\ \le &P\left(\{\sup_{h\in \Lambda\cap K\cap F}\mathbb{M}(h)\ge \sup_{h\in \Lambda\cap K}\mathbb{M}(h)\}\cap \{\hat h\in K\}\right)+P(\hat h\notin K)+\limsup_{n\rightarrow\infty}P^*\left(\hat h_n\notin K\right)\nonumber\\ \le &P\left(\hat h\in F\right)+P\left(\hat h\notin K\right)+\limsup_{n\rightarrow\infty}P^*\left(\hat h_n\notin K\right), \nonumber \end{align} where the second inequality follows from the assumed condition on $\hat h_n$, the third inequality follows from an argument in the next paragraph, and the fifth inequality follows because the event $\{\sup_{h\in \Lambda\cap K\cap F}\mathbb{M}(h)\ge \sup_{h\in \Lambda\cap K}\mathbb{M}(h)\}\cap \{\hat h\in K\}$ implies $\{\hat h\in F\}$ by ((ref)). The last two terms on the right can be made arbitrarily small by the choice of $K$. Apply the Portmanteau Theorem to conclude that $\hat h_n\rightsquigarrow \hat h$. The third inequality in the above expression follows from a subsequencing argument. Let $n_m$ be an arbitrary subsequence. Then, by Lemma (ref), there exists a further subsequence, $n_q$, and a compact set $K_1$, such that ((ref)) holds. By the Portmanteau Theorem and Slutsky's Lemma, \begin{align} &\limsup_{q\rightarrow\infty}P^*\left(\sup_{h\in \Lambda_{n_q}\cap K\cap F}\mathbb{M}_{n_q}(h)\ge \sup_{h\in \Lambda_{n_q}}\mathbb{M}_{n_q}(h)-o_P(1)\right)\\ &\le P\left(\sup_{h\in \Lambda\cap K_1}\mathbb{M}(h)\ge \sup_{h\in \Lambda}\mathbb{M}(h)\right)\nonumber\\ &\le P\left(\sup_{h\in \Lambda\cap K\cap F}\mathbb{M}(h)\ge \sup_{h\in \Lambda}\mathbb{M}(h)\right),\nonumber \end{align} where the second inequality follows from $\Lambda\cap K_1\subset \Lambda\cap K\cap F$. We have shown that for every subsequence, $n_m$, there exists a further subsequence, $n_q$, such that this inequality holds. The inequality must hold along the original sequence as well. \qed

Conclusion

This paper states a generalization of the argmax theorem that allows the maximization to take place over a sequence of subsets of the domain. If the sequence of subsets PK converges to a limiting subset, then the conclusion of the argmax theorem continues to hold. This paper demonstrates the usefulness of this generalization in three applications: structural break estimation, estimating a parameter on the boundary, and estimating a weakly identified parameter. The generalized argmax theorem simplifies the proofs for existing results and can be used to prove new results in these literatures.