Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
405,033 characters · 52 sections · 83 citation commands
Bayesian Penalized Empirical Likelihood and Markov Chain Monte Carlo Sampling
\if11 {
\affil[a]{\it Joint Laboratory of Data Science and Business Intelligence, Southwestern University of Finance and Economics, Chengdu, China} \affil[b]{\it Academy of Mathematics and Systems Science, Chinese Academy of Sciences, Beijing, China} \affil[c]{\it Department of Statistics, Operations, and Data Science, Temple University, Philadelphia, PA, USA }
} \fi
\if01 {
} \fi
{\it Key words:} Bayesian methods, Bernstein-von Mises theorem, Estimating equations, MCMC, Penalized empirical likelihood.
\baselineskip 24pt
EL Owen_2001 is a versatile and flexible tool for statistical inference, providing a framework that accommodates broadly defined model conditions. Unlike traditional likelihood approaches, EL does not require the explicit specification of probability distributions governing the data generation process. This inherent flexibility offers numerous practical advantages, such as the ability to incorporate a wide range of model specifications and prior knowledge, making it highly adaptable for integrating information from multiple data sources. Additionally, EL retains key benefits of its parametric likelihood counterpart, including efficiency (in the semiparametric sense) and the convenience of conducting hypothesis tests and estimating confidence sets through the Wilks-type likelihood ratio framework.
Recent developments in EL approaches have a focus on addressing the challenges posed by complex high-dimensional data. To handle the complexities arising from various model conditions, researchers have explored regularization techniques applied to the Lagrange multipliers associated with EL or the empirical versions of moment conditions, aiming to achieve enhanced model parsimony. In Shi2016, a two-step procedure is introduced. The first step involves employing a “relaxed" EL that incorporates specific inequality constraints in its formulation. The second step includes moment selection and bias correction. Chau2017 addresses a continuum of moment conditions where the numerical optimization problem becomes ill-conditioned. To resolve this, a penalty on the continuous version of the Lagrange multiplier's counterpart is proposed and investigated. Changetal_2018 proposes a method to penalize the magnitudes of both the Lagrange multiplier and the model parameters, specifically to tackle high-dimensional model parameters under complex conditions. More recently, Chang_2021 explores the projection of high-dimensional moment conditions onto lower-dimensional spaces to facilitate statistical inference for specific components of model parameters and to assess model specification validity. Besides addressing the challenge of handling many moment conditions, the development of EL approaches that incorporate penalties on model parameters to promote parsimonious structures can effectively manage high-dimensional problems, as discussed in TangLeng_2010_Bioka, LengTang_2010, ChangChenChen_2015, and Chang2023.
The synergy of Bayesian methodologies with traditional likelihoods has consistently demonstrated its effectiveness. Leveraging advances in sampling techniques, Bayesian approaches have established their significance in tackling a wide array of challenges across various domains. This is particularly valuable when dealing with intricate statistical problems where maximizing or even computing the objective function becomes infeasible. The amalgamation of Bayesian principles with EL shows great promise in practical applications. This integration enhances the adaptability and robustness of the Bayesian framework, enabling the creation of statistical models that can accommodate a wide range of scenarios. Recent developments in the realm of Bayesian EL (BEL) methods are evident in a growing body of literature; see Lazar2003, Rao2010, Chaudhuri2011, Yang2012, Mengersen2013, Chib2018, Cheng2019, Zhao2020, Tang2022, and Yu2023.
The class of EL approaches often encounters significant challenges due to substantial computational complexity, which frequently presents barriers in practice. These difficulties primarily arise from the nonconvex nature of the objective function and the potential nonconvexity of its support. As the complexity of the model increases with additional parameters and conditions, these computational obstacles become more severe. Thus, developing computationally efficient strategies is crucial to address these challenges. Indeed, as demonstrated in Chau2017 and related works, solving the associated optimization problem of penalized EL (PEL) can be a dauntingly difficult task. In our study, we demonstrate that, when combined with the Bayesian framework, sampling schemes offer promising alternatives. Once successfully drawn, samples from the posterior distribution can be used to develop the estimator.
In recent research, sampling techniques, often perceived as computationally demanding alternatives to optimization methods, demonstrate remarkable efficiency in approximating target distributions, outperforming optimization alternatives in handling nonconvex problems; see Ma2019. While sampling techniques offer a promising approach within the framework of BEL, there exist numerous challenges associated with devising these computational schemes. On one hand, EL has the potential to leverage information from various model conditions, leading to more precise estimates of unknown model parameters. However, the inclusion of a large number of these conditions introduces additional complexities, both in theory and practical implementation. Indeed, the dimensionality of the problem remains a central obstacle in EL approaches, as elaborated in Hjortetal_2008_AS. Furthermore, the incorporation of an increasing number of moment conditions can substantially amplify the nonconvex nature of the associated optimization problems, making the development of an effective sampling scheme increasingly more challenging. As underscored in Chaudhuri2017, traditional MCMC techniques encounter significant hurdles when applied to BEL due to the intricate and nonconvex characteristics of the parameter space in which new samples are generated.
Our research aims to establish an innovative methodological framework, guided by two primary objectives: (i) our approach maintains the inherent flexibility and adaptability of EL, allowing for the incorporation of broad model conditions; and (ii) our framework provides convenient access to well-established MCMC computing schemes, streamlining practical implementations. To address the first objective and mitigate challenges stemming from numerous model conditions, we propose a penalized approach. By penalizing the magnitudes of the Lagrange multipliers used in evaluating EL at specific model parameter values, we create an effective mechanism similar to moment selection. This approach reduces the problem's dimensionality while still leveraging the potential efficiency gains from a comprehensive set of model conditions. For the second objective, our approach effectively overcomes the obstacles associated with devising sampling schemes for applying Bayesian approaches, thanks to the efficient dimensionality reduction achieved through PEL. In our study, we demonstrate the practicality of our framework using two well-established sampling methods: the popular Metropolis-Hastings sampling and the influential adaptive multiple importance sampling technique for approximate Bayesian computations.
Our study makes several noteworthy contributions, in addition to the methodological advancement mentioned earlier. On a theoretical level, our analysis establishes the properties of the BPEL estimator, allowing for an exponentially increasing number of model conditions, thereby enabling unprecedented adaptability in practical applications. Furthermore, we develop theory that guarantees the convergence of the two showcased sampling schemes, thereby ensuring the validity of BPEL in statistical inference. Our study reinforces the observations made in a recent study by Ma2019 that sampling techniques offer compelling alternatives to optimization methods in addressing computationally demanding problems. Our theoretical results and numerical studies demonstrate that sampling schemes converge rapidly to stationary distributions centered around the true global optimizer. In contrast, optimization methods often require more time and can become trapped at local peaks, limiting their ability to locate the true optimum.
The rest of this article are structured as follows. Section (ref) delves into the framework of BPEL and introduces two MCMC algorithms. Numerical studies and real data analysis for an international trade dataset are presented in Sections (ref) and (ref), respectively. Section (ref) comprehensively develops the properties and theoretical guarantees of the proposed methods. Some discussions are provided in Section (ref), while all technical proofs are available in the supplementary material. The used real data and the code for implementing our proposed methods are available at the GitHub repository: {\tt https://github.com/JinyuanChang-Lab/BayesianPenalizedEL}.
{\it Notation.} For any positive integer $q$, write $[q]=\{1,\ldots,q\}$ and let ${\mathbf I}_q$ be the $q \times q$ identity matrix. Denote by $I(\cdot)$ the indicator function. Let ${\rm vech}(\cdot)$ be an operator that stacks the columns of the lower triangular part of its argument square matrix. For a $q$-dimensional vector ${\mathbf a} = (a_1,\ldots,a_q)^{\mathrm{\scriptscriptstyle \top} }$, we use $|{\mathbf a}|_2 = (\sum_{i=1}^q a_i^2)^{1/2}$ and ${\rm supp}({\mathbf a}) = \{i \in [q]: a_i \neq 0\}$ to denote its $L_2$-norm and support, respectively. Let $\mathcal{U}(a,b)$ be the uniform distribution among $(a,b)$, and $\mathcal{N}(\boldsymbol{\mu},\boldsymbol{\Sigma})$ be the Gaussian distribution with mean $\boldsymbol{\mu}$ and covariance matrix $\boldsymbol{\Sigma}$. Denote by $\mathcal{T}_k(\boldsymbol{\mu},\boldsymbol{\Sigma})$ the multivariate Student's distribution with $k$ degrees of freedom, mean $\boldsymbol{\mu}$, and covariance matrix $\boldsymbol{\Sigma}$. For two positive real-valued sequences $\{a_n\}$ and $\{b_n\}$, we write $a_n\lesssim b_n$ if $\limsup_{n\rightarrow\infty}a_n/b_n\leq c_0$ for some positive constant $c_0$, $a_n\asymp b_n$ if $a_n\lesssim b_n$ and $b_n\lesssim a_n$ hold simultaneously, and $a_n\ll b_n$ if $\limsup_{n\rightarrow\infty}a_n/b_n=0$.
Let $\mathcal{X}_n = \{\mathbf{x}_1, \ldots, \mathbf{x}_n\}$ represent a set of $d$-dimensional independent and identically distributed observations, and let $\boldsymbol{\theta} = (\theta_1, \ldots, \theta_p)^{\mathrm{\scriptscriptstyle \top} } \in \boldsymbol{\Theta}$ be a $p$-dimensional parameter. Here, the parameter space $\boldsymbol{\Theta} \subset \mathbb{R}^p$ is a compact set. The information regarding the model parameter $\boldsymbol{\theta}$ is gathered through a set of unbiased moment conditions ${\mathbb{E}}\{{\mathbf g}({\mathbf x}_{i}; \boldsymbol{\theta}_0)\}={\mathbf 0}$, where $\mathbf{g}(\cdot\,; \cdot) = \{g_1(\cdot\,; \cdot), \ldots, g_r(\cdot\,; \cdot)\}^{\mathrm{\scriptscriptstyle \top} }\in{\mathbb R}^r$ is referred to as the estimating function, and the true, yet unknown value $\boldsymbol{\theta}_0$ is situated within the interior of $\boldsymbol{\Theta}$.
In existing studies, it has been typically required that $r\geq p$ for the identification of $\boldsymbol{\theta}_0$. When $p$ and $r$ are fixed constants, the EL with the estimating function ${\mathbf g}(\cdot\,;\cdot)$ considered in QinLawless_1994_AS can be formulated as
where $\hat{\Lambda}_n(\boldsymbol{\theta})=\{\boldsymbol{\lambda} \in {\mathbb{R}}^{r}:\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g} ( {\mathbf x}_{i}; \boldsymbol{\theta}) \in \mathcal{V}~\textrm{for any}~ i\in[n] \} $ with an open interval $\mathcal{V}$ containing zero. The standard EL estimator for $\boldsymbol{\theta}_0$ is defined as $\tilde{\boldsymbol{\theta}}_n = \arg \max _{\boldsymbol{\theta} \in \boldsymbol{\Theta}}{\rm EL}(\boldsymbol{\theta})$, which is equivalent to solving the corresponding dual problem:
The estimator $\tilde{\boldsymbol{\theta}}_n$ exhibits several desirable properties: (i) it is $\sqrt{n}$-consistent, (ii) it possesses asymptotic normality, and (iii) it attains the semiparametric efficiency bound of GodambeHeyde_1987_ISR. However, in high-dimensional scenarios, the literature has highlighted the challenge of accommodating a diverging $r$. This issue is discussed in works such as Donald2003, ChenPengQin_2008, Hjortetal_2008_AS, LengTang_2010, and ChangChenChen_2015. To elaborate, it is generally required that $r \ll n^{1/2}$ for the consistency and $r \ll n^{1/3}$ for the asymptotic normality of $\tilde{\boldsymbol{\theta}}_n$. These constraints on the diverging rate of $r$ pose significant challenges when dealing with high-dimensional estimating equations.
To address scenarios where $r\gg n$ and $p$ remains fixed, we investigate the PEL estimator for $\boldsymbol{\theta}_0$ as follows:
where $\boldsymbol{\lambda} = (\lambda_{1},\ldots,\lambda_{r})^{\mathrm{\scriptscriptstyle \top} }$, and $P_{\nu}(\cdot)$ is a penalty function with the tuning parameter $\nu$. Given a penalty function $P_{\nu}(\cdot)$ with the tuning parameter $\nu$, we define $\rho(t;\nu)=\nu^{-1}P_\nu(t)$ for $t\in[0,\infty)$ and $\nu\in(0,\infty)$. For $P_\nu(\cdot)$ in (ref), we consider the following class of penalty functions:
The class $\mathscr{P}$ is broad and general, encompassing commonly used penalty functions. Theorem (ref) in Section (ref) demonstrates that the PEL estimator $\hat{\boldsymbol{\theta}}_n$ follows an asymptotically normal distribution and accommodates exponentially diverging $r$ with respect to $n$.
To practically implement (ref), we encounter a two-layer optimization problem for $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ and $\boldsymbol{\lambda} \in \mathbb{R}^r$. Let
Since $n^{-1}\sum_{i=1}^n\log\{1+\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} } \mathbf{g}(\mathbf{x}_i;\boldsymbol{\theta})\}$ is concave in $\boldsymbol{\lambda}$, the inner optimization layer of (ref), which seeks $\boldsymbol{\lambda}$ given $\boldsymbol{\theta}$ by maximizing $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$, can be efficiently implemented even for large $r$ when $P_{\nu}(\cdot)$ is chosen as a convex function, such as the $L_1$ penalty. The main challenge is the outer optimization layer of (ref), which seeks the optimizer $\hat{\boldsymbol{\theta}}_n$. This is difficulty due to the nonconvex nature of the problem, making it NP-hard to find global minima Jain2017. As a result, this complexity often leads to computational inefficiency and a higher likelihood of converging to local optima.
We are motivated to explore an alternative approach using sampling techniques to solve the nonconvex problem associated with PEL. Indeed, as an efficient alternative for addressing nonconvex optimization problems, Ma2019 has demonstrated that solving these issues with MCMC techniques can yield highly effective results. Their findings indicate that the computational complexity of sampling algorithms exhibits linear scalability with the model dimension, in contrast to the exponential scaling of optimization algorithms in nonconvex settings.
Applying sampling techniques to EL in conjunction with a Bayesian framework emerges as a compelling approach. For ${\rm EL}(\boldsymbol{\theta})$ defined as (ref), let $\pi_0(\cdot)$ represent a prior distribution for $\boldsymbol{\theta}$. Then, the posterior distribution $\pi(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ is proportional to $\pi_{0}(\boldsymbol{\theta})\times {\rm EL}(\boldsymbol{\theta})$. In cases where $r$ and $p$ are fixed constants, $\pi(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ converges to a Gaussian distribution with mean being the standard EL estimator $\tilde{\boldsymbol{\theta}}_n$ defined as (ref). Consequently, when samples are successfully drawn from the posterior distribution, their sample mean can serve as an estimator for $\boldsymbol{\theta}_0$.
As the model's complexity increases, BEL faces challenges. In this study, we explore a scenario with high-dimensional model conditions ($r\gg n$), while keeping $p$ fixed. The flexibility by allowing large number $r$ also brings significant challenges. For example, as demonstrated in Tsao_2004_AOS, as $n\rightarrow\infty$, $\mathbb{P}\{{\rm EL}(\boldsymbol{\theta})=0\}\rightarrow1$ for any $\boldsymbol{\theta}$ in a small neighborhood of $\boldsymbol{\theta}_0$ if $r/n\geq0.5$. Such degeneration renders ${\rm EL}(\boldsymbol{\theta})$ inapplicable in this scenario. To handle diverging $r$, we propose to replace ${\rm EL}(\boldsymbol{\theta})$ by
where $P_{\nu}(\cdot)$ is a penalty function with the tuning parameter $\nu$. Since adding the penalty term $P_{\nu}(\cdot)$ encourages sparse Lagrange multiplier $\boldsymbol{\lambda}$, the PEL effectively performs a selection of the model conditions at each given $\boldsymbol{\theta}$. We then consider the BPEL with a prior distribution $\pi_{0}(\cdot)$, which leads to a posterior distribution defined as
Our BPEL connects with and differs from the so-called Gibbs posterior in the literature of Bayesian methods Bissiri_2016, Tang2022, Frazier2023. On one hand, they share a common foundation with the Gibbs posterior in that both are built upon generic loss functions. The key difference lies in the device each utilizes: EL employs an appropriate multinomial likelihood, \((p_1, \dots, p_n)\) with \(p_i \ge 0\) and \(\sum^n_{i=1} p_i = 1\), subject to a broad class of model conditions. In contrast, the Gibbs posterior uses a “pseudo-likelihood” proportional to the exponential loss. Furthermore, the inclusion of the penalty on the Lagrange multiplier helps achieve substantial dimension reduction of the problem, which is key in handling high-dimensional problems with many moment conditions. As shown in our numerical studies in Section (ref) and Section (ref) of the supplementary material, MCMC schemes developed from the proposed BPEL demonstrate compelling performance in their finite sample accuracy in approximating the posterior distributions.
Our theory, as elaborated in Section (ref), establishes the fundamental properties of BPEL. Theorem (ref) in Section (ref) demonstrates that the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ defined as (ref) exhibits a Gaussian limiting distribution centered around the PEL estimator $\hat{\boldsymbol{\theta}}_n$ as defined in (ref). Additionally, we define the expected value as
Corollary (ref) in Section (ref) suggests that $\hat{\boldsymbol{\theta}}_n$ can be effectively approximated by $\mathbb{E}_{\boldsymbol{\theta}\sim\pi^\dag}(\boldsymbol{\theta})$ with an approximation error that diminishes faster than $n^{-1/2}$. This validates the approach to obtain $\hat{\boldsymbol{\theta}}_n$: generating samples from the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ and then using the associated sample mean to approximate $\hat{\boldsymbol{\theta}}_n$. In Section (ref), we will introduce two algorithms designed for implementing BPEL.
The impact of prior specification on the properties of resulting estimators is a notable area of research. For instance, Vexetal_2014_Bioka explores this in the context of EL. In various scenarios, the choice of prior can enhance desirable properties of the estimator derived from the posterior distribution, such as sparsity, as discussed in NarisettyHe2014, Castillo_2015, and OuyangBondell_2023. Given the two primary goals of our study -- developing BPEL and investigating it with MCMC -- we use a non-informative prior in our numerical demonstrations. As detailed in Section (ref) of the supplementary material, we examined the effects of different prior specifications. The overall finding is intuitive: when the prior is specified closer to the true value, the resulting estimator performs better compared to using a non-informative prior. Conversely, if the prior is specified further from the true value, the performance of the estimator deteriorates and becomes less competitive.
In recent decades, MCMC sampling methods have achieved significant success and have garnered influential applications across diverse fields. For an extensive overview of this body of work, we refer to the monograph by Brooks2011 and reference therein. The Metropolis-Hastings (M-H) algorithm family plays a central role in the practical implementation of MCMC techniques, serving as a cornerstone in the toolbox of statisticians and data scientists.
Our first algorithm explores the utilization of the M-H algorithm for BPEL. To accomplish this, we begin by specifying a proposal distribution with a density function denoted as $\phi(\cdot\,|\,{\mathbf x})$, where ${\mathbf x} \in \mathbb{R}^p$. Subsequently, we employ the M-H algorithm to generate samples from the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$, as defined in (ref). The specific steps for this process are detailed in Algorithm (ref).
At each iteration $k$, Algorithm (ref) begins with a state $\boldsymbol{\theta}^k \in \boldsymbol{\Theta}$. In the proposal step, it generates a new parameter $\boldsymbol{\vartheta}^{k+1}$ from the proposal distribution centered at $\boldsymbol{\theta}^k$, denoted by $\phi(\cdot\,|\,\boldsymbol{\theta}^k)$. Following this, in the accept-reject step, Algorithm (ref) decides whether to accept $\boldsymbol{\vartheta}^{k+1}$ with a probability denoted as $\alpha^{k+1}$. This crucial step ensures that the Markov chain, guided by Algorithm (ref), remains within the valid parameter space $\boldsymbol{\Theta}$. Consequently, it expedites the convergence of the resulting chain towards its stationary distribution, which is the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$. There exist various approaches for selecting the proposal distribution with density $\phi(\cdot\,|\,\cdot)$, including methods like the symmetric Metropolis algorithm, random walk M-H, and the independence sampler, as detailed by Roberts2004.
Another widely-used MCMC technique is Importance Sampling Ripley2006, Hesterberg1995. This method involves generating samples from a proposal distribution and then applying importance weights to these samples to account for the disparities between the proposal distribution and the target distribution. In practical applications, recycling successive samples often proves to be an effective strategy Marin2019, particularly when the computation of importance weights is computationally intensive. In this context, CORNUET2012 introduces the Adaptive Multiple Importance Sampling (AMIS) algorithm, which combines various importance sampling methods with adaptive techniques. The integration of the AMIS approach with EL, as shown in Mengersen2013, is particularly compelling. To ensure the consistency of AMIS, Marin2019 introduces a modified variant called Modified AMIS (MAMIS) with a simpler recycling strategy compared to AMIS.
We present and investigate an MAMIS algorithm, as outlined in Algorithm (ref), specifically designed for computing BPEL. This algorithm operates in a scenario where a density function $\varphi(\cdot\,; \boldsymbol{\zeta})$ is defined, with $\boldsymbol{\zeta}$ representing a parameter in ${\mathbb{R}}^s$, and where an explicit function ${\mathbf h}: {\mathbb{R}}^p \mapsto {\mathbb{R}}^s$ is known. This configuration allows us to generate weighted samples that effectively capture the characteristics of the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$, as defined in (ref).
Algorithm (ref) generates a sequence of samples while progressively adjusting the parameter $\boldsymbol{\zeta} \in {\mathbb{R}}^s$ involved in the proposal distribution. At each iteration $k$ of Algorithm (ref), the new value for the parameter $\boldsymbol{\zeta}$ of the proposal distribution is determined based on the most recent $N_k$ samples drawn. This represents the primary distinction between the MAMIS algorithm by Marin2019 and the AMIS algorithm by CORNUET2012. Specifically, MAMIS updates the proposal distribution parameter using only the last $N_k$ samples at iteration $k$, while AMIS updates this parameter by considering all past $\sum_{j=1}^{k}N_j$ samples. The end product output of Algorithm (ref) is generated by updating the importance weights for all samples produced during the recycling process.
We advocate the utilization of sampling techniques as a practical and efficient alternative to optimization methods for addressing computationally challenging PEL problems. Specifically for obtaining the estimator $\hat{\boldsymbol{\theta}}_n$ as defined in (ref), we can rely on samples $\boldsymbol{\theta}^1,\ldots,\boldsymbol{\theta}^K$ generated from the M-H algorithm (see Algorithm (ref)), estimating $\mathbb{E}_{\boldsymbol{\theta}\sim\pi^\dag}(\boldsymbol{\theta})$, as defined in (ref), by computing the sample mean, i.e., $K^{-1}\sum_{k=1}^K\boldsymbol{\theta}^k$. When employing the MAMIS algorithm (see Algorithm (ref)) and completing $K$ iterations, the estimator for $\mathbb{E}_{\boldsymbol{\theta}\sim\pi^\dag}(\boldsymbol{\theta})$ is determined as a weighted average:
where $S_K=N_1+\cdots+N_K$.
Our theory in Section (ref) supports the use of sampling algorithms as efficient alternatives. For the M-H algorithm, Theorem (ref) in Section (ref) demonstrates that, conditional on $\mathcal{X}_n$, the average $K^{-1}\sum_{k=1}^K\boldsymbol{\theta}^k$ converges almost surely to $\mathbb{E}_{\boldsymbol{\theta}\sim\pi^\dag}(\boldsymbol{\theta})$ as $K\rightarrow\infty$. For the MAMIS algorithm, Theorem (ref) in Section (ref) establishes that, conditional on $\mathcal{X}_n$, $\widehat{{\mathbb{E}}}_{\pi^\dag,K} (\boldsymbol{\theta})$ in (ref) converges almost surely to $\mathbb{E}_{\boldsymbol{\theta}\sim\pi^\dag}(\boldsymbol{\theta})$ as $K\rightarrow\infty$. These results, combined with Corollary (ref) in Section (ref), validate the properties of BPEL estimators obtained through these established sampling techniques. Another consideration in Algorithms (ref) and (ref) is the choice of the initial point, denoted, respectively, as $\boldsymbol{\theta}^{0}$ and $\hat{\boldsymbol{\zeta}}_1$. Our theoretical analyses only require $\boldsymbol{\theta}^0 \in \boldsymbol{\Theta}$ satisfying $\pi^\dag(\boldsymbol{\theta}^0\,|\,\mathcal{X}_n)>0$ and do not impose any restriction on $\hat{\boldsymbol{\zeta}}_1$; see Theorems (ref) and (ref) in Section (ref) for details. Our empirical simulation studies in Section (ref) consistently demonstrate the proposed algorithms' robust performance, irrespective of the initial value chosen. Notice that the performance of the optimization methods for the nonconvex optimization problems usually depends crucially on the choice of the initial point. The combination of theoretical analysis and empirical evidence underscores that, in comparison to competing optimization methods, these sampling-based approaches offer significant advantages in terms of convergence speed, stability across replications, and resilience to variations in initial values. This reaffirms the benefits of incorporating BPEL into the methodology.
The M-H and MAMIS algorithms each have their strengths. M-H is easy to implement, but high rejection rates can reduce its efficiency, especially with a poorly tuned proposal distribution. MAMIS, while requiring more effort -- particularly in computing importance weights -- offers improved sampling efficiency and is less sensitive to the proposal distribution, making it ideal for complex posterior distributions. Choosing between these algorithms depends on the specific problem and the balance between implementation ease and sampling efficiency.
We conduct simulation studies to empirically assess the performance of our proposed methods. For the data generation process (DGP), we adopt the structural equation $y_i = \hbar({\mathbf u}_i^{\mathrm{\scriptscriptstyle \top} } \boldsymbol{\theta}_0) + e_i^{(0)}$, $i \in [n]$, where $\hbar: {\mathbb{R}} \mapsto {\mathbb{R}}$ is a continuous function, $e_i^{(0)}$ is the error, and ${\mathbf u}_i=(u_{i,1},u_{i,2})^{\mathrm{\scriptscriptstyle \top} }$ represents two endogenous variables. The set of all instrumental variables (IVs) is denoted as ${\mathbf z}_{i}=(z_{i,1}, \ldots, z_{i,r})^{\mathrm{\scriptscriptstyle \top} }$ for $i \in [n]$. The true reduced-form equations for the endogenous variables are specified as $u_{i,1}=0.5z_{i,1}+0.5z_{i,2}+e_i^{(1)}$ and $ u_{i,2}=0.5z_{i,3}+0.5z_{i,4}+e_i^{(2)}$, where $(e_i^{(1)}, e_i^{(2)})$ represents the random errors. Essentially, each of the two endogenous variables is influenced by only two IVs. All IVs are selected orthogonal to the error term $e_i^{(0)}$. Hence, we have $ {\mathbb{E}} \{y_i - \hbar({\mathbf u}_i^{\mathrm{\scriptscriptstyle \top} } \boldsymbol{\theta}_0)\,|\, {\mathbf z}_{i}\} ={\mathbf 0} $, which implies that $\boldsymbol{\theta}_0$ can be identified by the $r$ unbiased moment conditions $\mathbb{E}\{{\mathbf g}({\mathbf x}_i;\boldsymbol{\theta}_0)\}={\mathbf 0}$, where ${\mathbf g}({\mathbf x}_i;\boldsymbol{\theta}) = \{y_i - \hbar({\mathbf u}_i^{\mathrm{\scriptscriptstyle \top} } \boldsymbol{\theta}) \}{\mathbf z}_i $ with ${\mathbf x}_i= (y_i, {\mathbf u}_i^{\mathrm{\scriptscriptstyle \top} }, {\mathbf z}_i^{\mathrm{\scriptscriptstyle \top} })^{\mathrm{\scriptscriptstyle \top} }$. In the DGP, we generate ${\mathbf z}_{i}\sim\mathcal{N}({\mathbf 0},{\mathbf I}_r)$, and
We set $\boldsymbol{\theta}_0=(0.5,0.5)^{\mathrm{\scriptscriptstyle \top} }$ and consider two selections for the link function $\hbar(\cdot)$: (i) the linear case with $\hbar(v)=v$, and (ii) the nonlinear case with $\hbar(v)=\sin v$.
We begin by demonstrating the improvement in sampling efficiency achieved through the use of PEL. In this context, we generate data following the DGP with linear link function $\hbar(\cdot)$ by setting $n=120$ and varying $r$ in the range $[50,1000]$. We aim to sample from two posterior distributions $\pi_{0}(\boldsymbol{\theta}) \times {\rm EL}(\boldsymbol{\theta})$ and $\pi_{0}(\boldsymbol{\theta}) \times {\rm PEL}_{\nu}(\boldsymbol{\theta})$, where ${\rm EL}(\boldsymbol{\theta})$ and ${\rm PEL}_{\nu}(\boldsymbol{\theta})$ are, respectively, given in (ref) and (ref). Evaluating ${\rm PEL}_{\nu}(\boldsymbol{\theta})$ involves an optimization problem that solves for $\boldsymbol{\lambda}$ by maximizing the objective function $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ defined as (ref) at given $\boldsymbol{\theta}$. To ensure the attainment of a sparse Lagrange multiplier and maintain the convexity of the objective function, we select $P_{\nu}(\cdot)$ as the $L_1$ penalty function. In practice, since the prior information about the true parameter $\boldsymbol{\theta}_0$ is typically unavailable, we select $\pi_0(\cdot)$ as the improper uniform prior. We implement Algorithm (ref) to sample from both posterior distributions using identical settings, employing a proposal distribution $\mathcal{N}(\boldsymbol{\theta}, \sigma^2{\mathbf I}_p)$ with $\sigma^2 = 10^{-4}$ and initializing from $\boldsymbol{\theta}^0 = (0.3, 0.3)^{\mathrm{\scriptscriptstyle \top} }$. In the case of PEL, we set the tuning parameter $\nu=0.03$ involved in ${\rm PEL}_{\nu}(\boldsymbol{\theta})$.
To compare efficiency, we measure the number of iterations required to obtain the same number of accepted samples. Figure (ref) illustrates the average number of iterations needed over 500 runs to accept 5 samples for different values of $r$, thereby providing a comparison between using EL and PEL within a Bayesian framework. The sampling efficiency of Algorithm (ref) when using ${\rm PEL}_{\nu}(\boldsymbol{\theta})$ is notably superior to that achieved with ${\rm EL}(\boldsymbol{\theta})$. The selection of $\sigma^2$ within the proposal distribution closely influences the acceptance rate in each step of the M-H algorithm. With our small choice of $\sigma^2$ in the simulation, the M-H algorithm should efficiently generate valid samples. It is worth highlighting that the acceptance rate remains consistently high and stable when using PEL across all $r$ settings. In contrast, when employing EL without any penalty, it may require thousands more iterations to achieve the same number of accepted samples. Additionally, it is evident that the M-H algorithm with EL becomes increasingly unstable as $r$ increases.
As we suggested in Section (ref), the computation of the PEL estimator $\hat{\boldsymbol{\theta}}_n$ defined as (ref) can be implemented using Algorithm (ref) (referred to as M-H) and Algorithm (ref) (referred to as MAMIS). In this part, we compare their performance with two optimization methods: (a) {\tt optim}: A versatile R function for general-purpose optimization of objective functions, supporting various optimization algorithms like Nelder-Mead, quasi-Newton, and conjugate-gradient; and (b) {\tt nlm}: An R function specialized in non-linear optimization, particularly designed for finding minima of objective functions using Newton-type algorithms.
The choice of the proposal distribution plays a crucial role in achieving efficient sampling with BPEL. Within the context of the M-H algorithm, one commonly used scheme is the random walk M-H, where the proposal distribution takes the form of a Gaussian distribution $\mathcal{N}(\boldsymbol{\theta}, \sigma^2 {\mathbf I}_p)$ with the current state denoted as $\boldsymbol{\theta}$. It is essential to carefully select an appropriate value for $\sigma^2$. A small value for $\sigma^2$ can result in slow exploration of the state space, while a large value can lead to decreased acceptance rates, subsequently slowing down the algorithm. To strike a balance between exploration and acceptance rates, we can monitor the acceptance rate of the algorithm. In the simulation for M-H, we set $\sigma^2=C(n\log r)^{-1}$ with some constant $C>0$. We adjust the value of $C$ until the acceptance rate closely matches the desired rate, typically aiming for approximately $0.234$, as suggested in Gelman1997. It is known that the M-H algorithm requires some time to converge to its stationary distribution, especially when the initial point $\boldsymbol{\theta}^0 \in \boldsymbol{\Theta}$ is situated in the tails of the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$. Considering this, we set a burn-in period of $500$ iterations. For the MAMIS algorithm, we adhere to recommendations from CORNUET2012 and Mengersen2013 that advocate for the adoption of $\mathcal{T}_3(\boldsymbol{\mu},\boldsymbol{\Sigma})$ as the proposal distribution. During each iteration $k$ of MAMIS, we calculate the updated value $\hat\boldsymbol{\zeta}_{k+1} = \{\hat{\boldsymbol{\mu}}_{k+1}^{\mathrm{\scriptscriptstyle \top} }, {\rm vech}(\widehat{\boldsymbol{\Sigma}}_{k+1})^{\mathrm{\scriptscriptstyle \top} }\}^{\mathrm{\scriptscriptstyle \top} }$ for the parameter vector $\boldsymbol{\zeta} = \{\boldsymbol{\mu}^{\mathrm{\scriptscriptstyle \top} }, {\rm vech}(\boldsymbol{\Sigma})^{\mathrm{\scriptscriptstyle \top} }\}^{\mathrm{\scriptscriptstyle \top} }$ involved in the proposal distribution $\mathcal{T}_3(\boldsymbol{\mu},\boldsymbol{\Sigma})$ as $\hat{\boldsymbol{\mu}}_{k+1} = N_k^{-1} \sum_{i=1}^{N_k} \omega^k_i \boldsymbol{\theta}^k_i$ and ${\rm vech}(\widehat{\boldsymbol{\Sigma}}_{k+1}) = N_k^{-1} \sum_{i=1}^{N_k} \omega^k_i {\rm vech}\{(\boldsymbol{\theta}^k_i - \hat{\boldsymbol{\mu}}_{k+1})(\boldsymbol{\theta}^k_i - \hat{\boldsymbol{\mu}}_{k+1})^{\mathrm{\scriptscriptstyle \top} }\}$, where $\omega^k_i$ represents the corresponding importance weights, as outlined in Algorithm (ref). In our simulations, we initialize $\widehat{\boldsymbol{\Sigma}}_{1} = {\mathbf I}_p$, and the selection of $\hat{\boldsymbol{\mu}}_{1}$ is described in the next paragraph.
We conduct 200 replications following the DGP and explore various combinations of dimensionalities. Specifically, we consider $n\in\{120, 240\}$ and $r\in\{80, 160, 320, 640\}$. To assess the robustness of these methods with respect to initial points, we select $49$ equally spaced grid points on the plane within the range of $[-3, 4] \times [-3, 4]$ as our chosen initial points. In the case of MAMIS, which is not an iterative algorithm, we set these initial points as the initial means $\hat{\boldsymbol{\mu}}_{1}$ for its proposal distribution $\mathcal{T}_3(\boldsymbol{\mu},\boldsymbol{\Sigma})$ to facilitate comparison. In our simulations, we identify the true global minima $\hat{\boldsymbol{\theta}}_n$ defined as (ref) through exhaustive search. To achieve this, in each replication of the simulation (indexed by $k$), we generate a grid of $10201$ equally spaced points within the range $[-0.5, 1.5]\times[-0.5, 1.5]$. We then compute the posterior probabilities for these points and selected $\boldsymbol{\theta}^{\rm mode}_{k}$ as the point with the highest probability. Since $\pi_0(\cdot)$ is selected as the improper uniform prior, $\boldsymbol{\theta}^{\rm mode}_k$ is actually the required true global minima in the $k$-th replication. We repeat this process for $k=1$ to $200$, and compare the outcomes obtained from both optimization and sampling methods by calculating the measure $$ {\rm MSE}_1 = \frac{1}{200 \times 49} \sum_{k=1}^{200} \sum_{l=1}^{49} |\check{\boldsymbol{\theta}}_{k}(l) - \boldsymbol{\theta}^{\rm mode}_{k}|_2^2\,.$$ Here, $\check{\boldsymbol{\theta}}_{k}(l)$ represents the related outcome in the $k$-th replication initiated from the $l$-th initial point.
In the context of BPEL sampling, we explore three scenarios with varying sample sizes of $1500$, $2500$, and $3500$, which we label as (M-H-1, M-H-2, M-H-3) and (MAMIS-1, MAMIS-2, MAMIS-3), respectively, for Algorithms 1 and 2. Additionally, we conduct an investigation into the influence of different values for the tuning parameter $\nu$. Table (ref) presents the simulation results. The overall performance of the sampling approaches surpasses that of the optimization methods. Notably, for the nonlinear model, the optimization using the R function {\tt nlm} is proven to be unreliable, resulting in highly unstable results. As the size of the generated samples increases, the performance of BPEL improves. Both M-H and MAMIS exhibit promising performance in both linear and nonlinear cases. For the nonlinear models, MAMIS significantly outperforms M-H, possibly owing to the advantages gained from employing importance weights for parameter estimation. The role of the tuning parameter $\nu$ is pivotal, underscoring the merits from using the PEL approach in achieving more parsimonious models by effectively selecting most useful model conditions within the constraints of the available data information. When using very small values of $\nu$, such as 0.01, the performance of the methods becomes less satisfactory. Overall, the BPEL performs satisfactorily for a reasonable range of choices for $\nu$.
In this part, we compare the PEL estimator $\hat{\boldsymbol{\theta}}_n$ defined as (ref) with two other estimators: the standard EL estimator $\tilde{\boldsymbol{\theta}}_n$ defined as (ref) and the relaxed EL (REL) estimator introduced by Shi2016. The REL is tailored for high-dimensional estimating equations, making it resilient to minor deviations from the equality constraints. Notice that the standard EL can only work for low-dimensional estimating equations. In line with our model specifications, where the two endogenous variables $u_{i,1}$ and $u_{i,2}$ are linked to IVs ($z_{i,1}, z_{i,2}$ and $z_{i,3}, z_{i,4}$, respectively) for each $i\in[n]$, we only use the first four moment conditions, that are related to the IVs $z_{i,1}, z_{i,2}, z_{i,3}$ and $z_{i,4}$, to produce the standard EL estimator $\tilde{\boldsymbol{\theta}}_n$. The computation of $\tilde{\boldsymbol{\theta}}_n$ can be implemented by the function {\tt gel} in the R-package {\tt gmm}. For both our PEL estimator $\hat{\boldsymbol{\theta}}_n$ and the REL estimator, we use all the $r$ moment conditions.
For the selection of the tuning parameter in the REL estimator, we follow the recommendation in Shi2016, using a consistent tuning parameter $0.5n^{-1/2}(\log r)^{1/2}$ throughout the simulations. Regarding the tuning parameter $\nu$ in our BPEL, we employ the Bayesian Information Criterion (BIC) defined as
for its selection, where $\hat{\boldsymbol{\theta}}_{n}^{(\nu)}$ is the associated PEL estimator with tuning parameter $\nu$ calculated by our sampling algorithm, and $\mathcal{R}_n^{(\nu)}={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n^{(\nu)})\}$ with $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n^{(\nu)})=(\hat\lambda_1^{(\nu)},\ldots,\hat\lambda_r^{(\nu)})^{\mathrm{\scriptscriptstyle \top} }=\arg\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_n(\hat{\boldsymbol{\theta}}_n^{(\nu)})}f_n(\boldsymbol{\lambda};\hat{\boldsymbol{\theta}}_n^{(\nu)})$ with $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ defined as (ref). In practice, we set $\mathcal{R}_n^{(\nu)} = \{j \in [r]: |\hat\lambda_j^{(\nu)}|>10^{-6}\}$ and restrict $\nu$ in the interval $[0.05n^{-1/2}(\log r)^{1/2}, 0.75n^{-1/2}(\log r)^{1/2}]$.
For the same $49$ initial points of the $200$ replications mentioned in Section (ref), we calculate the measure $$ {\rm MSE}_2 = \frac{1}{200 \times 49} \sum_{k=1}^{200} \sum_{l=1}^{49} |\check{\boldsymbol{\theta}}_{k}(l) - \boldsymbol{\theta}_0|_2^2 $$ to evaluate the performance of different estimators, where $\check{\boldsymbol{\theta}}_{k}(l)$ is the related estimator in the $k$-th replication initiated from the $l$-th initial point. Table (ref) compares the measure ${\rm MSE}_2$ for the three estimators: the PEL estimator (MAMIS, M-H), the standard EL estimator, and the REL estimator. The results for M-H and MAMIS are derived based on the generated samples of size $3500$. It becomes clear that BPEL demonstrates substantial performance improvements, clearly establishing its superiority over the other estimation methods. Particularly noteworthy is the effectiveness of MAMIS in addressing the challenges posed by nonlinear estimating equations, showing its promising performance.
We provide additional simulation studies in the supplementary material: Section (ref) examines the impact of prior specification, Section (ref) evaluates the performance of our method using an alternative data generation process with data from a Student's $t$-distribution instead of a normal distribution, Section (ref) assesses the finite sample accuracy of the MCMC algorithms in approximating the posterior distribution, Section (ref) compares the posterior distributions resulting from different Bayesian EL formulations, and Section (ref) presents the comparison between our method and two competing methods: approximate Bayesian computation and Bayesian synthetic likelihood. Overall, our findings confirm the highly competitive performance of the proposed BPEL with the MCMC framework in terms of finite sample performance and accuracy in approximating posterior distributions.
International trade refers to the cross-border exchange of capital, commodities, and services between nations or regions. This type of trade typically constitutes a substantial portion of a country's gross domestic product (GDP). Eaton2011, hereafter referred to as EKK, combined an empirical model with microeconomic principles to analyze France's international trade patterns. Additionally, Shi (2016) utilized EKK's microeconomic model to derive parameter estimates for Chinese exporting companies. In this section, we reexamine the dataset previously examined in Shi2016, employing the proposed BPEL approach.
The model proposed by EKK comprises five parameters denoted as $\boldsymbol{\theta} = (\theta_1, \ldots, \theta_5)^{\mathrm{\scriptscriptstyle \top} } \in \boldsymbol{\Theta}$. The first component, $\theta_1$, characterizes the distribution of production efficiency among firms, with a higher $\theta_1$ indicating a larger proportion of manufacturers with lower efficiency. The second component, $\theta_2$, quantifies the cost associated with accessing a fraction of potential buyers, where a higher $\theta_2$ corresponds to lower costs. Parameters $\theta_3$, $\theta_4$ and $\theta_5$ represent the standard deviation of the demand shock, the standard deviation of the entry cost shock, and the correlation coefficient between these two shocks, respectively. Each firm is identified by the index $i \in [n]$, while countries are represented by the index $j \in \{0\}\cup[r]$, with $j = 0$ denoting the home country.
According to the EKK's model, the sales of firm $i$ in country $j$ is $Z_{i,j}(\boldsymbol{\theta}; e^{(1)}_{i,j}, e^{(2)}_{i,j}, e^{(3)}_{i}) = \kappa \bar{Z}_j (1-\tau_{i,j})^{\theta_2/\theta_1} \tau_{i,j}^{-1/\theta_1} a^{(1)}_{i,j}/a^{(2)}_{i,j}$, where $a^{(1)}_{i,j}= \exp\{\theta_3(1-\theta_5^2)^{1/2}e^{(1)}_{i,j} + \theta_3\theta_5 e^{(2)}_{i,j}\}$, $a^{(2)}_{i,j} = \exp\{\theta_4e^{(2)}_{i,j}\}$, $\tau_{i,j} = \min\{1, e^{(3)}_{i} \bar{u}_{i} / \bar{u}_{i,j}\}$ and
with $\bar{u}_{i,j} = (a^{(2)}_{i,j})^{\theta_1} N_j$ and $ \bar{u}_{i} = \min\{\bar{u}_{i,0}, \max_{j\in[r]}\bar{u}_{i,j}\}$, and $(\bar{Z}_j, N_j)_{j\in\{0\}\cup[r]}$ are known constants. Here $e^{(1)}_{i,j} \sim \mathcal{N}(0,1)$, $e^{(2)}_{i,j} \sim \mathcal{N}(0,1)$ and $e^{(1)}_{i} \sim \mathcal{U}(0,1)$ are mutually independent. Furthermore, $Z_{i,j}(\boldsymbol{\theta}; e^{(1)}_{i,j}, e^{(2)}_{i,j}, e^{(3)}_{i}) = 0$ means that the firm $i$ is kept outside of the country $j$. As a pertinent economic indicator of our interest, the mean sale of all firms in country $j$ is $\mu_j(\boldsymbol{\theta}) = \mathbb{E}\{Z_{i,j}(\boldsymbol{\theta}; e^{(1)}_{i,j}, e^{(2)}_{i,j}, e^{(3)}_{i})\}$, where the expectation is taken respect to the random variables $\{e^{(1)}_{i,j}, e^{(2)}_{i,j}, e^{(3)}_{i}\}$. The dataset is sourced from the Chinese administrative databases, encompassing a total of $n=6754$ firms and their export data to $r=126$ foreign destination countries in 2006. Leveraging this dataset, we obtain the $r$-dimensional estimating function ${\mathbf g}({\mathbf x}_{i}; \boldsymbol{\theta}) =\{{g}_{1} ({\mathbf x}_{i}; \boldsymbol{\theta}),\ldots,{g}_{r} ({\mathbf x}_{i}; \boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} } $, $i\in[n]$, with ${\mathbf x}_{i} = (x_{i,1},\ldots,x_{i,r})^{\mathrm{\scriptscriptstyle \top} }$ and $g_j({\mathbf x}_{i}; \boldsymbol{\theta}) = x_{i,j} - \mu_j(\boldsymbol{\theta})$ for any $j\in[r]$ and $\boldsymbol{\theta} \in \boldsymbol{\Theta}$, where $x_{i,j}$ is the sale of firm $i$ in country $j$ from this dataset ($j=0$ is not considered in this dataset).
Since the model is highly nonlinear with respect to $\boldsymbol{\theta} \in \boldsymbol{\Theta}$, resulting in no closed-form expression for $\mu_j(\boldsymbol{\theta})$, we approximate it via numerical simulation Eaton2011, Shi2016. Specifically, in the estimation, we utilize the “artificial data" for another $5n = 33770$ firms from the dataset. This involves simulating the entry decisions and sales across various countries for each of these artificial firms. Subsequently, we calculate sample means to approximate $\mu_j(\boldsymbol{\theta})$ for any $j\in \{0\}\cup[r]$ and $\boldsymbol{\theta} \in \boldsymbol{\Theta}$. We generated samples of size 3500 from the posterior distribution for the BPEL. To select the tuning parameter $\nu$, we employed the BIC as defined in (ref). For the parameter space $\boldsymbol{\Theta}$, we adopted a compact range of values, specifically $\boldsymbol{\Theta} = [1.5, 10] \times [0.5, 5] \times [0.1, 5] \times [0.1, 5] \times [-0.9, 0.9]$, which is consistent with the economic context and aligns with the study of Shi2016. To initiate the analysis, we selected 15 samples uniformly distributed within the parameter space $\boldsymbol{\Theta}$. Figure (ref) presents the box-plots of the corresponding 15 estimates obtained by M-H and MAMIS from these initial values. The results for the REL with the same initial values are also included for comparative evaluation.
It is evident that for all five parameters, MAMIS exhibits the smallest variations in the resulting estimates, whereas the variations of M-H and REL are relatively similar. This consistency with the findings in Sections (ref) and (ref) reaffirms the robustness of MAMIS when considering different initial points. Such robustness is desirable for conducting more in-depth analyses. For instance, let us take $\theta_5$ into consideration which represents the correlation coefficient between the demand shock and the entry cost shock. The sign of its estimate carries the key implication. The 15 estimates of $\theta_5$ obtained by REL and M-H, from different initial values, fall within the ranges of $(-0.8738, 0.8507)$ and $(-0.8774, 0.8996)$, respectively. In contrast, the estimates of $\theta_5$ by MAMIS range in $(-0.7978, 0.1329)$, with the majority being negative, signaling a more assuring result.
We then proceed to examine the specific moments selected by the respective methods. For REL, we employ the greedy algorithm outlined in Section 3.2 of Shi2016. To assess the effectiveness of moment selection, we validate whether or not the top 10 trading partners of China in terms of export volume in this dataset, including the USA, Japan, Germany, etc., are either selected or partially selected. We find that, although REL selects at least some of these countries for 10 out of the 15 initial values, the number of selected countries does not exceed 3. In contrast, for 13 out of the 15 initial values, M-H identifies at least some of these countries, with 9 of them including more than 3. In the case of MAMIS, 13 out of the 15 initial values result in the identification of some of these countries, and all of them include more than 3 countries. Additionally, the robustness of MAMIS with respect to the initial points provides enhanced reliability in this context.
We introduce some additional notation first. For simplicity, write $\mathbb{E}_n(\cdot)=n^{-1}\sum_{i=1}^{n}\cdot$. For a $q \times q$ symmetric matrix ${\mathbf A}$, denote by $\lambda_{\min}({\mathbf A})$ and $\lambda_{\max}({\mathbf A})$ the smallest and largest eigenvalues of ${\mathbf A}$, respectively. For a $q_1 \times q_2$ matrix ${\mathbf B}=(b_{i,j})_{q_1 \times q_2}$, let $|{\mathbf B}|_{\infty} = \max_{i \in [q_1], j \in [q_2]} |b_{i,j}|$ be the super-norm. For the $r$-dimensional estimating function ${\mathbf g}(\cdot\,; \cdot)=\{g_1(\cdot\,;\cdot),\ldots,g_r(\cdot\,;\cdot)\}^{\mathrm{\scriptscriptstyle \top} }$ and $p$-dimensional parameter $\boldsymbol{\theta}=(\theta_1,\ldots,\theta_p)^{\mathrm{\scriptscriptstyle \top} }$, let $\nabla_{\boldsymbol{\theta}} {\mathbf g} (\cdot\,; \boldsymbol{\theta})=\{\partial g_j(\cdot\,;\boldsymbol{\theta})/\partial \theta_k\}_{j\in[r],k\in[p]}$, an $r\times p$ matrix, be the first-order partial derivative of ${\mathbf g} (\cdot\,; \boldsymbol{\theta})$ with respect to $\boldsymbol{\theta}$. Let ${\mathbf V} (\boldsymbol{\theta})= {\mathbb{E}} \{ {\mathbf g} ({\mathbf x}_{i}; \boldsymbol{\theta})^{\otimes2} \}$ and $\boldsymbol{\Gamma}(\boldsymbol{\theta}) ={\mathbb{E}}\{\nabla_{\boldsymbol{\theta}} {\mathbf g}({\mathbf x}_{i}; \boldsymbol{\theta}) \}$ for any $\boldsymbol{\theta}\in\boldsymbol{\Theta}$. For a given index set $\mathcal{F}$, let $|\mathcal{F}|$ be its cardinality. Denote by ${\mathbf g}_{\mathcal{F}}(\cdot\,; \cdot)$ the subvector of ${\mathbf g}(\cdot\,; \cdot)$ collecting the components indexed by $\mathcal{F}$. Let ${\mathbf V}_{\mathcal{F}} (\boldsymbol{\theta})= {\mathbb{E}} \{ {\mathbf g}_{\mathcal{F}} ( {\mathbf x}_{i}; \boldsymbol{\theta})^{\otimes2} \}$ and $\boldsymbol{\Gamma}_{\mathcal{F}}(\boldsymbol{\theta}) ={\mathbb{E}} \{\nabla_{\boldsymbol{\theta}} {\mathbf g}_{\mathcal{F}}({\mathbf x}_{i}; \boldsymbol{\theta}) \}$. Analogously, we also write ${\mathbf a}_{\mathcal{F}}$ as the corresponding subvector of vector ${\mathbf a}$. For any two probability measures $\mu$ and $\nu$, denote by $\mathcal{D}_{\mathrm{\scriptscriptstyle TV} }(\mu, \nu)$ the total variation distance between $\mu$ and $\nu$.
To investigate the asymptotic properties of $\hat{\boldsymbol{\theta}}_n$ in (ref), we assume some regularity conditions.
Detailed discussion on Conditions (ref) and (ref) are given in Section (ref) of the supplementary material. For any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$, define $$ \mathcal{M}_{\boldsymbol{\theta}}^*=\{j \in [r]: |\mathbb{E}_n\{ g_j ({\mathbf x}_{i}; \boldsymbol{\theta})\} | \geq C_*\nu \rho'(0^{+}) \} $$ for some $C_* \in (0,1)$. We assume the existence of a sequence $\ell_n\rightarrow\infty$ such that $$\mathbb{P}\bigg(\sup _{\boldsymbol{\theta} \in \boldsymbol{\Theta}:\, |\boldsymbol{\theta}-\boldsymbol{\theta}_0|_2 \leq c_n}|\mathcal{M}_{\boldsymbol{\theta}}^*|\leq \ell_n\bigg)\rightarrow1$$ as $n\rightarrow\infty$, with some $c_n\rightarrow 0$ satisfying $\nu c_n^{-1} \rightarrow 0$. Proposition (ref) shows that $\hat{\boldsymbol{\theta}}_n$ is consistent to the true parameter $\boldsymbol{\theta}_0$, allowing $r$ growing exponentially with the sample size $n$.
Proposition (ref) establishes the consistency of the PEL estimator with diverging \(r\), incorporating the impact of the penalty function. In particular, the convergence rate of \(\hat{\boldsymbol{\theta}}_n\) is \(\nu\), provided that the tuning parameter \(\nu\) in (ref) satisfies \(\nu \gg \ell_n n^{-1/2} (\log r)^{1/2}\). As a result, the convergence rate of \(\hat{\boldsymbol{\theta}}_n\) is slower than \(n^{-1/2}\), which can be viewed as the price paid for using the penalty in handling exponentially growing dimensionality \(r\).
Recall $\rho(t;\nu)=\nu^{-1}P_\nu(t)$. For $P_\nu(\cdot)\in\mathscr{P}$ with $\mathscr{P}$ defined as (ref), since $\rho'(0^{+};\nu)$ is independent of $\nu$, we write it as $\rho'(0^{+})$ for simplicity. Let $\mathcal{R}_n={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$ for the Lagrange multiplier $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)=(\hat\lambda_1,\ldots,\hat\lambda_r)^{\mathrm{\scriptscriptstyle \top} }=\arg\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_n(\hat{\boldsymbol{\theta}}_n)}f_n(\boldsymbol{\lambda};\hat{\boldsymbol{\theta}}_n)$ with $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ defined as (ref). Then $\hat{\boldsymbol{\theta}}_n$ and $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)$ satisfy the score equation:
where $\hat{\boldsymbol{\eta}}=(\hat{\eta}_{1},\ldots ,\hat{\eta}_{r})^{\mathrm{\scriptscriptstyle \top} }$ with $\hat{\eta}_j=\nu \rho' (|\hat{\lambda}_j|;\nu) {\mbox{\rm sgn}}(\hat{\lambda}_j)$ for $\hat{\lambda}_j\neq0$ and $\hat{\eta}_j \in [-\nu \rho'(0^{+}),\nu \rho'(0^{+})]$ for $\hat{\lambda}_j=0$. Here, an effective drastic dimension reduction is achieved with the associated sparse \(\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\). The use of the penalty function \(P_\nu(\cdot)\) leads to \(\hat{\boldsymbol{\eta}}\) in (ref), an extra term compared to that of the conventional EL. While \(P_\nu(\cdot)\) ensures the consistency of \(\hat{\boldsymbol{\theta}}_n\) as shown in Proposition (ref), as we will show in Theorem (ref) later, \(\hat{\boldsymbol{\eta}}\) leads to a bias of the PEL estimator \(\hat{\boldsymbol{\theta}}_n\).
We further remark that while penalizing the Lagrange multiplier in our PEL does effectively achieve the selection of moments, its properties in terms of the validity of the selected moments remain an interesting research question. On one hand, it is reasonable to expect that under appropriate conditions and with a suitably chosen tuning parameter, our PEL may correctly select the set of valid moments. On the other hand, the major challenge lies in the ambiguity of defining valid moments when the corresponding moment functions are evaluated at broad candidate values of the model parameters rather than the truth. This consideration opens the door to a research question of its own interest in the context of moment selection that we are interested in investigating in our future research.
To study the asymptotic distribution of $\hat{\boldsymbol{\theta}}_n$, we need the following regularity conditions.
Discussion of Conditions (ref) and (ref) are given in Section (ref) of the supplementary material. Write $\widehat{{\mathbf V}}_{{ \mathcal{R}_n}}(\hat{\boldsymbol{\theta}}_n)=\mathbb{E}_n\{{\mathbf g}_{{ \mathcal{R}_n}} ( {\mathbf x}_{i}; \hat{\boldsymbol{\theta}}_n)^{\otimes2}\} $ and $\widehat{\boldsymbol{\Gamma}}_{{ \mathcal{R}_n}}(\hat{\boldsymbol{\theta}}_n) =\mathbb{E}_n\{ \nabla_{\boldsymbol{\theta}} {\mathbf g}_{{ \mathcal{R}_n}}({\mathbf x}_{i}; \hat{\boldsymbol{\theta}}_n)\}$. Define
where $\hat{\boldsymbol{\eta}}=(\hat{\eta}_{1},\ldots ,\hat{\eta}_{r})^{\mathrm{\scriptscriptstyle \top} }$ is specified in (ref). We assume $(r,\ell_n,\nu)$ satisfy the following restrictions:
The asymptotic distribution of $\hat{\boldsymbol{\theta}}_n$ is stated in Theorem (ref), where the bias term $\hat{\boldsymbol{\psi}}_{ \mathcal{R}_n}$ comes from the penalty function $P_{\nu}(\cdot)$ imposed on the Lagrange multiplier $\boldsymbol{\lambda}$ in (ref).
Here, the estimated bias \(\hat{\boldsymbol{\psi}}_{ \mathcal{R}_n}\) can be easily calculated based on (ref). Theorem (ref) indicates that, upon correcting the bias by subtracting it from $\hat{\boldsymbol{\theta}}_n$, the resulting estimator $\hat{\boldsymbol{\theta}}_n-\hat{\boldsymbol{\psi}}_{ \mathcal{R}_n}$ is \(n^{1/2}\)-consistent and asymptotically normal.
For the proposed BPEL, we establish the Bernstein-von Mises theorem for the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$, as defined in (ref). Furthermore, we provide theoretical assurances for the performance of Algorithms (ref) and (ref) in Section (ref).
For any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$, write $\mathcal{R}(\boldsymbol{\theta})={\rm supp}\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})\}$ with $$\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})=\{\hat{\lambda}_1(\boldsymbol{\theta}),\ldots,\hat{\lambda}_r(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }=\arg\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_n(\boldsymbol{\theta})}f_n(\boldsymbol{\lambda};\boldsymbol{\theta})\,,$$ where $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ is defined as (ref). Then $\boldsymbol{\theta}$ and $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})$ satisfy the score equation:
where $\hat{\boldsymbol{\eta}}(\boldsymbol{\theta})=\{\hat{\eta}_{1}(\boldsymbol{\theta}),\ldots ,\hat{\eta}_{r}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$ with $\hat{\eta}_j(\boldsymbol{\theta})=\nu \rho' \{|\hat{\lambda}_j(\boldsymbol{\theta})|;\nu\} {\mbox{\rm sgn}}\{\hat{\lambda}_j(\boldsymbol{\theta})\}$ for $\hat{\lambda}_j(\boldsymbol{\theta}) \neq 0$ and $\hat{\eta}_j(\boldsymbol{\theta}) \in [-\nu \rho'(0^{+}),\nu \rho'(0^{+})]$ for $\hat{\lambda}_j(\boldsymbol{\theta})=0$. By the definition of the PEL estimator $\hat{\boldsymbol{\theta}}_n$, we have $f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\} \geq f_n\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n);\hat{\boldsymbol{\theta}}_n\}$ for any $\boldsymbol{\theta}\in\boldsymbol{\Theta}$. To investigate the asymptotic properties of the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ defined as (ref), we need to first study the asymptotic behavior of $ \aleph_n(\boldsymbol{\theta})=f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}-f_n\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n);\hat{\boldsymbol{\theta}}_n\}$ for $\boldsymbol{\theta}\in\boldsymbol{\Theta}$. Given $\alpha_n=n^{-1/2}(\log r)^{1/2}$ and some $\beta_n$ satisfying $\ell_n^{1/2} \nu \ll \beta_{n} \ll \min\{ \ell_n^{-1}n^{-1/\gamma}, \nu^{2/3}\ell_n^{-2/3} n^{-1/(3\gamma)}\}$, we split the whole parameter space $\boldsymbol{\Theta}$ into three regions: $\mathcal{C}_1=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2\leq \alpha_n\}$, $\mathcal{C}_2=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:\alpha_n<|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2\leq \beta_n\}$ and $\mathcal{C}_3=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2>\beta_n\}$. Proposition (ref) in the supplementary material shows that the asymptotic behavior of $\aleph_n(\boldsymbol{\theta})$ for $\boldsymbol{\theta}$ in these three regions are different.
Investigating the asymptotic behavior of $\aleph_n(\boldsymbol{\theta})$ calls some new technical arguments. Write
When $r$ is a fixed constant, we know $2n\tilde{f}_n\{\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$ is the conventional log-EL ratio in the literature. The asymptotic behavior of $2n\tilde{f}_n\{\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$ depends on the magnitude of $\mathbb{E}\{{\mathbf g}({\mathbf x}_i;\boldsymbol{\theta})\}$. More specifically, under some mild conditions, it holds that (i) $2n\tilde{f}_n\{\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$ is asymptotically chi-square distributed with degree of freedom $r$ if $|\mathbb{E}\{{\mathbf g}({\mathbf x}_i;\boldsymbol{\theta})\}|_2\ll n^{-1/2}$, (ii) $2n\tilde{f}_n\{\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$ converges to a noncentral chi-square distribution if $|\mathbb{E}\{{\mathbf g}({\mathbf x}_i;\boldsymbol{\theta})\}|_2\asymp n^{-1/2}$, and (iii) $2n\tilde{f}_n\{\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$ diverges to $\infty$ in probability if $|\mathbb{E}\{{\mathbf g}({\mathbf x}_i;\boldsymbol{\theta})\}|_2\gg n^{-1/2}$. See, for example, Proposition 1 and Theorem 1 of Changetal_2013_AOS for such results with $r=1$. In comparison to $\tilde{f}_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ defined in (ref), $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ involved in $\aleph_n(\boldsymbol{\theta})$ includes a penalty term imposed on the Lagrange multiplier $\boldsymbol{\lambda}$. This makes the standard technique for analyzing the conventional log-EL ratio inapplicable. To further establish the Bernstein-von Mises theorem for the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ defined as (ref), we assume the following regularity conditions.
Discussion of Conditions (ref) and (ref) are given in Section (ref) of the supplementary material. Let $\Pi^{\dag}_n(\cdot)$ be the measure which admits the posterior distribution $\pi^{\dag}(\cdot\,|\,\mathcal{X}_n)$. Denote by $\mathcal{N}_{\boldsymbol{\mu},\boldsymbol{\Sigma}}(\cdot)$ the Gaussian measure with mean $\boldsymbol{\mu}$ and covariance matrix $\boldsymbol{\Sigma}$. To establish the Bernstein-von Mises theorem for the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ as in Theorem (ref), we need to assume $(r,\ell_n,\nu)$ satisfy the following restrictions:
Theorem (ref) shows that $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ has a Gaussian limiting distribution and it concentrates on a $n^{-1/2}$-ball centered at the PEL estimator $\hat{\boldsymbol{\theta}}_n$ of interest, which indicates that $\hat{\boldsymbol{\theta}}_n$ can be approximated by the mean of the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$. More specifically, as shown in Corollary (ref), the approximation error is of order smaller than $n^{-1/2}$.
Theorems (ref) and (ref) state the theoretical guarantees for Algorithms (ref) and (ref), respectively.
In this paper, we explore BPEL and demonstrate its promising performance using MCMC sampling as a competitive alternative to optimization in addressing EL problems. This framework has the potential for further advancements in several areas. To maintain focus and avoid digressions, we have confined our study to fixed-dimensional model parameters and exponentially growing moment conditions. However, there is significant interest in extending this approach to tackle variable and model selection using BPEL, which could accommodate high-dimensional sparse model parameters and potentially a continuum of moment conditions, as considered in Chau2017. Incorporating specific priors in the context of concrete studies, particularly in high-dimensional problems, is another area of interest. Research in this direction presents additional challenges, especially in selecting appropriate priors, developing efficient sampling schemes, and conducting associated analyses.
In the broader context of Bayesian methodology, approximate Bayesian computation (ABC) and Bayesian synthetic likelihood (BSL) are two competitive methods for handling situations where the likelihood is difficult to evaluate or intractable. ABC and BSL have been extensively compared in the literature. We demonstrate that the rationale of ABC integrates well with our BPEL method, achieving both accuracy and computational efficiency. Our Algorithm 2, inspired by ABC, uses importance weights for samples drawn from an alternative distribution to address challenging sampling situations. Empirical evidence shows promising performance, particularly in difficult cases. BSL leverages the limiting distribution, such as the normal distribution, to handle intractable probability distributions, with the advantage of easy sampling from the normal distribution. We view our BPEL as a compelling alternative to BSL: EL uses a multinomial likelihood that incorporates model information without requiring a fully specified parametric model, making it a competitive option when the full likelihood is intractable.
Furthermore, we foresee the use of more sophisticated sampling schemes in conjunction with PEL as highly valuable for addressing complex problems with specific considerations. Examples include the Hamiltonian MCMC method examined in Chaudhuri2017 and the variational Bayesian approach explored in Yu2023. These avenues of research are part of our plans for future projects.
\setcounter{equation}0 \setcounter{section}0 \setcounter{table}0 \setcounter{figure}0
\numberwithin{equation}{section} \setcounter{page}{1} \rhead{\bfseriesS\arabic{page}} \lhead{ SUPPLEMENTARY MATERIAL}
In the sequel, we use the abbreviations “w.p.a.1” and “w.r.t” to denote, respectively, “with probability approaching one” and “with respect to”. Let $C$, $\bar{C}$ and $\tilde{C}$ be generic positive finite constants that may be different in different uses. Let $\lfloor a \rfloor$ represent the largest integer not greater than $a \in {\mathbb{R}}$. For any positive integer $q$, we write $[q]=\{1,\ldots,q\}$. Denote by $I(\cdot)$ the indicator function. Let ${\rm tr}({\mathbf A})$ be the trace of a $q \times q$ matrix ${\mathbf A}=(a_{i,j})_{q \times q}$. For a $q \times q$ symmetric matrix ${\mathbf A}$, denote by $\lambda_{\min}({\mathbf A})$ and $\lambda_{\max}({\mathbf A})$ the smallest and largest eigenvalues of ${\mathbf A}$, respectively. For a $q_1 \times q_2$ matrix ${\mathbf B}=(b_{i,j})_{q_1 \times q_2}$, let $\|{\mathbf B}\|_2=\lambda^{1/2}_{\max}({\mathbf B}^{ \otimes2})$ be the spectral norm with ${\mathbf B}^{\otimes2}={\mathbf B} {\mathbf B}^{\mathrm{\scriptscriptstyle \top} }$. Specifically, if $q_2=1$, we use $|{\mathbf B}|_{\infty}=\max_{i \in [q_1]} |b_{i,1}|$, $|{\mathbf B}|_1=\sum_{i=1}^{q_1}|b_{i,1}|$ and $|{\mathbf B}|_2=(\sum_{i=1}^{q_1}b^2_{i,1})^{1/2}$ to denote the $L_{\infty}$-norm, $L_1$-norm and $L_2$-norm of the $q_1$-dimensional vector ${\mathbf B}$, respectively. Given index sets $\mathcal{S}_1 \subset [q_1]$ and $\mathcal{S}_2 \subset [q_2]$, denote by $[{\mathbf B}]_{\mathcal{S}_1, \mathcal{S}_2}$ the $|\mathcal{S}_1| \times |\mathcal{S}_2|$ matrix that is obtained by extracting the rows of a $q_1 \times q_2$ matrix ${\mathbf B}$ indexed by $\mathcal{S}_1$ and columns indexed by $\mathcal{S}_2$. For simplicity and when no confusion arises, we use the notation ${\mathbf g}_i(\boldsymbol{\theta})=\{g_{i,1}(\boldsymbol{\theta}),\ldots,g_{i,r}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$ as the equivalence to ${\mathbf g}({\mathbf x}_i;\boldsymbol{\theta})$, and denote by $\mathbb{E}_n(\cdot)=n^{-1}\sum_{i=1}^{n}\cdot$. Let $\bar{{\mathbf g}}(\boldsymbol{\theta})=\mathbb{E}_n\{{\mathbf g}_i(\boldsymbol{\theta})\}$, and write its $j$-th component as $\bar{g}_j(\boldsymbol{\theta})=\mathbb{E}_n\{g_{i,j}(\boldsymbol{\theta})\} $. Denote by $\nabla^2_{\boldsymbol{\theta}} g_{i,j} ( \boldsymbol{\theta})=\{\partial^2 g_{i,j}(\boldsymbol{\theta})/\partial\theta_{k_1}\partial\theta_{k_2}\}_{k_1,k_2\in[p]}$, a $p\times p$ matrix, the second-order derivative of $g_{i,j} (\boldsymbol{\theta})$ with respect to $\boldsymbol{\theta}$. Let $\widehat{\boldsymbol{\Gamma}}(\boldsymbol{\theta}) =\nabla_{\boldsymbol{\theta}} \bar{{\mathbf g}}(\boldsymbol{\theta})$ and $\widehat{{\mathbf V}}(\boldsymbol{\theta})=\mathbb{E}_n\{{\mathbf g}_i(\boldsymbol{\theta})^{\otimes2} \}$. For a given set $\mathcal{F} \subset[r]$, we denote by ${\mathbf g}_{i,\mathcal{F}}(\boldsymbol{\theta})$ the subvector of ${\mathbf g}_i(\boldsymbol{\theta})$ collecting the components indexed by $\mathcal{F}$. Analogously, let $\bar{{\mathbf g}}_{\mathcal{F}}(\boldsymbol{\theta})=\mathbb{E}_n\{{\mathbf g}_{i,\mathcal{F}}(\boldsymbol{\theta})\}$, $\widehat{\boldsymbol{\Gamma}}_{\mathcal{F}}(\boldsymbol{\theta}) =\nabla_{\boldsymbol{\theta}} \bar{{\mathbf g}}_{\mathcal{F}}(\boldsymbol{\theta})$ and $\widehat{{\mathbf V}}_{\mathcal{F}}(\boldsymbol{\theta})=\mathbb{E}_n\{{\mathbf g}_{i,\mathcal{F}}(\boldsymbol{\theta})^{\otimes2} \} $. We also write ${\mathbf a}_{\mathcal{F}}$ as the corresponding subvector of vector ${\mathbf a}$. Recall $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})=\mathbb{E}_n[\log \{ 1+\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_i (\boldsymbol{\theta})\}]-\sum_{j=1}^{r} P_{\nu}(|\lambda_j|)$ and $\hat\boldsymbol{\lambda}(\boldsymbol{\theta})=\{\hat\lambda_1(\boldsymbol{\theta}),\ldots,\hat\lambda_r(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }=\arg\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})} f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$. Write $\mathcal{R}(\boldsymbol{\theta})={\rm supp}\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})\}$, $\mathcal{R}_n={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$, and $\mathcal{M}_{\boldsymbol{\theta}}^*=\{j\in [r]: |\bar{g}_j(\boldsymbol{\theta})| \geq C_*\nu \rho'(0^{+}) \}$ for some $C_*\in(0,1)$. Define $\mathcal{M}_{\boldsymbol{\theta}}(c)=\{ j\in[r]:|\bar{g}_j(\boldsymbol{\theta})| \geq c\nu\rho'(0^{+}) \}$ for $c \in (C_*,1)$. Recall $\alpha_n=n^{-1/2}(\log r)^{1/2}$.
In this section, we investigate the impact of the prior distribution $\pi_0(\boldsymbol{\theta})$ in (ref) on our proposed Bayesian penalized EL methods in estimating the true parameter $\boldsymbol{\theta}_0$. More specifically, we adopt the data generation process outlined in Section (ref) with the true parameter $\boldsymbol{\theta}_0=(0.5,0.5)^{\mathrm{\scriptscriptstyle \top} }$, and consider three choices for the prior:
For the $49$ initial points mentioned in Section (ref), we calculate the measure ${\rm MSE}_2$ defined in Section (ref) to evaluate the performance of these estimators. The results for M-H and MAMIS are derived based on the generated samples of size $3500$. Table (ref) summarizes the performance of our proposed methods with such selected three priors. In particular, we have observed that when the prior is specified “closer" to the truth, the resulting estimator has better performance in comparison to the one using a non-informative prior. Conversely, if a prior is specified “further away" from the truth, the performance of the resulting estimator deteriorates and becomes less competitive.
In this section, we further validate the efficacy of our proposed methods by conducting some additional simulation studies. For the simulation examples considered in Sections (ref) and (ref), we let all instrumental variables (IVs) $z_{i,j}$ be independently and identically distributed following the Student's $t$-distribution with three degrees of freedoms. For the $49$ initial points mentioned in Section (ref), we calculate the measure ${\rm MSE}_1$ defined in Section (ref) to evaluate the performance of these methods. The results for Algorithms (ref) and (ref) based on sample sizes of 1500, 2500, and 3500 are denoted by (M-H-1, M-H-2, M-H-3) and (MAMIS-1, MAMIS-2, MAMIS-3), respectively. The results for $n=120$ and $n=240$ are presented in Tables (ref) and (ref), respectively. Furthermore, we also compare the penalized empirical likelihood (EL) against the standard EL and the relaxed EL introduced by Shi2016_s for the Student's $t$ IVs. For the $49$ initial points mentioned in Section (ref), we calculate the measure ${\rm MSE}_2$ defined in Section (ref) to evaluate the performance of these estimators. The results are presented in Table (ref), where the results for M-H and MAMIS are derived based on the generated samples of size $3500$.
Overall, these simulation results for the Student's $t$ IVs align closely with those listed in Sections (ref) and (ref). These findings further validate the robustness and effectiveness of our proposed methods.
In this section, we conduct several numerical simulations to examine the performance of the normal approximation stated in Theorem (ref) to the posterior distribution $\pi^{\dag}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ defined as (ref) in finite samples. More specifically, we adopt the data generation process outlined in Section (ref) with linear link function $\hbar(v)=v$. As described in Section (ref), we identify the true global minima $\hat{\boldsymbol{\theta}}_n$ defined as (ref) through exhaustive search. Subsequently, we calculate its asymptotic covariance matrix $n^{-1}\widehat{{\mathbf H}}^{-1}_{\mathcal{R}_n}$ with $\widehat{{\mathbf H}}_{\mathcal{R}_n}$ defined as (ref). We then generate $5000$ samples, respectively, from the Gaussian distribution $\mathcal{N}(\hat{\boldsymbol{\theta}}_n, n^{-1}\widehat{{\mathbf H}}^{-1}_{\mathcal{R}_n})$ and the posterior distribution $\pi^{\dag}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ defined as (ref). To generate samples from $\mathcal{N}(\hat{\boldsymbol{\theta}}_n, n^{-1}\widehat{{\mathbf H}}^{-1}_{\mathcal{R}_n})$, we use the function {\tt mvrnorm} in the R-package {\tt MASS}. To generate samples from $\pi^{\dag}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$, we use Algorithm (ref) with the burn-in period of $1000$ iterations. Based on these samples, we compute the Wasserstein distance between the two distributions $\mathcal{N}(\hat{\boldsymbol{\theta}}_n, n^{-1}\widehat{{\mathbf H}}^{-1}_{\mathcal{R}_n})$ and $\pi^{\dag}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ using the function {\tt wasserstein} in the R-package {\tt transport}. Figure (ref) below illustrates the average Wasserstein distance between the two distributions across different sample sizes $n$ under $500$ replications. It can be observed that, the two distributions exhibit a relatively large differences at smaller sample sizes, with this distance diminishing notably as the sample size increases.
We further validate the efficacy of our proposed methods in approximating the true global minima $\hat{\boldsymbol{\theta}}_n$ defined as (ref). Table (ref) below presents the measure ${\rm MSE} = \frac{1}{500}\sum_{k=1}^{500} |\check{\boldsymbol{\theta}}_{k} - \hat{\boldsymbol{\theta}}_n|_2^2$ across different sample sizes $n$, where $\check{\boldsymbol{\theta}}_{k}$ is the mean of the $5000$ samples drawn from the posterior distribution $\pi^{\dag}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ in the $k$-th replication. It is evident that with small sample size $n$, although the normal approximation stated in Theorem (ref) to the posterior distribution $\pi^{\dag}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ may not work very well, our method can still effectively approximate the true global minimum $\hat{\boldsymbol{\theta}}_n$. As the sample size $n$ increases, the accuracy of the approximation exhibits a substantial improvement.
In this section, we conduct some additional numerical studies to further compare the performance of posteriors derived by different methods. Assume the observations ${\mathbf x}_1,\ldots,{\mathbf x}_n$ are drawn independently from the distribution $F(\boldsymbol{\theta}_0)$ with some unknown parameter $\boldsymbol{\theta}_0$. The likelihood function admits the form $L(\boldsymbol{\theta})=\prod_{i=1}^n f({\mathbf x}_i;\boldsymbol{\theta})$ where $f(\cdot\,;\boldsymbol{\theta})$ is the density function of $F_n(\boldsymbol{\theta})$. Write $\mathcal{X}_n=\{{\mathbf x}_1,\ldots,{\mathbf x}_n\}$. Let $\pi_0(\cdot)$ represent a prior distribution for $\boldsymbol{\theta}$. Then the traditional posterior is given by $\pi^{\rm L}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)\propto \pi_{0}(\boldsymbol{\theta}) \times L(\boldsymbol{\theta})$. To estimate the unknown parameter $\boldsymbol{\theta}_0$, we can also identify it by $\mathbb{E}\{\mathbf{g}({\mathbf x}_i;\boldsymbol{\theta}_0)\}={\mathbf 0}$ with some $r$-dimensional estimating function $\mathbf{g}(\cdot\,;\cdot)$. For given estimating function $\mathbf{g}(\cdot\,; \cdot)$, we can define the empirical likelihood ${\rm EL}(\boldsymbol{\theta})$ and penalized empirical likelihood ${\rm PEL}_\nu(\boldsymbol{\theta})$, respectively, as (ref) and (ref) in the main paper. Hence, the associated EL-based posterior distribution $\pi^{\rm EL}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ and PEL-based posterior distribution $\pi^{\dag}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ satisfy $\pi^{\rm EL}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)\propto \pi_{0}(\boldsymbol{\theta}) \times {\rm EL}(\boldsymbol{\theta}) $ and $\pi^{\dag}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)\propto \pi_{0}(\boldsymbol{\theta}) \times {\rm PEL}_\nu(\boldsymbol{\theta})$. We select $\pi_0(\cdot)$ as the improper uniform prior and compare these three posteriors via the following two models:
We generate $5000$ samples from each of these posterior distributions via the M-H algorithm with the burn-in period of $2000$ iterations. Based on these samples, we compute the Wasserstein distances between $\pi^{\rm L}(\theta\,|\,\mathcal{X}_n)$ and $\pi^{\rm EL}(\theta\,|\,\mathcal{X}_n)$, as well as between $\pi^{\rm L}(\theta\,|\,\mathcal{X}_n)$ and $\pi^{\dag}(\theta\,|\,\mathcal{X}_n)$, using the function {\tt wasserstein} in the R-package {\tt transport}. Figure (ref) below illustrates the average Wasserstein distances between these posterior distributions across different sample sizes $n$ under $500$ replications. Figures (ref) and (ref) show the density functions of $\pi^{\rm L}(\theta\,|\,\mathcal{X}_n)$, $\pi^{\rm EL}(\theta\,|\,\mathcal{X}_n)$ and $\pi^{\dag}(\theta\,|\,\mathcal{X}_n)$ across different sample sizes $n$ in one replication.
Overall, the numerical results indicate that the posterior distributions constructed the (un-penalized) empirical likelihood without the penalty on the Lagrange multiplier exhibit significant discrepancies from those derived from the likelihood function. However, this disparity can be markedly reduced through the introduction of a penalty term on the Lagrange multipliers in the empirical likelihood.
In this section, we compare the performance of our proposed methods with the approximate Bayesian computation (ABC) and Bayesian synthetic likelihood (BSL) methods, implemented as described below.
We conduct the comparisons using the same three models as described in Section (ref). Both the ABC and BSL methods require the selection of summary statistics. In our numerical experiments, for demonstration purposes, we chose sufficient statistics for the parameters of interest, thereby favoring the ABC and BSL methods. Specifically, for Model \({\rm I}\), we selected the sample standard deviation as the summary statistic. For Models \({\rm II}\) and \({\rm III}\), we selected the maximum likelihood estimator as the summary statistic.
When implementing {\tt abc}, we set the tolerance levels to \(0.05\), \(0.1\), and \(0.2\) to obtain \(5000\) valid approximate samples of the traditional posterior distributions for the three models. These tolerance levels correspond to \(100{,}000\), \(50{,}000\), and \(25{,}000\) MCMC iterations, respectively. The results are labeled as ({\tt abc}-0.05, {\tt abc}-0.1, {\tt abc}-0.2). For {\tt bsl}, we ran the MCMC sampler for \(7000\) iterations, discarding the first \(2000\) iterations for burn-in.
To evaluate the performance of {\tt abc} and {\tt bsl}, we used the {\tt wasserstein} function from the R package {\tt transport} to compute the Wasserstein distances between the approximate samples and those obtained directly via the Metropolis-Hastings (M-H) algorithm from the traditional posterior distribution. For our PEL method, we report results for the case with \(r=50\) and also evaluate the corresponding Wasserstein distances relative to those generated by the M-H algorithm. Figure (ref) below presents the results, showing the average Wasserstein distances calculated for different sample sizes \(n\) under \(500\) replications.
Overall, the numerical results indicate that our proposed method, along with {\tt abc} and {\tt bsl}, effectively approximates the posterior distribution, with the accuracy of the approximation improving as the sample size $n$ increases. For smaller sample sizes, {\tt abc} and {\tt bsl} exhibit better accuracy. However, as the sample size $n$ grows, our proposed method demonstrates substantial improvements, achieving satisfactory approximation accuracy. Additionally, it is notable that the ABC method often requires significantly higher computational costs to achieve comparable approximation accuracy.
Here, we also note that the settings favor the ABC and BSL methods in the choice of summary statistics, yet our method performs very competitively. Figure (ref)(b) presents results from a slightly modified setting of Model I, where the data generation process is changed from a normal distribution to a Student's $t$-distribution with $10$ degrees of freedom, and we are still interested in estimating the standard deviation parameter. In this case, the summary statistic based on the sample standard deviation for ABC and BSL is no longer sufficient. For side-by-side comparisons, the corresponding case with results from the normal distribution is shown in Figure (ref)(a). In this scenario, the performance of the PEL-based method is superior, demonstrating the compelling performance of our approach, owing to the merits of using empirical likelihood.
Conditions (ref) and (ref) are commonly used assumptions in the literature. Condition (ref) is the identification condition for the unknown true parameter $\boldsymbol{\theta}_0$. A similar condition can be found in Shi2016 and Changetal_2018. Condition (ref)(b) requires the covariance matrix of ${\mathbf g}({\mathbf x}_i;\boldsymbol{\theta}_0)$ behaves reasonably well. Conditions (ref)(a) and (ref)(c) impose the moments requirements on each estimating function $g_j(\cdot\,;\cdot)$ and its derivatives. If there exist functions $B_l(\cdot)$ with $\mathbb{E}\{B_l({\mathbf x}_i)\}<\infty$, $l=1,2,3$, such that $|g_j({\mathbf x};\boldsymbol{\theta})|^\gamma\leq B_1({\mathbf x})$, $|\partial g_j({\mathbf x};\boldsymbol{\theta})/\partial\theta_k|^2\leq B_2({\mathbf x})$ and $|\partial^2g_j({\mathbf x};\boldsymbol{\theta})/\partial\theta_{k_1}\partial\theta_{k_2}|^2\leq B_3({\mathbf x})$ for any $j\in[r]$ and $\boldsymbol{\theta}\in\boldsymbol{\Theta}$, then the second requirement in Condition (ref)(a) and the two requirements in Condition (ref)(c) hold automatically. More generally, if there exist functions $B_{l,j}(\cdot)$ such that $|g_j({\mathbf x};\boldsymbol{\theta})|^\gamma\leq B_{1,j}({\mathbf x})$, $|\partial g_j({\mathbf x};\boldsymbol{\theta})/\partial\theta_k|^2\leq B_{2,j}({\mathbf x})$ and $|\partial^2g_j({\mathbf x};\boldsymbol{\theta})/\partial\theta_{k_1}\partial\theta_{k_2}|^2\leq B_{3,j}({\mathbf x})$ for any $j\in[r]$ and $\boldsymbol{\theta}\in\boldsymbol{\Theta}$, and $\max_{j\in[r]}\mathbb{E}\{B_{l,j}^m({\mathbf x}_i)\}\leq Km!H^{m-2}$ for any $m\geq 2$ and $l=1,2,3$ with two universal constants $K, H>0$, it follows from Theorem 2.8 of Petrov_1995 that the second requirement in Condition (ref)(a) and the two requirements in Condition (ref)(c) hold automatically provided $\log(rp)=o(n)$. In fact, the order $O_{\rm p}(1)$ in Conditions (ref)(a) and (ref)(c) can be replaced by $O_{\rm p}(\varpi_n)$ with some diverging sequence $\varpi_n$, and our main results remain valid. We use $O_{\rm p}(1)$ here for ease of presentation. To establish the consistency of the penalized empirical likelihood estimator $\hat{\boldsymbol{\theta}}_n$, Conditions (ref), (ref)(a) and (ref)(b) are needed. Condition (ref)(c) is needed for establishing the asymptotic normality of $\hat{\boldsymbol{\theta}}_n$.
Condition (ref) is standard in the literature. Due to the penalty imposed on the Lagrange multiplier $\boldsymbol{\lambda}$ involved in the optimization (ref), the standard theoretical analysis of empirical likelihood cannot be applied here. Condition (ref)(a) is a technical assumption used to derive the convergence rate of the Lagrange multiplier $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)=\arg\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_n(\hat{\boldsymbol{\theta}}_n)}f_n(\boldsymbol{\lambda};\hat{\boldsymbol{\theta}}_n)$ associated with $\hat{\boldsymbol{\theta}}_n$; see the proof of Lemma (ref) in the supplementary material for details. Condition (ref)(b) requires that each $\hat{\eta}_j$ $(j\in \mathcal{R}_n^{\rm c})$ lies in the interior of $[-\nu \rho'(0^{+}),\nu \rho'(0^{+})]$ with probability approaching one, which is realistic in practice. If the distribution function of the random variable $\hat{\eta}_j$ is continuous at $\pm \nu\rho'(0^+)$, we then have $\mathbb{P}\{|\hat{\eta}_j|=\nu\rho'(0^+)\}=0$. Condition (ref) makes sure that $\hat\boldsymbol{\lambda}(\boldsymbol{\theta})=\arg\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_n(\boldsymbol{\theta})}f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ is continuously differentiable at $\hat{\boldsymbol{\theta}}_n$ with probability approaching one; see Lemma (ref) in Section (ref).
Condition (ref)(a) guarantees that the Lagrange multiplier $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})$ for $\boldsymbol{\theta} \in \mathcal{C}_1$ satisfies two properties: (i) $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})$ is continuously differentiable on $\mathcal{C}_1$ with probability approaching one, and (ii) $\mathcal{R}(\boldsymbol{\theta}) = \mathcal{R}(\hat{\boldsymbol{\theta}}_n)$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ with probability approaching one; see Lemmas (ref) and (ref) in Section (ref). When $\boldsymbol{\theta}\notin\mathcal{C}_1$, characterizing the asymptotic property of $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})$ is quite challenging. Due to $f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}\geq f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ for any $\boldsymbol{\lambda}\in\hat{\Lambda}_n(\boldsymbol{\theta})$, a feasible strategy to construct the lower bound of $f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$ with $\boldsymbol{\theta}\notin\mathcal{C}_1$ is to find a specific $\boldsymbol{\lambda}_*(\boldsymbol{\theta})\in\hat{\Lambda}_n(\boldsymbol{\theta})$ and then derive the lower bound of $f_n\{\boldsymbol{\lambda}_*(\boldsymbol{\theta});\boldsymbol{\theta}\}$ directly, where the asymptotic behavior of $\boldsymbol{\lambda}_*(\boldsymbol{\theta})$ can be well characterized even if $\boldsymbol{\theta}\notin\mathcal{C}_1$. Such strategy has been also used in Changetal2013, Changetal_2016_AOS to study the diverging rate of the conventional empirical likelihood ratio evaluated at a value not near the truth. In current setting, Condition (ref)(b) is applied to derive the lower bound of $f_n\{\boldsymbol{\lambda}_*(\boldsymbol{\theta});\boldsymbol{\theta}\}$ for $\boldsymbol{\theta}\in\mathcal{C}_2$; see Section (ref) for details. Condition (ref)(c) says that the sample covariance matrix of the estimating function and the gradient of the estimating function should behave reasonably well, which will be used to obtain the lower bound of $f_n\{\boldsymbol{\lambda}_*(\boldsymbol{\theta});\boldsymbol{\theta}\}$ for $\boldsymbol{\theta}\in\mathcal{C}_3$. See Section (ref) for details. Condition (ref) is a standard assumption concerning the prior distribution.
Write $\mathcal{R}_{0}=\mathcal{R}(\boldsymbol{\theta}_0)$. Then
where $\hat{\Lambda}^{\dag}_n(\boldsymbol{\theta}_0)=\{\boldsymbol{\eta}=(\eta_{1},\ldots, \eta_{|\mathcal{R}_{0}|})^{\mathrm{\scriptscriptstyle \top} } \in {\mathbb{R}}^{|\mathcal{R}_{0}|}:\boldsymbol{\eta}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{R}_{0}} (\boldsymbol{\theta}_0) \in \mathcal{V}~\textrm{for any}~i\in[n]\}$ for some open interval $\mathcal{V}$ containing zero. Our first step is to show $\max_{\boldsymbol{\eta} \in\hat{\Lambda}^{\dag}_n(\boldsymbol{\theta}_0)} \mathbb{E}_n[\log \{ 1+\boldsymbol{\eta}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{R}_{0}} (\boldsymbol{\theta}_0)\}]=O_{\rm p}(\ell_n \alpha_n^2)$. To do this, we need the following two lemmas whose proofs are given in Sections (ref) and (ref), respectively.
Define $\tilde{\boldsymbol{\eta}}=\arg \max _{\boldsymbol{\eta} \in \hat{\Lambda}^{\dag}_n(\boldsymbol{\theta}_0)} A_n(\boldsymbol{\theta}_0,\boldsymbol{\eta})$ with $A_n(\boldsymbol{\theta},\boldsymbol{\eta})=\mathbb{E}_n[\log \{ 1+\boldsymbol{\eta}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{R}_{0}} (\boldsymbol{\theta}) \}]$. By Lemma (ref), we have $|\mathcal{R}_{0}| \leq \ell_n$ w.p.a.1. Pick $\delta_n$ satisfying $\delta_n=o(\ell_n^{-1/2} n^{-1/\gamma})$ and $\ell_n^{1/2} \alpha_n=o(\delta_n) $ for $\gamma$ defined in Condition (ref)(a), which can be guaranteed by $\ell_n\alpha_n=o(n^{-1/\gamma})$. Let $\Lambda_{0}=\{\boldsymbol{\eta} \in {\mathbb{R}}^{|\mathcal{R}_{0}|}:|\boldsymbol{\eta}|_2\leq \delta_n \}$ and $\check{\boldsymbol{\eta}}=\arg \max _{\boldsymbol{\eta} \in \Lambda_{0}} A_n(\boldsymbol{\theta}_0,\boldsymbol{\eta})$. Condition (ref)(a) implies $\max _{i \in [n], \boldsymbol{\eta} \in \Lambda_{0}} |\boldsymbol{\eta}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{R}_{0}} (\boldsymbol{\theta}_0)| = o_{\rm p}(1)$. Then, by the Taylor expansion, we have
for some $C\in(0,1)$. By Condition (ref)(b) and the same arguments for deriving Lemma \rm(ref), if $\log r=o(n^{1/3})$ and $\ell_n\alpha_n=o(1)$, we have $\lambda_{\min}\{\widehat{{\mathbf V}}_{\mathcal{M}_{\boldsymbol{\theta}_0}} (\boldsymbol{\theta}_0)\}$ is uniformly bounded away from zero w.p.a.1. Thus $0\leq |\check{\boldsymbol{\eta}}|_2 |\bar{{\mathbf g}}_{\mathcal{R}_{0}}(\boldsymbol{\theta}_0)|_2 - 4^{-1}K_3|\check{\boldsymbol{\eta}}|_2^2 \{1+o_{\rm p}(1)\}$ w.p.a.1, where $K_3$ is specified in Condition (ref)(b). By the moderate deviation of self-normalized sums JingShaoWang2003, we have $ |\bar{{\mathbf g}}(\boldsymbol{\theta}_0)|_{\infty} =O_{\rm p}(\alpha_n) $, which implies $|\bar{{\mathbf g}}_{\mathcal{R}_{0}}(\boldsymbol{\theta}_0)|_2= O_{\rm p}(\ell_n^{1/2} \alpha_n)$ and $|\check{\boldsymbol{\eta}}|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n)=o_{\rm p}(\delta_n)$. Hence, $\check{\boldsymbol{\eta}} \in {\rm int} (\Lambda_{0}) $ w.p.a.1. Since $\Lambda_{0} \subset \hat{\Lambda}^{\dag}_n(\boldsymbol{\theta}_0)$ w.p.a.1, we have $\tilde{\boldsymbol{\eta}}=\check{\boldsymbol{\eta}}$ w.p.a.1 by the concavity of $A_n(\boldsymbol{\theta}_0,\boldsymbol{\eta})$ w.r.t $\boldsymbol{\eta}$. Then $ \max_{\boldsymbol{\eta} \in\hat{\Lambda}^{\dag}_n(\boldsymbol{\theta}_0)} A_n(\boldsymbol{\theta}_0,\boldsymbol{\eta})=O_{\rm p}(\ell_n \alpha_n^2)$. Let $b_n^2=\ell_n \alpha_n^2$ and $F_n(\boldsymbol{\theta})=\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})}f_n(\boldsymbol{\lambda};\boldsymbol{\theta}) $ for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$. Due to $\hat{\boldsymbol{\theta}}_n=\arg \min _{\boldsymbol{\theta} \in \boldsymbol{\Theta}}F_n(\boldsymbol{\theta}) $, we have $F_n(\hat{\boldsymbol{\theta}}_n) \leq F_n(\boldsymbol{\theta}_0)=O_{\rm p}(b_n^2)$.
Our second step is to show that for any $\epsilon_n \rightarrow \infty$ satisfying $b_n^2\epsilon_n^2 n^{2/\gamma}=o(1)$, there exists a universal constant $K>0$ independent of $\boldsymbol{\theta}$ such that ${\mathbb{P}}\{F_n(\boldsymbol{\theta})>K b_n^2 \epsilon_n^2\} \rightarrow 1$ as $n \rightarrow \infty$ for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ satisfying $|\boldsymbol{\theta}-\boldsymbol{\theta}_0|_{\infty} > \epsilon_n \nu$. Thus $|\hat{\boldsymbol{\theta}}_n-\boldsymbol{\theta}_0|_\infty=O_{\rm p}(\epsilon_n\nu)$. Due to $b_n^2=o(n^{-2/\gamma})$, we can select arbitrary slowly diverging $\epsilon_n$. We then have $|\hat{\boldsymbol{\theta}}_n-\boldsymbol{\theta}_0|_\infty=O_{\rm p}(\nu)$ by a standard result from probability theory. For any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ satisfying $|\boldsymbol{\theta}-\boldsymbol{\theta}_0|_{\infty} > \epsilon_n \nu$, let $j_{0}=\arg \max _{j \in [r]}|{\mathbb{E}}\{g_{i,j}(\boldsymbol{\theta})\}|$ and $\mu_{j_{0}}={\mathbb{E}}\{g_{i,j_{0}}(\boldsymbol{\theta})\} $. Define $\tilde{\boldsymbol{\lambda}}=\tau b_n \epsilon_n {\mathbf e}_{j_{0}}$, where $\tau>0$ is a constant to be determined later and ${\mathbf e}_{j_{0}}$ is an $r$-dimensional vector with the $j_{0}$-th component being $1$ and other components being $0$. Without loss of generality, we assume $\mu_{j_{0}}>0$. Condition (ref)(a) implies $\max_{i\in[n]} |\tilde{\boldsymbol{\lambda}}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i}(\boldsymbol{\theta})|= O_{\rm p}(b_n \epsilon_n n^{1/\gamma})=o_{\rm p}(1)$. Then $\tilde{\boldsymbol{\lambda}} \in \hat{\Lambda}_n(\boldsymbol{\theta})$ w.p.a.1. Write $\tilde{\boldsymbol{\lambda}}=(\tilde{\lambda}_{1},\ldots, \tilde{\lambda}_{r})^{\mathrm{\scriptscriptstyle \top} }$. By the Taylor expansion, it holds w.p.a.1 that
for some $\bar{C}\in(0,1)$, which implies $$ {\mathbb{P}}\big\{F_n(\boldsymbol{\theta}) \leq Kb_n^2 \epsilon_n^2 \big\} \leq {\mathbb{P}}\big\{\bar{g}_{j_{0}}(\boldsymbol{\theta})-\mu_{j_{0}} \leq b_n \epsilon_n[ K\tau^{-1} + \tau \mathbb{E}_n\{g^2_{i,j_{0}}(\boldsymbol{\theta})\}]+C\nu-\mu_{j_{0}}\big\}+o(1)\,.$$ From Condition (ref)(a) and Jensen's inequality, there exists a universal constant $L>0$ independent of $\boldsymbol{\theta}$ such that ${\mathbb{P}}[\mathbb{E}_n\{g^2_{i,j_{0}}(\boldsymbol{\theta})\} >L] \rightarrow 0$. Taking $\tau=(KL^{-1})^{1/2}$, we obtain $ {\mathbb{P}}\{F_n(\boldsymbol{\theta}) \leq K b_n^2 \epsilon_n^2 \} \leq {\mathbb{P}} \{\bar{g}_{j_{0}}(\boldsymbol{\theta})-\mu_{j_{0}} \leq 2 b_n\epsilon_n (KL)^{1/2} +C\nu -\mu_{j_{0}}\}+o(1)$. By Condition (ref), $\mu_{j_{0}} \geq \Delta(\epsilon_n \nu) \geq K_{1} \epsilon_n \nu/2$ with $K_1$ defined in Condition (ref) for sufficiently large $n$. We select sufficiently small $K>0$. Due to $b_n=o(\nu)$, when $n$ is sufficiently large, $2 b_n \epsilon_n (KL)^{1/2}+C\nu-\mu_{j_{0}} \leq -\check{C}\mu_{j_{0}}$ for some $\check{C}\in(0,1)$. Hence, we have $ \sqrt{n}\{2 b_n\epsilon_n (KL)^{1/2}+C\nu-\mu_{j_{0}}\} \leq -\check{C}\sqrt{n} \mu_{j_{0}} \rightarrow -\infty$. By the Central Limit Theorem, $\sqrt{n}\{\bar{g}_{j_{0}}( \boldsymbol{\theta})- \mu_{j_{0}}\}\rightarrow\mathcal{N}(0, \sigma^2)$ in distribution for some $\sigma>0$, which implies that ${\mathbb{P}}\{F_n(\boldsymbol{\theta}) \leq Kb_n^2 \epsilon_n^2\} \rightarrow 0$ as $n \rightarrow \infty$ for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ satisfying $|\boldsymbol{\theta}-\boldsymbol{\theta}_0|_{\infty} > \epsilon_n\nu$. We complete the proof of Proposition (ref). $\Box$
Recall $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})=\arg \max _{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})} f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ and $\mathcal{R}_n={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$. We first present two lemmas whose proofs are given in Sections (ref) and (ref), respectively.
For simplicity, we write $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)$ as $\hat{\boldsymbol{\lambda}}=(\hat\lambda_1,\ldots,\hat\lambda_r)^{\mathrm{\scriptscriptstyle \top} }$. Then we have
where $\hat{\boldsymbol{\eta}}=(\hat{\eta}_{1},\ldots, \hat{\eta}_{r})^{\mathrm{\scriptscriptstyle \top} }$ with $\hat{\eta}_j=\nu \rho' (|\hat{\lambda}_j|;\nu) {\mbox{\rm sgn}}(\hat{\lambda}_j)$ for $\hat{\lambda}_j\neq0$ and $\hat{\eta}_j \in [-\nu \rho'(0^{+}),\nu \rho'(0^{+})]$ for $\hat{\lambda}_j=0$. By the Taylor expansion, we know that
for some $C\in(0,1)$, which implies $\hat{\boldsymbol{\lambda}}_{ \mathcal{R}_n}= {\mathbf A}^{-1}(\hat{\boldsymbol{\theta}}_n)\{\bar{{\mathbf g}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)- \hat{\boldsymbol{\eta}}_{ \mathcal{R}_n}\}$. Since $\hat{\boldsymbol{\theta}}_n=\arg\min_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$, we have ${\mathbf 0} =\nabla_{\boldsymbol{\theta}} f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}|_{\boldsymbol{\theta}=\hat{\boldsymbol{\theta}}_n}$. Notice that
By Lemma (ref) and (ref), we have $[\nabla_{\boldsymbol{\theta}} \hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)]_{\mathcal{R}_n^{\rm c},[p]}={\mathbf 0}$ w.p.a.1 and $\partial f_n(\hat{\boldsymbol{\lambda}};\hat{\boldsymbol{\theta}}_n)/\partial\boldsymbol{\lambda}_{{ \mathcal{R}_n}} = {\mathbf 0}$. Thus, it holds w.p.a.1 that
We then obtain $ {\mathbf 0} = {\mathbf B}(\hat{\boldsymbol{\theta}}_n)^{\mathrm{\scriptscriptstyle \top} } {\mathbf A}^{-1}(\hat{\boldsymbol{\theta}}_n) \{\bar{{\mathbf g}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)- \hat{\boldsymbol{\eta}}_{ \mathcal{R}_n}\}$. To derive the limiting distribution of $\hat{\boldsymbol{\theta}}_n$, we need the following lemmas whose proofs are given in Sections (ref), (ref) and (ref), respectively.
For any ${\mathbf t} \in {\mathbb{R}}^p$ with $|{\mathbf t}|_2=1$, let $\boldsymbol{\delta}=\widehat{{\mathbf H}}_{ \mathcal{R}_n}^{-1/2} {\mathbf t}$ and ${\mathbf U} = \widehat{{\mathbf V}}^{-1/2}_{ \mathcal{R}_n} (\hat{\boldsymbol{\theta}}_n) \widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)$. Then $\widehat{{\mathbf H}}_{{ \mathcal{R}_n}} = {\mathbf U}^{{\mathrm{\scriptscriptstyle \top} }, \otimes2}$ and $ |\widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n) \boldsymbol{\delta}|_2^2 \leq \lambda_{\max} \{\widehat{{\mathbf V}}_{ \mathcal{R}_n} (\hat{\boldsymbol{\theta}}_n)\} |{\mathbf U} ({\mathbf U}^{{\mathrm{\scriptscriptstyle \top} }, \otimes2})^{-1/2} {\mathbf t}|_2^2= \lambda_{\max} \{\widehat{{\mathbf V}}_{ \mathcal{R}_n} (\hat{\boldsymbol{\theta}}_n)\}$. By Condition (ref)(b) and Lemma (ref), $|\widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n) \boldsymbol{\delta}|_2=O_{\rm p}(1)$. Under Conditions (ref)(b) and (ref), Lemmas (ref) and (ref) imply $|\boldsymbol{\delta}|_2^2 \leq \lambda_{\max} \{\widehat{{\mathbf V}}_{ \mathcal{R}_n} (\hat{\boldsymbol{\theta}}_n)\} \lambda_{\min}^{-1}\{\widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)^{{\mathrm{\scriptscriptstyle \top} },\otimes2} \} = O_{\rm p}(1)$. By Lemma (ref), we have w.p.a.1 that $\mathcal{R}_n \subset \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})$ and ${\mbox{\rm sgn}}(\hat{\lambda}_j) ={\mbox{\rm sgn}}\{\bar{g}_j(\hat{\boldsymbol{\theta}}_n)\}$ for any $j \in \mathcal{R}_n$. Since $P_{\nu}(\cdot) \in \mathscr{P}$ has bounded second-order derivative around $0$, by Lemma (ref), it holds w.p.a.1 that
for some $c_j\in(0,1)$. As shown in the proof of Lemma (ref), $ |\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})}(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})}(\hat{\boldsymbol{\theta}}_n)\}|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n)$. Then $|\bar{{\mathbf g}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)-\hat{\boldsymbol{\eta}}_{ \mathcal{R}_n}|_2 =O_{\rm p}(\ell_n^{1/2}\alpha_n)$. Due to $ {\mathbf B}(\hat{\boldsymbol{\theta}}_n)^{\mathrm{\scriptscriptstyle \top} } {\mathbf A}^{-1}(\hat{\boldsymbol{\theta}}_n) \{\bar{{\mathbf g}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)- \hat{\boldsymbol{\eta}}_{ \mathcal{R}_n}\}={\mathbf 0}$, by the triangle inequality,
By Lemma (ref),
Hence, $\boldsymbol{\delta}^{\mathrm{\scriptscriptstyle \top} } \widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf V}}_{ \mathcal{R}_n}^{-1} (\hat{\boldsymbol{\theta}}_n) \{\bar{{\mathbf g}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)-\hat{\boldsymbol{\eta}}_{ \mathcal{R}_n} \}= O_{\rm p}(\ell_n^{3/2} n^{1/\gamma} \alpha_n^2)$. By the Taylor expansion, we have
where $\tilde{\boldsymbol{\theta}}$ is on the line joining $\boldsymbol{\theta}_0$ and $\hat{\boldsymbol{\theta}}_n$. Write $\hat{\boldsymbol{\theta}}_n=(\hat{\theta}_{n,1},\ldots,\hat{\theta}_{n,p})^{\mathrm{\scriptscriptstyle \top} }$, $\boldsymbol{\theta}_0=(\theta_{0,1},\ldots,\theta_{0,p})^{\mathrm{\scriptscriptstyle \top} }$ and $\tilde{\boldsymbol{\theta}}=(\tilde{\theta}_{1},\ldots,\tilde{\theta}_{p})^{\mathrm{\scriptscriptstyle \top} }$. By the Taylor expansion, Jensen's inequality and Cauchy-Schwarz inequality, it holds that
where $\dot{\boldsymbol{\theta}}^{(j,k)}$ lies on the jointing line between $\tilde{\boldsymbol{\theta}}$ and $\hat{\boldsymbol{\theta}}_n$. Recall $p$ is fixed. By Proposition (ref), $|\hat{\boldsymbol{\theta}}_n-\boldsymbol{\theta}_0|_2=O_{\rm p}(\nu)$. Together with Condition {\rm (ref)(c)}, we have $| \{\widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\tilde{\boldsymbol{\theta}})- \widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)\} (\hat{\boldsymbol{\theta}}_n-\boldsymbol{\theta}_0)|_2= O_{\rm p}(\ell_n^{1/2} \nu^2)$. Recall $\hat{\boldsymbol{\psi}}_{ \mathcal{R}_n}= \widehat{{\mathbf H}}_{ \mathcal{R}_n}^{-1} \widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf V}}^{-1}_{ \mathcal{R}_n} (\hat{\boldsymbol{\theta}}_n) \hat{\boldsymbol{\eta}}_{ \mathcal{R}_n}$ and $\boldsymbol{\delta}=\widehat{{\mathbf H}}_{ \mathcal{R}_n}^{-1/2} {\mathbf t}$. Then (ref) leads to
By Lemma (ref), we have $ n^{1/2} {\mathbf t}^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf H}}_{ \mathcal{R}_n}^{1/2}(\hat{\boldsymbol{\theta}}_n-\boldsymbol{\theta}_0-\hat{\boldsymbol{\psi}}_{ \mathcal{R}_n})\rightarrow \mathcal{N}(0,1)$ in distribution as $n \rightarrow \infty$. $\Box$
Assume $(r,\ell_n,\nu)$ satisfy the following restrictions:
To construct Theorem (ref), we need the following proposition whose proof is given in Section (ref).
Recall the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) \propto \pi_{0}(\boldsymbol{\theta})\times \exp [-n\log n-n f_n\{\hat\boldsymbol{\lambda}(\boldsymbol{\theta});\boldsymbol{\theta} \}]I(\boldsymbol{\theta} \in \boldsymbol{\Theta})$. For any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$, let $w_n(\boldsymbol{\theta})=-n\log n-nf_n\{\hat\boldsymbol{\lambda}(\boldsymbol{\theta}); \boldsymbol{\theta}\}$ and write ${\mathbf t} = n^{1/2}(\boldsymbol{\theta}-\hat\boldsymbol{\theta}_n)$. Define $\mathcal{T}_n = \{{\mathbf t} \in {\mathbb{R}}^p: {\mathbf t} = n^{1/2}(\boldsymbol{\theta}-\hat\boldsymbol{\theta}_n), \boldsymbol{\theta} \in \boldsymbol{\Theta}\}$. Denote by $\pi_{0,{\mathbf t}}(\cdot)$ and $\pi^{\dag}_{{\mathbf t}}(\cdot\,|\,\mathcal{X}_n)$ the prior and the posterior distributions of ${\mathbf t}$, respectively. Then, $\pi_{0,{\mathbf t}}({\mathbf t})=n^{-p/2}\pi_{0}(\hat{\boldsymbol{\theta}}_n+n^{-1/2}{\mathbf t})$ and
To prove Theorem (ref), it is equivalent to show
in probability. It follows from the triangle inequality that
Notice that $ {\rm I} \geq |C_n - (2\pi)^{p/2} \pi_{0}(\hat{\boldsymbol{\theta}}_n) |\widehat{{\mathbf H}}_{ \mathcal{R}_n}|^{-1/2} | = {\rm II} $. To show (ref), it suffices to show $C_n^{-1} {\rm I} =o_{\rm p}(1)$. Recall $\widehat{{\mathbf H}}_{{ \mathcal{R}_n}}=\{\widehat{\boldsymbol{\Gamma}}_{{ \mathcal{R}_n}}(\hat\boldsymbol{\theta}_n)^{\mathrm{\scriptscriptstyle \top} }\widehat{{\mathbf V}}_{{ \mathcal{R}_n}}^{-1/2}(\hat\boldsymbol{\theta}_n)\}^{\otimes2}$. Under Conditions (ref)(b) and (ref), by Proposition (ref), Lemmas (ref) and (ref), if $\log r=o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n \nu^2=o(1)$, we know that the eigenvalues of $\widehat{{\mathbf H}}_{ \mathcal{R}_n}$ are uniformly bounded away from zero and infinity w.p.a.1. Notice that $\widehat{{\mathbf H}}_{{ \mathcal{R}_n}}$ is a $p \times p$ matrix with fixed $p$. Thus, $|\widehat{{\mathbf H}}_{ \mathcal{R}_n}|^{-1/2}$ is uniformly bounded away from zero w.p.a.1. Since $\pi_{0}(\boldsymbol{\theta})$ is bounded away from zero around $\boldsymbol{\theta}_0$ and $|\hat{\boldsymbol{\theta}}_n - \boldsymbol{\theta}_0|_\infty=O_{\rm p}(\nu)$, we know $\pi_0(\hat{\boldsymbol{\theta}}_n)$ is bounded away from zero w.p.a.1. If ${\rm I} =o_{\rm p}(1)$, then $| C_n - (2\pi)^{p/2} \pi_{0}(\hat{\boldsymbol{\theta}}_n) |\widehat{{\mathbf H}}_{ \mathcal{R}_n}|^{-1/2} | =o_{\rm p}(1)$, which implies $C_n^{-1}=O_{\rm p}(1)$. Hence, to show (ref), we only need to show ${\rm I} =o_{\rm p}(1)$. Recall $\ell_n \ll \min\{n^{(\gamma-2)/(9\gamma)}(\log r)^{-1/9},n^{1/3}(\log r)^{-1},n^{(\gamma-2)/(2\gamma)}(\log r)^{-3/2}\}$ and $\ell_nn^{-1/2}(\log r)^{1/2} \ll \nu \ll \min\{\ell_n^{-7/2}n^{-1/\gamma},(\log r)^{-1}\}$. We break the domain of integration into four regions:
Then $ {\rm I} = {\rm I}(1) + {\rm I}(2)+ {\rm I}(3)+{\rm I}(4) $ with
In the sequel, we will show each ${\rm I}(k) = o_{\rm p}(1)$.
For ${\rm I}(3)$, by the triangle inequality, we have
Due to $\int_{\mathcal{D}_3} \pi_{0}(\hat{\boldsymbol{\theta}}_n+n^{-1/2}{\mathbf t}) \,{\rm d}{\mathbf t} \leq n^{p/2} \int_{{\mathbb{R}}^{p}} \pi_{0}(\boldsymbol{\theta}) \,{\rm d}\boldsymbol{\theta} \leq n^{p/2}$, Proposition (ref)(iii) implies that
w.p.a.1 for any $\beta_n^{-1}\ell_n \alpha_n^2 \ll \xi_n \ll \beta_n$. Since $r\gg n$, we can select suitable $\xi_n$ satisfying $ n \beta_n \xi_n \gg \log n $. Then
Due to $n\beta_n^2 \rightarrow \infty$, Proposition 1.1 of Hsu2012 implies that
Therefore, ${\rm I}(3) = o_{\rm p}(1)$.
For ${\rm I}(2)$, it holds that
Since $n\alpha_n^2 \rightarrow \infty$, using the same arguments given above, we have $\pi_{0}(\hat{\boldsymbol{\theta}}_n)\int _{\mathcal{D}_2} \exp( -{\mathbf t}^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf H}}_{ \mathcal{R}_n} {\mathbf t}/2 ) \,{\rm d}{\mathbf t} = o_{\rm p}(1)$. By Proposition (ref)(ii), it then holds w.p.a.1 that
Since $\log n \ll n\kappa_{n}^2$, we have
Therefore, ${\rm I}(2) = o_{\rm p}(1)$.
For ${\rm I}(1)$, by the triangle inequality, we have
Due to $\mathcal{D}_1 \subset \mathcal{T}_n$, by Proposition (ref)(i), it holds that
where $\varpi_n=\max \{\ell_n^{3/2} \alpha_n, \nu, \ell_nn^{1/\gamma}\alpha_n\}$. By Condition (ref), we have $ \sup_{{\mathbf t} \in \mathcal{D}_{1}} |\pi_{0}(\hat{\boldsymbol{\theta}}_n+n^{-1/2}{\mathbf t})-\pi_{0}(\hat{\boldsymbol{\theta}}_n)| =o_{\rm p}(1)$, which implies
Due to $\varpi_nn \alpha_n^2=o(1)$, then $\sup_{{\mathbf t} \in \mathcal{D}_{1}}\{|{\mathbf t}|_2^2\cdot O_{\rm p}(\varpi_n)\}=o_{\rm p}(1)$. Notice that $|e^{x}-1|\leq |x|e^{x}$ for any $x \in {\mathbb{R}}$. Then $\sup_{{\mathbf t} \in \mathcal{D}_1}|\exp\{ |{\mathbf t}|_2^2\cdot O_{\rm p}(\varpi_n)\}-1| = o_{\rm p}(1)$, which implies that
Therefore, ${\rm I}(1) = o_{\rm p}(1)$.
For ${\rm I}(4)$, due to $\mathcal{D}_4 \cap \mathcal{T}_n = \emptyset$, we have $ {\rm I}(4) = \pi_{0}(\hat{\boldsymbol{\theta}}_n) \int _{\mathcal{D}_4} \exp( -{\mathbf t}^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf H}}_{ \mathcal{R}_n} {\mathbf t}/2 ) \,{\rm d}{\mathbf t} $. Since $\boldsymbol{\theta}_0$ is an interior point of $\boldsymbol{\Theta}$, there exists a constant $\iota>0$ such that $\boldsymbol{\Theta} \supset \mathcal{B}_2(\boldsymbol{\theta}_0,\iota) := \{\boldsymbol{\theta} \in{\mathbb{R}}^p:\,|\boldsymbol{\theta}-\boldsymbol{\theta}_0|_2 \leq \iota \}$, which implies $\mathcal{D}_4 = \mathcal{T}_n^{\rm c} \subset \mathcal{T}_n^{*,{\rm c}}$ with $\mathcal{T}_n^* = \{{\mathbf t} \in {\mathbb{R}}^p: {\mathbf t} = n^{1/2}(\boldsymbol{\theta}-\hat\boldsymbol{\theta}_n), \boldsymbol{\theta} \in \mathcal{B}_2(\boldsymbol{\theta}_0,\iota)\}$. By Proposition (ref), it holds w.p.a.1 that $n^{-1/2}|{\mathbf t}|_2 \geq |n^{-1/2}{\mathbf t} + \hat\boldsymbol{\theta}_n - \boldsymbol{\theta}_0|_2 - |\hat\boldsymbol{\theta}_n - \boldsymbol{\theta}_0|_2 \geq \iota/2$ for any ${\mathbf t} \in \mathcal{D}_4$. Together with Proposition 1.1 of Hsu2012, we have w.p.a.1 that
Therefore, ${\rm I}(4) = o_{\rm p}(1)$. $\Box$
Recall $\mathcal{R}(\boldsymbol{\theta})={\rm supp}\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})\}$ and $\mathcal{C}_1=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2\leq\alpha_n\}$ with $\alpha_n=n^{-1/2}(\log r)^{1/2}$. To prove part {\rm (i)} of Proposition (ref), we need the following lemmas whose proofs are given in Sections (ref), (ref) and (ref), respectively.
Notice that
for any $\boldsymbol{\theta} \in \mathcal{C}_1$. Due to $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})=\arg \max _{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})} f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$, then $\partial f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta}) ; \boldsymbol{\theta}\} / \partial \boldsymbol{\lambda}_{\mathcal{R}(\boldsymbol{\theta})} = {\mathbf 0}$. By Lemma (ref), we have $[\nabla_{\boldsymbol{\theta}} \hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})]_{{\mathcal{R}(\boldsymbol{\theta})}^{\rm c},[p]} = {\mathbf 0}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. Thus, it holds w.p.a.1 that
for any $\boldsymbol{\theta} \in \mathcal{C}_1$. By Lemma (ref),
for any $\boldsymbol{\theta} \in \mathcal{C}_1$. Lemma (ref) specifies the leading term of ${\mathbf t}^{\mathrm{\scriptscriptstyle \top} }[\nabla_{\boldsymbol{\theta}}^2 f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}]{\mathbf t}$ for ${\mathbf t}\in\mathbb{R}^p$, whose proof is given in Section (ref).
Let ${\mathbf t} =\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$. Since $\hat{\boldsymbol{\theta}}_n=\arg\min_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$, then $\nabla_{\boldsymbol{\theta}}f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}|_{\boldsymbol{\theta}=\hat{\boldsymbol{\theta}}_n} ={\mathbf 0}$. By the Taylor expansion, it holds that
for some $\tilde{\boldsymbol{\theta}}$ lying on the jointing line between $\hat{\boldsymbol{\theta}}_n$ and $\boldsymbol{\theta}$. Let $\widehat{{\mathbf H}}_{{\mathcal{R}(\boldsymbol{\theta})}} = \{\widehat{\boldsymbol{\Gamma}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf V}}^{-1/2}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})\}^{\otimes2}$ and recall $\widehat{{\mathbf H}}_{ \mathcal{R}_n}=\{\widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf V}}^{-1/2}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)\}^{\otimes2}$. By Lemma (ref), we know $\mathcal{R}(\boldsymbol{\theta})=\mathcal{R}_n$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. Under Conditions (ref)(b) and (ref), by Lemmas (ref) and (ref), if $\log r=o(n^{1/3})$, $\ell_n\nu^2=o(1)$ and $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$, we have $\|\widehat{{\mathbf V}}_{{ \mathcal{R}_n}} (\hat{\boldsymbol{\theta}}_n)\|_2=O_{\rm p}(1)$, $\|\widehat{{\mathbf V}}^{-1}_{{ \mathcal{R}_n}} (\hat{\boldsymbol{\theta}}_n)\|_2=O_{\rm p}(1)$ and $\|\widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)\|_2=O_{\rm p}(1)$. Using the same arguments in the proof of Lemmas (ref) and (ref), if $\log r=o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n\nu^2=o(1)$, we have
which implies $ \sup_{\boldsymbol{\theta} \in \mathcal{C}_1}\|\widehat{{\mathbf H}}_{{\mathcal{R}(\boldsymbol{\theta})}}-\widehat{{\mathbf H}}_{ \mathcal{R}_n}\|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n) $. Together with Lemma (ref), (ref) yields that $ \aleph_n(\boldsymbol{\theta}) - 2^{-1}(\boldsymbol{\theta}-\hat\boldsymbol{\theta}_n)^{\mathrm{\scriptscriptstyle \top} } \widehat{\mathbf H}_{\mathcal{R}_n} (\boldsymbol{\theta}-\hat\boldsymbol{\theta}_n) = |\boldsymbol{\theta}-\hat\boldsymbol{\theta}_n|_2^2 \cdot O_{\rm p}(\varpi_n)$ with $\varpi_n=\max\{\ell_n^{3/2}\alpha_n, \nu, \ell_nn^{1/\gamma}\alpha_n\}$, where $O_{\rm p}(\varpi_n)$ holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$. $\Box$
Recall $\mathcal{R}_n={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$ and $\mathcal{C}_2=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:\alpha_n<|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2\leq\beta_n\}$. Select $\delta_n$ satisfying $\delta_n=o(\ell_n^{-1/2}n^{-1/\gamma})$ and $\ell_n^{1/2}\beta_{n}=o(\delta_n)$, which can be guaranteed by $\ell_n \beta_{n}=o(n^{-1/\gamma})$. For any $\boldsymbol{\theta} \in \mathcal{C}_2$, let $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta}) = \arg\max_{\boldsymbol{\lambda} \in \tilde\Lambda}f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ and $\tilde{\mathcal{R}}(\boldsymbol{\theta})={\rm supp}\{\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta})\}$, where $\tilde\Lambda=\{\boldsymbol{\lambda}\in \mathbb{R}^r:\, |\boldsymbol{\lambda}_{\mathcal{R}_n}|_2\leq \delta_n, \boldsymbol{\lambda}_{\mathcal{R}_n^{\rm c}}={\mathbf 0}\}$. Write $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})=\{\tilde\lambda_1(\boldsymbol{\theta}),\ldots,\tilde\lambda_r(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$. By the Taylor expansion, we have
for some $C\in(0,1)$. By Lemma (ref), we have $|\tilde{\mathcal{R}}(\boldsymbol{\theta})|\leq |\mathcal{R}_n|\leq \ell_n$ w.p.a.1. Notice that $\nu=o(\beta_{n})$. By Proposition (ref), Lemma (ref) and Condition (ref)(b), if $\log r=o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n\beta_{n}=o(n^{-1/\gamma})$, we have $\inf_{\boldsymbol{\theta} \in \mathcal{C}_2} \lambda_{\min}\{\tilde{{\mathbf A}}(\boldsymbol{\theta})\}\geq K_3/2$ w.p.a.1, where $K_3$ is specified in Condition (ref)(b). Recall $\rho(t;\nu)$ is convex w.r.t $t$. Thus $$0 \leq \tilde\boldsymbol{\lambda}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} }\big[\bar{\mathbf g}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta}) - \nu\rho'(0^+){\mbox{\rm sgn}}\{\tilde\boldsymbol{\lambda}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})\}\big] - 4^{-1}K_3|\tilde\boldsymbol{\lambda}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})|_2^2$$ w.p.a.1. Then $|\tilde\boldsymbol{\lambda}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})|_2 \leq 4K_3^{-1} |\bar{\mathbf g}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta}) - \nu\rho'(0^+){\mbox{\rm sgn}}\{\tilde\boldsymbol{\lambda}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})\}|_2$ w.p.a.1. By the Taylor expansion and Condition (ref)(c), $\sup_{\boldsymbol{\theta} \in \mathcal{C}_2} |\bar{\mathbf g}(\boldsymbol{\theta}) - \bar{\mathbf g}(\hat{\boldsymbol{\theta}}_n)|_\infty = O_{\rm p}(\beta_{n})$. Together with the fact $|\tilde{\mathcal{R}}(\boldsymbol{\theta})|\leq \ell_n$ w.p.a.1, we have $\sup_{\boldsymbol{\theta} \in \mathcal{C}_2} |\bar{\mathbf g}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta}) - \bar{\mathbf g}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\hat{\boldsymbol{\theta}}_n)|_2 = O_{\rm p}(\ell_n^{1/2} \beta_{n})$. By the triangle inequality, Lemma (ref) and (ref), since $\alpha_n=o(\nu)$, it holds that $|\bar{\mathbf g}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\hat{\boldsymbol{\theta}}_n)|_2 = O_{\rm p}(\ell_n^{1/2} \nu)$. Due to $\nu=o(\beta_{n})$, we then have $$\sup_{\boldsymbol{\theta} \in \mathcal{C}_2}|\bar{\mathbf g}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta}) - \nu\rho'(0^+){\mbox{\rm sgn}}\{\tilde\boldsymbol{\lambda}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})\}|_2 = O_{\rm p}(\ell_n^{1/2} \beta_{n})\,,$$ which implies $\sup_{\boldsymbol{\theta} \in \mathcal{C}_2}|\tilde\boldsymbol{\lambda}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})|_2 = O_{\rm p}(\ell_n^{1/2} \beta_{n}) = o_{\rm p}(\delta_n)$. Recall $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})=\arg\max_{\boldsymbol{\lambda} \in \tilde\Lambda}f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ and $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta}) \in {\rm int}(\tilde\Lambda)$ for any $\boldsymbol{\theta} \in \mathcal{C}_2$ w.p.a.1. Write $\tilde\boldsymbol{\lambda}_{\mathcal{R}_n}(\boldsymbol{\theta})=\{\dot\lambda_1(\boldsymbol{\theta}),\ldots,\dot\lambda_{|\mathcal{R}_n|}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$. Restricted on $\tilde\Lambda$, by the first-order condition, we have $$ {\mathbf 0}=\frac{1}{n}\sum_{i=1}^{n} \frac{{\mathbf g}_{i,\mathcal{R}_n}(\boldsymbol{\theta})}{1+\tilde{\boldsymbol{\lambda}}_{\mathcal{R}_n}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} }{\mathbf g}_{i,\mathcal{R}_n}(\boldsymbol{\theta})} - \tilde\boldsymbol{\eta}(\boldsymbol{\theta})$$ holds for any $\boldsymbol{\theta} \in \mathcal{C}_2$ w.p.a.1, where $\tilde\boldsymbol{\eta}(\boldsymbol{\theta})=\{\tilde\eta_1(\boldsymbol{\theta}), \ldots,\tilde\eta_{|\mathcal{R}_n|}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$ with $\tilde\eta_j(\boldsymbol{\theta})=\nu\rho'\{|\dot{\lambda}_j(\boldsymbol{\theta})|;\nu\}{\mbox{\rm sgn}}\{\dot{\lambda}_j(\boldsymbol{\theta})\}$ for $\dot{\lambda}_j(\boldsymbol{\theta})\neq0$ and $\tilde\eta_j(\boldsymbol{\theta}) \in [-\nu\rho'(0^+), \nu\rho'(0^+)]$ for $\dot{\lambda}_j(\boldsymbol{\theta})=0$. By the Taylor expansion, we have
for some $\tilde{C}\in(0,1)$, where $\tilde\boldsymbol{\eta}_{*}(\boldsymbol{\theta})\in\mathbb{R}^{|\tilde{\mathcal{R}}(\boldsymbol{\theta})|}$ includes all elements $\tilde\eta_j(\boldsymbol{\theta})$'s in $\tilde\boldsymbol{\eta}(\boldsymbol{\theta})$ such that the associated $\dot \lambda_j(\boldsymbol{\theta}) \neq0$. Hence, $\tilde{\boldsymbol{\lambda}}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})=\check{\mathbf A}^{-1}(\boldsymbol{\theta})\{\bar{\mathbf g}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})-\tilde\boldsymbol{\eta}_{*}(\boldsymbol{\theta})\}$. Using the same arguments in the proof of Lemma (ref), we have $\sup_{\boldsymbol{\theta} \in \mathcal{C}_2}\|\check{\mathbf A}(\boldsymbol{\theta}) - \widehat{{\mathbf V}}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})\|_2=O_{\rm p}(\ell_n\beta_nn^{1/\gamma})=o_{\rm p}(1)$. Applying Proposition (ref), Lemma (ref) and Condition (ref)(b), we know $\sup_{\boldsymbol{\theta} \in \mathcal{C}_2}\lambda_{\max}\{\check{\mathbf A}(\boldsymbol{\theta})\}\leq 2K_4$ w.p.a.1 for $K_4$ specified in Condition (ref)(b). For $\tilde{{\mathbf A}}(\boldsymbol{\theta})$ specified in (ref), we can show $\sup_{\boldsymbol{\theta} \in \mathcal{C}_2}\|\tilde{\mathbf A}(\boldsymbol{\theta})- \check{\mathbf A}(\boldsymbol{\theta})\|_2=O_{\rm p}(\ell_n \beta_{n}n^{1/\gamma})$. By (ref) and Condition (ref)(b), we have
holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_2$ w.p.a.1, where the term $O_{\rm p}(\ell_n^2\beta_{n}^3 n^{1/\gamma})$ holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_2$. As we have shown in the proof of Proposition (ref), $f_n\{\hat\boldsymbol{\lambda}(\hat{\boldsymbol{\theta}}_n);\hat{\boldsymbol{\theta}}_n\}=O_{\rm p}(\ell_n \alpha_n^2)$. Notice that $f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}\geq f_n\{\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$ for any $\boldsymbol{\theta} \in \mathcal{C}_2$. If $\max\{\ell_n \alpha_n^2, \ell_n^2\beta_{n}^3 n^{1/\gamma}\} = o(\kappa_n^2)$, then
as $n \rightarrow \infty$. We complete the proof of part {\rm (ii)} of Proposition (ref). $\Box$
Recall $\mathcal{R}_n={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$ and $\mathcal{C}_3=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2>\beta_n\}$. For any $\boldsymbol{\theta} \in \mathcal{C}_3$, we consider $\check\boldsymbol{\lambda}(\boldsymbol{\theta})=\{\check\lambda_1(\boldsymbol{\theta}),\ldots,\check\lambda_r(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }\in \mathbb{R}^r$ with $\check\boldsymbol{\lambda}_{\mathcal{R}_n^{\rm c}}(\boldsymbol{\theta}) ={\mathbf 0}$ and $$\check\boldsymbol{\lambda}_{\mathcal{R}_n}(\boldsymbol{\theta}) = \frac{\xi_n\{\bar{\mathbf g}_{\mathcal{R}_n}(\boldsymbol{\theta})-\bar{\mathbf g}_{\mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)\}}{|\bar{\mathbf g}_{\mathcal{R}_n}(\boldsymbol{\theta})-\bar{\mathbf g}_{\mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)|_2}\,.$$ Due to $\ell_n^{1/2}\xi_n =o(n^{-1/\gamma})$, Condition (ref)(a) yields $\sup_{\boldsymbol{\theta} \in \mathcal{C}_3} \max_{ i \in [n]} |\check\boldsymbol{\lambda}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_i(\boldsymbol{\theta})|=o_{\rm p}(1)$, which implies the event $\bigcap_{\boldsymbol{\theta} \in \mathcal{C}_3} \{\check\boldsymbol{\lambda}(\boldsymbol{\theta}) \in \hat\Lambda_n(\boldsymbol{\theta})\}$ holds w.p.a.1. Then $$ {\mathbb{P}} \bigg[\inf_{\boldsymbol{\theta} \in \mathcal{C}_3} f_n\{\hat\boldsymbol{\lambda}(\boldsymbol{\theta});\boldsymbol{\theta}\} \geq \inf_{\boldsymbol{\theta} \in \mathcal{C}_3} f_n\{\check\boldsymbol{\lambda}(\boldsymbol{\theta});\boldsymbol{\theta}\}\bigg] =1-o(1)\,. $$ Recall $\sup_{\boldsymbol{\theta} \in \mathcal{C}_3} \lambda_{\max}\{\widehat{\mathbf V}_{\mathcal{R}_n}(\boldsymbol{\theta})\}\leq K_8$ w.p.a.1 for $K_8$ specified in Condition (ref)(c). By the Taylor expansion, we have
holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_3$ w.p.a.1, where $\check C,c_j\in(0,1)$. By Condition (ref)(c), we have $\inf_{\boldsymbol{\theta} \in \mathcal{C}_3}|\bar{\mathbf g}_{\mathcal{R}_n}(\boldsymbol{\theta}) - \bar{\mathbf g}_{\mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)|_2 = \inf_{\boldsymbol{\theta} \in \mathcal{C}_3} |\{\nabla_{\boldsymbol{\theta}} \bar{\mathbf g}_{\mathcal{R}_n}(\tilde \boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} } (\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n)|_2 \geq K_7^{1/2}\beta_{n}$ w.p.a.1. Since $\alpha_n=o(\nu)$, by Lemma (ref) and (ref), we have $|\bar{\mathbf g}_{\mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)|_2=O_{\rm p}(\ell_n^{1/2}\nu)$. Due to $\ell_n^{1/2}\nu=o(\beta_{n})$ and $\xi_n=o(\beta_{n})$, it holds that $$\inf_{\boldsymbol{\theta} \in \mathcal{C}_3} f_n\{\check\boldsymbol{\lambda}(\boldsymbol{\theta});\boldsymbol{\theta}\} \geq K_7^{1/2}\xi_n\beta_{n} + O_{\rm p}(\xi_n \ell_n^{1/2}\nu) + O_{\rm p}(\xi_n^2) \geq \frac{1}{2}K_7^{1/2} \xi_n\beta_{n}$$ w.p.a.1. Since $f_n\{\hat\boldsymbol{\lambda}(\hat{\boldsymbol{\theta}}_n);\hat{\boldsymbol{\theta}}_n\}=O_{\rm p}(\ell_n \alpha_n^2)=o_{\rm p}(\xi_n\beta_{n})$, then
as $n\rightarrow \infty$. We complete the proof of part {\rm (iii)} of Proposition (ref). $\Box$
Let ${\mathbb{E}}_{{\mathbf t} \sim \pi^{\dag}_{{\mathbf t}}}({\mathbf t}) = \int_{{\mathbb{R}}^p} {\mathbf t} \pi^{\dag}_{{\mathbf t}}({\mathbf t}\,|\,\mathcal{X}_n) \, {\rm d}{\mathbf t}$ for $\pi^{\dag}_{{\mathbf t}}({\mathbf t}\,|\,\mathcal{X}_n)$ given in (ref). Notice that ${\mathbb{E}}_{{\mathbf t} \sim \pi^{\dag}_{{\mathbf t}}}({\mathbf t}) = n^{1/2}\{{\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta}) - \hat{\boldsymbol{\theta}}_n\}$ and ${\mathbb{E}}_{{\mathbf t} \sim \mathcal{N}({\mathbf 0}, \widehat{{\mathbf H}}_{ \mathcal{R}_n}^{-1})}({\mathbf t}) = {\mathbf 0}$. To prove Corollary (ref), it is equivalent to show $|{\mathbb{E}}_{{\mathbf t} \sim \pi^{\dag}_{{\mathbf t}}}({\mathbf t}) - {\mathbb{E}}_{{\mathbf t} \sim \mathcal{N}({\mathbf 0}, \widehat{{\mathbf H}}_{ \mathcal{R}_n}^{-1})}({\mathbf t})|_\infty = o_{\rm p}(1)$. It follows from the triangle inequality that
As shown in Section (ref), we have $C_n^{-1} = O_{\rm p}(1)$. It suffices to show ${\rm III} = o_{\rm p}(1)$ and ${\rm IV}=o_{\rm p}(1)$. Notice that
Recall $\widehat{{\mathbf H}}_{{ \mathcal{R}_n}}=\{\widehat{\boldsymbol{\Gamma}}_{{ \mathcal{R}_n}}(\hat\boldsymbol{\theta}_n)^{\mathrm{\scriptscriptstyle \top} }\widehat{{\mathbf V}}_{{ \mathcal{R}_n}}^{-1/2}(\hat\boldsymbol{\theta}_n)\}^{\otimes2}$. Under Conditions (ref)(b) and (ref), by Proposition (ref), Lemmas (ref) and (ref), if $\log r=o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n \nu^2=o(1)$, we know that the eigenvalues of $\widehat{{\mathbf H}}_{ \mathcal{R}_n}$ are uniformly bounded away from zero and infinity w.p.a.1. Since $\widehat{{\mathbf H}}_{{ \mathcal{R}_n}}$ is a $p \times p$ matrix with fixed $p$, then
As shown in the proof of Theorem (ref), we have $|C_n - (2\pi)^{p/2} \pi_{0}(\hat{\boldsymbol{\theta}}_n) |\widehat{{\mathbf H}}_{ \mathcal{R}_n}|^{-1/2} | \leq {\rm I} = o_{\rm p}(1)$ for ${\rm I}$ defined in (ref), which implies ${\rm IV} \leq {\rm I} \cdot {\mathbb{E}}_{{\mathbf t} \sim \mathcal{N}({\mathbf 0}, \widehat{{\mathbf H}}_{ \mathcal{R}_n}^{-1})}(|{\mathbf t}|_\infty)= o_{\rm p}(1)$. In the sequel, we will show that ${\rm III} =o_{\rm p}(1)$. Recall $\ell_n \ll \min\{n^{(\gamma-2)/(9\gamma)}(\log r)^{-1/9},n^{1/3}(\log r)^{-1},n^{(\gamma-2)/(2\gamma)}(\log r)^{-3/2}\}$ and $\ell_nn^{-1/2}(\log r)^{1/2} \ll \nu \ll \min\{\ell_n^{-7/2}n^{-1/\gamma},(\log r)^{-1}\}$. For $(\mathcal{D}_1, \mathcal{D}_2, \mathcal{D}_3, \mathcal{D}_4)$ defined as (ref), it holds that ${\rm III} = {\rm III}(1) + {\rm III}(2)+ {\rm III}(3) + {\rm III}(4)$ with
For ${\rm III}(3)$, by the triangle inequality, we have
Since $\boldsymbol{\Theta} \subset {\mathbb{R}}^p$ is a compact set, then
By Proposition (ref)(iii),
w.p.a.1 for any $\beta_n^{-1}\ell_n \alpha_n^2 \ll \xi_n \ll \beta_n$. Since $r\gg n$, we can select suitable $\xi_n$ satisfying $ n \beta_n \xi_n \gg \log n $. Then
Recall that the eigenvalues of $\widehat{{\mathbf H}}_{ \mathcal{R}_n}$ are uniformly bounded away from zero and infinity w.p.a.1. Since $n\beta_n^2 \rightarrow \infty$ and $p$ is fixed, by the Cauchy-Schwarz inequality and Proposition 1.1 of Hsu2012, we have
which implies
Therefore, ${\rm III}(3) = o_{\rm p} (1)$.
For ${\rm III}(2)$, it holds that
Since $n\alpha_n^2 \rightarrow \infty$, using the same arguments for (ref), we have $$ \pi_{0}(\hat{\boldsymbol{\theta}}_n)\int _{\mathcal{D}_2} |{\mathbf t}|_\infty \exp( -{\mathbf t}^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf H}}_{ \mathcal{R}_n} {\mathbf t}/2 ) \,{\rm d}{\mathbf t} = o_{\rm p}(1) \,. $$ Due to $\log n \ll n\kappa_{n}^2$, by Proposition (ref)(ii), it then holds w.p.a.1 that
Therefore, ${\rm III}(2)=o_{\rm p}(1)$.
For ${\rm III}(1)$, by Proposition (ref)(i), we have
where $\varpi_n=\max \{\ell_n^{3/2} \alpha_n, \nu, \ell_nn^{1/\gamma}\alpha_n\}$. Under Condition (ref), we know $ \sup_{{\mathbf t} \in \mathcal{D}_{1}} |\pi_{0}(\hat{\boldsymbol{\theta}}_n+n^{-1/2}{\mathbf t})-\pi_{0}(\hat{\boldsymbol{\theta}}_n)| =o_{\rm p}(1)$. By (ref), we have
Due to $\varpi_nn \alpha_n^2=o(1)$, then $\sup_{{\mathbf t} \in \mathcal{D}_{1}}\{|{\mathbf t}|_2^2\cdot O_{\rm p}(\varpi_n)\}=o_{\rm p}(1)$. Notice that $|e^{x}-1|\leq |x|e^{x}$ for any $x \in {\mathbb{R}}$. Then $\sup_{{\mathbf t} \in \mathcal{D}_1}|\exp\{ |{\mathbf t}|_2^2\cdot O_{\rm p}(\varpi_n)\}-1| = o_{\rm p}(1)$, which implies that
Therefore, ${\rm III}(1)= o_{\rm p}(1)$.
For ${\rm III}(4)$, due to $\mathcal{D}_4 \cap \mathcal{T}_n = \emptyset$, we have $ {\rm III}(4) = \pi_{0}(\hat{\boldsymbol{\theta}}_n) \int _{\mathcal{D}_4} |{\mathbf t}|_\infty \exp( -{\mathbf t}^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf H}}_{ \mathcal{R}_n} {\mathbf t}/2 ) \,{\rm d}{\mathbf t} $. As shown in Section (ref), it holds w.p.a.1 that $n^{-1/2}|{\mathbf t}|_2 \geq \iota/2$ for any ${\mathbf t} \in \mathcal{D}_4$. Using the same arguments for (ref), we have ${\mathbb{E}}_{{\mathbf t} \sim \mathcal{N}({\mathbf 0}, \widehat{{\mathbf H}}_{ \mathcal{R}_n}^{-1})}\{|{\mathbf t}|_\infty I({\mathbf t} \in \mathcal{D}_4)\} = o_{\rm p}(1)$, which implies
Therefore, ${\rm III}(4) = o_{\rm p}(1)$. $\Box$
To prove Theorem (ref), we first introduce the following two concepts.
Denote by $\Psi(\boldsymbol{\theta}, \cdot)$ the transition probability of the Markov chain determined by Algorithm (ref) at $\boldsymbol{\theta} \in \boldsymbol{\Theta}$. For given $\boldsymbol{\theta} \in \boldsymbol{\Theta}$, $\alpha_{\boldsymbol{\theta}}(\boldsymbol{\vartheta}) = \min \{1, R_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})\}$ is the acceptance probability at $\boldsymbol{\vartheta} \in {\mathbb{R}}^p$, where
Then the transition probability of the associated Markov chain at $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ has a probability mass $\psi_{\boldsymbol{\theta}} = 1- \int_{\boldsymbol{\Theta}} \phi(\boldsymbol{\vartheta}\,|\,\boldsymbol{\theta}) \alpha_{\boldsymbol{\theta}}(\boldsymbol{\vartheta}) \,{\rm d}\boldsymbol{\vartheta}$. Define $\psi(\boldsymbol{\theta},\boldsymbol{\vartheta}) = \phi(\boldsymbol{\vartheta}\,|\,\boldsymbol{\theta}) \alpha_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})$ for any $\boldsymbol{\theta},\boldsymbol{\vartheta} \in \boldsymbol{\Theta}$. We have $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) \psi(\boldsymbol{\theta},\boldsymbol{\vartheta}) = \pi^\dag(\boldsymbol{\vartheta}\,|\,\mathcal{X}_n) \psi(\boldsymbol{\vartheta},\boldsymbol{\theta})$ for any $\boldsymbol{\theta}, \boldsymbol{\vartheta} \in \boldsymbol{\Theta}$. Since the Markov chain determined by Algorithm {\rm (ref)} always stays in $\boldsymbol{\Theta}$, its transition probability $\Psi(\cdot\,, \cdot): \boldsymbol{\Theta} \times \mathscr{B}(\boldsymbol{\Theta}) \mapsto {\mathbb{R}}_+$ satisfies $ \Psi(\boldsymbol{\theta}, {\rm d}\boldsymbol{\vartheta}) = \psi_{\boldsymbol{\theta}} \delta_{\boldsymbol{\theta}}({\rm d}\boldsymbol{\vartheta}) + \psi(\boldsymbol{\theta},\boldsymbol{\vartheta}) \, {\rm d}\boldsymbol{\vartheta} $, where $\mathscr{B}(\boldsymbol{\Theta})$ is the Borel $\sigma$-algebra on $\boldsymbol{\Theta}$, and $\delta_{\boldsymbol{\theta}}$ is the Dirac-delta function at $\boldsymbol{\theta}$ with $\delta_{\boldsymbol{\theta}}(A)=I(\boldsymbol{\theta} \in A)$. For any $A,B \in \mathscr{B}(\boldsymbol{\Theta})$, we have
Therefore, $\Pi^\dag_n(A) = \int_A \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)\,{\rm d}\boldsymbol{\theta} = \int_A \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) \Psi(\boldsymbol{\theta}, \boldsymbol{\Theta}) \,{\rm d}\boldsymbol{\theta} = \int_{\boldsymbol{\Theta}} \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) \Psi(\boldsymbol{\theta}, A) \,{\rm d}\boldsymbol{\theta}$ for any $A \in \mathscr{B}(\boldsymbol{\Theta})$, which implies that $\Pi^\dag_n$ is the stationary distribution of such Markov chain with transition probability $\Psi(\cdot\,,\cdot)$.
Denote by $\mathbb{L}(\cdot)$ the Lebesgue measure on ${\mathbb{R}}^p$. For any $A\in\mathscr{B}(\boldsymbol{\Theta})$ with $\Pi_n^\dag(A)>0$, due to $\Pi_n^\dag(A)=\int_A \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)\,{\rm d}\boldsymbol{\theta}$, we know $\mathbb{L}(A)>0$. Recall that $\boldsymbol{\Theta} \subset {\mathbb{R}}^p$ is a compact set. Since $\phi(\boldsymbol{\vartheta}\,|\,\boldsymbol{\theta})$ is positive and continuous on $(\boldsymbol{\theta}, \boldsymbol{\vartheta}) \in \boldsymbol{\Theta} \times \boldsymbol{\Theta}$, there exists a constant $C>0$ such that $\inf_{\boldsymbol{\theta}, \boldsymbol{\vartheta} \in \boldsymbol{\Theta}}\phi(\boldsymbol{\theta}\,|\,\boldsymbol{\vartheta}) > C$. On one hand, for any $A\in\mathscr{B}(\boldsymbol{\Theta})$ and $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ such that $\Pi_n^\dag(A)>0$ and $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) = 0$, we have $\Psi(\boldsymbol{\theta}, A) \geq \int_A \psi(\boldsymbol{\theta},\boldsymbol{\vartheta}) \, {\rm d}\boldsymbol{\vartheta} = \int_A \phi(\boldsymbol{\vartheta}\,|\,\boldsymbol{\theta}) \, {\rm d}\boldsymbol{\vartheta} >0$. On the other hand, for any $A\in\mathscr{B}(\boldsymbol{\Theta})$ and $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ such that $\Pi_n^\dag(A)>0$ and $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)>0$, we have
Since $\mathbb{L}(\{\boldsymbol{\vartheta} \in A: \pi^\dag(\boldsymbol{\vartheta}\,|\,\mathcal{X}_n) \geq \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)\})$ and $\int_{\boldsymbol{\vartheta} \in A:\, \pi^\dag(\boldsymbol{\vartheta}\,|\,\mathcal{X}_n) < \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)} \pi^\dag(\boldsymbol{\vartheta}\,|\,\mathcal{X}_n) \, {\rm d}\boldsymbol{\vartheta}$ cannot be zero simultaneously for any $A\in\mathscr{B}(\boldsymbol{\Theta})$ with $\Pi_n^\dag(A)>0$, then $\Psi(\boldsymbol{\theta}, A) >0$ for any $A\in\mathscr{B}(\boldsymbol{\Theta})$ and $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ such that $\Pi_n^\dag(A)>0$ and $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) > 0$. Therefore, it holds that $\Psi(\boldsymbol{\theta}, A) >0$ for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ and $A\in\mathscr{B}(\boldsymbol{\Theta})$ with $\Pi_n^\dag(A)>0$. By Definition (ref), the Markov chain with transition probability $\Psi(\cdot\,,\cdot)$ is $\Pi^\dag_n$-irreducible. Furthermore, by Definition (ref), we know the Markov chain $\{\boldsymbol{\theta}^k\}_{k\geq1}$ with transition probability $\Psi(\cdot\,,\cdot)$ and initial point $\boldsymbol{\theta}^0$ is aperiodic. Notice that $\mathscr{B}(\boldsymbol{\Theta})$ is a countably generated $\sigma$-algebra. Denote by $\mathcal{T}^k_{\boldsymbol{\theta}^{0}}(\cdot)$ the measure which admits the distribution of such Markov chain at $k$-th step with initial point $\boldsymbol{\theta}^{0}$. Conditional on $\mathcal{X}_n$, for any $\boldsymbol{\theta}^0 \in \boldsymbol{\Theta}$ such that $\pi^\dag(\boldsymbol{\theta}^0\,|\,\mathcal{X}_n)>0$, by Theorem 4 of RobertsRosenthal2004, we have $\mathcal{D}_{\mathrm{\scriptscriptstyle TV} }(\mathcal{T}^k_{\boldsymbol{\theta}^{0}} ,\, \Pi^{\dag}_n) \rightarrow 0$ as $k \rightarrow \infty$. Furthermore, notice that $\boldsymbol{\Theta} \subset {\mathbb{R}}^p$ is a compact set with fixed $p$. Conditional on $\mathcal{X}_n$, for any $\boldsymbol{\theta}^0 \in \boldsymbol{\Theta}$ such that $\pi^\dag(\boldsymbol{\theta}^0\,|\,\mathcal{X}_n)>0$, it follows from Fact 5 of RobertsRosenthal2004 that $|K^{-1}\sum_{k=1}^K \boldsymbol{\theta}^k - {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta})|_\infty \leq |K^{-1}\sum_{k=1}^K \boldsymbol{\theta}^k - {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta})|_1 \rightarrow 0 $ almost surely as $K \rightarrow \infty$, where $\{\boldsymbol{\theta}^k\}_{k\geq 1}$ are generated via Algorithm (ref) with the initial $\boldsymbol{\theta}^0$ and ${\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta})$ is defined in (ref). We complete the proof of Theorem (ref). $\Box$
For the function ${\mathbf h}: {\mathbb{R}}^p \mapsto {\mathbb{R}}^s$ involved in Algorithm {\rm (ref)}, let $\boldsymbol{\zeta}^* = {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}\{{\mathbf h}(\boldsymbol{\theta})\}$. Define
with $S_K=N_1+\cdots+N_K$, where $\{\boldsymbol{\theta}^1_1, \ldots, \boldsymbol{\theta}^1_{N_1}, \ldots,\boldsymbol{\theta}^K_1,\ldots, \boldsymbol{\theta}^K_{N_K}\}$ are generated via Algorithm {\rm (ref)}. To construct Theorem (ref), we need the following two lemmas whose proofs are given in Sections (ref) and (ref), respectively.
Denote by ${\mathbb{P}}_{\mathcal{X}_n}(\cdot)$ the conditional probability given $\mathcal{X}_n$. For some sufficiently large $M>0$, by Lemma (ref), we have that for any $\epsilon>0$, there exists a sufficiently large integer $k_\epsilon$ such that $\mathbb{P}_{\mathcal{X}_n}(\mathcal{A}) \leq \epsilon$ with $\mathcal{A} = \bigcup_{t=k_\epsilon}^\infty \{|\hat{\boldsymbol{\zeta}}_{t} - \boldsymbol{\zeta}^*|_\infty > M\}$. Define a compact set $\mathcal{B}=\{\boldsymbol{\zeta} \in {\mathbb{R}}^s: |\boldsymbol{\zeta} - \boldsymbol{\zeta}^*|_\infty \leq M \} $. Recall $\boldsymbol{\Theta} \subset {\mathbb{R}}^p$ is a compact set with fixed $p$. Since $\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})$ is positive and continuous on $(\boldsymbol{\theta}, \boldsymbol{\zeta}) \in \boldsymbol{\Theta} \times {\mathbb{R}}^s$, then $\inf_{\boldsymbol{\theta} \in \boldsymbol{\Theta}, \boldsymbol{\zeta} \in \mathcal{B}} \varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta}) \geq C_M$ for some constant $C_M>0$ and $\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})$ is uniformly continuous on $(\boldsymbol{\theta}\,; \boldsymbol{\zeta}) \in \boldsymbol{\Theta} \times \mathcal{B}$. For any $\varepsilon>0$, there exists $\delta(\varepsilon)>0$ such that $|\varphi(\boldsymbol{\theta}_1\,; \boldsymbol{\zeta}_1) - \varphi(\boldsymbol{\theta}_2\,; \boldsymbol{\zeta}_2)| < C_M (2+2C_M)^{-1}\varepsilon$ for any $(\boldsymbol{\theta}_1, \boldsymbol{\zeta}_1), (\boldsymbol{\theta}_2, \boldsymbol{\zeta}_2) \in \boldsymbol{\Theta} \times \mathcal{B}$ satisfying $|\boldsymbol{\theta}_1 - \boldsymbol{\theta}_2|_\infty \leq \delta(\varepsilon)$ and $|\boldsymbol{\zeta}_1 - \boldsymbol{\zeta}_2|_\infty \leq \delta(\varepsilon)$. For any $K \geq k_\epsilon$, it holds that $$ \inf_{\boldsymbol{\theta} \in \boldsymbol{\Theta}}\frac{1}{S_K} \sum_{k=1}^K N_k \varphi(\boldsymbol{\theta}\,;\boldsymbol{\zeta}) \geq \inf_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} \frac{1}{S_K} \sum_{k=k_\epsilon}^K N_k \varphi(\boldsymbol{\theta}\,;\boldsymbol{\zeta}) \geq \frac{C_M}{S_K}\bigg(S_K - \sum_{k=1}^{k_\epsilon-1} N_k\bigg) $$ for any $\boldsymbol{\zeta} \in \mathcal{B}$. Notice that $S_K \to \infty$ as $K \to \infty$. Given $k_\epsilon$, there exists a sufficiently large integer $K^*$ such that $S_K / (S_K - \sum_{k=1}^{k_\epsilon-1} N_k) \leq 1+C_M$ for all $K > K^*$. Recall $N_k\leq N_{k+1}$ for any $k\geq 1$ and $N_k\rightarrow\infty$ as $k\rightarrow\infty$. Since $\sup_{\boldsymbol{\theta}\in\boldsymbol{\Theta},\boldsymbol{\zeta}\in\mathbb{R}^s}\varphi(\boldsymbol{\theta}\,;\boldsymbol{\zeta})<\infty$, we have \[ \frac{1}{S_t}\sum_{k=1}^{\lfloor\sqrt{t}\rfloor}N_k\sup_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}|\varphi(\boldsymbol{\theta}\,;\boldsymbol{\zeta}^*)-\varphi(\boldsymbol{\theta}\,;\hat{\boldsymbol{\zeta}}_k)|\lesssim \frac{1}{S_t}\sum_{k=1}^{\lfloor\sqrt{t}\rfloor}N_k\rightarrow0 \] as $t\rightarrow\infty$. For given $\varepsilon>0$, there exists some sufficiently large integer $\tilde{K}$ such that \[ \frac{1}{S_t}\sum_{k=1}^{\lfloor\sqrt{t}\rfloor}N_k\sup_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}|\varphi(\boldsymbol{\theta}\,;\boldsymbol{\zeta}^*)-\varphi(\boldsymbol{\theta}\,;\hat{\boldsymbol{\zeta}}_k)|\leq \frac{C_M\varepsilon}{2(1+C_M)} \] for any $t\geq \tilde{K}$. It holds that
for any $t > \max(K^*, k_\epsilon,\tilde{K})$. We then have
where the second step is due to the fact that conditional on $\mathcal{X}_n$ we have $|\hat{\boldsymbol{\zeta}}_{k} - \boldsymbol{\zeta}^*|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$. Letting $\epsilon \to 0$, we know that conditional on $\mathcal{X}_n$,
almost surely as $K \rightarrow \infty$.
Define $$ \widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(|\boldsymbol{\theta}|_\infty) = \frac{1}{S_K} \sum_{k=1}^{K} \sum_{i=1}^{N_k} \frac{\pi^\dag(\boldsymbol{\theta}^k_i\,|\,\mathcal{X}_n)}{\varphi(\boldsymbol{\theta}^k_i\,;\boldsymbol{\zeta}^*)}|\boldsymbol{\theta}^k_i|_\infty ~~\textrm{and}~~ {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(|\boldsymbol{\theta}|_\infty) = \int_{{\mathbb{R}}^p} |\boldsymbol{\theta}|_\infty \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) \, {\rm d}\boldsymbol{\theta} \,, $$ where $\{\boldsymbol{\theta}^1_1, \ldots, \boldsymbol{\theta}^1_{N_1}, \ldots,\boldsymbol{\theta}^K_1,\ldots, \boldsymbol{\theta}^K_{N_K}\}$ are generated via Algorithm {\rm (ref)}. For $\widehat{{\mathbb{E}}}_{\pi^\dag,K}(\boldsymbol{\theta})$ defined in (ref) and $\widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(\boldsymbol{\theta})$ defined in (ref), since $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) |\boldsymbol{\theta}|_\infty / \varphi(\boldsymbol{\theta}\,;\boldsymbol{\zeta}^*) = 0$ for any $\boldsymbol{\theta} \notin \boldsymbol{\Theta}$, we then have
Using the same arguments for the proof of Lemma (ref) in Section (ref), it holds that conditional on $\mathcal{X}_n$ we have $|\widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(|\boldsymbol{\theta}|_\infty) - {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(|\boldsymbol{\theta}|_\infty)| \rightarrow 0$ almost surely as $K \rightarrow \infty$. Notice that $\boldsymbol{\Theta} \subset {\mathbb{R}}^p$ is a compact set with fixed $p$. Then ${\mathbb{E}}_{\boldsymbol{\theta} \sim\pi^\dag}(|\boldsymbol{\theta}|_\infty) < \infty$. Together with (ref), it holds that conditional on $\mathcal{X}_n$ we have $|\widehat{{\mathbb{E}}}_{\pi^\dag,K}(\boldsymbol{\theta})- \widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(\boldsymbol{\theta}) |_\infty \rightarrow 0$ almost surely as $K \rightarrow \infty$. By the triangle inequality and Lemma (ref), conditional on $\mathcal{X}_n$, $|\widehat{{\mathbb{E}}}_{\pi^\dag,K}(\boldsymbol{\theta})- {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta})|_\infty \leq |\widehat{{\mathbb{E}}}_{\pi^\dag,K}(\boldsymbol{\theta})- \widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(\boldsymbol{\theta})|_\infty + |\widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(\boldsymbol{\theta})- {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta})|_\infty \rightarrow 0$ almost surely as $K \rightarrow \infty$. We complete the proof of Theorem (ref). $\Box$
The proof is almost identical to that of Lemma 1 in Changetal2018. Recall $p$ is fixed. We only need to replace $\{\varrho_n,\omega_n,\xi_n,b_n^{1/(2\beta)},s\}$ appeared in the proof of Lemma 1 in Changetal2018 by $(1,1,1,\varphi_n,p)$ and all the arguments still hold. $\Box$
Due to the convexity of $P_{\nu}(\cdot)$, $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ is concave w.r.t $\boldsymbol{\lambda}$. We only need to show that there exists a local maximizer $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta}_0)$ satisfying the results stated in the lemma. Recall $\mathcal{M}_{\boldsymbol{\theta}_0}^*=\{j\in[r]: |\bar{g}_j(\boldsymbol{\theta}_0)| \geq C_*\nu \rho'(0^{+}) \}$ for some $C_* \in (0,1)$, and $\mathbb{P}(\max _{\boldsymbol{\theta} \in \boldsymbol{\Theta}:\,|\boldsymbol{\theta}-\boldsymbol{\theta}_0|_2 \leq c_n}|\mathcal{M}_{\boldsymbol{\theta}}^*|\leq\ell_n)\rightarrow1$ for some $c_n\rightarrow 0$ satisfying $\nu c_n^{-1} \rightarrow 0$. For any given $c\in(C_*,1)$, write $\mathcal{M}_{\boldsymbol{\theta}_0}:=\mathcal{M}_{\boldsymbol{\theta}_0}(c)=\{ j\in[r]:|\bar{g}_j(\boldsymbol{\theta}_0)| \geq c\nu\rho'(0^{+}) \}$. Then $\ell_n \geq |\mathcal{M}_{\boldsymbol{\theta}_0}^*| \geq |\mathcal{M}_{\boldsymbol{\theta}_0}|$ w.p.a.1. To prove Lemma (ref), we establish its validity separately with Case 1: $\mathcal{M}_{\boldsymbol{\theta}_0} \neq \emptyset$ and Case 2: $\mathcal{M}_{\boldsymbol{\theta}_0} = \emptyset$.
Restricted on $\mathcal{M}_{\boldsymbol{\theta}_0}$, we select $\delta_n$ satisfying $\delta_n=o(\ell_n^{-1/2} n^{-1/\gamma})$ and $\ell_n^{1/2} \alpha_n=o(\delta_n) $, which can be guaranteed by $\ell_n\alpha_n=o(n^{-1/\gamma})$. Let $\Lambda_{0}=\{\boldsymbol{\lambda}\in {\mathbb{R}}^{r}:|\boldsymbol{\lambda}_{\mathcal{M}_{\boldsymbol{\theta}_0}}|_2\leq \delta_n {\rm \ and \ }\boldsymbol{\lambda}_{\mathcal{M}^{\rm c}_{\boldsymbol{\theta}_0}}={\mathbf 0} \}$ and $\tilde{\boldsymbol{\lambda}}_{0}=\arg \max _{\boldsymbol{\lambda} \in \Lambda_{0}} f_n(\boldsymbol{\lambda}; \boldsymbol{\theta}_0)$. By Condition (ref)(a), we have $\max_{i\in[n],j\in[r]}|g_{i,j} (\boldsymbol{\theta}_0)| = O_{\rm p}(n^{1/\gamma})$, which implies $\max_{i\in[n]} |{\mathbf g}_{i,\mathcal{M}_{\boldsymbol{\theta}_0}} (\boldsymbol{\theta}_0)|_2 = O_{\rm p}(\ell_n^{1/2} n^{1/\gamma})$. Then $\max_{i\in[n]} |\tilde{\boldsymbol{\lambda}}^{\mathrm{\scriptscriptstyle \top} }_{0} {\mathbf g}_{i}(\boldsymbol{\theta}_0)|=o_{\rm p}(1)$. Write $\tilde{\boldsymbol{\lambda}}_{0}=(\tilde{\lambda}_{0,1},\ldots,\tilde{\lambda}_{0,r})^{\mathrm{\scriptscriptstyle \top} }$. By the Taylor expansion, we have
for some $C \in (0,1)$. By Condition (ref)(b) and the same arguments for deriving Lemma \rm(ref), if $\log r=o(n^{1/3})$ and $\ell_n\alpha_n=o(1)$, we have $\lambda_{\min}\{\widehat{{\mathbf V}}_{\mathcal{M}_{\boldsymbol{\theta}_0}} (\boldsymbol{\theta}_0)\}$ is uniformly bounded away from zero w.p.a.1. Thus $0 \leq|\tilde{\boldsymbol{\lambda}}_{0,\mathcal{M}_{\boldsymbol{\theta}_0}}|_2 |\bar{{\mathbf g}}_{\mathcal{M}_{\boldsymbol{\theta}_0}}(\boldsymbol{\theta}_0)|_2 - 4^{-1}K_3 |\tilde{\boldsymbol{\lambda}}_{0,\mathcal{M}_{\boldsymbol{\theta}_0}}|_2^2$ w.p.a.1, where $K_3$ is specified in Condition (ref)(b). By the moderate deviation of self-normalized sums JingShaoWang2003, $ |\bar{{\mathbf g}}(\boldsymbol{\theta}_0)|_{\infty} =O_{\rm p}(\alpha_n)$. Then $ |\bar{{\mathbf g}}_{\mathcal{M}_{\boldsymbol{\theta}_0}}(\boldsymbol{\theta}_0)|_2= O_{\rm p}(\ell_n^{1/2} \alpha_n) $ and $ |\tilde{\boldsymbol{\lambda}}_{0,\mathcal{M}_{\boldsymbol{\theta}_0}}|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n)=o_{\rm p}(\delta_n)$. Write $\tilde{\boldsymbol{\lambda}}_{0,\mathcal{M}_{\boldsymbol{\theta}_0}}=(\tilde{\lambda}_{1},\ldots,\tilde{\lambda}_{|\mathcal{M}_{\boldsymbol{\theta}_0}|})^{\mathrm{\scriptscriptstyle \top} }$. We then have w.p.a.1 that
where $\tilde{\boldsymbol{\eta}}=(\tilde{\eta}_1,\ldots,\tilde{\eta}_{|\mathcal{M}_{\boldsymbol{\theta}_0}|})^{\mathrm{\scriptscriptstyle \top} }$ with $\tilde{\eta}_j=\nu\rho'(|\tilde\lambda_{j}|;\nu){\mbox{\rm sgn}}(\tilde{\lambda}_{j})$ for $\tilde\lambda_{j}\neq0$ and $\tilde\eta_j\in[-\nu\rho'(0^+),\nu\rho'(0^+)]$ for $\tilde\lambda_{j}=0$. In the sequel, we will show that $\tilde{\boldsymbol{\lambda}}_{0}$ is a local maximizer for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1.
Firstly, define $\Lambda_{0}^* = \{\boldsymbol{\lambda}\in {\mathbb{R}}^r:|\boldsymbol{\lambda}_{\mathcal{M}_{\boldsymbol{\theta}_0}^*}|_2 \leq \varepsilon,\boldsymbol{\lambda}_{\mathcal{M}^{*,{\rm c}}_{\boldsymbol{\theta}_0}}={\mathbf 0}\}$ for some sufficiently small constant $\varepsilon>0$. For $\tilde{\boldsymbol{\lambda}}_{0}$ defined before, we will prove $\tilde\boldsymbol{\lambda}_{0}=\arg \max_{\boldsymbol{\lambda}\in\Lambda_{0}^*} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. Since $\tilde{\boldsymbol{\lambda}}_{0} \in\Lambda_{0}$ and $\mathcal{M}_{\boldsymbol{\theta}_0} \subset \mathcal{M}_{\boldsymbol{\theta}_0}^*$, we know $\tilde{\boldsymbol{\lambda}}_{0} \in \Lambda_{0}^*$ for sufficiently large $n$. Restricted on $\boldsymbol{\lambda} \in \Lambda_{0}^*$, by the concavity of $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.r.t $\boldsymbol{\lambda}$, it suffices to show that ${\mathbf w} = \tilde{\boldsymbol{\lambda}}_{0,\mathcal{M}^*_{\boldsymbol{\theta}_0}}=: (w_{1},\ldots,w_{|\mathcal{M}^*_{\boldsymbol{\theta}_0}|})^{\mathrm{\scriptscriptstyle \top} } \in {\mathbb{R}}^{|\mathcal{M}^*_{\boldsymbol{\theta}_0}|}$ satisfies the equation
w.p.a.1, where $\tilde{\boldsymbol{\eta}}^*=(\tilde\eta^*_1,\ldots,\tilde\eta^*_{|\mathcal{M}_{\boldsymbol{\theta}_0}^*|})^{\mathrm{\scriptscriptstyle \top} }$ with $\tilde{\eta}^*_j=\nu\rho'(|w_j|;\nu){\mbox{\rm sgn}}(w_j)$ for $w_j\neq0$ and $\tilde{\eta}^*_j \in [-\nu\rho'(0^+),\nu\rho'(0^+)]$ for $w_j=0$. By (ref), we know $0= n^{-1}\sum_{i=1}^{n} g_{i,j}(\boldsymbol{\theta}_0)/\{1+{\mathbf w} ^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{M}_{\boldsymbol{\theta}_0}^*}(\boldsymbol{\theta}_0)\} - \tilde{\eta}^*_j$ holds for any $j\in\mathcal{M}_{\boldsymbol{\theta}_0}$. For any $j \in \mathcal{M}_{\boldsymbol{\theta}_0}^*\backslash \mathcal{M}_{\boldsymbol{\theta}_0}$, since $\max_{i\in[n]} |{\mathbf w}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{M}_{\boldsymbol{\theta}_0}^*}(\boldsymbol{\theta}_0)|=\max_{i\in[n]} |\tilde\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} }_{0} {\mathbf g}_{i}(\boldsymbol{\theta}_0)|=o_{\rm p}(1)$, it holds that
with
Due to $|{\mathbf w}|_2 = |\tilde{\boldsymbol{\lambda}}_{0,\mathcal{M}_{\boldsymbol{\theta}_0}}|_2 = O_{\rm p}(\ell_n^{1/2} \alpha_n)$, by Conditions (ref)(a) and (ref)(b), $\max_{j\in[r]}|R_j|=O_{\rm p}(|{\mathbf w}|_2)=O_{\rm p}(\ell_n^{1/2} \alpha_n)$. Notice that $C_* \nu\rho'(0^+) \leq |\bar{g}_j(\boldsymbol{\theta}_0)| < c \nu\rho'(0^+)$ for any $j\in\mathcal{M}_{\boldsymbol{\theta}_0}^*\backslash \mathcal{M}_{\boldsymbol{\theta}_0}$, and $\ell_n^{1/2} \alpha_n=o(\nu)$. Then \[ \max_{j\in\mathcal{M}_{\boldsymbol{\theta}_0}^*\backslash \mathcal{M}_{\boldsymbol{\theta}_0}}\bigg|\frac{1}{n}\sum_{i=1}^n\frac{g_{i,j}(\boldsymbol{\theta}_0)}{1+{\mathbf w}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{M}_{\boldsymbol{\theta}_0}^*}(\boldsymbol{\theta}_0)}\bigg|\leq \nu\rho'(0^+) \] w.p.a.1, which implies (ref) holds. Thus $\tilde{\boldsymbol{\lambda}}_{0}=\arg\max_{\boldsymbol{\lambda}\in\Lambda_{0}^*}f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1.
Secondly, define $\tilde\Lambda_{0}=\{\boldsymbol{\lambda}\in {\mathbb{R}}^r: |\boldsymbol{\lambda}_{\mathcal{M}_{\boldsymbol{\theta}_0}^*}-\tilde\boldsymbol{\lambda}_{0,\mathcal{M}_{\boldsymbol{\theta}_0}^*}|_2 \leq O(\ell_n^{1/2} \alpha_n), |\boldsymbol{\lambda}_{\mathcal{M}^{*,{\rm c}}_{\boldsymbol{\theta}_0}}|_1\leq O(\ell_n \alpha_n) \}$. We will show $\tilde{\boldsymbol{\lambda}}_{0}=\arg \max_{\boldsymbol{\lambda}\in \tilde\Lambda_{0}} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. Recall $\max_{i\in[n],j\in[r]} |g_{i,j} (\boldsymbol{\theta}_0)| = O_{\rm p}(n^{1/\gamma})$ and $|\tilde\boldsymbol{\lambda}_{0}|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n)$. Since $\ell_n\alpha_n=o(n^{-1/\gamma})$, we have
For any $\boldsymbol{\lambda} \in \tilde\Lambda_{0}$, denote by $\mathring\boldsymbol{\lambda}=(\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} }_{\mathcal{M}_{\boldsymbol{\theta}_0}^*},{\mathbf 0}^{\mathrm{\scriptscriptstyle \top} })^{\mathrm{\scriptscriptstyle \top} }$ the projection of $\boldsymbol{\lambda}=(\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} }_{\mathcal{M}_{\boldsymbol{\theta}_0}^*},\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} }_{\mathcal{M}^{*,{\rm c}}_{\boldsymbol{\theta}_0}})^{\mathrm{\scriptscriptstyle \top} }$ onto $\Lambda_{0}^*$. Write $\boldsymbol{\lambda}=(\lambda_1,\ldots,\lambda_r)^{\mathrm{\scriptscriptstyle \top} }$. By the Taylor expansion, it holds that \[ f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)-f_n(\mathring\boldsymbol{\lambda};\boldsymbol{\theta}_0)=\frac{1}{n}\sum_{i=1}^n\frac{{\mathbf g}_{i}(\boldsymbol{\theta}_0)^{\mathrm{\scriptscriptstyle \top} } (\boldsymbol{\lambda}-\mathring\boldsymbol{\lambda})}{1+\boldsymbol{\lambda}_*^{\mathrm{\scriptscriptstyle \top} }{\mathbf g}_{i}(\boldsymbol{\theta}_0)}-\sum_{j\in\mathcal{M}^{*,{\rm c}}_{\boldsymbol{\theta}_0}}P_{\nu}(|\lambda_j|)\,, \] where $\boldsymbol{\lambda}_*$ is on the jointing line between $\boldsymbol{\lambda}$ and $\mathring\boldsymbol{\lambda}$. Let $\check{\boldsymbol{\lambda}}_{0}=\arg \max_{\boldsymbol{\lambda}\in \tilde\Lambda_{0}} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$. Due to $\tilde\boldsymbol{\lambda}_{0} \in {\rm int} (\tilde\Lambda_{0})$, then $f_n(\tilde{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)\leq f_n(\check{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)$. For any $\boldsymbol{\lambda}\in\tilde\Lambda_{0}$, due to $\sum_{j\in\mathcal{M}^{*,{\rm c}}_{\boldsymbol{\theta}_0}}P_{\nu}(|\lambda_j|) \geq \nu\rho'(0^+) |\boldsymbol{\lambda}_{\mathcal{M}^{*,{\rm c}}_{\boldsymbol{\theta}_0}}|_1$ and
then $ f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)-f_n(\mathring\boldsymbol{\lambda};\boldsymbol{\theta}_0) \leq \{-(1-C_*)\nu\rho'(0^+)+ O_{\rm p}(\ell_n\alpha_n)\}|\boldsymbol{\lambda}_{\mathcal{M}^{*,{\rm c}}_{\boldsymbol{\theta}_0}}|_1$ for any $\boldsymbol{\lambda}\in\tilde{\Lambda}_0$, where the term $O_{\rm p}(\ell_n\alpha_n)$ holds uniformly over $\boldsymbol{\lambda}\in\tilde{\Lambda}_0$. Since $\ell_n \alpha_n=o(\nu)$, we have $\check{\boldsymbol{\lambda}}_{0,\mathcal{M}_{\boldsymbol{\theta}_0}^{*,{\rm c}}}={\mathbf 0}$ w.p.a.1, which implies $\check{\boldsymbol{\lambda}}_0\in{\rm int}(\Lambda_0^*)$ w.p.a.1. Recall $\tilde \boldsymbol{\lambda}_{0}=\arg\max_{\boldsymbol{\lambda}\in\Lambda_{0}^*}f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. Then $f_n(\tilde{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)\geq f_n(\check{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)$ w.p.a.1. Therefore, $f_n(\tilde{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)= f_n(\check{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)$ w.p.a.1. By the concavity of $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.r.t $\boldsymbol{\lambda}$, we have $\check{\boldsymbol{\lambda}}_0=\tilde\boldsymbol{\lambda}_{0}$ w.p.a.1, which indicates that $\tilde{\boldsymbol{\lambda}}_0$ is a local maximizer for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. Then $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta}_0)=\tilde\boldsymbol{\lambda}_{0}$ and ${\rm supp}\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta}_0)\}\subset \mathcal{M}_{\boldsymbol{\theta}_0}$ w.p.a.1. $\Box$
In this case, we will show ${\mathbf 0} \in \mathbb{R}^r$ is a local maximizer for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. Due to the concavity of $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.r.t $\boldsymbol{\lambda}$, we then have $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta}_0) = {\mathbf 0}$ w.p.a.1, which implies ${\rm supp}\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta}_0)\} \subset \mathcal{M}_{\boldsymbol{\theta}_0}$ w.p.a.1. Let $\breve{\boldsymbol{\lambda}}_0 = \arg\max_{\boldsymbol{\lambda} \in \breve{\Lambda}_0} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$, where $\breve{\Lambda}_0 =\{ \boldsymbol{\lambda}\in\mathbb{R}^r: |\lambda_{j_0}|\leq (\log n)^{-1}n^{-1/\gamma}, \boldsymbol{\lambda}_{[r]\setminus \{j_0\}}={\mathbf 0} \}$ with $j_0 = \arg\max_{j\in[r]}\mathbb{E}\{g^2_{i,j}(\boldsymbol{\theta}_0)\}$. It follows from Condition (ref)(a) that $\max_{i\in[n]}|g_{i,j_0} (\boldsymbol{\theta}_0)| = O_{\rm p}(n^{1/\gamma})$. Hence, $\max_{i\in[n]} |\breve{\boldsymbol{\lambda}}_0^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i} (\boldsymbol{\theta}_0)|_2 = O_{\rm p}\{(\log n)^{-1}\} = o_{\rm p}(1)$. Write $\breve{\boldsymbol{\lambda}}_0 =(\breve{\lambda}_{0,1},\ldots,\breve{\lambda}_{0,r})^{{\mathrm{\scriptscriptstyle \top} }}$. By the Taylor expansion, we have
for some $C\in(0,1)$. Notice that $|\mathbb{E}_n\{g_{i,j_0}^2(\boldsymbol{\theta}_0)\} - \mathbb{E}\{g_{i,j_0}^2(\boldsymbol{\theta}_0)\}|= O_{\rm p}(n^{-1/2})$. By Condition (ref)(b), we have $\mathbb{E}_n\{g_{i,j_0}^2(\boldsymbol{\theta}_0)\} \geq \mathbb{E}\{g_{i,j_0}^2(\boldsymbol{\theta}_0)\} - o_{\rm p}(1)\geq 2 K_3/3$ w.p.a.1 for $K_3$ specified in Condition (ref)(b). Thus $0 \leq |\breve{\lambda}_{0,j_0}||\bar{g}_{j_0}(\boldsymbol{\theta}_0)| - 4^{-1} K_3 |\breve{\lambda}_{0,j_0}|^2$ w.p.a.1. Since $|\bar{g}_{j_0}(\boldsymbol{\theta}_0)| = O_{\rm p}(n^{-1/2})$, then $|\breve{\lambda}_{0,j_0}| = O_{\rm p}(n^{-1/2}) = o_{\rm p}\{(\log n)^{-1}n^{-1/\gamma}\}$. It then holds w.p.a.1 that
where $\breve{\eta}_{j_0} = \nu\rho'(|\breve\lambda_{j_0}|;\nu){\mbox{\rm sgn}}(\breve{\lambda}_{j_0})$ if $\breve\lambda_{j_0}\neq0$ and $\breve{\eta}_{j_0} \in[-\nu\rho'(0^+), \nu\rho'(0^+)]$ if $\breve\lambda_{j_0}=0$. Due to $|\breve{\lambda}_{0,j_0} g_{i,j_0}(\boldsymbol{\theta}_0)| = o_{\rm p}(1)$, then
with
By Conditions (ref)(a), we have $|R_{j_0}| = O_{\rm p}(|\breve{\lambda}_{0,j_0}|) = O_{\rm p}(n^{-1/2})$. Together with $|\bar{g}_{j_0}(\boldsymbol{\theta}_0)| = O_{\rm p}(n^{-1/2})$, (ref) leads to $|\breve{\eta}_{j_0}| = O_{\rm p}(n^{-1/2}) = o_{\rm p}(\nu)$. Then $\breve\lambda_{j_0}=0$ w.p.a.1, which implies $\breve{\boldsymbol{\lambda}}_0 = {\mathbf 0}$ w.p.a.1. In the sequel, we will show that $\breve{\boldsymbol{\lambda}}_0$ is a local maximizer for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1.
Firstly, define $\breve{\Lambda}_0^* = \{\boldsymbol{\lambda}\in {\mathbb{R}}^r:|\boldsymbol{\lambda}_{\mathcal{H}}|_2 \leq \varepsilon,\boldsymbol{\lambda}_{\mathcal{H}^{\rm c}}={\mathbf 0}\}$ for some sufficiently small constant $\varepsilon>0$, where $j_0 \in \mathcal{H} \subset [r]$ with $1 < |\mathcal{H}|\leq \ell_n$. For $\breve{\boldsymbol{\lambda}}_0$ defined before, we will prove $\breve{\boldsymbol{\lambda}}_0 =\arg \max_{\boldsymbol{\lambda}\in \breve{\Lambda}_0^*} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. Since $\breve{\boldsymbol{\lambda}}_0 \in \breve{\Lambda}_0$ and $j_0 \in \mathcal{H}$, we know $\breve{\boldsymbol{\lambda}}_0 \in \breve{\Lambda}_0^*$ for sufficiently large $n$. Restricted on $\boldsymbol{\lambda} \in \breve{\Lambda}_0^*$, by the concavity of $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.r.t $\boldsymbol{\lambda}$, it suffices to show that $\breve{{\mathbf w}} = \breve{\boldsymbol{\lambda}}_{0,\mathcal{H}}=: (\breve{w}_{1},\ldots,\breve{w}_{|\mathcal{H}|})^{\mathrm{\scriptscriptstyle \top} } \in {\mathbb{R}}^{|\mathcal{H}|}$ satisfies the equation
w.p.a.1, where $\breve{\boldsymbol{\eta}}^*=(\breve\eta^*_1,\ldots,\breve\eta^*_{|\mathcal{H}|})^{\mathrm{\scriptscriptstyle \top} }$ with $\breve{\eta}^*_j=\nu\rho'(|\breve{w}_j|;\nu){\mbox{\rm sgn}}(\breve{w}_j)$ for $\breve{w}_j\neq0$ and $\breve{\eta}^*_j \in [-\nu\rho'(0^+),\nu\rho'(0^+)]$ for $\breve{w}_j=0$. Recall $j_0\in\mathcal{H}$. Without loss of generality, we assume $j_0$ is the first component in $\mathcal{H}$. By (ref), we know $0= n^{-1}\sum_{i=1}^{n} g_{i,j_0}(\boldsymbol{\theta}_0)/\{1+\breve{{\mathbf w}}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{H}}(\boldsymbol{\theta}_0)\} - \breve{\eta}^*_{1}$ holds. Since $\breve{\boldsymbol{\lambda}}_0 = {\mathbf 0}$ w.p.a.1, then $\breve{{\mathbf w}}= {\mathbf 0}$ w.p.a.1, which implies it holds w.p.a.1 that
for any $j \in \mathcal{H} \backslash \{j_0\}$. By the moderate deviation of self-normalized sums JingShaoWang2003, $ |\bar{{\mathbf g}}(\boldsymbol{\theta}_0)|_{\infty} =O_{\rm p}(\alpha_n)$. Due to $\alpha_n = o(\nu)$, then \[ \max_{j \in \mathcal{H} \backslash \{j_0\}}\bigg|\frac{1}{n}\sum_{i=1}^n\frac{g_{i,j}(\boldsymbol{\theta}_0)}{1+\breve{{\mathbf w}}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{H}}(\boldsymbol{\theta}_0)}\bigg|\leq \nu\rho'(0^+) \] w.p.a.1, which implies (ref) holds. Thus $\breve{\boldsymbol{\lambda}}_0 =\arg \max_{\boldsymbol{\lambda}\in \breve{\Lambda}_0^*} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1.
Secondly, define $\bar{\Lambda}_{0}=\{\boldsymbol{\lambda}\in {\mathbb{R}}^r: |\boldsymbol{\lambda}_{\mathcal{H}}-\breve\boldsymbol{\lambda}_{0,\mathcal{H}}|_2 \leq O(\ell_n^{1/2} \alpha_n), |\boldsymbol{\lambda}_{\mathcal{H}^{\rm c}}|_1\leq O(\ell_n \alpha_n) \}$. We will show $\breve{\boldsymbol{\lambda}}_{0}=\arg \max_{\boldsymbol{\lambda}\in \bar{\Lambda}_{0}} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. Recall $\max_{i\in[n],j\in[r]} |g_{i,j} (\boldsymbol{\theta}_0)| = O_{\rm p}(n^{1/\gamma})$ and $|\breve{\boldsymbol{\lambda}}_{0}|_2=|\breve{\lambda}_{0,j_0}| = O_{\rm p}(n^{-1/2})$. Since $\ell_n\alpha_n=o(n^{-1/\gamma})$, we have
For any $\boldsymbol{\lambda} \in \bar{\Lambda}_{0}$, denote by $\mathring\boldsymbol{\lambda}=(\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} }_{\mathcal{H}},{\mathbf 0}^{\mathrm{\scriptscriptstyle \top} })^{\mathrm{\scriptscriptstyle \top} }$ the projection of $\boldsymbol{\lambda}=(\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} }_{\mathcal{H}},\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} }_{\mathcal{H}^{\rm c}})^{\mathrm{\scriptscriptstyle \top} }$ onto $\breve{\Lambda}_0^*$. Write $\boldsymbol{\lambda}=(\lambda_1,\ldots,\lambda_r)^{\mathrm{\scriptscriptstyle \top} }$. By the Taylor expansion, it holds that \[ f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)-f_n(\mathring\boldsymbol{\lambda};\boldsymbol{\theta}_0)=\frac{1}{n}\sum_{i=1}^n\frac{{\mathbf g}_{i}(\boldsymbol{\theta}_0)^{\mathrm{\scriptscriptstyle \top} } (\boldsymbol{\lambda}-\mathring\boldsymbol{\lambda})}{1+\boldsymbol{\lambda}_*^{\mathrm{\scriptscriptstyle \top} }{\mathbf g}_{i}(\boldsymbol{\theta}_0)}-\sum_{j\in\mathcal{H}^{\rm c}}P_{\nu}(|\lambda_j|)\,, \] where $\boldsymbol{\lambda}_*$ is on the jointing line between $\boldsymbol{\lambda}$ and $\mathring\boldsymbol{\lambda}$. Let $\check{\boldsymbol{\lambda}}_{0}=\arg \max_{\boldsymbol{\lambda}\in \bar\Lambda_{0}} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$. Due to $\breve\boldsymbol{\lambda}_{0} \in {\rm int} (\bar\Lambda_{0})$, then $f_n(\breve{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)\leq f_n(\check{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)$. For any $\boldsymbol{\lambda} \in \bar\Lambda_{0}$, due to $\sum_{j\in\mathcal{H}^{\rm c}}P_{\nu}(|\lambda_j|) \geq \nu\rho'(0^+) |\boldsymbol{\lambda}_{\mathcal{H}^{\rm c}}|_1$ and
then $ f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)-f_n(\mathring\boldsymbol{\lambda};\boldsymbol{\theta}_0) \leq \{-\nu\rho'(0^+)+ O_{\rm p}(\ell_n\alpha_n)\}|\boldsymbol{\lambda}_{\mathcal{H}^{\rm c}}|_1$ for any $\boldsymbol{\lambda}\in\bar{\Lambda}_0$, where the term $O_{\rm p}(\ell_n\alpha_n)$ holds uniformly over $\boldsymbol{\lambda}\in\bar{\Lambda}_0$. Since $\ell_n \alpha_n=o(\nu)$, we have $\check{\boldsymbol{\lambda}}_{0,\mathcal{H}^{\rm c}}={\mathbf 0}$ w.p.a.1, which implies $\check{\boldsymbol{\lambda}}_0\in{\rm int}(\breve{\Lambda}_0^*)$ w.p.a.1. Recall $\breve{\boldsymbol{\lambda}}_0 =\arg \max_{\boldsymbol{\lambda}\in \breve{\Lambda}_0^*} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. Then $f_n(\breve{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)\geq f_n(\check{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)$ w.p.a.1. Therefore, $f_n(\breve{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0) = f_n(\check{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)$ w.p.a.1. By the concavity of $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.r.t $\boldsymbol{\lambda}$, we have $\check{\boldsymbol{\lambda}}_0=\breve{\boldsymbol{\lambda}}_0$ w.p.a.1, which indicates that ${\mathbf 0}$ is a local maximizer for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. $\Box$
Same as the proof of Lemma (ref), we only need to show that there exists a local maximizer satisfying the results stated in the lemma. Recall $\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}^*=\{j\in[r]: |\bar{g}_j(\hat{\boldsymbol{\theta}}_n)| \geq C_*\nu \rho'(0^{+}) \}$ for some $C_* \in (0,1)$, and $\mathbb{P}(\max _{\boldsymbol{\theta} \in \boldsymbol{\Theta}:\,|\boldsymbol{\theta}-\boldsymbol{\theta}_0|_2 \leq c_n}|\mathcal{M}_{\boldsymbol{\theta}}^*|\leq\ell_n)\rightarrow1$ for some $c_n\rightarrow 0$ satisfying $\nu c_n^{-1} \rightarrow 0$. For $\tilde{c} \in (C_*,1)$ given in Condition (ref)(a), write $\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}:=\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})=\{ j\in[r]:|\bar{g}_j(\hat{\boldsymbol{\theta}}_n)| \geq \tilde{c}\nu\rho'(0^{+}) \}$. By Proposition (ref), $|\hat{\boldsymbol{\theta}}_n-\boldsymbol{\theta}_0|_\infty=O_{\rm p}(\nu)$. Notice that $p$ is fixed. Then $|\hat{\boldsymbol{\theta}}_n-\boldsymbol{\theta}_0|_2=O_{\rm p}(\nu)$ which implies $\ell_n \geq |\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}^*| \geq |\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}|$ w.p.a.1. Restricted on $\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}$, we select $\delta_n$ satisfying $\delta_n=o(\ell_n^{-1/2} n^{-1/\gamma})$ and $\ell_n^{1/2} \alpha_n=o(\delta_n) $, which can be guaranteed by $\ell_n\alpha_n=o(n^{-1/\gamma})$. Let $\Lambda_n=\{\boldsymbol{\lambda}\in {\mathbb{R}}^{r}:|\boldsymbol{\lambda}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}|_2\leq \delta_n, \boldsymbol{\lambda}_{\mathcal{M}^{\rm c}_{\hat{\boldsymbol{\theta}}_n}}={\mathbf 0} \}$ and $\tilde{\boldsymbol{\lambda}}_n=(\tilde{\lambda}_{n,1},\ldots,\tilde{\lambda}_{n,r})^{\mathrm{\scriptscriptstyle \top} }=\arg \max _{\boldsymbol{\lambda} \in \Lambda_n} f_n(\boldsymbol{\lambda}; \hat{\boldsymbol{\theta}}_n)$. By the Taylor expansion, we have
for some $C \in (0,1)$. By Proposition (ref), Lemma (ref) and Condition (ref)(b), if $\log r=o(n^{1/3})$, $\ell_n\nu^2=o(1)$ and $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$, we have $\lambda_{\min}\{\widehat{{\mathbf V}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}} (\hat{\boldsymbol{\theta}}_n)\}$ is uniformly bounded away from zero w.p.a.1. Therefore, it holds w.p.a.1 that $$ 0 \leq \tilde{\boldsymbol{\lambda}}_{n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}^{\mathrm{\scriptscriptstyle \top} } \big[\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)\}\big] - 4^{-1}K_3 |\tilde{\boldsymbol{\lambda}}_{n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}|_2^2$$ with $K_3$ specified in Condition (ref)(b), which implies $|\tilde\boldsymbol{\lambda}_{n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}|_2 \leq 4K_3^{-1} |\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)\}|_2$ w.p.a.1.
Select $\boldsymbol{\lambda}_n^* \in {\mathbb{R}}^r$ satisfying $\boldsymbol{\lambda}^*_{n,\mathcal{M}^{\rm c}_{\hat{\boldsymbol{\theta}}_n}}={\mathbf 0}$ and $$\boldsymbol{\lambda}^*_{n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}=\frac{\delta_n [\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)\}]}{|\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)\}|_2}\,.$$ Then $\boldsymbol{\lambda}_n^* \in \Lambda_n$. As shown in the proof of Proposition (ref), $\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta}_0)}f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0) = O_{\rm p}(\ell_n \alpha_n^2)=o_{\rm p}(\delta_n^2)$, which implies $\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\hat{\boldsymbol{\theta}}_n)}f_n(\boldsymbol{\lambda};\hat{\boldsymbol{\theta}}_n)=o_{\rm p}(\delta_n^2)$. Write $\boldsymbol{\lambda}_n^*=(\lambda_{n,1}^*,\ldots,\lambda_{n,r}^*)^{\mathrm{\scriptscriptstyle \top} }$. Notice that $\Lambda_n\subset\hat{\Lambda}_n(\hat{\boldsymbol{\theta}}_n)$ w.p.a.1. By the Taylor expansion, it holds w.p.a.1 that
for some $\bar C, c_j \in (0,1)$, where the last inequality follows from the condition that $P_{\nu}(\cdot)$ has bounded second-order derivative around $0$. For any $j \in \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}$, we have ${\mbox{\rm sgn}}(\lambda_{n,j}^*)={\mbox{\rm sgn}}\{\bar g_j(\hat{\boldsymbol{\theta}}_n)\}$ if $|\bar {g}_j(\hat{\boldsymbol{\theta}}_n)| > \nu \rho'(0^{+})$, and $\bar g_j(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar g_j(\hat{\boldsymbol{\theta}}_n)\} = 0 = \lambda_{n,j}^*$ if $|\bar {g}_j(\hat{\boldsymbol{\theta}}_n)| = \nu \rho'(0^{+})$. Thus, $$\lambda_{n,j}^* \big\{\bar g_j(\hat{\boldsymbol{\theta}}_n) - \nu \rho'(0^+) {\mbox{\rm sgn}}(\lambda_{n,j}^*)\big\} = \lambda_{n,j}^* \big[\bar g_j(\hat{\boldsymbol{\theta}}_n) - \nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar g_j(\hat{\boldsymbol{\theta}}_n)\}\big]$$ for any $j \in \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}$ with $|\bar {g}_j(\hat{\boldsymbol{\theta}}_n)| \geq \nu \rho'(0^{+})$. By Condition (ref)(a), $\{j \in [r]: \tilde{c} \nu\rho'(0^{+}) \leq |\bar {g}_j(\hat{\boldsymbol{\theta}}_n)| < \nu\rho'(0^{+})\} = \emptyset$ w.p.a.1. Recall $\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}=\{j \in [r]: |\bar {g}_j(\hat{\boldsymbol{\theta}}_n)| \geq \tilde{c} \nu \rho'(0^{+})\}$. We then have w.p.a.1 that
Thus, $ |\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)\}|_2 =O_{\rm p}(\delta_n) $. For any $\epsilon_n \rightarrow 0$, select $\boldsymbol{\lambda}_n^{**}$ such that $\boldsymbol{\lambda}_{n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}^{**}=\epsilon_n [\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)\}]$ and $\boldsymbol{\lambda}^{**}_{n,\mathcal{M}^{\rm c}_{\hat{\boldsymbol{\theta}}_n}}={\mathbf 0}$. Then $|\boldsymbol{\lambda}_n^{**}|_2=o_{\rm p}(\delta_n)$. Due to $f_n(\boldsymbol{\lambda}_n^{**};\hat{\boldsymbol{\theta}}_n) \leq \max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\hat{\boldsymbol{\theta}}_n)}f_n(\boldsymbol{\lambda};\hat{\boldsymbol{\theta}}_n) \leq \max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta}_0)}f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)=O_{\rm p}(\ell_n \alpha_n^2)$, using the same arguments given above, we have
Hence, $\epsilon_n |\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)\}|^2_2=O_{\rm p}(\ell_n \alpha_n^2)$. Since we can select arbitrary slow $\epsilon_n \rightarrow 0$, it holds that
which implies $ |\tilde{\boldsymbol{\lambda}}_n|_2 = |\tilde{\boldsymbol{\lambda}}_{n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n) = o_{\rm p}(\delta_n) $. Write $\tilde{\boldsymbol{\lambda}}_{n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}=(\tilde{\lambda}_{n,1},\ldots,\tilde{\lambda}_{n,|\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}|})^{\mathrm{\scriptscriptstyle \top} }$. We have w.p.a.1 that $${\mathbf 0}= \frac{1}{n}\sum_{i=1}^{n} \frac{{\mathbf g}_{i,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)}{1+\tilde{\boldsymbol{\lambda}}_{n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)} - \tilde{\boldsymbol{\eta}}\,,$$ where $\tilde{\boldsymbol{\eta}}=(\tilde\eta_1,\ldots,\tilde\eta_{|\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}|})^{\mathrm{\scriptscriptstyle \top} }$ with $\tilde\eta_j=\nu\rho'(|\tilde\lambda_{n,j}|;\nu){\mbox{\rm sgn}}(\tilde{\lambda}_{n,j})$ for $\tilde\lambda_{n,j}\neq0$ and $\tilde\eta_j\in[-\nu\rho'(0^+),\nu\rho'(0^+)]$ for $\tilde\lambda_{n,j}=0$. Identical to (ref), we have $\tilde{\boldsymbol{\eta}}=\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)+{\mathbf R}$ for some $|\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}|$-dimensional vector ${\mathbf R}$. Applying the same arguments for deriving the rate of $R_j$ in (ref), it holds that $|{\mathbf R}|_{\infty}=O_{\rm p}(\ell_n^{1/2} \alpha_n)$. Since $\ell_n\alpha_n=o(\nu) $, we then have ${\mbox{\rm sgn}} (\tilde{\lambda}_{n,j})={\mbox{\rm sgn}}\{\bar{g}_j(\hat{\boldsymbol{\theta}}_n)\}$ for any $j \in \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}$ with $\tilde{\lambda}_{n,j} \neq 0$ w.p.a.1. Using the arguments in Section (ref) for showing $\tilde\boldsymbol{\lambda}_{0}$ is a local maximizer for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1, we can prove $\tilde\boldsymbol{\lambda}_n$ is a local maximizer for $f_n(\boldsymbol{\lambda};\hat{\boldsymbol{\theta}}_n)$ w.p.a.1, which implies $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)=\tilde{\boldsymbol{\lambda}}_n$ w.p.a.1. We then have Lemma (ref). $\Box$
Recall $\hat\boldsymbol{\lambda}(\boldsymbol{\theta})=\arg\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})} f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$. Then $\hat{\boldsymbol{\theta}}_n$ and $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n) =(\hat\lambda_1,\ldots,\hat\lambda_r)^{\mathrm{\scriptscriptstyle \top} }$ satisfy
where $\hat{\boldsymbol{\eta}}=(\hat{\eta}_{1},\ldots,\hat{\eta}_{r})^{\mathrm{\scriptscriptstyle \top} }$ with $\hat{\eta}_j=\nu \rho' (|\hat{\lambda}_j|;\nu) {\mbox{\rm sgn}}(\hat{\lambda}_j)$ for $\hat{\lambda}_j\neq0$ and $\hat{\eta}_j \in [-\nu \rho'(0^{+}),\nu \rho'(0^{+})]$ for $\hat{\lambda}_j=0$. Recall $\mathcal{R}_n={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$. Restricted on $\mathcal{R}_n$, for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ and $\boldsymbol{\zeta}=(\zeta_{1},\ldots,\zeta_{|{ \mathcal{R}_n}|})^{\mathrm{\scriptscriptstyle \top} } \in {\mathbb{R}}^{|{ \mathcal{R}_n}|}$ with each $\zeta_j \neq 0$, define $$ {\mathbf m}(\boldsymbol{\zeta},\boldsymbol{\theta})=\frac{1}{n} \sum _{i=1}^{n} \frac{{\mathbf g}_{i,{ \mathcal{R}_n}}(\boldsymbol{\theta})}{1+\boldsymbol{\zeta}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,{ \mathcal{R}_n}}(\boldsymbol{\theta})}-{\mathbf w}\,,$$ where ${\mathbf w}=(w_{1},\ldots,w_{|{ \mathcal{R}_n}|})^{\mathrm{\scriptscriptstyle \top} }$ with $w_j=\nu \rho' (|\zeta_j|;\nu) {\mbox{\rm sgn}}(\zeta_j)$. From (ref), we know $\hat{\boldsymbol{\lambda}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)$ and $\hat{\boldsymbol{\theta}}_n$ satisfy ${\mathbf m}\{\hat{\boldsymbol{\lambda}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n), \hat{\boldsymbol{\theta}}_n\}= {\mathbf 0}$. By the implicit function theorem [Theorem 9.28 of Rudin1976], for all $\boldsymbol{\theta}$ in a small neighborhood of $\hat{\boldsymbol{\theta}}_n$, denoted by $\mathcal{U}(\hat{\boldsymbol{\theta}}_n)$, there exists a $\boldsymbol{\zeta}(\boldsymbol{\theta})$ such that ${\mathbf m}\{\boldsymbol{\zeta}(\boldsymbol{\theta}),\boldsymbol{\theta}\}={\mathbf 0}$, $\boldsymbol{\zeta}(\hat{\boldsymbol{\theta}}_n)=\hat{\boldsymbol{\lambda}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)$ and $\boldsymbol{\zeta}(\boldsymbol{\theta})$ is continuously differentiable in $\boldsymbol{\theta}\in \mathcal{U}(\hat{\boldsymbol{\theta}}_n)$. By Condition (ref)(b), the event ${\mathcal{E}}=\{\max_{j \in \mathcal{R}_n^{\rm c}} |\hat{\eta}_j| < \nu\rho'(0^+)\}$ holds w.p.a.1. Restricted on ${\mathcal{E}}$, let $\varsigma_n=\nu\rho'(0^+) - \max_{j \in \mathcal{R}_n^{\rm c}} |\hat{\eta}_j|$ and define $\boldsymbol{\Theta}_*=\{\boldsymbol{\theta} \in \mathcal{U}(\hat{\boldsymbol{\theta}}_n):\, |\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_1 \leq o[\min\{\varsigma_n, \chi_n\}], |\boldsymbol{\zeta}(\boldsymbol{\theta})-\boldsymbol{\zeta}(\hat{\boldsymbol{\theta}}_n)|_1 \leq o [\min\{\varsigma_n, \ell_n^{1/2} \alpha_n\}]\}$ for some $\chi_n>0$. Since all the components of $\boldsymbol{\zeta}(\hat{\boldsymbol{\theta}}_n)$ are nonzero and $\boldsymbol{\zeta}(\boldsymbol{\theta})$ is continuously differentiable in $\hat{\boldsymbol{\theta}}_n$, we can select sufficiently small $\chi_n$ such that all the components of $\boldsymbol{\zeta}(\boldsymbol{\theta})$ are nonzero for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}_*$. For any $\boldsymbol{\theta} \in \boldsymbol{\Theta}_*$, let $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})= \{\tilde\lambda_1(\boldsymbol{\theta}), \ldots, \tilde\lambda_r(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} } \in {\mathbb{R}}^r$ satisfy $\tilde\boldsymbol{\lambda}_{ \mathcal{R}_n}(\boldsymbol{\theta})=\boldsymbol{\zeta}(\boldsymbol{\theta})$ and $\tilde\boldsymbol{\lambda}_{\mathcal{R}_n^{\rm c}}(\boldsymbol{\theta})={\mathbf 0}$. Since ${\mathbf m}\{\boldsymbol{\zeta}(\boldsymbol{\theta}),\boldsymbol{\theta}\}={\mathbf 0}$, $\tilde\boldsymbol{\lambda}_{ \mathcal{R}_n}(\boldsymbol{\theta})=\boldsymbol{\zeta}(\boldsymbol{\theta})$ and $\tilde\boldsymbol{\lambda}_{\mathcal{R}_n^{\rm c}}(\boldsymbol{\theta})={\mathbf 0}$ for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}_*$, then $$ 0 = \frac{1}{n}\sum _{i=1}^{n} \frac{g_{i,j}(\boldsymbol{\theta})}{1+\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i}(\boldsymbol{\theta})} - \nu \rho' \{|\tilde{\lambda}_j(\boldsymbol{\theta})|;\nu\} {\mbox{\rm sgn}}\{\tilde{\lambda}_j(\boldsymbol{\theta})\}$$ for any $j \in \mathcal{R}_n$. For any $\boldsymbol{\theta} \in \boldsymbol{\Theta}_*$ and $j \in \mathcal{R}_n^{\rm c}$, by the Taylor expansion, we have
where $\check{\boldsymbol{\theta}}$ is lying on the jointing line between $\boldsymbol{\theta}$ and $\hat{\boldsymbol{\theta}}_n$, and $\check{\boldsymbol{\lambda}}$ is lying on the jointing line between $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})$ and $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)$. By Lemma (ref), $|\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n)$ and $|\mathcal{R}_n|\leq \ell_n$ w.p.a.1. Then $|\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta})|_2 = |\boldsymbol{\zeta}(\boldsymbol{\theta})|_2 \leq |\boldsymbol{\zeta}(\hat{\boldsymbol{\theta}}_n)|_2 + |\boldsymbol{\zeta}(\boldsymbol{\theta})-\boldsymbol{\zeta}(\hat{\boldsymbol{\theta}}_n)|_2 = O_{\rm p}(\ell_n^{1/2} \alpha_n)$, which implies $|\check{\boldsymbol{\lambda}}|_2 = O_{\rm p}(\ell_n^{1/2} \alpha_n)$. Together with Condition (ref)(a) and $\ell_n \alpha_n=o(n^{-1/\gamma})$, it yields that $\max_{i \in [n]} \{ |\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i}(\check{\boldsymbol{\theta}})| + |\check{\boldsymbol{\lambda}}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i}(\hat{\boldsymbol{\theta}}_n)|\} = o_{\rm p}(1)$. By Conditions (ref)(a) and (ref)(c), we have
It follows from the Cauchy-Schwarz inequality that
By (ref), for any $\boldsymbol{\theta}\in\boldsymbol{\Theta}_*$, we know
holds uniformly over $j \in \mathcal{R}_n^{\rm c}$. Due to ${\mathbb{P}}({\mathcal{E}}) \rightarrow 1$, we have $$ \max_{j \in \mathcal{R}_n^{\rm c}}\bigg| \frac{1}{n}\sum _{i=1}^{n} \frac{g_{i,j}(\boldsymbol{\theta})}{1+\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i}(\boldsymbol{\theta})} \bigg| \leq \nu\rho'(0^+)$$ w.p.a.1. Therefore, $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})$ and $\boldsymbol{\theta}$ satisfy the score equation $\nabla_{\boldsymbol{\lambda}} f_n\{\tilde\boldsymbol{\lambda}(\boldsymbol{\theta}); \boldsymbol{\theta}\}={\mathbf 0}$ for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}_*$ w.p.a.1. By the concavity of $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ w.r.t $\boldsymbol{\lambda}$, we have $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})=\arg\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})} f_n(\boldsymbol{\lambda};\boldsymbol{\theta})=\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})$ for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}_*$ w.p.a.1. Hence, $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})$ is continuously differentiable at $\hat{\boldsymbol{\theta}}_n$ and $ [\nabla_{\boldsymbol{\theta}} \hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)]_{\mathcal{R}_n^{\rm c},[p]}={\mathbf 0}$ w.p.a.1. $\Box$
The proof is almost identical to that of Lemma 2 in Changetal2018. From Lemma (ref), we have $|\hat{\boldsymbol{\lambda}}|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n)$. Recall $p$ is fixed in our current setting. We only need to replace the convergence rate of $|\hat{\boldsymbol{\lambda}}|_2$ in the proof of Lemma 2 in Changetal2018 by $O_{\rm p}(\ell_n^{1/2} \alpha_n)$ and also set $(s,\omega_n)$ there as $(p,1)$ and all the arguments still hold. $\Box$
The proof is almost identical to that of Lemma 3 in Changetal2018. Since $p$ is fixed, we only need to replace $\{\omega_n,\varpi_n,b_n^{1/(2\beta)},s\}$ in the proof of Lemma 3 in Changetal2018 by $(1,1,\nu,p)$ and all the arguments still hold. $\Box$
Recall $\boldsymbol{\Gamma}_{\mathcal{F}}(\boldsymbol{\theta}_0)={\mathbb{E}} \{\nabla_{\boldsymbol{\theta}} {\mathbf g}_{i,\mathcal{F}}(\boldsymbol{\theta}_0) \}$ and ${\mathbf V}_{\mathcal{F}} (\boldsymbol{\theta}_0)= {\mathbb{E}} \{ {\mathbf g}_{i,\mathcal{F}} (\boldsymbol{\theta}_0)^{\otimes2} \}$. For any ${\mathbf t} \in {\mathbb{R}}^p$ with $|{\mathbf t}|_2=1$, let $Z_{i,\mathcal{F}}={\mathbf t}^{\mathrm{\scriptscriptstyle \top} } {\mathbf H}_{\mathcal{F}}^{-1/2} \boldsymbol{\Gamma}_{\mathcal{F}}(\boldsymbol{\theta}_0)^{\mathrm{\scriptscriptstyle \top} } {\mathbf V}^{-1}_{\mathcal{F}} (\boldsymbol{\theta}_0) {\mathbf g}_{i,\mathcal{F}} (\boldsymbol{\theta}_0)$ with ${\mathbf H}_{\mathcal{F}}=\{\boldsymbol{\Gamma}_{\mathcal{F}}(\boldsymbol{\theta}_0)^{\mathrm{\scriptscriptstyle \top} } {\mathbf V}^{-1/2}_{\mathcal{F}} (\boldsymbol{\theta}_0)\}^{\otimes2}$. Write $G_{\mathcal{F}}= {\mathbb{E}}_n(Z_{i,\mathcal{F}})$ and $\hat{G}_{\mathcal{F}}= {\mathbf t}^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf H}}_{\mathcal{F}}^{-1/2} \widehat{\boldsymbol{\Gamma}}_{\mathcal{F}}(\hat{\boldsymbol{\theta}}_n)^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf V}}^{-1}_{\mathcal{F}} (\hat{\boldsymbol{\theta}}_n) \bar{{\mathbf g}}_{\mathcal{F}} (\boldsymbol{\theta}_0)$. It follows from the Berry-Esseen inequality that $$ \sup_{u \in {\mathbb{R}}} |{\mathbb{P}}(n^{1/2} G_{\mathcal{F}} \leq u) -\Phi(u)| \leq Cn^{-1/2} {\mathbb{E}}(|Z_{i,\mathcal{F}}|^3)$$ for some universal constant $C>0$. By the Cauchy-Schwarz inequality,
for $K_3$ given in Condition (ref)(b). By the Jensen's inequality, Condition (ref)(a) yields $ {\mathbb{E}}\{|{\mathbf g}_{i,\mathcal{F}}(\boldsymbol{\theta}_0)|_2^3\}\leq K_2^{3/\gamma}\ell_n^{3/2}$ for $K_2$ and $\gamma$ given in Condition (ref)(a), which implies $ {\mathbb{E}}(|Z_{i,\mathcal{F}}|^3) \leq K_3^{-3/2}{\mathbb{E}}\{|{\mathbf g}_{i,\mathcal{F}}(\boldsymbol{\theta}_0)|_2^3\} \leq K_2^{3/\gamma}K_3^{-3/2}\ell_n^{3/2} $. If $\ell_n=o(n^{1/3})$, we have $$ \sup_{\mathcal{F} \in \mathscr{F}} \sup_{u \in {\mathbb{R}}}|{\mathbb{P}}(n^{1/2} G_{\mathcal{F}} \leq u) -\Phi(u)| \rightarrow 0$$ as $n\rightarrow\infty$. By Conditions (ref)(b) and (ref) and Lemmas (ref) and (ref), it holds that $\sup_{\mathcal{F} \in \mathscr{F}}|n^{1/2}(\hat{G}_{\mathcal{F}}-G_{\mathcal{F}})|= O_{\rm p}\{\ell_n \nu (\log r)^{1/2}\} + O_{\rm p}\{\ell_n^{3/2} \alpha_n(\log r)^{1/2}\}$. For any constant $\delta>0$, due to $ {\mathbb{P}} (n^{1/2} \hat{G}_{\mathcal{F}} \leq u ) - \Phi(u) \leq {\mathbb{P}} (n^{1/2} G_{\mathcal{F}} \leq u+\delta) + {\mathbb{P}}\{|n^{1/2}(\hat{G}_{\mathcal{F}}-G_{\mathcal{F}})| \geq \delta \} - \Phi(u)$ and $ {\mathbb{P}}(n^{1/2} \hat{G}_{\mathcal{F}} \leq u) - \Phi(u) \geq {\mathbb{P}} (n^{1/2} G_{\mathcal{F}} \leq u-\delta) - {\mathbb{P}}\{|n^{1/2}(\hat{G}_{\mathcal{F}}-G_{\mathcal{F}})| \geq \delta \} - \Phi(u)$, it holds that
Notice that $\sup_{\mathcal{F} \in \mathscr{F}}|n^{1/2}(\hat{G}_{\mathcal{F}}-G_{\mathcal{F}})|=o_{\rm p}(1)$ and $\sup_{u \in {\mathbb{R}}} |\Phi(u+\delta)-\Phi(u-\delta)| \leq(2\pi^{-1})^{1/2}\delta$. Then it holds that $\limsup_{n\rightarrow\infty}\sup_{\mathcal{F} \in \mathscr{F}} \sup_{u \in {\mathbb{R}}} |{\mathbb{P}}(n^{1/2}\hat{G}_{\mathcal{F}} \leq u) - \Phi(u)|\leq (2\pi^{-1})^{1/2}\delta$. Due to the arbitrary selection of $\delta>0$, we have $\sup_{\mathcal{F} \in \mathscr{F}} \sup_{u \in {\mathbb{R}}} |{\mathbb{P}}(n^{1/2}\hat{G}_{\mathcal{F}} \leq u) - \Phi(u)|\rightarrow0$ as $n\rightarrow\infty$. $\Box$
Recall $\mathcal{C}_1=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2\leq \alpha_n\}$. For any $\boldsymbol{\theta} \in \mathcal{C}_1$, same as the proof of Lemma (ref), we only need to show that there exists a local maximizer satisfying the results stated in the lemma. For $\tilde{c}$ specified in Condition (ref)(a), we select $c \in (\tilde{c},1)$. We select $\delta_n$ satisfying $\delta_n=o(\ell_n^{-1/2} n^{-1/\gamma})$ and $\ell_n^{1/2} \alpha_n=o(\delta_n) $, which can be guaranteed by $\ell_n\alpha_n=o(n^{-1/\gamma})$. For each $\boldsymbol{\theta} \in \mathcal{C}_1$, define $\Lambda_{\boldsymbol{\theta}}=\{\boldsymbol{\lambda}\in {\mathbb{R}}^{r}:\,|\boldsymbol{\lambda}_{\mathcal{M}_{\boldsymbol{\theta}}(c)}|_2\leq \delta_n, \boldsymbol{\lambda}_{\mathcal{M}_{\boldsymbol{\theta}}(c)^{\rm c}}={\mathbf 0} \}$ and $\tilde \boldsymbol{\lambda}_{\boldsymbol{\theta}}=\arg \max _{\boldsymbol{\lambda} \in \Lambda_{\boldsymbol{\theta}}} f_n(\boldsymbol{\lambda}; \boldsymbol{\theta})$. Similar to (ref) and the arguments below (ref), if $\log r =o(n^{1/3})$, $\ell_n \nu^2=o(1)$ and $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$, we have $|\tilde \boldsymbol{\lambda}_{\boldsymbol{\theta},\mathcal{M}_{\boldsymbol{\theta}}(c)}|_2 \leq 4K_3^{-1} |\bar{{\mathbf g}}_{\mathcal{M}_{\boldsymbol{\theta}}(c)}(\boldsymbol{\theta})-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\boldsymbol{\theta}}(c)}(\boldsymbol{\theta})\}|_2$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1, where $K_3$ is specified in Condition (ref)(b). Notice that
By the Taylor expansion and Condition (ref)(c), we have $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1} |\bar{{\mathbf g}}(\boldsymbol{\theta}) - \bar{{\mathbf g}}(\hat{\boldsymbol{\theta}}_n)|_\infty = O_{\rm p}(\alpha_n)$, which implies $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})}(\boldsymbol{\theta}) - \bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})}(\hat{\boldsymbol{\theta}}_n)|_2 = O_{\rm p}(\ell_n^{1/2} \alpha_n)$. Due to $\alpha_n=o(\nu)$ and $|\bar{g}_j(\hat{\boldsymbol{\theta}}_n)| \geq \tilde{c} \nu \rho'(0^+)$ for any $j \in \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})$, we then have $ {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})}(\boldsymbol{\theta})\}={\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})}(\hat{\boldsymbol{\theta}}_n)\} $ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. By the triangle inequality and (ref), we have w.p.a.1 that
For any $j \in \mathcal{M}_{\boldsymbol{\theta}}(c) \bigcap \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})^{\rm c}$, we have $|\bar{g}_j(\boldsymbol{\theta})| \geq c \nu \rho'(0^+)$ and $|\bar{g}_j(\hat{\boldsymbol{\theta}}_n)| < \tilde{c} \nu \rho'(0^+)$. Due to $c \in (\tilde{c},1)$ and $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\bar{{\mathbf g}}(\boldsymbol{\theta}) - \bar{{\mathbf g}}(\hat{\boldsymbol{\theta}}_n)|_\infty = o_{\rm p}(\nu)$, then $ \mathcal{M}_{\boldsymbol{\theta}}(c) \bigcap \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})^{\rm c} = \emptyset$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1, which implies $T_{2,\boldsymbol{\theta}}=0$ for any $\boldsymbol{\theta}\in\mathcal{C}_1$ w.p.a.1. Hence,
Then $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\tilde \boldsymbol{\lambda}_{\boldsymbol{\theta}}|_2 = \sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\tilde \boldsymbol{\lambda}_{\boldsymbol{\theta},\mathcal{M}_{\boldsymbol{\theta}}(c)}|_2 = O_{\rm p}(\ell_n^{1/2} \alpha_n) =o_{\rm p}(\delta_n)$. Write $\tilde\boldsymbol{\lambda}_{\boldsymbol{\theta}}= (\tilde\lambda_{\boldsymbol{\theta},1},\ldots,\tilde\lambda_{\boldsymbol{\theta},r})^{\mathrm{\scriptscriptstyle \top} }$. Our next step is to show ${\mbox{\rm sgn}} (\tilde{\lambda}_{\boldsymbol{\theta},j})={\mbox{\rm sgn}}\{\bar{g}_j(\boldsymbol{\theta})\}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ and $j \in \mathcal{M}_{\boldsymbol{\theta}}(c)$ with $\tilde{\lambda}_{\boldsymbol{\theta},j} \neq 0$ w.p.a.1. Its proof is almost identical to that in Section (ref) for proving ${\mbox{\rm sgn}} (\tilde{\lambda}_{n,j})={\mbox{\rm sgn}}\{\bar{g}_j(\hat{\boldsymbol{\theta}}_n)\}$ for any $j \in \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})$ with $\tilde{\lambda}_{n,j} \neq 0$ w.p.a.1. We only need to replace $\{\tilde \boldsymbol{\lambda}_n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})\}$ there by $\{\tilde \boldsymbol{\lambda}_{\boldsymbol{\theta}},\mathcal{M}_{\boldsymbol{\theta}}(c)\}$ and all the arguments still hold uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$. Using the same arguments stated in the proof of Lemma (ref) for showing $\tilde\boldsymbol{\lambda}_0$ is a local maximizer for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1, we can also prove $\tilde\boldsymbol{\lambda}_{\boldsymbol{\theta}}$ is a local maximizer of $f_n(\boldsymbol{\lambda}; \boldsymbol{\theta})$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1, which implies $\hat \boldsymbol{\lambda}(\boldsymbol{\theta}) = \tilde{\boldsymbol{\lambda}}_{\boldsymbol{\theta}}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. We then have Lemma (ref). $\Box$
Recall $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})=\arg \max _{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})} f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ and $\mathcal{C}_1=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2\leq \alpha_n\}$. Then $\boldsymbol{\theta}$ and $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})=\{\hat{\lambda}_{1}(\boldsymbol{\theta}),\ldots,\hat{\lambda}_{r}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$ satisfy
where $\hat{\boldsymbol{\eta}}(\boldsymbol{\theta})=\{\hat{\eta}_{1}(\boldsymbol{\theta}),\ldots, \hat{\eta}_{r}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$ with $\hat{\eta}_j(\boldsymbol{\theta})=\nu \rho' \{|\hat{\lambda}_j(\boldsymbol{\theta})|;\nu\} {\mbox{\rm sgn}}\{\hat{\lambda}_j(\boldsymbol{\theta})\}$ for $\hat{\lambda}_j(\boldsymbol{\theta}) \neq 0$ and $\hat{\eta}_j(\boldsymbol{\theta}) \in [-\nu\rho'(0^{+}),\nu \rho'(0^{+})]$ for $\hat{\lambda}_j(\boldsymbol{\theta})=0$. Recall $\mathcal{R}(\boldsymbol{\theta})={\rm supp}\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})\}$. For any $\boldsymbol{\theta} \in \mathcal{C}_1$, restricted on $\mathcal{R}(\boldsymbol{\theta})$, define $$ {\mathbf m}_{\boldsymbol{\theta}}(\boldsymbol{\zeta},\boldsymbol{\vartheta})=\frac{1}{n}\sum _{i=1}^{n}\frac{{\mathbf g}_{i,{\mathcal{R}(\boldsymbol{\theta})}}(\boldsymbol{\vartheta})}{1+\boldsymbol{\zeta}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,{\mathcal{R}(\boldsymbol{\theta})}}(\boldsymbol{\vartheta})}-{\mathbf w} $$ for any $\boldsymbol{\vartheta} \in \boldsymbol{\Theta}$ and $\boldsymbol{\zeta}=\{\zeta_{1},\ldots,\zeta_{|{\mathcal{R}(\boldsymbol{\theta})}|}\}^{\mathrm{\scriptscriptstyle \top} } \in {\mathbb{R}}^{|{\mathcal{R}(\boldsymbol{\theta})}|}$ with each $\zeta_j \neq 0$, where ${\mathbf w}=\{w_{1},\ldots,w_{|{\mathcal{R}(\boldsymbol{\theta})}|}\}^{\mathrm{\scriptscriptstyle \top} }$ with $w_j=\nu \rho' (|\zeta_j|;\nu) {\mbox{\rm sgn}}(\zeta_j)$. From (ref), we know $\hat{\boldsymbol{\lambda}}_{\mathcal{R}(\boldsymbol{\theta})} (\boldsymbol{\theta}) $ and $\boldsymbol{\theta}$ satisfy ${\mathbf m}_{\boldsymbol{\theta}}\{\hat{\boldsymbol{\lambda}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta}), \boldsymbol{\theta}\}= {\mathbf 0}$. By the implicit function theorem [Theorem 9.28 of Rudin1976], for all $\boldsymbol{\vartheta}$ in a small neighborhood of $\boldsymbol{\theta}$, denoted by $\mathcal{U}(\boldsymbol{\theta})$, there exists a $\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})$ such that ${\mathbf m}_{\boldsymbol{\theta}}\{\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta}),\boldsymbol{\vartheta}\}={\mathbf 0}$, $\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\theta})=\hat{\boldsymbol{\lambda}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})$ and $\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})$ is continuously differentiable in $\boldsymbol{\vartheta} \in \mathcal{U}(\boldsymbol{\theta})$. By Condition (ref)(a), the event ${\mathcal{E}} = \bigcap_{\boldsymbol{\theta} \in \mathcal{C}_1} \{\max_{j \in {\mathcal{R}(\boldsymbol{\theta})}^{\rm c}} |\hat{\eta}_j(\boldsymbol{\theta})| < \nu\rho'(0^+)\}$ holds w.p.a.1. Restricted on ${\mathcal{E}}$, let $\varsigma_n=\nu\rho'(0^+) - \sup_{\boldsymbol{\theta} \in \mathcal{C}_1}\max_{j \in {\mathcal{R}(\boldsymbol{\theta})}^{\rm c}} |\hat{\eta}_j(\boldsymbol{\theta})|$ and define $\boldsymbol{\Theta}_*(\boldsymbol{\theta})=\{\boldsymbol{\vartheta} \in \mathcal{U}(\boldsymbol{\theta}): |\boldsymbol{\vartheta}-\boldsymbol{\theta}|_1 \leq o[\min\{\varsigma_n, \chi_n(\boldsymbol{\theta})\}], |\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})-\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\theta})|_1 \leq o [\min\{\varsigma_n, \ell_n^{1/2} \alpha_n\}] \}$ for some $\chi_n(\boldsymbol{\theta})>0$. Since all the components of $\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\theta})$ are nonzero and $\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})$ is continuously differentiable in $\boldsymbol{\vartheta}$, we can select sufficiently small $\chi_n(\boldsymbol{\theta})$ such that all the components of $\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})$ are nonzero for any $\boldsymbol{\vartheta} \in \boldsymbol{\Theta}_*(\boldsymbol{\theta})$. For any $\boldsymbol{\vartheta} \in \boldsymbol{\Theta}_*(\boldsymbol{\theta})$, let $\tilde\boldsymbol{\lambda}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta}) \in {\mathbb{R}}^r$ satisfy $\tilde\boldsymbol{\lambda}_{\boldsymbol{\theta},{\mathcal{R}(\boldsymbol{\theta})}}(\boldsymbol{\vartheta})=\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})$ and $\tilde\boldsymbol{\lambda}_{\boldsymbol{\theta},{\mathcal{R}(\boldsymbol{\theta})}^{\rm c}}(\boldsymbol{\vartheta})={\mathbf 0}$. By Lemma (ref), $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})|_2 =O_{\rm p}(\ell_n^{1/2} \alpha_n)$ and $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\mathcal{R}(\boldsymbol{\theta})| \leq \ell_n$ w.p.a.1, which imply
Using the same arguments in the proof of Lemma (ref) for proving that $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})$ and $\boldsymbol{\theta}$ satisfy the score equation $\nabla_{\boldsymbol{\lambda}} f_n\{\tilde\boldsymbol{\lambda}(\boldsymbol{\theta}); \boldsymbol{\theta}\}={\mathbf 0}$ w.p.a.1 there, we can prove $\nabla_{\boldsymbol{\lambda}} f_n\{\tilde\boldsymbol{\lambda}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta}); \boldsymbol{\vartheta}\}={\mathbf 0}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ and $\boldsymbol{\vartheta} \in \boldsymbol{\Theta}_*(\boldsymbol{\theta})$ w.p.a.1. By the concavity of $f_n(\boldsymbol{\lambda};\boldsymbol{\vartheta})$ w.r.t $\boldsymbol{\lambda}$, we have $\tilde\boldsymbol{\lambda}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})=\hat{\boldsymbol{\lambda}}(\boldsymbol{\vartheta})=\arg\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\vartheta})} f_n(\boldsymbol{\lambda};\boldsymbol{\vartheta})$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ and $\boldsymbol{\vartheta} \in \boldsymbol{\Theta}_*(\boldsymbol{\theta})$ w.p.a.1. Recall $\tilde\boldsymbol{\lambda}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})$ is continuously differentiable in $\boldsymbol{\vartheta}\in\mathcal{U}(\boldsymbol{\theta})$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$. Hence, $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})$ is continuously differentiable at $\boldsymbol{\theta}$ and $[\nabla_{\boldsymbol{\theta}} \hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})]_{{\mathcal{R}(\boldsymbol{\theta})}^{\rm c},[p]}={\mathbf 0}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. Write $\hat{\boldsymbol{\lambda}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta}) = \{\tilde{\lambda}_1(\boldsymbol{\theta}),\ldots,\tilde{\lambda}_{|{\mathcal{R}(\boldsymbol{\theta})}|}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$. Since $\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\theta})=\hat{\boldsymbol{\lambda}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})$, it holds that
We complete the proof of Lemma (ref). $\Box$
Recall $\hat\boldsymbol{\lambda}(\boldsymbol{\theta})=\arg\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})}f_n(\boldsymbol{\lambda};\boldsymbol{\theta} )$, $\mathcal{R}_n={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$ and $\mathcal{C}_1=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2\leq \alpha_n\}$. By Lemma (ref), $|\mathcal{R}_n|\leq \ell_n$ w.p.a.1. Select $\delta_n$ satisfying $\delta_n=o(\ell_n^{-1/2} n^{-1/\gamma})$ and $\ell_n^{1/2} \alpha_n=o(\delta_n) $, which can be guaranteed by $\ell_n\alpha_n=o(n^{-1/\gamma})$. For any $\boldsymbol{\theta} \in \mathcal{C}_1$, let $\tilde \boldsymbol{\lambda}(\boldsymbol{\theta})=\arg \max _{\boldsymbol{\lambda} \in \check\Lambda_n} f_n(\boldsymbol{\lambda}; \boldsymbol{\theta})$, where $\check\Lambda_n=\{\boldsymbol{\lambda}\in {\mathbb{R}}^{r}:|\boldsymbol{\lambda}_{ \mathcal{R}_n}|_2\leq \delta_n, \boldsymbol{\lambda}_{\mathcal{R}_n^{\rm c}}={\mathbf 0} \}$. Similar to (ref) and the arguments below (ref), if $\log r =o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n \nu^2=o(1)$, we have $|\tilde \boldsymbol{\lambda}_{\mathcal{R}_n}(\boldsymbol{\theta})|_2 \leq 4K_3^{-1} |\bar{{\mathbf g}}_{\mathcal{R}_n}(\boldsymbol{\theta})-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{R}_n}(\boldsymbol{\theta})\}|_2$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1, where $K_3$ is specified in Condition (ref)(b). By Lemma (ref), we have $\mathcal{R}_n \subset \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})$ w.p.a.1, where $\tilde{c}$ is specified in Condition {\rm (ref)(a)}. Using the arguments for deriving (ref), we have $$ \sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\bar{{\mathbf g}}_{ \mathcal{R}_n}(\boldsymbol{\theta})-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{ \mathcal{R}_n}(\boldsymbol{\theta})\}|_2 = O_{\rm p}(\ell_n^{1/2} \alpha_n)\,,$$ which implies $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\tilde \boldsymbol{\lambda}_{ \mathcal{R}_n}(\boldsymbol{\theta})|_2 = O_{\rm p}(\ell_n^{1/2} \alpha_n)=o_{\rm p}(\delta_n)$. Write $\tilde\boldsymbol{\lambda}_{\mathcal{R}_n}(\boldsymbol{\theta})=\{\dot\lambda_1(\boldsymbol{\theta}), \ldots, \dot\lambda_{|\mathcal{R}_n|}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$. By the first-order condition, for any $\boldsymbol{\theta} \in \mathcal{C}_1$, we have $$ {\mathbf 0} = \frac{1}{n}\sum_{i=1}^{n} \frac{{\mathbf g}_{i,\mathcal{R}_n}(\boldsymbol{\theta})}{1+\tilde \boldsymbol{\lambda}_{\mathcal{R}_n}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{R}_n}(\boldsymbol{\theta})} - \tilde\boldsymbol{\eta}(\boldsymbol{\theta})\,,$$ where $\tilde\boldsymbol{\eta}(\boldsymbol{\theta})=\{\tilde\eta_1(\boldsymbol{\theta}), \ldots,\tilde\eta_{|\mathcal{R}_n|}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$ with $\tilde\eta_j(\boldsymbol{\theta})=\nu\rho'\{|\dot\lambda_j(\boldsymbol{\theta})|;\nu\}{\mbox{\rm sgn}}\{\dot\lambda_j(\boldsymbol{\theta})\}$ for $\dot{\lambda}_j(\boldsymbol{\theta})\neq 0$ and $\tilde\eta_j(\boldsymbol{\theta}) \in [-\nu\rho'(0^+), \nu\rho'(0^+)]$ for $\dot{\lambda}_j(\boldsymbol{\theta})= 0$. Using the same arguments for addressing the remainder terms in (ref), for any $\boldsymbol{\theta} \in \mathcal{C}_1$ and $j\in\mathcal{R}_n^{\rm c}$, it holds that $$ \frac{1}{n}\sum _{i=1}^{n} \frac{g_{i,j}(\boldsymbol{\theta})}{1+\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i}(\boldsymbol{\theta})} =\frac{1}{n}\sum _{i=1}^{n} \frac{g_{i,j}(\hat{\boldsymbol{\theta}}_n)}{1+\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i}(\hat{\boldsymbol{\theta}}_n)} + o_{\rm p}(\nu)\,,$$ where the term $o_{\rm p}(\nu)$ holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$ and $j \in \mathcal{R}_n^{\rm c}$. Together with Condition (ref)(a), we have that $$ \sup_{\boldsymbol{\theta} \in \mathcal{C}_1} \max_{j \in \mathcal{R}_n^{\rm c}}\bigg|\frac{1}{n}\sum _{i=1}^{n} \frac{g_{i,j}(\boldsymbol{\theta})}{1+\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i}(\boldsymbol{\theta})}\bigg| \leq \nu\rho'(0^+)$$ w.p.a.1. Therefore, $\tilde \boldsymbol{\lambda}(\boldsymbol{\theta})$ and $\boldsymbol{\theta}$ satisfy the score equation $\nabla_{\boldsymbol{\lambda}} f_n\{\tilde\boldsymbol{\lambda}(\boldsymbol{\theta}); \boldsymbol{\theta}\}={\mathbf 0}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. By the concavity of $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ w.r.t $\boldsymbol{\lambda}$, it holds that $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})=\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})=\arg\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})} f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1, which implies ${\rm supp}\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})\} \subset {\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. Select $\boldsymbol{\theta}^* \in \mathcal{C}_1$ such that $|{\rm supp}\{\hat \boldsymbol{\lambda}(\boldsymbol{\theta}^*)\}| \leq |{\rm supp}\{\hat \boldsymbol{\lambda}(\boldsymbol{\theta})\}|$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$, and define $\mathcal{B}_2(\boldsymbol{\theta}^*,2\alpha_n) = \{\boldsymbol{\theta} \in\boldsymbol{\Theta}:\,|\boldsymbol{\theta}-\boldsymbol{\theta}^*|_2 \leq 2\alpha_n \}$. Using the same arguments above for proving ${\rm supp}\{\hat\boldsymbol{\lambda}(\boldsymbol{\theta})\} \subset {\rm supp}\{\hat \boldsymbol{\lambda}(\hat\boldsymbol{\theta}_n)\}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1, we have ${\rm supp} \{\hat\boldsymbol{\lambda}(\boldsymbol{\theta})\} \subset {\rm supp} \{\hat\boldsymbol{\lambda}(\boldsymbol{\theta}^*)\}$ for any $\boldsymbol{\theta} \in \mathcal{B}_2(\boldsymbol{\theta}^*,2\alpha_n)$ w.p.a.1. Since $\mathcal{C}_1 \subset \mathcal{B}_2(\boldsymbol{\theta}^*,2\alpha_n)$, then ${\rm supp} \{\hat\boldsymbol{\lambda}(\boldsymbol{\theta})\} \subset {\rm supp}\{\hat \boldsymbol{\lambda}(\boldsymbol{\theta}^*)\}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. Due to $|{\rm supp}\{\hat \boldsymbol{\lambda}(\boldsymbol{\theta}^*)\}| \leq |{\rm supp}\{\hat \boldsymbol{\lambda}(\boldsymbol{\theta})\}|$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$, we have ${\rm supp}\{\hat \boldsymbol{\lambda}(\boldsymbol{\theta})\}={\rm supp}\{\hat \boldsymbol{\lambda}(\boldsymbol{\theta}^*)\}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. We complete the proof of Lemma (ref). $\Box$
By Lemma (ref), we have $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\mathcal{R}(\boldsymbol{\theta})|\leq \ell_n$ w.p.a.1 and $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n)$. Under Condition (ref)(a), if $\ell_n\alpha_n=o(n^{-1/\gamma})$, then $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}\max _{ i \in [n]}|\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} }{\mathbf g}_{i}(\boldsymbol{\theta})|=o_{\rm p}(1)$. Write $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})=\{\hat\lambda_1(\boldsymbol{\theta}),\ldots,\hat\lambda_r(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$ and ${\mathbf t}=(t_1,\ldots,t_p)^{\mathrm{\scriptscriptstyle \top} }$. For $T_{\boldsymbol{\theta},1}$, by the Cauchy-Schwarz inequality and Condition (ref)(c), we have
holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$ and ${\mathbf t} \in \mathbb{R}^p$. For $T_{\boldsymbol{\theta},3}$, by Condition (ref)(c), we have
holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$ and ${\mathbf t} \in \mathbb{R}^p$. Let $\hat\boldsymbol{\lambda}_{{\mathcal{R}(\boldsymbol{\theta})}}(\boldsymbol{\theta})=\{\tilde{\lambda}_1(\boldsymbol{\theta}),\ldots,\tilde{\lambda}_{|{\mathcal{R}(\boldsymbol{\theta})}|}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$. By the Cauchy-Schwarz inequality, Conditions (ref)(a) and (ref)(c), if $\log r=o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n\nu^2=o(1)$, then
holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$ and ${\mathbf t} \in \mathbb{R}^p$. Recall $\alpha_n=o(\nu)$. By Proposition (ref), Lemma (ref) and Condition (ref)(b), if $\log r =o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n \nu^2=o(1)$, we know that $\inf_{\boldsymbol{\theta} \in \mathcal{C}_1}\lambda_{\min}\{\widehat{{\mathbf V}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})\}$ and $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}\lambda_{\max}\{\widehat{{\mathbf V}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})\}$ are uniformly bounded away from zero and infinity w.p.a.1. Using the same arguments in the proof of Lemma (ref), if $\log r=o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n\nu^2=o(1)$, it holds that
By Condition (ref) and the same arguments in the proof of Lemma (ref), if $\log r=o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n\nu^2=o(1)$, we have $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}\|\widehat{\boldsymbol{\Gamma}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})\|_2=O_{\rm p}(1)$. Notice that $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}\|\widehat{{\mathbf V}}^{-1}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})\|_2=O_{\rm p}(1)$ and $ \sup_{\boldsymbol{\theta} \in \mathcal{C}_1}\|\nu {\rm diag}[\rho''\{|\tilde\lambda_{1}(\boldsymbol{\theta})|;\nu \},\ldots,\rho''\{|\tilde\lambda_{|{\mathcal{R}(\boldsymbol{\theta})}|}(\boldsymbol{\theta})|;\nu \}]\|_2=O_{\rm p}(\nu) $. Thus,
Combining (ref), (ref) and (ref), by Lemma (ref), we know $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|[\nabla_{\boldsymbol{\theta}} \hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})]_{{\mathcal{R}(\boldsymbol{\theta})},[p]} {\mathbf t}|_2 = |{\mathbf t}|_2 \cdot O_{\rm p}(1)$ holds uniformly over ${\mathbf t} \in \mathbb{R}^p$, which implies
holds uniformly over ${\mathbf t} \in \mathbb{R}^p$. For $T_{\boldsymbol{\theta},4}$, by (ref), (ref) and (ref), Lemma (ref) implies
holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$ and ${\mathbf t} \in \mathbb{R}^p$. We then obtain the result by (ref). $\hfill\Box$
Denote by ${\mathbb{P}}_{\mathcal{X}_n}(\cdot)$ and $\mathbb{E}_{\mathcal{X}_n}(\cdot)$, respectively, the conditional probability and conditional expectation given $\mathcal{X}_n$. For any integer $k \geq 1$, recall $\hat{\boldsymbol{\zeta}}_{k+1} = N_{k}^{-1} \sum_{i=1}^{N_{k}} \omega^{k}_i {\mathbf h}(\boldsymbol{\theta}^{k}_i)$ only depends on the $N_k$ samples $\{\boldsymbol{\theta}^{k}_1,\ldots,\boldsymbol{\theta}^{k}_{N_k}\}$ generated from the proposal distribution with density $\varphi(\boldsymbol{\theta}\,;\hat{\boldsymbol{\zeta}}_{k})$, where $\omega^{k}_i = \pi^\dag(\boldsymbol{\theta}^{k}_i\,|\,\mathcal{X}_n) / \varphi(\boldsymbol{\theta}^{k}_i\,;\hat{\boldsymbol{\zeta}}_{k})$. Thus, the random sequence $\{\hat{\boldsymbol{\zeta}}_{k}\}_{k\geq1}$ forms a Markov chain. Recall that $\boldsymbol{\Theta} \subset {\mathbb{R}}^p$ is a compact set with fixed $p$. Since $\sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} |{\mathbf h}(\boldsymbol{\theta})|_\infty \leq K_9$ for some universal constant $K_9>0$, and $\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})$ is positive and continuous on $(\boldsymbol{\theta}, \boldsymbol{\zeta}) \in \boldsymbol{\Theta} \times {\mathbb{R}}^s$, there exists a positive and continuous function $\varrho(\cdot)$ such that $$ \sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} \frac{\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)|{\mathbf h}(\boldsymbol{\theta})|_\infty}{\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})} \leq \varrho(\boldsymbol{\zeta}) $$ for any $\boldsymbol{\zeta} \in {\mathbb{R}}^s$. Since $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) = 0$ for any $\boldsymbol{\theta} \notin \boldsymbol{\Theta}$, then $$ \sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}^{\rm c}} \frac{\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)|{\mathbf h}(\boldsymbol{\theta})|_\infty}{\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})} = 0 $$ for any $\boldsymbol{\zeta} \in {\mathbb{R}}^s$. Notice that ${\mathbb{E}}(\hat{\boldsymbol{\zeta}}_{k+1} \,|\, \hat{\boldsymbol{\zeta}}_{k}) = \boldsymbol{\zeta}^* = {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}\{{\mathbf h}(\boldsymbol{\theta})\}$. Write $\hat{\boldsymbol{\zeta}}_{k+1} = (\hat{\zeta}_{k+1, 1}, \ldots, \hat{\zeta}_{k+1, s})^{\mathrm{\scriptscriptstyle \top} }$ and $\boldsymbol{\zeta}^* = (\zeta^*_1,\ldots, \zeta^*_s)^{\mathrm{\scriptscriptstyle \top} } $. For any $\varepsilon>0$, by the Hoeffding's inequality, we have
Let $C_\varepsilon = \sup_{\boldsymbol{\zeta}\in{\mathbb{R}}^s:\, |\boldsymbol{\zeta} - \boldsymbol{\zeta}^*|_\infty \leq \varepsilon} \varrho^2(\boldsymbol{\zeta})$. By (ref), we have
which implies
For any integer $k'\geq k$, by the Markov property of $\{\hat{\boldsymbol{\zeta}}_{k}\}_{k\geq1}$, it then holds that
Letting $k' \rightarrow \infty$, then $$ {\mathbb{P}}_{\mathcal{X}_n}\bigg(\bigcap^\infty_{t= k} \{|\hat{\boldsymbol{\zeta}}_{t+1} - \boldsymbol{\zeta}^*|_\infty \leq \varepsilon\} \bigg) \geq {\mathbb{P}}_{\mathcal{X}_n}\big(|\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty \leq \varepsilon\big) \prod^\infty_{t= k+1} \bigg\{1 - 2s\exp\bigg(-\frac{N_{t}\varepsilon^2}{2 C_\varepsilon}\bigg) \bigg\} \,. $$ Since $s$ is fixed and $\sum_{k=1}^{\infty}\exp(-CN_k) < \infty$ for any $C>0$, we have
For any $z>0$, let $\bar{C}_z = \sup_{\boldsymbol{\zeta}\in {\mathbb{R}}^s:\,|\boldsymbol{\zeta}|_\infty \leq z} \varrho^2(\boldsymbol{\zeta})$. Using the same arguments for (ref), it holds that $$ {\mathbb{P}}_{\mathcal{X}_n}\big(|\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty > \varepsilon, |\hat{\boldsymbol{\zeta}}_{k}|_\infty \leq z\big) \leq 2s\exp\bigg(-\frac{N_{k}\varepsilon^2}{2C_z}\bigg) {\mathbb{P}}_{\mathcal{X}_n}\big(|\hat{\boldsymbol{\zeta}}_{k}|_\infty \leq z\big) \leq 2s\exp\bigg(-\frac{N_{k}\varepsilon^2}{2C_z}\bigg) \,. $$ By the Markov's inequality and triangle inequality,
which implies ${\mathbb{P}}_{\mathcal{X}_n}(|\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty > \varepsilon) \leq 2s\exp\{-(2C_z)^{-1}N_{k}\varepsilon^2\} + z^{-1}{\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}\{|{\mathbf h}(\boldsymbol{\theta})|_\infty\}$. Due to $N_k \rightarrow \infty$ as $k \rightarrow \infty$, we know $\limsup_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}(|\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty > \varepsilon) \leq z^{-1}{\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}\{|{\mathbf h}(\boldsymbol{\theta})|_\infty\}$. Notice that $\sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} |{\mathbf h}(\boldsymbol{\theta})|_\infty \leq K_9$ for some universal constant $K_9>0$. Letting $z \rightarrow \infty$, it holds that $\limsup_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}(|\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty > \varepsilon) = 0$, which implies $\liminf_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}(|\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty \leq \varepsilon) = 1$. Together with (ref), we have $$ \liminf_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}\bigg(\bigcap^\infty_{t = k} \{|\hat{\boldsymbol{\zeta}}_{t+1} - \boldsymbol{\zeta}^*|_\infty \leq \varepsilon\}\bigg) = 1 \,. $$ Since $\varepsilon>0$ is arbitrary, we then obtain that, conditional on $\mathcal{X}_n$, $|\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$. We complete the proof of Lemma (ref). $\Box$
Denote by ${\mathbb{P}}_{\mathcal{X}_n}(\cdot)$ and $\mathbb{E}_{\mathcal{X}_n}(\cdot)$, respectively, the conditional probability and conditional expectation given $\mathcal{X}_n$. Recall $\boldsymbol{\zeta}^* = {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}\{{\mathbf h}(\boldsymbol{\theta})\}$. Let $ \widehat{{\mathbf Z}}_{k+1}= N_k^{-1} \sum_{i=1}^{N_k} \boldsymbol{\theta}^k_i \pi^\dag(\boldsymbol{\theta}^k_i\,|\,\mathcal{X}_n) / \varphi(\boldsymbol{\theta}^k_i\,;\boldsymbol{\zeta}^*)$ for any integer $k \geq 2$ and $$ {\mathbf Z}(\boldsymbol{\zeta})={\mathbb{E}}_{\boldsymbol{\theta} \sim\varphi(\cdot\,;\,\boldsymbol{\zeta})} \bigg\{\frac{\boldsymbol{\theta} \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)}{\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta}^*)} \bigg\} = \int_{{\mathbb{R}}^p} \frac{\boldsymbol{\theta} \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)}{\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta}^*)} \varphi(\boldsymbol{\theta}\,;\boldsymbol{\zeta}) \,{\rm d} \boldsymbol{\theta} $$ for any $\boldsymbol{\zeta} \in {\mathbb{R}}^s$, where $\{\boldsymbol{\theta}^1_1, \ldots, \boldsymbol{\theta}^1_{N_1}, \ldots,\boldsymbol{\theta}^k_1,\ldots, \boldsymbol{\theta}^k_{N_k}\}$ are generated via Algorithm {\rm (ref)}.
Our first step is to show that conditional on $\mathcal{X}_n$, we have $|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k})|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$. Notice that $\boldsymbol{\Theta} \subset {\mathbb{R}}^p$ is a compact set with fixed $p$ and $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) =0$ for any $\boldsymbol{\theta} \notin \boldsymbol{\Theta}$. Since $\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})$ is positive and continuous on $(\boldsymbol{\theta}, \boldsymbol{\zeta}) \in \boldsymbol{\Theta} \times {\mathbb{R}}^s$, we know
for some universal constant $\tilde{C}>0$. Recall that $\widehat{{\mathbf Z}}_{k+1}$ depends on the $N_k$ samples $\{\boldsymbol{\theta}^k_1,\ldots,\boldsymbol{\theta}^k_{N_k}\}$ generated from the proposal distribution with density $\varphi(\boldsymbol{\theta}\,;\hat{\boldsymbol{\zeta}}_{k})$. Then ${\mathbb{E}}(\widehat{{\mathbf Z}}_{k+1} \,|\, \hat{\boldsymbol{\zeta}}_{k}) = {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k})$. For any $\varepsilon>0$, using the same arguments for (ref), we have
Define the event $A_{k+1} = \{|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k})|_\infty \leq \varepsilon, |\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty \leq \varepsilon\}$. Note that $A_{k} \in \sigma(\boldsymbol{\theta}_1^{k-1},\ldots, \boldsymbol{\theta}_{N_{k-1}}^{k-1}, \hat{\boldsymbol{\zeta}}_{k-1})$ and the conditional joint distribution of $(\widehat{{\mathbf Z}}_{k+1},\hat{\boldsymbol{\zeta}}_{k+1})$ given $\mathcal{X}_n$ is fully determined by $\hat{\boldsymbol{\zeta}}_{k}$. By (ref) and (ref), it holds that
where $C_\varepsilon = \sup_{\boldsymbol{\zeta}\in{\mathbb{R}}^s:\, |\boldsymbol{\zeta} - \boldsymbol{\zeta}^*|_\infty \leq \varepsilon} \varrho^2(\boldsymbol{\zeta})$ with the function $\varrho(\cdot)$ specified in the proof of Lemma (ref). Then $$ {\mathbb{P}}_{\mathcal{X}_n}\big(A^{\rm c}_{k+1} \,\big|\, A_{k} \big) \leq 2p\exp\bigg(-\frac{N_{k}\varepsilon^2}{2\tilde{C}^2} \bigg) + 2s\exp\bigg(-\frac{N_{k}\varepsilon^2}{2C_\varepsilon} \bigg) \leq 2(p+s)\exp\bigg(-\frac{N_{k}\varepsilon^2}{\check{C}_\varepsilon} \bigg) $$ for some $\check{C}_\varepsilon >0$ depending on $\varepsilon$. For any integer $k'\geq k$, by the Markov property of $\{(\widehat{{\mathbf Z}}_{k}, \hat{\boldsymbol{\zeta}}_{k})\}_{k\geq 2}$, it then holds that
Letting $k' \rightarrow \infty$, then $$ {\mathbb{P}}_{\mathcal{X}_n}\bigg(\bigcap^\infty_{t=k} A_{t+1} \bigg) \geq {\mathbb{P}}_{\mathcal{X}_n}(A_{k+1}) \prod^\infty_{t = k+1} \bigg\{1 - 2(p+s)\exp \bigg(-\frac{N_{t}\varepsilon^2}{\check{C}_\varepsilon}\bigg) \bigg\} \,. $$ Since $p$ and $s$ are fixed constants and $\sum_{k=1}^{\infty}\exp(-CN_k) < \infty$ for any $C>0$, we have
where the last step is due to the fact $\limsup_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}(|\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty > \varepsilon) = 0$ as shown in the proof of Lemma (ref). For any $z>0$, by (ref), it holds that
Together with (ref), we have $$ {\mathbb{P}}_{\mathcal{X}_n}\{|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k})|_\infty > \varepsilon\} \leq 2p\exp\bigg(-\frac{N_{k}\varepsilon^2}{2\tilde{C}^2} \bigg) + z^{-1}{\mathbb{E}}_{\boldsymbol{\theta} \sim\pi^\dag}\{|{\mathbf h}(\boldsymbol{\theta})|_\infty\} \,. $$ Due to $N_k \rightarrow \infty$ as $k \rightarrow \infty$, we know $\limsup_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}\{|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k})|_\infty > \varepsilon\} \leq z^{-1}{\mathbb{E}}_{\boldsymbol{\theta} \sim\pi^\dag}\{|{\mathbf h}(\boldsymbol{\theta})|_\infty\}$. Notice that $\sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} |{\mathbf h}(\boldsymbol{\theta})|_\infty \leq K_9$ for some universal constant $K_9>0$. Letting $z \rightarrow \infty$, it holds that $\limsup_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}\{|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k})|_\infty > \varepsilon\} = 0$. Together with (ref), we have $$ \liminf_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}\bigg[\bigcap^\infty_{t=k} \big\{ |\widehat{{\mathbf Z}}_{t+1} - {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{t})|_\infty \leq \varepsilon \big\}\bigg] \geq \liminf_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}\bigg(\bigcap^\infty_{t=k} A_{t+1} \bigg) =1 \,. $$ Since $\varepsilon>0$ is arbitrary, conditional on $\mathcal{X}_n$, we have $|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k})|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$.
Our second step is to show that conditional on $\mathcal{X}_n$, we have $|{\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k}) - {\mathbf Z}(\boldsymbol{\zeta}^*)|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$. By (ref), we have $|{\mathbf Z}(\boldsymbol{\zeta})|_\infty < \infty$ for any $\boldsymbol{\zeta} \in {\mathbb{R}}^s$, and $$ |{\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k}) - {\mathbf Z}(\boldsymbol{\zeta}^*)|_\infty \leq \int_{{\mathbb{R}}^p} \frac{\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)|\boldsymbol{\theta}|_\infty}{\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta}^*)} |\varphi(\boldsymbol{\theta}\,; \hat{\boldsymbol{\zeta}}_{k}) - \varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta}^*)| \,{\rm d}\boldsymbol{\theta} \leq \tilde{C}\int_{\boldsymbol{\Theta}}|\varphi(\boldsymbol{\theta}\,; \hat{\boldsymbol{\zeta}}_{k}) - \varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta}^*)| \,{\rm d}\boldsymbol{\theta} \,. $$ For some sufficiently large $M>0$, since conditional on $\mathcal{X}_n$, we have $|\hat{\boldsymbol{\zeta}}_{k} - \boldsymbol{\zeta}^*|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$, then for any $\epsilon >0$ there exists a sufficiently large integer $k_\epsilon$ such that $ {\mathbb{P}}_{\mathcal{X}_n}(\mathcal{A}) \leq \epsilon $ with $\mathcal{A} = \bigcup_{t=k_\epsilon}^\infty \{|\hat{\boldsymbol{\zeta}}_{t} - \boldsymbol{\zeta}^*|_\infty > M\}$. Define a compact set $\mathcal{B}=\{\boldsymbol{\zeta} \in {\mathbb{R}}^s: |\boldsymbol{\zeta} - \boldsymbol{\zeta}^*|_\infty \leq M \} $. Recall $\boldsymbol{\Theta} \subset {\mathbb{R}}^p$ is a compact set with fixed $p$. Due to the continuity of $\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})$, we know $\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})$ is uniformly continuous on $(\boldsymbol{\theta}\,; \boldsymbol{\zeta}) \in \boldsymbol{\Theta} \times \mathcal{B}$. For any $\varepsilon>0$, there exists $\delta(\varepsilon)>0$ such that $|\varphi(\boldsymbol{\theta}_1\,; \boldsymbol{\zeta}_1) - \varphi(\boldsymbol{\theta}_2\,; \boldsymbol{\zeta}_2)| < \tilde{C}^{-1} \varepsilon / \mathbb{L}(\boldsymbol{\Theta})$ for any $(\boldsymbol{\theta}_1, \boldsymbol{\zeta}_1), (\boldsymbol{\theta}_2, \boldsymbol{\zeta}_2) \in \boldsymbol{\Theta} \times \mathcal{B}$ satisfying $|\boldsymbol{\theta}_1 - \boldsymbol{\theta}_2|_\infty \leq \delta(\varepsilon)$ and $|\boldsymbol{\zeta}_1 - \boldsymbol{\zeta}_2|_\infty \leq \delta(\varepsilon)$, where $\mathbb{L}(\cdot)$ is the Lebesgue measure on ${\mathbb{R}}^p$. Since
we then have
where the second step is due to the fact that conditional on $\mathcal{X}_n$ we have $|\hat{\boldsymbol{\zeta}}_{k} - \boldsymbol{\zeta}^*|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$. Letting $\epsilon \to 0$, we know that conditional on $\mathcal{X}_n$, we have $|{\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k}) - {\mathbf Z}(\boldsymbol{\zeta}^*)|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$.
Our third step is to show that conditional on $\mathcal{X}_n$, we have $|\widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(\boldsymbol{\theta}) - {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta})|_\infty \rightarrow 0$ almost surely as $K \rightarrow \infty$. By the triangle inequality, $|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\boldsymbol{\zeta}^*)|_\infty \leq |\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k})|_\infty + |{\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k}) - {\mathbf Z}(\boldsymbol{\zeta}^*)|_\infty $. Based on the results shown in Steps 1 and 2 above, it holds that conditional on $\mathcal{X}_n$, we have $|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\boldsymbol{\zeta}^*)|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$. Notice that ${\mathbf Z}(\boldsymbol{\zeta}^*)={\mathbb{E}}_{\boldsymbol{\theta} \sim\pi^\dag}(\boldsymbol{\theta})$ and $$ \widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(\boldsymbol{\theta}) = \frac{1}{S_K} \sum_{k=1}^{K} \sum_{i=1}^{N_k} \frac{\pi^\dag(\boldsymbol{\theta}^k_i\,|\,\mathcal{X}_n)}{\varphi(\boldsymbol{\theta}^k_i\,;\boldsymbol{\zeta}^*) } \boldsymbol{\theta}^k_i = \frac{1}{S_K} \sum_{k=1}^{K} N_k \widehat{{\mathbf Z}}_{k+1} $$ with $S_K=N_1+\cdots+N_K$. Notice that conditional on $\mathcal{X}_n$, $|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\boldsymbol{\zeta}^*)|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$. Given a constant $\varepsilon>0$, for any $\epsilon>0$ there exists a sufficiently large integer $\tilde{k}_\epsilon$ such that ${\mathbb{P}}_{\mathcal{X}_n}(\mathcal{C}) \leq \epsilon$ with $\mathcal{C} = \bigcup_{k=\tilde{k}_\epsilon}^\infty\{|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\boldsymbol{\zeta}^*)|_