EconBase
← Back to paper

Bayesian penalized empirical likelihood and Markov Chain Monte Carlo sampling

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

405,033 characters · 52 sections · 83 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
This text was truncated for display. The citation measures were computed over the complete text.

Bayesian Penalized Empirical Likelihood and Markov Chain Monte Carlo Sampling

\if11 {

\affil[a]{\it Joint Laboratory of Data Science and Business Intelligence, Southwestern University of Finance and Economics, Chengdu, China} \affil[b]{\it Academy of Mathematics and Systems Science, Chinese Academy of Sciences, Beijing, China} \affil[c]{\it Department of Statistics, Operations, and Data Science, Temple University, Philadelphia, PA, USA }

} \fi

\if01 {

center[center omitted — 110 chars of source]

} \fi

abstractIn this study, we introduce a novel methodological framework called Bayesian Penalized Empirical Likelihood (BPEL), designed to address the computational challenges inherent in empirical likelihood (EL) approaches. Our approach has two primary objectives: (i) to enhance the inherent flexibility of EL in accommodating diverse model conditions, and (ii) to facilitate the use of well-established Markov Chain Monte Carlo (MCMC) sampling schemes as a convenient alternative to the complex optimization typically required for statistical inference using EL. To achieve the first objective, we propose a penalized approach that regularizes the Lagrange multipliers, significantly reducing the dimensionality of the problem while accommodating a comprehensive set of model conditions. For the second objective, our study designs and thoroughly investigates two popular sampling schemes within the BPEL context. We demonstrate that the BPEL framework is highly flexible and efficient, enhancing the adaptability and practicality of EL methods. Our study highlights the practical advantages of using sampling techniques over traditional optimization methods for EL problems, showing rapid convergence to the global optima of posterior distributions and ensuring the effective resolution of complex statistical inference challenges.

{\it Key words:} Bayesian methods, Bernstein-von Mises theorem, Estimating equations, MCMC, Penalized empirical likelihood.

\baselineskip 24pt

Introduction

EL Owen_2001 is a versatile and flexible tool for statistical inference, providing a framework that accommodates broadly defined model conditions. Unlike traditional likelihood approaches, EL does not require the explicit specification of probability distributions governing the data generation process. This inherent flexibility offers numerous practical advantages, such as the ability to incorporate a wide range of model specifications and prior knowledge, making it highly adaptable for integrating information from multiple data sources. Additionally, EL retains key benefits of its parametric likelihood counterpart, including efficiency (in the semiparametric sense) and the convenience of conducting hypothesis tests and estimating confidence sets through the Wilks-type likelihood ratio framework.

Recent developments in EL approaches have a focus on addressing the challenges posed by complex high-dimensional data. To handle the complexities arising from various model conditions, researchers have explored regularization techniques applied to the Lagrange multipliers associated with EL or the empirical versions of moment conditions, aiming to achieve enhanced model parsimony. In Shi2016, a two-step procedure is introduced. The first step involves employing a “relaxed" EL that incorporates specific inequality constraints in its formulation. The second step includes moment selection and bias correction. Chau2017 addresses a continuum of moment conditions where the numerical optimization problem becomes ill-conditioned. To resolve this, a penalty on the continuous version of the Lagrange multiplier's counterpart is proposed and investigated. Changetal_2018 proposes a method to penalize the magnitudes of both the Lagrange multiplier and the model parameters, specifically to tackle high-dimensional model parameters under complex conditions. More recently, Chang_2021 explores the projection of high-dimensional moment conditions onto lower-dimensional spaces to facilitate statistical inference for specific components of model parameters and to assess model specification validity. Besides addressing the challenge of handling many moment conditions, the development of EL approaches that incorporate penalties on model parameters to promote parsimonious structures can effectively manage high-dimensional problems, as discussed in TangLeng_2010_Bioka, LengTang_2010, ChangChenChen_2015, and Chang2023.

The synergy of Bayesian methodologies with traditional likelihoods has consistently demonstrated its effectiveness. Leveraging advances in sampling techniques, Bayesian approaches have established their significance in tackling a wide array of challenges across various domains. This is particularly valuable when dealing with intricate statistical problems where maximizing or even computing the objective function becomes infeasible. The amalgamation of Bayesian principles with EL shows great promise in practical applications. This integration enhances the adaptability and robustness of the Bayesian framework, enabling the creation of statistical models that can accommodate a wide range of scenarios. Recent developments in the realm of Bayesian EL (BEL) methods are evident in a growing body of literature; see Lazar2003, Rao2010, Chaudhuri2011, Yang2012, Mengersen2013, Chib2018, Cheng2019, Zhao2020, Tang2022, and Yu2023.

The class of EL approaches often encounters significant challenges due to substantial computational complexity, which frequently presents barriers in practice. These difficulties primarily arise from the nonconvex nature of the objective function and the potential nonconvexity of its support. As the complexity of the model increases with additional parameters and conditions, these computational obstacles become more severe. Thus, developing computationally efficient strategies is crucial to address these challenges. Indeed, as demonstrated in Chau2017 and related works, solving the associated optimization problem of penalized EL (PEL) can be a dauntingly difficult task. In our study, we demonstrate that, when combined with the Bayesian framework, sampling schemes offer promising alternatives. Once successfully drawn, samples from the posterior distribution can be used to develop the estimator.

In recent research, sampling techniques, often perceived as computationally demanding alternatives to optimization methods, demonstrate remarkable efficiency in approximating target distributions, outperforming optimization alternatives in handling nonconvex problems; see Ma2019. While sampling techniques offer a promising approach within the framework of BEL, there exist numerous challenges associated with devising these computational schemes. On one hand, EL has the potential to leverage information from various model conditions, leading to more precise estimates of unknown model parameters. However, the inclusion of a large number of these conditions introduces additional complexities, both in theory and practical implementation. Indeed, the dimensionality of the problem remains a central obstacle in EL approaches, as elaborated in Hjortetal_2008_AS. Furthermore, the incorporation of an increasing number of moment conditions can substantially amplify the nonconvex nature of the associated optimization problems, making the development of an effective sampling scheme increasingly more challenging. As underscored in Chaudhuri2017, traditional MCMC techniques encounter significant hurdles when applied to BEL due to the intricate and nonconvex characteristics of the parameter space in which new samples are generated.

Our research aims to establish an innovative methodological framework, guided by two primary objectives: (i) our approach maintains the inherent flexibility and adaptability of EL, allowing for the incorporation of broad model conditions; and (ii) our framework provides convenient access to well-established MCMC computing schemes, streamlining practical implementations. To address the first objective and mitigate challenges stemming from numerous model conditions, we propose a penalized approach. By penalizing the magnitudes of the Lagrange multipliers used in evaluating EL at specific model parameter values, we create an effective mechanism similar to moment selection. This approach reduces the problem's dimensionality while still leveraging the potential efficiency gains from a comprehensive set of model conditions. For the second objective, our approach effectively overcomes the obstacles associated with devising sampling schemes for applying Bayesian approaches, thanks to the efficient dimensionality reduction achieved through PEL. In our study, we demonstrate the practicality of our framework using two well-established sampling methods: the popular Metropolis-Hastings sampling and the influential adaptive multiple importance sampling technique for approximate Bayesian computations.

Our study makes several noteworthy contributions, in addition to the methodological advancement mentioned earlier. On a theoretical level, our analysis establishes the properties of the BPEL estimator, allowing for an exponentially increasing number of model conditions, thereby enabling unprecedented adaptability in practical applications. Furthermore, we develop theory that guarantees the convergence of the two showcased sampling schemes, thereby ensuring the validity of BPEL in statistical inference. Our study reinforces the observations made in a recent study by Ma2019 that sampling techniques offer compelling alternatives to optimization methods in addressing computationally demanding problems. Our theoretical results and numerical studies demonstrate that sampling schemes converge rapidly to stationary distributions centered around the true global optimizer. In contrast, optimization methods often require more time and can become trapped at local peaks, limiting their ability to locate the true optimum.

The rest of this article are structured as follows. Section (ref) delves into the framework of BPEL and introduces two MCMC algorithms. Numerical studies and real data analysis for an international trade dataset are presented in Sections (ref) and (ref), respectively. Section (ref) comprehensively develops the properties and theoretical guarantees of the proposed methods. Some discussions are provided in Section (ref), while all technical proofs are available in the supplementary material. The used real data and the code for implementing our proposed methods are available at the GitHub repository: {\tt https://github.com/JinyuanChang-Lab/BayesianPenalizedEL}.

{\it Notation.} For any positive integer $q$, write $[q]=\{1,\ldots,q\}$ and let ${\mathbf I}_q$ be the $q \times q$ identity matrix. Denote by $I(\cdot)$ the indicator function. Let ${\rm vech}(\cdot)$ be an operator that stacks the columns of the lower triangular part of its argument square matrix. For a $q$-dimensional vector ${\mathbf a} = (a_1,\ldots,a_q)^{\mathrm{\scriptscriptstyle \top} }$, we use $|{\mathbf a}|_2 = (\sum_{i=1}^q a_i^2)^{1/2}$ and ${\rm supp}({\mathbf a}) = \{i \in [q]: a_i \neq 0\}$ to denote its $L_2$-norm and support, respectively. Let $\mathcal{U}(a,b)$ be the uniform distribution among $(a,b)$, and $\mathcal{N}(\boldsymbol{\mu},\boldsymbol{\Sigma})$ be the Gaussian distribution with mean $\boldsymbol{\mu}$ and covariance matrix $\boldsymbol{\Sigma}$. Denote by $\mathcal{T}_k(\boldsymbol{\mu},\boldsymbol{\Sigma})$ the multivariate Student's distribution with $k$ degrees of freedom, mean $\boldsymbol{\mu}$, and covariance matrix $\boldsymbol{\Sigma}$. For two positive real-valued sequences $\{a_n\}$ and $\{b_n\}$, we write $a_n\lesssim b_n$ if $\limsup_{n\rightarrow\infty}a_n/b_n\leq c_0$ for some positive constant $c_0$, $a_n\asymp b_n$ if $a_n\lesssim b_n$ and $b_n\lesssim a_n$ hold simultaneously, and $a_n\ll b_n$ if $\limsup_{n\rightarrow\infty}a_n/b_n=0$.

Methodology

Penalized Empirical Likelihood

Let $\mathcal{X}_n = \{\mathbf{x}_1, \ldots, \mathbf{x}_n\}$ represent a set of $d$-dimensional independent and identically distributed observations, and let $\boldsymbol{\theta} = (\theta_1, \ldots, \theta_p)^{\mathrm{\scriptscriptstyle \top} } \in \boldsymbol{\Theta}$ be a $p$-dimensional parameter. Here, the parameter space $\boldsymbol{\Theta} \subset \mathbb{R}^p$ is a compact set. The information regarding the model parameter $\boldsymbol{\theta}$ is gathered through a set of unbiased moment conditions ${\mathbb{E}}\{{\mathbf g}({\mathbf x}_{i}; \boldsymbol{\theta}_0)\}={\mathbf 0}$, where $\mathbf{g}(\cdot\,; \cdot) = \{g_1(\cdot\,; \cdot), \ldots, g_r(\cdot\,; \cdot)\}^{\mathrm{\scriptscriptstyle \top} }\in{\mathbb R}^r$ is referred to as the estimating function, and the true, yet unknown value $\boldsymbol{\theta}_0$ is situated within the interior of $\boldsymbol{\Theta}$.

In existing studies, it has been typically required that $r\geq p$ for the identification of $\boldsymbol{\theta}_0$. When $p$ and $r$ are fixed constants, the EL with the estimating function ${\mathbf g}(\cdot\,;\cdot)$ considered in QinLawless_1994_AS can be formulated as

align[align omitted — 294 chars of source]

where $\hat{\Lambda}_n(\boldsymbol{\theta})=\{\boldsymbol{\lambda} \in {\mathbb{R}}^{r}:\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g} ( {\mathbf x}_{i}; \boldsymbol{\theta}) \in \mathcal{V}~\textrm{for any}~ i\in[n] \} $ with an open interval $\mathcal{V}$ containing zero. The standard EL estimator for $\boldsymbol{\theta}_0$ is defined as $\tilde{\boldsymbol{\theta}}_n = \arg \max _{\boldsymbol{\theta} \in \boldsymbol{\Theta}}{\rm EL}(\boldsymbol{\theta})$, which is equivalent to solving the corresponding dual problem:

align[align omitted — 322 chars of source]

The estimator $\tilde{\boldsymbol{\theta}}_n$ exhibits several desirable properties: (i) it is $\sqrt{n}$-consistent, (ii) it possesses asymptotic normality, and (iii) it attains the semiparametric efficiency bound of GodambeHeyde_1987_ISR. However, in high-dimensional scenarios, the literature has highlighted the challenge of accommodating a diverging $r$. This issue is discussed in works such as Donald2003, ChenPengQin_2008, Hjortetal_2008_AS, LengTang_2010, and ChangChenChen_2015. To elaborate, it is generally required that $r \ll n^{1/2}$ for the consistency and $r \ll n^{1/3}$ for the asymptotic normality of $\tilde{\boldsymbol{\theta}}_n$. These constraints on the diverging rate of $r$ pose significant challenges when dealing with high-dimensional estimating equations.

To address scenarios where $r\gg n$ and $p$ remains fixed, we investigate the PEL estimator for $\boldsymbol{\theta}_0$ as follows:

align[align omitted — 378 chars of source]

where $\boldsymbol{\lambda} = (\lambda_{1},\ldots,\lambda_{r})^{\mathrm{\scriptscriptstyle \top} }$, and $P_{\nu}(\cdot)$ is a penalty function with the tuning parameter $\nu$. Given a penalty function $P_{\nu}(\cdot)$ with the tuning parameter $\nu$, we define $\rho(t;\nu)=\nu^{-1}P_\nu(t)$ for $t\in[0,\infty)$ and $\nu\in(0,\infty)$. For $P_\nu(\cdot)$ in (ref), we consider the following class of penalty functions:

align[align omitted — 363 chars of source]

The class $\mathscr{P}$ is broad and general, encompassing commonly used penalty functions. Theorem (ref) in Section (ref) demonstrates that the PEL estimator $\hat{\boldsymbol{\theta}}_n$ follows an asymptotically normal distribution and accommodates exponentially diverging $r$ with respect to $n$.

To practically implement (ref), we encounter a two-layer optimization problem for $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ and $\boldsymbol{\lambda} \in \mathbb{R}^r$. Let

align[align omitted — 256 chars of source]

Since $n^{-1}\sum_{i=1}^n\log\{1+\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} } \mathbf{g}(\mathbf{x}_i;\boldsymbol{\theta})\}$ is concave in $\boldsymbol{\lambda}$, the inner optimization layer of (ref), which seeks $\boldsymbol{\lambda}$ given $\boldsymbol{\theta}$ by maximizing $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$, can be efficiently implemented even for large $r$ when $P_{\nu}(\cdot)$ is chosen as a convex function, such as the $L_1$ penalty. The main challenge is the outer optimization layer of (ref), which seeks the optimizer $\hat{\boldsymbol{\theta}}_n$. This is difficulty due to the nonconvex nature of the problem, making it NP-hard to find global minima Jain2017. As a result, this complexity often leads to computational inefficiency and a higher likelihood of converging to local optima.

Bayesian Penalized Empirical Likelihood

We are motivated to explore an alternative approach using sampling techniques to solve the nonconvex problem associated with PEL. Indeed, as an efficient alternative for addressing nonconvex optimization problems, Ma2019 has demonstrated that solving these issues with MCMC techniques can yield highly effective results. Their findings indicate that the computational complexity of sampling algorithms exhibits linear scalability with the model dimension, in contrast to the exponential scaling of optimization algorithms in nonconvex settings.

Applying sampling techniques to EL in conjunction with a Bayesian framework emerges as a compelling approach. For ${\rm EL}(\boldsymbol{\theta})$ defined as (ref), let $\pi_0(\cdot)$ represent a prior distribution for $\boldsymbol{\theta}$. Then, the posterior distribution $\pi(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ is proportional to $\pi_{0}(\boldsymbol{\theta})\times {\rm EL}(\boldsymbol{\theta})$. In cases where $r$ and $p$ are fixed constants, $\pi(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ converges to a Gaussian distribution with mean being the standard EL estimator $\tilde{\boldsymbol{\theta}}_n$ defined as (ref). Consequently, when samples are successfully drawn from the posterior distribution, their sample mean can serve as an estimator for $\boldsymbol{\theta}_0$.

As the model's complexity increases, BEL faces challenges. In this study, we explore a scenario with high-dimensional model conditions ($r\gg n$), while keeping $p$ fixed. The flexibility by allowing large number $r$ also brings significant challenges. For example, as demonstrated in Tsao_2004_AOS, as $n\rightarrow\infty$, $\mathbb{P}\{{\rm EL}(\boldsymbol{\theta})=0\}\rightarrow1$ for any $\boldsymbol{\theta}$ in a small neighborhood of $\boldsymbol{\theta}_0$ if $r/n\geq0.5$. Such degeneration renders ${\rm EL}(\boldsymbol{\theta})$ inapplicable in this scenario. To handle diverging $r$, we propose to replace ${\rm EL}(\boldsymbol{\theta})$ by

align[align omitted — 351 chars of source]

where $P_{\nu}(\cdot)$ is a penalty function with the tuning parameter $\nu$. Since adding the penalty term $P_{\nu}(\cdot)$ encourages sparse Lagrange multiplier $\boldsymbol{\lambda}$, the PEL effectively performs a selection of the model conditions at each given $\boldsymbol{\theta}$. We then consider the BPEL with a prior distribution $\pi_{0}(\cdot)$, which leads to a posterior distribution defined as

align[align omitted — 220 chars of source]

Our BPEL connects with and differs from the so-called Gibbs posterior in the literature of Bayesian methods Bissiri_2016, Tang2022, Frazier2023. On one hand, they share a common foundation with the Gibbs posterior in that both are built upon generic loss functions. The key difference lies in the device each utilizes: EL employs an appropriate multinomial likelihood, \((p_1, \dots, p_n)\) with \(p_i \ge 0\) and \(\sum^n_{i=1} p_i = 1\), subject to a broad class of model conditions. In contrast, the Gibbs posterior uses a “pseudo-likelihood” proportional to the exponential loss. Furthermore, the inclusion of the penalty on the Lagrange multiplier helps achieve substantial dimension reduction of the problem, which is key in handling high-dimensional problems with many moment conditions. As shown in our numerical studies in Section (ref) and Section (ref) of the supplementary material, MCMC schemes developed from the proposed BPEL demonstrate compelling performance in their finite sample accuracy in approximating the posterior distributions.

Our theory, as elaborated in Section (ref), establishes the fundamental properties of BPEL. Theorem (ref) in Section (ref) demonstrates that the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ defined as (ref) exhibits a Gaussian limiting distribution centered around the PEL estimator $\hat{\boldsymbol{\theta}}_n$ as defined in (ref). Additionally, we define the expected value as

align[align omitted — 228 chars of source]

Corollary (ref) in Section (ref) suggests that $\hat{\boldsymbol{\theta}}_n$ can be effectively approximated by $\mathbb{E}_{\boldsymbol{\theta}\sim\pi^\dag}(\boldsymbol{\theta})$ with an approximation error that diminishes faster than $n^{-1/2}$. This validates the approach to obtain $\hat{\boldsymbol{\theta}}_n$: generating samples from the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ and then using the associated sample mean to approximate $\hat{\boldsymbol{\theta}}_n$. In Section (ref), we will introduce two algorithms designed for implementing BPEL.

The impact of prior specification on the properties of resulting estimators is a notable area of research. For instance, Vexetal_2014_Bioka explores this in the context of EL. In various scenarios, the choice of prior can enhance desirable properties of the estimator derived from the posterior distribution, such as sparsity, as discussed in NarisettyHe2014, Castillo_2015, and OuyangBondell_2023. Given the two primary goals of our study -- developing BPEL and investigating it with MCMC -- we use a non-informative prior in our numerical demonstrations. As detailed in Section (ref) of the supplementary material, we examined the effects of different prior specifications. The overall finding is intuitive: when the prior is specified closer to the true value, the resulting estimator performs better compared to using a non-informative prior. Conversely, if the prior is specified further from the true value, the performance of the estimator deteriorates and becomes less competitive.

MCMC Algorithms

Algorithm 1

In recent decades, MCMC sampling methods have achieved significant success and have garnered influential applications across diverse fields. For an extensive overview of this body of work, we refer to the monograph by Brooks2011 and reference therein. The Metropolis-Hastings (M-H) algorithm family plays a central role in the practical implementation of MCMC techniques, serving as a cornerstone in the toolbox of statisticians and data scientists.

Our first algorithm explores the utilization of the M-H algorithm for BPEL. To accomplish this, we begin by specifying a proposal distribution with a density function denoted as $\phi(\cdot\,|\,{\mathbf x})$, where ${\mathbf x} \in \mathbb{R}^p$. Subsequently, we employ the M-H algorithm to generate samples from the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$, as defined in (ref). The specific steps for this process are detailed in Algorithm (ref).

algorithm[algorithm omitted — 2,039 chars of source]

At each iteration $k$, Algorithm (ref) begins with a state $\boldsymbol{\theta}^k \in \boldsymbol{\Theta}$. In the proposal step, it generates a new parameter $\boldsymbol{\vartheta}^{k+1}$ from the proposal distribution centered at $\boldsymbol{\theta}^k$, denoted by $\phi(\cdot\,|\,\boldsymbol{\theta}^k)$. Following this, in the accept-reject step, Algorithm (ref) decides whether to accept $\boldsymbol{\vartheta}^{k+1}$ with a probability denoted as $\alpha^{k+1}$. This crucial step ensures that the Markov chain, guided by Algorithm (ref), remains within the valid parameter space $\boldsymbol{\Theta}$. Consequently, it expedites the convergence of the resulting chain towards its stationary distribution, which is the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$. There exist various approaches for selecting the proposal distribution with density $\phi(\cdot\,|\,\cdot)$, including methods like the symmetric Metropolis algorithm, random walk M-H, and the independence sampler, as detailed by Roberts2004.

Algorithm 2

Another widely-used MCMC technique is Importance Sampling Ripley2006, Hesterberg1995. This method involves generating samples from a proposal distribution and then applying importance weights to these samples to account for the disparities between the proposal distribution and the target distribution. In practical applications, recycling successive samples often proves to be an effective strategy Marin2019, particularly when the computation of importance weights is computationally intensive. In this context, CORNUET2012 introduces the Adaptive Multiple Importance Sampling (AMIS) algorithm, which combines various importance sampling methods with adaptive techniques. The integration of the AMIS approach with EL, as shown in Mengersen2013, is particularly compelling. To ensure the consistency of AMIS, Marin2019 introduces a modified variant called Modified AMIS (MAMIS) with a simpler recycling strategy compared to AMIS.

We present and investigate an MAMIS algorithm, as outlined in Algorithm (ref), specifically designed for computing BPEL. This algorithm operates in a scenario where a density function $\varphi(\cdot\,; \boldsymbol{\zeta})$ is defined, with $\boldsymbol{\zeta}$ representing a parameter in ${\mathbb{R}}^s$, and where an explicit function ${\mathbf h}: {\mathbb{R}}^p \mapsto {\mathbb{R}}^s$ is known. This configuration allows us to generate weighted samples that effectively capture the characteristics of the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$, as defined in (ref).

algorithm[algorithm omitted — 1,963 chars of source]

Algorithm (ref) generates a sequence of samples while progressively adjusting the parameter $\boldsymbol{\zeta} \in {\mathbb{R}}^s$ involved in the proposal distribution. At each iteration $k$ of Algorithm (ref), the new value for the parameter $\boldsymbol{\zeta}$ of the proposal distribution is determined based on the most recent $N_k$ samples drawn. This represents the primary distinction between the MAMIS algorithm by Marin2019 and the AMIS algorithm by CORNUET2012. Specifically, MAMIS updates the proposal distribution parameter using only the last $N_k$ samples at iteration $k$, while AMIS updates this parameter by considering all past $\sum_{j=1}^{k}N_j$ samples. The end product output of Algorithm (ref) is generated by updating the importance weights for all samples produced during the recycling process.

Sampling vs Optimizations

We advocate the utilization of sampling techniques as a practical and efficient alternative to optimization methods for addressing computationally challenging PEL problems. Specifically for obtaining the estimator $\hat{\boldsymbol{\theta}}_n$ as defined in (ref), we can rely on samples $\boldsymbol{\theta}^1,\ldots,\boldsymbol{\theta}^K$ generated from the M-H algorithm (see Algorithm (ref)), estimating $\mathbb{E}_{\boldsymbol{\theta}\sim\pi^\dag}(\boldsymbol{\theta})$, as defined in (ref), by computing the sample mean, i.e., $K^{-1}\sum_{k=1}^K\boldsymbol{\theta}^k$. When employing the MAMIS algorithm (see Algorithm (ref)) and completing $K$ iterations, the estimator for $\mathbb{E}_{\boldsymbol{\theta}\sim\pi^\dag}(\boldsymbol{\theta})$ is determined as a weighted average:

align[align omitted — 179 chars of source]

where $S_K=N_1+\cdots+N_K$.

Our theory in Section (ref) supports the use of sampling algorithms as efficient alternatives. For the M-H algorithm, Theorem (ref) in Section (ref) demonstrates that, conditional on $\mathcal{X}_n$, the average $K^{-1}\sum_{k=1}^K\boldsymbol{\theta}^k$ converges almost surely to $\mathbb{E}_{\boldsymbol{\theta}\sim\pi^\dag}(\boldsymbol{\theta})$ as $K\rightarrow\infty$. For the MAMIS algorithm, Theorem (ref) in Section (ref) establishes that, conditional on $\mathcal{X}_n$, $\widehat{{\mathbb{E}}}_{\pi^\dag,K} (\boldsymbol{\theta})$ in (ref) converges almost surely to $\mathbb{E}_{\boldsymbol{\theta}\sim\pi^\dag}(\boldsymbol{\theta})$ as $K\rightarrow\infty$. These results, combined with Corollary (ref) in Section (ref), validate the properties of BPEL estimators obtained through these established sampling techniques. Another consideration in Algorithms (ref) and (ref) is the choice of the initial point, denoted, respectively, as $\boldsymbol{\theta}^{0}$ and $\hat{\boldsymbol{\zeta}}_1$. Our theoretical analyses only require $\boldsymbol{\theta}^0 \in \boldsymbol{\Theta}$ satisfying $\pi^\dag(\boldsymbol{\theta}^0\,|\,\mathcal{X}_n)>0$ and do not impose any restriction on $\hat{\boldsymbol{\zeta}}_1$; see Theorems (ref) and (ref) in Section (ref) for details. Our empirical simulation studies in Section (ref) consistently demonstrate the proposed algorithms' robust performance, irrespective of the initial value chosen. Notice that the performance of the optimization methods for the nonconvex optimization problems usually depends crucially on the choice of the initial point. The combination of theoretical analysis and empirical evidence underscores that, in comparison to competing optimization methods, these sampling-based approaches offer significant advantages in terms of convergence speed, stability across replications, and resilience to variations in initial values. This reaffirms the benefits of incorporating BPEL into the methodology.

The M-H and MAMIS algorithms each have their strengths. M-H is easy to implement, but high rejection rates can reduce its efficiency, especially with a poorly tuned proposal distribution. MAMIS, while requiring more effort -- particularly in computing importance weights -- offers improved sampling efficiency and is less sensitive to the proposal distribution, making it ideal for complex posterior distributions. Choosing between these algorithms depends on the specific problem and the balance between implementation ease and sampling efficiency.

Numerical Studies

Data Generation Process

We conduct simulation studies to empirically assess the performance of our proposed methods. For the data generation process (DGP), we adopt the structural equation $y_i = \hbar({\mathbf u}_i^{\mathrm{\scriptscriptstyle \top} } \boldsymbol{\theta}_0) + e_i^{(0)}$, $i \in [n]$, where $\hbar: {\mathbb{R}} \mapsto {\mathbb{R}}$ is a continuous function, $e_i^{(0)}$ is the error, and ${\mathbf u}_i=(u_{i,1},u_{i,2})^{\mathrm{\scriptscriptstyle \top} }$ represents two endogenous variables. The set of all instrumental variables (IVs) is denoted as ${\mathbf z}_{i}=(z_{i,1}, \ldots, z_{i,r})^{\mathrm{\scriptscriptstyle \top} }$ for $i \in [n]$. The true reduced-form equations for the endogenous variables are specified as $u_{i,1}=0.5z_{i,1}+0.5z_{i,2}+e_i^{(1)}$ and $ u_{i,2}=0.5z_{i,3}+0.5z_{i,4}+e_i^{(2)}$, where $(e_i^{(1)}, e_i^{(2)})$ represents the random errors. Essentially, each of the two endogenous variables is influenced by only two IVs. All IVs are selected orthogonal to the error term $e_i^{(0)}$. Hence, we have $ {\mathbb{E}} \{y_i - \hbar({\mathbf u}_i^{\mathrm{\scriptscriptstyle \top} } \boldsymbol{\theta}_0)\,|\, {\mathbf z}_{i}\} ={\mathbf 0} $, which implies that $\boldsymbol{\theta}_0$ can be identified by the $r$ unbiased moment conditions $\mathbb{E}\{{\mathbf g}({\mathbf x}_i;\boldsymbol{\theta}_0)\}={\mathbf 0}$, where ${\mathbf g}({\mathbf x}_i;\boldsymbol{\theta}) = \{y_i - \hbar({\mathbf u}_i^{\mathrm{\scriptscriptstyle \top} } \boldsymbol{\theta}) \}{\mathbf z}_i $ with ${\mathbf x}_i= (y_i, {\mathbf u}_i^{\mathrm{\scriptscriptstyle \top} }, {\mathbf z}_i^{\mathrm{\scriptscriptstyle \top} })^{\mathrm{\scriptscriptstyle \top} }$. In the DGP, we generate ${\mathbf z}_{i}\sim\mathcal{N}({\mathbf 0},{\mathbf I}_r)$, and

align*[align* omitted — 348 chars of source]

We set $\boldsymbol{\theta}_0=(0.5,0.5)^{\mathrm{\scriptscriptstyle \top} }$ and consider two selections for the link function $\hbar(\cdot)$: (i) the linear case with $\hbar(v)=v$, and (ii) the nonlinear case with $\hbar(v)=\sin v$.

Sampling Efficiency and Stability

We begin by demonstrating the improvement in sampling efficiency achieved through the use of PEL. In this context, we generate data following the DGP with linear link function $\hbar(\cdot)$ by setting $n=120$ and varying $r$ in the range $[50,1000]$. We aim to sample from two posterior distributions $\pi_{0}(\boldsymbol{\theta}) \times {\rm EL}(\boldsymbol{\theta})$ and $\pi_{0}(\boldsymbol{\theta}) \times {\rm PEL}_{\nu}(\boldsymbol{\theta})$, where ${\rm EL}(\boldsymbol{\theta})$ and ${\rm PEL}_{\nu}(\boldsymbol{\theta})$ are, respectively, given in (ref) and (ref). Evaluating ${\rm PEL}_{\nu}(\boldsymbol{\theta})$ involves an optimization problem that solves for $\boldsymbol{\lambda}$ by maximizing the objective function $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ defined as (ref) at given $\boldsymbol{\theta}$. To ensure the attainment of a sparse Lagrange multiplier and maintain the convexity of the objective function, we select $P_{\nu}(\cdot)$ as the $L_1$ penalty function. In practice, since the prior information about the true parameter $\boldsymbol{\theta}_0$ is typically unavailable, we select $\pi_0(\cdot)$ as the improper uniform prior. We implement Algorithm (ref) to sample from both posterior distributions using identical settings, employing a proposal distribution $\mathcal{N}(\boldsymbol{\theta}, \sigma^2{\mathbf I}_p)$ with $\sigma^2 = 10^{-4}$ and initializing from $\boldsymbol{\theta}^0 = (0.3, 0.3)^{\mathrm{\scriptscriptstyle \top} }$. In the case of PEL, we set the tuning parameter $\nu=0.03$ involved in ${\rm PEL}_{\nu}(\boldsymbol{\theta})$.

To compare efficiency, we measure the number of iterations required to obtain the same number of accepted samples. Figure (ref) illustrates the average number of iterations needed over 500 runs to accept 5 samples for different values of $r$, thereby providing a comparison between using EL and PEL within a Bayesian framework. The sampling efficiency of Algorithm (ref) when using ${\rm PEL}_{\nu}(\boldsymbol{\theta})$ is notably superior to that achieved with ${\rm EL}(\boldsymbol{\theta})$. The selection of $\sigma^2$ within the proposal distribution closely influences the acceptance rate in each step of the M-H algorithm. With our small choice of $\sigma^2$ in the simulation, the M-H algorithm should efficiently generate valid samples. It is worth highlighting that the acceptance rate remains consistently high and stable when using PEL across all $r$ settings. In contrast, when employing EL without any penalty, it may require thousands more iterations to achieve the same number of accepted samples. Additionally, it is evident that the M-H algorithm with EL becomes increasingly unstable as $r$ increases.

figure[figure omitted — 198 chars of source]

Comparison with the Optimization Methods

As we suggested in Section (ref), the computation of the PEL estimator $\hat{\boldsymbol{\theta}}_n$ defined as (ref) can be implemented using Algorithm (ref) (referred to as M-H) and Algorithm (ref) (referred to as MAMIS). In this part, we compare their performance with two optimization methods: (a) {\tt optim}: A versatile R function for general-purpose optimization of objective functions, supporting various optimization algorithms like Nelder-Mead, quasi-Newton, and conjugate-gradient; and (b) {\tt nlm}: An R function specialized in non-linear optimization, particularly designed for finding minima of objective functions using Newton-type algorithms.

The choice of the proposal distribution plays a crucial role in achieving efficient sampling with BPEL. Within the context of the M-H algorithm, one commonly used scheme is the random walk M-H, where the proposal distribution takes the form of a Gaussian distribution $\mathcal{N}(\boldsymbol{\theta}, \sigma^2 {\mathbf I}_p)$ with the current state denoted as $\boldsymbol{\theta}$. It is essential to carefully select an appropriate value for $\sigma^2$. A small value for $\sigma^2$ can result in slow exploration of the state space, while a large value can lead to decreased acceptance rates, subsequently slowing down the algorithm. To strike a balance between exploration and acceptance rates, we can monitor the acceptance rate of the algorithm. In the simulation for M-H, we set $\sigma^2=C(n\log r)^{-1}$ with some constant $C>0$. We adjust the value of $C$ until the acceptance rate closely matches the desired rate, typically aiming for approximately $0.234$, as suggested in Gelman1997. It is known that the M-H algorithm requires some time to converge to its stationary distribution, especially when the initial point $\boldsymbol{\theta}^0 \in \boldsymbol{\Theta}$ is situated in the tails of the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$. Considering this, we set a burn-in period of $500$ iterations. For the MAMIS algorithm, we adhere to recommendations from CORNUET2012 and Mengersen2013 that advocate for the adoption of $\mathcal{T}_3(\boldsymbol{\mu},\boldsymbol{\Sigma})$ as the proposal distribution. During each iteration $k$ of MAMIS, we calculate the updated value $\hat\boldsymbol{\zeta}_{k+1} = \{\hat{\boldsymbol{\mu}}_{k+1}^{\mathrm{\scriptscriptstyle \top} }, {\rm vech}(\widehat{\boldsymbol{\Sigma}}_{k+1})^{\mathrm{\scriptscriptstyle \top} }\}^{\mathrm{\scriptscriptstyle \top} }$ for the parameter vector $\boldsymbol{\zeta} = \{\boldsymbol{\mu}^{\mathrm{\scriptscriptstyle \top} }, {\rm vech}(\boldsymbol{\Sigma})^{\mathrm{\scriptscriptstyle \top} }\}^{\mathrm{\scriptscriptstyle \top} }$ involved in the proposal distribution $\mathcal{T}_3(\boldsymbol{\mu},\boldsymbol{\Sigma})$ as $\hat{\boldsymbol{\mu}}_{k+1} = N_k^{-1} \sum_{i=1}^{N_k} \omega^k_i \boldsymbol{\theta}^k_i$ and ${\rm vech}(\widehat{\boldsymbol{\Sigma}}_{k+1}) = N_k^{-1} \sum_{i=1}^{N_k} \omega^k_i {\rm vech}\{(\boldsymbol{\theta}^k_i - \hat{\boldsymbol{\mu}}_{k+1})(\boldsymbol{\theta}^k_i - \hat{\boldsymbol{\mu}}_{k+1})^{\mathrm{\scriptscriptstyle \top} }\}$, where $\omega^k_i$ represents the corresponding importance weights, as outlined in Algorithm (ref). In our simulations, we initialize $\widehat{\boldsymbol{\Sigma}}_{1} = {\mathbf I}_p$, and the selection of $\hat{\boldsymbol{\mu}}_{1}$ is described in the next paragraph.

We conduct 200 replications following the DGP and explore various combinations of dimensionalities. Specifically, we consider $n\in\{120, 240\}$ and $r\in\{80, 160, 320, 640\}$. To assess the robustness of these methods with respect to initial points, we select $49$ equally spaced grid points on the plane within the range of $[-3, 4] \times [-3, 4]$ as our chosen initial points. In the case of MAMIS, which is not an iterative algorithm, we set these initial points as the initial means $\hat{\boldsymbol{\mu}}_{1}$ for its proposal distribution $\mathcal{T}_3(\boldsymbol{\mu},\boldsymbol{\Sigma})$ to facilitate comparison. In our simulations, we identify the true global minima $\hat{\boldsymbol{\theta}}_n$ defined as (ref) through exhaustive search. To achieve this, in each replication of the simulation (indexed by $k$), we generate a grid of $10201$ equally spaced points within the range $[-0.5, 1.5]\times[-0.5, 1.5]$. We then compute the posterior probabilities for these points and selected $\boldsymbol{\theta}^{\rm mode}_{k}$ as the point with the highest probability. Since $\pi_0(\cdot)$ is selected as the improper uniform prior, $\boldsymbol{\theta}^{\rm mode}_k$ is actually the required true global minima in the $k$-th replication. We repeat this process for $k=1$ to $200$, and compare the outcomes obtained from both optimization and sampling methods by calculating the measure $$ {\rm MSE}_1 = \frac{1}{200 \times 49} \sum_{k=1}^{200} \sum_{l=1}^{49} |\check{\boldsymbol{\theta}}_{k}(l) - \boldsymbol{\theta}^{\rm mode}_{k}|_2^2\,.$$ Here, $\check{\boldsymbol{\theta}}_{k}(l)$ represents the related outcome in the $k$-th replication initiated from the $l$-th initial point.

In the context of BPEL sampling, we explore three scenarios with varying sample sizes of $1500$, $2500$, and $3500$, which we label as (M-H-1, M-H-2, M-H-3) and (MAMIS-1, MAMIS-2, MAMIS-3), respectively, for Algorithms 1 and 2. Additionally, we conduct an investigation into the influence of different values for the tuning parameter $\nu$. Table (ref) presents the simulation results. The overall performance of the sampling approaches surpasses that of the optimization methods. Notably, for the nonlinear model, the optimization using the R function {\tt nlm} is proven to be unreliable, resulting in highly unstable results. As the size of the generated samples increases, the performance of BPEL improves. Both M-H and MAMIS exhibit promising performance in both linear and nonlinear cases. For the nonlinear models, MAMIS significantly outperforms M-H, possibly owing to the advantages gained from employing importance weights for parameter estimation. The role of the tuning parameter $\nu$ is pivotal, underscoring the merits from using the PEL approach in achieving more parsimonious models by effectively selecting most useful model conditions within the constraints of the available data information. When using very small values of $\nu$, such as 0.01, the performance of the methods becomes less satisfactory. Overall, the BPEL performs satisfactorily for a reasonable range of choices for $\nu$.

sidewaystable\setlength\tabcolsep{3pt} \caption{Comparison of BPEL and optimization methods} \begin{spacing}{1} \begin{tabular} {cl|cccc|cccc|cccc|cccc} \hline & & \multicolumn{4}{c|}{{$\hbar(v) = v, n=120$} } & \multicolumn{4}{c|}{{$\hbar(v)=\sin v, n=120$}} & \multicolumn{4}{c|}{{$\hbar(v) = v, n=240$ }} & \multicolumn{4}{c}{{$\hbar(v)=\sin v, n=240$}} \\ {$\nu$} & {Methods} & {$r=80$} & {$r=160$} & {$r=320$} & {$r=640$} & {$r=80$} & {$r=160$} & {$r=320$} & {$r=640$} & {$r=80$} & {$r=160$} & {$r=320$} & {$r=640$} & {$r=80$} & {$r=160$} & {$r=320$} & {$r=640$} \\ \hline 0.01 & {MAMIS-1} & 0.0606 & 0.0432 & 0.0371 & 0.0336 & 8.8598 & 7.0955 & 6.7810 & 6.7583 & 0.2600 & 0.0467 & 0.0297 & 0.0250 & 9.7398 & 7.7113 & 6.9972 & 6.7786 \\ & {MAMIS-2} & 0.0062 & 0.0020 & 0.0016 & 0.0015 & 8.2760 & 6.4363 & 6.0974 & 6.0576 & 0.1139 & 0.0047 & 0.0004 & 0.0004 & 9.1874 & 7.1095 & 6.4394 & 6.1002 \\ & {MAMIS-3} & 0.0023 & 0.0014 & 0.0014 & 0.0013 & 7.8634 & 6.0010 & 5.6087 & 5.6073 & 0.0911 & 0.0018 & 0.0002 & 0.0003 & 8.8130 & 6.6946 & 6.0186 & 5.6339 \\ & {M-H-1} & 0.0457 & 0.0015 & 0.0016 & 0.0015 & 12.7310 & 12.2970 & 12.2665 & 12.3473 & 0.4498 & 0.0371 & 0.0082 & 0.0004 & 13.2222 & 12.2080 & 12.0538 & 11.9004 \\ & {M-H-2} & 0.0411 & 0.0015 & 0.0015 & 0.0015 & 12.6824 & 12.2486 & 12.1999 & 12.3289 & 0.4279 & 0.0311 & 0.0067 & 0.0003 & 13.2205 & 12.1876 & 12.0544 & 11.8679 \\ &{M-H-3} & 0.0389 & 0.0014 & 0.0015 & 0.0014 & 12.6477 & 12.2179 & 12.1564 & 12.3042 & 0.4146 & 0.0277 & 0.0053 & 0.0003 & 13.2281 & 12.1745 & 12.0461 & 11.8451 \\ & {optim} & 0.2822 & 0.0720 & 0.0133 & 0.0132 & 12.2067 & 12.0219 & 11.9113 & 11.9760 & 1.0834 & 0.2247 & 0.0582 & 0.0160 & 12.8270 & 12.2083 & 12.2197 & 12.1037 \\ & {nlm} & 0.0956 & 0.0341 & 0.0222 & 0.0151 & 117934.2 & 115037.8 & 85081.4 & 83013.3 & 6.2781 & 0.0674 & 0.0163 & 0.0076 & 121272.7 & 113263.9 & 104125.3 & 87304.6 \\ \hline 0.03 & {MAMIS-1} & 0.0465 & 0.0363 & 0.0317 & 0.0311 & 7.3502 & 6.2407 & 5.8879 & 6.2055 & 0.0707 & 0.0345 & 0.0279 & 0.0264 & 7.2303 & 6.6040 & 6.5472 & 6.3108 \\ & {MAMIS-2} & 0.0021 & 0.0011 & 0.0008 & 0.0009 & 6.5728 & 5.4791 & 5.0858 & 5.3622 & 0.0033 & 0.0012 & 0.0004 & 0.0004 & 6.3727 & 5.8076 & 5.8222 & 5.6557 \\ & {MAMIS-3} & 0.0008 & 0.0008 & 0.0007 & 0.0007 & 6.0018 & 4.9604 & 4.5459 & 4.7549 & 0.0007 & 0.0009 & 0.0002 & 0.0002 & 5.7517 & 5.2972 & 5.2865 & 5.1854 \\ & {M-H-1} & 0.0135 & 0.0009 & 0.0007 & 0.0008 & 12.5367 & 12.0546 & 12.1103 & 12.2611 & 0.0522 & 0.0034 & 0.0003 & 0.0002 & 12.8677 & 12.1031 & 12.1453 & 12.1239 \\ &{M-H-2} & 0.0116 & 0.0009 & 0.0007 & 0.0008 & 12.4304 & 11.9548 & 12.0239 & 12.1413 & 0.0450 & 0.0034 & 0.0003 & 0.0002 & 12.7587 & 12.0285 & 12.0758 & 12.0806 \\ &{M-H-3} & 0.0104 & 0.0008 & 0.0007 & 0.0007 & 12.3589 & 11.8816 & 11.9509 & 12.0655 & 0.0388 & 0.0034 & 0.0002 & 0.0002 & 12.6799 & 11.9762 & 12.0297 & 12.0417 \\ & {optim} & 0.0587 & 0.0109 & 0.0014 & 0.0014 & 12.8445 & 12.5361 & 12.3325 & 12.5015 & 0.2977 & 0.0454 & 0.0125 & 0.0005 & 13.2494 & 12.7951 & 12.7153 & 12.8332 \\ & {nlm} & 0.0385 & 0.0033 & 0.0015 & 0.0058 & 79382.0 & 73390.0 & 76096.0 & 94794.3 & 0.1321 & 0.0127 & 0.0057 & 0.0047 & 90321.8 & 70994.2 & 90414.5 & 79523.2 \\ \hline 0.05 & {MAMIS-1} & 0.0451 & 0.0359 & 0.0331 & 0.0295 & 6.0750 & 5.5727 & 5.3221 & 5.3884 & 0.0598 & 0.0387 & 0.0305 & 0.0261 & 5.8614 & 5.5769 & 5.6081 & 5.8757 \\ & {MAMIS-2} & 0.0025 & 0.0018 & 0.0011 & 0.0008 & 5.0407 & 4.6182 & 4.3750 & 4.3938 & 0.0043 & 0.0024 & 0.0010 & 0.0007 & 4.7997 & 4.6707 & 4.7641 & 4.9901 \\ & {MAMIS-3} & 0.0017 & 0.0016 & 0.0008 & 0.0007 & 4.4061 & 3.9666 & 3.7920 & 3.7718 & 0.0027 & 0.0019 & 0.0009 & 0.0005 & 4.0542 & 4.0482 & 4.1549 & 4.3903 \\ & {M-H-1} & 0.0076 & 0.0015 & 0.0010 & 0.0008 & 12.0832 & 11.7687 & 11.8703 & 11.8866 & 0.0063 & 0.0018 & 0.0010 & 0.0006 & 12.2347 & 11.8718 & 12.0358 & 12.0665 \\ &{M-H-2} & 0.0074 & 0.0014 & 0.0009 & 0.0007 & 11.8869 & 11.5660 & 11.6859 & 11.7368 & 0.0059 & 0.0016 & 0.0009 & 0.0005 & 12.0296 & 11.7100 & 11.9241 & 11.9669 \\ & {M-H-3} & 0.0074 & 0.0013 & 0.0009 & 0.0007 & 11.7189 & 11.4244 & 11.5409 & 11.6131 & 0.0056 & 0.0016 & 0.0009 & 0.0005 & 11.8578 & 11.5979 & 11.8305 & 11.8745 \\ & {optim} & 0.0209 & 0.0010 & 0.0001 & 0.0000 & 13.2051 & 12.8147 & 12.6047 & 12.7401 & 0.0827 & 0.0130 & 0.0073 & 0.0000 & 13.5813 & 13.0596 & 13.0502 & 13.0375 \\ & {nlm} & 0.0167 & 0.0002 & 0.0004 & 0.0009 & 69229.7 & 58220.3 & 61362.5 & 63215.1 & 0.0387 & 0.0023 & 0.0032 & 0.0052 & 48458.3 & 69014.4 & 62235.7 & 80570.4 \\ \hline 0.07 & {MAMIS-1} & 0.0454 & 0.0381 & 0.0323 & 0.0293 & 4.6585 & 4.7192 & 4.4856 & 4.5785 & 0.0534 & 0.0381 & 0.0341 & 0.0278 & 4.5237 & 4.5635 & 4.9180 & 5.0988 \\ & {MAMIS-2} & 0.0049 & 0.0031 & 0.0023 & 0.0015 & 3.5217 & 3.6260 & 3.5231 & 3.5504 & 0.0071 & 0.0058 & 0.0036 & 0.0023 & 3.2789 & 3.5005 & 3.8945 & 4.1681 \\ & {MAMIS-3} & 0.0042 & 0.0029 & 0.0019 & 0.0014 & 2.8438 & 2.9326 & 2.9355 & 2.9075 & 0.0066 & 0.0054 & 0.0033 & 0.0021 & 2.5653 & 2.8253 & 3.2818 & 3.5420 \\ & {M-H-1} & 0.0065 & 0.0031 & 0.0020 & 0.0016 & 11.6003 & 11.3310 & 11.3530 & 11.5029 & 0.0069 & 0.0054 & 0.0036 & 0.0022 & 11.7248 & 11.5753 & 11.6807 & 12.0016 \\ & {M-H-2} & 0.0063 & 0.0030 & 0.0020 & 0.0015 & 11.2135 & 10.9535 & 11.1190 & 11.2268 & 0.0067 & 0.0053 & 0.0035 & 0.0021 & 11.3384 & 11.2954 & 11.4157 & 11.8458 \\ & {M-H-3} & 0.0062 & 0.0029 & 0.0019 & 0.0015 & 10.9140 & 10.7093 & 10.8781 & 11.0160 & 0.0067 & 0.0053 & 0.0035 & 0.0021 & 11.0287 & 11.0815 & 11.2360 & 11.7241 \\ & {optim} & 0.0117 & 0.0000 & 0.0007 & 0.0000 & 13.2631 & 13.0115 & 12.8281 & 12.8492 & 0.0169 & 0.0033 & 0.0007 & 0.0000 & 13.5938 & 13.3118 & 13.2545 & 13.2196 \\ & {nlm} & 0.0119 & 0.0000 & 0.0017 & 0.0010 & 49184.4 & 62755.1 & 63796.5 & 71037.4 & 0.0076 & 0.0015 & 0.0061 & 0.0006 & 72324.7 & 59709.3 & 80154.4 & 62439.4 \\ \hline \end{tabular} \end{spacing}

Comparison with Competing Methods

In this part, we compare the PEL estimator $\hat{\boldsymbol{\theta}}_n$ defined as (ref) with two other estimators: the standard EL estimator $\tilde{\boldsymbol{\theta}}_n$ defined as (ref) and the relaxed EL (REL) estimator introduced by Shi2016. The REL is tailored for high-dimensional estimating equations, making it resilient to minor deviations from the equality constraints. Notice that the standard EL can only work for low-dimensional estimating equations. In line with our model specifications, where the two endogenous variables $u_{i,1}$ and $u_{i,2}$ are linked to IVs ($z_{i,1}, z_{i,2}$ and $z_{i,3}, z_{i,4}$, respectively) for each $i\in[n]$, we only use the first four moment conditions, that are related to the IVs $z_{i,1}, z_{i,2}, z_{i,3}$ and $z_{i,4}$, to produce the standard EL estimator $\tilde{\boldsymbol{\theta}}_n$. The computation of $\tilde{\boldsymbol{\theta}}_n$ can be implemented by the function {\tt gel} in the R-package {\tt gmm}. For both our PEL estimator $\hat{\boldsymbol{\theta}}_n$ and the REL estimator, we use all the $r$ moment conditions.

For the selection of the tuning parameter in the REL estimator, we follow the recommendation in Shi2016, using a consistent tuning parameter $0.5n^{-1/2}(\log r)^{1/2}$ throughout the simulations. Regarding the tuning parameter $\nu$ in our BPEL, we employ the Bayesian Information Criterion (BIC) defined as

align[align omitted — 229 chars of source]

for its selection, where $\hat{\boldsymbol{\theta}}_{n}^{(\nu)}$ is the associated PEL estimator with tuning parameter $\nu$ calculated by our sampling algorithm, and $\mathcal{R}_n^{(\nu)}={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n^{(\nu)})\}$ with $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n^{(\nu)})=(\hat\lambda_1^{(\nu)},\ldots,\hat\lambda_r^{(\nu)})^{\mathrm{\scriptscriptstyle \top} }=\arg\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_n(\hat{\boldsymbol{\theta}}_n^{(\nu)})}f_n(\boldsymbol{\lambda};\hat{\boldsymbol{\theta}}_n^{(\nu)})$ with $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ defined as (ref). In practice, we set $\mathcal{R}_n^{(\nu)} = \{j \in [r]: |\hat\lambda_j^{(\nu)}|>10^{-6}\}$ and restrict $\nu$ in the interval $[0.05n^{-1/2}(\log r)^{1/2}, 0.75n^{-1/2}(\log r)^{1/2}]$.

For the same $49$ initial points of the $200$ replications mentioned in Section (ref), we calculate the measure $$ {\rm MSE}_2 = \frac{1}{200 \times 49} \sum_{k=1}^{200} \sum_{l=1}^{49} |\check{\boldsymbol{\theta}}_{k}(l) - \boldsymbol{\theta}_0|_2^2 $$ to evaluate the performance of different estimators, where $\check{\boldsymbol{\theta}}_{k}(l)$ is the related estimator in the $k$-th replication initiated from the $l$-th initial point. Table (ref) compares the measure ${\rm MSE}_2$ for the three estimators: the PEL estimator (MAMIS, M-H), the standard EL estimator, and the REL estimator. The results for M-H and MAMIS are derived based on the generated samples of size $3500$. It becomes clear that BPEL demonstrates substantial performance improvements, clearly establishing its superiority over the other estimation methods. Particularly noteworthy is the effectiveness of MAMIS in addressing the challenges posed by nonlinear estimating equations, showing its promising performance.

table[table omitted — 1,412 chars of source]

Additional Numerical Studies

We provide additional simulation studies in the supplementary material: Section (ref) examines the impact of prior specification, Section (ref) evaluates the performance of our method using an alternative data generation process with data from a Student's $t$-distribution instead of a normal distribution, Section (ref) assesses the finite sample accuracy of the MCMC algorithms in approximating the posterior distribution, Section (ref) compares the posterior distributions resulting from different Bayesian EL formulations, and Section (ref) presents the comparison between our method and two competing methods: approximate Bayesian computation and Bayesian synthetic likelihood. Overall, our findings confirm the highly competitive performance of the proposed BPEL with the MCMC framework in terms of finite sample performance and accuracy in approximating posterior distributions.

Real Data Analysis

International trade refers to the cross-border exchange of capital, commodities, and services between nations or regions. This type of trade typically constitutes a substantial portion of a country's gross domestic product (GDP). Eaton2011, hereafter referred to as EKK, combined an empirical model with microeconomic principles to analyze France's international trade patterns. Additionally, Shi (2016) utilized EKK's microeconomic model to derive parameter estimates for Chinese exporting companies. In this section, we reexamine the dataset previously examined in Shi2016, employing the proposed BPEL approach.

The model proposed by EKK comprises five parameters denoted as $\boldsymbol{\theta} = (\theta_1, \ldots, \theta_5)^{\mathrm{\scriptscriptstyle \top} } \in \boldsymbol{\Theta}$. The first component, $\theta_1$, characterizes the distribution of production efficiency among firms, with a higher $\theta_1$ indicating a larger proportion of manufacturers with lower efficiency. The second component, $\theta_2$, quantifies the cost associated with accessing a fraction of potential buyers, where a higher $\theta_2$ corresponds to lower costs. Parameters $\theta_3$, $\theta_4$ and $\theta_5$ represent the standard deviation of the demand shock, the standard deviation of the entry cost shock, and the correlation coefficient between these two shocks, respectively. Each firm is identified by the index $i \in [n]$, while countries are represented by the index $j \in \{0\}\cup[r]$, with $j = 0$ denoting the home country.

According to the EKK's model, the sales of firm $i$ in country $j$ is $Z_{i,j}(\boldsymbol{\theta}; e^{(1)}_{i,j}, e^{(2)}_{i,j}, e^{(3)}_{i}) = \kappa \bar{Z}_j (1-\tau_{i,j})^{\theta_2/\theta_1} \tau_{i,j}^{-1/\theta_1} a^{(1)}_{i,j}/a^{(2)}_{i,j}$, where $a^{(1)}_{i,j}= \exp\{\theta_3(1-\theta_5^2)^{1/2}e^{(1)}_{i,j} + \theta_3\theta_5 e^{(2)}_{i,j}\}$, $a^{(2)}_{i,j} = \exp\{\theta_4e^{(2)}_{i,j}\}$, $\tau_{i,j} = \min\{1, e^{(3)}_{i} \bar{u}_{i} / \bar{u}_{i,j}\}$ and

equation*[equation* omitted — 262 chars of source]

with $\bar{u}_{i,j} = (a^{(2)}_{i,j})^{\theta_1} N_j$ and $ \bar{u}_{i} = \min\{\bar{u}_{i,0}, \max_{j\in[r]}\bar{u}_{i,j}\}$, and $(\bar{Z}_j, N_j)_{j\in\{0\}\cup[r]}$ are known constants. Here $e^{(1)}_{i,j} \sim \mathcal{N}(0,1)$, $e^{(2)}_{i,j} \sim \mathcal{N}(0,1)$ and $e^{(1)}_{i} \sim \mathcal{U}(0,1)$ are mutually independent. Furthermore, $Z_{i,j}(\boldsymbol{\theta}; e^{(1)}_{i,j}, e^{(2)}_{i,j}, e^{(3)}_{i}) = 0$ means that the firm $i$ is kept outside of the country $j$. As a pertinent economic indicator of our interest, the mean sale of all firms in country $j$ is $\mu_j(\boldsymbol{\theta}) = \mathbb{E}\{Z_{i,j}(\boldsymbol{\theta}; e^{(1)}_{i,j}, e^{(2)}_{i,j}, e^{(3)}_{i})\}$, where the expectation is taken respect to the random variables $\{e^{(1)}_{i,j}, e^{(2)}_{i,j}, e^{(3)}_{i}\}$. The dataset is sourced from the Chinese administrative databases, encompassing a total of $n=6754$ firms and their export data to $r=126$ foreign destination countries in 2006. Leveraging this dataset, we obtain the $r$-dimensional estimating function ${\mathbf g}({\mathbf x}_{i}; \boldsymbol{\theta}) =\{{g}_{1} ({\mathbf x}_{i}; \boldsymbol{\theta}),\ldots,{g}_{r} ({\mathbf x}_{i}; \boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} } $, $i\in[n]$, with ${\mathbf x}_{i} = (x_{i,1},\ldots,x_{i,r})^{\mathrm{\scriptscriptstyle \top} }$ and $g_j({\mathbf x}_{i}; \boldsymbol{\theta}) = x_{i,j} - \mu_j(\boldsymbol{\theta})$ for any $j\in[r]$ and $\boldsymbol{\theta} \in \boldsymbol{\Theta}$, where $x_{i,j}$ is the sale of firm $i$ in country $j$ from this dataset ($j=0$ is not considered in this dataset).

Since the model is highly nonlinear with respect to $\boldsymbol{\theta} \in \boldsymbol{\Theta}$, resulting in no closed-form expression for $\mu_j(\boldsymbol{\theta})$, we approximate it via numerical simulation Eaton2011, Shi2016. Specifically, in the estimation, we utilize the “artificial data" for another $5n = 33770$ firms from the dataset. This involves simulating the entry decisions and sales across various countries for each of these artificial firms. Subsequently, we calculate sample means to approximate $\mu_j(\boldsymbol{\theta})$ for any $j\in \{0\}\cup[r]$ and $\boldsymbol{\theta} \in \boldsymbol{\Theta}$. We generated samples of size 3500 from the posterior distribution for the BPEL. To select the tuning parameter $\nu$, we employed the BIC as defined in (ref). For the parameter space $\boldsymbol{\Theta}$, we adopted a compact range of values, specifically $\boldsymbol{\Theta} = [1.5, 10] \times [0.5, 5] \times [0.1, 5] \times [0.1, 5] \times [-0.9, 0.9]$, which is consistent with the economic context and aligns with the study of Shi2016. To initiate the analysis, we selected 15 samples uniformly distributed within the parameter space $\boldsymbol{\Theta}$. Figure (ref) presents the box-plots of the corresponding 15 estimates obtained by M-H and MAMIS from these initial values. The results for the REL with the same initial values are also included for comparative evaluation.

figure[figure omitted — 155 chars of source]

It is evident that for all five parameters, MAMIS exhibits the smallest variations in the resulting estimates, whereas the variations of M-H and REL are relatively similar. This consistency with the findings in Sections (ref) and (ref) reaffirms the robustness of MAMIS when considering different initial points. Such robustness is desirable for conducting more in-depth analyses. For instance, let us take $\theta_5$ into consideration which represents the correlation coefficient between the demand shock and the entry cost shock. The sign of its estimate carries the key implication. The 15 estimates of $\theta_5$ obtained by REL and M-H, from different initial values, fall within the ranges of $(-0.8738, 0.8507)$ and $(-0.8774, 0.8996)$, respectively. In contrast, the estimates of $\theta_5$ by MAMIS range in $(-0.7978, 0.1329)$, with the majority being negative, signaling a more assuring result.

We then proceed to examine the specific moments selected by the respective methods. For REL, we employ the greedy algorithm outlined in Section 3.2 of Shi2016. To assess the effectiveness of moment selection, we validate whether or not the top 10 trading partners of China in terms of export volume in this dataset, including the USA, Japan, Germany, etc., are either selected or partially selected. We find that, although REL selects at least some of these countries for 10 out of the 15 initial values, the number of selected countries does not exceed 3. In contrast, for 13 out of the 15 initial values, M-H identifies at least some of these countries, with 9 of them including more than 3. In the case of MAMIS, 13 out of the 15 initial values result in the identification of some of these countries, and all of them include more than 3 countries. Additionally, the robustness of MAMIS with respect to the initial points provides enhanced reliability in this context.

Theoretical Analysis

We introduce some additional notation first. For simplicity, write $\mathbb{E}_n(\cdot)=n^{-1}\sum_{i=1}^{n}\cdot$. For a $q \times q$ symmetric matrix ${\mathbf A}$, denote by $\lambda_{\min}({\mathbf A})$ and $\lambda_{\max}({\mathbf A})$ the smallest and largest eigenvalues of ${\mathbf A}$, respectively. For a $q_1 \times q_2$ matrix ${\mathbf B}=(b_{i,j})_{q_1 \times q_2}$, let $|{\mathbf B}|_{\infty} = \max_{i \in [q_1], j \in [q_2]} |b_{i,j}|$ be the super-norm. For the $r$-dimensional estimating function ${\mathbf g}(\cdot\,; \cdot)=\{g_1(\cdot\,;\cdot),\ldots,g_r(\cdot\,;\cdot)\}^{\mathrm{\scriptscriptstyle \top} }$ and $p$-dimensional parameter $\boldsymbol{\theta}=(\theta_1,\ldots,\theta_p)^{\mathrm{\scriptscriptstyle \top} }$, let $\nabla_{\boldsymbol{\theta}} {\mathbf g} (\cdot\,; \boldsymbol{\theta})=\{\partial g_j(\cdot\,;\boldsymbol{\theta})/\partial \theta_k\}_{j\in[r],k\in[p]}$, an $r\times p$ matrix, be the first-order partial derivative of ${\mathbf g} (\cdot\,; \boldsymbol{\theta})$ with respect to $\boldsymbol{\theta}$. Let ${\mathbf V} (\boldsymbol{\theta})= {\mathbb{E}} \{ {\mathbf g} ({\mathbf x}_{i}; \boldsymbol{\theta})^{\otimes2} \}$ and $\boldsymbol{\Gamma}(\boldsymbol{\theta}) ={\mathbb{E}}\{\nabla_{\boldsymbol{\theta}} {\mathbf g}({\mathbf x}_{i}; \boldsymbol{\theta}) \}$ for any $\boldsymbol{\theta}\in\boldsymbol{\Theta}$. For a given index set $\mathcal{F}$, let $|\mathcal{F}|$ be its cardinality. Denote by ${\mathbf g}_{\mathcal{F}}(\cdot\,; \cdot)$ the subvector of ${\mathbf g}(\cdot\,; \cdot)$ collecting the components indexed by $\mathcal{F}$. Let ${\mathbf V}_{\mathcal{F}} (\boldsymbol{\theta})= {\mathbb{E}} \{ {\mathbf g}_{\mathcal{F}} ( {\mathbf x}_{i}; \boldsymbol{\theta})^{\otimes2} \}$ and $\boldsymbol{\Gamma}_{\mathcal{F}}(\boldsymbol{\theta}) ={\mathbb{E}} \{\nabla_{\boldsymbol{\theta}} {\mathbf g}_{\mathcal{F}}({\mathbf x}_{i}; \boldsymbol{\theta}) \}$. Analogously, we also write ${\mathbf a}_{\mathcal{F}}$ as the corresponding subvector of vector ${\mathbf a}$. For any two probability measures $\mu$ and $\nu$, denote by $\mathcal{D}_{\mathrm{\scriptscriptstyle TV} }(\mu, \nu)$ the total variation distance between $\mu$ and $\nu$.

Properties of the Penalized Empirical Likelihood Estimator

To investigate the asymptotic properties of $\hat{\boldsymbol{\theta}}_n$ in (ref), we assume some regularity conditions.

conditionFor any $\varepsilon>0$, it holds that $$ \inf_{\boldsymbol{\theta} \in \boldsymbol{\Theta}:\, |\boldsymbol{\theta}-\boldsymbol{\theta}_0|_\infty>\varepsilon} |{\mathbb{E}}\{{\mathbf g} ( {\mathbf x}_{i}; \boldsymbol{\theta})\}|_{\infty} \geq \Delta(\varepsilon)\,,$$ where $\Delta(\cdot)$ is a nonnegative function satisfying $\lim \inf _{\varepsilon \rightarrow 0^{+}} \varepsilon^{-1} \Delta(\varepsilon) \geq K_{1}$ for some universal constant $K_{1}>0$.
condition{\rm(a)} There exist universal constants $K_2>0$ and $\gamma>4$ such that $$ \max_{j\in[r]}{\mathbb{E}}\bigg\{ \sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} |g_j ( {\mathbf x}_i; \boldsymbol{\theta})|^{\gamma}\bigg\} \leq K_2$$ and $\sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}}\max_{j\in[r]} \mathbb{E}_n \{|g_j ( {\mathbf x}_{i}; \boldsymbol{\theta})|^{\gamma}\} = O_{\rm p}(1)$. {\rm(b)} There exist universal constants $0<K_3<K_4$ such that $K_3<\lambda_{\min}\{{\mathbf V} (\boldsymbol{\theta}_0)\} \leq \lambda_{\max}\{{\mathbf V} (\boldsymbol{\theta}_0)\} < K_4$. {\rm(c)} For any ${\mathbf x}$ and $j\in[r]$, $g_j({\mathbf x};\boldsymbol{\theta})$ is twice continuously differentiable with respect to $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ satisfying $$ \sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} \max_{j \in [r], k \in [p]} \mathbb{E}_n \bigg\{\bigg|\frac{\partial g_j({\mathbf x}_{i}; \boldsymbol{\theta})}{\partial \theta_{k}}\bigg|^2\bigg\}=O_{\rm p}(1)= \sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} \max_{j \in [r], k_1, k_2 \in [p]} \mathbb{E}_n \bigg\{\bigg|\frac{\partial^2 g_j({\mathbf x}_{i}; \boldsymbol{\theta})}{\partial \theta_{k_1} \partial \theta_{k_2}}\bigg|^2\bigg\}\,.$$

Detailed discussion on Conditions (ref) and (ref) are given in Section (ref) of the supplementary material. For any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$, define $$ \mathcal{M}_{\boldsymbol{\theta}}^*=\{j \in [r]: |\mathbb{E}_n\{ g_j ({\mathbf x}_{i}; \boldsymbol{\theta})\} | \geq C_*\nu \rho'(0^{+}) \} $$ for some $C_* \in (0,1)$. We assume the existence of a sequence $\ell_n\rightarrow\infty$ such that $$\mathbb{P}\bigg(\sup _{\boldsymbol{\theta} \in \boldsymbol{\Theta}:\, |\boldsymbol{\theta}-\boldsymbol{\theta}_0|_2 \leq c_n}|\mathcal{M}_{\boldsymbol{\theta}}^*|\leq \ell_n\bigg)\rightarrow1$$ as $n\rightarrow\infty$, with some $c_n\rightarrow 0$ satisfying $\nu c_n^{-1} \rightarrow 0$. Proposition (ref) shows that $\hat{\boldsymbol{\theta}}_n$ is consistent to the true parameter $\boldsymbol{\theta}_0$, allowing $r$ growing exponentially with the sample size $n$.

propositionLet $P_{\nu}(\cdot) \in \mathscr{P}$ be a convex function for $\mathscr{P}$ defined as {\rm (ref)}. Under Conditions {\rm (ref), (ref)(a)} and {\rm(ref)(b)}, if $\log r\ll n^{1/3}$ and $\ell_nn^{-1/2}(\log r)^{1/2}\ll\min \{\nu,n^{-1/\gamma}\}$, then the PEL estimator $\hat{\boldsymbol{\theta}}_n$ defined as {\rm (ref)} satisfies $|\hat{\boldsymbol{\theta}}_n-\boldsymbol{\theta}_0|_{\infty}=O_{\rm p}(\nu)$.

Proposition (ref) establishes the consistency of the PEL estimator with diverging \(r\), incorporating the impact of the penalty function. In particular, the convergence rate of \(\hat{\boldsymbol{\theta}}_n\) is \(\nu\), provided that the tuning parameter \(\nu\) in (ref) satisfies \(\nu \gg \ell_n n^{-1/2} (\log r)^{1/2}\). As a result, the convergence rate of \(\hat{\boldsymbol{\theta}}_n\) is slower than \(n^{-1/2}\), which can be viewed as the price paid for using the penalty in handling exponentially growing dimensionality \(r\).

Recall $\rho(t;\nu)=\nu^{-1}P_\nu(t)$. For $P_\nu(\cdot)\in\mathscr{P}$ with $\mathscr{P}$ defined as (ref), since $\rho'(0^{+};\nu)$ is independent of $\nu$, we write it as $\rho'(0^{+})$ for simplicity. Let $\mathcal{R}_n={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$ for the Lagrange multiplier $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)=(\hat\lambda_1,\ldots,\hat\lambda_r)^{\mathrm{\scriptscriptstyle \top} }=\arg\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_n(\hat{\boldsymbol{\theta}}_n)}f_n(\boldsymbol{\lambda};\hat{\boldsymbol{\theta}}_n)$ with $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ defined as (ref). Then $\hat{\boldsymbol{\theta}}_n$ and $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)$ satisfy the score equation:

align[align omitted — 319 chars of source]

where $\hat{\boldsymbol{\eta}}=(\hat{\eta}_{1},\ldots ,\hat{\eta}_{r})^{\mathrm{\scriptscriptstyle \top} }$ with $\hat{\eta}_j=\nu \rho' (|\hat{\lambda}_j|;\nu) {\mbox{\rm sgn}}(\hat{\lambda}_j)$ for $\hat{\lambda}_j\neq0$ and $\hat{\eta}_j \in [-\nu \rho'(0^{+}),\nu \rho'(0^{+})]$ for $\hat{\lambda}_j=0$. Here, an effective drastic dimension reduction is achieved with the associated sparse \(\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\). The use of the penalty function \(P_\nu(\cdot)\) leads to \(\hat{\boldsymbol{\eta}}\) in (ref), an extra term compared to that of the conventional EL. While \(P_\nu(\cdot)\) ensures the consistency of \(\hat{\boldsymbol{\theta}}_n\) as shown in Proposition (ref), as we will show in Theorem (ref) later, \(\hat{\boldsymbol{\eta}}\) leads to a bias of the PEL estimator \(\hat{\boldsymbol{\theta}}_n\).

We further remark that while penalizing the Lagrange multiplier in our PEL does effectively achieve the selection of moments, its properties in terms of the validity of the selected moments remain an interesting research question. On one hand, it is reasonable to expect that under appropriate conditions and with a suitably chosen tuning parameter, our PEL may correctly select the set of valid moments. On the other hand, the major challenge lies in the ambiguity of defining valid moments when the corresponding moment functions are evaluated at broad candidate values of the model parameters rather than the truth. This consideration opens the door to a research question of its own interest in the context of moment selection that we are interested in investigating in our future research.

To study the asymptotic distribution of $\hat{\boldsymbol{\theta}}_n$, we need the following regularity conditions.

conditionLet ${\mathbf Q}_{\mathcal{F}}={\boldsymbol{\Gamma}_{\mathcal{F}}(\boldsymbol{\theta}_0)}^{{\mathrm{\scriptscriptstyle \top} }, \otimes2}$ for any $\mathcal{F}\subset[r]$. There exist universal constants $0<K_5<K_6$ such that $K_5<\lambda_{\min}({\mathbf Q}_{\mathcal{F}}) \leq \lambda_{\max}({\mathbf Q}_{\mathcal{F}}) < K_6$ for any $\mathcal{F} $ with $p\leq |\mathcal{F}| \leq \ell_n$.
condition{\rm(a)} For the PEL estimator $\hat{\boldsymbol{\theta}}_n$ defined as (ref), there exists a constant $\tilde{c} \in (C_*,1)$ such that $$ {\mathbb{P}}\bigg[\bigcup_{j \in [r]}\{\tilde{c} \nu\rho'(0^+)\leq |\mathbb{E}_n\{g_j ({\mathbf x}_{i}; \hat{\boldsymbol{\theta}}_n) \}| <\nu\rho'(0^+)\}\bigg] \rightarrow 0$$ as $n \rightarrow \infty$. {\rm(b)} It holds that $$ {\mathbb{P}}\bigg[\bigcup_{j \in \mathcal{R}_n^{\rm c}} \{|\hat{\eta}_j|=\nu \rho'(0^+)\}\bigg] \rightarrow 0 $$ as $n \rightarrow \infty$.

Discussion of Conditions (ref) and (ref) are given in Section (ref) of the supplementary material. Write $\widehat{{\mathbf V}}_{{ \mathcal{R}_n}}(\hat{\boldsymbol{\theta}}_n)=\mathbb{E}_n\{{\mathbf g}_{{ \mathcal{R}_n}} ( {\mathbf x}_{i}; \hat{\boldsymbol{\theta}}_n)^{\otimes2}\} $ and $\widehat{\boldsymbol{\Gamma}}_{{ \mathcal{R}_n}}(\hat{\boldsymbol{\theta}}_n) =\mathbb{E}_n\{ \nabla_{\boldsymbol{\theta}} {\mathbf g}_{{ \mathcal{R}_n}}({\mathbf x}_{i}; \hat{\boldsymbol{\theta}}_n)\}$. Define

align[align omitted — 608 chars of source]

where $\hat{\boldsymbol{\eta}}=(\hat{\eta}_{1},\ldots ,\hat{\eta}_{r})^{\mathrm{\scriptscriptstyle \top} }$ is specified in (ref). We assume $(r,\ell_n,\nu)$ satisfy the following restrictions:

align[align omitted — 290 chars of source]

The asymptotic distribution of $\hat{\boldsymbol{\theta}}_n$ is stated in Theorem (ref), where the bias term $\hat{\boldsymbol{\psi}}_{ \mathcal{R}_n}$ comes from the penalty function $P_{\nu}(\cdot)$ imposed on the Lagrange multiplier $\boldsymbol{\lambda}$ in (ref).

theoremLet $P_{\nu}(\cdot) \in \mathscr{P}$ be convex with bounded second-order derivative around $0$, where $\mathscr{P}$ is defined as (ref). Assume Conditions {\rm (ref)--(ref)} hold with $(r,\ell_n,\nu)$ satisfying (ref). For any ${\mathbf t} \in {\mathbb{R}}^{p}$ with $|{\mathbf t}|_2=1$, the PEL estimator $\hat{\boldsymbol{\theta}}_n$ defined as (ref) satisfies $ n^{1/2} {\mathbf t}^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf H}}_{ \mathcal{R}_n}^{1/2}(\hat{\boldsymbol{\theta}}_n-\boldsymbol{\theta}_0-\hat{\boldsymbol{\psi}}_{ \mathcal{R}_n}) \rightarrow \mathcal{N}(0,1) $ in distribution as $n\rightarrow\infty$, where $\widehat{{\mathbf H}}_{ \mathcal{R}_n}$ and $\hat{\boldsymbol{\psi}}_{ \mathcal{R}_n}$ are defined in (ref).

Here, the estimated bias \(\hat{\boldsymbol{\psi}}_{ \mathcal{R}_n}\) can be easily calculated based on (ref). Theorem (ref) indicates that, upon correcting the bias by subtracting it from $\hat{\boldsymbol{\theta}}_n$, the resulting estimator $\hat{\boldsymbol{\theta}}_n-\hat{\boldsymbol{\psi}}_{ \mathcal{R}_n}$ is \(n^{1/2}\)-consistent and asymptotically normal.

Properties of the Posterior Distribution and Algorithms

For the proposed BPEL, we establish the Bernstein-von Mises theorem for the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$, as defined in (ref). Furthermore, we provide theoretical assurances for the performance of Algorithms (ref) and (ref) in Section (ref).

For any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$, write $\mathcal{R}(\boldsymbol{\theta})={\rm supp}\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})\}$ with $$\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})=\{\hat{\lambda}_1(\boldsymbol{\theta}),\ldots,\hat{\lambda}_r(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }=\arg\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_n(\boldsymbol{\theta})}f_n(\boldsymbol{\lambda};\boldsymbol{\theta})\,,$$ where $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ is defined as (ref). Then $\boldsymbol{\theta}$ and $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})$ satisfy the score equation:

align[align omitted — 316 chars of source]

where $\hat{\boldsymbol{\eta}}(\boldsymbol{\theta})=\{\hat{\eta}_{1}(\boldsymbol{\theta}),\ldots ,\hat{\eta}_{r}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$ with $\hat{\eta}_j(\boldsymbol{\theta})=\nu \rho' \{|\hat{\lambda}_j(\boldsymbol{\theta})|;\nu\} {\mbox{\rm sgn}}\{\hat{\lambda}_j(\boldsymbol{\theta})\}$ for $\hat{\lambda}_j(\boldsymbol{\theta}) \neq 0$ and $\hat{\eta}_j(\boldsymbol{\theta}) \in [-\nu \rho'(0^{+}),\nu \rho'(0^{+})]$ for $\hat{\lambda}_j(\boldsymbol{\theta})=0$. By the definition of the PEL estimator $\hat{\boldsymbol{\theta}}_n$, we have $f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\} \geq f_n\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n);\hat{\boldsymbol{\theta}}_n\}$ for any $\boldsymbol{\theta}\in\boldsymbol{\Theta}$. To investigate the asymptotic properties of the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ defined as (ref), we need to first study the asymptotic behavior of $ \aleph_n(\boldsymbol{\theta})=f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}-f_n\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n);\hat{\boldsymbol{\theta}}_n\}$ for $\boldsymbol{\theta}\in\boldsymbol{\Theta}$. Given $\alpha_n=n^{-1/2}(\log r)^{1/2}$ and some $\beta_n$ satisfying $\ell_n^{1/2} \nu \ll \beta_{n} \ll \min\{ \ell_n^{-1}n^{-1/\gamma}, \nu^{2/3}\ell_n^{-2/3} n^{-1/(3\gamma)}\}$, we split the whole parameter space $\boldsymbol{\Theta}$ into three regions: $\mathcal{C}_1=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2\leq \alpha_n\}$, $\mathcal{C}_2=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:\alpha_n<|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2\leq \beta_n\}$ and $\mathcal{C}_3=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2>\beta_n\}$. Proposition (ref) in the supplementary material shows that the asymptotic behavior of $\aleph_n(\boldsymbol{\theta})$ for $\boldsymbol{\theta}$ in these three regions are different.

Investigating the asymptotic behavior of $\aleph_n(\boldsymbol{\theta})$ calls some new technical arguments. Write

equation[equation omitted — 425 chars of source]

When $r$ is a fixed constant, we know $2n\tilde{f}_n\{\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$ is the conventional log-EL ratio in the literature. The asymptotic behavior of $2n\tilde{f}_n\{\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$ depends on the magnitude of $\mathbb{E}\{{\mathbf g}({\mathbf x}_i;\boldsymbol{\theta})\}$. More specifically, under some mild conditions, it holds that (i) $2n\tilde{f}_n\{\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$ is asymptotically chi-square distributed with degree of freedom $r$ if $|\mathbb{E}\{{\mathbf g}({\mathbf x}_i;\boldsymbol{\theta})\}|_2\ll n^{-1/2}$, (ii) $2n\tilde{f}_n\{\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$ converges to a noncentral chi-square distribution if $|\mathbb{E}\{{\mathbf g}({\mathbf x}_i;\boldsymbol{\theta})\}|_2\asymp n^{-1/2}$, and (iii) $2n\tilde{f}_n\{\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$ diverges to $\infty$ in probability if $|\mathbb{E}\{{\mathbf g}({\mathbf x}_i;\boldsymbol{\theta})\}|_2\gg n^{-1/2}$. See, for example, Proposition 1 and Theorem 1 of Changetal_2013_AOS for such results with $r=1$. In comparison to $\tilde{f}_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ defined in (ref), $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ involved in $\aleph_n(\boldsymbol{\theta})$ includes a penalty term imposed on the Lagrange multiplier $\boldsymbol{\lambda}$. This makes the standard technique for analyzing the conventional log-EL ratio inapplicable. To further establish the Bernstein-von Mises theorem for the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ defined as (ref), we assume the following regularity conditions.

condition{\rm (a)} There exists a constant $\bar{c} \in (0,1)$ such that $$ {\mathbb{P}}\bigg\{\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}\max_{j \in \mathcal{R}(\boldsymbol{\theta})^{\rm c}} |\hat{\eta}_j(\boldsymbol{\theta})|\leq \bar{c} \nu \rho'(0^+)\bigg\} \rightarrow 1$$ as $n \rightarrow \infty$, where $\hat \eta_j(\boldsymbol{\theta})$ is specified in (ref). {\rm (b)} There exists $\kappa_n$ satisfying $\max\{\ell_n^{1/2}n^{-1/2}(\log r)^{1/2},$ $ \ell_n \beta_{n}^{3/2} n^{1/(2\gamma)}\} \ll \kappa_n \ll \nu$ such that $$ {\mathbb{P}}\bigg[\bigcup_{\boldsymbol{\theta} \in \mathcal{C}_2}\bigcup_{j \in \mathcal{R}_n} \{\nu\rho'(0^+)-\kappa_n < |\mathbb{E}_n\{g_j ({\mathbf x}_{i}; \boldsymbol{\theta})\}| < \nu \rho'(0^+)+\kappa_n\}\bigg] \rightarrow 0$$ as $n \rightarrow \infty$. {\rm (c)} There exist universal constants $K_7, K_8>0$ such that $$ {\mathbb{P}}\bigg\{\inf_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}\lambda_{\min}([\mathbb{E}_n\{\nabla_{\boldsymbol{\theta}}{\mathbf g}_{\mathcal{R}_n}({\mathbf x}_i;\boldsymbol{\theta})\}]^{{\mathrm{\scriptscriptstyle \top} },\otimes2})\geq K_7\bigg\} \rightarrow 1~~\textrm{and}~~{\mathbb{P}}\bigg[\sup_{\boldsymbol{\theta} \in \mathcal{C}_3}\lambda_{\max}\{\widehat{\mathbf V}_{\mathcal{R}_n}(\boldsymbol{\theta})\} \leq K_8\bigg] \rightarrow 1$$ as $n \rightarrow \infty$.
conditionThe prior density $\pi_{0}(\cdot)$ is continuously differentiable with bounded first-order derivatives and $\pi_{0}(\boldsymbol{\theta}_0)>0$.

Discussion of Conditions (ref) and (ref) are given in Section (ref) of the supplementary material. Let $\Pi^{\dag}_n(\cdot)$ be the measure which admits the posterior distribution $\pi^{\dag}(\cdot\,|\,\mathcal{X}_n)$. Denote by $\mathcal{N}_{\boldsymbol{\mu},\boldsymbol{\Sigma}}(\cdot)$ the Gaussian measure with mean $\boldsymbol{\mu}$ and covariance matrix $\boldsymbol{\Sigma}$. To establish the Bernstein-von Mises theorem for the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ as in Theorem (ref), we need to assume $(r,\ell_n,\nu)$ satisfy the following restrictions:

align[align omitted — 343 chars of source]
theoremLet $P_{\nu}(\cdot) \in \mathscr{P}$ be convex and assume $\rho(t;\nu)=\nu^{-1}P_{\nu}(t)$ has bounded second-order derivative with respect to $t$ around $0$, where $\mathscr{P}$ is defined in {\rm (ref)}. Assume Conditions {\rm(ref)--(ref)} hold with $(r,\ell_n,\nu)$ satisfying (ref). The posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ converges in total variation toward a Gaussian distribution $\mathcal{N}(\hat{\boldsymbol{\theta}}_n,n^{-1}\widehat{{\mathbf H}}_{\mathcal{R}_n}^{-1})$ in probability, that is, $\mathcal{D}_{\mathrm{\scriptscriptstyle TV} }(\Pi^{\dag}_n ,\, \mathcal{N}_{\hat{\boldsymbol{\theta}}_n,n^{-1}\widehat{{\mathbf H}}_{ \mathcal{R}_n}^{-1}}) \rightarrow 0$ in probability as $n\rightarrow \infty$, where $\hat{\boldsymbol{\theta}}_n$ is the PEL estimator in (ref), and $\widehat{{\mathbf H}}_{ \mathcal{R}_n}$ is defined in (ref).

Theorem (ref) shows that $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ has a Gaussian limiting distribution and it concentrates on a $n^{-1/2}$-ball centered at the PEL estimator $\hat{\boldsymbol{\theta}}_n$ of interest, which indicates that $\hat{\boldsymbol{\theta}}_n$ can be approximated by the mean of the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$. More specifically, as shown in Corollary (ref), the approximation error is of order smaller than $n^{-1/2}$.

corollaryUnder the conditions of Theorem {\rm (ref)}, we have $|{\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta}) - \hat{\boldsymbol{\theta}}_n|_\infty = o_{\rm p}(n^{-1/2})$, where $\hat{\boldsymbol{\theta}}_n$ is the PEL estimator defined as {\rm (ref)}, and ${\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta})$ is defined in (ref).

Theorems (ref) and (ref) state the theoretical guarantees for Algorithms (ref) and (ref), respectively.

theoremFor the density $\phi(\cdot\,|\,\cdot)$ of the proposal distribution in Algorithm {\rm (ref)}, we assume $\phi(\boldsymbol{\vartheta}\,|\,\boldsymbol{\theta})$ is positive and continuous on $(\boldsymbol{\theta}, \boldsymbol{\vartheta}) \in \boldsymbol{\Theta} \times \boldsymbol{\Theta}$. Conditional on $\mathcal{X}_n$, for any $\boldsymbol{\theta}^0 \in \boldsymbol{\Theta}$ such that $\pi^\dag(\boldsymbol{\theta}^0\,|\,\mathcal{X}_n)>0$ with $\pi^\dag(\cdot\,|\,\mathcal{X}_n)$ defined as (ref), it holds that $\mathcal{D}_{\mathrm{\scriptscriptstyle TV} } (\mathcal{T}^k_{\boldsymbol{\theta}^{0}},\, \Pi^{\dag}_n) \rightarrow 0$ as $k \rightarrow \infty$, where $\mathcal{T}^k_{\boldsymbol{\theta}^{0}}(\cdot)$ is the measure which admits the distribution of the Markov chain determined by Algorithm {\rm (ref)} at $k$-th step with initial point $\boldsymbol{\theta}^{0}$. Furthermore, conditional on $\mathcal{X}_n$, $|K^{-1}\sum_{k=1}^K \boldsymbol{\theta}^k - {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta})|_\infty \rightarrow 0$ almost surely as $K \rightarrow \infty$, where $\{\boldsymbol{\theta}^k\}_{k\geq1}$ are generated via Algorithm {\rm(ref)} with the initial point $\boldsymbol{\theta}^{0}$ satisfying $\pi^\dag(\boldsymbol{\theta}^0\,|\,\mathcal{X}_n)>0$.
theoremFor the density $\varphi(\cdot\,; \cdot)$ of the proposal distribution and the function ${\mathbf h}: {\mathbb{R}}^p \mapsto {\mathbb{R}}^s$ in Algorithm {\rm (ref)}, we assume $\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})$ is positive and continuous on $(\boldsymbol{\theta}, \boldsymbol{\zeta}) \in \boldsymbol{\Theta} \times {\mathbb{R}}^s$ and $\sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} |{\mathbf h}(\boldsymbol{\theta})|_\infty \leq K_9$ for some universal constant $K_9>0$. Conditional on $\mathcal{X}_n$, if $\sum_{k=1}^{\infty}\exp(-CN_k) < \infty$ for any $C>0$, then $|\widehat{{\mathbb{E}}}_{\pi^\dag,K}(\boldsymbol{\theta}) - {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta})|_\infty \rightarrow 0$ almost surely as $K \rightarrow \infty$, where $\widehat{{\mathbb{E}}}_{\pi^\dag,K}(\boldsymbol{\theta})$ is the MAMIS estimator defined as (ref).

Discussion

In this paper, we explore BPEL and demonstrate its promising performance using MCMC sampling as a competitive alternative to optimization in addressing EL problems. This framework has the potential for further advancements in several areas. To maintain focus and avoid digressions, we have confined our study to fixed-dimensional model parameters and exponentially growing moment conditions. However, there is significant interest in extending this approach to tackle variable and model selection using BPEL, which could accommodate high-dimensional sparse model parameters and potentially a continuum of moment conditions, as considered in Chau2017. Incorporating specific priors in the context of concrete studies, particularly in high-dimensional problems, is another area of interest. Research in this direction presents additional challenges, especially in selecting appropriate priors, developing efficient sampling schemes, and conducting associated analyses.

In the broader context of Bayesian methodology, approximate Bayesian computation (ABC) and Bayesian synthetic likelihood (BSL) are two competitive methods for handling situations where the likelihood is difficult to evaluate or intractable. ABC and BSL have been extensively compared in the literature. We demonstrate that the rationale of ABC integrates well with our BPEL method, achieving both accuracy and computational efficiency. Our Algorithm 2, inspired by ABC, uses importance weights for samples drawn from an alternative distribution to address challenging sampling situations. Empirical evidence shows promising performance, particularly in difficult cases. BSL leverages the limiting distribution, such as the normal distribution, to handle intractable probability distributions, with the advantage of easy sampling from the normal distribution. We view our BPEL as a compelling alternative to BSL: EL uses a multinomial likelihood that incorporates model information without requiring a fully specified parametric model, making it a competitive option when the full likelihood is intractable.

Furthermore, we foresee the use of more sophisticated sampling schemes in conjunction with PEL as highly valuable for addressing complex problems with specific considerations. Examples include the Hamiltonian MCMC method examined in Chaudhuri2017 and the variational Bayesian approach explored in Yu2023. These avenues of research are part of our plans for future projects.

thebibliography\baselineskip 16pt \bibitem[Bissiri et al., 2016]{Bissiri_2016} Bissiri, P. G., Holmes, C. C., & Walker, S. G. (2016). A general framework for updating belief distributions. {\em J. Roy. Statist. Soc. Ser. B}, 78, 1103--1130. \bibitem[Brooks et al., 2011]{Brooks2011} Brooks, S., Gelman, A., Jones, G., & Meng, X.-L. (2011). \newblock {\em Handbook of Markov Chain Monte Carlo}. \newblock New York, Chapman and Hall/CRC. \bibitem[Castillo et al., 2015]{Castillo_2015} Castillo, I., Schmidt-Hieber, J., & van der Vaart, A. (2015). Bayesian linear regression with sparse priors. {\em Ann. Statist.}, 43, 1986--2018. \bibitem[Chang et al., 2015]{ChangChenChen_2015} Chang, J., Chen, S. X., & Chen, X. (2015). \newblock High dimensional generalized empirical likelihood for moment restrictions with dependent data. \newblock {\em J. Econometrics}, 185, 283--304. \bibitem[Chang et al., 2021]{Chang_2021} Chang, J., Chen, S. X., Tang, C. Y., & Wu, T. T. (2021). \newblock High-dimensional empirical likelihood inference. \newblock {\em Biometrika}, 108, 127--147. \bibitem[Chang et al., 2023]{Chang2023} Chang, J., Shi, Z., and Zhang, J. (2023). \newblock Culling the herd of moments with penalized empirical likelihood. \newblock {\em J. Bus. Econ. Stat.}, 41, 791--805. \bibitem[Chang et al., 2018]{Changetal_2018} Chang, J., Tang, C. Y., & Wu, T. (2018). \newblock A new scope of penalized empirical likelihood with high-dimensional estimating equations. \newblock {\em Ann. Statist.}, 46, 3185--3216. \bibitem[Chang et al., 2013]{Changetal_2013_AOS} Chang, J., Tang, C. Y., & Wu, Y. (2013). \newblock Marginal empirical likelihood and sure independence feature screening. \newblock {\em Ann. Statist.}, 41, 2123--2148. \bibitem[Chaudhuri and Ghosh, 2011]{Chaudhuri2011} Chaudhuri, S., & Ghosh, M. (2011). \newblock Empirical likelihood for small area estimation. \newblock {\em Biometrika}, 98, 473--480. \bibitem[Chaudhuri et al., 2017]{Chaudhuri2017} Chaudhuri, S., Mondal, D., & Yin, T. (2017). \newblock Hamiltonian Monte Carlo sampling in Bayesian empirical likelihood computation. \newblock {\em J. Roy. Statist. Soc. Ser. B}, 79, 293--320. \bibitem[Chauss\'{e}, 2017]{Chau2017} Chauss\'{e}, P. (2017). \newblock Generalized empirical likelihood for a continuum of moment conditions. \newblock {\em Manuscript}. \bibitem[Chen et al., 2009]{ChenPengQin_2008} Chen, S. X., Peng, L., & Qin, Y. L. (2009). \newblock Effects of data dimension on empirical likelihood. \newblock {\em Biometrika}, 96, 711--722. \bibitem[Cheng and Zhao, 2019]{Cheng2019} Cheng, Y., & Zhao, Y. (2019). \newblock Bayesian jackknife empirical likelihood. \newblock {\em Biometrika}, 106, 981--988. \bibitem[Chib et al., 2018]{Chib2018} Chib, S., Shin, M., & Simoni, A. (2018). \newblock Bayesian estimation and comparison of moment condition models. \newblock {\em J. Amer. Statist. Assoc.}, 113, 1656--1668. \bibitem[Cornuet et al., 2012]{CORNUET2012} Cornuet, J.-M., Marin, J.-M., Mira, A., & Robert, C. P. (2012). \newblock Adaptive multiple importance sampling. \newblock {\em Scand. J. Stat.}, 39, 798--812. \bibitem[Donald et al., 2003]{Donald2003} Donald, S. G., Imbens, G. W., & Newey, W. K. (2003). \newblock Empirical likelihood estimation and consistent tests with conditional moment restrictions. \newblock {\em J. Econometrics}, 117, 55--93. \bibitem[Eaton et al.(2011)]{Eaton2011} Eaton, J., Kortum, S., & Kramarz, F. (2011). An anatomy of international trade: evidence from French firms. {\em Econometrica}, 79, 1453--1498. \bibitem[Frazier et al., 2023]{Frazier2023} Frazier, D. T., Drovandi, C., & Kohn, R. (2023). Calibrated generalized Bayesian inference. {\em arXiv:2311.15485}. \bibitem[Gelman et al., 1997]{Gelman1997} Gelman, A., Gilks, W. R., & Roberts, G. O. (1997). \newblock Weak convergence and optimal scaling of random walk metropolis algorithms. \newblock {\em Ann. Appl. Probab.}, 7, 110--120. \bibitem[Godambe and Heyde, 1987]{GodambeHeyde_1987_ISR} Godambe, V. P., & Heyde, C. C. (1987). \newblock Quasi-likelihood and optimal estimation. \newblock {\em Int. Stat. Rev.}, 55, 231--244. \bibitem[Hesterberg, 1995]{Hesterberg1995} Hesterberg, T. (1995). \newblock Weighted average importance sampling and defensive mixture distributions. \newblock {\em Technometrics}, 37, 185--194. \bibitem[Hjort et al., 2009]{Hjortetal_2008_AS} Hjort, N. L., McKeague, I., & Van Keilegom, I. (2009). \newblock Extending the scope of empirical likelihood. \newblock {\em Ann. Statist.}, 37, 1079--1111. \bibitem[Jain and Kar, 2017]{Jain2017} Jain, P., & Kar, P. (2017). \newblock Non-convex optimization for machine learning. \newblock {\em Found. Trends. Mach. Le}, 10, 142--336. \bibitem[Lazar, 2003]{Lazar2003} Lazar, N. A. (2003). \newblock Bayesian empirical likelihood. \newblock {\em Biometrika}, 90, 319--326. \bibitem[Leng and Tang, 2012]{LengTang_2010} Leng, C., & Tang, C. Y. (2012). \newblock Penalized empirical likelihood and growing dimensional general estimating equations. \newblock {\em Biometrika}, 99, 703--716. \bibitem[Ma et al., 2019]{Ma2019} Ma, Y.-A., Chen, Y., Jin, C., Flammarion, N., & Jordan, M. I. (2019). \newblock Sampling can be faster than optimization. \newblock {\em Proc. Natl. Acad. Sci.}, 116, 20881--20885. \bibitem[Marin et al., 2019]{Marin2019} Marin, J.-M., Pudlo, P., & Sedki, M. (2019). \newblock Consistency of adaptive importance sampling and recycling schemes. \newblock {\em Bernoulli}, 25, 1977--1998. \bibitem[Mengersen et al., 2013]{Mengersen2013} Mengersen, K. L., Pudlo, P., & Robert, C. P. (2013). \newblock Bayesian computation via empirical likelihood. \newblock {\em Proc. Natl. Acad. Sci.}, 110, 1321--1326. \bibitem[Narisetty and He(2014)]{NarisettyHe2014} Narisetty, N., & He, X. (2014). Bayesian variable selection with shrinking and diffusing priors. {\em Ann. Statist.}, 42, 789--817. \bibitem[Ouyang and Bondell, 2023]{OuyangBondell_2023} Ouyang, J., & Bondell, H. (2023). Bayesian analysis of longitudinal data via empirical likelihood. {\em Comput. Statist. Data Anal.}, 187, 107785. \bibitem[Owen, 2001]{Owen_2001} Owen, A. B. (2001). \newblock {\em Empirical {L}ikelihood}. \newblock New York, Chapman and Hall/CRC. \bibitem[Qin and Lawless, 1994]{QinLawless_1994_AS} Qin, J., & Lawless, J. (1994). \newblock Empirical likelihood and general estimating equations. \newblock {\em Ann. Statist.}, 22, 300--325. \bibitem[Rao and Wu, 2010]{Rao2010} Rao, J. N. K., & Wu, C. (2010). \newblock Bayesian pseudo-empirical-likelihood intervals for complex surveys. \newblock {\em J. Roy. Statist. Soc. Ser. B}, 72, 533--544. \bibitem[Ripley, 2006]{Ripley2006} Ripley, B. D. (2006). \newblock {\em Stochastic Simulation}. \newblock New York, Wiley-Interscience. \bibitem[Roberts and Rosenthal, 2004]{Roberts2004} Roberts, G. O., & Rosenthal, J. S. (2004). \newblock General state space markov chains and {MCMC} algorithms. \newblock {\em Probab. Surveys}, 1, 20--71. \bibitem[Shi, 2016]{Shi2016} Shi, Z. (2016). \newblock Econometric estimation with high-dimensional moment equalities. \newblock {\em J. Econometrics}, 195, 104--119. \bibitem[Tang and Yang, 2022]{Tang2022} Tang, R., & Yang, Y. (2022). \newblock Bayesian inference for risk minimization via exponentially tilted empirical likelihood. \newblock {\em J. Roy. Statist. Soc. Ser. B}, 84, 1257--1286. \bibitem[Tang and Leng, 2010]{TangLeng_2010_Bioka} Tang, C. Y., & Leng, C. (2010). \newblock Penalized high dimensional empirical likelihood. \newblock {\em Biometrika}, 97, 905--920. \bibitem[Tsao, 2004]{Tsao_2004_AOS} Tsao, M. (2004). \newblock Bounds on coverage probabilities of the empirical likelihood ratio confidence regions. \newblock {\em Ann. Statist.}, 32, 1215--1221. \bibitem[Vexler et al., 2014]{Vexetal_2014_Bioka} Vexler, A., Tao, G., & Hutson, A. D. (2014). \newblock Posterior expectation based on empirical likelihoods. \newblock {\em Biometrika}, 101, 711--718. \bibitem[Yang and He, 2012]{Yang2012} Yang, Y., & He, X. (2012). \newblock Bayesian empirical likelihood for quantile regression. \newblock {\em Ann. Statist.}, 40, 1102--1131. \bibitem[Yu and Bondell, 2024]{Yu2023} Yu, W., & Bondell, H. D. (2024). \newblock Variational bayes for fast and accurate empirical likelihood inference. \newblock {\em J. Amer. Statist. Assoc.}, 119, 1089--1101. \bibitem[Zhao et al., 2020]{Zhao2020} Zhao, P., Ghosh, M., Rao, J. N. K., & Wu, C. (2020). \newblock Bayesian empirical likelihood inference with complex survey data. \newblock {\em J. Roy. Statist. Soc. Ser. B}, 82, 155--174.

\setcounter{equation}0 \setcounter{section}0 \setcounter{table}0 \setcounter{figure}0

\numberwithin{equation}{section} \setcounter{page}{1} \rhead{\bfseriesS\arabic{page}} \lhead{ SUPPLEMENTARY MATERIAL}

center[center omitted — 171 chars of source]

In the sequel, we use the abbreviations “w.p.a.1” and “w.r.t” to denote, respectively, “with probability approaching one” and “with respect to”. Let $C$, $\bar{C}$ and $\tilde{C}$ be generic positive finite constants that may be different in different uses. Let $\lfloor a \rfloor$ represent the largest integer not greater than $a \in {\mathbb{R}}$. For any positive integer $q$, we write $[q]=\{1,\ldots,q\}$. Denote by $I(\cdot)$ the indicator function. Let ${\rm tr}({\mathbf A})$ be the trace of a $q \times q$ matrix ${\mathbf A}=(a_{i,j})_{q \times q}$. For a $q \times q$ symmetric matrix ${\mathbf A}$, denote by $\lambda_{\min}({\mathbf A})$ and $\lambda_{\max}({\mathbf A})$ the smallest and largest eigenvalues of ${\mathbf A}$, respectively. For a $q_1 \times q_2$ matrix ${\mathbf B}=(b_{i,j})_{q_1 \times q_2}$, let $\|{\mathbf B}\|_2=\lambda^{1/2}_{\max}({\mathbf B}^{ \otimes2})$ be the spectral norm with ${\mathbf B}^{\otimes2}={\mathbf B} {\mathbf B}^{\mathrm{\scriptscriptstyle \top} }$. Specifically, if $q_2=1$, we use $|{\mathbf B}|_{\infty}=\max_{i \in [q_1]} |b_{i,1}|$, $|{\mathbf B}|_1=\sum_{i=1}^{q_1}|b_{i,1}|$ and $|{\mathbf B}|_2=(\sum_{i=1}^{q_1}b^2_{i,1})^{1/2}$ to denote the $L_{\infty}$-norm, $L_1$-norm and $L_2$-norm of the $q_1$-dimensional vector ${\mathbf B}$, respectively. Given index sets $\mathcal{S}_1 \subset [q_1]$ and $\mathcal{S}_2 \subset [q_2]$, denote by $[{\mathbf B}]_{\mathcal{S}_1, \mathcal{S}_2}$ the $|\mathcal{S}_1| \times |\mathcal{S}_2|$ matrix that is obtained by extracting the rows of a $q_1 \times q_2$ matrix ${\mathbf B}$ indexed by $\mathcal{S}_1$ and columns indexed by $\mathcal{S}_2$. For simplicity and when no confusion arises, we use the notation ${\mathbf g}_i(\boldsymbol{\theta})=\{g_{i,1}(\boldsymbol{\theta}),\ldots,g_{i,r}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$ as the equivalence to ${\mathbf g}({\mathbf x}_i;\boldsymbol{\theta})$, and denote by $\mathbb{E}_n(\cdot)=n^{-1}\sum_{i=1}^{n}\cdot$. Let $\bar{{\mathbf g}}(\boldsymbol{\theta})=\mathbb{E}_n\{{\mathbf g}_i(\boldsymbol{\theta})\}$, and write its $j$-th component as $\bar{g}_j(\boldsymbol{\theta})=\mathbb{E}_n\{g_{i,j}(\boldsymbol{\theta})\} $. Denote by $\nabla^2_{\boldsymbol{\theta}} g_{i,j} ( \boldsymbol{\theta})=\{\partial^2 g_{i,j}(\boldsymbol{\theta})/\partial\theta_{k_1}\partial\theta_{k_2}\}_{k_1,k_2\in[p]}$, a $p\times p$ matrix, the second-order derivative of $g_{i,j} (\boldsymbol{\theta})$ with respect to $\boldsymbol{\theta}$. Let $\widehat{\boldsymbol{\Gamma}}(\boldsymbol{\theta}) =\nabla_{\boldsymbol{\theta}} \bar{{\mathbf g}}(\boldsymbol{\theta})$ and $\widehat{{\mathbf V}}(\boldsymbol{\theta})=\mathbb{E}_n\{{\mathbf g}_i(\boldsymbol{\theta})^{\otimes2} \}$. For a given set $\mathcal{F} \subset[r]$, we denote by ${\mathbf g}_{i,\mathcal{F}}(\boldsymbol{\theta})$ the subvector of ${\mathbf g}_i(\boldsymbol{\theta})$ collecting the components indexed by $\mathcal{F}$. Analogously, let $\bar{{\mathbf g}}_{\mathcal{F}}(\boldsymbol{\theta})=\mathbb{E}_n\{{\mathbf g}_{i,\mathcal{F}}(\boldsymbol{\theta})\}$, $\widehat{\boldsymbol{\Gamma}}_{\mathcal{F}}(\boldsymbol{\theta}) =\nabla_{\boldsymbol{\theta}} \bar{{\mathbf g}}_{\mathcal{F}}(\boldsymbol{\theta})$ and $\widehat{{\mathbf V}}_{\mathcal{F}}(\boldsymbol{\theta})=\mathbb{E}_n\{{\mathbf g}_{i,\mathcal{F}}(\boldsymbol{\theta})^{\otimes2} \} $. We also write ${\mathbf a}_{\mathcal{F}}$ as the corresponding subvector of vector ${\mathbf a}$. Recall $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})=\mathbb{E}_n[\log \{ 1+\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_i (\boldsymbol{\theta})\}]-\sum_{j=1}^{r} P_{\nu}(|\lambda_j|)$ and $\hat\boldsymbol{\lambda}(\boldsymbol{\theta})=\{\hat\lambda_1(\boldsymbol{\theta}),\ldots,\hat\lambda_r(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }=\arg\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})} f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$. Write $\mathcal{R}(\boldsymbol{\theta})={\rm supp}\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})\}$, $\mathcal{R}_n={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$, and $\mathcal{M}_{\boldsymbol{\theta}}^*=\{j\in [r]: |\bar{g}_j(\boldsymbol{\theta})| \geq C_*\nu \rho'(0^{+}) \}$ for some $C_*\in(0,1)$. Define $\mathcal{M}_{\boldsymbol{\theta}}(c)=\{ j\in[r]:|\bar{g}_j(\boldsymbol{\theta})| \geq c\nu\rho'(0^{+}) \}$ for $c \in (C_*,1)$. Recall $\alpha_n=n^{-1/2}(\log r)^{1/2}$.

Additional Numerical Results

The Impact of the Prior $\pi_0(\boldsymbol{\theta})$

In this section, we investigate the impact of the prior distribution $\pi_0(\boldsymbol{\theta})$ in (ref) on our proposed Bayesian penalized EL methods in estimating the true parameter $\boldsymbol{\theta}_0$. More specifically, we adopt the data generation process outlined in Section (ref) with the true parameter $\boldsymbol{\theta}_0=(0.5,0.5)^{\mathrm{\scriptscriptstyle \top} }$, and consider three choices for the prior:

itemize• the prior distribution $\mathcal{N}\{(-1, -1)^{\mathrm{\scriptscriptstyle \top} }, 0.5^2{\mathbf I}_2\}$, which contains no correct information about the truth. • the prior $\mathcal{N}\{(0.6, 0.6)^{\mathrm{\scriptscriptstyle \top} }, 0.5^2{\mathbf I}_2\}$, which concentrates around the true value. • the improper uniform prior, which provides no information about $\boldsymbol{\theta}_0$.

For the $49$ initial points mentioned in Section (ref), we calculate the measure ${\rm MSE}_2$ defined in Section (ref) to evaluate the performance of these estimators. The results for M-H and MAMIS are derived based on the generated samples of size $3500$. Table (ref) summarizes the performance of our proposed methods with such selected three priors. In particular, we have observed that when the prior is specified “closer" to the truth, the resulting estimator has better performance in comparison to the one using a non-informative prior. Conversely, if a prior is specified “further away" from the truth, the performance of the resulting estimator deteriorates and becomes less competitive.

table[table omitted — 2,215 chars of source]

Non-Gaussian Data Generation Process

In this section, we further validate the efficacy of our proposed methods by conducting some additional simulation studies. For the simulation examples considered in Sections (ref) and (ref), we let all instrumental variables (IVs) $z_{i,j}$ be independently and identically distributed following the Student's $t$-distribution with three degrees of freedoms. For the $49$ initial points mentioned in Section (ref), we calculate the measure ${\rm MSE}_1$ defined in Section (ref) to evaluate the performance of these methods. The results for Algorithms (ref) and (ref) based on sample sizes of 1500, 2500, and 3500 are denoted by (M-H-1, M-H-2, M-H-3) and (MAMIS-1, MAMIS-2, MAMIS-3), respectively. The results for $n=120$ and $n=240$ are presented in Tables (ref) and (ref), respectively. Furthermore, we also compare the penalized empirical likelihood (EL) against the standard EL and the relaxed EL introduced by Shi2016_s for the Student's $t$ IVs. For the $49$ initial points mentioned in Section (ref), we calculate the measure ${\rm MSE}_2$ defined in Section (ref) to evaluate the performance of these estimators. The results are presented in Table (ref), where the results for M-H and MAMIS are derived based on the generated samples of size $3500$.

Overall, these simulation results for the Student's $t$ IVs align closely with those listed in Sections (ref) and (ref). These findings further validate the robustness and effectiveness of our proposed methods.

table[table omitted — 4,517 chars of source]
table[table omitted — 4,532 chars of source]
table[table omitted — 1,575 chars of source]

The Normal Approximation in Finite Samples

In this section, we conduct several numerical simulations to examine the performance of the normal approximation stated in Theorem (ref) to the posterior distribution $\pi^{\dag}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ defined as (ref) in finite samples. More specifically, we adopt the data generation process outlined in Section (ref) with linear link function $\hbar(v)=v$. As described in Section (ref), we identify the true global minima $\hat{\boldsymbol{\theta}}_n$ defined as (ref) through exhaustive search. Subsequently, we calculate its asymptotic covariance matrix $n^{-1}\widehat{{\mathbf H}}^{-1}_{\mathcal{R}_n}$ with $\widehat{{\mathbf H}}_{\mathcal{R}_n}$ defined as (ref). We then generate $5000$ samples, respectively, from the Gaussian distribution $\mathcal{N}(\hat{\boldsymbol{\theta}}_n, n^{-1}\widehat{{\mathbf H}}^{-1}_{\mathcal{R}_n})$ and the posterior distribution $\pi^{\dag}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ defined as (ref). To generate samples from $\mathcal{N}(\hat{\boldsymbol{\theta}}_n, n^{-1}\widehat{{\mathbf H}}^{-1}_{\mathcal{R}_n})$, we use the function {\tt mvrnorm} in the R-package {\tt MASS}. To generate samples from $\pi^{\dag}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$, we use Algorithm (ref) with the burn-in period of $1000$ iterations. Based on these samples, we compute the Wasserstein distance between the two distributions $\mathcal{N}(\hat{\boldsymbol{\theta}}_n, n^{-1}\widehat{{\mathbf H}}^{-1}_{\mathcal{R}_n})$ and $\pi^{\dag}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ using the function {\tt wasserstein} in the R-package {\tt transport}. Figure (ref) below illustrates the average Wasserstein distance between the two distributions across different sample sizes $n$ under $500$ replications. It can be observed that, the two distributions exhibit a relatively large differences at smaller sample sizes, with this distance diminishing notably as the sample size increases.

figure[figure omitted — 466 chars of source]

We further validate the efficacy of our proposed methods in approximating the true global minima $\hat{\boldsymbol{\theta}}_n$ defined as (ref). Table (ref) below presents the measure ${\rm MSE} = \frac{1}{500}\sum_{k=1}^{500} |\check{\boldsymbol{\theta}}_{k} - \hat{\boldsymbol{\theta}}_n|_2^2$ across different sample sizes $n$, where $\check{\boldsymbol{\theta}}_{k}$ is the mean of the $5000$ samples drawn from the posterior distribution $\pi^{\dag}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ in the $k$-th replication. It is evident that with small sample size $n$, although the normal approximation stated in Theorem (ref) to the posterior distribution $\pi^{\dag}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ may not work very well, our method can still effectively approximate the true global minimum $\hat{\boldsymbol{\theta}}_n$. As the sample size $n$ increases, the accuracy of the approximation exhibits a substantial improvement.

table[table omitted — 561 chars of source]

Comparison of the Performance of Posteriors Derived by Different Methods

In this section, we conduct some additional numerical studies to further compare the performance of posteriors derived by different methods. Assume the observations ${\mathbf x}_1,\ldots,{\mathbf x}_n$ are drawn independently from the distribution $F(\boldsymbol{\theta}_0)$ with some unknown parameter $\boldsymbol{\theta}_0$. The likelihood function admits the form $L(\boldsymbol{\theta})=\prod_{i=1}^n f({\mathbf x}_i;\boldsymbol{\theta})$ where $f(\cdot\,;\boldsymbol{\theta})$ is the density function of $F_n(\boldsymbol{\theta})$. Write $\mathcal{X}_n=\{{\mathbf x}_1,\ldots,{\mathbf x}_n\}$. Let $\pi_0(\cdot)$ represent a prior distribution for $\boldsymbol{\theta}$. Then the traditional posterior is given by $\pi^{\rm L}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)\propto \pi_{0}(\boldsymbol{\theta}) \times L(\boldsymbol{\theta})$. To estimate the unknown parameter $\boldsymbol{\theta}_0$, we can also identify it by $\mathbb{E}\{\mathbf{g}({\mathbf x}_i;\boldsymbol{\theta}_0)\}={\mathbf 0}$ with some $r$-dimensional estimating function $\mathbf{g}(\cdot\,;\cdot)$. For given estimating function $\mathbf{g}(\cdot\,; \cdot)$, we can define the empirical likelihood ${\rm EL}(\boldsymbol{\theta})$ and penalized empirical likelihood ${\rm PEL}_\nu(\boldsymbol{\theta})$, respectively, as (ref) and (ref) in the main paper. Hence, the associated EL-based posterior distribution $\pi^{\rm EL}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ and PEL-based posterior distribution $\pi^{\dag}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)$ satisfy $\pi^{\rm EL}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)\propto \pi_{0}(\boldsymbol{\theta}) \times {\rm EL}(\boldsymbol{\theta}) $ and $\pi^{\dag}(\boldsymbol{\theta}\,|\,\mathcal{X}_n)\propto \pi_{0}(\boldsymbol{\theta}) \times {\rm PEL}_\nu(\boldsymbol{\theta})$. We select $\pi_0(\cdot)$ as the improper uniform prior and compare these three posteriors via the following two models:

itemize• Model ${\rm I}$: Let $x_1, \ldots, x_n$ be independent and identically distributed observations from the normal distribution $\mathcal{N}(0,\theta_0^2)$ with $\theta_0 = 1$. We can select $\mathbf{g}(\cdot\,; \cdot) = \{g_1(\cdot\,; \cdot), \ldots, g_r(\cdot\,; \cdot)\}^{\mathrm{\scriptscriptstyle \top} }\in{\mathbb R}^r$ with $g_j(x_{i}; \theta) = x_i^{2j} - (2j-1)!! \theta^{2j}$. • Model ${\rm II}$: Consider the linear regression model $y_i = z_i \theta_0 + e_i, ~i \in [n]$, where $\theta_0 = 0.5$, $z_i \sim \mathcal{N}(0,1)$ is the covariate variable and $e_i \sim \mathcal{N}(0,0.9)$ is the error orthogonal to $z_i$. We can select $\mathbf{g}(\cdot\,; \cdot) = \{g_1(\cdot\,; \cdot), \ldots, g_r(\cdot\,; \cdot)\}^{\mathrm{\scriptscriptstyle \top} }\in{\mathbb R}^r$ with $g_j({\mathbf x}_{i}; \theta) = (y_i - z_i \theta) z_i^j$ and ${\mathbf x}_{i} = (y_i, z_i)^{\mathrm{\scriptscriptstyle \top} }$. • Model ${\rm III}$: Consider the generalized linear model where the covariates $z_i, i \in [n]$, are drawn independently from the gamma distribution with shape parameter $2$ and rate parameter $1$. The response variables $y_i, i \in [n]$, are generated from the Bernoulli distribution such that $\mathbb{P}(y_i = 1\,|\,z_i) = \exp(z_i \theta_0)/\{1+\exp(z_i \theta_0)\}$ with the true parameter $\theta_0 = 0.2$. We can select $\mathbf{g}(\cdot\,; \cdot) = \{g_1(\cdot\,; \cdot), \ldots, g_r(\cdot\,; \cdot)\}^{\mathrm{\scriptscriptstyle \top} }\in{\mathbb R}^r$ with $g_j({\mathbf x}_{i}; \theta) = [y_i - \exp(z_i \theta)/\{1+\exp(z_i \theta)\}] z_i^j$ and ${\mathbf x}_{i} = (y_i, z_i)^{\mathrm{\scriptscriptstyle \top} }$.

We generate $5000$ samples from each of these posterior distributions via the M-H algorithm with the burn-in period of $2000$ iterations. Based on these samples, we compute the Wasserstein distances between $\pi^{\rm L}(\theta\,|\,\mathcal{X}_n)$ and $\pi^{\rm EL}(\theta\,|\,\mathcal{X}_n)$, as well as between $\pi^{\rm L}(\theta\,|\,\mathcal{X}_n)$ and $\pi^{\dag}(\theta\,|\,\mathcal{X}_n)$, using the function {\tt wasserstein} in the R-package {\tt transport}. Figure (ref) below illustrates the average Wasserstein distances between these posterior distributions across different sample sizes $n$ under $500$ replications. Figures (ref) and (ref) show the density functions of $\pi^{\rm L}(\theta\,|\,\mathcal{X}_n)$, $\pi^{\rm EL}(\theta\,|\,\mathcal{X}_n)$ and $\pi^{\dag}(\theta\,|\,\mathcal{X}_n)$ across different sample sizes $n$ in one replication.

figure[figure omitted — 880 chars of source]
figure[figure omitted — 1,363 chars of source]
figure[figure omitted — 1,363 chars of source]

Overall, the numerical results indicate that the posterior distributions constructed the (un-penalized) empirical likelihood without the penalty on the Lagrange multiplier exhibit significant discrepancies from those derived from the likelihood function. However, this disparity can be markedly reduced through the introduction of a penalty term on the Lagrange multipliers in the empirical likelihood.

Comparison to ABC and BSL

In this section, we compare the performance of our proposed methods with the approximate Bayesian computation (ABC) and Bayesian synthetic likelihood (BSL) methods, implemented as described below.

itemize• {\tt abc}: The R function performs parameter estimation using the approximate Bayesian computation (ABC) algorithm in the R-package {\tt abc}. • {\tt bsl}: The R function for performing Bayesian synthetic likelihood (BSL) to sample from the approximate posterior distribution in the R-package {\tt BSL}.

We conduct the comparisons using the same three models as described in Section (ref). Both the ABC and BSL methods require the selection of summary statistics. In our numerical experiments, for demonstration purposes, we chose sufficient statistics for the parameters of interest, thereby favoring the ABC and BSL methods. Specifically, for Model \({\rm I}\), we selected the sample standard deviation as the summary statistic. For Models \({\rm II}\) and \({\rm III}\), we selected the maximum likelihood estimator as the summary statistic.

When implementing {\tt abc}, we set the tolerance levels to \(0.05\), \(0.1\), and \(0.2\) to obtain \(5000\) valid approximate samples of the traditional posterior distributions for the three models. These tolerance levels correspond to \(100{,}000\), \(50{,}000\), and \(25{,}000\) MCMC iterations, respectively. The results are labeled as ({\tt abc}-0.05, {\tt abc}-0.1, {\tt abc}-0.2). For {\tt bsl}, we ran the MCMC sampler for \(7000\) iterations, discarding the first \(2000\) iterations for burn-in.

To evaluate the performance of {\tt abc} and {\tt bsl}, we used the {\tt wasserstein} function from the R package {\tt transport} to compute the Wasserstein distances between the approximate samples and those obtained directly via the Metropolis-Hastings (M-H) algorithm from the traditional posterior distribution. For our PEL method, we report results for the case with \(r=50\) and also evaluate the corresponding Wasserstein distances relative to those generated by the M-H algorithm. Figure (ref) below presents the results, showing the average Wasserstein distances calculated for different sample sizes \(n\) under \(500\) replications.

figure[figure omitted — 453 chars of source]

Overall, the numerical results indicate that our proposed method, along with {\tt abc} and {\tt bsl}, effectively approximates the posterior distribution, with the accuracy of the approximation improving as the sample size $n$ increases. For smaller sample sizes, {\tt abc} and {\tt bsl} exhibit better accuracy. However, as the sample size $n$ grows, our proposed method demonstrates substantial improvements, achieving satisfactory approximation accuracy. Additionally, it is notable that the ABC method often requires significantly higher computational costs to achieve comparable approximation accuracy.

Here, we also note that the settings favor the ABC and BSL methods in the choice of summary statistics, yet our method performs very competitively. Figure (ref)(b) presents results from a slightly modified setting of Model I, where the data generation process is changed from a normal distribution to a Student's $t$-distribution with $10$ degrees of freedom, and we are still interested in estimating the standard deviation parameter. In this case, the summary statistic based on the sample standard deviation for ABC and BSL is no longer sufficient. For side-by-side comparisons, the corresponding case with results from the normal distribution is shown in Figure (ref)(a). In this scenario, the performance of the PEL-based method is superior, demonstrating the compelling performance of our approach, owing to the merits of using empirical likelihood.

figure[figure omitted — 351 chars of source]

Discussion of the Technical Conditions

Conditions (ref) and (ref) are commonly used assumptions in the literature. Condition (ref) is the identification condition for the unknown true parameter $\boldsymbol{\theta}_0$. A similar condition can be found in Shi2016 and Changetal_2018. Condition (ref)(b) requires the covariance matrix of ${\mathbf g}({\mathbf x}_i;\boldsymbol{\theta}_0)$ behaves reasonably well. Conditions (ref)(a) and (ref)(c) impose the moments requirements on each estimating function $g_j(\cdot\,;\cdot)$ and its derivatives. If there exist functions $B_l(\cdot)$ with $\mathbb{E}\{B_l({\mathbf x}_i)\}<\infty$, $l=1,2,3$, such that $|g_j({\mathbf x};\boldsymbol{\theta})|^\gamma\leq B_1({\mathbf x})$, $|\partial g_j({\mathbf x};\boldsymbol{\theta})/\partial\theta_k|^2\leq B_2({\mathbf x})$ and $|\partial^2g_j({\mathbf x};\boldsymbol{\theta})/\partial\theta_{k_1}\partial\theta_{k_2}|^2\leq B_3({\mathbf x})$ for any $j\in[r]$ and $\boldsymbol{\theta}\in\boldsymbol{\Theta}$, then the second requirement in Condition (ref)(a) and the two requirements in Condition (ref)(c) hold automatically. More generally, if there exist functions $B_{l,j}(\cdot)$ such that $|g_j({\mathbf x};\boldsymbol{\theta})|^\gamma\leq B_{1,j}({\mathbf x})$, $|\partial g_j({\mathbf x};\boldsymbol{\theta})/\partial\theta_k|^2\leq B_{2,j}({\mathbf x})$ and $|\partial^2g_j({\mathbf x};\boldsymbol{\theta})/\partial\theta_{k_1}\partial\theta_{k_2}|^2\leq B_{3,j}({\mathbf x})$ for any $j\in[r]$ and $\boldsymbol{\theta}\in\boldsymbol{\Theta}$, and $\max_{j\in[r]}\mathbb{E}\{B_{l,j}^m({\mathbf x}_i)\}\leq Km!H^{m-2}$ for any $m\geq 2$ and $l=1,2,3$ with two universal constants $K, H>0$, it follows from Theorem 2.8 of Petrov_1995 that the second requirement in Condition (ref)(a) and the two requirements in Condition (ref)(c) hold automatically provided $\log(rp)=o(n)$. In fact, the order $O_{\rm p}(1)$ in Conditions (ref)(a) and (ref)(c) can be replaced by $O_{\rm p}(\varpi_n)$ with some diverging sequence $\varpi_n$, and our main results remain valid. We use $O_{\rm p}(1)$ here for ease of presentation. To establish the consistency of the penalized empirical likelihood estimator $\hat{\boldsymbol{\theta}}_n$, Conditions (ref), (ref)(a) and (ref)(b) are needed. Condition (ref)(c) is needed for establishing the asymptotic normality of $\hat{\boldsymbol{\theta}}_n$.

Condition (ref) is standard in the literature. Due to the penalty imposed on the Lagrange multiplier $\boldsymbol{\lambda}$ involved in the optimization (ref), the standard theoretical analysis of empirical likelihood cannot be applied here. Condition (ref)(a) is a technical assumption used to derive the convergence rate of the Lagrange multiplier $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)=\arg\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_n(\hat{\boldsymbol{\theta}}_n)}f_n(\boldsymbol{\lambda};\hat{\boldsymbol{\theta}}_n)$ associated with $\hat{\boldsymbol{\theta}}_n$; see the proof of Lemma (ref) in the supplementary material for details. Condition (ref)(b) requires that each $\hat{\eta}_j$ $(j\in \mathcal{R}_n^{\rm c})$ lies in the interior of $[-\nu \rho'(0^{+}),\nu \rho'(0^{+})]$ with probability approaching one, which is realistic in practice. If the distribution function of the random variable $\hat{\eta}_j$ is continuous at $\pm \nu\rho'(0^+)$, we then have $\mathbb{P}\{|\hat{\eta}_j|=\nu\rho'(0^+)\}=0$. Condition (ref) makes sure that $\hat\boldsymbol{\lambda}(\boldsymbol{\theta})=\arg\max_{\boldsymbol{\lambda}\in\hat{\Lambda}_n(\boldsymbol{\theta})}f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ is continuously differentiable at $\hat{\boldsymbol{\theta}}_n$ with probability approaching one; see Lemma (ref) in Section (ref).

Condition (ref)(a) guarantees that the Lagrange multiplier $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})$ for $\boldsymbol{\theta} \in \mathcal{C}_1$ satisfies two properties: (i) $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})$ is continuously differentiable on $\mathcal{C}_1$ with probability approaching one, and (ii) $\mathcal{R}(\boldsymbol{\theta}) = \mathcal{R}(\hat{\boldsymbol{\theta}}_n)$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ with probability approaching one; see Lemmas (ref) and (ref) in Section (ref). When $\boldsymbol{\theta}\notin\mathcal{C}_1$, characterizing the asymptotic property of $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})$ is quite challenging. Due to $f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}\geq f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ for any $\boldsymbol{\lambda}\in\hat{\Lambda}_n(\boldsymbol{\theta})$, a feasible strategy to construct the lower bound of $f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$ with $\boldsymbol{\theta}\notin\mathcal{C}_1$ is to find a specific $\boldsymbol{\lambda}_*(\boldsymbol{\theta})\in\hat{\Lambda}_n(\boldsymbol{\theta})$ and then derive the lower bound of $f_n\{\boldsymbol{\lambda}_*(\boldsymbol{\theta});\boldsymbol{\theta}\}$ directly, where the asymptotic behavior of $\boldsymbol{\lambda}_*(\boldsymbol{\theta})$ can be well characterized even if $\boldsymbol{\theta}\notin\mathcal{C}_1$. Such strategy has been also used in Changetal2013, Changetal_2016_AOS to study the diverging rate of the conventional empirical likelihood ratio evaluated at a value not near the truth. In current setting, Condition (ref)(b) is applied to derive the lower bound of $f_n\{\boldsymbol{\lambda}_*(\boldsymbol{\theta});\boldsymbol{\theta}\}$ for $\boldsymbol{\theta}\in\mathcal{C}_2$; see Section (ref) for details. Condition (ref)(c) says that the sample covariance matrix of the estimating function and the gradient of the estimating function should behave reasonably well, which will be used to obtain the lower bound of $f_n\{\boldsymbol{\lambda}_*(\boldsymbol{\theta});\boldsymbol{\theta}\}$ for $\boldsymbol{\theta}\in\mathcal{C}_3$. See Section (ref) for details. Condition (ref) is a standard assumption concerning the prior distribution.

Proof of Proposition (ref)

Write $\mathcal{R}_{0}=\mathcal{R}(\boldsymbol{\theta}_0)$. Then

align*[align* omitted — 655 chars of source]

where $\hat{\Lambda}^{\dag}_n(\boldsymbol{\theta}_0)=\{\boldsymbol{\eta}=(\eta_{1},\ldots, \eta_{|\mathcal{R}_{0}|})^{\mathrm{\scriptscriptstyle \top} } \in {\mathbb{R}}^{|\mathcal{R}_{0}|}:\boldsymbol{\eta}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{R}_{0}} (\boldsymbol{\theta}_0) \in \mathcal{V}~\textrm{for any}~i\in[n]\}$ for some open interval $\mathcal{V}$ containing zero. Our first step is to show $\max_{\boldsymbol{\eta} \in\hat{\Lambda}^{\dag}_n(\boldsymbol{\theta}_0)} \mathbb{E}_n[\log \{ 1+\boldsymbol{\eta}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{R}_{0}} (\boldsymbol{\theta}_0)\}]=O_{\rm p}(\ell_n \alpha_n^2)$. To do this, we need the following two lemmas whose proofs are given in Sections (ref) and (ref), respectively.

lemmaLet $\mathscr{F}=\{\mathcal{F} \subset [r]:|\mathcal{F}|\leq \ell_n\}$ and $\mathcal{B}_{\infty,{\rm p}}(\boldsymbol{\theta}_0,\varphi_n)=\{\boldsymbol{\theta} \in \boldsymbol{\Theta}: |\boldsymbol{\theta}-\boldsymbol{\theta}_0|_{\infty}\leq O_{\rm p}(\varphi_n)\}$ for $\varphi_n=o(\ell_n^{-1/2})$. Under Condition {\rm (ref)}, if $\log r=o(n^{1/3})$ and $\ell_n\alpha_n=o(1)$, then $$ \sup_{\boldsymbol{\theta} \in \mathcal{B}_{\infty,{\rm p}}(\boldsymbol{\theta}_0,\varphi_n)} \sup_{\mathcal{F} \in \mathscr{F}} \| \widehat{{\mathbf V}}_{\mathcal{F}} (\boldsymbol{\theta}) - {\mathbf V}_{\mathcal{F}} (\boldsymbol{\theta}_0)\|_2 = O_{\rm p}(\ell_n^{1/2} \varphi_n)+O_{\rm p}(\ell_n \alpha_n)\,.$$
lemmaLet $\log r=o(n^{1/3})$, $\ell_n \alpha_n=o[\min \{\nu,n^{-1/\gamma}\}]$, and $P_{\nu}(\cdot) \in \mathscr{P}$ be a convex function for $\mathscr{P}$ defined in {\rm (ref)}. Assume Conditions {\rm (ref)(a)} and {\rm (ref)(b)} hold. For any $c \in (C_*,1)$, the global maximizer $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta}_0)$ for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.r.t $\boldsymbol{\lambda}$ satisfies ${\rm supp}\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta}_0)\} \subset \mathcal{M}_{\boldsymbol{\theta}_0}(c)$ w.p.a.1.

Define $\tilde{\boldsymbol{\eta}}=\arg \max _{\boldsymbol{\eta} \in \hat{\Lambda}^{\dag}_n(\boldsymbol{\theta}_0)} A_n(\boldsymbol{\theta}_0,\boldsymbol{\eta})$ with $A_n(\boldsymbol{\theta},\boldsymbol{\eta})=\mathbb{E}_n[\log \{ 1+\boldsymbol{\eta}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{R}_{0}} (\boldsymbol{\theta}) \}]$. By Lemma (ref), we have $|\mathcal{R}_{0}| \leq \ell_n$ w.p.a.1. Pick $\delta_n$ satisfying $\delta_n=o(\ell_n^{-1/2} n^{-1/\gamma})$ and $\ell_n^{1/2} \alpha_n=o(\delta_n) $ for $\gamma$ defined in Condition (ref)(a), which can be guaranteed by $\ell_n\alpha_n=o(n^{-1/\gamma})$. Let $\Lambda_{0}=\{\boldsymbol{\eta} \in {\mathbb{R}}^{|\mathcal{R}_{0}|}:|\boldsymbol{\eta}|_2\leq \delta_n \}$ and $\check{\boldsymbol{\eta}}=\arg \max _{\boldsymbol{\eta} \in \Lambda_{0}} A_n(\boldsymbol{\theta}_0,\boldsymbol{\eta})$. Condition (ref)(a) implies $\max _{i \in [n], \boldsymbol{\eta} \in \Lambda_{0}} |\boldsymbol{\eta}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{R}_{0}} (\boldsymbol{\theta}_0)| = o_{\rm p}(1)$. Then, by the Taylor expansion, we have

align*[align* omitted — 571 chars of source]

for some $C\in(0,1)$. By Condition (ref)(b) and the same arguments for deriving Lemma \rm(ref), if $\log r=o(n^{1/3})$ and $\ell_n\alpha_n=o(1)$, we have $\lambda_{\min}\{\widehat{{\mathbf V}}_{\mathcal{M}_{\boldsymbol{\theta}_0}} (\boldsymbol{\theta}_0)\}$ is uniformly bounded away from zero w.p.a.1. Thus $0\leq |\check{\boldsymbol{\eta}}|_2 |\bar{{\mathbf g}}_{\mathcal{R}_{0}}(\boldsymbol{\theta}_0)|_2 - 4^{-1}K_3|\check{\boldsymbol{\eta}}|_2^2 \{1+o_{\rm p}(1)\}$ w.p.a.1, where $K_3$ is specified in Condition (ref)(b). By the moderate deviation of self-normalized sums JingShaoWang2003, we have $ |\bar{{\mathbf g}}(\boldsymbol{\theta}_0)|_{\infty} =O_{\rm p}(\alpha_n) $, which implies $|\bar{{\mathbf g}}_{\mathcal{R}_{0}}(\boldsymbol{\theta}_0)|_2= O_{\rm p}(\ell_n^{1/2} \alpha_n)$ and $|\check{\boldsymbol{\eta}}|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n)=o_{\rm p}(\delta_n)$. Hence, $\check{\boldsymbol{\eta}} \in {\rm int} (\Lambda_{0}) $ w.p.a.1. Since $\Lambda_{0} \subset \hat{\Lambda}^{\dag}_n(\boldsymbol{\theta}_0)$ w.p.a.1, we have $\tilde{\boldsymbol{\eta}}=\check{\boldsymbol{\eta}}$ w.p.a.1 by the concavity of $A_n(\boldsymbol{\theta}_0,\boldsymbol{\eta})$ w.r.t $\boldsymbol{\eta}$. Then $ \max_{\boldsymbol{\eta} \in\hat{\Lambda}^{\dag}_n(\boldsymbol{\theta}_0)} A_n(\boldsymbol{\theta}_0,\boldsymbol{\eta})=O_{\rm p}(\ell_n \alpha_n^2)$. Let $b_n^2=\ell_n \alpha_n^2$ and $F_n(\boldsymbol{\theta})=\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})}f_n(\boldsymbol{\lambda};\boldsymbol{\theta}) $ for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$. Due to $\hat{\boldsymbol{\theta}}_n=\arg \min _{\boldsymbol{\theta} \in \boldsymbol{\Theta}}F_n(\boldsymbol{\theta}) $, we have $F_n(\hat{\boldsymbol{\theta}}_n) \leq F_n(\boldsymbol{\theta}_0)=O_{\rm p}(b_n^2)$.

Our second step is to show that for any $\epsilon_n \rightarrow \infty$ satisfying $b_n^2\epsilon_n^2 n^{2/\gamma}=o(1)$, there exists a universal constant $K>0$ independent of $\boldsymbol{\theta}$ such that ${\mathbb{P}}\{F_n(\boldsymbol{\theta})>K b_n^2 \epsilon_n^2\} \rightarrow 1$ as $n \rightarrow \infty$ for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ satisfying $|\boldsymbol{\theta}-\boldsymbol{\theta}_0|_{\infty} > \epsilon_n \nu$. Thus $|\hat{\boldsymbol{\theta}}_n-\boldsymbol{\theta}_0|_\infty=O_{\rm p}(\epsilon_n\nu)$. Due to $b_n^2=o(n^{-2/\gamma})$, we can select arbitrary slowly diverging $\epsilon_n$. We then have $|\hat{\boldsymbol{\theta}}_n-\boldsymbol{\theta}_0|_\infty=O_{\rm p}(\nu)$ by a standard result from probability theory. For any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ satisfying $|\boldsymbol{\theta}-\boldsymbol{\theta}_0|_{\infty} > \epsilon_n \nu$, let $j_{0}=\arg \max _{j \in [r]}|{\mathbb{E}}\{g_{i,j}(\boldsymbol{\theta})\}|$ and $\mu_{j_{0}}={\mathbb{E}}\{g_{i,j_{0}}(\boldsymbol{\theta})\} $. Define $\tilde{\boldsymbol{\lambda}}=\tau b_n \epsilon_n {\mathbf e}_{j_{0}}$, where $\tau>0$ is a constant to be determined later and ${\mathbf e}_{j_{0}}$ is an $r$-dimensional vector with the $j_{0}$-th component being $1$ and other components being $0$. Without loss of generality, we assume $\mu_{j_{0}}>0$. Condition (ref)(a) implies $\max_{i\in[n]} |\tilde{\boldsymbol{\lambda}}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i}(\boldsymbol{\theta})|= O_{\rm p}(b_n \epsilon_n n^{1/\gamma})=o_{\rm p}(1)$. Then $\tilde{\boldsymbol{\lambda}} \in \hat{\Lambda}_n(\boldsymbol{\theta})$ w.p.a.1. Write $\tilde{\boldsymbol{\lambda}}=(\tilde{\lambda}_{1},\ldots, \tilde{\lambda}_{r})^{\mathrm{\scriptscriptstyle \top} }$. By the Taylor expansion, it holds w.p.a.1 that

align*[align* omitted — 685 chars of source]

for some $\bar{C}\in(0,1)$, which implies $$ {\mathbb{P}}\big\{F_n(\boldsymbol{\theta}) \leq Kb_n^2 \epsilon_n^2 \big\} \leq {\mathbb{P}}\big\{\bar{g}_{j_{0}}(\boldsymbol{\theta})-\mu_{j_{0}} \leq b_n \epsilon_n[ K\tau^{-1} + \tau \mathbb{E}_n\{g^2_{i,j_{0}}(\boldsymbol{\theta})\}]+C\nu-\mu_{j_{0}}\big\}+o(1)\,.$$ From Condition (ref)(a) and Jensen's inequality, there exists a universal constant $L>0$ independent of $\boldsymbol{\theta}$ such that ${\mathbb{P}}[\mathbb{E}_n\{g^2_{i,j_{0}}(\boldsymbol{\theta})\} >L] \rightarrow 0$. Taking $\tau=(KL^{-1})^{1/2}$, we obtain $ {\mathbb{P}}\{F_n(\boldsymbol{\theta}) \leq K b_n^2 \epsilon_n^2 \} \leq {\mathbb{P}} \{\bar{g}_{j_{0}}(\boldsymbol{\theta})-\mu_{j_{0}} \leq 2 b_n\epsilon_n (KL)^{1/2} +C\nu -\mu_{j_{0}}\}+o(1)$. By Condition (ref), $\mu_{j_{0}} \geq \Delta(\epsilon_n \nu) \geq K_{1} \epsilon_n \nu/2$ with $K_1$ defined in Condition (ref) for sufficiently large $n$. We select sufficiently small $K>0$. Due to $b_n=o(\nu)$, when $n$ is sufficiently large, $2 b_n \epsilon_n (KL)^{1/2}+C\nu-\mu_{j_{0}} \leq -\check{C}\mu_{j_{0}}$ for some $\check{C}\in(0,1)$. Hence, we have $ \sqrt{n}\{2 b_n\epsilon_n (KL)^{1/2}+C\nu-\mu_{j_{0}}\} \leq -\check{C}\sqrt{n} \mu_{j_{0}} \rightarrow -\infty$. By the Central Limit Theorem, $\sqrt{n}\{\bar{g}_{j_{0}}( \boldsymbol{\theta})- \mu_{j_{0}}\}\rightarrow\mathcal{N}(0, \sigma^2)$ in distribution for some $\sigma>0$, which implies that ${\mathbb{P}}\{F_n(\boldsymbol{\theta}) \leq Kb_n^2 \epsilon_n^2\} \rightarrow 0$ as $n \rightarrow \infty$ for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ satisfying $|\boldsymbol{\theta}-\boldsymbol{\theta}_0|_{\infty} > \epsilon_n\nu$. We complete the proof of Proposition (ref). $\Box$

Proof of Theorem (ref)

Recall $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})=\arg \max _{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})} f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ and $\mathcal{R}_n={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$. We first present two lemmas whose proofs are given in Sections (ref) and (ref), respectively.

lemmaLet $P_{\nu}(\cdot) \in \mathscr{P}$ be a convex function with bounded second-order derivative around $0$, where $\mathscr{P}$ is defined in {\rm (ref)}. Under the conditions of Proposition {\rm (ref)} and Conditions {\rm (ref)(c)} and {\rm (ref)(a)}, if $\ell_n\nu^2=o(1)$, it holds w.p.a.1 that the global maximizer $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)=(\hat{\lambda}_{1},\ldots,\hat{\lambda}_{r})^{\mathrm{\scriptscriptstyle \top} }$ for $f_n(\boldsymbol{\lambda};\hat{\boldsymbol{\theta}}_n)$ w.r.t $\boldsymbol{\lambda}$ satisfies: ${\rm (i)}$ $|\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n)$, ${\rm (ii)}$ $\mathcal{R}_n \subset \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})$ with $\tilde{c}$ given in Condition {\rm (ref)(a)}, and ${\rm (iii)}$ ${\mbox{\rm sgn}}(\hat{\lambda}_j) ={\mbox{\rm sgn}}\{\bar{g}_j(\hat{\boldsymbol{\theta}}_n)\}$ for any $j \in \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})$ with $\hat{\lambda}_j \neq 0$.
lemmaUnder the conditions of Lemma {\rm (ref)} and Condition {\rm (ref)(b)}, it holds w.p.a.1 that the global maximizer $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})$ for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ w.r.t $\boldsymbol{\lambda}$ is continuously differentiable at $\hat{\boldsymbol{\theta}}_n$ and $[\nabla_{\boldsymbol{\theta}} \hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)]_{\mathcal{R}_n^{\rm c},[p]}={\mathbf 0}$.

For simplicity, we write $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)$ as $\hat{\boldsymbol{\lambda}}=(\hat\lambda_1,\ldots,\hat\lambda_r)^{\mathrm{\scriptscriptstyle \top} }$. Then we have

align[align omitted — 266 chars of source]

where $\hat{\boldsymbol{\eta}}=(\hat{\eta}_{1},\ldots, \hat{\eta}_{r})^{\mathrm{\scriptscriptstyle \top} }$ with $\hat{\eta}_j=\nu \rho' (|\hat{\lambda}_j|;\nu) {\mbox{\rm sgn}}(\hat{\lambda}_j)$ for $\hat{\lambda}_j\neq0$ and $\hat{\eta}_j \in [-\nu \rho'(0^{+}),\nu \rho'(0^{+})]$ for $\hat{\lambda}_j=0$. By the Taylor expansion, we know that

align*[align* omitted — 645 chars of source]

for some $C\in(0,1)$, which implies $\hat{\boldsymbol{\lambda}}_{ \mathcal{R}_n}= {\mathbf A}^{-1}(\hat{\boldsymbol{\theta}}_n)\{\bar{{\mathbf g}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)- \hat{\boldsymbol{\eta}}_{ \mathcal{R}_n}\}$. Since $\hat{\boldsymbol{\theta}}_n=\arg\min_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$, we have ${\mathbf 0} =\nabla_{\boldsymbol{\theta}} f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}|_{\boldsymbol{\theta}=\hat{\boldsymbol{\theta}}_n}$. Notice that

align*[align* omitted — 899 chars of source]

By Lemma (ref) and (ref), we have $[\nabla_{\boldsymbol{\theta}} \hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)]_{\mathcal{R}_n^{\rm c},[p]}={\mathbf 0}$ w.p.a.1 and $\partial f_n(\hat{\boldsymbol{\lambda}};\hat{\boldsymbol{\theta}}_n)/\partial\boldsymbol{\lambda}_{{ \mathcal{R}_n}} = {\mathbf 0}$. Thus, it holds w.p.a.1 that

align*[align* omitted — 638 chars of source]

We then obtain $ {\mathbf 0} = {\mathbf B}(\hat{\boldsymbol{\theta}}_n)^{\mathrm{\scriptscriptstyle \top} } {\mathbf A}^{-1}(\hat{\boldsymbol{\theta}}_n) \{\bar{{\mathbf g}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)- \hat{\boldsymbol{\eta}}_{ \mathcal{R}_n}\}$. To derive the limiting distribution of $\hat{\boldsymbol{\theta}}_n$, we need the following lemmas whose proofs are given in Sections (ref), (ref) and (ref), respectively.

lemmaUnder the conditions of Lemma {\rm (ref)}, $ \| {\mathbf A}(\hat{\boldsymbol{\theta}}_n) - \widehat{{\mathbf V}}_{ \mathcal{R}_n} (\hat{\boldsymbol{\theta}}_n) \|_2 =O_{\rm p}(\ell_n n^{1/\gamma}\alpha_n) $, and $ | \{{\mathbf B}(\hat{\boldsymbol{\theta}}_n)- \widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n} (\hat{\boldsymbol{\theta}}_n) \}{\mathbf t} |_2=|{\mathbf t}|_2 \cdot O_{\rm p}(\ell_n\alpha_n) $ holds uniformly over ${\mathbf t} \in {\mathbb{R}}^p$.
lemmaAssume that the conditions of Proposition {\rm (ref)} and Condition {\rm (ref)(c)} hold. For $\mathscr{F}$ defined in Lemma {\rm(ref)}, $$ \sup_{\mathcal{F} \in \mathscr{F}} | \{\widehat{\boldsymbol{\Gamma}}_{\mathcal{F}}(\hat{\boldsymbol{\theta}}_n)- \boldsymbol{\Gamma}_{\mathcal{F}}(\boldsymbol{\theta}_0)\} {\mathbf t} |_2= |{\mathbf t}|_2 \cdot\{O_{\rm p}(\ell_n^{1/2} \nu)+O_{\rm p}(\ell_n^{1/2} \alpha_n)\} $$ holds uniformly over ${\mathbf t} \in {\mathbb{R}}^p$.
lemmaLet $\widehat{{\mathbf H}}_{\mathcal{F}}=\{\widehat{\boldsymbol{\Gamma}}_{\mathcal{F}}(\hat{\boldsymbol{\theta}}_n)^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf V}}^{-1/2}_{\mathcal{F}} (\hat{\boldsymbol{\theta}}_n)\}^{\otimes2}$ for any $\mathcal{F} \in \mathscr{F}$, where $\mathscr{F}$ is defined in Lemma {\rm(ref)}. Assume that the conditions of Proposition {\rm (ref)} and Conditions {\rm (ref)(c)} and {\rm (ref)} hold. If $\ell_n^2 \nu^2\log r=o(1)$ and $\ell_n^{3}\alpha_n^2\log r=o(1)$, for any ${\mathbf t} \in {\mathbb{R}}^p$ with $|{\mathbf t}|_2=1$, we have $$\sup_{\mathcal{F} \in \mathscr{F}} \sup_{u \in {\mathbb{R}}} |{\mathbb{P}}\{n^{1/2} {\mathbf t}^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf H}}_{\mathcal{F}}^{-1/2} \widehat{\boldsymbol{\Gamma}}_{\mathcal{F}}(\hat{\boldsymbol{\theta}}_n)^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf V}}^{-1}_{\mathcal{F}} (\hat{\boldsymbol{\theta}}_n) \bar{{\mathbf g}}_{\mathcal{F}} (\boldsymbol{\theta}_0) \leq u \} - \Phi(u) | \rightarrow 0$$ as $n \rightarrow \infty$, where $\Phi(\cdot)$ is the cumulative distribution function of the standard Gaussian distribution.

For any ${\mathbf t} \in {\mathbb{R}}^p$ with $|{\mathbf t}|_2=1$, let $\boldsymbol{\delta}=\widehat{{\mathbf H}}_{ \mathcal{R}_n}^{-1/2} {\mathbf t}$ and ${\mathbf U} = \widehat{{\mathbf V}}^{-1/2}_{ \mathcal{R}_n} (\hat{\boldsymbol{\theta}}_n) \widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)$. Then $\widehat{{\mathbf H}}_{{ \mathcal{R}_n}} = {\mathbf U}^{{\mathrm{\scriptscriptstyle \top} }, \otimes2}$ and $ |\widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n) \boldsymbol{\delta}|_2^2 \leq \lambda_{\max} \{\widehat{{\mathbf V}}_{ \mathcal{R}_n} (\hat{\boldsymbol{\theta}}_n)\} |{\mathbf U} ({\mathbf U}^{{\mathrm{\scriptscriptstyle \top} }, \otimes2})^{-1/2} {\mathbf t}|_2^2= \lambda_{\max} \{\widehat{{\mathbf V}}_{ \mathcal{R}_n} (\hat{\boldsymbol{\theta}}_n)\}$. By Condition (ref)(b) and Lemma (ref), $|\widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n) \boldsymbol{\delta}|_2=O_{\rm p}(1)$. Under Conditions (ref)(b) and (ref), Lemmas (ref) and (ref) imply $|\boldsymbol{\delta}|_2^2 \leq \lambda_{\max} \{\widehat{{\mathbf V}}_{ \mathcal{R}_n} (\hat{\boldsymbol{\theta}}_n)\} \lambda_{\min}^{-1}\{\widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)^{{\mathrm{\scriptscriptstyle \top} },\otimes2} \} = O_{\rm p}(1)$. By Lemma (ref), we have w.p.a.1 that $\mathcal{R}_n \subset \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})$ and ${\mbox{\rm sgn}}(\hat{\lambda}_j) ={\mbox{\rm sgn}}\{\bar{g}_j(\hat{\boldsymbol{\theta}}_n)\}$ for any $j \in \mathcal{R}_n$. Since $P_{\nu}(\cdot) \in \mathscr{P}$ has bounded second-order derivative around $0$, by Lemma (ref), it holds w.p.a.1 that

align*[align* omitted — 495 chars of source]

for some $c_j\in(0,1)$. As shown in the proof of Lemma (ref), $ |\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})}(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})}(\hat{\boldsymbol{\theta}}_n)\}|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n)$. Then $|\bar{{\mathbf g}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)-\hat{\boldsymbol{\eta}}_{ \mathcal{R}_n}|_2 =O_{\rm p}(\ell_n^{1/2}\alpha_n)$. Due to $ {\mathbf B}(\hat{\boldsymbol{\theta}}_n)^{\mathrm{\scriptscriptstyle \top} } {\mathbf A}^{-1}(\hat{\boldsymbol{\theta}}_n) \{\bar{{\mathbf g}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)- \hat{\boldsymbol{\eta}}_{ \mathcal{R}_n}\}={\mathbf 0}$, by the triangle inequality,

align*[align* omitted — 1,206 chars of source]

By Lemma (ref),

align*[align* omitted — 785 chars of source]

Hence, $\boldsymbol{\delta}^{\mathrm{\scriptscriptstyle \top} } \widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf V}}_{ \mathcal{R}_n}^{-1} (\hat{\boldsymbol{\theta}}_n) \{\bar{{\mathbf g}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)-\hat{\boldsymbol{\eta}}_{ \mathcal{R}_n} \}= O_{\rm p}(\ell_n^{3/2} n^{1/\gamma} \alpha_n^2)$. By the Taylor expansion, we have

align[align omitted — 839 chars of source]

where $\tilde{\boldsymbol{\theta}}$ is on the line joining $\boldsymbol{\theta}_0$ and $\hat{\boldsymbol{\theta}}_n$. Write $\hat{\boldsymbol{\theta}}_n=(\hat{\theta}_{n,1},\ldots,\hat{\theta}_{n,p})^{\mathrm{\scriptscriptstyle \top} }$, $\boldsymbol{\theta}_0=(\theta_{0,1},\ldots,\theta_{0,p})^{\mathrm{\scriptscriptstyle \top} }$ and $\tilde{\boldsymbol{\theta}}=(\tilde{\theta}_{1},\ldots,\tilde{\theta}_{p})^{\mathrm{\scriptscriptstyle \top} }$. By the Taylor expansion, Jensen's inequality and Cauchy-Schwarz inequality, it holds that

align*[align* omitted — 798 chars of source]

where $\dot{\boldsymbol{\theta}}^{(j,k)}$ lies on the jointing line between $\tilde{\boldsymbol{\theta}}$ and $\hat{\boldsymbol{\theta}}_n$. Recall $p$ is fixed. By Proposition (ref), $|\hat{\boldsymbol{\theta}}_n-\boldsymbol{\theta}_0|_2=O_{\rm p}(\nu)$. Together with Condition {\rm (ref)(c)}, we have $| \{\widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\tilde{\boldsymbol{\theta}})- \widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)\} (\hat{\boldsymbol{\theta}}_n-\boldsymbol{\theta}_0)|_2= O_{\rm p}(\ell_n^{1/2} \nu^2)$. Recall $\hat{\boldsymbol{\psi}}_{ \mathcal{R}_n}= \widehat{{\mathbf H}}_{ \mathcal{R}_n}^{-1} \widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf V}}^{-1}_{ \mathcal{R}_n} (\hat{\boldsymbol{\theta}}_n) \hat{\boldsymbol{\eta}}_{ \mathcal{R}_n}$ and $\boldsymbol{\delta}=\widehat{{\mathbf H}}_{ \mathcal{R}_n}^{-1/2} {\mathbf t}$. Then (ref) leads to

align*[align* omitted — 668 chars of source]

By Lemma (ref), we have $ n^{1/2} {\mathbf t}^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf H}}_{ \mathcal{R}_n}^{1/2}(\hat{\boldsymbol{\theta}}_n-\boldsymbol{\theta}_0-\hat{\boldsymbol{\psi}}_{ \mathcal{R}_n})\rightarrow \mathcal{N}(0,1)$ in distribution as $n \rightarrow \infty$. $\Box$

Proof of Theorem (ref)

Assume $(r,\ell_n,\nu)$ satisfy the following restrictions:

equation[equation omitted — 202 chars of source]

To construct Theorem (ref), we need the following proposition whose proof is given in Section (ref).

propositionLet $P_{\nu}(\cdot) \in \mathscr{P}$ be convex and assume $\rho(t;\nu)=\nu^{-1}P_{\nu}(t)$ has bounded second-order derivative with respect to $t$ around $0$, where $\mathscr{P}$ is defined as {\rm (ref)}. Assume $(r,\ell_n,\nu)$ satisfy (ref). {\rm (i)} Under Conditions {\rm (ref)--(ref)}, {\rm (ref)(a)} and {\rm (ref)(a)}, then $ \aleph_n(\boldsymbol{\theta})= 2^{-1} (\boldsymbol{\theta}-\hat\boldsymbol{\theta}_n)^{\mathrm{\scriptscriptstyle \top} } \widehat{\mathbf H}_{\mathcal{R}_n} (\boldsymbol{\theta}-\hat\boldsymbol{\theta}_n) + |\boldsymbol{\theta}-\hat\boldsymbol{\theta}_n|_2^2 \cdot O_{\rm p}(\varpi_n) $ holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$ with $\varpi_n=\max\{\ell_n^{3/2}\alpha_n, \nu, \ell_nn^{1/\gamma}\alpha_n\}$, where $\widehat{\mathbf H}_{\mathcal{R}_n}$ is defined in (ref) and the term $O_{\rm p}(\varpi_n) $ holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$. {\rm (ii)} Under Conditions {\rm (ref)}, {\rm (ref)}, {\rm (ref)(a)} and {\rm (ref)(b)}, then $ \inf_{\boldsymbol{\theta} \in \mathcal{C}_2} \aleph_n(\boldsymbol{\theta})\geq (8K_4)^{-1}\kappa_{n}^2 $ with probability approaching one, where $K_4$ and $\kappa_n$ are specified in Conditions {\rm(ref)(b)} and {\rm(ref)(b)}, respectively. {\rm (iii)} Under Conditions {\rm (ref)}, {\rm (ref)}, {\rm (ref)(a)} and {\rm (ref)(c)}, then $\inf_{\boldsymbol{\theta} \in \mathcal{C}_3} \aleph_n(\boldsymbol{\theta}) \geq 4^{-1}K_7^{1/2} \xi_n\beta_{n}$ with probability approaching one for any $\xi_n$ satisfying $\beta_n^{-1}\ell_n\alpha_n^2 \ll \xi_n \ll \beta_n$, where $K_7$ is specified in Condition {\rm(ref)(c)}.

Recall the posterior distribution $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) \propto \pi_{0}(\boldsymbol{\theta})\times \exp [-n\log n-n f_n\{\hat\boldsymbol{\lambda}(\boldsymbol{\theta});\boldsymbol{\theta} \}]I(\boldsymbol{\theta} \in \boldsymbol{\Theta})$. For any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$, let $w_n(\boldsymbol{\theta})=-n\log n-nf_n\{\hat\boldsymbol{\lambda}(\boldsymbol{\theta}); \boldsymbol{\theta}\}$ and write ${\mathbf t} = n^{1/2}(\boldsymbol{\theta}-\hat\boldsymbol{\theta}_n)$. Define $\mathcal{T}_n = \{{\mathbf t} \in {\mathbb{R}}^p: {\mathbf t} = n^{1/2}(\boldsymbol{\theta}-\hat\boldsymbol{\theta}_n), \boldsymbol{\theta} \in \boldsymbol{\Theta}\}$. Denote by $\pi_{0,{\mathbf t}}(\cdot)$ and $\pi^{\dag}_{{\mathbf t}}(\cdot\,|\,\mathcal{X}_n)$ the prior and the posterior distributions of ${\mathbf t}$, respectively. Then, $\pi_{0,{\mathbf t}}({\mathbf t})=n^{-p/2}\pi_{0}(\hat{\boldsymbol{\theta}}_n+n^{-1/2}{\mathbf t})$ and

align[align omitted — 736 chars of source]

To prove Theorem (ref), it is equivalent to show

align[align omitted — 512 chars of source]

in probability. It follows from the triangle inequality that

align[align omitted — 1,461 chars of source]

Notice that $ {\rm I} \geq |C_n - (2\pi)^{p/2} \pi_{0}(\hat{\boldsymbol{\theta}}_n) |\widehat{{\mathbf H}}_{ \mathcal{R}_n}|^{-1/2} | = {\rm II} $. To show (ref), it suffices to show $C_n^{-1} {\rm I} =o_{\rm p}(1)$. Recall $\widehat{{\mathbf H}}_{{ \mathcal{R}_n}}=\{\widehat{\boldsymbol{\Gamma}}_{{ \mathcal{R}_n}}(\hat\boldsymbol{\theta}_n)^{\mathrm{\scriptscriptstyle \top} }\widehat{{\mathbf V}}_{{ \mathcal{R}_n}}^{-1/2}(\hat\boldsymbol{\theta}_n)\}^{\otimes2}$. Under Conditions (ref)(b) and (ref), by Proposition (ref), Lemmas (ref) and (ref), if $\log r=o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n \nu^2=o(1)$, we know that the eigenvalues of $\widehat{{\mathbf H}}_{ \mathcal{R}_n}$ are uniformly bounded away from zero and infinity w.p.a.1. Notice that $\widehat{{\mathbf H}}_{{ \mathcal{R}_n}}$ is a $p \times p$ matrix with fixed $p$. Thus, $|\widehat{{\mathbf H}}_{ \mathcal{R}_n}|^{-1/2}$ is uniformly bounded away from zero w.p.a.1. Since $\pi_{0}(\boldsymbol{\theta})$ is bounded away from zero around $\boldsymbol{\theta}_0$ and $|\hat{\boldsymbol{\theta}}_n - \boldsymbol{\theta}_0|_\infty=O_{\rm p}(\nu)$, we know $\pi_0(\hat{\boldsymbol{\theta}}_n)$ is bounded away from zero w.p.a.1. If ${\rm I} =o_{\rm p}(1)$, then $| C_n - (2\pi)^{p/2} \pi_{0}(\hat{\boldsymbol{\theta}}_n) |\widehat{{\mathbf H}}_{ \mathcal{R}_n}|^{-1/2} | =o_{\rm p}(1)$, which implies $C_n^{-1}=O_{\rm p}(1)$. Hence, to show (ref), we only need to show ${\rm I} =o_{\rm p}(1)$. Recall $\ell_n \ll \min\{n^{(\gamma-2)/(9\gamma)}(\log r)^{-1/9},n^{1/3}(\log r)^{-1},n^{(\gamma-2)/(2\gamma)}(\log r)^{-3/2}\}$ and $\ell_nn^{-1/2}(\log r)^{1/2} \ll \nu \ll \min\{\ell_n^{-7/2}n^{-1/\gamma},(\log r)^{-1}\}$. We break the domain of integration into four regions:

align[align omitted — 412 chars of source]

Then $ {\rm I} = {\rm I}(1) + {\rm I}(2)+ {\rm I}(3)+{\rm I}(4) $ with

align*[align* omitted — 426 chars of source]

In the sequel, we will show each ${\rm I}(k) = o_{\rm p}(1)$.

For ${\rm I}(3)$, by the triangle inequality, we have

align*[align* omitted — 433 chars of source]

Due to $\int_{\mathcal{D}_3} \pi_{0}(\hat{\boldsymbol{\theta}}_n+n^{-1/2}{\mathbf t}) \,{\rm d}{\mathbf t} \leq n^{p/2} \int_{{\mathbb{R}}^{p}} \pi_{0}(\boldsymbol{\theta}) \,{\rm d}\boldsymbol{\theta} \leq n^{p/2}$, Proposition (ref)(iii) implies that

align*[align* omitted — 518 chars of source]

w.p.a.1 for any $\beta_n^{-1}\ell_n \alpha_n^2 \ll \xi_n \ll \beta_n$. Since $r\gg n$, we can select suitable $\xi_n$ satisfying $ n \beta_n \xi_n \gg \log n $. Then

align*[align* omitted — 234 chars of source]

Due to $n\beta_n^2 \rightarrow \infty$, Proposition 1.1 of Hsu2012 implies that

align*[align* omitted — 354 chars of source]

Therefore, ${\rm I}(3) = o_{\rm p}(1)$.

For ${\rm I}(2)$, it holds that

align*[align* omitted — 432 chars of source]

Since $n\alpha_n^2 \rightarrow \infty$, using the same arguments given above, we have $\pi_{0}(\hat{\boldsymbol{\theta}}_n)\int _{\mathcal{D}_2} \exp( -{\mathbf t}^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf H}}_{ \mathcal{R}_n} {\mathbf t}/2 ) \,{\rm d}{\mathbf t} = o_{\rm p}(1)$. By Proposition (ref)(ii), it then holds w.p.a.1 that

align*[align* omitted — 517 chars of source]

Since $\log n \ll n\kappa_{n}^2$, we have

align*[align* omitted — 234 chars of source]

Therefore, ${\rm I}(2) = o_{\rm p}(1)$.

For ${\rm I}(1)$, by the triangle inequality, we have

align*[align* omitted — 735 chars of source]

Due to $\mathcal{D}_1 \subset \mathcal{T}_n$, by Proposition (ref)(i), it holds that

align*[align* omitted — 693 chars of source]

where $\varpi_n=\max \{\ell_n^{3/2} \alpha_n, \nu, \ell_nn^{1/\gamma}\alpha_n\}$. By Condition (ref), we have $ \sup_{{\mathbf t} \in \mathcal{D}_{1}} |\pi_{0}(\hat{\boldsymbol{\theta}}_n+n^{-1/2}{\mathbf t})-\pi_{0}(\hat{\boldsymbol{\theta}}_n)| =o_{\rm p}(1)$, which implies

align*[align* omitted — 504 chars of source]

Due to $\varpi_nn \alpha_n^2=o(1)$, then $\sup_{{\mathbf t} \in \mathcal{D}_{1}}\{|{\mathbf t}|_2^2\cdot O_{\rm p}(\varpi_n)\}=o_{\rm p}(1)$. Notice that $|e^{x}-1|\leq |x|e^{x}$ for any $x \in {\mathbb{R}}$. Then $\sup_{{\mathbf t} \in \mathcal{D}_1}|\exp\{ |{\mathbf t}|_2^2\cdot O_{\rm p}(\varpi_n)\}-1| = o_{\rm p}(1)$, which implies that

align*[align* omitted — 592 chars of source]

Therefore, ${\rm I}(1) = o_{\rm p}(1)$.

For ${\rm I}(4)$, due to $\mathcal{D}_4 \cap \mathcal{T}_n = \emptyset$, we have $ {\rm I}(4) = \pi_{0}(\hat{\boldsymbol{\theta}}_n) \int _{\mathcal{D}_4} \exp( -{\mathbf t}^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf H}}_{ \mathcal{R}_n} {\mathbf t}/2 ) \,{\rm d}{\mathbf t} $. Since $\boldsymbol{\theta}_0$ is an interior point of $\boldsymbol{\Theta}$, there exists a constant $\iota>0$ such that $\boldsymbol{\Theta} \supset \mathcal{B}_2(\boldsymbol{\theta}_0,\iota) := \{\boldsymbol{\theta} \in{\mathbb{R}}^p:\,|\boldsymbol{\theta}-\boldsymbol{\theta}_0|_2 \leq \iota \}$, which implies $\mathcal{D}_4 = \mathcal{T}_n^{\rm c} \subset \mathcal{T}_n^{*,{\rm c}}$ with $\mathcal{T}_n^* = \{{\mathbf t} \in {\mathbb{R}}^p: {\mathbf t} = n^{1/2}(\boldsymbol{\theta}-\hat\boldsymbol{\theta}_n), \boldsymbol{\theta} \in \mathcal{B}_2(\boldsymbol{\theta}_0,\iota)\}$. By Proposition (ref), it holds w.p.a.1 that $n^{-1/2}|{\mathbf t}|_2 \geq |n^{-1/2}{\mathbf t} + \hat\boldsymbol{\theta}_n - \boldsymbol{\theta}_0|_2 - |\hat\boldsymbol{\theta}_n - \boldsymbol{\theta}_0|_2 \geq \iota/2$ for any ${\mathbf t} \in \mathcal{D}_4$. Together with Proposition 1.1 of Hsu2012, we have w.p.a.1 that

align*[align* omitted — 347 chars of source]

Therefore, ${\rm I}(4) = o_{\rm p}(1)$. $\Box$

Proof of Proposition (ref)

Proof of part {(i)} of Proposition (ref)

Recall $\mathcal{R}(\boldsymbol{\theta})={\rm supp}\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})\}$ and $\mathcal{C}_1=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2\leq\alpha_n\}$ with $\alpha_n=n^{-1/2}(\log r)^{1/2}$. To prove part {\rm (i)} of Proposition (ref), we need the following lemmas whose proofs are given in Sections (ref), (ref) and (ref), respectively.

lemmaLet $c \in (\tilde{c},1)$ be some constant with $\tilde{c}$ given in Condition {\rm (ref)(a)}. Under the conditions of Lemma {\rm (ref)}, it holds w.p.a.1 that the global maximizer $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})=\{\hat{\lambda}_{1}(\boldsymbol{\theta}),\ldots,\hat{\lambda}_{r}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$ for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ w.r.t $\boldsymbol{\lambda}$ satisfies the results: ${\rm (i)}$ $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n)$, ${\rm (ii)}$ $\mathcal{R}(\boldsymbol{\theta}) \subset \mathcal{M}_{\boldsymbol{\theta}}(c)$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$, and ${\rm (iii)}$ ${\mbox{\rm sgn}}\{\hat{\lambda}_j(\boldsymbol{\theta})\}={\mbox{\rm sgn}} \{\bar{g}_j(\boldsymbol{\theta})\}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ and $j \in \mathcal{M}_{\boldsymbol{\theta}}(c)$ with $\hat{\lambda}_j(\boldsymbol{\theta}) \neq 0$.
lemmaUnder the conditions of Lemma {\rm (ref)} and Condition {\rm(ref)(a)}, it holds w.p.a.1 that the global maximizer $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})$ for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ w.r.t $\boldsymbol{\lambda}$ is continuously differentiable in $\boldsymbol{\theta}\in\mathcal{C}_1$ with $[\nabla_{\boldsymbol{\theta}} \hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})]_{{\mathcal{R}(\boldsymbol{\theta})}^{\rm c},[p]}={\mathbf 0}$ and \begin{align*} [\nabla_{\boldsymbol{\theta}} \hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})]_{{\mathcal{R}(\boldsymbol{\theta})},[p]} =&\,\bigg(\frac{1}{n} \sum _{i=1}^{n}\frac{{\mathbf g}_{i,{\mathcal{R}(\boldsymbol{\theta})}}(\boldsymbol{\theta})^{\otimes2}}{\{1+\hat{\boldsymbol{\lambda}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} }{\mathbf g}_{i,{\mathcal{R}(\boldsymbol{\theta})}}(\boldsymbol{\theta})\}^2} +\nu {\rm diag}[\rho”\{|\tilde\lambda_{1}(\boldsymbol{\theta})|;\nu \},\ldots,\rho”\{|\tilde\lambda_{|{\mathcal{R}(\boldsymbol{\theta})}|}(\boldsymbol{\theta})|;\nu \}]\bigg)^{-1}\\ \qquad &\times \bigg\{\frac{1}{n} \sum_{i=1}^{n} \frac{[\nabla_{\boldsymbol{\theta}} {\mathbf g}_{i}(\boldsymbol{\theta})]_{{\mathcal{R}(\boldsymbol{\theta})},[p]}}{1+\hat{\boldsymbol{\lambda}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} }{\mathbf g}_{i,{\mathcal{R}(\boldsymbol{\theta})}}(\boldsymbol{\theta})}-\frac{1}{n} \sum_{i=1}^{n} \frac{{\mathbf g}_{i,{\mathcal{R}(\boldsymbol{\theta})}}(\boldsymbol{\theta})\hat{\boldsymbol{\lambda}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } [\nabla_{\boldsymbol{\theta}} {\mathbf g}_{i}(\boldsymbol{\theta})]_{{\mathcal{R}(\boldsymbol{\theta})},[p]}}{\{1+\hat{\boldsymbol{\lambda}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} }{\mathbf g}_{i,{\mathcal{R}(\boldsymbol{\theta})}}(\boldsymbol{\theta})\}^2} \bigg\} \,, \end{align*} where $\hat\boldsymbol{\lambda}_{{\mathcal{R}(\boldsymbol{\theta})}}(\boldsymbol{\theta})=\{\tilde{\lambda}_1(\boldsymbol{\theta}),\ldots,\tilde{\lambda}_{|{\mathcal{R}(\boldsymbol{\theta})}|}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$.
lemmaUnder the conditions of Lemma {\rm (ref)} and Condition {\rm(ref)(a)}, it holds that $\mathcal{R}(\boldsymbol{\theta})={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1.

Notice that

align*[align* omitted — 1,004 chars of source]

for any $\boldsymbol{\theta} \in \mathcal{C}_1$. Due to $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})=\arg \max _{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})} f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$, then $\partial f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta}) ; \boldsymbol{\theta}\} / \partial \boldsymbol{\lambda}_{\mathcal{R}(\boldsymbol{\theta})} = {\mathbf 0}$. By Lemma (ref), we have $[\nabla_{\boldsymbol{\theta}} \hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})]_{{\mathcal{R}(\boldsymbol{\theta})}^{\rm c},[p]} = {\mathbf 0}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. Thus, it holds w.p.a.1 that

align*[align* omitted — 449 chars of source]

for any $\boldsymbol{\theta} \in \mathcal{C}_1$. By Lemma (ref),

align[align omitted — 2,373 chars of source]

for any $\boldsymbol{\theta} \in \mathcal{C}_1$. Lemma (ref) specifies the leading term of ${\mathbf t}^{\mathrm{\scriptscriptstyle \top} }[\nabla_{\boldsymbol{\theta}}^2 f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}]{\mathbf t}$ for ${\mathbf t}\in\mathbb{R}^p$, whose proof is given in Section (ref).

lemmaLet $P_{\nu}(\cdot) \in \mathscr{P}$ be convex and assume $\rho(t;\nu)=\nu^{-1}P_{\nu}(t)$ has bounded second-order derivative w.r.t $t$ around $0$, where $\mathscr{P}$ is defined in {\rm (ref)}. Under the conditions of Lemma {\rm (ref)} and Conditions {\rm (ref)} and {\rm(ref)(a)}, then $${\mathbf t}^{\mathrm{\scriptscriptstyle \top} }\big[\nabla_{\boldsymbol{\theta}}^2 f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}\big]{\mathbf t}={\mathbf t}^{\mathrm{\scriptscriptstyle \top} } \big\{\widehat{\boldsymbol{\Gamma}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf V}}^{-1/2}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})\big\}^{\otimes2} {\mathbf t} + |{\mathbf t}|_2^2 \cdot \{O_{\rm p}(\ell_n^{3/2} \alpha_n) + O_{\rm p}(\ell_nn^{1/\gamma}\alpha_n) + O_{\rm p}(\nu) \}$$ holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$ and ${\mathbf t} \in \mathbb{R}^p$.

Let ${\mathbf t} =\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$. Since $\hat{\boldsymbol{\theta}}_n=\arg\min_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$, then $\nabla_{\boldsymbol{\theta}}f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}|_{\boldsymbol{\theta}=\hat{\boldsymbol{\theta}}_n} ={\mathbf 0}$. By the Taylor expansion, it holds that

align[align omitted — 905 chars of source]

for some $\tilde{\boldsymbol{\theta}}$ lying on the jointing line between $\hat{\boldsymbol{\theta}}_n$ and $\boldsymbol{\theta}$. Let $\widehat{{\mathbf H}}_{{\mathcal{R}(\boldsymbol{\theta})}} = \{\widehat{\boldsymbol{\Gamma}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf V}}^{-1/2}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})\}^{\otimes2}$ and recall $\widehat{{\mathbf H}}_{ \mathcal{R}_n}=\{\widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf V}}^{-1/2}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)\}^{\otimes2}$. By Lemma (ref), we know $\mathcal{R}(\boldsymbol{\theta})=\mathcal{R}_n$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. Under Conditions (ref)(b) and (ref), by Lemmas (ref) and (ref), if $\log r=o(n^{1/3})$, $\ell_n\nu^2=o(1)$ and $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$, we have $\|\widehat{{\mathbf V}}_{{ \mathcal{R}_n}} (\hat{\boldsymbol{\theta}}_n)\|_2=O_{\rm p}(1)$, $\|\widehat{{\mathbf V}}^{-1}_{{ \mathcal{R}_n}} (\hat{\boldsymbol{\theta}}_n)\|_2=O_{\rm p}(1)$ and $\|\widehat{\boldsymbol{\Gamma}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)\|_2=O_{\rm p}(1)$. Using the same arguments in the proof of Lemmas (ref) and (ref), if $\log r=o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n\nu^2=o(1)$, we have

align*[align* omitted — 535 chars of source]

which implies $ \sup_{\boldsymbol{\theta} \in \mathcal{C}_1}\|\widehat{{\mathbf H}}_{{\mathcal{R}(\boldsymbol{\theta})}}-\widehat{{\mathbf H}}_{ \mathcal{R}_n}\|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n) $. Together with Lemma (ref), (ref) yields that $ \aleph_n(\boldsymbol{\theta}) - 2^{-1}(\boldsymbol{\theta}-\hat\boldsymbol{\theta}_n)^{\mathrm{\scriptscriptstyle \top} } \widehat{\mathbf H}_{\mathcal{R}_n} (\boldsymbol{\theta}-\hat\boldsymbol{\theta}_n) = |\boldsymbol{\theta}-\hat\boldsymbol{\theta}_n|_2^2 \cdot O_{\rm p}(\varpi_n)$ with $\varpi_n=\max\{\ell_n^{3/2}\alpha_n, \nu, \ell_nn^{1/\gamma}\alpha_n\}$, where $O_{\rm p}(\varpi_n)$ holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$. $\Box$

Proof of part {(ii)} of Proposition (ref)

Recall $\mathcal{R}_n={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$ and $\mathcal{C}_2=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:\alpha_n<|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2\leq\beta_n\}$. Select $\delta_n$ satisfying $\delta_n=o(\ell_n^{-1/2}n^{-1/\gamma})$ and $\ell_n^{1/2}\beta_{n}=o(\delta_n)$, which can be guaranteed by $\ell_n \beta_{n}=o(n^{-1/\gamma})$. For any $\boldsymbol{\theta} \in \mathcal{C}_2$, let $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta}) = \arg\max_{\boldsymbol{\lambda} \in \tilde\Lambda}f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ and $\tilde{\mathcal{R}}(\boldsymbol{\theta})={\rm supp}\{\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta})\}$, where $\tilde\Lambda=\{\boldsymbol{\lambda}\in \mathbb{R}^r:\, |\boldsymbol{\lambda}_{\mathcal{R}_n}|_2\leq \delta_n, \boldsymbol{\lambda}_{\mathcal{R}_n^{\rm c}}={\mathbf 0}\}$. Write $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})=\{\tilde\lambda_1(\boldsymbol{\theta}),\ldots,\tilde\lambda_r(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$. By the Taylor expansion, we have

align[align omitted — 1,646 chars of source]

for some $C\in(0,1)$. By Lemma (ref), we have $|\tilde{\mathcal{R}}(\boldsymbol{\theta})|\leq |\mathcal{R}_n|\leq \ell_n$ w.p.a.1. Notice that $\nu=o(\beta_{n})$. By Proposition (ref), Lemma (ref) and Condition (ref)(b), if $\log r=o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n\beta_{n}=o(n^{-1/\gamma})$, we have $\inf_{\boldsymbol{\theta} \in \mathcal{C}_2} \lambda_{\min}\{\tilde{{\mathbf A}}(\boldsymbol{\theta})\}\geq K_3/2$ w.p.a.1, where $K_3$ is specified in Condition (ref)(b). Recall $\rho(t;\nu)$ is convex w.r.t $t$. Thus $$0 \leq \tilde\boldsymbol{\lambda}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} }\big[\bar{\mathbf g}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta}) - \nu\rho'(0^+){\mbox{\rm sgn}}\{\tilde\boldsymbol{\lambda}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})\}\big] - 4^{-1}K_3|\tilde\boldsymbol{\lambda}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})|_2^2$$ w.p.a.1. Then $|\tilde\boldsymbol{\lambda}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})|_2 \leq 4K_3^{-1} |\bar{\mathbf g}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta}) - \nu\rho'(0^+){\mbox{\rm sgn}}\{\tilde\boldsymbol{\lambda}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})\}|_2$ w.p.a.1. By the Taylor expansion and Condition (ref)(c), $\sup_{\boldsymbol{\theta} \in \mathcal{C}_2} |\bar{\mathbf g}(\boldsymbol{\theta}) - \bar{\mathbf g}(\hat{\boldsymbol{\theta}}_n)|_\infty = O_{\rm p}(\beta_{n})$. Together with the fact $|\tilde{\mathcal{R}}(\boldsymbol{\theta})|\leq \ell_n$ w.p.a.1, we have $\sup_{\boldsymbol{\theta} \in \mathcal{C}_2} |\bar{\mathbf g}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta}) - \bar{\mathbf g}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\hat{\boldsymbol{\theta}}_n)|_2 = O_{\rm p}(\ell_n^{1/2} \beta_{n})$. By the triangle inequality, Lemma (ref) and (ref), since $\alpha_n=o(\nu)$, it holds that $|\bar{\mathbf g}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\hat{\boldsymbol{\theta}}_n)|_2 = O_{\rm p}(\ell_n^{1/2} \nu)$. Due to $\nu=o(\beta_{n})$, we then have $$\sup_{\boldsymbol{\theta} \in \mathcal{C}_2}|\bar{\mathbf g}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta}) - \nu\rho'(0^+){\mbox{\rm sgn}}\{\tilde\boldsymbol{\lambda}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})\}|_2 = O_{\rm p}(\ell_n^{1/2} \beta_{n})\,,$$ which implies $\sup_{\boldsymbol{\theta} \in \mathcal{C}_2}|\tilde\boldsymbol{\lambda}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})|_2 = O_{\rm p}(\ell_n^{1/2} \beta_{n}) = o_{\rm p}(\delta_n)$. Recall $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})=\arg\max_{\boldsymbol{\lambda} \in \tilde\Lambda}f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ and $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta}) \in {\rm int}(\tilde\Lambda)$ for any $\boldsymbol{\theta} \in \mathcal{C}_2$ w.p.a.1. Write $\tilde\boldsymbol{\lambda}_{\mathcal{R}_n}(\boldsymbol{\theta})=\{\dot\lambda_1(\boldsymbol{\theta}),\ldots,\dot\lambda_{|\mathcal{R}_n|}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$. Restricted on $\tilde\Lambda$, by the first-order condition, we have $$ {\mathbf 0}=\frac{1}{n}\sum_{i=1}^{n} \frac{{\mathbf g}_{i,\mathcal{R}_n}(\boldsymbol{\theta})}{1+\tilde{\boldsymbol{\lambda}}_{\mathcal{R}_n}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} }{\mathbf g}_{i,\mathcal{R}_n}(\boldsymbol{\theta})} - \tilde\boldsymbol{\eta}(\boldsymbol{\theta})$$ holds for any $\boldsymbol{\theta} \in \mathcal{C}_2$ w.p.a.1, where $\tilde\boldsymbol{\eta}(\boldsymbol{\theta})=\{\tilde\eta_1(\boldsymbol{\theta}), \ldots,\tilde\eta_{|\mathcal{R}_n|}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$ with $\tilde\eta_j(\boldsymbol{\theta})=\nu\rho'\{|\dot{\lambda}_j(\boldsymbol{\theta})|;\nu\}{\mbox{\rm sgn}}\{\dot{\lambda}_j(\boldsymbol{\theta})\}$ for $\dot{\lambda}_j(\boldsymbol{\theta})\neq0$ and $\tilde\eta_j(\boldsymbol{\theta}) \in [-\nu\rho'(0^+), \nu\rho'(0^+)]$ for $\dot{\lambda}_j(\boldsymbol{\theta})=0$. By the Taylor expansion, we have

align*[align* omitted — 876 chars of source]

for some $\tilde{C}\in(0,1)$, where $\tilde\boldsymbol{\eta}_{*}(\boldsymbol{\theta})\in\mathbb{R}^{|\tilde{\mathcal{R}}(\boldsymbol{\theta})|}$ includes all elements $\tilde\eta_j(\boldsymbol{\theta})$'s in $\tilde\boldsymbol{\eta}(\boldsymbol{\theta})$ such that the associated $\dot \lambda_j(\boldsymbol{\theta}) \neq0$. Hence, $\tilde{\boldsymbol{\lambda}}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})=\check{\mathbf A}^{-1}(\boldsymbol{\theta})\{\bar{\mathbf g}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})-\tilde\boldsymbol{\eta}_{*}(\boldsymbol{\theta})\}$. Using the same arguments in the proof of Lemma (ref), we have $\sup_{\boldsymbol{\theta} \in \mathcal{C}_2}\|\check{\mathbf A}(\boldsymbol{\theta}) - \widehat{{\mathbf V}}_{\tilde{\mathcal{R}}(\boldsymbol{\theta})}(\boldsymbol{\theta})\|_2=O_{\rm p}(\ell_n\beta_nn^{1/\gamma})=o_{\rm p}(1)$. Applying Proposition (ref), Lemma (ref) and Condition (ref)(b), we know $\sup_{\boldsymbol{\theta} \in \mathcal{C}_2}\lambda_{\max}\{\check{\mathbf A}(\boldsymbol{\theta})\}\leq 2K_4$ w.p.a.1 for $K_4$ specified in Condition (ref)(b). For $\tilde{{\mathbf A}}(\boldsymbol{\theta})$ specified in (ref), we can show $\sup_{\boldsymbol{\theta} \in \mathcal{C}_2}\|\tilde{\mathbf A}(\boldsymbol{\theta})- \check{\mathbf A}(\boldsymbol{\theta})\|_2=O_{\rm p}(\ell_n \beta_{n}n^{1/\gamma})$. By (ref) and Condition (ref)(b), we have

align*[align* omitted — 1,025 chars of source]

holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_2$ w.p.a.1, where the term $O_{\rm p}(\ell_n^2\beta_{n}^3 n^{1/\gamma})$ holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_2$. As we have shown in the proof of Proposition (ref), $f_n\{\hat\boldsymbol{\lambda}(\hat{\boldsymbol{\theta}}_n);\hat{\boldsymbol{\theta}}_n\}=O_{\rm p}(\ell_n \alpha_n^2)$. Notice that $f_n\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}\geq f_n\{\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta});\boldsymbol{\theta}\}$ for any $\boldsymbol{\theta} \in \mathcal{C}_2$. If $\max\{\ell_n \alpha_n^2, \ell_n^2\beta_{n}^3 n^{1/\gamma}\} = o(\kappa_n^2)$, then

align*[align* omitted — 338 chars of source]

as $n \rightarrow \infty$. We complete the proof of part {\rm (ii)} of Proposition (ref). $\Box$

Proof of part {(iii)} of Proposition (ref)

Recall $\mathcal{R}_n={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$ and $\mathcal{C}_3=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2>\beta_n\}$. For any $\boldsymbol{\theta} \in \mathcal{C}_3$, we consider $\check\boldsymbol{\lambda}(\boldsymbol{\theta})=\{\check\lambda_1(\boldsymbol{\theta}),\ldots,\check\lambda_r(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }\in \mathbb{R}^r$ with $\check\boldsymbol{\lambda}_{\mathcal{R}_n^{\rm c}}(\boldsymbol{\theta}) ={\mathbf 0}$ and $$\check\boldsymbol{\lambda}_{\mathcal{R}_n}(\boldsymbol{\theta}) = \frac{\xi_n\{\bar{\mathbf g}_{\mathcal{R}_n}(\boldsymbol{\theta})-\bar{\mathbf g}_{\mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)\}}{|\bar{\mathbf g}_{\mathcal{R}_n}(\boldsymbol{\theta})-\bar{\mathbf g}_{\mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)|_2}\,.$$ Due to $\ell_n^{1/2}\xi_n =o(n^{-1/\gamma})$, Condition (ref)(a) yields $\sup_{\boldsymbol{\theta} \in \mathcal{C}_3} \max_{ i \in [n]} |\check\boldsymbol{\lambda}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_i(\boldsymbol{\theta})|=o_{\rm p}(1)$, which implies the event $\bigcap_{\boldsymbol{\theta} \in \mathcal{C}_3} \{\check\boldsymbol{\lambda}(\boldsymbol{\theta}) \in \hat\Lambda_n(\boldsymbol{\theta})\}$ holds w.p.a.1. Then $$ {\mathbb{P}} \bigg[\inf_{\boldsymbol{\theta} \in \mathcal{C}_3} f_n\{\hat\boldsymbol{\lambda}(\boldsymbol{\theta});\boldsymbol{\theta}\} \geq \inf_{\boldsymbol{\theta} \in \mathcal{C}_3} f_n\{\check\boldsymbol{\lambda}(\boldsymbol{\theta});\boldsymbol{\theta}\}\bigg] =1-o(1)\,. $$ Recall $\sup_{\boldsymbol{\theta} \in \mathcal{C}_3} \lambda_{\max}\{\widehat{\mathbf V}_{\mathcal{R}_n}(\boldsymbol{\theta})\}\leq K_8$ w.p.a.1 for $K_8$ specified in Condition (ref)(c). By the Taylor expansion, we have

align*[align* omitted — 1,689 chars of source]

holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_3$ w.p.a.1, where $\check C,c_j\in(0,1)$. By Condition (ref)(c), we have $\inf_{\boldsymbol{\theta} \in \mathcal{C}_3}|\bar{\mathbf g}_{\mathcal{R}_n}(\boldsymbol{\theta}) - \bar{\mathbf g}_{\mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)|_2 = \inf_{\boldsymbol{\theta} \in \mathcal{C}_3} |\{\nabla_{\boldsymbol{\theta}} \bar{\mathbf g}_{\mathcal{R}_n}(\tilde \boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} } (\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n)|_2 \geq K_7^{1/2}\beta_{n}$ w.p.a.1. Since $\alpha_n=o(\nu)$, by Lemma (ref) and (ref), we have $|\bar{\mathbf g}_{\mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)|_2=O_{\rm p}(\ell_n^{1/2}\nu)$. Due to $\ell_n^{1/2}\nu=o(\beta_{n})$ and $\xi_n=o(\beta_{n})$, it holds that $$\inf_{\boldsymbol{\theta} \in \mathcal{C}_3} f_n\{\check\boldsymbol{\lambda}(\boldsymbol{\theta});\boldsymbol{\theta}\} \geq K_7^{1/2}\xi_n\beta_{n} + O_{\rm p}(\xi_n \ell_n^{1/2}\nu) + O_{\rm p}(\xi_n^2) \geq \frac{1}{2}K_7^{1/2} \xi_n\beta_{n}$$ w.p.a.1. Since $f_n\{\hat\boldsymbol{\lambda}(\hat{\boldsymbol{\theta}}_n);\hat{\boldsymbol{\theta}}_n\}=O_{\rm p}(\ell_n \alpha_n^2)=o_{\rm p}(\xi_n\beta_{n})$, then

align*[align* omitted — 467 chars of source]

as $n\rightarrow \infty$. We complete the proof of part {\rm (iii)} of Proposition (ref). $\Box$

Proof of Corollary (ref)

Let ${\mathbb{E}}_{{\mathbf t} \sim \pi^{\dag}_{{\mathbf t}}}({\mathbf t}) = \int_{{\mathbb{R}}^p} {\mathbf t} \pi^{\dag}_{{\mathbf t}}({\mathbf t}\,|\,\mathcal{X}_n) \, {\rm d}{\mathbf t}$ for $\pi^{\dag}_{{\mathbf t}}({\mathbf t}\,|\,\mathcal{X}_n)$ given in (ref). Notice that ${\mathbb{E}}_{{\mathbf t} \sim \pi^{\dag}_{{\mathbf t}}}({\mathbf t}) = n^{1/2}\{{\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta}) - \hat{\boldsymbol{\theta}}_n\}$ and ${\mathbb{E}}_{{\mathbf t} \sim \mathcal{N}({\mathbf 0}, \widehat{{\mathbf H}}_{ \mathcal{R}_n}^{-1})}({\mathbf t}) = {\mathbf 0}$. To prove Corollary (ref), it is equivalent to show $|{\mathbb{E}}_{{\mathbf t} \sim \pi^{\dag}_{{\mathbf t}}}({\mathbf t}) - {\mathbb{E}}_{{\mathbf t} \sim \mathcal{N}({\mathbf 0}, \widehat{{\mathbf H}}_{ \mathcal{R}_n}^{-1})}({\mathbf t})|_\infty = o_{\rm p}(1)$. It follows from the triangle inequality that

align*[align* omitted — 1,735 chars of source]

As shown in Section (ref), we have $C_n^{-1} = O_{\rm p}(1)$. It suffices to show ${\rm III} = o_{\rm p}(1)$ and ${\rm IV}=o_{\rm p}(1)$. Notice that

align*[align* omitted — 281 chars of source]

Recall $\widehat{{\mathbf H}}_{{ \mathcal{R}_n}}=\{\widehat{\boldsymbol{\Gamma}}_{{ \mathcal{R}_n}}(\hat\boldsymbol{\theta}_n)^{\mathrm{\scriptscriptstyle \top} }\widehat{{\mathbf V}}_{{ \mathcal{R}_n}}^{-1/2}(\hat\boldsymbol{\theta}_n)\}^{\otimes2}$. Under Conditions (ref)(b) and (ref), by Proposition (ref), Lemmas (ref) and (ref), if $\log r=o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n \nu^2=o(1)$, we know that the eigenvalues of $\widehat{{\mathbf H}}_{ \mathcal{R}_n}$ are uniformly bounded away from zero and infinity w.p.a.1. Since $\widehat{{\mathbf H}}_{{ \mathcal{R}_n}}$ is a $p \times p$ matrix with fixed $p$, then

align[align omitted — 301 chars of source]

As shown in the proof of Theorem (ref), we have $|C_n - (2\pi)^{p/2} \pi_{0}(\hat{\boldsymbol{\theta}}_n) |\widehat{{\mathbf H}}_{ \mathcal{R}_n}|^{-1/2} | \leq {\rm I} = o_{\rm p}(1)$ for ${\rm I}$ defined in (ref), which implies ${\rm IV} \leq {\rm I} \cdot {\mathbb{E}}_{{\mathbf t} \sim \mathcal{N}({\mathbf 0}, \widehat{{\mathbf H}}_{ \mathcal{R}_n}^{-1})}(|{\mathbf t}|_\infty)= o_{\rm p}(1)$. In the sequel, we will show that ${\rm III} =o_{\rm p}(1)$. Recall $\ell_n \ll \min\{n^{(\gamma-2)/(9\gamma)}(\log r)^{-1/9},n^{1/3}(\log r)^{-1},n^{(\gamma-2)/(2\gamma)}(\log r)^{-3/2}\}$ and $\ell_nn^{-1/2}(\log r)^{1/2} \ll \nu \ll \min\{\ell_n^{-7/2}n^{-1/\gamma},(\log r)^{-1}\}$. For $(\mathcal{D}_1, \mathcal{D}_2, \mathcal{D}_3, \mathcal{D}_4)$ defined as (ref), it holds that ${\rm III} = {\rm III}(1) + {\rm III}(2)+ {\rm III}(3) + {\rm III}(4)$ with

align*[align* omitted — 479 chars of source]

For ${\rm III}(3)$, by the triangle inequality, we have

align*[align* omitted — 484 chars of source]

Since $\boldsymbol{\Theta} \subset {\mathbb{R}}^p$ is a compact set, then

align*[align* omitted — 278 chars of source]

By Proposition (ref)(iii),

align*[align* omitted — 581 chars of source]

w.p.a.1 for any $\beta_n^{-1}\ell_n \alpha_n^2 \ll \xi_n \ll \beta_n$. Since $r\gg n$, we can select suitable $\xi_n$ satisfying $ n \beta_n \xi_n \gg \log n $. Then

align*[align* omitted — 254 chars of source]

Recall that the eigenvalues of $\widehat{{\mathbf H}}_{ \mathcal{R}_n}$ are uniformly bounded away from zero and infinity w.p.a.1. Since $n\beta_n^2 \rightarrow \infty$ and $p$ is fixed, by the Cauchy-Schwarz inequality and Proposition 1.1 of Hsu2012, we have

align*[align* omitted — 649 chars of source]

which implies

align[align omitted — 546 chars of source]

Therefore, ${\rm III}(3) = o_{\rm p} (1)$.

For ${\rm III}(2)$, it holds that

align*[align* omitted — 482 chars of source]

Since $n\alpha_n^2 \rightarrow \infty$, using the same arguments for (ref), we have $$ \pi_{0}(\hat{\boldsymbol{\theta}}_n)\int _{\mathcal{D}_2} |{\mathbf t}|_\infty \exp( -{\mathbf t}^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf H}}_{ \mathcal{R}_n} {\mathbf t}/2 ) \,{\rm d}{\mathbf t} = o_{\rm p}(1) \,. $$ Due to $\log n \ll n\kappa_{n}^2$, by Proposition (ref)(ii), it then holds w.p.a.1 that

align*[align* omitted — 594 chars of source]

Therefore, ${\rm III}(2)=o_{\rm p}(1)$.

For ${\rm III}(1)$, by Proposition (ref)(i), we have

align*[align* omitted — 739 chars of source]

where $\varpi_n=\max \{\ell_n^{3/2} \alpha_n, \nu, \ell_nn^{1/\gamma}\alpha_n\}$. Under Condition (ref), we know $ \sup_{{\mathbf t} \in \mathcal{D}_{1}} |\pi_{0}(\hat{\boldsymbol{\theta}}_n+n^{-1/2}{\mathbf t})-\pi_{0}(\hat{\boldsymbol{\theta}}_n)| =o_{\rm p}(1)$. By (ref), we have

align*[align* omitted — 655 chars of source]

Due to $\varpi_nn \alpha_n^2=o(1)$, then $\sup_{{\mathbf t} \in \mathcal{D}_{1}}\{|{\mathbf t}|_2^2\cdot O_{\rm p}(\varpi_n)\}=o_{\rm p}(1)$. Notice that $|e^{x}-1|\leq |x|e^{x}$ for any $x \in {\mathbb{R}}$. Then $\sup_{{\mathbf t} \in \mathcal{D}_1}|\exp\{ |{\mathbf t}|_2^2\cdot O_{\rm p}(\varpi_n)\}-1| = o_{\rm p}(1)$, which implies that

align*[align* omitted — 634 chars of source]

Therefore, ${\rm III}(1)= o_{\rm p}(1)$.

For ${\rm III}(4)$, due to $\mathcal{D}_4 \cap \mathcal{T}_n = \emptyset$, we have $ {\rm III}(4) = \pi_{0}(\hat{\boldsymbol{\theta}}_n) \int _{\mathcal{D}_4} |{\mathbf t}|_\infty \exp( -{\mathbf t}^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf H}}_{ \mathcal{R}_n} {\mathbf t}/2 ) \,{\rm d}{\mathbf t} $. As shown in Section (ref), it holds w.p.a.1 that $n^{-1/2}|{\mathbf t}|_2 \geq \iota/2$ for any ${\mathbf t} \in \mathcal{D}_4$. Using the same arguments for (ref), we have ${\mathbb{E}}_{{\mathbf t} \sim \mathcal{N}({\mathbf 0}, \widehat{{\mathbf H}}_{ \mathcal{R}_n}^{-1})}\{|{\mathbf t}|_\infty I({\mathbf t} \in \mathcal{D}_4)\} = o_{\rm p}(1)$, which implies

align*[align* omitted — 520 chars of source]

Therefore, ${\rm III}(4) = o_{\rm p}(1)$. $\Box$

Proof of Theorem (ref)

To prove Theorem (ref), we first introduce the following two concepts.

definition[$\Pi$-irreducibility] For a distribution $\Pi$ on $\mathcal{D}$, a Markov chain is called $\Pi$-irreducible if for each $A \in \mathscr{B}(\mathcal{D})$ with $\Pi(A)>0$ and ${\mathbf x} \in \mathcal{D}$, there exists $k \in \mathbb{N}$ such that $\Psi^k({\mathbf x}, A)>0$, where $\mathscr{B}(\mathcal{D})$ is the Borel $\sigma$-algebra on $\mathcal{D}$, and $\Psi^k$ is the $k$-step transition probability defined recursively as $\Psi^{k}({\mathbf x}, {\rm d}{\mathbf y}) = \int_{{\mathbf z} \in \mathcal{D} } \Psi^{k-1}({\mathbf x}, {\rm d}{\mathbf z})\Psi({\mathbf z}, {\rm d}{\mathbf y})$.
definition[Aperiodic] A Markov chain with stationary distribution $\Pi$ on $\mathcal{D}$ and transition probability $\Psi(\cdot\,, \cdot)$ is aperiodic if there do not exist $T \geq 2$ and disjoint subsets $\mathcal{D}_1, \ldots, \mathcal{D}_T \subset \mathcal{D}$ with each $\Pi(\mathcal{D}_i)>0$ such that (i) $\Psi({\mathbf x}\,, \mathcal{D}_{i+1}) = 1$ for all ${\mathbf x} \in \mathcal{D}_{i}$ and $i= 1, \ldots, T-1$, and (ii) $\Psi({\mathbf x}\,, \mathcal{D}_{1}) = 1$ for all ${\mathbf x} \in \mathcal{D}_{T}$.

Denote by $\Psi(\boldsymbol{\theta}, \cdot)$ the transition probability of the Markov chain determined by Algorithm (ref) at $\boldsymbol{\theta} \in \boldsymbol{\Theta}$. For given $\boldsymbol{\theta} \in \boldsymbol{\Theta}$, $\alpha_{\boldsymbol{\theta}}(\boldsymbol{\vartheta}) = \min \{1, R_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})\}$ is the acceptance probability at $\boldsymbol{\vartheta} \in {\mathbb{R}}^p$, where

align*[align* omitted — 837 chars of source]

Then the transition probability of the associated Markov chain at $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ has a probability mass $\psi_{\boldsymbol{\theta}} = 1- \int_{\boldsymbol{\Theta}} \phi(\boldsymbol{\vartheta}\,|\,\boldsymbol{\theta}) \alpha_{\boldsymbol{\theta}}(\boldsymbol{\vartheta}) \,{\rm d}\boldsymbol{\vartheta}$. Define $\psi(\boldsymbol{\theta},\boldsymbol{\vartheta}) = \phi(\boldsymbol{\vartheta}\,|\,\boldsymbol{\theta}) \alpha_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})$ for any $\boldsymbol{\theta},\boldsymbol{\vartheta} \in \boldsymbol{\Theta}$. We have $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) \psi(\boldsymbol{\theta},\boldsymbol{\vartheta}) = \pi^\dag(\boldsymbol{\vartheta}\,|\,\mathcal{X}_n) \psi(\boldsymbol{\vartheta},\boldsymbol{\theta})$ for any $\boldsymbol{\theta}, \boldsymbol{\vartheta} \in \boldsymbol{\Theta}$. Since the Markov chain determined by Algorithm {\rm (ref)} always stays in $\boldsymbol{\Theta}$, its transition probability $\Psi(\cdot\,, \cdot): \boldsymbol{\Theta} \times \mathscr{B}(\boldsymbol{\Theta}) \mapsto {\mathbb{R}}_+$ satisfies $ \Psi(\boldsymbol{\theta}, {\rm d}\boldsymbol{\vartheta}) = \psi_{\boldsymbol{\theta}} \delta_{\boldsymbol{\theta}}({\rm d}\boldsymbol{\vartheta}) + \psi(\boldsymbol{\theta},\boldsymbol{\vartheta}) \, {\rm d}\boldsymbol{\vartheta} $, where $\mathscr{B}(\boldsymbol{\Theta})$ is the Borel $\sigma$-algebra on $\boldsymbol{\Theta}$, and $\delta_{\boldsymbol{\theta}}$ is the Dirac-delta function at $\boldsymbol{\theta}$ with $\delta_{\boldsymbol{\theta}}(A)=I(\boldsymbol{\theta} \in A)$. For any $A,B \in \mathscr{B}(\boldsymbol{\Theta})$, we have

align*[align* omitted — 991 chars of source]

Therefore, $\Pi^\dag_n(A) = \int_A \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)\,{\rm d}\boldsymbol{\theta} = \int_A \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) \Psi(\boldsymbol{\theta}, \boldsymbol{\Theta}) \,{\rm d}\boldsymbol{\theta} = \int_{\boldsymbol{\Theta}} \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) \Psi(\boldsymbol{\theta}, A) \,{\rm d}\boldsymbol{\theta}$ for any $A \in \mathscr{B}(\boldsymbol{\Theta})$, which implies that $\Pi^\dag_n$ is the stationary distribution of such Markov chain with transition probability $\Psi(\cdot\,,\cdot)$.

Denote by $\mathbb{L}(\cdot)$ the Lebesgue measure on ${\mathbb{R}}^p$. For any $A\in\mathscr{B}(\boldsymbol{\Theta})$ with $\Pi_n^\dag(A)>0$, due to $\Pi_n^\dag(A)=\int_A \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)\,{\rm d}\boldsymbol{\theta}$, we know $\mathbb{L}(A)>0$. Recall that $\boldsymbol{\Theta} \subset {\mathbb{R}}^p$ is a compact set. Since $\phi(\boldsymbol{\vartheta}\,|\,\boldsymbol{\theta})$ is positive and continuous on $(\boldsymbol{\theta}, \boldsymbol{\vartheta}) \in \boldsymbol{\Theta} \times \boldsymbol{\Theta}$, there exists a constant $C>0$ such that $\inf_{\boldsymbol{\theta}, \boldsymbol{\vartheta} \in \boldsymbol{\Theta}}\phi(\boldsymbol{\theta}\,|\,\boldsymbol{\vartheta}) > C$. On one hand, for any $A\in\mathscr{B}(\boldsymbol{\Theta})$ and $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ such that $\Pi_n^\dag(A)>0$ and $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) = 0$, we have $\Psi(\boldsymbol{\theta}, A) \geq \int_A \psi(\boldsymbol{\theta},\boldsymbol{\vartheta}) \, {\rm d}\boldsymbol{\vartheta} = \int_A \phi(\boldsymbol{\vartheta}\,|\,\boldsymbol{\theta}) \, {\rm d}\boldsymbol{\vartheta} >0$. On the other hand, for any $A\in\mathscr{B}(\boldsymbol{\Theta})$ and $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ such that $\Pi_n^\dag(A)>0$ and $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)>0$, we have

align*[align* omitted — 1,878 chars of source]

Since $\mathbb{L}(\{\boldsymbol{\vartheta} \in A: \pi^\dag(\boldsymbol{\vartheta}\,|\,\mathcal{X}_n) \geq \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)\})$ and $\int_{\boldsymbol{\vartheta} \in A:\, \pi^\dag(\boldsymbol{\vartheta}\,|\,\mathcal{X}_n) < \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)} \pi^\dag(\boldsymbol{\vartheta}\,|\,\mathcal{X}_n) \, {\rm d}\boldsymbol{\vartheta}$ cannot be zero simultaneously for any $A\in\mathscr{B}(\boldsymbol{\Theta})$ with $\Pi_n^\dag(A)>0$, then $\Psi(\boldsymbol{\theta}, A) >0$ for any $A\in\mathscr{B}(\boldsymbol{\Theta})$ and $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ such that $\Pi_n^\dag(A)>0$ and $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) > 0$. Therefore, it holds that $\Psi(\boldsymbol{\theta}, A) >0$ for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ and $A\in\mathscr{B}(\boldsymbol{\Theta})$ with $\Pi_n^\dag(A)>0$. By Definition (ref), the Markov chain with transition probability $\Psi(\cdot\,,\cdot)$ is $\Pi^\dag_n$-irreducible. Furthermore, by Definition (ref), we know the Markov chain $\{\boldsymbol{\theta}^k\}_{k\geq1}$ with transition probability $\Psi(\cdot\,,\cdot)$ and initial point $\boldsymbol{\theta}^0$ is aperiodic. Notice that $\mathscr{B}(\boldsymbol{\Theta})$ is a countably generated $\sigma$-algebra. Denote by $\mathcal{T}^k_{\boldsymbol{\theta}^{0}}(\cdot)$ the measure which admits the distribution of such Markov chain at $k$-th step with initial point $\boldsymbol{\theta}^{0}$. Conditional on $\mathcal{X}_n$, for any $\boldsymbol{\theta}^0 \in \boldsymbol{\Theta}$ such that $\pi^\dag(\boldsymbol{\theta}^0\,|\,\mathcal{X}_n)>0$, by Theorem 4 of RobertsRosenthal2004, we have $\mathcal{D}_{\mathrm{\scriptscriptstyle TV} }(\mathcal{T}^k_{\boldsymbol{\theta}^{0}} ,\, \Pi^{\dag}_n) \rightarrow 0$ as $k \rightarrow \infty$. Furthermore, notice that $\boldsymbol{\Theta} \subset {\mathbb{R}}^p$ is a compact set with fixed $p$. Conditional on $\mathcal{X}_n$, for any $\boldsymbol{\theta}^0 \in \boldsymbol{\Theta}$ such that $\pi^\dag(\boldsymbol{\theta}^0\,|\,\mathcal{X}_n)>0$, it follows from Fact 5 of RobertsRosenthal2004 that $|K^{-1}\sum_{k=1}^K \boldsymbol{\theta}^k - {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta})|_\infty \leq |K^{-1}\sum_{k=1}^K \boldsymbol{\theta}^k - {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta})|_1 \rightarrow 0 $ almost surely as $K \rightarrow \infty$, where $\{\boldsymbol{\theta}^k\}_{k\geq 1}$ are generated via Algorithm (ref) with the initial $\boldsymbol{\theta}^0$ and ${\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta})$ is defined in (ref). We complete the proof of Theorem (ref). $\Box$

Proof of Theorem (ref)

For the function ${\mathbf h}: {\mathbb{R}}^p \mapsto {\mathbb{R}}^s$ involved in Algorithm {\rm (ref)}, let $\boldsymbol{\zeta}^* = {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}\{{\mathbf h}(\boldsymbol{\theta})\}$. Define

align[align omitted — 290 chars of source]

with $S_K=N_1+\cdots+N_K$, where $\{\boldsymbol{\theta}^1_1, \ldots, \boldsymbol{\theta}^1_{N_1}, \ldots,\boldsymbol{\theta}^K_1,\ldots, \boldsymbol{\theta}^K_{N_K}\}$ are generated via Algorithm {\rm (ref)}. To construct Theorem (ref), we need the following two lemmas whose proofs are given in Sections (ref) and (ref), respectively.

lemmaAssume that the conditions of Theorem {\rm (ref)} hold. Conditional on $\mathcal{X}_n$, $|\hat{\boldsymbol{\zeta}}_{k} - \boldsymbol{\zeta}^*|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$, where $\hat{\boldsymbol{\zeta}}_{k}$ is defined in Algorithm {\rm (ref)}.
lemmaAssume that the conditions of Theorem {\rm (ref)} hold. Conditional on $\mathcal{X}_n$, $|\widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(\boldsymbol{\theta}) - {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta})|_\infty \rightarrow 0$ almost surely as $K \rightarrow \infty$, where $\widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(\boldsymbol{\theta})$ is defined in (ref).

Denote by ${\mathbb{P}}_{\mathcal{X}_n}(\cdot)$ the conditional probability given $\mathcal{X}_n$. For some sufficiently large $M>0$, by Lemma (ref), we have that for any $\epsilon>0$, there exists a sufficiently large integer $k_\epsilon$ such that $\mathbb{P}_{\mathcal{X}_n}(\mathcal{A}) \leq \epsilon$ with $\mathcal{A} = \bigcup_{t=k_\epsilon}^\infty \{|\hat{\boldsymbol{\zeta}}_{t} - \boldsymbol{\zeta}^*|_\infty > M\}$. Define a compact set $\mathcal{B}=\{\boldsymbol{\zeta} \in {\mathbb{R}}^s: |\boldsymbol{\zeta} - \boldsymbol{\zeta}^*|_\infty \leq M \} $. Recall $\boldsymbol{\Theta} \subset {\mathbb{R}}^p$ is a compact set with fixed $p$. Since $\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})$ is positive and continuous on $(\boldsymbol{\theta}, \boldsymbol{\zeta}) \in \boldsymbol{\Theta} \times {\mathbb{R}}^s$, then $\inf_{\boldsymbol{\theta} \in \boldsymbol{\Theta}, \boldsymbol{\zeta} \in \mathcal{B}} \varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta}) \geq C_M$ for some constant $C_M>0$ and $\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})$ is uniformly continuous on $(\boldsymbol{\theta}\,; \boldsymbol{\zeta}) \in \boldsymbol{\Theta} \times \mathcal{B}$. For any $\varepsilon>0$, there exists $\delta(\varepsilon)>0$ such that $|\varphi(\boldsymbol{\theta}_1\,; \boldsymbol{\zeta}_1) - \varphi(\boldsymbol{\theta}_2\,; \boldsymbol{\zeta}_2)| < C_M (2+2C_M)^{-1}\varepsilon$ for any $(\boldsymbol{\theta}_1, \boldsymbol{\zeta}_1), (\boldsymbol{\theta}_2, \boldsymbol{\zeta}_2) \in \boldsymbol{\Theta} \times \mathcal{B}$ satisfying $|\boldsymbol{\theta}_1 - \boldsymbol{\theta}_2|_\infty \leq \delta(\varepsilon)$ and $|\boldsymbol{\zeta}_1 - \boldsymbol{\zeta}_2|_\infty \leq \delta(\varepsilon)$. For any $K \geq k_\epsilon$, it holds that $$ \inf_{\boldsymbol{\theta} \in \boldsymbol{\Theta}}\frac{1}{S_K} \sum_{k=1}^K N_k \varphi(\boldsymbol{\theta}\,;\boldsymbol{\zeta}) \geq \inf_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} \frac{1}{S_K} \sum_{k=k_\epsilon}^K N_k \varphi(\boldsymbol{\theta}\,;\boldsymbol{\zeta}) \geq \frac{C_M}{S_K}\bigg(S_K - \sum_{k=1}^{k_\epsilon-1} N_k\bigg) $$ for any $\boldsymbol{\zeta} \in \mathcal{B}$. Notice that $S_K \to \infty$ as $K \to \infty$. Given $k_\epsilon$, there exists a sufficiently large integer $K^*$ such that $S_K / (S_K - \sum_{k=1}^{k_\epsilon-1} N_k) \leq 1+C_M$ for all $K > K^*$. Recall $N_k\leq N_{k+1}$ for any $k\geq 1$ and $N_k\rightarrow\infty$ as $k\rightarrow\infty$. Since $\sup_{\boldsymbol{\theta}\in\boldsymbol{\Theta},\boldsymbol{\zeta}\in\mathbb{R}^s}\varphi(\boldsymbol{\theta}\,;\boldsymbol{\zeta})<\infty$, we have \[ \frac{1}{S_t}\sum_{k=1}^{\lfloor\sqrt{t}\rfloor}N_k\sup_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}|\varphi(\boldsymbol{\theta}\,;\boldsymbol{\zeta}^*)-\varphi(\boldsymbol{\theta}\,;\hat{\boldsymbol{\zeta}}_k)|\lesssim \frac{1}{S_t}\sum_{k=1}^{\lfloor\sqrt{t}\rfloor}N_k\rightarrow0 \] as $t\rightarrow\infty$. For given $\varepsilon>0$, there exists some sufficiently large integer $\tilde{K}$ such that \[ \frac{1}{S_t}\sum_{k=1}^{\lfloor\sqrt{t}\rfloor}N_k\sup_{\boldsymbol{\theta}\in\boldsymbol{\Theta}}|\varphi(\boldsymbol{\theta}\,;\boldsymbol{\zeta}^*)-\varphi(\boldsymbol{\theta}\,;\hat{\boldsymbol{\zeta}}_k)|\leq \frac{C_M\varepsilon}{2(1+C_M)} \] for any $t\geq \tilde{K}$. It holds that

align*[align* omitted — 1,199 chars of source]

for any $t > \max(K^*, k_\epsilon,\tilde{K})$. We then have

align*[align* omitted — 658 chars of source]

where the second step is due to the fact that conditional on $\mathcal{X}_n$ we have $|\hat{\boldsymbol{\zeta}}_{k} - \boldsymbol{\zeta}^*|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$. Letting $\epsilon \to 0$, we know that conditional on $\mathcal{X}_n$,

align[align omitted — 264 chars of source]

almost surely as $K \rightarrow \infty$.

Define $$ \widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(|\boldsymbol{\theta}|_\infty) = \frac{1}{S_K} \sum_{k=1}^{K} \sum_{i=1}^{N_k} \frac{\pi^\dag(\boldsymbol{\theta}^k_i\,|\,\mathcal{X}_n)}{\varphi(\boldsymbol{\theta}^k_i\,;\boldsymbol{\zeta}^*)}|\boldsymbol{\theta}^k_i|_\infty ~~\textrm{and}~~ {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(|\boldsymbol{\theta}|_\infty) = \int_{{\mathbb{R}}^p} |\boldsymbol{\theta}|_\infty \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) \, {\rm d}\boldsymbol{\theta} \,, $$ where $\{\boldsymbol{\theta}^1_1, \ldots, \boldsymbol{\theta}^1_{N_1}, \ldots,\boldsymbol{\theta}^K_1,\ldots, \boldsymbol{\theta}^K_{N_K}\}$ are generated via Algorithm {\rm (ref)}. For $\widehat{{\mathbb{E}}}_{\pi^\dag,K}(\boldsymbol{\theta})$ defined in (ref) and $\widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(\boldsymbol{\theta})$ defined in (ref), since $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) |\boldsymbol{\theta}|_\infty / \varphi(\boldsymbol{\theta}\,;\boldsymbol{\zeta}^*) = 0$ for any $\boldsymbol{\theta} \notin \boldsymbol{\Theta}$, we then have

align*[align* omitted — 1,497 chars of source]

Using the same arguments for the proof of Lemma (ref) in Section (ref), it holds that conditional on $\mathcal{X}_n$ we have $|\widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(|\boldsymbol{\theta}|_\infty) - {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(|\boldsymbol{\theta}|_\infty)| \rightarrow 0$ almost surely as $K \rightarrow \infty$. Notice that $\boldsymbol{\Theta} \subset {\mathbb{R}}^p$ is a compact set with fixed $p$. Then ${\mathbb{E}}_{\boldsymbol{\theta} \sim\pi^\dag}(|\boldsymbol{\theta}|_\infty) < \infty$. Together with (ref), it holds that conditional on $\mathcal{X}_n$ we have $|\widehat{{\mathbb{E}}}_{\pi^\dag,K}(\boldsymbol{\theta})- \widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(\boldsymbol{\theta}) |_\infty \rightarrow 0$ almost surely as $K \rightarrow \infty$. By the triangle inequality and Lemma (ref), conditional on $\mathcal{X}_n$, $|\widehat{{\mathbb{E}}}_{\pi^\dag,K}(\boldsymbol{\theta})- {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta})|_\infty \leq |\widehat{{\mathbb{E}}}_{\pi^\dag,K}(\boldsymbol{\theta})- \widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(\boldsymbol{\theta})|_\infty + |\widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(\boldsymbol{\theta})- {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta})|_\infty \rightarrow 0$ almost surely as $K \rightarrow \infty$. We complete the proof of Theorem (ref). $\Box$

Proofs of auxiliary lemmas

Proof of Lemma (ref)

The proof is almost identical to that of Lemma 1 in Changetal2018. Recall $p$ is fixed. We only need to replace $\{\varrho_n,\omega_n,\xi_n,b_n^{1/(2\beta)},s\}$ appeared in the proof of Lemma 1 in Changetal2018 by $(1,1,1,\varphi_n,p)$ and all the arguments still hold. $\Box$

Proof of Lemma (ref)

Due to the convexity of $P_{\nu}(\cdot)$, $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ is concave w.r.t $\boldsymbol{\lambda}$. We only need to show that there exists a local maximizer $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta}_0)$ satisfying the results stated in the lemma. Recall $\mathcal{M}_{\boldsymbol{\theta}_0}^*=\{j\in[r]: |\bar{g}_j(\boldsymbol{\theta}_0)| \geq C_*\nu \rho'(0^{+}) \}$ for some $C_* \in (0,1)$, and $\mathbb{P}(\max _{\boldsymbol{\theta} \in \boldsymbol{\Theta}:\,|\boldsymbol{\theta}-\boldsymbol{\theta}_0|_2 \leq c_n}|\mathcal{M}_{\boldsymbol{\theta}}^*|\leq\ell_n)\rightarrow1$ for some $c_n\rightarrow 0$ satisfying $\nu c_n^{-1} \rightarrow 0$. For any given $c\in(C_*,1)$, write $\mathcal{M}_{\boldsymbol{\theta}_0}:=\mathcal{M}_{\boldsymbol{\theta}_0}(c)=\{ j\in[r]:|\bar{g}_j(\boldsymbol{\theta}_0)| \geq c\nu\rho'(0^{+}) \}$. Then $\ell_n \geq |\mathcal{M}_{\boldsymbol{\theta}_0}^*| \geq |\mathcal{M}_{\boldsymbol{\theta}_0}|$ w.p.a.1. To prove Lemma (ref), we establish its validity separately with Case 1: $\mathcal{M}_{\boldsymbol{\theta}_0} \neq \emptyset$ and Case 2: $\mathcal{M}_{\boldsymbol{\theta}_0} = \emptyset$.

Case 1: $ \mathcal{M}_{\boldsymbol{\theta}_0} \neq \emptyset$

Restricted on $\mathcal{M}_{\boldsymbol{\theta}_0}$, we select $\delta_n$ satisfying $\delta_n=o(\ell_n^{-1/2} n^{-1/\gamma})$ and $\ell_n^{1/2} \alpha_n=o(\delta_n) $, which can be guaranteed by $\ell_n\alpha_n=o(n^{-1/\gamma})$. Let $\Lambda_{0}=\{\boldsymbol{\lambda}\in {\mathbb{R}}^{r}:|\boldsymbol{\lambda}_{\mathcal{M}_{\boldsymbol{\theta}_0}}|_2\leq \delta_n {\rm \ and \ }\boldsymbol{\lambda}_{\mathcal{M}^{\rm c}_{\boldsymbol{\theta}_0}}={\mathbf 0} \}$ and $\tilde{\boldsymbol{\lambda}}_{0}=\arg \max _{\boldsymbol{\lambda} \in \Lambda_{0}} f_n(\boldsymbol{\lambda}; \boldsymbol{\theta}_0)$. By Condition (ref)(a), we have $\max_{i\in[n],j\in[r]}|g_{i,j} (\boldsymbol{\theta}_0)| = O_{\rm p}(n^{1/\gamma})$, which implies $\max_{i\in[n]} |{\mathbf g}_{i,\mathcal{M}_{\boldsymbol{\theta}_0}} (\boldsymbol{\theta}_0)|_2 = O_{\rm p}(\ell_n^{1/2} n^{1/\gamma})$. Then $\max_{i\in[n]} |\tilde{\boldsymbol{\lambda}}^{\mathrm{\scriptscriptstyle \top} }_{0} {\mathbf g}_{i}(\boldsymbol{\theta}_0)|=o_{\rm p}(1)$. Write $\tilde{\boldsymbol{\lambda}}_{0}=(\tilde{\lambda}_{0,1},\ldots,\tilde{\lambda}_{0,r})^{\mathrm{\scriptscriptstyle \top} }$. By the Taylor expansion, we have

align*[align* omitted — 596 chars of source]

for some $C \in (0,1)$. By Condition (ref)(b) and the same arguments for deriving Lemma \rm(ref), if $\log r=o(n^{1/3})$ and $\ell_n\alpha_n=o(1)$, we have $\lambda_{\min}\{\widehat{{\mathbf V}}_{\mathcal{M}_{\boldsymbol{\theta}_0}} (\boldsymbol{\theta}_0)\}$ is uniformly bounded away from zero w.p.a.1. Thus $0 \leq|\tilde{\boldsymbol{\lambda}}_{0,\mathcal{M}_{\boldsymbol{\theta}_0}}|_2 |\bar{{\mathbf g}}_{\mathcal{M}_{\boldsymbol{\theta}_0}}(\boldsymbol{\theta}_0)|_2 - 4^{-1}K_3 |\tilde{\boldsymbol{\lambda}}_{0,\mathcal{M}_{\boldsymbol{\theta}_0}}|_2^2$ w.p.a.1, where $K_3$ is specified in Condition (ref)(b). By the moderate deviation of self-normalized sums JingShaoWang2003, $ |\bar{{\mathbf g}}(\boldsymbol{\theta}_0)|_{\infty} =O_{\rm p}(\alpha_n)$. Then $ |\bar{{\mathbf g}}_{\mathcal{M}_{\boldsymbol{\theta}_0}}(\boldsymbol{\theta}_0)|_2= O_{\rm p}(\ell_n^{1/2} \alpha_n) $ and $ |\tilde{\boldsymbol{\lambda}}_{0,\mathcal{M}_{\boldsymbol{\theta}_0}}|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n)=o_{\rm p}(\delta_n)$. Write $\tilde{\boldsymbol{\lambda}}_{0,\mathcal{M}_{\boldsymbol{\theta}_0}}=(\tilde{\lambda}_{1},\ldots,\tilde{\lambda}_{|\mathcal{M}_{\boldsymbol{\theta}_0}|})^{\mathrm{\scriptscriptstyle \top} }$. We then have w.p.a.1 that

align[align omitted — 369 chars of source]

where $\tilde{\boldsymbol{\eta}}=(\tilde{\eta}_1,\ldots,\tilde{\eta}_{|\mathcal{M}_{\boldsymbol{\theta}_0}|})^{\mathrm{\scriptscriptstyle \top} }$ with $\tilde{\eta}_j=\nu\rho'(|\tilde\lambda_{j}|;\nu){\mbox{\rm sgn}}(\tilde{\lambda}_{j})$ for $\tilde\lambda_{j}\neq0$ and $\tilde\eta_j\in[-\nu\rho'(0^+),\nu\rho'(0^+)]$ for $\tilde\lambda_{j}=0$. In the sequel, we will show that $\tilde{\boldsymbol{\lambda}}_{0}$ is a local maximizer for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1.

Firstly, define $\Lambda_{0}^* = \{\boldsymbol{\lambda}\in {\mathbb{R}}^r:|\boldsymbol{\lambda}_{\mathcal{M}_{\boldsymbol{\theta}_0}^*}|_2 \leq \varepsilon,\boldsymbol{\lambda}_{\mathcal{M}^{*,{\rm c}}_{\boldsymbol{\theta}_0}}={\mathbf 0}\}$ for some sufficiently small constant $\varepsilon>0$. For $\tilde{\boldsymbol{\lambda}}_{0}$ defined before, we will prove $\tilde\boldsymbol{\lambda}_{0}=\arg \max_{\boldsymbol{\lambda}\in\Lambda_{0}^*} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. Since $\tilde{\boldsymbol{\lambda}}_{0} \in\Lambda_{0}$ and $\mathcal{M}_{\boldsymbol{\theta}_0} \subset \mathcal{M}_{\boldsymbol{\theta}_0}^*$, we know $\tilde{\boldsymbol{\lambda}}_{0} \in \Lambda_{0}^*$ for sufficiently large $n$. Restricted on $\boldsymbol{\lambda} \in \Lambda_{0}^*$, by the concavity of $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.r.t $\boldsymbol{\lambda}$, it suffices to show that ${\mathbf w} = \tilde{\boldsymbol{\lambda}}_{0,\mathcal{M}^*_{\boldsymbol{\theta}_0}}=: (w_{1},\ldots,w_{|\mathcal{M}^*_{\boldsymbol{\theta}_0}|})^{\mathrm{\scriptscriptstyle \top} } \in {\mathbb{R}}^{|\mathcal{M}^*_{\boldsymbol{\theta}_0}|}$ satisfies the equation

align[align omitted — 325 chars of source]

w.p.a.1, where $\tilde{\boldsymbol{\eta}}^*=(\tilde\eta^*_1,\ldots,\tilde\eta^*_{|\mathcal{M}_{\boldsymbol{\theta}_0}^*|})^{\mathrm{\scriptscriptstyle \top} }$ with $\tilde{\eta}^*_j=\nu\rho'(|w_j|;\nu){\mbox{\rm sgn}}(w_j)$ for $w_j\neq0$ and $\tilde{\eta}^*_j \in [-\nu\rho'(0^+),\nu\rho'(0^+)]$ for $w_j=0$. By (ref), we know $0= n^{-1}\sum_{i=1}^{n} g_{i,j}(\boldsymbol{\theta}_0)/\{1+{\mathbf w} ^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{M}_{\boldsymbol{\theta}_0}^*}(\boldsymbol{\theta}_0)\} - \tilde{\eta}^*_j$ holds for any $j\in\mathcal{M}_{\boldsymbol{\theta}_0}$. For any $j \in \mathcal{M}_{\boldsymbol{\theta}_0}^*\backslash \mathcal{M}_{\boldsymbol{\theta}_0}$, since $\max_{i\in[n]} |{\mathbf w}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{M}_{\boldsymbol{\theta}_0}^*}(\boldsymbol{\theta}_0)|=\max_{i\in[n]} |\tilde\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} }_{0} {\mathbf g}_{i}(\boldsymbol{\theta}_0)|=o_{\rm p}(1)$, it holds that

align[align omitted — 261 chars of source]

with

align[align omitted — 896 chars of source]

Due to $|{\mathbf w}|_2 = |\tilde{\boldsymbol{\lambda}}_{0,\mathcal{M}_{\boldsymbol{\theta}_0}}|_2 = O_{\rm p}(\ell_n^{1/2} \alpha_n)$, by Conditions (ref)(a) and (ref)(b), $\max_{j\in[r]}|R_j|=O_{\rm p}(|{\mathbf w}|_2)=O_{\rm p}(\ell_n^{1/2} \alpha_n)$. Notice that $C_* \nu\rho'(0^+) \leq |\bar{g}_j(\boldsymbol{\theta}_0)| < c \nu\rho'(0^+)$ for any $j\in\mathcal{M}_{\boldsymbol{\theta}_0}^*\backslash \mathcal{M}_{\boldsymbol{\theta}_0}$, and $\ell_n^{1/2} \alpha_n=o(\nu)$. Then \[ \max_{j\in\mathcal{M}_{\boldsymbol{\theta}_0}^*\backslash \mathcal{M}_{\boldsymbol{\theta}_0}}\bigg|\frac{1}{n}\sum_{i=1}^n\frac{g_{i,j}(\boldsymbol{\theta}_0)}{1+{\mathbf w}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{M}_{\boldsymbol{\theta}_0}^*}(\boldsymbol{\theta}_0)}\bigg|\leq \nu\rho'(0^+) \] w.p.a.1, which implies (ref) holds. Thus $\tilde{\boldsymbol{\lambda}}_{0}=\arg\max_{\boldsymbol{\lambda}\in\Lambda_{0}^*}f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1.

Secondly, define $\tilde\Lambda_{0}=\{\boldsymbol{\lambda}\in {\mathbb{R}}^r: |\boldsymbol{\lambda}_{\mathcal{M}_{\boldsymbol{\theta}_0}^*}-\tilde\boldsymbol{\lambda}_{0,\mathcal{M}_{\boldsymbol{\theta}_0}^*}|_2 \leq O(\ell_n^{1/2} \alpha_n), |\boldsymbol{\lambda}_{\mathcal{M}^{*,{\rm c}}_{\boldsymbol{\theta}_0}}|_1\leq O(\ell_n \alpha_n) \}$. We will show $\tilde{\boldsymbol{\lambda}}_{0}=\arg \max_{\boldsymbol{\lambda}\in \tilde\Lambda_{0}} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. Recall $\max_{i\in[n],j\in[r]} |g_{i,j} (\boldsymbol{\theta}_0)| = O_{\rm p}(n^{1/\gamma})$ and $|\tilde\boldsymbol{\lambda}_{0}|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n)$. Since $\ell_n\alpha_n=o(n^{-1/\gamma})$, we have

align*[align* omitted — 1,123 chars of source]

For any $\boldsymbol{\lambda} \in \tilde\Lambda_{0}$, denote by $\mathring\boldsymbol{\lambda}=(\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} }_{\mathcal{M}_{\boldsymbol{\theta}_0}^*},{\mathbf 0}^{\mathrm{\scriptscriptstyle \top} })^{\mathrm{\scriptscriptstyle \top} }$ the projection of $\boldsymbol{\lambda}=(\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} }_{\mathcal{M}_{\boldsymbol{\theta}_0}^*},\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} }_{\mathcal{M}^{*,{\rm c}}_{\boldsymbol{\theta}_0}})^{\mathrm{\scriptscriptstyle \top} }$ onto $\Lambda_{0}^*$. Write $\boldsymbol{\lambda}=(\lambda_1,\ldots,\lambda_r)^{\mathrm{\scriptscriptstyle \top} }$. By the Taylor expansion, it holds that \[ f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)-f_n(\mathring\boldsymbol{\lambda};\boldsymbol{\theta}_0)=\frac{1}{n}\sum_{i=1}^n\frac{{\mathbf g}_{i}(\boldsymbol{\theta}_0)^{\mathrm{\scriptscriptstyle \top} } (\boldsymbol{\lambda}-\mathring\boldsymbol{\lambda})}{1+\boldsymbol{\lambda}_*^{\mathrm{\scriptscriptstyle \top} }{\mathbf g}_{i}(\boldsymbol{\theta}_0)}-\sum_{j\in\mathcal{M}^{*,{\rm c}}_{\boldsymbol{\theta}_0}}P_{\nu}(|\lambda_j|)\,, \] where $\boldsymbol{\lambda}_*$ is on the jointing line between $\boldsymbol{\lambda}$ and $\mathring\boldsymbol{\lambda}$. Let $\check{\boldsymbol{\lambda}}_{0}=\arg \max_{\boldsymbol{\lambda}\in \tilde\Lambda_{0}} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$. Due to $\tilde\boldsymbol{\lambda}_{0} \in {\rm int} (\tilde\Lambda_{0})$, then $f_n(\tilde{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)\leq f_n(\check{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)$. For any $\boldsymbol{\lambda}\in\tilde\Lambda_{0}$, due to $\sum_{j\in\mathcal{M}^{*,{\rm c}}_{\boldsymbol{\theta}_0}}P_{\nu}(|\lambda_j|) \geq \nu\rho'(0^+) |\boldsymbol{\lambda}_{\mathcal{M}^{*,{\rm c}}_{\boldsymbol{\theta}_0}}|_1$ and

align*[align* omitted — 1,741 chars of source]

then $ f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)-f_n(\mathring\boldsymbol{\lambda};\boldsymbol{\theta}_0) \leq \{-(1-C_*)\nu\rho'(0^+)+ O_{\rm p}(\ell_n\alpha_n)\}|\boldsymbol{\lambda}_{\mathcal{M}^{*,{\rm c}}_{\boldsymbol{\theta}_0}}|_1$ for any $\boldsymbol{\lambda}\in\tilde{\Lambda}_0$, where the term $O_{\rm p}(\ell_n\alpha_n)$ holds uniformly over $\boldsymbol{\lambda}\in\tilde{\Lambda}_0$. Since $\ell_n \alpha_n=o(\nu)$, we have $\check{\boldsymbol{\lambda}}_{0,\mathcal{M}_{\boldsymbol{\theta}_0}^{*,{\rm c}}}={\mathbf 0}$ w.p.a.1, which implies $\check{\boldsymbol{\lambda}}_0\in{\rm int}(\Lambda_0^*)$ w.p.a.1. Recall $\tilde \boldsymbol{\lambda}_{0}=\arg\max_{\boldsymbol{\lambda}\in\Lambda_{0}^*}f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. Then $f_n(\tilde{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)\geq f_n(\check{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)$ w.p.a.1. Therefore, $f_n(\tilde{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)= f_n(\check{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)$ w.p.a.1. By the concavity of $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.r.t $\boldsymbol{\lambda}$, we have $\check{\boldsymbol{\lambda}}_0=\tilde\boldsymbol{\lambda}_{0}$ w.p.a.1, which indicates that $\tilde{\boldsymbol{\lambda}}_0$ is a local maximizer for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. Then $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta}_0)=\tilde\boldsymbol{\lambda}_{0}$ and ${\rm supp}\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta}_0)\}\subset \mathcal{M}_{\boldsymbol{\theta}_0}$ w.p.a.1. $\Box$

Case 2: $ \mathcal{M}_{\boldsymbol{\theta}_0} = \emptyset$

In this case, we will show ${\mathbf 0} \in \mathbb{R}^r$ is a local maximizer for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. Due to the concavity of $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.r.t $\boldsymbol{\lambda}$, we then have $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta}_0) = {\mathbf 0}$ w.p.a.1, which implies ${\rm supp}\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta}_0)\} \subset \mathcal{M}_{\boldsymbol{\theta}_0}$ w.p.a.1. Let $\breve{\boldsymbol{\lambda}}_0 = \arg\max_{\boldsymbol{\lambda} \in \breve{\Lambda}_0} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$, where $\breve{\Lambda}_0 =\{ \boldsymbol{\lambda}\in\mathbb{R}^r: |\lambda_{j_0}|\leq (\log n)^{-1}n^{-1/\gamma}, \boldsymbol{\lambda}_{[r]\setminus \{j_0\}}={\mathbf 0} \}$ with $j_0 = \arg\max_{j\in[r]}\mathbb{E}\{g^2_{i,j}(\boldsymbol{\theta}_0)\}$. It follows from Condition (ref)(a) that $\max_{i\in[n]}|g_{i,j_0} (\boldsymbol{\theta}_0)| = O_{\rm p}(n^{1/\gamma})$. Hence, $\max_{i\in[n]} |\breve{\boldsymbol{\lambda}}_0^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i} (\boldsymbol{\theta}_0)|_2 = O_{\rm p}\{(\log n)^{-1}\} = o_{\rm p}(1)$. Write $\breve{\boldsymbol{\lambda}}_0 =(\breve{\lambda}_{0,1},\ldots,\breve{\lambda}_{0,r})^{{\mathrm{\scriptscriptstyle \top} }}$. By the Taylor expansion, we have

align*[align* omitted — 783 chars of source]

for some $C\in(0,1)$. Notice that $|\mathbb{E}_n\{g_{i,j_0}^2(\boldsymbol{\theta}_0)\} - \mathbb{E}\{g_{i,j_0}^2(\boldsymbol{\theta}_0)\}|= O_{\rm p}(n^{-1/2})$. By Condition (ref)(b), we have $\mathbb{E}_n\{g_{i,j_0}^2(\boldsymbol{\theta}_0)\} \geq \mathbb{E}\{g_{i,j_0}^2(\boldsymbol{\theta}_0)\} - o_{\rm p}(1)\geq 2 K_3/3$ w.p.a.1 for $K_3$ specified in Condition (ref)(b). Thus $0 \leq |\breve{\lambda}_{0,j_0}||\bar{g}_{j_0}(\boldsymbol{\theta}_0)| - 4^{-1} K_3 |\breve{\lambda}_{0,j_0}|^2$ w.p.a.1. Since $|\bar{g}_{j_0}(\boldsymbol{\theta}_0)| = O_{\rm p}(n^{-1/2})$, then $|\breve{\lambda}_{0,j_0}| = O_{\rm p}(n^{-1/2}) = o_{\rm p}\{(\log n)^{-1}n^{-1/\gamma}\}$. It then holds w.p.a.1 that

align[align omitted — 186 chars of source]

where $\breve{\eta}_{j_0} = \nu\rho'(|\breve\lambda_{j_0}|;\nu){\mbox{\rm sgn}}(\breve{\lambda}_{j_0})$ if $\breve\lambda_{j_0}\neq0$ and $\breve{\eta}_{j_0} \in[-\nu\rho'(0^+), \nu\rho'(0^+)]$ if $\breve\lambda_{j_0}=0$. Due to $|\breve{\lambda}_{0,j_0} g_{i,j_0}(\boldsymbol{\theta}_0)| = o_{\rm p}(1)$, then

align*[align* omitted — 192 chars of source]

with

align*[align* omitted — 299 chars of source]

By Conditions (ref)(a), we have $|R_{j_0}| = O_{\rm p}(|\breve{\lambda}_{0,j_0}|) = O_{\rm p}(n^{-1/2})$. Together with $|\bar{g}_{j_0}(\boldsymbol{\theta}_0)| = O_{\rm p}(n^{-1/2})$, (ref) leads to $|\breve{\eta}_{j_0}| = O_{\rm p}(n^{-1/2}) = o_{\rm p}(\nu)$. Then $\breve\lambda_{j_0}=0$ w.p.a.1, which implies $\breve{\boldsymbol{\lambda}}_0 = {\mathbf 0}$ w.p.a.1. In the sequel, we will show that $\breve{\boldsymbol{\lambda}}_0$ is a local maximizer for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1.

Firstly, define $\breve{\Lambda}_0^* = \{\boldsymbol{\lambda}\in {\mathbb{R}}^r:|\boldsymbol{\lambda}_{\mathcal{H}}|_2 \leq \varepsilon,\boldsymbol{\lambda}_{\mathcal{H}^{\rm c}}={\mathbf 0}\}$ for some sufficiently small constant $\varepsilon>0$, where $j_0 \in \mathcal{H} \subset [r]$ with $1 < |\mathcal{H}|\leq \ell_n$. For $\breve{\boldsymbol{\lambda}}_0$ defined before, we will prove $\breve{\boldsymbol{\lambda}}_0 =\arg \max_{\boldsymbol{\lambda}\in \breve{\Lambda}_0^*} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. Since $\breve{\boldsymbol{\lambda}}_0 \in \breve{\Lambda}_0$ and $j_0 \in \mathcal{H}$, we know $\breve{\boldsymbol{\lambda}}_0 \in \breve{\Lambda}_0^*$ for sufficiently large $n$. Restricted on $\boldsymbol{\lambda} \in \breve{\Lambda}_0^*$, by the concavity of $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.r.t $\boldsymbol{\lambda}$, it suffices to show that $\breve{{\mathbf w}} = \breve{\boldsymbol{\lambda}}_{0,\mathcal{H}}=: (\breve{w}_{1},\ldots,\breve{w}_{|\mathcal{H}|})^{\mathrm{\scriptscriptstyle \top} } \in {\mathbb{R}}^{|\mathcal{H}|}$ satisfies the equation

align[align omitted — 286 chars of source]

w.p.a.1, where $\breve{\boldsymbol{\eta}}^*=(\breve\eta^*_1,\ldots,\breve\eta^*_{|\mathcal{H}|})^{\mathrm{\scriptscriptstyle \top} }$ with $\breve{\eta}^*_j=\nu\rho'(|\breve{w}_j|;\nu){\mbox{\rm sgn}}(\breve{w}_j)$ for $\breve{w}_j\neq0$ and $\breve{\eta}^*_j \in [-\nu\rho'(0^+),\nu\rho'(0^+)]$ for $\breve{w}_j=0$. Recall $j_0\in\mathcal{H}$. Without loss of generality, we assume $j_0$ is the first component in $\mathcal{H}$. By (ref), we know $0= n^{-1}\sum_{i=1}^{n} g_{i,j_0}(\boldsymbol{\theta}_0)/\{1+\breve{{\mathbf w}}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{H}}(\boldsymbol{\theta}_0)\} - \breve{\eta}^*_{1}$ holds. Since $\breve{\boldsymbol{\lambda}}_0 = {\mathbf 0}$ w.p.a.1, then $\breve{{\mathbf w}}= {\mathbf 0}$ w.p.a.1, which implies it holds w.p.a.1 that

align*[align* omitted — 221 chars of source]

for any $j \in \mathcal{H} \backslash \{j_0\}$. By the moderate deviation of self-normalized sums JingShaoWang2003, $ |\bar{{\mathbf g}}(\boldsymbol{\theta}_0)|_{\infty} =O_{\rm p}(\alpha_n)$. Due to $\alpha_n = o(\nu)$, then \[ \max_{j \in \mathcal{H} \backslash \{j_0\}}\bigg|\frac{1}{n}\sum_{i=1}^n\frac{g_{i,j}(\boldsymbol{\theta}_0)}{1+\breve{{\mathbf w}}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{H}}(\boldsymbol{\theta}_0)}\bigg|\leq \nu\rho'(0^+) \] w.p.a.1, which implies (ref) holds. Thus $\breve{\boldsymbol{\lambda}}_0 =\arg \max_{\boldsymbol{\lambda}\in \breve{\Lambda}_0^*} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1.

Secondly, define $\bar{\Lambda}_{0}=\{\boldsymbol{\lambda}\in {\mathbb{R}}^r: |\boldsymbol{\lambda}_{\mathcal{H}}-\breve\boldsymbol{\lambda}_{0,\mathcal{H}}|_2 \leq O(\ell_n^{1/2} \alpha_n), |\boldsymbol{\lambda}_{\mathcal{H}^{\rm c}}|_1\leq O(\ell_n \alpha_n) \}$. We will show $\breve{\boldsymbol{\lambda}}_{0}=\arg \max_{\boldsymbol{\lambda}\in \bar{\Lambda}_{0}} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. Recall $\max_{i\in[n],j\in[r]} |g_{i,j} (\boldsymbol{\theta}_0)| = O_{\rm p}(n^{1/\gamma})$ and $|\breve{\boldsymbol{\lambda}}_{0}|_2=|\breve{\lambda}_{0,j_0}| = O_{\rm p}(n^{-1/2})$. Since $\ell_n\alpha_n=o(n^{-1/\gamma})$, we have

align*[align* omitted — 935 chars of source]

For any $\boldsymbol{\lambda} \in \bar{\Lambda}_{0}$, denote by $\mathring\boldsymbol{\lambda}=(\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} }_{\mathcal{H}},{\mathbf 0}^{\mathrm{\scriptscriptstyle \top} })^{\mathrm{\scriptscriptstyle \top} }$ the projection of $\boldsymbol{\lambda}=(\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} }_{\mathcal{H}},\boldsymbol{\lambda}^{\mathrm{\scriptscriptstyle \top} }_{\mathcal{H}^{\rm c}})^{\mathrm{\scriptscriptstyle \top} }$ onto $\breve{\Lambda}_0^*$. Write $\boldsymbol{\lambda}=(\lambda_1,\ldots,\lambda_r)^{\mathrm{\scriptscriptstyle \top} }$. By the Taylor expansion, it holds that \[ f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)-f_n(\mathring\boldsymbol{\lambda};\boldsymbol{\theta}_0)=\frac{1}{n}\sum_{i=1}^n\frac{{\mathbf g}_{i}(\boldsymbol{\theta}_0)^{\mathrm{\scriptscriptstyle \top} } (\boldsymbol{\lambda}-\mathring\boldsymbol{\lambda})}{1+\boldsymbol{\lambda}_*^{\mathrm{\scriptscriptstyle \top} }{\mathbf g}_{i}(\boldsymbol{\theta}_0)}-\sum_{j\in\mathcal{H}^{\rm c}}P_{\nu}(|\lambda_j|)\,, \] where $\boldsymbol{\lambda}_*$ is on the jointing line between $\boldsymbol{\lambda}$ and $\mathring\boldsymbol{\lambda}$. Let $\check{\boldsymbol{\lambda}}_{0}=\arg \max_{\boldsymbol{\lambda}\in \bar\Lambda_{0}} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$. Due to $\breve\boldsymbol{\lambda}_{0} \in {\rm int} (\bar\Lambda_{0})$, then $f_n(\breve{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)\leq f_n(\check{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)$. For any $\boldsymbol{\lambda} \in \bar\Lambda_{0}$, due to $\sum_{j\in\mathcal{H}^{\rm c}}P_{\nu}(|\lambda_j|) \geq \nu\rho'(0^+) |\boldsymbol{\lambda}_{\mathcal{H}^{\rm c}}|_1$ and

align*[align* omitted — 1,455 chars of source]

then $ f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)-f_n(\mathring\boldsymbol{\lambda};\boldsymbol{\theta}_0) \leq \{-\nu\rho'(0^+)+ O_{\rm p}(\ell_n\alpha_n)\}|\boldsymbol{\lambda}_{\mathcal{H}^{\rm c}}|_1$ for any $\boldsymbol{\lambda}\in\bar{\Lambda}_0$, where the term $O_{\rm p}(\ell_n\alpha_n)$ holds uniformly over $\boldsymbol{\lambda}\in\bar{\Lambda}_0$. Since $\ell_n \alpha_n=o(\nu)$, we have $\check{\boldsymbol{\lambda}}_{0,\mathcal{H}^{\rm c}}={\mathbf 0}$ w.p.a.1, which implies $\check{\boldsymbol{\lambda}}_0\in{\rm int}(\breve{\Lambda}_0^*)$ w.p.a.1. Recall $\breve{\boldsymbol{\lambda}}_0 =\arg \max_{\boldsymbol{\lambda}\in \breve{\Lambda}_0^*} f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. Then $f_n(\breve{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)\geq f_n(\check{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)$ w.p.a.1. Therefore, $f_n(\breve{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0) = f_n(\check{\boldsymbol{\lambda}}_0;\boldsymbol{\theta}_0)$ w.p.a.1. By the concavity of $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.r.t $\boldsymbol{\lambda}$, we have $\check{\boldsymbol{\lambda}}_0=\breve{\boldsymbol{\lambda}}_0$ w.p.a.1, which indicates that ${\mathbf 0}$ is a local maximizer for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1. $\Box$

Proof of Lemma (ref)

Same as the proof of Lemma (ref), we only need to show that there exists a local maximizer satisfying the results stated in the lemma. Recall $\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}^*=\{j\in[r]: |\bar{g}_j(\hat{\boldsymbol{\theta}}_n)| \geq C_*\nu \rho'(0^{+}) \}$ for some $C_* \in (0,1)$, and $\mathbb{P}(\max _{\boldsymbol{\theta} \in \boldsymbol{\Theta}:\,|\boldsymbol{\theta}-\boldsymbol{\theta}_0|_2 \leq c_n}|\mathcal{M}_{\boldsymbol{\theta}}^*|\leq\ell_n)\rightarrow1$ for some $c_n\rightarrow 0$ satisfying $\nu c_n^{-1} \rightarrow 0$. For $\tilde{c} \in (C_*,1)$ given in Condition (ref)(a), write $\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}:=\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})=\{ j\in[r]:|\bar{g}_j(\hat{\boldsymbol{\theta}}_n)| \geq \tilde{c}\nu\rho'(0^{+}) \}$. By Proposition (ref), $|\hat{\boldsymbol{\theta}}_n-\boldsymbol{\theta}_0|_\infty=O_{\rm p}(\nu)$. Notice that $p$ is fixed. Then $|\hat{\boldsymbol{\theta}}_n-\boldsymbol{\theta}_0|_2=O_{\rm p}(\nu)$ which implies $\ell_n \geq |\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}^*| \geq |\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}|$ w.p.a.1. Restricted on $\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}$, we select $\delta_n$ satisfying $\delta_n=o(\ell_n^{-1/2} n^{-1/\gamma})$ and $\ell_n^{1/2} \alpha_n=o(\delta_n) $, which can be guaranteed by $\ell_n\alpha_n=o(n^{-1/\gamma})$. Let $\Lambda_n=\{\boldsymbol{\lambda}\in {\mathbb{R}}^{r}:|\boldsymbol{\lambda}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}|_2\leq \delta_n, \boldsymbol{\lambda}_{\mathcal{M}^{\rm c}_{\hat{\boldsymbol{\theta}}_n}}={\mathbf 0} \}$ and $\tilde{\boldsymbol{\lambda}}_n=(\tilde{\lambda}_{n,1},\ldots,\tilde{\lambda}_{n,r})^{\mathrm{\scriptscriptstyle \top} }=\arg \max _{\boldsymbol{\lambda} \in \Lambda_n} f_n(\boldsymbol{\lambda}; \hat{\boldsymbol{\theta}}_n)$. By the Taylor expansion, we have

align[align omitted — 633 chars of source]

for some $C \in (0,1)$. By Proposition (ref), Lemma (ref) and Condition (ref)(b), if $\log r=o(n^{1/3})$, $\ell_n\nu^2=o(1)$ and $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$, we have $\lambda_{\min}\{\widehat{{\mathbf V}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}} (\hat{\boldsymbol{\theta}}_n)\}$ is uniformly bounded away from zero w.p.a.1. Therefore, it holds w.p.a.1 that $$ 0 \leq \tilde{\boldsymbol{\lambda}}_{n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}^{\mathrm{\scriptscriptstyle \top} } \big[\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)\}\big] - 4^{-1}K_3 |\tilde{\boldsymbol{\lambda}}_{n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}|_2^2$$ with $K_3$ specified in Condition (ref)(b), which implies $|\tilde\boldsymbol{\lambda}_{n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}|_2 \leq 4K_3^{-1} |\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)\}|_2$ w.p.a.1.

Select $\boldsymbol{\lambda}_n^* \in {\mathbb{R}}^r$ satisfying $\boldsymbol{\lambda}^*_{n,\mathcal{M}^{\rm c}_{\hat{\boldsymbol{\theta}}_n}}={\mathbf 0}$ and $$\boldsymbol{\lambda}^*_{n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}=\frac{\delta_n [\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)\}]}{|\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)\}|_2}\,.$$ Then $\boldsymbol{\lambda}_n^* \in \Lambda_n$. As shown in the proof of Proposition (ref), $\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta}_0)}f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0) = O_{\rm p}(\ell_n \alpha_n^2)=o_{\rm p}(\delta_n^2)$, which implies $\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\hat{\boldsymbol{\theta}}_n)}f_n(\boldsymbol{\lambda};\hat{\boldsymbol{\theta}}_n)=o_{\rm p}(\delta_n^2)$. Write $\boldsymbol{\lambda}_n^*=(\lambda_{n,1}^*,\ldots,\lambda_{n,r}^*)^{\mathrm{\scriptscriptstyle \top} }$. Notice that $\Lambda_n\subset\hat{\Lambda}_n(\hat{\boldsymbol{\theta}}_n)$ w.p.a.1. By the Taylor expansion, it holds w.p.a.1 that

align*[align* omitted — 1,807 chars of source]

for some $\bar C, c_j \in (0,1)$, where the last inequality follows from the condition that $P_{\nu}(\cdot)$ has bounded second-order derivative around $0$. For any $j \in \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}$, we have ${\mbox{\rm sgn}}(\lambda_{n,j}^*)={\mbox{\rm sgn}}\{\bar g_j(\hat{\boldsymbol{\theta}}_n)\}$ if $|\bar {g}_j(\hat{\boldsymbol{\theta}}_n)| > \nu \rho'(0^{+})$, and $\bar g_j(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar g_j(\hat{\boldsymbol{\theta}}_n)\} = 0 = \lambda_{n,j}^*$ if $|\bar {g}_j(\hat{\boldsymbol{\theta}}_n)| = \nu \rho'(0^{+})$. Thus, $$\lambda_{n,j}^* \big\{\bar g_j(\hat{\boldsymbol{\theta}}_n) - \nu \rho'(0^+) {\mbox{\rm sgn}}(\lambda_{n,j}^*)\big\} = \lambda_{n,j}^* \big[\bar g_j(\hat{\boldsymbol{\theta}}_n) - \nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar g_j(\hat{\boldsymbol{\theta}}_n)\}\big]$$ for any $j \in \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}$ with $|\bar {g}_j(\hat{\boldsymbol{\theta}}_n)| \geq \nu \rho'(0^{+})$. By Condition (ref)(a), $\{j \in [r]: \tilde{c} \nu\rho'(0^{+}) \leq |\bar {g}_j(\hat{\boldsymbol{\theta}}_n)| < \nu\rho'(0^{+})\} = \emptyset$ w.p.a.1. Recall $\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}=\{j \in [r]: |\bar {g}_j(\hat{\boldsymbol{\theta}}_n)| \geq \tilde{c} \nu \rho'(0^{+})\}$. We then have w.p.a.1 that

align*[align* omitted — 691 chars of source]

Thus, $ |\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)\}|_2 =O_{\rm p}(\delta_n) $. For any $\epsilon_n \rightarrow 0$, select $\boldsymbol{\lambda}_n^{**}$ such that $\boldsymbol{\lambda}_{n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}^{**}=\epsilon_n [\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)\}]$ and $\boldsymbol{\lambda}^{**}_{n,\mathcal{M}^{\rm c}_{\hat{\boldsymbol{\theta}}_n}}={\mathbf 0}$. Then $|\boldsymbol{\lambda}_n^{**}|_2=o_{\rm p}(\delta_n)$. Due to $f_n(\boldsymbol{\lambda}_n^{**};\hat{\boldsymbol{\theta}}_n) \leq \max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\hat{\boldsymbol{\theta}}_n)}f_n(\boldsymbol{\lambda};\hat{\boldsymbol{\theta}}_n) \leq \max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta}_0)}f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)=O_{\rm p}(\ell_n \alpha_n^2)$, using the same arguments given above, we have

align*[align* omitted — 563 chars of source]

Hence, $\epsilon_n |\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)\}|^2_2=O_{\rm p}(\ell_n \alpha_n^2)$. Since we can select arbitrary slow $\epsilon_n \rightarrow 0$, it holds that

align[align omitted — 292 chars of source]

which implies $ |\tilde{\boldsymbol{\lambda}}_n|_2 = |\tilde{\boldsymbol{\lambda}}_{n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n) = o_{\rm p}(\delta_n) $. Write $\tilde{\boldsymbol{\lambda}}_{n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}=(\tilde{\lambda}_{n,1},\ldots,\tilde{\lambda}_{n,|\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}|})^{\mathrm{\scriptscriptstyle \top} }$. We have w.p.a.1 that $${\mathbf 0}= \frac{1}{n}\sum_{i=1}^{n} \frac{{\mathbf g}_{i,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)}{1+\tilde{\boldsymbol{\lambda}}_{n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)} - \tilde{\boldsymbol{\eta}}\,,$$ where $\tilde{\boldsymbol{\eta}}=(\tilde\eta_1,\ldots,\tilde\eta_{|\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}|})^{\mathrm{\scriptscriptstyle \top} }$ with $\tilde\eta_j=\nu\rho'(|\tilde\lambda_{n,j}|;\nu){\mbox{\rm sgn}}(\tilde{\lambda}_{n,j})$ for $\tilde\lambda_{n,j}\neq0$ and $\tilde\eta_j\in[-\nu\rho'(0^+),\nu\rho'(0^+)]$ for $\tilde\lambda_{n,j}=0$. Identical to (ref), we have $\tilde{\boldsymbol{\eta}}=\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}}(\hat{\boldsymbol{\theta}}_n)+{\mathbf R}$ for some $|\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}|$-dimensional vector ${\mathbf R}$. Applying the same arguments for deriving the rate of $R_j$ in (ref), it holds that $|{\mathbf R}|_{\infty}=O_{\rm p}(\ell_n^{1/2} \alpha_n)$. Since $\ell_n\alpha_n=o(\nu) $, we then have ${\mbox{\rm sgn}} (\tilde{\lambda}_{n,j})={\mbox{\rm sgn}}\{\bar{g}_j(\hat{\boldsymbol{\theta}}_n)\}$ for any $j \in \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}$ with $\tilde{\lambda}_{n,j} \neq 0$ w.p.a.1. Using the arguments in Section (ref) for showing $\tilde\boldsymbol{\lambda}_{0}$ is a local maximizer for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1, we can prove $\tilde\boldsymbol{\lambda}_n$ is a local maximizer for $f_n(\boldsymbol{\lambda};\hat{\boldsymbol{\theta}}_n)$ w.p.a.1, which implies $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)=\tilde{\boldsymbol{\lambda}}_n$ w.p.a.1. We then have Lemma (ref). $\Box$

Proof of Lemma (ref)

Recall $\hat\boldsymbol{\lambda}(\boldsymbol{\theta})=\arg\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})} f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$. Then $\hat{\boldsymbol{\theta}}_n$ and $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n) =(\hat\lambda_1,\ldots,\hat\lambda_r)^{\mathrm{\scriptscriptstyle \top} }$ satisfy

align[align omitted — 303 chars of source]

where $\hat{\boldsymbol{\eta}}=(\hat{\eta}_{1},\ldots,\hat{\eta}_{r})^{\mathrm{\scriptscriptstyle \top} }$ with $\hat{\eta}_j=\nu \rho' (|\hat{\lambda}_j|;\nu) {\mbox{\rm sgn}}(\hat{\lambda}_j)$ for $\hat{\lambda}_j\neq0$ and $\hat{\eta}_j \in [-\nu \rho'(0^{+}),\nu \rho'(0^{+})]$ for $\hat{\lambda}_j=0$. Recall $\mathcal{R}_n={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$. Restricted on $\mathcal{R}_n$, for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}$ and $\boldsymbol{\zeta}=(\zeta_{1},\ldots,\zeta_{|{ \mathcal{R}_n}|})^{\mathrm{\scriptscriptstyle \top} } \in {\mathbb{R}}^{|{ \mathcal{R}_n}|}$ with each $\zeta_j \neq 0$, define $$ {\mathbf m}(\boldsymbol{\zeta},\boldsymbol{\theta})=\frac{1}{n} \sum _{i=1}^{n} \frac{{\mathbf g}_{i,{ \mathcal{R}_n}}(\boldsymbol{\theta})}{1+\boldsymbol{\zeta}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,{ \mathcal{R}_n}}(\boldsymbol{\theta})}-{\mathbf w}\,,$$ where ${\mathbf w}=(w_{1},\ldots,w_{|{ \mathcal{R}_n}|})^{\mathrm{\scriptscriptstyle \top} }$ with $w_j=\nu \rho' (|\zeta_j|;\nu) {\mbox{\rm sgn}}(\zeta_j)$. From (ref), we know $\hat{\boldsymbol{\lambda}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)$ and $\hat{\boldsymbol{\theta}}_n$ satisfy ${\mathbf m}\{\hat{\boldsymbol{\lambda}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n), \hat{\boldsymbol{\theta}}_n\}= {\mathbf 0}$. By the implicit function theorem [Theorem 9.28 of Rudin1976], for all $\boldsymbol{\theta}$ in a small neighborhood of $\hat{\boldsymbol{\theta}}_n$, denoted by $\mathcal{U}(\hat{\boldsymbol{\theta}}_n)$, there exists a $\boldsymbol{\zeta}(\boldsymbol{\theta})$ such that ${\mathbf m}\{\boldsymbol{\zeta}(\boldsymbol{\theta}),\boldsymbol{\theta}\}={\mathbf 0}$, $\boldsymbol{\zeta}(\hat{\boldsymbol{\theta}}_n)=\hat{\boldsymbol{\lambda}}_{ \mathcal{R}_n}(\hat{\boldsymbol{\theta}}_n)$ and $\boldsymbol{\zeta}(\boldsymbol{\theta})$ is continuously differentiable in $\boldsymbol{\theta}\in \mathcal{U}(\hat{\boldsymbol{\theta}}_n)$. By Condition (ref)(b), the event ${\mathcal{E}}=\{\max_{j \in \mathcal{R}_n^{\rm c}} |\hat{\eta}_j| < \nu\rho'(0^+)\}$ holds w.p.a.1. Restricted on ${\mathcal{E}}$, let $\varsigma_n=\nu\rho'(0^+) - \max_{j \in \mathcal{R}_n^{\rm c}} |\hat{\eta}_j|$ and define $\boldsymbol{\Theta}_*=\{\boldsymbol{\theta} \in \mathcal{U}(\hat{\boldsymbol{\theta}}_n):\, |\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_1 \leq o[\min\{\varsigma_n, \chi_n\}], |\boldsymbol{\zeta}(\boldsymbol{\theta})-\boldsymbol{\zeta}(\hat{\boldsymbol{\theta}}_n)|_1 \leq o [\min\{\varsigma_n, \ell_n^{1/2} \alpha_n\}]\}$ for some $\chi_n>0$. Since all the components of $\boldsymbol{\zeta}(\hat{\boldsymbol{\theta}}_n)$ are nonzero and $\boldsymbol{\zeta}(\boldsymbol{\theta})$ is continuously differentiable in $\hat{\boldsymbol{\theta}}_n$, we can select sufficiently small $\chi_n$ such that all the components of $\boldsymbol{\zeta}(\boldsymbol{\theta})$ are nonzero for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}_*$. For any $\boldsymbol{\theta} \in \boldsymbol{\Theta}_*$, let $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})= \{\tilde\lambda_1(\boldsymbol{\theta}), \ldots, \tilde\lambda_r(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} } \in {\mathbb{R}}^r$ satisfy $\tilde\boldsymbol{\lambda}_{ \mathcal{R}_n}(\boldsymbol{\theta})=\boldsymbol{\zeta}(\boldsymbol{\theta})$ and $\tilde\boldsymbol{\lambda}_{\mathcal{R}_n^{\rm c}}(\boldsymbol{\theta})={\mathbf 0}$. Since ${\mathbf m}\{\boldsymbol{\zeta}(\boldsymbol{\theta}),\boldsymbol{\theta}\}={\mathbf 0}$, $\tilde\boldsymbol{\lambda}_{ \mathcal{R}_n}(\boldsymbol{\theta})=\boldsymbol{\zeta}(\boldsymbol{\theta})$ and $\tilde\boldsymbol{\lambda}_{\mathcal{R}_n^{\rm c}}(\boldsymbol{\theta})={\mathbf 0}$ for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}_*$, then $$ 0 = \frac{1}{n}\sum _{i=1}^{n} \frac{g_{i,j}(\boldsymbol{\theta})}{1+\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i}(\boldsymbol{\theta})} - \nu \rho' \{|\tilde{\lambda}_j(\boldsymbol{\theta})|;\nu\} {\mbox{\rm sgn}}\{\tilde{\lambda}_j(\boldsymbol{\theta})\}$$ for any $j \in \mathcal{R}_n$. For any $\boldsymbol{\theta} \in \boldsymbol{\Theta}_*$ and $j \in \mathcal{R}_n^{\rm c}$, by the Taylor expansion, we have

align[align omitted — 2,504 chars of source]

where $\check{\boldsymbol{\theta}}$ is lying on the jointing line between $\boldsymbol{\theta}$ and $\hat{\boldsymbol{\theta}}_n$, and $\check{\boldsymbol{\lambda}}$ is lying on the jointing line between $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})$ and $\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)$. By Lemma (ref), $|\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n)$ and $|\mathcal{R}_n|\leq \ell_n$ w.p.a.1. Then $|\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta})|_2 = |\boldsymbol{\zeta}(\boldsymbol{\theta})|_2 \leq |\boldsymbol{\zeta}(\hat{\boldsymbol{\theta}}_n)|_2 + |\boldsymbol{\zeta}(\boldsymbol{\theta})-\boldsymbol{\zeta}(\hat{\boldsymbol{\theta}}_n)|_2 = O_{\rm p}(\ell_n^{1/2} \alpha_n)$, which implies $|\check{\boldsymbol{\lambda}}|_2 = O_{\rm p}(\ell_n^{1/2} \alpha_n)$. Together with Condition (ref)(a) and $\ell_n \alpha_n=o(n^{-1/\gamma})$, it yields that $\max_{i \in [n]} \{ |\tilde{\boldsymbol{\lambda}}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i}(\check{\boldsymbol{\theta}})| + |\check{\boldsymbol{\lambda}}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i}(\hat{\boldsymbol{\theta}}_n)|\} = o_{\rm p}(1)$. By Conditions (ref)(a) and (ref)(c), we have

align*[align* omitted — 617 chars of source]

It follows from the Cauchy-Schwarz inequality that

align*[align* omitted — 1,326 chars of source]

By (ref), for any $\boldsymbol{\theta}\in\boldsymbol{\Theta}_*$, we know

align*[align* omitted — 669 chars of source]

holds uniformly over $j \in \mathcal{R}_n^{\rm c}$. Due to ${\mathbb{P}}({\mathcal{E}}) \rightarrow 1$, we have $$ \max_{j \in \mathcal{R}_n^{\rm c}}\bigg| \frac{1}{n}\sum _{i=1}^{n} \frac{g_{i,j}(\boldsymbol{\theta})}{1+\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i}(\boldsymbol{\theta})} \bigg| \leq \nu\rho'(0^+)$$ w.p.a.1. Therefore, $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})$ and $\boldsymbol{\theta}$ satisfy the score equation $\nabla_{\boldsymbol{\lambda}} f_n\{\tilde\boldsymbol{\lambda}(\boldsymbol{\theta}); \boldsymbol{\theta}\}={\mathbf 0}$ for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}_*$ w.p.a.1. By the concavity of $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ w.r.t $\boldsymbol{\lambda}$, we have $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})=\arg\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})} f_n(\boldsymbol{\lambda};\boldsymbol{\theta})=\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})$ for any $\boldsymbol{\theta} \in \boldsymbol{\Theta}_*$ w.p.a.1. Hence, $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})$ is continuously differentiable at $\hat{\boldsymbol{\theta}}_n$ and $ [\nabla_{\boldsymbol{\theta}} \hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)]_{\mathcal{R}_n^{\rm c},[p]}={\mathbf 0}$ w.p.a.1. $\Box$

Proof of Lemma (ref)

The proof is almost identical to that of Lemma 2 in Changetal2018. From Lemma (ref), we have $|\hat{\boldsymbol{\lambda}}|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n)$. Recall $p$ is fixed in our current setting. We only need to replace the convergence rate of $|\hat{\boldsymbol{\lambda}}|_2$ in the proof of Lemma 2 in Changetal2018 by $O_{\rm p}(\ell_n^{1/2} \alpha_n)$ and also set $(s,\omega_n)$ there as $(p,1)$ and all the arguments still hold. $\Box$

Proof of Lemma (ref)

The proof is almost identical to that of Lemma 3 in Changetal2018. Since $p$ is fixed, we only need to replace $\{\omega_n,\varpi_n,b_n^{1/(2\beta)},s\}$ in the proof of Lemma 3 in Changetal2018 by $(1,1,\nu,p)$ and all the arguments still hold. $\Box$

Proof of Lemma (ref)

Recall $\boldsymbol{\Gamma}_{\mathcal{F}}(\boldsymbol{\theta}_0)={\mathbb{E}} \{\nabla_{\boldsymbol{\theta}} {\mathbf g}_{i,\mathcal{F}}(\boldsymbol{\theta}_0) \}$ and ${\mathbf V}_{\mathcal{F}} (\boldsymbol{\theta}_0)= {\mathbb{E}} \{ {\mathbf g}_{i,\mathcal{F}} (\boldsymbol{\theta}_0)^{\otimes2} \}$. For any ${\mathbf t} \in {\mathbb{R}}^p$ with $|{\mathbf t}|_2=1$, let $Z_{i,\mathcal{F}}={\mathbf t}^{\mathrm{\scriptscriptstyle \top} } {\mathbf H}_{\mathcal{F}}^{-1/2} \boldsymbol{\Gamma}_{\mathcal{F}}(\boldsymbol{\theta}_0)^{\mathrm{\scriptscriptstyle \top} } {\mathbf V}^{-1}_{\mathcal{F}} (\boldsymbol{\theta}_0) {\mathbf g}_{i,\mathcal{F}} (\boldsymbol{\theta}_0)$ with ${\mathbf H}_{\mathcal{F}}=\{\boldsymbol{\Gamma}_{\mathcal{F}}(\boldsymbol{\theta}_0)^{\mathrm{\scriptscriptstyle \top} } {\mathbf V}^{-1/2}_{\mathcal{F}} (\boldsymbol{\theta}_0)\}^{\otimes2}$. Write $G_{\mathcal{F}}= {\mathbb{E}}_n(Z_{i,\mathcal{F}})$ and $\hat{G}_{\mathcal{F}}= {\mathbf t}^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf H}}_{\mathcal{F}}^{-1/2} \widehat{\boldsymbol{\Gamma}}_{\mathcal{F}}(\hat{\boldsymbol{\theta}}_n)^{\mathrm{\scriptscriptstyle \top} } \widehat{{\mathbf V}}^{-1}_{\mathcal{F}} (\hat{\boldsymbol{\theta}}_n) \bar{{\mathbf g}}_{\mathcal{F}} (\boldsymbol{\theta}_0)$. It follows from the Berry-Esseen inequality that $$ \sup_{u \in {\mathbb{R}}} |{\mathbb{P}}(n^{1/2} G_{\mathcal{F}} \leq u) -\Phi(u)| \leq Cn^{-1/2} {\mathbb{E}}(|Z_{i,\mathcal{F}}|^3)$$ for some universal constant $C>0$. By the Cauchy-Schwarz inequality,

align*[align* omitted — 544 chars of source]

for $K_3$ given in Condition (ref)(b). By the Jensen's inequality, Condition (ref)(a) yields $ {\mathbb{E}}\{|{\mathbf g}_{i,\mathcal{F}}(\boldsymbol{\theta}_0)|_2^3\}\leq K_2^{3/\gamma}\ell_n^{3/2}$ for $K_2$ and $\gamma$ given in Condition (ref)(a), which implies $ {\mathbb{E}}(|Z_{i,\mathcal{F}}|^3) \leq K_3^{-3/2}{\mathbb{E}}\{|{\mathbf g}_{i,\mathcal{F}}(\boldsymbol{\theta}_0)|_2^3\} \leq K_2^{3/\gamma}K_3^{-3/2}\ell_n^{3/2} $. If $\ell_n=o(n^{1/3})$, we have $$ \sup_{\mathcal{F} \in \mathscr{F}} \sup_{u \in {\mathbb{R}}}|{\mathbb{P}}(n^{1/2} G_{\mathcal{F}} \leq u) -\Phi(u)| \rightarrow 0$$ as $n\rightarrow\infty$. By Conditions (ref)(b) and (ref) and Lemmas (ref) and (ref), it holds that $\sup_{\mathcal{F} \in \mathscr{F}}|n^{1/2}(\hat{G}_{\mathcal{F}}-G_{\mathcal{F}})|= O_{\rm p}\{\ell_n \nu (\log r)^{1/2}\} + O_{\rm p}\{\ell_n^{3/2} \alpha_n(\log r)^{1/2}\}$. For any constant $\delta>0$, due to $ {\mathbb{P}} (n^{1/2} \hat{G}_{\mathcal{F}} \leq u ) - \Phi(u) \leq {\mathbb{P}} (n^{1/2} G_{\mathcal{F}} \leq u+\delta) + {\mathbb{P}}\{|n^{1/2}(\hat{G}_{\mathcal{F}}-G_{\mathcal{F}})| \geq \delta \} - \Phi(u)$ and $ {\mathbb{P}}(n^{1/2} \hat{G}_{\mathcal{F}} \leq u) - \Phi(u) \geq {\mathbb{P}} (n^{1/2} G_{\mathcal{F}} \leq u-\delta) - {\mathbb{P}}\{|n^{1/2}(\hat{G}_{\mathcal{F}}-G_{\mathcal{F}})| \geq \delta \} - \Phi(u)$, it holds that

align*[align* omitted — 453 chars of source]

Notice that $\sup_{\mathcal{F} \in \mathscr{F}}|n^{1/2}(\hat{G}_{\mathcal{F}}-G_{\mathcal{F}})|=o_{\rm p}(1)$ and $\sup_{u \in {\mathbb{R}}} |\Phi(u+\delta)-\Phi(u-\delta)| \leq(2\pi^{-1})^{1/2}\delta$. Then it holds that $\limsup_{n\rightarrow\infty}\sup_{\mathcal{F} \in \mathscr{F}} \sup_{u \in {\mathbb{R}}} |{\mathbb{P}}(n^{1/2}\hat{G}_{\mathcal{F}} \leq u) - \Phi(u)|\leq (2\pi^{-1})^{1/2}\delta$. Due to the arbitrary selection of $\delta>0$, we have $\sup_{\mathcal{F} \in \mathscr{F}} \sup_{u \in {\mathbb{R}}} |{\mathbb{P}}(n^{1/2}\hat{G}_{\mathcal{F}} \leq u) - \Phi(u)|\rightarrow0$ as $n\rightarrow\infty$. $\Box$

Proof of Lemma (ref)

Recall $\mathcal{C}_1=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2\leq \alpha_n\}$. For any $\boldsymbol{\theta} \in \mathcal{C}_1$, same as the proof of Lemma (ref), we only need to show that there exists a local maximizer satisfying the results stated in the lemma. For $\tilde{c}$ specified in Condition (ref)(a), we select $c \in (\tilde{c},1)$. We select $\delta_n$ satisfying $\delta_n=o(\ell_n^{-1/2} n^{-1/\gamma})$ and $\ell_n^{1/2} \alpha_n=o(\delta_n) $, which can be guaranteed by $\ell_n\alpha_n=o(n^{-1/\gamma})$. For each $\boldsymbol{\theta} \in \mathcal{C}_1$, define $\Lambda_{\boldsymbol{\theta}}=\{\boldsymbol{\lambda}\in {\mathbb{R}}^{r}:\,|\boldsymbol{\lambda}_{\mathcal{M}_{\boldsymbol{\theta}}(c)}|_2\leq \delta_n, \boldsymbol{\lambda}_{\mathcal{M}_{\boldsymbol{\theta}}(c)^{\rm c}}={\mathbf 0} \}$ and $\tilde \boldsymbol{\lambda}_{\boldsymbol{\theta}}=\arg \max _{\boldsymbol{\lambda} \in \Lambda_{\boldsymbol{\theta}}} f_n(\boldsymbol{\lambda}; \boldsymbol{\theta})$. Similar to (ref) and the arguments below (ref), if $\log r =o(n^{1/3})$, $\ell_n \nu^2=o(1)$ and $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$, we have $|\tilde \boldsymbol{\lambda}_{\boldsymbol{\theta},\mathcal{M}_{\boldsymbol{\theta}}(c)}|_2 \leq 4K_3^{-1} |\bar{{\mathbf g}}_{\mathcal{M}_{\boldsymbol{\theta}}(c)}(\boldsymbol{\theta})-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\boldsymbol{\theta}}(c)}(\boldsymbol{\theta})\}|_2$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1, where $K_3$ is specified in Condition (ref)(b). Notice that

align*[align* omitted — 969 chars of source]

By the Taylor expansion and Condition (ref)(c), we have $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1} |\bar{{\mathbf g}}(\boldsymbol{\theta}) - \bar{{\mathbf g}}(\hat{\boldsymbol{\theta}}_n)|_\infty = O_{\rm p}(\alpha_n)$, which implies $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})}(\boldsymbol{\theta}) - \bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})}(\hat{\boldsymbol{\theta}}_n)|_2 = O_{\rm p}(\ell_n^{1/2} \alpha_n)$. Due to $\alpha_n=o(\nu)$ and $|\bar{g}_j(\hat{\boldsymbol{\theta}}_n)| \geq \tilde{c} \nu \rho'(0^+)$ for any $j \in \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})$, we then have $ {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})}(\boldsymbol{\theta})\}={\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})}(\hat{\boldsymbol{\theta}}_n)\} $ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. By the triangle inequality and (ref), we have w.p.a.1 that

align*[align* omitted — 680 chars of source]

For any $j \in \mathcal{M}_{\boldsymbol{\theta}}(c) \bigcap \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})^{\rm c}$, we have $|\bar{g}_j(\boldsymbol{\theta})| \geq c \nu \rho'(0^+)$ and $|\bar{g}_j(\hat{\boldsymbol{\theta}}_n)| < \tilde{c} \nu \rho'(0^+)$. Due to $c \in (\tilde{c},1)$ and $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\bar{{\mathbf g}}(\boldsymbol{\theta}) - \bar{{\mathbf g}}(\hat{\boldsymbol{\theta}}_n)|_\infty = o_{\rm p}(\nu)$, then $ \mathcal{M}_{\boldsymbol{\theta}}(c) \bigcap \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})^{\rm c} = \emptyset$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1, which implies $T_{2,\boldsymbol{\theta}}=0$ for any $\boldsymbol{\theta}\in\mathcal{C}_1$ w.p.a.1. Hence,

align[align omitted — 319 chars of source]

Then $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\tilde \boldsymbol{\lambda}_{\boldsymbol{\theta}}|_2 = \sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\tilde \boldsymbol{\lambda}_{\boldsymbol{\theta},\mathcal{M}_{\boldsymbol{\theta}}(c)}|_2 = O_{\rm p}(\ell_n^{1/2} \alpha_n) =o_{\rm p}(\delta_n)$. Write $\tilde\boldsymbol{\lambda}_{\boldsymbol{\theta}}= (\tilde\lambda_{\boldsymbol{\theta},1},\ldots,\tilde\lambda_{\boldsymbol{\theta},r})^{\mathrm{\scriptscriptstyle \top} }$. Our next step is to show ${\mbox{\rm sgn}} (\tilde{\lambda}_{\boldsymbol{\theta},j})={\mbox{\rm sgn}}\{\bar{g}_j(\boldsymbol{\theta})\}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ and $j \in \mathcal{M}_{\boldsymbol{\theta}}(c)$ with $\tilde{\lambda}_{\boldsymbol{\theta},j} \neq 0$ w.p.a.1. Its proof is almost identical to that in Section (ref) for proving ${\mbox{\rm sgn}} (\tilde{\lambda}_{n,j})={\mbox{\rm sgn}}\{\bar{g}_j(\hat{\boldsymbol{\theta}}_n)\}$ for any $j \in \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})$ with $\tilde{\lambda}_{n,j} \neq 0$ w.p.a.1. We only need to replace $\{\tilde \boldsymbol{\lambda}_n,\mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})\}$ there by $\{\tilde \boldsymbol{\lambda}_{\boldsymbol{\theta}},\mathcal{M}_{\boldsymbol{\theta}}(c)\}$ and all the arguments still hold uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$. Using the same arguments stated in the proof of Lemma (ref) for showing $\tilde\boldsymbol{\lambda}_0$ is a local maximizer for $f_n(\boldsymbol{\lambda};\boldsymbol{\theta}_0)$ w.p.a.1, we can also prove $\tilde\boldsymbol{\lambda}_{\boldsymbol{\theta}}$ is a local maximizer of $f_n(\boldsymbol{\lambda}; \boldsymbol{\theta})$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1, which implies $\hat \boldsymbol{\lambda}(\boldsymbol{\theta}) = \tilde{\boldsymbol{\lambda}}_{\boldsymbol{\theta}}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. We then have Lemma (ref). $\Box$

Proof of Lemma (ref)

Recall $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})=\arg \max _{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})} f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ and $\mathcal{C}_1=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2\leq \alpha_n\}$. Then $\boldsymbol{\theta}$ and $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})=\{\hat{\lambda}_{1}(\boldsymbol{\theta}),\ldots,\hat{\lambda}_{r}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$ satisfy

align[align omitted — 295 chars of source]

where $\hat{\boldsymbol{\eta}}(\boldsymbol{\theta})=\{\hat{\eta}_{1}(\boldsymbol{\theta}),\ldots, \hat{\eta}_{r}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$ with $\hat{\eta}_j(\boldsymbol{\theta})=\nu \rho' \{|\hat{\lambda}_j(\boldsymbol{\theta})|;\nu\} {\mbox{\rm sgn}}\{\hat{\lambda}_j(\boldsymbol{\theta})\}$ for $\hat{\lambda}_j(\boldsymbol{\theta}) \neq 0$ and $\hat{\eta}_j(\boldsymbol{\theta}) \in [-\nu\rho'(0^{+}),\nu \rho'(0^{+})]$ for $\hat{\lambda}_j(\boldsymbol{\theta})=0$. Recall $\mathcal{R}(\boldsymbol{\theta})={\rm supp}\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})\}$. For any $\boldsymbol{\theta} \in \mathcal{C}_1$, restricted on $\mathcal{R}(\boldsymbol{\theta})$, define $$ {\mathbf m}_{\boldsymbol{\theta}}(\boldsymbol{\zeta},\boldsymbol{\vartheta})=\frac{1}{n}\sum _{i=1}^{n}\frac{{\mathbf g}_{i,{\mathcal{R}(\boldsymbol{\theta})}}(\boldsymbol{\vartheta})}{1+\boldsymbol{\zeta}^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,{\mathcal{R}(\boldsymbol{\theta})}}(\boldsymbol{\vartheta})}-{\mathbf w} $$ for any $\boldsymbol{\vartheta} \in \boldsymbol{\Theta}$ and $\boldsymbol{\zeta}=\{\zeta_{1},\ldots,\zeta_{|{\mathcal{R}(\boldsymbol{\theta})}|}\}^{\mathrm{\scriptscriptstyle \top} } \in {\mathbb{R}}^{|{\mathcal{R}(\boldsymbol{\theta})}|}$ with each $\zeta_j \neq 0$, where ${\mathbf w}=\{w_{1},\ldots,w_{|{\mathcal{R}(\boldsymbol{\theta})}|}\}^{\mathrm{\scriptscriptstyle \top} }$ with $w_j=\nu \rho' (|\zeta_j|;\nu) {\mbox{\rm sgn}}(\zeta_j)$. From (ref), we know $\hat{\boldsymbol{\lambda}}_{\mathcal{R}(\boldsymbol{\theta})} (\boldsymbol{\theta}) $ and $\boldsymbol{\theta}$ satisfy ${\mathbf m}_{\boldsymbol{\theta}}\{\hat{\boldsymbol{\lambda}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta}), \boldsymbol{\theta}\}= {\mathbf 0}$. By the implicit function theorem [Theorem 9.28 of Rudin1976], for all $\boldsymbol{\vartheta}$ in a small neighborhood of $\boldsymbol{\theta}$, denoted by $\mathcal{U}(\boldsymbol{\theta})$, there exists a $\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})$ such that ${\mathbf m}_{\boldsymbol{\theta}}\{\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta}),\boldsymbol{\vartheta}\}={\mathbf 0}$, $\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\theta})=\hat{\boldsymbol{\lambda}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})$ and $\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})$ is continuously differentiable in $\boldsymbol{\vartheta} \in \mathcal{U}(\boldsymbol{\theta})$. By Condition (ref)(a), the event ${\mathcal{E}} = \bigcap_{\boldsymbol{\theta} \in \mathcal{C}_1} \{\max_{j \in {\mathcal{R}(\boldsymbol{\theta})}^{\rm c}} |\hat{\eta}_j(\boldsymbol{\theta})| < \nu\rho'(0^+)\}$ holds w.p.a.1. Restricted on ${\mathcal{E}}$, let $\varsigma_n=\nu\rho'(0^+) - \sup_{\boldsymbol{\theta} \in \mathcal{C}_1}\max_{j \in {\mathcal{R}(\boldsymbol{\theta})}^{\rm c}} |\hat{\eta}_j(\boldsymbol{\theta})|$ and define $\boldsymbol{\Theta}_*(\boldsymbol{\theta})=\{\boldsymbol{\vartheta} \in \mathcal{U}(\boldsymbol{\theta}): |\boldsymbol{\vartheta}-\boldsymbol{\theta}|_1 \leq o[\min\{\varsigma_n, \chi_n(\boldsymbol{\theta})\}], |\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})-\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\theta})|_1 \leq o [\min\{\varsigma_n, \ell_n^{1/2} \alpha_n\}] \}$ for some $\chi_n(\boldsymbol{\theta})>0$. Since all the components of $\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\theta})$ are nonzero and $\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})$ is continuously differentiable in $\boldsymbol{\vartheta}$, we can select sufficiently small $\chi_n(\boldsymbol{\theta})$ such that all the components of $\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})$ are nonzero for any $\boldsymbol{\vartheta} \in \boldsymbol{\Theta}_*(\boldsymbol{\theta})$. For any $\boldsymbol{\vartheta} \in \boldsymbol{\Theta}_*(\boldsymbol{\theta})$, let $\tilde\boldsymbol{\lambda}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta}) \in {\mathbb{R}}^r$ satisfy $\tilde\boldsymbol{\lambda}_{\boldsymbol{\theta},{\mathcal{R}(\boldsymbol{\theta})}}(\boldsymbol{\vartheta})=\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})$ and $\tilde\boldsymbol{\lambda}_{\boldsymbol{\theta},{\mathcal{R}(\boldsymbol{\theta})}^{\rm c}}(\boldsymbol{\vartheta})={\mathbf 0}$. By Lemma (ref), $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})|_2 =O_{\rm p}(\ell_n^{1/2} \alpha_n)$ and $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\mathcal{R}(\boldsymbol{\theta})| \leq \ell_n$ w.p.a.1, which imply

align*[align* omitted — 621 chars of source]

Using the same arguments in the proof of Lemma (ref) for proving that $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})$ and $\boldsymbol{\theta}$ satisfy the score equation $\nabla_{\boldsymbol{\lambda}} f_n\{\tilde\boldsymbol{\lambda}(\boldsymbol{\theta}); \boldsymbol{\theta}\}={\mathbf 0}$ w.p.a.1 there, we can prove $\nabla_{\boldsymbol{\lambda}} f_n\{\tilde\boldsymbol{\lambda}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta}); \boldsymbol{\vartheta}\}={\mathbf 0}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ and $\boldsymbol{\vartheta} \in \boldsymbol{\Theta}_*(\boldsymbol{\theta})$ w.p.a.1. By the concavity of $f_n(\boldsymbol{\lambda};\boldsymbol{\vartheta})$ w.r.t $\boldsymbol{\lambda}$, we have $\tilde\boldsymbol{\lambda}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})=\hat{\boldsymbol{\lambda}}(\boldsymbol{\vartheta})=\arg\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\vartheta})} f_n(\boldsymbol{\lambda};\boldsymbol{\vartheta})$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ and $\boldsymbol{\vartheta} \in \boldsymbol{\Theta}_*(\boldsymbol{\theta})$ w.p.a.1. Recall $\tilde\boldsymbol{\lambda}_{\boldsymbol{\theta}}(\boldsymbol{\vartheta})$ is continuously differentiable in $\boldsymbol{\vartheta}\in\mathcal{U}(\boldsymbol{\theta})$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$. Hence, $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})$ is continuously differentiable at $\boldsymbol{\theta}$ and $[\nabla_{\boldsymbol{\theta}} \hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})]_{{\mathcal{R}(\boldsymbol{\theta})}^{\rm c},[p]}={\mathbf 0}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. Write $\hat{\boldsymbol{\lambda}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta}) = \{\tilde{\lambda}_1(\boldsymbol{\theta}),\ldots,\tilde{\lambda}_{|{\mathcal{R}(\boldsymbol{\theta})}|}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$. Since $\boldsymbol{\zeta}_{\boldsymbol{\theta}}(\boldsymbol{\theta})=\hat{\boldsymbol{\lambda}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})$, it holds that

align*[align* omitted — 2,115 chars of source]

We complete the proof of Lemma (ref). $\Box$

Proof of Lemma (ref)

Recall $\hat\boldsymbol{\lambda}(\boldsymbol{\theta})=\arg\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})}f_n(\boldsymbol{\lambda};\boldsymbol{\theta} )$, $\mathcal{R}_n={\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$ and $\mathcal{C}_1=\{\boldsymbol{\theta}\in\boldsymbol{\Theta}:|\boldsymbol{\theta}-\hat{\boldsymbol{\theta}}_n|_2\leq \alpha_n\}$. By Lemma (ref), $|\mathcal{R}_n|\leq \ell_n$ w.p.a.1. Select $\delta_n$ satisfying $\delta_n=o(\ell_n^{-1/2} n^{-1/\gamma})$ and $\ell_n^{1/2} \alpha_n=o(\delta_n) $, which can be guaranteed by $\ell_n\alpha_n=o(n^{-1/\gamma})$. For any $\boldsymbol{\theta} \in \mathcal{C}_1$, let $\tilde \boldsymbol{\lambda}(\boldsymbol{\theta})=\arg \max _{\boldsymbol{\lambda} \in \check\Lambda_n} f_n(\boldsymbol{\lambda}; \boldsymbol{\theta})$, where $\check\Lambda_n=\{\boldsymbol{\lambda}\in {\mathbb{R}}^{r}:|\boldsymbol{\lambda}_{ \mathcal{R}_n}|_2\leq \delta_n, \boldsymbol{\lambda}_{\mathcal{R}_n^{\rm c}}={\mathbf 0} \}$. Similar to (ref) and the arguments below (ref), if $\log r =o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n \nu^2=o(1)$, we have $|\tilde \boldsymbol{\lambda}_{\mathcal{R}_n}(\boldsymbol{\theta})|_2 \leq 4K_3^{-1} |\bar{{\mathbf g}}_{\mathcal{R}_n}(\boldsymbol{\theta})-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{\mathcal{R}_n}(\boldsymbol{\theta})\}|_2$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1, where $K_3$ is specified in Condition (ref)(b). By Lemma (ref), we have $\mathcal{R}_n \subset \mathcal{M}_{\hat{\boldsymbol{\theta}}_n}(\tilde{c})$ w.p.a.1, where $\tilde{c}$ is specified in Condition {\rm (ref)(a)}. Using the arguments for deriving (ref), we have $$ \sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\bar{{\mathbf g}}_{ \mathcal{R}_n}(\boldsymbol{\theta})-\nu \rho'(0^+) {\mbox{\rm sgn}}\{\bar{{\mathbf g}}_{ \mathcal{R}_n}(\boldsymbol{\theta})\}|_2 = O_{\rm p}(\ell_n^{1/2} \alpha_n)\,,$$ which implies $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\tilde \boldsymbol{\lambda}_{ \mathcal{R}_n}(\boldsymbol{\theta})|_2 = O_{\rm p}(\ell_n^{1/2} \alpha_n)=o_{\rm p}(\delta_n)$. Write $\tilde\boldsymbol{\lambda}_{\mathcal{R}_n}(\boldsymbol{\theta})=\{\dot\lambda_1(\boldsymbol{\theta}), \ldots, \dot\lambda_{|\mathcal{R}_n|}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$. By the first-order condition, for any $\boldsymbol{\theta} \in \mathcal{C}_1$, we have $$ {\mathbf 0} = \frac{1}{n}\sum_{i=1}^{n} \frac{{\mathbf g}_{i,\mathcal{R}_n}(\boldsymbol{\theta})}{1+\tilde \boldsymbol{\lambda}_{\mathcal{R}_n}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i,\mathcal{R}_n}(\boldsymbol{\theta})} - \tilde\boldsymbol{\eta}(\boldsymbol{\theta})\,,$$ where $\tilde\boldsymbol{\eta}(\boldsymbol{\theta})=\{\tilde\eta_1(\boldsymbol{\theta}), \ldots,\tilde\eta_{|\mathcal{R}_n|}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$ with $\tilde\eta_j(\boldsymbol{\theta})=\nu\rho'\{|\dot\lambda_j(\boldsymbol{\theta})|;\nu\}{\mbox{\rm sgn}}\{\dot\lambda_j(\boldsymbol{\theta})\}$ for $\dot{\lambda}_j(\boldsymbol{\theta})\neq 0$ and $\tilde\eta_j(\boldsymbol{\theta}) \in [-\nu\rho'(0^+), \nu\rho'(0^+)]$ for $\dot{\lambda}_j(\boldsymbol{\theta})= 0$. Using the same arguments for addressing the remainder terms in (ref), for any $\boldsymbol{\theta} \in \mathcal{C}_1$ and $j\in\mathcal{R}_n^{\rm c}$, it holds that $$ \frac{1}{n}\sum _{i=1}^{n} \frac{g_{i,j}(\boldsymbol{\theta})}{1+\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i}(\boldsymbol{\theta})} =\frac{1}{n}\sum _{i=1}^{n} \frac{g_{i,j}(\hat{\boldsymbol{\theta}}_n)}{1+\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i}(\hat{\boldsymbol{\theta}}_n)} + o_{\rm p}(\nu)\,,$$ where the term $o_{\rm p}(\nu)$ holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$ and $j \in \mathcal{R}_n^{\rm c}$. Together with Condition (ref)(a), we have that $$ \sup_{\boldsymbol{\theta} \in \mathcal{C}_1} \max_{j \in \mathcal{R}_n^{\rm c}}\bigg|\frac{1}{n}\sum _{i=1}^{n} \frac{g_{i,j}(\boldsymbol{\theta})}{1+\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} } {\mathbf g}_{i}(\boldsymbol{\theta})}\bigg| \leq \nu\rho'(0^+)$$ w.p.a.1. Therefore, $\tilde \boldsymbol{\lambda}(\boldsymbol{\theta})$ and $\boldsymbol{\theta}$ satisfy the score equation $\nabla_{\boldsymbol{\lambda}} f_n\{\tilde\boldsymbol{\lambda}(\boldsymbol{\theta}); \boldsymbol{\theta}\}={\mathbf 0}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. By the concavity of $f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ w.r.t $\boldsymbol{\lambda}$, it holds that $\tilde\boldsymbol{\lambda}(\boldsymbol{\theta})=\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})=\arg\max_{\boldsymbol{\lambda} \in \hat{\Lambda}_n(\boldsymbol{\theta})} f_n(\boldsymbol{\lambda};\boldsymbol{\theta})$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1, which implies ${\rm supp}\{\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})\} \subset {\rm supp}\{\hat{\boldsymbol{\lambda}}(\hat{\boldsymbol{\theta}}_n)\}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. Select $\boldsymbol{\theta}^* \in \mathcal{C}_1$ such that $|{\rm supp}\{\hat \boldsymbol{\lambda}(\boldsymbol{\theta}^*)\}| \leq |{\rm supp}\{\hat \boldsymbol{\lambda}(\boldsymbol{\theta})\}|$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$, and define $\mathcal{B}_2(\boldsymbol{\theta}^*,2\alpha_n) = \{\boldsymbol{\theta} \in\boldsymbol{\Theta}:\,|\boldsymbol{\theta}-\boldsymbol{\theta}^*|_2 \leq 2\alpha_n \}$. Using the same arguments above for proving ${\rm supp}\{\hat\boldsymbol{\lambda}(\boldsymbol{\theta})\} \subset {\rm supp}\{\hat \boldsymbol{\lambda}(\hat\boldsymbol{\theta}_n)\}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1, we have ${\rm supp} \{\hat\boldsymbol{\lambda}(\boldsymbol{\theta})\} \subset {\rm supp} \{\hat\boldsymbol{\lambda}(\boldsymbol{\theta}^*)\}$ for any $\boldsymbol{\theta} \in \mathcal{B}_2(\boldsymbol{\theta}^*,2\alpha_n)$ w.p.a.1. Since $\mathcal{C}_1 \subset \mathcal{B}_2(\boldsymbol{\theta}^*,2\alpha_n)$, then ${\rm supp} \{\hat\boldsymbol{\lambda}(\boldsymbol{\theta})\} \subset {\rm supp}\{\hat \boldsymbol{\lambda}(\boldsymbol{\theta}^*)\}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. Due to $|{\rm supp}\{\hat \boldsymbol{\lambda}(\boldsymbol{\theta}^*)\}| \leq |{\rm supp}\{\hat \boldsymbol{\lambda}(\boldsymbol{\theta})\}|$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$, we have ${\rm supp}\{\hat \boldsymbol{\lambda}(\boldsymbol{\theta})\}={\rm supp}\{\hat \boldsymbol{\lambda}(\boldsymbol{\theta}^*)\}$ for any $\boldsymbol{\theta} \in \mathcal{C}_1$ w.p.a.1. We complete the proof of Lemma (ref). $\Box$

Proof of Lemma (ref)

By Lemma (ref), we have $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\mathcal{R}(\boldsymbol{\theta})|\leq \ell_n$ w.p.a.1 and $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})|_2=O_{\rm p}(\ell_n^{1/2} \alpha_n)$. Under Condition (ref)(a), if $\ell_n\alpha_n=o(n^{-1/\gamma})$, then $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}\max _{ i \in [n]}|\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})^{\mathrm{\scriptscriptstyle \top} }{\mathbf g}_{i}(\boldsymbol{\theta})|=o_{\rm p}(1)$. Write $\hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})=\{\hat\lambda_1(\boldsymbol{\theta}),\ldots,\hat\lambda_r(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$ and ${\mathbf t}=(t_1,\ldots,t_p)^{\mathrm{\scriptscriptstyle \top} }$. For $T_{\boldsymbol{\theta},1}$, by the Cauchy-Schwarz inequality and Condition (ref)(c), we have

align*[align* omitted — 714 chars of source]

holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$ and ${\mathbf t} \in \mathbb{R}^p$. For $T_{\boldsymbol{\theta},3}$, by Condition (ref)(c), we have

align*[align* omitted — 946 chars of source]

holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$ and ${\mathbf t} \in \mathbb{R}^p$. Let $\hat\boldsymbol{\lambda}_{{\mathcal{R}(\boldsymbol{\theta})}}(\boldsymbol{\theta})=\{\tilde{\lambda}_1(\boldsymbol{\theta}),\ldots,\tilde{\lambda}_{|{\mathcal{R}(\boldsymbol{\theta})}|}(\boldsymbol{\theta})\}^{\mathrm{\scriptscriptstyle \top} }$. By the Cauchy-Schwarz inequality, Conditions (ref)(a) and (ref)(c), if $\log r=o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n\nu^2=o(1)$, then

align[align omitted — 1,917 chars of source]

holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$ and ${\mathbf t} \in \mathbb{R}^p$. Recall $\alpha_n=o(\nu)$. By Proposition (ref), Lemma (ref) and Condition (ref)(b), if $\log r =o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n \nu^2=o(1)$, we know that $\inf_{\boldsymbol{\theta} \in \mathcal{C}_1}\lambda_{\min}\{\widehat{{\mathbf V}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})\}$ and $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}\lambda_{\max}\{\widehat{{\mathbf V}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})\}$ are uniformly bounded away from zero and infinity w.p.a.1. Using the same arguments in the proof of Lemma (ref), if $\log r=o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n\nu^2=o(1)$, it holds that

align[align omitted — 1,114 chars of source]

By Condition (ref) and the same arguments in the proof of Lemma (ref), if $\log r=o(n^{1/3})$, $\ell_n\alpha_n=o[\min\{\nu,n^{-1/\gamma}\}]$ and $\ell_n\nu^2=o(1)$, we have $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}\|\widehat{\boldsymbol{\Gamma}}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})\|_2=O_{\rm p}(1)$. Notice that $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}\|\widehat{{\mathbf V}}^{-1}_{\mathcal{R}(\boldsymbol{\theta})}(\boldsymbol{\theta})\|_2=O_{\rm p}(1)$ and $ \sup_{\boldsymbol{\theta} \in \mathcal{C}_1}\|\nu {\rm diag}[\rho''\{|\tilde\lambda_{1}(\boldsymbol{\theta})|;\nu \},\ldots,\rho''\{|\tilde\lambda_{|{\mathcal{R}(\boldsymbol{\theta})}|}(\boldsymbol{\theta})|;\nu \}]\|_2=O_{\rm p}(\nu) $. Thus,

align[align omitted — 781 chars of source]

Combining (ref), (ref) and (ref), by Lemma (ref), we know $\sup_{\boldsymbol{\theta} \in \mathcal{C}_1}|[\nabla_{\boldsymbol{\theta}} \hat{\boldsymbol{\lambda}}(\boldsymbol{\theta})]_{{\mathcal{R}(\boldsymbol{\theta})},[p]} {\mathbf t}|_2 = |{\mathbf t}|_2 \cdot O_{\rm p}(1)$ holds uniformly over ${\mathbf t} \in \mathbb{R}^p$, which implies

align*[align* omitted — 984 chars of source]

holds uniformly over ${\mathbf t} \in \mathbb{R}^p$. For $T_{\boldsymbol{\theta},4}$, by (ref), (ref) and (ref), Lemma (ref) implies

align*[align* omitted — 1,096 chars of source]

holds uniformly over $\boldsymbol{\theta} \in \mathcal{C}_1$ and ${\mathbf t} \in \mathbb{R}^p$. We then obtain the result by (ref). $\hfill\Box$

Proof of Lemma (ref)

Denote by ${\mathbb{P}}_{\mathcal{X}_n}(\cdot)$ and $\mathbb{E}_{\mathcal{X}_n}(\cdot)$, respectively, the conditional probability and conditional expectation given $\mathcal{X}_n$. For any integer $k \geq 1$, recall $\hat{\boldsymbol{\zeta}}_{k+1} = N_{k}^{-1} \sum_{i=1}^{N_{k}} \omega^{k}_i {\mathbf h}(\boldsymbol{\theta}^{k}_i)$ only depends on the $N_k$ samples $\{\boldsymbol{\theta}^{k}_1,\ldots,\boldsymbol{\theta}^{k}_{N_k}\}$ generated from the proposal distribution with density $\varphi(\boldsymbol{\theta}\,;\hat{\boldsymbol{\zeta}}_{k})$, where $\omega^{k}_i = \pi^\dag(\boldsymbol{\theta}^{k}_i\,|\,\mathcal{X}_n) / \varphi(\boldsymbol{\theta}^{k}_i\,;\hat{\boldsymbol{\zeta}}_{k})$. Thus, the random sequence $\{\hat{\boldsymbol{\zeta}}_{k}\}_{k\geq1}$ forms a Markov chain. Recall that $\boldsymbol{\Theta} \subset {\mathbb{R}}^p$ is a compact set with fixed $p$. Since $\sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} |{\mathbf h}(\boldsymbol{\theta})|_\infty \leq K_9$ for some universal constant $K_9>0$, and $\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})$ is positive and continuous on $(\boldsymbol{\theta}, \boldsymbol{\zeta}) \in \boldsymbol{\Theta} \times {\mathbb{R}}^s$, there exists a positive and continuous function $\varrho(\cdot)$ such that $$ \sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} \frac{\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)|{\mathbf h}(\boldsymbol{\theta})|_\infty}{\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})} \leq \varrho(\boldsymbol{\zeta}) $$ for any $\boldsymbol{\zeta} \in {\mathbb{R}}^s$. Since $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) = 0$ for any $\boldsymbol{\theta} \notin \boldsymbol{\Theta}$, then $$ \sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}^{\rm c}} \frac{\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)|{\mathbf h}(\boldsymbol{\theta})|_\infty}{\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})} = 0 $$ for any $\boldsymbol{\zeta} \in {\mathbb{R}}^s$. Notice that ${\mathbb{E}}(\hat{\boldsymbol{\zeta}}_{k+1} \,|\, \hat{\boldsymbol{\zeta}}_{k}) = \boldsymbol{\zeta}^* = {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}\{{\mathbf h}(\boldsymbol{\theta})\}$. Write $\hat{\boldsymbol{\zeta}}_{k+1} = (\hat{\zeta}_{k+1, 1}, \ldots, \hat{\zeta}_{k+1, s})^{\mathrm{\scriptscriptstyle \top} }$ and $\boldsymbol{\zeta}^* = (\zeta^*_1,\ldots, \zeta^*_s)^{\mathrm{\scriptscriptstyle \top} } $. For any $\varepsilon>0$, by the Hoeffding's inequality, we have

align[align omitted — 410 chars of source]

Let $C_\varepsilon = \sup_{\boldsymbol{\zeta}\in{\mathbb{R}}^s:\, |\boldsymbol{\zeta} - \boldsymbol{\zeta}^*|_\infty \leq \varepsilon} \varrho^2(\boldsymbol{\zeta})$. By (ref), we have

align[align omitted — 714 chars of source]

which implies

align*[align* omitted — 276 chars of source]

For any integer $k'\geq k$, by the Markov property of $\{\hat{\boldsymbol{\zeta}}_{k}\}_{k\geq1}$, it then holds that

align*[align* omitted — 727 chars of source]

Letting $k' \rightarrow \infty$, then $$ {\mathbb{P}}_{\mathcal{X}_n}\bigg(\bigcap^\infty_{t= k} \{|\hat{\boldsymbol{\zeta}}_{t+1} - \boldsymbol{\zeta}^*|_\infty \leq \varepsilon\} \bigg) \geq {\mathbb{P}}_{\mathcal{X}_n}\big(|\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty \leq \varepsilon\big) \prod^\infty_{t= k+1} \bigg\{1 - 2s\exp\bigg(-\frac{N_{t}\varepsilon^2}{2 C_\varepsilon}\bigg) \bigg\} \,. $$ Since $s$ is fixed and $\sum_{k=1}^{\infty}\exp(-CN_k) < \infty$ for any $C>0$, we have

align[align omitted — 372 chars of source]

For any $z>0$, let $\bar{C}_z = \sup_{\boldsymbol{\zeta}\in {\mathbb{R}}^s:\,|\boldsymbol{\zeta}|_\infty \leq z} \varrho^2(\boldsymbol{\zeta})$. Using the same arguments for (ref), it holds that $$ {\mathbb{P}}_{\mathcal{X}_n}\big(|\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty > \varepsilon, |\hat{\boldsymbol{\zeta}}_{k}|_\infty \leq z\big) \leq 2s\exp\bigg(-\frac{N_{k}\varepsilon^2}{2C_z}\bigg) {\mathbb{P}}_{\mathcal{X}_n}\big(|\hat{\boldsymbol{\zeta}}_{k}|_\infty \leq z\big) \leq 2s\exp\bigg(-\frac{N_{k}\varepsilon^2}{2C_z}\bigg) \,. $$ By the Markov's inequality and triangle inequality,

align[align omitted — 766 chars of source]

which implies ${\mathbb{P}}_{\mathcal{X}_n}(|\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty > \varepsilon) \leq 2s\exp\{-(2C_z)^{-1}N_{k}\varepsilon^2\} + z^{-1}{\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}\{|{\mathbf h}(\boldsymbol{\theta})|_\infty\}$. Due to $N_k \rightarrow \infty$ as $k \rightarrow \infty$, we know $\limsup_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}(|\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty > \varepsilon) \leq z^{-1}{\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}\{|{\mathbf h}(\boldsymbol{\theta})|_\infty\}$. Notice that $\sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} |{\mathbf h}(\boldsymbol{\theta})|_\infty \leq K_9$ for some universal constant $K_9>0$. Letting $z \rightarrow \infty$, it holds that $\limsup_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}(|\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty > \varepsilon) = 0$, which implies $\liminf_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}(|\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty \leq \varepsilon) = 1$. Together with (ref), we have $$ \liminf_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}\bigg(\bigcap^\infty_{t = k} \{|\hat{\boldsymbol{\zeta}}_{t+1} - \boldsymbol{\zeta}^*|_\infty \leq \varepsilon\}\bigg) = 1 \,. $$ Since $\varepsilon>0$ is arbitrary, we then obtain that, conditional on $\mathcal{X}_n$, $|\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$. We complete the proof of Lemma (ref). $\Box$

Proof of Lemma (ref)

Denote by ${\mathbb{P}}_{\mathcal{X}_n}(\cdot)$ and $\mathbb{E}_{\mathcal{X}_n}(\cdot)$, respectively, the conditional probability and conditional expectation given $\mathcal{X}_n$. Recall $\boldsymbol{\zeta}^* = {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}\{{\mathbf h}(\boldsymbol{\theta})\}$. Let $ \widehat{{\mathbf Z}}_{k+1}= N_k^{-1} \sum_{i=1}^{N_k} \boldsymbol{\theta}^k_i \pi^\dag(\boldsymbol{\theta}^k_i\,|\,\mathcal{X}_n) / \varphi(\boldsymbol{\theta}^k_i\,;\boldsymbol{\zeta}^*)$ for any integer $k \geq 2$ and $$ {\mathbf Z}(\boldsymbol{\zeta})={\mathbb{E}}_{\boldsymbol{\theta} \sim\varphi(\cdot\,;\,\boldsymbol{\zeta})} \bigg\{\frac{\boldsymbol{\theta} \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)}{\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta}^*)} \bigg\} = \int_{{\mathbb{R}}^p} \frac{\boldsymbol{\theta} \pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)}{\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta}^*)} \varphi(\boldsymbol{\theta}\,;\boldsymbol{\zeta}) \,{\rm d} \boldsymbol{\theta} $$ for any $\boldsymbol{\zeta} \in {\mathbb{R}}^s$, where $\{\boldsymbol{\theta}^1_1, \ldots, \boldsymbol{\theta}^1_{N_1}, \ldots,\boldsymbol{\theta}^k_1,\ldots, \boldsymbol{\theta}^k_{N_k}\}$ are generated via Algorithm {\rm (ref)}.

Our first step is to show that conditional on $\mathcal{X}_n$, we have $|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k})|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$. Notice that $\boldsymbol{\Theta} \subset {\mathbb{R}}^p$ is a compact set with fixed $p$ and $\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n) =0$ for any $\boldsymbol{\theta} \notin \boldsymbol{\Theta}$. Since $\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})$ is positive and continuous on $(\boldsymbol{\theta}, \boldsymbol{\zeta}) \in \boldsymbol{\Theta} \times {\mathbb{R}}^s$, we know

align[align omitted — 460 chars of source]

for some universal constant $\tilde{C}>0$. Recall that $\widehat{{\mathbf Z}}_{k+1}$ depends on the $N_k$ samples $\{\boldsymbol{\theta}^k_1,\ldots,\boldsymbol{\theta}^k_{N_k}\}$ generated from the proposal distribution with density $\varphi(\boldsymbol{\theta}\,;\hat{\boldsymbol{\zeta}}_{k})$. Then ${\mathbb{E}}(\widehat{{\mathbf Z}}_{k+1} \,|\, \hat{\boldsymbol{\zeta}}_{k}) = {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k})$. For any $\varepsilon>0$, using the same arguments for (ref), we have

align[align omitted — 267 chars of source]

Define the event $A_{k+1} = \{|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k})|_\infty \leq \varepsilon, |\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty \leq \varepsilon\}$. Note that $A_{k} \in \sigma(\boldsymbol{\theta}_1^{k-1},\ldots, \boldsymbol{\theta}_{N_{k-1}}^{k-1}, \hat{\boldsymbol{\zeta}}_{k-1})$ and the conditional joint distribution of $(\widehat{{\mathbf Z}}_{k+1},\hat{\boldsymbol{\zeta}}_{k+1})$ given $\mathcal{X}_n$ is fully determined by $\hat{\boldsymbol{\zeta}}_{k}$. By (ref) and (ref), it holds that

align*[align* omitted — 1,484 chars of source]

where $C_\varepsilon = \sup_{\boldsymbol{\zeta}\in{\mathbb{R}}^s:\, |\boldsymbol{\zeta} - \boldsymbol{\zeta}^*|_\infty \leq \varepsilon} \varrho^2(\boldsymbol{\zeta})$ with the function $\varrho(\cdot)$ specified in the proof of Lemma (ref). Then $$ {\mathbb{P}}_{\mathcal{X}_n}\big(A^{\rm c}_{k+1} \,\big|\, A_{k} \big) \leq 2p\exp\bigg(-\frac{N_{k}\varepsilon^2}{2\tilde{C}^2} \bigg) + 2s\exp\bigg(-\frac{N_{k}\varepsilon^2}{2C_\varepsilon} \bigg) \leq 2(p+s)\exp\bigg(-\frac{N_{k}\varepsilon^2}{\check{C}_\varepsilon} \bigg) $$ for some $\check{C}_\varepsilon >0$ depending on $\varepsilon$. For any integer $k'\geq k$, by the Markov property of $\{(\widehat{{\mathbf Z}}_{k}, \hat{\boldsymbol{\zeta}}_{k})\}_{k\geq 2}$, it then holds that

align*[align* omitted — 355 chars of source]

Letting $k' \rightarrow \infty$, then $$ {\mathbb{P}}_{\mathcal{X}_n}\bigg(\bigcap^\infty_{t=k} A_{t+1} \bigg) \geq {\mathbb{P}}_{\mathcal{X}_n}(A_{k+1}) \prod^\infty_{t = k+1} \bigg\{1 - 2(p+s)\exp \bigg(-\frac{N_{t}\varepsilon^2}{\check{C}_\varepsilon}\bigg) \bigg\} \,. $$ Since $p$ and $s$ are fixed constants and $\sum_{k=1}^{\infty}\exp(-CN_k) < \infty$ for any $C>0$, we have

align[align omitted — 736 chars of source]

where the last step is due to the fact $\limsup_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}(|\hat{\boldsymbol{\zeta}}_{k+1} - \boldsymbol{\zeta}^*|_\infty > \varepsilon) = 0$ as shown in the proof of Lemma (ref). For any $z>0$, by (ref), it holds that

align*[align* omitted — 698 chars of source]

Together with (ref), we have $$ {\mathbb{P}}_{\mathcal{X}_n}\{|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k})|_\infty > \varepsilon\} \leq 2p\exp\bigg(-\frac{N_{k}\varepsilon^2}{2\tilde{C}^2} \bigg) + z^{-1}{\mathbb{E}}_{\boldsymbol{\theta} \sim\pi^\dag}\{|{\mathbf h}(\boldsymbol{\theta})|_\infty\} \,. $$ Due to $N_k \rightarrow \infty$ as $k \rightarrow \infty$, we know $\limsup_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}\{|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k})|_\infty > \varepsilon\} \leq z^{-1}{\mathbb{E}}_{\boldsymbol{\theta} \sim\pi^\dag}\{|{\mathbf h}(\boldsymbol{\theta})|_\infty\}$. Notice that $\sup_{\boldsymbol{\theta} \in \boldsymbol{\Theta}} |{\mathbf h}(\boldsymbol{\theta})|_\infty \leq K_9$ for some universal constant $K_9>0$. Letting $z \rightarrow \infty$, it holds that $\limsup_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}\{|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k})|_\infty > \varepsilon\} = 0$. Together with (ref), we have $$ \liminf_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}\bigg[\bigcap^\infty_{t=k} \big\{ |\widehat{{\mathbf Z}}_{t+1} - {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{t})|_\infty \leq \varepsilon \big\}\bigg] \geq \liminf_{k \rightarrow \infty} {\mathbb{P}}_{\mathcal{X}_n}\bigg(\bigcap^\infty_{t=k} A_{t+1} \bigg) =1 \,. $$ Since $\varepsilon>0$ is arbitrary, conditional on $\mathcal{X}_n$, we have $|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k})|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$.

Our second step is to show that conditional on $\mathcal{X}_n$, we have $|{\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k}) - {\mathbf Z}(\boldsymbol{\zeta}^*)|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$. By (ref), we have $|{\mathbf Z}(\boldsymbol{\zeta})|_\infty < \infty$ for any $\boldsymbol{\zeta} \in {\mathbb{R}}^s$, and $$ |{\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k}) - {\mathbf Z}(\boldsymbol{\zeta}^*)|_\infty \leq \int_{{\mathbb{R}}^p} \frac{\pi^\dag(\boldsymbol{\theta}\,|\,\mathcal{X}_n)|\boldsymbol{\theta}|_\infty}{\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta}^*)} |\varphi(\boldsymbol{\theta}\,; \hat{\boldsymbol{\zeta}}_{k}) - \varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta}^*)| \,{\rm d}\boldsymbol{\theta} \leq \tilde{C}\int_{\boldsymbol{\Theta}}|\varphi(\boldsymbol{\theta}\,; \hat{\boldsymbol{\zeta}}_{k}) - \varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta}^*)| \,{\rm d}\boldsymbol{\theta} \,. $$ For some sufficiently large $M>0$, since conditional on $\mathcal{X}_n$, we have $|\hat{\boldsymbol{\zeta}}_{k} - \boldsymbol{\zeta}^*|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$, then for any $\epsilon >0$ there exists a sufficiently large integer $k_\epsilon$ such that $ {\mathbb{P}}_{\mathcal{X}_n}(\mathcal{A}) \leq \epsilon $ with $\mathcal{A} = \bigcup_{t=k_\epsilon}^\infty \{|\hat{\boldsymbol{\zeta}}_{t} - \boldsymbol{\zeta}^*|_\infty > M\}$. Define a compact set $\mathcal{B}=\{\boldsymbol{\zeta} \in {\mathbb{R}}^s: |\boldsymbol{\zeta} - \boldsymbol{\zeta}^*|_\infty \leq M \} $. Recall $\boldsymbol{\Theta} \subset {\mathbb{R}}^p$ is a compact set with fixed $p$. Due to the continuity of $\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})$, we know $\varphi(\boldsymbol{\theta}\,; \boldsymbol{\zeta})$ is uniformly continuous on $(\boldsymbol{\theta}\,; \boldsymbol{\zeta}) \in \boldsymbol{\Theta} \times \mathcal{B}$. For any $\varepsilon>0$, there exists $\delta(\varepsilon)>0$ such that $|\varphi(\boldsymbol{\theta}_1\,; \boldsymbol{\zeta}_1) - \varphi(\boldsymbol{\theta}_2\,; \boldsymbol{\zeta}_2)| < \tilde{C}^{-1} \varepsilon / \mathbb{L}(\boldsymbol{\Theta})$ for any $(\boldsymbol{\theta}_1, \boldsymbol{\zeta}_1), (\boldsymbol{\theta}_2, \boldsymbol{\zeta}_2) \in \boldsymbol{\Theta} \times \mathcal{B}$ satisfying $|\boldsymbol{\theta}_1 - \boldsymbol{\theta}_2|_\infty \leq \delta(\varepsilon)$ and $|\boldsymbol{\zeta}_1 - \boldsymbol{\zeta}_2|_\infty \leq \delta(\varepsilon)$, where $\mathbb{L}(\cdot)$ is the Lebesgue measure on ${\mathbb{R}}^p$. Since

align*[align* omitted — 650 chars of source]

we then have

align*[align* omitted — 522 chars of source]

where the second step is due to the fact that conditional on $\mathcal{X}_n$ we have $|\hat{\boldsymbol{\zeta}}_{k} - \boldsymbol{\zeta}^*|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$. Letting $\epsilon \to 0$, we know that conditional on $\mathcal{X}_n$, we have $|{\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k}) - {\mathbf Z}(\boldsymbol{\zeta}^*)|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$.

Our third step is to show that conditional on $\mathcal{X}_n$, we have $|\widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(\boldsymbol{\theta}) - {\mathbb{E}}_{\boldsymbol{\theta} \sim \pi^\dag}(\boldsymbol{\theta})|_\infty \rightarrow 0$ almost surely as $K \rightarrow \infty$. By the triangle inequality, $|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\boldsymbol{\zeta}^*)|_\infty \leq |\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k})|_\infty + |{\mathbf Z}(\hat{\boldsymbol{\zeta}}_{k}) - {\mathbf Z}(\boldsymbol{\zeta}^*)|_\infty $. Based on the results shown in Steps 1 and 2 above, it holds that conditional on $\mathcal{X}_n$, we have $|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\boldsymbol{\zeta}^*)|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$. Notice that ${\mathbf Z}(\boldsymbol{\zeta}^*)={\mathbb{E}}_{\boldsymbol{\theta} \sim\pi^\dag}(\boldsymbol{\theta})$ and $$ \widehat{{\mathbb{E}}}^*_{\pi^\dag,K}(\boldsymbol{\theta}) = \frac{1}{S_K} \sum_{k=1}^{K} \sum_{i=1}^{N_k} \frac{\pi^\dag(\boldsymbol{\theta}^k_i\,|\,\mathcal{X}_n)}{\varphi(\boldsymbol{\theta}^k_i\,;\boldsymbol{\zeta}^*) } \boldsymbol{\theta}^k_i = \frac{1}{S_K} \sum_{k=1}^{K} N_k \widehat{{\mathbf Z}}_{k+1} $$ with $S_K=N_1+\cdots+N_K$. Notice that conditional on $\mathcal{X}_n$, $|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\boldsymbol{\zeta}^*)|_\infty \rightarrow 0$ almost surely as $k \rightarrow \infty$. Given a constant $\varepsilon>0$, for any $\epsilon>0$ there exists a sufficiently large integer $\tilde{k}_\epsilon$ such that ${\mathbb{P}}_{\mathcal{X}_n}(\mathcal{C}) \leq \epsilon$ with $\mathcal{C} = \bigcup_{k=\tilde{k}_\epsilon}^\infty\{|\widehat{{\mathbf Z}}_{k+1} - {\mathbf Z}(\boldsymbol{\zeta}^*)|_