The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
114,671 characters
Uniform Inference for Parameters Identified by Conditional Quantile Restrictions
\maketitle
\begin{abstract}
Many structural and dynamic economic models imply that key parameters are identified by conditional quantile restrictions. Building on the exponential-weighting approach of Bierens (1990) and recent advances in penalized maximum statistics for conditional moment restrictions (Chen et al., 2025), we develop a unified inference framework for such parameters. We propose an adaptive $\ell_1$-penalized supremum statistic that transforms the conditional restriction into a continuum of unconditional moment conditions and aggregates evidence across quantile indices. The penalty regularizes the maximization over the weighting direction. Under the stated uniformity conditions, the known-parameter adaptive selector has no lower maximin local power than the unpenalized test and yields a strict maximin local-power gain whenever some positive candidate penalty has a strictly larger population maximin criterion than the zero penalty. We extend the theory to settings with pre-estimated nuisance parameters, characterizing the additional terms induced by the plug-in step in the limiting process. We derive an analytically corrected variance estimator that accounts for plug-in estimation
uncertainty and establish the validity of a Gaussian multiplier bootstrap under the null and sequences of local alternatives. Monte Carlo simulations show that the proposed CvM--KS aggregation scheme has rejection rates relatively close to nominal size under pre-estimation in the linear design and approaches the nominal level in nonlinear designs as the sample size grows. The reported power comparisons between the adaptive and unpenalized procedures are design-specific, with higher empirical rejection frequencies for the adaptive procedure along some directions.
\\[1em]
\textbf{Keywords:} Conditional quantile restriction, parameter inference, adaptive penalization, local power, pre-estimation effect, multiplier bootstrap.
\end{abstract}
\section{Introduction}
\label{sec:introduction}
Conditional quantile methods describe heterogeneous and asymmetric economic
dynamics beyond the conditional mean. Since \citet{koenker1978}, these methods
have been widely used in empirical finance and macroeconomics; see \citet{koenker2005} for a general treatment. In time-series applications, parametric dynamic conditional quantile models and quantile autoregressions capture state-dependent persistence and nonlinear adjustment; see, among others, \citep{koenkerxiao2006}. These applications require valid tests and confidence sets for quantile-indexed parameters, often uniformly over a range of quantiles rather than at a single quantile. This paper develops inference procedures for parameters defined through conditional quantile restrictions. Specifically, we consider testing hypotheses on $\theta_0(\alpha)$ for a given $\alpha$ and, more generally, uniformly over $\alpha \in \mathcal{T}\subset(0,1)$. The true parameter $\theta_0(\alpha)$ is characterized by the conditional moment restriction
\begin{equation}\label{eq:cond_moment}
\mathbb{E}\big[\mathbf{1}\{Y_t \le m(I_{t-1}, \theta_0(\alpha))\} - \alpha \mid I_{t-1}\big] = 0 \quad \forall \alpha \in \mathcal{T},
\end{equation}
where $Y_t$ is the response variable, $I_{t-1}$ denotes the conditioning information set, and $m(\cdot,\theta)$ is a parametric conditional quantile function. Restriction \eqref{eq:cond_moment} is non-smooth due to the indicator function, and its enforcement over a continuum of conditioning events makes both estimation and inference nontrivial.
The literature on testing quantile restrictions has developed along several related lines. One line studies the quantile regression process and develops inference that is uniform across quantile levels. For example, \citet{koenkerxiao2002} develop inference based on the quantile regression process using a Khmaladze transformation, while \citet{chernozhukov2005} use subsampling to construct uniform inference procedures. Related work studies structural change in regression quantiles; \citet{okaqu2011}, for instance, develop methods for estimating multiple structural changes in conditional quantile functions.
These contributions are central to inference on the quantile index, but their main focus differs from testing conditional restrictions by searching over a rich class of conditioning directions.
A second line of work follows the integrated conditional moment (ICM) approach.
This literature transforms a conditional moment restriction into a continuum of unconditional moments indexed by test functions or instruments, and then aggregates the resulting empirical process. The exponential-weighting test of \citet{bierens1990} is a leading example. In dynamic conditional quantile models, \citet{escanciano2010} develop specification tests for parametric dynamic conditional quantiles, while \citet{escancianogoh2014} study specification analysis of linear quantile models using marked empirical processes. These papers motivate the use of indexed empirical processes to check conditional quantile restrictions, although their objects and aggregation schemes differ from those of the penalized uniform-in-quantile statistic studied here.
Our statistic combines these two perspectives. We construct a studentized empirical process indexed jointly by the quantile level $\alpha$ and by an exponential-weight direction $\gamma\in\Gamma$. The CvM--KS functional integrates over $\alpha$ and maximizes over $\gamma$, so it can detect both variation across quantiles and departures along conditioning directions. Since the null distribution is generally non-pivotal, we use a Gaussian multiplier bootstrap for feasible inference. The required weak convergence and bootstrap arguments are standard empirical-process arguments for indexed moment processes, but the present setting requires tracking quantile-index integration, directional maximization, and studentization.
To target local power along informative conditioning directions, we introduce an adaptive $\ell_1$-penalized maximum statistic. The statistic subtracts $\lambda\|\gamma\|_1$ inside the supremum over $\gamma$, generating a class of tests indexed by $\lambda\ge0$. The penalty regularizes the search over exponential-weight directions. Under the stated conditions, the known-parameter adaptive selector has no lower maximin local power than the unpenalized test because the candidate set contains $\lambda=0$. It yields a strict maximin local-power gain whenever the population maximin criterion is strictly larger at some positive candidate penalty than at zero. This tuning rule is closely related to the penalized generalized Bierens statistic of \citet{chen2025}, who use a data-dependent max--min rule to choose tuning parameters for conditional moment restrictions. Our contribution is to develop this local-power tuning idea in a uniform conditional-quantile setting with CvM aggregation over quantile indices and with an explicit treatment of pre-estimated nuisance parameters. The guarantee is maximin and does not imply pointwise power dominance for every local direction.
We also address the theoretical issues that arise when nuisance parameters are pre-estimated. In many empirical settings, inference is conducted on a subset of structural parameters, with other components replaced by preliminary estimates. Such plug-in steps can change the limiting covariance structure and can introduce additional terms in the limiting process, a phenomenon already visible in classical empirical-process results with estimated parameters; see \citet{durbin1973}. In our setting, we characterize the plug-in contribution to the quantile empirical process, derive a corrected variance estimator, and establish multiplier-bootstrap validity under the null and under sequences of local alternatives. We also compare aggregation orders under pre-estimation. In the reported designs, the CvM--KS statistic $T_n$ (CvM aggregation over $\alpha$ followed by KS aggregation over $\gamma$) has rejection rates closer to nominal size than the alternative aggregation orders considered below.
The remainder of the paper is organized as follows.
Section~\ref{sec:test_stat} introduces the test statistic and its asymptotic null theory;
Section~\ref{sec:bootstrap_power} describes multiplier-bootstrap inference and local-power analysis;
Section~\ref{sec:pre_estimation} develops the pre-estimation extension;
Sections~\ref{sec:simulation} and \ref{sec:empirical} report Monte Carlo and empirical results;
Section~\ref{sec:conclusion} concludes. Proofs are collected in the Supplementary Appendix.
\section{Test Statistics and Asymptotic Theory}
\label{sec:test_stat}
This section introduces the penalized Bierens-type statistic for conditional
quantile restrictions and the objects used for inference. The formal
regularity conditions are stated in the Supplementary Appendix. Under these
conditions, we establish weak convergence of the empirical process and the
resulting null limit distribution.
\subsection{The Penalized Bierens Test Statistic}
We are interested in testing whether the true parameter function $\theta_0(\cdot)$ equals a prespecified value $\bar{\theta}(\cdot)$:
\[
H_0: \theta_0(\alpha) = \bar{\theta}(\alpha) \;\text{ for all }\; \alpha \in \mathcal{T}
\qquad \text{vs.} \qquad
H_1: \theta_0(\alpha) \neq \bar{\theta}(\alpha) \;\text{ for some }\; \alpha \in \mathcal{T}.
\]
In Sections~\ref{sec:test_stat}--\ref{sec:bootstrap_power},
$\bar{\theta}(\cdot)$ is a fixed, nonrandom null-imposed curve, such as
$\bar{\theta}(\cdot)\equiv0$ in a significance test or a fixed candidate in a
confidence-set inversion. A null-imposed curve estimated from the same sample is
covered only after its influence-function contribution is included, as in
Section~\ref{sec:pre_estimation}. Testing $H_0$ via \eqref{eq:cond_moment} is
challenging because it corresponds to a continuum of unconditional moment
conditions; ICM methods address this via weighting and integration. Following the
exponential weighting approach of \citet{bierens1990} and the generalized
penalization framework of \citet{chen2025}, we transform the conditional
restriction into an equivalent continuum of unconditional ones using the weighting
function $w(I_{t-1},\gamma)=\exp(\gamma^\top\Phi(I_{t-1}))$. Here,
$\Phi:\mathbb{R}^{d_I}\to\mathbb{R}^{d_I}$ is a one-to-one transformation chosen
so that $\exp(\gamma^\top\Phi(I_{t-1}))$ has finite moments uniformly over
$\gamma\in\Gamma$ (e.g., identity for well-behaved covariates, or a bounded map
such as componentwise $\tan^{-1}(\cdot)$ for heavy-tailed covariates). For
notational simplicity, below we write $\exp(\gamma^\top I_{t-1})$ for
$\exp\{\gamma^\top\Phi(I_{t-1})\}$. Occurrences of $I_{t-1}$ in the conditional
quantile function $m(I_{t-1},\theta)$ continue to refer to the original
conditioning information.
Define the empirical process indexed by the quantile level $\alpha$ and the
weighting direction $\gamma$:
\begin{equation}
M_n(\alpha, \gamma)
:=
\frac{1}{n}\sum_{t=1}^{n}
\exp(\gamma^{\top} I_{t-1})
\big[\mathbf{1}\{Y_t \le m(I_{t-1}, \bar{\theta}(\alpha))\} - \alpha\big],
\label{eq:Mn}
\end{equation}
where $\gamma \in \Gamma$, a compact subset of $\mathbb{R}^{d_I}$ containing the origin. To studentize the process, we scale $M_n(\alpha,\gamma)$ by its null standard deviation $s_n(\alpha,\gamma)$, which is available in closed form under \eqref{eq:cond_moment}. The normalization term is given by the sample analog of the theoretical variance:
\begin{equation}
s_n^2(\alpha, \gamma)
:=
\frac{\alpha(1-\alpha)}{n}\sum_{t=1}^{n}\exp(2\gamma^{\top} I_{t-1}).
\label{eq:sn}
\end{equation}
The closed-form expression in \eqref{eq:sn} exploits the martingale difference
sequence (MDS) property of the centered indicator $\mathbf{1}\{Y_t \le m(I_{t-1},
\bar{\theta}(\alpha))\} - \alpha$ under $H_0$ (Assumption~S1(iii)).
Specifically, the conditional variance of this term given $I_{t-1}$ equals
$\alpha(1-\alpha)$ regardless of $I_{t-1}$, so the $\mathcal{F}_{t-1}$-conditional
variance of each summand in $\sqrt{n}\, M_n(\alpha, \gamma)$ is
$\alpha(1-\alpha)\exp(2\gamma^\top I_{t-1})$, and \eqref{eq:sn} is its sample
analog.
A more general alternative is the sample-second-moment estimator
\begin{equation}
\tilde{s}_n^2(\alpha, \gamma)
:= \frac{1}{n}\sum_{t=1}^n \exp(2\gamma^\top I_{t-1})
\big[\mathbf{1}\{Y_t \le m(I_{t-1}, \bar{\theta}(\alpha))\} - \alpha\big]^2,
\label{eq:sn-robust}
\end{equation}
constructed directly from the squared summands of $\sqrt{n}\, M_n(\alpha, \gamma)$.
Under the MDS property, the cross-terms in the long-run variance vanish, so
$\tilde{s}_n^2(\alpha, \gamma)$ and $s_n^2(\alpha, \gamma)$ share the same
probability limit $\sigma^2(\alpha, \gamma) = \alpha(1-\alpha)\,
E[\exp(2\gamma^\top I_{t-1})]$, and either can be used to studentize the empirical
process. We use $s_n^2$ in \eqref{eq:sn} because it is simple and stable in
finite samples. The same first-order results hold for $\tilde{s}_n^2$ when its
sample-second-moment analog is also used in the bootstrap.
The standardized process is
\[
Q_n(\alpha, \gamma)
:=
\frac{\sqrt{n}|M_n(\alpha, \gamma)|}{s_n(\alpha, \gamma)}.
\]
To conduct inference uniformly over \(\alpha\in\mathcal{T}\), we aggregate evidence
across quantile indices through a fixed probability measure \(\Pi\) on
\(\mathcal T\). Throughout the theoretical analysis, \(\mathcal T\) is a compact
subset of \((0,1)\), and \(\Pi\) is nonrandom. For instance, when
\(\mathcal T=[\underline\alpha,\overline\alpha]\), \(\Pi\) may be the uniform
probability measure on this interval. The finite-grid implementation is obtained
as the special case \(\Pi_A=A^{-1}\sum_{j=1}^A\delta_{\alpha_j}\), where
\(\mathcal T_A=\{\alpha_1,\ldots,\alpha_A\}\).
The standardized process is aggregated by a CvM-type functional over \(\alpha\)
and a penalized supremum over \(\gamma\in\Gamma\). We define the
\textbf{penalized Bierens statistic} as
\begin{equation}
T_{n,\Pi}(\lambda)
:=
\sup_{\gamma \in \Gamma}
\left[
\int_{\mathcal T} Q_n^2(\alpha,\gamma)\,d\Pi(\alpha)
-
\lambda\|\gamma\|_1
\right].
\label{eq:Tn}
\end{equation}
For the fixed-penalty asymptotic results below, \(\lambda\ge0\) is treated as a
deterministic index. The data-dependent selector \(\widehat\lambda\) is introduced
separately in Section~\ref{subsec:calibration}. When no confusion arises, we write
\(T_n(\lambda)\) for \(T_{n,\Pi}(\lambda)\). The statistic combines CvM
aggregation over \(\alpha\) with a penalized supremum over \(\gamma\). The
penalty discourages weighting directions with large \(\ell_1\)-norms.
\begin{remark}[Finite-grid implementation and aggregation schemes]
\label{rem:alternative-schemes}
Two clarifications about the aggregation in \eqref{eq:Tn} are in order.
First, the formulation in \eqref{eq:Tn} unifies continuous and finite-grid
aggregations over the quantile index. If \(\Pi\) is the uniform probability
measure on \([\underline\alpha,\overline\alpha]\), then
\[
\int_{\mathcal T} Q_n^2(\alpha,\gamma)\,d\Pi(\alpha)
=
\frac{1}{\overline\alpha-\underline\alpha}
\int_{\underline\alpha}^{\overline\alpha}
Q_n^2(\alpha,\gamma)\,d\alpha.
\]
If instead
\(\Pi_A=A^{-1}\sum_{j=1}^A\delta_{\alpha_j}\), then
\[
T_{n,\Pi_A}(\lambda)
=
\sup_{\gamma\in\Gamma}
\left[
\frac1A\sum_{j=1}^A Q_n^2(\alpha_j,\gamma)
-
\lambda\|\gamma\|_1
\right],
\]
which is exactly the finite-grid statistic used in the numerical implementation.
Thus the finite-grid statistic can be interpreted as a quadrature version of the
continuous CvM aggregation.
For comparison with alternative aggregation schemes, we also write
\[
T_{n,\Pi}^{CvM\text{-}KS}(\lambda)
:=
T_{n,\Pi}(\lambda),
\]
emphasizing that the proposed statistic applies CvM integration over
\(\alpha\) before taking the KS supremum over \(\gamma\).
For comparison, the numerical experiments use two alternative functionals.
Writing \(\eta\ge0\) for an optional penalty on the quantile index, define
\begin{equation*}
T_n^{KS\text{-}KS}(\eta,\lambda)
:=\sup_{\alpha\in\mathcal T,\gamma\in\Gamma}
\left[|Q_n(\alpha,\gamma)|-\eta|\alpha|-\lambda\|\gamma\|_1\right].
\end{equation*}
At \(\eta=\lambda=0\), this is the fully supremum-based counterpart.
The other comparator maximizes over \(\gamma\) at each quantile before
squaring and integrating:
\begin{equation*}
T_{n,\Pi}^{KS\text{-}CvM}(\lambda)
:=\int_{\mathcal T}
\left\{\sup_{\gamma\in\Gamma}
\left[|Q_n(\alpha,\gamma)|-\lambda\|\gamma\|_1\right]\right\}^{2}
d\Pi(\alpha).
\end{equation*}
These are the definitions used in the numerical experiments; the tables suppress
the penalty arguments. The penalty enters before squaring in the second
comparator, so penalized comparisons reflect both aggregation and penalty
placement. The main statistic \(T_{n,\Pi}\) and its theory are unchanged.
The implementation choices, including \(\eta\), are stated in
Section~\ref{sec:simulation}.
We do not separately consider the fully integrated CvM--CvM aggregation, which
would replace the outer supremum over \(\gamma\) by an integral. The reason is
that the local-separation result in Theorem~\ref{thm:power-enhance_main}
differentiates a value function at its maximizing \(\gamma\). It therefore
applies directly to statistics that retain a supremum over \(\gamma\), rather
than to an integral over \(\gamma\).
We use \(T_{n,\Pi}(\lambda)\) in \eqref{eq:Tn} as the main statistic because it
has rejection rates relatively close to nominal levels in the linear
pre-estimation design reported in Section~\ref{subsec:aggregation-comparison}.
This comparison is specific to the reported simulations.
\end{remark}
\subsection{Asymptotic Null Distribution}
The assumptions for the null theory are stated in the Supplementary Appendix.
The following theorem gives the null limit of the empirical process and test
statistic.
\begin{theorem}[Asymptotic null distribution]
\label{thm:null_main}
Suppose Assumption~S1\ and Assumption~S2\ hold. Under the null hypothesis $H_0: \theta_0(\alpha) = \bar{\theta}(\alpha)$ for all $\alpha \in \mathcal{T}$, the following results hold:
\begin{enumerate}
\item The empirical process converges weakly to a zero-mean Gaussian process:
\begin{equation*}
\sqrt{n} M_n(\alpha, \gamma) \rightsquigarrow \mathcal{M}(\alpha, \gamma) \quad \text{in } \ell^{\infty}(\mathcal{T} \times \Gamma).
\end{equation*}
The covariance kernel of the limiting process $\mathcal{M}(\alpha, \gamma)$ is given by:
\begin{equation*}
\mathrm{Cov}\big(\mathcal{M}(\alpha_1, \gamma_1), \mathcal{M}(\alpha_2, \gamma_2)\big)
= \big(\min(\alpha_1, \alpha_2) - \alpha_1 \alpha_2\big)
E\!\left[\exp\!\big((\gamma_1 + \gamma_2)^{\top} I_{t-1}\big)\right].
\end{equation*}
\item The variance estimator converges uniformly in probability:
\begin{equation*}
\sup_{\alpha\in \mathcal{T}, \gamma\in \Gamma}
\Big|
s_n^2(\alpha, \gamma)
- \sigma^2(\alpha, \gamma)
\Big|
\overset{p}{\longrightarrow} 0,
\end{equation*}
where $\sigma^2(\alpha, \gamma) := \alpha(1-\alpha) E[\exp(2\gamma^{\top} I_{t-1})]$.
\item The test statistic \(T_{n,\Pi}(\lambda)\) converges in distribution to
\begin{equation*}
T_{\Pi}(\lambda)
:=
\sup_{\gamma\in \Gamma}
\Bigg[
\int_{\mathcal T}
\left( \frac{\mathcal{M}(\alpha, \gamma)}{\sigma(\alpha, \gamma)} \right)^2
d\Pi(\alpha)
- \lambda \|\gamma\|_1
\Bigg].
\end{equation*}
\end{enumerate}
\end{theorem}
\subsection{Consistency under Global Alternatives}
To complement the null theory, we establish consistency against fixed alternatives that are visible under the aggregation measure. Define the population analog of the studentized moment:
\begin{equation}
h(\alpha, \gamma, \theta)
:= \frac{E\big[\exp(\gamma^\top I_{t-1})
\big(\mathbf{1}\{Y_t \le m(I_{t-1}, \theta(\alpha))\} - \alpha\big)\big]}
{\sigma(\alpha, \gamma)},
\label{eq:h-pop}
\end{equation}
where $\sigma(\alpha, \gamma) = \sqrt{\alpha(1-\alpha)\, E[\exp(2\gamma^\top I_{t-1})]}$ is the null standard deviation.
The exponential weighting function inherits the completeness property established
by \citet{bierens1990}: under \(H_0\), \(h(\alpha,\gamma,\bar\theta)=0\) for all
\((\alpha,\gamma)\in\mathcal T\times\Gamma\). Under a fixed alternative, it is
sufficient for consistency that the discrepancy be visible in \(L^2(\Pi)\) after
exponential weighting. This leads to the following condition.
\begin{theorem}[Consistency under fixed alternatives]
\label{thm:consistency_main}
Suppose Assumption~S1\ and Assumption~S2\ hold. If
\[
\sup_{\gamma\in\Gamma}
\int_{\mathcal T}
h^2(\alpha,\gamma,\bar\theta)
d\Pi(\alpha)
>0,
\]
then for any fixed \(\lambda \geq 0\),
\[
T_{n,\Pi}(\lambda) \xrightarrow{p} +\infty.
\]
\end{theorem}
To see this, note that
\[
\frac{1}{n}\,T_{n,\Pi}(\lambda)
=
\sup_{\gamma \in \Gamma}
\left[
\int_{\mathcal T}
\frac{M_n^2(\alpha,\gamma)}
{s_n^2(\alpha,\gamma)}
d\Pi(\alpha)
-
\frac{\lambda}{n}\|\gamma\|_1
\right].
\]
Under the alternative, the uniform law of large numbers yields
\[
\sup_{\alpha\in\mathcal T,\gamma\in\Gamma}
\left|
\frac{M_n^2(\alpha,\gamma)}
{s_n^2(\alpha,\gamma)}
-
h^2(\alpha,\gamma,\bar\theta)
\right|
\xrightarrow{p}0,
\]
while \(\lambda/n\to0\). Hence, by the continuous mapping theorem,
\[
\frac{1}{n}\,T_{n,\Pi}(\lambda)
\xrightarrow{p}
\sup_{\gamma\in\Gamma}
\int_{\mathcal T}
h^2(\alpha,\gamma,\bar\theta)
d\Pi(\alpha)
>0.
\]
Since the limit is a positive constant, \(T_{n,\Pi}(\lambda)\to+\infty\) in
probability.
If \(\Pi\) has full support on \(\mathcal T\) and
\(\alpha\mapsto h(\alpha,\gamma,\bar\theta)\) is continuous, the condition above
is implied by the existence of a pair \((\alpha',\gamma')\) such that
\(h(\alpha',\gamma',\bar\theta)\neq0\). In that case, the discrepancy persists on
a neighborhood of \(\alpha'\), which has positive \(\Pi\)-measure.
\section{Bootstrap Inference and Local Power Analysis}
\label{sec:bootstrap_power}
\subsection{Multiplier Bootstrap Implementation}
Because the covariance kernel of $\mathcal M$ depends on the data-generating
process, $T_{n,\Pi}(\lambda)$ is non-pivotal. We use a Gaussian multiplier
bootstrap. Conditional on the data, the bootstrap matches the contemporaneous
covariance of the scores; the MDS assumption makes the cross-lag covariances
zero. The procedure does not re-estimate the model in each replication.
Let $\{\omega_t\}_{t=1}^n$ be i.i.d.\ standard normal multipliers,
independent of the original sample $\mathcal Z_n$. We define the bootstrap
empirical process as
\begin{equation}
M_{n*}(\alpha, \gamma)
:=
\frac{1}{n} \sum_{t=1}^{n} \omega_t \exp(\gamma^{\top} I_{t-1})
\big[\mathbf{1}\{Y_t \le m(I_{t-1}, \bar{\theta}(\alpha))\} - \alpha\big].
\label{eq:Mn_star}
\end{equation}
The corresponding bootstrap variance estimator is given by:
\begin{equation}
s_{n*}^{2}(\alpha, \gamma)
:=
\frac{\alpha(1-\alpha)}{n}\sum_{t=1}^{n}\omega_t^2\,\exp(2\gamma^{\top} I_{t-1}).
\label{eq:sn_star}
\end{equation}
Mimicking the construction of the original test, the standardized bootstrap process
is defined as \(Q_{n*}(\alpha,\gamma):=\sqrt n\,|M_{n*}(\alpha,\gamma)|/
s_{n*}(\alpha,\gamma)\). The penalized bootstrap statistic uses the same
aggregation scheme as \(T_{n,\Pi}(\lambda)\):
\begin{equation}
T_{n*,\Pi}(\lambda)
:=
\sup_{\gamma \in \Gamma}
\left[
\int_{\mathcal T} Q_{n*}^2(\alpha,\gamma)\,d\Pi(\alpha)
-
\lambda\|\gamma\|_1
\right].
\label{eq:Tn_star}
\end{equation}
The following theorem establishes the conditional weak convergence needed for bootstrap critical values to be asymptotically valid under $H_0$.
\begin{theorem}[Multiplier bootstrap validity]
\label{thm:bootstrap_main}
Suppose Assumption~S1\ and Assumption~S2\ hold. Under the null hypothesis
$H_0:\theta_0(\alpha)=\bar{\theta}(\alpha)$ for all $\alpha\in\mathcal T$,
conditional on the data $\mathcal Z_n$, the following results hold in
probability:
\begin{enumerate}
\item The bootstrap process converges weakly to the same limiting Gaussian process as the original empirical process:
\begin{equation*}
\sqrt{n} M_{n*}(\alpha,\gamma) \rightsquigarrow^* \mathcal{M}(\alpha,\gamma) \quad \text{in } \ell^\infty(\mathcal T\times\Gamma).
\end{equation*}
\item The bootstrap variance estimator is uniformly consistent:
\begin{equation*}
\sup_{\alpha\in\mathcal{T},\gamma\in\Gamma} \big|\, s_{n*}^2(\alpha,\gamma) - \sigma^2(\alpha, \gamma) \,\big| \xrightarrow{p^*} 0.
\end{equation*}
\item The bootstrap test statistic consistently approximates the null distribution:
\begin{equation*}
T_{n*,\Pi}(\lambda)
\rightsquigarrow^*
\sup_{\gamma \in \Gamma}
\Bigg[
\int_{\mathcal T}
\left(
\frac{\mathcal{M}(\alpha, \gamma)}
{\sigma(\alpha, \gamma)}
\right)^2
d\Pi(\alpha)
-
\lambda\|\gamma\|_1
\Bigg].
\end{equation*}
\end{enumerate}
Here, $\rightsquigarrow^*$ and $\xrightarrow{p^*}$ denote weak convergence and convergence in probability conditional on the data, respectively.
\end{theorem}
The following corollary states explicitly the additional continuity condition
needed to translate bootstrap distributional consistency into rejection-probability
control.
\begin{corollary}[Fixed-penalty size control]
\label{cor:fixed-lambda-size_main}
Fix a deterministic $\lambda\ge0$, and let
$c_{n,1-\tau,\Pi}^*(\lambda)$ be the ideal conditional bootstrap
$(1-\tau)$-quantile. Under Theorem~\ref{thm:bootstrap_main}, if the distribution of
$T_\Pi(\lambda)$ is continuous at its $(1-\tau)$-quantile, then
\[
P\{T_{n,\Pi}(\lambda)>c_{n,1-\tau,\Pi}^*(\lambda)\}\to\tau.
\]
The same conclusion holds for the pre-estimated statistic under
Theorem~\ref{thm:boot-pre_main}, using its plug-in bootstrap critical value and
the corresponding continuity condition.
\end{corollary}
\subsection{Local Power Analysis}
\label{subsec:local-power}
To provide a theoretical foundation for selecting the penalty parameter
\(\lambda\), we study the behavior of the null-imposed statistic under a
sequence of Pitman local alternatives. We parameterize the local sequence as
\begin{equation}
H_{1,n}:\quad
\theta_{n,B}(\alpha)
=
\theta_0(\alpha)
-
\frac{B(\alpha)}{\sqrt n},
\qquad \alpha\in\mathcal T,
\label{eq:local-alt-theta}
\end{equation}
where \(B(\cdot)\) is an admissible bounded deterministic direction in the tangent
set specified in Assumption~S6. In particular, for all sufficiently large $n$, the
local curve remains in the parameter space and defines a nondecreasing conditional
quantile curve. Under this convention, \(B(\alpha)\) is the first-order displacement of the null-imposed quantile relative to the local true quantile. Let \(P_{n,B}\) and \(E_{n,B}\) denote probability and expectation under the local sequence \eqref{eq:local-alt-theta}. The data are generated under \(P_{n,B}\), but the implemented statistic is still evaluated at the null-imposed value \(\theta_0(\alpha)=\bar\theta(\alpha)\).
Define
\[
Z_{nt}(\alpha,\gamma)
:=
\exp(\gamma^\top I_{t-1})
\Big[
\mathbf 1\{Y_t\le m(I_{t-1},\theta_0(\alpha))\}
-
\alpha
\Big],
\]
and let
\[
M_{n,0}^{B}(\alpha,\gamma)
:=
\frac1n\sum_{t=1}^n Z_{nt}(\alpha,\gamma)
\]
denote the null-evaluated empirical process under \(P_{n,B}\). The following
identity separates the random fluctuation from the deterministic local mean
shift:
\begin{equation}
\label{eq:local-centered-decomp}
\sqrt n\,M_{n,0}^{B}(\alpha,\gamma)
=
\mathbb G_{n,B}(\alpha,\gamma)
+
\frac1{\sqrt n}\sum_{t=1}^n E_{n,B}Z_{nt}(\alpha,\gamma),
\end{equation}
where
\[
\mathbb G_{n,B}(\alpha,\gamma)
:=
\frac1{\sqrt n}\sum_{t=1}^n
\Big\{
Z_{nt}(\alpha,\gamma)
-
E_{n,B}Z_{nt}(\alpha,\gamma)
\Big\}.
\]
Thus the expectation inside the centered empirical process is not an additional
drift term; it only centers the stochastic fluctuation. The local alternative
enters through the second term in \eqref{eq:local-centered-decomp}. To see how
the drift arises, note that under $H_{1,n}$ the local true $\alpha$-quantile is
\[
m(I_{t-1},\theta_{n,B}(\alpha)) \approx m(I_{t-1},\theta_0(\alpha)) -
n^{-1/2}\nabla_\theta m(I_{t-1},\theta_0(\alpha))^\top B(\alpha).
\]
Thus $F_{Y_t\mid I_{t-1}}\!\big(m(I_{t-1},\theta_0(\alpha))\big)$ exceeds
$\alpha$ by approximately
\[
n^{-1/2}f_{Y_t\mid I_{t-1}}\!\big(m(I_{t-1},\theta_0(\alpha))\big)
\nabla_\theta m(I_{t-1},\theta_0(\alpha))^\top B(\alpha),
\]
uniformly over $(\alpha,\gamma)$. Taking the conditional expectation of $Z_{nt}$ under
$P_{n,B}$ and then the unconditional expectation therefore yields the drift
formula below; a formal derivation is given in the Supplementary Appendix. The local
drift expansion derived in the Supplementary Appendix gives
\begin{equation}
\label{eq:local-mean-expansion-main}
\frac1{\sqrt n}\sum_{t=1}^n E_{n,B}Z_{nt}(\alpha,\gamma)
=
d_B(\alpha,\gamma)+o(1),
\end{equation}
uniformly in \((\alpha,\gamma)\), where
\begin{equation}
d_B(\alpha,\gamma)
:=
E\!\left[
\exp(\gamma^\top I_{t-1})
f_{Y_t|I_{t-1}}\!\left(m(I_{t-1},\theta_0(\alpha))\right)
\nabla_\theta m(I_{t-1},\theta_0(\alpha))^\top B(\alpha)
\right].
\label{eq:local-drift}
\end{equation}
Under the local-stability and stochastic-equicontinuity conditions in
Assumption~S6, the centered process \(\mathbb G_{n,B}\) has the same first-order
Gaussian limit as under the null. Consequently,
\begin{equation}
\label{eq:local-expansion}
\sqrt n\,M_{n,0}^{B}(\alpha,\gamma)
\rightsquigarrow
\mathcal M(\alpha,\gamma)+d_B(\alpha,\gamma)
\qquad\text{in }\ell^\infty(\mathcal T\times\Gamma),
\end{equation}
where \(\mathcal M\) is the zero-mean Gaussian process in
Theorem~\ref{thm:null_main}.
To characterize the limiting distribution of the statistic, define
\[
\sigma(\alpha,\gamma)
:=
\sqrt{\alpha(1-\alpha)E[\exp(2\gamma^\top I_{t-1})]},
\]
and set
\begin{equation}
Q(\alpha,\gamma)
:=
\frac{\mathcal M(\alpha,\gamma)}{\sigma(\alpha,\gamma)},
\qquad
R_B(\alpha,\gamma)
:=
\frac{d_B(\alpha,\gamma)}{\sigma(\alpha,\gamma)}.
\label{eq:QR_def}
\end{equation}
By the continuous mapping theorem,
\begin{equation}
T_{n,\Pi}(\lambda)
\rightsquigarrow
T_{\Pi}(\lambda;B)
:=
\sup_{\gamma\in\Gamma}
\Bigg[
\int_{\mathcal T}
\big(Q(\alpha,\gamma)+R_B(\alpha,\gamma)\big)^2
d\Pi(\alpha)
-
\lambda\|\gamma\|_1
\Bigg].
\label{eq:limit_local}
\end{equation}
We next examine how penalization changes the separation between the null and
local-alternative limit experiments. The following theorem gives a sufficient
condition.
\begin{theorem}[Local-separation effect of penalization]
\label{thm:power-enhance_main}
For simplicity of exposition, suppose the local drift direction is constant,
\(B(\alpha)\equiv B_0\). Define the unpenalized limit objective
\begin{equation*}
J_\Pi(\gamma;B)
:=
\int_{\mathcal T}
\big(Q(\alpha,\gamma)+R_B(\alpha,\gamma)\big)^2
d\Pi(\alpha).
\end{equation*}
Assume that for both \(B=0\) and some \(B_0\neq0\), there exists a unique
maximizer \(\tilde\gamma_\Pi(B)\) of \(J_\Pi(\gamma;B)\) over the compact set
\(\Gamma\) almost surely. If
\begin{equation*}
\|\tilde\gamma_\Pi(0)\|_1-
\|\tilde\gamma_\Pi(B_0)\|_1>0
\end{equation*}
holds almost surely, then, for almost every realization of the Gaussian
process, there exists a realization-dependent \(\delta(Q,B_0)>0\) such that
for every \(0<\lambda<\delta(Q,B_0)\), the penalized limit statistic increases
the separation
\[
T_\Pi(\lambda;B_0)-T_\Pi(\lambda;0)
\]
relative to the unpenalized benchmark \(\lambda=0\).
\end{theorem}
Theorem~\ref{thm:power-enhance_main} compares the null and local-alternative
value functions along the same realization of the Gaussian process. A small
penalty increases their difference when the unpenalized null maximizer has a
larger \(\ell_1\)-norm than the alternative maximizer. The result is pathwise and
does not compare rejection probabilities because the critical value also depends
on \(\lambda\). The maximin result below provides a direct comparison of
penalty-specific rejection probabilities.
\begin{remark}[Scope and limitations of the structural condition]
\label{rem:power-scope}
The constant direction \(B(\alpha)\equiv B_0\) is used only for this sufficient
condition. The general local alternative in \eqref{eq:local-alt-theta} allows
\(B\) to vary with \(\alpha\), and the selector can be evaluated over a
prespecified class \(\mathcal B\).
The condition
\(\|\tilde\gamma_\Pi(0)\|_1>\|\tilde\gamma_\Pi(B_0)\|_1\) is sufficient, not
necessary. It depends on the realized Gaussian process and is therefore not
directly verifiable. The shifted bootstrap instead evaluates candidate penalties
for the chosen directions. Because \(0\in\Lambda\), the selector can retain the
unpenalized test.
\end{remark}
\begin{remark}[Strict power gains]
\label{rem:strict-gain_main}
The maximin guarantee in Theorem~\ref{thm:maximin-no-loss_main} is a no-loss
statement because \(0\in\Lambda\). Strict improvement requires
\(W(\lambda)>W(0)\) for some \(\lambda\in\Lambda\). Under the theorem's
unique-maximizer condition, this is equivalent to \(\lambda^*>0\).
Concentration of the drift on directions with smaller \(\ell_1\)-norm can
motivate penalization, but the pathwise separation in
Theorem~\ref{thm:power-enhance_main} does not account for the change in the
critical value. Neither that separation nor anti-concentration alone
establishes a strict power gain. Primitive sufficient conditions for strict
improvement remain to be developed.
\end{remark}
\subsection{Calibration of \texorpdfstring{$\lambda$}{lambda} via Local Bootstrap}
\label{subsec:calibration}
The adaptive rule requires an estimate of the local drift
$d_B(\alpha,\gamma)$, given by
\begin{equation}
d_B(\alpha,\gamma) = E\Big[\exp(\gamma^\top I_{t-1})\, f_{Y_t|I_{t-1}}\big(m(I_{t-1}, \theta_0(\alpha))\big)\, \nabla_\theta m(I_{t-1}, \theta_0(\alpha))^\top B(\alpha)\Big].
\label{eq:target_drift}
\end{equation}
Evaluating \eqref{eq:target_drift} requires the conditional density at the
fitted quantile, which we estimate using the method described below.
\subsubsection{Conditional Density Estimation}
For the drift and variance corrections, we use the fitted-quantile density
estimator of \citet{escanciano2019}. It smooths over fitted quantile values
rather than over the full conditioning vector $I_{t-1}$.
Let $\{\alpha_j\}_{j=1}^J$ be i.i.d.\ draws from $\mathrm{Unif}[a_1,a_2]$, independent of the data, where $\mathcal{T}\subset(a_1,a_2)$ and the distance from $\mathcal T$ to the boundary of $[a_1,a_2]$ is positive. The conditional density estimator is defined as:
\begin{equation}
\widehat{f}_h(I_{t-1}, \theta_0(\alpha)) := \frac{a_2-a_1}{J h} \sum_{j=1}^{J} K\left( \frac{m(I_{t-1}, \theta_0(\alpha)) - m(I_{t-1}, \theta_0(\alpha_j))}{h} \right),
\label{eq:density_est}
\end{equation}
where $K(\cdot)$ is a symmetric kernel function and $h$ is the bandwidth.
The null-imposed quantile curve must be correctly specified on the auxiliary
interval, as required by Assumption~S2(iv); this is a maintained condition
for density estimation in addition to the restriction tested on $\mathcal T$.
The following theorem gives the uniform consistency result used below.
\begin{theorem}[Uniform consistency and rate of $\widehat{f}_h$]
\label{thm:density-consistency_main}
Under Assumption~S1, Assumption~S2, and Assumption~S3(a)--(c), if
$\mathcal{T}\subset(a_1,a_2)$ has positive distance from the boundary of
$[a_1,a_2]$, the bandwidth satisfies
$h \asymp (\frac{\log n}{n})^{1/5}$, and the simulation size $J \asymp n$,
then:
\[
\sup_{\alpha \in \mathcal{T}}\max_{1\le t\le n} \Big| \widehat{f}_h(I_{t-1}, \theta_0(\alpha)) - f_{Y_t|I_{t-1}}\big(m(I_{t-1}, \theta_0(\alpha))\big) \Big| = o_p(1).
\]
\end{theorem}
The estimator smooths over the scalar fitted-quantile index
$m(I_{t-1},\theta_0(\alpha))$ rather than over the full conditioning vector.
The kernel condition in Assumption~S3\ permits either a smooth compactly supported
kernel or a smooth kernel with exponential tail decay, including the Gaussian
kernel. The Supplementary Appendix establishes the more explicit bound
$O_p\{(\log n/n)^{2/5}\}$ by a Bernstein inequality and a covering argument.
Writing $a_{J_n}$ for the deterministic lower endpoint of the admissible
bandwidth interval in Assumption~S3(b), the canonical choice
$a_{J_n}\asymp(\log n/n)^{1/5}$ also gives
\[
n^{-1/2}a_{J_n}^{-2}
=n^{-1/10}(\log n)^{-2/5}\longrightarrow0,
\]
which is the generated-parameter compatibility condition used by the
pre-estimation density plug-in expansion.
\subsubsection{Bootstrap Validity under Local Alternatives}
Using the uniformly consistent density estimator \(\widehat f_h\), we
construct a shifted multiplier bootstrap process. It reproduces the noncentral
limit in \eqref{eq:local-expansion}: the multiplier component approximates the
centered Gaussian fluctuation \(\mathbb G_{n,B}\), while a deterministic sample
drift estimates \(d_B(\alpha,\gamma)\). Because
\(E^*(\omega_t\mid\mathcal Z_n)=0\), the multiplier term is already conditionally
centered and no explicit subtraction of \(E_{n,B}Z_{nt}\) is needed in the
bootstrap stochastic component.
Let \(\{\omega_t\}_{t=1}^n\) be i.i.d. standard normal multipliers independent of
the data. For a candidate local direction \(B\), define
\begin{align}
M_{n*}(\alpha,\gamma;B)
&:=
\frac1n\sum_{t=1}^n
\omega_t
\exp(\gamma^\top I_{t-1})
\big[\mathbf 1\{Y_t\le m(I_{t-1},\bar\theta(\alpha))\}-\alpha\big]
\nonumber\\
&\quad+
\frac1n\sum_{t=1}^n
\exp(\gamma^\top I_{t-1})
\widehat f_h(I_{t-1},\bar\theta(\alpha))
\frac{B(\alpha)^\top}{\sqrt n}
\nabla_\theta m(I_{t-1},\bar\theta(\alpha)).
\label{eq:shifted_M}
\end{align}
The second line is of order \(n^{-1/2}\) at the summand level; after multiplying
by \(\sqrt n\), it converges to the local drift \(d_B(\alpha,\gamma)\).
The bootstrap variance estimator is
\begin{equation}
s_{n*,B}^2(\alpha,\gamma)
:=
\frac{\alpha(1-\alpha)}{n}
\sum_{t=1}^n \omega_t^2\exp(2\gamma^\top I_{t-1}).
\label{eq:shifted_s}
\end{equation}
\begin{remark}
If the original statistic is studentized by a sample second moment rather than
by the closed-form variance in \eqref{eq:shifted_s}, the shifted bootstrap
studentizer can be defined as the sample second moment of the shifted bootstrap
summands. The deterministic shift is \(O(n^{-1/2})\) at the summand level, so it
does not change the first-order variance limit.
\end{remark}
Define
\begin{equation}
Q_{n*,B}(\alpha,\gamma)
:=
\frac{\sqrt n\,|M_{n*}(\alpha,\gamma;B)|}{s_{n*,B}(\alpha,\gamma)}.
\label{eq:Qn_shifted}
\end{equation}
The shifted bootstrap statistic is
\begin{equation}
T_{n*,B,\Pi}(\lambda)
:=
\sup_{\gamma\in\Gamma}
\Bigg[
\int_{\mathcal T} Q_{n*,B}^2(\alpha,\gamma)d\Pi(\alpha)
-
\lambda\|\gamma\|_1
\Bigg].
\label{eq:Tn_shifted}
\end{equation}
The shifted bootstrap must estimate the power criterion for every candidate
direction $B$, whether the observed data are generated under the null or under
one of the contiguous local sequences. We therefore distinguish the candidate
shift $B$ from the actual data-generating direction $B_0$, with $B_0=0$
denoting the null.
\begin{theorem}[Shifted bootstrap validity]
\label{thm:MB-local_main}
Suppose Assumption~S1, Assumption~S2, Assumption~S3,
Assumption~S5, and Assumption~S6\ hold. Let the bandwidth \(h_n\) satisfy the
conditions in Theorem~\ref{thm:density-consistency_main}. For any admissible
candidate direction $B$, under the null or any actual local sequence
$P_{n,B_0}$ covered by Assumption~S6, the following results hold conditional on
the data \(\mathcal Z_n\) in probability:
\begin{enumerate}
\item The shifted bootstrap process replicates the noncentral limit process:
\[
\sqrt n\,M_{n*}(\alpha,\gamma;B)
\rightsquigarrow^*
\mathcal M(\alpha,\gamma)+d_B(\alpha,\gamma)
\quad\text{in }\ell^\infty(\mathcal T\times\Gamma).
\]
\item The variance estimator remains uniformly consistent:
\[
\sup_{\alpha,\gamma}
\big|s_{n*,B}^2(\alpha,\gamma)-\sigma^2(\alpha,\gamma)\big|
\xrightarrow{p^*}0.
\]
\item The shifted bootstrap statistic converges to the corresponding
noncentral limit:
\[
T_{n*,B,\Pi}(\lambda)
\rightsquigarrow^*
\sup_{\gamma\in\Gamma}
\Bigg[
\int_{\mathcal T}
\left(
\frac{\mathcal M(\alpha,\gamma)+d_B(\alpha,\gamma)}
{\sigma(\alpha,\gamma)}
\right)^2
d\Pi(\alpha)
-
\lambda\|\gamma\|_1
\Bigg].
\]
\end{enumerate}
\end{theorem}
\subsubsection{Adaptive Penalty Selection Rule}
Using the shifted bootstrap approximation, we select the penalty parameter
$\lambda$ to maximize estimated power against the least favorable local
alternative in a prespecified class $\mathcal B$.
For theoretical purposes, let
\[
\Lambda=[0,\bar\lambda]
\]
be a compact interval of candidate penalty values. Define
$\widehat W_n(\lambda)=\inf_{B\in\mathcal B}
\widehat{\mathcal R}_n(\lambda,B,\tau)$. For continuous search, use a measurable
approximate maximizer satisfying
\begin{equation}
\widehat W_n(\widehat\lambda)
\ge \sup_{\lambda\in\Lambda}\widehat W_n(\lambda)-\eta_n,
\qquad \eta_n=o_p(1),\quad \eta_n\ge0,
\label{eq:maxmin}
\end{equation}
with $\eta_n=0$ whenever a measurable exact maximizer is used. Here
\(\widehat{\mathcal R}_n(\lambda,B,\tau)\) denotes the shifted-bootstrap estimate of the rejection probability against the local direction \(B\), using the null critical value corresponding to the same penalty \(\lambda\).
In computation, \(\Lambda\) may be searched directly or approximated by a finite
grid \(\Lambda_G\). A fixed grid targets the grid maximizer. It approximates the
continuous population maximizer when the grid mesh tends to zero and the
population criterion is continuous. With finitely many bootstrap draws, the
estimated rejection-frequency criterion may be stepwise. We reuse the same
multiplier draws across penalties and directions to reduce simulation noise.
Algorithm~\ref{alg:adaptive-known} summarizes the procedure.
\begin{algorithm}[htbp]
\caption{Adaptive penalty selection without pre-estimation}
\label{alg:adaptive-known}
\begin{algorithmic}[1]
\Require Data \(\mathcal Z_n\), \(\Pi\), \(\Gamma\), search set
\(\mathcal S=\Lambda\) or \(\Lambda_G\), finite \(\mathcal B\), \(R\), \(\tau\),
and density-smoothing choices \((K,h,J,[a_1,a_2])\)
\State Estimate the conditional density by \eqref{eq:density_est} and, for each
\(B\in\mathcal B\), compute the sample drift
\(\widehat d_{n,B}=n^{-1}\sum_t e^{\gamma^\top I_{t-1}}
\widehat f_h(I_{t-1},\bar\theta(\alpha))
\nabla_\theta m(I_{t-1},\bar\theta(\alpha))^\top B(\alpha)\)
used in \eqref{eq:shifted_M}.
\State Draw \(R\) multiplier sequences
\(\Omega^{(r)}=\{\omega_t^{(r)}\}_{t=1}^n\); reuse them throughout.
\For{each penalty \(\lambda\) requested by the search routine}
\For{\(r=1,\ldots,R\)}
\State Compute the null bootstrap statistic
\(T_{n*,\Pi}^{(r)}(\lambda)\).
\EndFor
\State Set \(c_{n,1-\tau,\Pi}^*(\lambda)\) to their empirical
\((1-\tau)\)-quantile.
\For{each \(B\in\mathcal B\)}
\State Compute \(T_{n*,B,\Pi}^{(r)}(\lambda)\), \(r=1,\ldots,R\),
using the shifted bootstrap process.
\State Set
\(\widehat{\mathcal R}_n(\lambda,B,\tau)
=R^{-1}\sum_{r=1}^R
\mathbf 1\{T_{n*,B,\Pi}^{(r)}(\lambda)>
c_{n,1-\tau,\Pi}^*(\lambda)\}\).
\EndFor
\State Set \(\widehat W_n(\lambda)=
\min_{B\in\mathcal B}\widehat{\mathcal R}_n(\lambda,B,\tau)\).
\EndFor
\State On \(\Lambda_G\), select the smallest maximizer of \(\widehat W_n\);
for continuous search, use a measurable \(\eta_n\)-approximate maximizer,
\(\eta_n=o_p(1)\).
\State Compute \(T_{n,\Pi}(\widehat\lambda)\) and reject when it exceeds
\(c_{n,1-\tau,\Pi}^*(\widehat\lambda)\).
\end{algorithmic}
\end{algorithm}
For a grid search, the selection line in Algorithm~\ref{alg:adaptive-known}
reduces to
\begin{equation}
\widehat{\lambda}_G
=
\arg\max_{\lambda \in \Lambda_G}
\min_{B \in \mathcal{B}}
\widehat{\mathcal{R}}_n(\lambda, B, \tau),
\label{eq:maxmin-grid}
\end{equation}
again choosing the smallest maximizer. For a fixed finite grid, this rule targets
the grid optimum rather than the continuous optimum. Alternatively, the criterion
may be optimized directly over \([0,\bar\lambda]\) using a derivative-free method
such as particle swarm optimization (PSO).
The test rejects when \(T_{n,\Pi}(\widehat\lambda)\) exceeds
\(c_{n,1-\tau,\Pi}^*(\widehat\lambda)\). Here \(\widehat\lambda\) denotes the
continuous selector, while \(\widehat\lambda_G\) denotes the finite-grid
selector. The theoretical critical values are ideal conditional bootstrap
quantiles. Monte Carlo quantiles may be used when their simulation error is
negligible uniformly over the search set.
A sufficient condition for post-selection validity is that the selected
penalty converges to a deterministic limit. This condition follows, for example,
when the estimated maximin criterion converges uniformly to a population
criterion with a unique maximizer. The following corollary states the
continuous-\(\Lambda\) result and also applies to the plug-in statistic in
Section~\ref{sec:pre_estimation}.
\begin{corollary}[Validity for adaptive penalty selection]
\label{cor:adaptive-lambda_main}
Let \(\Lambda=[0,\bar\lambda]\) be compact. Suppose the relevant bootstrap approximation in Theorem~\ref{thm:bootstrap_main} holds uniformly over \(\lambda\in\Lambda\), and let \(c_{n,1-\tau,\Pi}^*(\lambda)\) denote the ideal conditional bootstrap critical value, or a Monte Carlo approximation with simulation error \(o_p(1)\) uniformly in \(\lambda\). Let \(c_{1-\tau,\Pi}(\lambda)\) be the corresponding \((1-\tau)\)-quantile of the null limit statistic. If
\[
\widehat\lambda\xrightarrow{p}\lambda_0
\qquad\text{for some deterministic }\lambda_0\in\Lambda,
\]
and the limiting distribution is continuous at \(c_{1-\tau,\Pi}(\lambda_0)\) with \(\lambda\mapsto c_{1-\tau,\Pi}(\lambda)\) continuous at \(\lambda_0\), then under \(H_0\),
\[
P\{T_{n,\Pi}(\widehat\lambda)>c_{n,1-\tau,\Pi}^*(\widehat\lambda)\}
\to\tau.
\]
The same conclusion holds for the pre-estimated statistic
\(\widehat T_{n,\Pi}\) under Theorem~\ref{thm:boot-pre_main}, with its
corresponding plug-in bootstrap critical value.
\end{corollary}
The preceding corollary establishes size control once the selector stabilizes.
The next theorem gives sufficient conditions for this convergence and a no-loss
result for maximin local power because the search set contains \(\lambda=0\).
Unlike Theorem~\ref{thm:power-enhance_main}, it compares rejection probabilities
computed with their penalty-specific null critical values.
For \(\lambda\in\Lambda\), let
\[
c_{1-\tau,\Pi}(\lambda)
:=q_{1-\tau}\{T_\Pi(\lambda;0)\}
\]
and define the local asymptotic rejection probability
\[
\mathcal R(\lambda,B,\tau)
:=P\{T_\Pi(\lambda;B)>c_{1-\tau,\Pi}(\lambda)\}.
\]
\begin{theorem}[Maximin local-power no-loss result]
\label{thm:maximin-no-loss_main}
Let \(\Lambda=[0,\bar\lambda]\) and suppose \(0\in\Lambda\). Let
\(\mathcal B=\{B_1,\ldots,B_J\}\) be a finite, nonempty collection of local
directions. Suppose the conclusions of Theorems~\ref{thm:bootstrap_main} and
\ref{thm:MB-local_main} hold for every fixed \(\lambda\in\Lambda\) and every
\(B\in\mathcal B\), with their process-level stochastic equicontinuity
conditions holding uniformly over this finite collection. Suppose also that
\[
\lim_{\varepsilon\downarrow0}
\max_{B\in\mathcal B\cup\{0\}}
\sup_{\lambda\in\Lambda}
P\!\left(
\left|T_\Pi(\lambda;B)-c_{1-\tau,\Pi}(\lambda)\right|
\le\varepsilon
\right)=0.
\]
Assume that the Monte Carlo errors in the estimated critical values and
rejection probabilities are \(o_p(1)\) uniformly in \(\lambda\), and define
\[
W(\lambda):=\min_{B\in\mathcal B}\mathcal R(\lambda,B,\tau).
\]
If \(W\) is continuous on \(\Lambda\) and has a unique maximizer
\(\lambda^*\), then
\[
\sup_{\lambda\in\Lambda,\,B\in\mathcal B}
\left|
\widehat{\mathcal R}_n(\lambda,B,\tau)
-\mathcal R(\lambda,B,\tau)
\right|\xrightarrow{p}0,
\qquad
\widehat\lambda\xrightarrow{p}\lambda^*.
\]
Moreover,
\[
\min_{B\in\mathcal B}\mathcal R(\lambda^*,B,\tau)
=\max_{\lambda\in\Lambda}
\min_{B\in\mathcal B}\mathcal R(\lambda,B,\tau)
\ge
\min_{B\in\mathcal B}\mathcal R(0,B,\tau).
\]
For every \(B\in\mathcal B\), under the corresponding contiguous local
sequence,
\[
P_{n,B}\!\left\{
T_{n,\Pi}(\widehat\lambda)
>c_{n,1-\tau,\Pi}^*(\widehat\lambda)
\right\}
\longrightarrow
\mathcal R(\lambda^*,B,\tau).
\]
Consequently, if \(\mathcal B=\{B^*\}\) is a singleton, then
\(\mathcal R(\lambda^*,B^*,\tau)\ge
\mathcal R(0,B^*,\tau)\). For a collection containing more than one direction,
the guarantee concerns worst-case local power and need not hold direction by
direction.
\end{theorem}
The anti-concentration condition rules out probability mass accumulating near
the penalty-specific critical values uniformly over \(\lambda\) and \(B\). It is
satisfied, for example, if the distributions of \(T_\Pi(\lambda;B)\) have
densities that are uniformly bounded in a neighborhood of the corresponding
critical values. On a fixed finite penalty grid, pointwise continuity at each
critical value is sufficient.
\begin{remark}[Fixed-grid no-loss result]
\label{rem:grid-no-loss_main}
Let $\Lambda_G$ be any fixed finite grid containing zero. At the population level,
\[
\max_{\lambda\in\Lambda_G}W(\lambda)\ge W(0).
\]
If the shifted-bootstrap power estimates converge uniformly over this finite grid
and the grid maximizer is unique, the selected grid value converges to that
maximizer and inherits the same maximin no-loss inequality relative to
$\lambda=0$. This result concerns the grid optimum and requires no approximation
to the continuous optimum.
\end{remark}
\begin{remark}[Distributional regularity of the limit statistic]
\label{rem:continuity_main}
Continuity at the critical quantile and uniform anti-concentration are
maintained distributional conditions in the preceding results. A sufficient
condition for uniform anti-concentration is a density uniformly bounded near
the relevant critical values. We do not infer that density bound from
continuity of the covariance kernel or uniqueness of the maximizing direction.
For the null Gaussian limit, continuity of the statistic as a functional of
the process gives connected distributional support; together with continuity
at the critical quantile, this yields the strict quantile crossing used in the
bootstrap proof (Lemma~\ref{lem:gaussian-quantile-crossing} in the Supplementary Appendix). This support argument applies to
both continuous and finite-grid aggregation measures $\Pi$.
\end{remark}
Theorem~\ref{thm:maximin-no-loss_main} is an asymptotic local-power result, not a
finite-sample dominance statement. Its unique-maximizer conclusion also verifies
the deterministic-stabilization condition used in
Corollary~\ref{cor:adaptive-lambda_main}. On a finite grid, uniform Monte Carlo
error follows over the grid as the number of draws increases. For a continuous
search, uniform simulation accuracy is a maintained condition of the theorem.
\section{Inference with Pre-Estimated Parameters}
\label{sec:pre_estimation}
We now allow the parameter vector to contain unknown nuisance components that
are estimated before the test statistic is constructed. Write
$\theta(\alpha)=(\theta_1(\alpha)^\top,\theta_2(\alpha)^\top)^\top$, where
$\theta_1(\alpha)$ is the component of interest and $\theta_2(\alpha)$ is a
nuisance parameter. We test the first component uniformly over the quantile
index:
\begin{equation}
H_0: \theta_{10}(\alpha) = \bar{\theta}_1(\alpha), \quad \forall \alpha \in \mathcal{T},
\end{equation}
while treating the true nuisance parameter $\theta_{20}(\alpha)$ as unknown. Under this null hypothesis, we evaluate the conditional moment restriction by replacing $\theta_{20}(\alpha)$ with a preliminary $\sqrt{n}$-consistent estimator.
\subsection{Test Statistic with Nuisance Parameters}
Let
\[
m_{t\alpha}(\theta_1,\theta_2)
:=m(I_{t-1},\theta_1(\alpha),\theta_2(\alpha))
\]
denote the conditional quantile function, and write its nuisance gradient as
\[
g_{t\alpha}(\theta_1,\theta_2)
:=\nabla_{\theta_2}m(I_{t-1},\theta_1(\alpha),\theta_2(\alpha)).
\]
We estimate $\theta_{20}(\alpha)$ under the null restriction
$\theta_{10}(\alpha)=\bar{\theta}_1(\alpha)$. The theory requires the preliminary
estimator $\hat{\theta}_2(\alpha)$ to be $\sqrt n$-consistent and to admit a
uniform linear influence-function expansion. As an example, consider the
restricted moment estimator that solves
\begin{equation}
\frac{1}{n} \sum_{t=1}^n g_{t\alpha}(\bar{\theta}_1(\alpha), \hat{\theta}_2(\alpha)) \,
\Big[ \mathbf{1}\big\{Y_t \le m_{t\alpha}(\bar{\theta}_1(\alpha), \hat{\theta}_2(\alpha))\big\} - \alpha \Big] = 0,
\label{eq:theta2-est}
\end{equation}
where $g_{t\alpha}$ is the gradient defined above, and the equation should be
interpreted as the subgradient/first-order condition of the restricted quantile
objective. An alternative estimator may be substituted if it satisfies the
uniform influence-function expansion, its influence function can be estimated
consistently, and the corrected-score conditions in Assumption~S4\ continue to hold.
We impose a uniform linear representation for the restricted quantile
estimator. Such representations are standard for regression quantiles; see
\citet{gutenbrunner1992}, \citet{mukherjee1999}, and the dynamic-quantile
framework of \citet{escanciano2010}. Specifically, we assume that
$\hat\theta_2(\alpha)$ has the following expansion uniformly over $\mathcal T$:
\begin{equation}
\sup_{\alpha \in \mathcal{T}} \left\| \sqrt{n}\big(\hat{\theta}_2(\alpha) - \theta_{20}(\alpha)\big) - \frac{1}{\sqrt{n}}\sum_{t=1}^n l_{t,\alpha}(\bar{\theta}_1, \theta_{20}) \right\| = o_p(1).
\label{eq:theta2-IF}
\end{equation}
For the restricted quantile estimator in \eqref{eq:theta2-est}, the familiar
quantile-score expansion gives the influence function
\begin{equation}
l_{t,\alpha} := -L_\alpha^{-1} \, g_{t\alpha}(\bar{\theta}_1, \theta_{20}) \, \Big[\mathbf{1}\{Y_t \le m_{t\alpha}(\bar{\theta}_1, \theta_{20})\} - \alpha\Big],
\label{eq:IF-def}
\end{equation}
where $L_\alpha := E[g_{t\alpha}(\bar{\theta}_1, \theta_{20}) g_{t\alpha}^\top(\bar{\theta}_1, \theta_{20}) f_{Y_t|I_{t-1}}(m_{t\alpha}(\bar{\theta}_1, \theta_{20}))]$ is the expected Jacobian. At the true parameters, the influence function is a martingale difference under the maintained conditions. Hence it has mean zero and is orthogonal across time; in particular,
$E[l_{t,\alpha}\{\mathbf 1(Y_s\le m_{s\alpha})-\alpha\}]=0$ for
$t\ne s$.
In the linear quantile-regression examples below, $g_{t\alpha}$ is the vector
of nuisance regressors, so $L_\alpha$ is their density-weighted second-moment
matrix. The required nonsingularity, uniform influence-function representation
on the full auxiliary quantile interval, and joint process conditions are
maintained in Assumption~S4; the nonlinear case is subject to the same interface.
For the plug-in statistic, we use sample-centered exponential weights. Because
the centered weight is zero at $\gamma=0$, the continuous theory uses a compact
index set separated from zero, for example
\[
\Gamma_\varepsilon
:=\{\gamma\in\Gamma_0:\|\gamma\|_2\ge\varepsilon\}
\]
for a compact $\Gamma_0$ and fixed $\varepsilon>0$, together with the
nondegeneracy conditions in Assumption~S4. For notational simplicity, we continue
to denote this set by $\Gamma$. A finite-grid implementation omits
$\gamma=0$ from the studentized search and requires positive corrected-score
variances at the remaining nodes, as in Assumption~S4. It may append a
deterministic zero baseline after this search, equivalently applying the
positive-part map to the supremum over the nonzero directions. The observed
and multiplier-bootstrap statistics must use the same convention. Since the
positive-part map is continuous, the stated weak-convergence and bootstrap
conclusions carry over by the continuous mapping theorem. This restriction
applies only to the centered-weight
plug-in statistic; the known-parameter statistic in
Sections~\ref{sec:test_stat}--\ref{sec:bootstrap_power} may include the origin.
Centering also changes the set of detectable directions. The pre-estimation
power results require the drift to be nonzero for at least one centered weight.
A departure whose conditional moment is constant in the conditioning variables
is orthogonal to every centered weight and is therefore not detected by this
component alone. This limitation does not affect null validity, and the
fixed-alternative consistency result in Theorem~\ref{thm:consistency_main}
concerns the known-parameter statistic. An additional constant instrument can
supply information only if its corrected score is nondegenerate and is not
absorbed by the nuisance estimating equations; a nuisance intercept can absorb
this component. Such an extension is outside the statistic studied here.
Let $\bar{e}_n(\gamma) := \frac{1}{n}\sum_{s=1}^n \exp(\gamma^\top I_{s-1})$ denote the sample mean of the weights, and define the \textbf{demeaned weight function} by
\begin{equation}
w_{t,n}(\gamma) := \exp(\gamma^\top I_{t-1}) - \bar{e}_n(\gamma).
\label{eq:weight-def}
\end{equation}
The \textbf{plug-in empirical process} is then constructed using the estimated nuisance parameters and the demeaned weights:
\begin{equation}
\widehat M_n(\alpha, \gamma) := \frac{1}{n}\sum_{t=1}^n w_{t,n}(\gamma) \Big[ \mathbf{1}\{Y_t\le m_{t\alpha}(\bar{\theta}_1, \hat{\theta}_2)\} - \alpha \Big].
\label{eq:Mhat-def}
\end{equation}
By stochastic equicontinuity and a Taylor expansion of the smoothed moment condition, the plug-in process admits the following asymptotic expansion under $H_0$.
Let
\[
w_t(\gamma)
:=
\exp(\gamma^\top I_{t-1})
-
E\!\left[\exp(\gamma^\top I_{t-1})\right]
\]
denote the population-centered weight, and define the corresponding centered-weight null process
\[
M_n^c(\alpha,\gamma)
:=
\frac1n\sum_{t=1}^n
w_t(\gamma)
\Big[
\mathbf 1\{Y_t\le m_{t\alpha}(\theta_{10},\theta_{20})\}
-
\alpha
\Big].
\]
Furthermore, define
\[
A_2(\alpha,\gamma)
:=
E\!\left[
w_t(\gamma)
f_{Y_t\mid I_{t-1}}
\!\left(
m_{t\alpha}(\theta_{10},\theta_{20})
\right)
\nabla_{\theta_2}
m_{t\alpha}(\theta_{10},\theta_{20})
\right].
\]
Then, uniformly over $(\alpha,\gamma)\in\mathcal T\times\Gamma$,
\begin{equation}
\sqrt{n}\,\widehat M_n(\alpha,\gamma)
=
\sqrt{n}\,M_n^c(\alpha,\gamma)
+
A_2(\alpha,\gamma)^\top
\sqrt{n}\{\widehat\theta_2(\alpha)-\theta_{20}(\alpha)\}
+
o_p(1).
\label{eq:Mhat-expansion}
\end{equation}
This expansion decomposes the plug-in process into a centered-weight Gaussian component and an additional term induced by nuisance-parameter estimation.
Define the population corrected score
\[
\Psi_t(\alpha,\gamma)
:=
w_t(\gamma)
\big[\mathbf 1\{Y_t\le m_{t\alpha}(\theta_{10},\theta_{20})\}-\alpha\big]
+l_{t,\alpha}^{\top}A_2(\alpha,\gamma).
\]
Under Assumption~S4(iii), this score is serially uncorrelated, so its long-run
variance is
\(\tilde\sigma^2(\alpha,\gamma)=E[\Psi_t(\alpha,\gamma)^2]\).
We estimate this variance using corrected scores. Here and below,
\[
\widehat f_{Y_t\mid I_{t-1}}
\!\left(m_{t\alpha}(\bar\theta_1,\widehat\theta_2)\right)
\]
denotes the fitted-quantile density estimator \(\widehat f_h\) from
Section~\ref{sec:bootstrap_power}, evaluated along the restricted plug-in
quantile curve \(u\mapsto m_{tu}(\bar\theta_1,\widehat\theta_2)\). The
corresponding sample Jacobian is
\[
\widehat L_\alpha
:=
\frac1n\sum_{t=1}^n
\widehat f_{Y_t\mid I_{t-1}}
\!\left(m_{t\alpha}(\bar\theta_1,\widehat\theta_2)\right)
g_{t\alpha}(\bar\theta_1,\widehat\theta_2)
g_{t\alpha}(\bar\theta_1,\widehat\theta_2)^\top .
\]
Using this density estimate and Jacobian, define
\[
\widehat l_{t,\alpha} := - \widehat L_\alpha^{-1} \, g_{t\alpha}(\bar{\theta}_1, \hat{\theta}_2) \, \Big[\mathbf{1}\{Y_t \le m_{t\alpha}(\bar{\theta}_1, \hat{\theta}_2)\} - \alpha\Big],
\]
\[
\widehat A_{2,n}(\alpha,\gamma) := \frac{1}{n}\sum_{t=1}^n \widehat f_{Y_t\mid I_{t-1}}(m_{t\alpha}(\bar{\theta}_1,\hat{\theta}_2))\, w_{t,n}(\gamma)\, \nabla_{\theta_2} m\!\big(I_{t-1},\bar{\theta}_1(\alpha),\hat\theta_2(\alpha)\big).
\]
Define the \textbf{corrected variance estimator} as the sample second moment of
the estimated corrected scores:
\begin{equation}
\begin{split}
\widehat s_n^2(\alpha, \gamma) := \frac{1}{n} \sum_{t=1}^n \Bigg( &
w_{t,n}(\gamma) \Big[ \mathbf{1}\{Y_t\le m_{t\alpha}(\bar{\theta}_1, \hat{\theta}_2)\} - \alpha \Big] \\
& + \widehat l_{t,\alpha}^\top \widehat A_{2,n}(\alpha,\gamma)
\Bigg)^2.
\end{split}
\label{eq:shat-def}
\end{equation}
We define the standardized plug-in process as
\[
\widehat Q_n(\alpha, \gamma) := \frac{\sqrt{n}\,|\widehat M_n(\alpha, \gamma)|}{\widehat s_n(\alpha, \gamma)}.
\]
Finally, the \textbf{penalized plug-in test statistic} uses the
same CvM--KS aggregation scheme:
\begin{equation}
\widehat T_{n,\Pi}(\lambda)
:=
\sup_{\gamma \in \Gamma}
\left[
\int_{\mathcal T}
\widehat Q_n^2(\alpha,\gamma)
d\Pi(\alpha)
-
\lambda\|\gamma\|_1
\right].
\label{eq:That-def}
\end{equation}
The next theorem gives the null limit.
\begin{theorem}[Pre-estimation null limit]
\label{thm:null-pre_main}
Suppose Assumption~S1--Assumption~S5\ hold, including the
generated-parameter compatibility condition
$n^{-1/2}a_{J_n}^{-2}\to0$ in Assumption~S3(d). Under the null hypothesis
$H_0: \theta_{10}(\alpha) = \bar{\theta}_1(\alpha)$ for all
$\alpha \in \mathcal{T}$, and with \(\Gamma\) interpreted as the
nondegenerate centered-weight index set described above and in Assumption~S4, the
following results hold:
\begin{enumerate}
\item The plug-in empirical process converges weakly:
\[
\sqrt{n} \widehat M_n(\alpha, \gamma) \rightsquigarrow \widetilde{\mathcal{M}}(\alpha, \gamma) \quad \text{in } \ell^{\infty}(\mathcal{T}\times\Gamma),
\]
where $\widetilde{\mathcal{M}}(\alpha, \gamma) := \mathcal{M}^c(\alpha, \gamma) + \mathcal{M}_{\mathrm{pre}}(\alpha, \gamma)$ is a tight, zero-mean Gaussian process. Here $\mathcal{M}^c$ denotes the Gaussian limit of the centered-weight process $\sqrt n M_n^c$, while $\mathcal{M}_{\mathrm{pre}}$ is the Gaussian process induced by the pre-estimation error. With $\mathcal{G}_{\theta_2}$ denoting the zero-mean Gaussian limit of the influence-function process $n^{-1/2}\sum_{t=1}^n l_{t,\alpha}$ (covariance kernel $K_{\theta_2}$; see Assumption~S4(ii) in the Supplementary Appendix), the pre-estimation component admits the representation
\[
\mathcal{M}_{\mathrm{pre}}(\alpha,\gamma) = A_2(\alpha,\gamma)^\top \mathcal{G}_{\theta_2}(\alpha),
\]
and the full covariance kernel of $\widetilde{\mathcal{M}}$ is
\[
\mathrm{Cov}\big(\widetilde{\mathcal{M}}(\alpha_1,\gamma_1),\widetilde{\mathcal{M}}(\alpha_2,\gamma_2)\big)
=
E\big[\Psi_t(\alpha_1,\gamma_1)\,\Psi_t(\alpha_2,\gamma_2)\big],
\]
where $\Psi_t$ is the population corrected score defined above.
\item The corrected variance estimator is uniformly consistent for the total asymptotic variance:
\[
\sup_{\alpha \in \mathcal{T}, \gamma \in \Gamma} \big|\, \widehat s_n^2(\alpha, \gamma) - \tilde{\sigma}^2(\alpha, \gamma) \,\big| \xrightarrow{p} 0,
\]
where
$\tilde{\sigma}^2(\alpha,\gamma)
=E[\Psi_t(\alpha,\gamma)^2]
=\mathrm{Var}\{\widetilde{\mathcal M}(\alpha,\gamma)\}$.
\item The test statistic therefore converges in distribution to
\[
\widehat T_{n,\Pi}(\lambda)
\rightsquigarrow
\sup_{\gamma \in \Gamma}
\Bigg[
\int_{\mathcal T}
\left(
\frac{\widetilde{\mathcal{M}}(\alpha,\gamma)}
{\tilde{\sigma}(\alpha,\gamma)}
\right)^2
d\Pi(\alpha)
-
\lambda\|\gamma\|_1
\Bigg].
\]
\end{enumerate}
\end{theorem}
\subsection{Bootstrap Critical Values with Pre-Estimation}
To approximate the null distribution of \(\widehat T_{n,\Pi}(\lambda)\), we keep
\(\hat\theta_2(\alpha)\) fixed and bootstrap the corrected score in the expansion
above.
Let $\{\omega_t\}_{t=1}^n$ be i.i.d.\ standard normal multipliers independent of
the data. Define the bootstrap analog of the plug-in empirical process by
\begin{equation}
\begin{split}
\widehat M_{n*}(\alpha,\gamma)
:= \frac{1}{n}\sum_{t=1}^n \omega_t \Bigg\{ &
\big[\mathbf 1\{Y_t\le m_{t\alpha}(\bar{\theta}_1,\hat\theta_2)\}-\alpha\big] w_{t,n}(\gamma) \\
& + \widehat l_{t,\alpha}^\top\, \widehat A_{2,n}(\alpha,\gamma)
\Bigg\}.
\end{split}
\label{eq:Mhat-star-def}
\end{equation}
Here $\widehat l_{t,\alpha}$ and
$\widehat A_{2,n}(\alpha,\gamma)$ are the estimates defined above.
Define the bootstrap variance estimator by
\begin{equation}
\begin{split}
\widehat s_{n*}^2(\alpha,\gamma)
:= \frac{1}{n}\sum_{t=1}^n \omega_t^2 \Bigg\{ &
\big[\mathbf 1\{Y_t\le m_{t\alpha}(\bar{\theta}_1,\hat\theta_2)\}-\alpha\big] w_{t,n}(\gamma) \\
& + \widehat l_{t,\alpha}^\top\, \widehat A_{2,n}(\alpha,\gamma)
\Bigg\}^2.
\end{split}
\label{eq:shat-star-def}
\end{equation}
We define the standardized bootstrap process as
\[
\widehat Q_{n*}(\alpha,\gamma)
:=
\frac{\sqrt n|\widehat M_{n*}(\alpha,\gamma)|}
{\widehat s_{n*}(\alpha,\gamma)}.
\]
The penalized bootstrap statistic is then constructed using the same
aggregation scheme as the original test:
\begin{equation}
\widehat T_{n*,\Pi}(\lambda)
:=
\sup_{\gamma\in\Gamma}
\Bigg[
\int_{\mathcal T}
\widehat Q_{n*}^2(\alpha,\gamma)
d\Pi(\alpha)
-
\lambda\|\gamma\|_1
\Bigg].
\label{eq:That-star-def}
\end{equation}
\begin{theorem}[Bootstrap validity with pre-estimation]
\label{thm:boot-pre_main}
\leavevmode\par\noindent
Suppose Assumption~S1--Assumption~S5\ hold, including
$n^{-1/2}a_{J_n}^{-2}\to0$ in Assumption~S3(d), and
$H_0:\theta_{10}(\alpha)=\bar{\theta}_1(\alpha)$ is true for all
$\alpha\in\mathcal T$.
Then, conditional on the data and in $\ell^\infty(\mathcal T\times\Gamma)$:
\begin{enumerate}
\item The bootstrap process converges to the limiting process:
\[
\sqrt n\,\widehat M_{n*}(\alpha,\gamma)
\ \overset{*}{\rightsquigarrow}\
\widetilde{\mathcal M}(\alpha,\gamma),
\]
where $\widetilde{\mathcal M}(\alpha,\gamma)$ is the same limiting Gaussian process as in Theorem~\ref{thm:null-pre_main}.
\item The bootstrap variance estimator consistently estimates the total asymptotic variance:
\[
\sup_{\alpha,\gamma}
\Big|
\widehat s_{n*}^2(\alpha,\gamma)
- \tilde{\sigma}^2(\alpha, \gamma)
\Big|
\ \xrightarrow{p^*}\ 0.
\]
\item The bootstrap statistic converges to the null limit:
\[
\widehat T_{n*,\Pi}(\lambda)
\ \overset{*}{\rightsquigarrow}\
\sup_{\gamma\in\Gamma}
\Bigg[
\int_{\mathcal T}
\left(
\frac{\widetilde{\mathcal M}(\alpha,\gamma)}
{\tilde{\sigma}(\alpha,\gamma)}
\right)^2
d\Pi(\alpha)
-
\lambda\|\gamma\|_1
\Bigg].
\]
\end{enumerate}
Here \(\xrightarrow{p^*}\) and \(\overset{*}{\rightsquigarrow}\) denote convergence in probability and weak convergence conditional on the data, respectively.
\end{theorem}
\subsection{Local Power with Pre-Estimation}
We now consider Pitman local alternatives for the parameter of interest:
\begin{equation}
\theta_{1n,B}(\alpha)
=
\theta_{10}(\alpha)
-
\frac{B(\alpha)}{\sqrt n},
\qquad \alpha\in\mathcal T,
\label{eq:pre-local-alt}
\end{equation}
where $B$ is an admissible bounded continuous direction in the tangent set of
Assumption~S6, so the local sequence remains a valid conditional-quantile model for
all sufficiently large $n$. The data are generated under
\eqref{eq:pre-local-alt}, but the implemented test is still constructed under
the null-imposed value \(\bar\theta_1(\alpha)=\theta_{10}(\alpha)\). Hence the
nuisance parameter is estimated under the null restriction by solving
\eqref{eq:theta2-est} with \(\theta_1(\alpha)\) fixed at \(\theta_{10}(\alpha)\).
We denote this restricted estimator under the local sequence by
\[
\hat\theta_{2,0n}(\alpha):=\hat\theta_2(\theta_{10},\alpha).
\]
The null-evaluated plug-in process under the local sequence is
\[
\widehat M_{n,0}^{B}(\alpha,\gamma)
=
\frac1n\sum_{t=1}^n
w_{t,n}(\gamma)
\Big[
\mathbf 1\{Y_t\le m_{t\alpha}(\theta_{10},\hat\theta_{2,0n})\}
-
\alpha
\Big].
\]
For notational clarity, write
\[
U_{nt}(\alpha,\gamma)
:=
w_t(\gamma)
\Big[
\mathbf 1\{Y_t\le m_{t\alpha}(\theta_{10},\theta_{20})\}-\alpha
\Big],
\]
and
\[
V_{nt}(\alpha)
:=
g_{t\alpha}(\theta_{10},\theta_{20})
\Big[
\mathbf 1\{Y_t\le m_{t\alpha}(\theta_{10},\theta_{20})\}-\alpha
\Big],
\]
where replacing the sample-centered weight \(w_{t,n}\) by its population version
\(w_t\) only changes the expansion by \(o_p(1)\). Define
\[
A_1(\alpha,\gamma)
:=
E\!\left[
f_{Y_t|I_{t-1}}\!\left(m_{t\alpha}(\theta_{10},\theta_{20})\right)
w_t(\gamma)
\nabla_{\theta_1}m_{t\alpha}(\theta_{10},\theta_{20})
\right],
\]
and
\[
D_\alpha
:=
E\!\left[
g_{t\alpha}(\theta_{10},\theta_{20})
f_{Y_t|I_{t-1}}\!\left(m_{t\alpha}(\theta_{10},\theta_{20})\right)
\nabla_{\theta_1}^{\top}m_{t\alpha}(\theta_{10},\theta_{20})
\right],
\qquad
H_\alpha:=-L_\alpha^{-1}D_\alpha.
\]
The corresponding local mean expansions derived in the Appendix give
\[
\frac1{\sqrt n}\sum_{t=1}^nE_{n,B}U_{nt}(\alpha,\gamma)
=
A_1(\alpha,\gamma)^\top B(\alpha)+o(1),
\]
and
\[
\frac1{\sqrt n}\sum_{t=1}^nE_{n,B}V_{nt}(\alpha)
=
D_\alpha B(\alpha)+o(1),
\]
uniformly in \((\alpha,\gamma)\). The restricted nuisance estimator satisfies
\begin{equation}
\label{eq:delta_2n_full}
\Delta_{2n}(\alpha)
:=
\sqrt n\{\hat\theta_{2,0n}(\alpha)-\theta_{20}(\alpha)\}
=
\mathbb L_{n,B}(\alpha)+H_\alpha B(\alpha)+o_p(1),
\end{equation}
where
\[
\mathbb L_{n,B}(\alpha)
:=
-L_\alpha^{-1}
\frac1{\sqrt n}\sum_{t=1}^n
\Big\{
V_{nt}(\alpha)-E_{n,B}V_{nt}(\alpha)
\Big\}.
\]
The first term is the centered nuisance-estimation fluctuation, while
\(H_\alpha B(\alpha)\) is the local bias induced by estimating the nuisance
component under the null restriction.
A Taylor expansion of the plug-in process around \((\theta_{10},\theta_{20})\)
then yields
\begin{equation}
\label{eq:pre-local-expansion}
\sqrt n\,\widehat M_{n,0}^{B}(\alpha,\gamma)
=
\mathbb G_{n,B}^{\mathrm{pre}}(\alpha,\gamma)
+
d_{\mathrm{pre},B}(\alpha,\gamma)
+o_p(1),
\end{equation}
where the centered Gaussian component is
\[
\mathbb G_{n,B}^{\mathrm{pre}}(\alpha,\gamma)
:=
\frac1{\sqrt n}\sum_{t=1}^n
\Big\{
U_{nt}(\alpha,\gamma)-E_{n,B}U_{nt}(\alpha,\gamma)
\Big\}
+
A_2(\alpha,\gamma)^\top\mathbb L_{n,B}(\alpha),
\]
and the deterministic local drift is
\begin{equation}
\label{eq:pre-local-drift}
d_{\mathrm{pre},B}(\alpha,\gamma)
:=
A_1(\alpha,\gamma)^\top B(\alpha)
+
A_2(\alpha,\gamma)^\top H_\alpha B(\alpha).
\end{equation}
The drift has two components: the direct term
\(A_1(\alpha,\gamma)^\top B(\alpha)\) and the additional term
\(A_2(\alpha,\gamma)^\top H_\alpha B(\alpha)\) induced by restricted nuisance
estimation. Under contiguity and the weak convergence assumptions in
Assumption~S4--Assumption~S6,
\begin{equation}
\label{eq:final_decomposition}
\sqrt n\,\widehat M_{n,0}^{B}(\alpha,\gamma)
\rightsquigarrow
\widetilde{\mathcal M}(\alpha,\gamma)
+
d_{\mathrm{pre},B}(\alpha,\gamma),
\end{equation}
where \(\widetilde{\mathcal M}\) is the zero-mean Gaussian process from
Theorem~\ref{thm:null-pre_main}. The deterministic drift
\(d_{\mathrm{pre},B}\) is the component that drives local power after
pre-estimation.
\begin{remark}[Effect of restricted nuisance estimation on local power]
\label{rem:pre-drift_main}
The second component \(A_2(\alpha,\gamma)^\top H_\alpha B(\alpha)\) of the
drift in \eqref{eq:pre-local-drift} arises because the nuisance parameter is
estimated under the null restriction
\(\theta_{10}(\alpha)=\bar\theta_1(\alpha)\) even though the data are generated
under the local sequence. Depending on the sign of the projection
\(A_2(\alpha,\gamma)^\top H_\alpha B(\alpha)\) relative to the direct drift
\(A_1(\alpha,\gamma)^\top B(\alpha)\), pre-estimation can either attenuate or
amplify the effective local signal at a given direction \((\alpha,\gamma)\); it
can also shift the set of \(\gamma\)-directions along which the studentized
drift concentrates. This directional distortion is one source of the
design-dependence of the adaptive-versus-unpenalized power comparisons in
Section~\ref{sec:simulation}: isolating a structural parameter by
pre-estimation changes the drift geometry in a way that the
\(\ell_1\)-penalization can exploit when the remaining drift is concentrated on
directions of small \(\ell_1\)-norm.
\end{remark}
\subsection{Construction of the Shifted Bootstrap Process}
\label{subsec:shifted-bootstrap}
To select \(\lambda\) with pre-estimated nuisance parameters, the bootstrap must
approximate the noncentral limit in \eqref{eq:final_decomposition}. The multiplier
term reproduces the centered Gaussian component \(\widetilde{\mathcal M}\), while
a deterministic shift estimates \(d_{\mathrm{pre},B}(\alpha,\gamma)\).
Let \(\widehat d_{\mathrm{pre},n,B}(\alpha,\gamma)\) denote a uniformly
consistent sample analog of \(d_{\mathrm{pre},B}(\alpha,\gamma)\). Define
\[
\widehat A_{1,n}(\alpha,\gamma)
:=
\frac1n\sum_{t=1}^n
\widehat f_{Y_t|I_{t-1}}
\!\left(m_{t\alpha}(\bar\theta_1,\hat\theta_2)\right)
w_{t,n}(\gamma)
\nabla_{\theta_1}m_{t\alpha}(\bar\theta_1,\hat\theta_2),
\]
\[
\widehat D_{\alpha,n}
:=
\frac1n\sum_{t=1}^n
g_{t\alpha}(\bar\theta_1,\hat\theta_2)
\widehat f_{Y_t|I_{t-1}}
\!\left(m_{t\alpha}(\bar\theta_1,\hat\theta_2)\right)
\nabla_{\theta_1}^{\top}m_{t\alpha}(\bar\theta_1,\hat\theta_2),
\qquad
\widehat H_{\alpha,n}:=-\widehat L_\alpha^{-1}\widehat D_{\alpha,n}.
\]
We set
\begin{equation}
\label{eq:dpre-hat}
\widehat d_{\mathrm{pre},n,B}(\alpha,\gamma)
:=
\widehat A_{1,n}(\alpha,\gamma)^\top B(\alpha)
+
\widehat A_{2,n}(\alpha,\gamma)^\top\widehat H_{\alpha,n}B(\alpha),
\end{equation}
which is obtained from \eqref{eq:pre-local-drift} by replacing population
quantities with their sample analogues.
For compactness, define the estimated corrected score
\[
\widehat\Psi_{t,n}(\alpha,\gamma)
:=
\big[\mathbf 1\{Y_t\le m_{t\alpha}(\bar\theta_1,\hat\theta_2)\}-\alpha\big]
w_{t,n}(\gamma)
+
\widehat l_{t,\alpha}^\top\widehat A_{2,n}(\alpha,\gamma).
\]
For a local direction \(B\), the shifted bootstrap process is
\begin{equation}
\label{eq:pre-local-bootstrap-M}
\widehat M_{n*,B}(\alpha,\gamma)
:=
\frac1n\sum_{t=1}^n
\left[
\omega_t\widehat\Psi_{t,n}(\alpha,\gamma)
+
\frac{\widehat d_{\mathrm{pre},n,B}(\alpha,\gamma)}{\sqrt n}
\right].
\end{equation}
The multiplier component is conditionally centered because \(E^*\omega_t=0\),
and the deterministic shift generates the local drift after multiplication by
\(\sqrt n\).
The shifted bootstrap variance estimator is
\begin{equation}
\label{eq:pre-local-bootstrap-s}
\widehat s_{n*,B}^2(\alpha,\gamma)
:=
\frac1n\sum_{t=1}^n
\left[
\omega_t\widehat\Psi_{t,n}(\alpha,\gamma)
+
\frac{\widehat d_{\mathrm{pre},n,B}(\alpha,\gamma)}{\sqrt n}
\right]^2.
\end{equation}
Define
\[
\widehat Q_{n*,B}(\alpha,\gamma)
:=
\sqrt n
\frac{|\widehat M_{n*,B}(\alpha,\gamma)|}
{\widehat s_{n*,B}(\alpha,\gamma)}.
\]
The shifted bootstrap statistic is
\begin{equation}
\widehat T_{n*,B,\Pi}(\lambda)
:=
\sup_{\gamma\in\Gamma}
\Bigg[
\int_{\mathcal T}\widehat Q_{n*,B}^2(\alpha,\gamma)d\Pi(\alpha)
-
\lambda\|\gamma\|_1
\Bigg].
\label{eq:That-star-local-def}
\end{equation}
As above, $B$ denotes the candidate shift used to estimate power, while
$B_0$ denotes the actual local data-generating direction; $B_0=0$ is the
null.
\begin{theorem}[Shifted multiplier bootstrap validity with pre-estimation]
\label{thm:boot-local-pre_main}
Suppose Assumption~S1--Assumption~S6\ hold, including
$n^{-1/2}a_{J_n}^{-2}\to0$ in Assumption~S3(d). For any admissible candidate direction
$B$, under the null or any actual local sequence $P_{n,B_0}$ covered by
Assumption~S6, the following results hold conditionally on the data, with process
convergence in \(\ell^\infty(\mathcal T\times\Gamma)\):
\begin{enumerate}
\item The shifted process converges to the noncentral limit:
\[
\sqrt n\,\widehat M_{n*,B}(\alpha,\gamma)
\overset{*}{\rightsquigarrow}
\widetilde{\mathcal M}(\alpha,\gamma)
+
d_{\mathrm{pre},B}(\alpha,\gamma).
\]
\item The variance estimator remains uniformly consistent:
\[
\sup_{\alpha,\gamma}
\big|\widehat s_{n*,B}^2(\alpha,\gamma)-\tilde\sigma^2(\alpha,\gamma)\big|
\xrightarrow{p^*}0.
\]
\item The penalized shifted bootstrap statistic captures the noncentral
distribution:
\[
\widehat T_{n*,B,\Pi}(\lambda)
\overset{*}{\rightsquigarrow}
\sup_{\gamma\in\Gamma}
\Bigg[
\int_{\mathcal T}
\left(
\frac{\widetilde{\mathcal M}(\alpha,\gamma)
+d_{\mathrm{pre},B}(\alpha,\gamma)}
{\tilde\sigma(\alpha,\gamma)}
\right)^2
d\Pi(\alpha)
-
\lambda\|\gamma\|_1
\Bigg].
\]
\end{enumerate}
\end{theorem}
\subsection{Adaptive Selection of the Penalty Parameter}
\label{subsec:lambda-selection}
We select $\lambda$ by maximizing estimated minimum local power over the
prespecified class $\mathcal B$.
As in Section~\ref{subsec:calibration}, the theoretical penalty set is a compact interval
\[
\Lambda=[0,\bar\lambda].
\]
Writing $\widehat W_n(\lambda)=\inf_{B\in\mathcal B}
\widehat{\mathcal R}_n(\lambda,B,\tau)$ for the corrected shifted-bootstrap
criterion, the continuous selector satisfies
\begin{equation}
\widehat W_n(\widehat\lambda)
\ge \sup_{\lambda\in\Lambda}\widehat W_n(\lambda)-\eta_n,
\qquad \eta_n=o_p(1),\quad\eta_n\ge0,
\label{eq:maxmin-pre}
\end{equation}
with a measurable choice of $\widehat\lambda$, as in \eqref{eq:maxmin}.
The interval
$\Lambda$ may be searched directly or approximated by a finite grid
$\Lambda_G\subset\Lambda$. A fixed grid targets the grid maximizer; approximation
to the continuous population maximizer additionally requires a vanishing mesh and
continuity of the population criterion. The criterion computed from finitely many
bootstrap draws may be stepwise. The multiplier sequences are drawn once and
reused across all $\lambda$ and $B$. Algorithm~\ref{alg:adaptive-pre} summarizes
the procedure.
\begin{algorithm}[htbp]
\caption{Adaptive penalty selection with pre-estimated nuisance parameters}
\label{alg:adaptive-pre}
\begin{algorithmic}[1]
\Require Data \(\mathcal Z_n\), \(\Pi\), \(\Gamma\), search set
\(\mathcal S=\Lambda\) or \(\Lambda_G\), finite \(\mathcal B\), \(R\), and \(\tau\)
\State Estimate the restricted nuisance curve \(\widehat\theta_2(\alpha)\)
under \(H_0\).
\State Estimate the fitted-quantile density and construct
\(\widehat L_\alpha\), \(\widehat l_{t,\alpha}\),
\(\widehat A_{2,n}(\alpha,\gamma)\), and \(\widehat\Psi_{t,n}(\alpha,\gamma)\).
\State Draw \(R\) multiplier sequences and reuse them for all \(\lambda\) and
\(B\).
\For{each penalty \(\lambda\) requested by the search routine}
\State Compute \(\widehat T_{n*,\Pi}^{(r)}(\lambda)\), \(r=1,\ldots,R\),
from \eqref{eq:Mhat-star-def}.
\State Set \(\widehat c_{n,1-\tau,\Pi}(\lambda)\) to their empirical
\((1-\tau)\)-quantile.
\For{each \(B\in\mathcal B\)}
\State Construct \(\widehat d_{\mathrm{pre},n,B}\) and compute
\(\widehat T_{n*,B,\Pi}^{(r)}(\lambda)\), \(r=1,\ldots,R\).
\State Set
\(\widehat{\mathcal R}_n(\lambda,B,\tau)
=R^{-1}\sum_{r=1}^R
\mathbf 1\{\widehat T_{n*,B,\Pi}^{(r)}(\lambda)>
\widehat c_{n,1-\tau,\Pi}(\lambda)\}\).
\EndFor
\State Set \(\widehat W_n(\lambda)=
\min_{B\in\mathcal B}\widehat{\mathcal R}_n(\lambda,B,\tau)\).
\EndFor
\State On \(\Lambda_G\), select the smallest maximizer of \(\widehat W_n\);
for continuous search, use a measurable \(\eta_n\)-approximate maximizer,
\(\eta_n=o_p(1)\).
\State Compute \(\widehat T_{n,\Pi}(\widehat\lambda)\) and reject when it
exceeds \(\widehat c_{n,1-\tau,\Pi}(\widehat\lambda)\).
\end{algorithmic}
\end{algorithm}
For the reported finite-grid implementation, the selection line is
\begin{equation}
\widehat{\lambda}_G
=
\arg\max_{\lambda \in \Lambda_G}
\min_{B \in \mathcal{B}}
\widehat{\mathcal{R}}_n(\lambda, B, \tau),
\label{eq:maxmin-pre-grid}
\end{equation}
with the smallest maximizer selected. The continuous derivative-free alternatives
described after Algorithm~\ref{alg:adaptive-known} apply without change.
The test rejects when \(\widehat T_{n,\Pi}(\widehat\lambda)\) exceeds
\(\widehat c_{n,1-\tau,\Pi}(\widehat\lambda)\), with \(\widehat\lambda_G\) used in
the finite-grid implementation. Corollary~\ref{cor:adaptive-lambda_main} gives
first-order validity when the bootstrap approximation and critical values are
uniform over \(\Lambda\) and the selector converges to a deterministic limit. On
a fixed finite grid, uniformity is required only over the grid. Convergence to the
continuous optimum additionally requires the grid mesh to vanish.
\begin{remark}[Pre-estimation analog of the maximin no-loss result]
Replace the ordinary local limit \(\mathcal M+d_B\) in
Theorem~\ref{thm:maximin-no-loss_main} by
\(\widetilde{\mathcal M}+d_{\mathrm{pre},B}\), and use the shifted plug-in
bootstrap in Algorithm~\ref{alg:adaptive-pre}. Under the corresponding uniform
bootstrap, anti-concentration, Monte Carlo, and unique-maximizer conditions, the
same argument gives the maximin local-power no-loss conclusion for the
pre-estimated procedure.
\end{remark}
\section{Monte Carlo Simulations}
\label{sec:simulation}
This section reports finite-sample rejection rates for \(T_{n,\Pi_A}\) and the
two alternative aggregation schemes. The main design is a Gaussian AR(2) model;
additional designs are reported in the Supplementary Appendix.
\subsection{Data-Generating Process and Quantile Specification}
The data are generated from a second-order autoregressive process (AR(2)) with Gaussian innovations:
\begin{equation}
Y_t = \mu_0 + \mu_1 Y_{t-1} + \mu_2 Y_{t-2} + \sigma \varepsilon_t, \quad t = 1, \dots, n,
\end{equation}
where $\{\varepsilon_t\}$ are i.i.d.\ $\mathcal N(0,1)$.
The conditional $\alpha$-quantile of $Y_t$ given
$I_{t-1}=(Y_{t-1},Y_{t-2})^\top$ is
\begin{equation}
m(I_{t-1}, \theta_0(\alpha)) = \beta_0(\alpha) + \beta_1(\alpha) Y_{t-1} + \beta_2(\alpha) Y_{t-2}.
\label{eq:quantile_model}
\end{equation}
The quantile-specific coefficient vector
$\theta_0(\alpha)=(\beta_0(\alpha),\beta_1(\alpha),\beta_2(\alpha))^\top$ is
given by
\[
\beta_0(\alpha) = \mu_0 + \sigma \Phi^{-1}(\alpha), \quad \beta_1(\alpha) = \mu_1, \quad \beta_2(\alpha) = \mu_2,
\]
where $\Phi^{-1}(\cdot)$ denotes the inverse cumulative distribution function of the standard normal distribution. Note that under this location-shift model, only the intercept $\beta_0(\alpha)$ varies across quantiles, while the autoregressive coefficients remain constant.
The simulation uses
\[
(\mu_0, \mu_1, \mu_2, \sigma)^\top = (0.5, 0.4, -0.2, 1.0)^\top.
\]
Power curves use nominal level \(\tau=0.10\); the size tables also report
\(\tau=0.05\). All reported simulation exercises use 1,000 Monte Carlo
replications and 1,000 multiplier draws per replication. Without
pre-estimation, the statistic is evaluated at nine equally spaced quantile
nodes on \([0.10,0.90]\). With pre-estimation, it is evaluated at seven
equally spaced test nodes on \([0.20,0.80]\), while the restricted quantile
curve is estimated on 100 equally spaced dense-grid nodes on \([0.15,0.85]\).
Sample-size labels refer to the generated series length.
\paragraph{Implementation of the Adaptive Test}
The main statistic uses the discrete measure
\(\Pi_A=A^{-1}\sum_{j=1}^A\delta_{\alpha_j}\) and a finite weighting grid
\(\Gamma_G\). Thus the quantile aggregation and directional search are
evaluated over the reported nodes. Here \(\Gamma_G\) is the grid of weighting
directions \(\gamma\), and is distinct from both the quantile nodes and the
reference class \(\mathcal B\). Writing \(E_m[a,b]\) for \(m\) equally spaced
points including the endpoints, the known-parameter main power exercise uses
\(\Gamma_G=E_5[-1,1]^2\), a \(5\times5\) grid. For the pre-estimation power
exercise and the corresponding size comparisons, the origin is removed:
\[
\Gamma_G^{\mathrm{pre}}
=E_5[-1,1]^2\setminus\{(0,0)^\top\}.
\]
This 24-point grid excludes the degenerate centered score at \(\gamma=0\).
The penalty is selected from
\(\{0,0.1,\ldots,1\}\); this grid contains the unpenalized value zero and
targets the grid maximizer in Algorithms~\ref{alg:adaptive-known}
and~\ref{alg:adaptive-pre}. The reference class used to select the penalty is
\(\mathcal B=\{(0.5,0.5,0.5)^\top\}\) for Figure~\ref{fig:power_no_pre},
in coordinates \((\beta_0(\alpha),\mu_1,\mu_2)\), and
\(\mathcal B=\{1\}\) in the \(\mu_2\) direction for
Figure~\ref{fig:power_pre} and Panel B of Table~\ref{tab:size_comparison}.
These reference directions are distinct from the fixed alternatives varied
along the power-curve axes. Common multiplier draws are reused across
candidate penalties within each replication. In the nonlinear simulations
reported in the Supplementary Appendix, the known-parameter runs use
\(\Gamma_G=E_{10}[-1,1]^2\), the penalty grid
\(\{0,0.1,\ldots,2.0\}\), and the base reference vector
\((0.2,0.2,0.2,0.2)^\top\) in coordinates
\((\beta_0(\alpha),\mu_1,\rho,\mu_2)\). The pre-estimation runs use the
same weighting grid, the penalty grid \(\{0,0.1,\ldots,0.9\}\), and
\(\mathcal B=\{1\}\) in the \(\rho\) direction. For the optional KS--KS
quantile-index penalty in the pre-estimation size comparisons, \(\eta\) is
searched over \(\{0,0.05,\ldots,0.25\}\).
\paragraph{Density Estimation in the Pre-Estimation Scenarios}
For implementation, the reported pre-estimation simulations estimate the
restricted quantile curve on the 100-point dense grid over \([0.15,0.85]\) and use linear
interpolation, with endpoint extrapolation when required, to evaluate the curve
at the test and auxiliary indices. The density routine draws \(J=1{,}000\)
auxiliary quantile indices from \([0.01,0.99]\) and applies a Gaussian kernel.
\subsection{Finite-Sample Power without Pre-Estimation
\texorpdfstring{($n=500$)}{(n=500)}}
We first consider a benchmark with no pre-estimated nuisance parameter: the
null-imposed quantile curve is treated as fixed. In this Gaussian design, the
full-curve null for
\(\theta_0(\alpha)=(\beta_0(\alpha),\mu_1,\mu_2)^\top\), where
\(\beta_0(\alpha)=\mu_0+\sigma\Phi^{-1}(\alpha)\), is equivalent to the joint
null for the four DGP primitives
\((\mu_0,\mu_1,\mu_2,\sigma)^\top\). The sample size is $n=500$. We report
rejection frequencies under fixed alternatives by perturbing one DGP primitive
at a time over $\delta\in[-0.2,0.2]$. In Figure~\ref{fig:power_no_pre}, the solid line shows the
adaptive test, the dashed line shows the unpenalized test, and the shaded area
marks parameter values for which the adaptive test rejects more often.
\begin{figure}[htbp]
\centering
\includegraphics[width=1.0\textwidth]{Figure_1.png}
\caption{Finite-sample rejection frequencies without pre-estimation
($n=500$, nominal level $0.10$). The four panels vary one DGP primitive at a
time from its null value. The gray shaded area marks parameter values for
which the adaptive statistic has a higher rejection frequency than the
unpenalized statistic.}
\label{fig:power_no_pre}
\end{figure}
Figure~\ref{fig:power_no_pre} shows that, for $n=500$, the adaptive test has higher empirical rejection probabilities for several intermediate deviations. The gray areas mark regions in which the adaptive statistic rejects more often than the unpenalized statistic.
\subsection{Finite-Sample Power with Pre-Estimation
\texorpdfstring{($n=1000$)}{(n=1000)}}
Next, we evaluate the test's performance when a quantile-specific nuisance
function must be estimated. We focus on testing the lag-2 coefficient \(\mu_2\).
Under \(H_0:\mu_2=\mu_{2,0}=-0.2\), the nuisance function is
\[
\theta_2(\alpha)
=
\big(\beta_0(\alpha),\beta_1(\alpha)\big)^\top.
\]
At each quantile, it is estimated by restricted quantile regression of
\(Y_t-\mu_{2,0}Y_{t-2}\) on \((1,Y_{t-1})^\top\). In the Gaussian location-shift
DGP, \(\beta_0(\alpha)=\mu_0+\sigma\Phi^{-1}(\alpha)\), so \(\mu_0\) and
\(\sigma\) are absorbed into the quantile-specific intercept rather than
estimated separately at a fixed \(\alpha\). The reported sample size is
\(n=1000\).
The restricted model is estimated on 100 equally spaced points in
\([0.15,0.85]\); the statistic is evaluated at seven equally spaced points in
\([0.20,0.80]\). The reported critical values use the plug-in bootstrap
implementation described above. The studentized search grid is
\[
\Gamma_G^{\mathrm{pre}}
=E_5[-1,1]^2\setminus\{(0,0)^\top\},
\]
which contains 24 nonzero directions. The origin is excluded because the
centered score is degenerate at \(\gamma=0\), and no variance is formed there.
The same 24-point grid is used for the observed and multiplier-bootstrap
statistics.
\begin{figure}[htbp]
\centering
\includegraphics[width=0.8\textwidth]{Figure_2.png}
\caption{Finite-sample rejection frequencies with pre-estimation
($n=1000$, nominal level $0.10$). The test concerns
$H_0:\mu_2=-0.2$, while the quantile-specific nuisance coefficients
$\{\beta_0(\alpha),\beta_1(\alpha)\}$ are estimated by restricted quantile
regression on the 100-point dense grid. The gray area marks parameter values
for which the adaptive statistic has a higher rejection frequency than the
unpenalized statistic.}
\label{fig:power_pre}
\end{figure}
Figure~\ref{fig:power_pre} reports rejection frequencies for deviations in
$\mu_2$. In the shaded region, the adaptive statistic rejects more often than
the unpenalized statistic in the reported design with $n=1000$.
Table~\ref{tab:size_comparison} concerns the linear AR(2) design.
Section~S5.3 of the Supplementary Appendix considers a nonlinear model with a
signed-power transformation. In that design, the pre-estimation correction has
substantial finite-sample distortion at $n=200$, with rejection rates moving
closer to the nominal level as $n$ increases. The same simulations also show
that the relative performance of the adaptive and unpenalized tests depends on
the hypothesis. When the full null quantile curve is tested without
pre-estimation, the unpenalized test has
slightly higher rejection frequencies for departures in $\rho$; when $\rho$ is
isolated through pre-estimation, the adaptive test rejects more often in the
reported design. These comparisons are design-specific and do not imply general
power dominance.
\subsection{Comparison of Aggregation Schemes under Pre-Estimation}
\label{subsec:aggregation-comparison}
We compare the size of \(T_{n,\Pi}^{CvM\text{-}KS}\) with
\(T_n^{KS\text{-}KS}\) and \(T_{n,\Pi}^{KS\text{-}CvM}\), defined in
Remark~\ref{rem:alternative-schemes}. Table~\ref{tab:size_comparison} reports
null rejection rates for the adaptive and unpenalized versions, both with and
without pre-estimated nuisance parameters.
\begin{table}[htbp]
\centering
\caption{Empirical Size Comparison of Different Aggregation Schemes}
\label{tab:size_comparison}
\resizebox{\textwidth}{!}{
\begin{tabular}{l c ccc c ccc}
\toprule
\multirow{2}{*}{$n$} & \multirow{2}{*}{Nominal $\tau$} & \multicolumn{3}{c}{Adaptive Penalization (Adapt)} & & \multicolumn{3}{c}{Unpenalized Benchmark (Unpenal)} \\
\cmidrule{3-5} \cmidrule{7-9}
& & \textbf{$T_{n,\Pi}$ (Ours)} & $T_n^{KS\text{-}KS}$ & $T_{n,\Pi}^{KS\text{-}CvM}$ & & \textbf{$T_{n,\Pi}$ (Ours)} & $T_n^{KS\text{-}KS}$ & $T_{n,\Pi}^{KS\text{-}CvM}$ \\
\midrule
\multicolumn{9}{l}{\textbf{Panel A: Without Pre-estimation (Fixed Null Quantile Curve)}} \\
\midrule
200 & 0.10 & 0.100 & 0.132 & 0.118 & & 0.095 & 0.159 & 0.126 \\
& 0.05 & 0.054 & 0.064 & 0.064 & & 0.044 & 0.083 & 0.065 \\
500 & 0.10 & 0.109 & 0.152 & 0.122 & & 0.104 & 0.176 & 0.123 \\
& 0.05 & 0.052 & 0.082 & 0.059 & & 0.056 & 0.098 & 0.066 \\
1000 & 0.10 & 0.109 & 0.136 & 0.126 & & 0.109 & 0.153 & 0.117 \\
& 0.05 & 0.057 & 0.065 & 0.062 & & 0.061 & 0.090 & 0.060 \\
\midrule
\multicolumn{9}{l}{\textbf{Panel B: With Pre-estimation (Estimated Nuisance Parameters)}} \\
\midrule
200 & 0.10 & \textbf{0.109} & 0.202 & 0.164 & & \textbf{0.104} & 0.162 & 0.131 \\
& 0.05 & \textbf{0.046} & 0.104 & 0.085 & & \textbf{0.042} & 0.089 & 0.065 \\
500 & 0.10 & \textbf{0.091} & 0.151 & 0.121 & & \textbf{0.081} & 0.153 & 0.102 \\
& 0.05 & \textbf{0.040} & 0.083 & 0.065 & & \textbf{0.034} & 0.083 & 0.052 \\
1000 & 0.10 & \textbf{0.096} & 0.128 & 0.126 & & \textbf{0.093} & 0.138 & 0.110 \\
& 0.05 & \textbf{0.043} & 0.066 & 0.049 & & \textbf{0.041} & 0.083 & 0.056 \\
\bottomrule
\end{tabular}
}
\end{table}
Panel A shows that the proposed \(T_{n,\Pi}\) statistic is close to the nominal
levels, whereas the two alternative aggregation schemes over-reject in several
cases even without pre-estimation. For example, at \(n=200\) and \(\tau=0.10\),
the adaptive rejection rates are \(0.100\), \(0.132\), and \(0.118\) for
\(T_{n,\Pi}\), \(T_n^{KS\text{-}KS}\), and
\(T_{n,\Pi}^{KS\text{-}CvM}\), respectively.
With pre-estimation, the two statistics that place a supremum earlier in the
aggregation order have larger rejection rates in this design. At \(n=200\) and
\(\tau=0.10\), the adaptive rejection rate is \(0.202\) for
\(T_n^{KS\text{-}KS}\) and \(0.164\) for
\(T_{n,\Pi}^{KS\text{-}CvM}\), compared with \(0.109\) for
\(T_{n,\Pi}^{CvM\text{-}KS}\). Because the comparator penalties also enter differently, these differences
do not isolate the effect of aggregation order alone.
In the linear AR(2) design, \(T_{n,\Pi}^{CvM\text{-}KS}\) is relatively close
to the nominal level in most reported comparisons. The nonlinear results in the
Supplementary Appendix show that this advantage does not eliminate finite-sample
distortion when the structural parameter is highly nonlinear.
\section{Empirical Study}\label{sec:empirical}
We apply the proposed penalized Bierens maximum statistic to the Growth-at-Risk
(GaR) framework of \citet{adrian2019vulnerable} to conduct uniform inference on
the association between financial conditions and the conditional distribution
of GDP growth. The original GaR analysis documents that tighter financial
conditions are especially informative about lower conditional quantiles of
future growth. Related work develops formal out-of-sample tests of conditional
quantile coverage and GaR model comparisons \citep{corradi2023outofsample}.
Our application addresses two in-sample restrictions: whether the NFCI
coefficient is uniformly zero and whether it can be represented by a common
constant over the quantile range under study.
\subsection{Data and Model Specification}
\subsubsection{Data}
Following \citet{adrian2019vulnerable}, we use U.S.\ macroeconomic data from the
Federal Reserve Economic Data (FRED) database; the files used for the reported
results were downloaded on March 6, 2026. The dependent variable is annualized
quarter-over-quarter real GDP growth (\texttt{A191RL1Q225SBEA}). The Chicago Fed
National Financial Conditions Index (NFCI) is observed weekly and converted to
quarterly frequency using the last weekly observation in each quarter. Higher
NFCI values correspond to tighter-than-average financial conditions. Current
GDP growth is included as a control for contemporaneous macroeconomic
conditions.
The dates index the quarter \(t\) of the regressors. The \(h=1\) sample
spans 1971:Q1--2025:Q3 with \(n=219\) observations; the \(h=4\) sample spans
1971:Q1--2025:Q1 with \(n=217\) observations.
\subsubsection{Model and Hypotheses}
Let \(y_t\) denote annualized GDP growth at quarter \(t\), and define the
\(h\)-quarter-ahead average growth rate as
\[
Y_{t+h}=\frac1h\sum_{j=1}^{h}y_{t+j}.
\]
For each horizon $h\in\{1,4\}$, we consider the conditional quantile specification
\begin{equation}\label{eq:gar_model}
Q_{Y_{t+h}}(\alpha \mid \mathcal{I}_t)
=
\beta_0(\alpha)
+
\beta_1(\alpha)\,\mathrm{NFCI}_t
+
\beta_2(\alpha)\,y_t,
\end{equation}
where $\beta_1(\alpha)$ is the coefficient of primary interest. It measures the
association between current financial conditions and the $\alpha$-quantile of
the corresponding growth outcome.
We consider two null hypotheses within the maintained linear conditional-quantile
model. The first imposes a uniformly zero NFCI coefficient:
\begin{equation}\label{eq:H0_zero}
H_0^{(1)}:\;\beta_1(\alpha)=0,\qquad \forall\, \alpha\in\mathcal{T}.
\end{equation}
The second asks whether the effect of financial conditions is quantile-invariant over the range under study:
\begin{equation}\label{eq:H0_const}
H_0^{(2)}:\;\beta_1(\alpha)=c,\qquad \forall\, \alpha\in\mathcal{T},
\end{equation}
where $c$ is a common constant. Inverting $H_0^{(2)}$ over candidate values
$c$ yields a grid-based confidence interval for the common effect.
We also report a fixed benchmark equal to the estimated OLS slope, where the
OLS slope is the coefficient from the linear projection of the growth outcome on
NFCI and the controls. This benchmark corresponds to a location-shift
approximation in which financial conditions move the conditional distribution
by a common amount across quantiles. The estimated OLS coefficients are
$\hat\beta_1^{\mathrm{OLS}}=-1.255$ for $h=1$ and
$\hat\beta_1^{\mathrm{OLS}}=-0.658$ for $h=4$. Tests using these values are
descriptive fixed-benchmark comparisons; common-constant confidence intervals
are obtained by inversion over fixed candidate values of $c$. Treating the
OLS benchmark as a same-sample random estimate would require adding its
influence function to the plug-in correction. In Table~\ref{tab:gar_results},
$c_{\mathrm{OLS}}$ denotes the reported numerical value treated in this fixed
manner.
\subsubsection{Implementation}
The empirical implementation uses discrete quadrature measures over two quantile
ranges. The main aggregation range covers the central part of the conditional
distribution,
\[
\mathcal{T}_{\mathrm{main}}
=
\{\alpha_1,\ldots,\alpha_{100}\}
\subset [0.10,0.90],
\]
with 100 equally spaced quantile levels and equal weights. Equivalently, the
implemented statistic uses
\[
\Pi_{\mathrm{main}}
=
\frac{1}{100}\sum_{j=1}^{100}\delta_{\alpha_j}.
\]
This range is chosen to focus on the economically relevant central quantiles
while avoiding the most weakly estimated extremes. To study downside risk more
directly, we also consider a left-tail aggregation range
\[
\mathcal{T}_{\mathrm{left}}
=
\{\alpha_1,\ldots,\alpha_{25}\}
\subset [0.05,0.30],
\qquad
\Pi_{\mathrm{left}}
=
\frac{1}{25}\sum_{j=1}^{25}\delta_{\alpha_j}.
\]
A 200-point grid on \([0.05,0.95]\) is used to estimate the restricted
quantile curve. For implementation, the fitted curve is evaluated at the test
and auxiliary indices by linear interpolation, with endpoint extrapolation when
required. The density routine uses \(J=1{,}000\) auxiliary draws from
\([0.01,0.99]\) and a Gaussian kernel. The aggregation grids lie strictly
inside this auxiliary interval.
The exponential weighting function is
\[
w_{t,n}(\gamma)=\exp(\gamma^\top \tilde{I}_t)-\bar{e}_n(\gamma),
\]
where $\tilde{I}_t=(\widetilde{\mathrm{NFCI}}_t, \tilde{y}_t)^\top$. Each
component is demeaned and divided by its within-sample standard deviation.
The weighting-direction space is $\Gamma=[-4,4]^2$, discretized on a
$40\times40$ grid with 1{,}600 candidate values of $\gamma$.
Under both null hypotheses, the nuisance parameters
$(\beta_0(\alpha),\beta_2(\alpha))$ are estimated by restricted quantile
regression. The multiplier bootstrap uses the influence-function correction in
Section~\ref{sec:pre_estimation}. The hypothesized value of $\beta_1$, either
$0$ under $H_0^{(1)}$ or a fixed constant $c$ under $H_0^{(2)}$, is held fixed
under the null and requires no additional correction. The row with
$c=\hat\beta_1^{\mathrm{OLS}}$ is the fixed-benchmark comparison described above.
We use $R = 2{,}000$ bootstrap multiplier sequences (reused across all candidate
$\lambda$), and approximate the theoretical interval $\Lambda=[0,2]$ by the grid
$\Lambda_G = \{0, 0.1, 0.2, \ldots, 2.0\}$. The reference class is the singleton
$\mathcal{B} = \{B^*\}$ with $B^*(\alpha) \equiv 0.50$ for adaptive penalty selection.
The GaR application has about 217 observations. The linear AR(2) simulation
gives reasonably accurate size at $n=200$, but it provides only limited guidance
for this application. The nonlinear results in the Supplementary Appendix show that
plug-in distortions can be larger in small samples. We therefore interpret
$p$-values near conventional thresholds cautiously.
\subsection{Uniform Test Results}\label{ssec:test_results}
Table~\ref{tab:gar_results} reports the test results under both quantile
aggregation measures for each horizon.
\begin{table}[ht]
\centering
\caption{Uniform Tests of the NFCI Coefficient in the GaR Model}\label{tab:gar_results}
\resizebox{\textwidth}{!}{
\begin{tabular}{lll ccccc ccccc}
\toprule
& & & \multicolumn{5}{c}{Adaptive Penalized Test} & \multicolumn{5}{c}{Unpenalized Test ($\lambda=0$)} \\
\cmidrule(lr){4-8} \cmidrule(lr){9-13}
Grid & $h$ & Null
& $\hat{T}_n$ & $\hat{c}_n^*$ & Margin & $p$ & $\hat\lambda$
& $\hat{T}_n$ & $\hat{c}_n^*$ & Margin & $p$ & $\lambda$ \\
\midrule
\multicolumn{13}{l}{\textit{Panel A: Main aggregation range $[0.10,0.90]$}} \\[2pt]
$[0.10,0.90]$ & $h=1$ & Fixed benchmark: $\beta_1=c_{\mathrm{OLS}}$
& 3.978 & 4.245 & $-$0.267 & 0.124 & 0.6
& 4.738 & 5.163 & $-$0.424 & 0.130 & 0 \\
$[0.10,0.90]$ & $h=1$ & $H_0^{(1)}:\beta_1=0$
& 8.008 & 4.904 & 3.104 & 0.025 & 0.5
& 8.213 & 5.803 & 2.410 & 0.034 & 0 \\[3pt]
$[0.10,0.90]$ & $h=4$ & Fixed benchmark: $\beta_1=c_{\mathrm{OLS}}$
& 2.404 & 3.860 & $-$1.456 & 0.261 & 0.9
& 3.753 & 4.956 & $-$1.204 & 0.210 & 0 \\
$[0.10,0.90]$ & $h=4$ & $H_0^{(1)}:\beta_1=0$
& 10.118 & 3.640 & 6.479 & 0.002 & 1.9
& 11.001 & 5.703 & 5.298 & 0.007 & 0 \\[3pt]
\midrule
\multicolumn{13}{l}{\textit{Panel B: Left-tail aggregation range $[0.05,0.30]$}} \\[2pt]
$[0.05,0.30]$ & $h=1$ & Fixed benchmark: $\beta_1=c_{\mathrm{OLS}}$
& 3.962 & 6.244 & $-$2.283 & 0.313 & 0.0
& 3.962 & 6.244 & $-$2.283 & 0.313 & 0 \\
$[0.05,0.30]$ & $h=4$ & Fixed benchmark: $\beta_1=c_{\mathrm{OLS}}$
& 6.666 & 4.826 & 1.840 & 0.032 & 0.2
& 7.521 & 5.282 & 2.239 & 0.024 & 0 \\
\bottomrule
\multicolumn{13}{l}{\footnotesize Notes: Bootstrap replications $R=2{,}000$. Nominal level $\tau=0.10$. Margin $=\hat{T}_n-\hat{c}_n^*$. $p$-values are bootstrap $p$-values.}\\
\multicolumn{13}{l}{\footnotesize Here $c_{\mathrm{OLS}}$ is the reported OLS estimate treated as a fixed benchmark; its same-sample estimation uncertainty is not included.}\\
\end{tabular}
}
\end{table}
The zero-effect null is rejected under the main quantile aggregation measure at
both reported horizons. For $h=1$, the adaptive and unpenalized $p$-values are
$0.025$ and $0.034$, respectively; for $h=4$, they are $0.002$ and $0.007$.
These are rejections of the zero-coefficient restriction within the maintained
linear GaR model using the reported bootstrap procedure.
Under the main quantile aggregation measure, the fixed OLS benchmark is not
rejected at either horizon: the adaptive $p$-values are $0.124$ for $h=1$ and
$0.261$ for $h=4$. The common-$c$ confidence interval reported below is also
nonempty at both horizons. Under the implemented nominal 10\% calibration, the composite
null in equation~\eqref{eq:H0_const} is therefore not rejected over
$[0.10,0.90]$. This conclusion does not establish that the coefficient is
constant; it states only that the reported procedure retains at least one
common value.
The left-tail rows have a narrower interpretation because no common-$c$
inversion is reported for $\mathcal{T}_{\mathrm{left}}$. They test only the
specific fixed OLS benchmark. That benchmark is not rejected for $h=1$, with
both $p$-values equal to $0.313$, and is rejected for $h=4$, with adaptive and
unpenalized $p$-values of $0.032$ and $0.024$. These results alone do not reject
the composite null that some other common constant fits the left tail.
The results are consistent with the qualitative left-tail pattern in
\citet{adrian2019vulnerable}, but the numerical conclusions are more limited. Over
the central range, the zero-effect null is rejected, whereas the existence of a
common negative effect is not rejected. Over the left tail, the $h=4$ test
rejects the particular OLS benchmark, not every possible constant effect.
\subsection{Common-Constant Confidence Intervals}\label{ssec:CI}
To complement the hypothesis tests, we construct a grid-based 90\% confidence
interval for the common constant by inverting $H_0^{(2)}$ over
$c\in\mathcal G_c:=\{-2,-1.95,\ldots,2\}$. Specifically, the reported set is
\[
\mathcal C_G=\{c\in\mathcal G_c:H_0(c)\text{ is not rejected}\},
\]
using the nominal 0.10 rule, where
\[
H_0(c):\beta_1(\alpha)=c\ \text{for all }\alpha\in\mathcal T_{\mathrm{main}}.
\]
Table~\ref{tab:gar_ci} reports the minimum and maximum retained grid values.
The accepted grid values are contiguous in the reported implementation, so
their endpoints define the displayed grid-based confidence intervals. These
intervals concern the common constant $c$ under $H_0(c)$; they are not
pointwise confidence intervals for the unrestricted coefficient path.
\begin{table}[ht]
\centering
\caption{Grid-Based 90\% Confidence Intervals for the Common Constant $c$}\label{tab:gar_ci}
\begin{tabular}{l cc cc}
\toprule
& \multicolumn{2}{c}{Adaptive} & \multicolumn{2}{c}{Unpenalized} \\
\cmidrule(lr){2-3} \cmidrule(lr){4-5}
Horizon & Confidence interval & Width & Confidence interval & Width \\
\midrule
$h=1$ & $[-1.25,\,-0.35]$ & 0.90 & $[-1.25,\,-0.25]$ & 1.00 \\
$h=4$ & $[-0.75,\,-0.40]$ & 0.35 & $[-0.70,\,-0.35]$ & 0.35 \\
\bottomrule
\multicolumn{5}{l}{\footnotesize Notes: Grid inversion of nominal 10\% tests of}\\
\multicolumn{5}{l}{\footnotesize $H_0:\beta_1(\alpha)=c$ for all $\alpha\in\mathcal{T}_{\mathrm{main}}$. Grid: $c\in[-2,2]$, step $0.05$.}
\end{tabular}
\end{table}
The confidence intervals are nonempty and contain only negative common
constants. They concern $c$ under the maintained constant-coefficient
restriction and do not imply pointwise significance at every quantile.
At the one-quarter horizon, the adaptive confidence interval
$[-1.25,-0.35]$ is slightly narrower than its unpenalized counterpart
$[-1.25,-0.25]$. The OLS value $-1.255$ differs slightly from the displayed
grid endpoint $-1.25$ because the inversion uses a $0.05$ grid, whereas the OLS
value is tested separately. At $h=4$, both procedures deliver intervals of the
same width, although the adaptive interval $[-0.75,-0.40]$ is shifted somewhat
toward more negative values relative to the unpenalized interval
$[-0.70,-0.35]$.
\subsection{Visualizing the Quantile Coefficient Paths}
Figures~\ref{fig:beta_path_h1} and~\ref{fig:beta_path_h4} display the
unrestricted quantile regression estimates $\hat\beta_1(\alpha)$ together with
the common-$c$ confidence intervals for $h=1$ and $h=4$, respectively. The
horizontal shaded bands are confidence intervals for the common constant $c$;
they are not pointwise confidence bands for $\beta_1(\alpha)$.
Descriptively, the estimated coefficient path $\hat\beta_1(\alpha)$ is strongly
negative at low quantiles and moves closer to zero at upper quantiles at both
horizons. The $h=1$ path changes relatively sharply around
$\alpha\approx0.30$, whereas the $h=4$ path is smoother but still varies in the
lower quantiles.
The common-$c$ confidence intervals do not provide pointwise uncertainty bands
around $\hat\beta_1(\alpha)$. A visually varying coefficient path can coexist
with a nonempty confidence interval for a common constant. Whether a plotted
point lies inside or outside the horizontal interval has no pointwise coverage
interpretation.
\begin{figure}[htbp]
\centering
\includegraphics[width=0.78\textwidth]{Figure_3.pdf}
\caption{NFCI coefficient path $\hat\beta_1(\alpha)$ and common-$c$ confidence interval,
$h=1$. The red curve plots unrestricted quantile regression estimates on
$\alpha\in[0.10,0.90]$. Horizontal shaded bands show the grid-based 90\%
confidence interval for the common constant;
they are not pointwise confidence bands. The
dashed green line marks the OLS estimate $\hat\beta_1^{\mathrm{OLS}}=-1.255$. The light shaded region highlights the left-tail range $\alpha\le0.30$, which is
the region of primary interest in the downside-risk analysis.}
\label{fig:beta_path_h1}
\end{figure}
\begin{figure}[htbp]
\centering
\includegraphics[width=0.78\textwidth]{Figure_4.pdf}
\caption{NFCI coefficient path $\hat\beta_1(\alpha)$ and common-$c$
confidence interval for the average growth outcome over the next four
quarters ($h=4$). The blue curve plots unrestricted quantile regression
estimates on $\alpha\in[0.10,0.90]$. Horizontal shaded bands show the
grid-based 90\% confidence interval for the common constant; it is not a
pointwise confidence band. The dashed green line marks the OLS estimate
$\hat\beta_1^{\mathrm{OLS}}=-0.658$.}
\label{fig:beta_path_h4}
\end{figure}
\subsection{Discussion}
The results distinguish evidence that financial conditions matter from
evidence that their association with growth varies across quantiles. Over
$[0.10,0.90]$, the zero-effect restriction is rejected, while the common-$c$
confidence intervals are nonempty. The left-tail rows answer a narrower question:
they test the fixed OLS benchmark and do not rule out other common constants.
The adaptive and unpenalized procedures lead to the same qualitative
conclusions in this application. Their numerical differences are mixed: the
adaptive test has smaller $p$-values in some rows and larger $p$-values in
others, while its common-$c$ confidence interval is narrower for $h=1$ and has
the same width for $h=4$. These results do not indicate a uniform finite-sample
improvement from penalization.
\section{Conclusion}
\label{sec:conclusion}
This paper develops a specification test for conditional quantile models based
on a penalized generalized Bierens maximum statistic. Exponential weights
transform the conditional restriction into a continuum of unconditional
moments. We then use an adaptive $\ell_1$ penalty to target local power. For
pre-estimated nuisance parameters, we derive the first-order plug-in term, a
corrected variance estimator, and a multiplier bootstrap that is valid under
the null and local alternatives. For known parameters, the adaptive selector
has no lower maximin local power than the unpenalized test under the stated
uniformity conditions.
The simulations show that finite-sample performance depends on the aggregation
order and on the complexity of the pre-estimation problem. In the linear AR(2)
design, the proposed CvM--KS statistic is relatively close to nominal size
in most of the reported comparisons. In the nonlinear design reported in the Supplementary Appendix, size distortion remains substantial at small sample sizes but
decreases as $n$ increases. The power comparisons are also design-specific: the
adaptive test need not outperform the unpenalized test for every parameter or
hypothesis.
The Growth-at-Risk application illustrates why zero-effect tests,
fixed-benchmark comparisons, and common-$c$ confidence intervals must be
interpreted separately. The zero-effect restriction is rejected at both
horizons, while the central-range common-$c$ confidence intervals are nonempty
and contain only negative values. In the left tail, the four-quarter result rejects the fixed
OLS benchmark; it does not address every possible common constant.
Allowing the dimension of the conditioning vector to grow with the sample size
would extend the method to richer information sets and turn the $\ell_1$
penalty into a direction-selection device. Oracle-type results for the
penalized supremum over $\gamma$ could then connect this framework to
high-dimensional specification testing. A sharper characterization of the
maximin selector, particularly in the Gaussian location-shift model where the
studentized drift has a closed form, would supplement the no-loss guarantee
with verifiable conditions for strict power gains; see
Remark~\ref{rem:strict-gain_main}.
On the numerical side, convolution-smoothed quantile regression provides a
smooth alternative for the nuisance-estimation step \citep{he2023smoothed}.
Its implications for the fitted-density plug-in correction, as well as those
of orthogonalizing the score with respect to the nuisance component, merit
separate finite-sample analysis. In applications such as
Growth-at-Risk, weighting $\Pi$ toward the tails or taking a
multiplicity-adjusted maximum over several aggregation ranges would focus the
analysis more directly on downside risk.
\clearpage
\phantomsection
\addcontentsline{toc}{section}{Appendix: Supplementary Material}