EconBase
← Back to paper

A Projection Framework for Testing Shape Restrictions That Form Convex Cones

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

91,509 characters

A Projection Framework for Testing Shape Restrictions That Form Convex Cones


\begin{bibunit}
\pdfbookmark[1]{Title}{title}
\title{A Projection Framework for Testing Shape Restrictions That Form Convex Cones\thanks{We thank three anonymous referees for valuable suggestions that have helped greatly improve this paper. We are indebted to Andres Santos for extensive discussions and constant support. We thank Xiaohong Chen, Denis Chetverikov, Stefan Hoderlein, Tatiana Komarova, Scott Kostyshak, Simon Lee, Sungwon Lee, Ulrich K. M\"{u}ller, Pepe Montiel Olea, Alex Poirier, Brandon Reeves, Kevin Song, Yanglei Song, Xun Tang and numerous seminar participants for their useful comments. We also thank Denis Chetverikov and Simon Lee for sharing their code. Advanced computing resources and consultation provided by Texas A\&M High Performance Research Computing and Marinus Pennings are gratefully acknowledged. A previous version of this paper was circulated under the title ``A General Framework for Inference on Shape Restrictions.''}
}
\author{
Zheng Fang \\ Department of Economics \\ Texas A\&M University\\ [email removed]
\and
Juwon Seo\\ Department of Economics\\National University of Singapore\\ [email removed]}
\date{First draft: October 16, 2019\\ This draft: May 6, 2021}
\maketitle


\begin{abstract}
This paper develops a uniformly valid and asymptotically nonconservative test based on projection for a class of shape restrictions. The key insight we exploit is that these restrictions form convex cones, a simple and yet elegant structure that has been barely harnessed in the literature. Based on a monotonicity property afforded by such a geometric structure, we construct a bootstrap procedure that, unlike many studies in nonstandard settings, dispenses with estimation of local parameter spaces, and the critical values are obtained in a way as simple as computing the test statistic. Moreover, by appealing to strong approximations, our framework accommodates nonparametric regression models as well as distributional/density-related and structural settings. Since the test entails a tuning parameter (due to the nonstandard nature of the problem), we propose a data-driven choice and prove its validity. Monte Carlo simulations confirm that our test works well.
\end{abstract}



\begin{center}
\textsc{Keywords:} Nonstandard inference, shape restrictions, convex cone, projection, strong approximations
\end{center}




\newpage

\section{Introduction}

Shape restrictions play a number of fundamental and prominent roles in economics. For example, they often arise as testable implications of economic theory, and may thus serve as plausible restrictions in specifying economic models \citep{Varian1982Demand,Varian1984Production}. In complementary empirical work, they help achieve point identification, tighten identification bounds, improve estimation precision, and develop powerful tests---see \citet{Matzkin1994Handbook} and \citet{ChetverikovSantosAzeem2018Shape} for more detailed discussions.

In this paper, we develop a uniformly valid and asymptotically nonconservative test for a class of shape restrictions. To illustrate, consider the regression model:
\begin{align}\label{Eqn: NP, intro}
Y = \theta_0(Z)+u~,
\end{align}
where $Z\in [0,1]$, $\theta_0:[0,1]\to\mathbf R$, and $E[u|Z]=0$. Suppose that we are interested in testing whether $\theta_0$ is nondecreasing. More formally, if $\theta_0\in \mathbf H\equiv \{\theta:[0,1]\to\mathbf R: \|\theta\|_{\mathbf H}\equiv \{\int_{[0,1]} |\theta(z)|^2\,\mathrm dz\}^{1/2}<\infty\}$ (a mild restriction) and $\Lambda$ is the class of nondecreasing functions in $\mathbf H$, then we may formulate the hypotheses as
\begin{align}\label{Eqn: hypotheses, projection}
\mathrm H_0: \phi(\theta_0)=0 \qquad\qquad \mathrm{vs.}\qquad\qquad \mathrm H_1: \phi(\theta_0)>0~,
\end{align}
where $\phi(\theta)\equiv\min_{\lambda\in\Lambda}\|\theta-\lambda\|_{\mathbf H}$ is the distance from $\theta$ to $\Lambda$. Thus, given an unconstrained (kernel or sieve) estimator $\hat\theta_n$ of $\theta_0$, we may employ a test that rejects $\mathrm H_0$ if $r_n\phi(\hat\theta_n)$ is ``large'' for a suitable $r_n\to\infty$. While conceptually intuitive, construction of critical values turns out to be a delicate and challenging matter. In particular, despite the well-established results on the rate $r_n$ and pointwise asymptotic normality, $r_n\{\hat\theta_n-\theta_0\}$ generically does not converge as a process indexed by $[0,1]$ \citep{ChernozhukovLeeRosen2013Intersection}, rendering the Delta method as in \citet{FangSantos2018HDD} inapplicable.

As a first step, we note that $\Lambda$ being a (closed) convex cone\footnote{By definition, $\Lambda$ is a convex cone if and only if $af+bg\in\Lambda$ whenever $a,b\ge 0$ and $f,g\in\Lambda$.} implies
\begin{align}\label{Eqn: test statistic}
r_n\phi(\hat\theta_n)=\|r_n\{\hat\theta_n-\theta_0\}+r_n\theta_0-\Pi_\Lambda(r_n\{\hat\theta_n-\theta_0\}+r_n\theta_0)\|_{\mathbf H}~,
\end{align}
where $\theta\mapsto\Pi_{\Lambda}(\theta)\equiv\operatornamewithlimits{arg\,min}_{\lambda\in\Lambda}\|\theta-\lambda\|_{\mathbf H}$ is the projection operator, i.e., $\Pi_{\Lambda}(\theta)$ is the closest (under $\|\cdot\|_{\mathbf H}$) nondecreasing function to $\theta$. Hence, \eqref{Eqn: test statistic} reveals that, in estimating the law of $r_n\phi(\hat\theta_n)$ in order to obtain critical values, it suffices to quantify the variation of $r_n\{\hat\theta_n-\theta_0\}$ and estimate the drift $r_n\theta_0$. Despite the lack of convergence (in $\mathbf H$), $r_n\{\hat\theta_n-\theta_0\}$ as a process may be approximated in law via strong approximations by a number of methods with only mild computation cost, including the simulation method in \citet{ChernozhukovLeeRosen2013Intersection}, the weighted bootstrap in \citet{BelloniChernozhukovChetverikovKato2015New}, and the sieve score bootstrap in \citet{ChenChristensen2018SupNormOptimal}. The treatment of $r_n\theta_0$, on the other hand, consists of the nonstandard step because, as in moment inequality models \citep{AndrewsSoares2010},  $r_n\theta_0$ cannot be consistently estimated in general. In this regard, the convex cone property (but not convexity alone) implies
\begin{align}\label{Eqn: stochastic domiance, intro}
r_n\phi(\hat\theta_n)\le \|r_n\{\hat\theta_n-\theta_0\}+\kappa_n\theta_0-\Pi_\Lambda(r_n\{\hat\theta_n-\theta_0\}+\kappa_n\theta_0)\|_{\mathbf H}~,
\end{align}
whenever $0\le \kappa_n\le r_n$. Hence, a valid critical value may be obtained by bootstrapping the upper bound in \eqref{Eqn: stochastic domiance, intro}, which is possible because $\kappa_n\theta_0$ may be consistently estimated by $\kappa_n\hat\theta_n$ provided $\kappa_n/r_n\to 0$. The very virtue that \eqref{Eqn: stochastic domiance, intro} is an inequality rather than equality may also raise concern for conservativeness of the resulting test. As an extreme case, the choice $\kappa_n=0$ leads to a least favorable test that may be viewed as assuming $\theta_0=0$, suggesting that $\kappa_n$ should not be too small. Along these lines, we develop a data-driven choice of $\kappa_n$ that delivers an asymptotically nonconservative test.




While we started with monotonicity in the univariate model \eqref{Eqn: NP, intro}, the main features of our test are not confined to this special problem. First, the convex cone property is in fact shared by a class of restrictions, e.g., nonnegativity, concavity, Slutsky restriction and supermodularity---see Appendix \ref{App: convex cone}. Regrettably, our framework is not directly applicable absent the property, e.g., quasi-concavity, log-concavity, and $r$-concavity \citep{Kostyshak2017U,KomarovaHidalgo2019Shape}. Second, our framework is applicable to distributional and structural settings \citep{PinkseSchurter2019Auction,Bhattacharya2020Empirical}---see Appendix \ref{App: convex cone}. Third, our framework allows for jointly testing multiple restrictions since intersections of convex cones remain convex cones. This is important because it is common that shape restrictions arise simultaneously. Fourth, while some convex cone restrictions (e.g., monotonicity and concavity) on $\theta_0$ can be characterized as inequalities on derivatives of $\theta_0$, we dispense with derivative estimation because it likely incurs power loss due to a slower convergence rate \citep{Stone1982Global,ChenChristensen2018SupNormOptimal}. Finally, our test does not rely on the least favorable configurations, while being asymptotically nonconservative (thus leading to improved power) and computationally tractable. We stress that, since the norm $\|\cdot\|_{\mathbf H}$ is of $L^2$ nature, our test cannot be inverted to obtain uniform confidence bands. In this sense, our paper complements \citet{HorowitzLee2017Shape}, \citet{FreybergerReeves2017Shape} and \citet{ChenChernozhukovFernandezKostyshakLuo2021Shape}.





The literature on shape restrictions was initiated in the 1950s \citep{BBBB1972}, with much attention since focused on estimation under solely shape restrictions---see, e.g., \citet{HanWangChatterjeeSamworth2019Isotonic} and references therein. There are also post-processing methods that enforce restrictions on unconstrained estimators---see \citet{ChenChernozhukovFernandezKostyshakLuo2021Shape} who studied a number of shape enforcing operators. The projection method in this paper is of post-processing nature. The convex cone structure was recognized in the 1960s as a device to generalize monotone regression, though the focus is on {\it analytic} properties of projections \citep{BBBB1972}. For testing, the structure has barely been exploited beyond identifying the least favorable distributions in parametric settings \citep{Wolak1987Exact,SilvapulleSen2005Constrained}, deriving minimax bounds in univariate Gaussian white noise models \citep{JuditskyNemirovski2002Shape}, and establishing minimaxity for the likelihood ratio test in Gaussian sequence models \citep{WeiWainwrightGuntuboyina2019Geometry}.


Much of the testing literature has been developed by exploiting {\it particular} structures of restrictions, with concavity and especially monotonicity in univariate models being the primary focuses. Despite a sizable literature, existing tests prevalently rely on the least favorable configurations, including \citet{GijbelsHallJonesKoch2000Monotonicity}, \citet{GhosalSenVaart2000Monotonicity}, \citet{HallHeckman2000Monotonicity}, \citet{Durot2003testing} and \citet{Gutknecht2016MonotoneTest} for monotonicity, and \citet{AbrevayaJiang2005Curvature} for concavity---see also
\citet{DumbgenSpokoiny2001Multiscale} and \citet{BaraudHuetLaurent2005Convex} who devised tests for both shapes. \citet{Chetverikov2018Monotonicity} developed, as far as we are aware, the first nonparametric uniformly valid tests designed specifically for monotonicity that avoid the least favorable configurations.


There is also a strand of literature motivated by the {\it common} structure of shape restrictions viewed as inequalities, including \citet{ChernozhukovNeweySantos2019CCMM}, \citet{LeeSongWhang2017Inequal},  \citet{BelloniChernozhukovChetverikovFernandez2019QR} and \citet{Zhu2020Shape}. If the inequalities are based on derivatives, the derivative estimation may then incur power loss. Moreover, all these (uniformly valid) tests, except the least favorable test of \citet{BelloniChernozhukovChetverikovFernandez2019QR}, involve estimation of local parameter spaces. \citet{ChernozhukovNeweySantos2019CCMM} and \citet{Zhu2020Shape}, unlike our setup, allowed for partial identification by working with moment restrictions. By the virtue of their setup, they required estimating the set of minimizers and the use of strong approximations may entail stronger regularity conditions. Similar in spirit to \citet{ChernozhukovNeweySantos2019CCMM} and \citet{Zhu2020Shape}, \citet{KomarovaHidalgo2019Shape} proposed a moment-based test in the univariate model \eqref{Eqn: NP, intro} for shape restrictions that may not form convex cones.


The remainder of the paper is structured as follows. Section \ref{Sec: setup} introduces the setup and some motivating examples. Section \ref{Sec: framework} presents our inferential framework, a data-driven choice of the tuning parameter, and an implementation guide. Section \ref{Sec: simulations} conducts simulation studies, while Section \ref{Sec: conclusion} concludes. Appendix \ref{Sec: main results proof} contains proofs of the main results. Due to space limitation, we relegate the discussions of the convex cone property, an investigation of the special case when $\hat\theta_n$ admits an asymptotic distribution, and presentations of some auxiliary results to the online supplement.





\section{The Setup and Examples}\label{Sec: setup}

Throughout, we denote by $\{X_i\}_{i=1}^n$ the sample with each $X_i$ living in some sample space $\mathcal X$, and by $P$ the joint law of $\{X_i\}_{i=1}^n$ that belongs to some family $\mathbf P$ of distributions on $\mathcal X^n$. The dependence of $P$ and $\mathbf P$ on $n$ is suppressed for notational simplicity. We stress that, under this configuration, $\{X_i\}_{i=1}^n$ need not be i.i.d. In turn, we let $\theta_0$ be the parameter of interest, and, whenever appropriate, make the dependence of $\theta_0$ on $P$ explicit by instead writing $\theta_P$. In order to accommodate Slutsky restriction, we shall work with an abstract Hilbert space (i.e., a complete inner product space) with inner product $\langle\cdot,\cdot\rangle_{\mathbf H}$ and induced norm $\|\cdot\|_{\mathbf H}$. Given the sample $\{X_i\}_{i=1}^n$, our objective is then to test whether $\theta_0$ satisfies the shape in question, i.e., the hypotheses in \eqref{Eqn: hypotheses, projection}. The setup \eqref{Eqn: hypotheses, projection} induces two models: $\mathbf P_0\equiv\{P\in\mathbf P: \phi(\theta_P)=0\}$ and $\mathbf P_1\equiv\mathbf P\backslash\mathbf P_0$.

Before proceeding further, we introduce additional notation. Set $\mathbf R_+\equiv\{x\in\mathbf R: x\ge 0\}$ and $L^2(\mathcal Z)\equiv\{f: \mathcal Z\to\mathbf R: \int_{\mathcal Z} |f(z)|^2\,\mathrm dz<\infty\}$ for $\mathcal Z\subset\mathbf R^{d_z}$. Let $\mathbf M^{m\times k}$ be the space of $m\times k$ matrices, and, for $A\in\mathbf M^{m\times k}$, write its transpose by $A^\text{\scalebox{0.8}{$\intercal$}}$, its trace by $\mathrm{tr}(A)$ if $m=k$, its Moore–Penrose inverse by $A^-$, and its Frobenius norm by $\|A\|\equiv\{\mathrm{tr}(A^\text{\scalebox{0.8}{$\intercal$}} A)\}^{1/2}$. For a sequence $\{h_j\}$ of functions, denote the vector $(h_1,\ldots,h_k)^\text{\scalebox{0.8}{$\intercal$}}$ by $h^k$. For generic families of distributions $\mathbf P_n$, a sequence $\{a_n\}$ of positive scalars, and a sequence $\{\mathbb X_n\}$ of random elements in a normed space $\mathbf D$ with norm $\|\cdot\|_{\mathbf D}$, write $\mathbb X_n=o_p(a_n)$ uniformly in $P\in\mathbf P_n$ if $\lim_{n\to\infty}\sup_{P\in\mathbf P_n}P(\|\mathbb X_n\|_{\mathbf D}>a_n\epsilon)= 0$ for any $\epsilon>0$, and $\mathbb X_n=O_p(a_n)$ uniformly in $P\in\mathbf P_n$ if $\lim_{M\to\infty}\limsup_{n\to\infty}\sup_{P\in\mathbf P_n}P(\|\mathbb X_n\|_{\mathbf D}>Ma_n)= 0$.


\subsection{Examples}\label{Sec: examples}

We now present examples where shape restrictions play important roles. The first example is concerned with nonparametric regression models.

\begin{ex}[Nonparametric Regression]\label{Ex: NP under shape}
Let $X\equiv(Y,Z)\in\mathbf R^{1+d_z}$ satisfy
\begin{align}\label{Eqn: Ex, NP under shape}
Y=\theta_0(Z)+u~,
\end{align}
where $\theta_0: \mathcal Z\subset\mathbf R^{d_z}\to\mathbf R$ with $\mathcal Z$ the support of $Z$, and $E[u|Z]=0$. Here, one may set $\mathbf H=L^2(\mathcal Z)$, let $\Lambda\subset \mathbf H$ consist of (say) monotonic functions, and obtain an unconstrained estimator $\hat\theta_n$ of $\theta_0$ by kernel methods such as local constant/linear/polynomial regression or sieve methods with various basis functions such as splines \citep{ChernozhukovLeeRosen2013Intersection}. The rate $r_n$ equals $(nh_n^{d_z})^{1/2}$ for kernel estimation with bandwidth $h_n$, and $(n/k_n)^{1/2}$ for sieve estimation based on, e.g., B-splines, with $k_n$ (henceforth) the sieve dimension. One may also set up this example as one on the conditional mean $z\mapsto E[Y|Z=z]$. \qed
\end{ex}

Our second example generalizes \eqref{Eqn: Ex, NP under shape} to its instrumental variable (IV) analog.

\begin{ex}[Nonparametric IV Regression]\label{Ex: NPIV under shape}
Let $X\equiv(Y,Z,V)\in\mathbf R^{1+d_z+d_v}$ satisfy \eqref{Eqn: Ex, NP under shape} but with $E[u|V]=0$. Then one may set $\mathbf H$ and $\Lambda$ as in Example \ref{Ex: NP under shape}, and employ the series two-stage least square (2SLS) estimation with, e.g., B-splines, resulting in a rate $r_n$ equal to $(n/k_n)^{1/2} s_n$, where $s_n$ is the smallest singular value of $E[b^{m_n}(V)h^{k_n}(Z)^\text{\scalebox{0.8}{$\intercal$}}]$ with $h^{k_n}$ and $b^{m_n}$ respectively $k_n\times 1$ and $m_n\times 1$ vectors (with $m_n\ge k_n$) of B-splines for $\theta_0$ and the instrument space \citep{NeweyPowell2003NPIV,BlundellChenKristensen2007Engel,ChenChristensen2018SupNormOptimal}. In practice, $s_n$ is unknown but can be replaced with its sample analog. One may conceivably also employ kernel-type estimators \citep{HallHorowitz2005NPIV,DarollesFanFlorensRenault2011NPIV}, though the strong approximation results appear to be lacking.  \qed
\end{ex}

The third example is concerned with nonparametric quantile regression.

\begin{ex}[Nonparametric Quantile Regression]\label{Ex: NPQR under shape}
Let $X\equiv(Y,Z)\in\mathbf R^{1+d_z}$ satisfy
\begin{align}\label{Eqn: Ex, NPQR under shape}
Y=\theta_0(Z,U)~,
\end{align}
where $\theta_0:\mathbf R^{d_z}\times[0,1]\to\mathbf R$, and $U$ is uniformly distributed on $[0,1]$. If $Z$ and $U$ are independent, then $\theta_0$ may be interpreted as the conditional quantile function of $Y$ given $Z$. Here, one may define $\mathbf H$ and $\Lambda$ as previously, and employ the recent sieve estimator of \citet{BelloniChernozhukovChetverikovFernandez2019QR} with $r_n=(n/k_n)^{1/2}$ for, e.g., B-splines.  \qed
\end{ex}

Our final example concerns a restriction in a possibly endogenous regression model.

\begin{ex}[Slutsky Restriction]\label{Ex: Slutsky}
Let $Q\in\mathbf R^{d_q}$ (quantities), $P\in\mathbf R^{d_q}$ (prices), $Y\in\mathbf R$ (income), and $Z\in\mathbf R^{d_z}$ (demographics)  satisfy the following $d_q$ equations:
\begin{align}
Q=g_0(P,Y)+\Gamma_0^\text{\scalebox{0.8}{$\intercal$}} Z + U~,
\end{align}
where $g_0: \mathbf R_+^{d_q+1}\to\mathbf R^{d_q}$ is differentiable, $\Gamma_0\in\mathbf M^{d_z\times d_q}$, and $U\in\mathbf R^{d_q}$ is the error term. The Slutsky matrix of $g_0$ is the mapping $\theta_0: \mathbf R_+^{d_q+1}\to\mathbf M^{d_q\times d_q}$ defined by:
\begin{align}\label{Eqn: Slutsky matrix in ex}
\theta_0(p,y)\equiv \frac{\partial g_0(p,y)}{\partial p^\text{\scalebox{0.8}{$\intercal$}}}+ \frac{\partial g_0(p,y)}{\partial y} g_0(p,y)^\text{\scalebox{0.8}{$\intercal$}}~.
\end{align}
The Slutsky restriction refers to $\theta_0(p,y)$ being negative semidefinite (nsd) at each pair $(p,y)$. For this example, we follow \citet{AguiarSerrano2017Slutsky} by letting $\mathbf H$ be the space of functions $\theta: \mathbf R_+^{d_q+1}\to \mathbf M^{d_q\times d_q}$ with inner product $\langle \cdot,\cdot\rangle_{\mathbf H}$ defined by
\begin{align}\label{Eqn: Slutsky norm in ex}
\langle \theta_1,\theta_2\rangle_{\mathbf H}\equiv \int \mathrm{tr}(\theta_1(p,y)^\text{\scalebox{0.8}{$\intercal$}}\theta_2(p,y))\, \mathrm dp\mathrm dy~,\,\forall\,\theta_1,\theta_2\in\mathbf H~,
\end{align}
and $\Lambda$ the family of mappings $\theta\in\mathbf H$ such that $\theta(p,y)$ is nsd and symmetric for all $(p,y)$. In turn, the nonlinear functional $\theta_0$ may be estimated based on the plug-in principle and a sieve estimator of $g_0$ \citep{DonaldNewey1994Semilinear,BlundellChenKristensen2007Engel}, resulting in a rate $r_n$ as in Example \ref{Ex: NP under shape} (without endogeneity) or \ref{Ex: NPIV under shape} (with endogeneity). \qed
\end{ex}

We refer the reader to Appendix \ref{App: convex cone} that illustrates how our framework also applies to distributional settings where shape restrictions often take the form of various dominance relations and structural models where shape restrictions may serve as testable implications. For ease of reference, we shall call settings where $\hat\theta_n$ admits an asymptotic distribution (in $\mathbf H$) {\it regular} and ones without this property {\it irregular}.



\section{The Inferential Framework}\label{Sec: framework}

\subsection{The General Framework}\label{Sec: general theory}

We commence with our main assumptions in this paper.


\begin{ass}\label{Ass: convex setup}
(i) $\Lambda$ is a known nonempty closed convex set in a Hilbert space $\mathbf H$ with known inner product $\langle \cdot,\cdot\rangle_{\mathbf H}$ and induced norm $\|\cdot\|_{\mathbf H}$; (ii) $\Lambda$ is a cone.
\end{ass}

\begin{ass}\label{Ass: strong approx}
(i) An estimator $\hat\theta_n: \{X_i\}_{i=1}^n\to\mathbf H$ satisfies $\|r_n\{\hat\theta_n-\theta_P\}-\mathbb Z_{n,P}\|_{\mathbf H}=o_p(c_n)$ uniformly in $P\in\mathbf P$ for some $r_n\to\infty$, $\mathbb Z_{n,P}\in\mathbf H$, and $c_n>0$ with $c_n=O(1)$; (ii) $\hat{\mathbb G}_n\in\mathbf H$ is a bootstrap estimator satisfying: $\|\hat{\mathbb G}_{n}-\bar{\mathbb Z}_{n,P}\|_{\mathbf H}=o_p(c_n)$ uniformly in $P\in\mathbf P$, for $\bar{\mathbb Z}_{n,P}$ a copy of $\mathbb Z_{n,P}$ that is independent of $\{X_i\}_{i=1}^n$.
\end{ass}

Assumption \ref{Ass: convex setup} simply abstracts the convex cone feature, where we single out the conic condition for ease of elucidating the roles it plays in this paper. Assumption \ref{Ass: strong approx}(i) requires that $r_n\{\hat\theta_n-\theta_P\}$ be approximated by $\mathbb Z_{n,P}$ (uniformly) at a rate $c_n$. In regular settings, it suffices to have $\sqrt n\{\hat\theta_n-\theta_P\}\xrightarrow{L} \mathbb G_P$ uniformly in $P\in\mathbf P$, so that $r_n=\sqrt n$, and $\mathbb Z_{n,P}$ equals $\mathbb G_P$ in law---see Appendix \ref{Sec: special case}. In irregular settings such as Example \ref{Ex: NP under shape}, one may obtain $\hat\theta_n$ by kernel or sieve methods. To illustrate, let $\{Y_i,Z_i\}_{i=1}^n$ be a sample generated by \eqref{Eqn: Ex, NP under shape}, and $h^{k_n}$ be a $k_n\times 1$ vector of B-splines on $\mathcal Z$. Then the sieve estimator of $\theta_P$ is given by $\hat\theta_n=\hat\beta_n^{\text{\scalebox{0.8}{$\intercal$}}}h^{k_n}$ with $\hat\beta_n= [\sum_{i=1}^{n}h^{k_n}(Z_i)h^{k_n}(Z_i)^{\text{\scalebox{0.8}{$\intercal$}}}]^{-}\sum_{i=1}^{n}h^{k_n}(Z_i)Y_i$. Under regularity conditions, one may obtain the linear expansion in $L^2(\mathcal Z)$:
\begin{align}\label{Eqn: implementation, aux1}
r_n \{\hat\theta_n-\theta_P\}=k_n^{-1/2}(h^{k_n})^\text{\scalebox{0.8}{$\intercal$}} (E_P[h^{k_n}(Z)h^{k_n}(Z)^\text{\scalebox{0.8}{$\intercal$}}])^{-1}\frac{1}{\sqrt n}\sum_{i=1}^nh^{k_n}(Z_i)u_i+o_p(\frac{1}{\log n})~,
\end{align}
uniformly in $P\in\mathbf P$. Here, undersmoothing is required in order to deliver \eqref{Eqn: implementation, aux1}. Following \citet{ChernozhukovLeeRosen2013Intersection}, one may then verify Assumption \ref{Ass: strong approx}(i) with $r_n=\sqrt{n/k_n}$, $c_n=1/\log n$, and $\mathbb Z_{n,P}= k_n^{-1/2}(h^{k_n})^\text{\scalebox{0.8}{$\intercal$}} (E_P[h^{k_n}(Z)h^{k_n}(Z)^\text{\scalebox{0.8}{$\intercal$}}])^{-1} G_{n,P}$ for some $G_{n,P}\sim N(0,E_P[u^2h^{k_n}(Z)h^{k_n}(Z)^\text{\scalebox{0.8}{$\intercal$}}])$ . The rate $c_n$ serves to cope with potential degeneracy of the test statistic, an issue inherent in the nonstandard nature of the problem.

Assumption \ref{Ass: strong approx}(ii) demands an analogous approximation for the bootstrap. In regular settings, typically $\hat{\mathbb G}_n=\sqrt n \{\hat\theta_n^*-\hat\theta_n\}$, where $\hat\theta_n^*$ is the same as $\hat\theta_n$ but based on samples drawn from the original data. In Example \ref{Ex: NP under shape}, the expansion \eqref{Eqn: implementation, aux1} suggests
\begin{align}\label{Eqn: implementation, aux2}
\hat{\mathbb G}_n=k_n^{-1/2}(h^{k_n})^\text{\scalebox{0.8}{$\intercal$}} (\frac{1}{n}\sum_{i=1}^nh^{k_n}(Z_i)h^{k_n}(Z_i)^\text{\scalebox{0.8}{$\intercal$}})^{-1}\frac{1}{\sqrt n}\sum_{i=1}^nW_ih^{k_n}(Z_i)\hat u_i~.
\end{align}
where $\hat u_i\equiv Y_i-\hat\theta_n(Z_i)$ and $\{W_i\}_{i=1}^n$ are (scalar) weights (e.g., standard normals). The verification of Assumption \ref{Ass: strong approx} for Example \ref{Ex: NP under shape} represents a general strategy: establishing an asymptotic linear expansion of $r_n\{\hat\theta_n-\theta_P\}$, and then verifying Assumption \ref{Ass: strong approx} based on the linear term---see \citet{ChernozhukovLeeRosen2013Intersection}, \citet{ChernozhukovNeweySantos2019CCMM}, \citet{ChenChristensen2018SupNormOptimal}, \citet{BelloniChernozhukovChetverikovFernandez2019QR}, and \citet{LiLiao2020Uniform} for more illustrations. We stress that Assumption \ref{Ass: strong approx}(ii) leaves the particular form of $\hat{\mathbb G}_n$ unspecified, and thus accommodates alternative resampling schemes.



\begin{figure}
\centering \scriptsize
\hspace*{\fill}
\begin{subfigure}[b]{0.45\linewidth}
\begin{tikzpicture}[scale=0.8, >=latex',dot/.style={circle,inner sep=1pt,fill,name=#1}]
\draw (-4,-2.5)--(4,-2.5) -- (4,3) -- (-4,3) -- (-4,-2.5);
\draw[loosely dotted,->] (-4,0)--(4,0);
\draw[loosely dotted,->] (0,-2.5)--(0,3);
\fill[red] (0,0) circle (0.8pt);
\draw (0,0) node[anchor=north west] {$0$};
\fill[fill=Honeydew3,semitransparent] (0,0)--(4/5,6/5)--(4,0)--(0,0);
\draw[line width=0.07mm,NavyBlue] (0,0)--(4/5,6/5)--(4,0)--(0,0);
\draw (2,1) node[anchor=west] {$\Lambda$};
\node[dot=theta,label={east: $\theta_0$}] at (2/5,3/5) {};
\node[dot=h,label={west:$h$}] at (-2.5,-0.5)  {};
\node (zero) at (-5/3,-2.5)  {};
\draw (-3.5,-2) -- (-1/6,3);
\draw (-1.7,2.3) node[anchor=south, semitransparent,SeaGreen4] {$h+a\theta_0:a\ge 0$};
\draw (-2,-1.5) node[anchor=north, semitransparent,SeaGreen4] {$h+a\theta_0:a\le 0$};
\draw[dashed] (h)--(0,0);
\draw[dashed] (-37/18,1/6)--(0,0);
\draw[dashed] (-3/2,1)--(0,0);
\draw[dashed] (-7/10,11/5)--(4/5,6/5);
\draw[dashed] (-1/4,23/8)--(4/5,6/5);
\fill[red] (-3/2,1) circle (0.8pt);
\draw (-3/2,1)  node[anchor=south east] {$h+a^*\theta_0$};
\end{tikzpicture}
\caption{$\Lambda$ is convex but not conic}\label{Fig: convex not cone}
\end{subfigure}
\hspace*{\fill}
\begin{subfigure}[b]{0.45\linewidth}
\begin{tikzpicture}[scale=0.8, >=latex',dot/.style={circle,inner sep=1pt,fill,name=#1}]
\draw (-4,-2.5)--(-4,3)--(2,3);
\draw (-4,-2.5)--(4,-2.5) -- (4,0);
\draw[loosely dotted,->] (-4,0)--(4,0);
\draw[loosely dotted,->] (0,-2.5)--(0,3);
\fill[red] (0,0) circle (0.8pt);
\draw (0,0) node[anchor=north west] {$0$};
\fill[fill=Honeydew3,semitransparent] (0,0)--(2,3)--(4,3)--(4,0)--(0,0);
\draw[line width=0.07mm,NavyBlue] (0,0)--(2,3);
\draw[line width=0.07mm,NavyBlue] (0,0)--(4,0);
\draw (2,1) node[anchor=west] {$\Lambda$};
\node[dot=theta,label={east: $\theta_0$}] at (4/5,6/5) {};
\node[dot=h,label={west:$h$}] at (-2.5,-0.5)  {};
\node (zero) at (-5/3,-2.5)  {};
\draw (-3.5,-2) -- (-1/6,3);
\draw (-1.7,2.3) node[anchor=south, semitransparent,SeaGreen4] {$h+a\theta_0:a\ge 0$};
\draw (-2,-1.5) node[anchor=north, semitransparent,SeaGreen4] {$h+a\theta_0:a\le 0$};
\draw[dashed] (h)--(0,0);
\draw[dashed] (-37/18,1/6)--(0,0);
\draw[dashed] (-3/2,1)--(0,0);
\fill[red] (-3/2,1) circle (0.8pt);
\draw (-3/2,1)  node[anchor=south east] {$h+a^*\theta_0$};
\end{tikzpicture}
\caption{$\Lambda$ is convex and conic}\label{Fig: convex cone}
\end{subfigure}
\hspace*{\fill}
\caption{$\psi_{a,P}(h)\equiv  \|h+a\theta_P-\Pi_\Lambda(h+a\theta_P)\|_{\mathbf H}$ is weakly decreasing in $a\in[0,\infty)$ if $\Lambda$ is convex and conic as in Figure \ref{Fig: convex cone}, but may not be so if it is convex but not conic as in Figure \ref{Fig: convex not cone}.}
\label{Fig: monotonicity}
\end{figure}

The next lemma lays out a number of important building blocks for our development.

\begin{lem}\label{Lem: main}
(i) If Assumption \ref{Ass: convex setup} holds, then any $\hat\theta_n\in\mathbf H$ and $r_n\in\mathbf R_+$ satisfy
\begin{align}
r_n\phi(\hat\theta_n)=\|r_n\{\hat\theta_n-\theta_P\}+r_n\theta_P-\Pi_\Lambda(r_n\{\hat\theta_n-\theta_P\}+r_n\theta_P)\|_{\mathbf H}~. \label{Eqn: test statistic, rep}
\end{align}
(ii) If Assumption \ref{Ass: convex setup} holds, $P\in\mathbf P_0$, and $\kappa_n\in[0,r_n]$, then it follows that
\begin{align}
r_n\phi(\hat\theta_n)\le \|r_n\{\hat\theta_n-\theta_P\}+\kappa_n\theta_P-\Pi_\Lambda(r_n\{\hat\theta_n-\theta_P\}+\kappa_n\theta_P)\|_{\mathbf H}~. \label{Eqn: stochastic domiance}
\end{align}
(iii) If Assumption \ref{Ass: strong approx}(i) holds, $\sup_{P\in\mathbf P} E[\|\mathbb Z_{n,P}\|_{\mathbf H}]<\infty$ uniformly in $n$, and $\kappa_n/r_n=o(c_n)$, then we have that $\kappa_n\hat\theta_n-\kappa_n\theta_P= o_p(c_n)$ uniformly in $P\in\mathbf P$.

\end{lem}

Lemma \ref{Lem: main}(i) highlights the standard, and critically, the nonstandard features of the problem. Specifically, while $r_n\{\hat\theta_n-\theta_P\}$ may be approximated in law by $\hat{\mathbb G}_n$ due to Assumption \ref{Ass: strong approx}, $r_n\theta_P$ cannot be consistently estimated in general, i.e., $r_n\hat\theta_n-r_n\theta_P\neq o_p(1)$. Lemma \ref{Lem: main}(ii) suggests that we may obtain critical values by instead estimating the law of the upper bound in \eqref{Eqn: stochastic domiance}. This is possible because $\kappa_n\theta_P$ can be consistently estimated by $\kappa_n\hat\theta_n$ as shown by Lemma \ref{Lem: main}(iii). We stress that \eqref{Eqn: stochastic domiance} is implied by the convex cone property but not convexity alone---see Figure \ref{Fig: monotonicity}.


To formalize the above discussions, we define maps $\psi_{a,P},\hat\psi_{\kappa_n}:\mathbf H\to\mathbf R$ by
\begin{align}\label{Eqn: psi function}
\psi_{a,P}(h)&\equiv  \|h+a\theta_P-\Pi_\Lambda(h+a\theta_P)\|_{\mathbf H}~,\\
\hat\psi_a(h)&\equiv \|h+a\Pi_\Lambda\hat\theta_n-\Pi_\Lambda(h+a\Pi_\Lambda\hat\theta_n)\|_{\mathbf H}~.
\end{align}
Thus, $\psi_{\kappa_n,P}(r_n\{\hat\theta_n-\theta_P\})$ is precisely the upper bound in \eqref{Eqn: stochastic domiance}, and $\hat\psi_{\kappa_n}(\hat{\mathbb G}_n)$ is its bootstrap analog in which the null hypothesis is enforced through $\Pi_\Lambda\hat\theta_n$ to improve power. Finally, for a significance level $\alpha\in(0,1)$, we define our critical value
\begin{align}\label{Eqn: critical value}
\hat c_{n,1-\alpha}\equiv \inf\{c\in\mathbf R: P(\hat\psi_{\kappa_n}(\hat{\mathbb G}_n)\le c|\{X_i\}_{i=1}^n)\ge 1-\alpha\}~.
\end{align}
As known in the literature \citep{ChernozhukovNeweySantos2019CCMM}, the validity of $\hat c_{n,1-\alpha}$, viewed as a mapping from distributions to the real line, additionally demands a suitable continuity condition, in accord with the continuous mapping theorem. To this end, we let $c_{n,P}(1-\alpha)$ be the $(1-\alpha)$-quantile of $\psi_{\kappa_n,P}(\mathbb Z_{n,P})$ and impose the following:




\begin{ass}\label{Ass: Gaussian}
(i) $\mathbb Z_{n,P}$ is tight and centered Gaussian in $\mathbf H$ for each $n\in\mathbf N$ and $P\in\mathbf P$; (ii) $\sup_{P\in\mathbf P} E[\|\mathbb Z_{n,P}\|_{\mathbf H}]<\infty$ uniformly in $n$; (iii) $c_{n,P}(1-\alpha-\varpi)\ge c_{n,P}(0.5)+\varsigma_n$ for some constant $\varpi,\varsigma_n>0$, each $n$ and $P\in\mathbf P_0$; (iv) $c_n/\varsigma_n^2=O(1)$ as $n\to\infty$.
\end{ass}

Assumption \ref{Ass: Gaussian}(i) formalizes Gaussianity and tightness of each $\mathbb Z_{n,P}$, which is fulfilled in Examples \ref{Ex: NP under shape}--\ref{Ex: Slutsky} as well as most regular settings. Assumption \ref{Ass: Gaussian}(ii) is a mild moment condition that, in our examples, is tantamount to uniform boundedness of eigenvalues of some matrices with growing dimensions. Assumptions \ref{Ass: Gaussian}(iii)(iv) are tied to the natures of densities of Gaussian functionals. Together with Assumptions \ref{Ass: convex setup} and \ref{Ass: Gaussian}(i)-(ii), they in effect amount to the aforementioned continuity condition.

We now state the first main result concerning the size of our test.

\begin{thm}\label{Thm: size control}
Let Assumptions \ref{Ass: convex setup}, \ref{Ass: strong approx}, and \ref{Ass: Gaussian} hold. If $0\le\kappa_n/r_n=o(c_n)$, then
\begin{align}\label{Thm: size control, aux1}
\limsup_{n\to\infty}\sup_{P\in\mathbf P_0} P( r_n \phi(\hat\theta_n)> \hat c_{n,1-\alpha})\le \alpha~,
\end{align}
and, for $\bar{\mathbf P}_0\equiv\{P\in\mathbf P_0: \langle\, \vartheta,\theta_P\rangle_{\mathbf H}=0\,\,\forall\,\vartheta\in\mathbf H\,\,\mathrm{ s.t.}\,\, \sup_{\lambda\in\Lambda}\langle\vartheta,\lambda\rangle_{\mathbf H}\le 0\}$,
\begin{align}\label{Thm: size control, aux2}
\limsup_{n\to\infty}\sup_{P\in\bar{\mathbf P}_0} |P( r_n \phi(\hat\theta_n)> \hat c_{n,1-\alpha})-\alpha|=0~.
\end{align}
\end{thm}

In addition to size control, Theorem \ref{Thm: size control} shows that the limiting rejection rate equals the nominal level, uniformly over $\bar{\mathbf P}_0$. Heuristically, $\bar{\mathbf P}_0$ may be viewed as consisting of the least favorable distributions in $\mathbf P_0$. Indeed, one can show that (i) $\bar{\mathbf P}_0=\{P\in\mathbf P_0: \theta_P=0\}$ if $\Lambda=\mathbf R_+$, and (ii) $\bar{\mathbf P}_0$ contains constant (resp.\ linear) functions if $\Lambda$ is the family of monotone (resp.\ concave) functions in $L^2(\mathcal Z)$. We stress that, the choice $\kappa_n=0$ leads to a least favorable test that amounts to assuming $\theta_P=0$ (or $\theta_P\in\bar{\mathbf P}_0$ by Lemma \ref{Lem: nonconservativeness}). As well documented, least favorable tests, while controlling size uniformly, can be substantially conservative, which has motivated the active development of more powerful tests in nonstandard settings \citep{AndrewsSoares2010,LeeSongWhang2017Inequal}. Intuitively, in view of Lemma \ref{Lem: main}(ii), it is desirable to have $\kappa_n\to\infty$ (to match $r_n\to\infty$), a condition recurrent in nonstandard problems for the sake of nonconservativeness---see \citet{FangSantos2018HDD} and Appendix \ref{Sec: special case}.

We emphasize that our test is in general asymptotically nonsimilar, i.e., there may exist a sequence $\{P_n\}$ of distributions from the null such that
\begin{align}\label{Eqn: similarity}
\liminf_{n\to\infty}P_n(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha})<\alpha~.
\end{align}
However, this is, in our view, neither an evidence against nonconservativeness nor a deficiency of our test, analogous to the one-sided $t$-test of $\mathrm H_0: \theta_0\le 0$ whose rejection rates tend to zero at $\theta_0<0$. At a deeper level, \eqref{Eqn: similarity} is in line with the fact that similarity is not a desirable criterion in nonstandard settings, and can lead to tests with very poor power \citep{Lehmann1952Multi,Andrews2012Similar}. Indeed, many powerful tests in these settings are asymptotically nonsimilar---see, e.g., \citet{AndrewsSoares2010}, \citet{LeeSongWhang2017Inequal} and \citet{Chetverikov2018Monotonicity}.







Turning to the power of our test, we define $\mathbf P_{1,n}^\Delta\equiv\{P\in\mathbf P_1: \phi(\theta_P)\ge \Delta/r_n\}$ for $\Delta>0$. The next theorem shows that our test has nontrivial power against $\mathbf P_{1,n}^\Delta$.

\begin{thm}\label{Thm: uniform power}
If Assumptions \ref{Ass: convex setup}, \ref{Ass: strong approx}, and \ref{Ass: Gaussian}(ii) hold and $\kappa_n\ge 0$, then
\begin{align}\label{Thm: uniform power eqn2}
\liminf_{\Delta\to\infty}\liminf_{n\to\infty}\inf_{P\in\mathbf P_{1,n}^\Delta} P( r_n \phi(\hat\theta_n)> \hat c_{n,1-\alpha})=1~.
\end{align}
\end{thm}

If we employ the kernel estimator in Example \ref{Ex: NP under shape}, then Theorem \ref{Thm: uniform power} predicts that our test is powerful against Pitman drifts of order $(nh_n^{d_z})^{-1/2}$. In comparison, the test of \citet{LeeSongWhang2017Inequal} is powerful against drifts of order $(nh_n^{2\nu})^{-1/2}$ or $(nh_n^{d_z/2+2\nu})^{-1/2}$ depending on the drifts, where $\nu=1$ for monotonicity and $\nu=2$ for concavity. Thus, for monotonicity, our convergence rate is faster if $d_z=1$ but may be slower if $d_z>4$; for concavity, our rate is faster if $d_z<4$ but may be slower if $d_z>8$. Theorem \ref{Thm: uniform power} also suggests that $\kappa_n$ affects power through higher order terms, and, we reiterate that asymptotic nonconservativeness requires $\kappa_n\to\infty$ subject to $\kappa_n/r_n=o(c_n)$.














\subsection{Selection of the Tuning Parameter}\label{Sec: tuning parameter}


To motivate, we note that Lemma \ref{Lem: main}(iii) implies: for any $\epsilon>0$,
\begin{align}
P(\frac{\kappa_n}{r_n}\|r_n\{\hat\theta_n-\theta_P\}\|_{\mathbf H}\le c_n\epsilon)=P(\|r_n\{\hat\theta_n-\theta_P\}\|_{\mathbf H}\le \frac{r_nc_n}{\kappa_n}\epsilon)\to 1~.
\end{align}
This suggests that we could choose $\kappa_n$ to be such that $r_nc_n/\kappa_n$ is the $(1-\gamma_n)$-quantile of $\|r_n\{\hat\theta_n-\theta_P\}\|_{\mathbf H}$ with $\gamma_n\downarrow 0$, or the $(1-\gamma_n)$ conditional quantile of $\|\hat{\mathbb G}_n\|_{\mathbf H}$ (given $\{X_i\}_{i=1}^n$) since $\|r_n\{\hat\theta_n-\theta_P\}\|_{\mathbf H}$ is unknown. Formally, we let $\hat\kappa_n\equiv r_nc_n/\hat\tau_{n,1-\gamma_n}$, where
\begin{align}\label{Eqn: tuning parameter, data driven}
\hat\tau_{n,1-\gamma_n}\equiv \inf\{c\in\mathbf R: P(\|\hat{\mathbb G}_n\|_{\mathbf H}\le c|\{X_i\}_{i=1}^n)\ge 1-\gamma_n\}~.
\end{align}
To justify the construction $\hat\kappa_n$, we need to introduce our final assumption.

\begin{ass}\label{Ass: tuning}
$\liminf_{n\to\infty}\inf_{P\in\mathbf P_0} \bar\sigma_{n,P}^2>0$ with $\bar\sigma_{n,P}^2\equiv \sup_{h\in\mathbf H: \|h\|_{\mathbf H}\le 1} E[\langle h,\mathbb Z_{n,P}\rangle_{\mathbf H}^2]$.
\end{ass}

Heuristically, Assumption \ref{Ass: tuning} requires that the coupling variable $\mathbb Z_{n,P}$ for $\hat\theta_n$ be asymptotically non-degenerate. In turn, we can now verify the validity of $\hat\kappa_n$.

\begin{pro}\label{Pro: tuning parameter}
Let Assumptions \ref{Ass: strong approx}, \ref{Ass: Gaussian}(i)-(ii), and \ref{Ass: tuning} hold, and set $\hat\kappa_n\equiv r_nc_n/\hat\tau_{n,1-\gamma_n}$ with $\gamma_n\in(0,1)$ and $\hat\tau_{n,1-\gamma_n}$ as in \eqref{Eqn: tuning parameter, data driven}. If $\gamma_n\to 0$, then $\hat\kappa_n/r_n = o_p(c_n)$ uniformly in $P\in\mathbf P_0$. If $(r_nc_n)^{-2}\log\gamma_n\to 0$, then $\hat\kappa_n\xrightarrow{p}\infty$ uniformly in $P\in\mathbf P_0$.
\end{pro}

Analogous rate conditions on $\gamma_n$ have appeared in parametric settings---see \citet{ChenFang2016Rank} and references therein. In nonparametric settings, the practice of obtaining data-driven tuning parameters through quantile estimation has appeared in \citet{ChernozhukovLeeRosen2013Intersection}, \citet{ChernozhukovNeweySantos2019CCMM}, and \citet{FangSantos2018HDD}, though a formal theory seems to be lacking. While the choice of $\gamma_n$ remains technically challenging, the situation somewhat improves because $\gamma_n$ is unit/scale-free, and prior studies such as \citet{FangSantos2018HDD} and \citet{ChenFang2016Rank} have shown that finite sample results are often insensitive to the choice of $\gamma_n$, as also confirmed in our simulations. We recommend $\gamma_n=0.01/\log n$ or $1/n$ for practical implementations.

Finally, the rate $c_n$ (in $\hat\kappa_n$) may be ignored in regular settings, and set to be $1/\log n$ in irregular settings without endogeneity. In Example \ref{Ex: NPIV under shape}, we may let $c_n$ be $(\log k_n)^{-\varsigma}$ for some $\varsigma\in[1/2,1]$ with the sieve 2SLS estimation and $k_n$ basis functions for $\theta_0$ \citep{ChenChristensen2018SupNormOptimal}. We note that $\kappa_n$ affects the critical value $\hat c_{n,1-\alpha}$ and hence the rejection rates monotonically: a smaller (resp.\ larger) $\kappa_n$ leads to less (resp.\ more) rejections---see Lemma \ref{Lem: projection monotonicity}. As a result, if one is uncertain about $k_n$ or $\varsigma$, he/she could simply take $c_n=1/\log n$, thereby only making $\hat c_{n,1-\alpha}$ larger.


\subsection{Implementation and Practical Issues}\label{Sec: implementation}


We next provide a guide for implementing our test. Computation of projections (needed in {\sc Steps} 1 and 2 below) will be discussed in the end.

\noindent\underline{\sc Step 1:} Compute the test statistic $r_n\phi(\hat\theta_n)=r_n\|\hat\theta_n-\Pi_\Lambda(\hat\theta_n)\|_{\mathbf H}$.

This step requires an unconstrained estimator $\hat\theta_n$ of $\theta_P$, which may be obtained by standard estimation procedures---see Examples \ref{Ex: NP under shape}-\ref{Ex: Slutsky} for specific estimators and their rates $r_n$. We stress that our framework imposes no additional structures on $\hat\theta_n$ as far as implementation is concerned. With $\hat\theta_n$ and its projection $\Pi_\Lambda(\hat\theta_n)$ in hand, one may approximate $r_n\phi(\hat\theta_n)$ by, e.g., the trapezoid rule in Examples \ref{Ex: NP under shape}--\ref{Ex: Slutsky} where the $\|\cdot\|_{\mathbf H}$-norm takes the form of an integral. Concretely, in Example \ref{Ex: NP under shape} with $\mathcal Z=[a,b]$ and $a<b$, we approximate $\phi(\hat\theta_n)$ by: for a large $N$ and $\Delta=(b-a)/N$,
\begin{align}\label{Eqn: trapezoid}
\{\frac{1}{2}\frac{b-a}{N} [f^2(z_0)+2f^2(z_1)+\cdots +2f^2(z_{N-1})+f^2(z_N)]\}^{1/2}~,
\end{align}
where $f=\hat\theta_n-\Pi_\Lambda(\hat\theta_n)$ and $z_j=a+j\Delta$.


\noindent\underline{\sc Step 2:} Construct the critical value $\hat c_{n,1-\alpha}$ defined in \eqref{Eqn: critical value}.

The construction requires a bootstrap analog $\hat{\mathbb G}_n$ of $r_n \{\hat\theta_n-\theta_P\}$ and a tuning parameter $\kappa_n$. In regular settings, one often has $\hat{\mathbb G}_n=\sqrt n \{\hat\theta_n^*-\hat\theta_n\}$, where $\hat\theta_n^*$ is the same as $\hat\theta_n$ but based on bootstrap samples. In Examples \ref{Ex: NP under shape}--\ref{Ex: Slutsky}, one may employ various bootstrap schemes in \citet{ChernozhukovLeeRosen2013Intersection}, \citet{ChernozhukovNeweySantos2019CCMM}, \citet{ChenChristensen2018SupNormOptimal}, and \citet{BelloniChernozhukovChetverikovFernandez2019QR}. For $\kappa_n$, we recommend $\hat\kappa_n$ proposed in Section \ref{Sec: tuning parameter} with a suitable $\gamma_n$ (e.g., $\gamma_n=0.01/\log n$ or $1/n$). Summarizing, the critical value $\hat c_{n,1-\alpha}$ may be then obtained as follows:
\begin{enumerate}[label=(\roman*)]
  \item Generate $B$ realizations $\{\hat{\mathbb G}_{n,b}\}_{b=1}^B$ of $\hat{\mathbb G}_n$ (e.g., $B=200$ or larger); e.g., in \eqref{Eqn: implementation, aux2}, generate $\{\{W_{i,b}\}_{i=1}^n\}_{b=1}^B$ that are i.i.d.\ across both $i$ and $b$, and then obtain each $\hat{\mathbb G}_{n,b}$ by evaluating $\hat{\mathbb G}_n$ at $\{W_{i,b}\}_{i=1}^n$, which only involves linear calculations.
  \item Set $\hat\kappa_n=r_nc_n/\hat\tau_{n,1-\gamma_n}$ where $\hat\tau_{n,1-\gamma_n}$ is the $(1-\gamma_n)$-quantile of the $B$ numbers $\|\hat{\mathbb G}_{n,1}\|_{\mathbf H},\ldots, \|\hat{\mathbb G}_{n,B}\|_{\mathbf H}$, and $c_n=1/\log n$ for Example \ref{Ex: NP under shape}. The $\|\cdot\|_{\mathbf H}$-norms (here and below) may be computed by the trapezoid rule as in {\sc Step} 1.
  \item Approximate $\hat c_{n,1-\alpha}$ by the $(1-\alpha)$-quantile of the $B$ numbers
\begin{multline}\label{Eqn: implementation, aux3}
\|\hat{\mathbb G}_{n,1}+\hat\kappa_n\Pi_\Lambda\hat\theta_n-\Pi_\Lambda(\hat{\mathbb G}_{n,1}+\hat\kappa_n\Pi_\Lambda\hat\theta_n)\|_{\mathbf H},\\
\ldots, \|\hat{\mathbb G}_{n,B}+\hat\kappa_n\Pi_\Lambda\hat\theta_n-\Pi_\Lambda(\hat{\mathbb G}_{n,B}+\hat\kappa_n\Pi_\Lambda\hat\theta_n)\|_{\mathbf H}~.
\end{multline}
\end{enumerate}

\noindent\underline{\sc Step 3:} Reject $\mathrm H_0$ if and only if $r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha}$.

\bigskip


Next, we illustrate the computation of projections. As described in Appendix \ref{App: convex cone}, when closed form expressions do not exist, the projection $\Pi_{\Lambda}(\theta)$ can be computed by solving a linearly constrained quadratic program: for some $A\in\mathbf M^{m\times k}$,
\begin{align}\label{Eqn: LCQP}
\min_{h\in\mathbf R^k}\|h-\vartheta\|\qquad \text{s.t.}\quad Ah\ge 0~,
\end{align}
where $\vartheta\in\mathbf R^k$ is the vector of values $\theta$ takes at the grid points, and $Ah\ge 0$ is the discretized version of the restriction in question. As an extensively studied problem, \eqref{Eqn: LCQP} admits polynomial-time algorithms; e.g., the iteration complexity of the interior point method is $O(\sqrt{m+k}\log(1/\epsilon))$ for an $\epsilon$-accurate solution.


Finally, since \eqref{Eqn: LCQP} is inherently more complicated to compute in higher dimensions, we note a number of strategies for ameliorating the situation. First, the recently developed open-source solver OSQP (\url{https://osqp.org}) is very robust in solving large scale quadratic programming problems. Second, while projection under convexity/concavity is computationally demanding in multivariate settings, the representation result in \citet{Kuosmanen2008Convex}, when coupled with the OSQP solver, can help significantly reduce the computation cost. Last, with regressors more than two or three, one may consider employing semiparametric models, instead of fully nonparametric ones, to alleviate the curse of dimensionality.





\section{Simulation Studies}\label{Sec: simulations}

This section examines the finite sample performance of our test. Due to space limitation, we focus on monotonicity, and defer concavity, monotonicity jointly with concavity, and Slutsky restriction to the online appendix. Throughout, the significance level is  5\%, and, unless otherwise specified, the number of Monte Carlo replications is 3000 while the number of bootstrap repetitions for each replication is 200. All integrals are approximated by the trapezoid rule, and quadratic programs are solved by the OSQP solver. Our test is implemented based on the data-driven choice $\hat\kappa_n$ with $\gamma_n\in\{n^{-1/2}, n^{-3/4}, 1/n, 0.1/\log n, 0.05/\log n, 0.01/\log n, 0.1, 0.05, 0.01\}$.


\subsection{Simulation Designs}

We aim to test whether $\theta_0$ is nondecreasing in Example \ref{Ex: NP under shape} with $d_z\in\{1,2\}$. For $d_z=1$, the regression function $\theta_0:[-1,1]\to\mathbf R$ is specified as: for some $(\mathsf a,\mathsf b,\mathsf c)\in\mathbf R^3$,
\begin{align}\label{Eqn: MC1,aux1}
\theta_0(z)=\mathsf az-\mathsf b\varphi(\mathsf cz)~,
\end{align}
where $\varphi$ is the standard normal pdf. We consider three choices for $(\mathsf a,\mathsf b,\mathsf c)$ under the null hypothesis, namely $(0,0,0)$, $(0.1,0.5,0.5)$ and $(0.5,2,1)$ that are labeled D1, D2 and D3 respectively, and the collection $\{(\mathsf a,\mathsf b,\mathsf c): \mathsf a =0,\mathsf b = 0.2\delta,\mathsf c = 5+0.1\delta,\delta=1,\ldots,10\}$ under the alternative---see Figure \ref{Fig: MC1,theta}. We then draw i.i.d.\ samples $\{Z_i^*,u_i\}_{i=1}^n$ with $n\in\{500,750,1000\}$ from the standard normal distribution in $\mathbf R^2$, and set $Z_i=-1+2\Phi(Z_i^*)\in[-1,1]$ with $\Phi$ the standard normal cdf. For $d_z=2$, the regression function $\theta_0:[0,1]^2\to\mathbf R$ is of the form: for some $(\mathsf a,\mathsf b,\mathsf c)\in\mathbf R^3$,
\begin{align}\label{Eqn: MC2,aux1}
\theta_0(z_1,z_2)=\mathsf a\big(\frac{1}{2}z_1^{\mathsf b}+\frac{1}{2}z_2^{\mathsf b}\big)^{1/\mathsf b}+\mathsf c\log(1+z_1+z_2)~.
\end{align}
We consider three choices of $(\mathsf a,\mathsf b,\mathsf c)$ for the null, namely $(0,0,0)$, $(0.2,1,0)$ and $(0.5,0,0.5)$ that are labeled D1, D2, and D3, respectively, and, for the alternative, set $\mathsf b=0$ and $\mathsf a=\mathsf c=-\Delta\delta $ with $\Delta=0.05$ and $\delta=1,\ldots,10$. Note that the first term on the right-hand side of \eqref{Eqn: MC2,aux1} collapses to $\mathsf a\sqrt{z_1z_2}$ whenever $\mathsf b=0$. In turn, we draw i.i.d.\ samples $\{Z_{1i}^*,Z_{2i}^*,u_i\}_{i=1}^n$ with $n\in\{500,750,1000\}$ from the standard normal distribution in $\mathbf R^3$, and set $Z_i\equiv(Z_{1i},Z_{2i})$ with $Z_{ji}=\Phi(Z_{ji}^*)\in[0,1]$ for all $i$ and $j=1,2$.

\begin{figure}
\centering
\scriptsize
\captionsetup[subfigure]{justification=centering}
\begin{subfigure}[b]{0.49\linewidth}
\resizebox{\textwidth}{!}{
\begin{tikzpicture}[baseline,>=latex']
\begin{axis}[axis equal image,
             domain=-1:1,
             samples=100,
             xmin=-1.1, xmax=1.1,
             ymin=-1.1, ymax=0.1,
             enlargelimits=false,
             xlabel=$z$,
             x label style={at={(axis description cs:0.5,-0.02)},anchor=north},
             xtick={-1,-0.5,...,1},
             xticklabels={-1,-0.5,...,1},
             xticklabel style={yshift=0.05cm,anchor=south},
             ytick={-1,-0.5,0},
             yticklabels={-1,-0.5,0},
             every axis plot/.append style={semithick},
            ]
\addplot[NavyBlue]{0};
\addplot[NavyBlue,dashed]{0.1*x-0.5*npdf(0.5*x, 0, 1)};
\addplot[NavyBlue,densely dotted]{0.5*x-2*npdf(x, 0, 1)};
\end{axis}
\end{tikzpicture}
}
\caption{$\mathrm H_0$: D1 (solid), D2 (dashed) and D3 (dotted)}\label{Fig: MC1a}
\end{subfigure}
\hspace*{\fill}
\begin{subfigure}[b]{0.49\linewidth}
\resizebox{\textwidth}{!}{
\begin{tikzpicture}[baseline,>=latex']
\begin{axis}[axis equal image,
             domain=-1:1,
             samples=100,
             xmin=-1.1, xmax=1.1,
             ymin=-1.1, ymax=0.1,
             enlargelimits=false,
             xlabel=$z$,
             x label style={at={(axis description cs:0.5,-0.02)},anchor=north},
             xtick={-1,-0.5,...,1},
             xticklabels={-1,-0.5,...,1},
             xticklabel style={yshift=0.05cm,anchor=south},
             ytick={-1,-0.5,0},
             yticklabels={-1,-0.5,0},
             every axis plot/.append style={thin},
             no markers,
            ]
\foreach \b in {0.2,0.4,...,2}{\pgfmathsetmacro{\k}{\b*120}
\edef\temp{\noexpand
\addplot+[solid,color=Blue!\k]{-\b*npdf((5+\b/2)*x, 0, 1)};
}\temp
}
\end{axis}
\end{tikzpicture}
}
\caption{$\mathrm H_1$: $\delta=1,2,\ldots,10$ (from top to bottom)}\label{Fig: MC1b}
\end{subfigure}
\hspace*{\fill}
\caption{The function $\theta_0$ in \eqref{Eqn: MC1,aux1} where, in Figure \ref{Fig: MC1b}, $\mathsf a=0$, $\mathsf b = 0.2\delta$, and $\mathsf c=5+0.1\delta$. Note that the standard deviation of the error is designed to be no smaller than the range of $\theta_0$.}
\label{Fig: MC1,theta}
\end{figure}



To implement our test for \eqref{Eqn: MC1,aux1}, we obtain $\hat\theta_n$ by sieve estimation with cubic B-splines with 3, 5, or 7 interior knots at the equispaced quantiles of $\{Z_i\}_{i=1}^n$, so that the sieve dimension $k_n$ equals $7$, $9$, or $11$, respectively. In multivariate settings, we construct the series functions via tensor product of univariate B-splines. Since the sieve dimension grows quickly as $d_z$ increases (e.g., $k_n=49$ with cubic B-splines, $3$ knots, and $d_z=2$), we employ univariate quadratic as well as cubic B-splines with one or zero knots along each dimension for \eqref{Eqn: MC2,aux1}. Thus, for example, $k_n=9$ with quadratic B-splines and zero knots. In both designs, we compute $\hat{\mathbb G}_{n,b}$ as in \eqref{Eqn: implementation, aux2} by drawing i.i.d.\ weights from the standard normal distribution, and let the coupling rate $c_n=1/\log n$. To alleviate the boundary effects, we evaluate the $L^2$-norms (here and below) over $[-0.9,0.9]$ for \eqref{Eqn: MC1,aux1} and $[0.1,0.9]^2$ for \eqref{Eqn: MC2,aux1}, with step size $0.05$. For ease of reference, we label our test with quadratic B-splines and $j$ knots as FS-Q$j$; similarly, FS-C$j$ is the implementation with cubic B-splines and $j$ knots.

To compare, we implement two alternative nonconservative tests, \citet{LeeSongWhang2017Inequal} and \citet{Chetverikov2018Monotonicity}. The latter also compares with some prior tests in simulations, which show marked power superiority of the author's three tests so that we take them as important benchmarks. For brevity, however, we only present the one-step test, labeled C-OS, and note that the results for the other two tests are very similar to those produced by C-OS, as also observed in \citet{Chetverikov2018Monotonicity}. The details for implementing C-OS in the univariate case are clearly laid out in \citet[p.749]{Chetverikov2018Monotonicity}. For the bivariate case, we compute the test statistic as on p.27 in the arXiv version of \citet{Chetverikov2018Monotonicity}, adopt equispaced empirical quantiles $0.1,0.15,...,0.9$ of each covariate as locations for the weighting function to make the computation feasible, and otherwise follow the implementation in the univariate case.

\citet{LeeSongWhang2017Inequal} considered $L^p$-type statistics. For the sake of comparison, we focus on $p =2$, since different choices of $p$ implicitly aim power at different alternatives. We estimate the first derivatives of $\theta_0$ by local quadratic regression \citep[p.59]{FanGijbels1996Local}, with a kernel $z\mapsto 1.5\max\{1-(2z)^2,0\}$ for \eqref{Eqn: MC1,aux1} and $z\mapsto 0.75\max\{1-z^2,0\}$ for \eqref{Eqn: MC2,aux1}. We choose two bandwidths: a ``large'' one $2s_nn^{-1/(d_z+2q+2)}$ (in the spirit of \citet{LeeSongWhang2017Inequal}) and a ``small'' one $2s_nn^{-1/(d_z+2(q-1)+2)}$ with $q=2$ and $s_n$ the standard derivation of $\{Z_i\}_{i=1}^n$, resulting in two tests labeled LSW-L and LSW-S, respectively. As studentization can be crucial for power (based on unreported simulations), we estimate the standard errors following \citet[p.115]{FanGijbels1996Local}, with the variance of the error estimated by a local polynomial regression of order $q+2$ with bandwidth $2s_nn^{-1/(d_z+2(q+2)+2)}$. Next, we construct critical values based on the empirical bootstrap (as in \citet{LeeSongWhang2017Inequal}) and the tuning parameter $\hat c_n$ in their Section 5.1 with $C_{\mathrm{cs}}=0.4$ (since their results are quite insensitive to other choices of $C_{\mathrm{cs}}$ there). Finally, to ease computation for the bivariate designs, the number of simulation replications is decreased to be 1000 (for LSW only).



{
\setlength{\tabcolsep}{6.5pt}
\begin{table}[!ht]
\caption{Empirical Size of Monotonicity Tests for $\theta_0$ in \eqref{Eqn: MC1,aux1} at $\alpha=5\%$} \label{Tab: MC1, size1, main}
\centering\footnotesize
\begin{threeparttable}
\sisetup{table-number-alignment = center, table-format = 1.3}
\begin{tabularx}{\linewidth}{@{} cc *{3}{S[round-mode = places,round-precision = 3]}  c *{3}{S[round-mode = places,round-precision = 3]}  c *{3}{S[round-mode = places,round-precision = 3]}@{}}
\hline
\hline
\multirow{2}{*}{$n$} & \multirow{2}{*}{$\gamma_n$}  & \multicolumn{3}{c}{FS-C3: $k_n=7$} & & \multicolumn{3}{c}{FS-C5: $k_n=9$} & & \multicolumn{3}{c}{FS-C7: $k_n=11$}\\
\cline{3-5} \cline{7-9} \cline{11-13}
 & & {D1} & {D2} & {D3}  & & {D1} & {D2} & {D3} & & {D1} & {D2} & {D3}\\
\hline
\multirow{3}{*}{$500$}   & $1/n$                  & 0.0530 & 0.0157 & 0.0030 & & 0.0577 & 0.0207 & 0.0033 & & 0.0583 & 0.0190 & 0.0027\\
                         & $0.01/\log n$          & 0.0530 & 0.0157 & 0.0030 & & 0.0577 & 0.0207 & 0.0033 & & 0.0583 & 0.0190 & 0.0027\\
                         & $0.01$                 & 0.0530 & 0.0157 & 0.0030 & & 0.0577 & 0.0207 & 0.0033 & & 0.0583 & 0.0193 & 0.0027\\
\cline{2-13}
 \multirow{3}{*}{$750$}  & $1/n$                  & 0.0523 & 0.0097 & 0.0013 & & 0.0560 & 0.0137 & 0.0017 & & 0.0590 & 0.0173 & 0.0030\\
                         & $0.01/\log n$          & 0.0523 & 0.0097 & 0.0013 & & 0.0560 & 0.0137 & 0.0017 & & 0.0590 & 0.0173 & 0.0030\\
                         & $0.01$                 & 0.0523 & 0.0097 & 0.0013 & & 0.0560 & 0.0140 & 0.0020 & & 0.0590 & 0.0173 & 0.0030\\
\cline{2-13}
 \multirow{3}{*}{$1000$} & $1/n$                  & 0.0557 & 0.0107 & 0.0007 & & 0.0560 & 0.0113 & 0.0003 & & 0.0560 & 0.0127 & 0.0007\\
                         & $0.01/\log n$          & 0.0557 & 0.0107 & 0.0007 & & 0.0560 & 0.0113 & 0.0003 & & 0.0560 & 0.0127 & 0.0007\\
                         & $0.01$                 & 0.0560 & 0.0107 & 0.0010 & & 0.0560 & 0.0117 & 0.0003 & & 0.0560 & 0.0130 & 0.0007\\
\hline
 \multicolumn{2}{c}{\multirow{2}{*}{$n$}}& \multicolumn{3}{c}{LSW-S} & & \multicolumn{3}{c}{LSW-L} & & \multicolumn{3}{c}{C-OS}\\
\cline{3-5} \cline{7-9} \cline{11-13}
 & & {D1} & {D2} & {D3}  & & {D1} & {D2} & {D3} & & {D1} & {D2} & {D3}\\
\hline
 \multicolumn{2}{c}{$500$}                        & 0.0600 & 0.0413 & 0.0083 & & 0.0660 & 0.0350 & 0.0037 & & 0.0597  & 0.0413  & 0.0120\\
 \multicolumn{2}{c}{$750$}                        & 0.0567 & 0.0363 & 0.0050 & & 0.0587 & 0.0303 & 0.0057 & & 0.0540  & 0.0337  & 0.0080\\
 \multicolumn{2}{c}{$1000$}                       & 0.0610 & 0.0347 & 0.0047 & & 0.0653 & 0.0347 & 0.0033 & & 0.0493  & 0.0357  & 0.0090\\
\hline
\hline
\end{tabularx}
\begin{tablenotes}[flushleft]
\item {\it Note:} The parameter $\gamma_n$ determines $\hat\kappa_n$ proposed in Section \ref{Sec: tuning parameter} with $c_n=1/\log n$ and $r_n=(n/k_n)^{1/2}$.
\end{tablenotes}
\end{threeparttable}
\end{table}
}

\subsection{Simulation Results}

Tables \ref{Tab: MC1, size1, main} and \ref{Tab: MC2, size} summarize the empirical sizes. Due to space limitation, we only present our tests with $\gamma_n\in\{1/n,0.01/\log n,0.01\}$, and relegate to Tables \ref{Tab: MC1Mon, size1, app}--H.\ref{Tab: MC2Mon, size, app} in Appendix \ref{App: full simulations} the complete set of results. Together, these tables show that our tests are insensitive to the choice of $\gamma_n$. For the univariate design, all tests control size well, though, relatively speaking, the two LSW tests, especially LSW-L, tend to over-reject under D1, while our tests tend to under-reject under D2 and D3. Note, however, that D1 is a least favorable case, D2 is in the ``interior'' (not in the topological sense), and D3 is further into the ``interior.'' Thus, the empirical sizes under D2 and D3 are expected to be smaller than 5\%. For the bivariate design, our tests are over-sized under D1 in small samples, though this feature is also shared by LSW-L and C-OS (to an overall lesser extent). The over-rejection of our tests, in particular FS-C1 (in which case $k_n=25$), is likely because the number of series functions is so ``large'' that the Gaussian approximation is somewhat inaccurate---note that the overall situation improves as $n$ increases. Thus, while undersmoothing requires a ``large'' $k_n$, there lies the tension that it should not be ``too large'' for the sake of distributional approximations.

{
\setlength{\tabcolsep}{3pt}
\begin{table}[!ht]
\caption{Empirical Size of Monotonicity Tests for $\theta_0$ in \eqref{Eqn: MC2,aux1} at $\alpha=5\%$} \label{Tab: MC2, size}
\centering\footnotesize
\begin{threeparttable}
\sisetup{table-number-alignment = center, table-format = 1.3}
\begin{tabularx}{\linewidth}{@{}cc *{3}{S[round-mode = places,round-precision = 3]} c *{3}{S[round-mode = places,round-precision = 3]} c *{3}{S[round-mode = places,round-precision = 3]} c *{3}{S[round-mode = places,round-precision = 3]}@{}}
\hline
\hline
 \multirow{2}{*}{$n$} & \multirow{2}{*}{$\gamma_n$}& \multicolumn{3}{c}{FS-Q0: $k_n=9$} & & \multicolumn{3}{c}{FS-Q1: $k_n=16$} & & \multicolumn{3}{c}{FS-C0: $k_n=16$} & & \multicolumn{3}{c}{FS-C1: $k_n=25$}\\
\cline{3-5} \cline{7-9} \cline{11-13} \cline{15-17}
&& {D1} & {D2} & {D3}  & & {D1} & {D2} & {D3} & & {D1} & {D2} & {D3} & & {D1} & {D2} & {D3}\\
\hline
\multirow{3}{*}{$500$} & $1/n$         & 0.0613 & 0.0183 & 0.0000 & & 0.0680 & 0.0300 & 0.0017 & & 0.0670 & 0.0303 & 0.0007 & & 0.0830 & 0.0437 & 0.0040\\
                       & $0.01/\log n$ & 0.0613 & 0.0183 & 0.0000 & & 0.0680 & 0.0300 & 0.0017 & & 0.0670 & 0.0303 & 0.0007 & & 0.0830 & 0.0437 & 0.0040\\
                       & $0.01$        & 0.0627 & 0.0193 & 0.0000 & & 0.0687 & 0.0300 & 0.0020 & & 0.0673 & 0.0307 & 0.0007 & & 0.0837 & 0.0447 & 0.0043\\
\cline{2-17}
 \multirow{3}{*}{$750$}& $1/n$         & 0.0583 & 0.0103 & 0.0000 & & 0.0653 & 0.0247 & 0.0003 & & 0.0640 & 0.0220 & 0.0000 & & 0.0737 & 0.0323 & 0.0003\\
                       & $0.01/\log n$ & 0.0583 & 0.0103 & 0.0000 & & 0.0653 & 0.0247 & 0.0003 & & 0.0640 & 0.0220 & 0.0000 & & 0.0737 & 0.0323 & 0.0003\\
                       & $0.01$        & 0.0597 & 0.0107 & 0.0000 & & 0.0660 & 0.0250 & 0.0003 & & 0.0650 & 0.0223 & 0.0000 & & 0.0740 & 0.0327 & 0.0007\\
\cline{2-17}
\multirow{3}{*}{$1000$}& $1/n$        & 0.0500 & 0.0107 & 0.0000 & & 0.0553 & 0.0210 & 0.0000 & & 0.0527 & 0.0203 & 0.0000 & & 0.0590 & 0.0227 & 0.0003\\
                       & $0.01/\log n$ & 0.0500 & 0.0107 & 0.0000 & & 0.0553 & 0.0210 & 0.0000 & & 0.0527 & 0.0203 & 0.0000 & & 0.0590 & 0.0227 & 0.0003\\
                       & $0.01$        & 0.0500 & 0.0107 & 0.0000 & & 0.0560 & 0.0213 & 0.0000 & & 0.0527 & 0.0203 & 0.0000 & & 0.0590 & 0.0230 & 0.0003\\
\hline
 \multicolumn{2}{c}{\multirow{2}{*}{$n$}}  & \multicolumn{3}{c}{LSW-S} & & \multicolumn{3}{c}{LSW-L} & & \multicolumn{3}{c}{C-OS} & & \multicolumn{3}{c}{ }\\
\cline{3-5} \cline{7-9} \cline{11-13}
 &     & {D1} & {D2} & {D3}  & & {D1} & {D2} & {D3} & & {D1} & {D2} & {D3}\\
\hline
\multicolumn{2}{c}{500}                & 0.0430 & 0.0200 & 0.0000 & & 0.0610  & 0.0260  & 0.0020 & & 0.0683  & 0.0607  & 0.0367\\
\multicolumn{2}{c}{750}                & 0.0540 & 0.0170 & 0.0000 & & 0.0660  & 0.0310  & 0.0000 & & 0.0570  & 0.0453  & 0.0230\\
\multicolumn{2}{c}{1000}               & 0.0440 & 0.0210 & 0.0000 & & 0.0540  & 0.0190  & 0.0000 & & 0.0553  & 0.0423  & 0.0187\\
\hline
\hline
\end{tabularx}
\begin{tablenotes}[flushleft]
\item {\it Note:} The parameter $\gamma_n$ determines $\hat\kappa_n$ proposed in Section \ref{Sec: tuning parameter} with $c_n=1/\log n$ and $r_n=(n/k_n)^{1/2}$.
\end{tablenotes}
\end{threeparttable}
\end{table}
}



\pgfplotstableread{
delta alpha MonFiveKn3 MonFiveKn5 MonFiveKn7 MonSevenKn3 MonSevenKn5 MonSevenKn7 MonTenKn3 MonTenKn5 MonTenKn7
0     0.05    0.0530     0.0577     0.0583     0.0523      0.0560      0.0590      0.0557   0.0560     0.0560
1     0.05    0.0697     0.0733     0.0707     0.0773      0.0717      0.0727      0.0857   0.0820     0.0780
2     0.05    0.1203     0.1183     0.1117     0.1563      0.1403      0.1337      0.1890   0.1683     0.1510
3     0.05    0.2143     0.2000     0.1800     0.3090      0.2790      0.2500      0.4087   0.3603     0.3220
4     0.05    0.3633     0.3197     0.2940     0.5107      0.4743      0.4337      0.6580   0.6180     0.5663
5     0.05    0.5357     0.4960     0.4467     0.7113      0.6853      0.6350      0.8597   0.8257     0.7837
6     0.05    0.6977     0.6717     0.6147     0.8787      0.8513      0.8190      0.9573   0.9487     0.9337
7     0.05    0.8323     0.8087     0.7667     0.9560      0.9460      0.9303      0.9890   0.9883     0.9823
8     0.05    0.9213     0.9077     0.8810     0.9873      0.9843      0.9773      0.9983   0.9980     0.9963
9     0.05    0.9640     0.9577     0.9440     0.9973      0.9973      0.9950      1.0000   1.0000     1.0000
10    0.05    0.9860     0.9843     0.9773     0.9993      0.9993      0.9993      1.0000   1.0000     1.0000
}\FirstMon



\pgfplotstableread{
delta alpha   Five1  Seven1   Ten1   Five2   Seven2   Ten2    Fifty2
0     0.05   0.0597  0.0540  0.0493  0.0683  0.0570  0.0553   0.0460
1     0.05   0.0623  0.0600  0.0620  0.0760  0.0600  0.0583   0.0653
2     0.05   0.0767  0.0873  0.1067  0.0797  0.0693  0.0690   0.1023
3     0.05   0.1173  0.1673  0.2260  0.0923  0.0783  0.0860   0.2077
4     0.05   0.1893  0.3010  0.4423  0.0980  0.0983  0.1077   0.4153
5     0.05   0.3210  0.5050  0.6797  0.1123  0.1190  0.1393   0.6890
6     0.05   0.4780  0.7067  0.8657  0.1283  0.1593  0.1877   0.9043
7     0.05   0.6383  0.8530  0.9613  0.1540  0.2120  0.2513   0.9810
8     0.05   0.7813  0.9470  0.9913  0.1810  0.2693  0.3413   0.9993
9     0.05   0.8807  0.9850  0.9987  0.2250  0.3410  0.4530   1.0000
10    0.05   0.9450  0.9957  1.0000  0.2737  0.4260  0.5753   1.0000
}\MainChetverikov

\pgfplotstableread{
delta alpha   FiveOp  FiveUn  SevenOp  SevenUn  TenOp    TenUn
0     0.05    0.0660  0.0600  0.0587   0.0567   0.0653   0.0610
1     0.05    0.0713  0.0607  0.0667   0.0563   0.0790   0.0653
2     0.05    0.0857  0.0660  0.0843   0.0677   0.0973   0.0737
3     0.05    0.1177  0.0787  0.1343   0.0783   0.1467   0.0933
4     0.05    0.1727  0.0993  0.2147   0.1053   0.2473   0.1217
5     0.05    0.2497  0.1293  0.3287   0.1453   0.3817   0.1693
6     0.05    0.3537  0.1717  0.4823   0.2027   0.5753   0.2360
7     0.05    0.4753  0.2280  0.6403   0.2870   0.7690   0.3387
8     0.05    0.6113  0.2973  0.7873   0.4017   0.8967   0.4620
9     0.05    0.7433  0.3777  0.9010   0.5153   0.9647   0.6140
10    0.05    0.8547  0.4803  0.9600   0.6527   0.9897   0.7613
}\MainFirstLSWMon



\pgfplotstableread{
delta alpha MonFiveQKn0 MonFiveQKn1 MonSevenQKn0 MonSevenQKn1 MonTenQKn0 MonTenQKn1
0     0.05    0.0613      0.0680       0.0583       0.0653      0.0500     0.0553
1     0.05    0.1110      0.0980       0.1270       0.1097      0.1267     0.0990
2     0.05    0.1987      0.1450       0.2530       0.1773      0.2607     0.1720
3     0.05    0.3173      0.2010       0.4100       0.2730      0.4573     0.2857
4     0.05    0.4597      0.2843       0.5907       0.3850      0.6733     0.4300
5     0.05    0.6147      0.3887       0.7643       0.5207      0.8473     0.5900
6     0.05    0.7457      0.5060       0.8843       0.6493      0.9493     0.7627
7     0.05    0.8573      0.6200       0.9587       0.7757      0.9823     0.8780
8     0.05    0.9313      0.7237       0.9853       0.8720      0.9963     0.9593
9     0.05    0.9790      0.8223       0.9973       0.9387      0.9993     0.9837
10    0.05    0.9937      0.8980       0.9993       0.9760      1.0000     0.9957
}\SecondMonQ

\pgfplotstableread{
delta alpha MonFiveCKn0 MonFiveCKn1 MonSevenCKn0 MonSevenCKn1 MonTenCKn0 MonTenCKn1
0     0.05    0.0670      0.0830       0.0640       0.0737      0.0527     0.0590
1     0.05    0.0960      0.1060       0.1090       0.1113      0.1027     0.0970
2     0.05    0.1413      0.1467       0.1827       0.1583      0.1800     0.1527
3     0.05    0.2070      0.1930       0.2737       0.2323      0.2917     0.2347
4     0.05    0.2860      0.2503       0.3917       0.3263      0.4400     0.3477
5     0.05    0.3947      0.3263       0.5243       0.4333      0.6017     0.4857
6     0.05    0.5073      0.4183       0.6517       0.5617      0.7630     0.6297
7     0.05    0.6240      0.5137       0.7750       0.6813      0.8787     0.7787
8     0.05    0.7307      0.6230       0.8763       0.7857      0.9557     0.8813
9     0.05    0.8303      0.7247       0.9413       0.8687      0.9837     0.9510
10    0.05    0.8970      0.8033       0.9737       0.9307      0.9940     0.9800
}\SecondMonC


\pgfplotstableread{
delta alpha   FiveOp  FiveUn  SevenOp  SevenUn  TenOp   TenUn
0     0.05    0.0610  0.0430  0.0660   0.0540   0.0540  0.0440
1     0.05    0.0920  0.0610  0.1000   0.0820   0.0960  0.0780
2     0.05    0.1300  0.0840  0.1670   0.1160   0.1520  0.1300
3     0.05    0.1960  0.1300  0.2430   0.1850   0.2380  0.1830
4     0.05    0.2600  0.1780  0.3450   0.2580   0.3800  0.2620
5     0.05    0.3460  0.2230  0.4750   0.3240   0.5210  0.3500
6     0.05    0.4560  0.2900  0.5810   0.4080   0.6530  0.4460
7     0.05    0.5590  0.3650  0.7200   0.4970   0.7980  0.5640
8     0.05    0.6780  0.4470  0.8230   0.5960   0.8970  0.6620
9     0.05    0.7780  0.5400  0.8900   0.7040   0.9720  0.7960
10    0.05    0.8410  0.6270  0.9410   0.7840   0.9930  0.8680
}\MainSecondLSWMon




Figure \ref{Fig: MC} depicts the power curves, where we only show our tests with $\gamma_n=0.01/\log n$ for brevity and the fact that other choices of $\gamma_n$ lead to very similar curves---see Figure \ref{Fig: MC1, FullMon} in Appendix \ref{App: full simulations}. For $d_z=1$, our tests are moderately more powerful than C-OS, across sample sizes and the number of interior knots, and they are all considerably more powerful than LSW-L and in particular LSW-S. For $d_z=2$, our tests remain competitive in terms of power, though LSW-L is more powerful than FS-C1 (the least powerful among our tests). Interestingly, C-OS is the least powerful of all tests, which may be explained by the fact that only part of the discordance between regressors and the outcome is being picked up through the indicator function in the test function $b(s)$---see p.27 in the arXiv version of \citet{Chetverikov2018Monotonicity}. We note that the power of our tests is overall decreasing in $k_n$, which is consistent with Theorem \ref{Thm: uniform power} since $r_n=\sqrt{n/k_n}$.

\begin{figure}[!h]
\centering\scriptsize
\begin{tikzpicture}
\begin{groupplot}[group style={group name=myplots,group size=3 by 2,horizontal sep= 0.8cm,vertical sep=1.1cm},
    grid = minor,
    width = 0.375\textwidth,
    xmax=10,xmin=0,
    ymax=1,ymin=0,
    every axis title/.style={below,at={(0.2,0.8)}},
    xlabel=$\delta$,
    x label style={at={(axis description cs:0.95,0.04)},anchor=south},
    xtick={0,2,...,10},
    ytick={0.05,0.5,1},
    tick label style={/pgf/number format/fixed},
    legend style={text=black,cells={align=center},row sep = 3pt,legend columns = -1, draw=none,fill=none},
    cycle list={
{smooth,tension=0.5,color=SpringGreen4, mark=halfsquare*,every mark/.append style={rotate=270},mark size=1.5pt,line width=0.5pt},
{smooth,tension=0.5,color=DarkOliveGreen4, mark=halfsquare*,every mark/.append style={rotate=90},mark size=1.5pt,line width=0.5pt},
{smooth,tension=0.5,color=Honeydew4, mark=10-pointed star,mark size=1.5pt,line width=0.5pt},
{smooth,tension=0.5,color=RoyalBlue1, mark=halfcircle*,every mark/.append style={rotate=90},mark size=1.5pt,line width=0.5pt},
{smooth,tension=0.5,color=RoyalBlue2, mark=halfcircle*,every mark/.append style={rotate=180},mark size=1.5pt,line width=0.5pt},
{smooth,tension=0.5,color=RoyalBlue3, mark=halfcircle*,every mark/.append style={rotate=270},mark size=1.5pt,line width=0.5pt},
{smooth,tension=0.5,color=RoyalBlue4, mark=halfcircle*,every mark/.append style={rotate=360},mark size=1.5pt,line width=0.5pt},
}
]
\nextgroupplot[legend style = {column sep = 7pt, legend to name = LegendMon1}]
\addplot[smooth,tension=0.5,color=NavyBlue, no markers,line width=0.25pt, densely dotted,forget plot] table[x = delta,y=alpha] from \FirstMon;
\addplot table[x = delta,y=FiveUn] from \MainFirstLSWMon;
\addplot table[x = delta,y=FiveOp] from \MainFirstLSWMon;
\addplot table[x = delta,y=Five1] from \MainChetverikov;
\addplot table[x = delta,y=MonFiveKn3] from \FirstMon;
\addplot table[x = delta,y=MonFiveKn5] from \FirstMon;
\addplot table[x = delta,y=MonFiveKn7] from \FirstMon;
\node[anchor=north] at (axis description cs: 0.25,  0.95) {\fontsize{5}{4}\selectfont $\begin{aligned} d_z&=1\\ n&=500 \end{aligned}$};
\addlegendentry{LSW-S};
\addlegendentry{LSW-L};
\addlegendentry{C-OS};
\addlegendentry{FS-C3};
\addlegendentry{FS-C5};
\addlegendentry{FS-C7};
\nextgroupplot
\node[anchor=north] at (axis description cs: 0.25,  0.95) {\fontsize{5}{4}\selectfont $\begin{aligned} d_z&=1\\ n&=750 \end{aligned}$};
\addplot[smooth,tension=0.5,color=NavyBlue, no markers,line width=0.25pt, densely dotted,forget plot] table[x = delta,y=alpha] from \FirstMon;
\addplot table[x = delta,y=SevenUn] from \MainFirstLSWMon;
\addplot table[x = delta,y=SevenOp] from \MainFirstLSWMon;
\addplot table[x = delta,y=Seven1] from \MainChetverikov;
\addplot table[x = delta,y=MonSevenKn3] from \FirstMon;
\addplot table[x = delta,y=MonSevenKn5] from \FirstMon;
\addplot table[x = delta,y=MonSevenKn7] from \FirstMon;
\nextgroupplot
\node[anchor=north] at (axis description cs: 0.25,  0.95) {\fontsize{5}{4}\selectfont $\begin{aligned} d_z&=1\\ n&=1000 \end{aligned}$};
\addplot[smooth,tension=0.5,color=NavyBlue, no markers,line width=0.25pt, densely dotted,forget plot] table[x = delta,y=alpha] from \FirstMon;
\addplot table[x = delta,y=TenUn] from \MainFirstLSWMon;
\addplot table[x = delta,y=TenOp] from \MainFirstLSWMon;
\addplot table[x = delta,y=Ten1] from \MainChetverikov;
\addplot table[x = delta,y=MonTenKn3] from \FirstMon;
\addplot table[x = delta,y=MonTenKn5] from \FirstMon;
\addplot table[x = delta,y=MonTenKn7] from \FirstMon;
\nextgroupplot[legend style = {column sep = 3.5pt, legend to name = LegendMon2}]
\node[anchor=north] at (axis description cs: 0.25,  0.95) {\fontsize{5}{4}\selectfont $\begin{aligned} d_z&=2\\ n&=500 \end{aligned}$};
\addplot[smooth,tension=0.5,color=NavyBlue, no markers,line width=0.25pt, densely dotted,forget plot] table[x = delta,y=alpha] from \FirstMon;
\addplot table[x = delta,y=FiveUn] from \MainSecondLSWMon;
\addplot table[x = delta,y=FiveOp] from \MainSecondLSWMon;
\addplot table[x = delta,y=Five2] from \MainChetverikov;
\addplot table[x = delta,y=MonFiveQKn0] from \SecondMonQ;
\addplot table[x = delta,y=MonFiveQKn1] from \SecondMonQ;
\addplot table[x = delta,y=MonFiveCKn0] from \SecondMonC;
\addplot table[x = delta,y=MonFiveCKn1] from \SecondMonC;
\addlegendentry{LSW-S};
\addlegendentry{LSW-L};
\addlegendentry{C-OS};
\addlegendentry{FS-Q0};
\addlegendentry{FS-Q1};
\addlegendentry{FS-C0};
\addlegendentry{FS-C1};
\nextgroupplot
\node[anchor=north] at (axis description cs: 0.25,  0.95) {\fontsize{5}{4}\selectfont $\begin{aligned} d_z&=2\\ n&=750 \end{aligned}$};
\addplot[smooth,tension=0.5,color=NavyBlue, no markers,line width=0.25pt, densely dotted,forget plot] table[x = delta,y=alpha] from \FirstMon;
\addplot table[x = delta,y=SevenUn] from \MainSecondLSWMon;
\addplot table[x = delta,y=SevenOp] from \MainSecondLSWMon;
\addplot table[x = delta,y=Seven2] from \MainChetverikov;
\addplot table[x = delta,y=MonSevenQKn0] from \SecondMonQ;
\addplot table[x = delta,y=MonSevenQKn1] from \SecondMonQ;
\addplot table[x = delta,y=MonSevenCKn0] from \SecondMonC;
\addplot table[x = delta,y=MonSevenCKn1] from \SecondMonC;
\nextgroupplot
\node[anchor=north] at (axis description cs: 0.25,  0.95) {\fontsize{5}{4}\selectfont $\begin{aligned} d_z&=2\\ n&=1000 \end{aligned}$};
\addplot[smooth,tension=0.5,color=NavyBlue, no markers,line width=0.25pt, densely dotted,forget plot] table[x = delta,y=alpha] from \FirstMon;
\addplot table[x = delta,y=TenUn] from \MainSecondLSWMon;
\addplot table[x = delta,y=TenOp] from \MainSecondLSWMon;
\addplot table[x = delta,y=Ten2] from \MainChetverikov;
\addplot table[x = delta,y=MonTenQKn0] from \SecondMonQ;
\addplot table[x = delta,y=MonTenQKn1] from \SecondMonQ;
\addplot table[x = delta,y=MonTenCKn0] from \SecondMonC;
\addplot table[x = delta,y=MonTenCKn1] from \SecondMonC;
\end{groupplot}
\node at ($(myplots c2r1) + (0,-2.25cm)$) {\ref{LegendMon1}};
\node at ($(myplots c2r2) + (0,-2.25cm)$) {\ref{LegendMon2}};
\end{tikzpicture}
\caption{Empirical power of our test (with $\gamma_n=0.01/\log n$), LSW-L, LSW-S, and C-OS for the designs \eqref{Eqn: MC1,aux1} and \eqref{Eqn: MC2,aux1}, where corresponding to $\delta=0$ are the empirical sizes under D1. Note that FS-Q1 and FS-C1 nearly overlap each other in the second row.} \label{Fig: MC}
\end{figure}



{
\setlength{\tabcolsep}{5.25pt}
\begin{table}[!ht]
\caption{Run-times (in Seconds) of Monotonicity Tests} \label{Tab: MC, runtime}
\centering\footnotesize
\sisetup{table-number-alignment = center, table-format = 1.2}
\begin{tabularx}{\linewidth}{@{}cc *{5}{S[round-mode = places,round-precision = 2]} c  *{5}{S[round-mode = places,round-precision = 2]} c@{}}
\hline
\hline
&\multirow{2}{*}{$n$}& \multicolumn{5}{c}{Design \eqref{Eqn: MC1,aux1}} & & \multicolumn{5}{c}{Design \eqref{Eqn: MC2,aux1}} & \\
\cline{3-7} \cline{9-13}
&& {FS-C3} & {FS-C7} & {LSW-L} & {LSW-S} & {C-OS} & & {FS-Q0} & {FS-C1} & {LSW-L} & {LSW-S} & {C-OS} &\\
\hline
&$500$        & 0.1236 & 0.1461 & 22.2861  & 22.2216  & 0.0822 & & 0.3984 & 0.4488 & 38.8922  & 39.2504  & 7.1872  &\\
&$750$        & 0.1280 & 0.1310 & 56.8640  & 56.6187  & 0.1566 & & 0.4126 & 0.5008 & 128.8489 & 128.1390 & 17.3866 &\\
&$1000$       & 0.1390 & 0.1558 & 103.5839 & 102.1905 & 0.3057 & & 0.3893 & 0.4221 & 219.2428 & 222.7889 & 29.4911 &\\
\hline
\hline
\end{tabularx}
\end{table}
}

Finally, we compare the run-times of a single replication based on the design D1 in both univariate and bivariate cases. For brevity, we only report our tests with the smallest and the largest $k_n$, based on $\gamma_n=0.01/\log n$. All numbers in Table \ref{Tab: MC, runtime} are obtained by running MATLAB R2019b in a Windows 10 PC with 16 GB RAM and an Intel\textsuperscript{\textregistered} Core\textsuperscript{\tiny TM} i7-7700 processor having 4 cores and 3.60 GHz base speed. Overall, Table \ref{Tab: MC, runtime} shows that our tests are relatively simple to implement, in both univariate and bivariate settings, and their computation cost increases only modestly as we move from $d_z=1$ to $d_z=2$. We stress that the computational complexity of quadratic programs involved in our tests depends on the fineness of discretization, not the sample size.

\section{Conclusion}\label{Sec: conclusion}

In this paper, we have developed a uniformly valid and asymptotically nonconservative test for a class of shape restrictions, which is applicable to nonparametric regression models as well as parametric, distributional, and structural settings. The key insight we exploit is that these restrictions form convex cones in Hilbert spaces, a structure that enables us to employ a projection-based test whose properties may be analyzed in an elegant, transparent, and unifying way. In particular, while the problem is inherently nonstandard, we are able to develop a bootstrap procedure that may be implemented in a way as simple as computing the test statistic.




\titleformat{\section}{\Large\center}{{\sc Appendix} \Alph{section}}{1em}{}
\setcounter{section}{0}
\numberwithin{equation}{section}
\numberwithin{ass}{section}
\numberwithin{figure}{section}
\numberwithin{table}{section}

\section{Proofs of Main Results}\label{Sec: main results proof}


\noindent{\sc Proof of Lemma \ref{Lem: main}:} Part (i) is immediate because $h\mapsto\Pi_\Lambda(h)$ is positively homogeneous by Assumption \ref{Ass: convex setup} and Theorem 5.6-(7) in \citet{Deutsch2012Best}, while part (ii) follows by Lemma \ref{Lem: projection monotonicity}, part (i), $\theta_P\in\Lambda$, and $\kappa_n\in[0,r_n]$. For part (iii), note that $\kappa_n\hat\theta_n-\kappa_n\theta_P=r_n \{\hat\theta_n-\theta_P\}\cdot (\kappa_n/r_n)$. By Assumption \ref{Ass: strong approx}(i), $\sup_{P\in\mathbf P}E[\|\mathbb Z_{n,P}\|_{\mathbf H}]<\infty$ uniformly in $n$ and Markov's inequality, we have: uniformly in $P\in\mathbf P$,
\begin{align}\label{Eqn: main lemma, aux2}
\| r_n \{\hat\theta_n-\theta_P\}\|_{\mathbf H}\le \| r_n \{\hat\theta_n-\theta_P\}-\mathbb Z_{n,P}\|_{\mathbf H}+\|\mathbb Z_{n,P}\|_{\mathbf H}=O_p(1)~.
\end{align}
Part (iii) now follows from combining \eqref{Eqn: main lemma, aux2} and  $\kappa_n/r_n=o(c_n)$. \qed


\noindent{\sc Proof of Theorem \ref{Thm: size control}:} We structure our proof in four steps.

\noindent\underline{Step 1:} Build up a strong approximation for $ r_n\phi(\hat\theta_n)$ that is valid uniformly in $P\in\mathbf P$.

We make use of the map $\psi_{a,P}$ defined in \eqref{Eqn: psi function}, and let $\mathbb G_{n,P}\equiv r_n\{\hat\theta_n-\theta_P\}$. By Lemma \ref{Lem: main} and the definition of $\psi_{a,P}$, we may rewrite the test statistic:
\begin{align}\label{Eqn: size control, aux1}
r_n\phi(\hat\theta_n)=\|\mathbb G_{n,P}+r_n\theta_P-\Pi_\Lambda(\mathbb G_{n,P}+r_n\theta_P)\|_{\mathbf H}=\psi_{r_n,P}(\mathbb G_{n,P})~.
\end{align}
By Theorem 3.16 in \citet{AliprantisandBorder2006} and Assumption \ref{Ass: strong approx}(i), we have:
\begin{align}\label{Eqn: size control, aux2}
|\psi_{r_n,P}(\mathbb G_{n,P})-\psi_{r_n,P}(\mathbb Z_{n,P})|\le \|\mathbb G_{n,P}- \mathbb Z_{n,P}\|_{\mathbf H}=o_p(c_n)~,
\end{align}
uniformly in $P\in\mathbf P$. Combining results \eqref{Eqn: size control, aux1} and \eqref{Eqn: size control, aux2}, we thus obtain the strong approximation for our test statistic: uniformly in $P\in\mathbf P$,
\begin{align}\label{Eqn: size control, aux3}
r_n\phi(\hat\theta_n)=\psi_{r_n,P}(\mathbb Z_{n,P})+o_p(c_n)~.
\end{align}


\noindent\underline{Step 2:} Build up a strong approximation of $ \hat\psi_{\kappa_n}(\hat{\mathbb G}_n)$ that is valid uniformly in $P\in\mathbf P_0$.

First, Theorem 3.16 in \citet{AliprantisandBorder2006} implies: for each $P\in\mathbf P_0$,
\begin{align}\label{Eqn: size control, aux5}
|\hat\psi_{\kappa_n}(\hat{\mathbb G}_n)- \psi_{\kappa_n,P}(\hat{\mathbb G}_n)|
\le \kappa_n\|\Pi_\Lambda\hat\theta_n-\theta_P\|_{\mathbf H}\le  \| \kappa_n \{\hat\theta_n-\theta_P\}\|_{\mathbf H} = o_p(c_n)~,
\end{align}
where the second inequality is due to $\theta_P=\Pi_\Lambda(\theta_P)$ for all $P\in\mathbf P_0$, Assumption \ref{Ass: convex setup}(i), Lemma 6.54-d in \citet{AliprantisandBorder2006}, and the final step is due to Lemma \ref{Lem: main}(iii). Again by Theorem 3.16 in \citet{AliprantisandBorder2006}, we have
\begin{align}\label{Eqn: size control, aux8}
| \psi_{\kappa_n,P}(\hat{\mathbb G}_n)-\psi_{\kappa_n,P}(\bar{\mathbb Z}_{n,P})|\le \|\hat{\mathbb G}_n- \bar{\mathbb Z}_{n,P}\|_{\mathbf H}=o_p(c_n)~,
\end{align}
uniformly in $P\in\mathbf P$, where the equality is due to Assumption \ref{Ass: strong approx}(ii). We thus obtain by \eqref{Eqn: size control, aux5} and \eqref{Eqn: size control, aux8} and the triangle inequality that: uniformly in $P\in\mathbf P_0$,
\begin{align}\label{Eqn: size control, aux9}
\hat\psi_{\kappa_n}(\hat{\mathbb G}_n) = \psi_{\kappa_n,P}(\bar{\mathbb Z}_{n,P}) +o_p(c_n)~.
\end{align}


\noindent\underline{Step 3:} Control the estimation error of $\hat c_{n,1-\alpha}$. By results \eqref{Eqn: size control, aux3} and \eqref{Eqn: size control, aux9}, we may select a sequence of positive scalars $\epsilon_n=o(c_n)$ (sufficiently slow) such that, as $n\to\infty$,
\begin{gather}
\sup_{P\in\mathbf P_0} P(|r_n\phi(\hat\theta_n)-\psi_{r_n,P}(\mathbb Z_{n,P})|>\epsilon_n)=o(1)~, \label{Eqn: size control, aux11a}\\
\sup_{P\in\mathbf P_0} P(|\hat\psi_{\kappa_n}(\hat{\mathbb G}_n) - \psi_{\kappa_n,P}(\bar{\mathbb Z}_{n,P})|>\epsilon_n)=o(1)~.\label{Eqn: size control, aux11b}
\end{gather}
By Markov's inequality, Fubini's theorem, and \eqref{Eqn: size control, aux11b}, we have: for each $\eta>0$,
\begin{multline}\label{Eqn: size control, aux12}
\sup_{P\in\mathbf P_0} P(P(|\hat\psi_{\kappa_n}(\hat{\mathbb G}_n) - \psi_{\kappa_n,P}(\bar{\mathbb Z}_{n,P})|>\epsilon_n|\{X_i\}_{i=1}^n)>\eta)\\
\le \sup_{P\in\mathbf P_0}\frac{1}{\eta} P(|\hat\psi_{\kappa_n}(\hat{\mathbb G}_n) - \psi_{\kappa_n,P}(\bar{\mathbb Z}_{n,P})|>\epsilon_n)=o(1)~.
\end{multline}
Thus, we may select a sequence of positive scalars $\eta_n\downarrow 0$ such that
\begin{align}\label{Eqn: size control, aux13}
P(|\hat\psi_{\kappa_n}(\hat{\mathbb G}_n) - \psi_{\kappa_n,P}(\bar{\mathbb Z}_{n,P})|>\epsilon_n|\{X_i\}_{i=1}^n)=o_p(\eta_n)~,
\end{align}
uniformly in $P\in\mathbf P_0$. Since $\bar{\mathbb Z}_{n,P}$ is independent of $\{X_i\}_{i=1}^n$ by Assumption \ref{Ass: strong approx}(ii), the conditional cdf of $\psi_{\kappa_n,P}(\bar{\mathbb Z}_{n,P})$ given $\{X_i\}_{i=1}^n$ is precisely its unconditional analog. Thus, we may conclude by Lemma 11 in \citet{ChernozhukovLeeRosen2013Intersection} and result \eqref{Eqn: size control, aux13} that
\begin{align}\label{Eqn: size control, aux14}
\liminf_{n\to\infty}\inf_{P\in\mathbf P_0} P(\hat c_{n,1-\alpha}+\epsilon_n\ge c_{n,P}(1-\alpha-\eta_n))=1~.
\end{align}


\noindent\underline{Step 4:} Conclude with the help of a partial anti-concentration inequality.

To begin with, note that by results \eqref{Eqn: size control, aux11a} and \eqref{Eqn: size control, aux14}, we have
\begin{align}\label{Eqn: size control, aux15}
\limsup_{n\to\infty}\sup_{P\in\mathbf P_0}& P(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha})\notag\\
&\le \limsup_{n\to\infty} \sup_{P\in\mathbf P_0}P(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha},|r_n\phi(\hat\theta_n)-\psi_{r_n,P}(\mathbb Z_{n,P})|\le \epsilon_n)\notag\\
&\le \limsup_{n\to\infty}\sup_{P\in\mathbf P_0}P(\psi_{r_n,P}(\mathbb Z_{n,P})>c_{n,P}(1-\alpha-\eta_n)-2\epsilon_n)\notag\\
&\le \limsup_{n\to\infty}\sup_{P\in\mathbf P_0}P(\psi_{\kappa_n,P}(\mathbb Z_{n,P})>c_{n,P}(1-\alpha-\eta_n)-2\epsilon_n)~,
\end{align}
where the final step follows by Lemma \ref{Lem: projection monotonicity} and $0\le\kappa_n\le r_n$ for all large $n$ (due to $\kappa_n/r_n=o(c_n)$ and $c_n=O(1)$). In turn, we note that, since $\eta_n\downarrow 0$ and $\epsilon_n=o(c_n)$, it follows by Assumption \ref{Ass: Gaussian}(iii), Proposition \ref{Pro: anti concentration}, and result \eqref{Eqn: size control, aux15} that
\begin{multline}\label{Eqn: size control, aux18}
\limsup_{n\to\infty}\sup_{P\in\mathbf P_0} P(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha})\le \limsup_{n\to\infty}\sup_{P\in\mathbf P_0}P(\psi_{\kappa_n,P}(\mathbb Z_{n,P})>c_{n,P}(1-\alpha-\eta_n))\\
\le\limsup_{n\to\infty}\sup_{P\in\mathbf P_0}\{\alpha+\eta_n\}=\alpha~,
\end{multline}
as desired for the first claim. For the second claim, it thus suffices to show
\begin{align}\label{Eqn: size control, aux19}
\liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0} P(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha})\ge \alpha~.
\end{align}
For this, we note that, by simple manipulations,
\begin{align}\label{Eqn: size control, aux20}
\liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0}& P(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha})\notag\\
&\ge \liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0}P(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha},|r_n\phi(\hat\theta_n)-\psi_{r_n,P}(\mathbb Z_{n,P})|\le \epsilon_n)\notag\\
& = \liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0}P(\psi_{r_n,P}(\mathbb Z_{n,P})-\epsilon_n>\hat c_{n,1-\alpha})~,
\end{align}
where the last step is due to result \eqref{Eqn: size control, aux11a} and $\bar{\mathbf P}_0\subset \mathbf P_0$. Moreover, another application of Lemma 11 in \citet{ChernozhukovLeeRosen2013Intersection} to \eqref{Eqn: size control, aux13} yields
\begin{align}\label{Eqn: size control, aux21}
\liminf_{n\to\infty}\inf_{P\in\mathbf P_0} P(\hat c_{n,1-\alpha}\le c_{n,P}(1-\alpha+\eta_n)+\epsilon_n)=1~.
\end{align}
By the definition of $\bar{\mathbf P}_0$ and Lemma \ref{Lem: nonconservativeness}, we also note that, for all $n$ and $P\in\bar{\mathbf P}_0$,
\begin{align}\label{Eqn: size control, aux22}
\psi_{r_n,P}(\mathbb Z_{n,P}) = \psi_{\kappa_n,P}(\mathbb Z_{n,P})=\|\mathbb Z_{n,P}-\Pi_\Lambda(\mathbb Z_{n,P})\|_{\mathbf H}~.
\end{align}
Combining results \eqref{Eqn: size control, aux20}, \eqref{Eqn: size control, aux21}, and \eqref{Eqn: size control, aux22} with Proposition \ref{Pro: anti concentration} yields
\begin{align}\label{Eqn: size control, aux23}
\liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0}& P(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha})\notag\\
&\ge \liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0}P(\psi_{r_n,P}(\mathbb Z_{n,P})-\epsilon_n>c_{n,P}(1-\alpha+\eta_n)+\epsilon_n)\notag\\
& = \liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0}P(\psi_{\kappa_n,P}(\mathbb Z_{n,P})>c_{n,P}(1-\alpha+\eta_n)+2\epsilon_n)\notag\\
& = \liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0}P(\psi_{\kappa_n,P}(\mathbb Z_{n,P})>c_{n,P}(1-\alpha+\eta_n)-2\epsilon_n)~.
\end{align}
Since $c_{n,P}(1-\alpha+\eta_n)-\epsilon_n< c_{n,P}(1-\alpha+\eta_n)$, we thus obtain by result \eqref{Eqn: size control, aux23}, the definition of quantiles, and $\eta_n=o(1)$ that
\begin{align}\label{Eqn: size control, aux24}
\liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0} P(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha})\ge \liminf_{n\to\infty}\{\alpha-\eta_n\}=\alpha~,
\end{align}
which, together with the first claim, establishes the second claim of the theorem. \qed


\noindent{\sc Proof of Theorem \ref{Thm: uniform power}:} First, by Assumption \ref{Ass: convex setup} and Lemma \ref{Lem: projection monotonicity}, we have
\begin{align}\label{Eqn: uniform power, aux1}
\|\hat{\mathbb G}_n+\kappa_n\Pi_\Lambda\hat\theta_n &-\Pi_\Lambda(\hat{\mathbb G}_n+\kappa_n\Pi_\Lambda\hat\theta_n)\|_{\mathbf H} \le  \|\hat{\mathbb G}_n-\Pi_\Lambda\hat{\mathbb G}_n\|_{\mathbf H} \notag\\
&\le \|\hat{\mathbb G}_n\|_{\mathbf H}\le \|\hat{\mathbb G}_n-\bar{\mathbb Z}_{n,P}\|_{\mathbf H}+\|\bar{\mathbb Z}_{n,P}\|_{\mathbf H}~,
\end{align}
where the second inequality follows from Assumption \ref{Ass: convex setup} and Theorem 5.6(5) in \citet{Deutsch2012Best}, and the third inequality is due to the triangle inequality. By Assumptions \ref{Ass: strong approx} and \ref{Ass: Gaussian}(ii), we in turn have from \eqref{Eqn: uniform power, aux1} that, uniformly in $P\in\mathbf P$,
\begin{align}\label{Eqn: uniform power, aux2}
\|\hat{\mathbb G}_n+\kappa_n\Pi_\Lambda\hat\theta_n-\Pi_\Lambda(\hat{\mathbb G}_n+\kappa_n\Pi_\Lambda\hat\theta_n)\|_{\mathbf H}
=o_p(c_n)+O_p(1)=O_p(1)~.
\end{align}
By the definition of $\hat c_{n,1-\alpha}$, we note that, for $M>0$ and uniformly in $P\in\mathbf P$,
\begin{align}\label{Eqn: uniform power, aux3}
P(\hat c_{n,1-\alpha}>M)&\le P(P(\|\hat{\mathbb G}_n+\kappa_n\Pi_\Lambda\hat\theta_n-\Pi_\Lambda(\hat{\mathbb G}_n+\kappa_n\Pi_\Lambda\hat\theta_n)\|_{\mathbf H}>M|\{X_i\}_{i=1}^n)>\alpha)\notag\\
&\le\frac{1}{\alpha} P(\|\hat{\mathbb G}_n+\kappa_n\Pi_\Lambda\hat\theta_n-\Pi_\Lambda(\hat{\mathbb G}_n+\kappa_n\Pi_\Lambda\hat\theta_n)\|_{\mathbf H}>M) ~,
\end{align}
where the second inequality holds by Markov's inequality and Fubini's theorem. It follows from results \eqref{Eqn: uniform power, aux2} and \eqref{Eqn: uniform power, aux3} that $\hat c_{n,1-\alpha}=O_p(1)$ uniformly in $P\in\mathbf P$.


Next, we bound $r_n\phi(\hat\theta_n)$ from below. By Theorem 3.16 in \citet{AliprantisandBorder2006} and the triangle inequality, we have: uniformly in $P\in\mathbf P$,
\begin{multline}\label{Eqn: uniform power, aux5}
|r_n\phi(\hat\theta_n)-r_n\phi(\theta_P)|\le \|r_n\{\hat\theta_n-\theta_P\}\|_{\mathbf H}\\
\le \|r_n\{\hat\theta_n-\theta_P\}-\mathbb Z_{n,P}\|_{\mathbf H}+ \|\mathbb Z_{n,P}\|_{\mathbf H}\le o_p(c_n)+O_p(1)=O_p(1)~,
\end{multline}
where the third inequality follows by Assumptions \ref{Ass: strong approx}(i) and \ref{Ass: Gaussian}(ii), and the last step is due to $c_n=O(1)$. It follows from result \eqref{Eqn: uniform power, aux5} and the definition of $\mathbf P_{1,n}^\Delta$ that
\begin{align}\label{Eqn: uniform power, aux6}
r_n\phi(\hat\theta_n)=r_n\phi(\theta_P)+r_n\phi(\hat\theta_n)-r_n\phi(\theta_P)\ge \Delta+O_p(1)~,
\end{align}
uniformly in $P\in\mathbf P_{1,n}^\Delta$. The theorem thus follows from combining result \eqref{Eqn: uniform power, aux6} and the order $\hat c_{n,1-\alpha}=O_p(1)$ uniformly in $P\in\mathbf P$ that we have established .\qed

\titleformat{\section}{\normalfont\Large\bfseries}{\Alph{section}}{1em}{}

\addcontentsline{toc}{section}{References}
\putbib
\end{bibunit}


\clearpage\newpage

\begin{bibunit}