EconBase
← Back to paper

A Projection Framework for Testing Shape Restrictions That Form Convex Cones

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

91,509 characters · 0 sections · 93 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

A Projection Framework for Testing Shape Restrictions That Form Convex Cones

bibunit\pdfbookmark[1]{Title}{title} \begin{abstract} This paper develops a uniformly valid and asymptotically nonconservative test based on projection for a class of shape restrictions. The key insight we exploit is that these restrictions form convex cones, a simple and yet elegant structure that has been barely harnessed in the literature. Based on a monotonicity property afforded by such a geometric structure, we construct a bootstrap procedure that, unlike many studies in nonstandard settings, dispenses with estimation of local parameter spaces, and the critical values are obtained in a way as simple as computing the test statistic. Moreover, by appealing to strong approximations, our framework accommodates nonparametric regression models as well as distributional/density-related and structural settings. Since the test entails a tuning parameter (due to the nonstandard nature of the problem), we propose a data-driven choice and prove its validity. Monte Carlo simulations confirm that our test works well. \end{abstract} \begin{center} Keywords: Nonstandard inference, shape restrictions, convex cone, projection, strong approximations \end{center} \section{Introduction} Shape restrictions play a number of fundamental and prominent roles in economics. For example, they often arise as testable implications of economic theory, and may thus serve as plausible restrictions in specifying economic models Varian1982Demand,Varian1984Production. In complementary empirical work, they help achieve point identification, tighten identification bounds, improve estimation precision, and develop powerful tests---see Matzkin1994Handbook and ChetverikovSantosAzeem2018Shape for more detailed discussions. In this paper, we develop a uniformly valid and asymptotically nonconservative test for a class of shape restrictions. To illustrate, consider the regression model: \begin{align} Y = \theta_0(Z)+u , \end{align} where $Z\in [0,1]$, $\theta_0:[0,1]\to\mathbf R$, and $E[u|Z]=0$. Suppose that we are interested in testing whether $\theta_0$ is nondecreasing. More formally, if $\theta_0\in \mathbf H\equiv \{\theta:[0,1]\to\mathbf R: \|\theta\|_{\mathbf H}\equiv \{\int_{[0,1]} |\theta(z)|^2\,\mathrm dz\}^{1/2}<\infty\}$ (a mild restriction) and $\Lambda$ is the class of nondecreasing functions in $\mathbf H$, then we may formulate the hypotheses as \begin{align} \mathrm H_0: \phi(\theta_0)=0 \qquad\qquad \mathrm{vs.}\qquad\qquad \mathrm H_1: \phi(\theta_0)>0 , \end{align} where $\phi(\theta)\equiv\min_{\lambda\in\Lambda}\|\theta-\lambda\|_{\mathbf H}$ is the distance from $\theta$ to $\Lambda$. Thus, given an unconstrained (kernel or sieve) estimator $\hat\theta_n$ of $\theta_0$, we may employ a test that rejects $\mathrm H_0$ if $r_n\phi(\hat\theta_n)$ is “large” for a suitable $r_n\to\infty$. While conceptually intuitive, construction of critical values turns out to be a delicate and challenging matter. In particular, despite the well-established results on the rate $r_n$ and pointwise asymptotic normality, $r_n\{\hat\theta_n-\theta_0\}$ generically does not converge as a process indexed by $[0,1]$ ChernozhukovLeeRosen2013Intersection, rendering the Delta method as in FangSantos2018HDD inapplicable. As a first step, we note that $\Lambda$ being a (closed) convex cone\footnote{By definition, $\Lambda$ is a convex cone if and only if $af+bg\in\Lambda$ whenever $a,b\ge 0$ and $f,g\in\Lambda$.} implies \begin{align} r_n\phi(\hat\theta_n)=\|r_n\{\hat\theta_n-\theta_0\}+r_n\theta_0-\Pi_\Lambda(r_n\{\hat\theta_n-\theta_0\}+r_n\theta_0)\|_{\mathbf H} , \end{align} where $\theta\mapsto\Pi_{\Lambda}(\theta)\equiv\operatornamewithlimits{arg\,min}_{\lambda\in\Lambda}\|\theta-\lambda\|_{\mathbf H}$ is the projection operator, i.e., $\Pi_{\Lambda}(\theta)$ is the closest (under $\|\cdot\|_{\mathbf H}$) nondecreasing function to $\theta$. Hence, (ref) reveals that, in estimating the law of $r_n\phi(\hat\theta_n)$ in order to obtain critical values, it suffices to quantify the variation of $r_n\{\hat\theta_n-\theta_0\}$ and estimate the drift $r_n\theta_0$. Despite the lack of convergence (in $\mathbf H$), $r_n\{\hat\theta_n-\theta_0\}$ as a process may be approximated in law via strong approximations by a number of methods with only mild computation cost, including the simulation method in ChernozhukovLeeRosen2013Intersection, the weighted bootstrap in BelloniChernozhukovChetverikovKato2015New, and the sieve score bootstrap in ChenChristensen2018SupNormOptimal. The treatment of $r_n\theta_0$, on the other hand, consists of the nonstandard step because, as in moment inequality models AndrewsSoares2010, $r_n\theta_0$ cannot be consistently estimated in general. In this regard, the convex cone property (but not convexity alone) implies \begin{align} r_n\phi(\hat\theta_n)\le \|r_n\{\hat\theta_n-\theta_0\}+\kappa_n\theta_0-\Pi_\Lambda(r_n\{\hat\theta_n-\theta_0\}+\kappa_n\theta_0)\|_{\mathbf H} , \end{align} whenever $0\le \kappa_n\le r_n$. Hence, a valid critical value may be obtained by bootstrapping the upper bound in (ref), which is possible because $\kappa_n\theta_0$ may be consistently estimated by $\kappa_n\hat\theta_n$ provided $\kappa_n/r_n\to 0$. The very virtue that (ref) is an inequality rather than equality may also raise concern for conservativeness of the resulting test. As an extreme case, the choice $\kappa_n=0$ leads to a least favorable test that may be viewed as assuming $\theta_0=0$, suggesting that $\kappa_n$ should not be too small. Along these lines, we develop a data-driven choice of $\kappa_n$ that delivers an asymptotically nonconservative test. While we started with monotonicity in the univariate model (ref), the main features of our test are not confined to this special problem. First, the convex cone property is in fact shared by a class of restrictions, e.g., nonnegativity, concavity, Slutsky restriction and supermodularity---see Appendix (ref). Regrettably, our framework is not directly applicable absent the property, e.g., quasi-concavity, log-concavity, and $r$-concavity Kostyshak2017U,KomarovaHidalgo2019Shape. Second, our framework is applicable to distributional and structural settings PinkseSchurter2019Auction,Bhattacharya2020Empirical---see Appendix (ref). Third, our framework allows for jointly testing multiple restrictions since intersections of convex cones remain convex cones. This is important because it is common that shape restrictions arise simultaneously. Fourth, while some convex cone restrictions (e.g., monotonicity and concavity) on $\theta_0$ can be characterized as inequalities on derivatives of $\theta_0$, we dispense with derivative estimation because it likely incurs power loss due to a slower convergence rate Stone1982Global,ChenChristensen2018SupNormOptimal. Finally, our test does not rely on the least favorable configurations, while being asymptotically nonconservative (thus leading to improved power) and computationally tractable. We stress that, since the norm $\|\cdot\|_{\mathbf H}$ is of $L^2$ nature, our test cannot be inverted to obtain uniform confidence bands. In this sense, our paper complements HorowitzLee2017Shape, FreybergerReeves2017Shape and ChenChernozhukovFernandezKostyshakLuo2021Shape. The literature on shape restrictions was initiated in the 1950s BBBB1972, with much attention since focused on estimation under solely shape restrictions---see, e.g., HanWangChatterjeeSamworth2019Isotonic and references therein. There are also post-processing methods that enforce restrictions on unconstrained estimators---see ChenChernozhukovFernandezKostyshakLuo2021Shape who studied a number of shape enforcing operators. The projection method in this paper is of post-processing nature. The convex cone structure was recognized in the 1960s as a device to generalize monotone regression, though the focus is on {\it analytic} properties of projections BBBB1972. For testing, the structure has barely been exploited beyond identifying the least favorable distributions in parametric settings Wolak1987Exact,SilvapulleSen2005Constrained, deriving minimax bounds in univariate Gaussian white noise models JuditskyNemirovski2002Shape, and establishing minimaxity for the likelihood ratio test in Gaussian sequence models WeiWainwrightGuntuboyina2019Geometry. Much of the testing literature has been developed by exploiting {\it particular} structures of restrictions, with concavity and especially monotonicity in univariate models being the primary focuses. Despite a sizable literature, existing tests prevalently rely on the least favorable configurations, including GijbelsHallJonesKoch2000Monotonicity, GhosalSenVaart2000Monotonicity, HallHeckman2000Monotonicity, Durot2003testing and Gutknecht2016MonotoneTest for monotonicity, and AbrevayaJiang2005Curvature for concavity---see also DumbgenSpokoiny2001Multiscale and BaraudHuetLaurent2005Convex who devised tests for both shapes. Chetverikov2018Monotonicity developed, as far as we are aware, the first nonparametric uniformly valid tests designed specifically for monotonicity that avoid the least favorable configurations. There is also a strand of literature motivated by the {\it common} structure of shape restrictions viewed as inequalities, including ChernozhukovNeweySantos2019CCMM, LeeSongWhang2017Inequal, BelloniChernozhukovChetverikovFernandez2019QR and Zhu2020Shape. If the inequalities are based on derivatives, the derivative estimation may then incur power loss. Moreover, all these (uniformly valid) tests, except the least favorable test of BelloniChernozhukovChetverikovFernandez2019QR, involve estimation of local parameter spaces. ChernozhukovNeweySantos2019CCMM and Zhu2020Shape, unlike our setup, allowed for partial identification by working with moment restrictions. By the virtue of their setup, they required estimating the set of minimizers and the use of strong approximations may entail stronger regularity conditions. Similar in spirit to ChernozhukovNeweySantos2019CCMM and Zhu2020Shape, KomarovaHidalgo2019Shape proposed a moment-based test in the univariate model (ref) for shape restrictions that may not form convex cones. The remainder of the paper is structured as follows. Section (ref) introduces the setup and some motivating examples. Section (ref) presents our inferential framework, a data-driven choice of the tuning parameter, and an implementation guide. Section (ref) conducts simulation studies, while Section (ref) concludes. Appendix (ref) contains proofs of the main results. Due to space limitation, we relegate the discussions of the convex cone property, an investigation of the special case when $\hat\theta_n$ admits an asymptotic distribution, and presentations of some auxiliary results to the online supplement. \section{The Setup and Examples} Throughout, we denote by $\{X_i\}_{i=1}^n$ the sample with each $X_i$ living in some sample space $\mathcal X$, and by $P$ the joint law of $\{X_i\}_{i=1}^n$ that belongs to some family $\mathbf P$ of distributions on $\mathcal X^n$. The dependence of $P$ and $\mathbf P$ on $n$ is suppressed for notational simplicity. We stress that, under this configuration, $\{X_i\}_{i=1}^n$ need not be i.i.d. In turn, we let $\theta_0$ be the parameter of interest, and, whenever appropriate, make the dependence of $\theta_0$ on $P$ explicit by instead writing $\theta_P$. In order to accommodate Slutsky restriction, we shall work with an abstract Hilbert space (i.e., a complete inner product space) with inner product $\langle\cdot,\cdot\rangle_{\mathbf H}$ and induced norm $\|\cdot\|_{\mathbf H}$. Given the sample $\{X_i\}_{i=1}^n$, our objective is then to test whether $\theta_0$ satisfies the shape in question, i.e., the hypotheses in (ref). The setup (ref) induces two models: $\mathbf P_0\equiv\{P\in\mathbf P: \phi(\theta_P)=0\}$ and $\mathbf P_1\equiv\mathbf P\backslash\mathbf P_0$. Before proceeding further, we introduce additional notation. Set $\mathbf R_+\equiv\{x\in\mathbf R: x\ge 0\}$ and $L^2(\mathcal Z)\equiv\{f: \mathcal Z\to\mathbf R: \int_{\mathcal Z} |f(z)|^2\,\mathrm dz<\infty\}$ for $\mathcal Z\subset\mathbf R^{d_z}$. Let $\mathbf M^{m\times k}$ be the space of $m\times k$ matrices, and, for $A\in\mathbf M^{m\times k}$, write its transpose by $A^\text{\scalebox{0.8}{$\intercal$}}$, its trace by $\mathrm{tr}(A)$ if $m=k$, its Moore–Penrose inverse by $A^-$, and its Frobenius norm by $\|A\|\equiv\{\mathrm{tr}(A^\text{\scalebox{0.8}{$\intercal$}} A)\}^{1/2}$. For a sequence $\{h_j\}$ of functions, denote the vector $(h_1,\ldots,h_k)^\text{\scalebox{0.8}{$\intercal$}}$ by $h^k$. For generic families of distributions $\mathbf P_n$, a sequence $\{a_n\}$ of positive scalars, and a sequence $\{\mathbb X_n\}$ of random elements in a normed space $\mathbf D$ with norm $\|\cdot\|_{\mathbf D}$, write $\mathbb X_n=o_p(a_n)$ uniformly in $P\in\mathbf P_n$ if $\lim_{n\to\infty}\sup_{P\in\mathbf P_n}P(\|\mathbb X_n\|_{\mathbf D}>a_n\epsilon)= 0$ for any $\epsilon>0$, and $\mathbb X_n=O_p(a_n)$ uniformly in $P\in\mathbf P_n$ if $\lim_{M\to\infty}\limsup_{n\to\infty}\sup_{P\in\mathbf P_n}P(\|\mathbb X_n\|_{\mathbf D}>Ma_n)= 0$. \subsection{Examples} We now present examples where shape restrictions play important roles. The first example is concerned with nonparametric regression models. \begin{ex}[Nonparametric Regression] Let $X\equiv(Y,Z)\in\mathbf R^{1+d_z}$ satisfy \begin{align} Y=\theta_0(Z)+u , \end{align} where $\theta_0: \mathcal Z\subset\mathbf R^{d_z}\to\mathbf R$ with $\mathcal Z$ the support of $Z$, and $E[u|Z]=0$. Here, one may set $\mathbf H=L^2(\mathcal Z)$, let $\Lambda\subset \mathbf H$ consist of (say) monotonic functions, and obtain an unconstrained estimator $\hat\theta_n$ of $\theta_0$ by kernel methods such as local constant/linear/polynomial regression or sieve methods with various basis functions such as splines ChernozhukovLeeRosen2013Intersection. The rate $r_n$ equals $(nh_n^{d_z})^{1/2}$ for kernel estimation with bandwidth $h_n$, and $(n/k_n)^{1/2}$ for sieve estimation based on, e.g., B-splines, with $k_n$ (henceforth) the sieve dimension. One may also set up this example as one on the conditional mean $z\mapsto E[Y|Z=z]$. \qed \end{ex} Our second example generalizes (ref) to its instrumental variable (IV) analog. \begin{ex}[Nonparametric IV Regression] Let $X\equiv(Y,Z,V)\in\mathbf R^{1+d_z+d_v}$ satisfy (ref) but with $E[u|V]=0$. Then one may set $\mathbf H$ and $\Lambda$ as in Example (ref), and employ the series two-stage least square (2SLS) estimation with, e.g., B-splines, resulting in a rate $r_n$ equal to $(n/k_n)^{1/2} s_n$, where $s_n$ is the smallest singular value of $E[b^{m_n}(V)h^{k_n}(Z)^\text{\scalebox{0.8}{$\intercal$}}]$ with $h^{k_n}$ and $b^{m_n}$ respectively $k_n\times 1$ and $m_n\times 1$ vectors (with $m_n\ge k_n$) of B-splines for $\theta_0$ and the instrument space NeweyPowell2003NPIV,BlundellChenKristensen2007Engel,ChenChristensen2018SupNormOptimal. In practice, $s_n$ is unknown but can be replaced with its sample analog. One may conceivably also employ kernel-type estimators HallHorowitz2005NPIV,DarollesFanFlorensRenault2011NPIV, though the strong approximation results appear to be lacking. \qed \end{ex} The third example is concerned with nonparametric quantile regression. \begin{ex}[Nonparametric Quantile Regression] Let $X\equiv(Y,Z)\in\mathbf R^{1+d_z}$ satisfy \begin{align} Y=\theta_0(Z,U) , \end{align} where $\theta_0:\mathbf R^{d_z}\times[0,1]\to\mathbf R$, and $U$ is uniformly distributed on $[0,1]$. If $Z$ and $U$ are independent, then $\theta_0$ may be interpreted as the conditional quantile function of $Y$ given $Z$. Here, one may define $\mathbf H$ and $\Lambda$ as previously, and employ the recent sieve estimator of BelloniChernozhukovChetverikovFernandez2019QR with $r_n=(n/k_n)^{1/2}$ for, e.g., B-splines. \qed \end{ex} Our final example concerns a restriction in a possibly endogenous regression model. \begin{ex}[Slutsky Restriction] Let $Q\in\mathbf R^{d_q}$ (quantities), $P\in\mathbf R^{d_q}$ (prices), $Y\in\mathbf R$ (income), and $Z\in\mathbf R^{d_z}$ (demographics) satisfy the following $d_q$ equations: \begin{align} Q=g_0(P,Y)+\Gamma_0^\scalebox{0.8}{$\intercal$} Z + U , \end{align} where $g_0: \mathbf R_+^{d_q+1}\to\mathbf R^{d_q}$ is differentiable, $\Gamma_0\in\mathbf M^{d_z\times d_q}$, and $U\in\mathbf R^{d_q}$ is the error term. The Slutsky matrix of $g_0$ is the mapping $\theta_0: \mathbf R_+^{d_q+1}\to\mathbf M^{d_q\times d_q}$ defined by: \begin{align} \theta_0(p,y)\equiv \frac{\partial g_0(p,y)}{\partial p^\scalebox{0.8}{$\intercal$}}+ \frac{\partial g_0(p,y)}{\partial y} g_0(p,y)^\scalebox{0.8}{$\intercal$} . \end{align} The Slutsky restriction refers to $\theta_0(p,y)$ being negative semidefinite (nsd) at each pair $(p,y)$. For this example, we follow AguiarSerrano2017Slutsky by letting $\mathbf H$ be the space of functions $\theta: \mathbf R_+^{d_q+1}\to \mathbf M^{d_q\times d_q}$ with inner product $\langle \cdot,\cdot\rangle_{\mathbf H}$ defined by \begin{align} \langle \theta_1,\theta_2\rangle_{\mathbf H}\equiv \int \mathrm{tr}(\theta_1(p,y)^\scalebox{0.8}{$\intercal$}\theta_2(p,y))\, \mathrm dp\mathrm dy ,\,\forall\,\theta_1,\theta_2\in\mathbf H , \end{align} and $\Lambda$ the family of mappings $\theta\in\mathbf H$ such that $\theta(p,y)$ is nsd and symmetric for all $(p,y)$. In turn, the nonlinear functional $\theta_0$ may be estimated based on the plug-in principle and a sieve estimator of $g_0$ DonaldNewey1994Semilinear,BlundellChenKristensen2007Engel, resulting in a rate $r_n$ as in Example (ref) (without endogeneity) or (ref) (with endogeneity). \qed \end{ex} We refer the reader to Appendix (ref) that illustrates how our framework also applies to distributional settings where shape restrictions often take the form of various dominance relations and structural models where shape restrictions may serve as testable implications. For ease of reference, we shall call settings where $\hat\theta_n$ admits an asymptotic distribution (in $\mathbf H$) {\it regular} and ones without this property {\it irregular}. \section{The Inferential Framework} \subsection{The General Framework} We commence with our main assumptions in this paper. \begin{ass} (i) $\Lambda$ is a known nonempty closed convex set in a Hilbert space $\mathbf H$ with known inner product $\langle \cdot,\cdot\rangle_{\mathbf H}$ and induced norm $\|\cdot\|_{\mathbf H}$; (ii) $\Lambda$ is a cone. \end{ass} \begin{ass} (i) An estimator $\hat\theta_n: \{X_i\}_{i=1}^n\to\mathbf H$ satisfies $\|r_n\{\hat\theta_n-\theta_P\}-\mathbb Z_{n,P}\|_{\mathbf H}=o_p(c_n)$ uniformly in $P\in\mathbf P$ for some $r_n\to\infty$, $\mathbb Z_{n,P}\in\mathbf H$, and $c_n>0$ with $c_n=O(1)$; (ii) $\hat{\mathbb G}_n\in\mathbf H$ is a bootstrap estimator satisfying: $\|\hat{\mathbb G}_{n}-\bar{\mathbb Z}_{n,P}\|_{\mathbf H}=o_p(c_n)$ uniformly in $P\in\mathbf P$, for $\bar{\mathbb Z}_{n,P}$ a copy of $\mathbb Z_{n,P}$ that is independent of $\{X_i\}_{i=1}^n$. \end{ass} Assumption (ref) simply abstracts the convex cone feature, where we single out the conic condition for ease of elucidating the roles it plays in this paper. Assumption (ref)(i) requires that $r_n\{\hat\theta_n-\theta_P\}$ be approximated by $\mathbb Z_{n,P}$ (uniformly) at a rate $c_n$. In regular settings, it suffices to have $\sqrt n\{\hat\theta_n-\theta_P\}\xrightarrow{L} \mathbb G_P$ uniformly in $P\in\mathbf P$, so that $r_n=\sqrt n$, and $\mathbb Z_{n,P}$ equals $\mathbb G_P$ in law---see Appendix (ref). In irregular settings such as Example (ref), one may obtain $\hat\theta_n$ by kernel or sieve methods. To illustrate, let $\{Y_i,Z_i\}_{i=1}^n$ be a sample generated by (ref), and $h^{k_n}$ be a $k_n\times 1$ vector of B-splines on $\mathcal Z$. Then the sieve estimator of $\theta_P$ is given by $\hat\theta_n=\hat\beta_n^{\text{\scalebox{0.8}{$\intercal$}}}h^{k_n}$ with $\hat\beta_n= [\sum_{i=1}^{n}h^{k_n}(Z_i)h^{k_n}(Z_i)^{\text{\scalebox{0.8}{$\intercal$}}}]^{-}\sum_{i=1}^{n}h^{k_n}(Z_i)Y_i$. Under regularity conditions, one may obtain the linear expansion in $L^2(\mathcal Z)$: \begin{align} r_n \{\hat\theta_n-\theta_P\}=k_n^{-1/2}(h^{k_n})^\scalebox{0.8}{$\intercal$} (E_P[h^{k_n}(Z)h^{k_n}(Z)^\text{\scalebox{0.8}{$\intercal$}}])^{-1}\frac{1}{\sqrt n}\sum_{i=1}^nh^{k_n}(Z_i)u_i+o_p(\frac{1}{\log n}) , \end{align} uniformly in $P\in\mathbf P$. Here, undersmoothing is required in order to deliver (ref). Following ChernozhukovLeeRosen2013Intersection, one may then verify Assumption (ref)(i) with $r_n=\sqrt{n/k_n}$, $c_n=1/\log n$, and $\mathbb Z_{n,P}= k_n^{-1/2}(h^{k_n})^\text{\scalebox{0.8}{$\intercal$}} (E_P[h^{k_n}(Z)h^{k_n}(Z)^\text{\scalebox{0.8}{$\intercal$}}])^{-1} G_{n,P}$ for some $G_{n,P}\sim N(0,E_P[u^2h^{k_n}(Z)h^{k_n}(Z)^\text{\scalebox{0.8}{$\intercal$}}])$ . The rate $c_n$ serves to cope with potential degeneracy of the test statistic, an issue inherent in the nonstandard nature of the problem. Assumption (ref)(ii) demands an analogous approximation for the bootstrap. In regular settings, typically $\hat{\mathbb G}_n=\sqrt n \{\hat\theta_n^*-\hat\theta_n\}$, where $\hat\theta_n^*$ is the same as $\hat\theta_n$ but based on samples drawn from the original data. In Example (ref), the expansion (ref) suggests \begin{align} \hat{\mathbb G}_n=k_n^{-1/2}(h^{k_n})^\text{\scalebox{0.8}{$\intercal$}} (\frac{1}{n}\sum_{i=1}^nh^{k_n}(Z_i)h^{k_n}(Z_i)^\text{\scalebox{0.8}{$\intercal$}})^{-1}\frac{1}{\sqrt n}\sum_{i=1}^nW_ih^{k_n}(Z_i)\hat u_i . \end{align} where $\hat u_i\equiv Y_i-\hat\theta_n(Z_i)$ and $\{W_i\}_{i=1}^n$ are (scalar) weights (e.g., standard normals). The verification of Assumption (ref) for Example (ref) represents a general strategy: establishing an asymptotic linear expansion of $r_n\{\hat\theta_n-\theta_P\}$, and then verifying Assumption (ref) based on the linear term---see ChernozhukovLeeRosen2013Intersection, ChernozhukovNeweySantos2019CCMM, ChenChristensen2018SupNormOptimal, BelloniChernozhukovChetverikovFernandez2019QR, and LiLiao2020Uniform for more illustrations. We stress that Assumption (ref)(ii) leaves the particular form of $\hat{\mathbb G}_n$ unspecified, and thus accommodates alternative resampling schemes. \begin{figure} \scriptsize \begin{subfigure}[b]{0.45\linewidth} \begin{tikzpicture}[scale=0.8, >=latex',dot/.style={circle,inner sep=1pt,fill,name=#1}] \draw (-4,-2.5)--(4,-2.5) -- (4,3) -- (-4,3) -- (-4,-2.5); \draw[loosely dotted,->] (-4,0)--(4,0); \draw[loosely dotted,->] (0,-2.5)--(0,3); \fill[red] (0,0) circle (0.8pt); \draw (0,0) node[anchor=north west] {$0$}; \fill[fill=Honeydew3,semitransparent] (0,0)--(4/5,6/5)--(4,0)--(0,0); \draw[line width=0.07mm,NavyBlue] (0,0)--(4/5,6/5)--(4,0)--(0,0); \draw (2,1) node[anchor=west] {$\Lambda$}; \node[dot=theta,label={east: $\theta_0$}] at (2/5,3/5) ; \node[dot=h,label={west:$h$}] at (-2.5,-0.5) ; \node (zero) at (-5/3,-2.5) ; \draw (-3.5,-2) -- (-1/6,3); \draw (-1.7,2.3) node[anchor=south, semitransparent,SeaGreen4] {$h+a\theta_0:a\ge 0$}; \draw (-2,-1.5) node[anchor=north, semitransparent,SeaGreen4] {$h+a\theta_0:a\le 0$}; \draw[dashed] (h)--(0,0); \draw[dashed] (-37/18,1/6)--(0,0); \draw[dashed] (-3/2,1)--(0,0); \draw[dashed] (-7/10,11/5)--(4/5,6/5); \draw[dashed] (-1/4,23/8)--(4/5,6/5); \fill[red] (-3/2,1) circle (0.8pt); \draw (-3/2,1) node[anchor=south east] {$h+a^*\theta_0$}; \end{tikzpicture} \caption{$\Lambda$ is convex but not conic} \end{subfigure} \begin{subfigure}[b]{0.45\linewidth} \begin{tikzpicture}[scale=0.8, >=latex',dot/.style={circle,inner sep=1pt,fill,name=#1}] \draw (-4,-2.5)--(-4,3)--(2,3); \draw (-4,-2.5)--(4,-2.5) -- (4,0); \draw[loosely dotted,->] (-4,0)--(4,0); \draw[loosely dotted,->] (0,-2.5)--(0,3); \fill[red] (0,0) circle (0.8pt); \draw (0,0) node[anchor=north west] {$0$}; \fill[fill=Honeydew3,semitransparent] (0,0)--(2,3)--(4,3)--(4,0)--(0,0); \draw[line width=0.07mm,NavyBlue] (0,0)--(2,3); \draw[line width=0.07mm,NavyBlue] (0,0)--(4,0); \draw (2,1) node[anchor=west] {$\Lambda$}; \node[dot=theta,label={east: $\theta_0$}] at (4/5,6/5) ; \node[dot=h,label={west:$h$}] at (-2.5,-0.5) ; \node (zero) at (-5/3,-2.5) ; \draw (-3.5,-2) -- (-1/6,3); \draw (-1.7,2.3) node[anchor=south, semitransparent,SeaGreen4] {$h+a\theta_0:a\ge 0$}; \draw (-2,-1.5) node[anchor=north, semitransparent,SeaGreen4] {$h+a\theta_0:a\le 0$}; \draw[dashed] (h)--(0,0); \draw[dashed] (-37/18,1/6)--(0,0); \draw[dashed] (-3/2,1)--(0,0); \fill[red] (-3/2,1) circle (0.8pt); \draw (-3/2,1) node[anchor=south east] {$h+a^*\theta_0$}; \end{tikzpicture} \caption{$\Lambda$ is convex and conic} \end{subfigure} \caption{$\psi_{a,P}(h)\equiv \|h+a\theta_P-\Pi_\Lambda(h+a\theta_P)\|_{\mathbf H}$ is weakly decreasing in $a\in[0,\infty)$ if $\Lambda$ is convex and conic as in Figure (ref), but may not be so if it is convex but not conic as in Figure (ref).} \end{figure} The next lemma lays out a number of important building blocks for our development. \begin{lem} (i) If Assumption (ref) holds, then any $\hat\theta_n\in\mathbf H$ and $r_n\in\mathbf R_+$ satisfy \begin{align} r_n\phi(\hat\theta_n)=\|r_n\{\hat\theta_n-\theta_P\}+r_n\theta_P-\Pi_\Lambda(r_n\{\hat\theta_n-\theta_P\}+r_n\theta_P)\|_{\mathbf H} . \end{align} (ii) If Assumption (ref) holds, $P\in\mathbf P_0$, and $\kappa_n\in[0,r_n]$, then it follows that \begin{align} r_n\phi(\hat\theta_n)\le \|r_n\{\hat\theta_n-\theta_P\}+\kappa_n\theta_P-\Pi_\Lambda(r_n\{\hat\theta_n-\theta_P\}+\kappa_n\theta_P)\|_{\mathbf H} . \end{align} (iii) If Assumption (ref)(i) holds, $\sup_{P\in\mathbf P} E[\|\mathbb Z_{n,P}\|_{\mathbf H}]<\infty$ uniformly in $n$, and $\kappa_n/r_n=o(c_n)$, then we have that $\kappa_n\hat\theta_n-\kappa_n\theta_P= o_p(c_n)$ uniformly in $P\in\mathbf P$. \end{lem} Lemma (ref)(i) highlights the standard, and critically, the nonstandard features of the problem. Specifically, while $r_n\{\hat\theta_n-\theta_P\}$ may be approximated in law by $\hat{\mathbb G}_n$ due to Assumption (ref), $r_n\theta_P$ cannot be consistently estimated in general, i.e., $r_n\hat\theta_n-r_n\theta_P\neq o_p(1)$. Lemma (ref)(ii) suggests that we may obtain critical values by instead estimating the law of the upper bound in (ref). This is possible because $\kappa_n\theta_P$ can be consistently estimated by $\kappa_n\hat\theta_n$ as shown by Lemma (ref)(iii). We stress that (ref) is implied by the convex cone property but not convexity alone---see Figure (ref). To formalize the above discussions, we define maps $\psi_{a,P},\hat\psi_{\kappa_n}:\mathbf H\to\mathbf R$ by \begin{align} \psi_{a,P}(h)&\equiv \|h+a\theta_P-\Pi_\Lambda(h+a\theta_P)\|_{\mathbf H} ,\\ \hat\psi_a(h)&\equiv \|h+a\Pi_\Lambda\hat\theta_n-\Pi_\Lambda(h+a\Pi_\Lambda\hat\theta_n)\|_{\mathbf H} . \end{align} Thus, $\psi_{\kappa_n,P}(r_n\{\hat\theta_n-\theta_P\})$ is precisely the upper bound in (ref), and $\hat\psi_{\kappa_n}(\hat{\mathbb G}_n)$ is its bootstrap analog in which the null hypothesis is enforced through $\Pi_\Lambda\hat\theta_n$ to improve power. Finally, for a significance level $\alpha\in(0,1)$, we define our critical value \begin{align} \hat c_{n,1-\alpha}\equiv \inf\{c\in\mathbf R: P(\hat\psi_{\kappa_n}(\hat{\mathbb G}_n)\le c|\{X_i\}_{i=1}^n)\ge 1-\alpha\} . \end{align} As known in the literature ChernozhukovNeweySantos2019CCMM, the validity of $\hat c_{n,1-\alpha}$, viewed as a mapping from distributions to the real line, additionally demands a suitable continuity condition, in accord with the continuous mapping theorem. To this end, we let $c_{n,P}(1-\alpha)$ be the $(1-\alpha)$-quantile of $\psi_{\kappa_n,P}(\mathbb Z_{n,P})$ and impose the following: \begin{ass} (i) $\mathbb Z_{n,P}$ is tight and centered Gaussian in $\mathbf H$ for each $n\in\mathbf N$ and $P\in\mathbf P$; (ii) $\sup_{P\in\mathbf P} E[\|\mathbb Z_{n,P}\|_{\mathbf H}]<\infty$ uniformly in $n$; (iii) $c_{n,P}(1-\alpha-\varpi)\ge c_{n,P}(0.5)+\varsigma_n$ for some constant $\varpi,\varsigma_n>0$, each $n$ and $P\in\mathbf P_0$; (iv) $c_n/\varsigma_n^2=O(1)$ as $n\to\infty$. \end{ass} Assumption (ref)(i) formalizes Gaussianity and tightness of each $\mathbb Z_{n,P}$, which is fulfilled in Examples (ref)--(ref) as well as most regular settings. Assumption (ref)(ii) is a mild moment condition that, in our examples, is tantamount to uniform boundedness of eigenvalues of some matrices with growing dimensions. Assumptions (ref)(iii)(iv) are tied to the natures of densities of Gaussian functionals. Together with Assumptions (ref) and (ref)(i)-(ii), they in effect amount to the aforementioned continuity condition. We now state the first main result concerning the size of our test. \begin{thm} Let Assumptions (ref), (ref), and (ref) hold. If $0\le\kappa_n/r_n=o(c_n)$, then \begin{align} \limsup_{n\to\infty}\sup_{P\in\mathbf P_0} P( r_n \phi(\hat\theta_n)> \hat c_{n,1-\alpha})\le \alpha , \end{align} and, for $\bar{\mathbf P}_0\equiv\{P\in\mathbf P_0: \langle\, \vartheta,\theta_P\rangle_{\mathbf H}=0\,\,\forall\,\vartheta\in\mathbf H\,\,\mathrm{ s.t.}\,\, \sup_{\lambda\in\Lambda}\langle\vartheta,\lambda\rangle_{\mathbf H}\le 0\}$, \begin{align} \limsup_{n\to\infty}\sup_{P\in\bar{\mathbf P}_0} |P( r_n \phi(\hat\theta_n)> \hat c_{n,1-\alpha})-\alpha|=0 . \end{align} \end{thm} In addition to size control, Theorem (ref) shows that the limiting rejection rate equals the nominal level, uniformly over $\bar{\mathbf P}_0$. Heuristically, $\bar{\mathbf P}_0$ may be viewed as consisting of the least favorable distributions in $\mathbf P_0$. Indeed, one can show that (i) $\bar{\mathbf P}_0=\{P\in\mathbf P_0: \theta_P=0\}$ if $\Lambda=\mathbf R_+$, and (ii) $\bar{\mathbf P}_0$ contains constant (resp.\ linear) functions if $\Lambda$ is the family of monotone (resp.\ concave) functions in $L^2(\mathcal Z)$. We stress that, the choice $\kappa_n=0$ leads to a least favorable test that amounts to assuming $\theta_P=0$ (or $\theta_P\in\bar{\mathbf P}_0$ by Lemma (ref)). As well documented, least favorable tests, while controlling size uniformly, can be substantially conservative, which has motivated the active development of more powerful tests in nonstandard settings AndrewsSoares2010,LeeSongWhang2017Inequal. Intuitively, in view of Lemma (ref)(ii), it is desirable to have $\kappa_n\to\infty$ (to match $r_n\to\infty$), a condition recurrent in nonstandard problems for the sake of nonconservativeness---see FangSantos2018HDD and Appendix (ref). We emphasize that our test is in general asymptotically nonsimilar, i.e., there may exist a sequence $\{P_n\}$ of distributions from the null such that \begin{align} \liminf_{n\to\infty}P_n(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha})<\alpha . \end{align} However, this is, in our view, neither an evidence against nonconservativeness nor a deficiency of our test, analogous to the one-sided $t$-test of $\mathrm H_0: \theta_0\le 0$ whose rejection rates tend to zero at $\theta_0<0$. At a deeper level, (ref) is in line with the fact that similarity is not a desirable criterion in nonstandard settings, and can lead to tests with very poor power Lehmann1952Multi,Andrews2012Similar. Indeed, many powerful tests in these settings are asymptotically nonsimilar---see, e.g., AndrewsSoares2010, LeeSongWhang2017Inequal and Chetverikov2018Monotonicity. Turning to the power of our test, we define $\mathbf P_{1,n}^\Delta\equiv\{P\in\mathbf P_1: \phi(\theta_P)\ge \Delta/r_n\}$ for $\Delta>0$. The next theorem shows that our test has nontrivial power against $\mathbf P_{1,n}^\Delta$. \begin{thm} If Assumptions (ref), (ref), and (ref)(ii) hold and $\kappa_n\ge 0$, then \begin{align} \liminf_{\Delta\to\infty}\liminf_{n\to\infty}\inf_{P\in\mathbf P_{1,n}^\Delta} P( r_n \phi(\hat\theta_n)> \hat c_{n,1-\alpha})=1 . \end{align} \end{thm} If we employ the kernel estimator in Example (ref), then Theorem (ref) predicts that our test is powerful against Pitman drifts of order $(nh_n^{d_z})^{-1/2}$. In comparison, the test of LeeSongWhang2017Inequal is powerful against drifts of order $(nh_n^{2\nu})^{-1/2}$ or $(nh_n^{d_z/2+2\nu})^{-1/2}$ depending on the drifts, where $\nu=1$ for monotonicity and $\nu=2$ for concavity. Thus, for monotonicity, our convergence rate is faster if $d_z=1$ but may be slower if $d_z>4$; for concavity, our rate is faster if $d_z<4$ but may be slower if $d_z>8$. Theorem (ref) also suggests that $\kappa_n$ affects power through higher order terms, and, we reiterate that asymptotic nonconservativeness requires $\kappa_n\to\infty$ subject to $\kappa_n/r_n=o(c_n)$. \subsection{Selection of the Tuning Parameter} To motivate, we note that Lemma (ref)(iii) implies: for any $\epsilon>0$, \begin{align} P(\frac{\kappa_n}{r_n}\|r_n\{\hat\theta_n-\theta_P\}\|_{\mathbf H}\le c_n\epsilon)=P(\|r_n\{\hat\theta_n-\theta_P\}\|_{\mathbf H}\le \frac{r_nc_n}{\kappa_n}\epsilon)\to 1 . \end{align} This suggests that we could choose $\kappa_n$ to be such that $r_nc_n/\kappa_n$ is the $(1-\gamma_n)$-quantile of $\|r_n\{\hat\theta_n-\theta_P\}\|_{\mathbf H}$ with $\gamma_n\downarrow 0$, or the $(1-\gamma_n)$ conditional quantile of $\|\hat{\mathbb G}_n\|_{\mathbf H}$ (given $\{X_i\}_{i=1}^n$) since $\|r_n\{\hat\theta_n-\theta_P\}\|_{\mathbf H}$ is unknown. Formally, we let $\hat\kappa_n\equiv r_nc_n/\hat\tau_{n,1-\gamma_n}$, where \begin{align} \hat\tau_{n,1-\gamma_n}\equiv \inf\{c\in\mathbf R: P(\|\hat{\mathbb G}_n\|_{\mathbf H}\le c|\{X_i\}_{i=1}^n)\ge 1-\gamma_n\} . \end{align} To justify the construction $\hat\kappa_n$, we need to introduce our final assumption. \begin{ass} $\liminf_{n\to\infty}\inf_{P\in\mathbf P_0} \bar\sigma_{n,P}^2>0$ with $\bar\sigma_{n,P}^2\equiv \sup_{h\in\mathbf H: \|h\|_{\mathbf H}\le 1} E[\langle h,\mathbb Z_{n,P}\rangle_{\mathbf H}^2]$. \end{ass} Heuristically, Assumption (ref) requires that the coupling variable $\mathbb Z_{n,P}$ for $\hat\theta_n$ be asymptotically non-degenerate. In turn, we can now verify the validity of $\hat\kappa_n$. \begin{pro} Let Assumptions (ref), (ref)(i)-(ii), and (ref) hold, and set $\hat\kappa_n\equiv r_nc_n/\hat\tau_{n,1-\gamma_n}$ with $\gamma_n\in(0,1)$ and $\hat\tau_{n,1-\gamma_n}$ as in (ref). If $\gamma_n\to 0$, then $\hat\kappa_n/r_n = o_p(c_n)$ uniformly in $P\in\mathbf P_0$. If $(r_nc_n)^{-2}\log\gamma_n\to 0$, then $\hat\kappa_n\xrightarrow{p}\infty$ uniformly in $P\in\mathbf P_0$. \end{pro} Analogous rate conditions on $\gamma_n$ have appeared in parametric settings---see ChenFang2016Rank and references therein. In nonparametric settings, the practice of obtaining data-driven tuning parameters through quantile estimation has appeared in ChernozhukovLeeRosen2013Intersection, ChernozhukovNeweySantos2019CCMM, and FangSantos2018HDD, though a formal theory seems to be lacking. While the choice of $\gamma_n$ remains technically challenging, the situation somewhat improves because $\gamma_n$ is unit/scale-free, and prior studies such as FangSantos2018HDD and ChenFang2016Rank have shown that finite sample results are often insensitive to the choice of $\gamma_n$, as also confirmed in our simulations. We recommend $\gamma_n=0.01/\log n$ or $1/n$ for practical implementations. Finally, the rate $c_n$ (in $\hat\kappa_n$) may be ignored in regular settings, and set to be $1/\log n$ in irregular settings without endogeneity. In Example (ref), we may let $c_n$ be $(\log k_n)^{-\varsigma}$ for some $\varsigma\in[1/2,1]$ with the sieve 2SLS estimation and $k_n$ basis functions for $\theta_0$ ChenChristensen2018SupNormOptimal. We note that $\kappa_n$ affects the critical value $\hat c_{n,1-\alpha}$ and hence the rejection rates monotonically: a smaller (resp.\ larger) $\kappa_n$ leads to less (resp.\ more) rejections---see Lemma (ref). As a result, if one is uncertain about $k_n$ or $\varsigma$, he/she could simply take $c_n=1/\log n$, thereby only making $\hat c_{n,1-\alpha}$ larger. \subsection{Implementation and Practical Issues} We next provide a guide for implementing our test. Computation of projections (needed in {\sc Steps} 1 and 2 below) will be discussed in the end. \underline{\sc Step 1:} Compute the test statistic $r_n\phi(\hat\theta_n)=r_n\|\hat\theta_n-\Pi_\Lambda(\hat\theta_n)\|_{\mathbf H}$. This step requires an unconstrained estimator $\hat\theta_n$ of $\theta_P$, which may be obtained by standard estimation procedures---see Examples (ref)-(ref) for specific estimators and their rates $r_n$. We stress that our framework imposes no additional structures on $\hat\theta_n$ as far as implementation is concerned. With $\hat\theta_n$ and its projection $\Pi_\Lambda(\hat\theta_n)$ in hand, one may approximate $r_n\phi(\hat\theta_n)$ by, e.g., the trapezoid rule in Examples (ref)--(ref) where the $\|\cdot\|_{\mathbf H}$-norm takes the form of an integral. Concretely, in Example (ref) with $\mathcal Z=[a,b]$ and $a<b$, we approximate $\phi(\hat\theta_n)$ by: for a large $N$ and $\Delta=(b-a)/N$, \begin{align} \{\frac{1}{2}\frac{b-a}{N} [f^2(z_0)+2f^2(z_1)+\cdots +2f^2(z_{N-1})+f^2(z_N)]\}^{1/2} , \end{align} where $f=\hat\theta_n-\Pi_\Lambda(\hat\theta_n)$ and $z_j=a+j\Delta$. \underline{\sc Step 2:} Construct the critical value $\hat c_{n,1-\alpha}$ defined in (ref). The construction requires a bootstrap analog $\hat{\mathbb G}_n$ of $r_n \{\hat\theta_n-\theta_P\}$ and a tuning parameter $\kappa_n$. In regular settings, one often has $\hat{\mathbb G}_n=\sqrt n \{\hat\theta_n^*-\hat\theta_n\}$, where $\hat\theta_n^*$ is the same as $\hat\theta_n$ but based on bootstrap samples. In Examples (ref)--(ref), one may employ various bootstrap schemes in ChernozhukovLeeRosen2013Intersection, ChernozhukovNeweySantos2019CCMM, ChenChristensen2018SupNormOptimal, and BelloniChernozhukovChetverikovFernandez2019QR. For $\kappa_n$, we recommend $\hat\kappa_n$ proposed in Section (ref) with a suitable $\gamma_n$ (e.g., $\gamma_n=0.01/\log n$ or $1/n$). Summarizing, the critical value $\hat c_{n,1-\alpha}$ may be then obtained as follows: \begin{enumerate}[label=(\roman*)] • Generate $B$ realizations $\{\hat{\mathbb G}_{n,b}\}_{b=1}^B$ of $\hat{\mathbb G}_n$ (e.g., $B=200$ or larger); e.g., in (ref), generate $\{\{W_{i,b}\}_{i=1}^n\}_{b=1}^B$ that are i.i.d.\ across both $i$ and $b$, and then obtain each $\hat{\mathbb G}_{n,b}$ by evaluating $\hat{\mathbb G}_n$ at $\{W_{i,b}\}_{i=1}^n$, which only involves linear calculations. • Set $\hat\kappa_n=r_nc_n/\hat\tau_{n,1-\gamma_n}$ where $\hat\tau_{n,1-\gamma_n}$ is the $(1-\gamma_n)$-quantile of the $B$ numbers $\|\hat{\mathbb G}_{n,1}\|_{\mathbf H},\ldots, \|\hat{\mathbb G}_{n,B}\|_{\mathbf H}$, and $c_n=1/\log n$ for Example (ref). The $\|\cdot\|_{\mathbf H}$-norms (here and below) may be computed by the trapezoid rule as in {\sc Step} 1. • Approximate $\hat c_{n,1-\alpha}$ by the $(1-\alpha)$-quantile of the $B$ numbers \begin{multline} \|\hat{\mathbb G}_{n,1}+\hat\kappa_n\Pi_\Lambda\hat\theta_n-\Pi_\Lambda(\hat{\mathbb G}_{n,1}+\hat\kappa_n\Pi_\Lambda\hat\theta_n)\|_{\mathbf H},\\ \ldots, \|\hat{\mathbb G}_{n,B}+\hat\kappa_n\Pi_\Lambda\hat\theta_n-\Pi_\Lambda(\hat{\mathbb G}_{n,B}+\hat\kappa_n\Pi_\Lambda\hat\theta_n)\|_{\mathbf H} . \end{multline} \end{enumerate} \underline{\sc Step 3:} Reject $\mathrm H_0$ if and only if $r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha}$. Next, we illustrate the computation of projections. As described in Appendix (ref), when closed form expressions do not exist, the projection $\Pi_{\Lambda}(\theta)$ can be computed by solving a linearly constrained quadratic program: for some $A\in\mathbf M^{m\times k}$, \begin{align} \min_{h\in\mathbf R^k}\|h-\vartheta\|\qquad \text{s.t.}\quad Ah\ge 0 , \end{align} where $\vartheta\in\mathbf R^k$ is the vector of values $\theta$ takes at the grid points, and $Ah\ge 0$ is the discretized version of the restriction in question. As an extensively studied problem, (ref) admits polynomial-time algorithms; e.g., the iteration complexity of the interior point method is $O(\sqrt{m+k}\log(1/\epsilon))$ for an $\epsilon$-accurate solution. Finally, since (ref) is inherently more complicated to compute in higher dimensions, we note a number of strategies for ameliorating the situation. First, the recently developed open-source solver OSQP (\url{https://osqp.org}) is very robust in solving large scale quadratic programming problems. Second, while projection under convexity/concavity is computationally demanding in multivariate settings, the representation result in Kuosmanen2008Convex, when coupled with the OSQP solver, can help significantly reduce the computation cost. Last, with regressors more than two or three, one may consider employing semiparametric models, instead of fully nonparametric ones, to alleviate the curse of dimensionality. \section{Simulation Studies} This section examines the finite sample performance of our test. Due to space limitation, we focus on monotonicity, and defer concavity, monotonicity jointly with concavity, and Slutsky restriction to the online appendix. Throughout, the significance level is 5%, and, unless otherwise specified, the number of Monte Carlo replications is 3000 while the number of bootstrap repetitions for each replication is 200. All integrals are approximated by the trapezoid rule, and quadratic programs are solved by the OSQP solver. Our test is implemented based on the data-driven choice $\hat\kappa_n$ with $\gamma_n\in\{n^{-1/2}, n^{-3/4}, 1/n, 0.1/\log n, 0.05/\log n, 0.01/\log n, 0.1, 0.05, 0.01\}$. \subsection{Simulation Designs} We aim to test whether $\theta_0$ is nondecreasing in Example (ref) with $d_z\in\{1,2\}$. For $d_z=1$, the regression function $\theta_0:[-1,1]\to\mathbf R$ is specified as: for some $(\mathsf a,\mathsf b,\mathsf c)\in\mathbf R^3$, \begin{align} \theta_0(z)=\mathsf az-\mathsf b\varphi(\mathsf cz) , \end{align} where $\varphi$ is the standard normal pdf. We consider three choices for $(\mathsf a,\mathsf b,\mathsf c)$ under the null hypothesis, namely $(0,0,0)$, $(0.1,0.5,0.5)$ and $(0.5,2,1)$ that are labeled D1, D2 and D3 respectively, and the collection $\{(\mathsf a,\mathsf b,\mathsf c): \mathsf a =0,\mathsf b = 0.2\delta,\mathsf c = 5+0.1\delta,\delta=1,\ldots,10\}$ under the alternative---see Figure (ref). We then draw i.i.d.\ samples $\{Z_i^*,u_i\}_{i=1}^n$ with $n\in\{500,750,1000\}$ from the standard normal distribution in $\mathbf R^2$, and set $Z_i=-1+2\Phi(Z_i^*)\in[-1,1]$ with $\Phi$ the standard normal cdf. For $d_z=2$, the regression function $\theta_0:[0,1]^2\to\mathbf R$ is of the form: for some $(\mathsf a,\mathsf b,\mathsf c)\in\mathbf R^3$, \begin{align} \theta_0(z_1,z_2)=\mathsf a\big(\frac{1}{2}z_1^{\mathsf b}+\frac{1}{2}z_2^{\mathsf b}\big)^{1/\mathsf b}+\mathsf c\log(1+z_1+z_2) . \end{align} We consider three choices of $(\mathsf a,\mathsf b,\mathsf c)$ for the null, namely $(0,0,0)$, $(0.2,1,0)$ and $(0.5,0,0.5)$ that are labeled D1, D2, and D3, respectively, and, for the alternative, set $\mathsf b=0$ and $\mathsf a=\mathsf c=-\Delta\delta $ with $\Delta=0.05$ and $\delta=1,\ldots,10$. Note that the first term on the right-hand side of (ref) collapses to $\mathsf a\sqrt{z_1z_2}$ whenever $\mathsf b=0$. In turn, we draw i.i.d.\ samples $\{Z_{1i}^*,Z_{2i}^*,u_i\}_{i=1}^n$ with $n\in\{500,750,1000\}$ from the standard normal distribution in $\mathbf R^3$, and set $Z_i\equiv(Z_{1i},Z_{2i})$ with $Z_{ji}=\Phi(Z_{ji}^*)\in[0,1]$ for all $i$ and $j=1,2$. \begin{figure} \scriptsize \captionsetup[subfigure]{justification=centering} \begin{subfigure}[b]{0.49\linewidth} \resizebox{\textwidth}{!}{ \begin{tikzpicture}[baseline,>=latex'] \begin{axis}[axis equal image, domain=-1:1, samples=100, xmin=-1.1, xmax=1.1, ymin=-1.1, ymax=0.1, enlargelimits=false, xlabel=$z$, x label style={at={(axis description cs:0.5,-0.02)},anchor=north}, xtick={-1,-0.5,...,1}, xticklabels={-1,-0.5,...,1}, xticklabel style={yshift=0.05cm,anchor=south}, ytick={-1,-0.5,0}, yticklabels={-1,-0.5,0}, every axis plot/.append style={semithick}, ] \addplot[NavyBlue]{0}; \addplot[NavyBlue,dashed]{0.1*x-0.5*npdf(0.5*x, 0, 1)}; \addplot[NavyBlue,densely dotted]{0.5*x-2*npdf(x, 0, 1)}; \end{axis} \end{tikzpicture} } \caption{$\mathrm H_0$: D1 (solid), D2 (dashed) and D3 (dotted)} \end{subfigure} \begin{subfigure}[b]{0.49\linewidth} \resizebox{\textwidth}{!}{ \begin{tikzpicture}[baseline,>=latex'] \begin{axis}[axis equal image, domain=-1:1, samples=100, xmin=-1.1, xmax=1.1, ymin=-1.1, ymax=0.1, enlargelimits=false, xlabel=$z$, x label style={at={(axis description cs:0.5,-0.02)},anchor=north}, xtick={-1,-0.5,...,1}, xticklabels={-1,-0.5,...,1}, xticklabel style={yshift=0.05cm,anchor=south}, ytick={-1,-0.5,0}, yticklabels={-1,-0.5,0}, every axis plot/.append style={thin}, no markers, ] \foreach \b in {0.2,0.4,...,2}{\pgfmathsetmacro{\k}{\b*120} \edef\temp{\noexpand \addplot+[solid,color=Blue!\k]{-\b*npdf((5+\b/2)*x, 0, 1)}; }\temp } \end{axis} \end{tikzpicture} } \caption{$\mathrm H_1$: $\delta=1,2,\ldots,10$ (from top to bottom)} \end{subfigure} \caption{The function $\theta_0$ in (ref) where, in Figure (ref), $\mathsf a=0$, $\mathsf b = 0.2\delta$, and $\mathsf c=5+0.1\delta$. Note that the standard deviation of the error is designed to be no smaller than the range of $\theta_0$.} \end{figure} To implement our test for (ref), we obtain $\hat\theta_n$ by sieve estimation with cubic B-splines with 3, 5, or 7 interior knots at the equispaced quantiles of $\{Z_i\}_{i=1}^n$, so that the sieve dimension $k_n$ equals $7$, $9$, or $11$, respectively. In multivariate settings, we construct the series functions via tensor product of univariate B-splines. Since the sieve dimension grows quickly as $d_z$ increases (e.g., $k_n=49$ with cubic B-splines, $3$ knots, and $d_z=2$), we employ univariate quadratic as well as cubic B-splines with one or zero knots along each dimension for (ref). Thus, for example, $k_n=9$ with quadratic B-splines and zero knots. In both designs, we compute $\hat{\mathbb G}_{n,b}$ as in (ref) by drawing i.i.d.\ weights from the standard normal distribution, and let the coupling rate $c_n=1/\log n$. To alleviate the boundary effects, we evaluate the $L^2$-norms (here and below) over $[-0.9,0.9]$ for (ref) and $[0.1,0.9]^2$ for (ref), with step size $0.05$. For ease of reference, we label our test with quadratic B-splines and $j$ knots as FS-Q$j$; similarly, FS-C$j$ is the implementation with cubic B-splines and $j$ knots. To compare, we implement two alternative nonconservative tests, LeeSongWhang2017Inequal and Chetverikov2018Monotonicity. The latter also compares with some prior tests in simulations, which show marked power superiority of the author's three tests so that we take them as important benchmarks. For brevity, however, we only present the one-step test, labeled C-OS, and note that the results for the other two tests are very similar to those produced by C-OS, as also observed in Chetverikov2018Monotonicity. The details for implementing C-OS in the univariate case are clearly laid out in Chetverikov2018Monotonicity. For the bivariate case, we compute the test statistic as on p.27 in the arXiv version of Chetverikov2018Monotonicity, adopt equispaced empirical quantiles $0.1,0.15,...,0.9$ of each covariate as locations for the weighting function to make the computation feasible, and otherwise follow the implementation in the univariate case. LeeSongWhang2017Inequal considered $L^p$-type statistics. For the sake of comparison, we focus on $p =2$, since different choices of $p$ implicitly aim power at different alternatives. We estimate the first derivatives of $\theta_0$ by local quadratic regression FanGijbels1996Local, with a kernel $z\mapsto 1.5\max\{1-(2z)^2,0\}$ for (ref) and $z\mapsto 0.75\max\{1-z^2,0\}$ for (ref). We choose two bandwidths: a “large” one $2s_nn^{-1/(d_z+2q+2)}$ (in the spirit of LeeSongWhang2017Inequal) and a “small” one $2s_nn^{-1/(d_z+2(q-1)+2)}$ with $q=2$ and $s_n$ the standard derivation of $\{Z_i\}_{i=1}^n$, resulting in two tests labeled LSW-L and LSW-S, respectively. As studentization can be crucial for power (based on unreported simulations), we estimate the standard errors following FanGijbels1996Local, with the variance of the error estimated by a local polynomial regression of order $q+2$ with bandwidth $2s_nn^{-1/(d_z+2(q+2)+2)}$. Next, we construct critical values based on the empirical bootstrap (as in LeeSongWhang2017Inequal) and the tuning parameter $\hat c_n$ in their Section 5.1 with $C_{\mathrm{cs}}=0.4$ (since their results are quite insensitive to other choices of $C_{\mathrm{cs}}$ there). Finally, to ease computation for the bivariate designs, the number of simulation replications is decreased to be 1000 (for LSW only). { {6.5pt} \begin{table}[!ht] \caption{Empirical Size of Monotonicity Tests for $\theta_0$ in (ref) at $\alpha=5\%$} \begin{threeparttable} \sisetup{table-number-alignment = center, table-format = 1.3} \begin{tabularx}{\linewidth}{@ cc *{3}{S[round-mode = places,round-precision = 3]} c *{3}{S[round-mode = places,round-precision = 3]} c *{3}{S[round-mode = places,round-precision = 3]}@} \hline \hline \multirow{2}{*}{$n$} & \multirow{2}{*}{$\gamma_n$} & \multicolumn{3}{c}{FS-C3: $k_n=7$} & & \multicolumn{3}{c}{FS-C5: $k_n=9$} & & \multicolumn{3}{c}{FS-C7: $k_n=11$}\\ \cline{3-5} \cline{7-9} \cline{11-13} & & {D1} & {D2} & {D3} & & {D1} & {D2} & {D3} & & {D1} & {D2} & {D3}\\ \hline \multirow{3}{*}{$500$} & $1/n$ & 0.0530 & 0.0157 & 0.0030 & & 0.0577 & 0.0207 & 0.0033 & & 0.0583 & 0.0190 & 0.0027\\ & $0.01/\log n$ & 0.0530 & 0.0157 & 0.0030 & & 0.0577 & 0.0207 & 0.0033 & & 0.0583 & 0.0190 & 0.0027\\ & $0.01$ & 0.0530 & 0.0157 & 0.0030 & & 0.0577 & 0.0207 & 0.0033 & & 0.0583 & 0.0193 & 0.0027\\ \cline{2-13} \multirow{3}{*}{$750$} & $1/n$ & 0.0523 & 0.0097 & 0.0013 & & 0.0560 & 0.0137 & 0.0017 & & 0.0590 & 0.0173 & 0.0030\\ & $0.01/\log n$ & 0.0523 & 0.0097 & 0.0013 & & 0.0560 & 0.0137 & 0.0017 & & 0.0590 & 0.0173 & 0.0030\\ & $0.01$ & 0.0523 & 0.0097 & 0.0013 & & 0.0560 & 0.0140 & 0.0020 & & 0.0590 & 0.0173 & 0.0030\\ \cline{2-13} \multirow{3}{*}{$1000$} & $1/n$ & 0.0557 & 0.0107 & 0.0007 & & 0.0560 & 0.0113 & 0.0003 & & 0.0560 & 0.0127 & 0.0007\\ & $0.01/\log n$ & 0.0557 & 0.0107 & 0.0007 & & 0.0560 & 0.0113 & 0.0003 & & 0.0560 & 0.0127 & 0.0007\\ & $0.01$ & 0.0560 & 0.0107 & 0.0010 & & 0.0560 & 0.0117 & 0.0003 & & 0.0560 & 0.0130 & 0.0007\\ \hline \multicolumn{2}{c}{\multirow{2}{*}{$n$}}& \multicolumn{3}{c}{LSW-S} & & \multicolumn{3}{c}{LSW-L} & & \multicolumn{3}{c}{C-OS}\\ \cline{3-5} \cline{7-9} \cline{11-13} & & {D1} & {D2} & {D3} & & {D1} & {D2} & {D3} & & {D1} & {D2} & {D3}\\ \hline \multicolumn{2}{c}{$500$} & 0.0600 & 0.0413 & 0.0083 & & 0.0660 & 0.0350 & 0.0037 & & 0.0597 & 0.0413 & 0.0120\\ \multicolumn{2}{c}{$750$} & 0.0567 & 0.0363 & 0.0050 & & 0.0587 & 0.0303 & 0.0057 & & 0.0540 & 0.0337 & 0.0080\\ \multicolumn{2}{c}{$1000$} & 0.0610 & 0.0347 & 0.0047 & & 0.0653 & 0.0347 & 0.0033 & & 0.0493 & 0.0357 & 0.0090\\ \hline \hline \end{tabularx} \begin{tablenotes}[flushleft] • {\it Note:} The parameter $\gamma_n$ determines $\hat\kappa_n$ proposed in Section (ref) with $c_n=1/\log n$ and $r_n=(n/k_n)^{1/2}$. \end{tablenotes} \end{threeparttable} \end{table} } \subsection{Simulation Results} Tables (ref) and (ref) summarize the empirical sizes. Due to space limitation, we only present our tests with $\gamma_n\in\{1/n,0.01/\log n,0.01\}$, and relegate to Tables (ref)--H.(ref) in Appendix (ref) the complete set of results. Together, these tables show that our tests are insensitive to the choice of $\gamma_n$. For the univariate design, all tests control size well, though, relatively speaking, the two LSW tests, especially LSW-L, tend to over-reject under D1, while our tests tend to under-reject under D2 and D3. Note, however, that D1 is a least favorable case, D2 is in the “interior” (not in the topological sense), and D3 is further into the “interior.” Thus, the empirical sizes under D2 and D3 are expected to be smaller than 5%. For the bivariate design, our tests are over-sized under D1 in small samples, though this feature is also shared by LSW-L and C-OS (to an overall lesser extent). The over-rejection of our tests, in particular FS-C1 (in which case $k_n=25$), is likely because the number of series functions is so “large” that the Gaussian approximation is somewhat inaccurate---note that the overall situation improves as $n$ increases. Thus, while undersmoothing requires a “large” $k_n$, there lies the tension that it should not be “too large” for the sake of distributional approximations. { {3pt} \begin{table}[!ht] \caption{Empirical Size of Monotonicity Tests for $\theta_0$ in (ref) at $\alpha=5\%$} \begin{threeparttable} \sisetup{table-number-alignment = center, table-format = 1.3} \begin{tabularx}{\linewidth}{@cc *{3}{S[round-mode = places,round-precision = 3]} c *{3}{S[round-mode = places,round-precision = 3]} c *{3}{S[round-mode = places,round-precision = 3]} c *{3}{S[round-mode = places,round-precision = 3]}@} \hline \hline \multirow{2}{*}{$n$} & \multirow{2}{*}{$\gamma_n$}& \multicolumn{3}{c}{FS-Q0: $k_n=9$} & & \multicolumn{3}{c}{FS-Q1: $k_n=16$} & & \multicolumn{3}{c}{FS-C0: $k_n=16$} & & \multicolumn{3}{c}{FS-C1: $k_n=25$}\\ \cline{3-5} \cline{7-9} \cline{11-13} \cline{15-17} && {D1} & {D2} & {D3} & & {D1} & {D2} & {D3} & & {D1} & {D2} & {D3} & & {D1} & {D2} & {D3}\\ \hline \multirow{3}{*}{$500$} & $1/n$ & 0.0613 & 0.0183 & 0.0000 & & 0.0680 & 0.0300 & 0.0017 & & 0.0670 & 0.0303 & 0.0007 & & 0.0830 & 0.0437 & 0.0040\\ & $0.01/\log n$ & 0.0613 & 0.0183 & 0.0000 & & 0.0680 & 0.0300 & 0.0017 & & 0.0670 & 0.0303 & 0.0007 & & 0.0830 & 0.0437 & 0.0040\\ & $0.01$ & 0.0627 & 0.0193 & 0.0000 & & 0.0687 & 0.0300 & 0.0020 & & 0.0673 & 0.0307 & 0.0007 & & 0.0837 & 0.0447 & 0.0043\\ \cline{2-17} \multirow{3}{*}{$750$}& $1/n$ & 0.0583 & 0.0103 & 0.0000 & & 0.0653 & 0.0247 & 0.0003 & & 0.0640 & 0.0220 & 0.0000 & & 0.0737 & 0.0323 & 0.0003\\ & $0.01/\log n$ & 0.0583 & 0.0103 & 0.0000 & & 0.0653 & 0.0247 & 0.0003 & & 0.0640 & 0.0220 & 0.0000 & & 0.0737 & 0.0323 & 0.0003\\ & $0.01$ & 0.0597 & 0.0107 & 0.0000 & & 0.0660 & 0.0250 & 0.0003 & & 0.0650 & 0.0223 & 0.0000 & & 0.0740 & 0.0327 & 0.0007\\ \cline{2-17} \multirow{3}{*}{$1000$}& $1/n$ & 0.0500 & 0.0107 & 0.0000 & & 0.0553 & 0.0210 & 0.0000 & & 0.0527 & 0.0203 & 0.0000 & & 0.0590 & 0.0227 & 0.0003\\ & $0.01/\log n$ & 0.0500 & 0.0107 & 0.0000 & & 0.0553 & 0.0210 & 0.0000 & & 0.0527 & 0.0203 & 0.0000 & & 0.0590 & 0.0227 & 0.0003\\ & $0.01$ & 0.0500 & 0.0107 & 0.0000 & & 0.0560 & 0.0213 & 0.0000 & & 0.0527 & 0.0203 & 0.0000 & & 0.0590 & 0.0230 & 0.0003\\ \hline \multicolumn{2}{c}{\multirow{2}{*}{$n$}} & \multicolumn{3}{c}{LSW-S} & & \multicolumn{3}{c}{LSW-L} & & \multicolumn{3}{c}{C-OS} & & \multicolumn{3}{c}\\ \cline{3-5} \cline{7-9} \cline{11-13} & & {D1} & {D2} & {D3} & & {D1} & {D2} & {D3} & & {D1} & {D2} & {D3}\\ \hline \multicolumn{2}{c}{500} & 0.0430 & 0.0200 & 0.0000 & & 0.0610 & 0.0260 & 0.0020 & & 0.0683 & 0.0607 & 0.0367\\ \multicolumn{2}{c}{750} & 0.0540 & 0.0170 & 0.0000 & & 0.0660 & 0.0310 & 0.0000 & & 0.0570 & 0.0453 & 0.0230\\ \multicolumn{2}{c}{1000} & 0.0440 & 0.0210 & 0.0000 & & 0.0540 & 0.0190 & 0.0000 & & 0.0553 & 0.0423 & 0.0187\\ \hline \hline \end{tabularx} \begin{tablenotes}[flushleft] • {\it Note:} The parameter $\gamma_n$ determines $\hat\kappa_n$ proposed in Section (ref) with $c_n=1/\log n$ and $r_n=(n/k_n)^{1/2}$. \end{tablenotes} \end{threeparttable} \end{table} } \pgfplotstableread{ delta alpha MonFiveKn3 MonFiveKn5 MonFiveKn7 MonSevenKn3 MonSevenKn5 MonSevenKn7 MonTenKn3 MonTenKn5 MonTenKn7 0 0.05 0.0530 0.0577 0.0583 0.0523 0.0560 0.0590 0.0557 0.0560 0.0560 1 0.05 0.0697 0.0733 0.0707 0.0773 0.0717 0.0727 0.0857 0.0820 0.0780 2 0.05 0.1203 0.1183 0.1117 0.1563 0.1403 0.1337 0.1890 0.1683 0.1510 3 0.05 0.2143 0.2000 0.1800 0.3090 0.2790 0.2500 0.4087 0.3603 0.3220 4 0.05 0.3633 0.3197 0.2940 0.5107 0.4743 0.4337 0.6580 0.6180 0.5663 5 0.05 0.5357 0.4960 0.4467 0.7113 0.6853 0.6350 0.8597 0.8257 0.7837 6 0.05 0.6977 0.6717 0.6147 0.8787 0.8513 0.8190 0.9573 0.9487 0.9337 7 0.05 0.8323 0.8087 0.7667 0.9560 0.9460 0.9303 0.9890 0.9883 0.9823 8 0.05 0.9213 0.9077 0.8810 0.9873 0.9843 0.9773 0.9983 0.9980 0.9963 9 0.05 0.9640 0.9577 0.9440 0.9973 0.9973 0.9950 1.0000 1.0000 1.0000 10 0.05 0.9860 0.9843 0.9773 0.9993 0.9993 0.9993 1.0000 1.0000 1.0000 }\FirstMon \pgfplotstableread{ delta alpha Five1 Seven1 Ten1 Five2 Seven2 Ten2 Fifty2 0 0.05 0.0597 0.0540 0.0493 0.0683 0.0570 0.0553 0.0460 1 0.05 0.0623 0.0600 0.0620 0.0760 0.0600 0.0583 0.0653 2 0.05 0.0767 0.0873 0.1067 0.0797 0.0693 0.0690 0.1023 3 0.05 0.1173 0.1673 0.2260 0.0923 0.0783 0.0860 0.2077 4 0.05 0.1893 0.3010 0.4423 0.0980 0.0983 0.1077 0.4153 5 0.05 0.3210 0.5050 0.6797 0.1123 0.1190 0.1393 0.6890 6 0.05 0.4780 0.7067 0.8657 0.1283 0.1593 0.1877 0.9043 7 0.05 0.6383 0.8530 0.9613 0.1540 0.2120 0.2513 0.9810 8 0.05 0.7813 0.9470 0.9913 0.1810 0.2693 0.3413 0.9993 9 0.05 0.8807 0.9850 0.9987 0.2250 0.3410 0.4530 1.0000 10 0.05 0.9450 0.9957 1.0000 0.2737 0.4260 0.5753 1.0000 }\MainChetverikov \pgfplotstableread{ delta alpha FiveOp FiveUn SevenOp SevenUn TenOp TenUn 0 0.05 0.0660 0.0600 0.0587 0.0567 0.0653 0.0610 1 0.05 0.0713 0.0607 0.0667 0.0563 0.0790 0.0653 2 0.05 0.0857 0.0660 0.0843 0.0677 0.0973 0.0737 3 0.05 0.1177 0.0787 0.1343 0.0783 0.1467 0.0933 4 0.05 0.1727 0.0993 0.2147 0.1053 0.2473 0.1217 5 0.05 0.2497 0.1293 0.3287 0.1453 0.3817 0.1693 6 0.05 0.3537 0.1717 0.4823 0.2027 0.5753 0.2360 7 0.05 0.4753 0.2280 0.6403 0.2870 0.7690 0.3387 8 0.05 0.6113 0.2973 0.7873 0.4017 0.8967 0.4620 9 0.05 0.7433 0.3777 0.9010 0.5153 0.9647 0.6140 10 0.05 0.8547 0.4803 0.9600 0.6527 0.9897 0.7613 }\MainFirstLSWMon \pgfplotstableread{ delta alpha MonFiveQKn0 MonFiveQKn1 MonSevenQKn0 MonSevenQKn1 MonTenQKn0 MonTenQKn1 0 0.05 0.0613 0.0680 0.0583 0.0653 0.0500 0.0553 1 0.05 0.1110 0.0980 0.1270 0.1097 0.1267 0.0990 2 0.05 0.1987 0.1450 0.2530 0.1773 0.2607 0.1720 3 0.05 0.3173 0.2010 0.4100 0.2730 0.4573 0.2857 4 0.05 0.4597 0.2843 0.5907 0.3850 0.6733 0.4300 5 0.05 0.6147 0.3887 0.7643 0.5207 0.8473 0.5900 6 0.05 0.7457 0.5060 0.8843 0.6493 0.9493 0.7627 7 0.05 0.8573 0.6200 0.9587 0.7757 0.9823 0.8780 8 0.05 0.9313 0.7237 0.9853 0.8720 0.9963 0.9593 9 0.05 0.9790 0.8223 0.9973 0.9387 0.9993 0.9837 10 0.05 0.9937 0.8980 0.9993 0.9760 1.0000 0.9957 }\SecondMonQ \pgfplotstableread{ delta alpha MonFiveCKn0 MonFiveCKn1 MonSevenCKn0 MonSevenCKn1 MonTenCKn0 MonTenCKn1 0 0.05 0.0670 0.0830 0.0640 0.0737 0.0527 0.0590 1 0.05 0.0960 0.1060 0.1090 0.1113 0.1027 0.0970 2 0.05 0.1413 0.1467 0.1827 0.1583 0.1800 0.1527 3 0.05 0.2070 0.1930 0.2737 0.2323 0.2917 0.2347 4 0.05 0.2860 0.2503 0.3917 0.3263 0.4400 0.3477 5 0.05 0.3947 0.3263 0.5243 0.4333 0.6017 0.4857 6 0.05 0.5073 0.4183 0.6517 0.5617 0.7630 0.6297 7 0.05 0.6240 0.5137 0.7750 0.6813 0.8787 0.7787 8 0.05 0.7307 0.6230 0.8763 0.7857 0.9557 0.8813 9 0.05 0.8303 0.7247 0.9413 0.8687 0.9837 0.9510 10 0.05 0.8970 0.8033 0.9737 0.9307 0.9940 0.9800 }\SecondMonC \pgfplotstableread{ delta alpha FiveOp FiveUn SevenOp SevenUn TenOp TenUn 0 0.05 0.0610 0.0430 0.0660 0.0540 0.0540 0.0440 1 0.05 0.0920 0.0610 0.1000 0.0820 0.0960 0.0780 2 0.05 0.1300 0.0840 0.1670 0.1160 0.1520 0.1300 3 0.05 0.1960 0.1300 0.2430 0.1850 0.2380 0.1830 4 0.05 0.2600 0.1780 0.3450 0.2580 0.3800 0.2620 5 0.05 0.3460 0.2230 0.4750 0.3240 0.5210 0.3500 6 0.05 0.4560 0.2900 0.5810 0.4080 0.6530 0.4460 7 0.05 0.5590 0.3650 0.7200 0.4970 0.7980 0.5640 8 0.05 0.6780 0.4470 0.8230 0.5960 0.8970 0.6620 9 0.05 0.7780 0.5400 0.8900 0.7040 0.9720 0.7960 10 0.05 0.8410 0.6270 0.9410 0.7840 0.9930 0.8680 }\MainSecondLSWMon Figure (ref) depicts the power curves, where we only show our tests with $\gamma_n=0.01/\log n$ for brevity and the fact that other choices of $\gamma_n$ lead to very similar curves---see Figure (ref) in Appendix (ref). For $d_z=1$, our tests are moderately more powerful than C-OS, across sample sizes and the number of interior knots, and they are all considerably more powerful than LSW-L and in particular LSW-S. For $d_z=2$, our tests remain competitive in terms of power, though LSW-L is more powerful than FS-C1 (the least powerful among our tests). Interestingly, C-OS is the least powerful of all tests, which may be explained by the fact that only part of the discordance between regressors and the outcome is being picked up through the indicator function in the test function $b(s)$---see p.27 in the arXiv version of Chetverikov2018Monotonicity. We note that the power of our tests is overall decreasing in $k_n$, which is consistent with Theorem (ref) since $r_n=\sqrt{n/k_n}$. \begin{figure}[!h] \scriptsize \begin{tikzpicture} \begin{groupplot}[group style={group name=myplots,group size=3 by 2,horizontal sep= 0.8cm,vertical sep=1.1cm}, grid = minor, width = 0.375\textwidth, xmax=10,xmin=0, ymax=1,ymin=0, every axis title/.style={below,at={(0.2,0.8)}}, xlabel=$\delta$, x label style={at={(axis description cs:0.95,0.04)},anchor=south}, xtick={0,2,...,10}, ytick={0.05,0.5,1}, tick label style={/pgf/number format/fixed}, legend style={text=black,cells={align=center},row sep = 3pt,legend columns = -1, draw=none,fill=none}, cycle list={ {smooth,tension=0.5,color=SpringGreen4, mark=halfsquare*,every mark/.append style={rotate=270},mark size=1.5pt,line width=0.5pt}, {smooth,tension=0.5,color=DarkOliveGreen4, mark=halfsquare*,every mark/.append style={rotate=90},mark size=1.5pt,line width=0.5pt}, {smooth,tension=0.5,color=Honeydew4, mark=10-pointed star,mark size=1.5pt,line width=0.5pt}, {smooth,tension=0.5,color=RoyalBlue1, mark=halfcircle*,every mark/.append style={rotate=90},mark size=1.5pt,line width=0.5pt}, {smooth,tension=0.5,color=RoyalBlue2, mark=halfcircle*,every mark/.append style={rotate=180},mark size=1.5pt,line width=0.5pt}, {smooth,tension=0.5,color=RoyalBlue3, mark=halfcircle*,every mark/.append style={rotate=270},mark size=1.5pt,line width=0.5pt}, {smooth,tension=0.5,color=RoyalBlue4, mark=halfcircle*,every mark/.append style={rotate=360},mark size=1.5pt,line width=0.5pt}, } ] \nextgroupplot[legend style = {column sep = 7pt, legend to name = LegendMon1}] \addplot[smooth,tension=0.5,color=NavyBlue, no markers,line width=0.25pt, densely dotted,forget plot] table[x = delta,y=alpha] from \FirstMon; \addplot table[x = delta,y=FiveUn] from \MainFirstLSWMon; \addplot table[x = delta,y=FiveOp] from \MainFirstLSWMon; \addplot table[x = delta,y=Five1] from \MainChetverikov; \addplot table[x = delta,y=MonFiveKn3] from \FirstMon; \addplot table[x = delta,y=MonFiveKn5] from \FirstMon; \addplot table[x = delta,y=MonFiveKn7] from \FirstMon; \node[anchor=north] at (axis description cs: 0.25, 0.95) {\fontsize{5}{4}\selectfont $\begin{aligned} d_z&=1\\ n&=500 \end{aligned}$}; \addlegendentry{LSW-S}; \addlegendentry{LSW-L}; \addlegendentry{C-OS}; \addlegendentry{FS-C3}; \addlegendentry{FS-C5}; \addlegendentry{FS-C7}; \nextgroupplot \node[anchor=north] at (axis description cs: 0.25, 0.95) {\fontsize{5}{4}\selectfont $\begin{aligned} d_z&=1\\ n&=750 \end{aligned}$}; \addplot[smooth,tension=0.5,color=NavyBlue, no markers,line width=0.25pt, densely dotted,forget plot] table[x = delta,y=alpha] from \FirstMon; \addplot table[x = delta,y=SevenUn] from \MainFirstLSWMon; \addplot table[x = delta,y=SevenOp] from \MainFirstLSWMon; \addplot table[x = delta,y=Seven1] from \MainChetverikov; \addplot table[x = delta,y=MonSevenKn3] from \FirstMon; \addplot table[x = delta,y=MonSevenKn5] from \FirstMon; \addplot table[x = delta,y=MonSevenKn7] from \FirstMon; \nextgroupplot \node[anchor=north] at (axis description cs: 0.25, 0.95) {\fontsize{5}{4}\selectfont $\begin{aligned} d_z&=1\\ n&=1000 \end{aligned}$}; \addplot[smooth,tension=0.5,color=NavyBlue, no markers,line width=0.25pt, densely dotted,forget plot] table[x = delta,y=alpha] from \FirstMon; \addplot table[x = delta,y=TenUn] from \MainFirstLSWMon; \addplot table[x = delta,y=TenOp] from \MainFirstLSWMon; \addplot table[x = delta,y=Ten1] from \MainChetverikov; \addplot table[x = delta,y=MonTenKn3] from \FirstMon; \addplot table[x = delta,y=MonTenKn5] from \FirstMon; \addplot table[x = delta,y=MonTenKn7] from \FirstMon; \nextgroupplot[legend style = {column sep = 3.5pt, legend to name = LegendMon2}] \node[anchor=north] at (axis description cs: 0.25, 0.95) {\fontsize{5}{4}\selectfont $\begin{aligned} d_z&=2\\ n&=500 \end{aligned}$}; \addplot[smooth,tension=0.5,color=NavyBlue, no markers,line width=0.25pt, densely dotted,forget plot] table[x = delta,y=alpha] from \FirstMon; \addplot table[x = delta,y=FiveUn] from \MainSecondLSWMon; \addplot table[x = delta,y=FiveOp] from \MainSecondLSWMon; \addplot table[x = delta,y=Five2] from \MainChetverikov; \addplot table[x = delta,y=MonFiveQKn0] from \SecondMonQ; \addplot table[x = delta,y=MonFiveQKn1] from \SecondMonQ; \addplot table[x = delta,y=MonFiveCKn0] from \SecondMonC; \addplot table[x = delta,y=MonFiveCKn1] from \SecondMonC; \addlegendentry{LSW-S}; \addlegendentry{LSW-L}; \addlegendentry{C-OS}; \addlegendentry{FS-Q0}; \addlegendentry{FS-Q1}; \addlegendentry{FS-C0}; \addlegendentry{FS-C1}; \nextgroupplot \node[anchor=north] at (axis description cs: 0.25, 0.95) {\fontsize{5}{4}\selectfont $\begin{aligned} d_z&=2\\ n&=750 \end{aligned}$}; \addplot[smooth,tension=0.5,color=NavyBlue, no markers,line width=0.25pt, densely dotted,forget plot] table[x = delta,y=alpha] from \FirstMon; \addplot table[x = delta,y=SevenUn] from \MainSecondLSWMon; \addplot table[x = delta,y=SevenOp] from \MainSecondLSWMon; \addplot table[x = delta,y=Seven2] from \MainChetverikov; \addplot table[x = delta,y=MonSevenQKn0] from \SecondMonQ; \addplot table[x = delta,y=MonSevenQKn1] from \SecondMonQ; \addplot table[x = delta,y=MonSevenCKn0] from \SecondMonC; \addplot table[x = delta,y=MonSevenCKn1] from \SecondMonC; \nextgroupplot \node[anchor=north] at (axis description cs: 0.25, 0.95) {\fontsize{5}{4}\selectfont $\begin{aligned} d_z&=2\\ n&=1000 \end{aligned}$}; \addplot[smooth,tension=0.5,color=NavyBlue, no markers,line width=0.25pt, densely dotted,forget plot] table[x = delta,y=alpha] from \FirstMon; \addplot table[x = delta,y=TenUn] from \MainSecondLSWMon; \addplot table[x = delta,y=TenOp] from \MainSecondLSWMon; \addplot table[x = delta,y=Ten2] from \MainChetverikov; \addplot table[x = delta,y=MonTenQKn0] from \SecondMonQ; \addplot table[x = delta,y=MonTenQKn1] from \SecondMonQ; \addplot table[x = delta,y=MonTenCKn0] from \SecondMonC; \addplot table[x = delta,y=MonTenCKn1] from \SecondMonC; \end{groupplot} \node at ($(myplots c2r1) + (0,-2.25cm)$) {(ref)}; \node at ($(myplots c2r2) + (0,-2.25cm)$) {(ref)}; \end{tikzpicture} \caption{Empirical power of our test (with $\gamma_n=0.01/\log n$), LSW-L, LSW-S, and C-OS for the designs (ref) and (ref), where corresponding to $\delta=0$ are the empirical sizes under D1. Note that FS-Q1 and FS-C1 nearly overlap each other in the second row.} \end{figure} { {5.25pt} \begin{table}[!ht] \caption{Run-times (in Seconds) of Monotonicity Tests} \sisetup{table-number-alignment = center, table-format = 1.2} \begin{tabularx}{\linewidth}{@cc *{5}{S[round-mode = places,round-precision = 2]} c *{5}{S[round-mode = places,round-precision = 2]} c@} \hline \hline &\multirow{2}{*}{$n$}& \multicolumn{5}{c}{Design (ref)} & & \multicolumn{5}{c}{Design (ref)} & \\ \cline{3-7} \cline{9-13} && {FS-C3} & {FS-C7} & {LSW-L} & {LSW-S} & {C-OS} & & {FS-Q0} & {FS-C1} & {LSW-L} & {LSW-S} & {C-OS} &\\ \hline &$500$ & 0.1236 & 0.1461 & 22.2861 & 22.2216 & 0.0822 & & 0.3984 & 0.4488 & 38.8922 & 39.2504 & 7.1872 &\\ &$750$ & 0.1280 & 0.1310 & 56.8640 & 56.6187 & 0.1566 & & 0.4126 & 0.5008 & 128.8489 & 128.1390 & 17.3866 &\\ &$1000$ & 0.1390 & 0.1558 & 103.5839 & 102.1905 & 0.3057 & & 0.3893 & 0.4221 & 219.2428 & 222.7889 & 29.4911 &\\ \hline \hline \end{tabularx} \end{table} } Finally, we compare the run-times of a single replication based on the design D1 in both univariate and bivariate cases. For brevity, we only report our tests with the smallest and the largest $k_n$, based on $\gamma_n=0.01/\log n$. All numbers in Table (ref) are obtained by running MATLAB R2019b in a Windows 10 PC with 16 GB RAM and an Intel\textsuperscript{\textregistered} Core\textsuperscript{\tiny TM} i7-7700 processor having 4 cores and 3.60 GHz base speed. Overall, Table (ref) shows that our tests are relatively simple to implement, in both univariate and bivariate settings, and their computation cost increases only modestly as we move from $d_z=1$ to $d_z=2$. We stress that the computational complexity of quadratic programs involved in our tests depends on the fineness of discretization, not the sample size. \section{Conclusion} In this paper, we have developed a uniformly valid and asymptotically nonconservative test for a class of shape restrictions, which is applicable to nonparametric regression models as well as parametric, distributional, and structural settings. The key insight we exploit is that these restrictions form convex cones in Hilbert spaces, a structure that enables us to employ a projection-based test whose properties may be analyzed in an elegant, transparent, and unifying way. In particular, while the problem is inherently nonstandard, we are able to develop a bootstrap procedure that may be implemented in a way as simple as computing the test statistic. \titleformat{\section}{\center}{{\sc Appendix} \Alph{section}}{1em} \setcounter{section}{0} \numberwithin{equation}{section} \numberwithin{ass}{section} \numberwithin{figure}{section} \numberwithin{table}{section} \section{Proofs of Main Results} {\sc Proof of Lemma (ref):} Part (i) is immediate because $h\mapsto\Pi_\Lambda(h)$ is positively homogeneous by Assumption (ref) and Theorem 5.6-(7) in Deutsch2012Best, while part (ii) follows by Lemma (ref), part (i), $\theta_P\in\Lambda$, and $\kappa_n\in[0,r_n]$. For part (iii), note that $\kappa_n\hat\theta_n-\kappa_n\theta_P=r_n \{\hat\theta_n-\theta_P\}\cdot (\kappa_n/r_n)$. By Assumption (ref)(i), $\sup_{P\in\mathbf P}E[\|\mathbb Z_{n,P}\|_{\mathbf H}]<\infty$ uniformly in $n$ and Markov's inequality, we have: uniformly in $P\in\mathbf P$, \begin{align} \| r_n \{\hat\theta_n-\theta_P\}\|_{\mathbf H}\le \| r_n \{\hat\theta_n-\theta_P\}-\mathbb Z_{n,P}\|_{\mathbf H}+\|\mathbb Z_{n,P}\|_{\mathbf H}=O_p(1) . \end{align} Part (iii) now follows from combining (ref) and $\kappa_n/r_n=o(c_n)$. \qed {\sc Proof of Theorem (ref):} We structure our proof in four steps. \underline{Step 1:} Build up a strong approximation for $ r_n\phi(\hat\theta_n)$ that is valid uniformly in $P\in\mathbf P$. We make use of the map $\psi_{a,P}$ defined in (ref), and let $\mathbb G_{n,P}\equiv r_n\{\hat\theta_n-\theta_P\}$. By Lemma (ref) and the definition of $\psi_{a,P}$, we may rewrite the test statistic: \begin{align} r_n\phi(\hat\theta_n)=\|\mathbb G_{n,P}+r_n\theta_P-\Pi_\Lambda(\mathbb G_{n,P}+r_n\theta_P)\|_{\mathbf H}=\psi_{r_n,P}(\mathbb G_{n,P}) . \end{align} By Theorem 3.16 in AliprantisandBorder2006 and Assumption (ref)(i), we have: \begin{align} |\psi_{r_n,P}(\mathbb G_{n,P})-\psi_{r_n,P}(\mathbb Z_{n,P})|\le \|\mathbb G_{n,P}- \mathbb Z_{n,P}\|_{\mathbf H}=o_p(c_n) , \end{align} uniformly in $P\in\mathbf P$. Combining results (ref) and (ref), we thus obtain the strong approximation for our test statistic: uniformly in $P\in\mathbf P$, \begin{align} r_n\phi(\hat\theta_n)=\psi_{r_n,P}(\mathbb Z_{n,P})+o_p(c_n) . \end{align} \underline{Step 2:} Build up a strong approximation of $ \hat\psi_{\kappa_n}(\hat{\mathbb G}_n)$ that is valid uniformly in $P\in\mathbf P_0$. First, Theorem 3.16 in AliprantisandBorder2006 implies: for each $P\in\mathbf P_0$, \begin{align} |\hat\psi_{\kappa_n}(\hat{\mathbb G}_n)- \psi_{\kappa_n,P}(\hat{\mathbb G}_n)| \le \kappa_n\|\Pi_\Lambda\hat\theta_n-\theta_P\|_{\mathbf H}\le \| \kappa_n \{\hat\theta_n-\theta_P\}\|_{\mathbf H} = o_p(c_n) , \end{align} where the second inequality is due to $\theta_P=\Pi_\Lambda(\theta_P)$ for all $P\in\mathbf P_0$, Assumption (ref)(i), Lemma 6.54-d in AliprantisandBorder2006, and the final step is due to Lemma (ref)(iii). Again by Theorem 3.16 in AliprantisandBorder2006, we have \begin{align} | \psi_{\kappa_n,P}(\hat{\mathbb G}_n)-\psi_{\kappa_n,P}(\bar{\mathbb Z}_{n,P})|\le \|\hat{\mathbb G}_n- \bar{\mathbb Z}_{n,P}\|_{\mathbf H}=o_p(c_n) , \end{align} uniformly in $P\in\mathbf P$, where the equality is due to Assumption (ref)(ii). We thus obtain by (ref) and (ref) and the triangle inequality that: uniformly in $P\in\mathbf P_0$, \begin{align} \hat\psi_{\kappa_n}(\hat{\mathbb G}_n) = \psi_{\kappa_n,P}(\bar{\mathbb Z}_{n,P}) +o_p(c_n) . \end{align} \underline{Step 3:} Control the estimation error of $\hat c_{n,1-\alpha}$. By results (ref) and (ref), we may select a sequence of positive scalars $\epsilon_n=o(c_n)$ (sufficiently slow) such that, as $n\to\infty$, \begin{gather} \sup_{P\in\mathbf P_0} P(|r_n\phi(\hat\theta_n)-\psi_{r_n,P}(\mathbb Z_{n,P})|>\epsilon_n)=o(1) , \\ \sup_{P\in\mathbf P_0} P(|\hat\psi_{\kappa_n}(\hat{\mathbb G}_n) - \psi_{\kappa_n,P}(\bar{\mathbb Z}_{n,P})|>\epsilon_n)=o(1) . \end{gather} By Markov's inequality, Fubini's theorem, and (ref), we have: for each $\eta>0$, \begin{multline} \sup_{P\in\mathbf P_0} P(P(|\hat\psi_{\kappa_n}(\hat{\mathbb G}_n) - \psi_{\kappa_n,P}(\bar{\mathbb Z}_{n,P})|>\epsilon_n|\{X_i\}_{i=1}^n)>\eta)\\ \le \sup_{P\in\mathbf P_0}\frac{1}{\eta} P(|\hat\psi_{\kappa_n}(\hat{\mathbb G}_n) - \psi_{\kappa_n,P}(\bar{\mathbb Z}_{n,P})|>\epsilon_n)=o(1) . \end{multline} Thus, we may select a sequence of positive scalars $\eta_n\downarrow 0$ such that \begin{align} P(|\hat\psi_{\kappa_n}(\hat{\mathbb G}_n) - \psi_{\kappa_n,P}(\bar{\mathbb Z}_{n,P})|>\epsilon_n|\{X_i\}_{i=1}^n)=o_p(\eta_n) , \end{align} uniformly in $P\in\mathbf P_0$. Since $\bar{\mathbb Z}_{n,P}$ is independent of $\{X_i\}_{i=1}^n$ by Assumption (ref)(ii), the conditional cdf of $\psi_{\kappa_n,P}(\bar{\mathbb Z}_{n,P})$ given $\{X_i\}_{i=1}^n$ is precisely its unconditional analog. Thus, we may conclude by Lemma 11 in ChernozhukovLeeRosen2013Intersection and result (ref) that \begin{align} \liminf_{n\to\infty}\inf_{P\in\mathbf P_0} P(\hat c_{n,1-\alpha}+\epsilon_n\ge c_{n,P}(1-\alpha-\eta_n))=1 . \end{align} \underline{Step 4:} Conclude with the help of a partial anti-concentration inequality. To begin with, note that by results (ref) and (ref), we have \begin{align} \limsup_{n\to\infty}\sup_{P\in\mathbf P_0}& P(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha})\notag\\ &\le \limsup_{n\to\infty} \sup_{P\in\mathbf P_0}P(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha},|r_n\phi(\hat\theta_n)-\psi_{r_n,P}(\mathbb Z_{n,P})|\le \epsilon_n)\notag\\ &\le \limsup_{n\to\infty}\sup_{P\in\mathbf P_0}P(\psi_{r_n,P}(\mathbb Z_{n,P})>c_{n,P}(1-\alpha-\eta_n)-2\epsilon_n)\notag\\ &\le \limsup_{n\to\infty}\sup_{P\in\mathbf P_0}P(\psi_{\kappa_n,P}(\mathbb Z_{n,P})>c_{n,P}(1-\alpha-\eta_n)-2\epsilon_n) , \end{align} where the final step follows by Lemma (ref) and $0\le\kappa_n\le r_n$ for all large $n$ (due to $\kappa_n/r_n=o(c_n)$ and $c_n=O(1)$). In turn, we note that, since $\eta_n\downarrow 0$ and $\epsilon_n=o(c_n)$, it follows by Assumption (ref)(iii), Proposition (ref), and result (ref) that \begin{multline} \limsup_{n\to\infty}\sup_{P\in\mathbf P_0} P(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha})\le \limsup_{n\to\infty}\sup_{P\in\mathbf P_0}P(\psi_{\kappa_n,P}(\mathbb Z_{n,P})>c_{n,P}(1-\alpha-\eta_n))\\ \le\limsup_{n\to\infty}\sup_{P\in\mathbf P_0}\{\alpha+\eta_n\}=\alpha , \end{multline} as desired for the first claim. For the second claim, it thus suffices to show \begin{align} \liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0} P(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha})\ge \alpha . \end{align} For this, we note that, by simple manipulations, \begin{align} \liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0}& P(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha})\notag\\ &\ge \liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0}P(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha},|r_n\phi(\hat\theta_n)-\psi_{r_n,P}(\mathbb Z_{n,P})|\le \epsilon_n)\notag\\ & = \liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0}P(\psi_{r_n,P}(\mathbb Z_{n,P})-\epsilon_n>\hat c_{n,1-\alpha}) , \end{align} where the last step is due to result (ref) and $\bar{\mathbf P}_0\subset \mathbf P_0$. Moreover, another application of Lemma 11 in ChernozhukovLeeRosen2013Intersection to (ref) yields \begin{align} \liminf_{n\to\infty}\inf_{P\in\mathbf P_0} P(\hat c_{n,1-\alpha}\le c_{n,P}(1-\alpha+\eta_n)+\epsilon_n)=1 . \end{align} By the definition of $\bar{\mathbf P}_0$ and Lemma (ref), we also note that, for all $n$ and $P\in\bar{\mathbf P}_0$, \begin{align} \psi_{r_n,P}(\mathbb Z_{n,P}) = \psi_{\kappa_n,P}(\mathbb Z_{n,P})=\|\mathbb Z_{n,P}-\Pi_\Lambda(\mathbb Z_{n,P})\|_{\mathbf H} . \end{align} Combining results (ref), (ref), and (ref) with Proposition (ref) yields \begin{align} \liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0}& P(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha})\notag\\ &\ge \liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0}P(\psi_{r_n,P}(\mathbb Z_{n,P})-\epsilon_n>c_{n,P}(1-\alpha+\eta_n)+\epsilon_n)\notag\\ & = \liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0}P(\psi_{\kappa_n,P}(\mathbb Z_{n,P})>c_{n,P}(1-\alpha+\eta_n)+2\epsilon_n)\notag\\ & = \liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0}P(\psi_{\kappa_n,P}(\mathbb Z_{n,P})>c_{n,P}(1-\alpha+\eta_n)-2\epsilon_n) . \end{align} Since $c_{n,P}(1-\alpha+\eta_n)-\epsilon_n< c_{n,P}(1-\alpha+\eta_n)$, we thus obtain by result (ref), the definition of quantiles, and $\eta_n=o(1)$ that \begin{align} \liminf_{n\to\infty}\inf_{P\in\bar{\mathbf P}_0} P(r_n\phi(\hat\theta_n)>\hat c_{n,1-\alpha})\ge \liminf_{n\to\infty}\{\alpha-\eta_n\}=\alpha , \end{align} which, together with the first claim, establishes the second claim of the theorem. \qed {\sc Proof of Theorem (ref):} First, by Assumption (ref) and Lemma (ref), we have \begin{align} \|\hat{\mathbb G}_n+\kappa_n\Pi_\Lambda\hat\theta_n &-\Pi_\Lambda(\hat{\mathbb G}_n+\kappa_n\Pi_\Lambda\hat\theta_n)\|_{\mathbf H} \le \|\hat{\mathbb G}_n-\Pi_\Lambda\hat{\mathbb G}_n\|_{\mathbf H} \notag\\ &\le \|\hat{\mathbb G}_n\|_{\mathbf H}\le \|\hat{\mathbb G}_n-\bar{\mathbb Z}_{n,P}\|_{\mathbf H}+\|\bar{\mathbb Z}_{n,P}\|_{\mathbf H} , \end{align} where the second inequality follows from Assumption (ref) and Theorem 5.6(5) in Deutsch2012Best, and the third inequality is due to the triangle inequality. By Assumptions (ref) and (ref)(ii), we in turn have from (ref) that, uniformly in $P\in\mathbf P$, \begin{align} \|\hat{\mathbb G}_n+\kappa_n\Pi_\Lambda\hat\theta_n-\Pi_\Lambda(\hat{\mathbb G}_n+\kappa_n\Pi_\Lambda\hat\theta_n)\|_{\mathbf H} =o_p(c_n)+O_p(1)=O_p(1) . \end{align} By the definition of $\hat c_{n,1-\alpha}$, we note that, for $M>0$ and uniformly in $P\in\mathbf P$, \begin{align} P(\hat c_{n,1-\alpha}>M)&\le P(P(\|\hat{\mathbb G}_n+\kappa_n\Pi_\Lambda\hat\theta_n-\Pi_\Lambda(\hat{\mathbb G}_n+\kappa_n\Pi_\Lambda\hat\theta_n)\|_{\mathbf H}>M|\{X_i\}_{i=1}^n)>\alpha)\notag\\ &\le\frac{1}{\alpha} P(\|\hat{\mathbb G}_n+\kappa_n\Pi_\Lambda\hat\theta_n-\Pi_\Lambda(\hat{\mathbb G}_n+\kappa_n\Pi_\Lambda\hat\theta_n)\|_{\mathbf H}>M) , \end{align} where the second inequality holds by Markov's inequality and Fubini's theorem. It follows from results (ref) and (ref) that $\hat c_{n,1-\alpha}=O_p(1)$ uniformly in $P\in\mathbf P$. Next, we bound $r_n\phi(\hat\theta_n)$ from below. By Theorem 3.16 in AliprantisandBorder2006 and the triangle inequality, we have: uniformly in $P\in\mathbf P$, \begin{multline} |r_n\phi(\hat\theta_n)-r_n\phi(\theta_P)|\le \|r_n\{\hat\theta_n-\theta_P\}\|_{\mathbf H}\\ \le \|r_n\{\hat\theta_n-\theta_P\}-\mathbb Z_{n,P}\|_{\mathbf H}+ \|\mathbb Z_{n,P}\|_{\mathbf H}\le o_p(c_n)+O_p(1)=O_p(1) , \end{multline} where the third inequality follows by Assumptions (ref)(i) and (ref)(ii), and the last step is due to $c_n=O(1)$. It follows from result (ref) and the definition of $\mathbf P_{1,n}^\Delta$ that \begin{align} r_n\phi(\hat\theta_n)=r_n\phi(\theta_P)+r_n\phi(\hat\theta_n)-r_n\phi(\theta_P)\ge \Delta+O_p(1) , \end{align} uniformly in $P\in\mathbf P_{1,n}^\Delta$. The theorem thus follows from combining result (ref) and the order $\hat c_{n,1-\alpha}=O_p(1)$ uniformly in $P\in\mathbf P$ that we have established .\qed \titleformat{\section}{\normalfont}{\Alph{section}}{1em} \addcontentsline{toc}{section}{References} \putbib
bibunit