EconBase
← Back to paper

Testing nonparametric shape restrictions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

131,623 characters · 16 sections · 137 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

TESTING NONPARAMETRIC SHAPE RESTRICTIONS

\address{Economics Department\\ London School of Economics\\ Houghton Street\\ London WC2A 2AE\\ U.K.} \email{[email removed]} \urladdr{http://www.authortwo.twouniv.edu} \address{Economics Department\\ London School of Economics\\ Houghton Street\\ London WC2A 2AE\\ U.K.} \email{[email removed]} \subjclass[2000]{Primary 05C38, 15A15; Secondary 05A15, 15A18}

abstractWe describe and examine a test for a general class of shape constraints, such as constraints on the signs of derivatives, U-(S-)shape, symmetry, quasi-convexity, log-convexity, $r$-convexity, among others, in a nonparametric framework using partial sums empirical processes. We show that, after a suitable transformation, its asymptotic distribution is a functional of the standard Brownian motion, so that critical values are available. However, due to the possible poor approximation of the asymptotic critical values to the finite sample ones, we also describe a valid bootstrap algorithm.

INTRODUCTION

Hypothesis testing is one of the most relevant tasks in empirical work. Such tests include situations when the null and alternative hypothesis are assumed to belong to a parametric family of models. In a second type of tests, known as diagnostic or lack-of-fit tests, only the null hypothesis is assumed to belong to a parametric family leaving the alternative nonparametric. The latter type of testing has a distinguished and long literature starting with the work of Kolmogorov, see Stephens Stephens, for testing the probability distribution function and in a time series context by Grenander and Rosenblatt GrenanderRos for testing the white noise hypothesis. In a regression model context a new avenue of work started in Stute Stute97, and Andrews Andrews97 with a more econometric emphasis, using partial sums empirical methodology, see also Stute et al. Stute_etal98b or Koul and Stute KoulStute99, among others. The methodology has attracted a lot of attention and it rivals tests based on a direct comparison between a parametric and a nonparametric fit to the regression model as first examined in H\"{a}rdle and Mammen HardleMammen or Hong and White HongWhite. One advantage of tests based on partial sums empirical methodology, when compared to the approach in HardleMammen, is that the former does not require the choice of a bandwidth parameter for its implementation, see also Nikabadze and Stute NikabadzeStute for some additional advantages. However, a possible drawback is that the asymptotic distribution depends, among other possible features, on the estimator of the parameters of the model under the null hypothesis in a nontrivial way, as it was shown in Durbin Durbin73, and hence its implementation requires either bootstrap algorithms, see Stute et al. Stute_etal98a, or the so-called Khmaladze's Khmaladze81 martingale transformation, see also earlier work and ideas in Brown at al. BrownDurbinEvans.

In this paper, though, we are interested in a third type of testing where neither the null hypothesis nor the alternative has a specific parametric form. This type of hypothesis testing can be denoted as testing for qualitative or shape restrictions. Many shape properties (including monotonicity, convexity/concavity, strong convexity, log-convexity) are widespread in economics and other disciplines. It is also often of interest to analyse statistical and economic relationships that do not have a persistent shape pattern on the whole domain but rather switch the patterns in the domain (for example, U-shaped or S-shaped relations).

To fix ideas, consider the nonparametric regression model

align[align omitted — 100 chars of source]

where $x$ has bounded support $\mathcal{X=}:\left[ \underline{x},\overline{x} \right] $ and $m\left( \cdot \right) $ is smooth. More specific conditions on the sequences $\left\{ u_{i}\right\} _{i\in \mathbb{Z}}$ and $\left\{ x_{i}\right\} _{i\in \mathbb{Z}}$ will be given in Condition $C1$ in Section (ref). Our main aim is testing whether the regression function $m\left( x\right) $ possesses shape properties captured by a general null hypothesis

equation[equation omitted — 72 chars of source]

where the class of interest $\mathcal{M}_{0}$ is a subset of smooth functions from $\mathcal{X}$ to $\mathbb{R}$. Our requirement on the class $ \mathcal{M}_{0}$ is given in Condition $C0$ below, which is formulated in a way to directly relate it to the methodology we employ later to approximate and estimate $m(\cdot )$.

Before we introduce Condition $C0$, let us introduce some useful notations. Let $\mathcal{B}_{q,L}$ denote the set of all B-splines of degree $q$ with knots that split $\mathcal{X}$ into $L^{\prime }$ equally spaced intervals.\footnote{ As it will be clear from our exposition later, the condition that these intervals are equally spaced is not important and is only imposed for the simplicity of the exposition. We only need that the system of knots has to become increasingly dense in $\mathcal{X}$. Also, in Section (ref) we shall motivate the choice of B-splines in comparison to other nonparametric estimation techniques.} A generic element in this set is written as a linear combination $m_{\mathcal{B}}(x)\equiv \sum_{\ell =1}^{L}\beta _{\ell }p_{\ell ,L}\left( x\right) $, where $L=L^{\prime }+q$, and $\left\{ p_{\ell ,L}\left( \cdot \right) \right\} $ is the collection of the B-splines base for the chosen system of knots (more details will be given in Section (ref)). Any element in $\mathcal{B}_{q,L}$ can be fully characterised by the vector $(\beta _{1},\ldots ,\beta _{L})\in \mathbb{R} ^{L}$ and constraints on this vector can be mapped into constraints on the B-spline. For any set $S_{q,L}\subseteq \mathbb{R}^{L}$ we can, thus, define

equation*[equation* omitted — 139 chars of source]

\vskip 0.1in

Condition C0. There is a set $S_{q,L} \subseteq \mathbb{R}^{L} $ for any $L'$ (and, hence, for any $L=L'+q$) that satisfies the following properties:

itemize$S_{q,L}$ does not depend on data $\left\{ x_{i}\right\} _{i\in \mathbb{Z}}$ and, thus, is non-stochastic; • the boundary of $S_{q,L}$ consists of a finite number of smooth surfaces; • \begin{equation} \mathcal{H}\left( \mathcal{M}_0,\mathcal{M}_{S_{q,L}}\right) \rightarrow 0\quad as\quad L\rightarrow \infty , \end{equation} where $\mathcal{H}$ is the Hausdorff distance in the supremum norm in the space of continuous functions from $\mathcal{X}$ to $\mathbb{R}$.

}

\vskip 0.1in

Condition $C0$ essentially states that the membership in $\mathcal{M}_0$ can be captured by restrictions on parameters $\{\beta_{\ell}\}_{\ell=1}^L$ which become necessary and sufficient restrictions as the system of knots becomes increasingly dense in $\mathcal{X}$. In addition, these restrictions on $\{\beta_{\ell}\}_{\ell=1}^L$ do not depend on the available data, which adds to the attractiveness of our approach for implementation purposes.

Our idea will be to test the null hypothesis

equation[equation omitted — 97 chars of source]

formulated in terms of the approximation for $m(\cdot )$. As may be expected, our methodology readily extends to situations when the domain $ \mathcal{X}$ is partitioned into several intervals with different null hypotheses in the spirit of ((ref)) formulated on different intervals in the partitioning. An illustration of this situation is given in Example (ref) below.

We shall now illustrate several examples of shape classes $\mathcal{M}_{0}$ that satisfy condition $C0$. In Section (ref), after a more thorough discussion of B-splines we indicate the corresponding sets $ S_{q,L}$ for each of these examples. As we shall discuss in Section (ref), the first two examples pertain to the case where in the constraints on the coefficients of the B-splines approximation are linear\footnote{Quasi-convexity in Example (ref) may be an exception depending on how exactly one approaches the testing, as is clear from further discussion in Section (ref).}, whereas in the last two examples these constraints are nonlinear. It is important to note that these examples are an illustration of the scope of our approach rather than a full list of properties that our methodology could test.

example[Constraints, may hold simultaneously, on the derivatives of $ m\left( \cdot \right) $] \begin{equation} H_{0}:d_{r}\cdot m^{\left( r\right) }\left( x\right) \geq c_{r}, \ \ \ \ r\in R, \end{equation} where $R$ is a finite subset of $\mathbb{N}^{+}$, $d_{r}\in \left\{ -1,1\right\} $ and $c_{r}$ are known constants. In this hypothesis, we specify conditions on inequalities for several derivatives simultaneously. If $c_{r}=0$, then we have familiar restrictions on the sign of the $r$-th derivative. Special cases in this example include testing for monotonicity ($r=1$ and $c_{1}=0$), testing for convexity/concavity ($r=2$ and $c_{2}=0$), testing for strong $ \lambda $-convexity ($r=2$ and $c_{2}=\lambda >0$), testing for monotonicity and concavity simultaneously, etc.

The properties of monotonicity, convexity/concavity or the ones in terms of higher-order derivatives are so commonplace in economics and other disciplines that the applications are too many to list. For instance, a demand function is expected to be a decreasing function of the price of a good, whereas a supply function is expected to be increasing. In single-object auctions, the equilibria analyzed commonly are those in which buyers play monotone strategy functions and mark-ups are often monotone functions of the bids (see e.g. Krishna). In other economic relationships it is often of importance whether the marginal returns are increasing or decreasing, which naturally amounts to convexity and concavity, respectively.

Note that in the context of isotone/monotone regressions the literature goes back to Brunk Brunk and Wright Wright, and in the context of convex regressions to Hildreth Hildreth. A more detailed coverage of this literature is given in Section (ref).

As mentioned above, our methodology can be used when different properties can be conjectured on different intervals of the partition $\mathcal{X}$. This is illustrated in the next example.

example[changing shape patterns on the domain; symmetry; quasi-convexity] In this case, we can take a partitioning of $\mathcal{ X}$ at points $\tau _{0}=\underline{x}<\tau _{1}<\ldots <\tau _{J-1}<\tau _{J}=\overline{x}$, and formulate the null hypothesis on the signs of various derivatives on the intervals in this partitioning. The null hypothesis in this case would be \begin{align} H_{0}& :d_{r_{j}}\cdot m^{\left( r_{j}\right) }\left( x\right) \geq 0,\quad x\in \left( \tau _{j},\tau _{j+1}\right] , \\ & for some fixed r_{j}\in \mathbb{N},\quad j=0,\ldots ,J-1, \notag \end{align} $d_{r_{j}}\in \left\{ -1,1\right\} $.\footnote{ It goes without saying that we can have several constraints on each $[\tau _{j},\tau _{j+1}],\;j=0,\ldots ,J-1$.} The partitioning points $\tau _{j}$ either have to be known or have to be consistently estimated at a suitable rate before the testing procedure. This type of hypotheses include, among others, the important scenarios of $U$ -shape, $S$-shape. For example, U-shape is the property of a function first decreasing and then increasing. So, we write the null hypothesis of U-shape with the switch at $s_{0}$ as \begin{equation} H_{0}:\text{ \ }\left\{ \begin{array}{c} m\left( x^{1}\right) >m\left( x^{2}\right) \quad \text{ when }\quad x^{1}<x^{2}\leq s_{0} \\ m\left( x^{1}\right) >m\left( x^{2}\right) \quad \text{ when }\quad x^{1}>x^{2}\geq s_{0}. \end{array} \right. \end{equation} The issues associated with testing properties outlined in this example relate to how functional pieces from different partition intervals are connected to each other. We can incorporate various scenarios -- e.g., we can require that different pieces are joint continuously, or smoothly at all or some partition points. It goes without saying that we can consider several different partitions and test simultaneously properties formulated on various intervals of these partitions. For example, one could consider a function defined on $[0,1]$ and conjecture that $m(\cdot)$ is convex on $(0,s_0$, concave on $(s_0,1]$, decreasing on $(0,0.3s_0]$, increasing on $(0.3s_0, 0.5s_0+0.5]$, and decreasing on $(0.5s_0, 1]$. Examples of \emph{U-shaped} relationships in economics and other disciplines can be found e.g. in CalabreseBaldwin, Corrao_etal, Goldin, Groes_etal. SuttonTrefler, Weiman, and in the discussions in Kostyshak, LindMelhum and Simonsohn. Inverse \emph{U-shaped} relationships include the case of the so-called single-peaked preferences, which is an important class of preferences in psychology and economics. Examples of \emph{S-shaped} relationships between unexpected earnings and stock price in accounting can be found in DasLev or FreemanTse, among others. \emph{ S-shaped} growth curves of the adopted population in a large society is a generally accepted empirical feature of innovation diffusion (see discussions in Rogers or Utterback). Thus, testing for an \emph{S-shape} in this case would allow one to conclude whether technology evolves as one would expect. Another shape property in this class of shape-changing patterns are the properties of \textbf{ quasi-convexity} and \textbf{ quasi-concavity}. Formally, the class $\mathcal{M}_{0}$ in ((ref)) of quasi-convex functions is defined as $$ \mathcal{M}_{0}=\bigg\{ \phi (\cdot ):\;\forall \,x_{1},x_{2}\in \mathcal{X} \; \; \forall \lambda \in \lbrack 0,1] \; \; \phi (\lambda x_{1}+(1-\lambda )x_{2})\leq \max {\big \{}\phi (x_{1}),\phi (x_{2}){\big \}} \bigg\}. $$ Since function $\phi $ is \textbf{quasi-concave} if and only if $-\phi $ is quasi-convex, we can easily write a description of the class of quasi-concave functions. A smooth function is quasi-convex if and only if it first decreases up to some point and then increases (giving a special case of monotonicity when such a switch point is located at one of the boundary points of the interval). Since this switch point may not be known a priori, it would have to be estimated. A further discussion of this is given later in Section (ref). The concept of quasi-convexity/quasi-concavity is quite important in economics and other disciplines. For example, it is known that indirect utility functions are quasi-convex provided that the direct utility function is continuous. Thus strict quasi-concavity (and continuity) of the production function provides a sufficient condition for the differentiability of the cost function with respect to input prices. Finally, we mention another property related to this framework, which is the \textbf{symmetry} of a function around some point $s_0$. Without a loss of generality, we can take $s_0$ to be the center of the domain $[\underline{x}, \overline{x}]$\footnote{For a given $s_0$, one can only test for symmetry only on the interval $[s_0 - \min\{\overline{x}-s_0, s_0-\underline{x}\}, s_0+\min\{\overline{x}-s_0, s_0-\underline{x}\}]$, which is centered around $s_0$.} and write the class of symmetric functions as \begin{equation*} \mathcal{M}_{0}=\left\{ \phi (\cdot ):\;\forall \,x \in [\underline{x},s_0] \quad \phi (s_0-x)=\phi (s_0+x) \right\}. \end{equation*}
example[$r$-convexity; $\protect\rho $-convexity] The null hypothesis of $\mathbf{r}$-convexity, $r \neq 0$, is described by the following $\mathcal{M}_0$; \begin{multline*} \mathcal{M}_0 = \bigg\{ \phi(\cdot): \; \phi(\cdot) \geq 0, \quad \forall \, x_1, x_2 \in \mathcal{X} \quad \forall \lambda \in [0,1] \\ \phi(\lambda x_1+(1-\lambda )x_2)\leq \log \left(\lambda e^{r\phi(x_1)}+(1-\lambda) e^{r\phi(x_2)}\right)^{\frac{1}{r}} \bigg\}, \end{multline*} and for $r=0$ this property is described by \begin{multline*} \mathcal{M}_0 = \bigg\{\phi(\cdot): \; \phi(\cdot) \geq 0, \quad \forall \, x_1, x_2 \in \mathcal{X} \quad \forall \lambda \in [0,1] \\ \phi(\lambda x_1+(1-\lambda )x_2)\leq \lambda \phi(x_1) + (1-\lambda) \phi(x_2) \bigg\}. \end{multline*} The definition of $\mathbf{r}$-concave functions would reverse the inequalities. If we consider functions that are twice continuously differentiable, then the class of $r$-convex functions can be described as (see e.g. Avriel Avriel72) \begin{equation*} \mathcal{M}_0 = \bigg\{\phi(\cdot): \; \phi(\cdot) \geq 0, \quad \forall \, x \in \mathcal{X} \quad r \left(\phi^{\prime }(x) \right)^2 + \phi^{{\prime \prime }}(x) \geq 0 \bigg\}. \end{equation*} A concept closely related to $r$-convexity/concavity is what is known as $ \mathbf{\rho}$-convexity/concavity in the economics literature. Following CaplinNalebuff91a, we can define the class of $\rho$-convex functions as follows: for $\rho \neq 0$, \begin{multline*} \mathcal{M}_0 = \bigg\{ \phi(\cdot): \; \phi(\cdot) > 0, \quad \forall \, x_1, x_2 \in \mathcal{X} \quad \forall \lambda \in [0,1] \\ \phi(\lambda x_1+(1-\lambda )x_2)\leq \left(\lambda \phi(x_1)^{\rho}+(1-\lambda) \phi(x_2)^{\rho}\right)^{\frac{1}{\rho}} \bigg\} , \end{multline*} and for $\rho=0$, \begin{multline*} \mathcal{M}_0 = \bigg\{ \phi(\cdot): \; \phi(\cdot) > 0, \quad \forall \, x_1, x_2 \in \mathcal{X} \quad \forall \lambda \in [0,1] \\ \phi(\lambda x_1+(1-\lambda )x_2)\leq \phi(x_1)^{\lambda} \phi(x_2)^{1-\lambda}\bigg\}. \end{multline*} The definition of $\mathbf{\rho}$-concave functions would reverse the inequalities. It is obvious that for any $r=\rho$, a real-valued function $\phi$ is $r$-convex if and only if $e^{\phi}$ is $\rho$-convex. It is also clear that $\rho$-convexity gives the standard definition of convexity when $\rho=1$ and gives log convexity for $\rho=0$. In economics, $\rho $-concavity has been used to measure the curvature of demand functions, production functions, and distributions (i.e., density and distribution functions) of individual characteristics. It has also been employed in the areas of imperfect competition, auctions and mechanism design and public economics.\footnote{ See, for instance, AndersonRenault03, CaplinNalebuff91b, Cowan, Ewerhart, Moyes, among others.} Log-convexity and log-concavity are of particular interest in economics.

Our final example illustrates further shape properties well developed in the mathematical literature. This example first requires defining a mean function.

definition[mean function; Niculescu Niculescu] A function $N: \mathbb{R}^{+} \times \mathbb{R}^{+} \rightarrow \mathbb{R}^{+}$ is called a mean function if \begin{itemize} • $N(x_1, x_2) =N(x_2, x_1)$$N(x, x) =x$$x_1<N(x_1, x_2) < x_2$ whenever $x_1<x_2$$N(ax_1, ax_2) =aN(x_1, x_2)$ for all $a>0$ \end{itemize}

Examples of mean functions include the arithmetic mean (A), the geometric mean (G), the harmonic mean (H), the logarithmic mean and the identric mean.

example[$MN$-convexity] For any two mean functions $M$ and $N$, the class of $\mathbf{MN}$- convex is defined as \begin{equation*} \mathcal{M}_0 = \bigg\{ \phi(\cdot): \; \phi(\cdot) > 0, \quad \forall \, x_1, x_2 \in \mathcal{X} \quad \phi(M(x_1, x_2)) \leq N(\phi(x_1), \phi(x_2)) \bigg\}. \end{equation*} The definition of $\mathbf{MN}$-concavity would reverse the inequalities. When we have different combinations of arithmetic (A), geometric (G) and harmonic (H) means, we end up with the following special cases of ${MN}$ -convex functions (see e.g. Anderson_etal): \begin{enumerate} • $m$ is AG-convex if and only if $\frac{m^{\prime }(x)}{m(x)}$ is increasing • $m$ is AH-convex if and only if $\frac{m^{\prime }(x)}{m^2(x) }$ is increasing • $m$ is GA-convex if and only if $x m^{\prime }(x)$ is increasing • $m$ is GG-convex if and only if $\frac{x m^{\prime }(x)}{m(x) }$ is increasing • $m$ is GH-convex if and only if $\frac{x m^{\prime }(x)}{ m^2(x)}$ is increasing • $m$ is HA-convex if and only if $x^2 m^{\prime }(x)$ is increasing • $m$ is HG-convex if and only if $\frac{x^2 m^{\prime }(x)}{ m(x)}$ is increasing • $m$ is HH-convex if and only if $\frac{x^{2}m^{\prime}(x)}{ m^{2}(x)}$ is increasing. \end{enumerate}

\vskip 0.1in

We want to emphasize again that the examples given in this introduction are meant to illustrate the scope of applicability of our testing methodology rather than give an exhaustive list of potential application.

\vskip 0.1in

LITERATURE REVIEW ON TESTING FOR SHAPES

$\left. {}\right. $

There is a range of the literature related to testing for monotonicity or convexity such as Bowman_etal, Hall and Heckman HallHeckman, Ghosal et al. GhosalSenVV, Juditsky and Nemirovski JudNem, Wang and Meyer WangMeyer, Chetverikov Chetverikov, Schlee Schlee, Diack and Thomas-Agnan DiackTA, Dumbgen and Spokoiny DumbgenSpok, Abrevaya and Jiang AbrevayaJiang, Baraud et al. Baraud_etal, Dette et al. DetteHN.

Some of the work referenced above (such as Baraud_etal, DumbgenSpok, JudNem) focuses on the regression function in the ideal Gaussian white noise model, while our framework does not require such assumptions. In the aforementioned literature some papers, such as Baraud_etal, DiackTA or HallHeckman, allow the explanatory variable to take only deterministic values. AbrevayaJiang, Chetverikov and GhosalSenVV treat the explanatory variable as random but either require its full stochastic independence with the unobserved regression error (Chetverikov, GhosalSenVV) or require the distribution of the error to be symmetric conditional on the explanatory variable, see AbrevayaJiang, whereas we require a weaker mean independence of error condition. Some of the above-mentioned papers, such as Bowman_etal, HallHeckman or WangMeyer, do not give any asymptotic theory. Many of the approaches are tailored to a specific type of shape and their extensions to more general shape properties do not appear straightforward, if at all possible. Moreover, these tests are often targeted to detecting specific deviations from the null hypothesis. The violations of the null hypothesis, however, can come from different sources. A parametric test for U-shapes was suggested in LindMelhum, whereas Kostyshak proposes a non-parametric test based on critical bandwidth giving sufficient conditions for its consistency. To the best of our knowledge, there is no formal approach in the literature to testing S-shapes or general shape properties formulated on intervals in a partitioning.

Also, to the best of our knowledge, this paper is the first to give a unified framework and a nonparametric methodology for testing (possibly simultaneously) a wide range of shape/qualitative constraints. Even for shape properties described in terms of the signs of the derivatives of the function, this methodology will, in particular, overcome some of the problems discussed above. But, of course, we want to make it clear that its applicability goes beyond such properties.

To be more specific, we propose a test based on a transformation, introduced in Khmaladze Khmaladze81, of the partial sums empirical process similar to that in Stute Stute97. Some of the properties of the test is that it converges to a standard Brownian motion, so that critical values of standard functionals such as Kolmogorov-Smirnov, Cram\'{e}r -von-Mises or Anderson-Darling are readily obtained. As a consequence, our testing procedure has the same asymptotic distribution regardless of the shape constraints under consideration.

Another feature of our testing procedure is its flexibility as it is able to test simultaneously for more than one shape constraint -- for instance, one could test monotonicity and log-convexity simultaneously. Finally, the test is easy to implement as it requires no more than \textquotedblleft recursive\textquotedblright\ least squares.

The remainder of the paper is organized as follows. Section (ref) introduces and motivates our nonparametric estimator of $ m\left( x\right) $ and it compares the methodology against rival nonparametric estimators. It also describes the form that the sets $S_{q,L}$ in Condition $C0$ take for the examples given in the introduction. In Section (ref), we describe the test and examine its statistical properties. Because the Monte-Carlo experiment suggests that the asymptotic critical values are not a good approximation of the finite sample ones, Section (ref) introduces a bootstrap algorithm. Section (ref) presents a Monte-Carlo experiment and some empirical examples, whereas Section (ref) concludes with a summary and possible extensions of the methodology. The proofs are confined to the Appendix A which employs a series of lemmas given in Appendix B.

METHODOLOGY AND REGULARITY CONDITIONS

Before we present the procedure for testing the hypothesis given in ((ref)), we introduce and motivate the type of nonparametric estimator that we shall use to estimate the regression function $m\left( x\right) $ under $H_{0}$ and/or $H_{1}$.

When testing for the null hypothesis of either monotonicity or convexity, several nonparameteric estimators have been considered in the literature. As mentioned in the introduction, early work on isotone/monotone regressions is Brunk Brunk and Wright Wright. Friedman and Tibshirani FT combines local averaging and isotonic regression, Mukerjee Mukerjee and Mammen Mammen91a propose a two-step approach, which involves smoothing the data by using kernel regression estimators in the first step, and then deals with the isotonization of the estimator by projecting it into the space of all isotonic functions in the second step. Hall and Huang HallHuang proposes an alternative method based on tilting estimation which preserves the optimal properties of the kernel regression estimator. Finally, a different approach to isotonization is based on rearrangement methods using a probability integral transformation, see Dette et al. DetteNP or Chernozhukov et al. ChernozhukovFVG. When the null hypothesis is that of convexity, Hildreth Hildreth approach is based on estimating $m\left( \cdot \right) $ by least squares approach. Consistency, rate of convergence properties and pointwise convergence results for this estimator are established, respectively, in HansonPledger, Mammen91b and GroeneboomJW. The global behaviour of such an estimator is examined in GuntuboyinaSen. Birke and Dette BirkeDette examine an estimator based on first obtaining unconstrained estimate of the derivative of the regression function which is isotonized and then integrated.

Our main aim is to propose a testing methodology which is not only applicable for a wide range of shape properties but also able to perform valid statistical inferences. The narrowness of the above mentioned techniques would make it extremely difficult to make them work in our more general context. Besides, the implementation of these techniques is not always trivial and/or they often lack any asymptotic theory useful for the purpose of inference.

A different approach might be based on series estimation using polynomial and, in particular, Bernstein polynomials basis, see the survey by Chen ChenHandbook, for results on sieve estimation in both nonparametric and semiparametric models. One motivation for the Bernstein polynomials is due to a well-known result that any continuous function can be approximated in the sup-norm by a Bernstein polynomial whose coefficients are the values of the function at uniform knot points. This property makes them an appealing candidate estimation method in our context. However, base Bernstein polynomials have an undesirable property of being highly correlated, making them difficult for our purposes -- in particular, to obtain a valid Khmaladze's transformation, which plays a key role in our results. This is discussed in more detail in Section (ref).

Due to many of the aforementioned reasons, in this paper we shall employ B-splines and/or penalized B-splines known as P-splines. Our testing will rely on the ability to write our null hypothesis in terms of the coefficients of the B-spline approximation, leading us to utilizing Condition $C0$ and testing ((ref)). Note that monotone regression splines (closely related to B-splines) were introduced by Ramsay Ramsay and later extended by Meyer Meyer to estimate convex/concave function or functions that are, e.g., both monotone and convex. A distinct feature in our approach is, first, that we now take the use of B-splines to a different level of generality and, second, that we derive the asymptotic theory as the number of knots goes to infinity, and, third, that we allow the number of constraints on the coefficients of \emph{B-splines} to increase with the number of knots. Both Ramsay and Meyer consider the number of knots, and hence the number of constraints, to be fixed.

B-SPLINES (or P-SPLINES)

$\left. {}\right. $

B-splines or P-splines are constructed from polynomial pieces joined at some specific points denoted knots. Their computation is obtained recursively, see deBoor, for any degree of the polynomial. It is well understood that the choice of the number of knots determines the trade-off between overfitting when there are too many knots, and underfitting when there are too few knots. The main difference between B-splines and P-splines is that the latter tend to employ a large number of knots but to avoid oversmoothing they incorporate a penalty function based on the second difference of the coefficients of adjacent B-splines, in contrast to the second derivative employed in OSullivan86, OSullivan88, see EilersMarx for details.

The methodology and applications of constrained B-splines and P-splines are discussed by many authors, too many to review here. For more detailed discussions, see, among others, the monographs deBoor and Dierckx for B-splines and EilersMarx, BollaertsEM for P-splines. Some literature on shape-preserving splines (for ordinary shapes such as monotonicity, convexity, etc.) includes, among others, LiNaikSwetits, MammenTA, Meyer and Ramsay.

In general, the B-spline basis of degree $q$

itemize• takes positive values on the domain spanned by $q+2$ adjacent knots, and is zero otherwise; • consists of $q+1$ polynomial pieces each of degree $q$, and the polynomial pieces join at $q$ inner knots; • at the joining points, the $q-1$th derivatives are continuous; • except at the boundaries, it overlaps with $2q$ polynomials pieces of its neighbours; • at a given $x$, only $q+1$ B-splines are nonzero.

Assume that one is interested in approximating the regression function $m\left( x\right) $ in the interval $\left[\underline{x},\overline{x}\right]$. Then we split the interval $\left[\underline{x},\overline{x}\right]$ into $L^{\prime }$ equal length subintervals with $L^{\prime }+1$ knots\footnote{ Although it is possible to have nonequidistant subintervals, for simplicity we consider equally spaced knots. An alternative way to locate the knots is based on the quantiles of the $x$ distribution.}, where each subinterval will be covered with $q+1$ B-splines of degree $q$. The total number of knots needed will be $L^{\prime }+2q+1$ (each boundary point $\underline{x}$, $\overline{x}$ is a knot of multiplicity $q+1$) and the number of B-splines is $ L=L^{\prime }+q$. Thus, $m\left( x\right) $ is approximated by a linear combination of B-splines of degree $q$ with coefficients $(\beta _{1},\ldots ,\beta _{L})$ as

equation[equation omitted — 135 chars of source]

and where henceforth we shall denote the knots as $\{z^{k}\}$, $ k=1,\ldots,L^{\prime }+2q+1$, where $\underline{x}=z^{1}=\ldots =z^{q+1}$ and $ \overline{x}=z^{L^{\prime }+q+1}=\ldots =z^{L^{\prime }+2q+1}$.

B-splines share some properties which turns out to be very useful for our purpose. Indeed,

eqnarray*[eqnarray* omitted — 373 chars of source]

where $\triangle \beta _{\ell }=\beta _{\ell }-\beta _{\ell -1}$. In particular, $\left( a\right) $ indicates that B-splines, as is the case with Bernstein polynomials, are a partition of $1$. The property (b) states that the derivative of a B-spline of degree $q$ becomes a B-spline of degree $q-1$. Using this expression for the derivative and taking into account that the knot system effectively changes with the first and the last knots now removed (thus, the multiplicity of $\underline{x}$ and $\overline{x}$ becomes $q$ rather than $q+1$), one can derive an expression for the second derivative, and so on.

We now describe the structure that the sets $S_{q,L}$ take in the Examples (ref)-(ref) given in the previous section. One general idea, especially useful when dealing with shape constraints in Examples (ref)-(ref), for constructing $S_{q,L}$ is to enforce the shape property of interest at the chosen knots (for simplicity we can take equidistant ones). As $L\rightarrow \infty $, the knot system becomes dense in $\mathcal{X}$ and, thus, the shape property of interest becomes satisfied on an increasingly dense system of points in $\mathcal{X}$.\footnote{On a related note, Dierckx1980 shows that for cubic splines, if the sign of the second derivative is enforced at the knot points, then the respective convexity/concavity property of the approximation will hold on the whole domain. A similar property can be established for the sign of the first derivative and monotonicity.}

\vskip 0.1in

Example (ref) (Cont.). For our general null hypothesis in ((ref)), the set $S_{q,L}$ is given by

multline*[multline* omitted — 195 chars of source]

When considering special cases of ((ref)), the set(s) $S_{q,L}$ takes much simpler and/or familiar structures. Indeed, consider $R=\{1\}$ and $c_1=0$, then the constraints become on the regression function to be (weakly) monotone and from the property $\left( b\right) $ of B-splines, it is easy to see that one can take

equation*[equation* omitted — 145 chars of source]

When $R=\{r\}$, $r>1$, $c_r=0$, then the conditions imposes on the set $S_{q,L}$ can be more elaborate than those in the monotonicity scenario. However, if the interval $\left[\underline{x},\overline{x}\right]$ is split into equidistance subintervals, the structure of $S_{q,L}$ can be described quite easily as

equation[equation omitted — 189 chars of source]

A more refined form of $S_{q,L}$, which may be beneficial for small $L$, would also involve constraints that capture the behaviour of $ m_{\mathcal{B}}\left( x;L\right) $ around the boundary in a more precise way. These constraints are linear inequalities and involve relations among coefficients corresponding to the B-splines around the boundaries. The form of these inequalities would, however, be slightly different from the inequalities in ((ref)). Just to give an example, for $r=2$, $d_{r}=1$ and $c_r=0$ (thus, convexity testing), additional inequalities around the boundary would have the following form:

align*[align* omitted — 267 chars of source]

As $L\rightarrow \infty $, these additional constraints becomes less and less important as constraints ((ref)) essentially capture the whole convexity property. However, in a finite sample these constraints on the coefficients around the boundary can be important for the power of the test.

\vskip 0.1in

Example (ref) (Cont.). In ((ref)), for each interval $[\tau _{j},\tau _{j+1}]$ in the partition, we have its own system of knots, so now for this interval we denote the number of unique knots as $ L_{j}^{\prime }+1$ and the number of base B-splines for that interval as $ L_{j}=L_{j}^{\prime }+q$. The approximating B-spline on that interval is denoted as $m_{\mathcal{B};j}\left( x;L_{j}\right) =\sum_{\ell =1}^{L_{j}}\beta _{j;\ell }p_{j;\ell ,L_{j}}\left( x;q\right) $ and the set of the constraints on coefficients is denoted as $S_{j;q,L_{j}}$. It goes without saying that we assume that $\tau _{j}$ and $\tau _{j+1}$ are included into the knot system for the interval as boundary knots with the multiplicity $q+1$. Denote the complete system of knots on the interval $ (\tau _{j},\tau _{j+1}]$ as $\left\{ z_{j}^{k}\right\} _{k=1}^{L_{j}+q+1}$.

Define

multline*[multline* omitted — 232 chars of source]

and consider this subset of $\mathbb{R}^{L_j}$ to be embedded into $\mathbb{R }^{\sum_{j=0}^{J-1} L_j}$, where coefficients from other intervals are taken into account too. Then without any additional constraints on how pieces on different partition intervals are joined together, the constraint set in $ \mathbb{R}^{\sum_{j=0}^{J-1} L_j} $ can be written as $\bigcap_{j=0}^{J-1} S_{j; q,L_j}$.

If one wants to impose restrictions that the curves on the adjacent subintervals are joined continuously, then the constraint set is $$S_{cont} = \left(\bigcap_{j=0}^{J-1} S_{j; q,L_j}\right) \bigcap \left(\bigcap_{j=0}^{J-2} \left\{ \beta_{j; L_j}= \beta_{j+1; 1}\right\} \right),$$ whereas under the restriction that the curves are smoothly joined, the constraint set becomes

equation*[equation* omitted — 160 chars of source]

To give a specific example, for the null hypothesis of the smooth U-shape function with switch at the known or estimated $s_{0}$, the constraint set is

multline*[multline* omitted — 371 chars of source]

For testing quasi-convexity, the set of constraints is

multline*[multline* omitted — 327 chars of source]

By formulating the constraints in this way, we enforce quasi-convexity at the system of knots. Note, however, that quasi-convexity can also be tested by using the above methodology for testing U-shapes but would require estimating the switch point $s_0$. In this case we can make use of the relatively large literature on estimation of the mode or the maximum of a regression function by nonparametric methods (see Eddy80, Eddy82 or Muller for the case when the function is continuous at $s_0$, and DelgadoHidalgo or Hidalgo for the case when the function is discontinuous at $s_0$).

Finally, for testing symmetry around the interval center $s_0$, given symmetric system of knots on $[\underline{x},s_0]$ and $[s_0, \overline{x}$ and, hence, the same number $L$ of B-splines of degree $q$ on these intervals,

multline*[multline* omitted — 177 chars of source]

\vskip 0.1in

Example (ref) (Cont.). For the sake of brevity, we write the constraints for $r$-convexity only (it would be straightforward for a reader to write the set of constraints for $\rho$-convexity as these two notions are related):

multline*[multline* omitted — 285 chars of source]

\vskip 0.1in

Example (ref) (Cont.). Convexity in mean for any two given mean functions $M$ and $N$ can be tested by using the set of constraints

multline*[multline* omitted — 363 chars of source]

This is a generic form of constraints for ${MN}$-convexity. Once we choose a specific form of mean functions $M$ and $N$, it is possible to write these constraints in a modified way. For illustration purposes, below we present alternative ways to write the constraints for GA-convex and HG-convex functions.

For GA-convexity, we can consider

multline*[multline* omitted — 315 chars of source]

For HG-convexity, we can consider

multline*[multline* omitted — 457 chars of source]

\vskip 0.1in

Note that in Examples (ref)-(ref) the constraints on the coefficients of B-splines happen to be linear\footnote{Quasi-convexity may be an exception depending on the exact approach.} and, therefore, especially easy to implement. We want to emphasize that it is a consequence of the type of the approximation basis that we employ. In other words, it is precisely the properties of B-splines that make the testing of these hypotheses particularly easy. In Examples (ref)-(ref) the constraints will in general be non-linear in such coefficients, even though in some special cases -- such as GA-convexity, HA-convexity, or AA-convexity (which of course results in the usual convexity) -- the constraints will still happen to be linear.

\vskip 0.1in

Bernstein polynomials and other sieve bases. We now discuss our motivation to not use other sieves bases in our testing approach. As already mentioned above, a strong candidate for approximating the regression function subject to qualitative properties is a Bernstein polynomial given by

equation*[equation* omitted — 209 chars of source]

where $\mathcal{B}_{\ell ,L}\left( \tilde{x}\right) =:\left(

array[array omitted — 25 chars of source]

\right) \left( 1-\tilde{x}\right) ^{L-\ell }\tilde{x}^{\ell }$, $\tilde{x} \in \lbrack 0,1]$, denotes the $\ell th$ base Bernstein polynomial. What makes it a strong candidate is the fact that we can take $\beta _{\ell }=m\left( x+\ell /L(\overline{x}-x)\right) $, and hence, translate constrains on $m\left( x\right) $ into constraints on coefficients $\{\beta _{\ell }\}_{\ell =1}^{L}$. However, our main motivation not to use \emph{Bernstein} polynomials comes from the observation that, contrary to \emph{B-splines }or \emph{P-splines}, \emph{ Bernstein} polynomials are highly correlated. Indeed, using results in \cite{LeeParkYoo}, the eigenvalues of the matrix $\left\{ E\left( \mathcal{B}_{\ell _{1},L}\left( x\right) \mathcal{B}_{\ell _{2},L}\left( x\right) \right) =a_{\ell _{1},\ell _{2}}\right\} _{\ell _{1},\ell _{2}=0}^{L}$ are

equation*[equation* omitted — 161 chars of source]

which implies that $\lambda _{L}\leq \left( 2L+1\right) ^{-1}L^{1/2}2^{1-2L}$ since $\left(

array[array omitted — 23 chars of source]

\right) \geq L^{-1/2}2^{2L-1}$. In addition, it is not difficult to see that $a_{\ell _{1},\ell _{2}}^{2}/a_{\ell _{1},\ell _{1}}a_{\ell _{2},\ell _{2}}\rightarrow e^{-1}$ if $\left\vert \ell _{1}-\ell _{2}\right\vert =o\left( L^{1/2}\right) $, which yields some adverse and important technical consequences for the test proposed in Section (ref).

Finally, other sieve bases could potentially be used but, the implementation of the test, even for case in Examples 1 and 2 would be laborious. One popular sieve basis is the power series $1,x,\ldots ,x^{L}$ for any $L$. However, the formulation of constraints becomes increasingly complicated when $L$ increases even if one restricts their attention to such cases as in Examples 1 and 2. For instance, for testing an increasing $m_{\mathcal{B} }\left( x,L\right) =:\sum_{\ell =0}^{L}\beta _{\ell }x^{\ell }$, one may require all coefficients in the first derivative polynomial $\sum_{\ell =1}^{L} l \beta _{\ell }x^{\ell -1 } $ to be non-negative. These conditions are clearly sufficient but do not become necessary even as $L \rightarrow \infty$. This could lead to the test rejecting with probability one the null hypothesis even if it is true, with the main reason for this being the imposition of constraints that are not true or suitable. A better approach would be to use the one-to-one relationship between the Bernstein polynomials basis and the power basis $1,x,\ldots ,x^{L}$ for any $L$. A similar comment applies if we want to employ the Legendre polynomial base -- as it can be seen from the relationship between the Legendre and Bernstein polynomials given e.g. in Propositions 1 and 2 in Farouki, and this, again, seems an unnecessary step.\footnote{ The relations between the Bernstein basis and some other polynomial bases have been addressed in the approximation literature. E.g., LiZhang discusses not only the relation between the Bernstein basis and the Legendre basis but also the relation between the Bernstein basis and Chebyshev orthogonal bases.}

We shall mention nonetheless that there is no reason to believe that the results or the methodology introduced in this paper cannot be implemented using a power series approximation or potentially some other approximation basis. However, this route seems more arduous than using B-splines. Moreover, in the context of testing shape properties we believe it is more natural to employ local objects like B-splines than the objects more a global type, like most of the polynomial bases.

First step in the testing methodology and regularity conditions

$\left. {}\right. $

In the remainder of the paper, we shall assume, without loss of generality, that $\mathcal{X}=\left[ 0,1\right]$.

We now describe the first step in our methodology of testing $H_{0}$ in ((ref)) by giving estimators of $m(\cdot )$ under the alternative and the null hypothesis. To that end, write the B-splines as a vector of functions

equation[equation omitted — 233 chars of source]

Then, the standard series estimator of $m\left( x\right) $ is defined as the projection of $y$ onto the space spanned by $\boldsymbol{P}_{L}\left( x\right) $, that is

align[align omitted — 377 chars of source]

where $B^{+}$ denotes the Moore-Penrose inverse of the matrix $B$.

To obtain an estimator under the null hypothesis, we conduct a linear projection subject to suitable constraints. For general null hypothesis ((ref)), we consider estimation under the constraints in ((ref)), where $S_{q,L}$ satisfies condition $C0$:

equation[equation omitted — 297 chars of source]

This gives us the estimator of $m\left( \cdot \right) $ under $\left( \ref {main_null}\right) $:

equation[equation omitted — 132 chars of source]

When the estimation in ((ref)) involves only constraints linear in the parameters, such as in the case of the null hypothesis of an increasing function:

equation[equation omitted — 303 chars of source]

then the constrained optimization problem becomes a quadratic programing problem with linear constraints. When the constraints are nonlinear, the constrained estimation may be slightly more challenging to implement but various global optimization techniques can be utilized or even some local optimization techniques as the range of plausible initial values of the coefficients of the B-spline approximation can be assessed from the data/model.\footnote{If the unconstrained least squares estimator is in the interior of $S_{q,L}$, then, of course, none of the constraints are binding and the constrained estimation is standard. The computational complications may only happen when some of the constraints are binding.}

To get a better idea of what nonlinear constraints may look like, take e.g. Example (ref). GA-convexity and HA-convexity there result in sets $S_{q,L}$ that have linear constraints only, whereas AG-, AH-, GG-, GH-, HG-, HH-convexities result in the quadratic constraints of the form

equation[equation omitted — 97 chars of source]

Indeed, take for example AH-convexity. Under enough smoothness, this property can be written $\frac{d\,}{d\,x}\left( \frac{m^{\prime }\left( x\right) }{m^{2}\left( x\right) }\right) \geq 0$, which is equivalent to $m^{\prime \prime }\left( x\right) \geq \frac{2\left( m^{\prime }\left( x\right) \right) ^{2}}{m\left( x\right) }$. Taking into account that $m\left( x\right)$ is known to be positive, as in many economics examples, rewrite the last displayed inequality as

equation*[equation* omitted — 128 chars of source]

Since all $m_{\mathcal{B}}(x;L)$, $m_{\mathcal{B}}^{\prime }(x;L)$, $m_{\mathcal{B}}^{\prime \prime }(x;L)$ are linear at in $\beta $, then for any $x$, this constraint is quadratic in $\beta $. When this inequality is enforced at $ L+1-q=L^{\prime }+1$ unique knots,\footnote{Recall an earlier discussion in Section (ref) of a general approach to constructing $S_{q,L}$ by enforcing the shape property of interest at the knots, which guarantees that the shape property becomes satisfied on an increasingly dense system of points in $\mathcal{X}$ as $L \rightarrow \infty$.} it gives us a system of quadratic inequalities ((ref)). It is worth noting that due to the nature of the B-splines, the matrices $Q_{j}$ are quite sparse. In fact, each $Q_{j}$ contains a non-zero $q\times q$ diagonal block while all the other elements in $Q_{j}$ are zero. This means that each inequality in ((ref)) will effectively contain only $q$ adjacent coefficients in $\beta $. The analogous conclusion will imply to other properties in Example (ref) which result in non-linear constraints.\footnote{ There is a rather rich literature on the quadratic optimization subject to quadratic constraints. A review can be found, e.g., in Park and Boyd (2017).}

\vskip 0.1in

We now introduce our regularity conditions.

description$\left\{ (x_{i},u_{i})^{\prime }\right\} _{i\in \mathbb{Z}}$ is a sequence of independent and identically distributed random vectors, where $x_{i}$ has support on $\mathcal{X}=:\left[ 0,1\right] $ and its probability density function, $f_{X}\left( x\right) $, is bounded away from zero. In addition, $E[u_{i}|x_{i}]=0$, $E[u_{i}^{2}|x_{i}]=\sigma _{u}^{2}$, and $u_{i}$ has finite $4th$ moments. \vskip 0.1in • $m\left( x\right) $ is three times continuously differentiable on $\left[ 0,1\right] $. \vskip0.1in • As $n\rightarrow \infty $, $L$ satisfies $ L^{2}/n+n/L^{4}=o\left( 1\right) $.

\vskip 0.1in

Condition $C1$ can be weakened to allow for heteroscedasticity, e.g. $E\left[ u_{i}^{2}\mid x\right] =\sigma _{u}^{2}\left( x\right) $ as it was done in Stute97. However, the latter condition complicates the technical arguments and for expositional simplicity we omit a detailed analysis of this case. However, in our empirical applications we present examples with heteroscedastic errors and illustrate how to deal with it in practice. Condition $C3$ bounds the rate at which $L$ increases to infinity with $n$.

Condition $C2$ is a smoothness condition on the regression function $m\left( x\right) $. It guarantees that the approximation error or bias

equation[equation omitted — 107 chars of source]

is $O(L^{-3})$, see see Agarwal and Studden's AgarwalStudden Theorems 3.1 and 4.1 or Zhou et al. Zhou. It can be weakened to say that the second derivatives are H\"{o}lder continuous of degree $\eta >0$. In that case, $C3$ had to be modified to $L^{2}/n+n/L^{2+2\eta }=o\left( 1\right) $. In case of using P-splines we refer to Claeskens Claeskens_etal Theorem 2.

THE TESTING METHODOLOGY

$\left. {}\right. $

Having presented our estimators of $m\left( x\right) $ using or not the constraints induced by the null hypothesis, we now discuss the main part of the methodology for testing shape restrictions outlined in the introduction.

We shall focus on the null hypothesis $\left( \ref{main_nullB}\right) $ in terms of the coefficients $\{ \beta _{\ell }\} _{\ell =1}^{L}$, with the alternative hypothesis being the negation of the null. Since the boundary of $S_{q,L}$ is taken to consist of a finite number of smooth surfaces, $ S_{q,L} $ can be described by a finite number of restrictions. In this way, our testing problem translates into the more familiar testing scenario when the null hypothesis is given as a set of constraints on the parameters of the model. However the main and key difference is that, in our scenario, the number of such constraints increases with the sample size.

When testing for constraints among the parameters in a regression model, one possibility is via the (Quasi) Likelihood Ratio principle which compares the fits obtained by constrained and unconstrained estimates $\widehat{m}_{ \mathcal{B}}\left( x_{i};L\right) $ and $\widetilde{m}_{\mathcal{B}}\left( x_{i};L\right) $, respectively. That is,

equation*[equation* omitted — 257 chars of source]

Although results in ChenChristensenJOE might be used to obtain its asymptotic distribution, this route appears difficult to implement, since there is a high probability that both estimators $\widehat{m }_{\mathcal{B}}\left( x_{i};L\right) $ and $\widetilde{m}_{\mathcal{B} }\left( x_{i};L\right) $ coincide numerically. The latter is the case when the constraints are not binding.

A second possibility is to employ the Wald principle, which involves checking if the constraints in $\left( \ref{main_nullB}\right) $ hold true for the unconstrained estimator $\widetilde{b}$ of the parameters $\beta =:\left\{ \beta _{\ell }\right\} _{\ell =1}^{L}$ given in $\left( \ref {uncons}\right) $, or, in other words, if the data supports the set of constraints

equation*[equation* omitted — 92 chars of source]

Generally $S_{q,L}$ will involve some inequalities, and even when $L$ is fixed, the test would be difficult to implement or even to compute the critical values based on its asymptotic distribution. In our scenario, however, we would have two potential further technical complications. First, the number of constraints increases with the sample size $n$, which makes, from both a theoretical and practical point of view, this route very arduous, if at all feasible. Second, when some constraints describing $ S_{q,L}$ are binding, we are then dealing with estimation at the boundary which implies that the asymptotic distribution cannot be Gaussian.

A third way to test for the null hypothesis $\left( \ref{main_nullB}\right) $ is to implement the Lagrange Multiplier,\ LM, test -- that is to check if the residuals

equation[equation omitted — 138 chars of source]

and $x_{i}$ satisfy the orthogonality condition imposed by Condition $C1$. That is, we might base our test on whether or not the set of moment conditions

equation[equation omitted — 171 chars of source]

are significantly different from zero, where $\mathcal{I}$ is the indicator function and we abbreviate $\mathcal{I}\left( x_{i}<x\right) $ as $\mathcal{I} _{i}\left( x\right) $. This approach was described and examined in Stute97 or Andrews97 with a more econometric emphasis. Tests based on $\mathcal{K}_{n}\left( x\right) $ are known as testing using partial sum empirical processes. Recall that in a standard regression model the LM test is based on the first order conditions

equation*[equation* omitted — 131 chars of source]

which has the interpretation of whether the residuals and regressors, $ p_{\ell ,L}\left( x_{i};q\right) $, satisfies the orthogonality moment condition induced by Condition $C1$.

However to motivate the reasons to employ a transformation of $\mathcal{K} _{n}\left( x\right) $, given in $\left( \ref{k_1}\right) $ or $\left( \ref {k_2}\right) $ below, as the basis for our test statistic, it is worth examining the structure of $\mathcal{K}_{n}\left( x\right) $ given in $ \left( \ref{T_n}\right) $. For that purpose, we observe that

equation[equation omitted — 325 chars of source]

where $m^{bias}\left( x_{i}\right) $ was given in $\left( \ref{m_bias} \right) $ and

equation[equation omitted — 168 chars of source]

Now, the third term on the right of $\left( \ref{T_n1}\right) $ is $O\left( L^{-3}\right) $ by Agarwal and Studden's AgarwalStudden Theorems 3.1 and 4.1, and then Condition $C1$. On the other hand, standard arguments and Condition $C1$ imply that

equation*[equation* omitted — 250 chars of source]

where $\mathcal{B}\left( z\right) $ denotes the standard Brownian motion and $F_{X}\left( x\right) $ the distribution function of $x_{i}$.

Next, we discuss the contribution due to $\sum_{\ell =1}^{L}\left( \widehat{b }_{\ell }-\beta _{\ell }\right) \mathcal{P}_{n,\ell }\left( x;q\right) $. In standard lack-of-fit testing problems with $L$ finite, when $\mathcal{K} _{n}\left( x\right) =O_{p}\left( n^{-1/2}\right) $, its contribution is nonnegligible as first showed by Durbin73 and later by Stute97 in a regression model context. However the proof of Theorem (ref) suggests that $\sum_{\ell =1}^{L}\left( \widehat{b} _{\ell }-\beta _{\ell }\right) \mathcal{P}_{n,\ell }\left( x;q\right) =O_{p}\left( \left( L/n\right) ^{1/2}\right) $, so that as $L$ increases with the sample size, it yields that the normalization factor for $\mathcal{K }_{n}\left( x\right) $ is of order $n^{\alpha }$ for some $\alpha <1/2$.

Thus the previous arguments suggest that under our conditions, we would have

equation*[equation* omitted — 267 chars of source]

Results by Newey Newey97, Lee and Robinson LeeRobinson or Chen and Christensen ChenChristensenJOE in a more general context, might suggest that the left side of the last displayed expression might converge to a Gaussian process when $\beta $ is in the interior of the set $ S_{q,L}$. However, when $\beta $ is at the boundary of $S_{q,L}$, then the asymptotic distribution is not Gaussian, and so to obtain the asymptotic distribution of $\left( n/L\right) ^{1/2}\sum_{\ell =1}^{L}\left( \widehat{b} _{\ell }-\beta _{\ell }\right) \mathcal{P}_{n,\ell }\left( x;q\right) $ for inference purposes appears quite difficult, if at all possible.

So, the purpose of the next section is to examine a transformation of $ \mathcal{K}_{n}\left( x\right) $ such that its statistical behaviour will be free from $\sum_{\ell =1}^{L}\left( \widehat{b}_{\ell }-\beta _{\ell }\right) \mathcal{P}_{n,\ell }\left( x;q\right) $. The consequence of the transformation would then be twofold. First, we would obtain that the transformation of $n^{1/2}\mathcal{K}_{n}\left( x\right) $ is $O_{p}\left( 1\right) $, which leads to better statistical properties of the test, and secondly and more importantly, the test will be pivotal in the sense that $ \sigma _{u}^{2}$ becomes the only unknown (although easy to estimate) of its asymptotic distribution. One consequence of our results is that the asymptotic distribution becomes independent of the null hypothesis under consideration.

KHMALADZE'S TRANSFORMATION

This section examines a transformation of $\mathcal{K}_{n}\left( x\right) $ whose asymptotic distribution is free from the statistical behaviour of $ \left\{ \widehat{b}_{\ell }\right\} _{\ell =1}^{L}$. To that end, we propose a \textquotedblleft martingale\textquotedblright\ transformation based on ideas by Khmaladze81, see also BrownDurbinEvans for earlier work. The (linear) transformation, denoted $ \mathcal{T}$, should satisfy that

eqnarray[eqnarray omitted — 526 chars of source]

where

equation*[equation* omitted — 310 chars of source]

with $\boldsymbol{P}_{L}\left( x\right) $ and $\mathcal{P}_{n,\ell }\left( x;q\right) $ given respectively in $\left( \ref{p_appro}\right) $ and $ \left( \ref{p_L}\right) $.

ALL THE CONSTRAINTS ON $\protect\beta_{\ell}$ ARE LINEAR

$\left. {}\right. $

We first present the form of Khmaladze's transformation when all the constraints characterizing $S_{q,L}$ are linear and, thus, the surfaces forming the boundary of $ S_{q,L}$ are hyperplanes.

For any $x<1$, denote

equation[equation omitted — 188 chars of source]

We then define the transformation $\mathcal{T}$ as

equation*[equation* omitted — 307 chars of source]

It is easy to see that the transformation $\mathcal{T}$ satisfies condition $ \left( \emph{ii}\right) $ in $\left( \ref{properties}\right) $, so that the main concern will be to show that $\left( \emph{i}\right) $ and $\left( \emph{iii}\right) $ hold true. It is important to note that the constraints that proved to be binding in the estimation under the null hypothesis have to be incorporated when conducting the Khmaladze's transformation and, thus, $\boldsymbol{P}_{L}$ in $\left( \ref{p_appro}\right) $ has to be redefined. To give a specific example, consider the monotonicity testing and, thus, the constrained estimation as in ((ref)). Suppose the estimation results in one binding constraint $\widehat{b}_{\ell _{0}}=\widehat{b}_{\ell _{0}+1} $. Then, we redefine $P_k$ in $\left( \ref{p_appro}\right)$ as

eqnarray*[eqnarray* omitted — 340 chars of source]

so that we effectively incorporate the binding constraint. The analogous methodology applies with several binding constraints.

In what follows then, we shall make no distinction among various cases of binding constraints to simplify the exposition of the arguments. However, we shall emphasize that the use of the \textquotedblleft correct\textquotedblright $P_{L}\left( x\right) $ is crucial for the power of the test. Using $P_{L}\left( x\right) $ as given in $\left( \ref{p_appro}\right)$ without taking into account the binding the constraints will make the test to have only trivial power (see the discussion in Section (ref)).

The transformation $\mathcal{T}$ has only a theoretical value and as such, from an inferential point of view, we need to replace it by its sample analogue, which we shall denote by $\mathcal{T}_{n}$. To that end, for any $x\in \mathcal{X}$, define

eqnarray[eqnarray omitted — 371 chars of source]

where $\mathcal{I}\left( x\leq x_{k}\right) =:\mathcal{J}_{k}\left( x\right) =1-\mathcal{I}_{k}\left( x\right) $.

In what follows, we shall abbreviate

equation[equation omitted — 143 chars of source]

where $\widetilde{x}_{i}=x_{i}$ if $x_{i}+n^{-\varsigma }<z^{k\left( x_{i}\right) }$ and $=z^{k\left( x_{i}\right) }$ otherwise, with $z^{k\left( x\right) }$ denoting the closest knot $z^{k}$, $k=1,\ldots ,L$, bigger than $ x$ and $1/2<\varsigma <1$. The motivation to make this \textquotedblleft trimming\textquotedblright\ is because when $x_{i}$ is too close to $ z^{k\left( x_{i}\right) }$, the B-spline is close but not equal to zero, which induces some technical complications in the proof of our main results. However, in small samples this\ \textquotedblleft trimming\textquotedblright\ does not appear to be needed, becoming a purely technical argument.

We define the sample analogue of $\left( \mathcal{TW}\right) \left( x\right) $ as

equation[equation omitted — 310 chars of source]

The transformation in $\left( \ref{Khma_n}\right) $ has a rather simple motivation. Suppose that we have ordered the observations according to $ x_{i} $, that is $x_{i-1}\leq x_{i}$, $i=2,...,n$, which would not affect the statistical behaviour of $\mathcal{K}_{n}\left( x\right) $. The latter follows by the well known argument that

equation[equation omitted — 123 chars of source]

where $x_{\left( i\right) }$ is the $i-th$ order statistic of $\left\{ x_{i}\right\} _{i=1}^{n}$. So, we have that

equation[equation omitted — 194 chars of source]

Now, $\widehat{u}_{i}=u_{i}-\boldsymbol{P}_{i}^{\prime }A_{n}^{+}\left( 0\right) C_{n}\left( 0\right) $, where

eqnarray*[eqnarray* omitted — 324 chars of source]

So, if instead of $C_{n}\left( 0\right) $ and $A_{n}\left( 0\right) $, we employed $C_{n,i}=n^{-1}\sum_{k=i}^{n}\boldsymbol{P}_{k}u_{k}$ and $ A_{n,i}=n^{-1}\sum_{k=i}^{n}\boldsymbol{P}_{k}\boldsymbol{P}_{k}^{\prime }$, we would replaced $\widehat{u}_{i}$ by

equation[equation omitted — 89 chars of source]

in $\left( \ref{cusum}\right) $, so that it has a martingale difference structure as $E\left[ v_{i}\mid past\right] =0$, in comparison with $ \widehat{u}_{i}$ where $E\left[ \widehat{u}_{i}\mid past\right] \neq 0$. This is the idea behind the so-called (recursive) Cusum statistic first examined in BrownDurbinEvans and developed and examined in length in Khmaladze81. Observe that $\left( \ref{u_rec}\right) $ becomes the \textquotedblleft prediction error\textquotedblright of $u_{i}$ when we use the \textquotedblleft last\textquotedblright\ $j=i,\ldots,n$ observations.

Thus, the preceding argument yields the Khmaladze's transformation

equation[equation omitted — 201 chars of source]

Observe that, using $\left( \ref{g_i}\right) $, we could write $\mathcal{M} _{n}\left( x\right) $ as

equation*[equation* omitted — 155 chars of source]

Now, Lemma (ref) in Appendix B implies that the condition $\left( \emph{ii}\right) $ in $ \left( \ref{properties}\right) $ holds true when

equation*[equation* omitted — 132 chars of source]

so the technical problem is to show that (asymptotically) conditions $\left( \emph{i}\right) $ and $\left( \emph{iii}\right) $ in $\left( \ref{properties} \right) $ also hold true. That is, to show that

eqnarray*[eqnarray* omitted — 271 chars of source]

Finally, it is worth mentioning that in $\left( \ref{Khma_n}\right) $ we might have employed $\mathcal{J}_{k}\left( x\right) =\mathcal{I}\left( x<x_{k}\right) $ instead of our definition $\mathcal{J}_{k}\left( x\right) = \mathcal{I}\left( x\leq x_{k}\right) $. However because by definition of B-splines, the matrix $A_{n,i}$, and hence $A_{L}\left( x_{i}\right) $ , might be singular, if we employed $\mathcal{J}_{k}\left( x\right) = \mathcal{I}\left( x<x_{k}\right) $, then it would not be guaranteed that

equation*[equation* omitted — 103 chars of source]

On the other hand, Theorem 12.3.4 in Harville yields that the last displayed equation holds true when $\mathcal{J}_{k}\left( x\right) =\mathcal{I}\left( x\leq x_{k}\right) $. Now, $\left( 1\right) $ will be shown in the next theorem, whereas $\left( 2\right) $ is shown in Theorem (ref).

theoremUnder Conditions $C1-C3$, we have that \begin{equation*} \mathcal{M}_{n}\left( x\right) \overset{weakly}{\Rightarrow }\mathcal{U} \left( x\right) ; \ \ x\in \left[ 0,1\right] . \end{equation*}

Unfortunately, we do not observe $u_{i}$, so that to implement the transformation we replace $v_{i}$ by $\widehat{v}_{i}$, where $\widehat{v} _{i}$ is defined as $v_{i}$ in $\left( \ref{u_rec}\right) $ but where we replace $u_{i}$ by $\widehat{u}_{i}$ as defined in $\left( \ref{res_con} \right) $, yielding the statistic

equation[equation omitted — 160 chars of source]
theoremAssuming that $H_{0}$ holds true, under Conditions $C1-C3$, we have that \begin{equation*} \widetilde{\mathcal{M}}_{n}\left( x\right) \overset{weakly}{\Rightarrow } \mathcal{U}\left( x\right) . \end{equation*}

Denote the estimator of the variance of $u_{i}$, $\sigma _{u}^{2}$, by

equation*[equation* omitted — 96 chars of source]
propositionUnder Conditions $C1-C3$, we have that $\widehat{\sigma } _{u}^{2}\overset{P}{\rightarrow }\sigma _{u}^{2}$.

We then have the following corollary.

corollaryUnder $H_{0}$ and assuming Conditions $C1-C3$, for any continuous functional $g:\mathbb{R\rightarrow R}^{+}$, \begin{equation*} g\left( \widetilde{\mathcal{M}}_{n}\left( x\right) /\widehat{\sigma } _{u}\right) \overset{d}{\rightarrow }g\left( \mathcal{U}\left( x\right) /\sigma _{u}\right) . \end{equation*}
proofThe proof is standard using Theorem $\ref{M_nest}$, Proposition (ref) and the continuous mapping theorem, so it is omitted.

Denoting $\tilde{n}=:n-L-2$ and $\widetilde{\mathcal{M}}_{n}\left( x^{q}\right) =\widetilde{\mathcal{M}}_{n,q}$, where $x^{q}=q/n$, standard functionals are the Kolmogorov-Smirnov, Cram\'{e}r-von-Mises and Anderson-Darling tests defined respectively as{\

eqnarray[eqnarray omitted — 553 chars of source]

}

equation*[equation* omitted — 347 chars of source]

\vskip 0.1in

NONLINEAR CONSTRAINTS ON $\protect\beta_{\ell}$

$\left. {}\right. $

We turn now our attention to describing the Khmaladze's transformation in situations when some constraints describing $S_{q,L}$ may be non-linear, as it happens to be in our Examples 3 and 4. If the constrained estimate $\widehat{b}$ lies in the interior of $S_{q,L}$ (and, thus, coincides with the unconstrained estimate), then the transformation is conducted in the same way as in Section (ref). However, the transformation will have a modified form if $ \widehat{b}$ lies on the boundary of $S_{q,L}$. We give its form and, in particular, discuss what objects plays the role of $\boldsymbol{P}_{L}\left( x\right)$ described in the previous section.

For expositional simplicity suppose that $\widehat{b}$ belongs only to one of the smooth surfaces describing the boundary of $S_{q,L}$. Usually these surfaces will be defined by implicit functions but, applying the implicit function theorem, the explicit representation of this surface can be obtained either analytically or numerically (even if local, which would suffice since $\widehat{b}$ is consistent) with respect to one parameter expressed as a function of other parameters: suppose that for some $\ell _{0}$ we can express it as $ \beta_{\ell _{0}}=h(\beta _{1},\ldots ,\beta _{\ell _{0}-1},\beta _{\ell _{0}+1},\ldots ,\beta _{L})$. In cases such as AG-, AH-convexity and other similar properties in Example (ref), for the reasons discussed in Section (ref) this would be a restriction $\beta_{\ell _{0}}=h(\beta _{\ell _{0}-2},\beta _{\ell _{0}-1})$ obtained from an implicit function which is a polynomial of degree 2. Then for the purpose of conducting the Khmaladze's transformation, instead of approximating $m(\cdot )$ by the linear function $\sum_{k=0}^{\ell -1}\beta _{k}p_{k}\left( x_{i}\right) $, we consider the approximation given by

equation[equation omitted — 243 chars of source]

where $\beta_{-\ell _{0}}=(\beta _{1},\ldots ,\beta _{\ell _{0}-1},\beta _{\ell _{0}+1},\ldots ,\beta _{L})$. Although this approximation function is nonlinear in parameters, it has a simple and useful structure, as it will become evident from our analysis below.

To define the Khmaladze's transformation, denote the vector of first derivatives of $g\left( x;\beta_{-\ell _{0}}\right) $ with respect to the parameters as

eqnarray*[eqnarray* omitted — 253 chars of source]

where

equation*[equation* omitted — 232 chars of source]

Then, the Khmaladze's transformation of the test statistic $\mathcal{K} _{n}\left( x\right) $ in ((ref)) has the following form:

equation*[equation* omitted — 119 chars of source]

where

equation*[equation* omitted — 279 chars of source]

$\widetilde{P}_{i}\left( \beta_{-\ell _{0}}\right) =:\widetilde{P} \left( x_{i};\beta_{-\ell _{0}}\right) $, and $\mathcal{D}_{n}\left(x; \beta _{-\ell _{0}}\right) =\sum_{k=1}^{n}\widetilde{P}_{k}\left( \beta _{-\ell _{0}}\right) \widetilde{P}_{k}^{\prime }\left( \beta_{-\ell _{0}}\right) \mathcal{J}_{k}\left(x\right) $, and $\mathcal{D}^{+}_{n}\left(x; \beta _{-\ell _{0}}\right)$ is the Moore-Penrose inverse of $\mathcal{D}_{n}\left(x; \beta _{-\ell _{0}}\right)$, and $\mathcal{D}_{n}\left( \beta _{-\ell _{0}};i\right) =\mathcal{D}_{n}\left(\widetilde{x}_{i}; \beta _{-\ell _{0}}\right) $ with $\widetilde{x}_{i}$ defined in the same way as in Section (ref). Note that by employing $\widetilde{p}_{\ell }\left( x_i;\beta_{-\ell _{0}}\right)$ instead of $p_{\ell }\left( x_i\right)$, we have automatically incorporated in the transformation our binding restriction.

In practice, instead of $v_{i}$ we use

equation*[equation* omitted — 309 chars of source]

and consider a feasible version of the Khmaladze transformation:

equation*[equation* omitted — 141 chars of source]

Just like in Section (ref), we have the following result.

theoremAssuming that $H_{0}$ holds true, under Conditions $ C1-C3$, we have that \begin{equation*} \widetilde{\mathcal{M}}_{n}\left( x\right) \overset{weakly}{ \Rightarrow }\mathcal{U}\left( x\right) . \end{equation*}

This methodology can be generalized, of course, to the situation when the constrained estimate belongs to the intersection of several boundary surfaces. In this case we would express several parameter values as functions of other parameters and plug these functions into the linear approximation obtaining a nonlinear expression in the remaining parameters. We would define the derivative of the new approximation and adjust the Khmaladze transformation to account now for several nonlinear equality constraints. The rest of the methodology and the asymptotic result would remain the same as above.

COMPUTATIONAL ISSUES

$\left. {}\right. $

This section is devoted at how we can compute our statistic. In view of the CUSUM interpretation, we shall rely on the standard recursive residuals. We will illustrate the computational issues using notations from the case of all linear constraints but, of course, the methodology in the case of non-linear constraints will be the same.

Note that since $f\left( x\right) $ is continuous the probability of a tie is zero, so that we can always consider the case $x_{i}<x_{i+1}$. Now with this view we have that

equation*[equation* omitted — 119 chars of source]

can be written with $v_{i}$ replaced by $v_{i}=u_{i}-\boldsymbol{P} _{i}^{\prime }A_{n}^{+}\left( x\right) C_{n}\left( x\right) $ and now

equation*[equation* omitted — 254 chars of source]

Then from a computational point of view is worth observing that

equation*[equation* omitted — 187 chars of source]

and

equation*[equation* omitted — 159 chars of source]

see BrownDurbinEvans for similar arguments. Alternatively, we could have considered the Cusum of backward recursive residuals, in which case we would have use the computational formulae,

equation*[equation* omitted — 219 chars of source]

and

equation*[equation* omitted — 199 chars of source]

Of course in the previous formulas one would replace $u_{i}$ by $\widehat{u} _{i}$ or $y_{i}$.

POWER AND LOCAL ALTERNATIVES

$\left. {}\right. $

Here we describe the local alternatives for which the tests based on $ \widetilde{\mathcal{M}}_{n}\left( x\right) $ have no trivial power. For that purpose, assume that the true model is such that

equation*[equation* omitted — 78 chars of source]

where $m\left( x\right) \in \mathcal{M}_0$ and $m_{1}\left( x\right) $ is incompatible with $\mathcal{M}_0$ at least on some subset $\mathcal{X} _{1}\subset \mathcal{X=}\left( 0,1\right) $ with the Lebesgue measure strictly greater than zero. Then, we have the result of Proposition (ref).

propositionAssuming Conditions $C1-C3$, under $H_{a}$ we have that \begin{equation*} \widetilde{\mathcal{M}}_{n}\left( x\right) +\mathcal{L}\left( x\right) \overset{weakly}{\Rightarrow }\mathcal{U}\left( x\right) , \ \ x\in \left[ 0,1\right] \end{equation*} where \begin{equation*} \mathcal{L}\left( x\right) =\int_{0}^{x}\left\{ m_{1}\left( v\right) - \mathbf{P}_{L}^{\prime }\left( v\right) A_{L}^{+}\left( v\right) \int_{v}^{1} \mathbf{P}_{L}\left( w\right) m_{1}\left( w\right) f_{X}\left( w\right) dw\right\} f_{X}\left( v\right) dv. \end{equation*}

One consequence of Proposition (ref) is that not only tests based on $\widetilde{\mathcal{M}}_{n}\left( x\right) $ are consistent since $\mathcal{ L}\left( x\right) $ is a nonzero function, but that they also have a nontrivial power against local alternatives converging to the null at the \textquotedblleft parametric\textquotedblright\ rate $n^{-1/2}$.

\paragraph*{Role of constraints for the power of the test} In our discussion of the Khmaladze's transformation we stated that in that transformation one has to enforce the binding constraints obtained in the constrained estimation under the null hypothesis. Here we present a simple example to illustrate what happens to the power of the test if such constraints are not enforced.

To keep arguments simple, consider the case where $L=2$ when we use the B-splines to approximate the function and the null hypothesis, and take the null hypothesis to be that of an increasing regression function. That is, we employ the approximation $m_{\mathcal{B}}\left( x;2\right) = \beta _{1}p_{1,2}\left( x;q\right) + \beta _{2}p_{2 ,2}\left( x;q\right)$ and rewrite it in the following way:

eqnarray*[eqnarray* omitted — 243 chars of source]

where $\alpha _{1}=\beta _{1}$ $\alpha _{2}=\beta _{2}-\beta _{1}$ and $ p\left( x;q\right) =:p_{1,2}\left( x;q\right) +p_{2,2}\left( x;q\right) $. Of course, because of the B-splines properties, we have $p\left( x;q\right)=1$ but for the sake of presenting a more general argument, we will keep the more general notation $p\left( x;q\right)$. The null hypothesis is then written as $\alpha _{2}\geq 0$. Suppose the null is not true and the true value is $\alpha _{2}<0$. When we estimate the model, we should expect that our estimator is of the form $\left( \widehat{ \alpha }_{1},0\right) $. We then define constrained residuals as

equation*[equation* omitted — 80 chars of source]

According to our methodology, at each step of the transformation we should only be projecting the residuals on $p\left( x;q\right)$.

Let us analyze what happens if in the transformation we use as our $\mathbf{P }_{k}$ that comes from the vector $\left( p\left( x_{k};q\right) ,{p} _{2,2}\left( x_{k};q\right) \right) $ instead of using only $p\left( x_{k};q\right)$. In other words, we use the test statistic

equation*[equation* omitted — 148 chars of source]

where $\widehat{u}_{i}=y_{i}-\widehat{\alpha }_{1}p\left( x;q\right) $ and the transformation is defined as follows:

eqnarray*[eqnarray* omitted — 257 chars of source]

Rewrite the test statistic as

equation*[equation* omitted — 228 chars of source]

and notice that the first term on the right-hand side converges to the Brownian motion. The second term is negligible under the null, as established earlier, and partly this is due to the result of Lemma (ref) in Appendix B. It is rather obvious that the power of the test, as usual, comes from that term. For the test to have power we need that under the alternative the term $\frac{1}{n}\sum_{i=1}^{n}\left( \widehat{v} _{i}-v_{i}\right) \mathcal{I}_{i}\left( x\right) $ converges somewhere different than zero on a set of a positive measure, so that once we multiply it by $n^{1/2}$, it diverges to $\pm \infty $. We have that

eqnarray[eqnarray omitted — 393 chars of source]

Lemma (ref) implies that when the transformation uses $\mathbf{P}_{k}$ that comes from the vector $\left( p\left( x_{k};q\right) ,{p}_{2,2}\left( x_{k};q\right) \right) $, any linear combination $\boldsymbol{ \mathring{p}}\left( x\right) =:a_{1}p\left( x;q\right) +a_{2}{p}_{2,2}\left( x;q\right) $ satisfies $\left( \mathcal{T}_{n}\mathcal{B}_{n,2}\right) \left( x\right) =0$, where $\mathcal{B}_{n,2}\left( x\right) =n^{-1}\sum_{k=1}^{n}\boldsymbol{\mathring{p}}\left( x_{k}\right) \mathcal{I} _{k}\left( x\right) $. This, taking into account representation ((ref)), implies that the part of $n^{-1}\sum_{i=1}^{n}\left( \widehat{ v}_{i}-v_{i}\right) \mathcal{I}_{i}\left( x\right) $ corresponding to $ \left( \widehat{\alpha }_{1}-\alpha _{1}\right) p\left( x_{i};q\right) -\alpha _{2}{p}_{2,2}\left( x_{i};q\right) $ in ((ref)) will be zero, and its part corresponding to the bias term $m_{\mathcal{B}}\left( x_{i};2\right) -m(x)$ in ((ref)) will be asymptotically negligible even when multiplied by $n^{1/2}$. Thus, the power of the test equals the trivial power because we failed to embed the binding constraint in the transformation.

On the other hand, if in the transformation we use only the “explanatory” polynomial $p\left( x;q\right) $, we would now have that the contribution due to $\left( \widehat{\alpha }_{1}-\alpha _{1}\right) p\left( x_{i};q\right) -\alpha _{2}{p}_{2,2}\left( x_{i};q\right) $ in $ n^{-1}\sum_{i=1}^{n}\left( \widehat{v}_{i}-v_{i}\right) \mathcal{I} _{i}\left( x\right) $ is exactly as that from $\alpha _{2}{p}_{2,2}\left( x_{i};q\right) $ and it is not equal to zero as ${p}_{2,2}\left( x_{i};q\right) $ is not in the space generated by $p\left( x_{i};q\right) $. In other words,

equation*[equation* omitted — 290 chars of source]

unless ${p}_{2,2}\left( x_{i};q\right) $ is proportional to ${p}_{2,2}\left( x_{i};q\right) $, which is not the case.

BOOTSTRAP ALGORITHM

One of our motivations to introduce a bootstrap algorithm for our test(s) is that although it is pivotal, our Monte Carlo experiment suggests that they suffer from small sample biases. When the asymptotic distribution does not provide a good approximation to the finite sample one, a standard approach to improve its performance is to employ bootstrap algorithms, as they provide small sample refinements. In fact, our Monte Carlo simulation does suggest that the bootstrap, to be described below, does indeed give a better finite sample approximation. The notation for the bootstrap is as usual and we shall implement the fast algorithm of WARP by GiacominiPW in the Monte Carlo experiment.

The bootstrap is based on the following 3 STEPS.

description• Compute the unconstrained residuals \begin{equation*} \widetilde{u}_{i}=y_{i}-\widetilde{m}_{\mathcal{B}}\left( x_{i};L\right) , \ \ \ i=1,...,n \end{equation*} with $\widetilde{m}_{\mathcal{B}}\left( x_{i};L\right) $ as defined in $ \left( \ref{uncons}\right) $. \vskip 0.1in • Obtain a random sample of size\ $n$ from the empirical distribution of $\left\{ \widetilde{u}_{i}-\frac{1}{n} \sum_{i=1}^{n}\widetilde{u}_{i}\right\} _{i=1}^{n}$. Denote such a sample as $\left\{ u_{i}^{\ast }\right\} _{i=1}^{n}$ and compute the bootstrap analogue of the regression model using $\widehat{m}_{\mathcal{B}}\left( x_{i};L\right) $, that is \begin{equation} y_{i}^{\ast }=\widehat{m}_{\mathcal{B}}\left( x_{i};L\right) +u_{i}^{\ast } , \ i=1,...,n. \end{equation} \vskip0.1in • Compute the bootstrap analogue of $\widetilde{\mathcal{M }}_{n}\left( x\right) $ as \begin{equation*} \widetilde{\mathcal{M}}_{n}^{\ast }\left( x\right) =:\frac{1}{n^{1/2}} \sum_{i=1}^{n}\widehat{v}_{i}^{\ast }\mathcal{I}_{i}\left( x\right) \end{equation*} where \begin{equation*} \widehat{v}_{i}^{\ast }=\widehat{u}_{i}^{\ast }-\boldsymbol{P}_{i}^{\prime }A_{n,i}^{+}C_{n,i}^{\ast };\text{ \ }C_{n,i}^{\ast }=:C_{n,i}^{\ast }\left( \widetilde{x}_{i}\right) =\frac{1}{n}\sum_{k=1}^{n}\boldsymbol{P}_{k} \widehat{u}_{k}^{\ast }\mathcal{J}_{k}\left( \widetilde{x}_{i}\right) \end{equation*} with $\widehat{u}_{i}^{\ast }=y_{i}^{\ast }-\boldsymbol{P}_{i}^{\prime }A_{n}^{+}\left( 0\right) C_{n}^{\ast }\left( 0\right) $, $i=1,...,n$.
theoremUnder Conditions $C1-C3$, we have that for any continuous function $g:\mathbb{R\rightarrow R}^{+}$, (in probability), \begin{equation*} g\left( \widetilde{\mathcal{M}}_{n}^{\ast }\left( x\right) \right) \overset{d }{\Rightarrow }g\left( \mathcal{U}\left( x\right) \right) . \end{equation*}

Finally, we can replace $\widehat{u}_{i}^{\ast }$ by $y_{i}^{\ast }$ in the computation of $\widetilde{\mathcal{M}}_{n}^{\ast }\left( x\right) $. That is,

corollaryUnder Conditions $C1-C3$, we have that \begin{equation*} \widetilde{\mathcal{M}}_{n}^{\ast }\left( x\right) -\widetilde{\widetilde{ \mathcal{M}}}_{n}^{\ast }\left( x\right) =0, \end{equation*} where \begin{equation*} \widetilde{\widetilde{\mathcal{M}}}_{n}^{\ast }\left( x\right) =:\frac{1}{ n^{1/2}}\sum_{i=1}^{n}\left( y_{i}^{\ast }-\boldsymbol{P}_{i}^{\prime }A_{n,i}^{+}\frac{1}{n}\sum_{k=1}^{n}\boldsymbol{P}_{k}y_{k}^{\ast }\mathcal{ J}_{k}\left( x_{i}\right) \right) \mathcal{I}_{i}\left( x\right) . \end{equation*}

The proof of Corollary (ref) is immediate by Lemma (ref) in the Appendix and therefore is omitted.

MONTE CARLO EXPERIMENTS AND EMPIRICAL EXAMPLES

MONTE CARLO EXPERIMENTS

$\left. {}\right. $

In this section we present the results of several computational experiments. All the results in this section are given for cubic splines with different number of knots. We present the results for B-splines as well as for P-splines with penalties on the second differences of coefficients. The penalty parameter is chosen by cross-validation in the unconstrained estimation. In the tables \textquotedblleft KS\textquotedblright\ refers to the Kolmogorov-Smirnov test statistic, \textquotedblleft CvM \textquotedblright\ refers to the Cram\'{e}r-von Mises test statistic and \textquotedblleft AD\textquotedblright\ to the Anderson-Darling integral test statistic. All three test statistics are based on a Brownian bridge. $L^{\prime }+1$ denotes the number of equidistant knots (including the boundary points) on the interval of interest. For example, when $ L^{\prime }=6$ and the interval is $[0,1]$, we consider knots $ 0,1/6,1/3,1/2,2/3,5/6,1$. In the implementation of P-splines in simulations, every simulation draw will give a different cross-validation parameter (we use ordinary cross validation described in Eilers and Marx, $1996$). In our simulation results for each $L^{\prime }$ we use a modal value of these cross-validation parameters.

In all the scenarios below

equation*[equation* omitted — 94 chars of source]

In Scenarios 1, 3-5 the interval of interest if $[0,1]$ whereas in Scenario 2 of U-shape we consider individually intervals $[0,s_0]$ and $[s_0,1]$ with $s_0$ being the switch point.

In the WARP bootstrap implementation, the demeaned residuals and $x$ are drawn independently.

\vskip 0.1in

Scenario 1 (test for monotonicity). We take the following regression function:

equation*[equation* omitted — 97 chars of source]

The results are summarized in Table (ref).

\vskip 0.05in

table[table omitted — 1,421 chars of source]

\vskip 0.1in

Scenario 2 (test for U-shape). The regression function is defined as

align*[align* omitted — 55 chars of source]

The graph of this function is U-shaped with the switch point at $ s_0=e^{0.33}-1$. In simulations $s_0$ is taken to be known.

The results are summarized in Table (ref). We use two different B-splines -- one on $[0,s_{0}]$ and the other on $[s_{0},1]$. We analyze the properties of the testing procedure in two approaches. In the first approach additional restrictions are imposed for the two B-splines to be joined continuously at $s_{0}$, and in the second approach these two B-splines are joined smoothly at $s_{0}$ (see details in Example 2).

table[table omitted — 1,404 chars of source]

\vskip 0.1in

Scenario 3 (analysis of power of the test). Take the regression function

align*[align* omitted — 147 chars of source]

and depicted in Figure (ref). As expected, the power of the test depends on the variance of the error. The results are summarized in Table (ref).

figure[figure omitted — 155 chars of source]
table[table omitted — 2,199 chars of source]

The power of monotonicity tests based on this regression function is considered in GhosalSenVV and a similar regression function is considered in HallHeckman. Note that GhosalSenVV considers smaller sample sizes and also smaller standard deviation of noise with $\sigma=0.1$.

\vskip 0.1in

Scenario 4 (analysis of power of the test). The regression function

equation*[equation* omitted — 56 chars of source]

and depicted in Figure (ref). The left-hand side graph in Figure (ref) is for the case $a=50$ and the right-hand side graph in Figure (ref). In the latter case the non-monotonicity dip is smaller. These situations are considered to be challenging for monotonicity tests as these functions are somewhat close to the set of monotone functions (in any conventional metric). As expected, the power of the test depends on the value of parameter $a$ and also depends on the variance of the error. The results are summarized in Table (ref).

figure[figure omitted — 464 chars of source]
table[table omitted — 4,229 chars of source]

The power of monotonicity tests based on this regression function is examined in GhosalSenVV and a similar regression function was considered in Bowman_etal. Note that GhosalSenVV uses smaller sample sizes and also only $a=50$ and $\sigma=0.1$ to analyze power implications.

\vskip 0.1in

Scenario 5 (test for log-convexity). We take the following regression function:

equation*[equation* omitted — 84 chars of source]

The results are summarized in Table (ref). In this case, the results for P-splines are the same as for B-splines as the cross-validation criterion indicated 0 as the optimal penalty parameter in the overwhelming majority of simulations.

\vskip 0.05in

table[table omitted — 1,006 chars of source]

APPLICATIONS

$\left. {}\right.$

1. Hospital data Here we use data on hospital finance for 332 hospitals in California in 2003.\footnote{ The dataset is from https://www.kellogg.northwestern.edu/faculty/dranove/\newline htm/dranove/coursepages/mgmt469.htm. The original dataset is for 333 hospitals but we had to remove the observation that had missing information about administrative expenses.} The data include many variables related to hospital finance and hospital utilization. We are interested in analyzing the effect of revenue derived from patients on administrative expenses.

Figure (ref) is a scatter plot of the logarithm of patient revenue and the logarithm of administrative expenses with the fitted curve obtained using cubic B-splines with $L^{\prime }+1=5$ uniform knots in the range of values of the log of patient revenue. The fitted cure is obtained under the monotonicity restriction.

figure[figure omitted — 356 chars of source]

We conduct tests for the following hypotheses: (a) monotonicity; ( b) convexity; (c) monotonicity and convexity.

In order to correct for heteroscedasticity of the errors, we estimate the scedastic function $\widehat{\sigma }^{2}(x)$ using residuals obtained in the unconstrained estimation using cubic B-splines with the same set of knots. The scedastic function $\widehat{\sigma }^{2}(x)$ is estimated by regressing the logarithm of the squared unconstrained residuals on a linear combination of first-order B-splines with $6$ knots in the domain of the log of patient revenue.

We then consider the constrained residuals divided by $\widehat{\sigma }(x)$ when calculating KS, CvM and AD test statistics and unconstrained residuals divided by $\widehat{\sigma }(x)$ when drawing bootstrap samples. After a bootstrap sample of residuals is drawn, we multiply each residual by the corresponding $\widehat{\sigma }(x)$ when generating a bootstrap sample of observations of the dependent variable.

We implement the testing procedure by conducting the Khmaladze transformation both from the right end of the support (as is described theoretically in this paper) and from the left end of the support and obtained extremely similar results. More specifically, we only report results when the transformation is conducted from the right end of the support.

In the case of $P$-splines, we use the same B-spline basis, take the second-order penalty and choose the penalization constant using the ordinary cross-validation criterion as in Eilers and Marx (1996). The penalty enters unconstrained optimization problems as well as constrained ones.

Tables (ref)-(ref) present results of our testing. Namely, Table (ref) shows test statistics for the null hypothesis of the monotonically increasing regression function and also bootstrap critical values using both B-splines and P-splines. Table (ref) presents analogous results for the null hypothesis of convexity of the regression function. Table (ref) gives results for the joint null hypothesis of monotonicity and convexity.

table[table omitted — 870 chars of source]
table[table omitted — 870 chars of source]
table[table omitted — 884 chars of source]

As we can see from Tables (ref)--(ref), we do not reject any of the three hypotheses even at 10% level.

$\left. {}\right.$

2. Energy consumption in the Southern region of Russia.

The data are on daily energy consumption (in MWh) and average daily temperature (in Celsius) in the Southern region of Russia in the period from February 1, 2016 till January 31, 2018. The data have been downloaded from the official website of System Operator of the Unified Energy System of Russia.\footnote{ http://so-ups.ru/}

We provide tests for U-shape with a switch at ${17.6}^{\circ }$ and convexity using the approaches outlined in section (ref) in the context of Examples 2 and 1, respectively. In order to correct for heteroscedasticity of the errors, we estimate the scedastic function $ \widehat{\sigma }^{2}(x)$ using residuals obtained in the unconstrained estimation using B-splines (or P-splines, respectively). The scedastic function is estimated using cubic B-splines with 6 uniform knots and in the form of

equation*[equation* omitted — 78 chars of source]

Figure (ref) gives scatter plots of the data together with fitted curves obtained under the U-shape constraint with the switch at $s_{0}={17.6} ^{\circ }$. This constraint fit is obtained in accordance with the technique in section (ref). Namely, we consider individual B-spline fits on intervals $[\underline{x},s_{0}]$ and $[s_{0},\overline{x} ] $, where $\underline{x}$ and $\overline{x}$ are respectively lowest and highest values of the temperature in the sample. On each interval we use $ L^{\prime }+1=5$ uniform knots. The left-hand side figure only imposes the continuity of the fitted curve at the switch point, whereas the right-hand side figure imposes continuous differentiability.

figure[figure omitted — 774 chars of source]

Tables (ref)-(ref) present results of our testing. Namely, Table (ref) shows test statistics for the null hypothesis of U-shaped regression function and also bootstrap critical values using both B-splines and P-splines in case when two B-spline curves are joined at the switch point in a continuous way. Table (ref) presents analogous results for the null hypothesis of U-shaped regression function when two B-spline curves are joined at the switch point in a continuously differentiable way. Table (ref) gives results for the null hypothesis of convexity. In all the cases Khmaladze's transformation is conducted from the right end of support. he bootstrap critical values obtained on the basis of $400$ bootstrap replications. As we can see, the null hypothesis of a \emph{U-shaped} relationship with the switch point at ${17.6}^{\circ }$ is not rejected at the $5\%$ level by any type of the test, whereas convexity is rejected. When testing convexity we use cubic splines with $L^{\prime }+1=7$ uniform knots on $[\underline{x},\overline{x}]$.

table[table omitted — 971 chars of source]
table[table omitted — 966 chars of source]
table[table omitted — 874 chars of source]

CONCLUSION

$\left. {}\right.$

This paper proposes a methodology for testing a wide range of shape properties of a regression function. The methodology relies on applying the Khmaladze's transformation to the partial sums empirical process in a nonparametric setting where B-splines have been used to approximate the functional space under the null hypothesis. We establish that the proposed Khmaladze's transformation eliminates the effect of nonparametric estimation and results in asymptotically pivotal testing. To the best of our knowledge, this paper is the first implementation of the Khmaladze's transformation in a nonparametric setting.

In our main examples we considered shape constraints that can be written as inequality constraints on the coefficients of the approximating regression splines. The generality of our procedure allows to test several shape properties simultaneously. The implementation is especially easy when the inequality constraints are linear, which is the case for shape properties expressed as linear inequality constraints on the derivatives.