Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
51,125 characters · 10 sections · 15 citation commands
Improving Point and Interval Estimates of Monotone Functions by Rearrangement
A common problem in statistics is the estimation of an unknown monotonic function. Examples of monotonic functions include biometric age-height charts, econometric demand functions, and quantile and distribution functions. If an original, potentially non-monotonic, estimate is available, then the rearrangement operation from variational analysis HLP52,lorentz:1953,villani can be used to monotonize the original estimate. The rearrangement has been shown to be useful in producing monotonized estimates of density functions fougeres1997, conditional mean functions DZ2005,dette1,dette3, and various conditional quantile and distribution functions, see, e.g., \citeasnoun{CFG-RE-ET} and the MIT working paper “Quantile and Probability Curves without Crossing” by the authors.
In this paper, we use Lorentz inequalities and their appropriate generalizations to show that the rearrangement of the original estimate is not only useful for producing monotonicity, but also always improves upon the original estimate, whenever the latter is not monotonic. Thus, the rearranged curves are always closer to the target curve being estimated. Furthermore, this improvement property does not depend on the nature of the original estimate and applies to both univariate and multivariate cases. The improvement property of the rearrangement also extends to the construction of confidence bands for monotone functions. We show that we can increase the coverage probabilities and reduce the lengths of the confidence bands for monotone functions by rearranging their upper and lower bounds.
Monotonization has a long history in the statistical literature, mostly in relation to isotone regression. We will not provide an extensive literature review, but reference a few other methods most related to the rearrangement. \citeasnoun{Mammen1991} studies two-step estimators, including one with smoothing in the first step and monotonization by isotone regression in the second. \citeasnoun{MMTW2001} show that this and many related procedures can be recast as projections with respect to a given norm. Another approach is the one-step procedure of \citeasnoun{Ramsay1988}, which projects on a class of monotone spline functions called I-splines. Later in the paper we will compare and combine these procedures with the rearrangement.
A basic problem in many areas of statistics is the estimation of an unknown target function $f_0: \Bbb{R}^d \to \Bbb{R}$. Suppose we know that $f_0$ is monotonic, namely weakly increasing, and an original estimate $\hat f$ is available, which is not necessarily monotonic, but is theoretically attractive and computationally tractable otherwise. Many common estimation methods do indeed produce such estimates. Can they always be improved with no harm? The answer is yes: the rearrangement method transforms the original estimate to a monotonic estimate $\hat f^*$, and this estimate is closer in common metrics to the true curve $f_0$ than the original estimate $\hat f$. Furthermore, the rearrangement is computationally tractable, and thus preserves the appeal of the original estimates.
Estimation methods used in regression analysis can be grouped into global methods and local methods. An example of a global method is the series estimator of $f_0$ taking the form $\hat f(x) = P_{k_n}(x)'\hat b,$ where $P_{k_n}(x)$ is a $k_n$-vector of suitable transformations of the variable $x$, such as B-splines, polynomials, and trigonometric functions, and $$ \hat b = \arg \min_{b \in \Bbb{R}^{{k}_n}} \sum_{i=1}^n \rho\{Y_i - P_{k_n}(X_i)' b\}, $$ where $\{(Y_i, X_i), i=1,\ldots,n\}$ denotes the data. In particular, using the square loss $\rho(u) = u^2$ produces estimates of the conditional mean of $Y_i$ given $X_i$ gallant:fourier,andrews:series,Stone:series,newey:series, while using the asymmetric absolute deviation loss $\rho(u) = \{u - 1(u<0)\} u$ produces estimates of the conditional $u$-quantile of $Y_i$ given $X_i$ koenker:1978,portnoy:splines,he:series. The series estimates $x \mapsto \hat f(x)= P_{k_n}(x)'\hat b$ are widely used in data analysis due to their desirable approximation and theoretical properties, and computational tractability. However, they need not be monotone, unless explicit constraints are added matzkin:handbook,silvapulle:book,koenker:inequality.
Examples of local methods include kernel and local polynomial estimators. A kernel estimator takes the form $$ \hat f(x) = \arg \min_{b \in {\Bbb{R}}} \ \sum_{i=1}^n w_i \rho(Y_i - b), \ \ w_i = K\left(\frac{X_i -x}{h}\right), $$ where the loss function $\rho$ plays the same role as above, $K(u)$ is a multivariate kernel function, and $h> 0$ is a vector of bandwidths wand:jones,silverman:book. The resulting estimate $x \mapsto \hat f(x)$ need not be monotone. \citeasnoun{dette1} show that the rearrangement transforms the kernel estimate into a monotonic one. We further show here that the rearranged estimate necessarily improves upon the original estimate, whenever the latter is not monotonic. Local polynomial regression is a related local method chaudhuri:1991,fan:book. In particular, the local linear estimator takes the form $$ \{\hat f(x), \hat d(x)\} = \underset{b \in {\Bbb{R}}, c \in \Bbb{R}^d}{\text{argmin}} \ \sum_{i=1}^n w_i \rho\{Y_i - b - c'(X_i -x)\}^2, \ \ w_i = K\left(\frac{X_i -x}{h}\right). $$ The resulting estimate $x \mapsto \hat f(x)$, while theoretically attractive and computationally tractable, may also be non-monotonic, as illustrated in Section 4.
In what follows, let $\mathcal{X}$ be a compact interval; without loss of generality we take $\mathcal{X}=[0,1]$. Let $f$ be a measurable function mapping $\mathcal{X}$ to $K$, a bounded subset of $\Bbb{R}$. The increasing rearrangement $f^*$ of $f$ is the quantile function of the random variable $f(X)$ when $X\sim U(0,1)$, that is, $$f^*(x) = \inf \left\{ y \in \Bbb{R}: \int_{\mathcal{X}} 1\{ f(u) \leq y \} du \geq x \right\}. $$ The rearrangement operator simply transforms a function $f$ to its quantile function $f^*$. For computing purposes when $f$ is continuous, we can think of the rearrangement as a sorting operation: given values of the function $f$ evaluated at $x$ in a fine enough net of equidistant points, we simply sort the values in increasing order to create the sorted, i.e., rearranged, function.
Proposition 1 establishes that the rearranged estimate $\hat f^*$ has a smaller, often strictly smaller, estimation error in the $L_p$ norm than the original estimate whenever the latter is not monotone. This very useful and generally applicable property is independent of the sample size and of the way the original estimate $\hat f$ is obtained. As follows from ((ref)), the reduction in estimation error is strict for $L^p$ norms with $p \in (1, \infty)$ if the original estimate $\hat f$ is decreasing on a subset of $\mathcal{X}$ having positive measure, while the target function $f_0$ is increasing on this subset. If $f_0$ is constant, then there is no reduction in estimation error; that is, the inequality ((ref)) becomes an equality, since the random variables $\hat f^*(X)$ and $\hat f(X)$ share the same quantile function $\hat f^*$ and hence the same distribution function, and $f_{0}(X)$ is constant.
The weak inequality ((ref)) is a direct, yet important, consequence of the classical rearrangement inequality due to \citeasnoun{lorentz:1953}: let $q$ and $g$ be two functions mapping $\mathcal{X}$ to $K$, and $q^{*}$ and $g^{*}$ be their corresponding increasing rearrangements, then $ \int_{\mathcal{X}} L\{q^*(x), g^*(x)\} d x \leq \int_{\mathcal{X}} L\{q(x), g(x) \} dx, $ for any submodular discrepancy function $L: \Bbb{R}^2 \mapsto \Bbb{R}_+$. We set $q = \hat f$, $q^* = \hat f^*$, $g = f_0$, and $g^* = f_0^*$. In our case $ f_0^* = f_0$ almost everywhere, that is, the target function is its own rearrangement. Further, recall that $L$ is submodular if for each pair of vectors $(v,t)$ and $(v',t')$ in $\Bbb{R}^2$, we have that
In other words, a function $L$ measuring the discrepancy between pairs of vectors is submodular if co-monotonization of the pair reduces the discrepancy. When the function $L$ is smooth, submodularity is equivalent to $\partial^2 L(v,t)/ (\partial v \partial t) \leq 0$ holding for each $(v, t)$ in $\Bbb{R}^2$. Thus, for example, power functions $L(v,t) = |v-t|^p$ for $p \in [1,\infty)$ and many other loss functions are submodular. The weak inequality ((ref)) then follows.
In this section we consider multivariate functions $f: \mathcal{X}^d \to K$, where $\mathcal{X}^d = [0,1]^d$ and $K$ is a bounded subset of $\Bbb{R}$. The notion of monotonicity we seek to impose on $f$ is the following: we say that the function $f$ is weakly increasing in the vector $x$ if $f(x') \leq f(x)$ whenever $x' \leq x$ (componentwise). In what follows, we use $f(x_j, x_{-j})$ to denote the dependence of $f$ on $x_j$, and all other arguments, $x_{-j}$, that exclude $x_j$. The notion of monotonicity above is equivalent to the requirement that for each $j$ in $1,\ldots,d$ the mapping $x_j \mapsto f(x_j, x_{-j})$ is weakly increasing in $x_j$, for each $x_{-j}$ in $\mathcal{X}^{d-1}$.
Define the rearrangement operator $R_j$ and the rearranged function $f_j^*$ with respect to $x_j$ as $$ f_j^*(x) = R_j f(x) = \inf \left \{ y: \left [\int_{\mathcal{X}} 1 \{ f(x_j', x_{-j}) \leq y \} d x_j' \right] \geq x_j \right \}.$$ This is the one-dimensional increasing rearrangement applied to the one-dimensional function $x_j \mapsto f(x_j, x_{-j})$, holding the other arguments $x_{-j}$ fixed. The rearrangement is applied for every value of the other arguments $x_{-j}$.
Let $\pi = (\pi_1,\ldots,\pi_d)$ be an ordering, i.e., a permutation, of the integers $1,\ldots,d$. Let us define the $\pi$-rearrangement operator $R_{\pi}$ and the $\pi$-rearranged function $f_\pi^*$ as $ f_{\pi}^* = R_\pi f = R_{\pi_1} \ldots R_{\pi_d} f.$ For any ordering $\pi$, the $\pi$-rearrangement operator rearranges the function with respect to all of its arguments. As shown below, the resulting function $f_\pi$ is weakly increasing in $x$. In general, two different orderings $\pi$ and $\pi'$ of $1,\ldots,d$ can yield different rearranged functions $f_{\pi}^*$ and $f_{\pi'}^*$. To resolve the conflict among rearrangements done with different orderings, we may consider averaging among them: letting $\Pi$ be any finite collection of orderings $\pi$, we can define the average rearrangement as $$ f^* = \frac{1}{|\Pi|} \sum_{\pi \in \Pi} f_\pi^*, $$ where $|\Pi|$ denotes the number of elements in the set of orderings $\Pi$. \citeasnoun{dette3} also proposed averaging all the possible orderings of a related smoothed procedure in the context of monotone conditional mean estimation. As shown below, the estimation error of the average rearrangement is weakly smaller than the average of estimation errors of individual $\pi$-rearrangements.
The following proposition describes the properties of multivariate $\pi$-rearrangements:
Proposition 2 generalizes Proposition 1 to the multivariate case, also demonstrating several features unique to the multivariate case. We see that the $\pi$-rearranged functions are monotonic in all of the arguments. \citeasnoun{dette3}, using a different argument, showed that their related smoothed procedure for conditional mean functions is monotonic in both arguments for the bivariate case in large samples. The rearrangement along any argument improves the estimation properties. Moreover, the improvement is strict when the rearrangement with respect to a $j$-th argument is performed on an estimate that is decreasing in the $j$-th argument, while the target function is increasing in the same $j$-th argument, in the sense precisely defined in the proposition. Averaging different $\pi$-rearrangements is better on average than using a single $\pi$-rearrangement chosen at random.
Here we informally explain why rearrangement provides the improvement property and compare rearrangement to isotonization.
We begin by noting that the proof of the improvement property can be first reduced to the case of step functions or, equivalently, functions with a finite domain, and then to the case of functions with a two-point domain. The improvement property for such functions then follows from the submodularity property ((ref)). In the left panel of Figure (ref) we illustrate this geometrically by plotting the original estimate $\hat f$, the rearranged estimate $\hat f^*$, and the true function $f_0$. In this example, the original estimate is decreasing and hence violates the monotonicity requirement. We see that the two-point rearrangement co-monotonizes $\hat f^*$ with $f_0$ and thus brings $\hat f^*$ closer to $f_0$. Also, we can view the rearrangement as a projection on the set of weakly increasing functions that have the same distribution as the original estimate $\hat f$.
In the right panel of Fig. (ref) we plot both the rearranged and isotonized estimates. The isotonized estimate $\hat f^I$ is a projection of the original estimate $\hat f$ on the set of weakly increasing functions, that only preserves the mean of the original estimate. We can compute the two values of the isotonized estimate $\hat f^I$ by assigning to both the average of the two values of the original estimate $\hat f$, whenever the latter violate the monotonicity requirement, and leaving the original values unchanged otherwise. In our example in Fig. (ref) this produces a flat function $\hat f^I$. This pool adjacent violators procedure extends to domains with more than two points by applying the procedure iteratively to any pair of points at which monotonicity is violated PAVA1955.
Using the computational definition of isotonization, one can show that, like rearrangement, isotonization also improves upon the original estimate, for any $p \in[1,\infty]$:
see, e.g., \citeasnoun{barlow:book}. Therefore, it follows that any function $\hat f^{\lambda}$ in the convex hull of the rearranged and isotonized estimate both (1) monotonizes and (2) improves upon the original estimate $\hat f$, that is, for any $p \in[1,\infty]$ and $\lambda \in [0,1]$,
where $\hat f^{\lambda} = \lambda \hat f^* + (1 - \lambda) \hat f^I$. The first property is obvious and the second follows from homogeneity and subadditivity of norms. By induction on the dimension, the improvement property extends to the sequential multivariate isotonization and to its convex hull with the sequential multivariate rearrangement.
Thus, we see that a rather rich class of procedures both monotonizes the original estimate and reduces the distance to the true target function. However, there is no single best distance-reducing monotonizing procedure. Indeed, whether the rearranged estimate $\hat f^*$ approximates the target function better than the isotonized estimate $\hat f^I$ depends on how steep or flat the target function is. We illustrate this point using the example plotted in the right panel of Fig. (ref): consider any increasing target function taking values in the shaded area between $\hat f^*$ and $\hat f^I$, and also the function $\hat f^{1/2}$, the average of the isotonized and the rearranged estimate, that passes through the middle of the shaded area. Suppose first that the target function is steeper than $\hat f^{1/2}$, then $\hat f^*$ has a smaller estimation error than $\hat f^I$. Now suppose instead that the target function is flatter than $\hat f^{1/2}$, then $\hat f^I$ has a smaller estimation error than $\hat f^*$. It is also clear that, if the target function is neither very steep nor very flat, $\hat f^{1/2}$ can outperform either $\hat f^*$ or $\hat f^I$. Thus, in practice we can choose rearrangement, isotonization, or, some combination of the two, depending on our beliefs about how steep or flat the target function is in a particular application.
In this section we propose to directly apply the rearrangement, univariate and multivariate, to simultaneous confidence intervals for monotone functions. We show that our proposal will necessarily improve the original intervals by decreasing their length while retaining the same or greater coverage level.
Suppose that we are given an initial simultaneous confidence interval
where $\ell$ and $u$ are the lower and upper end-point functions such that $\ell \leq u$ on $\mathcal{X}^d$, that is, $\ell(x) \leq u(x)$ for all $x \in \mathcal{X}^d$. We further suppose that the confidence interval $[\ell, u]$ has either the exact or the asymptotic confidence property for the estimand function $f$, namely, for a given $\alpha \in (0,1)$,
for all probability measures $P$ in some set $\mathcal{P}_n$ containing the true probability measure $P_0$. The statement $f \in [\ell, u]$ means that $\ell(x) \leq f(x) \leq u(x)$ for all $x \in \mathcal{X}^d$. We assume that property ((ref)) holds either in the finite sample sense, that is, for the given sample size $n$, or in the asymptotic sense, that is, for all but finitely many sample sizes $n$ LR2005.
A common confidence interval for functions specifies
where $\hat f(x)$ is a point estimate, $s(x)$ is the standard error of the point estimate, and $\hat c$ is a critical value chosen to attain the confidence property ((ref)). \citeasnoun{Wasserman06} provides an excellent overview of methods for constructing the critical value. The problem with such confidence intervals, as with the point estimates themselves, is that they need not be monotonic. Indeed, the end-point functions ((ref)) need not be monotonic, so the confidence interval may contain non-monotone functions excludable from it. Accordingly we can intersect the interval with the set of monotone functions to reduce its length without affecting its coverage level. In some cases, however, the initial interval may not contain any monotone function and the resulting intersected interval is empty, due, for example, to misspecification.
We say that confidence intervals are misspecified or incorrectly centered if the estimand $f$, being covered by $[\ell, u]$ in ((ref)), is not equal to the weakly increasing target function $f_0$, so that $f$ may not be monotone. Incorrect centering is rather common both in parametric and non-parametric estimation. In parametric estimation correct centering of confidence intervals requires perfect specification of functional forms, whereas in nonparametric estimation correct centering requires the so-called undersmoothing; both are difficult. In real applications with many regressors, researchers tend to use oversmoothing rather than undersmoothing. In a recent development, \citeasnoun{GW08} provide some formal justification for oversmoothing: targeting inference on functions $f$, that represent various smoothed versions of $f_0$ and thus summarize features of $f_0$, may be desirable to make inference more robust, or, equivalently, to enlarge the class of data-generating processes $\mathcal{P}_n$ for which ((ref)) holds. Regardless of the reasons for why the confidence intervals may target $f$ instead of $f_0$, our procedures will work for inference on the monotonized, hence improved, version $f^*$ of $f$.
Our proposal for improved interval estimates is to rearrange the entire simultaneous confidence interval into a monotonic interval
where the lower and upper end-point functions $\ell^*$ and $u^*$ are the increasing rearrangements of the original end-point functions $\ell$ and $u$. In the multivariate case, we use the symbols $\ell^*$ and $u^*$ to denote either multivariate $\pi$-rearrangements $\ell_\pi^*$ and $u_\pi^*$ or average multivariate rearrangements $\ell^*$ and $u^*$, whenever we do not need to emphasize specifically the dependence on $\pi$.
The following proposition describes the properties of the rearranged confidence intervals.
Proposition 3 shows that the rearranged confidence intervals are weakly shorter than the original confidence intervals, and also qualifies when the rearranged confidence intervals are strictly shorter. In particular, in the univariate case the inequality ((ref)) is necessarily strict for $p\in (1, \infty)$ if there is a region of positive measure in $\mathcal{X}$ over which the end-point functions $\ell$ and $u$ are not comonotonic. This weak shortening result follows for univariate cases directly from the Lorentz (1953) inequality, and the strong shortening by its strengthening. The shortening results for the multivariate case follow by induction on the dimension. Moreover, the order-preservation property of the univariate and multivariate rearrangements, demonstrated in the proof, implies that the rearranged confidence interval $[\ell^*, u^*]$ has a weakly higher coverage than the original confidence interval $[\ell, u]$. We do not quantify strict improvements in coverage, but demonstrate them through the examples in the next section.
Our idea of directly monotonizing the interval estimates also applies to other monotonization procedures. Indeed, the proof of Proposition 3 reveals that part 1 applies to any order-preserving monotonization operator $T$, such that
Furthermore, part 2 of Proposition 3 on the weak shortening of the confidence intervals applies to any distance-reducing operator $T$ such that
Rearrangements are instances of operators that have properties ((ref)) and ((ref)). Isotonization is another important instance RWD88. Moreover, convex combinations of order-preserving and distance-reducing operators, such as the average of rearrangement and isotonization, also have properties ((ref)) and ((ref)).
In this section we provide an empirical application to biometric age-height charts. We show how the rearrangement monotonizes and improves various nonparametric point and interval estimates for functions.
Since their introduction by Quetelet in the 19th century, reference growth charts have become common tools to assess an individual's health status. These charts describe the evolution of individual anthropometric measures, such as height, weight, and body mass index, across different ages. See \citeasnoun{cole} for a classical work on the subject, and \citeasnoun{koenker:charts} for a recent analysis from a quantile regression perspective and additional references. Here we consider an application of the rearrangement and other related methods to the estimation of growth charts for height. This makes sense since an individual's height should follow an increasing relationship with age up to adulthood. Our data consist of repeated cross sectional measurements of height in centimeters and age in months from the 2003-2004 US National Health and Nutrition Survey, and is further restricted to the subsample of US-born white males aged 2-20 to avoid other confounding factors, giving us a sample of 533 observations.
Let $Y$ and $X$ denote height and age, respectively. Let $E[Y\mid X=x]$ denote the conditional expectation of $Y$ given $X=x$, and $Q_Y[u\mid X=x]$ denote the conditional $u$-th quantile of $Y$ given $X=x$, where $u$ is the quantile index. The target functions of interests are the conditional expectation function, $x \mapsto E[Y\mid X=x]$, the conditional quantile functions for several quantile indices, $x\mapsto Q_Y[u\mid X=x]$, for $u =5\%$, $50\%$, and $95\% $, and the entire conditional quantile process for height given age, $(u,x)\mapsto Q_Y[u\mid X=x]$. The monotonicity requirements for these target functions are the following: the first two should be increasing in age $x$, and the third should be increasing in both age $x$ and the quantile index $u$.
We estimate the target functions using non-parametric ordinary least squares or quantile regression and then rearrange the estimates to satisfy the monotonicity requirements. We consider kernel, local linear, regression splines, and Fourier series methods. For the kernel and local linear methods, we choose a bandwidth of one year and a box kernel. For the regression splines method, we use cubic B-splines with a knot sequence $\{ 3, 5, 8, 10, 11.5, 13, 14.5, 16, 18 \}$ koenker:charts. For the Fourier method, we employ four sines and four cosines. For the estimation of the conditional quantile process, we use $\{ 0.005, 0.010, \ldots, 0.995 \}$ as a net of quantile indices.
Figure (ref) shows the original and rearranged estimates of the conditional quantile functions for the different methods. All the estimated curves have trouble capturing the slowdown in the growth of height after age fifteen and yield non-monotonic curves for the highest values of age. The Fourier series performs particularly poorly in approximating the aperiodic age-height relationship and has many non-monotonicities. The rearrangement delivers curves that improve upon the original estimates and that satisfy the natural monotonicity requirement. We quantify this improvement in the next subsection.
Figure (ref) (a,b) illustrates the multivariate rearrangement of the conditional quantile process along both the age and the quantile index arguments. We plot, in three dimensions, the original estimate and its average multivariate rearrangement (the average of the age-quantile and quantile-age rearrangements). We focus on the Fourier series estimates, which have the most severe non-monotonicity problems. Analogous figures for the other estimation methods are given in an MIT working paper containing an extended version of this article. We see that the estimated quantile process is non-monotone in age and in the quantile index at extremal values of this index. The average multivariate rearrangement fixes the non-monotonicity problem delivering an estimate of the quantile process that is monotone in both the age and the quantile index. Furthermore, by the theoretical results of the paper, the multivariate rearranged estimates necessarily improve upon the original estimates.
In Figures (ref) and (ref) (c,d), we plot original and rearranged 90% simultaneous confidence intervals. Fig. (ref) shows the intervals for the conditional expectation function and for the conditional 5%, 50%, and 95% quantile functions, based on Fourier series estimates. We obtain the original intervals of the form ((ref)) using the bootstrap with 200 repetitions to estimate the standard errors and critical values Hall1993. We then obtain the rearranged confidence intervals by rearranging the lower and upper end-point functions of the initial confidence intervals, following Section 3. In Fig. (ref) (c,d), we plot the original and the rearranged 90% simultaneous confidence intervals for the entire conditional quantile process, based on the Fourier series estimates. The rearranged confidence intervals correct the non-monotonicity of the original confidence intervals and reduce their integrated $L^p$ length.
In the following Monte Carlo experiment we quantify the improvement in the point and interval estimation that rearrangement can provide relative to the original estimates. We also compare it to isotonization and to its convex combinations with isotonization. Our experiment uses a model, described in detail in the Appendix, that mimics the empirical application very closely. This model implies a true conditional expectation function and quantile process that are monotone in age and in the quantile index.
In Table (ref) we report the average $L^p$ errors, for $p=1,2,$ and $\infty$, for the original estimates of the conditional expectation function. We also report the relative efficiency of the rearranged estimates, measured as the ratio of the average error of the rearranged estimate to the average error of the original estimate; together with relative efficiencies for alternative approaches based on isotonization of the original estimates Mammen1991 and on averaging the rearranged and isotonized estimates. For regression splines, we also consider the one-step monotone regression splines Ramsay1998.
For all of the methods and norms considered, the rearranged curves estimate the target function more accurately than the original curves. There is no uniform winner between rearrangement, isotonization, and the average of the two, which is consistent with the analysis of Section 2.4. For example, the rearrangement outperforms the other methods for kernel, local linear and splines, but performs worse than the average for Fourier in some norms. In numerical results not reported, we find that rearrangement performs worse than isotonization for global polynomials. This and other methods are available in the MIT working paper. For regression splines, the performance of the rearrangement is comparable to the computationally more intensive one-step monotone splines procedure.
In Table (ref) we report the average $L^p$ errors for the original estimates of the conditional quantile process. We also report the ratio of the average error of the multivariate rearranged estimate, with respect to the age and quantile index arguments, to the average error of the original estimate; together with the same ratios for isotonized and average rearranged-isotonized estimates. We obtain the multivariate isotonized estimates by sequentially applying the univariate isotonization to each argument, and then averaging for the two possible orderings age-quantile and quantile-age. For all the methods and norms considered, the multivariate rearranged curves estimate the target function more accurately than the original curves. There is again no uniform winner between rearrangement, isotonization, and their average.
Table (ref) reports Monte Carlo coverage frequencies and integrated lengths for the original and monotonized 90% confidence bands for the conditional expectation function. For a measure of length, we used the integrated $L^p$ length, as defined in Proposition 3, with $p = 1,2,$ and $\infty$. We construct the original confidence intervals of the form specified in equations ((ref)) by obtaining the pointwise standard errors of the original estimates using the bootstrap with 200 repetitions, and calibrate the critical value so that the original confidence bands cover the entire true function with the exact frequency of 90%. We construct monotonized confidence intervals by applying rearrangement, isotonization, and a rearrangement-isotonization average to the end-point functions of the original confidence intervals, as proposed in Section 3. In all cases the rearrangement and other monotonization methods increase the coverage of the confidence intervals while reducing their length. In particular, we see that monotonization increases coverage especially for the local estimation methods, whereas it reduces length most noticeably for the global estimation methods. For the most problematic Fourier estimates, there are large increases in coverage and reductions in length.