Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
443,289 characters · 55 sections · 0 citation commands
Quantile-regression methods for first-price auctions
\thispagestyle{empty}
The paper proposes a quantile-regression inference framework for first-price auctions with symmetric risk-neutral bidders under the independent private-value paradigm. It is first shown that a private-value quantile regression generates a quantile regression for the bids. The private-value quantile regression can be easily estimated from the bid quantile regression and its derivative with respect to the quantile level. This also allows to test for various specification or exogeneity null hypothesis using the observed bids in a simple way. A new local polynomial technique is proposed to estimate the latter over the whole quantile level interval. Plug-in estimation of functionals is also considered, as needed for the expected revenue or the case of CRRA risk-averse bidders, which is amenable to our framework. A quantile-regression analysis to USFS timber is found more appropriate than the homogenized-bid methodology and illustrates the contribution of each explanatory variables to the private-value distribution. Linear interactive sieve extensions are proposed and studied in the Appendices.
JEL:\ C14, L70
Keywords: First-price auction; independent private values; dimension reduction; quantile regression; local polynomial estimation; specification testing; boundary correction; sieve estimation.
{ A previous version of this paper has been circulated under the title "Quantile regression methods for first-price auction:a signal approach". The authors acknowledge useful discussions and comments from Xiaohong Chen, Valentina Corradi, Yanqin Fan, Phil Haile, Xavier d'Haultfoeuille, Vadim Marmer, Isabelle Perrigne, Martin Pesendorfer and Quang Vuong, and the audience of many conferences and seminars. Nathalie Gimenes also thanks Ying Fan and Ginger Jin for encouragements. Many thanks to Elie Tamer and three anonymous referees, who all have been extremely stimulating and helpful to enrich the paper. All remaining errors are our responsibility. Both authors would like to thank the School of Economics and Finance, Queen Mary University of London, for generous funding. Nathalie Gimenes would like to express her gratitude to PUC Rio for generous support.}
\setcounter{page}{1}
Since Paarsch (1992), many parametric methods have been proposed to estimate first-price auction models under the independent private-value paradigm. See Laffont, Ossard and Vuong (1995), Athey and Levin (2001), Hirano and Porter (2003), Li and Zheng (2012), Paarsch and Hong (2012) and the references therein to name just a few. Validating specification choice is difficult and seldom attempted.
On the other hand, the nonparametric approach is very flexible and less subject to misspecification of functional form, so that it is commonly considered in applications and theoretical studies. See Guerre, Perrigne and Vuong (2000, hereafter GPV), Lu and Perrigne (2008), Krasnokutskaya (2011), Marmer and Shneyerov (2012), Hubbard, Paarsch and Li (2012), Campo, Guerre, Perrigne and Vuong (2013), Marmer, Shneyerov and Xu (2013a,b), Hickman and Hubbard (2015), Enache and Florens (2017), Liu and Luo (2017), Liu and Vuong (2018), Luo and Wan (2018), Zincenko (2018) and Ma, Marmer and Shneyerov (2019) among others. But the nonparametric approach comes with the burden of the curse of dimensionality, which considerably limits its scope of applications.
Haile, Hong and Shum (2003, HSS hereafter) and Rezende (2008) have proposed to circumvent the curse of dimensionality using a regression specification that purges the bids from the covariate effects. The resulting homogenized bids are then used as in GPV to backup the density of their private-value counterparts. This approach can tackle linear dependence, but is not appropriate to capture more complex interactions. The present paper proposes to use instead a more flexible quantile-regression specification.
The use of quantile in first-price auctions is not new. Milgrom (2001, Theorem 4.7) reformulates the identification relation of Guerre, Perrigne and Vuong (2000, GPV afterwards) using quantile function. See Guerre, Perrigne, Vuong (2009) and Campo et al. (2013) for the use of quantile in risk-aversion identification and, for related estimation methods, Menzel and Morganti (2013), Enache and Florens (2017), Liu and Vuong (2018), Luo and Wan (2018). Marmer and Shneyerov (2012) have proposed a quantile-based estimator of the private-value probability density function (pdf), which is an alternative to the two-step GPV method. See also Marmer, Shneyerov and Xu (2013b) who consider a nonparametric single-index quantile model. Guerre and Sabbah (2012) have noted that the private-value quantile function can be estimated using a one-step procedure from the estimation of the bid quantile function and its first derivative. Gimenes (2017) has developed a flexible but parsimonious quantile-regression estimation strategy for ascending auction. The present paper is however the first to develop a quantile inference framework in a first-price auction setting allowing for many covariates.
Using Koenker and Bassett (1978) quantile regression framework is appealing for several reasons. First, the quantile-regression specification is flexible enough to capture economically relevant effects as in Gimenes (2017), which could be ignored using parametric ones or less interpretable nonparametric models. These parsimonious specifications can be estimated with reasonable nonparametric rates, allowing implementation in small samples with rich covariate environment. Compared to GPV, this estimation method is one-step and only requests one bandwidth parameter, which theoretical choice follows from standard bias variance expansion. As detailed in (ref), quantile-regression specification can be enriched to include more nonparametric features using sieve extensions ranging from the additive specification of Horowitz and Lee (2005) to fully nonparametric one as in Belloni, Chernozhukov, Chetverikov and Fern\'{a}ndez-Val (2019). Second, the quantile approach comes with a stability property of linear specifications, which ensures that a private-value quantile regression generates an bid quantile-regression. This is key for our estimation procedure and also for testing, as it transfers many null hypotheses of interest for the latent private-value distribution to the bid quantile-regression slopes. Tests derived from Koenker and Xiao (2002), Escanciano and Velasco (2010), Rothe and Wied (2013), Escanciano and Goh (2014) or Liu and Luo (2017) can be used to test correct specification of the quantile-regression or homogenized-bid models, or exogeneity of the auction format and of entry. Third, the quantile representation used in the paper can play the role of a reduced form generated by a more complex model, such as the random-coefficient model considered in Berry, Levinsohn and Pakes (1995), Hoderlein, Klemel\"{a} and Mammen (2010), or Backus and Lewis (2019) to name just a few. In particular, random coefficients drawn from an elliptical distribution generates a quantile specification which extends homogenized bid and can be easily estimated.
Fourth, the proposed augmented quantile-regression estimation methodology is based upon local polynomial for quantile levels, and is therefore not affected by asymptotic boundary bias. This permits better estimation of the upper tail distribution than most nonparametric methods, which is important as the winner's private value is high for a large number of bidders. Fifth, it can also be used to recover important parameters such as the probability or cumulative density functions (pdf and cdf hereafter), mitigating the curse of dimensionality that affects most nonparametric methods. Plug in estimation of the seller expected revenue, optimal reserve price and of agent constant relative risk-aversion parameter are also considered.
The rest of the paper is organized as follows. The next section (ref) introduces our stability result for linear quantile specification. Section (ref) considers the homogenized-bid and random-coefficient specifications. Section (ref) reviews some testing strategies based upon the bid quantile regression. Section (ref) explains how to use our quantile specification for estimating agent's risk-aversion, seller's expected revenue, and the cdf and pdf of the private values. A difficulty of the quantile approach for first-price auction is the need to estimate the bid quantile derivative with respect to quantile levels, see Guerre and Sabbah (2012) and the reference therein for related approaches. Section (ref) introduces our new augmented quantile regression estimators, which use a quantile-level local-polynomial approach to jointly estimate the bid quantile regression and its higher-order derivatives. Sections (ref) and (ref) group our main theoretical results, including Integral Mean Squared Error (IMSE), optimal bandwidth choice, optimal uniform convergence rate and Central Limit Theorem for the proposed private-value quantile-regression estimators.
Our theoretical results are illustrated with a simulation experiment and an application to USFS first-price auctions in Sections (ref) and (ref). Some simulation experiments illustrate how the new estimation procedure improves on the GPV two-step density estimator and homogenized bids. A preliminary quantile-regression analysis of the bid quantile function suggests that the homogenized-bid technique should not be applied here because the quantile-regression slopes are not constant. The private-value quantile-regression slope functions reveal the covariates impact, and how strongly bidders in the top of the distribution can differ from the bottom. Section (ref) concludes the paper.
(ref) details an interactive localized sieve quantile extension and related theoretical results which are the counterparts of the ones obtained for the quantile-regression specification. (ref) briefly sketches the main proof arguments and states some preliminary lemmas used for the proofs of the two key bias and linearization results in (ref) and (ref), from which our main results follow. The two remaining Appendices group the proofs of our main and intermediary results.
\setcounter{equation}{0}
A single and indivisible object with some characteristic $X\in\mathbb{R}^{D}$ is auctioned to $I\geq2$ buyers. The potential number of bidders $I$ and $X$ are known to the bidders and the econometrician. Bids $B_i$ are sealed so that a bidder does not know the other bids when forming his own bid. The object is sold to the highest bidder who pays his bid to the seller, and all the bids $B_i$ are then observed by the econometrician. Under the symmetric IPV paradigm, each potential bidder is assumed to have a private value $V_{i}$, $i=1,\ldots,I$ for the auctioned object. A buyer knows his private value but not the other ones, the common distribution of the independent $V_{i}$ being common knowledge. The private-value conditional cdf $F\left( \cdot|X,I\right)$ has a bounded support, or equivalently the conditional private-value quantile function \[ V\left( \alpha|X,I\right) =F^{-1}\left( \alpha|X,I\right) ,\quad \alpha\text{ in }\left[ 0,1\right] , \] is finite for $\alpha=0$ and $\alpha=1$.
The private-value quantile function $V(\alpha|x,I)$ plays an important economic role. The bidder's rent at quantile level $\alpha$ is $V(\alpha|x,I)-B(\alpha|x,I)$ where $B(\cdot|x,I)$ is the bid conditional quantile function, and assuming bids depend in a monotonous way on private values as considered below. The private-value quantile conditional function is important to compute counterfactuals, such as the bid quantile function in an alternative auction mechanism. In particular, it can be used to compute the seller expected revenue achieved with any reserve price, see ((ref)) below. It allows, as a consequence, to compute an optimal reserve price, or more generally to propose suitable auction designs.
It is well-known that the bidder $i$ private-value rank \[ A_{i}=F\left( V_{i}|X,I\right) \] has a uniform distribution over $\left[ 0,1\right] $ and is independent of $X$ and $I$. It also follows from the IPV paradigm that the private-value ranks $A_{i}=1,\ldots,I$ are independent. The dependence between the private value $V_{i}$ and the auction covariates $X$ and $I$ is therefore fully captured by the non separable quantile representation
which, when the private values are generated by an economic structural model, can be also viewed as a nonparametric reduced form. Following Milgrom and Weber (1982) or Milgrom (2001), $V\left( \cdot |X,I\right) $ can be also interpreted as a valuation function, the private-value rank $A_{i}$ being the associated signal. In what follows, $G\left( \cdot|X,I\right) $ and $g\left( \cdot|X,I\right) $ stand for respectively the bid conditional cdf and pdf.
Maskin and Riley (1984) have shown that Bayesian Nash Equilibrium bids $B_{i}=\sigma\left( V_{i};X,I\right) $ of symmetric risk-averse or risk-neutral bidders are strictly increasing and continuous in $V_i$. It follows that $B_{i}=B\left( A_{i} |X,i\right) $, where $B\left( \cdot;X,i\right) =\sigma\left( F\left( \cdot|X,I\right) ;X,I\right) $ can be viewed as a bidding strategy depending upon the rank $A_{i}$. If $F\left( \cdot|X,I\right) $ is also strictly increasing, so is $B\left( \cdot|X,I\right) $ and since $A_{i}$ is uniform it holds \[ G\left( b|X,I\right) =\mathbb{P}\left[ B\left( A_{i}|X,I\right) \leq b|X,I\right] =\mathbb{P}\left[ A_{i}\leq B^{-1}\left( b|X,I\right) |X,I\right] =B^{-1}\left( b|X,I\right) \] showing that the bidding strategy\ $B\left( \cdot|X,I\right) $ is also the bid quantile function.
A standard best response argument will show how to identify the private-value quantile function $V\left( \cdot|X,I\right) $ from $B\left( \cdot |X,I\right) $. Suppose bidder $i$ signal $A_{i}$ is equal to $\alpha$, but that her bid is a suboptimal $B\left( a|X,I\right) $, all other bidders bidding $B\left( A_{j}|X,I\right) $. Then the probability that bidder $i$ wins the auction is
because the $A_{j}$'s are independent $\mathcal{U}_{\left[ 0,1\right] }$ independent of $X$ and $I$. It follows that the expected revenue of such a bid is, for a risk-neutral bidder, $\left( V\left( \alpha|X,I\right) -B\left( a|X,I\right) \right) a^{I-1}$. If $B\left( \cdot|X,I\right) $ is a best-response bidding strategy, the optimal bid of a bidder with signal $\alpha$ is $B\left( \alpha|X,I\right) $, that is \[ \alpha=\arg\max_{a}\left\{ \left( V\left( \alpha|X,I\right) -B\left( a|X,I\right) \right) a^{I-1}\right\} . \] As $B\left( \cdot|X,I\right) $ is continuously differentiable, it follows that
or equivalently
Solving with the initial condition $B\left( 0|X,I\right) =V\left( 0|X,I\right) $ and rearranging the equation above gives Proposition (ref), which is the cornerstone of our estimation method. From now on $B^{\left( 1\right) }\left( \alpha|X,I\right) =\frac{d}{d\alpha}B\left( \alpha|X,I\right) $.
A key feature is the linearity with respect to $V\left( \cdot|X,I\right)$ of the private-value to bid quantile functions mapping ((ref)), which implies that a private value quantile linear model is mapped into a similar bid linear model, as detailed below for the well-known quantile regression. Proposition (ref)-(ii) shows that the private-value quantile function is identified from the bid quantile function and its derivative. It is a quantile version of the identification strategy of GPV, which is based on the identity\footnote{This can be recovered from ((ref)) taking $\alpha=A_{i}$ as $V_{i}=V\left( A_{i}|X,I\right) $, $B_{i}=B\left( A_{i}|X,I\right) $ implying that $A_{i}=G\left( A_{i}|X,I\right) $ and $B^{\left( 1\right) }\left( A_{i}|X,I\right) =1/g\left( B\left( A_{i}|X,I\right) |X,I\right) =1/g(B_{i}|X,I)$.}
Versions of ((ref)) with $B^{\left( 1\right) }\left( \alpha|X,I\right) $ changed into $1/g\left( B\left( \alpha|X,I\right) |X,I\right) $ can be found in Milgrom (2001, Theorem 4.7), Liu and Luo (2014), Liu and Vuong (2016), Luo and Wan (2016), Enache and Florens (2017) and, under risk-aversion, in Guerre et al. (2009) and Campo et al. (2011).
The linearity of ((ref)) has important model stability implications useful for practical implementation. Consider a private-value quantile given by the quantile-regression specification
As a linear regression is often viewed as an alternative to a nonparametric one which is difficult to estimate, this quantile regression is simpler to estimate than a general quantile function which must be estimated nonparametrically. As pointed by a Referee, the quantile level $\alpha$ can be viewed as a measure of the bidder efficiency and the slope function $\gamma(\cdot|I)$ indicates how this efficiency affects valuation in the covariate dimension. While more flexible than the homogenized bid specification detailed in Section (ref), the quantile approach only involves a unique signal: in particular, if each slope entries are increasing, then each covariate contribution to the value increases with efficiency. More flexibility is possible with the random coefficient model of Section (ref), which attaches a specific signal to each auction covariate.
Proposition (ref)-(i) implies that the conditional bid quantile function satisfies,
showing that $B\left( \alpha|X,I\right) $ belongs to the quantile-regression specification. Hence ((ref)) gives
so that estimating $\gamma\left( \alpha|I\right) $ amounts to estimate $\beta\left( \alpha|I\right) $ and $\beta^{\left( 1\right) }\left( \alpha|I\right) $.
This approach extends to more flexible nonparametric linear specifications, as developed in (ref) which considers a sieve extension
where $P(\cdot)$ is a localized sieve vector whose dimension grows with a smoothing parameter $h$. The choice of $P(\cdot)$ can be tailored to cover additivity or less stringent interaction restrictions. As for the quantile-regression estimators proposed below, the sieve approach developed in (ref) is not affected by asymptotic boundary issues.
HHS and Rezende (2008) consider a regression specification
where the iid $v_i$, the “homogenized” private values, are independent of $X$ and not centered.\footnote{Centering the $v_i$'s would amount to introduce an intercept parameter $\gamma_0$, which would be changed to a new intercept $\beta_0 (I)$ when turning to the bid regression when the $v_i$'s are independent of $I$. In contrast, the bid regression slope is $\gamma_1$, therefore unchanged. Hence ((ref)) does not include an intercept to better focus on the invariant parameter, the purpose being to estimate $\gamma_1$ and the distribution of $v_i$. Estimating an intercept in the bid regression is however necessary to consistently estimate $\gamma_1$ using OLS because the regression error term in ((ref)) is not centered.} The corresponding homogenized-bid quantile-regression specification is the following restriction of ((ref)) \[ V(\alpha|X,I) = X^{\prime} \gamma_1 + v(\alpha|I) \] where $v(\cdot|I)$ is the quantile function of the $v_i$'s. Since $\frac{I-1}{\alpha^{I-1}}\int_{0}^{\alpha}a^{I-2}da=1$, it follows that the associated bid quantile function is, by ((ref)) \[ B\left( \alpha|X,I\right) =X^{\prime}\gamma_{1}+b\left( \alpha|I\right) ,\text{ where }b\left( \alpha|I\right) =\frac{I-1} {\alpha^{I-1}}\int_{0}^{\alpha}a^{I-2}v\left( a|I\right) da. \] This gives the bid regression model
where the $b_{i}$ are the homogenized bids of HHS, which are independent of $X$ but depend upon $I$. Given a sample $X_{\ell},I_{\ell}, B_{1 \ell}, \ldots, B_{I_{\ell}\ell}$ of $\ell=1,\ldots,L$ first-price auctions, HSS and Rezende (2008) propose to backup the homogenized bids by regressing the bids $B_{i\ell}$ on $X_{1\ell}=[1, X_{\ell}^{\prime}]^{\prime}$, so that the estimation of the homogenized bids are $\widehat{b}_{i\ell} =B_{i\ell} - X_{\ell}^{\prime} \widehat{\gamma}_1$, where $\widehat{\gamma}_1$ is the OLS slope estimator. The pdf of $v_{i}$ can be estimated applying the GPV two-step method to the homogenized-bid estimates. An important feature of this model is that the dependence of the private values to the covariate is simple enough to allow for accurate estimation of $\gamma_1$. As noted in Paarsch and Hong (2006), a similar two-step procedure applies for the nonparametric regression model $V_i = m(X|I) +v_i$ where the $v_i$'s are independent of $X,I$, see also Marmer, Shneyerov and Xu (2013b).
However this approach requests independence between the regression error term $v_{i}$ and the covariate $X$, an assumption which may be too restrictive in practice as found by Gimenes (2017) and the application below. When $\gamma_{1}\left( \cdot\right) $ is not a constant and $V (\alpha|X,I) = X^{\prime} \gamma_1 (\alpha|I) + v(\alpha|I)$, it holds for $\beta_1 (\alpha|I) =\frac{I-1}{\alpha^{I-2}} \int_{0}^{\alpha} a^{I-2} \gamma_{1} (a|I) da$ and the OLS limit $\beta_1 (I) = \mathbb{E} [\beta_1 (A_i|I)]$ obtained when regressing the bids on the constant and $X$, \[ B_i = X^{\prime} \beta_1 (I) + b (A_i|X,I) \text{ where } b (A_i|X,I) = b(A_i|I) + X^{\prime} \left[ \beta_1 (\alpha|I) - \beta_1 (I) \right]. \] As $b (A_i|X,I)$ depends upon $X$, the homogenized-bid approach does not apply. As explained below, estimating the slope $\gamma_1 (\cdot)$ involves nonparametric techniques that cannot deliver the parametric rate feasible in the homogenized-bid model.
Consider $I$ private values from
where the random coefficients $\Gamma_i$ are iid $1\times(D+1)$ vectors independent of $X_1$. Since $V_i = \Gamma_{0i} + X^{\prime} \Gamma_{1i}$, taking $\Gamma_{1i}$ constant across bidders gives a homogenized-bid specification, which is therefore a particular case of random-coefficients regression. Compared to ((ref)) version which involves a unique signal $A_i$, ((ref)) allows for $D+1$ individual signals $\Gamma_{id}$ which models the impact of the common covariate $X_{d}$ on the private value $V_i$. How a quantile approach can be useful is first discussed when $\Gamma_i$ is drawn from an elliptical distribution.
\paragraph{Elliptical random coefficient.} $\Gamma_{i}$ is drawn from an elliptical distribution with translation parameter $\gamma(I)$ and symmetric nonnegative dispersion matrix $\Sigma_{\Gamma}(I)$ if the characteristic function $\mathbb{E} \left[ \exp \left( \mathsf{i} t^{\prime} (\Gamma_i-\gamma(I))\right)\right]$ only depends upon $t^{\prime} \Sigma_{\Gamma} (I) t$. Examples include the multivariate normal, lognormal or Student distribution, which can be truncated to satisfy our finite support restriction. A convenient representation of $\Gamma_i$ involves the Euclidean norm $R_i=\left\| \Sigma_{\Gamma}^{-1/2} (I) \left(\Gamma_i-\gamma(I)\right)\right\|$ and independent draws $\mathcal{S}_i$ from the uniform distribution over the $D+1$ dimensional unit sphere. Let $\mathcal{C}_i$ be the first coordinate of $\mathcal{S}_i$, noticing that $t'\mathcal{S}_i$ is distributed as $\| t \| \mathcal{C}_i$ for any $(D+1) \times 1$ vector $t$. Then by Fang, Kotz and Ng (1990, p.29), $\Gamma_i$ and $\gamma(I) + R_i \Sigma_{\Gamma}^{1/2} (I)\mathcal{S}_i$ have the same distribution, for independent $R_i$ and $\mathcal{S}_i$. It then follows by ((ref)), $\stackrel{d}{=}$ indicating random variables with identical distribution \[ V_i \stackrel{d}{=} X_1^{\prime} \gamma(I) + \left( \Sigma_{\Gamma}^{1/2} (I)X_1 \right)^{\prime} R_i \mathcal{S}_i \stackrel{d}{=} X_1^{\prime} \gamma (I)+ \left\| \Sigma_{\Gamma}^{1/2} (I) X_1 \right\| R_i \mathcal{C}_i. \] Hence the quantile specification generated by ((ref)) is
where the unknown quantile function $v(\alpha|I)$ is the one of $R_i \mathcal{C}_i$ given $I$. The generated bids have a common quantile function
by ((ref)). Using the normalization $b(1/2|I)=1$ for identification purpose gives that $ B(1/2|X,I) = X_1^{\prime} \gamma (I)+ \left\| \Sigma_{\Gamma}^{1/2} (I) X_1 \right\| $, so that the conditional bid median can be used to identify $\gamma(I)$ and $\Sigma_{\Gamma} (I)$. Identification of $v(\cdot|I)$ works as in Proposition (ref) as $v(\alpha|I) = b(\alpha|I) + \alpha b^{(1)} (\alpha|I)/(I-1)$, observing that $v(\cdot|I)$ identifies the common distribution of the $R_i$'s.
\paragraph{The general case.} Hoderlein et al. (2010) propose a nonparametric method that could be used to estimate the distribution of the random slope $\Gamma_i$ of ((ref)) if the private values were observed. This suggests to implement a two-step method using estimated private values. (ref) proposes a sieve method to estimate $V(\cdot|x,I)$, which is not subject to asymptotic boundary bias. Consider $S$ estimated private values $\widehat{V} (A_s|X_s,I)$ for arbitrary values $X_s$ of the covariate and independent uniform draws $A_s$, $s=1,\ldots,S$. Assuming that the $X_s/\| X_s \|$ are drawn from the uniform distribution on the unit sphere suggests to estimate the density $f_{\Gamma} (\gamma|I)$ of $\Gamma_{i}$ given $I$ using in the second step the Hoderlein et al (2010) kernel estimator
where $h>0$ is a bandwidth parameter and $0<r\leq \infty$.
The stability of private-value quantile-regression specification allows to use the bid one to test many hypothesis of interest, see Liu and Luo (2017) for a related point of view. This can be useful to obtain better performing tests as the presence of the derivative $\widehat{B}^{(1)} (\alpha|x,I)$ in the implementable private-value expression ((ref)) makes its use for testing harder. Examples of tests based on this idea are as follows.
\paragraph{Quantile-regression goodness of fit.} There is a recent literature that considers the null hypothesis of correct specification of a quantile-regression model over an subinterval $\mathcal{A}$ of $(0,1)$. See Escanciano and Velasco (2010), Rothe and Wied (2013), Escanciano and Goh (2014) and the references therein. These three papers propose test statistics of the form $\widehat{T} (\widehat{\beta}(\cdot|I))$, where $\widehat{\beta}(\cdot|I)$ is a quantile-regression estimator which converges to the true slope over $\mathcal{A}$ with a parametric rate, such as the standard quantile-regression estimator or the augmented ones proposed in Section (ref). See the Application Section 7 for the $\widehat{T}(\cdot)$ used by Rothe and Wied (2013). Liu and Luo (2017) based an entry exogeneity test on the integral of the squared difference of two quantile estimators, see ((ref)) below.
\paragraph{Homogenized bid and elliptical random coefficient.} The correct specification of ((ref)) or ((ref)) can be tested using Rothe and Wied (2013) without any restriction on the quantile alternative. If the alternative is restricted to a quantile-regression model, Koenker and Xiao (2002) or Escanciano and Goh (2014) can be used to test the homogenized-bid null hypothesis, as this specification coincides with the location-shift model considered by these authors. Following Liu and Luo (2017) suggests to consider, for the same null, a test statistic
where $L$ is the number of auctions in the sample, $X_{\ell}$ the auction covariate, $\widehat{\beta} (\cdot)$ a quantile-regression slope estimator $\sqrt{L}$-consistent over $[0,1]$ as the one proposed in the next section, and for instance $\widehat{\beta}_{H_0} (\cdot)=[\widehat{\beta}_0 (\cdot),\widehat{\beta}_{1,OLS},\ldots,\widehat{\beta}_{1D,OLS}]^{\prime}$. Confidence bands can also be used, see Gimenes (2017) and the theory developed in Fan, Guerre and Lazarova (2020).
\paragraph{Exogenous auction format.} Let $V_{j} (\alpha|x,I)=x^{\prime} \gamma_j (\alpha|I)$ be the private-value quantile function conditionally on participation to an ascending auction ($j=asc$) or a first-price one ($j=fp$). A null hypothesis of interest is exogeneity of the auction format, $H_0^{F}: V_{fp} (\cdot|\cdot,I) = V_{asc} (\cdot|\cdot,I)$. Gimenes (2017) gives a consistent quantile-regression estimator $\widehat{\gamma}_{asc} (\cdot|I)$ of $\gamma_{asc} (\cdot|I)$ using ascending auction data. It then follows by ((ref)) that $\widehat{\beta}_{H_0} (\alpha|I) = (I-1) \alpha^{-(I-1)} \int_0^{\alpha} a^{I-2} \widehat{\gamma}_{asc} (a|I) da$ is consistent under the null but not the alternative.\footnote{As the standard quantile-regression estimator may not be well-defined for quantile levels near $0$, it may be more suitable to use an augmented quantile-regression estimator as in Section (ref) to implement Gimenes (2017).} Then using first-price auction data to compute a test statistic $ \widehat{T} (\widehat{\beta}_{H_0} (\cdot|I))$ from Escanciano and Goh (2014) or Rothe and Wied (2013) for an arbitrary alternative, or using Liu and Luo (2017) statistic ((ref)) with a quantile-regression alternative, allow to test for auction format exogeneity.
\paragraph{Participation exogeneity.} The participation exogeneity null hypothesis states that the private values are independent of the number of bidder conditionally on the covariate $ H_0^{E}: V(\cdot|\cdot,I) = V(\cdot|\cdot) $ for all $I$, see also Gimenes (2017) for the ascending auction case. Liu and Luo (2017) use an integral version of $H_0^{E}$ to eliminate the bid quantile derivative in ((ref)). In a quantile-regression setup, Proposition (ref) implies under $H_0^{E}$,
Then tests for entry exogeneity can be obtained using the same construction than for the auction format exogeneity null, using a sample of first-price auction with $I_1$ bidders to estimate $\beta_{I_1} (\alpha|I_2)$ and another sample with $I_2$ bidders to compute a test statistic.
Under participation exogeneity, private value estimates can be averaged over $I$ to improve accuracy. Another important motivation for exogenous participation is risk-aversion estimation, see Guerre, et al. (2009). This approach can be modified to cope with an additional risk-aversion parameter which can be estimated with a parametric rate as shown in Section (ref).
\setcounter{equation}{0}
Many auction parameters of interest can be written using the private-value quantile function or, by ((ref)), the bid quantile function and its quantile derivative. We focus here on the conditional and unconditional integral functionals
where $\mathcal{F}\left( \alpha,x,b_{0I},b_{1I};I\in\mathcal{I}\right) $ is a real valued continuous function. Three illustrative examples are as follows.
\paragraph{Example 1: CRRA parameter.}
For symmetric risk-averse bidders with a concave utility function, the best-response condition ((ref)) becomes \[ \left. \frac{\partial}{\partial a}\left\{ U\left( V\left( \alpha |X,I\right) -B\left( a|X,I\right) \right) a^{I-1}\right\} \right\vert _{a=\alpha}=0. \] Rearranging as in Guerre et al. (2009) yields that $V\left( \alpha|X,I\right) =B\left( \alpha|X,I\right) +\lambda^{-1}\left( \frac{\alpha B^{\left( 1\right) }\left( \alpha|X,I\right) }{I-1}\right) $ where $\lambda\left( \cdot\right) =U\left( \cdot\right) /U^{\prime}\left( \cdot\right) $. For risk-averse bidders with a CRRA utility function $U\left( t\right) =t^{\nu}$, arguing as for Proposition (ref) shows
These two formulas show that the stability implications of Proposition (ref) for linear private-value and bid quantile functions are preserved under CRRA. Assuming as in Guerre et al. (2009) that the number of bidders is exogenous, i.e $V\left( \alpha|X,I\right) =V\left( \alpha|X\right) $ for all $I$, gives that the risk-aversion $\nu$ satisfies, for any pair $I_{0}\neq I_{1}$
which gives identification of $\nu$. Following Lu and Perrigne (2008), the risk-aversion parameter $\nu$ can also be identified combining ascending and first-price auctions data. As seen from Gimenes (2017), the private-value quantile function $V_{asc}\left( \alpha|X,I\right) $ can be easily estimated from ascending auctions. Equating $V_{asc}\left( \alpha|X,I\right) $ to $V\left( \alpha|X,I\right) $ in ((ref)) gives that $\nu$ satisfies
\paragraph{Example 2: Expected revenue.}
Suppose that the seller decides to reject bids lower than a reserve price $R$ and let $\alpha_{R}=\alpha_{R}\left( X,I\right) $ be the associated screening level, i.e. $\alpha_{R}=F\left( R|X,I\right) $. For CRRA bidders, the first-price auction seller expected revenue is\footnote{It is assumed for the sake of brevity that the seller value for the good is $0$. The expected revenue formula for the general case follows from Gimenes (2017).}
This expression includes an integral item \[ \theta\left( X;\alpha_{R}\right) =\int_{\alpha_{R}}^{1}a^{\frac{I-1}{\nu }-1}\left( 1-a^{\left( I-1\right) \frac{\nu-1}{\nu}+1}\right) V\left( a|X,I\right) da \] which can be estimated by plugging in a risk-aversion estimator $\widehat{\nu}$ and an estimator $\widehat{V}\left( \alpha|X,I\right) $ of the private-value quantile function, or estimators of the bid quantile function and its derivative by ((ref)).\footnote{Under risk-neutrality, integrating by parts gives that \[ \int_{\alpha_{R}}^{1}B^{\left( 1\right) }\left( \alpha|X,I\right) \alpha^{I-1}\left( 1-\alpha\right) d\alpha=B\left( \left. \alpha _{R}\right\vert X,I\right) \alpha_{R}^{I-1}\left( 1-\alpha_{R}\right) -\int_{\alpha_{R}}^{1}B\left( \alpha|X,I\right) \alpha^{I-1}\left( I-1-I\alpha\right) d\alpha, \] estimation of $\theta\left( X;\alpha_{R}\right) $ can also be done using only a bid quantile estimator.}
\paragraph{Example 3: Private-value distribution} Additional examples of conditional parameter $\theta (\cdot)$ are the private-value conditional cdf and pdf. Note first that ((ref)) shows that the conditional private-value cdf is an integral functional of the private-value quantile function
Dette and Volgushev (2008) have considered a smoothed version $\mathbb{I} _{\eta}\left( \cdot\right) $ of the indicator function \[ F_{\eta}\left( v|X,I\right) =\int_{0}^{1}\mathbb{I}_{\eta}\left[ v-V\left( \alpha|X,I\right) \right] d\alpha \] where $\mathbb{I}_{\eta}\left( t\right) =\int_{-\infty}^{t/\eta}k\left( u\right) du$, $k\left( \cdot\right) $ being a kernel function and $\eta$ a bandwidth parameter. Differentiating $F_{\eta}\left( v|X,I\right) $ gives \[ f_{\eta}\left( v|X,I\right) =\frac{1}{\eta}\int_{0}^{1}k\left( \frac{v-V\left( \alpha|X,I\right) }{\eta}\right) d\alpha \] which converges to the private-value pdf when $\eta$ goes to $0$. Note that $F_{\eta}\left( v|X,I\right) $ and $f_{\eta}\left( v|X,I\right) $ can be estimated by plugging in an estimator $\widehat{V}\left( \alpha|X,I\right) $ of $V\left( \alpha|X,I\right) $. The resulting cdf and pdf estimators inherit of the dimension reduction property of $\widehat{V}\left( \alpha|X,I\right) $. As the private-value estimator proposed in the next section is consistent over the whole $\left[ 0,1\right] $, no boundary trimming is needed. This contrasts with the GPV pdf estimator. As noted by Escanciano and Guo (2019) in a general context, the integral in $f_{\eta}\left( v|X,I\right)$ can be replaced by a sample average over iid uniform draws $A_s$, as used for the density estimator ((ref)).
\setcounter{equation}{0}
Proposition (ref) suggests to base the estimation of the private-value quantile function on estimations of $B\left( \alpha|x,I\right) $ and of its derivative $B^{\left( 1\right) }\left( \alpha|x,I\right) $ with respect to $\alpha$. The augmented methodology applies local polynomial expansion with respect to $\alpha$ for joint estimation of $B\left( \alpha|x,I\right) $ and $B^{\left( 1\right) }\left( \alpha|x,I\right) $. To ensure comparability with the auction literature which considers private-value pdf having $s$ continuous derivatives, we assume that the private-value quantile function $V\left( \alpha|x,I\right) $ has $s+1$ continuous derivatives with respect to $\alpha$. As seen from ((ref)), this implies that the bid quantile function $B\left( \alpha|x,I\right) $ has $s+2$ continuous derivatives with respect to $\alpha>0$. Let $\left(X_{\ell}, I_{\ell}, B_{1\ell},\ldots, B_{I_{\ell}\ell}\right)$, $\ell=1,\ldots,L$, be an iid first-price auction sample with $I_{\ell}$ bids $B_{i \ell}$ and good characteristics $X_{\ell}$.
\paragraph{Estimation.} Assume first that $V\left( \alpha|X,I\right) =V\left( \alpha|I\right) $ so that $B\left( \alpha|X,I\right) =B\left( \alpha|I\right) $. Let $\rho_{\alpha }\left( \cdot\right) $ be the check function \[ \rho_{\alpha}\left( q\right) =q\left( \alpha-\mathbb{I}\left( q\leq0\right) \right) . \] It is well known that \[ B\left( \alpha|I\right) =\arg\min_{q}\mathbb{E}\left[ \mathbb{I}\left( I_{\ell}=I\right) \rho_{\alpha}\left( B_{i\ell}-q\right) \right] ,\quad\alpha\in\left( 0,1\right) \text{.} \] We now exhibit a functional objective function which achieves its minimum at the restriction of $B(\cdot|I)$ over $\left[ \alpha-h,\alpha+h\right] \cap\left[ 0,1\right] $. It easily follows that, for a non negative kernel function $K\left( \cdot\right)$ with support $[-1,1]$ and a positive bandwidth $h=h_{L}$,
where the minimization is performed over the set of functions $q\left( \cdot \right) $ over $\left[ \alpha-h,\alpha+h\right] \cap\left[ 0,1\right] $. This can be used to estimate the derivative $B^{\left( 1\right) }\left( \alpha|I\right) $, using minimization over Taylor polynomial of order $s+1$ instead of $q(\cdot)$. A Taylor expansion of order $s+1$ gives
The $(s+2) \times 1$ vector $b(\cdot|I)$ stacks the successive bid quantile derivatives, and is the parameter to be estimated. Let $b=\left[ \beta_{0},\ldots,\beta_{s+1}\right] ^{\prime} \in \mathbb{R}^{s+2}$ be the generic coefficients of such a Taylor polynomial function. The sample version of the objective function ((ref)), restricted to local polynomial functions $\pi (\cdot)^{\prime} b$ instead of $q(\cdot)$, is
The augmented quantile estimator is $\widehat{b}\left( \alpha|I\right) =\arg\min_{b\in \mathbb{R}^{s+2}}\widehat{\mathcal{R}}\left( b;\alpha,I\right) $, $\widehat{\beta}_{0}\left( \alpha|I\right) $ and $\widehat{\beta} _{1}\left( \alpha|I\right) $ being estimators of $B\left( \alpha|I\right) $ and its first derivative $B^{\left( 1\right) }\left( \alpha|I\right) $, respectively.
\paragraph{Homogenized bid and elliptical random coefficients.} A two-step version of the augmented method presented above can be used to estimate the homogenized private-value quantile function $v(\cdot|I)$ from ((ref)). Regressing $B_{i \ell}$ on $X_{\ell}$ and an intercept for those auctions with $I_{\ell}=I$ gives a consistent estimator $\widehat{\gamma}_1 (I)$ of $\gamma_1 (I)$. Let $\widehat{B}_{i \ell}= B_{i \ell} - X_{\ell}^{\prime} \widehat{\gamma}_1$ be the estimated homogenized bids. Then replacing $B_{i \ell}$ with $\widehat{B}_{i \ell}$ in the objective function $\widehat{\mathcal{R}}\left( b;\alpha,I\right)$ gives estimators $\check{\beta} (\cdot|I)$ and $\check{\beta}_1 (\cdot|I)$ of the homogenized-bid quantile function and of its first derivative. The resulting estimator of the private-value quantile function is then \[ \widehat{V} (\alpha|X,I) = X^{\prime} \widehat{\gamma}_{1} (I) + \check{\beta} (\alpha|I) + \frac{\alpha \check{\beta}_1 (\alpha|I)}{I-1}. \] The elliptical random-coefficient quantile specification ((ref)) can be estimated similarly. Studying the asymptotic properties of this two-step procedures is outside the scope of this paper. Bhattacharya (2019) considers a related two-step procedure that can be useful for ascending auctions, where estimating quantile derivative is not needed.
An extension of this procedure is the augmented quantile-regression estimator, AQR hereafter, which assumes $ V\left( \alpha |x,I\right) =x_{1}^{\prime} \gamma\left( \alpha|I\right)$, recalling $x_1 = [1,x^{\prime}]^{\prime}$. Proposition (ref)-(i) then gives $B\left( \alpha|x,I\right) = x_{1}^{\prime} \beta\left( \alpha|I\right)$. Define now
so that the Taylor expansion of $B\left( \alpha|X,I\right) $ writes \[ B\left( \alpha+ht|x,I\right) = \sum_{k=0}^{s+1} x_1^{\prime} \beta^{(k)}(\alpha|I) \frac{(ht)^k}{k!} +O\left( h^{s+2}\right) = P\left( x,ht\right) ^{\prime}b\left( \alpha|I\right) +O\left( h^{s+2}\right). \] The corresponding generic parameter is the $(s+2)(D+1) \times 1$ column vector $ b= \left[ \beta_{0}^{\prime} , \beta_{1}^{\prime}, \ldots, \beta_{s+1}^{\prime} \right]^{\prime} $ where the $\beta_{j}$ are all of dimension $D+1$, and the objective function becomes
which accounts for the covariate $X_{\ell}$. The estimation of $b\left( \alpha|I\right) $ is \[ \widehat{b}\left( \alpha|I\right) = \arg \min_{b \in \mathbb{R}^{(s+2)(D+1)}} \widehat{\mathcal{R}}\left( b;\alpha,I\right) \] and the private-value quantile-regression estimator is \[ \widehat{V}\left( \alpha|x,I\right) = x_{1}^{\prime} \widehat{\gamma}\left( \alpha|I\right) \text{ with }\widehat{\gamma}\left( \alpha|I\right) =\widehat{\beta}_{0}\left( \alpha|I\right) +\frac {\alpha\widehat{\beta}_{1}\left( \alpha|I\right) }{I-1}. \] The bid quantile function and its derivatives can be estimated using $\widehat{B}\left( \alpha|x,I\right) =x_{1}^{\prime}\widehat{\beta}_{0}\left( \alpha|I\right) $ and $\widehat{B}^{\left( 1\right) }\left( \alpha|x,I\right) =x_1^{\prime}\widehat{\beta}_{1}\left( \alpha|I\right) $, so that $ \widehat{V} (\alpha|x,I) = \widehat{B} (\alpha|x,I) + \frac{\alpha \widehat{B}^{(1)} (\alpha|x,I)}{I-1} $. The rearrangement method of Chernozhukov, Fern\'{a}ndez-Val and Gallichon (2010) can be used to obtain increasing quantile estimators.
\paragraph{AQR estimator properties.} Bassett and Koenker (1982) report that standard quantile-regression estimators are not defined near the extreme quantile levels $\alpha=0$ or $\alpha=1$, mostly because the associated objective function has some flat parts. The AQR is better behaved because the objective function $\widehat{\mathcal{R}}\left( b;\alpha,I\right) $ averages the check function $\rho_{a}\left( \cdot\right) $ for quantile levels $a$ in $\left[ \alpha-h,\alpha+h\right] \cap\left[ 0,1\right] $, ensuring that the AQR objective function is not flat for extreme quantile levels, as illustrated in Figure (ref).\footnote{This averaging effect requests that $t\mapsto P\left( X_{\ell},ht\right) ^{\prime}b$ is not constant meaning that the derivative components of $b$ should not vanish.}
Therefore the AQR estimator is easier to define for the extreme quantile levels $\alpha=0$ and $\alpha=1$ than the standard quantile-regression estimator. This is especially relevant for estimating auction models as the winner is expected to belong to the upper tail as soon as the number of bidders is large enough. It also follows from the theoretical study of the objective function $\widehat{\mathcal{R}}\left( \cdot ;\cdot,I\right) $ that the AQR estimator is uniquely defined for all quantile levels with a probability tending to 1. The bid AQR estimator is also smoother than the standard quantile-regression one, see Figure (ref) in the Application Section and (ref) for a formal argument.
\setcounter{equation}{0}
Some additional notations are as follows. Let $S_{0} = [1,0,\ldots,0]$ and $S_{1}=\left[ 0,1,0,\ldots,0\right]$ be $1\times\left( s+2\right) $ selection vectors such that $S_0 \pi(t)=1$, $S_1 \pi(t)=t$. Let $\operatorname{Id}_{D+1}$ be the $(D+1)\times (D+1)$ identity matrix and set $\mathsf{S}_{j}=S_{j} \otimes\operatorname*{Id}_{D+1}$, $j=0,1$, so that $\mathsf{S}_{0} \widehat{b}\left( \alpha|I\right) =\widehat{\beta}_{0}\left( \alpha|I\right) $ and $\mathsf{S}_{1} \widehat{b}\left( \alpha|I\right) =\widehat{\beta}_{1}\left( \alpha|I\right) $ are respectively estimators of $\beta(\alpha)$ and its first derivative $\beta^{(1)} (\alpha)$. $\mathrm{Tr}(\cdot)$ is the trace of a square matrix and $\partial_u^n$ stands for $\frac{\partial^n}{ \partial u^n}$. For two sequences $\{a_L\}$ and $\{b_L\}$, $a_{L}\asymp b_{L}$ means that both $a_{L}/b_{L}=O\left( 1\right) $ and $b_{L}/a_{L}=O\left( 1\right) $. The norm $\left\Vert \cdot\right\Vert $ is the Euclidean one, i.e. $\left\Vert e\right\Vert =\left( e^{\prime}e\right) ^{1/2}$. For a matrix $A$, $\| A\| = \sup_{b:\|b\|=1} \| Ab \|$. Convergence in distribution is denoted as `$\stackrel{d}{\rightarrow}$'.
\setcounter{hp}{0}
\setcounter{hp}{18}
\setcounter{hp}{7}
\setcounter{hp}{5}
Assumption (ref)-(i) is standard. Assumption (ref)-(ii) allows for private values depending on the number of bidders. Recall that
where $f\left( v|x,I\right) $ is the conditional private-value pdf. Hence Assumption (ref)-(ii) amounts to assume that $f\left( v|x,I\right) $ is bounded away from $0$ and infinity on its support $\left[ V\left( 0|x,I\right) ,V\left( 1|x,I\right) \right] $ as assumed for instance in Riley and Samuelson (1981), Maskin and Riley (1984) or GPV. The condition $0<f\left( v|x,I\right) <\infty$ is also used for asymptotic normality of quantile-regression estimator, see Koenker (2005). Assumption (ref) is a standard smoothness condition which, by ((ref)), parallels GPV who assume that the pdf $f(v|x,I)$ is $s$-times differentiable.
The bandwidth rate in Assumption (ref) is unusual in kernel or local polynomial nonparametric estimation, where rate conditions as $1/(Lh)=o(1)$ are more common. This is due to a key linearization expansion for $\widehat{V} (\alpha|x,I)$, which holds with an $O_{\mathbb{P}} \left(\log L/(L \sqrt{h}) \right)$ error term that must go to $0$, see ((ref)) and ((ref)) in Theorem (ref) below.\footnote{See also Theorem (ref) in (ref), where it is shown more specifically that this bandwidth order is needed for the linearization of $\widehat{B}^{(1)} (\alpha|x,I)$.}
Assumption (ref) holds for most of the examples of functionals above. A notable exception is the cdf $F\left( v|x,I\right) $ in Example 3, which involves an indicator function which is not smooth. However it holds for the smoothed approximation $F_{\eta}\left( v|x,I\right) $ of the cdf, although Assumption (ref) implicitly rules out vanishing bandwidth $\eta$ in Example 3.
The next sections give our theoretical results for integrated mean squared error, uniform consistency and asymptotic distribution of the augmented estimator $\widehat{V} \left( \cdot|\cdot,I\right) $. These results are derived using a pseudo-true value framework, in which $\widehat{b} (\cdot|I)$ is viewed as an estimator of the minimizer $\overline{b} (\cdot|I)$ of the population counterpart of $\widehat{\mathcal{R}}\left( b;\alpha,I\right)$ \[ \overline{b}\left( \alpha|I\right) = \arg \min_{b \in \mathbb{R}^{(s+2)(D+1)}} \overline{\mathcal{R}}\left( b;\alpha,I\right) \text{ where } \overline{\mathcal{R}}\left( b;\alpha,I\right) = \mathbb{E} \left[\widehat{\mathcal{R}}\left( b;\alpha,I\right) \right] \] which asymptotic existence and uniqueness is established for the proofs of our main results. Define accordingly $\overline{\beta}_{0}\left( \alpha|I\right)=\mathsf{S}_{0}\overline{b}\left( \alpha|I\right) $ and $\overline{\beta}_{1}\left( \alpha|I\right) = \mathsf{S}_{1} \overline{b}\left( \alpha|I\right)$ and
The difference $\overline{V} (\alpha|x,I)- V(\alpha|x,I)$ can be interpreted as a bias term.
Because $\widehat{V} (\cdot|\cdot,I)$ is defined in an implicit way via the minimization of the objective function ((ref)), its asymptotic study relies on a linearization of $\widehat{b} (\alpha|I)-\overline{b} (\alpha|I)$ which, in a quantile setup, is called a Bahadur expansion, see Theorem (ref) in (ref) and Koenker (2005, Chap. 4). It is shown that, in a vicinity of $\overline{b} (\alpha|I)$, $b \mapsto \widehat{\mathcal{R}}\left( b;\alpha,I\right)$ is twice differentiable with a first derivative $\widehat{\mathcal{R}}^{(1)} \left( b;\alpha,I\right)$ satisfying $\mathbb{E} \left[\widehat{\mathcal{R}}^{(1)} \left( \overline{b} (\alpha|I);\alpha,I\right)\right] =0$, and with a Hessian matrix $\overline{\mathcal{R}}^{(2)}\left( \overline{b} (\alpha|I);\alpha,I\right)$ which is invertible. The leading term of $\widehat{b}(\alpha|I)-\overline{b}(\alpha|I)$ is $ - \left[ \overline{\mathcal{R}}^{(2)}\left( \overline{b}(\alpha|I);\alpha,I\right)\right ]^{-1} \widehat{\mathcal{R}}^{(1)} \left( \overline{b}(\alpha|I);\alpha,I\right) $ as shown in Theorem (ref), so that the leading term of $\widehat{V}(\alpha|x,I)$ is
see ((ref)) below. Because direct computations of the moments of $\widehat{V}(\alpha|x,I)$ are difficult due to its implicit definition and to potential nonlinearities, moments of its linear leading term $\widetilde{V}(\alpha|x,I)$ are used as an approximation.
Let us first introduce some notations for the integrated mean squared error (IMSE). Let $\Pi^{1}\left( \alpha\right) $ be the second column of the inverse of $\int\pi\left( t\right) \pi\left( t\right) ^{\prime}K\left( t\right) dt$, i.e., \[ \Pi^{1}\left( \alpha\right) =\left( \int\pi\left( t\right) \pi\left( t\right) ^{\prime}K\left( t\right) dt\right) ^{-1}S_{1}^{\prime} \] and consider the variance terms
where $\mathbb{E}^{-1} \left[ \frac{X_{1\ell} X_{1\ell}^{\prime} \mathbb{I}\left( I_{\ell}=I\right) }{B^{\left( 1\right) }\left( \alpha|X_{\ell},I_{\ell}\right) } \right] $ is the inverse matrix of $ \mathbb{E} \left[ \frac{X_{1\ell} X_{1\ell}^{\prime} \mathbb{I}\left( I_{\ell}=I\right) }{B^{\left( 1\right) }\left( \alpha|X_{\ell},I_{\ell}\right) } \right] $. That $v^{2}\left( \alpha\right) $, and then $\Sigma_{I}$, is strictly positive follows from the proof of Theorem (ref) below, see in particular Lemma (ref) in (ref). The bias, and integrated squared bias, of the estimator are asymptotically proportional to, respectively\footnote{The expression above depends upon the derivative $\beta^{(s+2)}(\alpha|I)$, which exists by ((ref)) for all $\alpha \neq 0$ since $\gamma(\cdot|I)$ is $(s+1)$-th differentiable. Proposition (ref)-(iii) in (ref) shows that $\alpha \beta^{(s+2)}(\alpha|I)$ can be defined over the whole quantile interval $[0,1]$ since $\lim_{\alpha\downarrow 0} \alpha \beta^{(s+2)}(\alpha|I) =0$. }
The next Theorem deals with the IMSE of $\widehat{V} (\cdot|\cdot,I)$ and with its difference to its linearization $\widetilde{V} (\cdot|\cdot,I)$ in ((ref)).
Theorem (ref) gives the IMSE of the linearization $\widetilde{V} (\cdot|\cdot,I)$ of $\widehat{V} (\cdot|\cdot,I)$ in ((ref)). Then ((ref)) gives the order of the linearization error in a uniform sense, which is negligible with the order $1/\sqrt{Lh} + O(h^{s+1})$ of the squared root IMSE under Assumption (ref). The linearization result ((ref)) requests $\log L / (h \sqrt{LI}) = o(1)$, which is the main motivation for the unusual rate of bandwidth rate of Assumption (ref). This condition is driven by the linearization of the bid quantile derivative estimator $\widehat{B}^{(1)} (\cdot|\cdot,I)$.
The bias and variance leading terms in the IMSE expansion ((ref)) are from the bid quantile derivative estimator $\alpha \widehat{B}^{(1)} (\alpha|x,I)/(I-1)$. As \[ B^{\left( 1\right) }\left( \alpha|x,I\right) = \frac{1}{g\left[ B\left(\alpha|x,I\right) |x,I\right] }, \] where $g\left( \cdot|\cdot\right) $ is the bid conditional pdf, estimation of this item is similar to estimating a pdf. The rate $1/Lh$ of the variance term $\Sigma_I/(LIh)$ is the rate of a kernel density estimator in the absence of covariate. This is due to the quantile-regression specification. Compared to GPV density estimation rate $1/\sqrt{Lh^{D+1}}$, the rate of the AQR estimator does not suffer from the curse of dimensionality. The order $O(h^{s+1})$ of the bias is given by ((ref)), implying that $B^{(1)} (\alpha|x,I)$ has as many derivatives as $V(\alpha|x,I)$, hence the exponent $s+1$.
Minimizing the leading term of the IMSE expansion ((ref)) yields the optimal bandwidth
As in kernel estimation, a pilot bandwidth can be computed using a simple private-value quantile-regression model to proxy $\Sigma_{I}$ and $\mathsf{Bias}_{I}^{2}$ in a parametric way. The corresponding square root IMSE rate is $ L^{\frac{s+1}{2s+3}} $ which corresponds to the optimal minimax rate given in GPV in the absence of covariate, up to a logarithmic term and an exponent $s+1$ due to estimation of the private-value quantile function, instead of $s$ appearing for pdf. In particular, it is $L^{-2/5}$ for $s=1$, with an exponent $2/5=.4$ close to $1/2$ suggesting potential good performances in small samples even in the presence of covariate.
A similar rate can also be derived for the uniform consistency of $\widehat{V}(\cdot|\cdot,I)$ stated in ((ref)). Note also that ((ref)) shows that the bid quantile estimator $\widehat{B} (\cdot|\cdot,I)$ converges uniformly to $B(\cdot|\cdot,I)$ with a rate which is nearly parametric.\footnote{The uniform consistency rate in ((ref)) includes a bias term $o(h^{s+1})$, which is $O(h^{s+2})$ for $\alpha \neq 0$ because $B(\cdot|x,I)=x_{1}^{\prime} \beta (\cdot|I)$ is $(s+2)$-th times continuously differentiable over $(0,1]$. Because $B(\cdot|x,I)$ may have only $(s+1)$ derivatives at $\alpha=0$, the bias order $h^{s+2}$ may not hold uniformly over $[0,1]$. } All these convergence results take place over the whole $[0,1] \times \mathcal{X}$, meaning that potential boundary biases disappear asymptotically.
While Theorem (ref) reviews the global performance of the AQR estimator, this section details some of its local features, and in particular its upper boundary behavior. Define \[ \Pi_{h}^{1}\left( \alpha\right) =\left( \int_{-\frac{\alpha}{h}} ^{\frac{1-\alpha}{h}}\pi\left( t\right) \pi\left( t\right) ^{\prime }K\left( t\right) dt\right) ^{-1} S_{1}^{\prime}, \] \[ v_{h}^{2}\left( \alpha\right) =\Pi_{h}^{1}\left( \alpha\right) ^{\prime }\int_{-\frac{\alpha}{h}}^{\frac{1-\alpha}{h}}\int_{-\frac{\alpha}{h}} ^{\frac{1-\alpha}{h}}\pi\left( t_{1}\right) \pi\left( t_{2}\right) ^{\prime}\min\left( t_{1},t_{2}\right) K\left( t_{1}\right) K\left( t_{2}\right) dt_{1}dt_{2}\Pi_{h}^{1}\left( \alpha\right) , \]
setting $\mathsf{Bias}_{h}(0|x,I)=0$, see Footnote (ref). The next Theorem gives some variance and bias expansions and the pointwise asymptotic distribution of the estimator. Recall that, in our pseudo true value framework, the bias of $\widehat{V} (\alpha|x,I)$ is given by $\overline{V} (\alpha|x,I)-V(\alpha|x,I)$. As the variance of $\widehat{V} (\alpha|x,I)$ is difficult to compute due to the implicit definition of this estimator, an expansion for the variance of its leading term $\widetilde{V} (\alpha|x,I)$ is given instead.
These expansions give a better understanding of potential boundary effects affecting $\widehat{V} (\alpha|x,I)$. For $\mathsf{Bias} (\alpha|x,I)$ and $\Sigma (\alpha|I)$ defined before Theorem (ref), it holds \[ \mathsf{Bias}_{h}(\alpha|x,I)=\mathsf{Bias}(\alpha|x,I)\text{ and }\Sigma _{h}\left( \alpha|I\right) =\Sigma\left( \alpha|I\right) \text{ for all }\alpha\text{ in }\left[ h,1-h\right] \] since the support of the kernel $K(\cdot)$ is $[-1,1]$. Hence a pointwise optimal bandwidth for central quantile levels is $h_{\ast} (\alpha|x,I)=\left( \frac{ \Sigma (\alpha|I)}{2\left( s+1\right) \mathsf{Bias}^{2} (\alpha|x,I)}\frac{1}{LI}\right) ^{\frac{1} {2s+3}} $, which is obtained by minimizing the leading term of the Mean Squared Error obtained from ((ref)) and ((ref)).
As the asymptotic bias and standard deviation are proportional to $\alpha$, more bias and variance are expected for higher quantile levels and smaller $I$. This follows from the fact that the leading term of $\widehat{V} (\alpha|x,I)$ is $\alpha \widehat{B}^{(1)} (\alpha|x,I) / (I-1)$, which is proportional to $\alpha/(I-1)$. Boundary effects can only occur for quantile levels in $[0,h]$ or $[1-h,1]$, which differ from this respect. As $\widehat{V} (\alpha|x,I) = \widehat{B} (\alpha|x,I) + \alpha \widehat{B}^{(1)} (\alpha|x,I) / (I-1)$, $\widehat{V} (\alpha|x,I)$ is close to $\widehat{B} (\alpha|x,I)$ when $\alpha$ is in $[0,h]$. In particular, $\widehat{V} (0|x,I)=\widehat{B} (0|x,I)$ which converges to $B(0|X,I)$ with the rate $1/\sqrt{LI}+o(h^{s+1})$ at least.
For upper quantile levels $\alpha$ in $[1-h,1]$, the consistency rate of $\widehat{V} (\alpha|x,I)$ is the slower $1/\sqrt{Lh}+O(h^{s+1})$ as for central quantile levels. Bias and variance also involve the matrix \[ \left( \int_{-\frac{\alpha}{h}} ^{\frac{1-\alpha}{h}}\pi\left( t\right) \pi\left( t\right) ^{\prime }K\left( t\right) dt\right)^{-1} = \left( \int_{-1} ^{\frac{1-\alpha}{h}}\pi\left( t\right) \pi\left( t\right) ^{\prime }K\left( t\right) dt\right) ^{-1} \] for $h$ small enough. This matrix increases with $\alpha$, so that higher $\left|\mathsf{Bias} (\alpha|x,I)\right|$ and $\Sigma_{h} (\alpha|I)$ can take place near $\alpha=1$ compared to central quantile levels. See Fan and Gijbels (1996) and the references therein for similar discussions. Accordingly, our simulations show an increase of the bias and variance for $\alpha$ approaching 1, see Figures (ref) and (ref) in Section (ref).
The plug-in estimators of $\theta\left( x\right) $ and $\theta$ in ((ref)) are \[ \widehat{\theta}\left( x\right) =\int_{0}^{1}\mathcal{F}\left[ \alpha,x,\widehat{B}\left( \alpha|x,I\right) ,\widehat{B}^{\left( 1\right) }\left( \alpha|x,I\right) ;I\in\mathcal{I}\right] d\alpha,\quad \widehat{\theta}=\int_{\mathcal{X}}\widehat{\theta}\left( x\right) dx, \] with AQR estimators $\widehat{B}\left( \alpha|x,I\right) $ and $\widehat{B} ^{\left( 1\right) }\left( \alpha|x,I\right) $. Let us now introduce the asymptotic variances of $\widehat{\theta}\left( x\right) $ and $\widehat{\theta}$. The variances depend upon the matrices
and of the functions, recalling $b_{0I}$ and $b_{1I}$ stand for $B\left( \alpha|x,I\right) $ and $B^{\left( 1\right) }\left( \alpha|x,I\right) $ respectively,
Let $A$ be a $\mathcal{U}_{[0,1]}$ random variable, $\mathbf{1} = [1,\ldots,1]^{\prime}$ a $(D+1) \times 1$ vector, and define, recalling $x_1= [1,x^{\prime}]^{\prime}$,
The proof of Theorem (ref) in (ref) shows that the asymptotic variances of $\widehat{\theta}\left( x\right) $ and $\widehat{\theta}$ are $\sigma_{L}^{2}\left( x\right) / L $ and $\sigma_{L}^{2}/L$ respectively provided they are bounded away from $0$. This holds if
It may be indeed that $\sigma_{L}^{2}\left( x|I\right) =0$ and $\sigma_{L}^{2}=0$, in which case $\widehat{\theta}\left( x\right) $ and $\widehat{\theta}$ can converge to $\theta\left( x\right) $ and $\theta$ with \textquotedblleft superefficient\textquotedblright\ rates, that is faster than $1/L^{1/2}$. Why it is possible is better understood in our quantile context, through an example of functionals for which ((ref)) does not hold. Consider, for some given $I_{0}$ of $\mathcal{I}$, \[ \mathcal{F}\left[ \alpha,x,B\left( \alpha|x,I\right) ,B^{\left( 1\right) }\left( \alpha|x,I\right) ;I\in\mathcal{I}\right] =2B\left( \alpha|x,I_{0}\right) B^{\left( 1\right) }\left( \alpha|x,I_{0}\right) \] which gives $\left( \varphi_{0I_{0}}\left( \alpha|x\right) ,\varphi _{1I_{0}}\left( \alpha|x\right) \right) =2\left( B^{\left( 1\right) }\left( \alpha|x,I_{0}\right) ,B\left( \alpha|x,I_{0}\right) \right) $. Hence ((ref)) does not hold and $\sigma_{L}^{2}\left( x\right) =\sigma_{L}^{2}=0$. Why $\widehat{\theta}\left( x\right) $ and $\widehat{\theta}$ can converge with superefficient rates for these functionals is in fact not surprising observing that they estimate \[ \theta \left( x\right) =B^{2}\left( 1|x,I_{0}\right) -B^{2}\left( 0|x,I_{0}\right) ,\quad\theta=\int_{\mathcal{X}}\theta\left( x\right) dx, \] respectively. For these examples, the parameters of interest only depend upon extreme quantiles, in which case superefficient estimation is possible, see e.g. Hirano and Porter (2003) and the references therein. The next Theorem establishes the asymptotic normality of $\widehat{\theta}\left( x\right) $ and $\widehat{\theta}$.
Note that the conditional functional estimator $\widehat{\theta} (x)$ converges with a parametric rate, under a bandwidth condition slightly stronger than in Assumption (ref). This bandwidth condition corresponds to the fact that the linearization error term $\widehat{B}^{(1)} (\alpha|x,I) -\widetilde{B} (\alpha|x,I)$ is of order $\log L/(Lh^{3/2})$ as guessed from ((ref)) and must be $o(1/\sqrt{L})$.
The bias term order $o\left(h^s\right)$ is given by the estimation of $B^{\left( 1\right) }\left( \alpha|x,I\right) $, and is of order $O(h^{s+1})$ when $\mathcal{F}\left( \cdot\right) $ depends upon $\alpha B^{\left( 1\right) }\left( \alpha|x,I\right) $ as in all the Examples. Let $\mathcal{G}_{b_{1I}}\left( \cdot\right) $ be the partial derivative of $\mathcal{F}\left( \cdot\right) $ with respect to $\alpha B^{\left( 1\right) }\left( \alpha|x,I\right) $, and $\mathsf{Bias}_{h}\left( \alpha|x,I\right) $ be as in ((ref)). Then
and $\mathsf{bias}_{L,\theta}=\int_{\mathcal{X}}\mathsf{bias}_{L,\theta\left( x\right) }dx$. The estimators $\widehat{\theta }\left( x\right) $ or $\widehat{\theta}$ are therefore asymptotically unbiased if $h^{s+1}\sqrt{Lh}=o\left( 1\right) $ or $h^{s+1}\sqrt{L}=o\left( 1\right) $ respectively.
Theorem (ref) applies to our functional Examples, but the resulting variance can be somehow involved, so that the use of the bootstrap can be preferred as discussed below. We first detail the variance obtained for the cdf estimator of Example 3 to illustrate the influence of the bandwidth $\eta$ on its variance.
\paragraph{Example 3 (cont'd).}
For the cdf estimator $\widehat{F}_{\eta}\left( v|x,I\right) =\int_{0}^{1}\mathbb{I}_{\eta}\left[ v-\widehat{V}\left( \alpha|x,I\right) \right] d\alpha$,
When $\eta$ goes to $0$, the dominant part of the variance is, for inner $v$, integrating by parts and setting $V_{x,I}=V\left( A|x,I\right) $
Hence the order of the variance of $\widehat{F}_{\eta}\left( v|x,I\right) $ is $1/\left( L\eta \right) $ when $\eta$ goes to $0$. Its bias has two components: the first is $\mathsf{bias} _{L,F_{\eta}\left( v|x,I\right) }$ due to the bias of $\widehat{V}\left( \alpha|x,I\right) $ and is of order $O\left( h^{s+1}\right) $, while the second is $F_{\eta}\left( v|x,I\right) -F\left( v|x,I\right) =O\left( \eta^{s+1}\right) $ if $k\left( \cdot\right) $ is a kernel of order $s$. Further work is needed to determine at which rate $\eta$ can go to $0$ in this heuristic.
\paragraph{Bootstrap inference.} Earlier theoretical works considering quantile-regression bootstrap inference are Rao and Zhao (1992), for the weighted bootstrap, and Hahn (1995) for the pairwise bootstrap. See also Liu and Luo (2017) for quantile-based auction testing procedures. For standard and sieve quantile-regression estimators, Belloni et al. (2019) establish consistency of several bootstrap procedures for functionals and uniformly with respect to the quantile levels. It is expected that it carries over to our AQR estimators when bias terms can be neglected, but is out of the scope of the present paper.
\setcounter{equation}{0}
The first simulation experiments compare the AQR estimation method with GPV and its homogenized-bid extension. Other simulations illustrate its performances for estimation of the private-value quantile function, expected seller revenue and optimal reserve price, or risk-aversion parameter. All experiments involve $L=100$ auctions with $I=2$ or $I=3$ bidders, so that small samples of $200$ or $300$ are considered. In most of the experiments, three auction-specific auction covariates are considered. The number of replications is $1,000$ in all experiments. AQR are computed over the estimation grid $\alpha =0,0.01,\ldots,0.99,1$. The AQR local polynomial order $s+1$ is set to $2$ and the AQR kernel is the Epanechnikov $K\left( t\right) =\frac{3}{4}\left( 1-t^2\right) \mathbb{I}\left( t\in\left[ -1,1\right] \right) $.
Since the asymptotic bias and variance of $\widehat{V} (\alpha|x,I)$ tend to decrease with $I$, choosing a small number $I$ of bidders is challenging. Hickman and Hubbard (2015) used $5$ bids while $I=3$ or $5$ in Marmer and Shneyerov (2012) and Ma, Marmer and Shneyerov (2019). The number of bids $LI$ ranges from $1,000$ for Hickman and Hubbard (2015) to $4,200$ for Marmer and Shneyerov (2012). In a simulation experiment focused on nonparametric estimation of the utility function of risk-averse bidders, Zincenko (2018) considers $I=2$ with $L=300$ and $I=4$ with $L=150$. These references do not consider covariate, with the exception of Zincenko (2018) for $L=900$ auctions with one or two covariates. Therefore our simulation setting correspond to rather demanding small sample situations.
The simulation experiments of this Section makes use of a “trigonometric” quantile function
whose probability density function has a compact support and is bounded away from $0$, with a shape similar to a peaked Gaussian one as seen from the left panel of Figure (ref).\footnote{The pdf graph is obtained noting that $T^{(1)} (\alpha)=1/f(T(\alpha))$, so that the graph $\alpha \in [0,1] \mapsto (T(\alpha),1/T^{(1)}(\alpha)$ is the one of the pdf $f(\cdot)$. The associated bid quantile function can be obtained using ((ref)), which is used to simulate bids via a quantile transformation of uniform draws, as performed elsewhere in our simulation experiments.}
\paragraph{Comparison with GPV.} This experiment considers private values drawn from $T(\cdot)$ and does not include covariate. It compares first the boundary bias corrected GPV two-step pdf estimator with triweight kernel and rule of thumb bandwidth of Hickman and Hubbard (2015) with the AQR pdf estimator \[ \widehat{f} (v) = \int_{0}^{1} \frac{1}{h_{AQR}} \widetilde{K}_{tri} \left(\frac{v-\widehat{V} (\alpha)}{h_{AQR}} \right) d \alpha, \quad h_{AQR} = \frac{3}{(LI)^{1/5}} \left( \int_{0}^{1} \left( \widehat{V} (\alpha) - \int_{0}^{1} \widehat{V} (t) dt \right)^2 d\alpha \right)^{1/2}, \] where the AQR bandwidth is $h=.3$, $I=2$ and $\widetilde{K}_{tri} (\cdot)$ is the Hickman and Hubbard (2015) boundary bias corrected triweight kernel.
The results are reported in Figure (ref), which illustrates how much harder estimating a pdf can be compared to estimating a quantile function. The variance and the bias of the two private-value pdf estimators look much higher, especially just after the density peak. This peak causes a small AQR bias for central quantiles.
Other features are common to the two pdf estimation procedures. The first stage of the GPV procedure is based upon an estimation of the private values from ((ref)), which is likely to have more bias and variance for small bids than for large ones. Accordingly, the left panel of Figure (ref) reveals that the GPV pdf estimator performs quite well before the density peak, but that both its bias and variance increase for higher private values. The AQR pdf estimator in the center panel looks less affected by these issues. By contrast, the quantile estimation procedure in the right panel of Figure (ref) is only affected by a variance increase for upper quantiles as expected.
\paragraph{Cdf estimation with homogenized bids and AQR} The bids considered in this experiment are associated with private values satisfying \[ V_{i\ell}= X_{1\ell} + X_{2\ell} + X_{3i\ell} + v_{i\ell} \] where the $X_{j\ell}$'s are independently drawn from the uniform, independent from $v_{i\ell}$, whose quantile function is $T(\cdot)$ in ((ref)). The bids are regressed on the covariate and the constant to obtain the homogenized bids $\widehat{b}_{i\ell}$. The latter are used as in the first-stage of Hickman and Hubbard (2015) to obtain homogenized pseudo private values $\widehat{v}_{i\ell}$, and the estimated private-value conditional cdf at $x_1=x_{2}=x_3=1/2$ is \[ \widehat{F} (v|x) = \frac{1}{LI} \sum_{\ell}^{L} \sum_{i=1}^I \mathbb{I} \left( \frac{1}{2}(\widehat{\beta}_1+\widehat{\beta}_{2}+\widehat{\beta}_3) + \widehat{v}_{i \ell} \leq v \right), \] where the $\widehat{\beta}_j$ are the OLS slope estimators computed in the homogenized-bid regression. Two AQR conditional estimators are computed. The first uses the homogenized bids and ((ref)) while the second is based on the standard AQR $\widehat{V}(\alpha|x,I)$, both using the bandwidth $h=.3$. The performances of these three cdf estimators are reported in Figure (ref).
As expected, considering estimation of the private-value cdf gives smaller bias and variance than estimating pdf. All procedures have a similar variability. However the homogenized-bid GPV procedure has a larger bias, dominating its variability in the right centre part of the private-value distribution, than its AQR counterparts. Applying the AQR to the homogenized bids seems to slightly dominate the other AQR procedure.
\paragraph{Quantile-regression model and estimation details.} The private-value quantile function is given by a quantile-regression model with an intercept and three independent covariates with the uniform distribution over $\left[ 0,1\right] $, \[ V\left( \alpha|X\right) =\gamma_{0}\left( \alpha\right) +\gamma_{1}\left( \alpha\right) X_{1}+\gamma_{2}\left( \alpha\right) X_{2}+\gamma_{3}\left( \alpha\right) X_{3} \] with
The coefficient $\gamma_{0}\left( \cdot\right) $ is flat near $0$ and strongly increases near $1$ while $\gamma_{2}\left( \cdot\right) $ strongly increases near $0$ and is flat after. The slope $\gamma_{3}\left( \cdot\right) $ is as the trigonometric quantile function ((ref)), but with stronger oscillations which makes it harder to estimate.
The performances of the private-value quantile estimation procedure are evaluated through the individual estimation of each slope function or estimation of $V(\alpha|x)$ when the $x_j$ are set to their median $1/2$. The curvature of the expected revenue is mostly due to $\gamma_{2} (\cdot)$, the other coefficients having a rather flat contribution. The performances of the expected revenue estimation procedure are therefore evaluated removing the intercept, setting $x_{1}$ and $x_{3}$ to $0$ and taking $x_{2}=0.8$. This choice gives a unique optimal reserve price achieved for $\alpha_{\ast}=.3$, which is not too close to the boundaries so that the expected revenue function has a substantial concave shape which is suppose to make estimation more difficult. This is also used for evaluating estimation of the optimal reserve price $R_{\ast}=.8 \gamma_{2} (\alpha_{\ast})$.
\paragraph{Simulation results.} Table (ref) summarizes the simulation results for the estimation of the private-value quantile function, the expected revenue and the optimal reserve price. The Bias and square Root Integrated Mean Squared Error (RIMSE) lines for $\widehat{V}\left( \cdot|\cdot\right) $ gives the simulation counterparts of, respectively \[ \left( \frac{1}{4}\sum_{j=0}^{3}\int_{0}^{1}\left( \mathbb{E}\left[ \widehat{\gamma}_{j}\left( \alpha\right) \right] -\gamma_{j}\left( \alpha\right) \right) ^{2}d\alpha\right) ^{1/2}\text{and }\left( \frac {1}{4}\sum_{j=0}^{3}\int_{0}^{1}\mathbb{E}\left[ \left( \widehat{\gamma} _{j}\left( \alpha\right) -\gamma_{j}\left( \alpha\right) \right) ^{2}\right] d\alpha\right) ^{1/2}. \] The Bias and RIMSE for the expected revenue are computed similarly. Table (ref) also gives the Bias and square Root Mean Squared Error (RMSE) of the optimal reserve price estimator. All these quantities are computed for bandwidths $.2,.3,\ldots,.9$.
Estimation of the private-value slope coefficients seems more sensitive to the bandwidth parameter than the expected revenue or optimal reserve price. It has also a much higher RIMSE. The bandwidth behavior of $\widehat{V}\left( \alpha|x\right) $ is also illustrated in Figure (ref) , which considers the small bandwidth $h=0.3$ and the larger $h=0.8$. As expected from Theorem (ref), the dispersion of $\widehat{V}\left( \alpha|x\right) $ increases with $\alpha$ and decreases with $h$, while the bias increases with $\alpha$ and $h$. Figure (ref) also suggests that choosing a large bandwidth as recommended by Table (ref) may lead to important bias issues, including underestimating the private-value quantile function for high $\alpha$. This is mostly due to the slope $\gamma_3(\cdot)$ which is an important source of bias.
This contrasts with estimation of the expected revenue and optimal reserve price, which seems mostly unaffected by the bandwidth. This is partly because the expected revenue depends upon $\left( 1-\alpha\right) V\left( \alpha|x\right) $: multiplying the private-value quantile function by $\left( 1-\alpha\right) $ mitigates larger bias and variance near the boundary $\alpha=1$, see also Figure (ref). For the considered experiment, the true expected revenue is always in the $95\%$ band of Figure (ref) while the true private-value quantile function is out for large $\alpha$ when $h=0.8$.
Two risk-aversion estimators are considered. The first estimator $\widehat{\nu}_{fp}$ is based upon ((ref)) and uses two independent samples of size $L=100$ with 2 and 3 bidders from the model above, which corresponds to a CRRA utility function $t^{\nu}$ with $\nu =1$.\footnote{The optimal bid functions can be computed explicitly under the risk-neutral case $\nu=1$. Considering other values of $\nu$ would request to use numerical computations of the bid functions.} Integrals with respect to $\alpha$ are computed using Riemann sums whereas integrals with respect to $x$ are replaced with sample means over the two auction samples. The second estimator $\widehat{\nu}_{asc}$ is based upon ((ref)) and uses an additional sample of size $L=100$ of ascending auctions with two bidders. In this case, it is possible to consider various values of $\nu$ and the simulation experiment considers the values $0.2$, $0.6$ and $1$. Indeed, if $B\left( \alpha|X\right) $ is the first-price auction quantile bid function with $I=2$, the observed bids drawn from $B\left( \alpha |X\right) $ are rationalized by a CRRA utility function $t^{\nu}$ if the private-value quantile function is set to \[ V_{\nu}\left( \alpha|X\right) =B\left( \alpha|X,2\right) +\nu\alpha B^{\left( 1\right) }\left( \alpha|X,2\right) \] provided $V_{\nu}^{\left( 1\right) }\left( \cdot|X\right) >0$ for all $X$. As $V_{\nu }^{\left( 1\right) }\left( \cdot|\cdot\right) >0$ holds in our case, we use $V_{\nu}\left( \alpha|X\right) $ to generate two ascending bids for each auction. Following Gimenes (2017), $V_{\nu}\left( \alpha|X\right) $ can be estimated from winning bids in these ascending auctions using AQR for quantile level $2\alpha-\alpha^{2}$.
Table (ref) shows that $\widehat{\nu}_{asc}$ dominates $\widehat{\nu}_{fp}$ in this experiment. While the RMSE\ and bias of $\widehat{\nu}_{asc}$ do not seem sensitive to $h$, this is not the case for $\widehat{\nu}_{fp}$ which has a high downward bias, and then RMSE, for small $h$. Further investigations suggest this is due to an unbalanced variable issue, the difference $\widehat{B}\left( \alpha|X,3\right) -\widehat{B}\left( \alpha|X,2\right) $ being very smooth while $\alpha\left( \widehat{B}^{\left( 1\right) }\left( \alpha|X,3\right) /2-\widehat{B}^{\left( 1\right) }\left( \alpha|X,2\right) \right) $ is more erratic, especially when $\alpha$ is close to $1$. This issue is addressed in the application by restricting $\alpha$ to $\left[ 0,.8\right] $ for risk-aversion estimation.
Timber auctions data have been used in several empirical studies (see Athey and Levin (2001), Athey, Levin and Seira (2011) Li and Zheng (2012), Aradillas-Lopez, Gandhi and Quint (2013) among others). Some other works have investigated risk-aversion in timber auctions (e.g., Lu and Perrigne (2008), Athey and Levin (2001), Campo et al. (2011)). This section uses data from timber auctions run by the US Forest Service (USFS) from Lu and Perrigne (2008) and Campo et al. (2011), which aggregates auctions of 1979 from the states covering the western half of the United States (regions 1--6 as labeled by the USFS). It contains bids and a set of variables characterizing each timber tract, including the estimated volume of the timber measured in thousands of board feet (mbf) and its estimated appraisal value given in dollars per unit of volume. We consider the 107 first-price auctions with two bidders, the first-price auctions with three bidders ($L=108$) and ascending auctions with two bidders ($L=241$). The considered covariates are the appraisal value and the timber volume taken in log. AQR is implemented with bandwidth $h=.3$ and the Epanechnikov kernel. Pointwise confidence intervals are computed using 10,000 pairwise bootstrap replications. As in the simulation experiments, the CRRA parameter estimator has a high variance, and risk-neutrality cannot be rejected, see Gimenes and Guerre (2019a). The rest of the application therefore assumes risk-neutral bidders.
\paragraph{Specification testing.} Table (ref) reports first the results of Rothe and Wied (2013) test for the four following null hypotheses: correct specification of the quantile-regression (QR), of the homogenized-bid (HHS) model, exogeneity of the auction format (Format), participation exogeneity (Entry). The three last null hypotheses are also tested using quantile-regression coefficients comparison tests. Quantile-regression coefficient test statistics are based upon a discretized version of Liu and Luo (2017) integral statistic ((ref)), using Riemann sum over a grid $\alpha=0,1/100,\ldots,1$. For HHS, the intercept of $\widehat{\beta}_{H_0} (\cdot )$ is from the AQR estimator while the slope are OLS. For the Format null hypothesis, $\widehat{\beta}_{H_0} (\cdot )= \alpha^{-1} \int_0^{\alpha} \widehat{\gamma}_{asc} (a|2) da$, $\widehat{\gamma}_{asc} (\cdot|2)$ being an AQR version of Gimenes (2017) using the ascending auction sample with two bidders. For Entry, $\widehat{\beta}_{H_0} (\cdot )$ is an AQR version of ((ref)) using first-price auction with three bidders data. The Rothe and Wied (2013) statistic uses the unconstrained and constrained cdf estimators computed from the two bidder sample
which are compared using the Cramer-von Mises statistic \[ \frac{1}{2L} \sum_{\ell=1}^{L} \sum_{i=1}^{2} \left( \widehat{G}_{H_0}(B_{i\ell},X_{\ell}) - \widehat{G}(B_{i\ell},X_{\ell}) \right)^2 . \]
The $p$-values of the tests based upon Rothe and Wied (2013) use 10,000 replications of the bootstrap procedure proposed by these authors, while the other $p$-values are from 10,000 pairwise bootstrap replications. The Rothe and Wied (2013) procedure does not reject the quantile-regression specification at the $5\%$ level. This test also gives very high $p$-values for the Format and Entry null hypotheses, which both correspond to quantile-regression models, estimated in a different way than from the null hypothesis QR.\footnote{Rothe and Wied (2013) testing procedure seems very sensitive to the estimation variance of the considered quantile model. Attempts not reported here show that it also holds for the Escanciano and Goh (2014) bootstrap procedure, which gives smaller $p$-values. However, it does not include a re-estimation of the quantile-regression specifications, which may underestimate $p$-values in small samples. While Rothe and Wied (2013) bootstrap combines pairwise bootstrap, which draws auctions with replacement, with a semiparametric one, which draws bids from the considered model, only the semiparametric bootstrap is implemented here. } Both tests reject the homogenized-bid specification at $5\%$ level. The coefficient-based test also rejects at this level exogeneity of the auction format, disagreeing with Rothe and Wied (2013).
\paragraph{Bid quantile functions.}
Table (ref) gives the results of bid OLS regressions. The dependent variables are the bids for first-price auctions and the winning bid for the ascending auction.
The appraisal value coefficient is close to $1$ in all auctions, but is found significantly distinct at the $5\%$ level when comparing the first-price auction with $I=2$ with the one with $I=3$ and the ascending auction. Similarly the volume coefficient of the first-price auction with $I=2$ differs from the one with $I=3$ at the $10\%$ level. The appraisal value and volume coefficients of the first-price auctions with $I=2$ and $I=3$ are statistically distinct at the $5\%$ level. This is not compatible with a homogenized-bid regression model assuming entry exogeneity.
Figure (ref) gives the estimated slope for the first-price auction bids with $I=2$. The volume OLS coefficient is consistently outside the pointwise $90\%$ bootstrap confidence interval of its AQR counterpart. The appraisal value OLS estimate lies outside the AQR confidence intervals for high quantile level in $[.9,1]$. Figure (ref) also reports standard quantile-regression estimators, which exhibit a similar pattern. The intercept function does not look significant. Therefore, the intercept will be kept constant and estimated using OLS in the rest of the application. Comparison of the augmented and standard quantile-regression estimation also shows that the former produces smoother slope coefficients.
\paragraph{Private value quantile function and expected revenue.}
Figure (ref) gives the private-value slope function of the volume and appraisal variables. The shape of the volume slope varies across the type of auctions: while convex and in the $[20,100]$ range for high $\alpha$ in the first-price case, it is in the $\left[ 8,15\right] $ range and more oscillating for ascending auctions. This suggests that the private-value distribution and the auction mechanism are not independent, as also reported in Table (ref) for the test based upon ((ref)).
The appraisal value slope seems statistically different from its OLS counterpart for ascending auctions. For all auctions, the estimated appraisal value slopes start at $1$ for $\alpha$ near $0$, suggesting that low type bidders do not get added value from the appraisal value. This contrasts with high type bidders with higher $\alpha$, which markup can be very high, in a significant way for the case of ascending auction. This illustrates the important difference between low type and high type bidders.
A possible discrepancy between first-price and ascending auctions with two bidders also appears in the expected revenue computed for median values of the two explanatory variables, see Figure (ref). The ascending auction expected revenue is always below the first-price one. This seems statistically significant for high screening levels. However, this may not be relevant for the seller as the optimal revenue is achieved for a wide range $\left[ 0,.5\right] $ of screening levels over which the two expected revenue curves seem flat.
This paper proposes a quantile-regression modeling strategy for first-price auction under the independent private value paradigm, which applies quantile-level local-polynomial to estimate the private-value quantile regression. This new framework can also be used to estimate some private-value random-coefficient models and to test some specification and exogeneity hypotheses of economic interest. This approach is found to work well both in simulations, and in a timber auction application where a strong low type/high type bidder heterogeneity is detected. Another empirical finding is that the seller expected revenue in a median auction is higher in first-price than in ascending auctions, but flat in a large zone around the optimal reserve prices. This suggests that the choice of a reserve price and of an auction mechanism may not be so important, at least for the median auction considered in the application.
Many aspects of the paper deserve further investigations. The estimated constant relative risk-aversion exhibits a quite large variance, suggesting that a better understanding of efficiency issues is needed. Various extensions can also be considered, such as endogenous entry as in Marmer, Shneyerov and Xu (2013a) or Gentry and Li (2014). Our quantile approach can be extended to exchangeable affiliated values as considered in Hubbard, Li and Paarsch (2012), see also Gimenes and Guerre (2019b) for the more involved case of interdependent values. The approach of Wei and Carroll (2009) can be used to tackle unobserved heterogeneity as in Krasnokutskaya (2011).
\setcounter{equation}{0}
\setcounter{section}{0}
\setcounter{section}{0} \setcounter{footnote}{0} \setcounter{equation}{0} \setcounter{theorem}{0}
We introduce here a sieve counterpart of AQR, the Augmented Sieve Quantile Regression, or in short ASQR, which gives primary conditions for ((ref)), see Proposition (ref) below.
The private-value quantile-regression model ((ref)) assumes linearity with respect to the covariate $X$. This may be too strong and can be relaxed using a quantile nonparametric additive specification, which was considered in Horowitz and Lee (2005). The latter writes, recalling $X=\left( X_{1},\ldots,X_{D}\right) $,
where each function $V_{j}\left( \alpha;X_{j},I\right) $ is specific to the entry $X_{j}$. The effective dimension\ involved in the nonparametric estimation of this model is $1$ because it can be estimated with the rate applying for a nonparametric model with a unique covariate as shown in Horowitz and Lee (2005). This parsimonious model can be generalized following Andrews and Whang (1990) to allow for more covariate interactions. This leads to the additive interactive quantile specification with $D_{\mathcal{M}}$ interactions
where each function $V_{\mathbf{j}}\left( \alpha ;X,I\right)=V_{\mathbf{j}}\left( \alpha;X_{j_{1} },\ldots X_{j_{D_{\mathcal{M} }}},I\right) $ depends upon only $ D_{\mathcal{M}}$ entries of $X$. Setting $D_{\mathcal{M}}=D$ gives the general, or saturated, quantile specification. As seen from Andrews and Whang (1990) for the regression case, such specification can be estimated with the nonparametric rate applying for a function of $D_{\mathcal{M}}$ variables, so that $D_{\mathcal{M}}$ can be viewed as the effective dimension of this model.
The stability property in Proposition (ref)-(i) ensures that a private-value quantile specification with $D_{\mathcal{M}}$ interactions will generate a bid one with the same interactions: if ((ref)) holds, then the bid quantile function satisfies \[ B\left( \alpha|X,I\right) = \sum_{\mathbf{j}:1\leq j_{1}<\cdots<j_{d}\leq D_{\mathcal{M}}} B_{\mathbf{j}} \left( \alpha ;X,I\right) \] and the private-value components of the specification can be recovered using Proposition (ref)-(ii).
The proposed estimation of the private-value quantile specification ((ref)) is based upon a localized sieve which depends upon a smoothing parameter analogous to a bandwidth, as in regressogram methods. It is tailored for the approximation of function $\mu (\cdot)$ with $D_{\mathcal{M}}$ interactions, for $\mathbf{j}=(j_1,\ldots,j_{D_{\mathcal{M}}})^{\prime}$ and $\mu_{\mathbf{j}} \left( x \right)=\mu_{\mathbf{j}} \left( x_{j_{1}},\ldots x_{j_{D_{\mathcal{M}}}} \right)$
For the sake of brevity, the localized sieve smoothing parameter will be taken identical to the AQR bandwidth $h$. Assume that the support of $X$ is $\mathcal{X}=[0,1]^D$. The localized sieves considered here for additive interactive quantile function of order $D_{\mathcal{M}}$ as in ((ref)) depend upon a real function $p(\cdot)$ with compact support. The considered localized sieve is a collection of product functions
where the entries of the $D_{\mathcal{M}}\times 1$ $\mathbf{i}=\left( i_1, \ldots, i_{D_{\mathcal{M}}}\right)^{\prime}$ consist in all positive and negative integer numbers such that the support of $P_{\mathbf{i},\mathbf{j},h} (\cdot)$ has a non empty intersection with $[0,1]^{D}$, and $\mathbf{j}$ is as in ((ref)).
The functions $P_{\mathbf{i},\mathbf{j},h} (\cdot)$ can be reordered into a collection $\{P_{n} (\cdot),n=1,\ldots,N\}$, with $N \asymp h^{-D_{\mathcal{M}}}$, using the lexicographic order for $\mathbf{j}$ and, for a given $\mathbf{j}$, defining the successor of $\mathbf{i}$ as $\arg \max_{\mathbf{k}} \int_{[0,1]^D} \left|P_{\mathbf{k},\mathbf{j},h} (x) P_{\mathbf{i},\mathbf{j},h} (x) \right| dx$. Note that this ordering does not depend upon the bandwidth $h$ and that the resulting $N \times N$ cross product matrix \[ \int_{[0,1]^D} P (x) P(x)^{\prime} dx , \quad P (x) = \left[ P_1 (x), \ldots, P(x) \right]^{\prime} \] is a band matrix, ie its entries satisfy $\int_{[0,1]^D} P_{n_1} (x) P_{n_{2}} (x) dx = 0$ provided $|n_{2}-n_1| \geq c$, where the band size $c$ only depends upon the support of $p (\cdot)$, but not upon $h$. The standardization by $h^{-D_{\mathcal{M}}/2}$ in ((ref)) ensures that its diagonal entries satisfy $1/C \leq \int_{[0,1]^D} P_{n}^2 (x) dx <C$ for all $h>0$ and $n=1,\ldots,N$. The Euclidean norm of the sieve vector satisfies $\max_{x \in \mathcal{X}} \| P(x) \| = O(h^{-{D_{\mathcal{M}}/2}})$ because there are only $C$ non $0$ entries $P_n (x)$, in which case the bound $|P_n (x)| \leq C h^{-{D_{\mathcal{M}}/2}}$ applies by ((ref)).
A simple choice of function $p(\cdot)$ is $p(t) = \mathbb{I} \left( t \in [0,1]\right)$, in which case the $\int_{[0,1]^D} P (x) P(x)^{\prime} dx$ is the $N \times N$ identity matrix $\mathrm{Id}_N$ provided $1/h$ is an integer number. This sieve corresponds to regressogram methods, which has good approximation properties for Lipshitz functions $\mu (\cdot)$ satisfying ((ref)). In particular, for $\mu = \mu_h= \int_{\mathcal{X}} \mu (x) P (x) dx$, it holds $\max_{x \in \mathcal{X}} \left| P(x)^{\prime} \mu - \mu (x)\right|=O(h)$. This rate can be improved for functions $\mu (\cdot)$ $(s+1)$-th differentiable using a proper choice of $p(\cdot)$, see the two examples below. Consider an extension of $\mu(\cdot)$, also denoted $\mu(\cdot)$ for the sake of brevity, defined over an enlargement $\mathcal{X}_{\epsilon}=[-\epsilon,1+\epsilon]^{D}$ of $\mathcal{X}$, $\epsilon>0$ . Define the modulus of continuity of the partial derivatives $\partial_{x_d}^r \mu(x) = \frac{\partial^r}{\partial x_d^{r}} \mu(x)$ as \[ \mathrm{mc}_{r} (\mu;h) = \sum_{d=1}^{D} \sup_{x \in \mathcal{X}_{\epsilon}} \sup_{t \in [-h,h]\cap[-\epsilon-x_{d},1+\epsilon-x_{d}]} \left| \partial^{r}_{x_d} \mu \left( x_1,\ldots,x_{d-1},x_{d}+t,x_{d+1},\ldots,x_{D} \right) - \partial^{r}_{x_d} \mu (x) \right| . \] We shall assume later that the choice of the function $p (\cdot)$ ensures that:
In other words, the sieve should allow to approximate interactive functions with $r$ bounded derivatives up to an $O(h^{r})$ error, which can be improved to $o(h^{r})$ for continuous derivatives of order $r$, for all $r \leq s+1$. This is sufficient to ensure a negligible bias contribution from the sieve component of the estimation compared to the one of quantile level smoothing as shown in Theorem (ref) below.
\paragraph{Sieve Example 1: wavelets with a compact support.} Wavelet methods are a natural extension of the regressogram. They are based upon a scaling function $p(\cdot)$ which generates a multiresolution analysis, i.e. a nested sequence of linear spaces $\mathcal{P}_k$, $k$ in $\mathbb{N}$, generated by the functions $2^{-k/2} p \left( 2^{k} (x-2^{-k}i)\right)$, $i \in \mathbb{Z}$, such that $\mathcal{P}_k$ is asymptotically dense in the space $L_{2} (\mathbb{R})$ of squared integrable functions.\footnote{This exposition differs from the literature, which builds on expansions \[ \sum_{i} a_i 2^{-k_0/2} p \left( 2^{k_0} (x-2^{-k_0}i)\right) + \sum_{k\geq k_0} \sum_{i} b_{ik} 2^{-k/2} q \left( 2^{k} (x-2^{-k}i)\right) \] where the so called “father wavelets” $\left\{2^{-k/2} p \left( 2^{k} (x-2^{-k}i)\right), i \in \mathbb{Z} \right\}$ is a basis of $\mathcal{P}_k$, and the bigger space $\mathcal{P}_{k+1}$ is generated by $\left\{2^{-k/2} p \left( 2^{k} (x-2^{-k}i)\right), i \in \mathbb{Z} \right\}$ and the “mother wavelets” $\left\{2^{-k/2} q \left( 2^{k} (x-2^{-k}i)\right), i \in \mathbb{Z} \right\}$. Most of the statistical applications build on the expansion above, with a fixed $k_0$ and a truncation of the mother wavelet expansion, see H\"{a}rdle et al. (1998), Chen (2007) and the references therein. We consider instead a father wavelet expansion $\sum_{i} a_i 2^{-k/2} p \left( 2^{k} (x-2^{-k}i)\right)$ where $k$ grows with the sample size and replaces $2^{-k}$ by a bandwidth $h$ as permitted for approximation purposes by the moment condition satisfied by $\mathcal{K} (\cdot,\cdot)$, see also H\"{a}rdle et al. (1998, Chap. 8). As these two kind of expansions are equivalent, our framework also applies to father/mother expansions. However the mother wavelets with different resolution have an overlapping support, which do not fit our disjoint support assumption (ref)-(ii) below, so that using a father wavelet expansion is better suited here. Although not detailed here, thresholding can be applied to the mother/father wavelet expansion associated to our estimated father one, see H\"{a}rdle et al. (1998) and the references therein. } As developed in H\"{a}rdle, Kerkyacharian, Picard and Tsybakov (1998), wavelet methods share many common features with kernel nonparametric estimation and, as explained here, can be applied for approximation of functions with compact support using a bandwidth $h$ instead of a dyadic power. H\"{a}rdle et al.(1998, Chap. 8) consider a scaling function $p (\cdot)$ which satisfies
We assume in addition that $p(\cdot)$ has a compact support, which ensures existence of all the integrals above as $\left|\mathcal{K}(t_1,t_{2})\right|\leq C \mathbb{I} \left(|t_{2}-t_1| \leq C \right)$. H\"{a}rdle et al. (1998, Theorem 8.3-(i)) show that ((ref)) and ((ref)) are satisfied by the Daubechies scaling function of order $s+2$, see Daubechies (1992, Chaps. 6-8) and H\"{a}rdle et al. (1998, Remark 8.1 and Chap. 7). Daubechies wavelet scaling functions do not have an explicit expression but can be implemented using standard scientific softwares such as Matlab. Part ((ref)) implies that the translated functions $p(\cdot - i)$, $i$ in $\mathbb{Z}$, form an orthonormal system while ((ref)) is the analogous of the vanishing moment condition used for the bias in kernel nonparametric estimation.
Wavelet sieve for additive interactive functions can be based upon ((ref)). To check that it satisfies the Approximation Property S, it is sufficient to consider $r=s+1$ and the saturated case $D=D_{\mathcal{M}}$ for any values of $D$ and to apply the approximation procedure detailed now to each function $\mu_{\mathbf{j}} (\cdot)$ in the decomposition ((ref)). Observe that, as $\mathbf{j}$ is set to $(1,\ldots,D)^{\prime}$, $P_{\mathbf{i},\mathbf{j},h} (\cdot)$ can be abbreviated into $P_{\mathbf{i},h} (\cdot)$. The multivariate counterpart of the kernel function $\mathcal{K}(t_1,t_{2})$ is $\mathcal{K}(x,y)=\prod_{d=1}^D \mathcal{K} \left(x_d,y_d\right)$. Observe that $\mathcal{K}_h (x,y)=h^{-1} \mathcal{K} (x/h,y/h)$ is such that
and that $\mathcal{K}_h (x,y) = \sum_{n=1}^N P_n (x) P_n (y)$ when $X$ belongs to $[0,1]^D$. In view of ((ref)), a natural approximation $P(x)^{\prime} \mu$ of $\mu(x)$ is its orthogonal projection on the sieve, given by
assuming that $h$ is small enough, so that the support of $P_n(\cdot)$ is in $\mathcal{X}_{\epsilon}$ for all $n$. This gives \[ \mu'P(x) = \int \mu (y) \mathcal{K}_h (x,y) dy. \] A Taylor expansion of order $s+1$ with integral remainder gives, for any $d$
Hence ((ref)) gives, starting with $d=1$,
Iterating over index $d$ then shows that $ \sup_{x \in \mathcal{X}} \left| P(x)^{\prime} \mu - \mu (x) \right| \leq C h^{s+1} \textrm{mc}_{s+1} (\mu;h) $. Hence the Approximation Property S holds.
\paragraph{Sieve Example 2: Cardinal B-spline.} For $m\geq s+2$, set $\left( t\right) _{+}^{m-1}=t^{m-1}$ if $t>0$ and $\left( t\right) _{+}^{m-1}=0$ otherwise. The cardinal B-spline sieve is based upon the uniformly spaced simple knots $B-$spline function of order $m$ (Schumaker (2007), p.135) \[ p\left( t\right) =\sum_{i=0}^{m}\frac{\left( -1\right) ^{i}\binom{m} {i}\left( t-i\right) _{+}^{m-1}}{(m-1)!} \] which has $m-2$ continuous derivatives over the straight line and whose support is $\left[ 0,m\right] $. Theorem 12.8 of Schumaker (2007) shows that the Approximation property S holds for this choice of $p(\cdot)$,\footnote{Following Schumaker (2007) shows that it is possible to modify $p(\cdot)$ for indices $\mathbf{i},\mathbf{j}$ such that the support of $P_{\mathbf{i},\mathbf{j},h} (\cdot)$ is not a subset of $\mathcal{X}$. However considering an extension of $\mu(\cdot)$ over the enlarged $\mathcal{X}_{\epsilon}$ shows it is not necessary. Note also that Schumaker(2007) Theorem 12.8 uses a different modulus of continuity. Equivalence with the one used here follows from Schumaker (2007), Theorem 13.24 and (13.62). } constructing a spline approximation of each $\mu_{\mathbf{j}} (\cdot)$ in ((ref)) as possible with the interactive sieve ((ref)). Note that the coefficients $\mu$ used for the spline approximation are not given explicitly here. Schumaker (2007, Theorem 12.8) gives integral expression of coefficients that can be used for splines. In both examples, it is however clear that these coefficients are not unique. For instance, in the Wavelet Example 1, using $\widetilde{\mu}_n$ close enough to $\mu_n$ in ((ref)) could also work.
\setcounter{equation}{0}
The set of Assumptions below extends the one used for the AQR to the ASQR estimation method. These extended Assumptions are identified by adding “.A” to the labels in the main section. For instance, Assumption (ref) below corresponds to Assumption (ref) in the main body of paper and extends it to cover both the AQR and ASQR cases. As a consequence, proofs refer to the Assumptions stated here. Assumption (ref) remains unchanged. Assumption (ref) is new and specific to the ASQR case. Hence the labels of these assumptions do not have an additional “.A”. Instead of $\max\left( a,b\right) $ and $\min\left( a,b\right) $, the notations $a\vee b$ and $a\wedge b$ are used, respectively. Recall that $\| \cdot \|$ is the Euclidean norm.
\setcounter{hp}{0}
\setcounter{hp}{18}
\setcounter{hp}{7}
\setcounter{hp}{17}
The bandwidth rate in Assumption (ref) is driven by the linearization of $\widehat{B}^{(1)} (\cdot|\cdot,I)$, see the discussion of Theorem (ref).
Assumption (ref) covers the sieve ((ref)) as seen from (i), but is slightly more general. In particular the product structure of the $P_{\mathbf{i},\mathbf{j},h} (\cdot)$ in ((ref)) is not needed. It includes the disjoint support property (ii) which is an important characteristic of localized sieves. It is useful to obtain bounds for scalar products of the form $\mathbf{1}_N^{\prime} P(x)$, where $\mathbf{1}_N$ is a $N\times 1$ vector with unit entries, so that $\| \mathbf{1}_N\| = N^{1/2}=O(h^{-{D_{\mathcal{M}}/2}})$. Hence, while the Cauchy-Schwarz inequality $| P(x)^{\prime}\mathbf{1}_N| \leq \|P(x)\| \|\mathbf{1}_N\| $ gives $\max_{x \in \mathcal{X}} | P(x)^{\prime}\mathbf{1}_N| = O(h^{-{D_{\mathcal{M}}}})$, the fact that $P_n (x) \neq 0$ for only $c$ entries of $P(x)$ implies the better bound $\max_{x \in \mathcal{X}} | P(x)^{\prime}\mathbf{1}_N| = O(h^{-{D_{\mathcal{M}}}/2})$ as $ \max_{n \leq N} \max_{x \in \mathcal{X}} \left| P_n (x )\right| \leq \max_{x \in \mathcal{X}} \left\| P(x) \right\| = O(h^{-{D_{\mathcal{M}}}/2}) $.\footnote{This also follows from the bound $\max_{x \in \mathcal{X}} | P(x)^{\prime}\beta| \leq \max_{x \in \mathcal{X}} \sum_n \| \beta \|_{\infty}|P_n(x)| \leq Ch^{-D_{\mathcal{M}}/2} \| \beta \|_{\infty}$, where $\| \beta \|_{\infty}$ is the largest entry of $\beta$ in absolute value. } For sieve as in ((ref)), Assumption (ref)-(iii) holds for bandwidth $h$ as in Assumption (ref), provided $p\left( \cdot\right) $ is H\"{o}lder with exponent $\eta$. This allows for cardinal B-splines for which $\eta=1$, but also for wavelets which are not always differentiable but H\"{o}lder with $\eta<1$, see Daubechies (1992).
While $ \max_{x\in\mathcal{X}}\left\Vert P\left( x\right) \right\Vert =O\left(h^{-D_{\mathcal{M}}/2}\right) $ and $ \max_{n\leq N}\left\{ \int_{\mathcal{X}}\left\vert P_{n}\left( x\right) \right\vert dx\right\} =O\left( h^{D_{\mathcal{M}}/2}\right) $ follows from ((ref)) for a compactly supported $p(\cdot)$, a less intuitive condition of Assumption (ref) is $ \min_{x\in\mathcal{X}}\left\Vert P\left( x\right) \right\Vert \geq C h^{-D_{\mathcal{M}}/2} $. This is may be easier to understand considering the Wavelet examples, for which $\mu$ in the Approximation property S can be set to $\int_{\mathcal{X}} \mu (x) P(x) dx$, which is such that $\sup_{n \leq N} |\mu_n| =O(h^{D_{\mathcal{M}}/2})$. Having $ \min_{x\in\mathcal{X}}\left\Vert P\left( x\right) \right\Vert = o( h^{-D_{\mathcal{M}}/2}) $ which gives for a localized sieve $\inf_{x \in \mathcal{X}} | P(x)^{\prime} \mu | \leq c h^{D_{\mathcal{M}}/2} o( h^{-D_{\mathcal{M}}/2}) = o(1)$, which would contradict the Approximation Property S taking for instance $\mu (\cdot)$ bounded away from $0$, such as $\mu(\cdot)=1$. This extends to general localized sieve using the least square choice of $\mu$ ((ref)) used in Proposition (ref) which satisfies $\sup_{n \leq N} |\mu_n| =O(h^{D_{\mathcal{M}}/2})$.\footnote{See the proof of Proposition (ref), which does not use $ \min_{x\in\mathcal{X}}\left\Vert P\left( x\right) \right\Vert \geq C h^{-D_{\mathcal{M}}/2} $ for some $C>0$.}
\subparagraph{Localized sieve and unknown covariate support.} Assumption (ref) imposes a known $[0,1]^{D}$ covariate support. As its purpose is to impose a finite number of localized sieve coefficients, it can be assumed that the covariate support is the closure of a bounded open set. A conjecture is that assuming a known support can be relaxed using a data-driven selection of the index $\mathbf{i}$ used in the localized sieve inspired by Cuevas and Fraiman (1997). These authors proposed to estimate the support $\mathcal{X}$ using $\{x; \widehat{f}_X (x) \geq \tau_L \}$ where $\widehat{f}_X (\cdot)$ is a covariate pdf kernel estimator and $\tau_L$ a threshold sequence that goes with $0$ with the sample size. In our framework, the sieve entry $P_{\mathbf{i},\mathbf{j},h} (\cdot)$ is useful if there are enough observations in the support of neighboring $P_{\mathbf{i}^{v},\mathbf{j},h} (\cdot)$, such that $P_{\mathbf{i},\mathbf{j},h} (\cdot)P_{\mathbf{i}^{v},\mathbf{j},h} (\cdot) \neq 0$. Following Cuevas and Fraiman (1997), this suggests to select, for each given $\mathbf{j}$, all $\mathbf{i}$ and $\mathbf{i}^{v}$ such that all $\frac{1}{L} \sum_{\ell=1}^{L} \left| P_{\mathbf{i}^{v},\mathbf{j},h} (X_{\ell}) \right| \mathbb{I} (I_{\ell}=I)$ are large enough compared to $ h^{D_{\mathcal{M}/2}}$. Note that, for $\mathbf{j}=(j_1,\ldots,j_{D_{\mathcal{M}}})$, the union of the support of the selected $P_{\mathbf{i},\mathbf{j},h} (\cdot)$ is an estimator for the support of $(x_{j_1},\ldots,x_{j_{D_{\mathcal{M}}}})$ given $I$. The sieve estimator should deliver its best performance over this estimated support.
The Approximation Property S ensures that there exists a sieve coefficient vector $\gamma (\alpha|I)=\gamma_{h} (\alpha|I)$ such that, for each $\alpha$, $I$ and $X$
for an interactive private-value quantile function $V(\alpha|X,I)$ satisfying some smoothness conditions. The LHS of the equation above is an asymptotic sieve quantile regression which can be estimated using an augmented quantile regression. Making this rigorous necessitates first to elicit the smoothness properties of the coefficients $\gamma (\cdot|I)$, which, as suggested by their expression ((ref)) for the Wavelet Example 1, should inherit the smoothness of the parent $V(\alpha|X,I)$. As the sieve coefficients used in the Approximation Property S are not unique, it is necessary to be more specific in their choice and, from now on, we use a Least Square choice of the sieve coefficient $\gamma (\cdot|I)$ of $V(\cdot|\cdot,I)$\footnote{In the AQR case, the slope coefficient in ((ref)) is identical to the quantile-regression slope as $P(X)=X_1$ and $V(\alpha|X,I)=X_1^{\prime} \gamma (\alpha|I)$.}
assuming that $ \mathbb{E} \left[ \left. P(X ) P(X)^{\prime} \right| I \right] $ is full rank. Second, a similar task should be carried for the asymptotic sieve bid quantile-regression model generated by ((ref)). These issues are addressed in the next Proposition, which proof is given in (ref). Note that it is also useful in the AQR case, in which case ((ref)) is identical to the slope in $V(\alpha|X,I)=X_{1}^{\prime} \gamma(\alpha|I)$ setting $P(X)=X_1$ and $B(\alpha|X,I)=X_1^{\prime} \beta(\alpha|I)$ in (ii) and (iii).
Note that the sieve bid quantile slope in ((ref)) can be defined through a least square formula similar to ((ref)), that is
Proposition (ref) gathers results on various rate of approximation which allows to extend the AQR approach to a sieve setup, starting from the asymptotic sieve quantile regression ((ref)) with slope $\gamma (\cdot)$ for the private value. Defining the sieve bid coefficients as in ((ref)) gives an asymptotic quantile regression ((ref)) \[ B(\alpha|x,I) = P(x)^{\prime} \beta (\alpha|I) + o \left(h^{s+1}\right) \] with a differentiable slope $\beta(\cdot|\cdot)$ which identifies the private-value slope using
as in the quantile-regression case, see ((ref)). In addition, Proposition (ref) establishes rates for the approximation of $B^{(p)} (\cdot|\cdot,I)$ using the $p$-th slope derivatives $\beta^{(p)} (\cdot|I)$ as necessary for the augmented approach presented in the next section. Note however that Proposition (ref)-(iii) does not give a rate for $ \max_{(\alpha,x) \in [0,1] \times \mathcal{X}} \left| P(x)^{\prime} \alpha \beta (\alpha|I) - \alpha B (\alpha|x,I) \right| $ which is at least $o(h^{s+1})$ by (ii). This is specific to the the ASQR case, as arguing with Taylor expansions as for the proof of Proposition (ref) gives a better rate $o(h^{s+2})$ for the AQR.
\paragraph{Augmented sieve quantile regression and functional estimation.} The AQR can be modified in a straightforward way to account for our sieve extension, leading to the Augmented sieve quantile regression (ASQR), which proceeds by using a vector $b$ of dimension $(s+2)N$ instead of $(s+2)(D+1)$ and to redefine $P(x,t)$ in ((ref)) as
The expression of the objective function $\widehat{\mathcal{R}}\left( b;\alpha,I\right)$ is unchanged and as in ((ref)). The estimation of $b\left( \alpha|I\right) = [\beta(\alpha|I)^{\prime}, \ldots, \beta^{(s+1)}(\alpha|I)^{\prime}] $ is $\widehat{b}\left( \alpha|I\right) = \arg\min_{b\in \mathbf{R}^{(s+2)N}} \widehat{\mathcal{R}}\left( b;\alpha,I\right) $ and the ASQR estimators are
The functional estimators use a plug-in construction as the AQR case.
\setcounter{equation}{0}
We give here the sieve version of Theorems (ref) and (ref). The pseudo true $\overline{V}(\cdot|\cdot)$ and the Bahadur leading term $\widetilde{V}(\cdot|\cdot)$ of $\widehat{V}(\cdot|\cdot)$ are also defined, respectively, as in ((ref)) and ((ref)) for the sieve case. While, in the ASQR case, the bias items $\mathsf{Bias}_{I}$ and $\mathsf{Bias}_h (\alpha|X,I)$ are defined as for the AQR, the variance items $\Sigma_{IL}$ and $\Sigma_{h}(\alpha|I)$ below are defined using $P(X)$ instead of the covariates in the variance items $\Sigma_{I}$ and $\Sigma_{h}(\alpha|I)$ introduced for Theorems (ref) and (ref). All these ASQR results are proved in (ref) with their AQR counterparts.
For the sieve counterpart of Theorem (ref), change the covariate $x_1$ into the sieve $P(x)$ in the definition of $\mathbf{P}_{0}\left( \alpha|I\right)$ and $\mathbf{P}\left( I\right)$. Recall $ \mathbf{1}_{N} = \left[1,\ldots,1 \right]^{\prime} $ is a $N\times 1$ column vector and redefine $\sigma_{L}^{2}\left( x|I\right)$ and $\sigma_{L}^{2}\left( I\right)$ as
with $\sigma_{L}^{2}\left( x\right) = \sum_{I\in\mathcal{I}} \frac{\sigma_{L}^{2}\left( x|I\right)}{I}$ and $ \sigma_{L}^{2}=\sum_{I\in\mathcal{I}}\frac{\sigma_{L}^{2}\left( I\right)}{I} $.
As expected, the number of interactions $D_{\mathcal{M}}$ plays the role of a covariate dimension in these sieve results, showing its ability to circumvent the curse of dimensionality. It is worth mentioning that, as in GPV, our results do not request a smoothness $s$ increasing with $D_{\mathcal{M}}$. This is due to the use of a localized sieve and specific proof techniques. When applied their results established for general sieve to cardinal B-spline,, Belloni, Chernozhukov, Chetverikov and Fern\'{a}ndez-Val (2019, Corollaries 1 and 2, Comment 4) have a $s \geq D$ and $s \geq 3D/2$ restrictions for quantile estimator, where their covariate dimension $D$ plays a role similar to $D_{\mathcal{M}}$. Aryal, Gabrielli and Vuong (2019, Assumption A2-(ii)) use a condition $s>D+1$ for a semiparametric version of GPV based on a local polynomial estimation of the private values. Such restrictions limit applications with many covariates or interactions, as $s$ is usually set to $2$, or even to 1 as in our applications.
Theorem (ref) holds under specific bandwidth conditions, which are driven by the order $\log L/(Lh^{4 D_{\mathcal{M}}+3} )$ of the linearization error term $\widehat{V} (\alpha|x,I) - \widetilde{V} (\alpha|x,I)$ in ((ref)), which must be $o(1/\sqrt{Lh^{D_{\mathcal{M}}}})$ for $\widehat{\theta} (x)$ and $o(1/\sqrt{L})$ for $\widehat{\theta}$.
\setcounter{equation}{0}
\setcounter{section}{1}
\setcounter{section}{0} \setcounter{footnote}{0} \setcounter{equation}{0} \setcounter{theorem}{0}
\subparagraph{Proof main arguments.} The proofs use some notations detailed below which allows to study the AQR and ASQR using the same framework. The intuition behind the various proofs are rather simple, although execution is long. Starting with Lemma (ref)-(i), it consists in finding a vicinity of the renormalized true parameter value $\mathsf{b}(\alpha|I)$ such that $t \mapsto P(x,t)\mathsf{b}$ is strictly increasing in a neighborhood of $0$, for all $x$. This ensures that the sample and population objective functions are both twice differentiable and strictly convex asymptotically over this vicinity, with a probability tending to 1 for the sample one. This vicinity is shown to be large enough to contain a local minimum of the objective functions, which is also the global minimum as they are convex over the full parameter space and strictly convex in the considered vicinity.\footnote{To see this, suppose there is another local minimum outside the considered vicinity, which has an inner local maximum. Consider the segment joining their locations. Strict convexity over the vicinity implies the objective function must increase on this segment near its inner local minimum. It cannot increase all along the segment. That it decreases or is constant later would violate global convexity, so that the vicinity local minimum should also be a global one.} The fact that the sample and population Hessians are Lipshitz then allows to obtain a bias approximation, see Theorem (ref), and a Bahadur representation for $\widehat{b}(\cdot|\cdot,I)$, (i.e. a linear expansion given in Theorem (ref)) which rate is important to cope with the fact that quantile and derivative slope estimators converge with different rates. Our main results follow from this two building blocks and from Lemma (ref) which studies the stochastic component of $\widehat{b}(\cdot|\cdot,I)$. For the sieve case, the proof arguments crucially rely on the vector supremum norm $\| \cdot \|_{\infty}$ and on various important properties induced by the disjoint support condition of Assumption (ref)-(ii) which implies that the matrix $P(x)P(x)^{\prime}$ is a band matrix. This starts with Lemma (ref) below which studies the supremum norm of the inverse of band matrices depending upon the quantile level, as the AQSR Hessian matrix which is a band one up to a basis permutation.
\subparagraph{Structure of appendices.} The preliminary lemmas of these appendices deals with the monotonicity of $t \mapsto P(x,t)^{\prime} \mathsf{b}$ and $\mathsf{b} \mapsto P(x,t)^{\prime} \mathsf{b}$, see Lemma (ref), which are key to get invertible population and sample Hessian, a distinctive feature of the augmented procedure. See Lemmas (ref) and (ref), which are used to study the bias in (ref) and derive a Bahadur expansion in (ref). Lemma (ref), which is used for the latter, deals with the score function. Lemma (ref) studies the leading term ((ref)) of $\widehat{V}(\cdot|\cdot)$. Proposition (ref) and Lemmas (ref)-(ref) are proven in (ref), which also gathers the proof of Lemmas (ref) and (ref), which are more specific to Theorems (ref) and (ref). These results are derived under the assumptions listed in Section (ref), which extends to the sieve case the ones of the main part of the paper.
The most important intermediary results are Theorem (ref) in (ref) for the bias, and the Bahadur expansion of Theorem (ref) in (ref). Our main results, Theorems (ref) and (ref), (ref) and (ref), (ref) and (ref) are established in corresponding Sections of (ref).
\subparagraph{Notations.} As mentioned earlier, $\wedge$ and $\vee$ stand for $\min$ and $\max$ respectively. The usual order for symmetric matrices is denoted $\preceq$, i.e. $A \preceq B$ if $B-A$ is a nonnegative symmetric matrix. Our default norms, $\|\cdot\|$, are the Euclidean norm of a vector or the associated operator norm of a matrix that we also denote $\|\cdot\|_{2}$, and the $\|\cdot\|_{\infty}$ norm for vectors, $\|\mathsf{b}\|_{\infty}=\max_{j} |\mathsf{b}_j|$. Note that for $\mathsf{b}$ of dimension $d_h=N$, $N(s+2)$ or $D$,
For the associated balls $\mathcal{B}_j (\mathsf{b},\epsilon) = \{\mathsf{a};\|\mathsf{a}-\mathsf{b}\|_j \leq \epsilon\}$, $j=2,\infty$, it therefore holds $\mathcal{B} (\mathsf{b},\epsilon)= \mathcal{B}_{2} (\mathsf{b},\epsilon)\subset \mathcal{B}_{\infty} (\mathsf{b},\epsilon)$. Note that for two conformable vectors $\mathsf{b}_1$, $\mathsf{b}_{2}$, it holds $| \mathsf{b}_1^{\prime} \mathsf{b}_{2} | \leq \| \mathsf{b}_1 \| \| \mathsf{b}_{2} \|$ which is the standard Cauchy-Schwarz inequality, but also $| \mathsf{b}_1^{\prime} \mathsf{b}_{2} | \leq \| \mathsf{b}_1 \|_1 \| \mathsf{b}_{2} \|_{\infty}$ with $\| \mathsf{b}_1 \|_1=\sum_{j} |\mathsf{b}_{1j}|$. In particular, $| P(x)^{\prime} \beta | \leq \| P(x) \|_1 \|\beta\|_{\infty}$, noting that the disjoint support property in Assumption (ref)-(ii) ensures $\| P(x) \|_1 \leq c^{1/2} \| P(x) \|_{2} \leq c \| P(x) \|_{\infty}$, so that $ \max_{x \in \mathcal{X} } | P(x)^{\prime} \beta | \leq C h^{-D_{\mathcal{M}}/2}\|\beta\|_{\infty} $ by Assumption (ref)-(i).
The operator norms associated to $\|\cdot\|_{2}$ and $\|\cdot\|_{\infty}$ are, for a matrix $A$, $\|A \|_{j}=\sup_{\mathsf{b}:\|\mathsf{b}\|_{j}=1} \|A \mathsf{b} \|_{j}$, so that $\|A \mathsf{b} \|_{j} \leq \|A \|_{j} \|\mathsf{b}\|_{j}$, $j=2,\infty$. We also use $|A|_{\infty}=\max_{i,j}|A_{ij}|$. As well-known, $\|A\|_{2}$ is the largest eigenvalue in absolute value of a symmetric $A$.
\subparagraph{Matrix norm equivalence for band matrix.} When restricted to $c$-band matrices with $A_{ij}=0$ for $|i-j|\geq c/2$, the considered matrix norms satisfy\footnote{ The inequalities for $\| A \|_{\infty}$ follow $ \| A \|_{\infty} = \max_{\|\mathsf{b}\|_{\infty}=1} \max_{i} \left| \sum_{j:|j-i|< c/2} A_{ij} \mathsf{b}_j \right| = \max_{i} \sum_{j:|j-i|< c/2} \left| A_{ij} \right| $. For $\| A \|_{2}$, if $\{ s_i \}$ is the canonical basis, $| A_{ij} | = |s_i^{\prime} A s_j|\leq \|s_i \|_{2} \| A \|_{2} \| s_j \|_{2} = \| A \|_{2}$, so that $| A |_{\infty} \leq \| A \|_{2}$. The inequality $\| A \|_{2} \leq c | A |_{\infty}$ follows from $\|A\|_{2} = \sup_{\mathsf{b}:\|\mathsf{b}\|_{2}=1} \| A\mathsf{b} \|_{2}$ and
}
which establishes equivalence of the three matrix norms over the linear space of $c$-band matrices. Note that this holds independently of the matrix dimension, and this also holds for permutations $A_{\sigma} = \Big[A_{\sigma (i),\sigma (j)} \Big]$ of $c$-band matrices, where $\sigma (\cdot)$ is an index permutation. This in particular covers $\widehat{\mathsf{R}}^{(2)} (\mathsf{b};\alpha,I)$, $\overline{\mathsf{R}}^{(2)} (\mathsf{b};\alpha,I)$ in addition to $\mathbb{E} \left[ \mathbb{I} \left(I_{\ell}=I\right)P(X_{\ell})P(X_{\ell})^{\prime}\right]$ as $P(x,t)P(x,t)^{\prime}=\pi(t) \pi (t)^{\prime} \otimes P(x) P(x)^{\prime}$ is derived through an index permutation from the $c(s+2)/2$-band matrix $P(x) P(x)^{\prime} \otimes \pi(t) \pi (t)^{\prime}$. We shall also make use of the following Lemma, which deals with the $\|\cdot\|_{\infty}$ norm of the inverse of band matrices. The proof of Lemma (ref) can be found in (ref).
We start with additional notations used all along the proof section and some preliminary lemmas which are established in (ref). Set $N=D+1$ and $D_{\mathcal{M}}=0$ in the AQR case, $N$ being the dimension of the sieve $P(\cdot)$ in the ASQR case. Recall that the AQR and ASQR estimators write \[ \widehat{b} (\alpha|I) = \left[ \widehat{\beta}_0 (\alpha|I)^{\prime},\widehat{\beta}_1 (\alpha|I)^{\prime}, \ldots, \widehat{\beta}_{s+1} (\alpha|I)^{\prime} \right]^{\prime} \] where the $\widehat{\beta}_{j} (\alpha|I)^{\prime}$ are $1\times (D+1)$ in the AQR case and $1 \times N$ in the ASQR case. Recall that $\widehat{\beta}_{j} (\alpha|I)$ estimate $\beta^{(j)} (\alpha|I)$ in the AQR, see ((ref)) in Proposition (ref) for the ASQR case.\footnote{The notation $\widehat{\beta}^{(j)} (\alpha|I)^{\prime}$ is avoided because $\widehat{\beta}_{j} (\alpha|I)^{\prime}$ is not the $j$-th derivative of $\widehat{\beta}_{0} (\alpha|I)$.} Set \[ P\left( x\right) =\left\{
\right. \] and recall $P(x,t) =\pi (t) \otimes P(x)$ in both case, allowing an unified treatment of the two estimators, although the proof focus is on the more difficult ASQR case. Recall that $\left\Vert P\left( x\right) \right\Vert =\left( P\left( x\right) ^{\prime}P\left( x\right) \right) ^{1/2}$ is the standard Euclidean norm and that, under Assumptions (ref) and (ref)-(i), $ \max_{\left( x,t\right) \in\mathcal{X\times}\left[ -1,1\right] }\left\Vert P\left( x,t\right) \right\Vert =O\left( h^{-D_{\mathcal{M}}/2}\right)$.
\subparagraph{Parameter renormalization.} Since \[ P\left( x,ht\right) =\pi\left( ht\right) \otimes P\left( x\right) ,\quad\pi\left( ht\right) ^{\prime}=\left[ 1,ht,\ldots,\frac{\left( ht\right) ^{s+1}}{\left( s+1\right) !}\right] \text{ with $\lim_{h \downarrow 0} \pi\left( ht\right) =[1,0,\ldots,0]^{\prime}$,} \] the \textquotedblleft design\textquotedblright\ matrix $\mathbb{E} \left[ P\left( X,ht\right) P\left( X,ht\right) ^{\prime }\right] $ degenerates asymptotically. To avoid this, consider the change of parameters $\mathsf{b}=Hb$ with $H=\operatorname*{Diag}\left( 1,\ldots ,h^{s+1}\right) \otimes\operatorname*{Id}_{N}$,
so that $P\left( x,ht\right) ^{\prime}\beta=P\left( x,t\right) ^{\prime}\mathsf{b}$. Let $S_0 = [1,0,\ldots,0]$ and $S_1 = [0,1,0,\ldots,0]$ be $1 \times (s+2)$ row vectors and define $\mathsf{S}_0 = S_0 \otimes \mathrm{Id}_N$, $\mathsf{S}_1 = S_1 \otimes \mathrm{Id}_N$, with $N=D+1$ in the AQR case. These selection matrices are such that
Define accordingly
which are such that \[ \widehat{\mathcal{R}}\left( b;\alpha,I\right)=\widehat{\mathsf{R}}\left( Hb;\alpha,I\right), \quad \mathbb{E} \left[ \widehat{\mathcal{R}}\left( b;\alpha,I\right) \right] = \overline{\mathsf{R}}\left( Hb;\alpha,I\right). \] Note that $\mathsf{b}\mapsto\mathsf{\int_{-\frac{\alpha}{h}} ^{\frac{1-\alpha}{h}}}\rho_{a+ht}\left( B_{i\ell}-P\left( X_{\ell},t\right) ^{\prime}\mathsf{b}\right) K\left( t\right) dt$ is convex as an integral of convex functions. It follows that $\widehat{\mathsf{R}}\left( \mathsf{b} ;\alpha,I\right) $ and $\overline{\mathsf{R}}\left( \mathsf{b} ;\alpha,I\right) $ have minimizers,
which uniqueness will be established in the next section. Set $\overline {b}\left( \alpha|I\right) =H^{-1}\overline{\mathsf{b}}\left( \alpha |I\right) $ recalling $\overline{b}\left( \alpha|I\right) =\left[ \overline{\beta}_{0}\left( \alpha|I\right) ^{\prime},\ldots,\overline{\beta }_{s+1}^{\prime}\left( \alpha|I\right) \right] ^{\prime}$ and define $\overline{B}\left( \alpha|x,I\right) =P\left( x\right) ^{\prime} \overline{\beta}_{0}\left( \alpha|I\right) ,$ \[ \overline{\gamma}_{0}\left( \alpha|I\right) =\overline{\beta}_{0}\left( \alpha|I\right) +\frac{\alpha\overline{\beta}_{1}\left( \alpha|I\right) }{I-1},\quad\overline{V}\left( \alpha|x,I\right) =P\left( x\right) ^{\prime}\overline{\gamma}_{0}\left( \alpha|I\right) . \]
\subparagraph{Pseudo-true value and value of interest.} To sum up, we interpret $\overline{\mathsf{b}}\left( \alpha|I\right)$ as a pseudo-true value, while the aim is to estimate a slope $\mathsf{b}\left( \cdot|\cdot\right)$ derived from ((ref)) in the AQR case or ((ref)) for the ASQR. Observe there exists some $b \left( \cdot|\cdot\right) $ in $\mathbb{R}^{N(s+2)}$, \[ b\left( \alpha|I\right) ^{\prime}=\left[ \beta\left( \alpha|I\right) ^{\prime},\beta^{(1)} \left( \alpha|I\right) ^{\prime },\ldots,\beta^{(s+1)} \left( \alpha|I\right) ^{\prime}\right] , \] such that \[ \sup_{\left( \alpha,x\right) \in\left[ 0,1\right] \times\mathcal{X} }\left\vert P\left( x\right) b\left( \alpha|I\right) -B\left( \alpha|x,I\right) \right\vert =o\left( h^{s+1}\right) . \] by Proposition (ref) with $\beta(\alpha|I)$ as in ((ref)) for the ASQR case and, for the AQR one, with $\beta(\alpha|I)$ as in ((ref)) and a remainder term $o\left( h^{s+1}\right)$ set to $0$. An intermediate aim of the proof section is to estimate $\mathsf{b}\left( \cdot|\cdot\right) =Hb\left( \cdot|\cdot\right) $.
\subparagraph{The score function.} The next notations deal with the two first derivatives of the objective functions $\widehat{\mathsf{R}}\left( \mathsf{\cdot};\alpha,I\right) $. Since \[ \partial_{\mathsf{b}} \left[ \rho_{\alpha+ht}\left( B-P\left( X_{\ell},t\right) ^{\prime }\mathsf{b}\right) \right] = \left\{ \mathbb{I}\left( B_{i\ell}\leq P\left( X_{\ell},t\right) ^{\prime }\mathsf{b}\right) -\left( \alpha+ht\right) \right\} P\left( X_{\ell },t\right) , \] almost everywhere, it follows that $\widehat{\mathsf{R}}\left( \mathsf{\cdot };\alpha,I\right) \ $is differentiable with
by the Dominated Convergence Theorem. Note also that $\widehat{\mathsf{R}}^{\left( 1\right) }\left( \mathsf{b};\alpha,I\right)$ is equal to \[ \frac{1}{LI}\sum_{\ell=1}^{L}\mathbb{I}\left( I_{\ell}=I\right) \sum_{i=1}^{I_{\ell}} \int_{0}^{1} \left\{ \mathbb{I}\left( B_{i\ell}\leq P\left( X_{\ell},\frac{a-\alpha}{h}\right) ^{\prime}\mathsf{b}\right) -a \right\} P\left( X_{\ell},\frac{a-\alpha}{h}\right) K\left( \frac{a-\alpha}{h}\right) da. \]
\subparagraph{The Hessian.} To compute the Hessian $\widehat{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b};\alpha,I\right)$, it is convenient to remove the indicator $\mathbb{I}\left( B_{i\ell}\leq P\left( X_{\ell},t\right)^{\prime}\mathsf{b}\right)$ in ((ref)) by showing that $t \mapsto P\left( X_{\ell},t\right)^{\prime}\mathsf{b}$ is strictly increasing for suitable $\mathsf{b}$. The relevant interval for $t$ in the integral expression of $\widehat{\mathsf{R}}^{\left( 1\right) }\left( \mathsf{b};\alpha,I\right)$ is \[ \mathcal{T}_{\alpha,h}=\left[ \underline{t}_{\alpha,h},\overline{t}_{\alpha,h}\right] =\left[ -\min\left( 1,\frac{\alpha}{h}\right) ,\min\left( 1,\frac{1-\alpha}{h}\right) \right] =\left[ -1,1\right] \cap\left[ -\frac{\alpha}{h},\frac{1-\alpha}{h}\right], \] as the support of $K(\cdot)$ is $[-1,1]$. It is convenient to redefine $P(x,t)^{\prime} \mathsf{b}$ as a constant function outside $\mathcal{T}_{\alpha,h}$, that is\footnote{In principle $\Psi\left( \cdot|\cdot\right) $ should be denoted $\Psi_{\alpha,h}\left( \cdot |\cdot\right) $ to acknowledge that its definition depends upon $\alpha$ and $h$. Instead, $t$ is restricted to lie in $\mathcal{T}_{\alpha,h}$ in the sequel. The same comment applies for the functions $\Psi\left( \cdot |\cdot\right) $ and $\Delta\left( \cdot|\cdot\right) $ introduced below.} \[ \Psi\left( t|x,\mathsf{b}\right) =\left\{
\right. . \] When $\mathsf{b}=\mathsf{b}\left( \alpha|I\right) $, $\Psi\left( t|x,\mathsf{b}\left( \alpha|I\right)\right) =P\left( x,ht\right) ^{\prime}b\left( \alpha|I\right) $ is close to $B\left( \alpha+ht|x,I\right) $, which inverse as a function of $t$ is \[ \frac{G\left( u|x,I\right) -\alpha}{h}\text{,\quad}u\in\left[ B\left( \alpha+h\underline{t}_{\alpha,h}|x,I\right) ,B\left( \alpha+h \overline{t}_{\alpha,h}|x,I\right) \right] . \] When $h$ is small enough, Lemma (ref) below shows that $\Psi\left( \cdot|x,\mathsf{b}\left( \alpha|I\right)\right)$ is strictly increasing provided $\mathsf{b}$ is in a suitable vicinity of $\mathsf{b}\left( \alpha|I\right)$. In such case, define
which is such that, as seen above, the central part of $\Phi\left( u|x,\mathsf{b}\left( \alpha|I\right) \right) $ is close to $G\left( u|x,I\right) $ when $u$ is in $\Psi\left( \mathcal{T}_{\alpha,h} |x,\mathsf{b}\right) $. Observe now that, provided $\Psi\left( \cdot|x,\mathsf{b}\right) $ is increasing and since the support of $K\left( \cdot\right) $ is $\left[ -1,1\right] $
which is differentiable with respect to $\mathsf{b}$, with by the Implicit Function Theorem and for $B_{i\ell}$\ in $\Psi\left( \mathcal{T}_{\alpha,h}|x,\mathsf{b}\right) $ \[ \partial_{\mathsf{b}} \Phi\left( B_{i\ell}|X_{\ell},\mathsf{b}\right) =-\frac{P\left( x,\Delta\left( B_{i\ell }|X_{\ell},\mathsf{b}\right) \right) }{\Psi^{\left( 1\right) }\left( \Delta\left( B_{i\ell}|X_{\ell},\mathsf{b}\right) |X_{\ell},\mathsf{b} \right) /h}\mathbb{I}\left[ B_{i\ell}\in\Psi\left( \mathcal{T}_{\alpha ,h}|X_{\ell},\mathsf{b}\right) \right] . \] Since $K(\cdot)$ must vanish at its frontier boundaries and is continuous, $\widehat{\mathsf{R}}\left( \mathsf{b};\alpha,I\right) $ and $\overline{\mathsf{R}}\left( \mathsf{b} ;\alpha,I\right) $ are twice continuously differentiable over a vicinity of $\mathsf{b}\left( \alpha|I\right) $ for all $h$ small enough with,
This existence of a sample second derivative contrasts with standard quantile-regression methods.
\subparagraph{Properties of $\Psi\left( \cdot|x,\mathsf{b}\right) $ and $\Phi\left( \cdot|x,\mathsf{b}\right) $.} Lemma (ref)-(i) below gives necessary conditions ensuring $\Psi\left( \cdot|x,\mathsf{b}\right)$ is strictly increasing for all $\mathsf{b}$ in a suitable vicinity, and then existence of $\widehat{\mathsf{R}}^{\left( 2\right) }\left( \cdot;\alpha ,I\right)$. Part (ii) recalls derivatives of $\Phi\left( \cdot|x,\mathsf{b}\right)$ derived from the Implicit Function Theorem, as used above to sketch existence of the Hessian. Lemma (ref)-(iii) completes Proposition (ref), and gives an expansion for $\alpha \left(P(x,t)^{\prime} \mathsf{b} (\alpha|I)-B(\alpha+ht|x,I)\right)$ which is used for the bias of the AQR and ASQR estimators. Lemma (ref)-(iv) establishes that $\mathsf{b} \mapsto \Psi\left( t|x,\mathsf{b}\right), \Phi\left( t|x,\mathsf{b}\right)$ are Lipshitz, with a Lipshitz factor diverging with the sample size.
Define
where $\underline{f}$ and $\overline{f}$ will be taken large enough later one. While $\mathcal{BI}_{\alpha,h}$ is used to bound the first derivative of $\Psi\left( \cdot|x,\mathsf{b}\right) $ away from $0$ to get strict monotonicity, $\underline{\mathcal{BI}}_{\alpha,h}$ is used to bound the successive derivatives $\Psi^{\left( p\right) }\left( \cdot|x,\mathsf{b} \right) $, $p=1,\ldots,s+1$, away from infinity. As made possible by Lemma (ref)-(i), below, a ball $\mathcal{B}_{\infty} \left( \mathsf{b} \left( \alpha|I\right) ,Ch^{D_{\mathcal{M}}/2+1}\right) $ with a small enough constant $C>0$ can be considered instead of the sets $\mathcal{BI}_{\alpha,h}$ and $\underline{\mathcal{BI}}_{\alpha,h}$.
Lemma (ref)-(i) implies that the sample and population objective function are twice continuously differentiable in the vicinity of the true value $\mathsf{b} (\alpha|I)$. This will allow to use standard first-order linearization technique to study the AQR and ASQR estimators. The order $h^{D_{\mathcal{M}}/2+1}$ for the ball radius in (i) can be understood from ((ref)), assuming that $\mathbb{E} [ P(X)P(X)^{\prime}|I]$ is the identity matrix to simplify the discussion.\footnote{If not, Lemma (ref) can also be used to show that $\max_{\alpha \in [0,1]}\| \beta^{(p)} (\alpha) \|_{\infty} = O (h^{D_{\mathcal{M}}/2})$ for all $p\leq s+1$.} If so $\max_{1 \leq n \leq N} \int_{\mathcal{X}} |P_n (x)| dx = O (h^{D_{\mathcal{M}}/2})$ implies that $\max_{\alpha \in [0,1]}\| \beta^{(p)} (\alpha) \|_{\infty} = O (h^{D_{\mathcal{M}}/2})$. This gives $\Psi^{(1)}(t|x,\mathbf{b}(\alpha|I)) = h P(x)^{\prime} \beta^{(1)} (\alpha|I)+O(h^2)$, with a leading term $h P(x)^{\prime} \beta^{(1)} (\alpha|I)$ which is positive for all $\alpha,x$ provided $h$ is small enough by Proposition (ref)-(ii). Therefore, $\Psi^{(1)}(t|x,\mathsf{b})$ will be positive and $\Psi(\cdot|x,\mathsf{b})$ increasing if $\|\mathsf{b} - \mathbf{b}(\alpha|I) \|_{\infty}$ is small compared to $h\beta^{(1)} (\alpha|I)$, which is of order $h^{D_{\mathcal{M}}/2+1}$.
\subparagraph{Population Hessian.} Let $\Omega_{h}\left( \alpha\right) $, $\Omega\left( 0\right) $, $\Omega\left( 1\right) $, $\Omega=\Omega\left( 0\right) +\Omega\left( 1\right) $ and $\Omega_{1h}\left( \alpha\right) $ be the $\left( s+2\right) \times\left( s+2\right) $ matrices
While $\Omega_{h}\left( \alpha\right) \preceq\Omega$ for all $\alpha$ and $h$, it holds that for $h$ small enough $\Omega_{h}\left( \alpha\right) \succeq\Omega\left( 0\right) $ for all $\alpha$ in $\left[ 0,1/2\right] $ and $\Omega_{h}\left( \alpha\right) \succeq\Omega\left( 1\right) $ for all $\alpha$ in $\left[ 1/2,1\right] $, ensuring that the eigenvalues of these matrices stay bounded away from $0$ and infinity when $h$ goes to $0$. Define also
Note that Assumptions (ref)-(i), (ref) and Proposition (ref)-(i) ensures that $\mathbf{P}_{0}\left( \alpha |I\right)$ is well-conditioned for all $\alpha$ in $[0,1]$, as $\int_{\mathcal{X}} P(x) P(x)^{\prime} dx$ is and $f(x,I)/B^{(1)} (\alpha|x,I)$ is bounded away from $0$ and infinity so that
These two inequalities yield that for any $N \times 1$ vector $S$, $ S^{\prime} \mathbf{P}_{0}\left( \alpha |I\right) S \asymp S^{\prime} \int_{\mathcal{X}} P(x) P(x)^{\prime} dx S $ uniformly in $\alpha$ and $S$. Taking the infimum and supremum over those $S$ with $\left\| S \right\|=1$ gives that the smallest and largest eigenvalues of $\mathbf{P}_{0}\left( \alpha |I\right)$ are, up to some constants, between the smallest and largest ones of $\int_{\mathcal{X}} P(x) P(x)^{\prime} dx$ for all $\alpha$, so that the eigenvalues of $\mathbf{P}_{0}\left( \alpha |I\right)$ are bounded away from $0$ and infinity for all $\alpha$ and all $h$ small enough. Similarly, $\mathbf{P}\left( I\right)$ is well-conditioned provided $h$ is small enough.
In (i) the normalization item $\alpha(1-\alpha)+h$ indicates a rate dependence with respect to the quantile level $\alpha$, and more specifically some boundary effects. For instance, if $\alpha = O(h)$ or $1-\alpha = O(h)$, then for all $\mathsf{b}^{0}$, $\mathsf{b}^{1}$ as in the Lemma, \[ \left\Vert \overline{\mathsf{R} }^{\left( 2\right) }\left( \mathsf{b}^{1};\alpha,I\right) \mathsf{-} \overline{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b}^{0} ;\alpha,I\right) \right\Vert_{\infty} \leq \frac{O\left( h^{-D_{\mathcal{M}}/2}\right)}{h} \left\Vert \mathsf{b}^{1}-\mathsf{b}^{0}\right\Vert _{\infty} \] while for $\alpha$ in the vicinity of $1/2$ \[ \left\Vert \overline{\mathsf{R} }^{\left( 2\right) }\left( \mathsf{b}^{1};\alpha,I\right) \mathsf{-} \overline{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b}^{0} ;\alpha,I\right) \right\Vert_{\infty} \leq O\left( h^{-D_{\mathcal{M}}/2}\right) \left\Vert \mathsf{b}^{1}-\mathsf{b}^{0}\right\Vert_{\infty} \] indicating a Lipschitz constant inflated by a $1/h$ factor for extreme quantile levels compared to central ones. The use of the factor $\alpha(1-\alpha)+h$ is repeatedly used below to capture some quantile level boundary effects.
Lemma (ref)-(i) yields, for any $C>0$ and by the matrix norm equivalence ((ref)),
noting that the bandwidth condition above holds under Assumption (ref).
It then follows that the eigenvalues of $\overline{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b};\alpha,I\right) $ stays bounded away from $0$ and infinity uniformly in $\alpha$ and in $\mathsf{b}$ in the two neighborhoods considered above. The fact that this holds for all $\alpha$, including $0$ and $1$, is useful to establish existence and uniqueness of the pseudo-true value $\mathsf{b} (\alpha|I)$. Note however that the matrix expansion leading term in Lemma (ref)-(ii) involves $\Omega_h (\alpha)$, which is such that $\Omega_h (\alpha)$, which is constant over $[h,1-h]$ for $h$ small enough, but depends upon $\alpha$ and $h$ otherwise on $[0,h]$ and $[1-h,1]$, another example of boundary effects.
\subparagraph{Sample Hessian and score function.}
The two next Lemmas study the first and second derivatives of $\widehat{\mathsf{R}}\left( \mathsf{\cdot};\alpha,I\right) $ in a shrinking vicinity of $\mathsf{b}\left( \alpha|I\right) $. In particular, Lemma (ref) implies that $\widehat{\mathsf{R}}\left( \mathsf{\cdot} ;\alpha,I\right) $ is strictly convex over such a vicinity with a probability tending to $1$.
Since $\overline{\mathsf{R}}^{\left( 1\right) }\left( \overline{\mathsf{b} }\left( \alpha|I\right) ;\alpha,I\right) =0$ and assuming $\sup_{\alpha\in\left[ 0,1\right] }\left\Vert \overline{\mathsf{b}}\left( \alpha|I\right) -\mathsf{b}\left( \alpha|I\right) \right\Vert =o\left( ^{D_{\mathcal{M}}/2+1}\right) $ as established in ((ref)), it holds that \[ \max_{\alpha\in\left[ 0,1\right] }\left\Vert \frac{\widehat{\mathsf{R} }^{\left( 1\right) }\left( \overline{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right) }{\left( h+\alpha\left( 1-\alpha\right) \right) ^{1/2} }\right\Vert_{\infty} =O_{\mathbb{P}}\left( \left( \frac{\log L}{L}\right) ^{1/2}\right) . \] The normalization factor $h+\alpha\left( 1-\alpha\right)$ gives that $ \widehat{\mathsf{R}}^{\left( 1\right) } \left( \overline{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right) $ has the fast rate $\left(\frac{h \log L}{L} \right)^{1/2}$ for boundary $\alpha$ in $[0,h]$ and $[1-h,1]$, while it is of order $\left( \frac{\log L}{L} \right)^{1/2}$ for central quantile levels.
Lemmas (ref) and (ref) implies that the second derivative $\widehat{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b};\alpha,I\right)$ is well-conditioned for $\mathsf{b}$ in a vicinity of $\mathsf{b} (\alpha|I)$, implying local strict convexity of the sample objective function for all $\alpha$ in $[0,1]$ with a probability tending to 1. This is used later on to show existence and uniqueness of the AQR and ASQR for all $\alpha$ including extreme quantile levels, with a probability tending to $1$.
\subparagraph{Bahadur leading term.} The next Lemma studies the leading term $\widehat{\mathsf{e}}\left( \alpha|I\right) $ of $\widehat{\mathsf{b}}\left( \alpha|I\right) -\overline{\mathsf{b}}\left( \alpha|I\right) $, \[ \widehat{\mathsf{e}}\left( \alpha|I\right) =-\left[ \overline{\mathsf{R} }^{\left( 2\right) }\left( \overline{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right) \right] ^{-1}\widehat{\mathsf{R}}^{\left( 1\right) }\left( \overline{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right) \] see Theorem (ref) below. Note that $\overline{\mathsf{R}}^{\left( 2\right) }\left( \overline{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right) $ has, by Lemma (ref), an inverse provided $\sup_{\alpha \in\left[ 0,1\right] }\left\Vert \overline{\mathsf{b}}\left( \alpha |I\right) -\mathsf{b}\left( \alpha|I\right) \right\Vert =o\left( h^{s+1+D_{\mathcal{M}}/2}\right)=o\left(h^{D_{\mathcal{M}}/2+1}\right) $ as therefore assumed and established in the proof of Theorem (ref) below, see ((ref)). Let $\mathbf{P}\left( I\right)$ and $\mathbf{P}_0\left( \alpha|I\right)$ be as for Lemma (ref) and recall $\Sigma_h (\alpha|I)=\alpha^2 v_{h}^{2}\left( \alpha\right) \mathbf{P}_0 (\alpha|I)^{-1} \mathbf{P}\left( I\right) \mathbf{P}_0 (\alpha|I)^{-1}/(I-1)$.
As explained now, Lemma (ref) improves Lemma (ref), which gives taking $\mathsf{b} = \overline{\mathsf{b}} (\alpha|I)$
by Lemmas (ref) and (ref), $\| A b \|_{\infty} \leq \|A \|_{\infty} \| b \|_{\infty}$ for a matrix $A$ and a conformable $b$. This implies that $\widehat{e}_1 (\cdot|I) = \mathsf{S}_{1}\widehat{e} (\cdot|I)/h$ is of order $\frac{1}{h} (\log L /L)^{1/2}$ in $\| \cdot \|_{\infty}$ norm. Lemma (ref) improves this order to $\frac{1}{h^{1/2}} (\log L /L)^{1/2}$. This could be seen here noting that the variance of $\widehat{e}_1 (\alpha|I)$ is of order $1/(Lh)$ instead of $1/(Lh^2)$.
A precise evaluation of this variance is the key tool to obtain this improvement, which is important to relate the estimation rate of $\beta^{(1)} (\cdot)$ with the nonparametric density estimation rate. To see this, consider the AQR case, ie $D_{\mathcal{M}}=0$ and $P(x)=[1,x]^{\prime}$. The leading term of the estimation error for $\beta(\alpha|I)$ is, by ((ref)) $\mathsf{e}_{0}\left( \alpha|I\right)$, which has a parametric rate $1/\sqrt{L}$ by Lemma (ref)-(i). The leading term of the estimation error of $\beta^{(1)}(\alpha|I)$ is $\mathsf{e}_{1}\left( \alpha|I\right)/h$ is of order $1/\sqrt{Lh}$, as it would hold for a kernel density estimator with bandwidth $h$.
\setcounter{section}{2}
\setcounter{section}{0} \setcounter{footnote}{0} \setcounter{equation}{0} \setcounter{theorem}{0}
The study of the bias $\overline{V}\left( \alpha|X,I\right) -V\left( \alpha|X,I\right) $ and $\overline{B}\left( \alpha|X,I\right) -B\left( \alpha|X,I\right) $ is based on the following Lemma which is a consequence of the Kantorovitch-Newton Theorem, see e.g. Gragg and Tapia (1974).
The next lemma gathers results for intermediary bias terms. Define, $\mathbf{P}_0 (\alpha|I)$ being as for Lemma (ref),
Recall $\mathsf{S}_0 = S_0 \otimes \mathrm{Id}_{N}$ and $\mathsf{S}_1 = S_1 \otimes \mathrm{Id}_{N}$, where the row vectors $S_0 = [1,0,\ldots,0]$ and $S_1 = [0,1,0,\ldots,0]$ have dimension $s+2$, see ((ref)).
Lemma (ref) will be used in the proof of Theorem (ref) after having established ((ref)). Lemma (ref)-(i) shows that the approximation error $P(x)^{\prime} \beta (\alpha|I) - B(\alpha|x,I)$, which is $o(h^{s+1})$ by Proposition (ref)-(ii), does not contribute to the bias for the estimation of $\alpha B^{(1)} (\alpha|x,I)$ as $P(x)^{\prime} \mathsf{S}_1 \overline{\mathfrak{bias}}_{h} (\alpha|I)$ has the better order $o(h^{s+2})$. The proof of Lemma (ref) is given at the end of this Appendix, after the proof of Theorem (ref).
Theorem (ref) gives a bias expansion for $\widehat{V} (\alpha|x,I)$ and $\alpha \widehat{B}^{(1)} (\alpha|x,I)$, and the order of the bias of $\widehat{B} (\alpha|x,I)$. It shows that the leading term of the bias of $\widehat{V} (\alpha|x,I)$ is the one of $\alpha \widehat{B}^{(1)} (\alpha|x,I)$ for central quantile levels, as the leading term of the bias of $\alpha \widehat{B}^{(1)} (\alpha|x,I)/(I-1)$ vanishes for $\alpha =0$ by ((ref)) and $\lim_{\alpha \downarrow 0} \alpha B^{(s+2)} (\alpha|x,I)=0$ by Proposition (ref)-(i). It follows that the bias of $\widehat{V} (\alpha|x,I)$ is expected to be smaller for quantile levels close to $0$ than for central or upper ones.
The proof of Theorem (ref) is given after discussing some properties of the ASQR estimator.
\subparagraph{Local polynomial and sieve bias.} The bias $\overline{V} (\cdot|\cdot) - V(\cdot|\cdot)$ is given by the bias of the bid quantile derivative, which has two components: one which is due to the sieve procedure and a second induced by local polynomials. That the AQR and ASQR have the same asymptotic bias suggests that the bias is due to the local polynomial method. This can be seen noticing that the sieve component of the bias is due to the item $P(x)^{\prime} \alpha \beta^{(1)} (\alpha|I) - \alpha B^{(1)} (\alpha|I)$, which is $o(h^{s+1})$ by Proposition (ref)-(iii) and is negligible compared to the Bias leading term $h^{s+1} \mathsf{Bias}_{h}\left( \alpha|x,I\right)$. The expression of $\mathsf{Bias} _{h}\left( \alpha|x,I\right)$ given in ((ref)) depends upon the local polynomial vector $\pi(t)$ and upon the kernel, which are associated to the local polynomial component of the ASQR estimator.
\subparagraph{Asymptotic uniqueness of the estimator.} The proof of Theorem (ref) establishes that $\sup_{\alpha\in\left[ 0,1\right] }\left\Vert \overline{\mathsf{b}}\left( \alpha|I\right) -\mathsf{b}\left( \alpha|I\right) \right\Vert_{\infty} =o\left( h^{s+1+D_{\mathcal{M}}/2}\right) $, see ((ref)) below. Hence by a second order Taylor expansion, Lemmas (ref) and (ref), it holds
Then, as the eigenvalues of $\overline{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b} \left( \alpha|I\right) ;\alpha,I\right)$ stay bounded away from $0$ and infinity by Lemma (ref)-(ii), the linear approximation of the sample objective function $\widehat{\mathsf{R}}\left( \mathsf{b} ;\alpha,I\right) -\widehat{\mathsf{R}}\left( \overline{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right)$, \[ \left( \mathsf{b-}\overline{\mathsf{b}}\left( \alpha|I\right) \right)^{\prime} \widehat{\mathsf{R}}^{\left(1\right) } \left( \overline{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right) + \frac{1}{2} \left( \mathsf{b-}\overline{\mathsf{b}}\left( \alpha|I\right) \right) ^{\prime }\overline{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b} \left( \alpha|I\right) ;\alpha,I\right) \left( \mathsf{b-}\overline {\mathsf{b}}\left( \alpha|I\right) \right) \] has a unique minimizer $\overline{\mathsf{b}}^{\ast} (\alpha|x,I)$, \[ \overline{\mathsf{b}}^{\ast} (\alpha|x,I) = \overline{\mathsf{b}} (\alpha|x,I) - \left[ \overline{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b} \left( \alpha|I\right) ;\alpha,I\right) \right]^{-1} \widehat{\mathsf{R}}^{\left(1\right) } \left( \overline{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right), \] which belongs to $\mathcal{B}_{\infty} \left(\overline{\mathsf{b}}\left( \alpha|I\right) ,C_0h^{D_{\mathcal{M}}/2+1}\right)$ for all $\alpha$ in $[0,1]$ with a probability tending to 1 by Lemmas (ref) and (ref) as $(\log L/L)^{1/2} = o \left(h^{D_{\mathcal{M}}/2+1}\right)$ under Assumption (ref). Then, since $\widehat{\mathsf{R}}\left( \mathsf{\cdot};\alpha,I\right) $ is also strictly convex over $\mathcal{B}_{\infty} \left(\overline{\mathsf{b}}\left( \alpha|I\right) ,C_0h^{D_{\mathcal{M}}/2+1}\right)$ for all $\alpha$ in $[0,1]$ with a probability tending to 1, its minimizer $\widehat{\mathsf{b}}^{\ast} (\alpha|I)=\arg\min_{\mathsf{b} \in \mathcal{B}_{\infty} \left(\overline{\mathsf{b}}\left( \alpha|I\right) ,C_0h^{D_{\mathcal{M}}/2+1}\right)} \widehat{\mathsf{R}}\left( \mathsf{b};\alpha,I\right)$ and by the Argmax Theorem 2.1 in Newey and McFadden (1994)\footnote{It is easy to extend this result to the case where the objective functions depend upon a parameter $\alpha$.} is unique and satisfies \[ \sup_{ \alpha \in [0,1] } \left\| \widehat{\mathsf{b}}^{\ast} (\alpha|I) - \overline{\mathsf{b}} (\alpha|I) \right\|_{\infty} = o_{\mathbb{P}} \left(h^{D_{\mathcal{M}}/2+1}\right). \] Hence $\widehat{\mathsf{b}}^{\ast} (\alpha|I)$ is an interior point of the considered ball, over which the sample objective function is strictly convex. Arguing as in Footnote (ref) implies \[ \lim_{L\uparrow\infty} \mathbb{P} \left( \widehat{\mathsf{b}} (\alpha|I) = \widehat{\mathsf{b}}^{\ast} (\alpha|I) \text{ for all $\alpha \in [0,1]$}\right)=1. \] It follows that $\widehat{\mathsf{b}} (\alpha|I)$ is unique for all $\alpha$ of $[0,1]$ with a probability tending to 1.
\subparagraph{Asymptotic smoothness of $\widehat{\mathsf{b}}(\cdot|I)$.} As $\widehat{\mathsf{b}}(\alpha|I)$ belongs to $\mathcal{B}_{\infty} \left(\overline{\mathsf{b}}\left( \alpha|I\right) ,C_0h^{D_{\mathcal{M}}/2+1}\right)$ for all $\alpha$ in $[0,1]$ with a probability tending to 1, Lemmas (ref)-(i), (ref) and (ref) allow to apply the Implicit Function Theorem to the FOC $\widehat{\mathsf{R}}^{\left(1\right) } \left( \widehat{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right)=0$. As noted in Fernandes, Guerre and Horta (2019), it then follows that $\widehat{\mathsf{b}}(\cdot|I)$ is differentiable with a probability tending to $1$ with derivative \[ \widehat{\mathsf{b}}^{(1)}(\alpha|I) = - \left[ \widehat{\mathsf{R}}^{\left(2\right) } \left( \widehat{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right) \right]^{-1} \partial_{\alpha} \widehat{\mathsf{R}}^{\left(1\right) } \left( \mathsf{b} ;\alpha,I\right) \bigg|_{\mathsf{b}=\widehat{\mathsf{b}}\left( \alpha|I\right)} \] where, by ((ref)) \[ \partial_{\alpha} \widehat{\mathsf{R}}^{\left(1\right) } \left( \mathsf{b} ;\alpha,I\right) = - \frac{1}{L} \sum_{\ell=1}^{L} \mathbb{I} \left( I_{\ell} = I \right) \int_{-\frac{\alpha}{h}}^{\frac{1-\alpha}{h}} P \left( X_{\ell},t \right) K(t) dt. \] As a consequence, $\widehat{\beta}_0 (\cdot|I)$ and $B(\cdot|x,I)$ are differentiable with a probability tending to $1$, contrasting with the standard quantile-regression estimator, as also illustrated in Figure (ref) in the application Section. If the derivative of $B(\cdot|x,I)$ can be used as an alternative estimator of $B^{(1)} (\cdot|x,I)$ as in Fernandes et al. (2019) deserves further work.
That $\sup_{\left( \alpha,x,I\right) \in\left[ 0,1\right] \times \mathcal{X\times I}}\left| \alpha \mathsf{Bias} _{h}\left( \alpha|x,I\right)\right| =O\left( 1\right) $ follows from ((ref)) and Proposition (ref)-(i).
\subparagraph{Step 1.} The first step of the proof works by establishing that there is a solution of the first-order condition in a open ball where $\overline{\mathsf{R}}\left( \mathsf{b} ;\alpha,I\right) $ is strictly convex by checking the conditions of Lemma (ref), which will also gives the rate stated in the Theorem and the uniqueness of $\overline{\mathsf{b}}\left( \alpha|I\right) $. Observe first that the definition of $ \overline{\mathsf{R}}^{\left( 1\right) }\left( \mathsf{b}\left( \alpha|I\right) ;\alpha,I\right) $ gives, by the Taylor inequality as $\alpha+ht = G \left(\left. B(\alpha+ht|x,I) \right| x,I\right)$,
where $\epsilon_{L}=o\left( h^{s+1+D_{\mathcal{M}}/2}\right) $ follows from Lemma (ref)-(iii) and Assumptions (ref)-(i) and (ref) which gives $\max_{1\leq n \leq N} \mathbb{E} \left[ \left| P_n (X) \right| \right] \leq C \max_{1\leq n \leq N} \int_{\mathcal{X}} |P_n(x)| dx = O (h^{D_{\mathcal{M}}/2}) $. The second part of Condition (i) in Lemma (ref) follows from Lemma (ref)-(ii) and the norm equivalence ((ref)), which ensures that there is a $C_{0}>0$ such that, for all $h$ small enough, \[ \sup_{\left( \alpha,I\right) \in\left[ 0,1\right] \times\mathcal{I} }\left\Vert \left[ \overline{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b}\left( \alpha|I\right) ;\alpha,I\right) \right] ^{-1}\right\Vert_{\infty} \leq C_{0}. \] Condition (ii) in Lemma (ref) follows from Lemma (ref)-(i), which ensures that for $C_{1L}=O\left( h^{-D_{\mathcal{M}}/2-1}\right) $, \[ \left\Vert \overline{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b}^{1};\alpha,I\right) - \overline{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b}^{0};\alpha,I\right) \right\Vert_{\infty} \leq C_{1L}\left\Vert \mathsf{b}^{1}-\mathsf{b}^{0}\right\Vert_{\infty} \] for all $\mathsf{b}^{1}$, $\mathsf{b}^{0}$ in $\mathcal{B}_{\infty} \left( \mathsf{b}\left( \alpha|I\right) ,2C_{0}\epsilon_{L}\right) $ and all $\alpha$, $I$. For condition (iii) in Lemma (ref), $\epsilon_{L}=o\left( h^{s+1+D_{\mathcal{M}}/2}\right) $ implies $C_{0}^{2}C_{1L}\epsilon_{L}=o\left( h^{s}\right) =o\left( 1\right) <1/2$ for $h$ small enough. Hence Lemma (ref) ensures that, for $h$ small enough, all $\alpha$ and all $I$, there is a unique $\overline{\mathsf{b}}\left( \alpha|I\right) $ in $\mathcal{B}_{\infty} \left( \mathsf{b}\left( \alpha|I\right) ,2C_{0}\epsilon_{L}\right) $ such that \[ \overline{\mathsf{R}}^{\left( 1\right) }\left( \overline{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right) =0 \] and is therefore the unique minimizer of $\overline{\mathsf{R}}\left( \cdot;\alpha,I\right) $ over $\mathcal{B}_{\infty} \left( \mathsf{b}\left( \alpha|I\right) ,2C_{0}\epsilon_{L}\right) $. Since the convex function $\overline{\mathsf{R}}\left( \cdot;\alpha,I\right) $ cannot have distinct separated minimizers as discussed in Footnote (ref), $\overline{\mathsf{b}}\left( \alpha|I\right) $ is also the unique global minimizer of $\overline{\mathsf{R}}\left( \cdot;\alpha ,I\right) $. Since $\epsilon_{L}=o\left( h^{s+1+D_{\mathcal{M}}/2}\right) $, it follows that
Note that this together with Proposition (ref)-(ii) gives, as $\mathsf{S}_0 \mathsf{b}\left( \alpha|I\right) = \beta (\alpha|I)$ and $\mathsf{S}_1 \mathsf{b}\left( \alpha|I\right) = h \beta^{(1)} (\alpha|I)$, $\overline{B} (\alpha|x,I) = P(x)^{\prime} \mathsf{S}_0 \mathsf{b}\left( \alpha|I\right)$ and $\overline{B}^{(1)} (\alpha|x,I) = P(x)^{\prime} \mathsf{S}_1 \mathsf{b}\left( \alpha|I\right)/h$,
so that the results for $B(\cdot|\cdot)$ and $B^{(1)} (\cdot|\cdot)$ are now established.
\subparagraph{Step 2.} For $V(\alpha|x,I)$ and $\alpha B^{(1)} (\alpha|x,I)$, we shall obtain a more accurate expansion of $\alpha\overline{\mathsf{b}}\left( \alpha|I\right) -\alpha\mathsf{b}\left( \alpha|I\right) $. Observe that, for $\overline{g} (\alpha|t,x,I)$ as in Lemma (ref),
The first-order condition ((ref)) then gives
where, for $\check{\mathsf{R}}^{(2)} \left(\alpha|I\right)$ as in Lemma (ref),
For the other integral, the definition of $\Psi (t|x,\mathsf{b})$ gives, as $\mathsf{b}_0 (\alpha|I) = \beta(\alpha|I)$,
by Lemma (ref)-(iii) with a $o\left(h^{s+2}\right)$ which is uniform in $\alpha$ and $x$. This gives, for $\check{b}_{s+2} (\alpha|I)$ and $\overline{\mathfrak{bias}}_{h} (\alpha|I)$ as in Lemma (ref),
Hence the first-order condition ((ref)) gives
by Assumption (ref)-(i) and Lemma (ref)-(i).
As $\overline{B}^{(1)} (\alpha|x,I) = P(x)^{\prime} \mathsf{S}_1 \mathsf{b}\left( \alpha|I\right)/h$ with $\sup_{(\alpha,x,I) \in [0,1] \times \mathcal{X} \times \mathcal{I}} \left| P(x)^{\prime} \mathsf{S}_1 \overline{\mathfrak{bias}}_{h} (\alpha|I)/h \right| = o(h^{s+1}) $ and $\sup_{(\alpha,x,I) \in [0,1] \times \mathcal{X} \times \mathcal{I}} \left| P(x)^{\prime} \mathsf{S}_1 \check{b}_{s+2} (\alpha|I) - (I-1) \mathsf{Bias}_h (\alpha|x,I) \right|=o(1)$ by Lemma (ref), Proposition (ref)-(iii) gives
As $\overline{V} (\alpha|x,I) = \overline{B} (\alpha|x,I) + \alpha \overline{B}^{(1)} (\alpha|x,I)/(I-1)$ with $ \sup_{(\alpha,x,I) \in [0,1] \times \mathcal{X} \times \mathcal{I}} \left| \overline{B} (\alpha|x,I) - B(\alpha|x,I) \right| = o(h^{s+1})$, this also establishes the claimed result for $V(\cdot|\cdot)$. $\Box$
\subparagraph{Proof of (i).} Set
Define, for $\Omega_h (\alpha)$ as in Lemma (ref) and $S_{0} =[1,0,\ldots,0]^{\prime}$, so that $\Omega_h (\alpha) S_{0} = \int_{\underline{t}_{\alpha,h}}^{\overline{t}_{\alpha,h}} \pi (t) K(t) dt$,
so that $ \overline{\mathfrak{bias}}_{h} (\alpha|I) = \left[\check{\mathsf{R}}^{(2)} \left(\alpha|I\right)\right]^{-1} \overline{\mathsf{r}} (\alpha|I) $ and $ \mathfrak{bias}_{h} (\alpha|I) = \left[ \Omega_h (\alpha) \otimes \mathbf{P}_0 (\alpha|I) \right]^{-1} \overline{\mathsf{r}} (\alpha|I) $.
Under ((ref)) $ \max_{(\alpha,x,I) \in [0,1] \times \mathcal{I} \times \mathcal{I}} \max_{t \in \mathcal{T}_{\alpha,h}} \left| \Psi\left( t|x,\overline{\mathsf{b}}\left( \alpha|I\right) \right) - B\left(\alpha+ht|x,I\right) \right| = o \left(h^{s+1}\right) = o(h^2) $ by Lemma (ref)-(iii,iv) and $s\geq 1$ by Assumption (ref). Set
which therefore satisfies $\epsilon_{BL}=o(h^2)$. It follows that, for $C$ large enough, all $u \in [0,1]$ and all $t$ in $ \left[ \underline{t}_{\alpha,h} + C \epsilon_{BL}/h , \overline{t}_{\alpha,h} - C \epsilon_{BL}/h \right] $, \[ \Psi\left( t|x,\overline{\mathsf{b}}\left( \alpha|I\right) \right) +u\left( B\left( \alpha+ht|x,I\right) -\Psi\left( t|x,\mathsf{b}\left( \alpha |I\right) \right) \right) \in \left[ B(0|x,I),B(1|x,I) \right]. \] Hence since $g(\cdot|x,I)$ is differentiable and $g\left( \left. B(\alpha|x,I) \right|x,I\right)=1/B^{(1)}(\alpha|x,I)$
It follows that, as $P(x,t)=\pi (t) \otimes P(x)$ and for $\Omega_h (\alpha)$ as for Lemma (ref),
where the latter follows from ((ref)), $\max_{(\alpha,x) \in [0,1] \times \mathcal{X}} |P(x)^{\prime}\beta(\alpha)-B(\alpha|x,I)|=O(h^{s+1})$ by Proposition (ref)-(ii), and $\max_{1\leq n \leq N} \mathbb{E} \left[ |P_n(X)| \right]=O(h^{D_{\mathcal{M}}/2})$ by Assumptions (ref)-(i) and (ref). As $ \check{\mathsf{R}}^{(2)} \left(\alpha|I\right) - \Omega_h (\alpha|I) \otimes \mathbf{P}_0 (\alpha|I)$ is an index-permuted $c(s+2)/2$-band matrix, and since $\Omega_h (\alpha|I) \otimes \mathbf{P}_0 (\alpha|I)$ which has an inverse with a $\|\cdot\|_{2}$ norm bounded away from infinity by Assumption (ref)-(i) and definition of $\Omega_h (\alpha|I)$, it holds\footnote{The details, which are standard, are as follows. Set $ \mathcal{O}_h = \check{\mathsf{R}}^{(2)} \left(\alpha|I\right) - \Omega_h (\alpha|I) \otimes \mathbf{P}_0 (\alpha|I) $ which is such that $\|\mathcal{O}_h \|_{2}$ and $\|\mathcal{O}_h \|_{\infty}$ are $O(h)$ by ((ref)). Since $\left\|\left[\Omega_h (\alpha|I) \otimes \mathbf{P}_0 (\alpha|I) \right]^{-1} \right\|_{2}$ stays bounded away from infinity,
where the series converges in the $\|\cdot\|_{2}$ sense for all $h$ small enough. Now Lemma (ref) gives, for a constant $C$ independent of $h$, since the eigenvalues of $\int_{\mathcal{X}}P(x)P(x)^{\prime} dx$ and $f(\cdot,I)$ are bounded away from $0$ and infinity,
} \[ \max_{\alpha \in [0,1]} \left\| \left[ \check{\mathsf{R}}^{(2)} \left(\alpha|I\right) \right]^{-1} - \Omega_h (\alpha|I)^{-1} \otimes \mathbf{P}_0 (\alpha|I)^{-1} \right\|_{\infty} = O(h). \] This implies $\sup_{(\alpha,x,I) \in [0,1] \times \mathcal{X} \times \mathcal{I}} \left\| \left[\check{\mathsf{R}}^{(2)} \left(\alpha|I\right)\right]^{-1} \right\|_{\infty} =O(1) $ by Lemma (ref), Assumptions (ref)-(i) and (ref). As $ \max_{\alpha \in [0,1]} \left\| \overline{\mathsf{r}}(\alpha|I) \right\|_{\infty} = o \left( h^{s+1+D_{\mathcal{M}}/2} \right) $, $ \overline{\mathfrak{bias}}_{h} (\alpha|I) = \left[\check{\mathsf{R}}^{(2)} \left(\alpha|I\right)\right]^{-1} \overline{\mathsf{r}} (\alpha|I) $ and $ \mathfrak{bias}_{h} (\alpha|I) = \left[ \Omega_h (\alpha) \otimes \mathbf{P}_0 (\alpha|I) \right]^{-1} \overline{\mathsf{r}} (\alpha|I) $, it follows
Hence, for $\mathsf{S}=\mathsf{S}_0$ or $\mathsf{S}_1$,
This gives the results of the Lemma since $\sup_{(\alpha,x,I) \in [0,1] \times \mathcal{X} \times \mathcal{I}} \left| P(x)^{\prime} \mathsf{S}_0 \mathfrak{bias}_{h} (\alpha|I) \right| = o(h^{s+1})$ and $\mathsf{S}_1 \mathfrak{bias}_{h} (\alpha|I)=0$.
\subparagraph{Proof of (ii).} By the Approximation Property S and since $(\alpha,x) \in [0,1] \times \mathcal{X} \mapsto \alpha B^{(s+2)} (\alpha|x,I)$ is continuous by Proposition (ref)-(i), there exists for each $\alpha$ of $[0,1]$ a $\beta^{\ast}(\alpha|I)$ satisfying \[ \sup_{(\alpha,x) \in [0,1] \times \mathcal{X}} \left| P(x)^{\prime} \beta^{\ast} (\alpha|I) - \alpha B^{(s+2)} \left(\alpha|x,I\right) \right| = o(1). \] Define \[ \beta^{\ast}_h (\alpha|I) = \beta^{\ast} (\alpha|I) S_1 \Omega_h (\alpha)^{-1} \int_{-\frac{\alpha}{h}}^{\frac{1-\alpha}{h}} \frac{t^{s+2}}{(s+2)!} K(t) dt \] so that $ \sup_{(\alpha,x) \in [0,1] \times \mathcal{X}} \left| P(x)^{\prime} \beta^{\ast}_h (\alpha|I) - (I-1) \mathsf{Bias}_h \left(\alpha|x,I\right) \right| = o(1) $. Define also
which is such that $\mathsf{S}_1 b^{\ast}_h (\alpha|I) = \beta^{\ast}_h (\alpha|I)$. Now, arguing as above gives
Hence $ \sup_{\alpha \in [0,1]} \left\| \check{b}_{s+2} (\alpha|I) - b^{\ast}_h (\alpha|I) \right\|_{\infty} = o \left( h^{D_{\mathcal{M}}/2} \right) $ which gives uniformly in $\alpha$ and $x$
This ends the proof of the Lemma. $\Box$
\setcounter{section}{3}
\setcounter{section}{0} \setcounter{footnote}{0} \setcounter{equation}{0} \setcounter{theorem}{0}
Let $\widehat{\mathsf{e}}\left( \alpha|I\right) $ be a candidate linearization leading term for $\widehat{\mathsf{b}}\left( \alpha|I\right) -\overline{\mathsf{b}}\left( \alpha|I\right) $ and $\widehat{\mathsf{d} }\left( \alpha|I\right) $ the associate linearization error term, or Bahadur remainder term,
This section goal is to study the order of $\widehat{\mathsf{d}}\left( \alpha|I\right) $ and of the remainder term $P^{\prime}\left( x\right) \widehat{\mathsf{d}}_{0}\left( \alpha|I\right) $ and $P^{\prime }\left( x\right) \widehat{\mathsf{d}}_{1}\left( \alpha|I\right) /h$, which are the differences of $\widehat{B} (\alpha|x,I)$ and $\widehat{B}^{(1)} (\alpha|x,I)$ to their respective linear expansions $P(x)^{\prime}\widehat{\mathsf{e}}_0 (\alpha|I)$ and $P(x)^{\prime}\widehat{\mathsf{e}}_1 (\alpha|I)/h$, with $\widehat{\mathsf{e}}_0 (\alpha|I) = \mathsf{S}_0 \widehat{\mathsf{e}} (\alpha|I)$ and $\widehat{\mathsf{e}}_1 (\alpha|I) = \mathsf{S}_1 \widehat{\mathsf{e}} (\alpha|I)$, $\mathsf{S}_0$ and $\mathsf{S}_1$ as in ((ref)).
The order $\left( \frac{\log^2 L}{L h^{3 D_{\mathcal{M}}+2}} \right)^{1/2}$ obtained for $\left(L h^{D_{\mathcal{M}}+1} \right)^{1/2} \sup_{(\alpha,x) \in [0,1] \times \mathcal{X}} \left| \widehat{B}^{(1)} (\alpha|x,I) - \frac{P(x)^{\prime} \widehat{\mathsf{e}}_{1} (\alpha|I)}{h} \right| $ is the reason of the bandwidth condition in Assumption (ref). It holds when $\frac{\log^2 L}{Lh^2} = o(1)$ for the AQR case ($D_{\mathcal{M}}=0$), see also Assumption (ref). For a purely local-polynomial version of the ASQR conditional quantile estimator, Guerre and Sabbah (2017) allow for a better condition $\log L/(Lh^{D+1}) = o(1)$, where $D$ plays the role of $D_{\mathcal{M}}$.
As Theorem (ref) and Lemma (ref) gives, for $\widehat{\mathsf{d}} (\alpha|I)$ as in ((ref)),
and the order for $\widehat{\mathsf{e}} (\alpha|I)$ is, up to a $\log^{1/2} L$ term, sharp as its variance is proportional to $1/L$, $\widehat{\mathsf{e}} (\alpha|I)$ is the leading term of $\widehat{\mathsf{b}} (\alpha|I)-\overline{\mathsf{b}} (\alpha|I)$. Note that all these rates above are driven by central quantiles and can be improved for extreme $\alpha$ with $\alpha = O(h)$ or $\alpha=1-O(h)$.
As $\frac{\log L}{Lh^{ D_{\mathcal{M}}+2} } =o(1)$ by Assumption (ref), it follows that $\left( \frac{\log L}{L} \right)^{1/2} = o \left( h^{D_{\mathcal{M}}/2+1} \right) $, Lemma (ref), Theorem (ref) and ((ref)) show that \[ \sup_{ \alpha \in [0,1] } \left\| \widehat{\mathsf{b}} (\alpha|I)-\mathsf{b} (\alpha|I) \right\|_{\infty} = o \left( h^{s+1+D_{\mathcal{M}}/2} \right) + O_{\mathbb{P}} \left( \left( \frac{\log L}{L} \right)^{1/2} \right) = o_{\mathbb{P}} \left( h^{D_{\mathcal{M}}/2+1} \right), \] so that Lemma (ref)-(i) shows that, uniformly in $\alpha$, the ASQR objective function $\widehat{\mathsf{R}}(\cdot;\alpha,I)$ is twice continuously differentiable in a vicinity $\mathcal{B}_{\infty}\left(\widehat{\mathsf{b}} (\alpha|I),Ch^{D_{\mathcal{M}}/2+1} \right)$, with a probability tending to $1$. In addition of simplifying the proof of Theorem (ref) compared to Guerre and Sabbah (2012, 2014), it can have practical implications for the numerical computation of the estimator as well as for inference.
\paragraph{Proof of Theorem (ref). } Let $\widehat{\mathsf{d}} (\alpha|I)$ be as in ((ref)). We first introduce some normalizations. Let, for $\widehat{\mathsf{e}}\left( \alpha|I\right) $ as in ((ref)),
The definition of $\widehat{\boldsymbol{R}}\left( \mathsf{d};\alpha,I\right)$ ensures that \[ \frac{\widehat{\mathsf{d}}\left( \alpha|I\right) }{\varrho_{\alpha L}} =\arg\min_{\mathsf{d}}\widehat{\boldsymbol{R}}\left( \mathsf{d};\alpha,I\right) . \] It follows that
since $\inf_{\left\Vert \mathsf{d}\right\Vert_{\infty} \leq t}\widehat{\boldsymbol{R}}\left( \mathsf{d};\alpha,I\right) \leq\widehat{\boldsymbol{R}}\left( 0;\alpha,I\right) =0$. The next step uses a convexity argument that can be found in Pollard (1991). For any $\mathsf{d}$ with $\left\Vert \mathsf{d}\right\Vert_{\infty} \geq t$, convexity yields
so that $\inf_{\left\Vert \mathsf{d}\right\Vert_{\infty} \geq t}\widehat{\boldsymbol{R}}\left( \mathsf{d};\alpha,I\right) \leq0$ implies $\inf_{\left\Vert \mathsf{d}\right\Vert_{\infty} =t}\widehat{\boldsymbol{R}}\left( \mathsf{d};\alpha,I\right) \leq0$ and then
Thus it is sufficient to consider those $\mathsf{d}$ with $\left\Vert \mathsf{d}\right\Vert_{\infty} =t$, as done from now on.
Note that, for any $t>0$, $\left(\frac{\log L}{L}\right)^{1/2} + t \varrho_{\alpha L} \leq 2 \left(\frac{\log L}{L}\right)^{1/2} = o\left(h^{D_{\mathcal{M}}/2+1}\right)$. It follows that $ \lim_{L\uparrow \infty} \mathbb{P} \left( \max_{\alpha \in [0,1]} \max_{\mathsf{d}: \|\mathsf{d}\|_{\infty}=t} \left\| \widehat{\mathsf{e}} (\alpha|I) + \varrho_{\alpha L} \mathsf{d} \right\|_{\infty} \leq C_0 h^{D_{\mathcal{M}}/2+1} \right) =1 $ for $C_0$ as in Lemma (ref)-(i), ensuring with ((ref)) that $\mathsf{d} \mapsto \widehat{\boldsymbol{R}}\left( \mathsf{d};\alpha,I\right)$ is twice continuously differentiable for $L$ large enough with a large probability.
Under this condition and by ((ref)), the expression of $\widehat{\boldsymbol{R}}\left( \mathsf{d};\alpha,I\right) $ gives for all $\mathsf{d}$ with $\|\mathsf{d}\|_{\infty}=t$, with a probability tending to $1$ for any $t$,
Since $\widehat{\mathsf{R}}^{\left( 1\right) }\left( \overline{\mathsf{b} }\left( \alpha|I\right) ;\alpha,I\right) +\widehat{\mathsf{R}}^{\left( 2\right) }\left( \overline{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right) \widehat{\mathsf{e}}\left( \alpha|I\right) =0$ by ((ref)), it follows that
We now give an upper bound for $\widehat{\boldsymbol{R}}_1 \left( \mathsf{d};\alpha,I\right)$ and a lower one for $\widehat{\boldsymbol{R}}_{2} \left( \mathsf{d};\alpha,I\right)$.
\subparagraph{Upper bound for $\left| \widehat{\boldsymbol{R}}_1 \left( \mathsf{d};\alpha,I\right) \right|$.} The Cauchy-Schwarz inequality, the norm equivalences ((ref)) and ((ref)) under the disjoint support property of Assumption (ref)-(ii) give, for all $\alpha$ of $[0,1]$ and $\mathsf{d}$ with $\| \mathsf{d} \|_{\infty} = t$,
Observe that, by Assumption (ref)-(i) and ((ref)), \[ \max_{\alpha \in [0,1]} \frac{ (s+2)N \| \mathsf{d} \|_{\infty} \left\| \widehat{\mathsf{e}} (\alpha|I) \right\|_{\infty}}{ \left(h+\alpha (1-\alpha) \right)^{1/2} } = t \cdot h^{-D_{\mathcal{M}}} O_{\mathbb{P}} \left( \left( \frac{\log L}{L} \right)^{1/2} \right). \] For the matrix norm, it holds by ((ref)), Lemmas (ref) and (ref)-(i) with ((ref)),
Combining this bound then gives, uniformly in $\alpha$ in $[0,1]$, $t>0$ and $\mathsf{d}$ with $\|\mathsf{d}\|_{\infty}=t$,
\subparagraph{Lower bound for $\widehat{\boldsymbol{R}}_{2} \left( \mathsf{d};\alpha,I\right)$.} Note that there is a diverging $t_{L}>0$ such that \[ t_{L} \max_{\alpha \in [0,1]} \varrho_{\alpha L} \asymp t_{L} \frac{\log L}{L h^{3 D_{\mathcal{M}}/2+1/2}} = o \left( h^{D_{\mathcal{M}}/2+1} \right). \] Lemma (ref), the matrix norm equivalence ((ref)), Lemma (ref)-(i,ii) and ((ref)) then give, uniformly in $\alpha \in [0,1]$, $t \in [0,t_{L}]$ and $\mathsf{d} \in \mathcal{B}_{\infty} (0,t)$
\subparagraph{Order of $\widehat{\mathsf{d}}\left( \alpha|I\right)$.} The bounds for $\widehat{\boldsymbol{R}}_{1} \left( \mathsf{d};\alpha,I\right)$ and $\widehat{\boldsymbol{R}}_{2} \left( \mathsf{d};\alpha,I\right)$ then give for $L$ large enough, which allows to take $t\leq t_{L}$ large enough
Hence ((ref)) shows that
\subparagraph{Estimation of bid quantile function and its first derivatives.} As $\max_{\alpha \in [0,1]} \varrho_{\alpha L} \asymp \frac{\log L}{L h^{3 D_{\mathcal{M}}/2+1/2}}$ and $ \max_{x \in \mathcal{X}} \left| P(x)^{\prime} \beta \right| \leq C h^{-D_{\mathcal{M}}/2} \| \beta \|_{\infty} $ under Assumption (ref)-(i,ii), ((ref)) implies, using also Assumption (ref),
This ends the proof of the Theorem. $\Box$
\setcounter{section}{4}
\setcounter{section}{0} \setcounter{footnote}{0} \setcounter{equation}{0} \setcounter{theorem}{0} \setcounter{footnote}{0}
Recall that $S_{1}$ is the row vector $\left[ 0,1,0,\ldots,0\right] $ of dimension $s+2$ and that $S_{0}=\left[ 1,0,\ldots,0\right] $, $\mathsf{S}_{0}=S_{0}\otimes\operatorname*{Id}_{N}$, $\mathsf{S}_{1}=S_{1}\otimes\operatorname*{Id}_{N}$ so that $\widehat{\beta} _{j}\left( \alpha|I\right) =\mathsf{S}_{j}\widehat{\beta}\left( \alpha|I\right) $, $j=0,1$ and
see ((ref)). Recall that ((ref)) gives, for $\widehat{\mathsf{e}}\left( \alpha|I\right) $ as in ((ref))
which is such, for $\widehat{\mathsf{d}}\left( \alpha|I\right) $ as in ((ref)), \[ \widehat{V}\left( \alpha|x,I\right) -\widetilde{V}\left( \alpha|x,I\right) =P\left( x\right) ^{\prime}\left[ \mathsf{S}_{0}+\frac{\alpha\mathsf{S} _{1}}{h\left( I-1\right) }\right] \widehat{\mathsf{d}}\left( \alpha|I\right) . \] Then Theorem (ref) implies, as $\widehat{V} (\alpha|x,I) = \widehat{B} (\alpha|x,I) + \alpha \widehat{B}^{(1)} (\alpha|x,I)/(I-1)$, \[ \sup_{(\alpha,x) \in [0,1] \times \mathcal{X}} \left(L h^{D_{\mathcal{M}}+1}\right)^{1/2} \left| \widehat{V}\left( \alpha|x,I\right) -\widetilde{V}\left( \alpha|x,I\right) \right| = O_{\mathbb{P}} \left( \left( \frac{\log^2 L}{Lh^{3D_{\mathcal{M}}+2}} \right)^{1/2} \right) \] which gives ((ref)) and ((ref)). ((ref)) and ((ref)), ((ref)) and ((ref)) follow from Lemma (ref)-(ii) together with Theorem (ref) and Theorem (ref).
Consider now ((ref)) and ((ref)). It holds since $\mathbb{E}\left[ \widehat{\mathsf{e}}\left( \alpha|I\right) \right] =\overline{\mathsf{R}}^{\left( 2\right) }\left( \overline {\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right) ^{-1}\overline {\mathsf{R}}^{\left( 1\right) }\left( \overline{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right) =0$ for all $\alpha$ in $\left[ 0,1\right] $
For the bias part, Theorem (ref) gives
For the stochastic part, Lemma (ref)-(i) yields, recalling $\mathsf{S}_{1} \widehat{\mathsf{e}}\left( \alpha|I\right) = \widehat{\mathsf{e}}_1 \left( \alpha|I\right)$,
Note that $\mathrm{Var} \left( P(x)^{\prime} \widehat{\mathsf{e}}_1 (\alpha|I) /h\right) \asymp 1/(L h^{D_{\mathcal{M}}+1})$ uniformly in $\alpha$ and $x$ gives $\Sigma_{LI}^{2}=O(1)$. Substituting in the bias-variance decomposition of the integrated mean squared error ends the proof of the Theorem. $\hspace*{\fill}\square$
Note that ((ref)) and ((ref)) follow from Theorem (ref), while ((ref)) and ((ref)) follow from Lemma (ref)-(i).
Lemma (ref)-(i) implies that $h^{D_{\mathcal{M}}} P(x)^{\prime} \Sigma_h (\alpha|I) P(x) \asymp \alpha^2$ uniformly in $x$, for all $\alpha>0$. To prove the CLT part, Lemma (ref)-(i) and Theorem (ref) give \[ \sqrt{LIh^{D_{\mathcal{M}}+1}} \left( \widehat{V} (\alpha|x,I) - \overline{V} (\alpha|x,I) \right) = \sqrt{LIh^{D_{\mathcal{M}}+1}} \frac{\alpha P(x)^{\prime} \widehat{\mathsf{e}}_1 (\alpha|I)}{h(I-1)} + o_{\mathbb{P}} (1). \] Hence it remains to show that \[ \left( \frac{LIh}{P\left( x\right) ^{\prime}\Sigma_{h}\left( \alpha|I\right) P\left( x\right) }\right) ^{1/2}\frac{\alpha P\left( x\right) ^{\prime}\mathsf{S}_{1}\widehat{\mathsf{e}}\left( \alpha|I\right) }{h\left( I-1\right) }\overset{d}{\rightarrow}\mathcal{N}\left( 0,1\right) . \] Write \[ \left( \frac{LIh}{P\left( x\right) ^{\prime}\Sigma_{h}\left( \alpha|I\right) P\left( x\right) }\right) ^{1/2}\frac{\alpha P\left( x\right) ^{\prime}\mathsf{S}_{1}\widehat{\mathsf{e}}\left( \alpha|I\right) }{h\left( I-1\right) }=\sum_{\ell=1}^{L}r_{\ell}\left( \alpha|x,I\right) \] with $r_{\ell}\left( \alpha|x,I\right) =\mathbb{I}\left( I_{\ell}=I\right) \sum_{i=1}^{I_{\ell}}r_{i\ell}\left( \alpha|x,I\right) $ and
Since the first-order condition gives $\mathbb{E}\left[ r_{\ell}\left( \alpha|x,I\right) \right] = \overline{\mathsf{R}}^{(1)}\left(\overline{\mathsf{b}}(\alpha|I);\alpha,I\right)= 0$ and\\ $\left\vert \operatorname*{Var}\left( r_{\ell}\left( \alpha|x,I\right) \right) -1\right\vert =o\left( 1\right) $, it is sufficient to show that $\left\vert \mathbb{E}\left[ r_{\ell}^{3}\left( \alpha|x,I\right) \right] \right\vert =o\left( 1\right) $ holds, see e.g. Theorem $<$ 19 $>$ p.179 in Pollard (2002). But Assumption (ref)-(i) and Proposition (ref)-(i), Lemma (ref) and ((ref)) give \[ \left\vert r_{i\ell}\left( \alpha|x,I\right) \right\vert \leq\frac {C}{\left( Lh\right) ^{1/2}}\frac{\left\Vert P\left( x\right) \right\Vert }{\left\Vert P\left( x\right) \right\Vert } \times\max_{x\in\mathcal{X} }\left\Vert P\left( x\right) \right\Vert =O\left( \frac{1}{\left( Lh^{D_{\mathcal{M}}+1}\right) ^{1/2}}\right) . \] It follows that by Assumption (ref)
This ends the proof of the Theorem.$\hfill\square$
The proof of these theorems requests some specific additional results. The next Lemma gives an expansion for, $f(\cdot)$ and $g(\cdot)$ being two $(s+2)\times 1$ functions and $A$ a $\mathcal{U}_{[0,1]}$ random variable,
A similar item appears when computing the covariance \[ LI \mathrm{Cov} \left[ \int_{0}^{1} \left[f(\alpha) \otimes P(x) \right]^{\prime} \widehat{\mathsf{R}}^{(1)} (\mathsf{b}(\alpha|I);\alpha,I) d\alpha, \int_{0}^{1} \left[g(\alpha) \otimes P(x) \right]^{\prime} \widehat{\mathsf{R}}^{(1)} (\mathsf{b}(\alpha|I);\alpha,I) d\alpha \right], \] see the score expression below ((ref)), noting that $ \mathbb{I} \left( B_{i\ell} \leq P \left( X_{\ell}, \frac{a-\alpha}{h} \right)^{\prime} \mathsf{b} (\alpha|I_{\ell}) \right) $ is close to \[ \mathbb{I} \left( B_{i\ell} \leq B(a|X_{\ell},I_{\ell}) \right) = \mathbb{I} \left( A_{i\ell} \leq a \right). \] Recall that $S_{0}=\left[ 1,0,\ldots,0\right] $, $S_{1}=\left[ 0,1,0,\ldots,0\right] $ and $S_{2}=\left[ 0,0,1,0,\ldots ,0\right] $ are row vectors of dimension $s+2$. The next lemma describes the limit of $\mathcal{C}_{h}$ when $h$ goes to $0$, in relation with the direction $S_0$ and $S_1$ as in the variance study of Lemma (ref).
The smoothness assumption on $f(\cdot)$ and $g(\cdot)$ accounts for vector functions proportional to $\Omega_h (\alpha)$, which is constant over $[h,1-h]$ with $\Omega_h^{(1)} (\alpha) = O(1/h)$ over $[0,h]$ and $[1-h,1]$.
\paragraph{Proof of Lemma (ref):} See (ref).
Consider two functions $\varphi_{0}\left( \alpha,x\right) $ and $\varphi _{1}\left( \alpha|x\right) $ and consider the real random variable
The purpose of the next Lemma is to compute the variance of this integral. Recall
and set \[ \mathsf{M}_{0}\left( \alpha\right) =\Omega_{h}\left( \alpha\right) \otimes\mathbf{P}_{0}\left( \alpha\right) ,\text{ }\mathsf{M}_{1}\left( \alpha\right) =\Omega_{1h}\left( \alpha\right) \otimes\mathbf{P}_{1}\left( \alpha\right) . \] Recall $\mathbf{1}_N = [1,\ldots,1]$ is a $N \times 1$ column vector.
\paragraph{Proof of Lemma (ref).}
Abbreviate $\overline{\mathsf{R}}^{\left( 2\right) }\left( \overline {\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right) $, $\widehat{\mathsf{R} }^{\left( 1\right) }\left( \overline{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right) $ into $\overline{\mathsf{R}}^{\left( 2\right) }\left( \alpha\right) $ and $\widehat{\mathsf{R}}^{\left( 1\right) }\left( \alpha\right) $ respectively. We now give a suitable expansion for $\overline{\mathsf{R}}^{\left( 2\right) }\left( \alpha\right) ^{-1}$. Lemma (ref) and ((ref)) with $s\geq 1$ give, uniformly over $[0,1]$
with respect to $\|\cdot\|_2$, $\|\cdot\|_{\infty}$ or $|\cdot|_{\infty}$ by ((ref)) and since the matrices above are band ones up to a basis permutation. It then follows, uniformly over $\left[ 0,1\right] $
Now $\mathsf{M}_{0}\left( \alpha\right) ^{-1}=\Omega_{h}\left( \alpha\right) ^{-1}\otimes\mathbf{P}_{0}\left( \alpha\right) ^{-1}$ and \[ \mathsf{M}_{0}\left( \alpha\right) ^{-1}\mathsf{M}_{1}\left( \alpha\right) \mathsf{M}_{0}\left( \alpha\right) ^{-1}=\left[ \Omega_{h}\left( \alpha\right) ^{-1}\Omega_{1h}\left( \alpha\right) \Omega_{h}\left( \alpha\right) ^{-1}\right] \otimes\left[ \mathbf{P}_{0}\left( \alpha\right) ^{-1}\mathbf{P}_{1}\left( \alpha\right) \mathbf{P}_{0}\left( \alpha\right) ^{-1}\right] \] with
where $S_{s+1}=[0,\ldots,0,1]$, $c\left( \alpha\right) =c_{h}\left( \alpha\right) $ satisfying the smoothness conditions of Lemma (ref).
Define, for any $(s+2) \times 1$ vector functions $f\left( \cdot\right) $ and $g\left( \cdot\right) $ satisfying the conditions of Lemma (ref), and for two $x_f$ and $x_g$ of $\mathcal{X}$,
so that
Now ((ref)), $\max_{\left( x,t\right) \in\mathcal{X\times}\left[ -1,1\right] }\left\Vert P\left( x,t\right) \right\Vert =O\left( h^{-D_{\mathcal{M}}/2}\right) $ and Lemma (ref)-(iii) gives \[ P\left( X_{\ell},\frac{a-\alpha}{h}\right) \overline{\mathsf{b}}\left( \alpha|I\right) =B\left( a|X_{\ell},I\right) +o\left( h^{s+1}\right) \] uniformly in $a$, $\alpha$ and $X_{\ell}$ with $\frac{a-\alpha}{h}$ in the support of $K\left( \cdot\right) $, that is $\left\vert a-\alpha\right\vert \leq h$. This gives under Assumption (ref) -(ii) and by definition of $\mathbf{P}$
Set \[ \mathfrak{P}_h (x) = h^{D_{\mathcal{M}}/2} \mathbf{P}^{1/2}P(x) . \] Let $\mathcal{C} (f,g)$ be as in Lemma (ref). The expansion ((ref)) of $\overline{\mathsf{R}}^{\left( 2\right) }\left( \alpha\right) ^{-1}$ then gives, under Assumption (ref)-(i),
As $s \geq 2$, Proposition (ref)-(i), Lemma (ref) and the disjoint support condition in Assumption (ref)-(ii) ensure that $\alpha \in [0,1] \mapsto \mathbf{1}_N^{\prime}\mathbf{P}_0 (\alpha )^{-1}\mathfrak{P}_h(x)$, $\mathbf{1}_N^{\prime} \mathbf{P}_0 (\alpha )^{-1} \mathbf{P}_{1} (\alpha ) \mathbf{P}_0 (\alpha )^{-1} \mathfrak{P}_h(x)$ are continuously differentiable with bounded derivatives independently of $h$. $\Omega_h(\alpha)^{-1} $ and $\Omega_h(\alpha)^{-1} \Omega_{1h} (\alpha) \Omega_h(\alpha)^{-1}$ are constant over $[h,1-h]$ and have $O(1/h)$ derivatives outside this interval.
Now, abbreviating $\varphi_0(\alpha|x)$ and $\varphi_1(\alpha|x)$ by removing $x$, recall
and then
Lemma (ref) gives, $A$ being $\mathcal{U}_{[0,1]}$,
and by ((ref)) which gives $ S_1 \Omega_h(\alpha)^{-1} \Omega_{1h} (\alpha) S_0^{\prime}=1 $ and $ S_0 \Omega_h(\alpha)^{-1} \Omega_{1h} (\alpha) S_0^{\prime}=0 $,
Collecting the items then gives
This gives
Observe now that \[ \partial_{\alpha}\left[ \varphi_{1}\left( \alpha\right) \mathbf{P}_{0}\left( \alpha\right) ^{-1}\right] = \left[ \partial_{\alpha} \varphi_{1}\left( \alpha\right) \right] \mathbf{P}_{0} \left(\alpha\right)^{-1} - \varphi_{1}\left( \alpha\right) \mathbf{P}_{0}\left(\alpha\right)^{-1} \mathbf{P}_{1}\left( \alpha\right) \mathbf{P}_{0}\left(\alpha\right)^{-1} \] so that
It then follows
which is equal to $ \sigma_L^2(x)+o(1) $ as stated in the Lemma. The study of $ \operatorname*{Var}\left( \sqrt{LI} \int_{\mathcal{X}} \widehat{\boldsymbol{I} }_{\varphi}\left( x|I\right) dx \right) $ is similar, changing $h^{D_{\mathcal{M}}/2} P(x)$ into $P(x)$ and allowing for vector functions $f(\alpha|x_f)$ and $g(\alpha|x_g)$ and integrating with respect to $(x_f,x_g)$ in the key equation ((ref)), which would then hold with a remainder term $o(h^2) \left\| \int_{\mathcal{X}} |P(x)| dx \right\|_2^2 = o(h^2)$ under Assumption (ref)-(i). $\Box$
Consider two real valued continuous functions $\mathcal{F}_{0}\left( b_{0},b_{1}\right) $ and $\mathcal{F}_{1}\left( b_{0},b_{1}\right) $. Define
A condition ensuring that the variances $\sigma_{L}^{2}\left( x|I\right) $ and $\sigma_{L}^{2}\left( I\right) $ of Lemma (ref) do not vanish is ((ref)), that is for all $x$ in $\mathcal{X}$, \[ \varphi_{0}\left( \alpha|x,I\right) -\partial_{\alpha} \varphi_{1}\left( \alpha|x,I\right) \neq0. \]
\paragraph{Proof of Proposition (ref).}
The eigenvalues of $\mathbf{P}_{0}\left( \alpha\right) ^{-1}$, $\mathbf{P}_{1}\left( \alpha\right) $ and $\mathbf{P}$ are bounded uniformly in $h$ and $\alpha$ by Assumptions (ref) and (ref), and $\left\Vert h^{D_{\mathcal{M}}/2}P\left( x\right) \right\Vert $ is bounded away from $0$ and infinity by Assumptions (ref) and (ref). Then if ((ref)) holds for some $\alpha$, $\sigma_{L}^{2}\left( x|I\right) $ is bounded away from $0$ and infinity and the exact order of $\operatorname*{Var}\left( \widehat{\boldsymbol{I}}_{\mathcal{F}}\left( x|I\right) \right) $ is $1/LIh^{D_{\mathcal{M}}}$. We now check the Lyapounov condition. Write $\widehat{\mathsf{R}}^{\left( 1\right) }\left( \alpha\right) =\frac{1}{LI}\sum_{\ell=1}^{L}\mathbb{I}\left[ I_{\ell }=I\right] r_{\ell}\left( \alpha\right) $, with \[ r_{\ell}\left( \alpha\right) =\sum_{i=1}^{I_{\ell}}\mathsf{\int _{-\frac{\alpha}{h}}^{\frac{1-\alpha}{h}}}\left\{ \mathbb{I}\left( B_{i\ell }\leq P\left( X_{\ell},t\right) ^{\prime}\overline{\mathsf{b}}\left( \alpha|I\right) \right) -\left( \alpha+ht\right) \right\} \pi\left( t\right) \otimes P\left( X_{\ell}\right) K\left( t\right) dt. \] This gives, since the eigenvalues of $\overline{\mathsf{R}}^{\left( 2\right) }\left( \alpha\right) $ are asymptotically bounded from $0$ by Lemma (ref) and ((ref)),
$Lh^{D_{\mathcal{M}}+2}\rightarrow\infty$ implies that the Lyapounov condition holds since \[ \frac{C}{Lh^{D_{\mathcal{M}}+1}\operatorname*{Var}^{3/2}\left( \widehat{\boldsymbol{I}}_{\mathcal{F}}\left( x|I\right) \right) }\operatorname*{Var}\left( \widehat{\boldsymbol{I}}_{\mathcal{F}}\left( x|I\right) \right) =O\left( \frac{1}{\left( Lh^{D_{\mathcal{M}}+2}\right) ^{1/2}}\right) \rightarrow0 \] This implies that $\widehat{\boldsymbol{I}}_{\mathcal{F}}\left( x|I\right) /\operatorname*{Var}^{1/2}\left( \widehat{\boldsymbol{I}}_{\mathcal{F} }\left( x|I\right) \right) $ is asymptotically $\mathcal{N}\left( 0,1\right) $, and then the stated asymptotic normality.
For $\sqrt{LI}\int_{\mathcal{X}}\widehat{\boldsymbol{I}}_{\mathcal{F}}\left( x|I\right) dx$, recall that $\left\Vert \int\left\vert P\left( x\right) \right\vert dx\right\Vert =O\left( 1\right) $ by Assumption (ref). This also gives
Therefore the Lyapounov condition holds since $Lh^{2}$ diverges, because \[ \frac{C}{Lh\operatorname*{Var}^{3/2}\left( \int_{\mathcal{X}} \widehat{\boldsymbol{I}}_{\mathcal{F}}\left( x|I\right) dx\right) }\operatorname*{Var}\left( \int_{\mathcal{X}}\widehat{\boldsymbol{I} }_{\mathcal{F}}\left( x|I\right) dx\right) =\frac{C}{\left( Lh^{2}\right) ^{1/2}}\rightarrow0 \] The rest of the proof is as above.$\hfill\square$
\paragraph{Proof of Theorems (ref) and (ref).}
Let $\widehat{\mathsf{d}}\left( \alpha|I\right) $ and $\widehat{\mathsf{e} }\left( \alpha|I\right) $ be as in ((ref)) and ((ref)),
Let $\widehat{\boldsymbol{I}}_{\mathcal{F}}\left( x|I\right) $ be as above, replacing $\varphi_{j}\left( \cdot\right) $ with $\varphi_{jI}\left( \cdot\right) $, $j=0,1$. Then the second-order Taylor inequality gives
The second-order Taylor inequality, Theorems (ref) and (ref), Lemma (ref) give,
as $\frac{\log^2 L}{L h^{3D_{\mathcal{M}}+3 } }=o(1)$.
Proposition (ref) then gives the result since the $\widehat{\boldsymbol{I}}_{\mathcal{F}}\left( x|I\right) $ are independent. The asymptotic normality of $\widehat{\theta}$ similarly follows from Assumption (ref), which gives $\left\Vert \int_{\mathcal{X}}\left\vert P\left( x\right) \right\vert dx\right\Vert =O\left( 1\right) $, and Theorem (ref) which implies
Theorems (ref) and (ref) follow from Proposition (ref). $\Box$
\setcounter{section}{5}
\setcounter{footnote}{0} \setcounter{section}{0} \setcounter{equation}{0} \setcounter{theorem}{0}
These two results make use of Theorem 2.2 in Demko (1977):
As using Lemma (ref) shortens the proof of Proposition (ref), this lemma is established in the first place.
Note that the rows and the columns of the permutation matrix $S$ and $S^{-1}$ have a unique entry equal to $1$, all other being set to $0$, so that for $k=2,\infty$, $\| S^{-1} M(\alpha) S^{-1} \|_{k} = \| M(\alpha) \|_{k}$ and $\| S M(\alpha)^{-1} S \|_{k} = \| M(\alpha)^{-1} \|_{k}$. Hence one can assume without loss of generality that the $M(\alpha)$'s are $c$-band matrix. Let $m^{ij} (\alpha)$ be the entries of $M(\alpha)^{-1}$. Then Theorem (ref) gives that there are a $C$ and a $\varrho \in (0,1)$, which only depend upon $c$ and $C_0$, such that $|m^{ij} (\alpha)| \leq C \varrho^{|i-j|}$. Recall now that \[ \left\| M(\alpha)^{-1} \right\| = \max_{i} \sum_{j} \left| m^{ij} (\alpha)\right|. \] But $\sum_{j} \left| m^{ij} (\alpha)\right| \leq C\left(1+2\sum_{k=1}^{\infty} \varrho^k \right)=C \left(1+\frac{2\varrho}{1-\varrho}\right)$, and the Lemma is proven. $\Box$
\subparagraph{Proof of ((ref)).} Note that $\gamma (\cdot)$ is differentiable with, by the Lebesgue Dominated Convergence Theorem \[ \gamma^{(p)} (\alpha|I) = \left( \mathbb{E} \left[ \left. P(X ) P(X)^{\prime} \right| I \right] \right)^{-1} \mathbb{E} \left[ \left. P(X ) V^{(p)}(\alpha|X,I) \right| I \right] , \quad p=0,\ldots,s+1, \] so that $\gamma^{(s+1)} (\cdot|I)$ is also continuous.
By the Approximation Property S, there exists some $N \times 1$ $\gamma_p (\cdot)$ such that $r_p (\alpha|x) = P(x)^{\prime} \gamma_p (\alpha)-V^{(p)} (\alpha|x,I) $ satisfies\footnote{Using for instance a suitable expansion of $V(\cdot|\cdot,I)$ over $[0,1] \times \mathcal{X}_{\epsilon}$.}, for $p=0,\ldots,s+1$, \[ \sup_{\alpha \in [0,1]} \max_{x \in [0,1]^D} \left| r_p (\alpha|x) \right| = h^{s+1-p} O \left( \sum_{\mathbf{j}} \sup_{\alpha \in [0,1]} \textrm{mc}_{s+1-p} \left(V^{(p)}_{\mathbf{j}}(\alpha|x,I);h \right) \right) = o(h^{s+1-p}). \] Hence, for all $x$ of $\mathcal{X}$ and by the sieve disjoint support property,
As the eigenvalues of $\int_{\mathcal{X}} P(x)P(x)^{\prime}dx$ are bounded away from $0$ and $\infty$ by Assumption (ref)-(i), so are the ones of the $c/2$ band matrix $ \mathbb{E} \left[ \left. P(X ) P(X)^{\prime} \right| I \right] $ by Assumption (ref), and Lemma (ref) gives $ \left\| \left( \mathbb{E} \left[ \left. P(X ) P(X)^{\prime} \right| I \right] \right)^{-1} \right\|_{\infty} \leq C $. As $ \left\| \mathbb{E} \left[ \left. P(X ) r_p (\alpha|X) \right| I \right] \right\|_{\infty} \leq o \left(h^{s+1-p}\right) \sup_{1 \leq n \leq N} \int_{\mathcal{X}} |P_n(x)| dx$ with $\sup_{1 \leq n \leq N} \int_{\mathcal{X}} |P_n(x)| dx \leq C h^{D_{\mathcal{M}}/2}$ by Assumption (ref)-(i), ((ref)) follows from: \[ \max_{(\alpha,x)\in [0,1] \times \mathcal{X}} \left| P(x)^{\prime} \gamma^{(p)} (\alpha|I) - P(x)^{\prime} \gamma_p (\alpha) \right| \leq Ch^{-D_{\mathcal{M}}/2} h^{D_{\mathcal{M}}/2} o \left(h^{s+1-p}\right) = o \left(h^{s+1-p}\right). \]
\subparagraph{Proof of (i).} By ((ref)), $B\left( \alpha|x,I\right) =\left( I-1\right) \int_{0} ^{1}u^{I-2}V\left( \alpha u|x,I\right) du$, so that $B^{\left( 1\right) }\left( \alpha|x,I\right) =\left( I-1\right) \int_{0}^{1}u^{I-1}V^{\left( 1\right) }\left( \alpha u|x,I\right) du$ which implies the two first statements in (i) about lower and upper bounds for $B^{\left( 1\right) }\left( \alpha|x,I\right) $ and that $B\left( \cdot|\cdot,I\right) $ is $\left( s+1\right) $th continuously differentiable. That $B\left( \cdot|x,I\right) $ is $(s+2)$th continuously differentiable over $\left( 0,1\right] $ follows from its integral expression ((ref)). Observe now that for $p=1,\ldots,s+2$ \[ \partial_{\alpha}^{p} \left[ \alpha B\left( \alpha|x,I\right) \right] =\alpha B^{\left( p\right) }\left( \alpha|x,I\right) +pB^{\left( p-1\right) }\left( \alpha|x,I\right) \] with, for $p=1,\ldots,s+1$
Hence, when $\alpha$ goes to $0$
uniformly with respect to $x$. As $\partial^{s+2} [\alpha B(\alpha|x,I)] = \alpha B^{(s+2)} (\alpha|x,I) + (s+2) B^{(s+1)} (\alpha|x,I)$, it follows that $\alpha \in [0,1] \mapsto \alpha B(\alpha|x,I)$ is $(s+2)$ times continuously differentiable.
\subparagraph{Proof of (ii).} For $\beta(\cdot|I)$ as in ((ref)) \[ \beta^{\left( p\right) }\left( \alpha|I\right) =\left( I-1\right) \int_{0}^{1}u^{I+p-2}\gamma^{\left( p\right) }\left( \alpha u|I\right) du,\quad p=0,\ldots,s+1 \] and
by ((ref)), which gives the sieve approximation result for $B\left( \alpha|x,I\right) $ in (ii).
\subparagraph{Proof of (iii).} That $\alpha \beta^{(s+2)} (\alpha)$ is continuous over $(0,1]$ with $\lim_{\alpha \downarrow 0 } \alpha \beta^{(s+2)} (\alpha)=0$ can be established as for $\alpha B^{(s+2)} (\alpha|x,I)$. Observe that $\alpha\beta^{(1)} (\alpha|I)= (I-1) \left(\gamma(\alpha|I)-\beta(\alpha|I)\right)$ by ((ref)). It follows
by ((ref)), which gives the approximation result for $\alpha B^{\left( p+1\right) }\left( \alpha|x,I\right) $, $p=0,\ldots,s+1$ in (iii) using $\partial_{\alpha}^p\left[\alpha B^{(1)} (\alpha|x,I)\right] = \alpha B^{(p+1)} (\alpha|X,I)+(p+1)B^{(p)} (\alpha|x,I) $ and ((ref)). $\hfill\square$
Consider the harder $ASQR$ case. We start with (iii). Recall that \[ \Psi(\alpha|x,\mathsf{b}) = P(x,t)^{\prime} \mathsf{b} = P(x)^{\prime} \sum_{p=0}^{s+1} \frac{t^p}{p!} \mathsf{b}_p = P(x)^{\prime} \sum_{p=0}^{s+1} \frac{\left(ht\right)^p}{p!} \beta_p \] where the $1 \times N$ $\beta_p(\alpha|I)$ is equal to $\beta^{(p)} (\alpha|I)$, $\beta (\cdot|I)$ as in ((ref)).
\subparagraph{Proof of (iii).} Proposition (ref)-(ii,i) gives, uniformly in $\alpha$, $t$ in $\mathcal{T}_{\alpha,h}$ and $x$ in $\mathcal{X}$
by the Taylor expansion formula of order $s+1$, so that
For the next result in (iii) and since $\alpha \mapsto \alpha B(\alpha|x,I)$ is $(s+2)$ times continuously differentiable over $[0,1]$ by Proposition (ref)-(i), a Taylor expansion with integral remainder shows, observing that $\partial^{p}_{\alpha} \left[\alpha B(\alpha|x,I)\right]=\alpha B^{(p)}(\alpha|x,I) + pB^{(p-1)}(\alpha|x,I)$
for all $t$ in $\mathcal{T}_{\alpha,h}$. It also holds
Combining the two Taylor expansions then gives
The Dominated Convergence Theorem then gives \[ \max_{(\alpha,x)\in[0,1]\times\mathcal{X}} \max_{t \in \mathcal{T}_{\alpha,h}} \left| \alpha B(\alpha+ht|x,I) - \alpha B(\alpha|x,I) - \sum_{p=1}^{s+2} \frac{(ht)^p}{p!} \alpha B^{(p)}(\alpha|x,I) \right| = o(h^{s+2}). \] As Proposition (ref)-(iii) gives $ P(x)^{\prime} \alpha \mathsf{b}_p (\alpha) = h^{p} \alpha B^{(p)} (\alpha|x,I) + o(h^{s+2})$ uniformly, the second result of (iii) is proven.
The third result in (iii) follows from Proposition (ref)-(iii). The fourth equality of (iii) follows from the first result of (iii), which implies
by Proposition (ref)-(i) which gives that $B^{(1)} (\cdot|\cdot)$ is bounded away from $0$.
\subparagraph{Proof of (i).} We first show that $\mathsf{b} (\alpha|I)$ belongs to $\underline{\mathcal{BI} }_{\alpha,h}$ for all $\alpha$. Observe that Proposition (ref)-(ii) gives, uniformly in $\alpha$, $x$ and $t$ in $\mathcal{T}_{\alpha,h} \subset [-1,1]$
Hence $ \min_{\alpha \in [0,1]} \min_{(t,x) \in \mathcal{T}_{\alpha,h} \times \mathcal{X}} \Psi^{(1)} \left(t|x,\mathsf{b}\left( \alpha|I\right)\right) = h \min_{(\alpha,x) \in [0,1] \times \mathcal{X}} B^{\left( 1\right) }\left( \alpha|x,I\right) +O\left( h^2\right) \geq h/\underline{f} $ for any $\underline{f}$ large enough so that $\min_{(\alpha,x) \in [0,1] \times \mathcal{X}} B^{\left( 1\right) }\left( \alpha|x,I\right) > 1/\underline{f}$ and $h$ small enough as $B^{(1)} (\cdot|\cdot)$ is bounded away from $0$ by Proposition (ref)-(i). Proposition (ref)-(ii) also implies
provided $\overline{f}$ is large enough, since $B^{\left( 1\right) }\left( \cdot|\cdot,\cdot\right) $ is bounded away from infinity by Proposition (ref), and $h$ small enough, so that $\mathsf{b}\left( \alpha|I\right) $ is in $\underline{\mathcal{BI} }_{\alpha,h}$ for all $\alpha$.
Suppose now that $\left\Vert \mathsf{b}-\mathsf{b}\left( \alpha|I\right) \right\Vert_{\infty} \leq C_0 h^{D_{\mathcal{M}}/2+1}$. Then since $\max_{x \in \mathcal{X}} \|P(x)\|=O (h^{-D_{\mathcal{M}}/2})$ by Assumption (ref)-(i),
and $\mathcal{B}_{\infty} \left( \mathsf{b}\left( \alpha|I\right) ,Ch^{D_{\mathcal{M}}/2+1}\right) \subset\underline{\mathcal{BI}}_{\alpha,h}$ for all $\alpha$ when $h\geq 1$ provided $C_0$ is small enough. Hence (i) holds.
\subparagraph{Proof of (ii).} (ii) follows from the Implicit Function Theorem and the definition of $\mathcal{BI}_{\alpha,h}$.
\subparagraph{Proof of (iv).} The first bound follows from the Cauchy-Schwarz inequality. This bound implies for all $u$ in $\Psi\left[ \mathcal{T}_{\alpha ,h}|x,\mathsf{b}^{1}\right] \cap\Psi\left[ \mathcal{T}_{\alpha ,h}|x,\mathsf{b}^{0}\right] $
By definition of $\underline{\mathcal{BI}}_{\alpha,h}$
and substituting shows that the second bound of (iv) holds.
The expression in (ii) of $\partial_u \Phi\left(u|x,\mathsf{b} \right) $ and the definition of $\underline{\mathcal{BI}}_{\alpha,h}$ yield the third inequality, using the fourth bound which is established now. Note that
But, by definition of $\underline{\mathcal{BI}}_{\alpha,h}$ \[ \max_{t\in\mathcal{T}_{\alpha,h}}\left\vert \partial_{t}^{2}\Psi\left( t|x,\mathsf{b}^{1}\right) \right\vert \leq Ch\max _{p=2,\ldots,s+1}\left\vert \frac{P\left( x\right) \mathsf{b}^{1}_{p}} {h}\right\vert =O\left( h\right) \] so that substituting and using the bound for $\Phi\left( u|x,\mathsf{b}^{1}\right) -\Phi\left( u|x,\mathsf{b}^{0}\right) $ give, for all $\alpha$, $x$, $u$, $\mathsf{b}^{1}$ and $\mathsf{b}^{0}$ \[ \left\vert \Psi^{(1)}\left[ \Delta\left( u|x,\mathsf{b} _{1}\right) |x,\mathsf{b}^{1}\right] - \Psi^{(1)}\left[ \Delta\left( u|x,\mathsf{b}^{0}\right) |x,\mathsf{b}^{0}\right] \right\vert \leq Ch^{-D_{\mathcal{M}}/2}\left\Vert \mathsf{b}^{1}-\mathsf{b}^{0}\right\Vert_{\infty} , \] which is the fourth inequality. $\hspace*{\fill}\square$
Recall
In the equation above, $\Delta (u|x,I)$ is such that $\Delta\left[ \Psi\left[ t|x,\mathsf{b}\right] |x,\mathsf{b}\right] =t$ for all $t$ in $\mathcal{T}_{\alpha,h}$. Let \[ \overline{t}_{\alpha,h}\left( x,I;\mathsf{b}\right) =\overline{t}_{\alpha ,h}\wedge\Delta\left[ B\left( 1|x,I\right) |x,\mathsf{b}\right] ,\quad\underline{t}_{\alpha,h}\left( x,I;\mathsf{b}\right) = \underline{t}_{\alpha,h}\vee\Delta\left[ B\left( 0|x,I\right) |x,\mathsf{b}\right] . \] The change of variable $y=\Psi\left( t|x,\mathsf{b}\right) $ yields that \[ \overline{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b};\alpha,I\right) =\int\left[ \int_{\underline{t}_{\alpha,h}\left( x,I;\mathsf{b}\right) }^{\overline{t}_{\alpha,h}\left( x,I;\mathsf{b}\right) }P\left( x,t\right) P\left( x,t\right) ^{\prime}K\left( t\right) g\left( \Psi\left( t|x,\mathsf{b}\right) ,x,I\right) dt\right] dx. \] The Dominated Convergence Theorem, Proposition (ref) -(i) and $s\geq1$, \footnote{Recall this implies that $g\left( \cdot,\cdot,I\right) $ is bounded away from $0$ and infinity and differentiable with respect to its first variable. Hence so is $\mathsf{b} \mapsto g\left( \Psi\left( t|x,\mathsf{b}\right) ,x,I\right)$ for $t$ in the open interval $(\underline{t}_{\alpha,h}\left( x,I;\mathsf{b}\right),\overline{t}_{\alpha,h}\left( x,I;\mathsf{b}\right)$.} yield that $\overline{\mathsf{R} }^{\left( 2\right) }\left( \mathsf{\cdot};\alpha,I\right) $ is continuously differentiable over $\underline{\mathcal{BI}}_{\alpha,h}$ with a differential operator $\mathsf{R}^{\left( 3\right) }\left( \mathsf{b};\alpha,I\right) \left[ \cdot\right]$ which maps a vector $\mathsf{d}$ of $\mathbb{R}^{N(s+2)}$ to a $N(s+2) \times N(s+2)$ matrix. By the Leibniz integral rule
Consider any $\mathsf{b}$ in $\mathcal{B}_{\infty} \left(\mathsf{b}(\alpha|I),C_0 h^{D_{\mathcal{M}}/2+1}\right)$, the inequalities below being uniform over such $\mathsf{b}$. Proposition (ref)-(i), ((ref)) and Assumption (ref)-(i), (ref) imply \[ \left\Vert \overline{\mathsf{R}}_{0}^{\left( 3\right) }\left( \mathsf{b};\alpha,I\right) \left[ \mathsf{d}\right] \right\Vert_{\infty} \leq C\max_{x\in\mathcal{X}}\left\Vert P\left( x\right) \right\Vert \left\Vert \mathsf{d}\right\Vert_{\infty} \leq Ch^{-D_{\mathcal{M}}/2}\left\Vert \mathsf{d} \right\Vert . \] The operators $\overline{\mathsf{R}}_{i}^{\left( 3\right) }\left( \mathsf{b};\alpha,I\right) \left[ \mathsf{d}\right] $, $i=1,2$, can be studied in a similar way so that only $\overline{\mathsf{R}}_{1}^{\left( 3\right) }\left( \mathsf{b};\alpha,I\right) \left[ \mathsf{d}\right] $ is considered. Observe \[ \partial_{\mathsf{b}}\overline{t}_{\alpha,h}\left( x,I;\mathsf{b}\right) =\left\{
\right. . \] But by Lemma (ref)-(iii,iv) and $\left\Vert \mathsf{b}-\mathsf{b}\left( \alpha|I\right) \right\Vert_{\infty} \leq Ch^{ D_{\mathcal{M}}/2+1}$ , for $h$ small enough,
uniformly in $\alpha$, $x$ and $\mathsf{b}$ in $\mathcal{B}_{\infty}\left( \mathsf{b}\left( \alpha|I\right) ,Ch^{D_{\mathcal{M}}/2+1}\right) $. Hence, if $\alpha\leq1-C^{\prime}h$ with $C^{\prime }>0$ large enough \[ \Delta\left[ B\left( 1|x,I\right) |x,\mathsf{b}\right] \geq\frac {\min\left\{ \alpha+h,1-Ch\right\} -\alpha}{h}\geq1\geq\overline{t}_{\alpha,h} \] so that $\partial_{\mathsf{b}}\overline{t}_{\alpha,h}\left( x,I;\mathsf{b}\right) =0$. Hence since $\mathcal{B}_{\infty}\left( \mathsf{b}\left( \alpha|I\right) ,Ch^{D_{\mathcal{M}}/2+1}\right) \subset\underline{\mathcal{BI}}_{\alpha,h}$, the definition of $\underline{\mathcal{BI}}_{\alpha,h}$ which implies $\Psi^{(1)}\left( \Delta\left( B\left( 1|x,I\right) |x,\mathsf{b}\right) |x,\mathsf{b}\right) \geq h/\underline{f}$, and using $\frac{\mathbb{I} \left[ \alpha \geq 1-C' h \right]}{h} \leq \frac{C}{\alpha(1-\alpha)+h}$ give for $\alpha \geq 1-C' h $
Substituting in the expression of $\overline{\mathsf{R}}^{\left( 3\right) }\left( \mathsf{b};\alpha,I\right) \left[ \mathsf{d}\right] $ then gives uniformly in $\mathsf{d}$ \[ \max_{\alpha\in\left[ 0,1\right] }\max_{\mathsf{b\in}\mathcal{B}_{\infty} \left( \mathsf{b}\left( \alpha|I\right) ,Ch^{D_{\mathcal{M}}/2+1}\right) }\left( \alpha\left( 1-\alpha\right) +h\right) \left\Vert \overline{\mathsf{R}}^{\left( 3\right) }\left( \mathsf{b};\alpha,I\right) \left[ \mathsf{d}\right] \right\Vert_{\infty} \leq Ch^{-D_{\mathcal{M}}/2} \left\Vert \mathsf{d}\right\Vert_{\infty} . \] The Taylor inequality then shows that the bound in (i) holds.
For (ii), Lemma (ref)-(iii) and Proposition (ref)-(i) give that, uniformly in $\alpha$ and $x$
This gives for $\overline{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b}\left( \alpha|I\right) ;\alpha,I\right)$ the expansion
where all remainder terms are with respect to the matrix norm. This together the fact that the eigenvalues of the matrices $\Omega_{h}\left( \alpha\right) $ and $\int_{\mathcal{X}}P\left( x\right) P\left( x\right) ^{\prime}dx$ are bounded away from $0$ and infinity, the fact that $B^{\left( 1\right) }\left( \alpha|X,I\right) $ is bounded away from $0$ and infinity shows that (ii) holds.$\hspace*{\fill}\square$
The proofs of the lemmas grouped here make use of a deviation inequality from Massart (2007). Consider $n$ independent random variables $Z_{\ell}$ and, for a known real function $\xi\left( z,\theta\right) $ separable with respect to $\theta\in\Theta$, $Z_{\ell}\left( \theta\right) =\xi\left( Z_{\ell} ,\theta\right) $ where $\theta$ is a parameter. Let $\underline{\xi}\left( \cdot\right) \leq\overline{\xi}\left( \cdot\right) $ be two functions. A bracket $\left[ \underline{\xi},\overline{\xi}\right] $ is the set of all functions $\xi\left( \cdot\right) $ such that $\underline{\xi}\left( z\right) \leq\xi\left( z\right) \leq\overline{\xi}\left( z\right) $ for all $z$. The next Theorem is from Massart (2007, Theorem 6.8 and Corollary 6.9).
Note that $\widehat{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b} ;\alpha,I\right) \mathsf{-}\overline{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b};\alpha,I\right) $ is an index permutation of a $c\left( s+2\right) $-band matrix, so that the order of its matrix norm is the same than the order of its largest entry by ((ref)). The generic entry of $\widehat{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b};\alpha,I\right) -\overline{\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b};\alpha,I\right) $ can be written as \[ \widehat{\mathsf{r}}^{(2)}\left( \mathsf{b};\alpha,I\right) =\frac{1}{LIh^{\left( D_{\mathcal{M}}+1\right) /2}}\sum_{\ell=1}^{L}\xi_{\ell}\left( \mathsf{b};\alpha\right) \] where the $\xi_{\ell}\left( \mathsf{b};\alpha\right) $ are centered iid with
The proof of the Lemma follows from Theorem (ref). Observe \[ \left\vert \xi_{\ell}\left( \mathsf{b};\alpha\right) \right\vert \leq C\frac{h^{D_{\mathcal{M}}/2}\max_{x\in\mathcal{X}}\left\Vert P\left( x\right) \right\Vert ^{2}}{h^{1/2}}\leq M_{\infty}\text{ with }M_{\infty }\asymp h^{-\left( D_{\mathcal{M}}+1\right) /2}. \] for all $\alpha$ in $\left[ 0,1\right] $ and all admissible $\mathsf{b}$. For the variance, Lemma (ref)-(iii,iv) gives
uniformly. It follows that, $U_{i\ell}=G\left( B_{i\ell}|X_{\ell},I_{\ell }\right) $ being a uniform random variable independent of $\left( x_{\ell },I_{\ell}\right) $
under Assumption (ref)-(i), uniformly in $\mathsf{b}$ and $\alpha$.
Consider now the brackets covering. The key observation is that, due to the localized sieve construction and the definition of $\Delta(\cdot|x,\mathsf{b})=\Psi^{-1} (\cdot|x,\mathsf{b})$, $\xi_{\ell }\left( \mathsf{b};\alpha\right) $ only depends on a finite dimension subvector of $\mathsf{b}$, $\mathsf{b}^{\left( n_{1},n_{2}\right) }$ which groups the entries of $\mathsf{b}$ corresponding to those $P_{n}\left( \cdot\right) $ such that $P_{n}\left( \cdot\right) P_{n_{1}}\left( \cdot\right) \neq0$ or $P_{n}\left( \cdot\right) P_{n_{2}}\left( \cdot\right) \neq0$, so that the dimension of $\mathsf{b}^{\left( n_{1},n_{2}\right) }$ is less than $2c\left( s+2\right) $ under Assumption (ref)-(ii). Consequently the class to be bracketed is \[ \mathcal{F}\mathcal{=}\left\{ \xi_{\ell}\left( \mathsf{b}^{\left( n_{1},n_{2}\right) };\alpha\right) ;\alpha\in\left[ 0,1\right] ,\mathsf{b}^{\left( n_{1},n_{2}\right) }\mathsf{\in}\mathcal{B}_{\infty} \left( \mathsf{b}^{\left( n_{1},n_{2}\right) }\left( \alpha|I\right) ,Ch^{D_{\mathcal{M}}/2+1}\right) \right\} . \] Lemma (ref)-(iii), $1/\left( Lh^{D_{\mathcal{M}}+1}\right) =o\left( 1\right) $, van de Geer (1999, p.20) and arguing as in Lemma B.2 of Guerre and Sabbah (2012,2014) , Lemma (ref)-(iii), differentiability of $K(\cdot)$ and $h \asymp L^{-C}$ imply that $\mathcal{F}$ can be bracketed with a number of brackets \[ \exp\left( H_{L}\left( \epsilon\right) \right) \asymp\left( \frac{L^{C} }{\epsilon}\right) ^{C} \] so that \[ \int_{0}^{M_{2}/2}\sqrt{\min\left( L,H_{L}\left( \epsilon\right) \right) }d\epsilon\leq\left( \frac{M_{2}}{2}\right) ^{1/2}\left( \int_{0}^{M_{2} /2}H_{L}\left( \epsilon\right) d\epsilon\right) ^{1/2}=O\left( \log L\right) ^{1/2} \] and for the item $\mathcal{H}_{L}$ of Theorem (ref), \[ \mathcal{H}_{L}=O\left( \log L\right) ^{1/2}+O\left( \frac{\log L}{Lh^{D_{\mathcal{M}}+1}}\right) ^{1/2}=O\left( \log L\right) ^{1/2} \] since $1/\left( Lh^{D_{\mathcal{M}}+1}\right) $ is bounded. Hence, by Theorem (ref) for $t\leq10L^{1/2}M_{2}/M_{\infty}$
uniformly over all the non zero entries $\widehat{\mathsf{r}}\left( \mathsf{b};\alpha,I\right) $ of the band matrix $\widehat{\mathsf{R} }^{\left( 2\right) }\left( \mathsf{b};\alpha,I\right) -\overline {\mathsf{R}}^{\left( 2\right) }\left( \mathsf{b};\alpha,I\right) $. This gives, by the Bonferroni inequality,
which, as $N \asymp h^{-D_{\mathcal{M}}} \leq L^{C}$, implies the result of the lemma by the matrix norm equivalence ((ref)) and since $t\leq10L^{1/2}M_{2}/M_{\infty }=O\left( Lh^{D_{\mathcal{M}}+1}\right) ^{1/2}$ can be set to $t=5\tau \log^{1/2}L$ for an arbitrary large $\tau$ as $\log L/\left( Lh^{D_{\mathcal{M}}+1}\right) =o\left( 1\right) $, so that $ N(s+2)\exp\left( -\tau \log L\right)= o(1)$ .$\hspace*{\fill}\square$
That $\max_{\alpha \in [0,1]}\left\| \mathrm{Var} \left[ \widehat{\mathsf{R}}^{(1)} \left(\overline{\mathsf{b}} (\alpha|I);\alpha,I\right) \right] \right\|_{j}=O\left(\frac{1}{LI}\right)$, $j=2,\infty$, noting that $\mathrm{Var} \left[ \widehat{\mathsf{R}}^{(1)} \left(\overline{\mathsf{b}} (\alpha|I);\alpha,I\right) \right]$ is a band matrix up to a basis permutation follow from the expansion ((ref)) obtained in the next section. The rest of the proof of Lemma (ref) is similar to the one of Lemma (ref). The generic entry of $\widehat{\mathsf{R}}^{\left( 1\right) }\left( \mathsf{b};\alpha,I\right) -\overline{\mathsf{R}}^{\left( 1\right) }\left( \mathsf{b};\alpha,I\right) $ writes \[ \widehat{\mathsf{r}}^{(1)}\left( \mathsf{b};\alpha,I\right) =\frac{1}{LI}\sum _{\ell=1}^{L}\xi_{\ell}\left( \mathsf{b};\alpha\right) \] where the $\xi_{\ell}\left( \mathsf{b};\alpha\right) $ are centered iid with, for $K_{p}\left( t\right) =t^{p}K\left( t\right) /p!$,
This gives \[ \left\vert \frac{\xi_{\ell}\left( \mathsf{b};\alpha\right) }{\left( h+\alpha\left( 1-\alpha\right) \right) ^{1/2}}\right\vert \leq Ch^{-1/2}\max_{x\in\mathcal{X}}\left\Vert P\left( x\right) \right\Vert \leq M_{\infty}\text{ with }M_{\infty}\asymp h^{-\left( D_{\mathcal{M}}+1\right) /2}. \] For the computation of the variance, Lemma (ref)-(iii,iv) and Proposition (ref)-(i) give uniformly in $\alpha$, $t$ in $\mathcal{T}_{\alpha,h}$ the admissible $\mathsf{b}$ and $X_{\ell}$, and for the uniform $U_{i\ell}=G\left( B_{i\ell}|X_{\ell},I_{\ell}\right) $,
It then follows, since $U_{i\ell}\sim \mathcal{U}_{[0,1]}$ is independent of $\left( X_{\ell} ,I_{\ell}\right) $
uniformly in $\alpha$ and $\mathsf{b}$. Hence, uniformly in $\alpha$ and $\mathsf{b}$ \[ \operatorname*{Var}\left( \frac{\xi_{\ell}\left( \mathsf{b};\alpha\right) }{\left( h+\alpha\left( 1-\alpha\right) \right) ^{1/2}}\right) \leq M_{2}^{2}\text{ with }M_{2}<\infty. \]
The bracketing part of the proof is similar to the one in Lemma (ref) noticing that $\zeta_{\ell} (\mathsf{b};\alpha)$ only depends upon a subvector $\mathsf{b}^{\left( n_{1},n_{2}\right) }$ of $\mathsf{b}$ of finite dimension, and similar to Guerre and Sabbah (2012,2014, Lemma B.2). This gives, for $\mathcal{H}_{L}$ as in Theorem (ref) \[ \mathcal{H}_{L}=O\left( \log L\right) ^{1/2}+O\left( \frac{\log L}{Lh^{D_{\mathcal{M}}+1}}\right) ^{1/2}=O\left( \log L\right) ^{1/2}. \] Arguing with Theorem (ref) then shows that the order of the largest entry in $\widehat{\mathsf{R}}^{\left( 1\right) }\left( \mathsf{b} ;\alpha,I\right) -\overline{\mathsf{R}}^{\left( 1\right) }\left( \mathsf{b};\alpha,I\right) $ is $O_{\mathbb{P}}\left( \log L/L\right) ^{1/2}$, which gives \[ \max_{\alpha\in\left[ 0,1\right] } \max_{\mathsf{b\in}\mathcal{B}_{\infty} \left( \mathsf{b}\left( \alpha|I\right) ,Ch^{D_{\mathcal{M}}/2+1}\right)} \frac{ \left\Vert \widehat{\mathsf{R}}^{\left( 1\right) }\left( \mathsf{b} ;\alpha,I\right) -\overline{\mathsf{R}}^{\left( 1\right) }\left( \mathsf{b};\alpha,I\right) \right\Vert_{\infty} }{ \left( h + \alpha(1-\alpha) \right)^{1/2} } =O_{\mathbb{P}}\left( \frac{\log L}{L}\right)^{1/2}, \] which also gives the result for the $\|\cdot\|_{2}$ by ((ref)), $ \mathcal{B} \left( \mathsf{b} ,\epsilon \right) \subset \mathcal{B}_{\infty} \left( \mathsf{b} ,\epsilon \right) $ and $N \asymp h^{-D_{\mathcal{M}}}$.$\hspace*{\fill}\square$
For (i), we first give a convenient expression for $v_h^2 (\alpha)$. Recall that $S_0$ and $S_1$ are $1\times (s+2)$ vectors with $S_0 = [1,0,\ldots,0]$ and $S_1 = [0,1,0,\ldots,0]$, see ((ref)). Abbreviate $\Omega_{h}\left( \alpha\right) $, $\Omega_{1h}\left( \alpha\right) $ in $\Omega$, $\Omega_{1}$. Define
Note that $v_{h}^{2}\left( \alpha\right) =S_{1}\Omega^{-1}\mathbf{\Pi} _{m}\Omega^{-1}S_{1}^{\prime}$. Let $\mathbb{W}_0 (\cdot)$ be a Brownian motion over the straight line and $\mathbb{W} (\cdot)$ a Brownian motion over $[0,\infty)$, and set \[ \mathbf{v}_h^2 (\alpha) = \mathrm{Var} \left[ \int_{\underline{t}_{\alpha,h}}^{\overline{t}_{\alpha,h}} \mathbb{W}_0 (t) S_1 \Omega^{-1} \pi(t) K(t) dt \right]. \] As $S_1 \Omega^{-1} \int_{\underline{t}_{\alpha,h}}^{\overline{t}_{\alpha,h}} \pi(t) K(t) dt = S_1 \Omega^{-1} \omega_0 =S_{1} S_{0}^{\prime} =0$, it follows
As $\overline{t}_{\alpha,h} - \underline{t}_{\alpha,h} \geq 1 $ for $h$ small enough, it follows that $\inf_{\alpha \in [0,1]} v_{h}^2 (\alpha) \geq 1/C$ while $\sup_{\alpha \in [0,1]} v_{h}^2 (\alpha) \leq C$ since the eigenvalues of $\Omega = \Omega_h (\alpha)$ stay bounded away from $0$ and infinity when $h$ goes to $0$, uniformly in $\alpha \in [0,1]$.
Abbreviate $\mathbf{P}(I)$, $\mathbf{P}_0(\alpha|I)$ and $\mathbf{P}_1 (\alpha|I)$ as
and abbreviate $\Omega_{h}\left( \alpha\right) $, $\Omega_{1h}\left( \alpha\right) $ in $\Omega$, $\Omega_{1}$. It holds \[ \operatorname*{Var}\left( \widehat{\mathsf{e}}\left( \alpha|I\right) \right) =\left[ \overline{\mathsf{R}}^{\left( 2\right) }\left( \overline{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right) \right] ^{-1}\operatorname*{Var}\left[ \widehat{\mathsf{R}}^{\left( 1\right) }\left( \overline{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right) \right] \left[ \overline{\mathsf{R}}^{\left( 2\right) }\left( \overline{\mathsf{b}}\left( \alpha|I\right) ;\alpha,I\right) \right] ^{-1} \] with by Lemma (ref)