EconBase
← Back to paper

Inference for First-Price Auctions with Guerre, Perrigne, and Vuong's Estimator

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

111,339 characters · 11 sections · 72 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference for First-Price Auctions with Guerre, Perrigne, and Vuong's Estimator 2019. This manuscript version is made available under the Creative Commons CC-BY-NC-ND 4.0 license http://creativecommons.org/licenses/by-nc-nd/4.0/. First version: March 3, 2016, This version: .

abstractWe consider inference on the probability density of valuations in the first-price sealed-bid auctions model within the independent private value paradigm. We show the asymptotic normality of the two-step nonparametric estimator of \citet*{Guerre_Perrigne_Vuong_Auction_Econometrica_2000} (GPV), and propose an easily implementable and consistent estimator of the asymptotic variance. We prove the validity of the pointwise percentile bootstrap confidence intervals based on the GPV estimator. Lastly, we use the intermediate Gaussian approximation approach to construct bootstrap-based asymptotically valid uniform confidence bands for the density of the valuations. \\ Keywords: Asymptotic Normality, Bootstrap, First-Price Auctions, Gaussian Approximation, Independent Private Values, Two-Step Nonparametric Estimators, Uniform Confidence Bands JEL classification: C14, C57

Introduction

The structural estimation of auctions is an important and rapidly growing subfield at the junction of econometrics and industrial organization. Since the seminal work of \citet*[GPV hereafter]{Guerre_Perrigne_Vuong_Auction_Econometrica_2000}, much of theoretical and applied work has focused on nonparametric estimation of first-price, sealed-bid auctions.\footnote{\citet*{hendricks2007empirical} survey the empirical auction literature, while \citet*{athey2007nonparametric} survey the nonparametric identification approaches. hickman2012structural provide a recent review.} The object of interest is the probability density function (PDF) of latent valuations, which can then be used for a variety of policy counterfactuals such as the optimal reserve price (paarsch1997deriving and li2003semiparametric).\footnote{In auctions, the valuation of a bidder is simply his or her willingness to pay for the object.}

The focus on nonparametric estimation is due to several reasons. First, in empirical applications, large auction datasets are often available.\footnote{E.g., kawai2014detecting utilize a dataset of $40,000$ auctions in their study of collusion in Japan, while augenblick2015sunk employs a dataset $160,000$ penny auctions.} Second, nonparametric methods are flexible since no functional form assumptions are needed. Third, in auctions, nonparametric estimators are built directly from the identification arguments, and are often easy to implement.\footnote{See athey2002identification for a number of additional identification results.}

GPV proposed a two-step nonparametric estimator of the PDF of valuations in first-price auctions, and showed that it is uniformly consistent and attains the minimax optimal uniform convergence rate. However, it has been an open question whether this estimator also converges in a distributional sense, which would allow empirical researchers to perform inferences, thereby increasing the scope of applications.

Recently, Marmer_Shneyerov_Quantile_Auctions developed an alternative quantile-based estimator of the PDF of valuations and showed its asymptotic normality.\footnote{More recently, quantile methods have been used in the context of auctions in gimenes2016quantile, gimenes2017econometrics, liu2017nonparametric, and luo2017integrated.} However, the GPV estimator is well established in the literature and is used in all the empirical applications we are aware of. Moreover, our results imply that the GPV estimator has a smaller asymptotic variance than that of the quantile-based estimator, as long as the two estimators use the same second-order kernel.

Inference and the closely related problem of nonparametric testing in structural auction models are important and have been receiving increasing attention in the literature. Beginning with the fundamental haile2003nonparametric's test for common values, recent contributions include testing for the monotonicity of bidding strategies (liuvuong2013), endogenous entry (li2009entry and marmer2013model), common versus private values (hill2013there), the affiliation of bidder valuations (jun2010consistent, li2010testing and de2010testing), and inferences on bidder risk attitudes (fang2014inference). In the absence of the asymptotic distribution framework for the GPV estimator, these papers have adopted problem-specific approaches in each case.

The first main result we show in this paper is that the GPV estimator is asymptotically normal. The key difficulty is the presence of the nonparametric first step, which provides nonparametric estimates of the valuations in each auction. In the second step, the kernel density estimator is applied to those estimates, rather than the true valuations. This creates a unique challenge, to our knowledge not previously addressed in the econometrics literature. Our main insight is that the leading term in an asymptotic expansion of the estimator can be viewed as a V-statistic with a kernel that depends on the bandwidth. A projection argument shows that the distribution of this V-statistic is asymptotically normal. Using maximal inequalities for empirical processes and U-processes developed in recent literature, we show that the remainder term is uniformly negligible. The proof is rather long due to an intricate nature of the estimator, and involves some delicate steps.

Note that a working paper version of GPV GPV1995 also has an asymptotic normality result for the GPV estimator. However, the result therein is of a limited nature as it relies on a particular choice of tuning parameters, which insures that only the second stage of the estimating procedure contributes to the asymptotic variance. Thus, in their approach, the uncertainty due to the estimation of valuations in the first stage can be ignored asymptotically, which is achieved by applying different rates of smoothing of auction-specific covariates at both stages. The approach is restrictive in two respects: (i) While GPV's smoothing strategy makes first-stage estimation errors negligible asymptotically, in finite samples their contribution to the variance may still be significant. Our approach takes into account the contribution of both stages and, as a result, is more accurate in finite samples. (ii) Equally importantly, their approach cannot be applied in cases with no auction-specific covariates or when covariates are modeled semi-parametrically as in haile2003nonparametric. Note that treating auction-specific characteristics semi-parametrically is particularly appealing to practitioners.

One unusual feature of our asymptotic normality result concerns the form of the asymptotic variance of the GPV estimator. Typically, the asymptotic variances of kernel density estimators depend on the integral of the squared kernel, which is a known constant that can be easily computed analytically or numerically.\footnote{See, e.g., li2007net.} However, in the case of the GPV estimator, this constant is replaced by a convoluted integral transformation that involves the kernel function, its derivative, and the derivatives of the bidding strategy. This is a consequence of the two-step nature of the GPV estimator and happens due to the impact of the estimation errors from the first stage of the procedure on the distribution of the estimator. Since the bidding strategy is unknown, this fact complicates estimation of the asymptotic variance of the GPV estimator.

Our second contribution is to propose a consistent estimator of the asymptotic variance that avoids estimation of the bidding strategy and its derivative. Its uniform rate of convergence is established using the maximal inequalities. Our third contribution is to show the validity of the percentile bootstrap for the GPV estimator, which allows constructing confidence intervals without estimation of the asymptotic variance.

Our pointwise asymptotic normality results can be used for inference on the optimal reserve price, as the latter is determined by a nonlinear equation in the PDF of valuations haile2003iim. In our fourth contribution, however, we extend the pointwise results and develop valid uniform confidence bands for the PDF. The uniform confidence bands can be used, e.g., for specification of valuations' density. The extension utilizes the uniform rates of convergence of the remainder terms in our V-statistic approximation of the GPV estimator and its Hoeffding decomposition; it also relies on Gaussian anti-concentration inequalities and Gaussian coupling theorems developed in recent literature. This approach, referred to as the Intermediate Gaussian Approximation (IGA, hereafter) in the literature, is based on the seminal work of chernozhukov2014anti,chernozhukov2014gaussian,chernozhukov2016empirical. chernozhukov2014gaussian showed that although a random function based on nonparametric estimation errors does not typically weakly converge to any tight Gaussian random element, under certain conditions the supremum of its studentized version can be often approximated by the supremum of a tight Gaussian random element, the distribution of which changes with the sample size. chernozhukov2016empirical,chernozhukov2014anti showed that under certain conditions the distribution of the Gaussian supremum can be approximated by bootstrapping, and the bootstrap consistency can be shown by applying the coupling theorems and the Gaussian anti-concentration inequality developed in these papers.\footnote{See, e.g., Kato_Sasaki_2,Kato_Sasaki_1 for recent applications of these theorems for constructing confidence bands for different nonparametric curves. } Our paper is one of the first applications of these results. Our Monte Carlo simulation results show that the IGA approach produces confidence bands with excellent finite-sample coverage properties.

Our paper is also related to the recent literature on nonparametrically generated regressors in nonparametric regression. See, e.g., rilstone1996nonparametric, pinkse2001nonparametric, and mammen2012nonparametric. Note, however, that while that literature is concerned with nonparametrically estimated exogenous covariates, we deal with kernel estimation of the density of a nonparametrically generated “dependent” variable, potentially in presence of observable conditioning variables.

The rest of the paper proceeds as follows. Section (ref) introduces the data-generating process (DGP) and describes the GPV estimator in detail. Due to complexity of the estimator, in Section (ref) we show the asymptotic normality of the GPV estimator in a simplified model that has a constant number of bidders across auctions and no auction-specific heterogeneity. Such a simplification allows us to present the main ideas in a more transparent fashion. In Section (ref), we derive an estimator for the asymptotic variance and establish its uniform rate of convergence. We also show consistency of the percentile bootstrap confidence intervals. Section (ref) provides results on constructing valid confidence bands within the same simplified framework. Proofs of the results in Sections (ref)\textendash (ref) are given in the Appendix. Section (ref) provides corresponding theorems in the general model with a random number of bidders and auction-specific heterogeneity. The proofs of these results can be found in the Supplement (included). Section (ref) discusses how our approach can be extended to auctions with binding reserve prices. We report the results from our Monte Carlo study in Section (ref). Section (ref) concludes.

ntnbold$a\coloneqq b$” is understood as “$a$ is defined by $b$”. “$a\eqqcolon b$” is understood as “$b$ is defined by $a$”. $\mathbbm{1}\left(\cdot\right)$ denotes the indicator function, and we also denote $\mathbbm{1}_{A}\coloneqq\mathbbm{1}\left(\cdot\in A\right)$. Let $\mathbb{H}\left(\boldsymbol{c},r\right)$ denote a closed hypercube centered at some real vector $\boldsymbol{c}$ with edge $r$. Let $\boldsymbol{c}^{\mathrm{T}}$ denote the transpose of $\boldsymbol{c}$. “$\overset{d}{=}$” means “equal in distribution”. Let $\ell^{\infty}\left(A\right)$ be the class of bounded functions defined on $A$. For any $f\in\ell^{\infty}\left(A\right)$, let $\left\Vert f\right\Vert _{A}\coloneqq\underset{x\in A}{\mathrm{sup}}\left|f\left(x\right)\right|$ be the sup-norm.

Data-Generating Process (DGP) and the GPV Estimator

The econometrician observes data from $L$ auctions. Let $\boldsymbol{X}_{l}$ denote the $d$-dimensional relevant characteristics for the object in the $l$-th auction. Let $N_{l}$ denote the number of bidders in the $l$-th auction. Let $B_{il}$ denote the bid submitted by the $i$-th bidder in the $l$-th auction. The data observed by the econometrician is given by $\left\{ \left(B_{il},\boldsymbol{X}_{l},N_{l}\right):i=1,...,N_{l},\,l=1,...,L\right\} .$ Unobserved bidders' valuations of the $l$-th auctioned object are denoted by $\left\{ V_{il}:i=1,...,N_{l},\,l=1,...,L\right\} .$ The following assumption describes the DGP.\footnote{Assumption (ref) is similar to Assumptions A1 and A2 of GPV and Marmer_Shneyerov_Quantile_Auctions. Part (f) imposes the condition that the valuations and the random number of bidders are independent conditionally on the characteristics. See Footnote 14 of GPV.}

assumption[DGP] (a) $\left\{ \left(\boldsymbol{X}_{l},N_{l}\right):l=1,...,L\right\} $ are i.i.d. (b) The marginal PDF of $\boldsymbol{X}_{1}$, denoted by $\varphi\left(\cdot\right)$, is strictly positive and continuous on its support $\mathcal{X}\coloneqq\left[\underline{x},\overline{x}\right]^{d}$ for some $\underline{x}<\overline{x}$ assumed to be known and admits up to $R+1$ ($R\geq2$) continuous partial derivatives. (c) The conditional probability mass function of $N_{1}$ given $\boldsymbol{X}_{1}=\boldsymbol{x}$, denoted by $\pi\left(\cdot|\boldsymbol{x}\right)$, has a known support $\mathcal{N}\coloneqq\left\{ \underline{n},...,\overline{n}\right\} $ for all $\boldsymbol{x}\in\mathcal{X}$, $\underline{n}\geq2$. (d) For all $n\in\mathcal{N}$, $\pi\left(n|\cdot\right)$ is strictly positive and admits up to $R+1$ continuous partial derivatives. (e) $\left(n,\boldsymbol{x}\right)\mapsto\pi\left(n|\boldsymbol{x}\right)\varphi\left(\boldsymbol{x}\right)$ is bounded above and away from zero on its support $\mathcal{N}\times\mathcal{X}$. \textup{(f) For each $l=1,...,L$, given $\boldsymbol{X}_{l}=\boldsymbol{x}$ and $N_{l}=n$, $\left\{ V_{il}:i=1,...,n\right\} $ are i.i.d. with conditional PDF $f\left(\cdot|\boldsymbol{x}\right)$ and conditional cumulative distributional function (CDF) $F\left(\cdot|\boldsymbol{x}\right)$.} \textup{(g) For each $n\in\mathcal{N}$, the support of $\left(V_{11},\boldsymbol{X}_{1}\right)$ is $\mathcal{S}_{V,\boldsymbol{X}}\coloneqq\left\{ \left(v,\boldsymbol{x}\right):\boldsymbol{x}\in\mathcal{X},\,v\in\left[\underline{v}\left(\boldsymbol{x}\right),\overline{v}\left(\boldsymbol{x}\right)\right]\right\} ,$ with some positive boundary functions $\underline{v}\left(\cdot\right)$ and $\overline{v}\left(\cdot\right)$.} \textup{(h) $f\left(\cdot|\cdot\right)$ is strictly positive and bounded away from zero and admits up to $R$ continuous partial derivatives on $\mathcal{S}_{V,\boldsymbol{X}}$.}

Bidders' valuations are not directly observable. Following GPV, we assume that $B_{il}$ is the equilibrium bid of risk-neutral bidder $i$ submitted in the $l$-th auction. Therefore the valuations are linked to the observed bids through the Bayesian Nash equilibrium (BNE) bidding strategy:

equation[equation omitted — 322 chars of source]

\footnotetext{See Equations (1) and (8) of GPV.}Under Assumptions (ref)(a) and (ref)(f), $\left\{ B_{il}:i=1,...,N_{l}\right\} $ are conditionally i.i.d. draws given $\boldsymbol{X}_{l}$ and $N_{l}$. Let $\overline{b}\left(\boldsymbol{x},n\right)\coloneqq s\left(\overline{v}\left(\boldsymbol{x}\right),\boldsymbol{x},n\right)$ and $\underline{b}\left(\boldsymbol{x}\right)\coloneqq\underline{v}\left(\boldsymbol{x}\right)$. Proposition 1(i) of GPV shows that the support of $\left(B_{il},\boldsymbol{X}_{l},N_{l}\right)$ is $\left\{ \left(b,\boldsymbol{x},n\right):n\in\mathcal{N},\,\left(b,\boldsymbol{x}\right)\in\mathcal{S}_{B,\boldsymbol{X}}^{n}\right\} ,$ where $\mathcal{S}_{B,\boldsymbol{X}}^{n}\coloneqq\left\{ \left(b,\boldsymbol{x}\right):\boldsymbol{x}\in\mathcal{X},\,b\in\left[\underline{b}\left(\boldsymbol{x}\right),\overline{b}\left(\boldsymbol{x},n\right)\right]\right\} .$

Let $G\left(\cdot|\boldsymbol{x},n\right)$ denote the conditional CDF of $B_{il}$ given $\boldsymbol{X}_{l}=\boldsymbol{x}$ and $N_{l}=n$. Let $g\left(\cdot|\boldsymbol{x},n\right)$ be the corresponding conditional PDF. GPV established identification of the inverse bidding strategy:

equation[equation omitted — 239 chars of source]

By replacing $G\left(\cdot|\cdot,\cdot\right)$ and $g\left(\cdot|\cdot,\cdot\right)$ in ((ref)) with their nonparametric estimators, GPV proposed an estimator of $\xi\left(\cdot,\cdot,\cdot\right)$, denoted by $\widehat{\xi}\left(\cdot,\cdot,\cdot\right)$. The GPV estimator of $f\left(v|\boldsymbol{x}\right)$ is the kernel density estimator that, in place of the true valuations, uses the so-called pseudo valuations $\left\{ \widehat{V}_{il}\coloneqq\widehat{\xi}\left(B_{il},\boldsymbol{X}_{l},N_{l}\right):i=1,...,N_{l},\,l=1,...,L\right\} $.

Below we provide the details of GPV's estimation procedure. Let $K_{0}$ and $K_{1}$ be univariate kernel functions of different orders satisfying the following assumption:

assumption[Kernel] (a) $K_{0}$ and $K_{1}$ are symmetric, compactly supported on $[-1,1]$ and twice continuously differentiable on $\mathbb{R}$ with Lipschitz derivatives. (b) $K_{0}$ is of order $R$, and $K_{1}$ is of order $1+R$: $\int K_{0}(u)\mathrm{d}u=1$ and $\int u^{k}K_{0}(u)\mathrm{d}u=0$ for $k=1,...,R-1$; $\int K_{1}(u)\mathrm{d}u=1$ and $\int u^{k}K_{1}(u)\mathrm{d}u=0$ for $k=1,...,R$.

With $K_{g}\coloneqq K_{1}$ and the multi-dimensional product kernels \[ K_{f}\left(v,\boldsymbol{x}\right)\coloneqq K_{0}\left(v\right)\cdot\prod_{k=1}^{d}K_{0}\left(x_{k}\right)\textrm{ and }K_{\boldsymbol{X}}\left(\boldsymbol{x}\right)\coloneqq\prod_{k=1}^{d}K_{1}\left(x_{k}\right),\textrm{ for }v\in\mathbb{R},\,\boldsymbol{x}=\left(x_{1},...,x_{d}\right)\in\mathbb{R}^{d}, \] define the following nonparametric estimators: \[ \widehat{\varphi}\left(\boldsymbol{x}\right)\coloneqq\frac{1}{L}\sum_{l=1}^{L}\frac{1}{h^{d}}K_{\boldsymbol{X}}\left(\frac{\boldsymbol{X}_{l}-\boldsymbol{x}}{h}\right)\textrm{ and }\widehat{\pi}\left(n|\boldsymbol{x}\right)\coloneqq\frac{1}{\widehat{\varphi}\left(\boldsymbol{x}\right)L}\sum_{l=1}^{L}\mathbbm{1}\left(N_{l}=n\right)\frac{1}{h^{d}}K_{\boldsymbol{X}}\left(\frac{\boldsymbol{X}_{l}-\boldsymbol{x}}{h}\right), \] where $\widehat{\varphi}\left(\cdot\right)$ is the kernel density estimator of $\varphi$ and $\widehat{\pi}\left(\cdot|\cdot\right)$ is the Nadaraya-Watson estimator of the conditional probability mass function $\pi\left(\cdot|\cdot\right)$. Based on these, we define below the nonparametric estimators of the conditional CDF and PDF of the bids:

gather*[gather* omitted — 724 chars of source]

Consider a partition of $\mathbb{R}^{d}$ with generic half-open hypercubes of side $h_{\partial}>0$: \[ \Pi_{k_{1},...,k_{d}}\coloneqq\left[k_{1}h_{\partial},\left(k_{1}+1\right)h_{\partial}\right)\times\cdots\times\left[k_{d}h_{\partial},\left(k_{d}+1\right)h_{\partial}\right), \] where $\left(k_{1},...,k_{d}\right)$ runs over $\mathbb{Z}^{d}$. Let $\Pi_{h_{\partial}}\left(\boldsymbol{x}\right)$ denote the hypercube that contains $\boldsymbol{x}$ in this partition. Define

gather[gather omitted — 459 chars of source]

to be the estimators of the boundaries of the support. Note that the estimators of the boundaries are super consistent. Let $\mathcal{\widehat{S}}_{B,\boldsymbol{X}}^{n}\coloneqq\left\{ \left(b,\boldsymbol{x}\right):\boldsymbol{x}\in\mathcal{X},\,b\in\left[\widehat{\underline{b}}\left(\boldsymbol{x}\right),\widehat{\overline{b}}\left(\boldsymbol{x},n\right)\right]\right\} .$ The support of $\left(B_{il},\boldsymbol{X}_{l},N_{l}\right)$ then can be estimated by $\left\{ \left(b,\boldsymbol{x},n\right):n\in\mathcal{N},\,\left(b,\boldsymbol{x}\right)\in\mathcal{\widehat{S}}_{B,\boldsymbol{X}}^{n}\right\} .$

The kernel density estimator $\widehat{g}\left(b|\boldsymbol{x},n\right)$ is asymptotically biased when $\left(b,\boldsymbol{x}\right)$ is near the boundaries of the support. GPV suggested that trimming should be applied to the observations near the estimated boundaries using the trimming factor $\mathbb{T}_{il}\coloneqq\mathbbm{1}\left(\mathbb{H}\left(\left(B_{il},\boldsymbol{X}_{l}\right),2h\right)\subseteq\mathcal{\widehat{S}}_{B,\boldsymbol{X}}^{N_{l}}\right).$ The two-step nonparametric estimator of $f\left(v|\boldsymbol{x}\right)$ developed by GPV is

equation[equation omitted — 327 chars of source]

For deriving the asymptotic properties of the GPV estimator, we make the following assumption on the bandwidths $h$ and $h_{\partial}$.\footnote{Assumption (ref)(a) is the same as the assumption on the rate of bandwidth for Marmer_Shneyerov_Quantile_Auctions's quantile-based estimator. See Assumption 3 therein. }

assumption[Bandwidth] (a) The bandwidth $h$ is of the form $h=\lambda_{1}L^{-\gamma_{1}}$, for some strictly positive constants $\lambda_{1}$ and $\gamma_{1}$ satisfying $\nicefrac{1}{\left(2R+3+d\right)}\leq\gamma_{1}<\nicefrac{1}{\left(3+d\right)}$. (b) When $d>0$, the “boundary” bandwidth is of the form $h_{\partial}=\lambda_{\partial}\left(\nicefrac{\mathrm{log}\left(L\right)}{L}\right)^{\nicefrac{1}{\left(1+d\right)}}$, where $\lambda_{\partial}$ is a strictly positive constant.

GPV showed that the optimal uniform convergence rate of their estimator is attained when the bandwidth $h$ is of order $O\left(\left(\nicefrac{\mathrm{log}\left(L\right)}{L}\right)^{\nicefrac{1}{\left(2R+3+d\right)}}\right)$. Note that the bandwidth in Assumption (ref) is of smaller order. Under-smoothing imposed in Assumption (ref) is needed to control the asymptotic bias of the GPV estimator, which is important for the validity of inference.

Asymptotic Normality of the GPV Estimator

For clarity of the presentation of the main ideas and results, in this section we first establish pointwise asymptotic normality of the GPV estimator in a simplified version of the model that has a fixed number of bidders and no auction-specific heterogeneity. When there are covariates capturing auction-specific heterogeneity present, these results can be used by treating the covariates additively semi-parametrically as in haile2003nonparametric. In that case, there is no kernel smoothing over the covariates, as the GPV procedure would be applied to the “homogenized” bids, which are constructed as residuals from the parametric regression of the bids against the covariates.

In the simplified model, the econometrician observes data on bids in $L$ identical auctions, with a fixed number of bidders $N$ in each auction: $\left\{ B_{il}:i=1,\ldots,N,\,l=1,\ldots,L\right\} .$ Under Assumption (ref), the valuations $\left\{ V_{il}:i=1,\ldots,N,\,l=1,\ldots,L\right\} $ are i.i.d. with a compact support $\left[\underline{v},\overline{v}\right]\subseteq\mathbb{R}_{+}$, PDF $f$ and CDF $F$. The object of interest is the PDF of the valuation at interior points of $\left[\underline{v},\overline{v}\right]$. Suppose that $v_{l}>\underline{v}$, $v_{u}<\overline{v}$ and $I\coloneqq\left[v_{l},v_{u}\right]$ is an inner closed sub-interval of $\left[\underline{v},\overline{v}\right]$. Fix \[ \overline{\delta}\coloneqq\mathrm{min}\left\{ \nicefrac{\left(\overline{v}-v_{u}\right)}{2},\nicefrac{\left(v_{l}-\underline{v}\right)}{2}\right\} . \]

Under Assumption (ref), $f$ is strictly positive and bounded away from zero on its support and admits at least $R$ continuous derivatives. Lemma A1 of GPV showed that under Assumption (ref), the BNE bidding strategy is strictly increasing and $R+1$ times continuously differentiable. In this simplified framework, the inverse of the BNE bidding strategy is

equation[equation omitted — 127 chars of source]

where $G$ and $g$ are the CDF and PDF of bids respectively. Denote $\overline{b}\coloneqq s\left(\overline{v}\right)$ and $\underline{b}\coloneqq s\left(\underline{v}\right)$. Proposition 1(ii) of GPV shows that under Assumption (ref), $g$ is also bounded away from zero on its support $\left[\underline{b},\overline{b}\right]$:

equation[equation omitted — 170 chars of source]

The inverse bidding strategy ((ref)) can be estimated by \[ \widehat{\xi}\left(b\right)\coloneqq b+\frac{1}{N-1}\frac{\widehat{G}\left(b\right)}{\widehat{g}\left(b\right)}, \] where we use the usual nonparametric estimators of $G$ and $g$: \[ \widehat{G}\left(b\right)\coloneqq\frac{1}{N\cdot L}\sum_{i,l}\mathbbm{1}\left(B_{il}\leq b\right)\textrm{ and }\widehat{g}\left(b\right)\coloneqq\frac{1}{N\cdot L}\sum_{i,l}\frac{1}{h}K_{g}\left(\frac{B_{il}-b}{h}\right), \] where $\sum_{i,l}$ is understood as $\sum_{l=1}^{L}\sum_{i=1}^{N}$.

Let $\widehat{\overline{b}}\coloneqq\mathrm{max}\left\{ B_{il}:i=1,\ldots,N,\,l=1,...,L\right\} $, and $\widehat{\underline{b}}\coloneqq\mathrm{min}\left\{ B_{il}:i=1,\ldots,N,\,l=1,...,L\right\} $. The trimming factor is now simply $\mathbb{T}_{il}\coloneqq\mathbbm{1}\left(\widehat{\underline{b}}+h\leq B_{il}\leq\widehat{\overline{b}}-h\right)$. The GPV estimator of $f\left(v\right)$ is now given by \[ \widehat{f}_{GPV}\left(v\right)=\frac{1}{N\cdot L}\sum_{i,l}\mathbb{T}_{il}\frac{1}{h}K_{f}\left(\frac{\widehat{V}_{il}-v}{h}\right), \] where $K_{f}=K_{0}$ in this simplified framework.

We derive the following stochastic expansion of $\widehat{f}_{GPV}\left(v\right)$ around $f\left(v\right)$:

equation[equation omitted — 334 chars of source]

where $\widetilde{\mathbb{T}}_{il}\coloneqq\mathbbm{1}\left(\left|V_{il}-v\right|\leq\overline{\delta}\right)$ is an infeasible trimming factor and the remainder term is uniform in $v\in I$. In the above expression, the derivative $K_{f}'$ of the kernel function appears due to the linearization of $K_{f}\left(\nicefrac{\left(\widehat{V}_{il}-v\right)}{h}\right)$ around $K_{f}\left(\nicefrac{\left(V_{il}-v\right)}{h}\right)$. The result in ((ref)) shows that the distribution of the GPV estimator depends not only on the variation in $V_{il}$'s, but also on the estimation errors of pseudo valuations. In other words, the errors from estimation of the inverse bidding strategy affect the asymptotic distribution of the GPV estimator.

Since $\widehat{G}$ has a faster rate of convergence than $\widehat{g}$, the discrepancy between $V_{il}$ and $\widehat{V}_{il}$ depends on that between the true PDF $g\left(B_{il}\right)$ and the estimated PDF $\widehat{g}\left(B_{il}\right)$, which in turn depends on the averaged discrepancy between $K_{g}\left(\nicefrac{\left(B_{jm}-B_{il}\right)}{h}\right)$ and $g\left(B_{il}\right)$, where the averaging is across $B_{jm}$'s. Lemma (ref) establishes a further asymptotic expansion for the GPV estimator:

equation[equation omitted — 275 chars of source]

where the remainder term is uniform in $v\in I$, and

equation[equation omitted — 270 chars of source]

For any fixed $v$, the leading term in ((ref)) is a V-statistic with a kernel that depends on the bandwidth $h$. We now apply Hoeffding decomposition to this leading term. Define

eqnarray[eqnarray omitted — 362 chars of source]

and further, \[ \mathcal{M}_{2}\left(b;v\right)\coloneqq\int\mathcal{M}\left(b',b;v\right)\mathrm{d}G\left(b'\right),\textrm{ and }\mu_{\mathcal{M}}\left(v\right)\coloneqq\int\int\mathcal{M}\left(b,b';v\right)\mathrm{d}G\left(b\right)\mathrm{d}G\left(b'\right). \] Note that $\mu_{\mathcal{M}}\left(v\right)=\mathrm{E}\left[\mathcal{M}_{1}\left(B_{11};v\right)\right]=\mathrm{E}\left[\mathcal{M}_{2}\left(B_{11};v\right)\right]$. The Hoeffding decomposition yields

align[align omitted — 941 chars of source]

In the proof of Theorem (ref) below, we use results for empirical processes and U-processes to show that the terms in the third and fourth lines of ((ref)) and $\left(N\cdot L\right)^{-1}\sum_{i,l}\mathcal{M}_{1}\left(B_{il};v\right)$ are asymptotically negligible uniformly in $v\in I$. As is apparent from the definition of $\mathcal{M}_{1}$ in ((ref)), the contribution of the $\mathcal{M}_{1}\left(b;v\right)$ terms is negligible because they depend on the difference between the expectation $\mathrm{E}\left[h^{-1}K_{g}\left(\nicefrac{\left(B_{il}-b\right)}{h}\right)\right]$ and $g\left(b\right)$, i.e., the bias of the kernel density estimator, which is of order $O\left(h^{1+R}\right)$. Thus, the asymptotic distribution of the GPV estimator is driven solely by

equation[equation omitted — 211 chars of source]

where $\mathcal{M}_{2}\left(B_{il};v\right)-\mu_{\mathcal{M}}\left(v\right)$, $i=1,...,N$, $l=1,...,L$ are independent, zero-mean and depend on the bandwidth.

We also show in the proof of Theorem (ref) that the rescaled variance of ((ref)) satisfies

align[align omitted — 591 chars of source]

where the remainder term is uniform in $v\in I$. Thus, the asymptotic variance of the GPV estimator is the limit of the leading term in ((ref)) as $h\downarrow0$. Note that

multline[multline omitted — 441 chars of source]

We have the following result.

thm[Asymptotic Normality] Suppose Assumptions (ref) - (ref) hold. Then for any interior point $v\in\left(\underline{v},\overline{v}\right)$, $\left(Lh^{3}\right)^{\nicefrac{1}{2}}\left(\widehat{f}_{GPV}\left(v\right)-f\left(v\right)\right)\rightarrow_{d}\mathrm{N}\left(0,\mathrm{V}_{GPV}\left(v\right)\right),$ where \begin{equation} \mathrm{V}_{GPV}\left(v\right)\coloneqq\frac{1}{N\left(N-1\right)^{2}}\frac{F\left(v\right)^{2}f\left(v\right)^{2}}{g\left(s\left(v\right)\right)^{3}}\int\left\{ \int K_{f}'\left(u\right)K_{g}\left(w-s'\left(v\right)u\right)\mathrm{d}u\right\} ^{2}\mathrm{d}w. \end{equation}
remboldThe asymptotic variance $\mathrm{V}_{GPV}\left(v\right)$ has an interesting and non-standard feature. Typically, the asymptotic variance of a kernel estimator involves a constant which is a simple functional of the kernel and does not involve any elements of the DGP. However, in the case of the GPV estimator, the constant has a convolution form. Moreover, this constant also involves the derivative of the unknown bidding function. As can be seen from ((ref)), the convolution form is due to the fact that the variance of the GPV estimator is determined not only by the variation of $V_{il}$, but also by the estimation errors $\widehat{V}_{il}-V_{il}$. The $s'(v)$ term appears due to averaging of $\xi(b)$'s in a small neighborhood of $v$, as one can see from ((ref)). The presence of the $s'(v)$ term inside the integral in ((ref)) can cause additional complications in estimation of the asymptotic variance $\mathrm{V}_{GPV}(v)$, as the econometrician now also has to estimate the functional of the kernel function. Potentially, one could estimate $s'(v)$ and then estimate the convolution by the plug-in approach. However, we show in the next section that the asymptotic variance can be estimated directly without separate estimation of $s'(v)$ and the convolution by using the sample analogue of the expression in the second line of ((ref)).
rembold[Extensions]Theorem (ref) also implies asymptotic normality of some modified GPV estimators. henderson_2012_JOE proposed to modify the standard GPV approach by estimating $\xi$ under a monotonicity constraint implied by the structural model. However, if the auction model is correctly specified, the unconstrained estimator of $\xi$ will be monotone with probability approaching one. Hence, we expect the standard GPV estimator and the estimator proposed in henderson_2012_JOE to be first-order asymptotically equivalent. hickman_hubbard_2014_JAE proposed another modified GPV estimator by replacing sample trimming used in the standard GPV procedure with a boundary correction. Their method uses a boundary-bias-corrected kernel estimator in estimation of the density of bids. The corrected estimator is uniformly consistent over the entire support of the distribution of bids. When estimating the PDF of valuations away from the boundaries, our proof of ((ref)) can be adapted to show that the same V-statistic approximation holds for the estimator in hickman_hubbard_2014_JAE. Hence, their estimator is first-order asymptotically equivalent to the standard GPV estimator, when $v$ is chosen away from the boundaries.
rembold[Asymptotic Bias]In the proof of Theorem (ref), we incorporate the bias term: \begin{multline} \left(Lh^{3}\right)^{\nicefrac{1}{2}}\left(\widehat{f}_{GPV}\left(v\right)-f\left(v\right)-\frac{1}{R!}f^{\left(R\right)}\left(v\right)\left(\int K_{f}\left(u\right)u^{R}\mathrm{d}u\right)h^{R}+o\left(h^{R}\right)\right)\\ \rightarrow_{d}\mathrm{N}\left(0,\mathrm{V}_{GPV}\left(v\right)\right). \end{multline} The leading bias term of the GPV estimator is the same as that of the infeasible estimator constructed using the unobserved true valuations. This is due to the following feature of the first-price auction model: at interior points, the bid density has $1+R$ continuous derivatives rather than $R$. It is natural to incorporate this structural feature in the estimation procedure by using a higher-order kernel in the first stage.
rembold[Comparison with the Quantile-Based Estimator]The quantile-based (QB) estimator of Marmer_Shneyerov_Quantile_Auctions does not require estimation of latent valuations, and instead relies on a direct representation of the PDF of valuations using the distribution functions of observable bids. While the two estimators have the same rate of convergence, the limiting distribution of the GPV estimator shows that it indeed improves on the QB estimator in the following sense. One can show that the GPV estimator has a smaller asymptotic variance than that of the quantile-based estimator, as long as the two estimators use the same second-order kernel functions. Suppose now $K_{f}=K_{g}=K$, for some second-order kernel function $K$. Let $\mathrm{V}_{QB}\left(v\right)$ be the asymptotic variance of the QB estimator in the simplest setting without auction-specific heterogeneity.\footnote{See Marmer_Shneyerov_Quantile_Auctions for the expression of the asymptotic variance of the quantile-based estimator.} By Jensen's inequality and standard calculus techniques, it can be easily shown that $\nicefrac{\mathrm{V}_{QB}\left(v\right)}{\mathrm{V}_{GPV}\left(v\right)}\geq1$. The details of the proof can be found in the Supplement. Consider the following family of PDF's: $f_{\theta}(v)=\theta v^{\theta-1}\cdot\mathbbm{1}\left(0\leq v\leq1\right)$ for some $\theta>0$. The corresponding BNE bidding strategy is $s(v)=\left(1-\left(\theta(N-1)+1\right)^{-1}\right)v$. Note that in this case $s'(v)$ is constant, and we can compute the ratio $\nicefrac{\mathrm{V}_{QB}\left(v\right)}{\mathrm{V}_{GPV}\left(v\right)}$ analytically. In the case of the triweight kernel $K$, ratio is found to be, e.g., 1.3259 when $\left(\theta,N\right)=\left(1,2\right)$ and 2.3038 when $\left(\theta,N\right)=\left(2,7\right)$. Thus, depending on the model, the GPV estimator could be substantially more precise than the QB estimator.\footnote{We believe that this finding is of interest in a more general context outside of the empirical auctions literature. It illustrates that two-step nonparametric estimators can outperform more direct estimators that avoid first-stage estimation of latent variables.}

Pointwise Confidence Intervals

If the asymptotic variance $\mathrm{V}_{GPV}(v)$ can be consistently estimated by some estimator $\widehat{\mathrm{V}}_{GPV}\left(v\right)$, one can construct an asymptotically valid pointwise confidence interval for $f\left(v\right)$ as

equation[equation omitted — 333 chars of source]

where $z_{1-\nicefrac{\alpha}{2}}$ denotes the $1-\nicefrac{\alpha}{2}$ quantile of the standard normal distribution.

While the formula for the asymptotic variance in ((ref)) can be used for plug-in estimation of $\mathrm{V}_{GPV}(v)$, such an estimator would be difficult to implement in practice. Firstly, it would require estimating the bidding strategy and its derivative. Secondly, even with an estimate of $s'\left(v\right)$, it is not always easy to compute analytically the double integral in the definition of $\mathrm{V}_{GPV}\left(v\right)$. This issue becomes even more severe when there is auction-specific heterogeneity, as we discuss in Section (ref). In that case, one would need to evaluate a multidimensional integral.

To avoid those issues, we propose an alternative approach to estimation of the asymptotic variance. As we discuss in the previous section, the asymptotic variance of the GPV estimator is the limit of the expression in ((ref)). The leading term on the right-hand side of ((ref)) can be estimated using a U-type-statistic, while replacing the unknown $G$, $g$, and $\xi$ with $\widehat{G}$, $\widehat{g}$, and $\widehat{\xi}$ respectively. The resulting estimator is given by

eqnarray[eqnarray omitted — 668 chars of source]

The estimator avoids estimation of the bidding strategy and its derivative and evaluation of multidimensional integrals. It is very easily implementable in practice since it depends only on the bids, $\widehat{G}$, $\widehat{g}$, and the pseudo valuations. The next theorem shows consistency of the proposed estimator and provides an estimate of its uniform convergence rate.

thm[Variance Estimation] Suppose Assumptions (ref) - (ref) hold. Then, \[ \underset{v\in I}{\mathrm{sup}}\left|\widehat{\mathrm{V}}_{GPV}\left(v\right)-\mathrm{V}_{\mathcal{M}}\left(v\right)\right|=O_{p}\left(\left(\frac{\mathrm{log}\left(L\right)}{Lh^{3}}\right)^{\nicefrac{1}{2}}+h^{R}\right). \]

An alternative to the confidence interval ((ref)) is the bootstrap. We show below that the bootstrap approximation to the distribution of $S\left(v\right)\coloneqq\left(Lh^{3}\right)^{\nicefrac{1}{2}}\left(\widehat{f}_{GPV}\left(v\right)-f\left(v\right)\right)$ is asymptotically valid. Our focus is on the percentile bootstrap as it does not require estimation of the asymptotic variance, which makes it fairly popular among practitioners.

Let $\left\{ B_{il}^{*}:i=1,\ldots,N,l=1,\ldots,L\right\} $ denote the bootstrap sample, i.e., a set of independent random variables drawn from the distribution $\widehat{G}$ conditionally on the original sample of bids. Let $\widehat{G}^{*}$ and $\widehat{g}^{*}$ denote the bootstrap analogues of $\widehat{G}$ and $\widehat{g}$ respectively: they are constructed by following exactly the same procedure as that for constructing $\widehat{G}$ and $\widehat{g}$, however using the (empirical) bootstrap sample instead of the original sample. Let $\widehat{\xi}^{*}$ be the bootstrap analogue of $\widehat{\xi}$ defined using $\widehat{G}^{*}$ and $\widehat{g}^{*}$ in place of $\widehat{G}$ and $\widehat{g}$. We generate bootstrap samples of pseudo values as $\widehat{V}_{il}^{*}\coloneqq\widehat{\xi}^{*}\left(B_{il}^{*}\right)$. Lastly, we construct a bootstrap analogue of $\widehat{f}_{GPV}\left(v\right)$: \[ \widehat{f}_{GPV}^{*}\left(v\right)\coloneqq\frac{1}{N\cdot L}\sum_{i,l}\mathbb{T}_{il}^{*}\frac{1}{h}K_{f}\left(\frac{\widehat{V}_{il}^{*}-v}{h}\right), \] where $\mathbb{T}_{il}^{*}\coloneqq\mathbbm{1}\left(\widehat{\underline{b}}+h\leq B_{il}^{*}\leq\widehat{\overline{b}}-h\right)$.

Let $q_{\tau}^{*}(v)$ be the $\tau$-th quantile of the conditional distribution of $\widehat{f}_{GPV}^{*}\left(v\right)$ given the original sample. The percentile bootstrap confidence interval is \[ CI^{*}\left(v\right)\coloneqq\left[q_{\nicefrac{\alpha}{2}}^{*}(v),\,q_{1-\nicefrac{\alpha}{2}}^{*}(v)\right]=\left[\widehat{f}_{GPV}\left(v\right)+\frac{s_{\nicefrac{\alpha}{2}}^{*}(v)}{\sqrt{Lh^{3}}},\,\widehat{f}_{GPV}\left(v\right)+\frac{s_{\nicefrac{1-\alpha}{2}}^{*}(v)}{\sqrt{Lh^{3}}}\right], \] where $s_{\tau}^{*}(v)$ is the $\tau-$th quantile of the conditional distribution of \[ S^{*}\left(v\right)\coloneqq\left(Lh^{3}\right)^{\nicefrac{1}{2}}\left(\widehat{f}_{GPV}^{*}\left(v\right)-\widehat{f}_{GPV}\left(v\right)\right) \] given the original sample. The conditional distributions of the bootstrap statistics $\widehat{f}_{GPV}^{*}\left(v\right)$ and $S^{*}\left(v\right)$ given the original sample can be easily approximated by Monte Carlo methods. We show below that the bootstrap estimator of the finite-sample distribution of $S\left(v\right)$ is consistent. Let $\mathrm{P}^{*}\left[\cdot\right]$ denote the conditional probability given the original sample of bids.

thm[Bootstrap Consistency] Suppose Assumptions (ref) - (ref) hold. Then for any interior point $v\in\left(\underline{v},\overline{v}\right)$, \[ \underset{z\in\mathbb{R}}{\mathrm{sup}}\left|\mathrm{P}^{*}\left[S^{*}\left(v\right)\leq z\right]-\mathrm{P}\left[S\left(v\right)\leq z\right]\right|\rightarrow_{p}0,\,\textrm{as \ensuremath{L\uparrow\infty}}. \]
remboldTheorems (ref), (ref) and P\'{o}lya's theorem yield \[ \underset{z\in\mathbb{R}}{\mathrm{sup}}\left|\mathrm{P}^{*}\left[S^{*}\left(v\right)\leq z\right]-\mathrm{P}\left[\mathrm{N}\left(0,\mathrm{V}_{GPV}\left(v\right)\right)\leq z\right]\right|\rightarrow_{p}0,\,\textrm{as \ensuremath{L\uparrow\infty}}, \] for each $v\in\left(\underline{v},\overline{v}\right)$. The above result and standard arguments (see, e.g., VanDerVaart_Asymptotic_Statistics_Book) yield the asymptotic validity (consistency) of the percentile bootstrap confidence interval $CI^{*}\left(v\right)$, i.e., $\mathrm{P}\left[f\left(v\right)\in CI^{*}\left(v\right)\right]\rightarrow1-\alpha$ as $\textrm{\ensuremath{L\uparrow\infty}}.$
remboldOne can also studentize $S\left(v\right)$ using our estimator of the asymptotic variance: \begin{equation} Z\left(v\right)\coloneqq\frac{\widehat{f}_{GPV}\left(v\right)-f\left(v\right)}{\left(Lh^{3}\right)^{-\nicefrac{1}{2}}\widehat{\mathrm{V}}_{GPV}\left(v\right)^{\nicefrac{1}{2}}}. \end{equation} Since $\widehat{\mathrm{V}}_{GPV}\left(v\right)$ is consistent for the asymptotic variance, $Z\left(v\right)$ is asymptotically distributed as a standard normal random variable. Let $\widehat{\mathrm{V}}_{GPV}^{*}\left(v\right)$ be the bootstrap analogue of $\widehat{\mathrm{V}}_{GPV}\left(v\right)$. The bootstrap analogue of $Z\left(v\right)$ is $Z^{**}\left(v\right)\coloneqq\nicefrac{S^{*}\left(v\right)}{\sqrt{\widehat{\mathrm{V}}_{GPV}^{*}\left(v\right)}}$. Let $z_{\tau}^{**}(v)$ be the $\tau-$th quantile of the conditional distribution of $Z^{**}(v)$ given the original sample. The “bootstrap-$t$” (or studentized bootstrap) confidence intervals can be obtained by replacing the critical value $z_{1-\nicefrac{\alpha}{2}}$ in $CI^{\dagger}\left(v\right)$ with their bootstrap counterpart $z_{\tau}^{**}(v)$. The asymptotic validity of this alternative bootstrap confidence interval easily follows as a corollary to Theorem (ref).

Uniform Confidence Bands

Consider the stochastic process

equation[equation omitted — 391 chars of source]

Note that $\mathrm{E}\left[\mathit{\Gamma}\left(v\right)\right]=0$ and $\mathrm{E}\left[\mathit{\Gamma}\left(v\right)^{2}\right]=1$ for all $v\in I$.

The following theorem shows that (a version of) the centered Gaussian process with index set $I$ and covariance function $\mathrm{E}\left[\mathit{\Gamma}\left(v\right)\mathit{\Gamma}\left(v'\right)\right]$, for $\left(v,v'\right)\in I^{2}$, is a tight random element in $\ell^{\infty}\left(I\right)$. This Gaussian process, denoted by $\left\{ \mathit{\Gamma}_{G}\left(v\right):v\in I\right\} $, is the intermediate Gaussian process. The tightness of $\varGamma_{G}$ as a random element in $\ell^{\infty}\left(I\right)$ can be established using standard results (see, e.g., chernozhukov2014gaussian). The following theorem also shows that one can approximate the distribution of the sup-norm $\left\Vert Z\right\Vert _{I}=\underset{v\in I}{\mathrm{sup}}\left|Z\left(v\right)\right|$ with that of $\mathit{\Gamma}_{G}$. The result follows from uniform approximations of ((ref)) and $\widehat{\mathrm{V}}_{GPV}\left(v\right)$ and uses the coupling theorem for suprema of empirical processes of chernozhukov2014gaussian and the Gaussian anti-concentration inequality of chernozhukov2014anti.

thmSuppose Assumptions (ref) - (ref) hold. Then there exists a tight Gaussian random element $\varGamma_{G}$ in $\ell^{\infty}\left(I\right)$ that has mean zero and the same covariance structure as that of $\mathit{\Gamma}$. Moreover, \[ \underset{z\in\mathbb{R}}{\mathrm{sup}}\left|\mathrm{P}\left[\left\Vert Z\right\Vert _{I}\leq z\right]-\mathrm{P}\left[\left\Vert \mathit{\Gamma}_{G}\right\Vert _{I}\leq z\right]\right|\rightarrow0,\,\textrm{as \ensuremath{L\uparrow\infty}}. \]

The next result shows that the distribution of the sup-norm of the (empirical) bootstrap process

equation[equation omitted — 250 chars of source]

can be similarly approximated by that of $\mathit{\Gamma}_{G}$.

thmSuppose Assumptions (ref) - (ref) hold. Then, \[ \underset{z\in\mathbb{R}}{\mathrm{sup}}\left|\mathrm{P}^{*}\left[\left\Vert Z^{*}\right\Vert _{I}\leq z\right]-\mathrm{P}\left[\left\Vert \mathit{\Gamma}_{G}\right\Vert _{I}\leq z\right]\right|\rightarrow_{p}0,\,\textrm{as \ensuremath{L\uparrow\infty}}. \]

Since the distributions of the suprema of (the absolute values of) $\left\{ Z\left(v\right):v\in I\right\} $ and that of $\left\{ Z^{*}\left(v\right):v\in I\right\} $ are both well approximated by that of $\left\{ \varGamma_{G}\left(v\right):v\in I\right\} $, one can use the bootstrap critical values based on $\left\Vert Z^{*}\right\Vert _{I}$ for construction of uniform confidence bands. Let

equation[equation omitted — 196 chars of source]

be the $\left(1-\alpha\right)$-quantile of the conditional distribution of $\left\Vert Z^{*}\right\Vert _{I}$ given the original sample. The uniform confidence band is given by \[ CB^{*}\left(v\right)\coloneqq\left[\widehat{f}_{GPV}\left(v\right)-\zeta_{L,\alpha}^{*}\sqrt{\frac{\widehat{\mathrm{V}}_{GPV}\left(v\right)}{Lh^{3}}},\,\widehat{f}_{GPV}\left(v\right)+\zeta_{L,\alpha}^{*}\sqrt{\frac{\widehat{\mathrm{V}}_{GPV}\left(v\right)}{Lh^{3}}}\right],\,\textrm{for \ensuremath{v\in I}}. \] The following corollary establishes its asymptotic validity and provides an estimate of the order of the bootstrap critical value $\zeta_{L,\alpha}^{*}$.\footnote{It also implies that the (supremum) width of the band $CB^{*}$ is of order $O_{p}\left(\mathrm{log}\left(h^{-1}\right)^{\nicefrac{1}{2}}\left(Lh^{3}\right)^{-\nicefrac{1}{2}}\right)$.}

cor[Validity of Bootstrap Confidence Band] Suppose Assumptions (ref) - (ref) hold. Then, $\mathrm{P}\left[f\left(v\right)\in CB^{*}\left(v\right),\,\textrm{for all \ensuremath{v\in I}}\right]\rightarrow1-\alpha$ as $L\uparrow\infty$. Moreover, $\zeta_{L,\alpha}^{*}=O_{p}\left(\mathrm{log}\left(h^{-1}\right)^{\nicefrac{1}{2}}\right)$.
rembold[Limiting Distribution of Uniform Error]Since the seminal work of bickel1973some, it has been found that in many cases a suitable normalization of the supremum of a studentized absolute difference between a nonparametric curve and its kernel-based estimator converges in distribution to standard Gumbel distribution. Asymptotically valid uniform confidence bands can be based on the Gumbel approximation. We do not pursue such an approach in this paper, since it is known that the accuracy of such approximation is poor. See gine2015mathematical and chernozhukov2014anti for discussion. On the other hand, for the auction model one can show that (a suitable normalization of) $\left\Vert Z\right\Vert _{I}$ converges in distribution to standard Gumbel distribution when the true bidding strategy is linear. Derivation of the limiting distribution for the general case is interesting but beyond the scope of this paper. See the Supplement for more discussion.
remboldTaking the IGA approach, we establish consistency of bootstrap uniform confidence bands by showing that the distributions of both $\left\Vert Z\right\Vert _{I}$ and its empirical bootstrap counterpart can be approximated by the distribution of $\left\Vert \mathit{\Gamma}_{G}\right\Vert _{I}$. The proof hinges on using the coupling theorems of chernozhukov2014gaussian,chernozhukov2016empirical and the Gaussian anti-concentration inequality of chernozhukov2014anti. In the literature, the multiplier bootstrap is used when the IGA approach is taken to construct confidence bands for nonparametric curves. See, e.g., chernozhukov2014anti and Kato_Sasaki_2,Kato_Sasaki_1. Here, we chose the empirical bootstrap since it is practically convenient given the two-step nature of the GPV estimator.

Auction-Specific Heterogeneity

Asymptotic Normality and Estimation of the Asymptotic Variance

We now turn to the general model with auction-specific heterogeneity and a random number of bidders. Firstly, we establish the asymptotic normality of the GPV estimator by following the same approach and steps as in the case of the simplified model in Section (ref). While handling the general case is complicated by much heavier notations, all the results provided in this section can be viewed as straightforward generalizations of the results in Sections (ref)-(ref). The proofs of the results for the general case can be found in the Supplement.

In comparison with the simplified model, one of the main differences is in the form of the asymptotic variance. Recall that in the simplified case, the asymptotic variance of the GPV estimator depends on the derivative of the bidding strategy. As we show below, when there is auction-specific heterogeneity, the asymptotic variance also involves the partial derivatives of the bidding strategy with respect to the auction-specific characteristics. This is in addition to the partial derivative with respect to the valuation.

For some fixed $\boldsymbol{x}$ which is an interior point of $\mathcal{X}$, let $I\left(\boldsymbol{x}\right)\coloneqq\left[v_{l}\left(\boldsymbol{x}\right),v_{u}\left(\boldsymbol{x}\right)\right]$ be an inner closed sub-interval of $\left[\underline{v}\left(\boldsymbol{x}\right),\overline{v}\left(\boldsymbol{x}\right)\right]$. The fact that the conditional density of the valuations given $\boldsymbol{X}=\boldsymbol{x}$ and $N=n$ is $f\left(\cdot|\boldsymbol{x}\right)$ under Assumption (ref) motivates the following two-step estimator of $f\left(v|\boldsymbol{x}\right)$: \[ \widehat{f}_{GPV}\left(v|\boldsymbol{x},n\right)\coloneqq\frac{1}{\widehat{\pi}\left(n|\boldsymbol{x}\right)\widehat{\varphi}\left(\boldsymbol{x}\right)L}\sum_{l=1}^{L}\mathbbm{1}\left(N_{l}=n\right)\frac{1}{N_{l}}\sum_{i=1}^{N_{l}}\mathbb{T}_{il}\frac{1}{h^{1+d}}K_{f}\left(\frac{\widehat{V}_{il}-v}{h},\frac{\boldsymbol{X}_{l}-\boldsymbol{x}}{h}\right). \] Note that the above estimator only uses data from auctions with $N_{l}=n$. Since the PDF of valuations does not depend on the number of bidders, an estimator for $f(v|\boldsymbol{x})$ can be constructed as a weighted average of $\left\{ \widehat{f}_{GPV}\left(v|\boldsymbol{x},n\right):n\in\mathcal{N}\right\} $. E.g., GPV suggested using estimates of the conditional probabilities of drawing $N_{l}=n$ as the weights:

equation[equation omitted — 210 chars of source]

Note that this gives an expression that is the same as the right hand side of ((ref)).

By repeating the steps from Section (ref), one can show that the following analogue of the linearization results in ((ref)) holds for the general model:

multline[multline omitted — 501 chars of source]

where $\boldsymbol{B}_{\cdot l}\coloneqq\left(B_{1l},...,B_{N_{l}l}\right)$, and the remainder term is uniform in $v\in I\left(\boldsymbol{x}\right)$. For $\boldsymbol{b}_{\cdot}\coloneqq(b_{1},\ldots,b_{m})$ , the kernel function $\mathcal{M}^{n}$ is given by:

eqnarray[eqnarray omitted — 778 chars of source]

where $K_{f}'\left(\cdot,\cdot\right)$ denotes the partial derivative function of $K_{f}$ with respect to its first argument, \[ G\left(b,\boldsymbol{z},m\right)\coloneqq G\left(b|\boldsymbol{z},m\right)\pi\left(m|\boldsymbol{z}\right)\varphi\left(\boldsymbol{z}\right)\text{ and }g\left(b,\boldsymbol{z},m\right)\coloneqq g\left(b|\boldsymbol{z},m\right)\pi\left(m|\boldsymbol{z}\right)\varphi\left(\boldsymbol{z}\right). \]

Note that the leading term on the right-hand side of ((ref)) involves a V-statistic (with a kernel that depends on the bandwidth) and, therefore, can be analyzed using the Hoeffding decomposition. Thus, ((ref)) can be generalized as

align*[align* omitted — 673 chars of source]

where the remainder term is uniform in $v\in I\left(\boldsymbol{x}\right)$,

align*[align* omitted — 495 chars of source]

and $\mu_{\mathcal{M}^{n}}\left(v\right)\coloneqq\mathrm{E}\left[\mathcal{M}^{n}\left(\left(\boldsymbol{B}_{\cdot1},\boldsymbol{X}_{1},N_{1}\right),\left(\boldsymbol{B}_{\cdot2},\boldsymbol{X}_{2},N_{2}\right);v\right)\right]$.

The projection term $\mathcal{M}_{1}^{n}$ is the expectation of the kernel $\mathcal{M}^{n}\left(\left(\boldsymbol{b}.,\boldsymbol{z},m\right),\left(\boldsymbol{B}_{\cdot1},\boldsymbol{X}_{1},N_{1}\right);v\right)$ with the first argument fixed at $\left(\boldsymbol{b}.,\boldsymbol{z},m\right)$. The expression for $\mathcal{M}_{1}^{n}$ is:

eqnarray*[eqnarray* omitted — 869 chars of source]

As in the case of the simplified model, the contribution of $\mathcal{M}_{1}^{n}$ is asymptotically negligible. This happens for the same reason as in Section (ref): $\mathcal{M}_{1}^{n}$ depends on the difference between the expectation of the kernel function and the true density. Hence, the asymptotic distribution of the GPV estimator is driven solely by the $\mathcal{M}_{2}^{n}$ term, which is the expectation of the kernel $\mathcal{M}^{n}\left(\left(\boldsymbol{B}_{\cdot1},\boldsymbol{X}_{1},N_{1}\right),\left(\boldsymbol{b}.,\boldsymbol{z},m\right);v\right)$ with the second argument fixed at $\left(\boldsymbol{b}.,\boldsymbol{z},m\right)$.

We show in the supplement that a generalized version of ((ref)) holds:

align[align omitted — 1,019 chars of source]

where the remainder term is uniform in $v\in I\left(\boldsymbol{x}\right)$. Moreover, similarly to ((ref)),

align[align omitted — 1,120 chars of source]

where $s_{v}$ and $s_{\boldsymbol{x}}$ denote the partial derivatives of the bidding function:

equation[equation omitted — 376 chars of source]

The asymptotic variance of the GPV estimator is the limit of $\left(\pi\left(n|\boldsymbol{x}\right)\varphi\left(\boldsymbol{x}\right)\right)^{-2}\mathrm{V}_{\mathcal{M}}\left(v|\boldsymbol{x},n\right)$. After using a change of variable argument and ((ref)), the variance is shown to be

multline[multline omitted — 639 chars of source]

The following theorem is a generalization of Theorem (ref). The proof of the theorem as well as the proofs of all other results provided in Section (ref) are in the Supplement.

thmSuppose Assumptions (ref) - (ref) hold. Then, for any interior point $\left(v,\boldsymbol{x}\right)\in\mathcal{S}_{V,\boldsymbol{X}},$ \[ \left(Lh^{3+d}\right)^{\nicefrac{1}{2}}\left(\widehat{f}_{GPV}\left(v|\boldsymbol{x},n\right)-f\left(v|\boldsymbol{x}\right)\right)\rightarrow_{d}\mathrm{N}\left(0,\mathrm{V}_{GPV}\left(v|\boldsymbol{x},n\right)\right). \] Moreover, $\left\{ \widehat{f}_{GPV}\left(v|\boldsymbol{x},n\right):n\in\mathcal{N}\right\} $ are asymptotically independent.
remboldAs in the simplified model, the asymptotic variance of the GPV estimator depends on a convoluted integral transformation involving the kernel function, its derivative and the derivatives of the bidding strategy. Auction-specific heterogeneity complicates the expression in two ways. Firstly, the dimension of the integral depends on the number of auction characteristics (covariates), and analytical calculation of the integral becomes cumbersome when there are many covariates. Secondly, the expression now contains the derivatives of the bidding strategy with respect to the covariates. The reason for that is apparent from the expression for the second moment of $\mathcal{M}_{2}^{n}\left(\boldsymbol{B}_{\cdot1},\boldsymbol{X}_{1},N_{1};v\right)$ in ((ref)) as it involves averaging of the inverse bidding function $\xi(u,\boldsymbol{z}',n)$ over the auction characteristics $\boldsymbol{z}'$ in a shrinking neighborhood of $\boldsymbol{x}$. While the GPV estimator and the quantile-based estimator have the same rate of convergence (see Marmer_Shneyerov_Quantile_Auctions for the rate of convergence and the expression of the asymptotic variance $\mathrm{V}_{QB}\left(v|\boldsymbol{x},n\right)$ of the quantile-based estimator in the general model), one can show that the GPV estimator has a smaller asymptotic variance than that of the quantile-based estimator, as long as the two estimators use the same second-order kernel function. It can be shown that the ratio $\nicefrac{\mathrm{V}_{GPV}\left(v|\boldsymbol{x},n\right)}{\mathrm{V}_{QB}\left(v|\boldsymbol{x},n\right)}\leq1$ by applying Jensen's inequality and standard techniques for multi-dimensional integration. A detailed proof can be found in the Supplement.
rembold[Optimal Weights]Since the distribution of valuations does not depend on the number of bidders, i.e., $f\left(v|\boldsymbol{x},n\right)=f\left(v|\boldsymbol{x}\right)$ for all $\left(v,\boldsymbol{x},n\right)\in\mathcal{S}_{V,\boldsymbol{X}}\times\mathcal{N}$, one can average the estimators $\left\{ \widehat{f}_{GPV}\left(v|\boldsymbol{x},n\right):n\in\mathcal{N}\right\} $ to obtain a weighted estimator of the density $f\left(v|\boldsymbol{x}\right)$: \[ \widehat{f}_{GPV}^{w}\left(v|\boldsymbol{x}\right)\coloneqq\sum_{n\in\mathcal{N}}\widehat{w}\left(n,v,\boldsymbol{x}\right)\widehat{f}_{GPV}\left(v|\boldsymbol{x},n\right), \] where the weights $\widehat{w}\left(n,v,\boldsymbol{x}\right)$, $n\in\mathcal{N}$ should satisfy $\widehat{w}\left(n,v,\boldsymbol{x}\right)\rightarrow_{p}w\left(n,v,\boldsymbol{x}\right)$, $n\in\mathcal{N}$ and $\sum_{n\in\mathcal{N}}w\left(n,v,\boldsymbol{x}\right)=1$. As in Marmer_Shneyerov_Quantile_Auctions, the optimal weights that minimize the asymptotic variance of of the resulting weighted estimator are inversely related to $\mathrm{V}_{GPV}\left(v|\boldsymbol{x},n\right)$ and given by \[ w^{\mathrm{opt}}\left(n,v,\boldsymbol{x}\right)\coloneqq\frac{\frac{1}{\mathrm{V}_{GPV}\left(v|\boldsymbol{x},n\right)}}{\sum_{n\in\mathcal{N}}\frac{1}{\mathrm{V}_{GPV}\left(v|\boldsymbol{x},n\right)}}. \] These weights can be consistently estimated by the plug-in principle using an estimator of $\mathrm{V}_{GPV}\left(v|\boldsymbol{x},n\right)$. The original GPV estimator uses the weights $\widehat{w}\left(v,\boldsymbol{x},n\right)=\widehat{\pi}(n|\boldsymbol{x})$, $n\in\mathcal{N}$. See ((ref)) and ((ref)). Note, however, that such weights would be sub-optimal from the point of view of minimizing the asymptotic variance. We provide the asymptotic normality of $\widehat{f}_{GPV}\left(v|\boldsymbol{x}\right)$ in the corollary below.
corSuppose Assumptions (ref) - (ref) hold. Then, for any interior point $\left(v,\boldsymbol{x}\right)\in\mathcal{S}_{V,\boldsymbol{X}}$, \[ \left(Lh^{3+d}\right)^{\nicefrac{1}{2}}\left(\widehat{f}_{GPV}\left(v|\boldsymbol{x}\right)-f\left(v|\boldsymbol{x}\right)\right)\rightarrow_{d}\mathrm{N}\left(0,\mathrm{V}_{GPV}\left(v|\boldsymbol{x}\right)\right), \] where $\mathrm{V}_{GPV}\left(v|\boldsymbol{x}\right)\coloneqq\sum_{n\in\mathcal{N}}\pi\left(n|\boldsymbol{x}\right)^{2}\mathrm{V}_{GPV}\left(v|\boldsymbol{x},n\right)$.

For practical purposes, it is important to have a consistent estimator of the asymptotic variance $\mathrm{V}_{GPV}\left(v|\boldsymbol{x}\right)$ that avoids estimation of the bidding strategy $s\left(\cdot,\cdot,n\right)$ and its derivatives. It is also highly desirable to avoid analytical or numerical evaluation of a multidimensional integral in the definition of the asymptotic variance. Following the same approach we used in the case of the simplified model (see ((ref))), we rely on the sample analogue of ((ref)):

multline[multline omitted — 512 chars of source]

where

multline*[multline* omitted — 446 chars of source]

The following result is a generalization of Theorem (ref).

thmSuppose Assumptions (ref) - (ref) hold. Then, for any interior point $\boldsymbol{x}$, \[ \underset{v\in I\left(\boldsymbol{x}\right)}{\mathrm{sup}}\left|\widehat{\mathrm{V}}_{GPV}\left(v|\boldsymbol{x},n\right)-\left(\pi\left(n|\boldsymbol{x}\right)\varphi\left(\boldsymbol{x}\right)\right)^{-2}\mathrm{V}_{\mathcal{M}}\left(v|\boldsymbol{x},n\right)\right|=O_{p}\left(\left(\frac{\mathrm{log}\left(L\right)}{Lh^{3+d}}\right)^{\nicefrac{1}{2}}+h^{R}\right). \]
remboldAn estimator for the asymptotic variance $\mathrm{V}_{GPV}(v|\boldsymbol{x})$ of the estimator $\widehat{f}_{GPV}\left(v|\boldsymbol{x}\right)$ in Corollary (ref) can be constructed using the plug-in approach: \[ \mathrm{\widehat{V}}_{GPV}\left(v|\boldsymbol{x}\right)\coloneqq\sum_{n\in\mathcal{N}}\widehat{\pi}\left(n|\boldsymbol{x}\right)^{2}\widehat{\mathrm{V}}_{GPV}\left(v|\boldsymbol{x},n\right). \] Its rate of convergence is the same as in Theorem (ref).

Bootstrap-based Inference and Uniform Confidence Bands

To generate bootstrap samples, we apply the same resampling procedure as that proposed in Marmer_Shneyerov_Quantile_Auctions. First, we randomly draw $L$ observations from $\left\{ \left(\boldsymbol{X}_{l},N_{l}\right):l=1,...,L\right\} $ (i.e., the auction-specific characteristics) with replacement. Next, we randomly draw bids with replacement from the bids corresponding to each selected auction. Given $\left(\boldsymbol{X}_{l}^{*},N_{l}^{*}\right)=\left(\boldsymbol{X}_{l'},N_{l'}\right)$ in the first step, in the second step $\{B_{il}^{*}:i=1,...,N_{l}^{*}\}$ is generated as an empirical bootstrap sample drawn from $\{B_{il'}:i=1,...,N_{l'}\}$. Let $\widehat{\xi}^{*}\left(\cdot,\cdot,\cdot\right)$ and $\widehat{\varphi}^{*}\left(\cdot\right)$ be the bootstrap analogues of $\widehat{\xi}\left(\cdot,\cdot,\cdot\right)$ and $\widehat{\varphi}\left(\cdot\right)$ respectively. Let $\widehat{f}_{GPV}^{*}\left(v|\boldsymbol{x}\right)$ denote the bootstrap version of the GPV estimator: \[ \widehat{f}_{GPV}^{*}\left(v|\boldsymbol{x}\right)\coloneqq\frac{1}{\widehat{\varphi}^{*}\left(\boldsymbol{x}\right)L}\sum_{l=1}^{L}\frac{1}{N_{l}^{*}}\sum_{i=1}^{N_{l}^{*}}\mathbb{T}_{il}^{*}\frac{1}{h^{1+d}}K_{f}\left(\frac{\widehat{V}_{il}^{*}-v}{h},\frac{\boldsymbol{X}_{l}^{*}-\boldsymbol{x}}{h}\right), \] where $\widehat{V}_{il}^{*}\coloneqq\widehat{\xi}^{*}\left(B_{il}^{*},\boldsymbol{X}_{l}^{*},N_{l}^{*}\right)$ and the bootstrap version of the trimming factor is given by \[ \mathbb{T}_{il}^{*}\coloneqq\mathbbm{1}\left(\mathbb{H}\left(\left(B_{il}^{*},\boldsymbol{X}_{l}^{*}\right),2h\right)\subseteq\mathcal{\widehat{S}}_{B,\boldsymbol{X}}^{N_{l}}\right). \]

Consider the scaled deviation of the GPV estimator from the true PDF, and its bootstrap analogue:

gather*[gather* omitted — 392 chars of source]

The following result, which is a generalization of Theorem (ref), establishes the validity of the percentile bootstrap for $f\left(v|\boldsymbol{x}\right)$.

thmSuppose Assumptions (ref) - (ref) hold. Then, for any interior point $\left(v,\boldsymbol{x}\right)\in\mathcal{S}_{V,\boldsymbol{X}}$, \[ \underset{z\in\mathbb{R}}{\mathrm{sup}}\left|\mathrm{P}^{*}\left[S^{*}\left(v|\boldsymbol{x}\right)\leq z\right]-\mathrm{P}\left[S\left(v|\boldsymbol{x}\right)\leq z\right]\right|\rightarrow_{p}0,\,\textrm{as \ensuremath{L\uparrow\infty}}. \]

We now turn to construction of uniform confidence bands for $\left\{ f\left(v|\boldsymbol{x}\right):v\in I\left(\boldsymbol{x}\right)\right\} $ given a fixed interior point $\boldsymbol{x}$. Consider the following processes: \[ Z\left(v|\boldsymbol{x}\right)\coloneqq\frac{\widehat{f}_{GPV}\left(v|\boldsymbol{x}\right)-f\left(v|\boldsymbol{x}\right)}{\left(Lh^{3+d}\right)^{-\nicefrac{1}{2}}\widehat{\mathrm{V}}_{GPV}\left(v|\boldsymbol{x}\right)^{\nicefrac{1}{2}}}\textrm{ and }Z^{*}\left(v|\boldsymbol{x}\right)\coloneqq\frac{\widehat{f}_{GPV}^{*}\left(v|\boldsymbol{x}\right)-\widehat{f}_{GPV}\left(v|\boldsymbol{x}\right)}{\left(Lh^{3+d}\right)^{-\nicefrac{1}{2}}\widehat{\mathrm{V}}_{GPV}\left(v|\boldsymbol{x}\right)^{\nicefrac{1}{2}}},\textrm{ \ensuremath{v\in I\left(\boldsymbol{x}\right)}}. \] Similarly to the simplified model, the distribution of $\left\Vert Z\left(\cdot|\boldsymbol{x}\right)\right\Vert _{I\left(\boldsymbol{x}\right)}$ can be approximated by the conditional distribution of $\left\Vert Z^{*}\left(\cdot|\boldsymbol{x}\right)\right\Vert _{I\left(\boldsymbol{x}\right)}$. Let $\zeta_{L,\alpha}^{*}$ be the $\left(1-\alpha\right)$-quantile of the conditional distribution of $\left\Vert Z^{*}\left(\cdot|\boldsymbol{x}\right)\right\Vert _{I\left(\boldsymbol{x}\right)}$ given the original sample. Consider the following confidence band: for $\ensuremath{v\in I\left(\boldsymbol{x}\right)}$, \[ CB^{*}\left(v|\boldsymbol{x}\right)\coloneqq\left[\widehat{f}_{GPV}\left(v|\boldsymbol{x}\right)-\zeta_{L,\alpha}^{*}\sqrt{\frac{\widehat{\mathrm{V}}_{GPV}\left(v|\boldsymbol{x}\right)}{Lh^{3+d}}},\,\widehat{f}_{GPV}\left(v|\boldsymbol{x}\right)+\zeta_{L,\alpha}^{*}\sqrt{\frac{\widehat{\mathrm{V}}_{GPV}\left(v|\boldsymbol{x}\right)}{Lh^{3+d}}}\right]. \] The following result, which is a generalization of Corollary (ref), establishes the validity of $CB^{*}\left(\cdot|\boldsymbol{x}\right)$.

thmSuppose Assumptions (ref) - (ref) hold. Then, \[ \mathrm{P}\left[f\left(v|\boldsymbol{x}\right)\in CB^{*}\left(v|\boldsymbol{x}\right),\,\textrm{for all \ensuremath{v\in I\left(\boldsymbol{x}\right)}}\right]\rightarrow1-\alpha,\,\textrm{as \ensuremath{L\uparrow\infty}}. \]

Binding Reserve Price

Section 4 of GPV shows how to modify their identification and estimation strategy when there is a binding reserve price. Here, we discuss how our approach can be applied in that case.

When there is a binding reserve price, it is assumed that only bidders with valuations exceeding the reserve price submit bids. Thus, one has to distinguish between the numbers of potential and actual (active) bidders. Let $N$ denote the number of potential bidders, which is assumed to be known to players. The bidding strategy depends on the number of potential bidders instead of the number of active bidders. As discussed in GPV, $N$ can be estimated by taking the maximum of the observed numbers of actual bidders across the auctions: $\widehat{N}\coloneqq\mathrm{max}\left\{ N_{l}:l=1,\ldots,L\right\} ,$ where $N_{l}$ is the number of actual bidders in auction $l$.

GPV assume that the reserve price in auction $l$, denoted $P_{0l}$, is some unknown deterministic function of the auction characteristics $\boldsymbol{X}_{l}$: $P_{0l}=p_{0}\left(\boldsymbol{X}_{l}\right)$. The probability of drawing a valuation below the reserve price is given by $\Phi\left(\boldsymbol{X}_{l}\right)\coloneqq F\left(P_{0l}|\boldsymbol{X}_{l}\right)$. The conditional CDF and PDF of the distribution of valuations given participation (submitting a bid) are

equation[equation omitted — 313 chars of source]

respectively. The third displayed equation on page 550 of GPV shows that $\Phi(\cdot)$ can be estimated using a nonparametric regression of the number of actual bidders: \[ \widehat{\Phi}\left(\boldsymbol{x}\right)\coloneqq1-\frac{1}{\widehat{N}\widehat{\varphi}\left(\boldsymbol{x}\right)L}\sum_{l=1}^{L}\frac{1}{h^{d}}N_{l}K_{\boldsymbol{X}}\left(\frac{\boldsymbol{X}_{l}-\boldsymbol{x}}{h}\right). \]

GPV point out that the density of bids is unbounded at the reserve price $p_{0}\left(\boldsymbol{x}\right)$ and, in its neighborhood, behaves as $\nicefrac{1}{\sqrt{b-p_{0}(\boldsymbol{x})}}$ . To avoid technical problems due to the unbounded density, they propose to transform the bids as \[ B_{\dagger il}=\left(B_{il}-P_{0l}\right)^{\nicefrac{1}{2}}. \] The support of $\left(B_{\dagger11},\boldsymbol{X}_{1}\right)$ is given by $\mathcal{S}_{B_{\dagger},\boldsymbol{X}}\coloneqq\left\{ \left(b,\boldsymbol{x}\right):\boldsymbol{x}\in\mathcal{X},\,b\in\left[0,\overline{b}_{\dagger}\left(\boldsymbol{x}\right)\right]\right\} $, where $\overline{b}_{\dagger}\left(\boldsymbol{z}\right)\coloneqq\left(\overline{b}_{\dagger}\left(\boldsymbol{z}\right)-p_{0}\left(\boldsymbol{z}\right)\right)^{\nicefrac{1}{2}}$. The support can be estimated by $\mathcal{\widehat{S}}_{B_{\dagger},\boldsymbol{X}}\coloneqq\left\{ \left(b,\boldsymbol{x}\right):\boldsymbol{z}\in\mathcal{X},\,b\in\left[0,\widehat{\overline{b}}_{\dagger}\left(\boldsymbol{x}\right)\right]\right\} $, where $\widehat{\overline{b}}_{\dagger}\left(\boldsymbol{x}\right)\coloneqq\mathrm{\mathrm{max}}\left\{ B_{\dagger pl}:p=1,...,N_{l},\,\boldsymbol{X}_{l}\in\Pi_{h_{\partial}}\left(\boldsymbol{x}\right),\,l=1,...,L\right\} $, see page 550 in GPV.

Let $G_{\dagger}\left(\cdot|\cdot\right)$ and $g_{\dagger}\left(\cdot|\cdot\right)$ denote respectively the conditional CDF and PDF of the transformed bids $B_{\dagger11}$ given $\boldsymbol{X}_{1}$. Let $G_{\dagger}\left(b_{\dagger},\boldsymbol{z}\right)\coloneqq G_{\dagger}\left(b_{\dagger}|\boldsymbol{z}\right)\varphi\left(\boldsymbol{z}\right)$ and $g_{\dagger}\left(b_{\dagger},\boldsymbol{z}\right)\coloneqq g_{\dagger}\left(b_{\dagger}|\boldsymbol{z}\right)\varphi\left(\boldsymbol{z}\right)$. GPV show that $g_{\dagger}\left(\cdot,\cdot\right)$ is bounded on its support, and that latent valuations can be recovered using \[ V_{il}=\xi_{\dagger}\left(B_{\dagger il},\boldsymbol{X}_{l}\right)\coloneqq P_{0l}+B_{\dagger il}^{2}+\frac{2B_{\dagger il}}{N-1}\frac{\left(1-\Phi\left(\boldsymbol{X}_{l}\right)\right)G_{\dagger}\left(B_{\dagger il},\boldsymbol{X}_{l}\right)+\Phi\left(\boldsymbol{X}_{l}\right)\varphi\left(\boldsymbol{X}_{l}\right)}{\left(1-\Phi\left(\boldsymbol{X}_{l}\right)\right)g_{\dagger}\left(B_{\dagger il},\boldsymbol{X}_{l}\right)}. \]

In the modified GPV procedure, one first estimates $\xi_{\dagger}\left(\cdot,\cdot\right)$ by replacing $N$, $\Phi\left(\cdot\right)$, $G_{\dagger}\left(\cdot,\cdot\right)$, $g_{\dagger}\left(\cdot,\cdot\right)$ and $\varphi\left(\cdot\right)$ with their estimators. $G_{\dagger}\left(\cdot,\cdot\right)$ and $g_{\dagger}\left(\cdot,\cdot\right)$ can be estimated using the transformed bids $B_{\dagger il}$:

gather*[gather* omitted — 540 chars of source]

In the second step of the modified GPV procedure, one uses the pseudo valuations \[ \left\{ \widehat{V}_{il}\coloneqq\widehat{\xi}_{\dagger}\left(B_{\dagger il},\boldsymbol{X}_{l}\right):i=1,\dots,N_{l},l=1,\ldots,L\right\} , \] where $\widehat{\xi}_{\dagger}\left(\cdot,\cdot\right)$ is the estimated version of $\xi_{\dagger}(\cdot,\cdot)$, in place of latent valuations to construct a kernel density estimator of $f^{\star}\left(v|\boldsymbol{x}\right)$: \[ \widehat{f}_{GPV}^{\star}(v|\boldsymbol{x})\coloneqq\frac{1}{\widehat{\varphi}(\boldsymbol{x})L}\sum_{l=1}^{L}\frac{1}{N_{l}}\sum_{i=1}^{N_{l}}\mathbb{T}_{il}\frac{1}{h^{1+d}}K_{f}\left(\frac{\widehat{V}_{il}-v}{h},\frac{\boldsymbol{X}_{l}-\boldsymbol{x}}{h}\right), \] where the trimming factors $\mathbb{T}_{il}$ can be defined analogously to the case with no binding reserve price: $\mathbb{T}_{il}\coloneqq\mathbbm{1}\left(\mathbb{H}\left(\left(B_{\dagger il},\boldsymbol{X}_{l}\right),2h\right)\subseteq\mathcal{\widehat{S}}_{B_{\dagger},\boldsymbol{X}}\right).$

Our approach can be used to obtain the asymptotic distribution of the modified GPV estimator as follows. In view of the definitions of $\xi_{\dagger}\left(\cdot.\cdot\right)$ and its estimator, the $\widehat{V}_{il}-V_{il}$ term in the analogue of ((ref)) can be expanded as \[ \widehat{V}_{il}-V_{il}=\frac{2B_{\dagger il}}{N-1}\frac{\left(1-\Phi\left(\boldsymbol{X}_{l}\right)\right)G_{\dagger}\left(B_{\dagger il},\boldsymbol{X}_{l}\right)+\Phi\left(\boldsymbol{X}_{l}\right)\varphi\left(\boldsymbol{X}_{l}\right)}{\left(1-\Phi\left(\boldsymbol{X}_{l}\right)\right)g_{\dagger}\left(B_{\dagger il},\boldsymbol{X}_{l}\right)^{2}}\left(\widehat{g}_{\dagger}\left(B_{\dagger il},\boldsymbol{X}_{l}\right)-g_{\dagger}\left(B_{\dagger il},\boldsymbol{X}_{l}\right)\right)+s.o., \] where “$s.o.$” stands for smaller order terms. Hence, the GPV estimator of $f^{\star}\left(v|\boldsymbol{x}\right)$ still has a representation of the same form as in ((ref)):

multline*[multline* omitted — 436 chars of source]

where $\boldsymbol{B}_{\dagger\cdot l}\coloneqq\left(B_{\dagger1l},...,B_{\dagger N_{l}l}\right)$, and

eqnarray*[eqnarray* omitted — 868 chars of source]

Similarly to the case with no reserve price, one can apply the Hoeffding decomposition with only $\mathcal{M}_{\dagger2}\left(\boldsymbol{b}.,\boldsymbol{z},m;v\right)\coloneqq\mathrm{E}\left[\mathcal{M}_{\dagger}\left(\left(\boldsymbol{B}_{\dagger\cdot1},\boldsymbol{X}_{1},N_{1}\right),\left(\boldsymbol{b}.,\boldsymbol{z},m\right);v\right)\right]$ contributing to the asymptotic variance:

eqnarray*[eqnarray* omitted — 825 chars of source]

Similarly to ((ref)), the asymptotic variance of the GPV estimator is now given by the limit of

align[align omitted — 909 chars of source]

as $h\downarrow0$, where $\overline{\pi}\left(\boldsymbol{x}\right)\coloneqq\mathrm{E}\left[N_{1}^{-1}\mid\boldsymbol{X}_{1}=\boldsymbol{x}\right]$. Note that conditionally on $\boldsymbol{X}_{1}$, the number of active bidders $N_{1}$ has a binomial distribution with parameters $N$ and $1-\Phi\left(\boldsymbol{X}_{1}\right)$. Lastly, similarly to Theorem (ref), the expression in ((ref)), and for any interior point $v\in\left(p_{0}\left(\boldsymbol{x}\right),\bar{v}\left(\boldsymbol{x}\right)\right)$, the GPV estimator of $f^{\star}\left(v\mid\boldsymbol{x}\right)$ is asymptotically normal with the asymptotic variance given by

eqnarray*[eqnarray* omitted — 950 chars of source]

where $s_{\dagger}(\cdot,\boldsymbol{x})\coloneqq\xi_{\dagger}^{-1}(\cdot,\boldsymbol{x})$, and the partial derivatives $s_{\dagger v}$ and $s_{\dagger\boldsymbol{x}}$ are defined similarly to $s_{v}$ and $s_{\boldsymbol{x}}$ in ((ref)). Similarly to Corollary (ref), one can show: \[ \left(Lh^{3+d}\right)^{\nicefrac{1}{2}}\left(\widehat{f}^{\star}\left(v\mid\boldsymbol{x}\right)-f^{\star}\left(v\mid\boldsymbol{x}\right)\right)\rightarrow_{d}\mathrm{N}\left(0,\mathrm{V}_{\dagger GPV}\left(v,\boldsymbol{x}\right)\right). \]

To estimate the asymptotic variance $\mathrm{V}_{\dagger GPV}(v,\boldsymbol{x})$, one can use the sample analogue of ((ref)) in the same way as that used to construct the estimator of $\mathrm{V}_{GPV}(v,\boldsymbol{x})$ defined by ((ref)) from ((ref)). As before, the approach does not require estimation of the bidding strategy $s_{\dagger}(v,\boldsymbol{x})$ or its derivatives. The analogue estimator is given by

multline*[multline* omitted — 450 chars of source]

where $\widehat{\overline{\pi}}\left(\boldsymbol{x}\right)$ is the Nadaraya-Watson estimator of $\overline{\pi}\left(\boldsymbol{x}\right)$, and

multline*[multline* omitted — 723 chars of source]

Suppose $I\left(\boldsymbol{x}\right)$ is an inner closed sub-interval of $\left[p_{0}\left(\boldsymbol{x}\right),\bar{v}\left(\boldsymbol{x}\right)\right]$. The uniform convergence rate of $\mathrm{\widehat{V}}_{\dagger GPV}\left(v,\boldsymbol{x}\right)$ to ((ref)) can be shown to be the same as that in the statement of Theorem (ref). In view of the definitions in ((ref)), the nonparametric estimator for $f\left(v|\boldsymbol{x}\right)$ is $\left(1-\widehat{\Phi}\left(\boldsymbol{x}\right)\right)\widehat{f}^{\star}\left(v|\boldsymbol{x}\right)$. Since $\widehat{\Phi}\left(\boldsymbol{x}\right)$ converges at a faster rate than the PDF estimator $\widehat{f}^{\star}\left(v\mid\boldsymbol{x}\right)$, one can see that \[ \left(Lh^{3+d}\right)^{\nicefrac{1}{2}}\left(\left(1-\widehat{\Phi}\left(\boldsymbol{x}\right)\right)\widehat{f}^{\star}\left(v\mid\boldsymbol{x}\right)-f\left(v\mid\boldsymbol{x}\right)\right)\rightarrow_{d}\mathrm{N}\left(0,\left(1-\Phi\left(\boldsymbol{x}\right)\right)^{2}\mathrm{V}_{\dagger GPV}\left(v,\boldsymbol{x}\right)\right). \]

A valid uniform confidence band of $\left\{ f\left(v|\boldsymbol{x}\right):v\in I\left(\boldsymbol{x}\right)\right\} $ can be constructed by adapting the methods described in Section (ref).

Monte Carlo Simulations

In this section, we assess the finite-sample coverage accuracy of the uniform confidence bands. Our simulation design follows \citet*{Marmer_Shneyerov_Quantile_Auctions}, and the DGP is described in Remark (ref). We consider $\theta\in\left\{ 1,2\right\} $ and draw valuations from $f_{\theta}$. We choose the triweight kernel when implementing the two-step estimator. We used the second-order triweight kernel in the second step and used the fourth-order triweight kernel in the first step.

We need to choose the bandwidths in the first step when we construct the pseudo valuations and the second step when we implement kernel density estimation using the pseudo valuations. We follow GPV (see Section 2.4) and use $h_{g}=3.72\cdot\widehat{\sigma}_{b}\cdot\left(N\cdot L\right)^{-\nicefrac{1}{5}}$as the first-step bandwidth, where $\widehat{\sigma}_{b}$ is the estimated standard deviation of the observed bids. We use $h_{f}=3.15\cdot\widehat{\sigma}_{v}\cdot\left(\left(N\cdot L\right)_{\mathbb{T}}\right)^{-\nicefrac{1}{5}}$as the second-step bandwidth, where $\widehat{\sigma}_{v}$ is the estimated standard deviation of the trimmed pseudo valuations and $\left(N\cdot L\right)_{\mathbb{T}}$ is the number of bids remaining after the trimming. The constants $3.72$ and $3.15$ are Silverman's rule-of-thumb constants corresponding to fourth-order and second-order triweight kernels.\footnote{See li2007net for a description of the Silverman approach; see also li2003semiparametric.} We consider different numbers of bidders $N\in\left\{ 3,5,7\right\} $, and also the density function over different ranges: $v\in\left[0.2,0.8\right]$ and $v\in\left[0.3,0.7\right]$.\footnote{We use grid maximization, where the grid is chosen as $[v_{l}:0.001:v_{u}]$. We have also tried a finer grid $[v_{l}:0.0001:v_{u}],$ which produced similar results.} The number of auctions $L$ is chosen so the total number of observations is fixed as $N\cdot L=2100$.

table[table omitted — 1,483 chars of source]

In Table (ref), we report our simulation results for the bootstrap-based IGA uniform confidence band $CB^{*}$. We find that the IGA bootstrap approach provides accurate coverage probabilities. Additional simulation results are reported in the Supplement.

Conclusion

The GPV estimator has proven to be the essential input in virtually all nonparametric structural auction models. By proving the asymptotic normality and the first-order validity of the bootstrap uniform confidence bands, this paper completes the econometric theory of the GPV estimator and opens way to new applications.

Our pointwise asymptotic normality results can be used for inference on an important policy variable: the optimal reserve price. As discussed, e.g., in haile2003iim, the optimal reserve price $r(\boldsymbol{x})$ in auctions with $\boldsymbol{X}_{l}=\boldsymbol{x}$ satisfies the following equation: \[ r(\boldsymbol{x})-\frac{1-F(r(\boldsymbol{x})|\boldsymbol{x})}{f(r(\boldsymbol{x})|\boldsymbol{x})}=c(\boldsymbol{x}), \] where $c(\boldsymbol{x})$ is the seller's own valuation. Suppose that the estimator $\widehat{r}(\boldsymbol{x})$ is constructed by solving an estimated version of the above equation with $f(\cdot|\boldsymbol{x})$ replaced by its GPV estimator $\widehat{f}_{GPV}\left(\cdot|\boldsymbol{x}\right)$. In that case, our pointwise normality results imply the asymptotic normality of the estimated optimal reserve price:

multline*[multline* omitted — 420 chars of source]

Our results for the validity of the percentile bootstrap of the GPV estimator naturally carry over to the above estimator of the optimal reserve price. Thus in practice, one can use the percentile bootstrap for inference on the optimal reserve price.

Our uniform confidence bands can be used for specification of the density of valuations.

In future research, our results could be extended in several directions. Below, we briefly describe some of potentially interesting extensions. These extensions would address the limitations of the independent private values model that underlies the GPV estimator.

First, in GPV the bidders are treated symmetrically. Empirically this is not always the case. See, e.g., flambard2006asymmetry for an application to snow removal contracts. Second, we abstract from the empirically relevant issue of unobserved heterogeneity, as in krasnokutskaya2011identification, hu2013identification and roberts2013unobserved. Third, correlated values may also be important empirically. li2002structural and hubbard2012semiparametric extend the GPV estimator to the affiliated value environment. Fourth, guerre2009nonparametric and zincenko2018nonparametric provide extensions to an environment with risk-averse bidders. Fifth, the literature on endogenous entry in auctions has developed rapidly. See, e.g., li2009entry, krasnokutskaya2011bid, marmer2013model, roberts2013should and gentry2014identification.

In the above models, the basic GPV estimator is often adapted to suit the needs of a particular application. But most of these estimators share the underlying two-step structure of the GPV estimator. We conjecture that our main results will also prove useful for establishing the asymptotic normality and validity of certain uniform confidence bands for these GPV-like estimators.

A remaining unresolved important practical issue is bandwidth selection for the GPV estimator. It is possible that the recent advances in that area (e.g., calonico2014robust and armstrong2016simple) can be adapted to the framework of GPV.