Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
111,339 characters · 11 sections · 72 citation commands
Inference for First-Price Auctions with Guerre, Perrigne, and Vuong's Estimator 2019. This manuscript version is made available under the Creative Commons CC-BY-NC-ND 4.0 license http://creativecommons.org/licenses/by-nc-nd/4.0/. First version: March 3, 2016, This version: .
The structural estimation of auctions is an important and rapidly growing subfield at the junction of econometrics and industrial organization. Since the seminal work of \citet*[GPV hereafter]{Guerre_Perrigne_Vuong_Auction_Econometrica_2000}, much of theoretical and applied work has focused on nonparametric estimation of first-price, sealed-bid auctions.\footnote{\citet*{hendricks2007empirical} survey the empirical auction literature, while \citet*{athey2007nonparametric} survey the nonparametric identification approaches. hickman2012structural provide a recent review.} The object of interest is the probability density function (PDF) of latent valuations, which can then be used for a variety of policy counterfactuals such as the optimal reserve price (paarsch1997deriving and li2003semiparametric).\footnote{In auctions, the valuation of a bidder is simply his or her willingness to pay for the object.}
The focus on nonparametric estimation is due to several reasons. First, in empirical applications, large auction datasets are often available.\footnote{E.g., kawai2014detecting utilize a dataset of $40,000$ auctions in their study of collusion in Japan, while augenblick2015sunk employs a dataset $160,000$ penny auctions.} Second, nonparametric methods are flexible since no functional form assumptions are needed. Third, in auctions, nonparametric estimators are built directly from the identification arguments, and are often easy to implement.\footnote{See athey2002identification for a number of additional identification results.}
GPV proposed a two-step nonparametric estimator of the PDF of valuations in first-price auctions, and showed that it is uniformly consistent and attains the minimax optimal uniform convergence rate. However, it has been an open question whether this estimator also converges in a distributional sense, which would allow empirical researchers to perform inferences, thereby increasing the scope of applications.
Recently, Marmer_Shneyerov_Quantile_Auctions developed an alternative quantile-based estimator of the PDF of valuations and showed its asymptotic normality.\footnote{More recently, quantile methods have been used in the context of auctions in gimenes2016quantile, gimenes2017econometrics, liu2017nonparametric, and luo2017integrated.} However, the GPV estimator is well established in the literature and is used in all the empirical applications we are aware of. Moreover, our results imply that the GPV estimator has a smaller asymptotic variance than that of the quantile-based estimator, as long as the two estimators use the same second-order kernel.
Inference and the closely related problem of nonparametric testing in structural auction models are important and have been receiving increasing attention in the literature. Beginning with the fundamental haile2003nonparametric's test for common values, recent contributions include testing for the monotonicity of bidding strategies (liuvuong2013), endogenous entry (li2009entry and marmer2013model), common versus private values (hill2013there), the affiliation of bidder valuations (jun2010consistent, li2010testing and de2010testing), and inferences on bidder risk attitudes (fang2014inference). In the absence of the asymptotic distribution framework for the GPV estimator, these papers have adopted problem-specific approaches in each case.
The first main result we show in this paper is that the GPV estimator is asymptotically normal. The key difficulty is the presence of the nonparametric first step, which provides nonparametric estimates of the valuations in each auction. In the second step, the kernel density estimator is applied to those estimates, rather than the true valuations. This creates a unique challenge, to our knowledge not previously addressed in the econometrics literature. Our main insight is that the leading term in an asymptotic expansion of the estimator can be viewed as a V-statistic with a kernel that depends on the bandwidth. A projection argument shows that the distribution of this V-statistic is asymptotically normal. Using maximal inequalities for empirical processes and U-processes developed in recent literature, we show that the remainder term is uniformly negligible. The proof is rather long due to an intricate nature of the estimator, and involves some delicate steps.
Note that a working paper version of GPV GPV1995 also has an asymptotic normality result for the GPV estimator. However, the result therein is of a limited nature as it relies on a particular choice of tuning parameters, which insures that only the second stage of the estimating procedure contributes to the asymptotic variance. Thus, in their approach, the uncertainty due to the estimation of valuations in the first stage can be ignored asymptotically, which is achieved by applying different rates of smoothing of auction-specific covariates at both stages. The approach is restrictive in two respects: (i) While GPV's smoothing strategy makes first-stage estimation errors negligible asymptotically, in finite samples their contribution to the variance may still be significant. Our approach takes into account the contribution of both stages and, as a result, is more accurate in finite samples. (ii) Equally importantly, their approach cannot be applied in cases with no auction-specific covariates or when covariates are modeled semi-parametrically as in haile2003nonparametric. Note that treating auction-specific characteristics semi-parametrically is particularly appealing to practitioners.
One unusual feature of our asymptotic normality result concerns the form of the asymptotic variance of the GPV estimator. Typically, the asymptotic variances of kernel density estimators depend on the integral of the squared kernel, which is a known constant that can be easily computed analytically or numerically.\footnote{See, e.g., li2007net.} However, in the case of the GPV estimator, this constant is replaced by a convoluted integral transformation that involves the kernel function, its derivative, and the derivatives of the bidding strategy. This is a consequence of the two-step nature of the GPV estimator and happens due to the impact of the estimation errors from the first stage of the procedure on the distribution of the estimator. Since the bidding strategy is unknown, this fact complicates estimation of the asymptotic variance of the GPV estimator.
Our second contribution is to propose a consistent estimator of the asymptotic variance that avoids estimation of the bidding strategy and its derivative. Its uniform rate of convergence is established using the maximal inequalities. Our third contribution is to show the validity of the percentile bootstrap for the GPV estimator, which allows constructing confidence intervals without estimation of the asymptotic variance.
Our pointwise asymptotic normality results can be used for inference on the optimal reserve price, as the latter is determined by a nonlinear equation in the PDF of valuations haile2003iim. In our fourth contribution, however, we extend the pointwise results and develop valid uniform confidence bands for the PDF. The uniform confidence bands can be used, e.g., for specification of valuations' density. The extension utilizes the uniform rates of convergence of the remainder terms in our V-statistic approximation of the GPV estimator and its Hoeffding decomposition; it also relies on Gaussian anti-concentration inequalities and Gaussian coupling theorems developed in recent literature. This approach, referred to as the Intermediate Gaussian Approximation (IGA, hereafter) in the literature, is based on the seminal work of chernozhukov2014anti,chernozhukov2014gaussian,chernozhukov2016empirical. chernozhukov2014gaussian showed that although a random function based on nonparametric estimation errors does not typically weakly converge to any tight Gaussian random element, under certain conditions the supremum of its studentized version can be often approximated by the supremum of a tight Gaussian random element, the distribution of which changes with the sample size. chernozhukov2016empirical,chernozhukov2014anti showed that under certain conditions the distribution of the Gaussian supremum can be approximated by bootstrapping, and the bootstrap consistency can be shown by applying the coupling theorems and the Gaussian anti-concentration inequality developed in these papers.\footnote{See, e.g., Kato_Sasaki_2,Kato_Sasaki_1 for recent applications of these theorems for constructing confidence bands for different nonparametric curves. } Our paper is one of the first applications of these results. Our Monte Carlo simulation results show that the IGA approach produces confidence bands with excellent finite-sample coverage properties.
Our paper is also related to the recent literature on nonparametrically generated regressors in nonparametric regression. See, e.g., rilstone1996nonparametric, pinkse2001nonparametric, and mammen2012nonparametric. Note, however, that while that literature is concerned with nonparametrically estimated exogenous covariates, we deal with kernel estimation of the density of a nonparametrically generated “dependent” variable, potentially in presence of observable conditioning variables.
The rest of the paper proceeds as follows. Section (ref) introduces the data-generating process (DGP) and describes the GPV estimator in detail. Due to complexity of the estimator, in Section (ref) we show the asymptotic normality of the GPV estimator in a simplified model that has a constant number of bidders across auctions and no auction-specific heterogeneity. Such a simplification allows us to present the main ideas in a more transparent fashion. In Section (ref), we derive an estimator for the asymptotic variance and establish its uniform rate of convergence. We also show consistency of the percentile bootstrap confidence intervals. Section (ref) provides results on constructing valid confidence bands within the same simplified framework. Proofs of the results in Sections (ref)\textendash (ref) are given in the Appendix. Section (ref) provides corresponding theorems in the general model with a random number of bidders and auction-specific heterogeneity. The proofs of these results can be found in the Supplement (included). Section (ref) discusses how our approach can be extended to auctions with binding reserve prices. We report the results from our Monte Carlo study in Section (ref). Section (ref) concludes.
The econometrician observes data from $L$ auctions. Let $\boldsymbol{X}_{l}$ denote the $d$-dimensional relevant characteristics for the object in the $l$-th auction. Let $N_{l}$ denote the number of bidders in the $l$-th auction. Let $B_{il}$ denote the bid submitted by the $i$-th bidder in the $l$-th auction. The data observed by the econometrician is given by $\left\{ \left(B_{il},\boldsymbol{X}_{l},N_{l}\right):i=1,...,N_{l},\,l=1,...,L\right\} .$ Unobserved bidders' valuations of the $l$-th auctioned object are denoted by $\left\{ V_{il}:i=1,...,N_{l},\,l=1,...,L\right\} .$ The following assumption describes the DGP.\footnote{Assumption (ref) is similar to Assumptions A1 and A2 of GPV and Marmer_Shneyerov_Quantile_Auctions. Part (f) imposes the condition that the valuations and the random number of bidders are independent conditionally on the characteristics. See Footnote 14 of GPV.}
Bidders' valuations are not directly observable. Following GPV, we assume that $B_{il}$ is the equilibrium bid of risk-neutral bidder $i$ submitted in the $l$-th auction. Therefore the valuations are linked to the observed bids through the Bayesian Nash equilibrium (BNE) bidding strategy:
\footnotetext{See Equations (1) and (8) of GPV.}Under Assumptions (ref)(a) and (ref)(f), $\left\{ B_{il}:i=1,...,N_{l}\right\} $ are conditionally i.i.d. draws given $\boldsymbol{X}_{l}$ and $N_{l}$. Let $\overline{b}\left(\boldsymbol{x},n\right)\coloneqq s\left(\overline{v}\left(\boldsymbol{x}\right),\boldsymbol{x},n\right)$ and $\underline{b}\left(\boldsymbol{x}\right)\coloneqq\underline{v}\left(\boldsymbol{x}\right)$. Proposition 1(i) of GPV shows that the support of $\left(B_{il},\boldsymbol{X}_{l},N_{l}\right)$ is $\left\{ \left(b,\boldsymbol{x},n\right):n\in\mathcal{N},\,\left(b,\boldsymbol{x}\right)\in\mathcal{S}_{B,\boldsymbol{X}}^{n}\right\} ,$ where $\mathcal{S}_{B,\boldsymbol{X}}^{n}\coloneqq\left\{ \left(b,\boldsymbol{x}\right):\boldsymbol{x}\in\mathcal{X},\,b\in\left[\underline{b}\left(\boldsymbol{x}\right),\overline{b}\left(\boldsymbol{x},n\right)\right]\right\} .$
Let $G\left(\cdot|\boldsymbol{x},n\right)$ denote the conditional CDF of $B_{il}$ given $\boldsymbol{X}_{l}=\boldsymbol{x}$ and $N_{l}=n$. Let $g\left(\cdot|\boldsymbol{x},n\right)$ be the corresponding conditional PDF. GPV established identification of the inverse bidding strategy:
By replacing $G\left(\cdot|\cdot,\cdot\right)$ and $g\left(\cdot|\cdot,\cdot\right)$ in ((ref)) with their nonparametric estimators, GPV proposed an estimator of $\xi\left(\cdot,\cdot,\cdot\right)$, denoted by $\widehat{\xi}\left(\cdot,\cdot,\cdot\right)$. The GPV estimator of $f\left(v|\boldsymbol{x}\right)$ is the kernel density estimator that, in place of the true valuations, uses the so-called pseudo valuations $\left\{ \widehat{V}_{il}\coloneqq\widehat{\xi}\left(B_{il},\boldsymbol{X}_{l},N_{l}\right):i=1,...,N_{l},\,l=1,...,L\right\} $.
Below we provide the details of GPV's estimation procedure. Let $K_{0}$ and $K_{1}$ be univariate kernel functions of different orders satisfying the following assumption:
With $K_{g}\coloneqq K_{1}$ and the multi-dimensional product kernels \[ K_{f}\left(v,\boldsymbol{x}\right)\coloneqq K_{0}\left(v\right)\cdot\prod_{k=1}^{d}K_{0}\left(x_{k}\right)\textrm{ and }K_{\boldsymbol{X}}\left(\boldsymbol{x}\right)\coloneqq\prod_{k=1}^{d}K_{1}\left(x_{k}\right),\textrm{ for }v\in\mathbb{R},\,\boldsymbol{x}=\left(x_{1},...,x_{d}\right)\in\mathbb{R}^{d}, \] define the following nonparametric estimators: \[ \widehat{\varphi}\left(\boldsymbol{x}\right)\coloneqq\frac{1}{L}\sum_{l=1}^{L}\frac{1}{h^{d}}K_{\boldsymbol{X}}\left(\frac{\boldsymbol{X}_{l}-\boldsymbol{x}}{h}\right)\textrm{ and }\widehat{\pi}\left(n|\boldsymbol{x}\right)\coloneqq\frac{1}{\widehat{\varphi}\left(\boldsymbol{x}\right)L}\sum_{l=1}^{L}\mathbbm{1}\left(N_{l}=n\right)\frac{1}{h^{d}}K_{\boldsymbol{X}}\left(\frac{\boldsymbol{X}_{l}-\boldsymbol{x}}{h}\right), \] where $\widehat{\varphi}\left(\cdot\right)$ is the kernel density estimator of $\varphi$ and $\widehat{\pi}\left(\cdot|\cdot\right)$ is the Nadaraya-Watson estimator of the conditional probability mass function $\pi\left(\cdot|\cdot\right)$. Based on these, we define below the nonparametric estimators of the conditional CDF and PDF of the bids:
Consider a partition of $\mathbb{R}^{d}$ with generic half-open hypercubes of side $h_{\partial}>0$: \[ \Pi_{k_{1},...,k_{d}}\coloneqq\left[k_{1}h_{\partial},\left(k_{1}+1\right)h_{\partial}\right)\times\cdots\times\left[k_{d}h_{\partial},\left(k_{d}+1\right)h_{\partial}\right), \] where $\left(k_{1},...,k_{d}\right)$ runs over $\mathbb{Z}^{d}$. Let $\Pi_{h_{\partial}}\left(\boldsymbol{x}\right)$ denote the hypercube that contains $\boldsymbol{x}$ in this partition. Define
to be the estimators of the boundaries of the support. Note that the estimators of the boundaries are super consistent. Let $\mathcal{\widehat{S}}_{B,\boldsymbol{X}}^{n}\coloneqq\left\{ \left(b,\boldsymbol{x}\right):\boldsymbol{x}\in\mathcal{X},\,b\in\left[\widehat{\underline{b}}\left(\boldsymbol{x}\right),\widehat{\overline{b}}\left(\boldsymbol{x},n\right)\right]\right\} .$ The support of $\left(B_{il},\boldsymbol{X}_{l},N_{l}\right)$ then can be estimated by $\left\{ \left(b,\boldsymbol{x},n\right):n\in\mathcal{N},\,\left(b,\boldsymbol{x}\right)\in\mathcal{\widehat{S}}_{B,\boldsymbol{X}}^{n}\right\} .$
The kernel density estimator $\widehat{g}\left(b|\boldsymbol{x},n\right)$ is asymptotically biased when $\left(b,\boldsymbol{x}\right)$ is near the boundaries of the support. GPV suggested that trimming should be applied to the observations near the estimated boundaries using the trimming factor $\mathbb{T}_{il}\coloneqq\mathbbm{1}\left(\mathbb{H}\left(\left(B_{il},\boldsymbol{X}_{l}\right),2h\right)\subseteq\mathcal{\widehat{S}}_{B,\boldsymbol{X}}^{N_{l}}\right).$ The two-step nonparametric estimator of $f\left(v|\boldsymbol{x}\right)$ developed by GPV is
For deriving the asymptotic properties of the GPV estimator, we make the following assumption on the bandwidths $h$ and $h_{\partial}$.\footnote{Assumption (ref)(a) is the same as the assumption on the rate of bandwidth for Marmer_Shneyerov_Quantile_Auctions's quantile-based estimator. See Assumption 3 therein. }
GPV showed that the optimal uniform convergence rate of their estimator is attained when the bandwidth $h$ is of order $O\left(\left(\nicefrac{\mathrm{log}\left(L\right)}{L}\right)^{\nicefrac{1}{\left(2R+3+d\right)}}\right)$. Note that the bandwidth in Assumption (ref) is of smaller order. Under-smoothing imposed in Assumption (ref) is needed to control the asymptotic bias of the GPV estimator, which is important for the validity of inference.
For clarity of the presentation of the main ideas and results, in this section we first establish pointwise asymptotic normality of the GPV estimator in a simplified version of the model that has a fixed number of bidders and no auction-specific heterogeneity. When there are covariates capturing auction-specific heterogeneity present, these results can be used by treating the covariates additively semi-parametrically as in haile2003nonparametric. In that case, there is no kernel smoothing over the covariates, as the GPV procedure would be applied to the “homogenized” bids, which are constructed as residuals from the parametric regression of the bids against the covariates.
In the simplified model, the econometrician observes data on bids in $L$ identical auctions, with a fixed number of bidders $N$ in each auction: $\left\{ B_{il}:i=1,\ldots,N,\,l=1,\ldots,L\right\} .$ Under Assumption (ref), the valuations $\left\{ V_{il}:i=1,\ldots,N,\,l=1,\ldots,L\right\} $ are i.i.d. with a compact support $\left[\underline{v},\overline{v}\right]\subseteq\mathbb{R}_{+}$, PDF $f$ and CDF $F$. The object of interest is the PDF of the valuation at interior points of $\left[\underline{v},\overline{v}\right]$. Suppose that $v_{l}>\underline{v}$, $v_{u}<\overline{v}$ and $I\coloneqq\left[v_{l},v_{u}\right]$ is an inner closed sub-interval of $\left[\underline{v},\overline{v}\right]$. Fix \[ \overline{\delta}\coloneqq\mathrm{min}\left\{ \nicefrac{\left(\overline{v}-v_{u}\right)}{2},\nicefrac{\left(v_{l}-\underline{v}\right)}{2}\right\} . \]
Under Assumption (ref), $f$ is strictly positive and bounded away from zero on its support and admits at least $R$ continuous derivatives. Lemma A1 of GPV showed that under Assumption (ref), the BNE bidding strategy is strictly increasing and $R+1$ times continuously differentiable. In this simplified framework, the inverse of the BNE bidding strategy is
where $G$ and $g$ are the CDF and PDF of bids respectively. Denote $\overline{b}\coloneqq s\left(\overline{v}\right)$ and $\underline{b}\coloneqq s\left(\underline{v}\right)$. Proposition 1(ii) of GPV shows that under Assumption (ref), $g$ is also bounded away from zero on its support $\left[\underline{b},\overline{b}\right]$:
The inverse bidding strategy ((ref)) can be estimated by \[ \widehat{\xi}\left(b\right)\coloneqq b+\frac{1}{N-1}\frac{\widehat{G}\left(b\right)}{\widehat{g}\left(b\right)}, \] where we use the usual nonparametric estimators of $G$ and $g$: \[ \widehat{G}\left(b\right)\coloneqq\frac{1}{N\cdot L}\sum_{i,l}\mathbbm{1}\left(B_{il}\leq b\right)\textrm{ and }\widehat{g}\left(b\right)\coloneqq\frac{1}{N\cdot L}\sum_{i,l}\frac{1}{h}K_{g}\left(\frac{B_{il}-b}{h}\right), \] where $\sum_{i,l}$ is understood as $\sum_{l=1}^{L}\sum_{i=1}^{N}$.
Let $\widehat{\overline{b}}\coloneqq\mathrm{max}\left\{ B_{il}:i=1,\ldots,N,\,l=1,...,L\right\} $, and $\widehat{\underline{b}}\coloneqq\mathrm{min}\left\{ B_{il}:i=1,\ldots,N,\,l=1,...,L\right\} $. The trimming factor is now simply $\mathbb{T}_{il}\coloneqq\mathbbm{1}\left(\widehat{\underline{b}}+h\leq B_{il}\leq\widehat{\overline{b}}-h\right)$. The GPV estimator of $f\left(v\right)$ is now given by \[ \widehat{f}_{GPV}\left(v\right)=\frac{1}{N\cdot L}\sum_{i,l}\mathbb{T}_{il}\frac{1}{h}K_{f}\left(\frac{\widehat{V}_{il}-v}{h}\right), \] where $K_{f}=K_{0}$ in this simplified framework.
We derive the following stochastic expansion of $\widehat{f}_{GPV}\left(v\right)$ around $f\left(v\right)$:
where $\widetilde{\mathbb{T}}_{il}\coloneqq\mathbbm{1}\left(\left|V_{il}-v\right|\leq\overline{\delta}\right)$ is an infeasible trimming factor and the remainder term is uniform in $v\in I$. In the above expression, the derivative $K_{f}'$ of the kernel function appears due to the linearization of $K_{f}\left(\nicefrac{\left(\widehat{V}_{il}-v\right)}{h}\right)$ around $K_{f}\left(\nicefrac{\left(V_{il}-v\right)}{h}\right)$. The result in ((ref)) shows that the distribution of the GPV estimator depends not only on the variation in $V_{il}$'s, but also on the estimation errors of pseudo valuations. In other words, the errors from estimation of the inverse bidding strategy affect the asymptotic distribution of the GPV estimator.
Since $\widehat{G}$ has a faster rate of convergence than $\widehat{g}$, the discrepancy between $V_{il}$ and $\widehat{V}_{il}$ depends on that between the true PDF $g\left(B_{il}\right)$ and the estimated PDF $\widehat{g}\left(B_{il}\right)$, which in turn depends on the averaged discrepancy between $K_{g}\left(\nicefrac{\left(B_{jm}-B_{il}\right)}{h}\right)$ and $g\left(B_{il}\right)$, where the averaging is across $B_{jm}$'s. Lemma (ref) establishes a further asymptotic expansion for the GPV estimator:
where the remainder term is uniform in $v\in I$, and
For any fixed $v$, the leading term in ((ref)) is a V-statistic with a kernel that depends on the bandwidth $h$. We now apply Hoeffding decomposition to this leading term. Define
and further, \[ \mathcal{M}_{2}\left(b;v\right)\coloneqq\int\mathcal{M}\left(b',b;v\right)\mathrm{d}G\left(b'\right),\textrm{ and }\mu_{\mathcal{M}}\left(v\right)\coloneqq\int\int\mathcal{M}\left(b,b';v\right)\mathrm{d}G\left(b\right)\mathrm{d}G\left(b'\right). \] Note that $\mu_{\mathcal{M}}\left(v\right)=\mathrm{E}\left[\mathcal{M}_{1}\left(B_{11};v\right)\right]=\mathrm{E}\left[\mathcal{M}_{2}\left(B_{11};v\right)\right]$. The Hoeffding decomposition yields
In the proof of Theorem (ref) below, we use results for empirical processes and U-processes to show that the terms in the third and fourth lines of ((ref)) and $\left(N\cdot L\right)^{-1}\sum_{i,l}\mathcal{M}_{1}\left(B_{il};v\right)$ are asymptotically negligible uniformly in $v\in I$. As is apparent from the definition of $\mathcal{M}_{1}$ in ((ref)), the contribution of the $\mathcal{M}_{1}\left(b;v\right)$ terms is negligible because they depend on the difference between the expectation $\mathrm{E}\left[h^{-1}K_{g}\left(\nicefrac{\left(B_{il}-b\right)}{h}\right)\right]$ and $g\left(b\right)$, i.e., the bias of the kernel density estimator, which is of order $O\left(h^{1+R}\right)$. Thus, the asymptotic distribution of the GPV estimator is driven solely by
where $\mathcal{M}_{2}\left(B_{il};v\right)-\mu_{\mathcal{M}}\left(v\right)$, $i=1,...,N$, $l=1,...,L$ are independent, zero-mean and depend on the bandwidth.
We also show in the proof of Theorem (ref) that the rescaled variance of ((ref)) satisfies
where the remainder term is uniform in $v\in I$. Thus, the asymptotic variance of the GPV estimator is the limit of the leading term in ((ref)) as $h\downarrow0$. Note that
We have the following result.
If the asymptotic variance $\mathrm{V}_{GPV}(v)$ can be consistently estimated by some estimator $\widehat{\mathrm{V}}_{GPV}\left(v\right)$, one can construct an asymptotically valid pointwise confidence interval for $f\left(v\right)$ as
where $z_{1-\nicefrac{\alpha}{2}}$ denotes the $1-\nicefrac{\alpha}{2}$ quantile of the standard normal distribution.
While the formula for the asymptotic variance in ((ref)) can be used for plug-in estimation of $\mathrm{V}_{GPV}(v)$, such an estimator would be difficult to implement in practice. Firstly, it would require estimating the bidding strategy and its derivative. Secondly, even with an estimate of $s'\left(v\right)$, it is not always easy to compute analytically the double integral in the definition of $\mathrm{V}_{GPV}\left(v\right)$. This issue becomes even more severe when there is auction-specific heterogeneity, as we discuss in Section (ref). In that case, one would need to evaluate a multidimensional integral.
To avoid those issues, we propose an alternative approach to estimation of the asymptotic variance. As we discuss in the previous section, the asymptotic variance of the GPV estimator is the limit of the expression in ((ref)). The leading term on the right-hand side of ((ref)) can be estimated using a U-type-statistic, while replacing the unknown $G$, $g$, and $\xi$ with $\widehat{G}$, $\widehat{g}$, and $\widehat{\xi}$ respectively. The resulting estimator is given by
The estimator avoids estimation of the bidding strategy and its derivative and evaluation of multidimensional integrals. It is very easily implementable in practice since it depends only on the bids, $\widehat{G}$, $\widehat{g}$, and the pseudo valuations. The next theorem shows consistency of the proposed estimator and provides an estimate of its uniform convergence rate.
An alternative to the confidence interval ((ref)) is the bootstrap. We show below that the bootstrap approximation to the distribution of $S\left(v\right)\coloneqq\left(Lh^{3}\right)^{\nicefrac{1}{2}}\left(\widehat{f}_{GPV}\left(v\right)-f\left(v\right)\right)$ is asymptotically valid. Our focus is on the percentile bootstrap as it does not require estimation of the asymptotic variance, which makes it fairly popular among practitioners.
Let $\left\{ B_{il}^{*}:i=1,\ldots,N,l=1,\ldots,L\right\} $ denote the bootstrap sample, i.e., a set of independent random variables drawn from the distribution $\widehat{G}$ conditionally on the original sample of bids. Let $\widehat{G}^{*}$ and $\widehat{g}^{*}$ denote the bootstrap analogues of $\widehat{G}$ and $\widehat{g}$ respectively: they are constructed by following exactly the same procedure as that for constructing $\widehat{G}$ and $\widehat{g}$, however using the (empirical) bootstrap sample instead of the original sample. Let $\widehat{\xi}^{*}$ be the bootstrap analogue of $\widehat{\xi}$ defined using $\widehat{G}^{*}$ and $\widehat{g}^{*}$ in place of $\widehat{G}$ and $\widehat{g}$. We generate bootstrap samples of pseudo values as $\widehat{V}_{il}^{*}\coloneqq\widehat{\xi}^{*}\left(B_{il}^{*}\right)$. Lastly, we construct a bootstrap analogue of $\widehat{f}_{GPV}\left(v\right)$: \[ \widehat{f}_{GPV}^{*}\left(v\right)\coloneqq\frac{1}{N\cdot L}\sum_{i,l}\mathbb{T}_{il}^{*}\frac{1}{h}K_{f}\left(\frac{\widehat{V}_{il}^{*}-v}{h}\right), \] where $\mathbb{T}_{il}^{*}\coloneqq\mathbbm{1}\left(\widehat{\underline{b}}+h\leq B_{il}^{*}\leq\widehat{\overline{b}}-h\right)$.
Let $q_{\tau}^{*}(v)$ be the $\tau$-th quantile of the conditional distribution of $\widehat{f}_{GPV}^{*}\left(v\right)$ given the original sample. The percentile bootstrap confidence interval is \[ CI^{*}\left(v\right)\coloneqq\left[q_{\nicefrac{\alpha}{2}}^{*}(v),\,q_{1-\nicefrac{\alpha}{2}}^{*}(v)\right]=\left[\widehat{f}_{GPV}\left(v\right)+\frac{s_{\nicefrac{\alpha}{2}}^{*}(v)}{\sqrt{Lh^{3}}},\,\widehat{f}_{GPV}\left(v\right)+\frac{s_{\nicefrac{1-\alpha}{2}}^{*}(v)}{\sqrt{Lh^{3}}}\right], \] where $s_{\tau}^{*}(v)$ is the $\tau-$th quantile of the conditional distribution of \[ S^{*}\left(v\right)\coloneqq\left(Lh^{3}\right)^{\nicefrac{1}{2}}\left(\widehat{f}_{GPV}^{*}\left(v\right)-\widehat{f}_{GPV}\left(v\right)\right) \] given the original sample. The conditional distributions of the bootstrap statistics $\widehat{f}_{GPV}^{*}\left(v\right)$ and $S^{*}\left(v\right)$ given the original sample can be easily approximated by Monte Carlo methods. We show below that the bootstrap estimator of the finite-sample distribution of $S\left(v\right)$ is consistent. Let $\mathrm{P}^{*}\left[\cdot\right]$ denote the conditional probability given the original sample of bids.
Consider the stochastic process
Note that $\mathrm{E}\left[\mathit{\Gamma}\left(v\right)\right]=0$ and $\mathrm{E}\left[\mathit{\Gamma}\left(v\right)^{2}\right]=1$ for all $v\in I$.
The following theorem shows that (a version of) the centered Gaussian process with index set $I$ and covariance function $\mathrm{E}\left[\mathit{\Gamma}\left(v\right)\mathit{\Gamma}\left(v'\right)\right]$, for $\left(v,v'\right)\in I^{2}$, is a tight random element in $\ell^{\infty}\left(I\right)$. This Gaussian process, denoted by $\left\{ \mathit{\Gamma}_{G}\left(v\right):v\in I\right\} $, is the intermediate Gaussian process. The tightness of $\varGamma_{G}$ as a random element in $\ell^{\infty}\left(I\right)$ can be established using standard results (see, e.g., chernozhukov2014gaussian). The following theorem also shows that one can approximate the distribution of the sup-norm $\left\Vert Z\right\Vert _{I}=\underset{v\in I}{\mathrm{sup}}\left|Z\left(v\right)\right|$ with that of $\mathit{\Gamma}_{G}$. The result follows from uniform approximations of ((ref)) and $\widehat{\mathrm{V}}_{GPV}\left(v\right)$ and uses the coupling theorem for suprema of empirical processes of chernozhukov2014gaussian and the Gaussian anti-concentration inequality of chernozhukov2014anti.
The next result shows that the distribution of the sup-norm of the (empirical) bootstrap process
can be similarly approximated by that of $\mathit{\Gamma}_{G}$.
Since the distributions of the suprema of (the absolute values of) $\left\{ Z\left(v\right):v\in I\right\} $ and that of $\left\{ Z^{*}\left(v\right):v\in I\right\} $ are both well approximated by that of $\left\{ \varGamma_{G}\left(v\right):v\in I\right\} $, one can use the bootstrap critical values based on $\left\Vert Z^{*}\right\Vert _{I}$ for construction of uniform confidence bands. Let
be the $\left(1-\alpha\right)$-quantile of the conditional distribution of $\left\Vert Z^{*}\right\Vert _{I}$ given the original sample. The uniform confidence band is given by \[ CB^{*}\left(v\right)\coloneqq\left[\widehat{f}_{GPV}\left(v\right)-\zeta_{L,\alpha}^{*}\sqrt{\frac{\widehat{\mathrm{V}}_{GPV}\left(v\right)}{Lh^{3}}},\,\widehat{f}_{GPV}\left(v\right)+\zeta_{L,\alpha}^{*}\sqrt{\frac{\widehat{\mathrm{V}}_{GPV}\left(v\right)}{Lh^{3}}}\right],\,\textrm{for \ensuremath{v\in I}}. \] The following corollary establishes its asymptotic validity and provides an estimate of the order of the bootstrap critical value $\zeta_{L,\alpha}^{*}$.\footnote{It also implies that the (supremum) width of the band $CB^{*}$ is of order $O_{p}\left(\mathrm{log}\left(h^{-1}\right)^{\nicefrac{1}{2}}\left(Lh^{3}\right)^{-\nicefrac{1}{2}}\right)$.}
We now turn to the general model with auction-specific heterogeneity and a random number of bidders. Firstly, we establish the asymptotic normality of the GPV estimator by following the same approach and steps as in the case of the simplified model in Section (ref). While handling the general case is complicated by much heavier notations, all the results provided in this section can be viewed as straightforward generalizations of the results in Sections (ref)-(ref). The proofs of the results for the general case can be found in the Supplement.
In comparison with the simplified model, one of the main differences is in the form of the asymptotic variance. Recall that in the simplified case, the asymptotic variance of the GPV estimator depends on the derivative of the bidding strategy. As we show below, when there is auction-specific heterogeneity, the asymptotic variance also involves the partial derivatives of the bidding strategy with respect to the auction-specific characteristics. This is in addition to the partial derivative with respect to the valuation.
For some fixed $\boldsymbol{x}$ which is an interior point of $\mathcal{X}$, let $I\left(\boldsymbol{x}\right)\coloneqq\left[v_{l}\left(\boldsymbol{x}\right),v_{u}\left(\boldsymbol{x}\right)\right]$ be an inner closed sub-interval of $\left[\underline{v}\left(\boldsymbol{x}\right),\overline{v}\left(\boldsymbol{x}\right)\right]$. The fact that the conditional density of the valuations given $\boldsymbol{X}=\boldsymbol{x}$ and $N=n$ is $f\left(\cdot|\boldsymbol{x}\right)$ under Assumption (ref) motivates the following two-step estimator of $f\left(v|\boldsymbol{x}\right)$: \[ \widehat{f}_{GPV}\left(v|\boldsymbol{x},n\right)\coloneqq\frac{1}{\widehat{\pi}\left(n|\boldsymbol{x}\right)\widehat{\varphi}\left(\boldsymbol{x}\right)L}\sum_{l=1}^{L}\mathbbm{1}\left(N_{l}=n\right)\frac{1}{N_{l}}\sum_{i=1}^{N_{l}}\mathbb{T}_{il}\frac{1}{h^{1+d}}K_{f}\left(\frac{\widehat{V}_{il}-v}{h},\frac{\boldsymbol{X}_{l}-\boldsymbol{x}}{h}\right). \] Note that the above estimator only uses data from auctions with $N_{l}=n$. Since the PDF of valuations does not depend on the number of bidders, an estimator for $f(v|\boldsymbol{x})$ can be constructed as a weighted average of $\left\{ \widehat{f}_{GPV}\left(v|\boldsymbol{x},n\right):n\in\mathcal{N}\right\} $. E.g., GPV suggested using estimates of the conditional probabilities of drawing $N_{l}=n$ as the weights:
Note that this gives an expression that is the same as the right hand side of ((ref)).
By repeating the steps from Section (ref), one can show that the following analogue of the linearization results in ((ref)) holds for the general model:
where $\boldsymbol{B}_{\cdot l}\coloneqq\left(B_{1l},...,B_{N_{l}l}\right)$, and the remainder term is uniform in $v\in I\left(\boldsymbol{x}\right)$. For $\boldsymbol{b}_{\cdot}\coloneqq(b_{1},\ldots,b_{m})$ , the kernel function $\mathcal{M}^{n}$ is given by:
where $K_{f}'\left(\cdot,\cdot\right)$ denotes the partial derivative function of $K_{f}$ with respect to its first argument, \[ G\left(b,\boldsymbol{z},m\right)\coloneqq G\left(b|\boldsymbol{z},m\right)\pi\left(m|\boldsymbol{z}\right)\varphi\left(\boldsymbol{z}\right)\text{ and }g\left(b,\boldsymbol{z},m\right)\coloneqq g\left(b|\boldsymbol{z},m\right)\pi\left(m|\boldsymbol{z}\right)\varphi\left(\boldsymbol{z}\right). \]
Note that the leading term on the right-hand side of ((ref)) involves a V-statistic (with a kernel that depends on the bandwidth) and, therefore, can be analyzed using the Hoeffding decomposition. Thus, ((ref)) can be generalized as
where the remainder term is uniform in $v\in I\left(\boldsymbol{x}\right)$,
and $\mu_{\mathcal{M}^{n}}\left(v\right)\coloneqq\mathrm{E}\left[\mathcal{M}^{n}\left(\left(\boldsymbol{B}_{\cdot1},\boldsymbol{X}_{1},N_{1}\right),\left(\boldsymbol{B}_{\cdot2},\boldsymbol{X}_{2},N_{2}\right);v\right)\right]$.
The projection term $\mathcal{M}_{1}^{n}$ is the expectation of the kernel $\mathcal{M}^{n}\left(\left(\boldsymbol{b}.,\boldsymbol{z},m\right),\left(\boldsymbol{B}_{\cdot1},\boldsymbol{X}_{1},N_{1}\right);v\right)$ with the first argument fixed at $\left(\boldsymbol{b}.,\boldsymbol{z},m\right)$. The expression for $\mathcal{M}_{1}^{n}$ is:
As in the case of the simplified model, the contribution of $\mathcal{M}_{1}^{n}$ is asymptotically negligible. This happens for the same reason as in Section (ref): $\mathcal{M}_{1}^{n}$ depends on the difference between the expectation of the kernel function and the true density. Hence, the asymptotic distribution of the GPV estimator is driven solely by the $\mathcal{M}_{2}^{n}$ term, which is the expectation of the kernel $\mathcal{M}^{n}\left(\left(\boldsymbol{B}_{\cdot1},\boldsymbol{X}_{1},N_{1}\right),\left(\boldsymbol{b}.,\boldsymbol{z},m\right);v\right)$ with the second argument fixed at $\left(\boldsymbol{b}.,\boldsymbol{z},m\right)$.
We show in the supplement that a generalized version of ((ref)) holds:
where the remainder term is uniform in $v\in I\left(\boldsymbol{x}\right)$. Moreover, similarly to ((ref)),
where $s_{v}$ and $s_{\boldsymbol{x}}$ denote the partial derivatives of the bidding function:
The asymptotic variance of the GPV estimator is the limit of $\left(\pi\left(n|\boldsymbol{x}\right)\varphi\left(\boldsymbol{x}\right)\right)^{-2}\mathrm{V}_{\mathcal{M}}\left(v|\boldsymbol{x},n\right)$. After using a change of variable argument and ((ref)), the variance is shown to be
The following theorem is a generalization of Theorem (ref). The proof of the theorem as well as the proofs of all other results provided in Section (ref) are in the Supplement.
For practical purposes, it is important to have a consistent estimator of the asymptotic variance $\mathrm{V}_{GPV}\left(v|\boldsymbol{x}\right)$ that avoids estimation of the bidding strategy $s\left(\cdot,\cdot,n\right)$ and its derivatives. It is also highly desirable to avoid analytical or numerical evaluation of a multidimensional integral in the definition of the asymptotic variance. Following the same approach we used in the case of the simplified model (see ((ref))), we rely on the sample analogue of ((ref)):
where
The following result is a generalization of Theorem (ref).
To generate bootstrap samples, we apply the same resampling procedure as that proposed in Marmer_Shneyerov_Quantile_Auctions. First, we randomly draw $L$ observations from $\left\{ \left(\boldsymbol{X}_{l},N_{l}\right):l=1,...,L\right\} $ (i.e., the auction-specific characteristics) with replacement. Next, we randomly draw bids with replacement from the bids corresponding to each selected auction. Given $\left(\boldsymbol{X}_{l}^{*},N_{l}^{*}\right)=\left(\boldsymbol{X}_{l'},N_{l'}\right)$ in the first step, in the second step $\{B_{il}^{*}:i=1,...,N_{l}^{*}\}$ is generated as an empirical bootstrap sample drawn from $\{B_{il'}:i=1,...,N_{l'}\}$. Let $\widehat{\xi}^{*}\left(\cdot,\cdot,\cdot\right)$ and $\widehat{\varphi}^{*}\left(\cdot\right)$ be the bootstrap analogues of $\widehat{\xi}\left(\cdot,\cdot,\cdot\right)$ and $\widehat{\varphi}\left(\cdot\right)$ respectively. Let $\widehat{f}_{GPV}^{*}\left(v|\boldsymbol{x}\right)$ denote the bootstrap version of the GPV estimator: \[ \widehat{f}_{GPV}^{*}\left(v|\boldsymbol{x}\right)\coloneqq\frac{1}{\widehat{\varphi}^{*}\left(\boldsymbol{x}\right)L}\sum_{l=1}^{L}\frac{1}{N_{l}^{*}}\sum_{i=1}^{N_{l}^{*}}\mathbb{T}_{il}^{*}\frac{1}{h^{1+d}}K_{f}\left(\frac{\widehat{V}_{il}^{*}-v}{h},\frac{\boldsymbol{X}_{l}^{*}-\boldsymbol{x}}{h}\right), \] where $\widehat{V}_{il}^{*}\coloneqq\widehat{\xi}^{*}\left(B_{il}^{*},\boldsymbol{X}_{l}^{*},N_{l}^{*}\right)$ and the bootstrap version of the trimming factor is given by \[ \mathbb{T}_{il}^{*}\coloneqq\mathbbm{1}\left(\mathbb{H}\left(\left(B_{il}^{*},\boldsymbol{X}_{l}^{*}\right),2h\right)\subseteq\mathcal{\widehat{S}}_{B,\boldsymbol{X}}^{N_{l}}\right). \]
Consider the scaled deviation of the GPV estimator from the true PDF, and its bootstrap analogue:
The following result, which is a generalization of Theorem (ref), establishes the validity of the percentile bootstrap for $f\left(v|\boldsymbol{x}\right)$.
We now turn to construction of uniform confidence bands for $\left\{ f\left(v|\boldsymbol{x}\right):v\in I\left(\boldsymbol{x}\right)\right\} $ given a fixed interior point $\boldsymbol{x}$. Consider the following processes: \[ Z\left(v|\boldsymbol{x}\right)\coloneqq\frac{\widehat{f}_{GPV}\left(v|\boldsymbol{x}\right)-f\left(v|\boldsymbol{x}\right)}{\left(Lh^{3+d}\right)^{-\nicefrac{1}{2}}\widehat{\mathrm{V}}_{GPV}\left(v|\boldsymbol{x}\right)^{\nicefrac{1}{2}}}\textrm{ and }Z^{*}\left(v|\boldsymbol{x}\right)\coloneqq\frac{\widehat{f}_{GPV}^{*}\left(v|\boldsymbol{x}\right)-\widehat{f}_{GPV}\left(v|\boldsymbol{x}\right)}{\left(Lh^{3+d}\right)^{-\nicefrac{1}{2}}\widehat{\mathrm{V}}_{GPV}\left(v|\boldsymbol{x}\right)^{\nicefrac{1}{2}}},\textrm{ \ensuremath{v\in I\left(\boldsymbol{x}\right)}}. \] Similarly to the simplified model, the distribution of $\left\Vert Z\left(\cdot|\boldsymbol{x}\right)\right\Vert _{I\left(\boldsymbol{x}\right)}$ can be approximated by the conditional distribution of $\left\Vert Z^{*}\left(\cdot|\boldsymbol{x}\right)\right\Vert _{I\left(\boldsymbol{x}\right)}$. Let $\zeta_{L,\alpha}^{*}$ be the $\left(1-\alpha\right)$-quantile of the conditional distribution of $\left\Vert Z^{*}\left(\cdot|\boldsymbol{x}\right)\right\Vert _{I\left(\boldsymbol{x}\right)}$ given the original sample. Consider the following confidence band: for $\ensuremath{v\in I\left(\boldsymbol{x}\right)}$, \[ CB^{*}\left(v|\boldsymbol{x}\right)\coloneqq\left[\widehat{f}_{GPV}\left(v|\boldsymbol{x}\right)-\zeta_{L,\alpha}^{*}\sqrt{\frac{\widehat{\mathrm{V}}_{GPV}\left(v|\boldsymbol{x}\right)}{Lh^{3+d}}},\,\widehat{f}_{GPV}\left(v|\boldsymbol{x}\right)+\zeta_{L,\alpha}^{*}\sqrt{\frac{\widehat{\mathrm{V}}_{GPV}\left(v|\boldsymbol{x}\right)}{Lh^{3+d}}}\right]. \] The following result, which is a generalization of Corollary (ref), establishes the validity of $CB^{*}\left(\cdot|\boldsymbol{x}\right)$.
Section 4 of GPV shows how to modify their identification and estimation strategy when there is a binding reserve price. Here, we discuss how our approach can be applied in that case.
When there is a binding reserve price, it is assumed that only bidders with valuations exceeding the reserve price submit bids. Thus, one has to distinguish between the numbers of potential and actual (active) bidders. Let $N$ denote the number of potential bidders, which is assumed to be known to players. The bidding strategy depends on the number of potential bidders instead of the number of active bidders. As discussed in GPV, $N$ can be estimated by taking the maximum of the observed numbers of actual bidders across the auctions: $\widehat{N}\coloneqq\mathrm{max}\left\{ N_{l}:l=1,\ldots,L\right\} ,$ where $N_{l}$ is the number of actual bidders in auction $l$.
GPV assume that the reserve price in auction $l$, denoted $P_{0l}$, is some unknown deterministic function of the auction characteristics $\boldsymbol{X}_{l}$: $P_{0l}=p_{0}\left(\boldsymbol{X}_{l}\right)$. The probability of drawing a valuation below the reserve price is given by $\Phi\left(\boldsymbol{X}_{l}\right)\coloneqq F\left(P_{0l}|\boldsymbol{X}_{l}\right)$. The conditional CDF and PDF of the distribution of valuations given participation (submitting a bid) are
respectively. The third displayed equation on page 550 of GPV shows that $\Phi(\cdot)$ can be estimated using a nonparametric regression of the number of actual bidders: \[ \widehat{\Phi}\left(\boldsymbol{x}\right)\coloneqq1-\frac{1}{\widehat{N}\widehat{\varphi}\left(\boldsymbol{x}\right)L}\sum_{l=1}^{L}\frac{1}{h^{d}}N_{l}K_{\boldsymbol{X}}\left(\frac{\boldsymbol{X}_{l}-\boldsymbol{x}}{h}\right). \]
GPV point out that the density of bids is unbounded at the reserve price $p_{0}\left(\boldsymbol{x}\right)$ and, in its neighborhood, behaves as $\nicefrac{1}{\sqrt{b-p_{0}(\boldsymbol{x})}}$ . To avoid technical problems due to the unbounded density, they propose to transform the bids as \[ B_{\dagger il}=\left(B_{il}-P_{0l}\right)^{\nicefrac{1}{2}}. \] The support of $\left(B_{\dagger11},\boldsymbol{X}_{1}\right)$ is given by $\mathcal{S}_{B_{\dagger},\boldsymbol{X}}\coloneqq\left\{ \left(b,\boldsymbol{x}\right):\boldsymbol{x}\in\mathcal{X},\,b\in\left[0,\overline{b}_{\dagger}\left(\boldsymbol{x}\right)\right]\right\} $, where $\overline{b}_{\dagger}\left(\boldsymbol{z}\right)\coloneqq\left(\overline{b}_{\dagger}\left(\boldsymbol{z}\right)-p_{0}\left(\boldsymbol{z}\right)\right)^{\nicefrac{1}{2}}$. The support can be estimated by $\mathcal{\widehat{S}}_{B_{\dagger},\boldsymbol{X}}\coloneqq\left\{ \left(b,\boldsymbol{x}\right):\boldsymbol{z}\in\mathcal{X},\,b\in\left[0,\widehat{\overline{b}}_{\dagger}\left(\boldsymbol{x}\right)\right]\right\} $, where $\widehat{\overline{b}}_{\dagger}\left(\boldsymbol{x}\right)\coloneqq\mathrm{\mathrm{max}}\left\{ B_{\dagger pl}:p=1,...,N_{l},\,\boldsymbol{X}_{l}\in\Pi_{h_{\partial}}\left(\boldsymbol{x}\right),\,l=1,...,L\right\} $, see page 550 in GPV.
Let $G_{\dagger}\left(\cdot|\cdot\right)$ and $g_{\dagger}\left(\cdot|\cdot\right)$ denote respectively the conditional CDF and PDF of the transformed bids $B_{\dagger11}$ given $\boldsymbol{X}_{1}$. Let $G_{\dagger}\left(b_{\dagger},\boldsymbol{z}\right)\coloneqq G_{\dagger}\left(b_{\dagger}|\boldsymbol{z}\right)\varphi\left(\boldsymbol{z}\right)$ and $g_{\dagger}\left(b_{\dagger},\boldsymbol{z}\right)\coloneqq g_{\dagger}\left(b_{\dagger}|\boldsymbol{z}\right)\varphi\left(\boldsymbol{z}\right)$. GPV show that $g_{\dagger}\left(\cdot,\cdot\right)$ is bounded on its support, and that latent valuations can be recovered using \[ V_{il}=\xi_{\dagger}\left(B_{\dagger il},\boldsymbol{X}_{l}\right)\coloneqq P_{0l}+B_{\dagger il}^{2}+\frac{2B_{\dagger il}}{N-1}\frac{\left(1-\Phi\left(\boldsymbol{X}_{l}\right)\right)G_{\dagger}\left(B_{\dagger il},\boldsymbol{X}_{l}\right)+\Phi\left(\boldsymbol{X}_{l}\right)\varphi\left(\boldsymbol{X}_{l}\right)}{\left(1-\Phi\left(\boldsymbol{X}_{l}\right)\right)g_{\dagger}\left(B_{\dagger il},\boldsymbol{X}_{l}\right)}. \]
In the modified GPV procedure, one first estimates $\xi_{\dagger}\left(\cdot,\cdot\right)$ by replacing $N$, $\Phi\left(\cdot\right)$, $G_{\dagger}\left(\cdot,\cdot\right)$, $g_{\dagger}\left(\cdot,\cdot\right)$ and $\varphi\left(\cdot\right)$ with their estimators. $G_{\dagger}\left(\cdot,\cdot\right)$ and $g_{\dagger}\left(\cdot,\cdot\right)$ can be estimated using the transformed bids $B_{\dagger il}$:
In the second step of the modified GPV procedure, one uses the pseudo valuations \[ \left\{ \widehat{V}_{il}\coloneqq\widehat{\xi}_{\dagger}\left(B_{\dagger il},\boldsymbol{X}_{l}\right):i=1,\dots,N_{l},l=1,\ldots,L\right\} , \] where $\widehat{\xi}_{\dagger}\left(\cdot,\cdot\right)$ is the estimated version of $\xi_{\dagger}(\cdot,\cdot)$, in place of latent valuations to construct a kernel density estimator of $f^{\star}\left(v|\boldsymbol{x}\right)$: \[ \widehat{f}_{GPV}^{\star}(v|\boldsymbol{x})\coloneqq\frac{1}{\widehat{\varphi}(\boldsymbol{x})L}\sum_{l=1}^{L}\frac{1}{N_{l}}\sum_{i=1}^{N_{l}}\mathbb{T}_{il}\frac{1}{h^{1+d}}K_{f}\left(\frac{\widehat{V}_{il}-v}{h},\frac{\boldsymbol{X}_{l}-\boldsymbol{x}}{h}\right), \] where the trimming factors $\mathbb{T}_{il}$ can be defined analogously to the case with no binding reserve price: $\mathbb{T}_{il}\coloneqq\mathbbm{1}\left(\mathbb{H}\left(\left(B_{\dagger il},\boldsymbol{X}_{l}\right),2h\right)\subseteq\mathcal{\widehat{S}}_{B_{\dagger},\boldsymbol{X}}\right).$
Our approach can be used to obtain the asymptotic distribution of the modified GPV estimator as follows. In view of the definitions of $\xi_{\dagger}\left(\cdot.\cdot\right)$ and its estimator, the $\widehat{V}_{il}-V_{il}$ term in the analogue of ((ref)) can be expanded as \[ \widehat{V}_{il}-V_{il}=\frac{2B_{\dagger il}}{N-1}\frac{\left(1-\Phi\left(\boldsymbol{X}_{l}\right)\right)G_{\dagger}\left(B_{\dagger il},\boldsymbol{X}_{l}\right)+\Phi\left(\boldsymbol{X}_{l}\right)\varphi\left(\boldsymbol{X}_{l}\right)}{\left(1-\Phi\left(\boldsymbol{X}_{l}\right)\right)g_{\dagger}\left(B_{\dagger il},\boldsymbol{X}_{l}\right)^{2}}\left(\widehat{g}_{\dagger}\left(B_{\dagger il},\boldsymbol{X}_{l}\right)-g_{\dagger}\left(B_{\dagger il},\boldsymbol{X}_{l}\right)\right)+s.o., \] where “$s.o.$” stands for smaller order terms. Hence, the GPV estimator of $f^{\star}\left(v|\boldsymbol{x}\right)$ still has a representation of the same form as in ((ref)):
where $\boldsymbol{B}_{\dagger\cdot l}\coloneqq\left(B_{\dagger1l},...,B_{\dagger N_{l}l}\right)$, and
Similarly to the case with no reserve price, one can apply the Hoeffding decomposition with only $\mathcal{M}_{\dagger2}\left(\boldsymbol{b}.,\boldsymbol{z},m;v\right)\coloneqq\mathrm{E}\left[\mathcal{M}_{\dagger}\left(\left(\boldsymbol{B}_{\dagger\cdot1},\boldsymbol{X}_{1},N_{1}\right),\left(\boldsymbol{b}.,\boldsymbol{z},m\right);v\right)\right]$ contributing to the asymptotic variance:
Similarly to ((ref)), the asymptotic variance of the GPV estimator is now given by the limit of
as $h\downarrow0$, where $\overline{\pi}\left(\boldsymbol{x}\right)\coloneqq\mathrm{E}\left[N_{1}^{-1}\mid\boldsymbol{X}_{1}=\boldsymbol{x}\right]$. Note that conditionally on $\boldsymbol{X}_{1}$, the number of active bidders $N_{1}$ has a binomial distribution with parameters $N$ and $1-\Phi\left(\boldsymbol{X}_{1}\right)$. Lastly, similarly to Theorem (ref), the expression in ((ref)), and for any interior point $v\in\left(p_{0}\left(\boldsymbol{x}\right),\bar{v}\left(\boldsymbol{x}\right)\right)$, the GPV estimator of $f^{\star}\left(v\mid\boldsymbol{x}\right)$ is asymptotically normal with the asymptotic variance given by
where $s_{\dagger}(\cdot,\boldsymbol{x})\coloneqq\xi_{\dagger}^{-1}(\cdot,\boldsymbol{x})$, and the partial derivatives $s_{\dagger v}$ and $s_{\dagger\boldsymbol{x}}$ are defined similarly to $s_{v}$ and $s_{\boldsymbol{x}}$ in ((ref)). Similarly to Corollary (ref), one can show: \[ \left(Lh^{3+d}\right)^{\nicefrac{1}{2}}\left(\widehat{f}^{\star}\left(v\mid\boldsymbol{x}\right)-f^{\star}\left(v\mid\boldsymbol{x}\right)\right)\rightarrow_{d}\mathrm{N}\left(0,\mathrm{V}_{\dagger GPV}\left(v,\boldsymbol{x}\right)\right). \]
To estimate the asymptotic variance $\mathrm{V}_{\dagger GPV}(v,\boldsymbol{x})$, one can use the sample analogue of ((ref)) in the same way as that used to construct the estimator of $\mathrm{V}_{GPV}(v,\boldsymbol{x})$ defined by ((ref)) from ((ref)). As before, the approach does not require estimation of the bidding strategy $s_{\dagger}(v,\boldsymbol{x})$ or its derivatives. The analogue estimator is given by
where $\widehat{\overline{\pi}}\left(\boldsymbol{x}\right)$ is the Nadaraya-Watson estimator of $\overline{\pi}\left(\boldsymbol{x}\right)$, and
Suppose $I\left(\boldsymbol{x}\right)$ is an inner closed sub-interval of $\left[p_{0}\left(\boldsymbol{x}\right),\bar{v}\left(\boldsymbol{x}\right)\right]$. The uniform convergence rate of $\mathrm{\widehat{V}}_{\dagger GPV}\left(v,\boldsymbol{x}\right)$ to ((ref)) can be shown to be the same as that in the statement of Theorem (ref). In view of the definitions in ((ref)), the nonparametric estimator for $f\left(v|\boldsymbol{x}\right)$ is $\left(1-\widehat{\Phi}\left(\boldsymbol{x}\right)\right)\widehat{f}^{\star}\left(v|\boldsymbol{x}\right)$. Since $\widehat{\Phi}\left(\boldsymbol{x}\right)$ converges at a faster rate than the PDF estimator $\widehat{f}^{\star}\left(v\mid\boldsymbol{x}\right)$, one can see that \[ \left(Lh^{3+d}\right)^{\nicefrac{1}{2}}\left(\left(1-\widehat{\Phi}\left(\boldsymbol{x}\right)\right)\widehat{f}^{\star}\left(v\mid\boldsymbol{x}\right)-f\left(v\mid\boldsymbol{x}\right)\right)\rightarrow_{d}\mathrm{N}\left(0,\left(1-\Phi\left(\boldsymbol{x}\right)\right)^{2}\mathrm{V}_{\dagger GPV}\left(v,\boldsymbol{x}\right)\right). \]
A valid uniform confidence band of $\left\{ f\left(v|\boldsymbol{x}\right):v\in I\left(\boldsymbol{x}\right)\right\} $ can be constructed by adapting the methods described in Section (ref).
In this section, we assess the finite-sample coverage accuracy of the uniform confidence bands. Our simulation design follows \citet*{Marmer_Shneyerov_Quantile_Auctions}, and the DGP is described in Remark (ref). We consider $\theta\in\left\{ 1,2\right\} $ and draw valuations from $f_{\theta}$. We choose the triweight kernel when implementing the two-step estimator. We used the second-order triweight kernel in the second step and used the fourth-order triweight kernel in the first step.
We need to choose the bandwidths in the first step when we construct the pseudo valuations and the second step when we implement kernel density estimation using the pseudo valuations. We follow GPV (see Section 2.4) and use $h_{g}=3.72\cdot\widehat{\sigma}_{b}\cdot\left(N\cdot L\right)^{-\nicefrac{1}{5}}$as the first-step bandwidth, where $\widehat{\sigma}_{b}$ is the estimated standard deviation of the observed bids. We use $h_{f}=3.15\cdot\widehat{\sigma}_{v}\cdot\left(\left(N\cdot L\right)_{\mathbb{T}}\right)^{-\nicefrac{1}{5}}$as the second-step bandwidth, where $\widehat{\sigma}_{v}$ is the estimated standard deviation of the trimmed pseudo valuations and $\left(N\cdot L\right)_{\mathbb{T}}$ is the number of bids remaining after the trimming. The constants $3.72$ and $3.15$ are Silverman's rule-of-thumb constants corresponding to fourth-order and second-order triweight kernels.\footnote{See li2007net for a description of the Silverman approach; see also li2003semiparametric.} We consider different numbers of bidders $N\in\left\{ 3,5,7\right\} $, and also the density function over different ranges: $v\in\left[0.2,0.8\right]$ and $v\in\left[0.3,0.7\right]$.\footnote{We use grid maximization, where the grid is chosen as $[v_{l}:0.001:v_{u}]$. We have also tried a finer grid $[v_{l}:0.0001:v_{u}],$ which produced similar results.} The number of auctions $L$ is chosen so the total number of observations is fixed as $N\cdot L=2100$.
In Table (ref), we report our simulation results for the bootstrap-based IGA uniform confidence band $CB^{*}$. We find that the IGA bootstrap approach provides accurate coverage probabilities. Additional simulation results are reported in the Supplement.
The GPV estimator has proven to be the essential input in virtually all nonparametric structural auction models. By proving the asymptotic normality and the first-order validity of the bootstrap uniform confidence bands, this paper completes the econometric theory of the GPV estimator and opens way to new applications.
Our pointwise asymptotic normality results can be used for inference on an important policy variable: the optimal reserve price. As discussed, e.g., in haile2003iim, the optimal reserve price $r(\boldsymbol{x})$ in auctions with $\boldsymbol{X}_{l}=\boldsymbol{x}$ satisfies the following equation: \[ r(\boldsymbol{x})-\frac{1-F(r(\boldsymbol{x})|\boldsymbol{x})}{f(r(\boldsymbol{x})|\boldsymbol{x})}=c(\boldsymbol{x}), \] where $c(\boldsymbol{x})$ is the seller's own valuation. Suppose that the estimator $\widehat{r}(\boldsymbol{x})$ is constructed by solving an estimated version of the above equation with $f(\cdot|\boldsymbol{x})$ replaced by its GPV estimator $\widehat{f}_{GPV}\left(\cdot|\boldsymbol{x}\right)$. In that case, our pointwise normality results imply the asymptotic normality of the estimated optimal reserve price:
Our results for the validity of the percentile bootstrap of the GPV estimator naturally carry over to the above estimator of the optimal reserve price. Thus in practice, one can use the percentile bootstrap for inference on the optimal reserve price.
Our uniform confidence bands can be used for specification of the density of valuations.
In future research, our results could be extended in several directions. Below, we briefly describe some of potentially interesting extensions. These extensions would address the limitations of the independent private values model that underlies the GPV estimator.
First, in GPV the bidders are treated symmetrically. Empirically this is not always the case. See, e.g., flambard2006asymmetry for an application to snow removal contracts. Second, we abstract from the empirically relevant issue of unobserved heterogeneity, as in krasnokutskaya2011identification, hu2013identification and roberts2013unobserved. Third, correlated values may also be important empirically. li2002structural and hubbard2012semiparametric extend the GPV estimator to the affiliated value environment. Fourth, guerre2009nonparametric and zincenko2018nonparametric provide extensions to an environment with risk-averse bidders. Fifth, the literature on endogenous entry in auctions has developed rapidly. See, e.g., li2009entry, krasnokutskaya2011bid, marmer2013model, roberts2013should and gentry2014identification.
In the above models, the basic GPV estimator is often adapted to suit the needs of a particular application. But most of these estimators share the underlying two-step structure of the GPV estimator. We conjecture that our main results will also prove useful for establishing the asymptotic normality and validity of certain uniform confidence bands for these GPV-like estimators.
A remaining unresolved important practical issue is bandwidth selection for the GPV estimator. It is possible that the recent advances in that area (e.g., calonico2014robust and armstrong2016simple) can be adapted to the framework of GPV.