EconBase
← Back to paper

Nonparametric inference on counterfactuals in first-price auctions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

70,991 characters · 18 sections · 63 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Nonparametric inference on counterfactuals in first-price auctions

abstractIn a classical model of the first-price sealed-bid auction with independent private values, we develop nonparametric estimators for several policy-relevant targets, such as the bidder's surplus and auctioneer's revenue under counterfactual reserve prices. Motivated by the linearity of these targets in the quantile function of bidders' values, we propose an estimator of the latter and derive its Bahadur-Kiefer expansion. This makes it possible to construct exact uniform confidence bands and test complex hypotheses about the auction design. Using the data on U.S. Forest Service timber auctions, we test whether setting zero reserve prices in these auctions was revenue maximizing. JEL Classfication: C57, D44 Keywords: first-price auction, uniform inference, quantile density, spacings, counterfactual reserve price, Bahadur-Kiefer expansion, USFS auctions

Introduction

In the empirical studies of first-price auctions, a structural approach to estimation and inference is often used. This approach exploits restrictions derived from economic theory to recover bidders' latent valuations from the observed bids. With these valuations, the researcher can make predictions about the effects of changes in auction rules or the composition of bidders. Various methods, in both parametric and nonparametric frameworks, have been developed; see, e.g., paarsch2006introduction, athey2007nonparametric, and perrigne2019econometrics for an overview.

Since the seminal papers by elyakime1994first, guerre2000optimal and li2000conditionally, it is the probability density function (PDF) of bidders' values that has been considered a default target of nonparametric analysis. This choice is natural since it allows for constructive identification matzkin2013nonparametric when valuations are independent\footnote{With correlated valuations, nonparametric identification is partial, see aradillas2013identification.}. However, researchers are often interested in revenue and other targets, which are nonlinear functionals of the density. Even with the asymptotic theory for the density estimator developed, constructing the confidence intervals and bands for these targets is not simple. This often leads researchers to report confidence intervals based on simulation from the estimated PDF li2003timber or none at all.

Our primary focus is testing complex hypotheses about several natural targets, such as total surplus, bidders' surplus, and auctioneer's revenue. For example, before introducing a reserve price into the auction, a policy-maker might want to test whether any reserve price would yield at least a 5% increase in the auctioneer's revenue or the median bidder's surplus.\footnote{The auctioneer might be concerned about the bidder's welfare as much as about his own revenue, see, e.g., andreyanov2023past. The auctioneer's expected revenue and the median bidder's surplus are typically maximized at a positive reserve price.}

To capture a wide variety of data-generating processes, we allow for binding reserve prices, risk aversion in bidders' preferences, and random (i.e., unknown) number of bidders -- a delicate feature often neglected in the literature. Indeed, many auctions, including procurement auctions, are sealed-bid with bids submitted online, so the participants do not know the number of active bidders.

Our approach relies on the quantile function of valuations --- an alternative candidate for constructive identification --- and on the observation that many targets, including revenue, are continuous linear functionals of this quantile function. Since the value quantile function is the key ingredient of our counterfactual evaluation, we provide its complete first-order asymptotic analysis. Namely, we derive the uniform, asymptotically linear (Bahadur-Kiefer, or BK) expansion for the kernel estimator $\hat v_h$ of the value quantile function $v$, where $h$ is a smoothing bandwidth. This expansion implies that, despite converging to a Gaussian distribution pointwise, the estimator does not admit a functional central limit theorem, which calls for alternative ways of conducting uniform inference. Luckily, the linear term of the studentized estimator is known and pivotal, allowing us to suggest simple simulation-based confidence bands -- a viable alternative to bootstrap -- and establish their validity using the anti-concentration theory of chernozhukov2014gaussian.

With the asymptotic theory of the value quantile function at hand, we move towards analyzing the estimators of our targets at counterfactual reserve prices. We show that these can be divided into two broad classes. One class contains the “smoother” (w.r.t. the value quantile function) counterfactuals that are estimable at the parametric rate $n^{-1/2}$ and converge weakly to a Gaussian process in $\linfzeroone$. The other class contains the “less smooth” counterfactuals that are only estimable at the slower rate $(nh)^{-1/2}$ and do not converge weakly in $\linfzeroone$. For each class, we develop a distinct protocol for constructing confidence intervals and bands and establish their validity. Under the appropriate choice of bandwidth, the convergence rate of the “less smooth” estimator is MSE-optimal. On the other hand, for inference, undersmoothing needs to be used.\footnote{The undersmoothing approach is standard in this literature, see, e.g., Assumption 4.1 in zincenko2024estimation and Equation (5.11) in gimenes2021quantile.}

To illustrate our methodology, we use Philip Haile's data on U.S. Forest Service timber auctions, where a reserve price was set to zero, and assess the optimality of this auction design. Namely, we test whether the seller's expected revenue could have increased if the auction designer chose a nonzero reserve price.

While the basic properties of the classical two-step estimator of the density of latent valuations were stated in guerre2000optimal, an exhaustive theoretical analysis was completed much later in ma2019inference. However, even this analysis did not automatically extend to the functionals of interest. The complete first-order analysis of a single functional -- revenue -- was performed only in zincenko2024estimation and required a significant effort. At the same time, a competing, quantile-regression-based approach was developed in guerre2012uniform, allowing for a wide range of targets. However, their analysis is missing uniformity over quantiles (i.e., counterfactual reserve prices), which is needed to test the kinds of hypotheses that are the focus of our study. Thus, our work is most similar to zincenko2024estimation in spirit, but we also consider targets other than revenue.

More broadly, our work contributes to the expanding literature on quantile methods in first-price auctions, see marmer2012quantile and enache2017quantile for kernel-based estimators, luo2018integrated for isotone regression-based estimators, and guerre2012uniform and gimenes2021quantile for local polynomial estimators. Interestingly, our estimator of valuation quantiles is a weighted sum of the differences of ordered bids, often referred to as bid spacings. The latter has been used for collision detection in ingraham2005test, for set identification of bidders' rents in PAUL2004103 and marra2020sample, and in the prior-free clock auction design in loer2020spacings.

The rest of the paper is organized as follows. In (ref), we set up the theoretical and econometric framework for our analysis. In (ref) and (ref), we develop estimation and inference procedures for the value quantile function and the counterfactual quantities of interest, respectively. In (ref), we provide the Monte Carlo simulations of the finite-sample coverage of our confidence bands. In (ref), we use the timber auction data to test whether counterfactual reserve prices increase the auctioneer's revenue. (ref) contains a discussion of some practical aspects of our methodology. (ref) concludes the paper. Proofs of theoretical results are provided in the Appendix.

Framework

Baseline model

We seek a symmetric Bayes-Nash equilibrium in a first-price auction with $M \ge 2$ ex-ante identical and risk-neutral bidders. Denote the valuation of a bidder by $v$. We impose the following assumption on the value distribution guerre2009nonparametric.

assumption[Distribution of values] The values $v_1,\dots,v_M$ of bidders are drawn independently from a common CDF $G$ with support $[0, \bar v]$ that is twice continuously differentiable and has a strictly positive density $g(v) = G'(v)$ for all $v \in [0,\bar v]$.

The auctioneer announces a reserve price $r^*>0$, which is necessarily binding. Every bidder then submits a sealed bid of $b$. The equilibrium bidding strategy $\beta(v)$ can be characterized by applying the Envelope Theorem to the maximization of $(b-v)G^{M-1}(\beta^{-1}(b|r^*))$:

equation[equation omitted — 172 chars of source]

for all $v \geqslant r^* \geqslant 0$ see, e.g., rileysamuelson or krishna2009auction.

This strategy is strictly increasing and twice continuously differentiable. Moreover, by the first-order conditions\footnote{$\beta'(v|r^*) = \frac{(v-\beta(v|r^*))g(v)}{G(v)/(M-1)}>0$ for all $v > r^*$ and $\beta'(r^*|r^*) = \frac{1}{1+(G/g)'(r^*)/(M-1)} > 0$ by L'H\^{o}pital's rule.}, it has a strictly positive derivative on $[0,\bar v]$. Denoting by $F$ the CDF of the equilibrium bid and $f=F'$, the inverse bidding strategy can be written as

equation[equation omitted — 89 chars of source]

allowing the recovery of the latent values from the observed bids. This suggests a nonparametric estimation approach popularized by guerre2000optimal and li2000conditionally.

Alternatively, we can rewrite the equation (ref) in terms of the quantiles of the participating values. Denote by $Q(u) \bydef F^{-1}(u)$ the bid quantile function and by $q(u)\bydef Q'(u)$ the associated bid quantile density. Let $v(u)\bydef G^{-1}(u)$ be the $u$-th quantile of the participating value distribution $G$. Then equation (ref) can be rewritten as

align[align omitted — 69 chars of source]

where we use the change of variables $b=Q(u)$ along with the identities $F(Q(u))=u$ and $f(Q(u))q(u) = 1$. Since, by definition, $Q(u) = \beta(G^{-1}(u)|r^*)$, and both $(G^{-1})'(v)$ and $\beta'(v|r^*)$ are strictly positive for all $v \in [r^*, \bar v]$, we arrive at the following property.

prop[Distribution of bids] Under (ref), the equilibrium bids are drawn independently from a distribution with a twice continuously differentiable quantile function $Q$ such that $q(u) = Q'(u) > 0$ for all $u \in [0,1]$.

Counterfactuals

We show that a variety of counterfactual metrics can be written in terms of the counterfactual reserve price $r^*$ and the distribution $G$ of bids submitted under the original reserve price $\underline r$. We then show that, in our model, these counterfactual metrics can be rewritten as linear functionals of the value quantile function $v(\cdot)$, the key observation enabling simple inference procedures in (ref).

One such counterfactual is the total expected (ex ante) surplus. In a symmetric equilibrium, it is ex post equal to the highest valuation if it exceeds $r^*$, and zero otherwise, which is a random variable with CDF $G^M(\cdot)$. Hence the total surplus is its expectation

equation[equation omitted — 74 chars of source]

Another counterfactual is bidder's expected surplus. By the revenue equivalence principle krishna2009auction, the interim surplus of a bidder is related to her equilibrium probability of winning, equal to $G^{M-1}(v)$, via the envelope conditions

equation[equation omitted — 72 chars of source]

To derive the bidder's expected (ex ante) surplus $\textit{BS}$, we need to take the expectation of $\pi(v|r^*)$ w.r.t. the distribution of $v$. Integration by parts yields the formula

equation*[equation* omitted — 204 chars of source]

Finally, we consider the seller's expected revenue under the counterfactual reserve price $r^*$, which is equal to the difference between the total expected surplus and $M$ times the bidder's expected surplus,

align*[align* omitted — 303 chars of source]
table[table omitted — 1,705 chars of source]

It can be seen that all the aforementioned counterfactuals are complicated, nonlinear functionals of the primitives. However, using change of variables $z=G(x)$ (i.e., passing to the ranks of valuations from their levels) and denoting $u^*=G(r^*)$ yields expressions that are linear in the quantile function $v(\cdot)$, see (ref).

Interestingly, since the bid $\beta(v|r^*)$ and the bidder's interim surplus $\pi(v|r^*)$ are monotone in $v$, their median values can be easily computed as $$v\left(\frac{1+u^*}{2}\right) - \int_{u^*}^{\frac{1+u^*}{2}} \frac{z^{M-1}}{(\frac{1+u^*}{2})^{M-1}} \ dv(z), \quad \int_{u^*}^{\frac{1+u^*}{2}} z^{M-1} d v(z),$$ which are also linear in $v(u)$. This makes $v(u)$ a key object for the counterfactual analysis.

Random number of bidders and binding reserve prices

Let there be $M \geqslant 2$ potential bidders, and $m \leqslant M$ active bidders in the auction. A potential bidder becomes active if she passes an exogenous and anonymous selection procedure, such as, for example, an (existing) binding reserve price $\underline r$. We can interpret the active bidders as the ones observed by the econometrician and $\underline r$ as the lower end of the support of bids in the data. Denote the resulting probability of observing $m$ active bidders by $p_m$ so that the expected number of active bidders equals $\tilde M := \sum_{m=1}^M m p_m$. Crucially, bidders do not observe how many opponents they face. We are interested in equilibrium behavior when switching to a (counterfactual) higher reserve price $r^* > \underline r$.\footnote{If the counterfactual reserve price exceeds the valuation of an active bidder, she is still considered active.}

A subtle difficulty with this approach is that while the likelihood of observing $m$ active bidders, as perceived by the econometrician, is equal to $p_m$, the same likelihood, as perceived by the participant (i.e., conditional on being active), is equal to $\tilde p_{m} := m p_m / \tilde M$. We will refer to $p_m$ and $\tilde p_m$ as the objective and subjective probabilities. To stay within the rational behavior framework, the bidder's beliefs have to be aligned to the subjective probability.\footnote{Although the equilibrium beliefs $(\tilde p_m)_{m=1}^{M}$ depend on the reserve price $\underline r$ as well as the unspecified selection procedure, its nature is irrelevant as long as the beliefs are identical and do not depend on the identity of the bidder nor his value, see, e.g., krishna2009auction.} This would guarantee coherent formulas for the bidders's surplus and the auctioneer's revenue.

We impose a modified version of (ref), with the only difference being that the relevant support of $G$ is $[\underline r, \bar v]$ rather than $[0, \bar v]$ and the beliefs of the bidders satisfy $\tilde p_0 + \tilde p_1 \neq 1$. The primitives $\underline r, G,(p_m)_{m=0}^{M}$ of the model are common knowledge.

The ex-ante total surplus, using objective probabilities, is $\textit{TS}(r^*) \bydef \int_{r^*}^{\bar v} {x} d \left( A_2(G(x)) \right)$, where $A_2(u) \bydef \sum_{m=1}^{M} p_m u^m$. To derive the interim bidder's surplus, conditional on being active, using subjective probabilities, denote $A_1(u) \bydef \sum_{m=1}^M \tilde p_{m} u^{m-1}$ and write $\pi(v|r^*) \bydef \int_{r^{\ast}}^{v} A_1(G(x)) \, d x$.\footnote{Note that, with correctly specified beliefs, the coefficients of $A_1(u)$ are proportional to that of $A'_2(u)$.} Taking expectation over the distribution of $v$ and integrating by parts, the ex-ante bidder's surplus, conditional on being active, is $\textit{BS}(r^*) \bydef \int_{r^{\ast}}^{\bar v}A_3(G(x)) dx$, where $A_3(u) \bydef (1-u) A_1(u)$.

The seller's expected revenue under the counterfactual reserve price is equal to the difference between the total surplus and the bidder's surplus (conditional on being active) times the expected number of active bidders:

align*[align* omitted — 317 chars of source]

The exact same formula can be obtained via revenue equivalence with the second price auction or via expectation of the strategy over the highest value of the active bidder.

The equilibrium bidding strategy $\beta(v|r^*)$ with reserve price $r^*$ can be derived, using subjective probabilities, via the envelope conditions:

gather*[gather* omitted — 146 chars of source]

for all $v \geqslant r^* \geqslant \underline r$. This strategy is strictly increasing and twice continuously differentiable. Moreover, it has a strictly positive derivative on $[r^*,\bar v].$\footnote{$\beta'(v|r^*) = \frac{(v-\beta(v|r^*))g(v)}{A(G(v))}$ for all $v > r^*$ and $\beta'(r^*|r^*) = \frac{1}{1 + (A(G)/g)'(r^*)} > 0$, where $A(u) \bydef A_1(u)/A_1'(u)$.} Similar to the baseline model, the inverse bidding strategy can be written either directly or in the quantile form:

eqnarray[eqnarray omitted — 165 chars of source]

and (ref) also holds under the modified version of (ref).

table[table omitted — 1,727 chars of source]

Finally, the counterfactual participation pattern using objective probabilities is $$p_m(r^*) := \sum_{i = m}^{M} \binom{i}{m} (1-G(r^{*}))^m G(r^{*})^{i-m}p_{i}.$$

Using the change of variables $z=G(x)$ and denoting $u^*=G(r^*)$ yields expressions that are linear in the quantile function $v(\cdot)$, see (ref).

Risk aversion and opportunity cost

Suppose the number of bidders entering an auction is random. Suppose also that the bidders have constant relative risk aversion (CRRA) utility with risk-aversion parameter $\eta$, and let the auctioneer have the opportunity cost $c$. The equilibrium bidding strategy $\beta(\cdot|r^*)$ of the active bidder with reserve price $r^*$ can be characterized by applying the Envelope Theorem to the maximization of $(v - b)A^{\frac{1}{1-\eta}}_1(G(\beta^{-1}(b)))$:

gather*[gather* omitted — 238 chars of source]

for all $v \geqslant r^* > \underline r$. Clearly, the strategy is still linear in $v(\cdot)$.

Due to risk aversion, we cannot derive revenue as in the previous sections, nor can we use the Revenue Equivalence between the first-price and the second-price auction, see krishna2009auction. Instead, similar to zincenko2024estimation, we can derive revenue as the expectation of the highest bid net the opportunity cost over the distribution of the highest active bidder's value, $\textit{RE}(r^*) = \int^{\overline v}_{r^*} (\beta(x) - c) d A_2(G(x))$. Then,

equation[equation omitted — 166 chars of source]

where $A_4(x) \bydef \int_0^{x} A^{\frac{-1}{1-\eta}}_1(x)d A_2(x) = \tilde M \int_0^{x} A^{\frac{-\eta}{1-\eta}}_1(x) dx$ because $A'_2(x) = \tilde M A_1(x)$. Finally, using integration by parts similar to the risk-neutral case,

gather[gather omitted — 311 chars of source]

Therefore, with risk aversion, expected revenue is still linear in $v(\cdot)$.

Finally, although bidder's interim surplus $\pi(v|r^*) = (\int_{r^*}^v A^{\frac{1}{1-\eta}}_1(G(x)) dx)^{1-\eta}$ is not linear in $v(\cdot)$, this problem can be circumvented since tests about the maximum of $\pi(v|r^*)$ over $r^*$ can be constructed from tests about the maximum of $(\pi(v|r^*))^{\frac{1}{1-\eta}}$.

Data generating process

The observed data is a random sample of bids $\{b_{il},\,\, i=1,\dots,m_l,\,\, l=1,\dots,L \}$, where $b_{il}$ denotes the bid submitted by the $i$-th participant in the $l$-th auction. All the auctions are ex-ante symmetric and independent, and $m_l$ is the number of participants in the $l$-th auction. For brevity, we denote the (random) sample size by $n=n(L) = \sum_{l=1}^L m_l$ and define

align[align omitted — 89 chars of source]

Note that since the bidders do not know the realizations of the number of active bidders, the samples $\{b_1,\dots,b_n\}$ and $\{m_1,\dots,m_L\}$ are independent. Besides, as $L \to \infty$, the sample size $n(L) \to \infty$ with probability one. Therefore, without loss of generality, we condition our subsequent exposition on a realization $\{m_l\}_{l=1}^\infty$ such that $n(L) \to \infty$. This has the following important implication: although the auxiliary functions $A_1,A_2,A_3,A$ and the constant $a$ need to be estimated from the data, we can assume that they are known since their estimators only depend on the conditioning variables $m_1,\dots,m_L$, see equations (ref)-(ref) below.

Estimation and inference for value quantiles

As explained in (ref), the value quantile function $v(\cdot)$ is the key object needed for the counterfactual analysis. In this section we develop the asymptotic theory for its natural (plug-in) estimator. To define the estimator, we need to introduce two auxiliary objects.

The first object is the kernel estimator of the bid quantile density $q(u)$, defined by

align[align omitted — 87 chars of source]

Here $K$ is a compactly supported kernel, $K_h(z) \bydef h^{-1}K\left(h^{-1}z\right)$, $h>0$ is a bandwidth, and $\hat Q(u)$ is the empirical bid quantile function,

align[align omitted — 166 chars of source]

where $b_{(1)}\le\dots\le b_{(n)}$ are the order statistics of the observed bids $b_1,\dots,b_n.$ We note that $\hat q_h$ takes the form of a weighted sum of bid spacings $b_{(i+1)}-b_{(i)}$,

equation[equation omitted — 101 chars of source]

This estimator was previously studied by siddiqui1960distribution and bloch1968simple for the case of rectangular kernel, and by falk1986estimation, welsh1988asymptotically, csorgHo1991estimating, and jones1992estimating for general kernels.

The second auxiliary object is the plug-in estimators of $A_1,A_2,A_3,A$, and $\tilde M$ defined by

align[align omitted — 362 chars of source]

where $\check p_m$ is the empirical frequency of auctions with $m$ bidders,

align[align omitted — 103 chars of source]

We use the “check” (as opposed to “hat”) notation here to highlight that $\check A_1, \check A_2, \check A_3, \check A, \check M$ are treated as known since, as explained in (ref), our analysis is conditional on $m_1,\dots,m_L$.

Given $\check A$ and $\hat q_h$, we define our estimator of the value quantile $v(u)$ by

align[align omitted — 110 chars of source]

We note that $\hat v_h$ consists of two parts: (i) the empirical quantile function $\hat Q$ that is uniformly consistent and converges to a Gaussian process in $\linfzeroone$ at the parametric rate $n^{-1/2}$, and (ii) the kernel quantile density $\hat q_h$ that is uniformly consistent only away from the boundary $\{0,1\}$ and does not converge to a (tight) limit in $\ell^\infty[\varepsilon,1-\varepsilon]$ even if $\varepsilon>0$, but converges pointwise to a Gaussian limit at the nonparametric rate $(nh)^{-1/2}$, see the proof of (ref). Therefore, the first-order asymptotic properties of $\hat v_h$ are determined by the kernel quantile density $\hat q_h$.

We impose the following assumptions.

assumption[Kernel function] \begin{enumerate} • $K: \R \to \R$ is a nonnegative function such that \begin{align} \int_\R K(z)\, dz=1 and R_K\bydef \int_{\R} K(z)^2 \, dz < \infty. \end{align} • $K$ is a Lipschitz function supported on the interval $[-1,1]$. \end{enumerate}
assumption[Bandwidth, estimation] The bandwidth $h=h_n$ is such that $h\to 0$ and there exist $c>0$ and $\alpha>0$ such that $h_n \ge c n^{-1/2+\alpha}$ for all $n$.
assumption[Bandwidth, inference] The bandwidth $h=h_n$ is such that there exist $C>0$ and $\beta>0$ such that $h_n \le C n^{-1/3-\beta}$ for all $n$.

(ref).(ref) states that $K$ is a valid, square-integrable PDF. (ref).(ref) is standard in the literature on strong approximations of local empirical processes rio1994local. In particular, it implies that $K$ is a function of bounded variation, which is crucial in our derivation of the BK expansion. (ref) is sufficient if the goal is estimation, because it leads to consistency of $\hat v_h$ and allows for MSE-optimal estimation of $\hat v$, where the optimal bandwidth is $h=O(n^{-1/5})$, see Theorem 2.2 of csorgHo1991estimating. On the other hand, if the goal is inference, then our approach is to impose (ref) (undersmoothing) to eliminate the bias in $\hat v_h$ and related nonsmooth counterfactuals.

The Bahadur--Kiefer expansion

In this section, we derive the Bahadur--Kiefer (i.e. almost sure, uniform, asymptotically linear) representation of the form

align[align omitted — 144 chars of source]

where

align[align omitted — 159 chars of source]

and the remainder $R_n(u)$ converges to zero a.s. uniformly in $u\in[h,1-h]$ with an explicit rate.

The key feature of this representation is that the main term is fully known and pivotal: its distribution does not depend on the data generating process since $U_i \bydef F(b_i) \sim \text{iid Uniform}[0,1]$. Heuristically, this suggests that the distribution of the left-hand side under any DGP is a valid approximation for its true distribution. Indeed, in (ref), we show the validity of such approximation by combining pivotality with the anti-concentration theory of chernozhukov2014gaussian. This leads to a simple algorithm for the confidence bands on the quantile function $v(\cdot).$

To derive this representation, we rely on the classical BK expansion of the quantile function bahadur1966note,kiefer1967bahadur,

align[align omitted — 201 chars of source]

Here $\ell(n)=(\log n)^{1/2}(\log \log n)^{1/4}$ is a logarithmic offset factor that arises due to the uniform nature of the approximation and may often be disregarded in practice. Note that the BK expansion represents a nonlinear estimator $\hat Q(u)$ as a sum of the linear estimator --- the empirical distribution function $\hat F(Q(u))$ --- and the remainder $r_n(u)$ that converges to zero a.s. uniformly at a nonparametric (slow) rate $n^{-3/4}\ell(n)$.

thm[Bahadur-Kiefer expansion for value quantiles] Under the Assumptions (ref) and (ref), the estimator $\hat v_h(u)$ has the representation \begin{align} Z_n(u) = Z_n^*(u) + R_n(u), \quad u\in[h,1-h], \end{align} where \begin{align} Z_n(u) &\bydef \frac{\sqrt{nh}\left(\hat v_h(u)-v(u)\right)}{\hat q_h(u)}, \quad Z_n^*(u) \bydef -\check A(u) \eG_{n,h}(u), \\ R_n(u) &= O_{a.s.}\left(n^{1/2}h^{3/2} + h^{1/2} + h^{-1/2}n^{-1/4}\ell(n)\right) uniformly in u\in[h,1-h]. \end{align}
rem[BK expansion for quantile density] The proof of the preceding theorem also implies the BK expansion for the normalized quantile density $\sqrt{nh}\left(\hat q_h(u)-q_h(u)\right)$, which may be of independent interest. In this case, the right-hand side does not have the factor $\check A(u)$, while the term $h^{1/2}$ in the remainder rate can be replaced by the faster term $h\log h$.

We note that two types of biases arise in the estimation of $v(\cdot)$. The first type of bias is the smoothing bias $\E \hat v_h(u)-v(u)$ which manifests in the term $n^{1/2}h^{3/2}$ in the remainder rate. This bias can be eliminated by undersmoothing $h=o(n^{-1/3})$, i.e. choosing a (suboptimally) small bandwidth such that $\sqrt{nh}(\E \hat v_h(u)-v(u)) \to 0$, which is our (ref).

The other type of bias is the boundary bias, stemming from the estimator $\hat v_h(u)$ being inconsistent when $u$ is close to the boundary $\{0,1\}$ of its domain $[0,1]$. Because our interest is in valid hypothesis testing, and not the confidence bands per se, we can eliminate this bias by introducing the trimming $u \in [h,1-h]$ while maintaining the validity of inference procedures based on the representation (ref).

Inference on value quantiles

(ref) allows us to construct pointwise confidence intervals and uniform confidence bands for the value quantile function. In particular, the following corollary provides the asymptotic distribution of the estimator of a fixed valuation quantile.

corUnder the Assumptions (ref), (ref), (ref), and (ref), we have, for every $u\in(0,1)$, \begin{align} &\sqrt{nh}\left( \hat v_h(u)-v(u) \right) \weakto N(0, V(u)),\\ &V(u) \bydef A^2(u)q^2(u)R_K. \end{align}

Special cases of this result for the quantile density estimator $\hat q_h$ were derived by siddiqui1960distribution and bloch1968simple. It implies that a confidence interval of nominal confidence level $1-\alpha$ for $v_h(u)$ can be constructed as

align[align omitted — 220 chars of source]

where $z_{1-\frac{\alpha}{2}}$ is the standard normal quantile of level $1-\frac{\alpha}{2}$.

We now turn to the problem of uniform inference on $v(\cdot)$.

Note that if the process $Z_n$ converged weakly in $\linfh$ to a known (or estimable) process $Z$, this would have enabled the construction of asymptotically valid confidence bands by using the quantiles of $\sup_u | Z(u) |$ as critical values.\footnote{For one-sided confidence bands, one would use the quantiles of $\sup_u Z_n(u)$ instead.} Unfortunately, although $Z_n(u)$ is asymptotically Gaussian at each point $u \in (0,1)$, it does not converge in $\linfh$. This follows from the fact that the main term $Z_n^*$ in the BK expansion is the scaled kernel density process, which is known to lack functional convergence rio1994local.\footnote{For an example of a sequence of stochastic processes on $[0,1]$ that weakly converges pointwise, but not in $\linfzeroone$, consider $X_n(u)=h_n^{-1/2}\left(B\bigPar{u+h_n}-B\bigPar{u}\right)$, where $B$ is the Brownian motion, $h_n \to 0$ and $u\in(0,1)$. Clearly, $X_n(u) \weakto N(0,1)$ for all $u$, but, by L\'{e}vy's modulus of continuity theorem, $\sup_u \left|X_n(u)\right| \to \infty$ a.s., and so there is no convergence in $\linfzeroone$.} In such a case, there are two common ways to circumvent the problem and derive valid confidence intervals.

One approach is to derive the asymptotic distribution of a normalized version of $\sup_u Z_n(u)$ using extreme value theory, and then rely on the knowledge of the normalizing constants to construct the confidence band. In the case of kernel and histogram density estimation, this approach was pioneered by smirnov1950construction and bickel1973some.\footnote{For a nonasymptotic version of Smirnov-Bickel-Rosenblatt extreme value theorem, see rio1994local.} However, convergence to the asymptotic distribution turns out to be very slow, leading to the coverage error of the resulting confidence band to be $O(1/\log n)$), as shown by hall1991convergence.

The other approach is to rely on finite-sample approximations for (the distribution of) the supremum

align[align omitted — 83 chars of source]

If such an approximation admits simulation, it can be used for the construction of confidence bands. This is the approach we take in this paper.

We consider two types of approximations, both of which are pivotal, and hence allow for simulation. One is simply the supremum of the linear term $Z_n^*$, viz.

align[align omitted — 95 chars of source]

The other is the supremum of $Z_n$ under an alternative, uniform[0,1] distribution of bids, viz.

align[align omitted — 104 chars of source]

where $Z_n^{U[0,1]}(u)$ is the process $Z_n(u)$ calculated using the pseudo-sample

align[align omitted — 97 chars of source]

This approximation is nonstandard and makes use of the asymptotic pivotality of $W_n$. In principle, any distribution of the pseudo-bids rationalized by a value distribution satisfying Assumption (ref) can be chosen; however, the uniform distribution is convenient since, in this case, we have, for all $u\in [0,1]$,

align[align omitted — 80 chars of source]

and hence

align[align omitted — 115 chars of source]

We emphasize that it is not immediate that the distributions of $W_n^*$ and $W_n^{U[0,1]}$ approximate the distribution of $W_n$ in a way that guarantees the validity of the associated confidence bands

align[align omitted — 200 chars of source]

where $c_{n,1-\alpha/2}$ is the $(1-\alpha/2)$-quantile of either $W_n^*$ or $W_n^{U[0,1]}$. Indeed, note that (ref) implies the inequality

align[align omitted — 157 chars of source]

and hence the coupling

align[align omitted — 62 chars of source]

where $r_n$ tends to zero a.s. at a known rate. If one could show that this implies Kolmogorov convergence

align[align omitted — 112 chars of source]

then the confidence bands based on $W_n^*$ would be valid. However, (ref) need not follow from (ref) even if the a.s. convergence rate of $r_n$ is very fast, unless further conditions are imposed on $W_n^*$.

As an illustration of this phenomenon, consider an abstract example $W_n=n^{-1} U$, $W_n^*=n^{-1} (U-1)$, where $U\sim \text{Uniform}[0,1]$. Then $r_n=W_n - W_n^*= n^{-1}$, but

align*[align* omitted — 71 chars of source]

and so (ref) does not hold. On the other hand, if $W_n^*$ had an absolutely continuous asymptotic distribution $\mathcal{D}$, then the CDF of $W_n$ would converge to the CDF of $\mathcal{D}$ pointwise, and hence the quantiles of $\mathcal{D}$ would serve as valid critical values.

Therefore, intuitively, a certain degree of anti-concentration of $W_n^*$ is needed to guarantee that the coupling (ref) implies Kolmogorov convergence (ref) and hence validity of simulated critical values. The anti-concentration literature mainly focuses on Gaussian processes, while the process $Z_n^*$ is non-Gaussian. Fortunately, $Z_n^*$ is the normalized kernel density estimator for uniform data, which is a well-studied process. In particular, we rely on the seminal work chernozhukov2014gaussian to establish a coupling of $W_n$ with the supremum of a Gaussian process and show that the latter exhibits sufficient anti-concentration. We then argue that an identical argument works for $W_n^*$. Finally, the pivotality of $W_n^*$ and the coupling (ref) imply the Kolmogorov convergence for $W_n^{U[0,1]}$. Formally, we have the following result.

thmUnder the Assumptions (ref), (ref), (ref), and (ref), \begin{align} &\sup_{x \in \R} \left| \Pb(W_n\le x) - \Pb(W_n^*\le x) \right| \to 0,\\ &\sup_{x \in \R} \left| \Pb(W_n\le x) - \Pb(W_n^{U[0,1]}\le x) \right| \to 0, \end{align} and hence the confidence bands (ref) are asymptotically valid and exact.
remTo construct the one-sided confidence bands, we note that the same result holds with $W_n$, $W_n^*$, and $W_n^{U[0,1]}$ replaced by $\sup_{u\in[h,1-h]} Z_n(u)$, $\sup_{u\in[h,1-h]} Z_n^*(u)$, and \newline $\sup_{u\in[h,1-h]} Z_n^{U[0,1]}(u)$, respectively.

Estimation and inference for counterfactuals

In this section, we develop the asymptotic theory for the counterfactuals in (ref) which heavily relies on the analysis of the estimator of value quantiles in the previous section.

Clearly, every such counterfactual has the general form

align[align omitted — 102 chars of source]

where $\varphi$ and $\psi$ are continuously differentiable functions on $[0,1]$ that only depend on the auxiliary objects $A_1,A_2,A_3,A$ (or their derivatives). As an example, for the total expected surplus $\varphi(x) \equiv 0$ and $\psi(x) = A_2'(x)$, while for the expected revenue $\varphi(u^*)=\tilde M A_3(u^*)$ and $\psi(x) = A_2'(x)+\tilde M A_3'(x)$.

The representation (ref) implies $T(u^*)$ is a (weighted) sum of two continuous linear functionals of $v$ of different smoothness: (i) evaluation at a point $v(u^*)$ and (ii) integration

align[align omitted — 92 chars of source]

The natural estimators of the two components have fundamentally different asymptotic properties. Namely, in (ref) we showed that the less smooth functional $v(u^*)$ is only estimable at the nonparametric rate $(nh)^{-1/2}$ and does not converge in $\ell^\infty[\varepsilon,1-\varepsilon]$ even for $\varepsilon>0$. On the other hand, we will show in (ref) that the smoother functional $S_{\psi}(u^*)$ is estimable at the parametric rate $n^{-1/2}$ and converges to a Gaussian process in $\linfzeroone$. We will combine the two results in (ref) to show that, whenever $\varphi \neq 0$, inference on $T$ can be performed similarly to that on $v$.

Smooth ($S$-type) counterfactuals

First, let us consider estimation and inference for functionals (ref), where $\psi:\,[0,1]\to \R$ is a known, continuously differentiable function.

To motivate our estimator, use integration by parts to rewrite

align[align omitted — 498 chars of source]

where we denote

align[align omitted — 100 chars of source]

Note that the latter formula expresses $S_\psi(u^*)$ as a continuous linear functional of the quantile function $Q$, which is estimable at a parametric rate. This leads to a natural estimator of $S_\psi$ that does not contain tuning parameters. Namely, for any $u^*\in [0,1]$, we define the estimator

align[align omitted — 375 chars of source]

The following theorem establishes standard Gaussian process asymptotics for $\hat S_\psi$.

thmUnder the Assumptions (ref), (ref), (ref), and (ref), \begin{align} \sqrt{n}(\hat S_\psi(\cdot) - S_\psi(\cdot)) \weakto \G_{\psi,q}(\cdot) in \linfzeroone, \end{align} where $\G_{\psi,q}(\cdot)$ is a tight, centered Gaussian process on $[0,1]$ with the covariance function \begin{align} \E \G_{\psi,q}(u^*)\G_{\psi,q}(v^*) = Cov\left( f_{u^*}(U),f_{v^*}(U)\right),\quad U\sim uniform[0,1], \quad u^*,v^*\in[0,1], \end{align} and \begin{align} f_{u^*}(U) \bydef - \int_{u^*}^1 \chi_\psi(u)q(u)1(U\le u)\,du + \check A(u^*)\psi(u^*) q(u^*) 1(U\le u^*). \end{align}
remThe integral in the expression for $\hat S_\psi(u^*)$ can be replaced by $\chi_\psi(\zeta(i))/n$ for any $\zeta(i) \in \left[ \frac{i}{n}, \frac{i+1}{n} \right]$. This would have no impact on the statement of (ref).

Nonsmooth ($T$-type) counterfactuals

Given the asymptotic results for the two components of the generic counterfactual (ref), we may now turn to estimation and inference on the latter.

To this end, define the estimator

align[align omitted — 116 chars of source]

where $\hat v_h(u^*)$ is defined in (ref) and $\hat S_{\check \psi}(u^*)$ is the estimator (ref) with $\psi=\check \psi$. Since $\hat S_{\check \psi}(u^*)$ converges fast, while $\hat v_h(u^*)$ converges slowly, the asymptotics of $\hat T_h(u^*)$ is dominated by the latter, as illustrated by the following theorem.

thmUnder the Assumptions (ref) and (ref), we have the representation \begin{align} Z_n^T(u^*) = Z_n^{T*}(u^*) + R_n^T(u^*), \quad u^*\in [h,1-h], \end{align} where \begin{align} Z_n^T(u^*) &\bydef \frac{\sqrt{nh}\left( \hat T_h(u^*)-T(u^*) \right)}{ \hat q_h(u^*)}, \quad Z_n^{T*}(u^*) \bydef -\check\varphi(u^*) \check A(u^*) \eG_{n,h}(u^*),\\ R_n^T(u^*) &\bydef O_{a.s.}\left(n^{1/2}h^{3/2} + h^{1/2} + h^{-1/2} n^{-1/4} \ell(n)\right) uniformly in u^*\in[h,1-h]. \end{align}

(ref) immediately yields the following result on the asymptotic distribution of $\hat T_h(u^*)$ at a fixed point $u^* \in (0,1)$.

corUnder the Assumptions (ref), (ref), (ref), and (ref), we have, for every $u^*\in(0,1),$ \begin{align} &\sqrt{nh}(\hat T_h(u^*) - T(u^*)) \weakto N(0,V(u^*)),\\ &V(u^*) = \left(A(u^*) q(u^*) \varphi(u^*) \right)^2 R_K. \end{align}

We now consider uniform inference on $T(\cdot)$. Since the estimator $\hat T_h(\cdot)$ does not converge in $\linfh$, but the approximating process $Z_n^{T*}$ is known and pivotal, the methodology will be similar to the case of the valuation quantile function, see (ref). In particular, we show that valid confidence bands for $T(\cdot)$ can be based on simulation from either (i) the approximating process $Z_n^{T*}$, or (ii) the process $Z_n^T$ under an alternative distribution of bids. To this end, define

align[align omitted — 234 chars of source]

where $Z_n^{T,U[0,1]}(u^*)$ is the process $Z_n^T(u^*)$ calculated using the pseudo-sample

align[align omitted — 69 chars of source]

cf. equations (ref)-(ref). Define the confidence bands by

align[align omitted — 210 chars of source]

where $c_{n,1-\alpha/2}$ is the $(1-\alpha/2)$-quantile of either $W_n^{T*}$ or $W_n^{T,U[0,1]}$.

thmUnder the Assumptions (ref), (ref), (ref), and (ref), we have \begin{align} &\sup_{t\in \R} \left|\Pb(W_n^T \le t) - \Pb(W_n^{T*} \le t) \right| \to 0,\\ &\sup_{t\in \R} \left|\Pb(W_n^T \le t) - \Pb(W_n^{T,U[0,1]} \le t) \right| \to 0, \end{align} and hence the confidence bands (ref) are asymptotically valid and exact.
remFor the purpose of constructing the one-sided confidence bands, we note that the same result holds with $W_n^T$, $W_n^{T*}$, and $W_n^{T,U[0,1]}$ replaced by $\sup_{u\in[h,1-h]} Z_n^T(u)$, $\sup_{u\in[h,1-h]} Z_n^{T*}(u)$, and $\sup_{u\in[h,1-h]} Z_n^{T,U[0,1]}(u)$, respectively.
rem[Shape of confidence bands] Note that, for any function $\iota(\cdot)$ bounded away from zero on $[0,1]$, the representation (ref) is equivalent to \begin{align} Z_n^T(u^*)/\iota(u^*) = Z_n^{T*}(u^*)/\iota(u^*) + \tilde R_n^T(u^*), \quad u^*\in [h,1-h], \end{align} where $\tilde R_n^T(u^*)$ has the same uniform convergence rate as $R_n^T(u^*)$. Similarly to (ref), the two-sided confidence bands based on such representation are \begin{align} \left[ \hat T_h(u^*) - \frac{\iota(u^*) \hat q_h(u^*) c_{n,1-\alpha/2}}{\sqrt{nh}}, \, \hat T_h(u^*) + \frac{\iota(u^*) \hat q_h(u^*) c_{n,1-\alpha/2}}{\sqrt{nh}} \right], \quad u\in[h,1-h], \end{align} where $c_{n,1-\alpha/2}$ is the $(1-\alpha/2)$-quantile of either of the random variables \begin{align} W_n^{T,\iota} &\bydef \sup_{u^*\in[h,1-h]} \left| Z_n^{T*}(u^*)/\iota(u^*) \right|, \\ W_n^{T,U[0,1],\iota} &\bydef \sup_{u^*\in[h,1-h]} \left| Z_n^{T,U[0,1]}(u^*)/\iota(u^*) \right|, \end{align} These bands stay asymptotically exact for any $\iota$, but have the shape $\iota(u^*) \hat q_h(u^*)$, which may affect their finite-sample performance and asymptotic power.\footnote{montiel2019simultaneous discuss a similar issue with simultaneous confidence bands for a vector (rather than functional) parameter.}

Monte Carlo experiments

While our theoretical results establish the asymptotic validity of the confidence bands, they do not rule out a substantial finite-sample size distortion. In this section, we evaluate the extent of this distortion in a set of Monte Carlo experiments.

For simplicity, we simulate an auction with exactly two bidders ($M = 2$ and $p_2 = \check p_2 = 1$) and a non-binding (original) reserve price $\underline r = 0$. We consider three choices for the distributions of the observed bids: uniform, beta and power-law; all of which are supported on the interval $[0,1]$. The simplest choice is the uniform distribution, since it has a strictly positive density with $q(u) = 1$. The beta$(\alpha, \beta)$ distribution features a bell-shaped density for $\alpha, \beta > 1$, with varying skewness. In contrast, the density $f(x) = \alpha x^{\alpha-1}$ of the power-law distribution increases on the support $[0,1]$ for $\alpha > 1$. We censor these distributions at top 5% and bottom 5% quantile levels, so that the quantile density is strictly positive and satisfies the statement of (ref).\footnote{The censoring of the distribution in the Monte Carlo simulations is achieved by replacing their true quantile function $Q(u)$ with $(Q(0.05 + 0.9 u)-Q(0.05))/(Q(0.95)-Q(0.05))$. We emphasize that every simulated bid distribution can be rationalized by some value distribution satisfying (ref).}

table[table omitted — 1,516 chars of source]

The estimation targets are (i) the bid quantile function $q$, (ii) the value quantile function $v$; and the following quantities as functions of the counterfactual reserve price: (iii) the potential bidder's expected surplus, (iv) the expected revenue, and (v) the total expected surplus, see (ref).

For the non-counterfactual targets (i), (ii) and for the $T$-type counterfactuals (iii), (iv), we calculate the confidence bands by simulation from the left-hand side of the respective BK expansions under the uniform[0,1] bid distribution, see (ref). For the $S$-type functional (v), the confidence bands are constructed by simulation from the estimated process

align[align omitted — 139 chars of source]

where $U_i \sim \text{iid Uniform}[0,1]$ and $\hat f_{u^*}$ is equal to $f_{u^*}$ with the true values $q$ replaced by their estimates $\hat q_h$, see (ref). We use the undersmoothing bandwidth $h = 1.06 s \cdot n^{-0.34}$, where $s$ is the standard deviation of bid spacings, and set both the number of DGP simulations and the number of simulations for the critical values to 500.

The results are shown in (ref). The simulated coverage can be seen to be close to the nominal level of $0.95$ for larger sample sizes, which validates our theoretical results in (ref).

Empirical application

In this section, we apply our methodology to test the hypothesis about the optimality of the auction design employed in timber sales held by the U.S. Forest Service in between 1974 and 1989, see, e.g., haile2001auctions. These auctions did not feature a reserve price (i.e., $\underline r=0$), which raises the question of whether the collected revenue could have been higher had the reserve price been set at a positive level. We use the data kindly provided by Phil Haile on his website.\footnote{\url{http://www.econ.yale.edu/ pah29/}}

We select a subsample of auctions that are sealed-bid and have at least 2 bidders, see (ref). As is common in the literature, we residualize the log-bids using available auction-level characteristics: year and location dummies, the logarithms of the tract advertised value and the Herfindahl index.\footnote{The exponentiated log-bid residuals (to which we will refer simply as bid residuals) are interpreted as estimates of the idiosyncratic component of bids, while the exponentiated fitted values are interpreted as estimates of the common component of bids, see haile2003nonparametric.} The latter is a measure of the homogeneity of the tract with respect to the timber species. This procedure is consistent with a multiplicative model of observed auction heterogeneity.\footnote{The asymptotic distributions of the test statistics are not affected by the error in the residualization procedure, as long as the estimates of the common component of bids converge at a faster (in our case parametric) rate, see haile2003nonparametric and athey2007nonparametric.} The distribution of bid residuals is truncated at the 5th percentile on each end, leaving 60758 observations, see (ref).

figure[figure omitted — 207 chars of source]

To assess the change $\Delta(u^{\ast})$ in the expected revenue, we use the quantile version of the revenue formula in (ref), which yields

align*[align* omitted — 188 chars of source]

where $u^{\ast} > 0$ is the counterfactual exclusion level (rank of the counterfactual reserve price in the distribution of valuations) and we used the fact that in our application $v(0) = \underline r$. Since $\Delta(u^*)$ is similar to the $T$-type functional (ref), its estimator is

align[align omitted — 198 chars of source]

where $\varphi(x) \bydef \tilde M \check A_3(x)$, $\psi(x) \bydef A_2'(x) + \tilde M A_3'(x)$ and $\chi_\psi$ is defined in (ref).

We use the undersmoothing bandwidth $h = 1.06 s \cdot n^{-0.34}$, where $s$ is the standard deviation of spacings of bid residuals, and evaluate $\hat \Delta_h(u^{\ast})$ on the evenly spaced grid $\{i/n\}_{i=0}^n$. This bandwidth is slightly smaller than the Silverman rule of thumb bandwidth $h = 1.06 s \cdot n^{-1/5}$.

To construct the confidence bands, we use the representation in (ref) and (ref). First, 1000 realizations of the bid quantile density $\hat q_{h}^{U}(\cdot)$ are simulated, independently from the data, based on pseudo-bids from the uniform[0,1] distribution. Second, for a nominal confidence level $(1-\alpha)$, the critical value $c_{n,1-\alpha}$ is computed as the $(1-\alpha)$-quantile of $\sup_{u \in [h,1-h]}(\hat q_{h}^{U}(u) - 1)$. Finally the one-sided confidence band is computed as

align[align omitted — 146 chars of source]
figure[figure omitted — 207 chars of source]

We test the hypothesis of the nonexistence of a counterfactual (positive) reserve price that would increase the seller's expected revenue. Formally, we test the hypotheses $H_0$ against $H_1$, where

equation[equation omitted — 135 chars of source]

The corresponding test statistic is the maximal (over the grid) value of the lower end point function of the one-sided confidence band, and $H_0$ is rejected whenever this maximum is positive. We denote by $\tilde u$ the point at which the maximum of the estimated function $\textit{RE}(u^*)$ is attained, i.e. the optimal exclusion level.

We test the hypothesis using subsamples of auctions with different numbers of bidders, see (ref). We use subsamples with 2 and 3 bidders, and also with 2-5 (small auctions), 5-9 (large auctions), and 2-9 (all auctions) bidders.\footnote{Typically, a researcher would pick, for the sake of simplicity, a subsample of auctions with the same number of bidders. However, our methodology allows for a random number of bidders, so we can pool auctions with different numbers of bidders together.} Under all specifications, $H_0$ is rejected at a 95% confidence level, see (ref), meaning that the revenue gains at the optimal reserve price are statistically significant, albeit relatively small. We also point out that the sample size can be significantly increased by pooling together the auctions with a different number of bidders. For a full sample, the results of the estimation are illustrated in (ref).

table[table omitted — 593 chars of source]

Practical considerations

In this section, we briefly discuss some important technical aspects of our methodology.

\paragraph{Choice of the grid.} While it is theoretically possible to evaluate our estimators at any quantile level, choosing the evenly-spaced grid $\{i/n\}_{i=2}^n$ has a massive impact on the computational complexity of the estimation procedure and its performance.

Note that, with this grid, the estimate of $\hat q_h(u)$ becomes a discrete convolution of the vector of spacings $\{b_{(i)} - b_{(i-1)}\}_{i=2}^n$ with a discrete filter corresponding to $K_h$. The discrete convolution is a remarkably fast and reliable procedure. Moreover, the counterfactual estimators can be well approximated with the weighted cumulative sums of the vectors of spacings. Consequently, all our estimators can be thought of as combinations of elementary vector operations with sorting and convolution.

\paragraph{Shape restrictions.} Since $v(\cdot)$ is a quantile function, one may want to impose monotonicity on $\hat v_h(\cdot)$ and the associated confidence bands. As suggested in chernozhukov2009improving,rearrangement, an effective way of doing so is smooth rearrangement of the estimate and the confidence bands, whose discrete counterpart is merely a sorting algorithm. We leave the analysis of such shape-restricted estimators for future work. We note that there is emerging literature exploiting shape restrictions for auction counterfactuals, e.g. henderson2012empirical,luo2018integrated,pinkse2019estimation,ma2021monotonicity.

\paragraph{Alternative estimator of $q(u)$.} A close competitor to our estimator $\hat q_h(\cdot)$ of the bid quantile density is the reciprocal of the kernel estimator $\hat f_l(\cdot)$ of the bid density, as in the first step of the procedure of guerre2000optimal,

equation*[equation* omitted — 186 chars of source]

An insightful comparison of $\hat q_h$ and $\tilde q_l$ was carried out by jones1992estimating who showed that the variance components of the mean squared errors (MSE) of the estimators are equal if so-called scale match-up bandwidths are used. Namely, for a fixed point $u \in (0,1)$, the MSE of $\hat q(u)$ is only less than or equal to that of $\tilde q_l(u)$ when $q(u) q''(u) \le 1.5 q'(u)^2$ or, equivalently, $f(b) f''(b) \ge 1.5 f'(b)^2.$ Therefore, the reciprocal kernel density $\tilde q_l$ performs better close to the center of the distribution, while the kernel quantile density $\hat q_h$ is preferable at the tails.

Finally, we note that the algorithms to construct the estimates of $\tilde q_l$ and $\hat q_h$ on the grid, seem to have different computational complexity, which can be heuristically shown to be $O(n^2)$ and $O(n \log n)$, respectively. This is due to the fact that the convolution algorithm, which the estimator $\hat q_l(u)$ relies on, has the complexity of roughly $O(n \log n)$ due to its usage of the fast Fourier transform.

Conclusion

In this paper, we develop a novel approach to estimation and inference on counterfactual functionals of interest (such as the expected revenue as a function of the rank of the counterfactual reserve price) in the standard nonparametric model of the first-price sealed-bid auction. We show that these counterfactuals can be written as continuous linear functionals of the quantile function of bidders' valuations, which can be recovered from the observed bids using a well-known explicit formula. We suggest natural estimators of the counterfactuals and show that their asymptotic behavior depends on their structure. In particular, we classify the counterfactuals into two types, one allowing for parametric (fast) convergence rates and standard inference, and the other exhibiting nonparametric (slow) convergence rates and lack of uniform convergence. For each of the types of counterfactuals, we develop simple, simulation-based algorithms for constructing pointwise confidence intervals and uniform confidence bands.

In terms of efficiency, our estimators of the “less smooth” targets, such as the bidder's surplus and auctioneer's revenue, perform on par with the best that the literature can offer, achieving the optimal rate of estimation while using undersmoothing for inference. However, there may be subtle differences in implementation. Indeed, our approach constructs the estimator on the grid of quantiles simultaneously (rather than separately) for each counterfactual reserve price. Combined with the Fast Fourier Transform for discrete convolution, this leads to significant computational gains. The latter can only be seen in massive simulations, however, so the final choice of the estimator should be ultimately based on specific goals and computational resources.

Avenues for further research include boundary correction, data-driven bandwidth selection, shape-restricted estimation of the valuation quantile function and associated functionals, and auction heterogeneity.

Acknowledgements

We are grateful to Karun Adusumilli, Tim Armstrong, John Asker, Yuehao Bai, Zheng Fang, Sergio Firpo, Antonio Galvao, Andreas Hagemann, Vitalijs Jascisens, Hiroaki Kaido, Nail Kashaev, Michael Leung, Tong Li, Vadim Marmer, Hyungsik Roger Moon, Hashem Pesaran, Guillaume Pouliot, Geert Ridder, Pedro Sant'Anna, Yuya Sasaki, Matthew Shum, Liyang Sun, Takuya Ura, Quang Vuong, Kaspar W\"{u}thrich, and seminar participants at USC for valuable comments. All errors and omissions are our own.

\part*{Appendix}