Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
159,009 characters · 39 sections · 60 citation commands
Estimation of Auction Models with Shape Restrictions
We develop several new estimators that leverage shape constraints implied by the bidder's incentive--compatibility condition in auction models. Unlike most existing methods, the basic approach applies broadly in that it works both for a wide range of auction formats and allows for asymmetric bidders. For the case of first price auctions, we establish asymptotic results for multiple unsmoothed (piecewise--constant) and smoothed estimators of the inverse strategy function, the first direct estimator of the bid function, and estimators of a variety of other objects, including a bidder's value density function, her expected surplus, and the mean of her value distribution. We consider each of these objects separately in order to provide guidance to applied researchers on ways in which our estimators can be optimized for the specific objects of their interest. For objects like the expected surplus, our approach (unlike most existing methods) does not require the researcher to choose an input parameter (e.g. a kernel bandwidth) and achieves the semiparametric efficiency bound. For objects like the valuation distribution, we use a boundary--corrected kernel smoothing method so that our estimators converge at the same optimal nonparametric rate as popular alternatives. We purposely propose a relatively large number of estimation options because different choices work better depending on the nature of a given problem, as is borne out by our simulation study. We further provide simulation evidence to confirm our theoretical finding that, because our approach imposes shape restrictions a priori, it is more robust to the choice of inputs compared with alternative estimation strategies and also appears robust to the choice of design.
The key insight behind our approach is that, despite the great diversity of auction formats we might consider, the fundamental nature of a bidder's decision problem is the same. Specifically, given the strategies of a bidder's competitors and the distribution of their private values, the bidder chooses her bid to optimally trade off the probability of winning with her expected payment to the seller. Though the details of this trade--off as a function of the bid are complicated and depend on the specifics of the auction rules, an envelope theorem argument demonstrates that the equilibrium payment function $e$ must be convex in the probability with which the bidder expects to win. Moreover, the first--order condition of the bidder's problem requires that the slope of the equilibrium expected payment function at the optimally chosen win probability is equal to her private valuation. Thus, the derivative $\alpha$ of the convex payment function is equivalently viewed as the inverse strategy function, which maps optimally chosen win--probabilities to values. These facts have been used to establish the revenue equivalence theorem myerson1981optimal,Milgrom1982 and were subsequently invoked as a generic nonparametric identification strategy Larsen2018. Our paper exploits this change of variables from bids to win--probabilities further in order to relate the literature on nonparametric estimation in auctions to the large literature on nonparametric estimation under shape constraints.
The main benefits of this change of variables are threefold. First, by reformulating the target of estimation as the slope of a convex function, we open the door to a variety of well--known estimation strategies such as (shape--)constrained (nonparametric) least squares and (a new version of) nonparametric maximum likelihood estimation (MLE),\footnote{See e.g.\ brunk1955maximum for an early example of nonparametric estimation subject to shape constraints.} as well as some more adventurous estimators like a jackknife estimator. Second, this framework generalizes the large and growing toolkit for nonparametric estimation and testing in first--price auctions to generic auction mechanisms. And, finally, it allows the econometrician to easily impose the structure of symmetric equilibria, namely that the marginal distribution of a bidder's optimally chosen win--probability is a known function that only depends on the number of bidders. Importantly, the distribution function does not depend on the unknown distribution of the bidders' private valuations. This a priori knowledge of the win--probability distribution yields sizable improvements in the asymptotic distribution of our estimators.
The latter observation is especially useful when we consider the estimation of objects that are less primitive than the value density $f_v$ but may be of more direct interest to the researcher. For example, the bidder's ex ante expected surplus can be expressed as an integral of $\alpha(p) p - e(p)$ with respect to the win--probability distribution. The win--probability distribution can be precisely estimated in symmetric or asymmetric equilibria. In the symmetric case, however, we show that one can significantly reduce the asymptotic variance by substituting the win--probability distribution which is known to prevail in any symmetric equilibrium rather than an estimate thereof.
Similar to the bidder's surplus, the mean of the bidder's valuations and the seller's expected profit as a function of the number of bidders can also be expressed as a (weighted) integral of the inverse strategy function. We show that such objects can be estimated at a square--root rate of convergence when ${\alpha}$ is replaced with an unsmoothed estimate. The reason is that although the first step estimator of ${\alpha}$ converges at a cube--root rate, it has little bias.\footnote{Indeed, the asymptotic distribution is centered at zero.} The act of integration in the second stage acts as an average and hence reduces the asymptotic variance. Moreover, the resultant estimators achieve the semiparametric efficiency bound when one fully exploits the symmetric structure of the equilibrium. If the researcher does not assume bidders are symmetric or only observes one competitor bid per auction, the semiparametric efficiency bound on the asymptotic variance is larger, but we again show that the unsmoothed plug--in estimators for the mean valuation and bidder's surplus attain the efficiency bound.
Thus, smoothing the estimate of ${\alpha}$ does not necessarily improve the asymptotic performance of the desired object. Indeed, it can be detrimental. For example, in order to achieve square--root consistency of the mean valuation using a smoothed estimate of the inverse strategy function, one would need to “undersmooth” by choosing an input parameter to slow down the pointwise rate of convergence of the inverse strategy function and reduce its bias. However, one must avoid too much undersmoothing using methods that do not impose monotonicity a priori (e.g.\ marmer2012quantile,ma2019monotonicity, and Guerre2000), because letting the bandwidth go to zero for a fixed sample size would yield an inconsistent estimator for the inverse bid function and also produce an inconsistent estimator of the mean valuation. In contrast, there is no risk of too much undersmoothing using our approach, because our undersmoothed and unsmoothed estimators of objects like the mean valuation are asymptotically equivalent and attain the efficiency bound. This asymptotic efficiency result for both our unsmoothed and undersmoothed estimators therefore provides a large degree of robustness in the choice of bandwidths relative to existing methods.
More generally, we identify several ways in which the estimator can be tailored to the ultimate target of the empirical analysis. Though it would be possible to first estimate the value density and then obtain, e.g.\ bidder one's expected surplus, there is no benefit of taking this intermediate step. Indeed, as noted in the previous paragraph, making choices that optimize accuracy of an estimator of ${\alpha}$ or $f_v$ is usually harmful in terms of estimation of the eventual object of interest. As another example, the researcher might select inputs to minimize the integrated mean square error of the estimator for the quantile function of a bidder's valuations, which may be written as ${\alpha}$ evaluated at the quantiles of the win--probability distribution. Because the density of the win--probabilities is often unbounded at the left boundary, the researcher might have to smooth ${\alpha}$ less near the left boundary than away from it in order to ensure integrability of the mean square error. We implement this idea by applying a transformation to the data in conjunction with a kernel--based smoothing method.
Our paper relates to recent work on identification in trading mechanisms and estimation of monotone bidding strategies in first--price auctions. Larsen2018 operationalize a similar change of variables to prove generic nonparametric identification results in settings where the researcher does not observes the rules of the mechanism and may not directly observe the agent's actions, either, but is willing to assume the data are generated in a Bayes--Nash equilibrium. Thus, their analysis begins one step behind ours in the sense that they estimate the mapping from actions (e.g. bids) to outcomes (payments and allocations) in a first stage. Not surprisingly, their simulations demonstrate their approach suffers from a large loss of precision compared to estimation strategies that take advantage of prior knowledge of the auction mechanism. We therefore view our respective contributions as complementary advances in identification and estimation of auctions and auction--like mechanisms under shape constraints.
Three recent papers have also considered shape--constrained estimation in first--price auctions. henderson2012empirical impose monotonicity on a nonparametric estimator of the inverse bidding strategy---which is equivalent to convexity of the expected payment function---by `tilting' the empirical distribution of the bids, luo2018integrated consider an alternative approach that imposes convexity of the integrated quantile function of the bidders' valuations, and ma2019monotonicity apply a rearrangement technique to the first step estimator in GPV. The constrained least squares estimator in this paper may be viewed as an extension of luo2018integrated to more general auction models with possibly asymmetric bidders, which we achieve by considering the equilibrium expected payment instead of the integrated quantile function. Indeed, both our constrained least squares estimator and the one in luo2018integrated can be characterized as the (slope of) greatest convex minorants (GCM), albeit of different functions. In the case of first--price auctions with two symmetric bidders, the integrated quantile function coincides with the equilibrium expected payment function, and our constrained least squares estimator is numerically equivalent to Luo and Wan's estimator. More generally, however, the integrated quantile function differs from the expected payment function when there are more than two bidders, and it need not be convex when bidders are asymmetric, because a bidder with a valuation equal to the $\tau$--quantile of its distribution will generally not submit a bid equal to the $\tau$--quantile of its highest competing bid. Thus, the Luo and Wan approach does not apply to asymmetric auctions. The approach pursued in ma2019monotonicity uses the bids instead of the probabilities and hence does not readily extend to other auction mechanisms.
Like the estimator proposed in luo2018integrated, neither our constrained least squares estimator nor our nonparametric MLE requires the choice of an input parameter. Both of our unsmoothed estimators of the equilibrium expenditure function $e$ converge as a process to the same (tight) Gaussian limit process. The inverse strategy function ${\alpha}$, if the choice variable is the probability of winning, is the derivative of $e$. Both of our unsmoothed estimators of ${\alpha}$ converge at a $\sqrt[3]{T}$ rate, where $T$ is the number of auctions, and both have a Chernoff limit distribution; this is also true for the estimator in luo2018integrated. Computation of both of our unsmoothed estimators is simple: the constrained least squares estimator can be computed using an off--the--shelf algorithm and we show that our nonparametric maximum likelihood estimator can be easily computed using a simple pooled adjacent violators algorithm, also.
Although the MLE is asymptotically equivalent to our least--squares alternative, the MLE exhibits finite--sample advantages over the least--squares estimator when the true expected payment function is more convex for large values of $p$, which tends to be the case when bidder one's valuations are relatively strong compared to the maximum of its competitors'. Loosely speaking, the least-squares estimator is biased upward near $p$ equal to one because it is the slope of the GCM of an unconstrained estimator for $e$, with the result that finite--sample noise in the unconstrained estimator for $e$ forces the GCM to “bow” outward.\footnote{Illustration of the bowing out issue: \tikz[baseline=8ex,scale=0.3]{ \draw[thick] (0,5)--(0,0)--(5,0); \fill[Blue] (0,0) circle(2mm) (1,0.6) circle(2mm) (2,0.75) circle(2mm) (3,2) circle(2mm) (4,2.25) circle(2mm) (5,5) circle(2mm); \draw (0,0) -- (2,0.75) -- (4,2.25)--(5,5); \draw[->,Red] (3,4)--(4.4,3.62); \draw[->,Red] (1.48,2.14)--(2.48,1.14); \draw (4.5,3.6) node[right]{slope is steep here}; } } This finite--sample bias is greater when $e$ is more convex. On the other hand, the MLE is less negatively impacted because the MLE for $e$ is not forced to “bow” as much as the least--squares estimator. Hence, the MLE can be expected to outperform the least-squares estimator in finite samples in auctions with a small number of symmetric (or approximately symmetric) bidders.
Our nonparametric maximum likelihood estimator can alternatively be characterized as a two--step estimator, in which the first step is an inverse isotonic regression function estimator that yields bidder one's bid function. To our knowledge, this is the first direct estimator of the equilibrium bid function itself.\footnote{The estimator that comes closest is bierens2012semi, which assumes symmetry and independence, parameterizes the density of valuations and then matches the bid distribution implied by candidate parameter values to the observed bid distribution. The estimation method is (semi)nonparametric in that the dimension of the parameter vector increases with the sample size, like it is in sieve estimation.}$^,$\footnote{To avoid ambiguity, we refer to the mapping from (to) values to (from) bids as the “(inverse) bid function” and the mapping from (to) values to (from) win--probabilities as the “(inverse) strategy function.”}
Although it is not our primary objective, in the interest of completeness and to facilitate comparison with earlier work, we provide estimators of the quantile function, the distribution function, and the density $f_v$ of valuations. The quantile function and distribution functions can be estimated using routine operations (such as the delta method) on our estimates of $\alpha$. As noted by GPV and others, estimating $f_v$ requires nonparametric derivative estimation and the optimal convergence rate is a leisurely $T^{2/7}$.\footnote{To obtain a fourfold improvement requires a data set that's 128 times as large, compared to 32 for typical nonparametric estimators of univariate objects and 16 for parametric estimators.} We provide estimation results for the derivative ${\alpha}'$ of ${\alpha}$, which indeed converges at the $T^{2/7}$ rate. There are two ways of estimating $f_v$ using our approach: a two--step procedure in the spirit of GPV and a one--step procedure like marmer2012quantile.\footnote{luo2018integrated raise the interesting possibility of using a first step unsmoothed estimator as an input to the second stage of GPV, but do not provide asymptotic results. It is likely that consistency obtains, but the$f_v$ convergence rate and indeed the asymptotic distribution are unknown.} We do not see any reason to prefer either the one--step or two--step procedure. We provide asymptotic linear expansions of our first step estimators to allow readers to make up their own mind.
Because first--price auction models can be identified when bidders are asymmetric and when only a subset of the bids is observed Athey2002, Campo2003, we provide separate results depending on assumptions made about the bids observed by the econometrician. The reason for this flexibility is that, in a first--price auction, bidder one's equilibrium expenditure function $e$ only depends on the distribution of the maximum competitor bid. We therefore assume that the data are sufficient to obtain an estimate of this distribution and use this (unconstrained) estimator as the starting point for our analysis. If bidders are symmetric and their bids are independent, such an estimator could be $G_T^{n-1}$, where $G_T$ is the empirical distribution of the bids. If bidders are asymmetric then the product of the competitors' marginal empirical bid distributions would be a natural choice. If there is possible dependence among the competitors' bids, either arising from dependence in the competitor's valuations or coordination in their bids, then the empirical distribution of the maximum of the competitors' bids can be used. We show how the asymptotic properties of our constrained estimator improve as we add independence and symmetry assumptions: such improvements can be substantial and depend on the object being estimated.
Our paper addresses many issues and is fairly exhaustive in several dimensions. Nevertheless, there are several issues that we do not address in the paper. First, we ignore the potential presence of a (binding) reserve price. A binding reserve price would affect identification of certain objects,\footnote{For instance, the value distribution below the reserve price.} but for many other objects, allowing for a reserve price would pose a minor, not especially interesting (from an econometric perspective), nuisance. Further, we do not allow for endogenous entry. Although endogenous entry can be an important concern in empirical work and raises interesting modeling and identification questions Levin1994, Li2009, Marmer2013, Gentry2014, there are many ways of modeling this and it would be beyond the scope of this paper. The same comment applies to possible risk aversion, albeit that risk aversion would likely pose a tougher problem because nonparametric identification of the bidders' utility functions requires an exclusion restriction Guerre2009. Finally, there can be auction--level heterogeneity. Correcting for observed heterogeneity is relatively routine; unobserved heterogeneity might be addressed using methods similar to Krasnokutskaya2011 or Roberts2013. We leave these questions for future work.
We analyze the performance of our estimators in a fairly extensive simulation study. In general, we find that our estimators perform well and exhibit considerable robustness, both with respect to the design of the simulation study and to the choice of input parameter. However, no clear winner emerges and our various methods differ in systematic ways that are consistent with our asymptotic theory and with intuition. Hence, we do not offer empirical researchers a specific recommendation; rather, we provide a collection of tools, asymptotic results, and general insights that can be applied on a case--by--case basis.
Our paper is organized as follows. In (ref) we describe our model. (ref) contains the description of our unsmoothed estimators of the equilibrium expenditure function $e$ and its derivative ${\alpha}$, including a description of their computation. (ref) also contains a description of the direct estimate of the bid function. (ref) introduces and provides results for the smoothed versions of our estimates of ${\alpha}$ and its first derivative ${\alpha}'$, including boundary correction and transformation schemes, plus a description of jackknife estimators. Results for the estimation of the probability distribution of probabilities under various (a)symmetry and (in)dependence assumptions, can be found in (ref). Then, (ref) contains results on the estimation of several objects of potential interest. We present our simulation results in (ref). Finally, (ref) concludes.
Let $i=1,\dots,n$ index the bidders competing for an object in a first--price, sealed--bid auction. A bidder's value $v_{i}$ is drawn from a distribution $F_{i}$, which takes support on a compact interval in the nonnegative reals. We assume the seller sets a nonbinding reserve price of zero. We further assume each distribution is absolutely continuous with a density $f_{i}$ that is bounded away from zero on its support.
Each risk--neutral bidder chooses her bid to maximize her expected surplus taking her competitors' strategies as given. We will take bidder one (1) to be the bidder whose value is to be recovered and use a subscript $c$ to denote her competitors. Thus, bidder one solves
where $G_c\parens{b}$ denotes the probability that bidder one's competitors all bid no more than $b$.\footnote{In principle one can accommodate dependence among bidder one's competitors' bids by treating groups of bidders as individual bidders. Doing so requires additional assumptions to ensure monotonicity of equilibrium strategies.}
We can equivalently formulate the bidder's problem as a choice of her equilibrium probability of winning:
where $e_1(p) = Q_c\parens{p} p$ is bidder $i$'s equilibrium expected payment to the seller with $Q_c=G_c^{-1}$ the function.
The well--known fact that $e_1$ must be convex in $p$ can be seen as a consequence of monotonicity of the equilibrium strategies or incentive compatibility of the direct revelation selling mechanism that implements the Bayes--Nash equilibrium of the first--price auction Maskin2000. In any case, the solution to bidder one's problem is illustrated in (ref). As noted by Larsen2018 and Milgrom1982, bidder one's indifference curves in $(p,e)$-space are represented by straight lines with a slope equal to $v_1$. The optimal expected surplus is therefore attained where $\alpha_1(p) = e'_1\parens{p}=v_1$.
From here on we drop the subscript from the function $e_1$ and write $e$ to mean $e_1$.
In order to eliminate conditioning variables in our notation for the competing distribution of bids and equilibrium expected payment function, we assume valuations are independent across bidders, there is no auction--level heterogeneity, and the same set of bidders compete in each auction.
Under the assumption that bidders' valuations are independent across auctions, each auction is an independent realization of the same game. Therefore, the probability of winning and the expected payment as a function of $b$ can be estimated by $G_{cT}(b)$ and $G_{cT}(b)\,b$, where $G_{cT}$ is a suitable estimate of the distribution function $G_c$ of the maximum of bidder one's competitors' bids.
Though a piecewise linear function $e_T$ whose graph contains
converges to $e$ at a $\sqrt{T}$--rate, it is generally non--convex in finite samples, and the slope of the menu between two nearby points is a poor approximation of its derivative.\footnote{$\sqrt{T} \cparens{e_T\parens{\cdot} - e\parens{\cdot}}$ converges weakly to a Gaussian limit process.} We will show that the greatest convex minorant of $e_T$ can be used to estimate the expected payment function and its derivative in a single step.\footnote{luo2018integrated consider a greatest convex minorant estimator of a different function.} Moreover, we show that this estimator can be justified by a least--squares criterion and estimated by isotonic regression.
To motivate the least--squares criterion, suppose that a differentiable estimate of the quantile function for bidder one's highest competing bid were available. Multiplying this hypothetical estimator by $p$ would yield a differentiable, though possibly non--convex, estimator, $e_T$. A shape--constrained estimate of the derivative of the expected payment function could then be obtained by solving the following problem \[ \min_{{\alpha}\in\mathscr{A}} \parens[\bigg]{ \frac{1}{2}\,\int_0^1 \parens[\big]{{\alpha}\parens{p} - e_T'\parens{p}}^2 \:\mathrm{d} p}, \] where $\mathscr{A}$ is the set of nondecreasing nonnegative functions defined on $[0,1]$. This least--squares objective can be rewritten as \[ \frac{1}{2}\,\int_0^1 \parens[\big]{{\alpha}\parens{p} - e_T'\parens{p}}^2\:\mathrm{d} p = \frac{1}{2} \int_0^1 {\alpha}^2\parens{p} \:\mathrm{d} p - \int_0^1 {\alpha}(p)\,e'_T(p) \:\mathrm{d} p + \frac{1}{2}\int_0^1 e'_T(p)^2 \:\mathrm{d} p\,. \] The last term does not depend on ${\alpha}$ and may therefore be dropped from the criterion without affecting the shape--constrained estimator. The problem becomes
which can be solved for any $e_T$, differentiable or not, provided that the second integral in (ref) exists.
In a first-price auction, we use an unconstrained estimate of the empirical quantile function for bidder one's highest competing bid, $Q_{cT}(p)$, and set $e_T(p) = Q_{cT}(p)\,p$ in (ref). As we noted above, this $e_T$ will generally be non--convex in finite samples and piecewise linear. If $Q_{cT}(p)$ is the empirical quantile function of the highest rival bid, $e_T$ will be discontinuous at $t/T$ for $t = 1,\dots,T$, and the least--squares criterion can be rewritten as
where the second integral exists because ${\alpha}$ is bounded and increasing, and $e_T$ is left--continuous.
Given this representation, we show that the minimizer of (ref) over all ${\alpha}\in\mathscr{A}$ is a right--continuous step--function.
Using (ref), and as illustrated in (ref), the least--squares problem can be further simplified to \< \min_{{\alpha}_{T1}\leq \dots \leq {\alpha}_{Tt}} \sum_{t=1}^T\parens[\Big]{\frac{1}{2} {\alpha}_{Tt}^2 - Q_{cTt}{\alpha}_{Tt} - \parens[\big]{Q_{cTt}-Q_{cT,t-1}} \parens{t-1} {\alpha}_{Tt}}\,, \> where ${\alpha}_{Tt}= {\alpha}\cparens{\parens{t-1}/T}$ and $Q_{cTt}= Q_{cT}\parens{t/T}$.
The first term in the summand in (ref) comes from the integral of ${\alpha}^2$; the second term comes from the integral of ${\alpha}$ with respect to the linear portions of $e_T$; and the third term comes from the integral of ${\alpha}$ at the discontinuities of $Q_{cT}$.
Define ${\alpha}_T\parens{p}= {\alpha}_{T\lceil Tp \rceil}$. We can integrate up ${\alpha}_T$ to obtain a convex estimator $\breve e_T$ of $e$, \[ \breve e_T(p)= \int_0^p \alpha_T(u)\:\mathrm{d} u. \] As it turns out, $\breve e_T$ is simply the greatest convex minorant of $e_T$, and ${\alpha}_T$ is its left--derivative: see (ref). Note that $\breve e_T$ is both piecewise linear and continuous by construction.
An alternative way to arrive at an estimator ${\alpha}_T^*$ and thence an estimator $\breve e_T^*$ is by defining the problem as an inverse isotonic regression problem. The idea is to define \< {\Theta}_T\parens{\tilde {\alpha}} = \inf \operatorname*{argmin}_p \sum_{t=1}^T \cparens[\Big]{ e_T\parens[\Big]{\frac{t}{T}} - e_T\parens[\Big]{\frac{t-1}{T}} - \frac{\tilde {\alpha}}{T}} \mathbb{1}\parens[\Big]{\frac{t-1}{T} \leq p}, \> The rationale for (ref) is that (ref) essentially imposes monotonicity of the derivative of $e$: it is the natural analog to inverse isotonic regression estimators for the current context.\footnote{An inverse isotonic regression estimator can for a given $m$ be characterized as a minimizer $x$ of $\sum_{i=1}^n \parens{y_i-m} \mathbb{1}\parens{x_i\leq x}$.} An alternative way of thinking about it is that the population objective function corresponding to (ref) is \[ \frac1T \sum_{t=1}^T \sparens[\Big]{ T \cparens[\Big]{e\parens[\Big]{\frac tT} - e\parens[\Big]{\frac{t-1}T}} - \tilde {\alpha} } \mathbb{1}\parens[\Big]{\frac{t-1}T \leq p} \simeq \frac1T \sum_{t=1}^T \cparens[\Big]{ {\alpha}\parens[\Big]{\frac {t-1}T}- \tilde {\alpha} } \mathbb{1}\parens[\Big]{\frac{t-1}T \leq p}, \] which is optimized at the value of $p=\parens{t-1}/T$ for which ${\alpha}\cparens{\parens{t-1}/T}$ is the largest value less than $\tilde {\alpha}$.
Returning to (ref), an estimator ${\alpha}_T^*$ can be defined as \[ {\alpha}_T^*\parens{p}= \sup \Set{\tilde {\alpha} \colon {\Theta}_T\parens{\tilde {\alpha}} \leq p}. \] There is no a priori reason to prefer ${\alpha}_T^*$ to ${\alpha}_T$ or vice versa, albeit that ${\alpha}_T$ may be easier to compute. In fact, they are numerically equivalent because ${\Theta}_{T}(\tilde \alpha) = \sup\{p:\alpha_{T}(p) < \tilde \alpha\}$.
If bidders are symmetric then a more efficient unconstrained estimator for the expected payment function is given by $p Q_{T}\parens[\big]{p^{1/(n-1)}}$, where $Q_T$ is the empirical quantile function for the pooled sample of bids $\Set{b_\ell}$ for $\ell = 1,\dots, nT$. In this case, the solution to the least--squares problem in (ref) is found via a weighted isotonic regression of $\cparens[\big]{b_{\parens{\ell}} \ell^{n-1} - b_{\parens{\ell-1}} \parens[\big]{\ell-1}^{n-1}}/\cparens[\big]{\ell^{n-1} - \parens{\ell-1}^{n-1}}$ on $\cparens[\big]{\parens{\ell-1}/{nT}}^{n-1}$ with weights given by $\parens{\ell/nT}^{n-1} - \cparens[\big]{\parens{\ell-1}/nT}^{n-1}$, where $b_{(\ell)}$ denotes the $\ell$--th order statistic. The corresponding constrained estimator for $e$ is the GCM of $p Q_T\parens[\big]{p^{1/(n-1)}}$, as before.
We now develop some asymptotic estimation results for our convex estimator $\breve e_T$. Before we do so, we will make several assumptions and discuss conditions under which they would hold.
(ref) is a standard assumption in the auctions literature. It is sufficient to ensure the existence of monotone bid functions Lebrun2006. The common support assumption embedded in (ref) is unnecessary, but is imposed to make our analysis more wieldy. We note that (ref) is stronger than we need: we do not use independence among the competitors' valuations in the proofs of any of our theorems. The assumption can be relaxed provided that bidder one's unique best reply to its competitors' bids is a monotone pure strategy.
Similarly, (ref) is slightly stronger than we need, because our results only require bidder one's decision problem to be of the form in (ref). Thus, (ref) and (ref) are merely one set of sufficient conditions on the primitives of the model under which our results may be proven.
A consequence of the assumptions made thus far is that $Q_c'$ is continuous and bounded on any closed interval that does not contain zero. Indeed, the first order condition corresponding to (ref) implies that \[ Q_c'\parens{p}= \frac{v-Q_c\parens{p}}{p}. \] We now make a high level assumption on the convergence of an estimator of the bid distribution functions and develop conditions under which it is known to hold.
(ref) is a relatively weak assumption, and as previously discussed, is the starting point for our analysis. It is for instance satisfied if we take $G_{cT}$ to be the empirical distribution function of the maximum competitor bid, in which case \< H^*\cparens{Q_c\parens{p},Q_c\parens{p^*}}=\min\parens{p,p^*}-pp^*. \> It would also be satisfied if, instead, we assumed symmetry and took $G_{cT}= G_T^{n-1}$, i.e.\ the empirical bid distribution estimated off all bids raised to the power $n-1$, in which case\footnote{ Note that $\sqrt{T}\parens{G_T - G}$ converges to a Gaussian limit process with covariance kernel $\sparens{G\cparens{\min\parens{b,b^*}} - G\parens{b}G\parens{b^*}}/n$. Hence $\sqrt{T}\parens{G_T^{n-1}-G^{n-1}}$ converges to a Gaussian limit process with covariance kernel $G^{n-2}\parens{b} G^{n-2}\parens{b^*} \sparens{G\cparens{\min\parens{b,b^*}} - G\parens{b}G\parens{b^*}}/n$. Insert $Q_c\parens{p}=Q\parens{p^{1/\parens{n-1}}}$ to get the stated result. } \< H^*\cparens{Q_c\parens{p},Q_c\parens{p^*}} = \frac{\parens{n-1}^2}{n} \cparens[\big]{ \min\parens{ p,p^*}^{1/\parens{n-1}} - \parens{pp^*}^{1/\parens{n-1}} } \parens{pp^*}^{\parens{n-2}/\parens{n-1}}. \> A final example is one in which there is asymmetry plus independence and all bids are observed in which case\footnote{Note that \[ \tag{*} \sqrt{T}\parens{G_{cT} - G_c} = \sqrt{T} \parens[\Big]{\prod_{i=2}^n G_{Ti} - \prod_{i=2}^n G_i} \simeq \sum_{i=2}^n \sqrt{T}\parens{G_{Ti}-G_i } G_{-i1}, \label{eq:Hstar asymmetric derivation} \] where $G_{-i1}= \prod_{j\neq i,1}^n G_j$. The right hand side in (ref) converges weakly to a Gaussian limit process with covariance kernel $\sum_{i=2}^n G_{-i1}\parens{b} G_{-i1}\parens{b^*} \sparens{G_i\cparens{\min\parens{b,b^*}}- G_i\parens{b}G_i\parens{b^*}}$, which produces (ref), after noting that $G_{-i1}\cparens{Q_c\parens{p}}= p/G_i\cparens{Q_c\parens{p}}$. } \< H^*\cparens{Q_c\parens{p,Q_c\parens{p^*}}} = \sum_{i=2}^n \parens[\bigg]{ \frac1{G_i\sparens{Q_c\cparens{\max\parens{p,p^*}}}} - 1} pp^*, \> where $G_i$ is the bid distribution of bidder $i$. Note that for $n=2$ (ref) collapses to (ref) divided by two since bidder one's bids can also be used in the case of symmetry. We make (ref) to avoid having to hard--wire a particular set of distributional assumptions into the problem.
It should be noted that weak convergence of quantile processes for distributions with compact support is usually stated on $\parens{0,1}$; see e.g.\ vandervaart2000asymptotic. The reason is that Hadamard--differentiability only obtains on the interior of $\sparens{0,1}$. We prove that that weak convergence of our quantile--related process in fact obtains on $\sparens{0,1}$.
Cube--root--$T$ convergence of ${\alpha}_T$ is not surprising in view of e.g.\ kim1990cube. Indeed, $\sqrt[3]{T}\cparens{{\alpha}_T\parens{p}-{\alpha}\parens{p}}$ has a Chernoff limit distribution at each fixed $p$. Further, equations (15) and (16) in jun2015classical suggest that \< \sqrt[3]{T} \cparens{{\alpha}_T\parens{p}-{\alpha}\parens{p}} \stackrel{d}{\to} {\alpha}'\parens{p} \operatorname*{arg\,max}_{t\in \mathbb{R}} \cparens{ \mathbb{G}^\circ\parens{t} - {\alpha}'\parens{p}t^2/2}. \> where $\mathbb{G}^\circ$ is a Gaussian process. A justification and description of the properties of $\mathbb{G}^\circ$ can be found in (ref). If (ref) holds then (ref) simplifies to \< \sqrt[3]{T} \cparens{{\alpha}_T\parens{p}-{\alpha}\parens{p}} \stackrel{d}{\to} \sqrt[3]{4 \zeta^{2}\parens{p} \alpha'\parens{p}}\,\mathbb{C}\,, \> where $\mathbb{C}$ is a standard Chernoff--distributed random variable. (ref) is also justified in (ref).
Our result for ${\alpha}_T$ in (ref) extends the convergence rate result to uniform convergence. Note that this is in contrast to e.g.\ nonparametric kernel regression or density estimation where uniform convergence only obtains at a slower rate.
To this point we have relied on the least--squares criterion used to motivate the GCM estimator for $e$ and the isotonic regression estimator for its derivative. The GCM has the feature that it may be broadly applied in any auction or auction--like setting as long as an appropriate unconstrained estimate $e_T$ is available. In this section, we develop an estimator based upon a nonparametric likelihood criterion that specifically exploits the structure of a first--price auction. For ease of exposition and notation, we assume that only the maximum competitor bid is used to construct the likelihood, though we note how more data may be used to produce a more efficient estimate in (ref).
We rearrange the familiar formula for a bidder's inverse strategy function \[ {\alpha}\cparens{G_{c}\parens{b}} = b + \frac{G_c\parens{b}}{g_c\parens{b}}, \] in order to relate the density of a bidder's highest competing bid to her expected payment function: \[ g_c\parens{b} = \frac{e\cparens{G_c\parens{b}} / b }{{\alpha}\cparens{G_c\parens{b}} - b}\,. \] The loglikelihood of an independent sample of highest competing bids may then be written as
where the $b_t$'s are the maximum competitor bid and the shorthand forms $\tilde e_t$ and $\tilde {\alpha}_t$ are candidate values for $e\cparens{G_c\parens{b_t}}$ and the left--derivative of $e$ evaluated at $G_c\parens{b_t}$, respectively. Before (ref) may be used as the basis for a nonparametric maximum likelihood estimator, a few remarks are in order. First, nondecreasing convex real--valued functions defined on $[0,1]$ are continuous on $[0,1)$,\footnote{A nondecreasing convex function defined on a compact interval can jump discontinuously at the right boundary.} which implies that $e$ is uniquely determined on $[0,1)$ by its left-derivative $\tilde {\alpha}$. We will therefore replace $\tilde e$ with a function of $\tilde {\alpha}$ in what follows. Second, for a given $\tilde e$, the implied $G_{c}$ will be a proper distribution function if $\tilde e$ is convex and $\tilde e\parens{1}$ equals the highest order statistic among the rivals' observed bids. Third, because the loglikelihood contribution of $b_{t}$ is increasing in $\tilde e_{t}$ and decreasing in $\tilde {\alpha}_{t}$, the shape--constrained MLE should be piecewise linear in order to minimize the density at values of $b$ between realizations of the competitors' bids while maximizing the density at the observed bids. In particular, kinks in the MLE $\breve e_T^{\text{MLE}}$ occur precisely where $\breve e_T^{\text{MLE}}\parens{p}/p=b_{t}$ for some observed bid $b_{t}$. We can therefore maximize (ref) by searching over the space of left--continuous, nondecreasing step functions $\tilde {\alpha}$ defined on the unit interval.
To facilitate the numerical optimization of the maximum likelihood objective in (ref), we introduce the notation $b_{(t)}$ for the $t$-th lowest order statistic and use the fact that $\breve e_T^{\text{MLE}}$ is linear on the interval $[e_{(t-1)}/b_{(t-1)}, e_{(t)}/b_{(t)}]$ to express $e_{\parens{t}}$ in terms of its derivative and $e_{\parens{t-1}}$ \[ e_{(t)} = e_{(t-1)} + {\alpha}_{(t)}\parens[\Big]{\frac{e_{(t)}}{b_{(t)}} - \frac{e_{(t-1)}}{b_{(t-1)}}} = e_{(t-1)}\frac{{\alpha}_{(t)}/b_{(t-1)}-1}{{\alpha}_{(t)}/b_{(t)}-1}\,. \] Combining this recursive relationship with the constraint that $e_{(T)} = b_{(T)}$, we may write $e_{(t)}$ as
where the product $\prod_{s = t+1}^{T} a_s$ is defined equal to one for $t=T$ for any sequence $\{a_s\}$. Using this expression to replace $e_{(t)}$ in (ref), the loglikelihood becomes \< \mathscr{L}(\tilde {\alpha};b) = \sum_{t=1}^{T} \parens[\Bigg]{\log b_{(T)} + \sum_{s=t+1}^{T} \cparens[\Big]{\log \parens{\tilde {\alpha}_{(s)}/b_{(s)} - 1} - \log \parens{\tilde {\alpha}_{(s)}/b_{(s-1)}-1}} - \log b_{(t)} - \log \parens{\tilde {\alpha}_{(t)}-b_{(t)}}}\,. \>
By inspection of the above display, the MLE must satisfy ${\alpha}_{\parens{1}} = b_{\parens{1}}$ and ${\alpha}_{\parens{t}}\geq b_{\parens{t}}$. Problematically, this implies an unbounded density at the lower end of the competitor bids' support, which in turn implies that the solution to the maximum likelihood problem is not unique, since the loglikelihood criterion is infinite for any ${\alpha}$ with ${\alpha}_{\parens{1}} = b_{\parens{1}}$ and ${\alpha}_{\parens{2}}>b_2$. Nonetheless, one maximizer of the likelihood distinguishes itself from the rest because, for a fixed ${\alpha}_{\parens{1}}$ and ${\alpha}_{\parens{2}}$ with $b_{\parens{1}} < {\alpha}_{\parens{1}},\, b_{\parens{2}} <{\alpha}_{\parens{2}}$ and ${\alpha}_{\parens{2}}< 2\,b_{\parens{3}} - b_{\parens{2}}$, the solution for $\{{\alpha}_{(t)}\}$ for $t=3,\dots,T$ is unique. Furthermore, this unique solution does not depend on the values of ${\alpha}_{\parens{1}}$ and $\alpha_{\parens{2}}$ because the loglikelihood is additively separable in ${\alpha}_{(t)}$ and the monotonicity constraints on $\alpha_{\parens{t}}$ do not bind for $t=1$, 2, and 3. Thus, we may first maximize $\mathscr{L}\parens{\tilde {\alpha};b_1,\dots,b_T} - \log \parens{\tilde {\alpha}_{(1)}-b_{(1)}}$ over $\{\tilde{{\alpha}}_{\parens{t}} : t>2\}$. We may then separately define $\breve{{\alpha}}_{T,\parens{1}}^\text{MLE} = b_{\parens{1}}$ and choose any ${\alpha}_{\parens{2}}\in(b_{\parens{2}},\breve{{\alpha}}_{T,\parens{3}}^\text{MLE} ]$. Within this (shrinking) interval, the likelihood contribution of the second--lowest observed competitor bid is strictly decreasing in $\tilde {\alpha}_{\parens{2}}$. In practice, we suggest defining the MLE equal to the boundary value $\breve{{\alpha}}_{T,\parens{2}}^\text{MLE}=b_{\parens{2}}$.
By adding and subtracting $\log b_{(s)}$ and $\log b_{(s-1)}$ and canceling terms in the inner summation of (ref), we can rewrite the loglikelihood as
The Lagrangian for the isotonic maximum likelihood problem is then\footnote{We omit the constraint $\tilde {\alpha}_{\parens{3}}>\tilde {\alpha}_{\parens{2}}$ from the Lagrangian because this constraint is always slack.}
This problem can be solved using a pool--adjacent--violators algorithm (PAVA), which divides the large optimization problem into a sequence of at most $T-3$ one-dimensional optimizations. To see this, we observe that the Karush--Kuhn--Tucker (KKT) conditions for this problem are
Let $t_{j}$ be the subsequence of starting points of “blocks” for which the nondecreasing constraint binds. By construction, $\tilde {\alpha}_{(t_{j}-1)} < \tilde {\alpha}_{(t_{j})} = \dots = \tilde {\alpha}_{(t_{j+1} -1)} < \tilde {\alpha}_{(t_{j+1})}$. Complementary slackness then implies $\lambda_{t_{j}} = \lambda_{t_{j+1}} = 0$. Within each block $j$, the value $\tilde {\alpha}$ that satisfies the KKT conditions can then be found by solving for $\tilde {\alpha}$ in
To find the solution to the constrained NPMLE, we initially assign each $\tilde {\alpha}_{t}$ to its own block and set $\Set{\tilde {\alpha}_{t}}$ equal to the unconstrained solution $\tilde {\alpha}_{(t)}= \parens{t-1} b_{(t)} - \parens{t-2}b_{(t-1)}$ for $t>1$ and $\tilde \alpha_{(1)} = b_{(1)}$. This initial guess satisfies the constraint $\alpha_{\parens{t}}\geq b_{\parens{t}}$ but might not produce a monotonic sequence. Beginning with $t = 4$, the PAVA proceeds sequentially by finding the smallest $t$ such that $\tilde {\alpha}_{t}<\tilde {\alpha}_{t-1}$. If such a $t$ exists, we pool $\tilde {\alpha}_{t}$ together with the left adjacent block and recalculate $\tilde {\alpha}$ in the above first--order condition for that block. We then set $\tilde {\alpha}_{(s)}=\tilde {\alpha}$ for all $s$ in the block and repeat until no further violations are found. This algorithm will converge in no more than $T-3$ steps, because exactly one more of the $T-3$ monotonicity constraints are made to bind with equality in each step and no constraints are ever made slack again.
Importantly, every iterate satisfies the dual feasibility KKT condition $\lambda_t\geq0$ because violations of the primal feasibility condition $\tilde{{\alpha}}_{\parens{t}}\geq \tilde {\alpha}_{\parens{t-1}}$ are resolved by imposing $\tilde{{\alpha}}_{\parens{t}}=\tilde{{\alpha}}_{\parens{t-1}}$. Though primal feasibility may also be satisfied, for instance, by setting $\tilde {\alpha}_{\parens{t}}=\tilde{{\alpha}}_{\parens{t+1}}$, these deviations from the PAVA algorithm typically lead to a violation of dual feasibility unless the PAVA algorithm would have pooled these values in a later iteration. Thus, the final iterate of $\Set{\tilde {\alpha}_t}$ will satisfy the KKT conditions. (ref) formally establishes this claim in an appendix. Moreover, (ref) demonstrates that the KKT conditions are both necessary and sufficient for the constrained global maximum of the loglikelihood objective. Thus, the algorithm converges to the MLE for $\alpha$.
We next obtain the NPMLE of the equilibrium payment function by substituting $\breve \alpha_T^{\text{MLE}}$ into equation (ref). Figure (ref) depicts the maximum likelihood estimator for $e$ in comparison with $\breve e_T$ for a sample of five rival bids. In larger samples, the differences in the estimators for $e$ are not visually apparent.
As can be seen in (ref), the MLE is invariably above the GCM estimator. This is no coincidence. The nodes of the GCM are positioned at integer multiples of $1/T$ by construction, whereas the MLE can move the position of the nodes as well as the values at the nodes. The MLE can therefore achieve both convexity and proximity to the original nonconvex estimator without having to duck below the original estimator everywhere.
An estimate ${\beta}_T\parens{v}$ of bidder one's bid function at $v$ can be obtained as the minimizer of \[ \mathbb{S}_{T}\parens{b,v} = \sum_{t=1}^{T}\parens[\Big]{\frac{t-2}{v-b_{\parens{t}}} - \frac{t-1}{v-b_{\parens{t-1}}}}\mathbb{1}\parens{b_{(t)}\leq b}\,. \] The estimator ${\beta}_T$ is monotonic and it can be inverted to obtain an estimator ${\alpha}_T^\text{mle}$ of ${\alpha}$ at $p=G_c\parens{b}$. Indeed, it turns out that both \( \sqrt[3]{T} \cparens{{\beta}_T\parens{v}-{\beta}\parens{v}} \) and \( \sqrt[3]{T} \cparens{{\alpha}_T\parens{p}-{\alpha}\parens{p}} \) have limiting Chernoff distributions. Indeed, we have \< \forall 0<p<1: \sqrt[3]{T} \cparens{{\alpha}_T\parens{p}-{\alpha}\parens{p}} \stackrel{d}{\to} \sqrt[3]{4 {\zeta}^2\parens{p} \cparens[\big]{2Q_c'\parens{p}+Q_c”\parens{p}p}} \mathbb{C}, \> where $\mathbb{C}$ is a standard Chernoff distribution. A sketch of the proof and a derivation of the limit distribution can be found in (ref). The limit distribution in (ref) is the same as that in (ref) under (ref).
Presumably the limit distribution of the least--squares and maximum likelihood estimators would also coincide when we use all bids, not just the maximum rival bid. In the case of $n$ symmetric bidders, the likelihood of the pooled sample of bids is obtained from $\parens{n-1} g(b) = G\parens{b}/\parens{v-b}$, which becomes $\cparens{e\parens{p}/p}^{1/\parens{n-1}}/\cparens{{\alpha}\parens{p} - Q\parens{p^{1/\parens{n-1}}}}$ after a change of variables. The MLE can then be computed by applying the PAVA to the objective \[ \sum_{\ell=1}^{nT}\cparens[\Bigg]{ \frac{\ell-n}{n-1} \log\parens[\big]{\alpha_{\parens{\ell}} - b_{\parens{\ell}}} - \frac{\ell-1}{n-1}\log\parens[\big]{\alpha_{\parens{\ell}} - b_{\parens{\ell-1}}}}\,. \] The maximum likelihood estimator for the bid function at a fixed $v$ is then given by the minimizer over $b$ of \[ \sum_{\ell=1}^{nT}\cparens[\Bigg]{ \frac{\ell-n}{\parens{n-1}\parens{v - b_{\parens{\ell}}}} - \frac{\ell-1}{\parens{n-1}\parens{v - b_{\parens{\ell-1}}}}}\mathbb{1}\parens{b_{\parens{\ell}}\leq b}\,. \] The maximum likelihood estimator using the full vector of bids in the asymmetric case is considerably more complicated.
Though not specific to shape--constrained estimation, this paper would be incomplete if it did not also address smoothing because, for objects of interest such as the valuation distribution, smoothing improves the convergence rate of estimators of $\alpha$. Indeed, if $\alpha$ is twice continuously differentiable then the $\sqrt[3]{T}$ convergence rate of the unsmoothed estimators of $\alpha$ can be improved to the standard nonparametric $T^{2/5}$ rate. In this section, we first introduce our basic smoothing method, which is similar to that in luo2018integrated, then develop two important enhancements: boundary correction and transformation.
As noted by hickman2015replacing, boundary correction can be important in the estimation of auction models, especially if the objective is to estimate the density of valuations. The reason is that the bid distribution (in hickman2015replacing) or the distribution of win probabilities (here) has compact support and it is well--known that, absent a boundary correction, most nonparametric density estimators are inconsistent at the boundaries. The situation is more favorable in our case since we know that probabilities vary from zero to one whereas the top of the bid distribution must be estimated, albeit that this can be done super--consistently. We provide two distinct boundary correction methods, one based on boundary kernels, and one on a boundary correction scheme in the spirit of hickman2015replacing. As expected, both methods yield vast improvements on the performance of our uncorrected estimators near the boundary. In developing these methods, we have identified an improvement in the choice of the bandwidth sequence recommended in karunamuni2008some, which improves the performance of hickman2015replacing's version of the Guerre2000 estimator substantially. This improvement is described in a separate paper, pinkse2019actual.
Our smoothed estimators for $\alpha$ can be further improved by applying a transformation $\psi$ to the win--probabilities as part of the smoothing method. Indeed, we show that such transformations $\psi$ can improve the first order asymptotic mean square error, though not the convergence rate, of our smoothed estimators of $\alpha$. The effect of such transformations on first order asymptotics help explain the feature noted in ma2019inference that the asymptotic variance of the marmer2012quantile quantile--based--estimator of $f_v$ is often greater than that of the corresponding GPV estimator. These transformation methods are complements, not substitutes, to our boundary correction methods.
The main limitation of our method above is that $\alpha_T$ converges at a $\sqrt[3]{T}$ rate. This is due to the fact that $\alpha_T$ is discontinuous and hence that $\breve e_T$ is kinky. To obtain convergence at the typical nonparametric $n^{2/5}$ rate, we can replace $\breve e_T$ with a smoothed version $\hat e_T$, defined by, \[ \hat e_T\parens{p} = \frac1h\int_{-\infty}^{\infty} \breve e_T\parens{s}\, k\parens[\Big]{\frac{s-p}{h}} \:\mathrm{d} s, \] where $k$ is a twice continuously differentiable kernel with compact support for which $\int_{-\infty}^\infty k\parens{s}s^2 \:\mathrm{d} s=1$,\footnote{$\int k\parens{s} s^2 \:\mathrm{d} s=1$ is a normalization.} and $h=h_T$ is a bandwidth such that $\Xi=\lim_{T\to \infty} \sqrt{Th^5} <\infty$. The restriction on the bandwidth sequence is not necessary for consistency of $\hat e_T$: unlike kernel--estimators employed by others, $\hat e_T$ is a consistent estimator of $e$ for all bandwidth sequences that converge to zero and the same is true for $\hat\alpha_T$ defined below.
This definition of $\hat e_T$ requires modification near the boundaries because $\breve e_T$ is not defined outside $\sparens{0,1}$. We address this issue in (ref). Before providing results for smoothed estimates of $\alpha$ evaluated away from the boundary, we need one further assumption.
(ref) is essentially equivalent to assuming that $g_c=G_c'$ is twice continuously differentiable, which is implied by continuous differentiability of the value densities Guerre2000. Thus, assuming one more continuous derivative in (ref) is sufficient for (ref). Assuming that a density is twice continuously differentiable is standard in the nonparametric kernel estimation literature. The compact subset requirement is needed since $g_c\parens{0}$ may be infinite.
The variance formula in (ref) is intimidating, but in many cases it simplifies substantially. First, as noted following (ref), if $G_{cT}$ is taken to be the empirical distribution function of the maximum rival bid then (ref) holds.
Since $k$ is chosen, $\mathscr{V}$ is easy to estimate.
A second simplification obtains under full symmetry, i.e.\ when (ref) holds.
Finally, we provide a result for the asymmetric IPV case with $n$ bidders.
Note that for $n=2$, the result in (ref) reduces to that in (ref). For $n>2$, the function $H$ in (ref) is generally more favorable, i.e.\ it is more efficient to estimate each rival bid distribution separately than to estimate the distribution of the maximum rival bid using only the maximum rival bids.\footnote{For the case in which rival distributions happen to coincide but this fact is not used in the estimation, $\mathscr{V}$ in (ref) reduces to ${\kappa}_2 {\zeta}^2\parens{p} p^{\parens{n-2}/\parens{n-1}}$ which equals $\mathscr{V}$ in (ref) if $n=2$ or $p\in \Set{0,1}$ but is otherwise less. More generally, note that $\mathscr{V}$ in (ref) is ${\kappa}_2{\zeta}^2 \sum_{i=2}^n G_{-i1}^2 g_i / \sum_{i=2}^n G_{-i1} g_i$ which is equal to ${\kappa}_2 {\zeta}^2$ and hence to $\mathscr{V}$ in (ref) if $G_{-i1}=1$, i.e.\ if $n=2$ or $p=1$. }
A more interesting comparison is that of the formulas for $\mathscr{V}$ in (ref) if there is symmetry. Indeed, the ratio of variances is $\parens{n-1}/n$ in favor of exploiting symmetry. This result is intuitive since exploiting symmetry means that one can also use the bids of bidder one to estimate $G_c$: one then uses data on $n$ bids per auction instead of $n-1$.
One limitation of (ref) compared to (ref) is that (ref) does not extend to all of $\sparens{0,1}$. A second issue is that the bias of $\hat\alpha_T$ can be large for small values of $p$ as the following example illustrates.
The GPV estimator also has the unbounded bias at zero problem.
Below, we address each of these limitations.
Let ${\psi}$ be an increasing function such that for $j=1,2,3$, $e^{\parens{j}}\parens{p} / {\psi}'^j\parens{p}$ and ${\psi}^{\parens{j}}\parens{p} /{\psi}'^j\parens{p}$ are uniformly bounded on $(0,1]$ and for which $\lim_{p \downarrow 0}$ of each of these functions is finite, also. Then define
To see how (ref) solves the exploding bias near zero problem, consider the following. The reason we needed $e$ to be three times boundedly differentiable in (ref) is that its proof contains a second order (bias) expansion of both $e\parens{p+sh}-e\parens{p}$ and ${\alpha}\parens{p+sh}-{\alpha}\parens{p}$: the former for $\breve e_T$, the latter for $\hat\alpha_T$. If one uses $\hat e_{T\psi}$ then the corresponding expansions become $e\sparens{{\psi}^{-1}\cparens{{\psi}\parens{p}+sh}} - e\parens{p}$ and ${\psi}'\parens{p}\parens[\big]{{\alpha}\sparens{{\psi}^{-1}\cparens{{\psi}\parens{p}+sh}} - {\alpha}\parens{p}}$. This makes all the difference since the second derivative of the first difference with respect to $s$ evaluated at $s=0$ is
The corresponding expression for the second difference in the preceding paragraph is
Note that the asymptotic bias in (ref) can be made to equal zero by choosing ${\psi}=e$. Unfortunately, we do not know $e$, so making that choice is infeasible.
Consider (ref).
The bias formula in (ref) is somewhat complicated and a downside of the formula for $\hat {\alpha}_{T{\psi}}$ in (ref) is that, depending on the choices of $k,{\psi}$, it may require numerical integration. This is an inconvenience more than a serious problem since ${\alpha}_T$ is piecewise constant. However, both issues can be addressed by replacing $\hat {\alpha}_{T{\psi}}$ in (ref) with
which produces the simpler form
in lieu of (ref). The asymptotic bias is zero if one chooses ${\psi}={\alpha}$, which is again infeasible.\footnote{Recall that choosing ${\psi}=e$ was infeasible for $\hat {\alpha}_{T{\psi}}$.}
When computing $\hat e_T$ at values of $p$ near the boundary or using a kernel with infinite support, the locally weighted average of $\breve e_T(s)$ attempts to put positive weight on values of $\breve e_T$ for which $\breve e_T$ is undefined. If one does not make adjustments to the kernel $k$ or the definition of $\breve e_T$ outside of $\sparens{0,1}$ then ${\alpha}\parens{1}$ will not be consistently estimated, as the following example illustrates for $\bar {\alpha}_{T{\psi}}$.
If the lower end of the valuation's support is zero then ${\alpha}\parens{0}=0$ and the estimator $\bar {\alpha}_{T{\psi}}^\mathrm{bad}\parens{0}$ is a consistent estimator of ${\alpha}\parens{0}=0$. Note that the problem is true whether ${\psi}\parens{p}=p$ or not.
There are many solutions to this problem. The traditional approach is to use a `boundary kernel,' i.e.\ a kernel that scales the kernel to make up for the lost mass if a function is estimated near a boundary. We discuss this possibility in (ref). A second possibility is to make use of techniques similar to those espoused in karunamuni2008some in order to “make up” values of $e$ and $\alpha$ beyond $\sparens{0,1}$. This approach is investigated in (ref).\footnote{gimenes2019quantile smooth the quantile function using a local polynomial approach. The problem studied therein is otherwise unrelated.}
The boundary bias problem can be addressed by the use of boundary kernels. We replace (ref) with
where, for each $p\in[0,1]$, the function $k_{{\psi} h}(\cdot \bkdelim p)$ is a boundary kernel defined now. Let $\bar {\upsilon}_{\psi} =\cparens{ {\psi}\parens{1}-{\psi}\parens{p}} / h $ and $\underaccent{\bar}{\upsilon}_{\psi}=\cparens{ {\psi}\parens{0}-{\psi}\parens{p}} / h $. Then we require $k_{{\psi}h}$ to be such that for all $p\in \sparens{0,1}$, $j=0,1,2$,
where the requirement for $j=2,p\in \Set{0,1}$ is replaced with boundedness. The requirement that the kernel integrate to one is to ensure consistency in view of (ref). We also want it to integrate to zero if multiplied by $s$ to kill the `$h$ term' in a bias expansion.
Boundary kernels are easy to construct as (ref) demonstrates.
The cut--out for $j=2$ and $p\in \Set{0,1}$ in the requirements for the boundary kernel is there because the requirements on $\int s^2 k_{{\psi}h}\parens{-s}$ only affect the `bias' in the asymptotic distribution and because it simplifies the formula for the boundary kernel. A formula for a boundary kernel that does not require this exception is provided in (ref) in (ref).
We now turn to making $\bar {\alpha}_{T{\psi}}$ boundary--compliant, also. We use
which produces the following theorem.
In both (ref) the kernel used was taken to be the kernel constructed in (ref). This is inessential. Indeed, the results go through with ${\phi}$ replaced with a second--order kernel $k$ if $k_{{\psi}h}$ is chosen as $k_{{\psi}h}\parens{s\bkdelim p} = (\omega_{{\psi}k1} - {\omega}_{{\psi}k2}s)k(s)$ where ${\omega}_{{\psi}k1}$ and ${\omega}_{{\psi}k2}$ \[ {\omega}_{{\psi}k1} = \frac{{\Omega}_{{\psi}k2}}{{\Omega}_{{\psi}k0}{\Omega}_{{\psi}k2}- {\Omega}_{{\psi}k1}^2}, \quad {\omega}_{{\psi}k2} = \frac{-{\omega}_{{\psi}k1}{\Omega}_{{\psi}k1}}{{\Omega}_{{\psi}k2}}\,, \] with ${\Omega}_{{\psi}kj} = \int_{\underaccent{\bar}{\upsilon}_{\psi}}^{\bar {\upsilon}_{\psi}} u^j k(-u) \:\mathrm{d} u$.
One advantage of $\bar {\alpha}_{T{\psi}}$ over $\hat {\alpha}_{T{\psi}}$ is computation, as the following lemma demonstrates.
A second way of implementing boundary corrections is to create artificial values of ${\alpha}_T\parens{p}$ for $p$ outside $\sparens{0,1}$. Our approach is loosely motivated by the KZ method for kernel estimators, but it is a bit cleaner because of our specific circumstances: we are trying to smooth out an existing estimator which means that we already have values of ${\alpha}_T\parens{p}$ between zero and one.
Here, we restrict $k$ to be the Epanechnikov kernel $\tikz \draw[smooth,Red,domain=-1:1,scale=0.35] plot({\x},{(0.75-0.75*\x*\x)});$ which is a quadratic on $\sparens{-1,1}$; indeed it is $3\parens{1-x^2}/4$.\footnote{Earlier, we had taken $\int k\parens{s}s^2 \:\mathrm{d} s$ to equal one, which is not true for an Epanechnikov kernel. We adjust the asymptotic bias expression accordingly.} Consequently, any boundary correction procedure will be immaterial if the distance between ${\psi}\parens{p}$ and ${\psi}\parens{1},{\psi}\parens{0}$ exceeds $h$. We focus on correcting estimates near the upper bound, $p=1$. Impose the scale and location normalizations ${\psi}\parens{1}=0$ and ${\psi}'\parens{1}=1$.
Define
where \( {\rho}\parens{s} = s + d s^2 + \cparens{d^2 - {\psi}''\parens{1} d/6}s^3, \) with $d={\alpha}'\parens{1}/{\alpha}\parens{1}$. Then it is straightforward but unpleasant to verify that the thus extended version of ${\alpha}$ is twice continuously differentiable at one. We extend ${\alpha}_T$ analogously to (ref) using a suitable estimator $\hat d$ in lieu of $d$, defining $\hat {\rho}$ to be like ${\rho}$ but with $\hat d$ replacing $d$. We can then obtain a smoothed estimate of ${\alpha}$ by defining
where the superscript $R$ stands for `reflection.' As noted, away from the boundary, the behavior of $\bar {\alpha}_{T{\psi}}^R$ is no different than that of the estimator without boundary bias correction. So we only analyze its behavior in an $h$--neighborhood of the boundary, as formulated in (ref).
As noted, the conditions on ${\psi}$ are normalizations: without them ${\psi}\parens{1},{\psi}'\parens{1}$ would pop up in various places. Our assumption of the Epanechnikov kernel is not essential but the proofs do make use of the fact that the kernel has bounded support. Moreover, the polynomial portion of the second term in (ref) would be more complicated.
Since $\hat d$ is essentially a nonparametric kernel derivative estimator, achieving a $T^{1/5}$ rate is feasible under (ref).\footnote{If the function whose derivative is estimated is twice differentiable then it is well--known that the bias is $O\parens{h_d}$ and the variance $O\parens{1/Th_d^3}$, where $h_d$ is the bandwidth used for the estimation of $d$. Here, ${\alpha}$ is the function whose derivative is to be estimated, which is twice differentiable under (ref) since ${\alpha}''=Q_c'''p + 3 Q_c''$.} If one assumes $Q_c$ to have one more derivative at $1$ then ${\alpha}$ is thrice differentiable at 1, which would imply that picking a bandwidth $h_d$ for $\hat d$ that converges faster than $T^{-1/10}$ and slower than $T^{-1/5}$ would make the second term in (ref) disappear: $h_d \sim T^{-1/7}$ would be optimal.
So, here we advocate picking a bandwidth for $\hat d$ which tends to zero more slowly than $T^{-1/5}$ whereas KZ advocates making the bandwidth go to zero faster than $T^{-1/5}$. In a separate note pinkse2019actual we show that there is a bug in both karunamuni2005boundary and KZ and that there one needs to assume the existence of an extra derivative and choose a bandwidth that converges more slowly in order to obtain their claimed results.
Near the left boundary, we apply an analogous reflection method based upon \[ {\alpha}\parens{s} = {\alpha}\sparens[\bigg]{{\rho}_0\parens[\bigg]{\frac{{\psi}\parens{0}-{\psi}\parens{s}}{{\psi}'\parens{0}}}}{\rho}_0'\parens[\bigg]{\frac{{\psi}\parens{0}-{\psi}\parens{s}}{{\psi}'\parens{0}}}\,, \] where ${\rho}_0(s) = s - d_0 s^2 + \cparens[\big]{d_0^2 - d_0 {\psi}''\parens{0}/\sparens{6{\psi}'\parens{0}}}s^3$ and $d_0 = {\alpha}'\parens{0}/{\alpha}\parens{0}$. The formula is messier simply because we had already normalized the location and scale of ${\psi}$ at $p=1$ to simplify the expressions near the right boundary.
One caveat to our boundary kernel estimators and `reflection' procedure is that they can undo monotonicity near the boundaries in finite samples, although for different reasons. The boundary kernels are nonpositive near the boundary and are therefore capable of producing nonmonotonic estimates when ${\alpha}_T$ is relatively flat near the boundary. On the other hand, the transformation--and--reflection procedure in (ref) continuously extends ${\alpha}$ and its first two derivatives such that $\alpha(1+s)$ is generally decreasing in $s$ for large enough $s>0$. Indeed, this is inevitable when ${\alpha}'$ is close to zero and ${\alpha}''$ is negative. In any case, we may easily remedy this by redefining the smoothed estimator for ${\alpha}$ as the “cumulative maximum” of the objects defined in (ref), (ref), and (ref), for example $\bar \alpha_{T\psi}(p) = \max\cparens[\big]{\sparens{{\psi}'\parens{p}/h} \int_0^1 {\alpha}_T\parens{s} k_{{\psi}h}\cparens{\parens{{\psi}\parens{p}-{\psi}\parens{s}}/h \bkdelim p} \:\mathrm{d} s,\ \sup_{q<p}\hat {\alpha}_{T{\psi}}(q)}$.
Alternatively, in the case of the transformation-and-reflection procedure, we may apply this monotonization device to the definition of the extended ${\alpha}_{T}$. The kernel--smoothed estimator of the resulting monotonic function will then be increasing on $[0,1]$ because $k$ is a nonnegative kernel. Such a procedure will continuously extend ${\alpha}'$ and ${\alpha}''$ at one, but may introduce a discontinuity in ${\alpha}''$ at a point $p>1$ for which ${\alpha}'(p)=0$. We tolerate this discontinuity, however, because $d>0$ and a finite ${\alpha}''$ imply that the discontinuity is at a location bounded away from one. As a result, (ref) does not require any modification.
As we will see in (ref), the density of the value distribution depends on ${\alpha}'$, not on ${\alpha}$ itself. Although the primary objective in our paper concerns estimation of derived objects like the bidder surplus, we include results for the value density in the interest of completeness. For that purpose, we derive some results for an estimator of ${\alpha}'$, both away from and near the boundary.
The first result, (ref), is for the case in which we are trying to estimate ${\alpha}'$ away from the boundary, whereas (ref) applies to a neighborhood of the (upper) boundary.
Observe that the optimal convergence rate is the same as that for nonparametric kernel derivative estimators, namely $T^{2/7}$ for $h \sim T^{-1/7}$, as expected. Note further that, like before, the scale of ${\psi}$ and the bandwidth $h$ are interchangeable. Again, the optimal yet infeasible choice of ${\psi}$ in terms of the asymptotic bias is ${\psi} \propto {\alpha}$.
Although the asymptotic variances are formulated differently, the asymptotic distributions in the two theorems coincide if one takes $t=1$ in (ref). Indeed, if $t=1$ then the correction via $\hat d$ becomes immaterial since there is no boundary bias concern then. Note that if $h \sim T^{-1/7}$ then the convergence rate is still $T^{2/7}$ irrespective of the value of $t$.
A perhaps puzzling finding is that the asymptotic variance is zero if $t=0$. However, note that this is not the asymptotic variance of $\hat {\alpha}_T^{R}\,\!'\parens{1}$ itself. Indeed, the (variation in the) asymptotic distribution of $\hat {\alpha}_T^{R}\,\!'\parens{1}$ is then determined by the estimation of $d$. To get the asymptotic distribution of $\hat {\alpha}_T^{R}\,\!'\parens{1}$ itself requires us to commit to a specific estimator of $\hat d$ and derive the joint distribution. This is neither difficult nor interesting.
(ref) motivate still more estimators of ${\alpha}$. Note that ${\alpha}\parens{p}= Q_c'\parens{p} p + Q_c\parens{p} = \zeta\parens{p}+Q_c\parens{p}$. Since $Q_c$ can be estimated at a rate of $\sqrt{T}$, its estimation is of secondary concern. But $\zeta\parens{p}$ enters the variance formulas in (ref).
We will assume for the purpose of this discussion that $G_c$ is estimated using the empirical distribution function of the maximum rival bid, such that $H\parens{p,p^*}= \zeta\parens{p}\zeta\parens{p^*}\cparens{\min\parens{p,p^*}-pp^*}$ and the conditions of (ref) are satisfied.
We present two versions, one based on (ref) and one on (ref): \[ \left\{
\right. \] where the $-t$ subscripts denote leave--one--out estimators, i.e.\ the identical estimator without using observation $t$. Note that $\breve\alpha_{TJ}$ is only defined on $0<p<1$ albeit that it can be defined to equal zero at zero and one. This is precisely the reason for having $\breve e_T\parens{1}-\breve e_{T,-t}\parens{1}$ in the numerator even though it could be left out without affecting the result for fixed $0<p<1$.\footnote{$\sqrt{T}\cparens{\breve e_T\parens{1}-e\parens{1}}=\ensuremath{o_p\parens{1}}$.}
We inserted a generic estimator $\hat Q_c$ into the definitions of $\breve\alpha_{TJ},\hat\alpha_{TJ}$. Its form is largely immaterial, but natural choices would be respectively $\breve e_T\parens{p}/p$ and $\hat e_T\parens{p}/p$ for $p>0$ and zero for $p=0$.
There are three downsides to the use of these jackknife estimators. The first issue is that in their current incarnation it is assumed that $H^*$ has a specific form. But the formulas can be generalized or derived for other forms of $H^*$. Second, the jackknife estimators are costlier to compute since each estimator has to be computed $T+1$ times. This may be of little practical relevance since computation of $\hat\alpha_T$ is fast. Finally, the jackknife estimators are not guaranteed to be monotonic. This is a property they share with other estimators, including GPV, and which can be addressed by the use of a monotonization procedure, which is not difficult but admittedly cumbersome.\footnote{See ma2019monotonicity for a monotonization procedure of the GPV estimator.} We do not study the asymptotic properties of jackknife estimators in this paper.
We now turn to the much simpler problem of estimating the distribution of a bidder's equilibrium win--probabilities.
In a symmetric equilibrium with $n$ bidders, the probability that a bidder with a valuation of $v$ wins is simply given by the probability that all other bidders have a valuation less than $v$. Accordingly, the distribution of a bidder's optimally chosen win--probabilities is \[ F_{p}(p) = p^{1/\parens{n-1}}\,, \] No estimation is necessary if $n$ is known because the distribution of equilibrium win--probabilities does not depend on the unknown distribution $F_{v}$.
We can accommodate endogenous entry as long as the screening value, i.e. the lowest valuation $v^*$ for which a bidder is willing to participate, is observed. We would simply define $F_{p}(p)$ to equal zero for all $p < \alpha^{-1}(v^{*})$. For example, in a first--price auction with a reserve price $r>\underbar{$v$}$, $v^{*}=r$ and \[ F_{p}(p) = \left\{
\right.\,. \] We will use this fact when we discuss estimation of counterfactual expected revenues for the seller in (ref).
In a high--bid auction\footnote{A high--bid auction is one in which the highest bidder wins with probability one.} with bidders whose valuation distributions are not identically distributed, the equilibrium distribution of win--probabilities for bidder is $F_{p}(p) = G\cparens{Q_{c}(p)}$. The distribution $F_{p}$ can then be estimated in a straightforward fashion as $F_{pT}(p) = G_{T}\cparens{Q_{cT}(p)}$, where $G_{T}$ and $Q_{cT}$ are the empirical distribution of bidder 1's bid and an estimate of the quantile function of its highest competing bid. The weak convergence of this process on $(0,1)$ is closely related to the extensively studied “copula process” and the fact that the marginal bid densities are strictly positive on their compact support.
Recall that $ H^*\cparens{Q_c\parens{p},Q_c\parens{p^*}}$ can be as simple as $\min\parens{p,p^*}-pp^*$ in case only the maximum competitor bid is used: see (ref).
In a particular application, the true marginal distributions of valuations might not differ substantially, even when the econometrician is unwilling to impose bidder symmetry in the estimation. Thus, a minimum relative entropy estimator for the distribution of win--probabilities may be an attractive alternative. Define $\hat f_{pT}$ as the minimizer of \[ \int_{0}^{1} f_{p}\parens{p} \log\parens[\Big]{ f_{p}\parens{p} p^{\frac{n-2}{n-1}}}\:\mathrm{d} p\, \qquad \text{subject to} \qquad \int_{0}^{1} \delta_{T}\parens{p} f_{p}\parens{p} \:\mathrm{d} p = \int_0^1 \delta_{T}\parens{p}\:\mathrm{d} F_{pT}\parens{p}, \] where $\Set{\delta_{T}}$ is a user--specified sequence of functions.\footnote{Elsewhere, we use $k$ to denote kernel and $h$ to denote bandwidth. In view of the similar meaning and limited scope for confusion we duplicate notation here to make better use of other symbols.} A natural choice would be $\delta_T(p) = \sparens{1,p,p^{2},\dots, p^{\iota_{T}}}'$ for some growing sequence of natural numbers $\iota_{T}$. The solution to this problem is \( f_{p}\parens{p} = \exp\cparens{\mu'\delta_{T}\parens{p}}p^{\parens{2-n}/\parens{n-1}}\,, \) where $\mu$ solves \( \int_{0}^{1} \delta_{T}\parens{p} \exp\cparens{\mu'\iota_{T}\parens{p}}p^{\parens{2-n}/\parens{n-1}}\:\mathrm{d} p = \int_0^1 \delta_{T}\parens{p}\:\mathrm{d} F_{pT}\parens{p}\,. \) Given our choice of $\delta_{T}$, the estimate $\hat{f}_{pT}$ is the nearest density (in the sense of Kullback--Leibler divergence) to the symmetric case that matches the first $\iota_{T}$ sample moments of $p$. Since this yields something similar to a sieve estimator, we do not provide asymptotic results here and refer to chen2007large for details of such estimators.
Applied researchers typically are not directly interested in the private values that rationalize a particular sample of bids, but may estimate these so--called pseudo values in order to construct other estimates. For instance, the sample of pseudo values may be used to obtain estimates of the private value distribution.
The same comment applies to the density of the private value distribution: because the marginal value distributions are the primitives of the model, i.e. any counterfactual outcomes or other objects of interest may be computed using the private value distribution, estimating the (density of the) pseudo values at an optimal rate is considered a goal itself in a good chunk of the literature. This intermediate step may be unnecessary or undesirable when the ultimate target of estimation can be written in terms of higher level objects or when the distribution of equilibrium win--probabilities is known. Below are some examples.
In each case, the object of interest may be expressed as $\theta\parens{{\alpha}, F_p}$, where ${\theta}$ is a known function, and we estimate the object by plugging in some combination of estimates of ${\alpha}$ and $F_p$. There are two overarching themes in the following discussion. First, the asymptotic derivations are greatly simplified by the fact that $F_p$ is known in any symmetric equilibrium, and we can expect significant improvements in finite--sample (and often also asymptotic) performance when we plug in the true $F_p$ as opposed to an estimated distribution and pool bids across bidders to more accurately estimate the rival bid distributions. Second, we may expect the plug--in estimator for ${\theta}$ to be $\sqrt{T}$--consistent and asymptotically unbiased when ${\theta}$ takes the form ${\theta}\parens{{\alpha}, F_p} = \int {\theta}_1(\alpha) \:\mathrm{d} F_p$ for an appropriately differentiable function ${\theta}_1$, as is often the case when integrating over nonparametrically estimated objects.\footnote{By `asymptotically unbiased' we mean that the limit distribution has mean zero.}
We now turn to a discussion of individual objects to be estimated. Although not the primary objective in our exercise, we briefly discuss how to estimate the value distribution function, quantiles, and density function in (ref). We then turn to some objects of greater interest to us, namely the bidder surplus, the mean of the value distribution, profit as a function of the number of bidders, and profit as a function of a hypothetical reserve price.
There are different attributes of the value distribution that can be estimated. The easiest object to recover is the quantile function. Note that since $v = {\alpha}\parens{p} = {\alpha}\cparens{G_c\parens{b}}$, \[ Q_v\parens{{\tau}}= {\alpha}\cparens{Q_p\parens{{\tau}}}={\alpha}\sparens{G_c\cparens{Q_b\parens{{\tau}}}}, \qquad {\tau} \in \sparens{0,1}, \] which simplifies to ${\alpha}\parens{{\tau}^{n-1}}$ in the symmetric case.\footnote{In the symmetric case, a direct estimator of the quantile function like the one proposed in gimenes2019quantile may be preferable.} The functions $G_c,Q_b$ can be estimated $\sqrt{T}$--con\-sis\-tent\-ly, but not so for ${\alpha}$ as our results thus far have shown. So even though we are estimating quantiles, namely quantiles of the value distribution, these quantiles cannot be estimated at the parametric rate because the values are not observed. Indeed, the limit distribution of an estimator $\hat Q_v$ of $Q_v$ is simply the limit distribution of whatever estimator of ${\alpha}$ is used evaluated at $p=G_c\cparens{Q_b\parens{{\tau}}}$. Likewise, the value distribution function is simply \[ F_v\parens{v} = F_p\cparens{{\alpha}^{-1}\parens{v}}. \] With symmetric bidders, $F_p= p^{1/\parens{n-1}}$. In the case of asymmetry, $F_p$ can be estimated $\sqrt{T}$--consistently, such that the limit distribution is by the delta method given by $\parens{f_p/{\alpha}'}\cparens{{\alpha}^{-1}\parens{v}}=f_v\parens{v}$ times the limit distribution of the estimator of ${\alpha}$. Note that the delta method is only valid for $v \neq 0,\bar v$, which is of little consequence since we already know the values of $F_v\parens{0},F_v\parens{1}$, albeit that uniformity arguments would suggest that the implied asymptotic distribution would not reflect the finite sample performance near 0 and 1, either, although the convergence rate is still $T^{2/5}$ for the same reason that $\breve e_T$ converges at the $\sqrt{T}$ rate on the entire interval $\sparens{0,1}$: see the comments in the paragraph following (ref).
There are two ways to estimate the value density: one--step and two--step. With the two--step estimator, one first generates valuation estimates by doing e.g.\ $\hat v_t = \bar {\alpha}_{T{\psi}}\cparens{\hat G_{cT}\parens{b_{t1}}}$ and then plugs those estimates into a nonparametric kernel density estimator. The two--step estimator is analogous to GPV except that our first step is different. It can be shown\footnote{Derivation not provided here.} that both the first step in GPV and our smoothed estimates of ${\alpha}$ permit asymptotic linear expansions of estimator minus expectation at $b=Q_c\parens{p}$, \< \left\{
\right. \> The first two formulas in (ref) are similar, but note the different arguments to the kernel and the fact that one denominator has a square on $g_c$ and the other one does not. The formula with ${\psi}$ simplifies to the one without for ${\psi}\parens{p}=p$ and to the GPV expansion for ${\psi}\parens{p}=Q_c\parens{p}$. However, the bias of our estimator with ${\psi}=Q_c$ does not coincide with that for the first step GPV bias: either can be greater.
We only provide asymptotics for the one--step estimator. For the one--step estimator, note that the value density function is \[ f_v\parens{v} = \parens{f_p/{\alpha}'}\cparens{{\alpha}^{-1}\parens{v}}, \] and hence requires an estimate of ${\alpha}'$, which we provided in (ref) In the symmetric case, $f_p\parens{p}=p^{\parens{2-n}/\parens{n-1}}/\parens{n-1}$ and one would need to use an estimate of $G_c$ (and hence $Q_c$) that fully exploits symmetry. With asymmetric bidders it also requires an estimate of $f_p$, but density estimates converge faster than do their derivatives so the estimate of ${\alpha}'$ determines the asymptotic distribution of $\hat f_v\parens{v}$, which is $- \parens{f_p/{\alpha}'^2}\cparens{{\alpha}^{-1}\parens{v}}$ times the limit distribution of one's estimate of ${\alpha}'$, again by the delta method. From (ref) it follows that the bias and variance of our estimator of the value density in the symmetric case are given by \[ \mathscr{B}_{f}^\mathrm{symm}\parens{p}= - \frac{f_p\parens{p}}{{\alpha}'^2\parens{p}}\mathscr{B}^R\parens{p}, \] and
For ${\psi}\parens{p}=Q_c\parens{p}$ (or indeed a suitable estimate thereof) our variance coincides with that of marmer2012quantile, theorem 2. ma2019inference note that the variance of the MS estimator exceeds that of GPV for the same choice of kernel and bandwidth if one undersmooths, i.e.\ if one chooses a bandwidth which makes the bias disappear faster than the variance. We recommend against undersmoothing for the purpose of estimating ${\alpha}$ and note that ${\psi}=Q_c$ is not optimal.\footnote{For the purpose of inference undersmoothing makes sense but for estimation it is better to choose a bandwidth that converges at the optimal rate since it results in a better convergence rate of the estimator than if one undersmooths. Second, (ref) suggests that the observation in ma2019inference is due to the use of a one--step instead of a two--step estimator. Finally, ma2019inference do not employ transformations like ${\psi}$, which can yield a smaller variance. Indeed, for the infeasible choice ${\psi}\parens{p}=c {\alpha}\parens{p}$ for $c>0$ one obtains a bias of zero and a variance equal to \[ \frac{c^3K_1 G^2\parens{b} g\parens{b}}{n^2\parens{n-1}\cparens{ g^2\parens{b} - g'\parens{b}G\parens{b}/n}}, \] which can be made small by choosing $c$ small. Thus, any gains one obtains from doing a two--step procedure can be obtained by making a different choice of ${\psi}$ and kernel or bandwidth. }
\todo[inline,color=Magenta]{The above results should be checked rigorously.}
Note that the bid function at $v$ is simply \( Q_c\cparens{{\alpha}^{-1}\parens{v}}, \) and that $Q_c$ can be estimated $\sqrt{T}$--consistently. Hence the limit distribution of our bid function estimate is simply $Q_c'/{\alpha}'$ times the limit distribution of the estimate of ${\alpha}$ used. Since the bid function estimate uses an estimate of the inverse of ${\alpha}$, the estimate of ${\alpha}$ had better be monotonic: this is yet another advantage of imposing monotonicity from the outset.
We now turn our attention to estimation of the bidder's surplus. The surplus for bidder one is given by
where $A\parens{p}={\alpha}\parens{p}p - e\parens{p} = Q_c'\parens{p} p^2$.
There are two important and separate cases. First, in the case of symmetry $F_p$ is known to be $p^{1/\parens{n-1}}$ and does not need to be estimated. If $F_p$ is unknown then it can be replaced with the empirical distribution function.
Regardless, one would expect $\sqrt{T}$--consistency despite the presence of nonparametric objects in the definition of $\mathrm{BS}$. This is a common theme in the semiparametric econometrics literature robinson1988root,powell1989semiparametric. Even though nonparametric estimators, other than e.g.\ the empirical distribution function, typically converge at a rate slower than $\sqrt{T}$, integrating them often restores the parametric $\sqrt{T}$ rate. The reason is that integrating is like averaging and hence reduces the variance, which opens up the possibility of undersmoothing to make the bias vanish at a rate faster than $\sqrt{T}$. Note that if the unsmoothed estimator ${\alpha}_T$ is used, no adjustment of smoothing parameters is needed at all since no smoothing is conducted in the first place. It does not appear to matter for the asymptotic distribution of our estimator of $\mathrm{BS}$ whether or not a smoothed estimator of $\mathrm{BS}$ is used, as long as it is undersmoothed. Symmetry matters a lot, however.
In (ref) we discuss the symmetric case and in (ref) the asymmetric case.
With symmetry the situation simplifies in that then
This simplifies the asymptotic theory since integration by parts and (ref) turns (ref) into \[ \operatorname{BS} = \frac{e\parens{1}}{n-1} - \int_0^1 e\parens{p} \cparens{ f_p'\parens{p} p + 2 f_p\parens{p}} \:\mathrm{d} p = \frac{e\parens{1}}{n-1} - \frac n{n-1} \int_0^1 e\parens{p} p^{\frac{2-n}{n-1}} \:\mathrm{d} p, \] which can be estimated by \[ \operatorname{\widehat{\operatorname{BS}}}^\mathrm{symm} = \frac{\breve e_T\parens{1}}{n-1} - \frac{n}{\parens{n-1}^2} \int_0^1 \breve e_T\parens{p} p^{\frac{2-n}{n-1}} \:\mathrm{d} p. \] The asymptotic theory for $\operatorname{\widehat{\operatorname{BS}}}^\mathrm{symm}$ is trivial in view of (ref).
It should be pointed out that if symmetry is fully exploited then the function $H$ is different, also. Indeed, from (ref) it follows that then
which equals \< \frac{n}{\parens{n-1}^2} \int_0^{\bar b} \int_0^{\bar b} G^{n-1}\parens{b} G^{n-1}\parens{b^*} \sparens{G\cparens{\min\parens{b,b^*}} - G\parens{b} G\parens{b^*}} \:\mathrm{d} b \:\mathrm{d} b^*, \> where we provide (ref) if readers would like to compare it to a future GPV--based estimator.
A natural generic estimator of $\mathrm{BS}$ in the absence of a symmetry assumption is \< \operatorname{\widehat{\operatorname{BS}}} = \int_0^1 \cparens{{\alpha}_T\parens{p}p - \breve e_T\parens{p}} \:\mathrm{d} F_{pT}\parens{p}. \> Naturally, ${\alpha}_T$ can be replaced with a smoothed version in which case it would be advisable to replace $\breve e_T$ with the estimator of $e$ corresponding to the smoothed estimate of ${\alpha}$, also.
The asymptotic variance in (ref) is intimidating but it simplifies considerably in an important special case, as (ref) illustrates. The fact that our estimator achieves the semiparametric efficiency bound should come as no surprise since our estimator is asymptotically linear and imposing shape restrictions is well--known not to help in reducing the asymptotic variance in many cases.\footnote{See newey1990semiparametric for a discussion and tripathi2000local for results on the semiparametric efficiency bound subject to shape restrictions in the partially linear model of robinson1988root.}
Note that the semiparametric efficiency bound is defined only relative to the amount of information available. For instance, if one uses all bids instead of only the maximum rival bid then (ref) is less than (ref) but still achieves the semiparametric efficiency bound. We do not show this. Nevertheless, it is reassuring that no other regular estimator exists with a smaller asymptotic variance under the same assumptions.
It is not immediately obvious that the variance in (ref) is worse than that in (ref), albeit that the fact that a more efficient estimate of $G_c$ can be used should tip the balance. We provide a comparison in the least favorable case for symmetry, namely that of two bidders.\footnote{With more than two bidders, the gain in efficiency of estimating $G_c$ is greater.}
The conclusion from (ref) must be that symmetry should be imposed whenever reasonable. If the researcher is not willing to assume symmetry and pool bids in estimating $e_T$, the minimum relative entropy estimator for $f_p$ may close part of the gap between $\mathscr{V}_\mathrm{BS}^a$ and $\mathscr{V}_\mathrm{BS}^\mathrm{symm}$.
As it turns out, smoothing does not improve the asymptotic distribution. Too much smoothing can introduce an asymptotic bias and slow down convergence. The most important consideration is that the implied estimator of $e$ converges (after norming and scaling) to the same Gaussian limit process as $\breve e_T$ which for the smoothed estimator simply requires that $h \to 0$ fast enough as $T\to\infty$. We state the theorem for $\bar {\alpha}_{T{\psi}}$ but the result applies with minor modifications to any estimators which satisfy the aforementioned desiderata.
Note that the bandwidth can tend to zero arbitrarily fast since a bandwidth of zero simply takes us back to the unsmoothed estimator. This is in sharp contrast to other approaches, e.g.\ one based on the estimator of the inverse bid function in GPV, where taking the bandwidth to zero before taking the sample size to infinity would blow up the asymptotic variance: letting $h \downarrow 0$ with GPV does not produce a consistent estimator of $g_c,g,{\alpha}$. Consequently, it is not clear a priori that using a second order kernel and undersmoothing GPV would produce a consistent estimator of $\mathrm{BS}$, let alone a $\sqrt{T}$--consistent estimator. We have no theoretical results on this, though our simulation results suggest that letting the bandwidth go to zero in a GPV--based estimator of $\mathrm{BS}$ would break $\sqrt{T}$--consistency.
The mean of the value distribution of bidder one is
There are several ways of estimating $\mathrm{MV}$. For instance, one can estimate it using estimates of the bid distributions directly since \[ \mathrm{MV} = \int_{0}^{\bar b} \parens[\Big]{b + \frac{G_{c}\parens{b}}{g_{c}\parens{b}}} \:\mathrm{d} G_{1}\parens{b}\,, \] which would be most natural if one used the GPV machinery. However, we will present results for \[ \widehat{\mathrm{MV}} = \int_{0}^{1}\alpha_{T}\parens{p} \:\mathrm{d} F_{pT}\parens{p}\,. \] The asymptotics for $\widehat{\mathrm{MV}}$ are similar to those for $\widehat{\mathrm{BS}}$ and the proof is therefore mercifully short.
Note that the asymptotic variance in (ref) equals zero if $n=2$. This is intuitive since then the mean of the value distribution is simply $\bar b$, which can be estimated super--consistently. Naturally, some of these properties evaporate once we examine the asymmetric case.
Note that the symmetry assumption is also easy to exploit without using our machinery, because $G_c(b)/g_c(b)= G(b)/\cparens{(n-1)g(b)}$ and \[ \mathrm{MV} = \int_0^{\bar b} b \:\mathrm{d} G(b) + \int_0^{\bar b} \frac{G(b)}{n-1} \:\mathrm{d} b\ = \frac{\bar b + \parens{n-2} \mathbb{E} b}{n-1}. \] Since the upper bound of the bid distribution can be estimated at a rate faster than $\sqrt{T}$, estimation of $\mathrm{MV}$ by replacing $\bar b$ and $\mathbb{E} b$ with their sample counterparts would work, also. So for the purpose of estimating the mean of the value distribution in the symmetric case, our methodology is probably overkill.
Estimating the seller's profit is a trivial exercise (if the seller's valuation is zero as we assume throughout) since profit is simply the sum of the winning bids. A more interesting object is profit as a function of a hypothetical reserve price $r$, $\mathrm{PR}\parens{r}$, or number of bidders $\mathrm{PR}^*\parens{n}$.
In the asymmetric case, this is a complicated endeavor. Indeed, using the machinery described earlier in the paper we can recover the value distributions for each bidder. However, there is generally no analytical solution for the bid function in the asymmetric case like there is in the symmetric case. We must therefore numerically solve for the counterfactual equilibrium bid distributions in order to compute the counterfactual revenue. Because this method does not depart from the existing literature, we limit our discussion to the symmetric case, where we do have suggestions for how to exploit the symmetry assumption in the counterfactual analysis. Since the theoretical results here are similar to those obtained earlier in terms of method of proof, we state the results in the text instead of enunciating them.
We have repeatedly made use of the fact that $F_{p}$ is a known function of $n$, the number of bidders. $F_p$ is still a known function of any counterfactual number of bidders $m$, possibly different from $n$. In addition, the counterfactual expected payment function is a known function of the factual expected payment function if the distribution of valuations is held fixed. Specifically, the equilibrium $\alpha$ for a given number of bidders $n$ satisfies $\alpha\parens{\tau^{n-1};n} = Q_{v}(\tau)$ for all $n,{\tau}$, which implies that for ${\xi}={\xi}_{mn}=\parens{n-1}/\parens{m-1}$, $\alpha(p;m) = \alpha\parens{p^{\xi};n}$ for all $p\in[0,1]$ and $n,m\geq 2$.\footnote{Note that the bid distribution changes with $n$ but the value distribution remains constant. Indeed, $Q_v\parens{{\tau}} = Q\parens{{\tau};n} + {\tau} Q'\parens{{\tau};n}/ \parens{n-1}$ defines $Q\parens{\cdot;n}$ as an implicit function of $Q_v$, indeed $Q\parens{{\tau};n}= \int_0^1 Q_v\parens{t^{1/\parens{n-1}}{\tau}} \:\mathrm{d} t$. }
The counterfactual expected payment function is then \( e(p;m) = \int_{0}^{p}\alpha\parens{t^{\xi}} \:\mathrm{d} t \) and the expected revenue is given by
where ${\chi}_j= \parens{m-2n+j}/\parens{n-1}$.
Following the previous examples, a plug--in estimator for $\mathrm{PR}^*\parens{m}$ in which we substitute an estimate of $e$ and the known $F_{p}(\cdot;m)$ converges at a $\sqrt{T}$--rate. Indeed, the limit distribution is a mean zero normal with variance \[ \int_0^1 \int_0^1 H\parens{p,p^*} \parens[\Big]{\frac{{\chi}_2+1}{{\xi}} p^{{\chi}_2} - \frac{{\chi}_1+1}{{\xi}} p^{{\chi}_1}}\parens[\Big]{\frac{{\chi}_2+1}{{\xi}} {p^*}^{{\chi}_2} - \frac{{\chi}_1+1}{{\xi}} {p^*}^{{\chi}_1}} \:\mathrm{d} p^* \:\mathrm{d} p p^* }. \]
By (1) in jun2019information, we have \[ \mathrm{PR}\parens{r} = \bar v - r F_v^n\parens{r} + \int_r^{\bar v} \cparens[\big]{F_v^n\parens{v} - n F_v^{n-1}\parens{v}} \:\mathrm{d} v. \] Using the substitution $p = F_v \parens{v}$ the problem then entails finding $p^*=F_v^{n-1}\parens{r}$ for which \< \mathrm{PR}\cparens{{\alpha}\parens{p^*}} = n\parens[\bigg]{ p^*{\alpha}\parens{p^*}\parens[\big]{1-p^{*1/(n-1)}} + \frac1{n-1} \int_{p^*}^1\int_{p^*}^p {\alpha}(u) \:\mathrm{d} u \; p^{\frac{2-n}{n-1}} \:\mathrm{d} p} \>
It should be apparent from our earlier discussion that since ${\alpha}$ is estimated at a slower--than--parametric rate, the first right hand side term in (ref) is estimated at a rate less than $\sqrt{T}$ but the second right hand side term in (ref) can be estimated at the typical parametric rate. The asymptotic distribution is hence determined by the estimation of $F_v\parens{r}$, which was already discussed in (ref). The choice of bandwidth should therefore be made with an eye toward the precise value or range of values of counterfactual reserve prices under consideration.
We provide a simulation study to compare the performance of our estimators. Our goal is not to crown a winner but to highlight systematic ways in which various methodological choices impact the bias and mean squared error of the estimator.
We parameterize bidder one's maximum competitor bid distribution as $G_{c}\parens{b} \propto \parens{\theta/b + \gamma - \theta}^{-1/\theta}$ for $b\in\sparens{0,\bar b}$ with $\bar b = 2 \big/ \parens[\big]{1+\theta + \sqrt{4 \gamma +(\theta-1)^{2}}}$ and $\gamma,\,\theta>0$. Bidder one's inverse bid function is then $\beta^{-1}(b) = (1+\theta)b + (\gamma-\theta)b^{2}$. Note that $\beta^{-1}$ is strictly increasing on $[0,\bar b]$, which implies convexity of the expected payment function $e(p) = \theta c p^{\theta + 1} \big/ \cparens[\big]{1 + c (\gamma-\theta)p^{\theta}}$, where $c$ is a constant that depends on $\gamma$ and $\theta$.
The maximum competitor bid distribution is chosen such that the support of bidder one's valuations is $\sparens{0,1}$ regardless of the values of $\gamma$ and $\theta$. We can then fix bidder one's valuation distribution and independently vary the maximum competitor bid distribution to achieve various shapes of the inverse strategy function $\alpha$ and competitor bid density, which are the respective targets of estimators based on our approach or estimators based on the inverse bid function (IBF) like GPV. (ref) plot these functions. If $\gamma=\theta$ then the competitor bid distribution is a power distribution and the inverse bid function is simply linear. As $\theta$ approaches zero, the competitor bid distribution approaches a truncated Fr\'echet distribution and the inverse bid function a convex quadratic.
For every combination of $\gamma = 3/2,\, 3/4, \,1/3,\, 1/7$ and $\theta=3/4, 1/3, 1/7, 1/9$, we draw independent and identically distributed samples of $T=100$, 250, and 500 maximum competitor bids. Thus, $T$ represents the number of auctions as well as the number of bids used to estimate bidder one's expected payment function and inverse strategy. We then independently sample $T$ draws of bidder one's valuations according to a power distribution $F_{v_1}(v) = v^{3/2}$, compute her optimal bid, and apply our various methodologies to estimate several objects of interest. We compare our estimators with an estimator based on an approach similar to GPV in which only the independent sample of highest competitor bids are used to estimate the inverse bid function. This estimator is labeled “IBF” to indicate the estimates of the various objects were constructed from a nonparametric estimate of the inverse bid function. The IBF estimator does not perform any boundary correction or trimming, hence cannot be expected to perform well near the boundary. To be clear, this is not a critique of Guerre2000: the goal in their paper is to estimate the inverse bid function and valuation density at an optimal rate in the interior of the support of the valuations. Here, we use the estimator for the inverse bid function as an input into other objects of interest. A fairer comparison with our boundary--corrected estimators can be found in “IBF--BC.” This estimator uses a boundary correction routine similar to hickman2015replacing and is hence better suited for estimating objects that require integration over the entire support of the bids or valuations.
All simulations employ an Epanechnikov kernel and a rule--of--thumb bandwidth sequence, multiplied by an additional scaling factor of $1/5, 1/2, 1,$ or $3/2$ in order to explore sensitivity to the choice of bandwidth. We use a Gaussian reference distribution for choosing bandwidths for methods that use nonparametric kernel density estimators and $\alpha(p) = \bar v p^\gamma$ as our reference function for bandwidths using our procedure. Specifically, we use the sample mean and variance of ${\alpha}_T\cparens{G_{cT}(b_{1})}$ to estimate the parameters $\bar v$ and ${\gamma}$ in the parametric reference model, then choose the bandwidth that would minimize the mean integrated squared error of the estimator for ${\alpha}$ under the reference model. This optimal bandwidth also depends on $\psi$.\footnote{For some choices of ${\psi}$, the squared error is not integrable on $[0,1]$. In these cases, our rule of thumb minimizes the integrated squared error on $[0.05,1]$.}\
We consider five different choices of ${\psi}$. The first is the identity transformation ${\psi}_1(p) = p$ and the second is the infeasible zero--bias transformation ${\psi}_2(p) = {\alpha}(p)$. The next transformation ${\psi}_3(p) = \log(p)$ ensures that the asymptotic bias is vanishingly small for $p$ close to zero, though the asymptotic variance can be large. The transformation ${\psi}_4=\sqrt{p}$ minimizes the MISE in ${\alpha}$ if ${\alpha}$ is a power function with exponent greater than $1/2$.\footnote{If ${\alpha}$ is a power function and the exponent is less than $1/2$, the optimal ${\psi}$ would be the infeasible choice ${\psi}_2$.} Finally, ${\psi}_5(p) = \sqrt[5]{p}$ balances the integrated asymptotic bias and variance of the estimator for ${\alpha}$ when ${\alpha}$ is a power function, regardless of the exponent.
For ${\psi}_2$, the rule of thumb suggests that an infinite bandwidth would minimize the integrated MSE because the first--order bias is always zero. A better rule of thumb would suggest a bandwidth sequence on the order of $T^{-1/7}$, resulting in a faster rate of convergence than the other estimators. For the sake of comparison, we do not take this route and instead use the same bandwidth as we do for ${\psi}_5$. For the undersmoothed estimates, we simply multiply the rule--of--thumb bandwidths by $T^{-2/15}$ so that the sequences are on the order of $T^{-1/3}$.
For the boundary corrected estimates that use reflection---IBF--BC and $\bar {\alpha}_{T}^R$---the auxiliary bandwidth is proportional to the main bandwidth. In pinkse2019actual, we show that choosing bandwidths converging at a rate of $T^{-1/7}$ would be optimal for both estimators if ${\alpha}$ (equivalently $g_c$) has three continuous derivatives near the boundary. In this paper, we do not choose a bandwidth sequence to capitalize on this extra smoothness because doing so would put the reflection methods at an advantage relative to the boundary kernel estimators. Unlike the reflection methods, our boundary kernel method does not involve any auxiliary input parameters that could be modified to take advantage of this extra smoothness. That said, we note that if the researcher is willing to strengthen the smoothness assumption, the reflection method or a boundary kernel method that takes advantage of the extra smoothness could be more attractive in practice precisely for this reason.\footnote{In fact, if the target of the estimation were the valuation density, the researcher might assume three derivatives of $g_c$, anyway, in order to attain the typical $T^{2/7}$ rate of convergence.}
All simulations use a thousand replications.
We first review the simulation results for the integrated objects MV and BS. (ref) illustrates the relative root mean squared error (RMSE) of our unsmoothed and smoothed, boundary--kernel--based estimators along with the IBF estimators. The bandwidths are chosen proportional to $T^{-1/3}$ so that the resulting $\sqrt{T}$--consistent estimator is asymptotically unbiased. For lack of a better rule, the constant of proportionality in the bandwidth sequence is simply the rule--of--thumb constant multiplied by our additional scale factor. The various estimators are arranged in columns, and each row represents a different combination of the target object, bandwidth scaling factor, ${\gamma}$, ${\theta}$, and $T$. The value in each cell is colored to reflect the value of the RMSE divided by the minimum RMSE across the columns. The lightest green indicates the best performing estimator, while the darkest purple indicates the RMSE was at least three times as large as that of the best performing estimator. The color scale is top--coded because some estimators performed extremely poorly.
The unsmoothed MLE consistently performs well across a variety of parameter values, while the unsmoothed isotonic regression estimator (LS) has difficulty for some parameter values because it suffers from finite--sample bias for values of $p$ close to one. Intuitively, this bias arises from the fact that the GCM, by definition, must lie below the estimate of the true expected payment function, which leads to an upward bias in its slope near $p=1$. This bias is more pronounced when ${\gamma}$ and ${\theta}$ are both relatively large, because the true expected payment is more convex near the right boundary. In contrast, the unsmoothed MLE is not as badly biased when ${\gamma}$ and ${\theta}$ are large, because the graph of the MLE for $e$ does not have to lie below the unconstrained estimator. The finite sample bias in estimates of ${\alpha}$ for large values of $p$ more negatively affects the relative performance in estimating the bidder's expected surplus because the values of $\alpha(p)$ for large $p$ are weighted relatively more in the integral formula for BS than for MV. We expect these differences in the unsmoothed estimators to vanish as $T$ increases because all our estimators in (ref) are asymptotically equivalent.
When ${\gamma}$ is small, the undersmoothed, boundary--corrected IBF estimator for BS and MV appears to under--perform in small samples, but otherwise has a relatively small RMSE. As expected, however, the relative performance of the IBF approach is sensitive to the scale of the bandwidth sequence. We cannot conclude from (ref) that our approach is robust to the choice of bandwidth, however, because it does not compare the relative performance of different bandwidth scaling factors. (ref) makes this comparison in the estimation of the bidder's surplus using $T=500$ auctions. The results demonstrate that the asymptotic behavior of our undersmoothed estimators is fairly similar across bandwidth sequences, whereas the IBF--BC estimator can be the best performing for some parameter values or worst performing estimator depending on the choice of bandwidth. We consider this robust performance, and indeed not having to choose an input parameter, a valuable characteristic of our approach.
The IBF--BC estimator for $\mathrm{BS}$ performs particularly well when $\gamma$ equals $\theta$, in which case the maximum competitor bid is a power distribution and the inverse bid function is linear. Even without undersmoothing, the asymptotic bias in the bidder's expected surplus would be zero when $\gamma = \theta = 1$ or $\gamma= \theta = 1/2$ and is fairly small at intermediate values. This fact helps explain why the RMSE for BS using LS and MLE relative to IBF--BC is larger, for example, in the $(1/3, 1/3)$ column compared to the columns on either side.
(ref) represents less than 6% of the information contained in (ref). Many more tables (available online \href{http://personal.psu.edu/kes380/files/shape_supp.pdf}{here}) provide further quantitative comparisons of the RMSE, as well as the bias of these and other estimators discussed below.
The next set of results compare the root mean integrated squared error (RMISE) of the value distribution and quantile function. Based upon the results in the first two columns, the IBF approach without boundary correction is (unsurprisingly) dominated by the boundary--corrected estimator. Next, comparing the second column with the columns to the right, we find that our estimators are again more robust to the choice of bandwidth and tend to outperform the boundary--corrected IBF approach when $\gamma$ and $\theta$ are small, i.e. the highest competing bid is stronger. Although all bandwidth sequences are proportional to $T^{-1/5}$, the finite--sample behavior of our smoothed estimators is less (negatively) impacted by a small bandwidth. This finding is related to the fact that our estimator for $\alpha$ is consistent even when the bandwidth tends to zero. Again, robustness is a virtue.
Comparing the rows for which either $\gamma$ or $\theta$ is small, we also find that our transformation method significantly reduces the RMISE in the estimate of the quantile function. In particular, the IBF--BC and the “no transformation” ($\psi_{1}$) estimators have greater RMISE compared to the estimators that employ the transformations $\psi_{3}$, $\psi_{4}$, or $\psi_{5}$. Looking across all columns, the smoothed least--squares estimator in conjunction with the transformations $\psi_{3}$ or $\psi_{5}$ appears to consistently perform best or near the best when the highest competing bid is relatively stronger ($\gamma$ and $\theta$ are small), while the the IBF--BC estimator and the smoothed MLE estimator with $\psi_{5}$ or $\psi_{3}$ perform better in terms of RMISE when the highest competing bid is relatively weak. We note that in auctions with three or more bidders one would expect the highest competing bid to be relatively strong and hence the smoothed least squares estimator to outperform the alternatives.
The differences between the estimators of the value distribution are less striking. This is due in part to the fact that some of the differences in the estimates of the quantile function are driven by the behavior near the left boundary. Using the delta method, one can show that the asymptotic distribution of the estimators for $F_{v}$ are scaled by the value density at $v$. Under our simulation design, $f_{v}$ approaches zero for small $v$. The relative differences in the estimators are dampened as a result. This also explains the fact that the log--transformation $\psi_{3}$ tends to be the best in terms of RMISE($F_{v}$) for small values of $\gamma$ and $\theta$. Under these parameters, $\alpha''$ diverges as $p$ approaches zero, which produces a large bias for small values of $p$ absent a transformation. The log--transformation ensures the asymptotic bias vanishes at the low end, while the fact that $f_{v}(v)$ is small mitigates the detrimental effects on the asymptotic variance. The transformation $\psi_{5}$ yields similar results, though it does not reduce the bias and increase the variance by as much.
We now compare the various methods of estimation of the density of bidder one's valuations. For the indirect methods, a bandwidth proportional to $T^{-1/5}$ is used in both the first and second steps. For the direct method, ${\alpha}'$ is estimated using a bandwidth proportional to $T^{-1/7}$. We compare the direct method using an estimate of $f_p$ as well as the true $f_p$, which the econometrician would know under the assumption the data were generated in a symmetric equilibrium.\footnote{When ${\gamma} = {\theta} = 1/3$, the data could be generated by a three-bidder auction with $F_{v}(v) = v^{3/2}$. These simulated data might also be generated in a two-bidder auction in which bidder one's competitor's valuations are distributed according to $F_{v_2}(v)=(4 v/ 5)^3$. When only the highest competitor bid is used, our approach does not depend on which of these models is correct except when we consider that $f_p$ is known a priori in the former case but not in the latter.}\ The boundary-corrected kernel density estimate $f_p$ is obtained from the sample of $p_t = G_{cT}(b_{1t})$ using a bandwidth proportional to $T^{-1/5}$. Note, however, that the true density $f_p$ is unbounded near $p=0$ in a symmetric equilibrium with more than two bidders, which may result in poor performance of the density estimate for small values of $p>0$. Thus, even though the pointwise rate of convergence of our estimate of $f_p$ is faster than the rate of convergence of our estimate of ${\alpha}'$, we would expect this estimator to perform poorly in finite samples. Indeed, the simulation results indicate that the direct method combined with the true $f_p$ compares favorably with the indirect estimates, but the direct method combined with an estimate of $f_p$ can be relatively poor. In such cases, better results might be achieved by estimating the density of an appropriate transformation of $p$ and using a change of variable formula to recover an estimate of $f_p$. Alternatively, the minimum relative entropy estimator in (ref) could be used.
(ref) illustrates the simulation results for estimates of the inverse strategy function at the right boundary. The reflection--based boundary correction methods tend to perform better when the target of smoothing---$\alpha$ or $g_{c}$---is relatively flat and linear near the boundary. In this case, we would expect the error in the estimate of the auxiliary parameter $\hat d$ to be relatively small. On the other hand, the boundary kernel method tends to perform better when $\alpha'$ is large near the boundary, which would be more likely to happen if the number of bidders is small.
This paper reformulates the empirical analysis of auction models as an isotonic estimation problem by treating the probability of winning as the choice variable in the bidders' decision problem. The nonparametric least--squares and nonparametric maximum likelihood estimators for a bidder's inverse strategy function are shown to converge at the optimal nonparametric rate. As a complementary set of results, we prove the asymptotic behavior of two boundary correction methods that can be combined with transformation to better control the bias--variance tradeoff in the kernel smoothed versions of our estimators. While these smoothing methods are important when estimating some objects of potential interest to the researcher, smoothing is not necessary for others. We prove that using our unsmoothed estimator as an input to a simple plug--in estimator of parameters such as the bidder's expected surplus achieves the semiparametric efficiency bound.
Though the results in this paper can guide several important methodological choices in empirical research on auctions, our theorems are silent regarding several extensions to the baseline model that have become standard in the empirical auction literature. Namely, we do not address the possibility of affiliation among the bidders' valuations, unobserved auction--level heterogeneity, or risk aversion. We leave these considerations for future work.