EconBase
← Back to paper

Order Statistics Approaches to Unobserved Heterogeneity in Auctions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

76,831 characters · 14 sections · 57 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Order Statistics Approaches to Unobserved Heterogeneity in Auctions

\def\spacingset#1{ {#1}} \spacingset{1}

\if00 \fi

\if10 {

center[center omitted — 97 chars of source]

} \fi

abstractWe establish nonparametric identification of auction models with continuous and nonseparable unobserved heterogeneity using three consecutive order statistics of bids. We then propose sieve maximum likelihood estimators for the joint distribution of unobserved heterogeneity and the private value, as well as their conditional and marginal distributions. Lastly, we apply our methodology to a novel dataset from judicial auctions in China. Our estimates suggest substantial gains from accounting for unobserved heterogeneity when setting reserve prices. We propose a simple scheme that achieves nearly optimal revenue by using the appraisal value as the reserve price.

{\it Keywords:} Sieve Estimation, Nonseparable, Measurement Error, Consecutive Order Statistics, Judicial Auctions

\spacingset{1.8}

Introduction

comment\blue{In our Lemma (ref), we try to show that an operator $T: \mathcal{X} \longrightarrow \mathcal{X}$ is surjective given that its adjoint operator is injective. hu2008instrumental restricts $\mathcal{X}$ to be $\mathcal{L}^1$ in their Lemma 1. We, however, choose $\mathcal{X}$ to be $\mathcal{L}^2$, which is a Hilbert space. This treatment enables us to define the adjoint operator in more convenient way and take advantage of rich theories developed for operators defined over a Hilbert space, thus facilitating theoretical derivations. Note that in infinite-dimensional spaces, there exist operators which are injective but not surjective. For example, let $\{e_i\}_{i = 1}^{\infty}$ denote an orthonormal basis of the space $\mathcal{X}$ and the operator $A$ is defined as $A(e_i) = e_{i + 1}, i = 1, \ldots$. Obviously, $A$ is injective but not surjective. This feature dist inguishes operators defined over an infinite-dimensional space from linear transformations defined over a finite-dimensional linear space. } \blue{In terms of contributions, we proposed using nonparametric methods to estimate the joint distribution of latent variables and observed variables. As a result, we obtain nonparametric representations for both the conditional density of the observed variable and its marginal density. In particular, we consider two types of sieve space for estimation: B-spline functions and mixtures of Beta distributions. The complexity of these sieve spaces is characterized using the bracketing numbers. With the established results for these bracketing numbers, we employ theories of empirical process to show consistency of our proposed sieve MLEs. Therefore, we obtain theoretical guarantees for implementations with the two commonly used sieve spaces in nonparametric estimations of density functions in the context of auction.}

In empirical auction analysis, to estimate bidder value distributions, the analyst usually needs to pool data from auctions for similar but not identical items. However, the data available for these auctions often lack a precise description of the auctioned item. This situation results in auction-level unobserved heterogeneity (UH), which leads to inaccurate estimates of bidder value distributions and, therefore, misleading policy implications. For instance, HQT2019 finds that UH accounts for two-thirds of price variation after controlling for information provided in the eBay Motors auctions, and that ignoring this feature would dramatically mis-estimate the welfare measures. The existing literature adapts measurement error approaches to tackle such an issue. Suppose the analyst observes all bids. The analyst could then identify the value distribution using observed bids as measurements for the unobserved characteristics, since these bids are independent conditional on such unobserved characteristics.

However, the conditional independence condition fails when the analyst only observes incomplete bid data. This could occur for various reasons. First, in English or ascending outcry auctions, the bidder with the highest value only needs to outbid the bidder with the second-highest value to win, which means the recorded bids do not contain the highest value. Moreover, even in first-price sealed-bid auctions, where all bids are supposed to be submitted to the auctioneer, the auctioneer may still not record all the bids in practice: sometimes the auctioneer only records the most competitive bids, such as the top three bids in regular auctions or apparent low bids in procurement auctions. Thus, the econometrician can only observe a few order statistics of the bids, i.e., incomplete bid information. For instance, the U.S. Forest Service timber auctions only record at most the top 12 bids regardless of the number of bidders. The Washington State Department of Transportation provides an online archive of bid opening results that are six months or older, but only for the top three apparent low bids. Even if the auctioneer records all bids, the most competitive bids are often more accessible to the public. For instance, The Federal Deposit Insurance Corporation resolves insolvent banks using first-price auctions but only publishes the top two bids and bidders' identities allen2019resolving. The three apparent low bids are one-click downloadable on the website of the California Department of Transportation. These order statistics are naturally dependent, invalidating conventional identification strategies.

We make three contributions in this paper. First, our paper is the first to study identification of auction models with continuous and nonseparable UH using incomplete bid data. Our specification allows for flexibility in how UH affects both bidder value and the equilibrium bidding strategy, i.e., the mapping from a bidder's private value to his/her bid.\footnote{Even if one assumes separable UH in the value, separability passing to the bid often requires additional institutional features or assumptions. See, e.g., andreyanov2020secret.}

Our identification strategy adapts hu2008instrumental for nonclassical measurement error models to the auction setting. This extension is nontrivial in that we only observe order statistics of UH-contaminated bids. As a result, we cannot achieve a parsimonious conditional independence structure as in their work.\footnote{They assume that the outcome variable is independent of the observed independent variable and an instrument conditional on the unobserved true regressor.} Instead, we follow luo2020identification and consider the most common case of incomplete bid data: consecutive order statistics of bids. Their main insight is that consecutive order statistics have a semi-multiplicatively separable joint distribution with a simple indicator function capturing the correlation. Unlike both papers using two measurements with an instrument, we use three consecutive order statistics of bids. Given a partition on the range of the measurements, we again obtain a separable structure traditionally achieved under conditional independence. This turns the identification problem into an operator diagonalization problem, allowing constructive identification arguments using linear operator tools. Moreover, we use these tools differently by considering bounded linear operators defined on a Hilbert space and taking values in another Hilbert space. This space is smaller than the $\mathcal{L}^1$ space adopted in hu2008instrumental, which focuses on a Banach space. While we could also work with Banach space, using Hilbert space simplifies the analysis of relevant operators and thus our proofs thanks to many existing theoretical results.\footnote{For instance, it is straightforward to define the adjoint operator by using the concept of inner product in Hilbert spaces.}

Second, we propose sieve maximum likelihood estimators (MLE) of the model primitives and provide conditions that guarantee their consistency. The estimation of auction models allows for counterfactual policy analysis, such as computing the optimal reserve price. If UH is common knowledge among agents in the auction, it is a critical control in policy analysis. Therefore, optimal policy recommendation requires estimating the joint distribution of UH and bidder private value.\footnote{Since there is a known mapping between the bid distribution and the value distribution, we will use the two terms interchangeably. See GPV2000 and AH2002 for this mapping.} In particular, we approximate the joint density of bids and UH using the tensor product of two univariate sieve bases. We then represent the marginal density of the UH and the conditional distribution of the value using the sieve-approximated joint distribution. Therefore, these distributions are all estimated nonparametrically.\footnote{In contrast, previous research only focuses on the estimation of the joint distribution using a semiparametric structure chen2006efficient and hu2008instrumental or a nonparametric structure wu2012partially.} hu2008instrumental proposes sieve approximations to the conditional distribution and marginal distribution. Our sieve approximation to the joint distribution is more convenient as we just need to impose the normalization assumption on the joint distribution approximation once.

The consistency of our estimator relies on the condition that the sieve space approximates well the joint distribution of bids and UH. To formalize this intuition, we quantify the complexity of this space using bracket entropy and prove consistency of the sieve MLEs for the joint, conditional, and marginal densities. We establish a concentration inequality based on the bracketing number, a similar notation to covering numbers used in hu2008instrumental. The online supplement Section S.2.2 further investigates the properties of B-splines and Bernstein polynomials, both of which are popular in empirical applications.

Lastly, we apply our identification and estimation method to a novel dataset from judicial auctions conducted by a municipal court in China. By default, this court uses 70% of the appraisal value as the starting price, which also serves as a reserve price. Our estimation results suggest substantial gains from accounting for UH when designing reserve prices. The court can gain $5.81\%$ more revenue using an optimal reserve price for each item. However, this scheme is complex; the seller would need to know UH and recover the conditional density of bidder values. Instead, we propose a simple scheme that achieves nearly optimal revenue by using the appraisal value as the reserve price. Specifically, using the estimated model, we find that using the appraisal value as the reserve price achieves $98.85\%$ of the potential gains from the optimal reserve prices.

Literature Review

The auction literature has widely applied techniques developed in the measurement error literature for identifying auction models with UH. If the UH is continuous and has a separable structure on bidder valuations, identification relies on the deconvolution approach and requires two random bids for each auction. See LV1998, li2000conditionally, and K2011, among others. If the UH is finite and discrete, which by nature is nonseparable, identification relies on the condition that the bids are independent conditional on the UH and requires three random bids for each auction. See H2008, HMS2013, and Luo2018.

Moreover, the literature has seen rapid growth in identifying and estimating auctions models using order statistics of bids. AH2002 shows that symmetric independent private value (IPV) auctions are identifiable by the transaction price and the number of bidders using the one-to-one mapping between the distribution of an order statistic and its parent distribution; komarova2013new identifies asymmetric second-price auctions using the winner's identity and the transaction price; GL2018 shows that IPV first-price auctions without observable competition is identifiable using the transaction price; menzel studies large sample properties for nonparametric estimators using order statistics of bids.

A growing literature tackles the identification of auction models with UH and incomplete bid information. Assuming the UH is finite and discrete, M2017 provides identification results from (any) five order statistics to restore the conditional independence condition by the Markov property of order statistics. luo2020identification provides an alternative identification strategy using two consecutive order statistics of bids and an instrument. Finiteness simplifies their identification arguments because model restrictions can be written in matrix algebra. In contrast, we use linear operators, which is not a trivial extension of the matrix operations. Moreover, we extend our identification results to allow for binding reserve prices and apply them in our empirical application.

In the framework of additively separable continuous UH, HQT2019 achieves point identification using English auction models, assuming piecewise real analytic density functions and using variations in the number of bidders across auctions; FL2017 provides identification results for ascending auctions, relying on reserve prices and two order statistics of bids. CLX2022 studies deconvolution using two order statistics. Our paper is the first to show point identification of auction models with continuous and nonseperable UH using incomplete bid data.

The remainder of this paper is organized as follows. Section (ref) presents our main identification results. Section (ref) proposes sieve maximum likelihood estimators. Section (ref) presents an application to judicial auctions in China. Section (ref) concludes. The online supplement contains detailed proofs for the identification results and the asymptotic properties and the finite-sample properties of the proposed estimators.

Main Identification Results

Consider a standard IPV auction model for items with a scalar heterogeneous characteristic $\mathtt{T}$ that is observable to all bidders but unobserved to the analyst. For simplicity, we abstract from observable (to the analyst) characteristics. Suppose $n \geq 2$ symmetric bidders participate in an auction with zero reserve price.\footnote{We assume the number of potential bidders is known. Otherwise, we can treat it as an additional dimension of UH, as in luo2020identification, or construct it through alternative data sources. In procurement auctions, we can construct it using the number of qualified firms in the local market via public information, such as the list of qualified firms and their contact information.} All bidders observe the characteristics $\mathtt{T}$ before they submit bids.\footnote{While we focus on regular auctions here, our results extend trivially to procurement auctions.} Our identification strategy applies regardless of whether the seller observes $\mathtt{T}$ or not. Among $n$ potential bidders, bidder $i$, where $i=1,...,n$, draws his/her value $V_i$ from the conditional value distribution $f^{V|\mathtt{T}}(v|\tau)$ and submits a bid $X_i$. We consider the situation wherein the latent auction characteristic and bids/values are continuous. We denote the marginal distribution of the latent characteristic $\mathtt{T}$ as $f^{\mathtt{T}}(\tau)$ and the optimal conditional bid distribution as $f^{X|\mathtt{T}}(x|\tau)$, where $x$ is the optimal bid.

We first introduce the standard assumption regarding the value distribution.

ass(Conditional Independence) Bidder values, $V_1$,..., $V_n$, are i.i.d. conditional on the auction-level heterogeneity $\mathtt{T}$.

In a first-price auction, the bidder with the highest bid wins and pays his own bid price. GPV2000 provides a one-to-one mapping between the conditional value distribution $f^{V|\mathtt{T}}(v|\tau)$ and the conditional bid distribution $f^{X|\mathtt{T}}(x|\tau)$ given that the competition $n$ is known. Thus, the identification of the conditional value distribution boils down to recover the conditional bid distribution $f^{X|\mathtt{T}}(x|\tau)$ from the bid data. If the data record all bids in each auction, the conditional independence property passes from values to bids. Consequently, the joint distribution of three independent bids, e.g., $X_1$, $X_2$, and $X_3$, denoted as $f(x, y, z)$, has the following multiplicatively separable structure:

eqnarray[eqnarray omitted — 222 chars of source]

based on which the conditional densities $f^{X|\mathtt{T}}(x|\tau)$ can be identified via eigenfunction decomposition hu2008instrumental. The main idea is to exploit that the bids are repeated measurements of UH. Under Assumption (ref), their correlation reveals how UH affects the bids. Specifically, the observed joint distribution on the left-hand side of (ref) identifies the conditional and marginal distributions on the right-hand side.

Unfortunately, the auctioneer often does not record all bid information, and instead only records the most competitive bids. That is, the data essentially record a few order statistics of all bids, under which the conditional independence condition fails to hold. This is because order statistics are ordered by definition.

In an ascending auction, the bidder with the highest bid wins and pays the second highest submitted price, so a weakly dominant strategy is to continue bidding until the standing bid reaches one's own value. Therefore, all bidders bid their own values except the one with the highest value, who can simply outbid the second highest value by a small amount. That is, the highest bid and the second highest bid reveal essentially the same information regarding the second highest value, indicating that the highest bid is redundant. Because of this particular auction format, it is impossible to observe the highest value from the bids. Equivalently, we can view the auction as every one bids her/his value, but the auction fails to observe the highest bid/value. Consequently, one cannot follow the aforementioned identification results to recover the conditional value distribution $f^{V|\mathtt{T}}(v|\tau)$, because the conditional independence condition fails.

Facing the data limitation of incomplete bids, this paper focuses on identifying the conditional bid distribution $f^{X|\mathtt{T}}(x|\tau)$ for both first-price and ascending auctions from any three consecutive order statistics of all bids, i.e., $\{X_{r-2:n}, X_{r-1:n}, X_{r:n}\}$, where $X_{r-2:n}\le X_{r-1:n}\le X_{r:n}$. Once the conditional bid distribution is identified, the conditional value distribution can be identified using the one-to-one mapping between the bid and the value.

Let $\mathcal{V}$, $\mathcal{X}$, and $\mathcal{T}$ denote the supports of the distributions of the random variables $V$, $X$, and $\mathtt{T}$, respectively. We first introduce the following regularity assumption.

ass(Bound and Continuity) The joint density of $X$ and $\mathtt{T}$ admits a bounded and continuous density with respect to the product measure of some dominating measure $\mu$ (defined on $\mathcal{X}$) and the Lebesgue measure on $\mathcal{T}$. All marginal and conditional densities are also bounded and positive.

We use $f_{r-2,r-1,r:n}(\cdot)$ and $f_{r-2,r-1,r:n}(\cdot|\tau)$ ($r \geq 3$) to represent the unconditional and conditional joint probability density functions (PDF) of the three order statistics, respectively, and $f^{X}_{r:s}(\cdot)$ and $f^{X \mid \mathtt{T}}_{r:s}(\cdot|\tau)$ represent the unconditional and conditional PDF of the $r$th order statistic of measurements ${X}$ out of a sample of size $s$ ($r \leq s$).

The identification exploits the fact that the conditional joint distribution of three consecutive order statistics has a multiplicative separable structure. Specifically, the unconditional joint distribution, which can be estimated from the data, can be expressed as

align[align omitted — 421 chars of source]

where $c_{r,n}=\frac{n!}{(r-2)! \cdot (n-r+1)!}$, and $\mathbbm{1}(\cdot)$ is the indicator function. The first equality holds by the law of total probability, and the second extends luo2020identification's Lemma 1 to three consecutive order statistics.\footnote{The joint distribution of any three order statistics does not have such a multiplicatively separable structure, i.e., $f_{r,s,t:n}(x,y,z) \sim f(x)f(y)f(z) [F(x)]^{r-1} [F(y)-F(x)]^{s-r-1}[F(z)-F(y)]^{t-s-1}[1-F(z)]^{n-t}$, where $r<s<t$, see david2004order. We derive Equation (ref) in the online supplement Section S.1.1.} This joint distribution of the consecutive order statistics has a semi-separable structure in the sense that we can separate the observed joint density function into the integration of three density functions, which is similar to (ref) in the measurement error literature, but it has an extra restriction by the nature of order statistics, $\mathbbm{1}(x \le y \le z)$, which cannot be separated. This semi-separable structure precludes us from readily borrowing the same identification procedure in the existing literature to identify the conditional latent distributions directly.

Fortunately, the restriction by the indicator function can be safely circumvented if we divide the original support by two cutoff points $c_1$ and $c_2$, where $c_1 < c_2$, to separate the support into three parts, referred to as “low," “middle," and “high," and denote them as $\mathcal{X}_l \equiv \{x: x \le c_1\}, \mathcal{X}_m \equiv[c_1, c_2]$, and $\mathcal{X}_h \equiv \{x: x \ge c_2\}$, respectively. Our context of three order statistics calls for three-part discretization, which extends luo2020identification's two-part discretization using two order statistics and an IV. The separable structure of the joint distribution $f_{r-2,r-1,r:n}(x, y,z)$ reappears if we always restrict $x \in \mathcal{X}_l$, $y\in \mathcal{X}_m$, and $z\in \mathcal{X}_h$. Specifically, if $x \in \mathcal{X}_l $, $y \in \mathcal{X}_m$, and $z \in \mathcal{X}_h$, the joint distribution can be expressed as

eqnarray[eqnarray omitted — 220 chars of source]

which has the same structure as the measurement error models but a different conceptual interpretation for each component. Figure (ref) provides a visualization of the discretization.

figure[figure omitted — 1,198 chars of source]

Following the identification strategy developed in hu2008instrumental, we introduce the following integral operator that associates a function of two variables.

definitionLet $L_{x|\tau}$ denote an operator that maps function $g$, where $g \in \mathcal{G}(\mathcal{T})$, to $L_{x|\tau}g \in \mathcal{G}(\mathcal{X}_l)$; and $H_{x|\tau}$ maps function $g$, where $g \in \mathcal{G}(\mathcal{X}_h) $, to $H_{x|\tau}g \in \mathcal{G}(\mathcal{T})$. Specifically, the two operators are defined as $$[L_{x|\tau}g](x) \equiv \int_{\mathcal{T}} f^{X \mid \mathtt{T} }(x|\tau) g(\tau) d\tau ~\quad\mbox{and}\quad~ [H_{x|\tau}g](\tau) \equiv \int_{\mathcal{X}_h} f^{X \mid \mathtt{T} }(x|\tau) g(x) dx. $$

Note that both operators involve a segment of bid support $\mathcal{X}$. We further introduce another linear operator based on the joint distribution and the diagonal operator defined as follows. In particular, for a given $y \in \mathcal{X}_m$, let $J_y$ denote an operator mapping $g \in \mathcal{G} (\mathcal{X}_h)$ to $J_{y} g \in \mathcal{G}(\mathcal{X}_l)$:

equation*[equation* omitted — 90 chars of source]

Given a particular partition $\{\mathcal{X}_l, \mathcal{X}_m, \mathcal{X}_h\}$, $J_y$ is defined for every given $y$ in $\mathcal{X}_m$. Let $\Delta_{X=y,\mathtt{T}}$ denote the diagonal operator mapping $g\in \mathcal{G}(\mathcal{T})$ to $\Delta_{X=y,\mathtt{T}}g \in \mathcal{G}(\mathcal{T})$:

equation*[equation* omitted — 122 chars of source]

We derive the equivalence of operators in the online supplement Section S.1.2 as follows:

eqnarray[eqnarray omitted — 119 chars of source]

based on Equation (ref) and by exploiting the following features: (i) an interchange of the order of integrations (justified by Fubini's theorem), (ii) the definition of $H_{X_{1:n-r+1}|\mathtt{T}}$, (iii) the definition of $\Delta_{X=y,\mathtt{T}}$ operating on $H_{X_{1:n-r+1}|\mathtt{T}}g$, and (iv) the definition of $L_{X_{r-2:r-2}|\mathtt{T}}$ operating on $[\Delta_{X=y,\mathtt{T}}H_{X_{1:n-r+1}|\mathtt{T}}g]$. Note that such equivalence between the operators holds for any value of $y \in \mathcal{X}_m$.

For identification, we impose the following injective assumption.

ass(Injective) There exists one division of the domain such that the operators $L_{\mathtt{T}|X_{r-2:r-2}}$ and $H_{X_{1:n-r+1}|\mathtt{T}}$ are injective for $\mathcal{G}=\mathcal{L}^2$, where $\mathcal{L}^2(\mathcal{X})$ denotes the set of all square integrable functions with domain $\mathcal{T}$ and $\mathcal{X}_h$, respectively.

An operator $A$ is injective if $Af = Ag$ implies $f = g$ for any $f, g$ in the domain of $A$. A linear operator being injective is equivalent to the family of kernel functions used to define the operator being complete; see hu2008instrumental. In our context, if the family of distributions $\{f^{X|\mathtt{T}}_{r-2:r-2}(x| \tau): x \in \mathcal{X}_l \}$ is complete over $\mathcal{L}^2(\mathcal{T})$, that is, the unique solution $\tilde{g}$ to the equation $\int_{\mathcal{T}} g(\tau)f^{X|\mathtt{T}}_{r-2:r-2}(x| \tau) d \tau = 0$ for all $x \in \mathcal{X}_l$ is $\tilde{g}(\cdot) = 0$, then $L_{\mathtt{T}|X_{r-2:r-2}}$ is injective under Assumption (ref). We further provide conditions on the parental distributions under which the family of the order statistics' distributions is complete in the online supplement Section S.1.3. However, the equivalence between the injectiveness of operator $H_{X_{1:n-r+1}|\mathtt{T}}$ and the completeness of the kernel function family $\{f^{X|\mathtt{T}}_{1:n-r+1}(x| \tau):\tau \in \mathcal{T} \}$ over $\mathcal{L}^2(\mathcal{X}_h)$ is not straightforward, because the operator is defined only in a segment of the support. We prove that as long as the original distribution family is complete, i.e., $\{f^{X|\mathtt{T}}_{1:n-r+1}(x| \tau):\tau \in \mathcal{T} \}$ over $\mathcal{L}^2(\mathcal{X})$ is complete, there exists at least one division of the support such that operator $H_{X_{1:n-r+1}|\mathtt{T}}$ is injective. See the online supplement Section S.1.3.

Completeness of the relevant family of distributions provides one way to characterize the injectivity of an operator. Intuitively, the family of distributions $\{f^{X | \mathtt{T}} (x |\tau): x \in \mathcal{X} \}$ being complete implies there is {\it sufficient variation} in the conditional density of $X$ across different values of $\mathtt{T}$. An example for such a complete distribution is a normal distribution with mean $\tau$ and variance 1. On the other hand, if the conditional density of $X$ does not vary sufficiently across $\tau$, such as the standard normal distribution, the distribution family is not complete. Obviously in such a scenario, $X$ is independent of $\mathtt{T}$, and hence we can easily find $g \neq 0$ such that $\int g(\tau) f^{X | \mathtt{T}}(x | \tau) d \tau = 0$ for any $x$.

Assumption (ref) also specifies that we consider the identification with $\mathcal{G} =\mathcal{L}^2$. Such consideration is due to the following two reasons. First, this space is sufficiently large such that the density can be sampled everywhere, which ensures a one-to-one mapping between a density function and its corresponding operator. Thus, the density function can be uniquely determined by the associated operator with such a choice of $\mathcal{G}$.\footnote{The space $\mathcal{G}=\mathcal{L}^2$ is sufficiently rich, because $ f^{X}_{r-2:r-2}(x|\tau_0) = \lim\limits_{n \rightarrow \infty} [L_{X_{r-2:r-2}|\mathtt{T}}g_{n, \tau_0}](x),$ where $g_{n, \tau_0}(\tau) = n\mathbbm{1}(|\tau - \tau_0| \leq n^{-1} )$, a sequence of bounded and square-integrable functions. } Second, it is a Hilbert space if equipped with the norm $\|g\|_{\mathcal{L}^2} = \left(\int_{\mathcal{X}} g^2(x) dx\right)^{1/2}$ for any $g \in \mathcal{G}(\mathcal{X})$. One advantage of considering Hilbert spaces is that it is easier to use properties of the operators such as $L_{X_{r-2:r-2}|\mathtt{T}}$ and $H_{X_{1:n-r+1}|\mathtt{T}}$ later, because there are many existing theoretical results developed for operators defined in Hilbert spaces. For instance, it is straightforward to define the adjoint operator by using the concept of inner product in Hilbert spaces. It is also worth noting that this space is smaller than the $\mathcal{L}^1$ space adopted in hu2008instrumental, which is a Banach space.

If an operator is injective, its inverse is well-defined, but may be defined over a restricted domain. We further prove that $L_{X_{r-2:r-2}|\mathtt{T}}$ is surjective in addition to being injective, so that the domain of its inverse is the whole space $\mathcal{L}^2(\mathcal{X})$. This is important for proving the equivalence of operators defined in the data and in the distributions to be identified. We summarize this result in the following lemma and relegate the proof to the online supplement Section S.1.4.

lemmaIf Assumptions (ref)-(ref) hold, then $L^{-1}_{X_{r-2:r-2}|\mathtt{T}}$ exists and is densely defined over $\mathcal{L}^2(\mathcal{X}_l)$.

Lemma (ref) essentially indicates that operator $L_{X_{r-2:r-2}|\mathtt{T}}$ is surjective if it is injective. We use the following simple example to facilitate understanding the necessity of the surjective property and the difference between linear operators and matrices. Suppose that $\mathcal{D}^1$ and $\mathcal{D}^2$ are two linear spaces, and $L$ is a linear transformation from $\mathcal{D}^1$ to $\mathcal{D}^2$. If both $\mathcal{D}^1$ and $\mathcal{D}^2$ are finite-dimensional, $L$ is injective if and only if it is surjective. In particular, if dim($\mathcal{D}^1$) = dim($\mathcal{D}^2$) and $L$ is associated with a square matrix $A$, then $L$ is both injective and surjective if and only if $A$ has full rank. But this relationship does not trivially hold in infinite-dimensional cases. For example, let $\{e_i\}^\infty_{i=1}$ be the basis of $\mathcal{D}^1$ as well as $\mathcal{D}^2$. We assume that $Le_i = e_{i+1}$ for every $i \ge 1$. Such an operator $L$ is obviously injective but not surjective, {because the base $e_1$ is missing in its range.}

Since $H_{X_{1:n-r+1}|\mathtt{T}}$ is injective under Assumption (ref), we can eliminate the common operator $H_{X_{1:n-r+1}|\mathtt{T}}$ by equivalence of operators specified in Equation (ref) for any two different values of $y$, i.e., $y_1$ and $y_2$, leading to the following main equation for identification:

eqnarray[eqnarray omitted — 173 chars of source]
commentThe surjective property is important to derive Equation (ref).\textcolor{red}{[RX: does this mean $H_{X_{1:n-r+1}|\mathtt{T}}$ has to be surjective?]} \blue{Using the same techniques as in proving Lemma 1, we can show this operator [RX: this means L only, or this means H and L both?] is both injective and surjective.} If $J_{y_2}$ is not surjective, then the domain of $J^{-1}_{y_2}$ is smaller than $\mathcal{L}^2(\mathcal{X}_m)$,[\red{should it be $\mathcal{L}^2(\mathcal{X}_l)$?}] which is vital for establishing the equivalent relationship between the operators on both sides of the equation.\footnote{Note that operator $J_{y_2}$ is injective, guaranteed by the injection of operators $L_{X_{r-2:r-2}|\mathtt{T}}$ and $H_{X_{1:n-r+1}|\mathtt{T}}$.}

By Lemma (ref), the relation (ref) is established over a dense subset of $\mathcal L ^2(\mathcal X_l)$. In fact, it can be further extended to the full space $\mathcal L ^2(\mathcal X_l)$ by leveraging the extension procedure of linear operators. This equation ensures that operator $J_{y_1} J^{-1}_{y_2}$ can be represented as an eigenvalue-eigenfunction decomposition with the two unknown operators $L_{X_{r-2:r-2}|\mathtt{T}}$ and $\Delta_{X=y_1,\mathtt{T}}\Delta^{-1}_{X=y_2,\mathtt{T}} $ being the eigenfunctions and eigenvalues, respectively. Consequently, diagonalizing operator $J_{y_1} J^{-1}_{y_2}$, which can be computed from the data directly since it is defined using observable densities, provides the eigenfunctions $L_{X_{r-2:r-2}|\mathtt{T}}$, indexed by the latent UH, and further provides the unobserved densities of order statistic $X_{r-2:r-2}|\mathtt{T}$.

Note that there are three features prevalent in identification using decomposition: The identification may not be unique; the identification is up to scales; the identification is up to ordering and location. We tackle the three issues one at a time below.

Unique Decomposition

To guarantee unique decomposition, we impose restrictions on the relationship between observed measurement $X$ and UH $\mathtt{T}$ in segment $\mathcal{X}_m$.

ass(Distinct) there exists one division of the domain such that, for all $\tau_1, \tau_2 \in \mathcal{T}$, the set $\{(y_1,y_2): \frac{f^{X|\mathtt{T}}(y_1|\tau_1)}{f^{X|\mathtt{T}}(y_2|\tau_1)} \neq \frac{f^{X|\mathtt{T}}(y_1|\tau_2)}{f^{X|\mathtt{T}}(y_2|\tau_2)}, where ~ (y_1,y_2) \in \mathcal{X}_m \times \mathcal{X}_m\}$ has positive probability whenever $\tau_1 \neq \tau_2$.

This assumption is weaker than assuming that the associated operator is injective in segment $\mathcal{X}_m$. Note that we just need one division where such an assumption holds. This assumption fails only if the distribution of the measurement conditional on the latent factor is the same at the two distinct values $\tau_1$ and $\tau_2$.

Assumption (ref) guarantees unique eigenvalues, so that conducting the decomposition to operator $J_{y_1} J^{-1}_{y_2}$ identifies operator $L_{X_{r-2:r-2}|\mathtt{T}}$, and thus identifies the conditional density $f^{X|\mathtt{T}}_{r-2:r-2}(x|\tau)$, for $x \in \mathcal{X}_l$. However, such identification is up to scales. That is, the conditional density $f^{X|\mathtt{T}}_{r-2:r-2}(x|\tau)$ is identified as the true density multiplied by an unknown constant, which could differ for each UH. The existing literature relies on the property that the total probability is equal to 1 for each conditional distribution to pin down the scales. Such an approach is not feasible in our framework because, from the decomposition, we only identify the conditional distribution in one segment of the full support, i.e., $\mathcal{X}_l$. Mover, one can neither pin down the ordering or the actual values of UH, which calls for extra restrictions. To proceed, we propose to leave the ordering of the UH and the scales in the low segment as undetermined and proceed to identify the conditional distributions in the other two segments first. In this procedure we mainly use Equation (ref). One main feature worth noting during this process is that we keep the value of the UH consistently matched across the three segments. Furthermore, these scales are the same for the same UH in the same segment but may vary across UH or segments. Given these, we can then pin down the scales and ordering in what follows.

Unique Scale

Note that we can identify the conditional distributions in all three segments up to different scales. That is, each segment of the conditional distribution is associated with one scale parameter, so together there are three scale parameters to pin down for each conditional distribution. These scales can then be pinned down by invoking the continuity of the component PDFs and the total probability argument. First, the PDFs identified separately in the three segments should be the same at the cutoff points due to the continuity of the true conditional distributions. Second, the fact that each conditional distribution should integrate to 1 provides the third restriction on the scales. These restrictions uniquely identify the scales.

Unique Ordering and Location

Given that the conditional distributions are identified in the full support, we provide a condition using the auction setting to pin down the exact location of the UH. Specifically, letting UH be the unobserved quality of the auctioned item, we would expect that bidders' values/bids are, on average, higher and of better quality. For instance, in second-hand automobile auctions, omitted details from the car description, such as dents and scratches, are revealed upon pre-auction inspection and enter bidder values.

ass(Monotonicity and Location) The expected value/bid is strictly monotone with UH; that is, $E(X|\mathtt{T}=\tau)$ is strictly monotone with $\tau$ for all $\tau \in \mathcal{T}$. Moreover, we assume that the support of UH is [0, 1].

The monotonicity assumption is useful to pin down UH's relative ordering. However, its exact location/value is still unidentified. That is, one could always apply a monotone transformation to the UH and obtain an observationally equivalent model that satisfies all assumptions. To pin down UH's exact location, we normalize its support to be $[0,1]$, which is without loss of generality. Such a normalization is similar to the mean zero normalization.

theoremIf Assumptions (ref)-(ref) are satisfied, conditional bid distribution $f^{X|\mathtt{T}}(x|\tau)$ for $x\in \mathcal{X}$ and $\tau \in \mathcal{T}$ and UH's distribution $f^{\mathtt{T}}(\tau)$ for any $\tau \in \mathcal{T}$ are identified using any three consecutive order statistics of bids.

We summarize the main steps of the proofs below and leave the details to the online supplement Section S.1.6.\footnote{We thank Yingyao Hu and Ji-Liang Shiu for valuable insights about proving the theorem.} First, we identify operator $L_{X_{r-2:r-2}|\mathtt{T}}$ from the decomposition of Equation ((ref)). Such identification is unique by Assumption (ref), but up to scales and location. Second, we identify the operator $H_{X_{1:n-r+1}|\mathtt{T}}$ up to different scales, similar to the identification of $L_{X_{r-2:r-2}|\mathtt{T}}$. Third, for any value $y\in \mathcal{X}_m$, we can identify operator $\Delta_{X=y,\mathtt{T}}$ up to the same scales for all $y$ once we plug the identified operators $L_{X_{r-2:r-2}|\mathtt{T}}$ and $H_{X_{1:n-r+1}|\mathtt{T}}$ into Equation (ref). Using the one-to-one mapping between operators and the associated densities, we then identify the unobserved densities $f^{X|\mathtt{T}}_{r-2:r-2}(x|\tau)$ for $x \in \mathcal{X}_l$, $f^{X|\mathtt{T}}(y|\tau) f^{\mathtt{T}}(\tau)$ for $y\in \mathcal{X}_m$, and $f^{X|\mathtt{T}}_{1:n-r+1}(z|\tau)$ for $z \in \mathcal{X}_h$ up to scales. The scales are the same in the same segment but may vary across different segments. Furthermore, we show that the one-to-one mapping between the distribution of an order statistic and its parent distribution can be extended from the full support to a segment. Thus, we identify the conditional distribution up to different scales in all three segments. Lastly, the scales are then pinned down using three restrictions.

Once the conditional bid distributions are identified as in Theorem (ref), we can exploit the one-to-one mapping between the conditional value and bid distributions to recover the conditional value distributions, which are the target of interest. Specifically, for ascending auctions, where bidders' weakly dominant strategy is to bid their values, the conditional value distribution is the same as the conditional bid distribution;\footnote{Many empirical studies adopt the same assumption in ascending auctions; see, e.g., lu2008estimating, aradillas2013identification, and hortaccsu2021empirical. We exclude other possible bidding strategies such as jump bidding allowed in haile2003inference. Such abstraction is a good approximation for online auctions and button auctions. For instance, eBay allows bidders to set up a proxy bid.} for first-price auctions, we can identify the conditional value distribution by exploiting the one-to-one mapping established in GPV2000. We summarize this result in the following Corollary.

corollaryIf Assumptions (ref)-(ref) are satisfied, the conditional value distribution \\ $f^{v|\mathtt{T}}(v|\tau)$ for $v \in \mathcal{V}$ and $\tau \in \mathcal{T}$ and the latent variable's distribution $f^{\mathtt{T}}(\tau)$ for $\tau \in \mathcal{T}$ are identified using any three consecutive order statistics of bids.

The identification results in Theorem (ref) are achieved under the assumption that the reserve price is not binding. However, in practice, the reserve price appears to be binding in many cases, leading to a truncation in the observed bid distribution. We show in the following corollary that we can still identify the bid/value distribution with a truncation. We can also identify the conditional probability of the truncation when the number of potential bidders is observed.

Reserve Price for Ascending Auctions

If the reserve price is binding, the optimal bidding strategy for any bidder is to submit the optimal bid computed without reserve prices when such an optimal bid is above the reserve price, and to not bid otherwise. Therefore, the presence of a binding reserve price $R$ creates a truncation in the observed bid distribution, i.e., $\tilde F^{X|\mathtt{T}}(x|\tau) \equiv \frac{F^{X|\mathtt{T}}(x|\tau)-F^{X|\mathtt{T}}(R|\tau)}{1-F^{X|\mathtt{T}}(R|\tau)}$, where $ x\in [R,\overline{x}]$. Let $n$ denote the number of actual bidders and $N$ denote the number of potential bidders. In first-price auctions, even if entry is exogenous, the observed bid distribution depends on both $N$ and $n$, while in ascending auctions, it only depends on $n$. Therefore, to illustrate the intuition, we focus on ascending auctions.

Under such a situation, even with a truncation caused by a binding reserve price, we can still follow the identification strategy in Theorem (ref) to identify the truncated CDF $\tilde F^{X|\mathtt{T}}(x|\tau)$, PDF $\tilde f^{X|\mathtt{T}}(x|\tau)$, and the marginal distribution of the UH without information on $N$ as long as $n$ is known. Specifically, the joint distribution of three consecutive active bids with a bidding reserve price can be expressed as

eqnarray*[eqnarray* omitted — 272 chars of source]

A few features are worth noticing. First, identification using eigen-decomposition applies regardless of whether $N$ is observed, as the bidding strategy does not vary with $N$ under exogenous entry. Second, without observing bids below the reserve price, there is no information to identify the bid/value distribution for this segment. Lastly, we establish that we can identify the conditional probability of the truncation $F^{X|\mathtt{T}}(R|\tau)$.

corollaryIn ascending auctions, when $N$ is observed and has a large support, the conditional probability of truncation $F^{X|\mathtt{T}}(R|\tau)$ is identified using the distribution of the number of actual bidders conditional on the potential bidders. Therefore, for all $x\geq R$, $F^{X|\mathtt{T}}(x|\tau)$ is identified from $\tilde F^{X|\mathtt{T}}(x|\tau) \equiv \frac{F^{X|\mathtt{T}}(x|\tau)-F^{X|\mathtt{T}}(R|\tau)}{1-F^{X|\mathtt{T}}(R|\tau)}$.

The detailed proof for Corollary (ref) can be found in the online supplement Section S.1.7. Intuitively, the distribution of $n$ conditional on $N$ is a mixture of binomial distributions with the success probability being the conditional truncated probability. That is,

eqnarray[eqnarray omitted — 151 chars of source]

where $\Pr(n|N)$ is estimable from the data, $C_{N,n}$ is a constant, $F^{\mathtt{T}}(\tau)$ can be treated as known, and conditional truncation probability $F^{X|\mathtt{T}}(R|\tau)$ is the object of interest. This is similar in structure to but differs conceptually from the identification in the mixture literature gut2005probability, where the goal is to identify the mixture distribution with the success probability taking any value in $[0,1]$. We show that our identification problem can be viewed as the dual problem of that by changing variables in the integral.

The Number of Order Statistics

Our discussion so far assumes that three consecutive order statistics of bids are available. There are various ways to extend this main identification result. First, the required number of consecutive order statistics reduces to two if there exists an instrument that is independent of the bids conditional on UH; see luo2020identification.\footnote{Measurement error approaches are inapplicable when only one order statistic, such as the winning bid, is observed. This calls for alternative strategies, such as density discontinuity approaches first proposed by GL2018.} Second, while consecutiveness barely restricts the data with incomplete bids, exploiting the Markov property of order statistics relaxes this requirement. In the online supplement Section S.3, we show that any four order statistics identify the model.\footnote{The idea of using Markov property for dealing with UH and incomplete bid data simultaneously is first explored in M2017, who uses five order statistics in finite UH framework.}

Sieve Maximum Likelihood Estimation

Note that conducting counterfactual policy analysis requires one to estimate the joint distribution of UH and bidder private values. In principal, the conditional bid distribution and UH's marginal distribution could be estimated fully nonparametrically by following the constructive identification argument step-by-step. Specifically, one could do a partition in the full support and conduct eigenfunction decomposition to estimate the distribution of the order statistics in the three segments, then use the one-to-one mapping between the distribution of an order statistic and its parent distribution to estimate the parent distribution. Such a fully nonparametric estimator not only poses a high demand on the data but is also of low efficiency, as it depends critically on the partition of the support and involves sequential estimation.

Considering the fact that, in applications, the analyst oftentimes can only access modest-sized data, we propose to estimate these two densities using the method of sieves grenander1981abstract, shen1997methods, chen1998sieve, chen2007large to fully exploit variations in the data instead of relying on a particular partition. We establish consistency and convergence rates for such estimators.

Our strategy is to first provide some regularity assumptions on the sieve approximation for consistency, which usually depends on the smoothness of the function to be approximated and the complexity of the sieve space. Such complexity is characterized by its upper bound and bracketing numbers.\footnote{In contrast, hu2008instrumental uses a covering number to characterize complexity.} To further understand the scope of our general results, the online supplement Section S.2.2 proves that the sieve space constructed by either B-spline or Bernstein basis functions, which are popular sieve spaces in auctions, satisfies the regularity assumptions, and thus, the estimator is consistent.

We represent the log likelihood function of the joint distribution of the three consecutive order statistics, i.e., $\text{data}\equiv\{X_{r-2:n}=x^i, X_{r-1:n}=y^i, X_{r:n}=z^i\}^m_{i=1}$, as follows:

eqnarray[eqnarray omitted — 330 chars of source]

As both the conditional density and the marginal density can be derived from a joint density, we propose to approximate joint distribution $f^{X, \mathtt{T}}(x,\tau)$ by using tensor product bases of univariate series. Specifically, let $\mathcal{B}_m$ be the finite-dimensional sieve space and $\xi_1, \ldots, \xi_{p_m}$ be its basis, where $p_m$ is the number of basis functions in the sieve space.

With slight abuse of notation, we denote the sieve representation of this joint distribution as $\mathfrak{f}$. We then represent the marginal distribution, the conditional distribution, and CDF of such a conditional distribution as follows:

eqnarray[eqnarray omitted — 523 chars of source]

Consequently, the sieve estimator for the joint distribution of the three observed consecutive bids can be represented as

eqnarray[eqnarray omitted — 263 chars of source]

Next, we show that under some regularity conditions the proposed sieve estimator for the joint distribution in Equation (ref) is consistent. Once the joint distribution is consistently estimated, the conditional and marginal distributions, specified in Equations (ref) and (ref) respectively, are also consistently estimated. Let $ f_0^{X, \mathtt{T}}(x, \tau)$ denote the true joint density, and let $ f_0^{X | \mathtt{T}}(x | \tau)$ and $f_0^{\mathtt{T}}(\tau)$ denote the true conditional density of $X$ given $\mathtt{T} = \tau$ and the marginal density of the latent variable, respectively. We introduce some regularity conditions.

ass(Compactness) $X$ has a compact support. Without loss of generality, we assume that its support is [0, 1].

This compact support assumption is standard in the auction literature. Moreover, we can linearly transform random variables with compact support to ones that have support on $[0,1]$. Note that such a transformation has to be linear, rather than an arbitrary monotone transformation. The linear transformation is for convenience of using the observed data in estimation. The support of the two random variables, $X$ and $\mathtt{T}$, plays an important role in choosing an appropriate sieve space to perform maximum likelihood estimation. For example, the trigonometric sieve is inapplicable when the support is $\mathbb{R}$. In this case, Hermite polynomials and B-splines are preferable. B-spline approximation is also useful when the support is compact. It is worth emphasizing that our identification results hold regardless of this normalization.

ass(Sieve approximation) There exists $f_m^{X, \mathtt{T}}(x , \tau)$, which is represented in terms of the bases $\xi_1, \ldots, \xi_{p_m}$ in the sieve space, for some $\beta > 0$, such that \begin{align*} \|f_m ^{X, \mathtt{T}}(x, \tau) - f_0^{X, \mathtt{T}} (x, \tau)\|_{L_{\infty}([0, 1]^2)} = O(p_m^{-\beta}). \end{align*}

Assumption (ref) ensures that the joint density can be approximated sufficiently well in the sieve space. Consequently, by Equations (ref) and (ref), both the conditional density and the marginal density can be approximately sufficiently well by functions in the sieve space. That is, with Assumption (ref), there exist $f_m^{X | \mathtt{T}}(x | \tau)$ and $f_m^{\mathtt{T}}(\tau)$, both represented in terms of $\xi_1, \ldots, \xi_{p_m}$ in the sieve space, such that

align*[align* omitted — 235 chars of source]

To study the asymptomatic properties of the proposed estimator, we first establish the relationship among the sieve estimator, the sieve representation, and the underlying true densities. Let $G(x, y, z; f^{X | \mathtt{T}}, f^{\mathtt{T}})$ be the log-likelihood function from one single observation that depends on the conditional density of $X$ given $\mathtt{T} = \tau$ and the marginal density of $\mathtt{T}$.

lemmaLet $\hat{f}_m^{X | \mathtt{T}} (x | \tau)$ and $\hat{f}_m^{\mathtt{T}} (\tau)$ denote the estimated conditional density of $X$ and the marginal density of the latent variable $\mathtt{T}$, respectively. We have \begin{align} \nonumber \frac{1}{\sqrt m} \boldsymbol{G}_m\left[\log \frac{G(x, y, z; \hat{f}_m^{X | \mathtt{T}}, \hat{f}_m^{\mathtt{T}})}{G(x, y, z; f_m^{X | \mathtt{T}}, f_m^{\mathtt{T}})} \right] & \ge \boldsymbol{P}\left[\log \frac{G(x, y, z; f_m^{X | \mathtt{T}}, f_m^{\mathtt{T}})}{G(x, y, z; f_0^{X | \mathtt{T}}, f_0^{\mathtt{T}})} \right] \\ & + \boldsymbol{P}\left[\log \frac{G(x, y, z; f_0^{X | \mathtt{T}}, f_0^{\mathtt{T}})}{G(x, y, z; \hat{f}_m^{X | \mathtt{T}}, \hat{f}_m^{\mathtt{T}})} \right], \end{align} where $\boldsymbol{G}_m = \sqrt{m}(\boldsymbol{P}_m - \boldsymbol{P})$, $\boldsymbol{P}_m$ denotes the empirical measure of data $(x_i, y_i, z_i)_{i = 1}^m$, and $\boldsymbol{P}$ denotes the true distribution.

Lemma (ref) holds by definition of $f_0$ and $\hat{f}$. The proof can be found in the online supplement Section S.2.1. To show consistency of the sieve estimator, we need to bind the left-hand side of Equation (ref). To accomplish this, we resort to empirical process theories and impose restrictions on the complexity of the sieve space. We first introduce the following two assumptions to characterize its complexity.

ass[Bound of sieve space] The logarithm of the upper bound over $\mathcal{B}_m$, denoted by $Q_m$, satisfies $\log\{\sup_{\mathfrak{f} \in \mathcal{B}_m} \|\mathfrak{f}\|_{L_{\infty}([0, 1]^2)} \} \leq Q_m = O(\log \log m)$.
ass[Bracketing number] The $\epsilon$ bracketing number of the sieve space $\mathcal{B}_m$ is of order $O\left((e^{2Q_m}/\epsilon)^{p_m + 2}\right)$ for some constant $p_m = O(m^{\alpha})$ with $0 < \alpha < 1/2$.

Intuitively, $Q_m$ would be larger for a larger space. We define the bracketing number following van199. Specifically, given two functions $l$ and $u$, the bracket $[l ,u]$ is the set of all functions $f$ with $l \leq f \leq u$. An $\epsilon$-bracket is a bracket $[l, u]$ with $\|u - l\| \leq \epsilon$ under a certain norm $\|\cdot\|$. The $\epsilon$ bracketing number $N_{[]}(\epsilon, \mathcal{B}, \|\cdot\|)$ is the minimum number of $\epsilon$-brackets needed to cover $\mathcal{B}$. A larger $\epsilon$ bracketing number corresponds to a more complex sieve space.

To guarantee consistency, we consider the function class $\mathcal{F}_m$, defined by

eqnarray*[eqnarray* omitted — 353 chars of source]

where $\mathfrak{f}_m$ is represented in terms of $\xi_1, \ldots, \xi_{p_m}$ in sieve space $\mathcal{B}_m$. If the complexity of sieve space $\mathcal B_m$ satisfies Assumptions (ref)-(ref), we are able to quantify the upper bound on $\mathcal{F}_m$, which is the upper bound on the left-hand side of Equation (ref).

We now establish consistency of the proposed sieve estimator.

theoremUnder Assumptions (ref)-(ref), the proposed sieve MLE for the joint distribution is consistent. Moreover, both the conditional and marginal distributions are consistently estimated. That is, \begin{align*} \|\hat{f}_m^{X | \mathtt{T}} (x | \tau) - f_0^{X | \mathtt{T}}(x | \tau)\|_{L_2} \overset{p}{{\longrightarrow}} 0, & and \|\hat{f}_m^{\mathtt{T}} (\tau) - f_0^{\mathtt{T}}(\tau)\|_{L_2} \overset{p}{{\longrightarrow}} 0. \end{align*} The convergence rate for these estimators is derived to be $B(m, p_m, Q_m)^{1/2}$, where\\ $B(m, K_m, Q_m) = e^{c_2Q_m} p_m \log p_m / \sqrt{m} + e^{c_2Q_m} / p_m^{\beta}$, with $c_2$ being a constant.

The detailed proof is given in the online supplement Section S.2.1. Note that we consider $L_2$ convergence of our proposed estimator. Establishing the (uniform) convergence rate is beyond the scope of this paper and thus left for future research. As pointed out in menzel, the uniform rate depends on $r$, and the nonparametric MLE of the parent distribution obtained using order statistics may have a slower convergence rate near the tail of the parent distribution. The primary reason for the latter is that the mapping from the distribution of order statistics to the corresponding parent distribution may not be Lipschitz continuous. The derivative of this mapping may diverge near the tail. In this context, we found similar issues with respect to the proposed sieve MLE using Berstein polynomials from simulation studies. See the online supplement Section S.2.3.

\paragraph{The Conditional Value Distributions }Theorem (ref) concerns the distribution of UH and the conditional bid distributions. While the bid equals the value in ascending auctions, recovering the value distributions in first-price auctions requires several additional steps. First, we estimate the conditional bid quantile functions $\widehat{b}(\alpha|\tau)$ by inverting the estimated conditional bid distribution. That is, $\widehat{b}(\alpha|\tau)=\widehat{F}^{-1}(\alpha|\tau)$, where we have omitted supscript $X|\mathtt{T}$ for simplicity. Second, following GPV2000, we can recover the conditional value quantile function

eqnarray[eqnarray omitted — 119 chars of source]

which allows constructing the conditional value density and distribution. By the continuous mapping theorem chung2001course, the estimated conditional value quantile function, density, and distribution are also consistent. Moreover, if we impose higher-order smoothness assumptions on the value distribution, we may achieve a faster convergence rate, which is similar to the results in menzel.

Empirical Application

In this section, we apply our methodology to an empirical analysis of judicial auctions in China. Chinese courts began holding online auctions in 2012 through taobao.com, the shopping site of Chinese e-commerce giant Alibaba. As of 2022, almost all of China's courts have registered on this judicial sales platform, auctioning assets ranging from cars, diamonds, property, land use rights, and Boeing 747s to company shares. As of December 2019, over 500,000 items have been sold, with turnover reaching about 1.3 trillion yuan on the Taobao judicial sales platform.\footnote{\href{https://global.chinadaily.com.cn/a/201912/26/WS5e0411dda310cf3e35580b5c_2.html}{Source: China Daily.}}

The court first posts the property-related information on taobao.com, including the appraisal value, obtained through a third-party appraisal company, and a starting price. Potential buyers can view the information page online and visit the property physically before the auction starts. Interested bidders can register to participate in the bidding by paying a security deposit and then bid in an ascending fashion. They can also set up automatic bidding.\footnote{On average, a sold item receives 55 bids from 3 bidders, suggesting that jump bidding may not be a big concern.} The highest bidder wins the object and pays his/her bid.

Data

We collect a sample of residential property auctions from taobao.com, which contains all sales by the court in Jiangmen city of Guangdong Province between January 2018 and June 2020. We drop a few sales that are below ten thousand RMB or above five million RMB. In total, we have 477 auctions with 329 successful sales. By default, this court uses 70% of the appraisal value as the starting price, which also serves as a reserve price.

These auctions are subject to UH for many reasons. A third party provides appraisal based on available information at hand but may miss important details that become revealed upon careful study of the listing and a physical visit. For example, any unpaid electricity bills or property management fees of a sold property are the responsibility of the winning bidder. Some condos may have defects that are unknown to the appraisal firm. These unobserved factors constitute a significant portion of potential bidders' values. But how they enter bidder value is unknown. Therefore, it is preferable to retain flexibility when specifying how bidder value depends on UH and private information.

Following the literature, we homogenize the bids by dividing them by the appraisal value.\footnote{For the homogenization to be valid, we need either 1) the appraisal value to be realized before the realization of UH or 2) the seller or the third-party appraisal company to have the same access to UH but choose to ignore the additional knowledge.} We further rescale the homogenized bids by dividing them by the maximum value in estimation but report the results in homogenized terms for convenience. As usual, the highest and second-highest bids are close to each other, both revealing information about the second-highest value among all bidders. To avoid redundant information, we use the highest bid as the second-highest value among all the bidders and exclude the second-highest bid from the data.\footnote{We obtained almost identical estimation and counterfactual results using the second highest bid as the second highest valuation.}

Table (ref) provides some summary statistics of our data. On average, each property is worth one million RMB, which is approximately \$140,000 USD. Only about $70\%$ of listings are sold successfully, at a transaction price close to the appraisal value on average.

table[table omitted — 576 chars of source]

Empirical Model with a Binding Reserve Price

Our empirical model accounts for the binding reserve price. Upon arrival, $N$ potential bidders observe the realization of UH $\tau$ and draw i.i.d. private values from $F^{X|\mathtt{T}}(\cdot | \tau)$. Those with a valuation higher than reserve price $R$ submit a bid equal to their value. As a result, the amount of truncation for a given UH is $F^{X|\mathtt{T}}(R|\tau)$, where $R=0.7$.

Conditional on the number of potential bidders $N$, the probability of observing the bid vector $\boldsymbol{b}_{n} \equiv \{b_{1:n},....,b_{n-1:n-1}\}$ is \[ \int f^{\mathtt{T}}(\tau)p(n|N,\tau)g(\boldsymbol{b}_{n}|n,\tau)d\tau , \] where $p(n|N,\tau)=C_{N,n}\left[1-F^{X|\mathtt{T}}(R|\tau)\right]^{n}\left[F^{X|\mathtt{T}}(R|\tau)\right]^{N-n}$ is the probability of observing $n$ active bidders given the number of potential bidders $N$ and UH $\tau$, and $g(\boldsymbol{b}_{n}|n,\tau)$ represents the joint PDF of the bid vector including all active bids.\footnote{Note that the identification requires the number of active bidders to be at least four. We pool bids from all auctions, including those with fewer than four active bidders, to improve estimation efficiency but rely on the auctions with $n\ge 4$ for identification.} If $n=0$, $g( \boldsymbol{0} |n,\tau)=1$ because there is no bid. If $n=1$, $g(R|1,\tau)=1$ because the bid will be $R$, as there is no reason to bid higher than the reserve price when there is only one bidder. If $2\leq n\leq N$, the joint PDF simply becomes

eqnarray[eqnarray omitted — 154 chars of source]

To estimate the model, we ignore the fact that we cannot identify $F^{X|\mathtt{T}}(\cdot|\tau)$ below the reserve price\footnote{Fortunately, this abstraction is barely binding for calculating the optimal reserve prices. In fact, haile2003inference shows that as long as the existing reserve price is below the optimal, we obtain the same optimal $p^*$ by replacing $F_0$ and $f_0$ with the truncated version $F$ and $f$, respectively.} and approximate the joint density function using Berstein polynomials, $f(x,\tau; \theta) \approx \sum_{i,j}\theta_{ij}\beta_{i}(x)\beta_{j}(\tau)$, and solve the following optimization problem:\footnote{We approximate the integration by Monte Carlo simulations \[ \int\beta_{j}(\tau)p(n_{\ell}|N_{\ell},\tau)g(\boldsymbol{b}_{\ell}|n_{\ell},\tau)d\tau\approx\frac{1}{S_{j}}\sum_{i=1}^{S_{j}}p(n_{\ell}|N_{\ell},\tau_{ij})g(\boldsymbol{b}_{\ell}|n_{\ell},\tau_{ij}), \] where $\tau_{ij}$ represent i.i.d. random draws from the beta density function $\beta_{j}(\cdot)$. By fixing the random draws, we make the maximization smooth in the sieve parameters $\theta$.}

eqnarray[eqnarray omitted — 227 chars of source]

Empirical Findings

We let the number of sieve bases $J=3$. Figure (ref) shows the estimated joint density function of bidder value $X$ and UH $\mathtt{T}$ in homogenized and rescaled terms. Two important features are worth noting. First, the conditional densities are skewed to the left. This suggests an abundance of low willingness-to-pay amongst the potential bidders in the market, consistent with the observation that the number of registered bidders exceeds the number of actual bidders. Second, UH has important effects on bidder value. The higher $\mathtt{T}$ is, the more skewed (to the left) the density becomes.

figure[figure omitted — 157 chars of source]

To demonstrate the practical use of our estimation results, we use the distribution estimated allowing for UH to calculate the optimal reserve price for each UH. Given the number of potential bidders, the optimal reserve price maximizes

eqnarray[eqnarray omitted — 175 chars of source]

where $v_0$ is the seller's reserve value for keeping the item. The first term represents the seller's expected gain due to selling at the reserve price when only one value is higher than $r$, and the second term represents the gain due to selling at the second highest value when two values are higher than $r$. Its FOC leads to the following optimal reserve price

eqnarray[eqnarray omitted — 111 chars of source]

which is strictly increasing in the reserve value. We can infer the auctioneer's reserve value from the series of judicial rules for judicial auctions issued by the supreme court. Specifically, one important rule says that the reserve price cannot be lower than $50\%$ of the appraisal value. This seems a reasonable proxy for $v_0$, i.e., $v_0 = 0.5$.

Figure (ref) shows the optimal reserve price for different levels of UH. The reserve price is strictly monotone in UH, which is consistent with the monotonicity assumption (ref) and the estimated joint density in Figure (ref). It is also reassuring that the optimal reserve prices are well above the current reserve price, which means that underidentification below the reserve does not prevent us from calculating the optimal reserve price.\footnote{The optimal reserve prices are still above $0.7$ with more conservative values as low as $v_0=0.1$.} In Figure (ref), the blue dashed line shows the optimal expected seller gain as a function of UH. The unconditional optimal gain $\sum_{N}p_{N}\pi(r^*, N)$ is $36.61\%$ of the appraisal value, which is $5.81\%$ higher than the current one ($34.60\%$ of the appraisal value).

figure[figure omitted — 148 chars of source]

Of course, it is difficult to imagine that the seller adopts such a complex strategy. To achieve the optimal gain, the seller would need to know the UH and recover the conditional density of bidder values. Simpler strategies that require less knowledge of the value distributions are often preferable.\footnote{coey2020scalable makes a similar point. They provide an approach to calculate optimal reserve prices without fully recovering value distributions.} We observe that the optimal reserve price is almost constant and close to one when UH is above $0.4$. Moreover, the density of UH is heavily skewed to the right (near 1). Therefore, a simple alternative to a complex UH-specific reserve price is to use the appraisal value as the reserve price. We calculate the expected revenue in this simple scheme. In this case, the unconditional expected gain is $36.59\%$ of the appraisal value, which achieves $98.85\%$ of the potential gains from the optimal reserve prices.\footnote{The appraisal value as the reserve price is nearly optimal; this finding is robust to “large” auctions, different seller reserve values, and alternative tuning parameters.}

figure[figure omitted — 157 chars of source]

Conclusion

Auction data often contain incomplete bids and miss some payoff-relevant covariates. The conventional measurement error approaches to UH are inapplicable. In this paper, we extend the analysis of hu2008instrumental to auctions with continuous UH while accounting for incomplete bid data. Specifically, we provide point identification results for auctions with nonseparable continuous UH using consecutive order statistics of bids. We then propose sieve maximum likelihood estimators jointly for the value distribution conditional on UH and its marginal distribution. We illustrate our methodology using a novel dataset from judicial auctions conducted by a municipal court in China. After recovering the model primitives, we propose a simple scheme that achieves nearly optimal revenue by using the appraisal value as the reserve price.