Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
76,148 characters · 14 sections · 72 citation commands
Identification of Auction Models Using Order Statistics
JEL: C14, D44
Key Words: Consecutive Order Statistics, Finite Mixture, Unobserved Competition, Multidimensional Unobserved Heterogeneity
Recent years have witnessed a rapid growth in the literature combining auction theory with econometric analysis to understand auction markets and inform policies. While the key elements in auction theory generally match well with the empirics, this may not be the case when there is auction-level unobserved heterogeneity, i.e., factors affecting bidder values that are common knowledge among bidders but unobserved by the analyst. Ignoring such unobserved heterogeneity (UH) would lead to erroneous estimates and misinformed policy conclusions. See, e.g., HK2019 for a survey.
Measurement error approaches exploit the multiplicity of bids in each auction to account for UH. See, e.g., li2000conditionally, K2011 and HMS2013.\footnote{The earliest approach to allow for UH in auctions exploits an auxiliary variable that is monotone in UH. See, e.g., CGPV2003 and HHS2003.} They require observed bids to be independent conditioning on UH. In practice, however, this conditional independence assumption may not hold due to auction format or data truncation. Often, we may only observe multiple order statistics or a subset of bids rather than all bids themselves. For instance, the highest bid is never observed in ascending auctions.\footnote{For instance, KL2014 observes the second, third, and fourth highest bids in ascending used-car auctions. FL2017 uses the second and third order statistics of bids in eBay auctions for used iPhones.} Moreover, the low/high bids are much easier to access in many settings.\footnote{For instance, three apparent low bids in all auctions since 1999 are one-click downloadable on the California DOT website. The Washington State DOT archives bid opening results of six months or older online, but only for three apparent low bids. allen2019resolving studies FDIC auction data, which contain the bid and the identity of the individuals associated with the winning and second-highest bids. Unfortunately, the data miss the names associated with all other bids. U.S. Forest Service timber auctions only record at most the top 12 bids regardless of the number of bidders.} Naturally, these order statistics are dependent even though the bids themselves are independent, thereby rendering the conventional measurement error approaches inapplicable.
This paper provides a set of positive results on the identification of auction models with discrete UH using multiple order statistics of bids.\footnote{In the classical measurement error setting, AH2002 suggests “there may be sufficient structure to identify the model from only two order statistics... However, we have not obtained such a result.” Recently, luo2022two obtains this result. Another companion paper, luo2020order, considers nonseparable continuous UH using three consecutive order statistics.} Rather than circumventing the correlation among order statistics, we propose new identification strategies that exploit this very correlation structure with the same number of bids as the conventional measurement error approaches. Our results provide new perspectives into the identification of auction models with UH when the conditional independence assumption fails. Moreover, this opens a new window to explore the use of statistical/model structure to restore identification.
We consider a common form of incomplete bid data: consecutive order statistics of bids. All the examples mentioned above have this form. Despite their correlation, we derive that the joint distribution of consecutive order statistics has a semi-separable structure with a type-independent indicator function capturing the correlation. Based on this finding, we show that two consecutive order statistics and an instrument or three consecutive ones identify independent private value (IPV) auction models with nonseparable finite UH.
Our first result considers a benchmark case with symmetric bidders and observed competition (i.e., known number of bidders). We allow the cardinality of UH's support to be unknown and show its identification using two consecutive ordered bids. Given this cardinality, we first identify the component bid distributions and then link them to the model primitives --- value distributions.
We propose a novel multi-step procedure using two consecutive order statistics and an instrument. First, exploiting the joint probability of observed bids and their being in two ordered intervals, we turn the semi-separable joint probability into an eigendecomposition structure and identify the distribution of an order statistic up to scale in each interval. Second, using its relationship with order statistics, we identify the parent bid distribution up to (different) scales in these intervals. Third, we pin down the scales exploiting the continuity of the component bid density functions. In ascending auctions, the bid distributions equal the value distributions. In first-price auctions, we follow GPV2000 to recover the value distributions using the bid distribution and the observed competition. When an instrument is difficult to find, our method extends to the scenario of three consecutive order statistics.
Our second result concerning ascending auctions allows competition as an additional dimension of UH, which is common with incomplete bid data. Missing information on competition leads to multiple sources of UH: auction state and number of bidders. This has rarely been dealt with in the auction literature. In ascending auctions, bidding strategies and bid distributions are independent of competition intensity. We extend S2004 by allowing for UH. On the one hand, we follow the paper to “concentrate out” unknown competition intensity. This transforms our problem into a mixture problem with just an unknown auction state and identifies the component bid distributions up to scales. On the other hand, we argue that S2004's identification of the scales at the limit is at odds with our setting. Instead, we propose a novel approach to pin down competition intensity and the scales by combining the joint and marginal distributions of the observed order statistics. Without UH, we identify the value distribution and the distribution of competition using two consecutive order statistics; with UH, we can do the same by using an additional instrument or order statistic. In contrast, S2004's approach cannot identify the distribution of competition under UH.
There has been growing interest in using order statistics for identification in the auction literature. Many markets where only the transaction price is observed can be treated as auctions. AH2002 shows that the auction model with symmetric IPV is identifiable with the transaction price and the number of bidders. This exploits a one-to-one mapping between the distribution of an order statistic and its parent distribution. komarova2013new shows that the asymmetric second‐price auction model is identifiable solely using the winner's identity and the transaction price. GL2018 shows that the first-price auction model with symmetric IPV and unknown competition is identifiable solely using the transaction price and exploiting discontinuities in the density function of the winning bid due to changing competition. Our paper is the first to consider consecutive order statistics for identification; it allows a general form of UH as opposed to unknown competition in GL2018. In contrast to partial identification approaches (see, e.g., haile2003inference and AGQ2013), our approach leads to point identification and thus allows for the calculation of optimal reserve prices conditional on the realization of UH.
Our paper focuses on discrete nonseparable UH. The existing literature that allows for finite UH uses the eigendecomposition approach, as in HMS2013. Identification is achieved using the joint distribution of at least three bidders per auction, following H2008. This approach has been adapted to various settings. See, e.g., an2010estimating for unknown competition, hortaccsu2018empirical for auctions of multiple objects, and Luo2018 for multiple sources of UH such as the auction state and the number of bidders. Related to our first result, M2017 takes an alternative approach that attempts to restore the conditional independence structure needed in the usual approaches. In particular, he exploits the Markov property of order statistics and shows that the bidder's UH-specific distribution and the distribution of UH is point identified from (any) five order statistics. In contrast, we take advantage of the structure provided by consecutive order statistics and achieve identification using two consecutive order statistics and an instrument or three consecutive ones.
Our extension contributes to the literature dealing with unobservability of competition to the analyst. This literature begins with laffont1995econometrics, which treats unknown competition as an unknown parameter in first-price auctions. Recent contributions treat unknown competition as random (see, e.g., an2010estimating, shneyerov2011identification, and GL2018 for first-price auctions). Regarding ascending auctions, S2004 considers eBay auctions where the number of bidders is random and unobservable, showing that any two order statistics identify the symmetric IPV model without UH. KL2014 applies the identification results of S2004 to wholesale used-car auctions in Korea, for which only a small subset of bids is observable (due to the ascending auction format). FL2017 studies eBay auctions for used iPhones with auction-specific UH and an unknown number of bidders. The issue of correlated order statistics is circumvented using observed reserve prices, returning the identification problem to its standard form. Our paper extends this literature on unknown competition by introducing multiple sources of UH: auction state and the number of bidders.
Another literature considers continuous UH using the deconvolution approach, as in LV1998, li2000conditionally, and K2011, among others.\footnote{chen2011nonlinear provides an excellent survey of measurement error models.} Two random bids are required in each auction, and UH is restricted to have a separable structure on bidder valuation. HQT2019 achieves point identification using English auction models with additively separable UH, assuming piecewise real analytic density functions and using variations in the number of bidders across auctions. We focus on finite UH, which does not directly extend to the deconvolution approach.
The remainder of this paper is organized as follows. Section 2 describes the auction models. Section 3 explains some useful properties of order statistics. Section 4 presents our main identification results. Section 5 concludes. The Appendix contains longer proofs.
We consider symmetric standard independent private value (IPV) auction models. In practice, the auctioned items can be heterogeneous. While bidders observe such heterogeneity, the analyst may not observe all auction characteristics, resulting in auction-level unobserved heterogeneity (UH). Hereafter, we refer to the realization of the auction-level UH as “state.” For the remainder of this section and Section (ref), we present results that condition on a realization of UH. In Section (ref), we return to dealing with the distribution of UH.
Suppose that, for a given auction, the bidder values are i.i.d.\ draws from the same distribution $\Phi^k(v)$, where auction state $k$ is finite and discrete, i.e., $k=1,...,K$, and state $k$ is realized with probability $p_k>0$, and $\sum^K_{k=1} p_k=1$. Hereafter, we abstract from observable (to the analyst) characteristics for simplicity. Our results are applicable to models with both types of heterogeneity.
Suppose $n \geq 2$ symmetric bidders participate in an auction with a zero reserve price. Bidders are risk-neutral. Conditioning on auction heterogeneity state $k$, bidder valuations are $i.i.d.$ draws from the same distribution $\Phi^k(v)$ with support $[\underline{v},\overline{v}]$. All bidders know auction heterogeneity $k$ before they submit bids. We denote the optimal bid distribution for state $k$ as $F^k(x)$, where $x$ is the optimal bid with support $[\underline{x},\overline{x}_k]$. For the remainder of this section and section 3, we present results that condition on a realization of the unobserved state. In Section 4, we then return to how we identify the state-specific bid distribution, i.e., the bid distribution conditioning on the UH, and the marginal distribution of UH.
Once the auction starts, bidders place their respective bids in ascending order. This process continues until there are no higher bids. Optimal bidding behaviors in these auctions are straightforward: a weakly dominant strategy is to continue bidding until the standing bid reaches your value.\footnote{Ascending auction in our content means ascending button auction, that is, we abstract from other possible behaviors, such as jump bidding in haile2003inference.} In other words, once the next-to-last bidder drops out, the bidder with the highest value wins at a price equal to the second-highest value. We call the “planned" dropout price the bid, which equals the value. Therefore,
However, the highest bid is never observed. As a result, the distribution of the transaction price given auction state $k$, denoted as $F_{n-1:n}^k(\cdot)$, is the distribution of the $(n-1)^{th}$ order statistic out of an $n$ ordered sample. If the state $k$ and the number of bidders $n$ are known, the unknown state-specific bid distribution $F^k(\cdot)$ can be identified because there is a one-to-one mapping between the distribution of the $(n-1)^{th}$ order statistic and the underlying parent distribution itself.\footnote{That is, $F_{n-1:n}^k(x) = n(n-1)\int_0^{F^k(x)} t^{n-2} (1-t) dt.$}
In a first-price auction, the bidder with the highest bid wins and pays his own bid price. In a state-$k$ auction, a bidder with valuation $v$ chooses his bid $x$ to maximize his expected payoff, as follows: \[ \max_x \quad (v-x) \cdot \underbrace{\Phi^k \big(\xi^k(x) \big)^{n-1}}_{\text{probability of winning}}, \] where $\xi^k(\cdot)$ is the inverse of his optimal bidding strategy in state-$k$ auctions, and $\Phi^k \big(\xi^k(x) \big)^{n-1}$ is the probability of winning.
GPV2000 studies the identification of the state-specific value distribution $\Phi^k(\cdot)$ when the competition level $n$ and the state-specific bid distribution $F^k(\cdot)$ are known. In particular, the identification comes from a one-to-one mapping between the unknown state-specific value distribution and the observed state-specific bid distribution:
where $x$ is any arbitrary bid in its support, and $F^k(\cdot)$ and $f^k(\cdot)$ are the state-specific bid distribution and density function, respectively. Once the state-specific bid distribution $f^k(\cdot)$ and competition $n$ is known, we can recover values corresponding to bids in the data and, in this way, identify the value distribution.
\paragraph{Remark} In both models, we need two elements to identify the state-specific value distributions $\Phi^k(\cdot)$: state-specific bid distributions $F^k(\cdot)$ and number of bidders $n$. In this paper, we will provide identification results for each case of when the number of bidders is known and unknown (to the analyst). For purposes of exposition, we start with assuming known competition and then extend the result to unknown competition. Since (ref) and (ref) provide explicit mapping from the bid distribution to the value distribution, we can claim identification as soon as we identify state-specific bid distribution $F^k(\cdot)$ and number of bidders $n$.
We now derive a useful property of order statistics that we exploit to obtain our identification results. In particular, we show that the joint distribution of consecutive order statistics has a semi-separable structure with a state-independent indicator function that captures the correlation. We omit UH in this section.
We first introduce some notation. Let $X_{1:n} \le X_{2:n}\le...\le X_{n:n}$ represent the $n$ order statistics out of an $n$ ordered sample. Let $f_{r,s:n}(\cdot,\cdot)$ denote the joint PDF of order statistics $\{X_{r:n}, X_{s:n}\}$; and let $f_{r:n}(\cdot)$ denote the marginal distribution of order statistics $X_{r:n}$. CDFs are defined analogously.
Suppose that, out of a sample of $n$ bids in the same auction, we observe two order statistics, $X_{r:n}$ and $X_{s:n}$, where $r<s$. First, following david2004order, we visualize the event $(x<X_{r:n} \leq x+\delta_r, y< X_{s:n} \leq y+\delta_s)$ with $x < x+\delta_r \le y$ in Figure (ref). When $\delta_r$ and $\delta_s$ are both small, the likelihood of this event is approximately proportional to the product of the following components: (1) the probability that $r-1$ draws are less than $x$, i.e., $[F(x)]^{r-1}$; (2) the probability that one draw is in $(x, x+\delta_r)$, i.e., $f(x) \delta_r$; (3) the probability that $n-s$ draws are greater than $y$, i.e., $[1-F(y)]^{n-s}$; (4) the probability that one draw is in $(y, y+\delta_s)$, i.e., $f(y) \delta_s$; (5) the probability that $s-r-1$ draws are in $(x + \delta_r, y)$, i.e., $[F(y) - F(x)]^{s-r-1}$. Second, we observe that the likelihood of this event can be represented in a semi-separable structure with respect to $x$ and $y$ if and only if $s-r-1=0$; i.e., the two order statistics are consecutive.
This lemma formalizes the “separability by consecutiveness” shown in Figure (ref). $f_{r-1:r-1}(x)$ is the PDF of the top-order statistic of a sample of size $r-1$, and $f_{1:n-r+1}(y)$ is the PDF of the bottom-order statistic of a sample of size $n-r+1$. Compared to the joint PDF of (any) two random variables, a simple indicator function captures the correlation among two consecutive order statistics. As a result, their joint PDF is separable on the segment of $x \le y$.
\paragraph{Remark} It is worth noting that the multiplicatively separable structure of the joint distribution of two consecutive order statistics, represented in Equation (ref), is different from the conditional independence of unordered bids. First, the joint distribution of any two unordered bids can be represented as the product of the marginal distribution of the two unordered bids. In contrast, the separable structure provided by the consecutiveness of the order statistics is represented by the product of the marginal distribution of two newly constructed random variables instead of the marginal distribution of the two ordered statistics. That is, $f_{r-1,r:n}(x,y) \neq c_{r-1,n} f_{r-1:n}(x) f_{r:n}(y).$ Second, the multiplicatively separable structure only holds when $x\leq y$ and not at all times. If $x >y$, the left-hand side equals 0 by definition while the two marginal distributions do not necessarily equal 0. Therefore, we need to use the indicator function to capture the ordered relationship. However, we show that conditional independence is not necessary for identification with UH. Instead, the correlation structure by the consecutiveness of the order statistics is sufficient for identification.
In general, the distribution function of the $r^{th}$ order statistic $X_{r:n}$ is \[ F_{r:n}(x) = \frac{n!}{(n-r)! (r-1)!} \int_0^{F(x)} t^{r-1} (1-t)^{n-r} dt , \] which is strictly increasing in $F(x)$ and thus invertible. Therefore, there exists a one-to-one mapping between the distribution of any order statistic and its parent distribution $F(x)$. That is to say, the distribution of any order statistic is sufficient to identify its parent distribution. This property has been used in several papers including AH2002 and S2004. Of course, this relationship between the distribution of an order statistic and its parent distribution depends on $n$, which is unobserved if we consider unknown competition.
As mentioned above, auction data often only capture a subset of all bids. In sealed-bid first-price auctions, we typically observe the most competitive bids, be it the highest bids in regular auctions or the lowest bids in procurement auctions. Note that we define the order statistics as $X_{1:n} \le X_{2:n}\le...\le X_{n:n}$, indicating that there are some subtle differences in notation when it comes to auctions and procurements. Specifically, using the language of order statistics, the winning bid is $X_{n:n}$ in regular auctions, i.e., $r=n$, but $X_{1:n}$ in procurement auctions, i.e., $r=1$.
These subtleties are more evident when the number of bidders is unknown, i.e., $n$ is unobserved. Specifically, the two most competitive bids should be denoted as $\{X_{1:n},X_{2:n}\}$ in a procurement auction but as $\{X_{n:n}, X_{n-1:n}\}$ in a regular auction. As a result, we know the exact value of $r$ in the former, i.e., $r=1,2$, but only the relation between $r$ and $n$ in the latter, i.e., $r=n, n-1$.
In this section, we consider identification of symmetric auction models with finite nonseparable UH using only order statistics of bids instead of all bids. First, conditional on the level of competition, we show that the cardinality of UH's support is identified. We then provide sufficient conditions to identify the state-specific value distributions using two consecutive order statistics and one instrumental variable. This result can be extended to the case with an additional order statistic but no instrumental variable. Second, we extend the identification result to allow for unobserved competition for ascending auctions. Last, we discuss estimation and inference.
Suppose that, for a given auction, the bidder values are i.i.d.\ draws from the same distribution $\Phi^k(v)$, where auction state $k$ is finite and discrete, i.e., $k=1,...,K$, and state $k$ is realized with probability $p_k>0$, and $\sum^K_{k=1} p_k=1$. With slight abuse of notation, we will consistently use $c$ to denote a known constant and $\eta$ to denote an unknown constant throughout the paper. The econometrician observes two consecutive bids and one instrumental variable, i.e., $\{(X_{r-1:n}, X_{r:n}),W\}$, while the auction-level state $k$ is unobserved. Therefore, any component that only involves the order statistics of bids are directly estimable from the data and thus can be treated as known while any state-specific component is to be identified.
The unconditional joint PDF of two consecutive order statistics can be estimated directly from the data. By total probability, it can be represented as
where superscript $k$ indicates the state-$k$ parent distribution. Note that $c_{r-1,n} = \frac{n!}{(r-1)!(n-r+1)!} $, $f_{r-1,r:n}(x,y)$ can be directly estimated from the data while $p_k$, $f^k_{r-1:r-1}(x)$, and $f^k_{1:n-r+1}(y)$ are the unobserved component to be identified. This indicator function, capturing the correlation of the order statistics, precludes us from directly following the existing procedure of eigenvalue-eigenvector decompositions to identify the state-specific bid distribution. Fortunately, such a correlation is known and does not depend on the unobserved state, which enables us to modify the existing identification result to achieve identification in our context.
Specifically, if we limit variation in the order statistics in predetermined no-overlapping intervals, the ordering correlation trivially holds. Essentially, to account for this correlation/ordering, we introduce a discretization of bids before proceeding to identification analysis. We divide the original support into two segments, “low" and “high," using one cutoff $\chi$, where $\underline x<\chi<\bar x = \max_k \{\bar x_k\}$. Denote the two segments as $l = [\underline{x}, \chi]$ and $h = [\chi, \bar x]$. Therefore, if we only exploit the joint variation of $X_{r-1:n}=x$ and $X_{r:n}=y$ in the two ordered interval $x\in l$ and $y\in h$, the separable structure of the joint PDF $f_{r-1,r:n}(x,y)$ reappears, which has three important features. First, this representation has a similar structure to finite mixture models. This feature allows us to turn this structure into an eigendecomposition to study identification. Second, the variations of the two observed ordered bids are restricted to their respective intervals. This restriction requires us to study the identification of component distributions interval by interval. Third, each term in the multiplicatively separable structure is the PDF of a newly constructed order statistic associated with the parent distribution. This allows us to translate the identified distributions of order statistics into the parent distributions.
Similar to the finite mixture and measurement error literature (H2008), our identification uses matrix algebra. We now further discretize the bid support and introduce some matrix notation. We select $\tilde K$ exclusive intervals from “low" segment $l$ and “high" segment $h$, denoted as $l_i$ and $h_i$, respectively, where $i=1,...,\tilde K$. Note that these intervals do not have to be fully exhaustive, and they can simply be points. Figure (ref) provides a visualization of the discretization.
Following Equation ((ref)), we express the probability of the event $\{X_{r-1:n} \in l_i,X_{r:n} \in h_j\}$ as
We first introduce the following matrix notation:
where $\boldsymbol{J}_{l,h}$ is a probability matrix with $(i,j)^{th}$ element representing the probability of event $\{X_{r-1:n} \in l_i,X_{r:n}\in h_j\}$, which can be identified and estimated directly from the data. $\boldsymbol{L}$ is a probability matrix with $(i,k)^{th}$ element representing the probability of event $\{X^k_{r-1:r-1} \in l_i\}$; $\boldsymbol{D}_p$ is a diagonal matrix with $k^{th}$ diagonal element representing the weight of state $k$; $\boldsymbol{H} $ is a probability matrix with $(j,k)^{th}$ element representing the probability of event $\{X^k_{1:n-r+1} \in h_j\}$. Note that superscript $k$ indicates the random variable is associated with the state-$k$ parent distribution. Also, matrices $\boldsymbol{L}$, $\boldsymbol{D}_p$, and $ \boldsymbol{H}$ are unobserved and yet to be identified.
With the matrix notation, we have the following matrix representation connecting the observed joint probability with the unknown component matrices:
Following the existing literature (H2008), identification of models with UH using the mixture features generally requires a rank condition, which is stated as follows.
The linearly independent assumption basically requires that there is sufficient variation in the bid distribution for different unobserved states. The smaller $K$ is, the easier this condition holds. For example, if $K=2$, this assumption trivially holds for any two bid distributions except that both state-specific bid distributions are uniform distributions.
\paragraph{Identification of $K$} We first show that we can identify the unobserved cardinality of the support of UH; that is, $K$ is identified. In contrast, the existing literature usually assumes $K$ is known, as in HMS2013. We explicitly use superscript $\tilde K$ to denote the matrices associated with a $\tilde K$-interval discretization of $l$ and $h$. That is, the joint distribution of the two order statistics with $\tilde K$-interval discretization can be represented as
where the observed matrix $\boldsymbol{J}^{\tilde K}_{l,h}$ has dimensions of $\tilde K \times \tilde K$, the matrices $ \boldsymbol{L}^{\tilde K} $ and $\boldsymbol{H}^{\tilde K}$ have dimensions $ \tilde K \times K$, and $\boldsymbol{D}_p$ are diagonal matrices with dimension of $K \times K$.
We now derive the relationship between the unknown $K$ and the rank of the joint probability matrix $ \boldsymbol{J}^{\tilde K}_{l,h}$, which is directly estimable from the data.
Lemma (ref) says that the maximum of $rank(\boldsymbol{J}^{\tilde K}_{l,h})$ among all $\tilde K$-interval discretization is strictly increasing in $\tilde K$ when $\tilde K < K$ and constant when $\tilde K \geq K$. Intuitively, one can start with $\tilde K=2$ and stops until the rank of the observed matrix $\boldsymbol{J}^{\tilde K}_{l,h}$ stops growing. Once the unobserved $K$ is identified, we treat it as known and suppress the discretization of $K$ as superscript hereafter.
\paragraph{Identification of the State-specific Bid Distribution}Once the cardinality of UH, i.e., $K$, is identified, we show that the state-specific bid distribution can also be nonparametrically identified if there exists an instrumental variable, which we denote as $W$. The requirement for the instrument is mild --- as long as there is variation in the instrument; that is, the instrumental variable can be a binary variable.\footnote{This is similar to the hu2017econometrics 2.1 measurement model.} We present the assumptions for a valid binary instrumental variable. All assumptions and identification results here can be readily extended to the situation with general instrumental variables.
In empirical applications, we can often construct such an instrument from supplementary data sources. For example, one possible instrumental variable in timber auctions is slope (i.e., the incline of the land) or aspect (i.e., the compass direction that a terrain surface faces). Because the department of natural resources geocodes the locations of timber lots, we can derive both slope and aspect using GIS and public elevation data. For example, in the Northern Hemisphere, timber on the southern side receives more sunlight than on the northern side. As a result, timber on the southern side tends to be of higher quality $k$, i.e., satisfying the relevance condition. On the other hand, the aspect does not directly affect bidders’ valuations, satisfying the exogeneity condition. Similar examples include average rainfall or soil quality (HMS2013). In highway procurement auctions, projects in areas with high traffic volume tend to be of higher difficulty. On the other hand, traffic volume does not directly affect bidders' costs.
With the instrumental variable's exogeneity condition being satisfied, the joint distribution of the two consecutive order statistics, i.e., $X_{r-1:n}=x$ and $X_{r:n}=y$ in the two ordered interval $x\in l$ and $y\in h$, and the instrumental variable can be represented as
Fixing $W=0$, using the $K$ intervals in the $l$ and $h$ segments, we can rewrite the above equations connecting the unknown state distribution with the observed joint probability into a matrix representation:
where $\boldsymbol{J}_{l,h,0}$ is constructed similar to $\boldsymbol{J}_{l,h}$ with $W=0$, and $\boldsymbol{D}_0$ is a diagonal matrix with the $l^{th}$ diagonal element being $\Pr(W=0|k)$.
If Assumption (ref) holds, there exists a discretization so that matrices $\boldsymbol{J}_{l,h}$ are invertible because both $\boldsymbol{L}$ and $\boldsymbol{H}$ are full rank. Therefore, combinig Equations (ref) and (ref) leads to the following main equation for identification:
which indicates that the observed matrix on the left-hand side and the unknown matrices on the right-hand side are similar. Specifically, probability matrix $ \boldsymbol{L}$ of the “low" segment and the conditional probability $ \boldsymbol{D}_{0} $ can be identified as the eigenvector and eigenvalue matrices of the observed joint probability matrix, respectively.
The relevance of the instrumental variable guarantees that the decomposition admits distinct eigenvalues, which guarantees the eigendecompositon to be unique. The identification of the state-specific bid distribution then proceeds in several steps. First, an eigenvalue decomposition argument identifies a key matrix that governs the finite mixture structure in our order statistic setting. This allows for the identification of component bid distributions using the one-to-one mapping between the distribution of an order statistic and its parent distribution in the low and high segment of the support. Second, we pin down the scales.
Identification of the components up to permutation is a prevalent feature of identification via eigendecomposition, which requires additional information about the UH to pin down its exact value. If one is unwilling to impose such an assumption, we can keep the UH as an index without providing any economic meaning to the labeling. Continuing the example of aspect, suppose it maps to UH one-to-one. Since timber sales are a mixture of the two unobserved states, Lemma (ref) means that we can identify two sets of objects --- each with the probabilities associated with the same state in “low” and “high” segments --- but we cannot determine which set is for the southern or northern side. That is, the ordering of the UH is the same in both matrices. Typically, a monotonicity condition can pin down the true labeling of the UH, which depends on the economic context of the UH. Knowing that timber on the southern side tends to be more valuable allows us to assign aspect to the identified objects.
Identification up to scales means that the component distribution identified from the decomposition is not the distribution itself but a scalar multiplication of it. This feature is also prevalent in identification using the mixture feature. Pinning down the scale requires a normalization condition. In the existing literature of independent measurements, such as HMS2013, since the identification exploits the variation of each bid in its full support, such a feature enables direct pinning down the scales. Specifically, each column of their eigenvector matrix $\boldsymbol{L}$ sums to one because this sum represents the cumulative distribution over the full support. However, such a normalization condition is not applicable in our order statistics context. This is because, to guarantee the multiplicative separable structure, we can only exploit the variation of ordered bids in two exclusive segments of the full support. Therefore, each column of our eigenvector matrix $\boldsymbol{L}$ sums to an unknown quantity strictly less than one, representing the probability of observing $X_{r-1:r-1}^k$ within the “low” segment, i.e., $F^k_{r-1:r-1}(\chi)$.
To summarize, we identify the state-specific bid distributions in the “low" and “high" segments up to the same ordering but different scales:
where $\check{f}_l^{k}(\cdot)$ and $\check{f}^{k}_h(\cdot)$ are the components identified for the “low" and “high" segments with the associated scales $\eta_l^k$ and $\eta_h^k$, respectively.
To identify the state-specific value function, we need to pin down the scales, which requires two restrictions for the two unknowns for each $k$. We invoke the continuity of the component PDFs and the total probability argument. First, the PDFs identified separately in the “low" and “high" segments should be the same at cutoff point $\chi$ due to the continuity of the true component PDF. Second, the fact that each component PDF integrates to 1 provides the second restriction on the scales. In the Appendix, we show that these restrictions uniquely identify the scales. We summarize the result that scales are identified in the following lemma.
Once the scales have been pinned down, we can identify state-specific weight $p_k$ using the marginal PDF/CDF of one order statistic. To summarize, we can identify the state-specific bid distribution using only two consecutive order statistics of the bids and one binary instrumental variable with suitable rank conditions. After identifying the state-specific bid distribution, we then identify the state-specific value function for first-price and ascending IPV auctions since the number of bidders is known. The following theorem summarizes our results.
\paragraph{Remark}Lemma (ref) implies that $X_{r-1:n}$ and $X_{r:n}$ are independent conditioning on the event $\{X_{r-1:n} \in l, X_{r:n} \in h\}$. Therefore, we can also consider $\Pr(X_{r-1:n} \in l_i, X_{r:n} \in h_j | X_{r-1:n} \in l, X_{r:n} \in h)$ and follow H2008, as in HMS2013, in conducting eigendecomposition to identify the two probability matrices and then the conditional distribution of the associated order statistics. HMS2013 does not require further treatment for scales. In contrast, there are two unresolved issues in our context. First, the decompositions can only identify the conditional distribution of the order statistics, and one needs to invoke the one-to-one mapping between the distribution of an order statistic and its parent distribution. Second, while the decompositions identify the conditional distribution of the order statistics, the probability of the conditional event is unknown. Pinning down the latter requires additional treatment, such as a strategy similar to Lemma (ref).
It is also worth noting that while the literature mainly relies on the independence property for identification (H2008), the eigendecomposition approach essentially exploits the implication of such a property --- multiplicative separability. That is, multiplicative separability is a condition weaker than independence that is sufficient for identification. This observation makes our finding “separability by consecutiveness” essential in solving the long-standing identification problem. Therefore, we work directly with the separability of the unconditional joint distribution of the consecutive order statistics, obviating the additional conditioning in the proofs.
\paragraph{Remark} When it is difficult to find a valid instrumental variable, we can use an additional order statistic of bids instead. Specifically, a similar logic can be applied to three consecutive order statistics $\{X_{r-2:n}, X_{r-1:n},X_{r:n}\}$; their joint PDF has the following semi-separable structure:
Intuitively, we can treat the additional order statistic as the instrumental variable. See our companion paper luo2020order for a formal discussion.
In this subsection, we consider two sources of unobserved auction-level characteristics: unobserved auction state $k$ and unobserved competition $n$. For simplicity, we focus on ascending auctions in this subsection.\footnote{See our working paper luo2020identification for discussion on first-price auctions with unknown competition.} We build on S2004, which considers online English auctions with unobserved competition. Related to our result, FL2017 studies the problem in the classical setting (i.e., separable and continuous UH) using reserve price and two order statistics of bids; without identifying the value distributions, coey2021scalable focuses on the identification of the optimal reserve price using the top two bids for online auctions under a second-price-auction-like format with sequential arrival of bidders. Our approach uses two consecutive order statistics with an instrument, and achieves point identification under nonseparable and finite UH.
We assume exogenous participation, that is, the value distribution only varies with the auction state but not the number of bidders, and the number of bidders takes values from a known set, i.e., $n \in \{\underline N, \underline N+1, ..., \bar N\}$, where $\underline N$ and $\bar N$ are known.\footnote{In principle, we can further exploit model restrictions to identify the support. For instance, GL2018 proposes a density discontinuity approach in first-price auctions using the fact that the upper boundary of the bid distribution is strictly increasing in the number of bidders. However, this approach fails in ascending auctions. An alternative is to obtain estimates from other sources. For instance, the seller may require at least two bidders. Therefore, $\underline{N} = 2$. If the data contain firm locations, the maximum number of bidders $\overline{N}$ can be estimated by the number of firms in the local market.} Moreover, $p_{k,n} \in (0,1), \sum_{k,n} p_{k,n} =1$. Denote the number of possible competition levels as $|n|$, so $|n|=\bar N-\underline N+1$. Under exogenous participation, the value distribution differs with auction state $k$ but is the same regardless of competition $n$. That is, there are $K$ different value distributions to be identified, i.e., $\Phi^k(v)$, where $k=1,...,K$. In ascending auctions, bidding their true values is a weakly dominant strategy regardless of the competition level. That is, the optimal bidding strategy is the same for any number of bidders. Consequently, the state-competition-specific bid distribution is independent of $n$, i.e., $F^{k,n}(\cdot) = \Phi^{k}(\cdot)$.
As discussed in Subsection (ref), there are some notation subtleties when competition is unknown. Here, we represent the $n$ order statistics as $X_{1:n}\geq X_{2:n} \geq \cdots \geq X_{n:n}$ in regular auctions and $X_{1:n}\leq X_{2:n} \leq \cdots \leq X_{n:n}$ in procurement auctions, which condenses the unknowns in the subscript to one element. Additionally, we can translate these results to procurement auctions by replacing all mention of “value” with “cost.”
To understand the source of identification, we first demonstrate the identification abstracted from unobserved states and then extend the identification to allow for both unobserved competition and unobserved state. \paragraph{Unobserved Competition} When the only unobserved factor is the number of bidders, the joint distribution of order statistics $X_{r-1}$ and $X_r$ can be represented as
which follows from Lemma (ref). Consecutive order statistics have a much simpler joint distribution than arbitrary ones, as employed in S2004. This joint distribution reveals that the unconditional joint distribution from the data only has a mixture of order statistics larger than the $r^{th}$ while the information of the $(r-1)^{th}$ order statistic is invariant to the level of competition. S2004 makes this important observation. This is because the order statistics observed in the data are from the bottom. Intuitively, if we fix the $r^{th}$ order statistic to be $\chi$, the joint density of the two order statistics will be equivalent to the marginal distribution of the top order statistics from an $r-1$ sample scaled by an unknown constant, i.e., $\bar f(y) \equiv \sum_n p_{n} c_{r-1,n} f_{1:n-r+1}(y) $. Consequently, we can identify the marginal distribution for such a top order statistic out of an $r-1$ sample up to an unknown scale; i.e., $f_{r-1:r-1}(x)$ is identified with an unknown constant for $x\le \chi \le \bar x$. Furthermore, we also identify its parent distribution up to a scale using its connection with the distribution of its order statistic. That is,
where $\check {F}(x)$ is identified and $\eta$ is the unknown scale.
To fully identify the bid distribution, we need to tackle the following two issues: (1) we need to pin down the scale $\eta$; (2) we need to identify the bid distribution for $y \ge \chi$. S2004 solves both problems simultaneously by taking $\chi$ to the upper limit so that the whole distribution is identified, because scale $\eta$ can be pinned down using the fact that $F(\bar x)=1$.
We propose an alternative approach that simultaneously identifies the scales and distribution of competition $p_n$. Specifically, we combine both Equation (ref) and the one-to-one mapping between the CDF of the observed $r^{th}$ order statistic and its parent distribution. In fact, the distribution of the observed $r^{th}$ order statistic is a mixture \[ F_r(x) = \sum_n p_n \sum_{i=r}^{n} c_{i,n} [\eta \check{F}(x)]^{i}[1-\eta \check{F}(x)]^{n-i}, \forall x \le \chi, \] whose left-hand side is known and right-hand side is linear in $p_n$ and polynomial in $\eta$. It allows us to construct a system of $|n|+1$ equations that provides identifying restrictions on both $p_n$s and $\eta$. We summarize this result in the following lemma.
The distribution of competition is of interest in many settings. Moreover, our approach avoids taking $z$ to the upper limit, which brings advantages in the setting with two sources of unobserved factors.
\paragraph{Two Sources of Unobserved Factors} We now examine the identification of ascending auctions allowing for both unobserved state and unobserved competition. The identification requires additional variation such as a binary instrumental variable besides the two consecutive order statistics because of the additional unobserved auction state $k$.
The identification exploits the following properties: (1) bids are independent across bidders in the same auction; (2) the joint density of consecutive order statistics admits a semi-separable structure; (3) we can lump the impact of unobserved competition into one component that only involves information about the $r^{th}$ order statistic due to the auction format. That is to say, we can treat the mixture with two sources of unobserved factors $(k,n)$ as a mixture with one source of unobserved state $k$ by lumping together all competition effects.
We first identify the state-specific density for the “low" segment and the component mixture of all competition in the “high" segment up to scales using eigendecomposition. We then follow the intuition for the scenario without unobserved states to pin down the scales and weight distribution and to identify the state-specific bid/value distributions.
Specifically, the joint distribution of these two consecutive order statistics together with the instrumental variable can be represented as a mixture of state $k$ and competition $n$. Specifically, if we only exploit the variation of the two order statistics $X_{r-1:n}=x$ and $X_{r:n}=y$ in the two ordered intervals $x\in l$ and $y\in h$, we have
We first present the condition for identifying the unobserved $K$.
Following Lemma (ref), if Assumption (ref) holds, then the unknown $K$ is identified by the rank of the observed joint distribution of the two consecutive order statistics even if competition is unknown. Therefore, we treat $K$ as known from now on and focus on identifying the state-specific bid/value distributions.
Following Lemmas (ref), we can identify state-specific bid functions $f^{k}_{r-1:r-1}(x)$ and $\bar f^k(y) \equiv \sum_n p_{k,n} c_{r-1,n} f^{k}_{1:n-r+1}(y)$ up to different scales and the same ordering, where $x\in l$ and $y\in h$. Consequently, we can identify the state-specific bid distribution $f^{k}(x)$, where $x \le \chi$, but cannot identify it in the “high" segment because competition is unobserved. To summarize, we identify the state-specific bid distributions up to the same ordering but different scales:
where $\check{f}_l^{k}(\cdot)$ and $ \check{\bar f}_h^{k}(\cdot)$ are the state-specific distributions identified from the above analysis for the “low" and “high" segments, respectively.
To fully identify the state-specific bid distributions, we need to tackle the following two issues: (1) pin down scales, i.e., $\eta_l^k$ and $\eta_h^k$; (2) identify the state-specific bid distribution for $y \ge \chi$. One possible solution to both issues is to let $\chi$ go to upper bound $\bar x$. However, taking $\chi$ to the limit is at odds with our setting. First, the density of the bottom order statistic $f_{1:n-r+1}^k(y)$, and hence mixture $\bar f^k(y)$, is zero at the limit, which leaves us little data to pin down the scale. Second, our identification strategy relies on the variation of $\bar f^k(y)$ in the $K$ exclusive points/intervals above $\chi$. Specifically, we need sufficient variation in the distribution at those $K$ points/intervals so that matrix $ \boldsymbol{H}$ is invertible and eigendecomposition is feasible. Lastly, the distribution of unobserved factors $p_{k,n}$ is of interest in our setting. Taking $\chi$ to the limit is only useful for identification of the value distributions.
Instead, we propose addressing the first issue of pinning down the state-specific scales and the probability of the combined unobserved factors, $p_{k,n}$, by plugging the identified marginal densities back into the marginal distributions of the $r^{th}$ order statistic. Specifically, we plug in the identified component distributions with the associated scales, i.e., $\eta_l^k \check{f}_l^{k}(x)$, into the CDF of order statistics $X_{r}$:
Intuitively, we can construct polynomial equations with those to-be-identified components as the multivariate unknowns. Altogether, we need to solve for $(K+K\times |n|-1$) unknowns (minus one due to the fact that $\sum_{k,n} p_{k,n}=1$). Note that we can construct a continuum of equations because the bids are continuous. We impose the following assumption to guarantee there exists a unique solution for the scales and weights.
Once the distribution of state and competition $p_{k,n}$ is identified, we can address the second issue of identifying the bid distributions for $y \ge \chi$ in the following two steps. First, the distribution of the mixture of competition is identified from Equation (ref) because density $f_{r-1:r-1}(x)$ is identified for $ x \le \chi$, given that scale $\eta_l^k$ is known. Second, we identify the state-specific bid distribution for $y \ge \chi$ using the fact that the identified state-specific mixture component is a monotonic function of the to-be-identified state-specific bid distribution (see Equation (ref)).
We summarize the identification result in the following theorem.
In this subsection, we briefly discuss a two-step estimation procedure involving estimating $K$ and the state-specific value distributions sequentially.\footnote{Another possibility is to estimate the value distributions jointly with the number of unknown $K$ by MLE that penalizes a larger $K$, which is also out of the scope of this paper.} To estimate $K$, we construct a finite set of discretizations to approximate the continuum counterpart and estate the rank of a general matrix via a sequential test. Once $K$ is estimated, we can estimate the state-specific bid distribution via eigenvalue decomposition or a semiparametric sieve Maximum Likelihood Estimation. We also briefly discuss the consistency of the two-step estimator.
\paragraph{Estimation of Cardinality} The identification of $K$ suggests an intuitive path to estimate the unknown $K$. We first introduce some notation. Let $D_{\tilde K}$ collect all possible discretizations with the number of intervals being $\tilde K$, let $d$ denote any discretization in the set, and let $\boldsymbol{J}^{d}_{l,h}$ denote the matrix constructed based on discretization $d$. From Lemma (ref), we have
Therefore, we can start by letting $\tilde K=2$. If $\max_{d \in D_{\tilde K} } \{rank( \boldsymbol{J}^{d}_{l,h})\}=\tilde K$, we should increase $\tilde K$ by 1 and then re-estimate the rank. This whole process continues until the maximum matrix rank stops growing with $\tilde K$. That is, we stop when $\max \{rank( \boldsymbol{J}^{\tilde K}_{l,h})\}=\tilde K-1$ and conclude that $K=\tilde K-1$.
However, two challenges arise for estimating the unobserved $K$. First, note that the set $D_{\tilde K}$ is continuous, so it is impossible to exhaust the list. Moreover, for identification purposes, we only need to show that there exists one discretization where the full column rank condition holds. However, the identification result is nonconstructive in the sense that it does not provide a straightforward path regarding how to construct such a discretization. One can only try as many discretizations as possible for each $\tilde K$ and hope that we can estimate the $\max_{d \in D_{\tilde K} } \{rank( \boldsymbol{J}^{d}_{l,h})\}$ consistently. The second challenge lies in the estimation of the rank of a general matrix given any discretization. In what follows, we deal with both challenges one by one.
In practice, we propose approximating the continuous set $D_{\tilde K}$ for each $\tilde K$ using a finite but large set involving the choices of the middle cutoff and grid points in both the $l$ and $h$ segments. We first determine the set of all possible values that the cutoff point $\chi$ can take on. Specifically, suppose we want to try $R_1$ different values of the cutoff $\chi$. Naturally, we should spread out those values in the full support to try to capture its variation, which increases the chance of satisfying the full rank condition. Therefore, we propose to determine these values using the quantiles of the bids. Specifically, we construct the set of cutoff points as $\Omega_{\chi}(R_1) = \{\tau(1/(R_1+1)),....\tau(R_1/(R_1+1))\}$, where $\tau(\cdot)$ indicates the quantile function of the observed bids. Given any cutoff from $\Omega_{\chi}$, we now determine the set of all possible discretizations in each segment by first determining the possible choices of the $\tilde K$ intervals, which is equivalent to choosing $\tilde K-1$ grid points. We denote the set including all possible values of the grid points as $\Omega_{l}$ and $\Omega_{h}$. Let us use segment $l$ as an illustration. Suppose we decide there are $R_l$ possible values that the $\tilde K-1$ grid points can choose from, where $R_l \ge \tilde K-1$. Once again, intuitively, it is informative to use quantiles to determine these grid points. That is, $\Omega_{l}\equiv \{\tau_l(1/(R_{l}+1)),....,\tau_l(R_{l}/(R_{l}+1))\}$, where $\tau_l(\cdot)$ is the quantile function adjusted based on bids in segment $l$, which generates $C^{R_{l}}_{\tilde K-1}$ possible combinations of the discretization in segment $l$. The same procedure can be conducted for segment $h$. Therefore, we have created $R \equiv R_1\times C^{R_{l}}_{\tilde K-1} \times C^{R_{h}}_{\tilde K-1}$ possible discretizations and denote the approximated set as $\tilde D_{\tilde K} \equiv \{d_r, j=1,...,R\}$, where $d$ denotes any discretization in the set.
Example: We now provide a simple example of the construction of $\tilde D_{\tilde K}$. Let $\tilde K=2$. Suppose $R_1=3$, indicating that we would like to try three different cutoff points: quantiles 0.25, 0.5, 0.75, respectively (i.e., $\Omega_{\chi}(R_1=3) =\{\tau(0.25), \tau(0.5), \tau(0.75)\}$). We use the median cutoff as an example for discretization of the segment $l$. Since $\tilde K=2$, we only need to determine one grid point, which will produce two exclusive intervals. Specifically, suppose we allow three different values of the grid points, $R_l=R_h=3$, so $\Omega_l=\{\tau_l(0.25), \tau_l(0.5), \tau_l(0.75)\}=\{\tau(0.125), \tau(0.25), \tau(0.375)\}$, and $\Omega_h=\{\tau_h(0.25), \tau_h(0.5), \tau_h(0.75)\}=\{\tau(0.625), \tau(0.75), \tau(0.875)\}$. Therefore, we have constructed the approximation of the original continuum set as $\tilde D_{\tilde K}=\{(i,j,k)\}_{i,j,k}$, where $i$ refers to the middle cutoff, and $j,k$ the grid point chosen for segment $l$ and $h$, respectively. The cardinality of the constructed set can be computed as $R=3*3*3=27$, because each $i,j,k$ can be chosen from three possible values.
We now discuss how to estimate the rank of matrix $ \boldsymbol{J}^{d}_{l,h}$ for a given discretization $d$. Specifically, we first estimate every element in the joint probability matrix with $\tilde K$ intervals in each segment using a simple frequency estimator. That is,
Given the probability matrix, we estimate the rank of matrix $ \boldsymbol{ J}^{d}_{l,h}$ using a sequence of tests, following robin2000tests. Specifically, we construct the hypotheses as: $H^r_0: rank(\boldsymbol{J}^{d}_{l,h})=r$ against the alternatives $H^r_1: rank(\boldsymbol{J}^{d}_{l,h}) > r$ with $r=1,2,...,\tilde K$. The sequence of tests proceeds as follows. First, we start with the null hypothesis that the rank of matrix $\boldsymbol{J}^{d}_{l,h}$ is 1. If such a null hypothesis is rejected, we augment $r$ by one and repeat the test. When we fail to reject the null of $rank(\boldsymbol{J}^{d}_{l,h})=r$ for the first time, the rank of $\boldsymbol{J}^{d}_{l,h}$ is estimated as $r$.
Therefore, the unknown $K$ can be estimated as
Note that such an estimator does not require an optimization algorithm, so it is easy to compute. Such an estimator is unconventional, and the consistency of such an estimator is not trivial. Two issues need to be resolved for consistency. The first issue is determining how to guarantee the rank estimator is consistent. This problem is well studied in the existing literature (see robin2000tests). The consistency of such an estimator relies on the selection of the significance levels for the sequential tests. Specifically, as sample size $M$ increases, significance level $\alpha_{M}$ should go to infinity but at a lower rate. The second issue is ascertaining how well our proposed simplified set $\tilde D_{\tilde K}$ approximates the set of all possible discretization $D_{\tilde K}$. The answer to this question depends on the specification of the discretization. Intuitively, the more experiments we try, i.e., the larger of $R_1$, $R_{l}$, and $R_{h}$, the better the approximation is. Note that given one discretization, we can estimate the rank fairly quickly, since no optimization algorithm is needed. Thus, in practice, one could try very large $R_1$, $R_{l}$, and $R_{h}$. Intuitively, if $R_1$, $R_{l}$, and $R_{h}$ go to infinite as the sample size goes to infinity, $\tilde D_{\tilde K}$ should approximate the true $D_{\tilde K}$ very well. However, establishing such a theoretical result is outside the scope of this paper, so we leave it for future research. Note that the estimation of $K$ serves as a model selection procedure. Therefore, “if the selection is consistent, i.e., $prob(\hat K_n=K) \rightarrow 1$ as $M \rightarrow \infty$, the asymptotic properties of any statistic based on the true model and the selected model are identical and hence asymptotic inference is unaffected,"(Lemma 1, potscher1991effects).
\paragraph{Estimation of the State-specific Value Distribution} Once $K$ is estimated, we can estimate the state-specific bid/value distribution, given the discretization, by following the identification steps one by one since it is constructive. Suppose there is a subset of $R$ experiments where the rank of the joint probability matrix is the same as $\hat K$. This suggests that the model is over-identified. We can estimate the state-specific bid/value distribution by every discretization with the joint probability matrix being full rank. In this case, we can leverage such over-identification power and improve estimation efficiency by averaging over all those estimators.
However, estimation following the identification strategy involves multiple steps, with the eigendecomposition being the first step. Once the decomposition is achieved, we need to estimate the density of the order statistics, then invoke the one-to-one mapping to estimate the parent distribution from the density of its order statistic, and then pin down the scales. Such a sequential estimation procedure is usually inefficient. Therefore, we propose to directly estimate the state-specific bid/value distribution using a semi-parametric sieve estimator. Specifically, we can approximate each state-specific value distribution using a sieve base. Therefore, we just need to estimate $\hat K \times L$ sieve coefficients, where $L$ is the number of base functions chosen. Specifically, we first introduce the Bernstein series for semiparametrically estimating the underlying state-specific bid distribution $f^k(x)$. We use the following specification to approximate the unknown bid distribution, which is the value distribution in the scenario of ascending auctions. That is,
where $L_k$ is a smoothing parameter that increases with sample size, and $b_k \equiv \{b_{1k},...,b_{L_kk}\}$ is the vector of sieve parameters for state $k$ to be estimated. As $ f^k(x;b_k)$ is a density function, it has to be non-negative, and its integration over the domain $[0,1]$ is 1. That is, $b_{1k} \ge 0$ and $\sum_l b_{lk}=1$. The CDF of the state-specific bid distribution can then be approximated as: \[\check F^k(x;b_k) \equiv \int^x_{0} \check f^k(x;b_k)d x \simeq \int^x_{0} f^k(x)d x =F^k(x).\]
Let $\theta_{sieve}$ collect all sieve parameters, i.e., $\theta_{sieve} \equiv \{b_{1k},...,b_{L_kk}, p_k, \Pr(W|k)\}_{k}$, where $k=1,...,K$. The semi-parametric sieve estimator $\hat \theta_{sieve}$ is the one that maximizes the log likelihood of the joint distribution of the two observed order statistics and the instrumental variable. That is,
ghosal2001convergence proves that the convergence for semi-parametric sieve MLE with Bernstein polynomial base functions is at “nearly parametric rate" $\sqrt {\frac{\log M}{M}}$ for Hellinger distance, when the true density is indeed a Bernstein density. When the true density is not of the Bernstein type, the sieve estimator converges at rate $(\frac{\log M}{M})^{1/3}$ if the true density is twice differentiable and bounded away from 0.
Auction data often fail to record all bids or all relevant auction-specific characteristics that shift bidder values. Instead, they may contain only a subset or order statistics of the bids and suffer from unobserved heterogeneity (UH). In this paper, we present a set of new identification results for auction models with discrete UH using consecutive order statistics. In particular, we show that despite correlation between order statistics, employing the same number of measurements is sufficient for achieving identification.
Mixture models for UH usually rely on independence assumptions. Our results show that exploring the statistical/model structure is a fruitful approach to restore identification when independence fails. While this paper focuses on IPV auction models and finite UH, natural extensions include affiliated private value and continuous UH, as seen in li2002structural, LV1998, and K2011. Moreover, UH and data truncation arise in other settings, such as beauty contests and war of attrition models, where many players compete for multiple prizes. We leave these for future research.