EconBase
← Back to paper

Deconvolution from two order statistics

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

60,050 characters · 11 sections · 72 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Deconvolution from two order statisticsWe thank Gaurab Aryal, Christian Gouri\'eroux, Philip Haile, Lixiong Li, Jos\'e Luis Montiel Olea, Chen Qiu, Cailin Slattery, and conference participants for helpful comments. We thank Cristi\'an Hern\'andez, Daniel Quint, and Christopher Turansick for sharing their code.

abstractEconomic data are often contaminated by measurement errors and truncated by ranking. This paper shows that the classical measurement error model with independent and additive measurement errors is identified nonparametrically using only two order statistics of repeated measurements. The identification result confirms a hypothesis by athey2002identification for a symmetric ascending auction model with unobserved heterogeneity. Extensions allow for heterogeneous measurement errors, broadening the applicability to additional empirical settings, including asymmetric auctions and wage offer models. We adapt an existing simulated sieve estimator and illustrate its performance in finite samples.

\noindentKeywords: Measurement error, Order statistics, Nonparametric identification, Spacing, Cross-sum

\thispagestyle{empty}

Introduction

We consider the classical measurement error model with repeated measurements:

align*[align* omitted — 73 chars of source]

where the latent variable of interest $\xi$ is measured $n$ times with i.i.d. measurement errors $(\varepsilon_j)_{j=1,\ldots,n}$ that are independent of $\xi$. The identification result under such a model is well-established when at least two repeated measurements are observed, known as Kotlarski's lemma. The result has been widely applied in econometrics since its introduction by li1998nonparametric.\footnote{Examples include bonhomme2010generalized, krasnokutskaya2011identification and grundl2019identification. An outline of uses of Kotlarski's lemma in econometrics is in schennach2016recent.} In practice, the researcher may observe only a subset of order statistics of the measurements, i.e., $(X_{(j)})_{j \in J}$, where $X_{(j)}$ is the $j$\textsuperscript{th} smallest order statistic from a sample of size $n$ and $J \subset \{1,\ldots,n\}$. This paper shows the underlying probability distributions of the latent variable and the measurement errors are identified from the joint distribution of the ordered measurements $(X_{(j)})_{j \in J}$.

The problem took its shape in athey2002identification as a nonparametric identification problem in ascending auctions, which has particular relevance in auctions with electronic bidding. They conjecture that the model consists of enough structure to attain point identification using two order statistics. However, the question has been long-standing for two decades.\footnote{Some positive findings are made recently in the finite mixture context; see mbakop2017identification using five order statistics, luo2023identification using two and an instrument and luo2021order using three. However, these papers do not tackle the original conjecture in the spirit of Kotlarski's lemma.} Importantly, the standard approach using Kotlarski's lemma fails because of dependence in the order statistics of measurement errors. Further, the attempt to difference out the latent variable and exploit the spacing distribution is shown to be insufficient for point identification, evidenced by rossberg1972characterization's counterexample.

While the literature has documented the increasing importance of unobserved heterogeneity in auctions,\footnote{For example, see aradillas2013identification, krasnokutskaya2011identification, and li2000conditionally.} the lack of identification results in the existing literature has hindered allowance for such heterogeneity in the classical fashion, unless relying on additional external variations in the data.\footnote{For example, hernandez2020estimation uses variation in the number of bidders across auctions. freyberger2022identification circumvents the issue of dependence between order statistics using observed reserve prices, rendering the identification problem classical.} We fill this important gap by showing that the underlying distributions are identified with only data on two order statistics, which need not be consecutive or extreme, without relying on extra variations.

We propose a new identification strategy that exploits both features in kotlarski1967characterizing and rossberg1972characterization. In particular, we make use of the within independence of the latent variable and the additive-separable measurement errors as in kotlarski1967characterizing to derive that the model imposes an additional restriction beyond the spacing of order statistics.\footnote{The term spacing usually refers to the difference between two consecutive order statistics. Here, we use it more broadly as the difference between any two order statistics.} More precisely, we find that two measurement error distributions both rationalize the observed distribution of the ordered measurements only if the joint distributions of spacing and cross-sum are identical.\footnote{The term cross-sum refers to the sum of two order statistics of two different ranks, where the order statistics are of two independent random samples from distinct underlying parent distributions.} Exploiting the structure of order statistics and commonly seen conditions (a support or tail restriction), we show that the spacing in conjunction with this additional restriction is indeed sufficient to point identify the underlying distribution using any two order statistics. The latent variable distribution is subsequently identified by a standard deconvolution argument.

We extend our main result to the setting where measurement errors are independent but nonidentically distributed (i.n.i.d.). We show that the underlying distributions are identified when two order statistics are observed, provided some measurement errors are known to be identically distributed and the data consist of group identities of high-order statistics. The result applies to asymmetric ascending auctions, where only high-order dropout bids are recorded, and asymmetric first-price auctions and wage offer settings, where consecutive and extreme order statistics are observed.

The outline of the paper is as follows. Section (ref) introduces two motivating examples and Section (ref) presents the identification results. To complement the identification result, in Section (ref) we adapt bierens2012semi's simulated sieve estimator to consistently estimate the unknown distributions. Section (ref) concludes.

Motivating Examples

To motivate the identification problem, we introduce two examples: an ascending auction model with auction-specific unobserved heterogeneity and a model of wage offers where wage is determined by worker-specific labor productivity. Throughout the paper, we focus only on unobserved heterogeneity since adding exogenous covariates does not require novel considerations.

example[Ascending auction] Consider an ascending auction with one indivisible good and $n$ bidders. Each bidder's valuation for the item is determined by a set of item features---summarized and denoted by $\xi \in \mathbb{R}$---and a private value component $\varepsilon_j$ that is independent of $\xi$ and across bidders. The valuation of bidder $j$ is given by $X_j = \xi + \varepsilon_j$. Note that the valuation of each bidder can be cast as a measurement of the value $\xi$ with measurement error $\varepsilon_j$. See, e.g., athey2002identification. In an ascending auction, the price rises until only one bidder remains. Suppose the auctioneer records prices $P_1 \le P_2 \le \cdots$ at which bidders drop out. Assuming that the bidders play the dominant strategy of remaining in the auction until the price reaches their valuations, the dropout bids reveal the ordered valuations of the bidders, i.e., $P_j = X_{(j)}$. Importantly, the highest valuation $X_{(n)}$ is never observed since the auction ends when the price reaches $P_{n-1} = X_{(n-1)}$, and the bidder with the highest valuation wins. Therefore, in an ascending auction, we observe only an incomplete set of order statistics on bidders' valuations. For example, kim2014nonparametric observe at least the three highest dropout prices in ascending used-car auctions; see also larsen2021efficiency for used-car auctions in the U.S. Data on timber auctions by the U.S. Forest Service---analyzed in a number of empirical papers, e.g., haile2003inference, athey2011comparing, and aradillas2013identification to name a few---records at most top twelve bids regardless of the number of bidders. Both the distributions of $\xi$ and $\varepsilon_j$ are of interest in an empirical analysis of auctions. For example, the counterfactual expected revenue to the auctioneer under a hypothetical auction rule requires the knowledge of both distributions.
example[Wage offer determination] Consider a simple wage offer model $X_j = \xi + w_j + \varepsilon_j$ where $X_j$ is the $j$\textsuperscript{th} (log-)wage offer an individual receives. $X_j$ is assumed to be composed of the worker's productivity $\xi$, the wage rate $w_j$, and measurement error $\varepsilon_j$ (on, e.g., labor productivity) associated with the offer. It is assumed that the period in consideration is relatively short such that the worker's productivity remains unchanged, but the worker may receive multiple offers; see guo2021identification. Data from the Survey of Consumer Expectations (SCE) Labor Market Survey records salary-related responses of all job offers for those who received at most three offers and responses to three best offers for those who received more than three offers within the last four months. In this setting, the researcher observes $X_{(n)},\ldots,X_{(n-k+1)}$ where $k = \min\{3,n\}$ for each respondent. The identification results in the current paper show that the distributions $F_\xi$ and $F_{w_j+\varepsilon_j}$ are nonparametrically identified.

The Model and Main Results

In this section, we first formalize the i.i.d. framework considered throughout the paper and provide sufficient conditions to identify the latent variable and measurement error distributions. In Section (ref), we consider i.n.i.d. measurement errors.

The Setup

We begin by stating the sampling process.

assumption[Sampling process] For each $1 \le j \le n$, $X_{j} := \xi + \varepsilon_{j}$, where the random variable $\xi$ is independent of the random vector $(\varepsilon_{1},\ldots,\varepsilon_{n})$.\footnote{The identification results herein also applies to a setting with multiplicatively separable measurement error by taking logs when $\xi$ and $\varepsilon_j$'s are positive, as in Example (ref).}

The researcher observes $(X_{(r)},X_{(s)})$, the $r$\textsuperscript{th} and $s$\textsuperscript{th} order statistics from $n$ observations, where $r$, $s$, and $n$ are known and $1 \le r < s \le n$. That is, while the total number of measurements is known, one only observes two measurements of known ranks. In this section, we identify the unknown distributions of $\xi$ and $(\varepsilon_{1},\ldots,\varepsilon_{n})$ using $F_{X_{(r,s)}}$, the joint distribution of $(X_{(r)},X_{(s)})$, which is estimable from a random sample of the two order statistics.

For the benchmark case, we make the following distributional assumptions.

assumption[i.i.d. errors] \leavevmode \begin{enumerate}[label=(\alph*)] • $\varepsilon_{1},\ldots,\varepsilon_{n}$ are i.i.d. with a common distribution $F_\varepsilon$ on $\mathbb{R}$. • $F_\varepsilon$ is absolutely continuous with a probability density function $f_\varepsilon$ that is light-tailed, i.e., for some $C > 0$, $f_\varepsilon(\epsilon) = O(e^{-C\lvert\epsilon\rvert})$ as $\lvert\epsilon\rvert \rightarrow \infty$. \end{enumerate}

The tail condition in Assumption (ref)(b) is trivially satisfied when the support is bounded. If the support is unbounded on either side, the assumption restricts the tail of the density function to decay at an exponential rate.\footnote{Throughout the paper, we use the term “light-tailed” to refer to densities with exponentially decaying tails as stated in the assumption. } The same assumption is found in evdokimov2012some. Alternatively, we may follow kotlarski1967characterizing and miller1970characterizing to assume (a.e.-)nonvanishing or analytic characteristic functions (ch.f.) of both $\xi$ and $\varepsilon$. Assumption (ref)(b) is a sufficient condition for the ch.f. of $F_\varepsilon$ to be analytic. However, to our knowledge, there exists no result regarding its implication on the joint ch.f. of order statistics, the property we need for our identification results. In Lemma (ref), we show that the joint ch.f. of the order statistics $(\varepsilon_{(r)},\varepsilon_{(s)})$ (and thus its marginals) is also analytic under this assumption.

In addition to Assumption (ref), we further restrict the support of the measurement errors as follows. For any distribution $F$ on $\mathbb{R}$, let $S(F) \subseteq \mathbb{R}$ denote its support.\footnote{The support of $F$ is defined as the smallest closed set $K \subseteq \mathbb{R}$ such that $P_F(K) = 1$, where $P_F$ is the Borel probability measure induced by $F$. Thus, the density $f(x)$ may be zero for some values $x \in K$, including points on the boundary of $K$.}

assumption[Support condition] The measurement errors are bounded from below, which is normalized to zero, i.e., $\inf S(F_\varepsilon) = 0$.

Two aspects are worth mentioning regarding the above assumption. First, we require the measurement error to be bounded at one end; nevertheless, we allow the support to be possibly unbounded from above.\footnote{The results herein apply to the case when only the upper bound is finite, e.g., $S(F_\varepsilon) = (-\infty,0]$, by considering $(X_{(r')}',X_{(s')}^\prime) = (-X_{(s)},-X_{(r)})$ where $r' = n-s+1$ and $s'=n-r+1$. } We do not assume the upper bound of the support is known nor assume any additional structure on the support, e.g., connectedness. Second, as is typical with measurement error models, the underlying distributions are identified only up to location. Although one may alternatively normalize the mean and have an unknown but finite lower bound, normalizing the lower bound turns out to be more convenient for our identification argument. Our identification strategy heavily utilizes the condition that one boundary of the support of $F_\varepsilon$ is finite. On the other hand, the distribution of $\xi$ is left completely unspecified.

The support restriction is nonrestrictive in many applications. The literature on games with incomplete information typically assumes bounded support of agent types (athey2001single), such as bidder valuation in empirical auctions (guerre2000optimal, athey2007nonparametric), private costs of exerting efforts in contest games (huang2021structural), and private information about variable costs in Cournot games (aryal2021empirical). Wage offers are also bounded below in job search models (burdett1998wage, guo2021identification).

Nonparametric Identification

Throughout the section, we treat the observable joint distribution $F_{X_{(r,s)}}$ as known (i.e., as a datum), and show that the distribution $F_\xi$ of the latent variable $\xi$ and $F_\varepsilon$ are identified under the aforementioned assumptions. The bulk of our identification result is concerned with showing that the measurement error distribution that can rationalize the observed distribution $F_{X_{(r,s)}}$ must be unique. Then the latent variable distribution is identified by a standard deconvolution argument. To achieve such a goal, we begin by isolating the empirical content on measurement errors from the datum. In Lemma (ref) below, we do so by exploiting the multiplicative structure of the ch.f. of a sum of independent random variables.\footnote{Similar ideas are also exploited in the classical repeated measurement error setting. See hall2003inference and evdokimov2012some.}

In the following lemmas, let $\eta_1,\ldots,\eta_n$ be $n$ i.i.d. copies of $\eta \in \mathbb{R}$ and let $(\eta_{(r)},\eta_{(s)})$ be order statistics. $\psi_{\eta_{(r,s)}}$ and $\psi_{\eta_{(r)}}$ denote the joint and marginal ch.f., respectively.

lemmaIf the sampling process is described as in Assumption (ref) and the measurement errors satisfy Assumption (ref), a measurement error distribution $F$ on $\mathbb{R}$ is data-consistent\footnote{We say a pair of distribution functions $(G_\xi,G_\varepsilon)$ rationalizes the data, or is data-consistent, if $\xi' \sim G_\xi$ and $(\varepsilon_j')_{j=1,\ldots,n} \sim \times_{j=1}^n G_\varepsilon$, independent of $\xi'$, implies $(\xi'+\varepsilon_{(r)}', \xi'+\varepsilon_{(s)}') =_d (X_{(r)},X_{(s)})$. We say $G_\xi$ (resp., $G_\varepsilon$) is data-consistent if $(G_\xi,G_\varepsilon)$ is data-consistent for some $G_\varepsilon$ (resp., $G_\xi$).} only if \begin{align} \frac{\psi_{X_{(r,s)}}(t_r,t_s)}{\psi_{X_{(j)}}(t_r+t_s)} = \frac{\psi_{\eta_{(r,s)}}(t_r,t_s)}{\psi_{\eta_{(j)}}(t_r+t_s)} , \quad for all $(t_r,t_s) \in B_0$, j \in \{r,s\} , \end{align} where $\eta \sim F$ and $B_0$ is an open ball in $\mathbb{R}^2$ centered at zero. Further, $F$ induces a unique data-consistent latent variable distribution.

Suppose $F$ and $G$ are two measurement error distributions that are data-consistent. Lemma (ref) states that the order statistics $(\eta_{(r)},\eta_{(s)})$ from the parent distribution $F$ and $(\eta'_{(r)},\eta'_{(s)})$ from the parent distribution $G$ must have the same ratio of joint and marginal ch.f.s. To obtain identification, it remains to show that such a measurement error distribution is unique, i.e., $F=G$. In contrast to the prevalent identification strategy in the existing literature that expands on kotlarski1967characterizing, which begins with identifying the distribution of the latent variable $\xi$, Lemma (ref) suggests an identification strategy that first focuses on the measurement error distribution.

Note that without any assumption on the latent variable $\xi$, the ch.f. of $\xi$ may vanish on a nontrivial interval bounded away from the origin.\footnote{Recall that a ch.f. takes value one at the origin and is uniformly continuous, implying that it does not vanish on some neighborhood about the origin.} Hence, in general, the ratio is only guaranteed to be well-defined on a neighborhood about the origin. Analyticity of the ch.f. of measurement error order statistics (Lemma (ref)) thus plays a crucial role in pinning down the entire distribution based only on the information of the ch.f. about the origin. In particular, Lemma (ref) implies that any two data-consistent measurement error distributions must satisfy

align*[align* omitted — 155 chars of source]

for all $(t_r,t_s) \in B_0$. The equality extends to all of $\mathbb{R}^2$ because a product of two analytic functions is analytic. Finally, noticing that the ch.f. of the sum of two independent random vectors is multiplicatively separable, the following lemma establishes necessary conditions for any two data-consistent measurement error distributions.

lemmaLet $\eta_{(r)}$ and $\eta_{(s)}$ be two order statistics of a random sample with size $n$ from $F$, and $\eta_{(r)}^\prime$ and $\eta_{(s)}^\prime$ be two order statistics of a random sample with size $n$ from $G$, where $(\eta_{(r)},\eta_{(s)})$ and $(\eta'_{(r)},\eta'_{(s)})$ are independent. Under Assumptions (ref), if measurement error distributions $F$ and $G$ admit light-tailed densities,\footnote{Cf. Assumption (ref)(b) and footnote (ref).} they both rationalize the data only if \begin{align} Z_{1j} := \begin{pmatrix} \eta_{(j)}^\prime + \eta_{(r)} \\ \eta_{(j)}^\prime + \eta_{(s)} \end{pmatrix} \stackrel{d}{=} \begin{pmatrix} \eta_{(j)} + \eta_{(r)}^\prime \\ \eta_{(j)} + \eta_{(s)}^\prime \end{pmatrix} =: Z_{2j} , \quad j \in \{r,s\} . \end{align}

The equivalence of ratios of ch.f.s (Lemma (ref)) implies an equivalence of two joint distributions of particular linear combinations of order statistics that arise from the two measurement error distributions. The linear combinations not only expose the empirical content contained in the spacing of errors (see Corollary (ref)) as investigated by hall2003inference among many others in the classical repeated measurements setting and by rossberg1972characterization in the context of order statistics, but they also imposes restrictions on the sum of order statistics from potentially distinct parent distributions, which we call a cross-sum.

Because both $Z_{1r}$ and $Z_{2r}$ involve three order statistics that are associated with potentially different parent distributions $F$ and $G$, it appears to be difficult to establish the equivalence of $F$ and $G$ directly from the distributional equivalence of $Z_{1r}$ and $Z_{2r}$. A reasonable attempt to tackle the problem would be to explore features of the distributions of $Z_{1r}$ and $Z_{2r}$ that depend on the parent distributions in a simple manner. Under the support condition in Assumption (ref), we show that the distributions of $Z_{1r}$ and $Z_{2r}$ depend on only one of the two parent distributions upon conditioning on an extreme event.\footnote{The identification argument here and below investigates the implications of the condition $Z_{1r} =_d Z_{2r}$ only. $Z_{1s} =_d Z_{2s}$ in (ref) is the relevant condition to exploit under the assumption, in place of Assumption (ref), that the measurement error is bounded from above (cf. footnote (ref)).} Thus the joint distribution of the cross-sum and spacing together with Assumption (ref) provide information sufficient to claim equivalence of two data-consistent measurement error distributions and point identify $F_\varepsilon$.

lemmaLet $\eta_{(r)}$, $\eta_{(s)}$, $\eta_{(r)}^\prime$ and $\eta_{(s)}^\prime$ be as specified in Lemma (ref). If measurement error distributions $F$ and $G$ admit light-tailed densities and $\inf S(F) = \inf S(G) = 0$, we have, for all $c \in \mathbb{R}$, \begin{align} \lim_{\delta \downarrow 0} \mathbb{P}(\eta_{(r)}^\prime + \eta_{(s)} \le c \ | \ \eta_{(r)}^\prime + \eta_{(r)} \le \delta) &= F_{s-r:n-r}(c) , \\ \lim_{\delta \downarrow 0} \mathbb{P}(\eta_{(r)} + \eta_{(s)}^\prime \le c \ | \ \eta_{(r)} + \eta_{(r)}^\prime \le \delta) &= G_{s-r:n-r}(c) , \end{align} where $F_{s-r:n-r}$ (resp., $G_{s-r:n-r}$) denotes the distribution of the $(s-r)$\textsuperscript{th} order statistic of a random sample with size $n-r$ from $F$ (resp., $G$).

For the sake of intuition, suppose that conditioning on the event $\{\eta_{(r)}^\prime + \eta_{(r)} = 0\}$ is well-defined. The condition implies that both $\eta_{(r)}^\prime = 0$ and $\eta_{(r)} = 0$. Thus, the independence between the pairs of order statistics simplifies the conditional distribution to $\mathbb{P}( \eta_{(s)} \le c \ | \ \eta_{(r)} = 0)$. Standard arguments in order statistics (cf. Theorem 2.4.2 in arnold2008first) suggest that this conditional distribution has a simple form: \[ \mathbb{P}( \eta_{(s)} \le c \ | \ \eta_{(r)} = 0) = F_{s-r:n-r}(c) . \] In other words, the equality above says that if the lowest $r$ draws all take the minimal possible value, then the distribution of any higher order statistic must equal the distribution of order statistic when the $r$ draws at the minimal value are discarded.

However, the event $\{\eta_{(r)}^\prime + \eta_{(r)} = 0\}$ may not be well-defined as it occurs with zero probability. In particular, the event $\{\eta_{(r)} = 0\}$ always has zero density (regardless of the density of $\eta$) unless $r=1$, i.e., the minimum order statistic. Further, because the true density $f_\varepsilon(0)$ may be zero, even the minimum order statistic may have zero density. In Lemma (ref), we make the intuitive claim made above more rigorous by considering the limiting argument as in (ref) and (ref), which is well-defined.

Lemma (ref) implies that if $F$ and $G$ are both data-consistent, then the conditional distributions (ref) and (ref) in Lemma (ref) must be equal, i.e., for all $c \in \mathbb{R}$, \[ F_{s-r:n-r}(c) = G_{s-r:n-r}(c) . \] As the distribution of an order statistic uniquely identifies the parent distribution, we can conclude that $F = G$.\footnote{The one-to-one mapping between the distribution of an order statistic and the parent distribution is a standard result, e.g., see david2003order, p.10, (2.1.5). The mapping in (2.1.5) is an invertible map of the c.d.f. $F$ because $p \mapsto I_p(a, b)$ is strictly monotone.} Combining Lemmas (ref)--(ref) leads to the following main identification result for the i.i.d. case.

theoremUnder Assumption (ref), if the measurement errors satisfy Assumptions (ref) and (ref), both $F_\xi$ and $F_\varepsilon$ are identified.

Discussion on rossberg1972characterization's Counterexample

athey2002identification pointed out that an identification strategy based on the spacing between order statistics does not yield a positive result. The discussion therein relied on a counterexample by rossberg1972characterization. A straightforward corollary to Lemma (ref) highlights the model restrictions we exploit in addition to the spacing.

corollaryUnder the same assumptions as in Lemma (ref), measurement error distributions $F$ and $G$ both rationalize the data only if \begin{align*} \begin{pmatrix} \eta_{(s)} - \eta_{(r)} \\ \eta_{(r)}^\prime + \eta_{(s)} \end{pmatrix} \stackrel{d}{=} \begin{pmatrix} \eta_{(s)}^\prime - \eta_{(r)}^\prime \\ \eta_{(r)} + \eta_{(s)}^\prime \end{pmatrix} \quad and \quad \begin{pmatrix} \eta_{(s)} - \eta_{(r)} \\ \eta_{(r)} + \eta_{(s)}^\prime \end{pmatrix} \stackrel{d}{=} \begin{pmatrix} \eta_{(s)}^\prime - \eta_{(r)}^\prime \\ \eta_{(r)}^\prime + \eta_{(s)} \end{pmatrix} . \end{align*}

In words, the observational equivalence between any two measurement error distributions in our setting implies the equivalence of the joint distributions of the spacing and cross-sum of order statistics. This finding complements the nonidentification result in rossberg1972characterization that the distribution of spacing between two order statistics alone cannot identify their parent distribution. The additional constraint on the distribution of cross-sum absent in rossberg1972characterization emerges from the within independence assumption. Figure (ref) in Appendix (ref) shows that, while the spacing distributions for the exponential and Rossberg's counterexample overlap, the cross-sum distributions differ over nontrivial regions. See Appendix (ref) for more discussion.

Extensions of the Identification Result

We extend our identification result to the case of independent but nonidentically distributed (i.n.i.d.) measurement errors. To motivate the extension to the problem, we expand on Example (ref) by introducing bidder asymmetry and discuss additional features available in some auction data.

example[Asymmetric auction] In empirical auctions, bidders are said to be asymmetric if the bidders' private values $\varepsilon_1,\ldots,\varepsilon_n$ have different marginal (or parent, in the case of order statistics) distributions. Such asymmetry arises naturally in procurement auctions, where contractors differ in cost efficiency and productivity (see, e.g., flambard2006asymmetry), timber auctions, where mills have the manufacturing capacity and loggers do not (see, e.g., athey2011comparing)\footnote{In a wage offer setting as in Example (ref) employers may value the same productivity differently, which may result in heterogeneous measurement errors.}. Due to asymmetry, identifying bidder-specific valuation distributions requires knowing their identities, i.e., knowing who participated in each auction. Fortunately, the auctioneer often publishes such identities despite the highest valuation being missing; see, e.g., athey2011comparing. Another common practice in the asymmetric auction literature is grouping bidders by commonly known bidder types, such as mills and loggers in athey2011comparing and strong and weak bidders in luo2018auctions, which leads to more tractable theory and empirics. Such grouping uses additional information about the bidders, typically available as a public record, such as bidders' manufacturing capacity in a timber auction and pre-qualified contractors' experience in Department of Transportation procurement auctions.

In this section, we first show that any two order statistics suffice to point identify the underlying distributions when there is a relatively small number of groups and the group identities of high-order statistics are observable. We then discuss several relevant applications in empirical auctions.

assumption[i.n.i.d. errors] \leavevmode \begin{enumerate}[label=(\alph*)] • $\varepsilon_{1},\ldots,\varepsilon_{n}$ are independent with distributions $F_{\varepsilon_1},\ldots,F_{\varepsilon_n}$ on $\mathbb{R}$, respectively. • For each $j \in \{1,\ldots,n\}$, $F_{\varepsilon_j}$ admits a probability density function $f_{\varepsilon_j}$ that is light-tailed, i.e., $f_{\varepsilon_j}(\epsilon) = O(e^{-C_j\lvert\epsilon\rvert})$ as $\lvert\epsilon\rvert \rightarrow \infty$ for some $C_j > 0$. • For each $j \in \{1,\ldots,n\}$, $\inf S(F_{\varepsilon_j}) = 0$. \end{enumerate}

Assumption (ref)(c) assumes both finite and common support lower bound across distinct measurement error distributions, the latter of which holds trivially in the i.i.d. case. Note that apart from the lower boundary, the supports may differ (cf. Corollary (ref)). The cost of relaxing the homogeneity assumption translates to more stringent data requirements. We require observing indices or types associated with all high-order statistics (for ranks above $r$). The following assumption records any prior information on the different types of measurement errors.

assumption[Group structure] There exists a partition $g_1,\ldots,g_p$ of $\{1,\ldots,n\}$ such that $\varepsilon_j =_d \varepsilon_k$ if $j,k \in g_q$ for some $q \in \{1,\ldots,p\}$. In addition, the measurement errors $\varepsilon_1,\ldots,\varepsilon_n$ have common support.

Assumption (ref) posits there are at most $p$ distinct measurement error distributions. The case $p=1$ corresponds to the i.i.d. case in Section (ref); and $p=n$ corresponds to the case with no prior information on homogeneity. We do not preclude the possibility that two groups $g_q$ and $g_{q'}$ have the same distribution beyond the researcher's knowledge. E.g., the extreme case $p=n$ subsumes the i.i.d. setting as a special case. Assumption (ref) implicitly assumes that the group structure is held constant across observations. In an application to auctions, the assumption may be imposed by restricting to auctions with the same composition of bidder types when all bidder types of participants are available in the dataset. In such a case, the analysis should be interpreted conditionally on the bidder composition.

The common support assumption is a sufficient condition to guarantee that all measurement error distributions are identified on the entirety of their support. Otherwise, some distributions may be identified only on a strict subset of their support, as we highlight in a discussion below. Let $R_{(j)} = \{q : X_{(j)} = X_k \text{ for some } k \in g_q\}$ be the group identity of the $j$\textsuperscript{th} order statistic.\footnote{Under Assumption (ref)(b), $R_{(j)}$ is singleton with probability $1$. Thus, we abuse notation and use $R_{(j)}$ to denote both the set and the a.s. unique element in $\{1,\ldots,p\}$.}

theoremSuppose Assumptions (ref), (ref), and (ref) hold. Further, suppose two order statistics and the group identities of high-order statistics $(X_{(r)},X_{(s)},R_{(r+1)},\ldots,R_{(n)})$ are observed. If there exists a group $g_q$ with at least $n-r$ members, both $F_\xi$ and $(F_{\varepsilon_j} : j=1,\ldots,n)$ are identified.

Note that the identification result does not require the researcher to observe the group identity $R_{(r)}$ of the $r$\textsuperscript{th} order statistic or those of lower order statistics. An heuristic explanation is provided in a discussion below. Theorem (ref) has an important implication when the measurement errors are left completely heterogeneous, i.e., $p=n$. In particular, the theorem implies that the observed order statistics should be consecutive and extreme because there is no group of a larger size, i.e., $n-r=1$.\footnote{The support condition in Assumption (ref) plays no role in this corollary because the observed order statistics are the largest among all measurement errors.}

corollarySuppose Assumptions (ref) and (ref) hold. Further, suppose $(X_{(n-1)},X_{(n)},R_{(n)})$ is observed. Both $F_\xi$ and $(F_{\varepsilon_j} : j=1,\ldots,n)$ are identified.

So as to appreciate the additional conditions assumed in Theorem (ref), we first illustrate the identification strategy in the setting of Corollary (ref), which delivers a simpler argument. Suppose results similar to Lemmas (ref)--(ref) hold in the i.n.i.d. case, in the sense that two sets of data-consistent measurement error distributions $(F_j : j=1,\ldots,n)$ and $(G_j : j=1,\ldots,n)$ must satisfy

align[align omitted — 174 chars of source]

for every constant $c$. Had the measurement errors been i.i.d., (ref) corresponds the equivalence of two parent distributions. In the i.n.i.d. case, the top-order statistic may arise from any of the parent distributions. Intuition suggests that this should be a mixture of measurement errors from different parent distributions. Since the mixing weights depend on $(F_j : j=1,\ldots,n)$, it appears formidable to show its equivalence with $(G_j : j=1,\ldots,n)$ from (ref) without additional assumptions. Alternatively, suppose---as we formally show in the proof of Theorem (ref)---data-consistency implies a condition similar to (ref) conditional on the member index of the highest order statistic. Heuristically speaking, because the index is known, say $j$, the conditional distribution simplifies to the $j$\textsuperscript{th} marginal distribution (i.e., a mixture with trivial weights). Thus $F_j = G_j$ and the $j$\textsuperscript{th} measurement error distribution is identified. Under the assumption of a common support lower bound, each measurement error has a nonzero chance of being the largest, which implies all $n$ measurement error distributions are identified.

Corollary (ref) does not grant identification of an ascending auction model when the data on dropout bids is incomplete and only a few high-order dropout bids as in, e.g., freyberger2022identification. Nonetheless, group structures as in Theorem (ref) are commonly present in the empirical auction literature.

To illustrate the identification strategy in the setting of Theorem (ref), suppose out of $n=4$ measurements, one observes $(X_{(2)},X_{(3)},R_{(3)},R_{(4)})$ and it is known a priori that two of the measurement errors have a common distribution (call it group 1 and the other groups 2 and 3). By conditioning on the event that $\{R_{(3)} = R_{(4)} = 1\}$, we can homogenize the conditional distribution of the 3\textsuperscript{rd}-order statistic of measurement errors conditional on the 2\textsuperscript{nd}-order statistic taking the smallest possible value, which allows us to identify the measurement error distribution for group 1. This rests on (i) being able to observe the group identities and (ii) group 1 being large enough to condition on such an event. Provided the measurement errors have a common support, the remaining group distributions can be identified sequentially. For example, the group-2 distribution can be identified by conditioning on $\{R_{(3)} = 2, R_{(4)} = 1\}$.\footnote{If one measurement error has larger support than another, the distribution may not be fully identified. If group-1 measurement errors have support $[0,1]$ but group 2 has support $[0,2]$, regardless of whether one conditions on $\{R_{(3)} = 1, R_{(4)} = 2\}$ or $\{R_{(3)} = 2, R_{(4)} = 1\}$, the 3\textsuperscript{rd}-order statistic only has support $[0,1]$. Thus group-2 measurement error distribution is not identified on $(1,2]$.}

remarkCorollary (ref) has natural applications in first-price auctions and wage offer settings, where consecutive and extreme order statistics are common and identities are observed. For instance, the Washington State Department of Transportation archives three apparent low bids and bidder identities of six months or older online. FDIC auction data contain the winning and second-highest bids and the associated identities (allen2023resolving). U.S. Forest Service timber auctions only record the top twelve bids and bidder identities regardless of the number of bidders. Lastly, data from the Survey of Consumer Expectations (SCE) Labor Market Survey records salary-related responses on the three best offers for those who received more than three offers within the last four months.
remarkIf all dropout bids are observed, Corollary (ref) is applicable to ascending auctions. Specifically, if private values have a common upper bound, i.e., $\sup S(F_{\varepsilon_j}) = \overline{\varepsilon} < +\infty$, the lowest order statistics $(X_{(1)},X_{(2)})$ identify the underlying distributions.

Nonparametric Estimation

We propose a simulated sieve estimator for the i.i.d. measurement error framework, which can be easily modified for the i.n.i.d. case, albeit with the curse of dimensionality. The procedure, motivated by the sieve estimator proposed in bierens2012semi, estimates the distribution functions of the latent variable and of the measurement error simultaneously.

A Simulated Sieve Estimator

We follow closely the development in bierens2008semi to specify the parameter space and its sieve space. It has the advantage that we can incorporate prior information on the support without having to choose different orthogonal bases. This is an attractive feature for our purpose because, unlike the latent variable $\xi$, the measurement errors are restricted to be nonnegative-valued. The sieve space is constructed using only Legendre polynomials instead of using, e.g., Hermite polynomials for the latent variable and Laguerre polynomials for the measurement errors.

The construction of the sieve space begins with the observation that any absolutely continuous c.d.f. $F$ on $\mathbb{R}$ can be expressed as $F(\cdot) = (H \circ G)(\cdot)$ where $H$ is an absolutely continuous c.d.f. on $[0,1]$ and $G$ is an absolutely continuous c.d.f. that is strictly increasing on $S(F)$. Equivalently,

align[align omitted — 123 chars of source]

where $h = \pi^2$ is the density of $H$ and $\pi$ is a Borel-measurable square-integrable function on $[0,1]$.\footnote{The density of $F$ may then also expressed as $f(\cdot) = (h \circ G)(\cdot) g(\cdot)$ where $g$ is the density of $G$.} Thus, e.g., with a fixed $G$ with support $S(G) = [0,\infty)$, a large enough set $\mathcal{P}$ of Borel-measurable square-integrable functions maps via (ref) to a large enough set of distribution functions with support contained in $[0,\infty)$.

The sieve space is constructed by using orthonormal polynomials to approximate $\pi$. bierens2008semi considers the following compact set of square-integrable functions

align[align omitted — 289 chars of source]

for some large constant $c > 0$, where $\rho_k$, defined on $[0,1]$, is a recentered and rescaled Legendre polynomial of order $k$: if $\ell_k$ is a Legendre polynomial of order $k$ on $[-1,1]$, then $\rho_k(u) := \sqrt{2k+1}\ell_k(2u-1)$ for $u \in [0,1]$ (see Section 2.2 in bierens2008semi). This implicitly defines a compact parameter space $\mathcal{F}$ for the unknown c.d.f. $F$ by (ref) for some fixed $G$ and $\pi \in \mathcal{P}$. The sieve $\{\mathcal{F}_k\}_k$ is constructed by truncating the series in (ref) at order $k$. Since we estimate two distribution functions $F_\xi$ and $F_\varepsilon$, we consider a product sieve space $\{\mathcal{F}^\xi_k \times \mathcal{F}^\varepsilon_k\}_k$ for two choices of $G$, denoted $G_\xi$ and $G_\varepsilon$.

As bierens2008semi notes, the c.d.f. $G$ not only restricts the support but also acts as an initial guess of the unknown c.d.f.\footnote{The flexibility of choosing a base distribution $G$ in bierens2008semi is more a norm than an exception in nonparametric methods using orthogonal bases. The standard polynomial sieve on the unit interval may be seen to have the uniform distribution as an initial guess. The Hermite and Laguerre polynomial sieves take normal and exponential distributions as initial guesses, respectively. While these “initial guesses” are determined naturally by the weight function associated with the orthogonal bases, the construction in bierens2008semi allows for an explicit choice by the researcher.} With prior information on the shape of the distribution functions, an educated initial guess helps reduce the approximation error resulting from low-order sieve spaces. Without any prior information, one clearly cannot expect any a priori advantage to choosing a particular distribution. In light of the often-used standard, Hermite, and Laguerre polynomial sieves, it may be reasonable to choose a uniform distribution if the support is known to be contained in a bounded interval, a normal distribution if one is agnostic about the support, and an exponential distribution if the support is known to be contained in $[0,\infty)$.

We consider a sieve extremum estimator that minimizes the average (squared) contrast between two empirical ch.f.s: one based on the factual sample and the other based on simulated draws. The population criterion function is devised as follows. For any candidate pair of distribution functions $F = (F_\xi,F_\varepsilon)$, let \[ X_{(r,s)}(F) = (X_{(r)}(F), X_{(s)}(F)) := (\xi(F_\xi) + \varepsilon_{(r)}(F_\varepsilon), \xi(F_\xi) + \varepsilon_{(s)}(F_\varepsilon)) , \] where $\xi(F_\xi)$ is a random draw from $F_\xi$ and $\varepsilon_{(j)}(F_\varepsilon)$ is the $j$\textsuperscript{th} order statistic of $n$ i.i.d. draws from $F_\varepsilon$. For any $t=(t_r,t_s) \in \mathbb{R}^2$, let $\varphi(t;F) \coloneqq \mathbb{E} e^{it^\top X_{(r,s)}(F)}$ denote its ch.f. where the expectation is induced by the $n+1$ uniform draws in the simulation process described below. We consider the following population criterion function:

align[align omitted — 224 chars of source]

where $\kappa > 0$ is a tuning parameter that determines the integration region. In place of the box weight $\operatorname*{1}\{(t_r,t_s) \in (-\kappa,\kappa)^2\}$, one may also consider a smooth weight. We choose the latter because it has a closed-form expression for the empirical criterion function (see Appendix (ref)).\footnote{While a data-driven choice of $\kappa$ (as well as the polynomial order $k_N$ in Theorem (ref) below) may be of interest, we do not explore its possibility here.} We define a simulated sieve extremum estimator:

align[align omitted — 232 chars of source]

where $\{k_N\}$ is an arbitrary sequence of positive integers such that $k_N \rightarrow \infty$ and $\widehat{Q}_N(F)$ is the empirical criterion function with the empirical ch.f. $\widehat\psi_N$ and the simulated ch.f. $\widehat\varphi_N(\cdot;F)$ in place of $\psi_{X_{(r,s:n)}}$ and $\varphi(\cdot;F)$, respectively. $\widehat\varphi_N(\cdot;F)$ is constructed from $N$ repeated draws of $(n+1)$ uniform random variables $(V,U_1,\ldots,U_n)$ and computing the inverse transform $X_{(j)}(F_k) = F_{\xi,k}^{-1}(V) + F_{\varepsilon,k}^{-1}(U_{(j)})$ for $j \in \{r,s\}$.

We show that the estimator is consistent under the following set of assumptions.

assumption[Consistency] \leavevmode \begin{enumerate}[label=(\alph*)] • $\{(X_{(r),i},X_{(s),i})\}_{i=1,\ldots,N}$ are $N$ i.i.d. realizations of $(X_{(r)},X_{(s)})$, where $(X_{(r)},X_{(s)})$ has a bounded support. • $G_\xi$ and $G_\varepsilon$ are known absolutely continuous c.d.f.s with support on $\mathbb{R}$ and $[0,\infty)$, respectively. • Given $G_\xi$ and $G_\varepsilon$ in part (b), the pair of true underlying distributions $(F_\xi,F_\varepsilon)$ is in the closure of $\bigcup_k \mathcal{F}^\xi_k \times \mathcal{F}^\varepsilon_k$. • $\{(V_i,U_{1,i},\ldots,U_{n,i})\}_{i=1,\ldots,N}$ are $N$ i.i.d. draws from $\mathcal{U}(0,1)^{n+1}$, independent of the sampling process. \end{enumerate}

The support restriction in Assumption (ref)(a) is a sufficient condition to ensure that the measurement errors have light tails. Despite being restrictive, this appears to be the most straightforward assumption to impose on the observables to guarantee identification.\footnote{Note that as remarked in Section (ref), one may consider alternatives to Assumption (ref)(b), e.g., (a.e.-)nonvanishing or analytic ch.f.s, and still achieve the same identification result. Thus, one may also consider estimating the ch.f.s over a space of analytic functions and then estimate the densities via inverse Fourier transform. Clearly, various other estimation methods can be considered. For example, one may consider a sequential approach where one estimates the measurement error distribution via (ref) in the first stage and then estimate the latent variable distribution using nonparametric deconvolution in the second stage. Alternatively, one may consider estimating the densities via a sieve maximum likelihood. We consider the proposed simulated sieve extremum estimator as it both allows to estimate the underlying distributions simultaneously and admits a closed-form objective function.} Note that the researcher does not have to know a priori the true support of the observables, nor the implied bounded supports for $F_\xi$ and $F_\varepsilon$.

Assumption (ref)(b) restricts the support of the base distribution for $F_\varepsilon$ to be in line with the normalization in Assumption (ref)(b). Assumption (ref)(c) is a standard assumption that the model is correctly specified. Assumption (ref)(d) specifies the simulation process. Note that the random draws $(V_i,U_{j,i})_{j,i}$ are obtained once and not repeatedly drawn across different candidate parameter values $F_{\xi,k}$ and $F_{\varepsilon,k}$ in order to ensure that the criterion function is continuous with respect to the parameter.

theoremLet $\kappa > 0$ and let $\{k_N\}_N$ be any sequence of positive integers such that $k_N \rightarrow \infty$. Under Assumptions (ref)--(ref) and (ref), the estimator in (ref) is strongly uniformly consistent, i.e., \[ \max\left\{ \lVert \widehat{F}_{\xi,N} - F_\xi \rVert_\infty, \lVert \widehat{F}_{\varepsilon,N} - F_\varepsilon \rVert_\infty \right\} \stackrel{\text{a.s.}}{\longrightarrow} 0 . \]

Steps for implementing the estimator is described in Appendix (ref) and its finite sample performance is illustrated below. As is standard in nonparametric estimation, the choice of the sieve order $k_N$ affects the approximation bias and variance of the estimator. An information-criterion-based approach analogous to that found in bierens2012semi may be used to choose the sieve order $k_N$, although the procedure may be computationally intensive. With larger $\kappa$, the estimator is expected to perform better as it accounts for more discrepancy between the two ch.f.s. There appears to be no theoretical reason to restrict $\kappa$ except to ensure that the criterion function is well-defined, but limited simulation suggests the aggregation may come with larger variance. We do not have a useful criterion, but in light of the discussion, one may consider minimizing $\widehat{Q}_N$ over $\kappa$ as well as $F$ with a large upper bound on the parameter space for $\kappa$. From simulation studies, we find that the estimator is less sensitive to the choice of $\kappa$ as long as $\kappa$ is not too small. So we set $\kappa=1$ as its baseline value in this paper, as is done in bierens2012semi, and explore its properties in a separate paper. We also leave inference procedures for future research.\footnote{Valid inference procedures exist in similar settings, for example, with independent measurement errors kato2021robust or order statistics without unobserved heterogeneity menzel2013large. A common feature they tackle is a certain lack of continuity in the inverse problems. Similar irregularity concerns may have to be addressed here.}

Monte Carlo Evidence

To illustrate the performance of the proposed estimator, we conduct a simple Monte Carlo experiment. We begin by describing the data generating process (DGP). For observation $i$, the variable of interest $\xi_i$ is measured $n=3$ times with i.i.d. measurement errors $\varepsilon_{1,i},\ldots,\varepsilon_{3,i}$. The measurement is constructed as $X_{j,i} = \xi_i + \varepsilon_{j,i}$. We assume that only two order statistics of ranks $r=1$ and $s=2$ remain. That is, only $X_{(1),i}$ and $X_{(2),i}$ are recorded. We repeat the process to obtain $N$ pairs of observations. The experiment is replicated $R = 500$ times to obtain $500$ random samples of size $N$ of the form $\{x_{(1),i}^{(r)},x_{(2),i}^{(r)}\}_{i=1,\ldots,N}$.

For the distributions of $\xi_i$ and $\varepsilon_{j,i}$, we set up a design that resembles hernandez2020estimation's application on eBay Motors auctions. Specifically, we use their estimated distributions of unobserved heterogeneity ($F_\xi$ in our set-up) and private values ($F_\varepsilon$ in our set-up) to calibrate the DGP for our simulation exercise. To construct these two distributions, we approximate the estimates in Figure 4 of hernandez2020estimation with a sieve of order $6$, which resulted in almost identical distributions to those in the original article.

We then simulate data from the DGP and investigate finite-sample performance of our estimator. We consider sample sizes $N = 1000$, $2000$, and $4000$ with $k = 4$, $5$, and $6$, respectively. Note that by construction, there is no sieve approximation error when $N=4000$, i.e., any estimation error is associated only with sampling error. Furthermore, we set the bandwidth $\kappa =1$, $3.14$, and $5$, and the base distribution $G_\xi = \mathcal{N}(0,1/4)$ for $\xi$ and $G_\varepsilon = \mathcal{N}(2,1)_+$ for $\varepsilon$, where $\mathcal{N}(2,1)_+$ denotes the truncated normal between 0 and $\infty$.

We present the estimation results for $F_\xi$ and $F_\varepsilon$ with $\kappa=1$ in Figure (ref). Simulation results with $\kappa=3.14$ and $5$ are similar and are presented in Figures (ref) and (ref) in Appendix (ref). The true distribution functions are shown in black. We also plot some randomly selected estimates along with some box plots that illustrate the pointwise sampling error of $\widehat{F}_\xi$ and $\widehat{F}_\varepsilon$ at various evaluation points. The figure suggest that the estimator performs reasonably well under all three sample sizes, and the performance improves with larger sample size. Although the choice of $\kappa$ does not seem to affect the behavior of the estimator in any ill-behaved manner, there appear to be larger pointwise variance when $\kappa = 3.14$ and $5$ relative to the case when $\kappa = 1$.

figure[figure omitted — 1,506 chars of source]

Conclusion

This paper shows that distributions of the latent variable and measurement errors are identified nonparametrically under mild assumptions when two or more order statistics are recorded from repeated measurements with independent errors, providing a positive answer to the hypothesis in athey2002identification for an ascending auction with unobserved heterogeneity. Our results are also applicable to other applications with unobserved heterogeneity when order statistics are observed, survey data on wage offers being a notable example. More examples include repeated experiments with type II censoring, such as in reliability testing, where consecutive low-order failure times are recorded, and estimating the effects and damages of collusion in auctions, a setting in which asker2010study emphasizes the importance of accounting for unobserved heterogeneity. Relatedly, the identification result may be applied to extend the framework in marmer2017identifying for testing collusion in ascending auctions.