The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
60,050 characters
Deconvolution from two order statisticsWe thank Gaurab Aryal, Christian Gouri\'eroux, Philip Haile, Lixiong Li, Jos\'e Luis Montiel Olea, Chen Qiu, Cailin Slattery, and conference participants for helpful comments. We thank Cristi\'an Hern\'andez, Daniel Quint, and Christopher Turansick for sharing their code.
\maketitle
\begin{abstract}
Economic data are often contaminated by measurement errors and truncated by ranking. This paper shows that the classical measurement error model with independent and additive measurement errors is identified nonparametrically using only two order statistics of repeated measurements. The identification result confirms a hypothesis by \cite{athey2002identification} for a symmetric ascending auction model with unobserved heterogeneity. Extensions allow for heterogeneous measurement errors, broadening the applicability to additional empirical settings, including asymmetric auctions and wage offer models. We adapt an existing simulated sieve estimator and illustrate its performance in finite samples.
\end{abstract}
\medskip
\noindent\emph{Keywords}: Measurement error, Order statistics, Nonparametric identification, Spacing, Cross-sum
\thispagestyle{empty}
\clearpage
\newpage
\section{Introduction} \label{section:introduction}
We consider the classical measurement error model with repeated measurements:
\begin{align*}
X_j = \xi + \varepsilon_j , \qquad j \in \{1,\ldots,n\} ,
\end{align*}
where the latent variable of interest $\xi$ is measured $n$ times with i.i.d.~measurement errors $(\varepsilon_j)_{j=1,\ldots,n}$ that are independent of $\xi$. The identification result under such a model is well-established when at least \emph{two} repeated measurements are observed, known as Kotlarski's lemma. The result has been widely applied in econometrics since its introduction by \cite{li1998nonparametric}.\footnote{Examples include \cite{bonhomme2010generalized}, \cite{krasnokutskaya2011identification} and \cite{grundl2019identification}. An outline of uses of Kotlarski's lemma in econometrics is in \cite{schennach2016recent}.} In practice, the researcher may observe only a subset of order statistics of the measurements, i.e., $(X_{(j)})_{j \in J}$, where $X_{(j)}$ is the $j$\textsuperscript{th} smallest order statistic from a sample of size $n$ and $J \subset \{1,\ldots,n\}$. This paper shows the underlying probability distributions of the latent variable and the measurement errors are identified from the joint distribution of the ordered measurements $(X_{(j)})_{j \in J}$. \par
The problem took its shape in \cite{athey2002identification} as a nonparametric identification problem in ascending auctions, which has particular relevance in auctions with electronic bidding. They conjecture that the model consists of enough structure to attain point identification using \emph{two} order statistics. However, the question has been long-standing for two decades.\footnote{Some positive findings are made recently in the finite mixture context; see \cite{mbakop2017identification} using five order statistics, \cite{luo2023identification} using two and an instrument and \cite{luo2021order} using three. However, these papers do not tackle the original conjecture in the spirit of Kotlarski's lemma.} Importantly, the standard approach using Kotlarski's lemma fails because of dependence in the order statistics of measurement errors. Further, the attempt to difference out the latent variable and exploit the spacing distribution is shown to be insufficient for point identification, evidenced by \cite{rossberg1972characterization}'s counterexample. \par
While the literature has documented the increasing importance of unobserved heterogeneity in auctions,\footnote{For example, see \cite{aradillas2013identification}, \cite{krasnokutskaya2011identification}, and \cite{li2000conditionally}.} the lack of identification results in the existing literature has hindered allowance for such heterogeneity in the classical fashion, unless relying on additional external variations in the data.\footnote{For example, \cite{hernandez2020estimation} uses variation in the number of bidders across auctions. \cite{freyberger2022identification} circumvents the issue of dependence between order statistics using observed reserve prices, rendering the identification problem classical.} We fill this important gap by showing that the underlying distributions are identified with only data on \emph{two} order statistics, which need not be consecutive or extreme, without relying on extra variations. \par
We propose a new identification strategy that exploits both features in \cite{kotlarski1967characterizing} and \cite{rossberg1972characterization}. In particular, we make use of the \emph{within} independence of the latent variable and the additive-separable measurement errors as in \cite{kotlarski1967characterizing} to derive that the model imposes an additional restriction beyond the spacing of order statistics.\footnote{The term \emph{spacing} usually refers to the difference between two \emph{consecutive} order statistics. Here, we use it more broadly as the difference between any two order statistics.} More precisely, we find that two measurement error distributions both rationalize the observed distribution of the ordered measurements only if the joint distributions of \emph{spacing} and \emph{cross-sum} are identical.\footnote{The term \emph{cross-sum} refers to the sum of two order statistics of two different ranks, where the order statistics are of two independent random samples from distinct underlying parent distributions.} Exploiting the structure of order statistics and commonly seen conditions (a support or tail restriction), we show that the spacing in conjunction with this additional restriction is indeed sufficient to point identify the underlying distribution using any two order statistics. The latent variable distribution is subsequently identified by a standard deconvolution argument. \par
We extend our main result to the setting where measurement errors are independent but nonidentically distributed (i.n.i.d.). We show that the underlying distributions are identified when two order statistics are observed, provided \emph{some} measurement errors are known to be identically distributed and the data consist of group identities of high-order statistics. The result applies to asymmetric ascending auctions, where only high-order dropout bids are recorded, and asymmetric first-price auctions and wage offer settings, where consecutive and extreme order statistics are observed. \par
The outline of the paper is as follows. Section \ref{section:motivating_examples} introduces two motivating examples and Section \ref{section:model_and_identification} presents the identification results. To complement the identification result, in Section \ref{section:estimator_main} we adapt \cite{bierens2012semi}'s simulated sieve estimator to consistently estimate the unknown distributions. Section \ref{section:conclusion} concludes. \par
\section{Motivating Examples} \label{section:motivating_examples}
To motivate the identification problem, we introduce two examples: an ascending auction model with auction-specific unobserved heterogeneity and a model of wage offers where wage is determined by worker-specific labor productivity. Throughout the paper, we focus only on unobserved heterogeneity since adding exogenous covariates does not require novel considerations. \par
\begin{example}[Ascending auction]
\label{example:ascending_auction}
Consider an ascending auction with one indivisible good and $n$ bidders. Each bidder's valuation for the item is determined by a set of item features---summarized and denoted by $\xi \in \mathbb{R}$---and a private value component $\varepsilon_j$ that is independent of $\xi$ and across bidders. The valuation of bidder $j$ is given by $X_j = \xi + \varepsilon_j$. Note that the valuation of each bidder can be cast as a measurement of the value $\xi$ with measurement error $\varepsilon_j$. See, e.g., \cite{athey2002identification}. \par
In an ascending auction, the price rises until only one bidder remains. Suppose the auctioneer records prices $P_1 \le P_2 \le \cdots$ at which bidders drop out. Assuming that the bidders play the dominant strategy of remaining in the auction until the price reaches their valuations, the dropout bids reveal the ordered valuations of the bidders, i.e., $P_j = X_{(j)}$. Importantly, the highest valuation $X_{(n)}$ is never observed since the auction ends when the price reaches $P_{n-1} = X_{(n-1)}$, and the bidder with the highest valuation wins. Therefore, in an ascending auction, we observe only an incomplete set of order statistics on bidders' valuations. For example, \cite{kim2014nonparametric} observe at least the three highest dropout prices in ascending used-car auctions; see also \cite{larsen2021efficiency} for used-car auctions in the U.S. Data on timber auctions by the U.S.~Forest Service---analyzed in a number of empirical papers, e.g., \cite{haile2003inference}, \cite{athey2011comparing}, and \cite{aradillas2013identification} to name a few---records at most top twelve bids regardless of the number of bidders. \par
Both the distributions of $\xi$ and $\varepsilon_j$ are of interest in an empirical analysis of auctions. For example, the counterfactual expected revenue to the auctioneer under a hypothetical auction rule requires the knowledge of both distributions. \par
\end{example}
\begin{example}[Wage offer determination]
\label{example:wage_offer}
Consider a simple wage offer model $X_j = \xi + w_j + \varepsilon_j$ where $X_j$ is the $j$\textsuperscript{th} (log-)wage offer an individual receives. $X_j$ is assumed to be composed of the worker's productivity $\xi$, the wage rate $w_j$, and measurement error $\varepsilon_j$ (on, e.g., labor productivity) associated with the offer. It is assumed that the period in consideration is relatively short such that the worker's productivity remains unchanged, but the worker may receive multiple offers; see \cite{guo2021identification}. \par
Data from the Survey of Consumer Expectations (SCE) Labor Market Survey records salary-related responses of all job offers for those who received at most three offers and responses to three best offers for those who received more than three offers within the last four months. In this setting, the researcher observes $X_{(n)},\ldots,X_{(n-k+1)}$ where $k = \min\{3,n\}$ for each respondent. The identification results in the current paper show that the distributions $F_\xi$ and $F_{w_j+\varepsilon_j}$ are nonparametrically identified. \par
\end{example}
\section{The Model and Main Results} \label{section:model_and_identification}
In this section, we first formalize the i.i.d.~framework considered throughout the paper and provide sufficient conditions to identify the latent variable and measurement error distributions. In Section~\ref{subsection:extensions}, we consider i.n.i.d.~measurement errors. \par
\subsection{The Setup} \label{subsection:model_setup}
We begin by stating the sampling process. \par
\begin{assumption}[Sampling process]
\label{assumption:sampling_process}
For each $1 \le j \le n$, $X_{j} := \xi + \varepsilon_{j}$, where the random variable $\xi$ is independent of the random vector $(\varepsilon_{1},\ldots,\varepsilon_{n})$.\footnote{The identification results herein also applies to a setting with multiplicatively separable measurement error by taking logs when $\xi$ and $\varepsilon_j$'s are positive, as in Example~\ref{example:wage_offer}.}
\end{assumption}
The researcher observes $(X_{(r)},X_{(s)})$, the $r$\textsuperscript{th} and $s$\textsuperscript{th} order statistics from $n$ observations, where $r$, $s$, and $n$ are known and $1 \le r < s \le n$. That is, while the total number of measurements is known, one only observes two measurements of known ranks. In this section, we identify the unknown distributions of $\xi$ and $(\varepsilon_{1},\ldots,\varepsilon_{n})$ using $F_{X_{(r,s)}}$, the joint distribution of $(X_{(r)},X_{(s)})$, which is estimable from a random sample of the two order statistics. \par
For the benchmark case, we make the following distributional assumptions.
\begin{assumption}[i.i.d.~errors]
\label{assumption:iid_case}
\leavevmode
\begin{enumerate}[label=(\alph*)]
\item $\varepsilon_{1},\ldots,\varepsilon_{n}$ are i.i.d.~with a common distribution $F_\varepsilon$ on $\mathbb{R}$.
\item $F_\varepsilon$ is absolutely continuous with a probability density function $f_\varepsilon$ that is light-tailed, i.e., for some $C > 0$, $f_\varepsilon(\epsilon) = O(e^{-C\lvert\epsilon\rvert})$ as $\lvert\epsilon\rvert \rightarrow \infty$.
\end{enumerate}
\end{assumption}
The tail condition in Assumption~\ref{assumption:iid_case}(b) is trivially satisfied when the support is bounded. If the support is unbounded on either side, the assumption restricts the tail of the density function to decay at an exponential rate.\footnote{Throughout the paper, we use the term ``light-tailed'' to refer to densities with exponentially decaying tails as stated in the assumption. \label{fn:light_tailed}} The same assumption is found in \cite{evdokimov2012some}. Alternatively, we may follow \cite{kotlarski1967characterizing} and \cite{miller1970characterizing} to assume (a.e.-)nonvanishing or analytic characteristic functions (ch.f.) of both $\xi$ and $\varepsilon$. Assumption~\ref{assumption:iid_case}(b) is a sufficient condition for the ch.f.~of $F_\varepsilon$ to be analytic. However, to our knowledge, there exists no result regarding its implication on the joint ch.f.~of order statistics, the property we need for our identification results. In Lemma~\ref{lemma:appendix_joint_analyticity}, we show that the joint ch.f.~of the order statistics $(\varepsilon_{(r)},\varepsilon_{(s)})$ (and thus its marginals) is also analytic under this assumption. \par
In addition to Assumption~\ref{assumption:iid_case}, we further restrict the support of the measurement errors as follows. For any distribution $F$ on $\mathbb{R}$, let $S(F) \subseteq \mathbb{R}$ denote its support.\footnote{The support of $F$ is defined as the smallest closed set $K \subseteq \mathbb{R}$ such that $P_F(K) = 1$, where $P_F$ is the Borel probability measure induced by $F$. Thus, the density $f(x)$ may be zero for some values $x \in K$, including points on the boundary of $K$.}
\begin{assumption}[Support condition]
\label{assumption:bound}
The measurement errors are bounded from below, which is normalized to zero, i.e., $\inf S(F_\varepsilon) = 0$.
\end{assumption}
Two aspects are worth mentioning regarding the above assumption. First, we require the measurement error to be bounded at one end; nevertheless, we allow the support to be possibly unbounded from above.\footnote{The results herein apply to the case when only the upper bound is finite, e.g., $S(F_\varepsilon) = (-\infty,0]$, by considering $(X_{(r')}',X_{(s')}^\prime) = (-X_{(s)},-X_{(r)})$ where $r' = n-s+1$ and $s'=n-r+1$. \label{footnote:support_upper_bound}} We do not assume the upper bound of the support is known nor assume any additional structure on the support, e.g., connectedness. Second, as is typical with measurement error models, the underlying distributions are identified only up to location. Although one may alternatively normalize the mean and have an unknown but finite lower bound, normalizing the lower bound turns out to be more convenient for our identification argument. Our identification strategy heavily utilizes the condition that one boundary of the support of $F_\varepsilon$ is finite. On the other hand, the distribution of $\xi$ is left completely unspecified. \par
The support restriction is nonrestrictive in many applications. The literature on games with incomplete information typically assumes bounded support of agent types (\cite{athey2001single}), such as bidder valuation in empirical auctions (\cite{guerre2000optimal}, \cite{athey2007nonparametric}), private costs of exerting efforts in contest games (\cite{huang2021structural}), and private information about variable costs in Cournot games (\cite{aryal2021empirical}). Wage offers are also bounded below in job search models (\cite{burdett1998wage}, \cite{guo2021identification}). \par
\subsection{Nonparametric Identification} \label{subsection:identification}
Throughout the section, we treat the observable joint distribution $F_{X_{(r,s)}}$ as known (i.e., as a datum), and show that the distribution $F_\xi$ of the latent variable $\xi$ and $F_\varepsilon$ are identified under the aforementioned assumptions. The bulk of our identification result is concerned with showing that the measurement error distribution that can rationalize the observed distribution $F_{X_{(r,s)}}$ must be unique. Then the latent variable distribution is identified by a standard deconvolution argument. To achieve such a goal, we begin by isolating the empirical content on measurement errors from the datum. In Lemma~\ref{lemma:rationalizability} below, we do so by exploiting the multiplicative structure of the ch.f.~of a sum of independent random variables.\footnote{Similar ideas are also exploited in the classical repeated measurement error setting. See \cite{hall2003inference} and \cite{evdokimov2012some}.} \par
In the following lemmas, let $\eta_1,\ldots,\eta_n$ be $n$ i.i.d.~copies of $\eta \in \mathbb{R}$ and let $(\eta_{(r)},\eta_{(s)})$ be order statistics. $\psi_{\eta_{(r,s)}}$ and $\psi_{\eta_{(r)}}$ denote the joint and marginal ch.f., respectively. \par
\begin{lemma}
\label{lemma:rationalizability}
If the sampling process is described as in Assumption~\ref{assumption:sampling_process} and the measurement errors satisfy Assumption~\ref{assumption:iid_case}, a measurement error distribution $F$ on $\mathbb{R}$ is data-consistent\footnote{We say a pair of distribution functions $(G_\xi,G_\varepsilon)$ rationalizes the data, or is data-consistent, if $\xi' \sim G_\xi$ and $(\varepsilon_j')_{j=1,\ldots,n} \sim \times_{j=1}^n G_\varepsilon$, independent of $\xi'$, implies $(\xi'+\varepsilon_{(r)}', \xi'+\varepsilon_{(s)}') =_d (X_{(r)},X_{(s)})$. We say $G_\xi$ (resp., $G_\varepsilon$) is data-consistent if $(G_\xi,G_\varepsilon)$ is data-consistent for some $G_\varepsilon$ (resp., $G_\xi$).} only if
\begin{align}
\frac{\psi_{X_{(r,s)}}(t_r,t_s)}{\psi_{X_{(j)}}(t_r+t_s)} = \frac{\psi_{\eta_{(r,s)}}(t_r,t_s)}{\psi_{\eta_{(j)}}(t_r+t_s)} , \quad \text{for all $(t_r,t_s) \in B_0$, } j \in \{r,s\} , \label{eq:identify_eps}
\end{align}
where $\eta \sim F$ and $B_0$ is an open ball in $\mathbb{R}^2$ centered at zero. Further, $F$ induces a unique data-consistent latent variable distribution.
\end{lemma}
Suppose $F$ and $G$ are two measurement error distributions that are data-consistent. Lemma~\ref{lemma:rationalizability} states that the order statistics $(\eta_{(r)},\eta_{(s)})$ from the parent distribution $F$ and $(\eta'_{(r)},\eta'_{(s)})$ from the parent distribution $G$ must have the same ratio of joint and marginal ch.f.s. To obtain identification, it remains to show that such a measurement error distribution is unique, i.e., $F=G$. In contrast to the prevalent identification strategy in the existing literature that expands on \cite{kotlarski1967characterizing}, which begins with identifying the distribution of the latent variable $\xi$, Lemma~\ref{lemma:rationalizability} suggests an identification strategy that first focuses on the measurement error distribution. \par
Note that without any assumption on the latent variable $\xi$, the ch.f.~of $\xi$ may vanish on a nontrivial interval bounded away from the origin.\footnote{Recall that a ch.f.~takes value one at the origin and is uniformly continuous, implying that it does not vanish on some neighborhood about the origin.} Hence, in general, the ratio is only guaranteed to be well-defined on a neighborhood about the origin. Analyticity of the ch.f.~of measurement error order statistics (Lemma \ref{lemma:appendix_joint_analyticity}) thus plays a crucial role in pinning down the entire distribution based only on the information of the ch.f.~about the origin. In particular, Lemma~\ref{lemma:rationalizability} implies that any two data-consistent measurement error distributions must satisfy
\begin{align*}
\psi_{\eta_{(r,s)}}(t_r,t_s) \psi_{\eta'_{(j)}}(t_r+t_s) = \psi_{\eta'_{(r,s)}}(t_r,t_s) \psi_{\eta_{(j)}}(t_r+t_s) , \quad j \in \{r,s\} ,
\end{align*}
for all $(t_r,t_s) \in B_0$. The equality extends to all of $\mathbb{R}^2$ because a product of two analytic functions is analytic. Finally, noticing that the ch.f.~of the sum of two independent random vectors is multiplicatively separable, the following lemma establishes necessary conditions for any two data-consistent measurement error distributions. \par
\begin{lemma}
\label{lemma:observational_equivalence_eps}
Let $\eta_{(r)}$ and $\eta_{(s)}$ be two order statistics of a random sample with size $n$ from $F$, and $\eta_{(r)}^\prime$ and $\eta_{(s)}^\prime$ be two order statistics of a random sample with size $n$ from $G$, where $(\eta_{(r)},\eta_{(s)})$ and $(\eta'_{(r)},\eta'_{(s)})$ are independent. Under Assumptions~\ref{assumption:sampling_process}, if measurement error distributions $F$ and $G$ admit light-tailed densities,\footnote{Cf.~Assumption~\ref{assumption:iid_case}(b) and footnote~\ref{fn:light_tailed}.} they both rationalize the data only if
\begin{align}
Z_{1j}
:= \begin{pmatrix} \eta_{(j)}^\prime + \eta_{(r)} \\ \eta_{(j)}^\prime + \eta_{(s)} \end{pmatrix}
\stackrel{d}{=}
\begin{pmatrix} \eta_{(j)} + \eta_{(r)}^\prime \\ \eta_{(j)} + \eta_{(s)}^\prime \end{pmatrix}
=: Z_{2j} , \quad j \in \{r,s\} . \label{eq:observational_equivalence}
\end{align}
\end{lemma}
The equivalence of ratios of ch.f.s~(Lemma~\ref{lemma:rationalizability}) implies an equivalence of two joint distributions of particular linear combinations of order statistics that arise from the two measurement error distributions. The linear combinations not only expose the empirical content contained in the spacing of errors (see Corollary~\ref{corollary:observational_equivalence_eps}) as investigated by \cite{hall2003inference} among many others in the classical repeated measurements setting and by \cite{rossberg1972characterization} in the context of order statistics, but they also imposes restrictions on the sum of order statistics from potentially distinct parent distributions, which we call a cross-sum. \par
Because both $Z_{1r}$ and $Z_{2r}$ involve three order statistics that are associated with potentially different parent distributions $F$ and $G$, it appears to be difficult to establish the equivalence of $F$ and $G$ directly from the distributional equivalence of $Z_{1r}$ and $Z_{2r}$. A reasonable attempt to tackle the problem would be to explore features of the distributions of $Z_{1r}$ and $Z_{2r}$ that depend on the parent distributions in a simple manner. Under the support condition in Assumption~\ref{assumption:bound}, we show that the distributions of $Z_{1r}$ and $Z_{2r}$ depend on only one of the two parent distributions upon conditioning on an extreme event.\footnote{The identification argument here and below investigates the implications of the condition $Z_{1r} =_d Z_{2r}$ only. $Z_{1s} =_d Z_{2s}$ in \eqref{eq:observational_equivalence} is the relevant condition to exploit under the assumption, in place of Assumption~\ref{assumption:bound}, that the measurement error is bounded from above (cf.~footnote~\ref{footnote:support_upper_bound}).} Thus the joint distribution of the cross-sum and spacing together with Assumption~\ref{assumption:bound} provide information sufficient to claim equivalence of two data-consistent measurement error distributions and point identify $F_\varepsilon$. \par
\begin{lemma}\label{lemma:limit_of_conditional_distribution}
Let $\eta_{(r)}$, $\eta_{(s)}$, $\eta_{(r)}^\prime$ and $\eta_{(s)}^\prime$ be as specified in Lemma~\ref{lemma:observational_equivalence_eps}. If measurement error distributions $F$ and $G$ admit light-tailed densities and $\inf S(F) = \inf S(G) = 0$, we have, for all $c \in \mathbb{R}$,
\begin{align}
\lim_{\delta \downarrow 0} \mathbb{P}(\eta_{(r)}^\prime + \eta_{(s)} \le c \ | \ \eta_{(r)}^\prime + \eta_{(r)} \le \delta) &= F_{s-r:n-r}(c) , \label{eqn:limitF} \\
\lim_{\delta \downarrow 0} \mathbb{P}(\eta_{(r)} + \eta_{(s)}^\prime \le c \ | \ \eta_{(r)} + \eta_{(r)}^\prime \le \delta) &= G_{s-r:n-r}(c) , \label{eqn:limitG}
\end{align}
where $F_{s-r:n-r}$ (resp., $G_{s-r:n-r}$) denotes the distribution of the $(s-r)$\textsuperscript{th} order statistic of a random sample with size $n-r$ from $F$ (resp., $G$).
\end{lemma}
For the sake of intuition, suppose that conditioning on the event $\{\eta_{(r)}^\prime + \eta_{(r)} = 0\}$ is well-defined. The condition implies that both $\eta_{(r)}^\prime = 0$ and $\eta_{(r)} = 0$. Thus, the independence between the pairs of order statistics simplifies the conditional distribution to $\mathbb{P}( \eta_{(s)} \le c \ | \ \eta_{(r)} = 0)$. Standard arguments in order statistics (cf.~Theorem~2.4.2 in \cite{arnold2008first}) suggest that this conditional distribution has a simple form:
\[
\mathbb{P}( \eta_{(s)} \le c \ | \ \eta_{(r)} = 0) = F_{s-r:n-r}(c) .
\]
In other words, the equality above says that if the lowest $r$ draws all take the minimal possible value, then the distribution of any higher order statistic must equal the distribution of order statistic when the $r$ draws at the minimal value are discarded. \par
However, the event $\{\eta_{(r)}^\prime + \eta_{(r)} = 0\}$ may not be well-defined as it occurs with zero probability. In particular, the event $\{\eta_{(r)} = 0\}$ always has zero density (regardless of the density of $\eta$) unless $r=1$, i.e., the minimum order statistic. Further, because the true density $f_\varepsilon(0)$ may be zero, even the minimum order statistic may have zero density. In Lemma~\ref{lemma:limit_of_conditional_distribution}, we make the intuitive claim made above more rigorous by considering the limiting argument as in \eqref{eqn:limitF} and \eqref{eqn:limitG}, which is well-defined. \par
Lemma~\ref{lemma:observational_equivalence_eps} implies that if $F$ and $G$ are both data-consistent, then the conditional distributions \eqref{eqn:limitF} and \eqref{eqn:limitG} in Lemma~\ref{lemma:limit_of_conditional_distribution} must be equal, i.e., for all $c \in \mathbb{R}$,
\[
F_{s-r:n-r}(c) = G_{s-r:n-r}(c) .
\]
As the distribution of an order statistic uniquely identifies the parent distribution, we can conclude that $F = G$.\footnote{The one-to-one mapping between the distribution of an order statistic and the parent distribution is a standard result, e.g., see \cite{david2003order}, p.10, (2.1.5). The mapping in (2.1.5) is an invertible map of the c.d.f.~$F$ because $p \mapsto I_p(a, b)$ is strictly monotone.} Combining Lemmas~\ref{lemma:rationalizability}--\ref{lemma:limit_of_conditional_distribution} leads to the following main identification result for the i.i.d.~case. \par
\begin{theorem}
\label{theorem:identification_iid_case}
Under Assumption \ref{assumption:sampling_process}, if the measurement errors satisfy Assumptions~\ref{assumption:iid_case} and \ref{assumption:bound}, both $F_\xi$ and $F_\varepsilon$ are identified.
\end{theorem}
\subsubsection*{Discussion on \cite{rossberg1972characterization}'s Counterexample}
\cite{athey2002identification} pointed out that an identification strategy based on the spacing between order statistics does not yield a positive result. The discussion therein relied on a counterexample by \cite{rossberg1972characterization}. A straightforward corollary to Lemma~\ref{lemma:observational_equivalence_eps} highlights the model restrictions we exploit in addition to the spacing. \par
\begin{corollary}
\label{corollary:observational_equivalence_eps}
Under the same assumptions as in Lemma~\ref{lemma:observational_equivalence_eps}, measurement error distributions $F$ and $G$ both rationalize the data only if
\begin{align*}
\begin{pmatrix} \eta_{(s)} - \eta_{(r)} \\ \eta_{(r)}^\prime + \eta_{(s)} \end{pmatrix}
\stackrel{d}{=} \begin{pmatrix} \eta_{(s)}^\prime - \eta_{(r)}^\prime \\ \eta_{(r)} + \eta_{(s)}^\prime \end{pmatrix}
\quad \text{and} \quad
\begin{pmatrix} \eta_{(s)} - \eta_{(r)} \\ \eta_{(r)} + \eta_{(s)}^\prime \end{pmatrix}
\stackrel{d}{=} \begin{pmatrix} \eta_{(s)}^\prime - \eta_{(r)}^\prime \\ \eta_{(r)}^\prime + \eta_{(s)} \end{pmatrix} .
\end{align*}
\end{corollary}
In words, the observational equivalence between any two measurement error distributions in our setting implies the equivalence of the joint distributions of the spacing and cross-sum of order statistics. This finding complements the nonidentification result in \cite{rossberg1972characterization} that the distribution of spacing between two order statistics alone cannot identify their parent distribution. The additional constraint on the distribution of cross-sum absent in \cite{rossberg1972characterization} emerges from the within independence assumption. Figure~\ref{fig:df_fig} in Appendix~\ref{appendix:rossberg} shows that, while the spacing distributions for the exponential and Rossberg's counterexample overlap, the cross-sum distributions differ over nontrivial regions. See Appendix~\ref{appendix:rossberg} for more discussion. \par
\subsection{Extensions of the Identification Result} \label{subsection:extensions}
We extend our identification result to the case of independent but nonidentically distributed (i.n.i.d.) measurement errors. To motivate the extension to the problem, we expand on Example~\ref{example:ascending_auction} by introducing bidder asymmetry and discuss additional features available in some auction data.
\begin{example}[Asymmetric auction]
In empirical auctions, bidders are said to be asymmetric if the bidders' private values $\varepsilon_1,\ldots,\varepsilon_n$ have different marginal (or parent, in the case of order statistics) distributions. Such asymmetry arises naturally in procurement auctions, where contractors differ in cost efficiency and productivity (see, e.g., \cite{flambard2006asymmetry}), timber auctions, where mills have the manufacturing capacity and loggers do not (see, e.g., \cite{athey2011comparing})\footnote{In a wage offer setting as in Example~\ref{example:wage_offer} employers may value the same productivity differently, which may result in heterogeneous measurement errors.}. \par
Due to asymmetry, identifying bidder-specific valuation distributions requires knowing their identities, i.e., knowing who participated in each auction. Fortunately, the auctioneer often publishes such identities despite the highest valuation being missing; see, e.g., \cite{athey2011comparing}. Another common practice in the asymmetric auction literature is grouping bidders by commonly known bidder types, such as mills and loggers in \cite{athey2011comparing} and strong and weak bidders in \cite{luo2018auctions}, which leads to more tractable theory and empirics. Such grouping uses additional information about the bidders, typically available as a public record, such as bidders' manufacturing capacity in a timber auction and pre-qualified contractors' experience in Department of Transportation procurement auctions.
\end{example}
In this section, we first show that any two order statistics suffice to point identify the underlying distributions when there is a relatively small number of groups and the group identities of high-order statistics are observable. We then discuss several relevant applications in empirical auctions. \par
\begin{assumption}[i.n.i.d.~errors] \label{assumption:noniid_case}
\leavevmode
\begin{enumerate}[label=(\alph*)]
\item $\varepsilon_{1},\ldots,\varepsilon_{n}$ are independent with distributions $F_{\varepsilon_1},\ldots,F_{\varepsilon_n}$ on $\mathbb{R}$, respectively.
\item For each $j \in \{1,\ldots,n\}$, $F_{\varepsilon_j}$ admits a probability density function $f_{\varepsilon_j}$ that is light-tailed, i.e., $f_{\varepsilon_j}(\epsilon) = O(e^{-C_j\lvert\epsilon\rvert})$ as $\lvert\epsilon\rvert \rightarrow \infty$ for some $C_j > 0$.
\item For each $j \in \{1,\ldots,n\}$, $\inf S(F_{\varepsilon_j}) = 0$.
\end{enumerate}
\end{assumption}
Assumption~\ref{assumption:noniid_case}(c) assumes both finite and common support lower bound across distinct measurement error distributions, the latter of which holds trivially in the i.i.d.~case. Note that apart from the lower boundary, the supports may differ (cf.~Corollary~\ref{corollary:identification_noniid}).
The cost of relaxing the homogeneity assumption translates to more stringent data requirements. We require observing indices or types associated with all high-order statistics (for ranks above $r$). The following assumption records any prior information on the different types of measurement errors.
\begin{assumption}[Group structure] \label{assumption:group_structure}
There exists a partition $g_1,\ldots,g_p$ of $\{1,\ldots,n\}$ such that $\varepsilon_j =_d \varepsilon_k$ if $j,k \in g_q$ for some $q \in \{1,\ldots,p\}$. In addition, the measurement errors $\varepsilon_1,\ldots,\varepsilon_n$ have common support.
\end{assumption}
Assumption~\ref{assumption:group_structure} posits there are at most $p$ distinct measurement error distributions. The case $p=1$ corresponds to the i.i.d.~case in Section~\ref{subsection:identification}; and $p=n$ corresponds to the case with no prior information on homogeneity. We do not preclude the possibility that two groups $g_q$ and $g_{q'}$ have the same distribution beyond the researcher's knowledge. E.g., the extreme case $p=n$ subsumes the i.i.d.~setting as a special case. Assumption~\ref{assumption:group_structure} implicitly assumes that the group structure is held constant across observations. In an application to auctions, the assumption may be imposed by restricting to auctions with the same composition of bidder types when all bidder types of participants are available in the dataset. In such a case, the analysis should be interpreted conditionally on the bidder composition. \par
The common support assumption is a sufficient condition to guarantee that all measurement error distributions are identified on the entirety of their support. Otherwise, some distributions may be identified only on a strict subset of their support, as we highlight in a discussion below. Let $R_{(j)} = \{q : X_{(j)} = X_k \text{ for some } k \in g_q\}$ be the group identity of the $j$\textsuperscript{th} order statistic.\footnote{Under Assumption~\ref{assumption:noniid_case}(b), $R_{(j)}$ is singleton with probability $1$. Thus, we abuse notation and use $R_{(j)}$ to denote both the set and the a.s.~unique element in $\{1,\ldots,p\}$.} \par
\begin{theorem} \label{theorem:identification_noniid_group}
Suppose Assumptions~\ref{assumption:sampling_process}, \ref{assumption:noniid_case}, and \ref{assumption:group_structure} hold. Further, suppose two order statistics and the group identities of high-order statistics $(X_{(r)},X_{(s)},R_{(r+1)},\ldots,R_{(n)})$ are observed. If there exists a group $g_q$ with at least $n-r$ members, both $F_\xi$ and $(F_{\varepsilon_j} : j=1,\ldots,n)$ are identified.
\end{theorem}
Note that the identification result does not require the researcher to observe the group identity $R_{(r)}$ of the $r$\textsuperscript{th} order statistic or those of lower order statistics. An heuristic explanation is provided in a discussion below. Theorem~\ref{theorem:identification_noniid_group} has an important implication when the measurement errors are left completely heterogeneous, i.e., $p=n$. In particular, the theorem implies that the observed order statistics should be consecutive and extreme because there is no group of a larger size, i.e., $n-r=1$.\footnote{The support condition in Assumption~\ref{assumption:group_structure} plays no role in this corollary because the observed order statistics are the largest among all measurement errors.}
\begin{corollary}
\label{corollary:identification_noniid}
Suppose Assumptions~\ref{assumption:sampling_process} and \ref{assumption:noniid_case} hold. Further, suppose $(X_{(n-1)},X_{(n)},R_{(n)})$ is observed. Both $F_\xi$ and $(F_{\varepsilon_j} : j=1,\ldots,n)$ are identified.
\end{corollary}
So as to appreciate the additional conditions assumed in Theorem~\ref{theorem:identification_noniid_group}, we first illustrate the identification strategy in the setting of Corollary~\ref{corollary:identification_noniid}, which delivers a simpler argument. Suppose results similar to Lemmas~\ref{lemma:rationalizability}--\ref{lemma:limit_of_conditional_distribution} hold in the i.n.i.d.~case, in the sense that two sets of data-consistent measurement error distributions $(F_j : j=1,\ldots,n)$ and $(G_j : j=1,\ldots,n)$ must satisfy
\begin{align}
\mathbb{P}(\eta_{(n)} \le c \,\vert\, \eta_{(n-1)} = 0) = \mathbb{P}(\eta'_{(n)} \le c \,\vert\, \eta'_{(n-1)} = 0) , \label{eq:conditional_implication_noniid}
\end{align}
for every constant $c$. Had the measurement errors been i.i.d., \eqref{eq:conditional_implication_noniid} corresponds the equivalence of two parent distributions. In the i.n.i.d.~case, the top-order statistic may arise from any of the parent distributions. Intuition suggests that this should be a mixture of measurement errors from different parent distributions. Since the mixing weights depend on $(F_j : j=1,\ldots,n)$, it appears formidable to show its equivalence with $(G_j : j=1,\ldots,n)$ from \eqref{eq:conditional_implication_noniid} without additional assumptions. Alternatively, suppose---as we formally show in the proof of Theorem~\ref{theorem:identification_noniid_group}---data-consistency implies a condition similar to \eqref{eq:conditional_implication_noniid} conditional on the member index of the highest order statistic. Heuristically speaking, because the index is known, say $j$, the conditional distribution simplifies to the $j$\textsuperscript{th} marginal distribution (i.e., a mixture with trivial weights). Thus $F_j = G_j$ and the $j$\textsuperscript{th} measurement error distribution is identified. Under the assumption of a common support lower bound, each measurement error has a nonzero chance of being the largest, which implies all $n$ measurement error distributions are identified. \par
Corollary~\ref{corollary:identification_noniid} does not grant identification of an ascending auction model when the data on dropout bids is incomplete and only a few high-order dropout bids as in, e.g., \cite{freyberger2022identification}. Nonetheless, group structures as in Theorem~\ref{theorem:identification_noniid_group} are commonly present in the empirical auction literature. \par
To illustrate the identification strategy in the setting of Theorem~\ref{theorem:identification_noniid_group}, suppose out of $n=4$ measurements, one observes $(X_{(2)},X_{(3)},R_{(3)},R_{(4)})$ and it is known \emph{a priori} that two of the measurement errors have a common distribution (call it group 1 and the other groups 2 and 3). By conditioning on the event that $\{R_{(3)} = R_{(4)} = 1\}$, we can homogenize the conditional distribution of the 3\textsuperscript{rd}-order statistic of measurement errors conditional on the 2\textsuperscript{nd}-order statistic taking the smallest possible value, which allows us to identify the measurement error distribution for group 1. This rests on (i) being able to observe the group identities and (ii) group 1 being large enough to condition on such an event. Provided the measurement errors have a common support, the remaining group distributions can be identified sequentially. For example, the group-2 distribution can be identified by conditioning on $\{R_{(3)} = 2, R_{(4)} = 1\}$.\footnote{If one measurement error has larger support than another, the distribution may not be fully identified. If group-1 measurement errors have support $[0,1]$ but group 2 has support $[0,2]$, regardless of whether one conditions on $\{R_{(3)} = 1, R_{(4)} = 2\}$ or $\{R_{(3)} = 2, R_{(4)} = 1\}$, the 3\textsuperscript{rd}-order statistic only has support $[0,1]$. Thus group-2 measurement error distribution is not identified on $(1,2]$.} \par
\begin{remark}
Corollary~\ref{corollary:identification_noniid} has natural applications in first-price auctions and wage offer settings, where consecutive and extreme order statistics are common and identities are observed. For instance, the Washington State Department of Transportation archives three apparent low bids and bidder identities of six months or older online. FDIC auction data contain the winning and second-highest bids and the associated identities (\cite{allen2023resolving}). U.S.~Forest Service timber auctions only record the top twelve bids and bidder identities regardless of the number of bidders. Lastly, data from the Survey of Consumer Expectations (SCE) Labor Market Survey records salary-related responses on the three best offers for those who received more than three offers within the last four months.
\end{remark}
\begin{remark} \label{remark:identification_noniid}
If \emph{all} dropout bids are observed, Corollary~\ref{corollary:identification_noniid} is applicable to ascending auctions. Specifically, if private values have a common upper bound, i.e., $\sup S(F_{\varepsilon_j}) = \overline{\varepsilon} < +\infty$, the lowest order statistics $(X_{(1)},X_{(2)})$ identify the underlying distributions.
\end{remark}
\section{Nonparametric Estimation} \label{section:estimator_main}
We propose a simulated sieve estimator for the i.i.d.~measurement error framework, which can be easily modified for the i.n.i.d.~case, albeit with the curse of dimensionality. The procedure, motivated by the sieve estimator proposed in \cite{bierens2012semi}, estimates the distribution functions of the latent variable and of the measurement error simultaneously. \par
\subsection{A Simulated Sieve Estimator} \label{subsection:estimator_sub}
We follow closely the development in \cite{bierens2008semi} to specify the parameter space and its sieve space. It has the advantage that we can incorporate prior information on the support without having to choose different orthogonal bases. This is an attractive feature for our purpose because, unlike the latent variable $\xi$, the measurement errors are restricted to be nonnegative-valued. The sieve space is constructed using only Legendre polynomials instead of using, e.g., Hermite polynomials for the latent variable and Laguerre polynomials for the measurement errors. \par
The construction of the sieve space begins with the observation that any absolutely continuous c.d.f.~$F$ on $\mathbb{R}$ can be expressed as $F(\cdot) = (H \circ G)(\cdot)$ where $H$ is an absolutely continuous c.d.f.~on $[0,1]$ and $G$ is an absolutely continuous c.d.f.~that is strictly increasing on $S(F)$. Equivalently,
\begin{align} \label{eq:cdf_representation_G}
F(\cdot) = \int_0^{G(\cdot)} h(u) \, du = \int_0^{G(\cdot)} \pi^2(u) \, du ,
\end{align}
where $h = \pi^2$ is the density of $H$ and $\pi$ is a Borel-measurable square-integrable function on $[0,1]$.\footnote{The density of $F$ may then also expressed as $f(\cdot) = (h \circ G)(\cdot) g(\cdot)$ where $g$ is the density of $G$.} Thus, e.g., with a fixed $G$ with support $S(G) = [0,\infty)$, a large enough set $\mathcal{P}$ of Borel-measurable square-integrable functions maps via \eqref{eq:cdf_representation_G} to a large enough set of distribution functions with support contained in $[0,\infty)$. \par
The sieve space is constructed by using orthonormal polynomials to approximate $\pi$. \cite{bierens2008semi} considers the following compact set of square-integrable functions
\begin{align} \label{eq:legendre_series_pi}
\mathcal{P} \coloneqq \left\{ \pi(\cdot) = \frac{1 + \sum_{\ell=1}^\infty \delta_\ell \rho_\ell(\cdot)}{\sqrt{1 + \sum_{\ell=1}^\infty \delta_\ell^2}} : \lvert\delta_\ell\rvert \le \frac{c}{1+\sqrt{\ell}\ln \ell} , \, \ell=1,2,\ldots \right\} ,
\end{align}
for some large constant $c > 0$, where $\rho_k$, defined on $[0,1]$, is a recentered and rescaled Legendre polynomial of order $k$: if $\ell_k$ is a Legendre polynomial of order $k$ on $[-1,1]$, then $\rho_k(u) := \sqrt{2k+1}\ell_k(2u-1)$ for $u \in [0,1]$ (see Section~2.2 in \cite{bierens2008semi}). This implicitly defines a compact parameter space $\mathcal{F}$ for the unknown c.d.f.~$F$ by \eqref{eq:cdf_representation_G} for some fixed $G$ and $\pi \in \mathcal{P}$. The sieve $\{\mathcal{F}_k\}_k$ is constructed by truncating the series in \eqref{eq:legendre_series_pi} at order $k$. Since we estimate two distribution functions $F_\xi$ and $F_\varepsilon$, we consider a product sieve space $\{\mathcal{F}^\xi_k \times \mathcal{F}^\varepsilon_k\}_k$ for two choices of $G$, denoted $G_\xi$ and $G_\varepsilon$. \par
As \cite{bierens2008semi} notes, the c.d.f.~$G$ not only restricts the support but also acts as an initial guess of the unknown c.d.f.\footnote{The flexibility of choosing a base distribution $G$ in \cite{bierens2008semi} is more a norm than an exception in nonparametric methods using orthogonal bases. The standard polynomial sieve on the unit interval may be seen to have the uniform distribution as an initial guess. The Hermite and Laguerre polynomial sieves take normal and exponential distributions as initial guesses, respectively. While these ``initial guesses'' are determined naturally by the weight function associated with the orthogonal bases, the construction in \cite{bierens2008semi} allows for an explicit choice by the researcher.} With prior information on the shape of the distribution functions, an educated initial guess helps reduce the approximation error resulting from low-order sieve spaces. Without any prior information, one clearly cannot expect any \emph{a priori} advantage to choosing a particular distribution. In light of the often-used standard, Hermite, and Laguerre polynomial sieves, it may be reasonable to choose a uniform distribution if the support is known to be contained in a bounded interval, a normal distribution if one is agnostic about the support, and an exponential distribution if the support is known to be contained in $[0,\infty)$. \par
We consider a sieve extremum estimator that minimizes the average (squared) contrast between two empirical ch.f.s: one based on the factual sample and the other based on simulated draws. The population criterion function is devised as follows. For any candidate pair of distribution functions $F = (F_\xi,F_\varepsilon)$, let
\[
X_{(r,s)}(F)
= (X_{(r)}(F), X_{(s)}(F))
:= (\xi(F_\xi) + \varepsilon_{(r)}(F_\varepsilon), \xi(F_\xi) + \varepsilon_{(s)}(F_\varepsilon)) ,
\]
where $\xi(F_\xi)$ is a random draw from $F_\xi$ and $\varepsilon_{(j)}(F_\varepsilon)$ is the $j$\textsuperscript{th} order statistic of $n$ i.i.d.~draws from $F_\varepsilon$. For any $t=(t_r,t_s) \in \mathbb{R}^2$, let $\varphi(t;F) \coloneqq \mathbb{E} e^{it^\top X_{(r,s)}(F)}$ denote its ch.f.~where the expectation is induced by the $n+1$ uniform draws in the simulation process described below. We consider the following population criterion function:
\begin{align}
Q(F)
\coloneqq \frac{1}{4\kappa^2} \int \operatorname*{1}\left\{t \in (-\kappa,\kappa)^2\right\} \left\lvert \psi_{X_{(r,s)}}(t) - \varphi(t;F) \right\rvert^2 \, dt , \label{eq:population_criterion_function}
\end{align}
where $\kappa > 0$ is a tuning parameter that determines the integration region. In place of the box weight $\operatorname*{1}\{(t_r,t_s) \in (-\kappa,\kappa)^2\}$, one may also consider a smooth weight. We choose the latter because it has a closed-form expression for the empirical criterion function (see Appendix~\ref{appendix:estimation}).\footnote{While a data-driven choice of $\kappa$ (as well as the polynomial order $k_N$ in Theorem~\ref{theorem:estimator_consistency} below) may be of interest, we do not explore its possibility here.} We define a simulated sieve extremum estimator:
\begin{align}
\widehat F_N \coloneqq (\widehat F_{\xi,N}, \widehat F_{\varepsilon,N}) \in \operatorname*{arg\,min}_{F \in \mathcal{F}^\xi_{k_N} \times \mathcal{F}^\varepsilon_{k_N}} \widehat{Q}_N(F) , \label{eq:sieve_estimator_def}
\end{align}
where $\{k_N\}$ is an arbitrary sequence of positive integers such that $k_N \rightarrow \infty$ and $\widehat{Q}_N(F)$ is the empirical criterion function with the empirical ch.f.~$\widehat\psi_N$ and the simulated ch.f.~$\widehat\varphi_N(\cdot;F)$ in place of $\psi_{X_{(r,s:n)}}$ and $\varphi(\cdot;F)$, respectively. $\widehat\varphi_N(\cdot;F)$ is constructed from $N$ repeated draws of $(n+1)$ uniform random variables $(V,U_1,\ldots,U_n)$ and computing the inverse transform $X_{(j)}(F_k) = F_{\xi,k}^{-1}(V) + F_{\varepsilon,k}^{-1}(U_{(j)})$ for $j \in \{r,s\}$. \par
We show that the estimator is consistent under the following set of assumptions.
\begin{assumption}[Consistency]
\label{assumption:consistency}
\leavevmode
\begin{enumerate}[label=(\alph*)]
\item $\{(X_{(r),i},X_{(s),i})\}_{i=1,\ldots,N}$ are $N$ i.i.d.~realizations of $(X_{(r)},X_{(s)})$, where $(X_{(r)},X_{(s)})$ has a bounded support.
\item $G_\xi$ and $G_\varepsilon$ are known absolutely continuous c.d.f.s with support on $\mathbb{R}$ and $[0,\infty)$, respectively.
\item Given $G_\xi$ and $G_\varepsilon$ in part (b), the pair of true underlying distributions $(F_\xi,F_\varepsilon)$ is in the closure of $\bigcup_k \mathcal{F}^\xi_k \times \mathcal{F}^\varepsilon_k$.
\item $\{(V_i,U_{1,i},\ldots,U_{n,i})\}_{i=1,\ldots,N}$ are $N$ i.i.d.~draws from $\mathcal{U}(0,1)^{n+1}$, independent of the sampling process.
\end{enumerate}
\end{assumption}
The support restriction in Assumption~\ref{assumption:consistency}(a) is a sufficient condition to ensure that the measurement errors have light tails. Despite being restrictive, this appears to be the most straightforward assumption to impose on the \emph{observables} to guarantee identification.\footnote{Note that as remarked in Section~\ref{section:model_and_identification}, one may consider alternatives to Assumption~\ref{assumption:iid_case}(b), e.g., (a.e.-)nonvanishing or analytic ch.f.s, and still achieve the same identification result. Thus, one may also consider estimating the ch.f.s over a space of analytic functions and then estimate the densities via inverse Fourier transform. Clearly, various other estimation methods can be considered. For example, one may consider a sequential approach where one estimates the measurement error distribution via \eqref{eq:identify_eps} in the first stage and then estimate the latent variable distribution using nonparametric deconvolution in the second stage. Alternatively, one may consider estimating the densities via a sieve maximum likelihood. We consider the proposed simulated sieve extremum estimator as it both allows to estimate the underlying distributions simultaneously and admits a closed-form objective function.} Note that the researcher does not have to know \emph{a priori} the true support of the observables, nor the implied bounded supports for $F_\xi$ and $F_\varepsilon$. \par
Assumption~\ref{assumption:consistency}(b) restricts the support of the base distribution for $F_\varepsilon$ to be in line with the normalization in Assumption~\ref{assumption:iid_case}(b). Assumption~\ref{assumption:consistency}(c) is a standard assumption that the model is correctly specified. Assumption~\ref{assumption:consistency}(d) specifies the simulation process. Note that the random draws $(V_i,U_{j,i})_{j,i}$ are obtained once and not repeatedly drawn across different candidate parameter values $F_{\xi,k}$ and $F_{\varepsilon,k}$ in order to ensure that the criterion function is continuous with respect to the parameter. \par
\begin{theorem}
\label{theorem:estimator_consistency}
Let $\kappa > 0$ and let $\{k_N\}_N$ be any sequence of positive integers such that $k_N \rightarrow \infty$. Under Assumptions \ref{assumption:sampling_process}--\ref{assumption:bound} and \ref{assumption:consistency}, the estimator in \eqref{eq:sieve_estimator_def} is strongly uniformly consistent, i.e.,
\[
\max\left\{ \lVert \widehat{F}_{\xi,N} - F_\xi \rVert_\infty, \lVert \widehat{F}_{\varepsilon,N} - F_\varepsilon \rVert_\infty \right\} \stackrel{\text{a.s.}}{\longrightarrow} 0 .
\]
\end{theorem}
Steps for implementing the estimator is described in Appendix~\ref{appendix:estimation} and its finite sample performance is illustrated below. As is standard in nonparametric estimation, the choice of the sieve order $k_N$ affects the approximation bias and variance of the estimator. An information-criterion-based approach analogous to that found in \cite{bierens2012semi} may be used to choose the sieve order $k_N$, although the procedure may be computationally intensive. With larger $\kappa$, the estimator is expected to perform better as it accounts for more discrepancy between the two ch.f.s. There appears to be no theoretical reason to restrict $\kappa$ except to ensure that the criterion function is well-defined, but limited simulation suggests the aggregation may come with larger variance. We do not have a useful criterion, but in light of the discussion, one may consider minimizing $\widehat{Q}_N$ over $\kappa$ as well as $F$ with a large upper bound on the parameter space for $\kappa$. From simulation studies, we find that the estimator is less sensitive to the choice of $\kappa$ as long as $\kappa$ is not too small. So we set $\kappa=1$ as its baseline value in this paper, as is done in \cite{bierens2012semi}, and explore its properties in a separate paper. We also leave inference procedures for future research.\footnote{Valid inference procedures exist in similar settings, for example, with independent measurement errors \citep{kato2021robust} or order statistics without unobserved heterogeneity \citep{menzel2013large}. A common feature they tackle is a certain lack of continuity in the inverse problems. Similar irregularity concerns may have to be addressed here.} \par
\subsection{Monte Carlo Evidence} \label{subsection:montecarlo}
To illustrate the performance of the proposed estimator, we conduct a simple Monte Carlo experiment. We begin by describing the data generating process (DGP). For observation $i$, the variable of interest $\xi_i$ is measured $n=3$ times with i.i.d.~measurement errors $\varepsilon_{1,i},\ldots,\varepsilon_{3,i}$. The measurement is constructed as $X_{j,i} = \xi_i + \varepsilon_{j,i}$. We assume that only two order statistics of ranks $r=1$ and $s=2$ remain. That is, only $X_{(1),i}$ and $X_{(2),i}$ are recorded. We repeat the process to obtain $N$ pairs of observations. The experiment is replicated $R = 500$ times to obtain $500$ random samples of size $N$ of the form $\{x_{(1),i}^{(r)},x_{(2),i}^{(r)}\}_{i=1,\ldots,N}$. \par
For the distributions of $\xi_i$ and $\varepsilon_{j,i}$, we set up a design that resembles \cite{hernandez2020estimation}'s application on eBay Motors auctions. Specifically, we use their estimated distributions of unobserved heterogeneity ($F_\xi$ in our set-up) and private values ($F_\varepsilon$ in our set-up) to calibrate the DGP for our simulation exercise. To construct these two distributions, we approximate the estimates in Figure~4 of \cite{hernandez2020estimation} with a sieve of order $6$, which resulted in almost identical distributions to those in the original article. \par
We then simulate data from the DGP and investigate finite-sample performance of our estimator. We consider sample sizes $N = 1000$, $2000$, and $4000$ with $k = 4$, $5$, and $6$, respectively. Note that by construction, there is no sieve approximation error when $N=4000$, i.e., any estimation error is associated only with sampling error. Furthermore, we set the bandwidth $\kappa =1$, $3.14$, and $5$, and the base distribution $G_\xi = \mathcal{N}(0,1/4)$ for $\xi$ and $G_\varepsilon = \mathcal{N}(2,1)_+$ for $\varepsilon$, where $\mathcal{N}(2,1)_+$ denotes the truncated normal between 0 and $\infty$. \par
We present the estimation results for $F_\xi$ and $F_\varepsilon$ with $\kappa=1$ in Figure~\ref{fig:montecarlo_maintext}. Simulation results with $\kappa=3.14$ and $5$ are similar and are presented in Figures~\ref{fig:montecarlo1} and \ref{fig:montecarlo2} in Appendix~\ref{appendix:estimation}. The true distribution functions are shown in black. We also plot some randomly selected estimates along with some box plots that illustrate the pointwise sampling error of $\widehat{F}_\xi$ and $\widehat{F}_\varepsilon$ at various evaluation points. The figure suggest that the estimator performs reasonably well under all three sample sizes, and the performance improves with larger sample size. Although the choice of $\kappa$ does not seem to affect the behavior of the estimator in any ill-behaved manner, there appear to be larger pointwise variance when $\kappa = 3.14$ and $5$ relative to the case when $\kappa = 1$. \par
\begin{figure}[h!]
\centering
\begin{subfigure}[!t]{0.33\textwidth}
\centering
\includegraphics[width=\textwidth]{figures/cdfxi_N=1000_k=4_kappa=100_R=500.pdf}
\caption*{\emph{Panel 1(a)}}
\end{subfigure}
\hfill
\begin{subfigure}[!t]{0.33\textwidth}
\centering
\includegraphics[width=\textwidth]{figures/cdfxi_N=2000_k=5_kappa=100_R=500.pdf}
\caption*{\emph{Panel 1(b)}}
\end{subfigure}
\hfill
\begin{subfigure}[!t]{0.33\textwidth}
\centering
\includegraphics[width=\textwidth]{figures/cdfxi_N=4000_k=6_kappa=100_R=500.pdf}
\caption*{\emph{Panel 1(c)}}
\end{subfigure}
\begin{subfigure}[!t]{0.33\textwidth}
\centering
\includegraphics[width=\textwidth]{figures/cdfepsilon_N=1000_k=4_kappa=100_R=500.pdf}
\caption*{\emph{Panel 2(a)}}
\end{subfigure}
\hfill
\begin{subfigure}[!t]{0.33\textwidth}
\centering
\includegraphics[width=\textwidth]{figures/cdfepsilon_N=2000_k=5_kappa=100_R=500.pdf}
\caption*{\emph{Panel 2(b)}}
\end{subfigure}
\hfill
\begin{subfigure}[!t]{0.33\textwidth}
\centering
\includegraphics[width=\textwidth]{figures/cdfepsilon_N=4000_k=6_kappa=100_R=500.pdf}
\caption*{\emph{Panel 2(c)}}
\end{subfigure}
\caption{Monte Carlo simulation results ($\kappa=1$). Panel 1 displays results for $F_\varepsilon$ and Panel 2 for $F_\xi$. Subpanels \emph{(a)}, \emph{(b)}, and \emph{(c)} correspond to sample sizes $N=1000$, $2000$, and $4000$, respectively.} \label{fig:montecarlo_maintext}
\end{figure}
\section{Conclusion} \label{section:conclusion}
This paper shows that distributions of the latent variable and measurement errors are identified nonparametrically under mild assumptions when two or more order statistics are recorded from repeated measurements with independent errors, providing a positive answer to the hypothesis in \cite{athey2002identification} for an ascending auction with unobserved heterogeneity. Our results are also applicable to other applications with unobserved heterogeneity when order statistics are observed, survey data on wage offers being a notable example. More examples include repeated experiments with type II censoring, such as in reliability testing, where consecutive low-order failure times are recorded, and estimating the effects and damages of collusion in auctions, a setting in which \cite{asker2010study} emphasizes the importance of accounting for unobserved heterogeneity. Relatedly, the identification result may be applied to extend the framework in \cite{marmer2017identifying} for testing collusion in ascending auctions. \par
\newpage
\bibliographystyle{apalike}
\bibliography{bibfile}
\newpage