EconBase
← Back to paper

Estimating Nonseparable Selection Models: A Functional Contraction Approach

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

102,148 characters · 13 sections · 65 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimating Nonseparable Selection Models: A Functional Contraction Approach

\onehalfspacing

abstract\singlespacing We propose a novel method for estimating nonseparable selection models. We show that, for a given selection function, the potential outcome distributions are nonparametrically identified from the selected outcome distributions and can be recovered using a simple iterative algorithm based on a contraction mapping. This result enables a full-information approach to estimating selection models without imposing parametric or separability assumptions on the outcome equation. We propose a two-step estimation strategy for the potential outcome distributions and the parameters of the selection function and establish the consistency and asymptotic normality of the resulting estimators. Monte Carlo simulations demonstrate that our approach performs well in finite samples. The method is applicable to a wide range of empirical settings, including consumer demand models with only transaction prices, auctions with incomplete bid data, and Roy models with data on accepted wages. Keywords: Sample Selection, Nonseparble Models, Functional Contraction, Potential Outcome Distribution, Semiparametric Estimation, Demand Estimation, Auction, Roy Models. JEL Codes: C14, C24, C51, L11, D44, J31.

Introduction

Sample selection issues arise in many empirical settings. In consumer demand studies, researchers often observe only the transaction prices of chosen products goldberg1996dealer, cicala2015does, crawford2018asymmetric, allen2019search, d2019automobile, sagl2023dispersion, cosconati2024competing. In auctions, available data may consist solely of the winning bids athey2002identification, komarova2013new, guerre2019nonparametric,allen2024resolving. In labor economics, wage data are typically observed only for individuals who choose to work gronau1974wage, heckman1974shadow, and in the Roy model roy1951some, earnings within an occupation are observed only for those who self-select into that sector.

Observing only a selected sample of outcomes---such as prices, bids, or wages---poses significant challenges for estimating two key elements. First is the selection function, which specifies how agents choose among alternatives, for example, through a consumer demand system, an auction’s winning rule, or a labor force participation decision. Second is the distribution of outcomes prior to selection, often referred to as “potential outcomes" in the literature. Flexibly estimating potential outcome distributions is crucial in many empirical contexts, such as analyzing price distributions to understand firms' pricing strategies and wage distributions to examine inequality.

Our paper proposes a new approach to estimating nonseparable selection models by exploiting a previously unrecognized one-to-one mapping between the outcome distributions before and after selection. Our key contribution is a constructive identification result showing that, for a given selection function, the potential outcome distributions are nonparametrically identified from the selected outcome distributions and can be recovered using a simple iterative algorithm. Consequently, the only remaining object to be estimated is the selection function, which can be recovered from observed choice patterns. Our method enables a full-information approach to estimating selection models without imposing parametric or separability assumptions on the outcome equation. In addition, it does not rely on an excluded variable in the selection equation.

Formally, we consider a discrete choice problem in which each alternative is associated with a potential outcome distribution. A selection function maps a vector of realized potential outcomes to a probability distribution over the alternatives. For example, in the consumer demand setting, each alternative represents a product, and the potential outcome is the offered price, with the selection function micro-founded by the consumer's utility maximization problem. We allow the outcome equations to be fully nonparametric with nonseparable error terms and to vary flexibly across different alternatives. We assume that potential outcomes across alternatives are independent conditional on observable and unobservable characteristics, which allows for correlation across outcomes when conditioning only on observables. In addition, we allow the unobservable component in the outcome equation to enter directly into agents’ preferences over alternatives, thereby capturing selection on unobservables.

We analyze how the selection model maps the potential outcome distributions to the distributions of selected outcomes and seek to invert the mapping. The key insight of our approach is that, given the selection model and potential outcome distributions across all alternatives, we can derive the likelihood of an outcome being selected at each price. Conversely, if this selection likelihood were known, we could recover the potential outcome distributions from the selected outcome distributions using Bayes’ rule. This two-way relationship characterizes a fixed-point problem.

Building on this intuition, we construct an operator whose fixed point is the potential outcome distributions and establish sufficient conditions for it to be a functional contraction (Theorems (ref) and (ref)). Our results imply that, given the selection function and the distributions of selected outcomes, we can nonparametrically identify the potential outcome distributions. Moreover, this identification result is constructive: starting with any initial guess for the potential outcome distributions, we iteratively apply the operator. This process converges to the potential outcome distributions associated with the selection function.

We then embed this identification result into a two-step estimation strategy for the unobserved potential outcome distributions and the parameters of the selection function. In the first step, we estimate the selected outcome distribution conditional on both observable and unobservable covariates. In the second step, we propose a nested fixed-point algorithm to estimate the parameters of the selection function: in the inner loop, for any candidate selection function, we recover the potential outcome distributions by iterating the operator, while in the outer loop we search for the parameter values that maximize the likelihood of the observed choice patterns. The potential outcome distributions are then obtained by reapplying the fixed-point algorithm at the estimated parameters of the selection function.

We establish the consistency and asymptotic normality of the proposed estimators in Theorems (ref) and (ref). To examine their finite sample properties, we conduct Monte Carlo simulations across various designs of the outcome equation. Our results show that the biases in our estimators are generally small, and the standard deviation decreases as the sample size increases across all simulation designs. Our nonparametric estimation of the potential outcome distributions outperforms the classic Heckman parametric two-step approach and the quantile selection model of arellano2017quantile with linear quantile functions and a Gaussian copula, particularly when the outcome equation contains nonseparable error terms. In addition, we show that our approach does not require an excluded variable in the selection equation and remains robust even when the selection function is misspecified by the econometrician.

In a companion paper cosconati2024competing, we apply our method to estimate consumer demand for auto insurance products when only transaction prices are observed. We nonparametrically estimate the offered price distribution for each insurance company and allow these distributions to vary fully flexibly across firms. The substantial heterogeneity in the recovered price distributions reflects differences in firms’ information technologies and cost structures, which are key primitives we estimate through a supply-side competition model. We omit the details of this application here and refer readers to cosconati2024competing for the full empirical setup and results.

\paragraph{Related Literature}

Our paper contributes to the extensive theoretical literature on sample selection models. An early solution to sample selection bias is full information maximum likelihood (FIML) estimation based on parametric assumptions, as in heckman1974shadow and lee1982some (lee1982some, lee1983generalized). More commonly employed methods for sample selection models are the two-step control function approach pioneered by heckman1976common (heckman1976common, heckman1979sample). A substantial body of theoretical work has been developed to relax the distributional assumptions in the two stages of the estimation procedure ahn1993semiparametric,andrews1998semiparametric,chen2003semiparametric, das2003nonparametric, newey2007nonparametric,newey2009two,fernandez2024nonseparable,chernozhukov2025distribution. For a comprehensive survey of semiparametric two-step estimation methods for selection models, see vella1998estimating.

Compared with the existing methods, our approach offers several key advantages. First, we allow the outcome equation to be nonparametric and nonseparable in the error terms, and we exploit the full information in the selected outcome distribution to recover the entire distribution of potential outcomes. newey2007nonparametric and fernandez2024nonseparable use control function approaches to correct for sample selection in nonseparable models with binary and censored selection rules, respectively, and they focus on identifying certain global and local parameters of the outcome distribution.

Second, our method accommodates fully heterogeneous effects of covariates on outcomes, whereas most existing approaches that estimate conditional mean models restrict covariates to affecting only the location of the outcome distribution.\footnote{An exception is the recent paper by chernozhukov2025distribution, which proposes a semiparametric generalization of the Heckman selection model that allows for rich forms of heterogeneity in the effects of covariates on both outcomes and selection.} More recently, arellano2017quantile propose a method to correct for sample selection in quantile regression models by modeling the copula of the error terms in the outcome and selection equations. Although identification under more general settings is discussed, their estimation strategy primarily considers linear quantile models and copulas characterized by a low-dimensional set of parameters.

Third, our approach does not require an instrument that shifts the choice probability without entering the outcome equation, which is central to identification in two-step methods. In practice, finding such an instrument can be challenging (see vella1998estimating for further discussion).\footnote{d2013another and d2018extremal develop estimation methods for semiparametric sample selection models without an instrument or a large-support regressor, leveraging the independence-at-infinity assumption.} Furthermore, our method does not rely on identification-at-infinity arguments. Instead, when the conditioning set of variables includes an unobservable component, our method requires instruments for estimating the distribution of selected outcomes conditional on the unobservable in the first step. We adopt the measurement error framework in hu2008identification and hu2008instrumental, where the key requirement is to find instruments such that, conditional on the latent variable, the outcome and the instruments are independent.

Estimation of the selection function in our model is closely related to the demand estimation literature following the seminal work of berry1994estimating and berry1995automobile. In particular, observed choice patterns play the same role as market shares in recovering consumer preference parameters. Our method addresses the problem of missing full price menu that arises in many demand estimation contexts goldberg1996dealer,cicala2015does,crawford2018asymmetric,allen2019search, d2019automobile, sagl2023dispersion, cosconati2024competing, an issue that is especially relevant in the presence of price discrimination or personalized pricing.\footnote{A recent paper by d2019automobile addresses a related challenge in demand estimation under unobserved price discrimination by imposing supply-side restrictions, such as assumptions about firm conduct (e.g., Bertrand competition), and assuming identical costs across consumers. }

At a broader conceptual level, our reliance on the structural restrictions implied by the selection model resonates with the nonparametric identification literature on auction models with missing bids. athey2002identification show that the symmetric independent private values (IPV) models are identified with the transaction price by exploiting a one-to-one mapping between an order statistic and its parent distribution. komarova2013new analyzes asymmetric second-price auctions where only the winning bids and the winner's identity are observed. A related result for generalized competing risks models can be found in meilijson1981estimation. More recently, guerre2019nonparametric examine nonparametric identification of symmetric IPV first-price auctions with only winning bids, accounting for unobserved competition. In these auction models, the selection rule is deterministic conditional on bids (the highest bidder wins), which allows order-statistic arguments to be applied. In contrast, our selection model assigns a probability distribution over alternatives and is therefore closer in spirit to multi-attribute auction environments (see e.g., krasnokutskaya2020role). Moreover, our framework can flexibly accommodate asymmetries across alternatives, whereas bidder asymmetries are known to pose significant challenges in auction models (see the discussion in the handbook chapter by athey2007nonparametric).

\paragraph{Outline} The rest of the paper is organized as follows. Section (ref) formally introduces our model and provides an illustrative example. Section (ref) presents the main theoretical results. In Section (ref), we describe our estimation strategy and establish the asymptotic properties of the estimators. Section (ref) reports results from our Monte Carlo simulations, and Section (ref) discusses the empirical application in cosconati2024competing. Section (ref) concludes. All proofs are collected in the appendix.

Model

In Sections (ref)--(ref), all analyses are conditional on a vector of characteristics $(x, x^\ast)$, where $x$ denotes observables and $x^\ast$ denotes unobservables. The structure of the model and the main theoretical results do not depend on whether the conditioning set includes unobserved components, although the presence of unobservables introduces additional challenges for estimation, which we address in Section (ref). Because all results in these two sections are stated conditional on $(x, x^\ast)$, we omit these variables from the notation to simplify exposition.

Throughout the paper, we use a consumer demand example to illustrate the main results and clarify key ideas. In this context, potential outcomes are offered prices, while selected outcomes are selected (or transaction) prices. We use these terms interchangeably when discussing the demand example. Nevertheless, our approach is broadly applicable to a wide class of selection models.

Consider a discrete choice problem. There is a finite set of alternatives $\mathcal{J}=\{1, \cdots, J\}$. Each alternative is associated with a price distribution. Let $G_j \in \Delta ([\underline p_j, \overline p_j])$ represent the price distribution associated with alternative $j$, where $\Delta (Y)$ denotes the set of all cumulative distribution functions over a set $Y\subset \mathbb R$. We assume that $p_j\sim G_j$ are independently distributed across alternatives (conditional on $x$ and $x^\ast$). The collection of $G_j$ is denoted by $G=\prod_{j\in\mathcal J} G_j$. We refer to $G$ as the offered price distribution.

A selection function is denoted by $f=(f_1,f_2,\cdots,f_J)$ where $f_j$ maps the prices of alternatives $\boldsymbol{p}=(p_1, \cdots, p_J)$ to a strictly positive probability of selecting alternative $j \in \mathcal{J}$.\footnote{The assumption that the probability of selecting each alternative is strictly positive is analogous to the overlap assumption in the treatment effect literature, which requires each individual to have a positive probability of receiving each treatment level. This assumption is crucial for recovering the offered price distribution. To illustrate, consider a scenario where $f_j = 0$ whenever $p_j$ falls within a certain subset of $[\underline{p}_j, \overline{p}_j]$. In this case, any $p_j$ within that subset would not be observed in the data, making it impossible to identify $G_j$ within that subset without introducing additional assumptions. } We assume that the selection function is continuously differentiable, $$f_j\in \mathscr{C}^1\colon \prod_j [\underline p_j, \overline p_j]\to (0,1],$$ with $\sum_{j\in\mathcal{J}}f_j\leq 1$. Here, the inequality allows for the case with an outside option. The selection function is a primitive of the model. To provide a microfoundation, for example, $f$ might be derived from a consumer's utility maximization problem as illustrated in Section (ref).

Let $\boldsymbol{p}_{-j}=(p_1, \cdots, p_{j-1}, p_{j+1}, \cdots, p_J)$ denote the vector of prices excluding $j$'s price. The probability of selecting $j$ conditional on $p_j$ is given by

align[align omitted — 128 chars of source]

where $Pr_j(\cdot;G)$ is a function defined on $[\underline p_j, \overline{p}_j]$. Independent prices across different alternatives (conditional on $x$ and $x^\ast$) allow us to express the joint distribution of $\boldsymbol{p}_{-j}$ as the product of their individual marginal distribution functions.

Let $\tilde G_j\in \Delta ([\underline p_j, \overline p_j])$ represent the price distribution conditional on selecting alternative $j$. We derive $\tilde G_j$ using Bayes' rule:

equation[equation omitted — 162 chars of source]

Note that $G_j$ and $\tilde G_j$ share the same support, as selection function $f_j$ is strictly positive. Let $\tilde G=\prod_{j\in\mathcal J} \tilde G_j$ and we call $\tilde G$ selected price distribution. Equations (ref) and (ref) define a mapping from $G$ to $\tilde G$. Let $F\colon \prod_j \Delta([\underline p_j, \overline p_j])\to \prod_j \Delta([\underline p_j, \overline p_j])$ denote this mapping, i.e., $\tilde G=F(G)$.

In many empirical settings, researchers have access only to the selected price distribution, such as the distribution of transaction prices, accepted wages, or winning bids. However, the key primitives of interest are often the offered price distribution, such as the distributions of posted prices, wage offers, or submitted bids. How to recover the offered price distribution $G$ from the selected price distribution $\tilde{G}$? Note that both $G$ and $\tilde{G}$ are collections of $J$ cumulative distribution functions. Therefore, the cardinality of unknowns and constraints are exactly the same in Equation (ref) (assuming the selection function is known). Since a cumulative distribution function is an infinite-dimensional object, the key challenge is solving for a collection of infinite-dimensional objects entangled in a nonlinear system. We will explore this in detail in Section (ref).

An Illustrative Example

We now present a simple example to illustrate the key assumptions of our model and compare them to the standard assumptions in the literature. Consider a consumer choosing between two products, $j=1, 2$, to maximize her utility. The utilities from products 1 and 2 are:

align[align omitted — 164 chars of source]

where $p_j$ represents the price of product $j$ for this consumer, and $\varepsilon_j$ is an idiosyncratic utility shock. We also allow an unobservable characteristic $x^\ast$ to enter directly into the utility from product 1. This captures settings in which unobserved traits affect preferences across alternatives. For example, in insurance markets, high-risk consumers may prefer products offered by certain firms; similarly, in labor markets, workers with higher productivity may prefer certain types of jobs. Our model can flexibly allow utility to depend on observable consumer attributes as well, but we omit these terms here for simplicity of presentation.

In this model, the price sensitivity parameter $\gamma$, coefficient $\kappa$, and the distribution of $\varepsilon_j$ determine the selection function $f$, for any fixed $x^\ast$. If $\varepsilon_1-\varepsilon_2 \sim \mathcal{N}(0, 1)$, the selection function for product 1 takes the standard binary probit form:

align*[align* omitted — 104 chars of source]

where $\Phi_{\mathcal{N}}$ denotes the CDF for standard normal distribution. For simplicity, we denote the difference in unobservables across the two utilities as $\tilde{\varepsilon}=x^\ast\kappa + (\varepsilon_1-\varepsilon_2)$.

In this illustrative example, we consider a simple linear outcome equation with an additive error term. For each product $j=1, 2$, the price is generated by the following equation:

align[align omitted — 115 chars of source]

where $x$ denotes observable characteristics, $x^\ast$ denotes unobservable characteristics, and $\eta_j$ is an idiosyncratic shock. We define the composite error in the pricing equation as $\eta_j^\ast \equiv x^\ast\delta_{j} + \eta_j$, which, for simplicity, is assumed to be independent of $x$. We assume that the true underlying price shocks $\eta_1$ and $\eta_2$ are independent. However, when $x^\ast$ is unobserved by the econometrician, the composite errors $\eta_1^\ast$ and $\eta_2^\ast$ may be correlated through $x^\ast$.

Suppose the econometrician observes the price of product 1 only when it is chosen by the consumer. We derive the conditional mean of $p_1$ given that it is observed:

align[align omitted — 484 chars of source]

The conditioning term $x\beta^\ast + \varepsilon^\ast >0$ in Equation (ref) represents the reduced-form selection model typically seen in the literature. Sample selection issue arises when $\eta_1^\ast$ and $\varepsilon^\ast$ are correlated, so that $E(\eta_1^\ast | x\beta^\ast + \varepsilon^\ast >0) \neq 0$. In the two-step estimation literature, researchers often impose assumptions on the joint distribution of $(\varepsilon^\ast, \eta_1^\ast, \eta_2^\ast)$. For example,

align*[align* omitted — 401 chars of source]

We now take a closer look at the correlation between the error in the selection model ($\varepsilon^\ast$) and the error in the outcome equation ($\eta_1^\ast$). Specifically,

align[align omitted — 289 chars of source]

Equation (ref) shows that the error term $\eta_1^\ast$ directly enters the composite error $\varepsilon^\ast$, generating the first term $\gamma var(\eta_1^\ast) \neq 0$ unless $\gamma = 0$. This correlation is by construction in selection models, as agents make decisions after observing the potential outcomes. The second term in Equation (ref) reflects the correlation between the composite errors in the two outcome equations. It is important to emphasize that our independence assumption—that $p_j$ are independently distributed across alternatives conditional on $(x, x^\ast)$—requires independence of the underlying shocks $\eta_1$ and $\eta_2$, not of the composite errors $\eta_1^\ast$ and $\eta_2^\ast$. Thus, our framework accommodates settings in which prices across alternatives are correlated conditional on observables. For instance, if $x^\ast$ captures a worker’s unobserved productivity, then wage offers from two firms may appear correlated when $x^\ast$ is not conditioned on.

Another common concern regarding selection bias arises from potential correlation between errors in the outcome equation (e.g., $\eta_1^\ast$) and those in the structural selection model (e.g., $\tilde{\varepsilon}$), as represented by the third term in Equation (ref). For example, unobserved productivity factors may create correlation between a worker's willingness to work and their wage. Our model also accommodates this type of correlation by allowing unobservable characteristics $x^\ast$ to enter both the selection equation and the outcome equation.

This simple two-product example illustrates how our notation for the selection function and the offered price distribution maps to the conventions commonly used in the existing literature, and highlights how our setting accommodates all key sources of correlation that give rise to selection bias. The general framework introduced in Section (ref) is substantially more flexible than this illustrative case. In particular, our model allows the outcome equation in (ref) to be fully flexible and nonparametrically specified with a nonseparable error term. Moreover, we impose minimal assumptions on the selection function. It can accommodate nonparametric, nonseparable relationships between observable and unobserved errors, offering much greater flexibility than the utility specification in Equations (ref) and (ref); in fact, it does not even need to be derived from a utility maximization problem. Our framework also allows for alternative-specific unobserved heterogeneity (such as product quality or job amenities), which is a desirable feature in many empirical contexts.

Main Results

We now present our main theoretical results on how to recover the offered price distribution from the selected price distribution. As the selected price distribution is derived from the offered price distribution through Bayes' rule in Equation (ref), we can first invert Equation (ref):

equation[equation omitted — 176 chars of source]

Note that if the selection probability $Pr_j(\cdot;G)$—that is, the probability of selecting product $j$ conditional on its offered price—were known, then recovering the offered price distribution from Equation ((ref)) would be straightforward.

We illustrate this inversion process using the simulated example in Figure (ref). The red solid line plots the selected price density for an alternative. Dividing this density by the probability that the alternative is chosen at each price, $Pr_j(p;G)$, yields the unnormalized offered price density shown by the blue dashed line. The gap between these two densities captures the selection mechanism: when a lower price is offered, agents are more likely to accept it, whereas higher prices make them more likely to choose other alternatives. The offered price distribution shown by the blue solid line is then obtained after normalization, which corresponds to the denominator in Equation (ref).

figure[figure omitted — 360 chars of source]

However, $Pr_j(\cdot;G)$ is not known, because it depends on the offered price distribution $G$, which we seek to recover. A tentative solution is to start with a conjecture $\Psi$ for the offered price distribution and use it to compute the implied selection probability $Pr_j(\cdot;\Psi)$. Equation ((ref)) then delivers an updated conjecture of the offered price distribution. This procedure, which maps a conjectured offered price distribution into its updated version, defines an operator $T\colon \prod_j \Delta([\underline p_j, \overline p_j])\to \prod_j \Delta([\underline p_j, \overline p_j])$ as follows.

equation[equation omitted — 180 chars of source]

where $\Psi=(\Psi_1,\Psi_2,\cdots,\Psi_J)\in \prod_j \Delta([\underline p_j, \overline p_j])$. Importantly, if the conjecture $\Psi$ is correct, i.e., $\Psi=G$, then the selection probability $Pr_j(\cdot;\Psi)$ is correctly specified, ensuring that the updated conjecture $T\Psi$ also equals $G$. Thus, the offered price distribution $G$ is a fixed point of the operator $T$.

The operator $T$ is a contraction if there exists some real number $0\leq \rho<1$ such that for all $\Psi, \Phi\in \prod_j \Delta([\underline p_j, \overline p_j])$, $$D(T\Psi, T\Phi)\leq \rho D(\Psi, \Phi),$$ given some metric $D$.\footnote{We adopt the convention that $+\infty$ and $+\infty$ are not comparable, but $c<+\infty $ for any $c\in \mathbb R_+$.} In the remainder of this section, we first construct the metric $D$ and then characterize the modulus $\rho$. We discuss several special cases of our model at the end.

Constructing the Metric

We begin by defining a metric in the set of all cumulative distribution functions for alternative $j$. Let $\Psi_j$ and $\Phi_j$ denote two probability measures in $\Delta ([\underline p_j, \overline p_j])$. Recall that two probability measures $\Psi_j$ and $\Phi_j$ are equivalent, denoted $\Psi_j\sim \Phi_j$, if they are absolutely continuous with respect to each other. Since $f_j>0$, $G_j\sim \tilde G_j$. When $\Psi_j\sim \Phi_j$, the Radon-Nikodym derivative, $$\frac{d\Psi_j}{d\Phi_j}\colon [\underline p_j, \overline p_j] \to \mathbb (0,\infty),$$ exists, as guaranteed by the Radon-Nikodym Theorem. If both $\Psi_j$ and $\Phi_j$ have continuous densities, the Radon-Nikodym derivative simplifies to the ratio of densities: $$\frac{d\Psi_j}{d\Phi_j}(p)=\frac{\Psi_j'(p)}{\Phi_j'(p)}.$$ Note that $$\Psi_j=\Phi_j\quad\Leftrightarrow \quad \frac{d\Psi_j}{d\Phi_j}(p)=1 \quad\quad \Phi_j\text{-a.e. }$$

In the space $\Delta([\underline p_j, \overline p_j])$, we define a metric $d\colon \Delta([\underline p_j, \overline p_j])\times \Delta([\underline p_j, \overline p_j])\to [0,+\infty]$ to simplify the analysis.\footnote{This metric is a variant of the Thompson metric thompson1963certain. The Thompson metric between two functions $s,q\in \mathbb R^{Y}$ is $$d_{Thompson}(s,q)=\max\{\ln\sup\frac{s(y)}{q(y)},\ln\sup\frac{q(y)}{s(y)}\}.$$} \[d(\Psi_j,\Phi_j)=

cases\ln\operatorname*{ess\,sup}_{y\in [\underline p_j, \overline p_j]}\frac{d\Psi_j}{d\Phi_j}(y)+\ln\operatorname*{ess\,sup}_{y\in [\underline p_j, \overline p_j]}\frac{d\Phi_j}{d\Psi_j}(y), & if \Psi_j \sim \Phi_j,\\ +\infty & otherwise.

\] Given our operator $T$ in Equation (ref), for all $\Psi_j,\Phi_j\in \Delta([\underline p_j, \overline p_j])$, $$(T\Psi)_j\sim \tilde G_j\sim (T\Phi)_j.$$ Thus, $$d((T\Psi)_j,(T\Phi)_j)= \ln\operatorname*{ess\,sup}_{p_j}\frac{d(T\Psi)_j}{d(T\Phi)_j}(p_j)+\ln\operatorname*{ess\,sup}_{p_j}\frac{d(T\Phi)_j}{d(T\Psi)_j}(p_j).$$ The selected price distribution $\tilde{G}_j$ appears in both $(T\Psi)_j$ and $ (T\Phi)_j $. As a result, $\tilde{G}_j$ cancels out in the distance above. Moreover, the denominator in our operator is a normalizing factor, which is also canceled out after we take the sum of log ratios. Consequently, the distance between $(T\Psi)_j$ and $(T\Phi)_j$ relies only on the ratio between selection probabilities: $$d((T\Psi)_j,(T\Phi)_j)\leq \sup_{p_j}\ln \frac{Pr_j(p_j;\Psi)}{Pr_j(p_j;\Phi)}+\sup_{p_j} \ln \frac{Pr_j(p_j;\Phi)}{Pr_j(p_j;\Psi)}.$$ The equality holds when $\tilde G_j$ admits full support on $[\underline p_j,\overline{p}_j]$. Since $f_j>0$ is continuous with compact support, $Pr_j$ is bounded away from $0$. Thus, $d((T\Psi)_j,G_j)$, $d((T\Psi)_j,\tilde G_j)$ and $d((T\Psi)_j,(T\Phi)_j)$ are all finite.

Next, we define a metric in the space $ \prod_j \Delta([\underline p_j, \overline p_j])$ by taking the maximum distance among all alternatives: $$D(\Psi, \Phi)=\max_{j\in\mathcal{J}} d(\Psi_j, \Phi_j)$$ for any $\Psi, \Phi\in \prod_j \Delta([\underline p_j, \overline p_j])$. From now on, we work with the metric space $(\prod_j \Delta([\underline p_j, \overline p_j]), D)$.

Functional Contraction

For $j\in\mathcal{J}$, we define the maximum semi-elasticity difference as

equation[equation omitted — 231 chars of source]

The quantity $\frac{\partial\ln f_j}{\partial p_j}$ measures the sensitivity of the log choice probability to price and is therefore referred to as the semi-elasticity. Let $$\rho=\frac{J-1}{4}\max_{j\in\mathcal{J}} (\overline p_j-\underline p_j) M_{j}.$$

thmIf $\rho<1$, the operator $T$ is a contraction with modulus less than $\rho$.
proofSee Appendix (ref).

Theorem (ref) establishes a key identification result for selection models. By the Banach fixed point theorem, whenever $\rho<1$, the operator $T$ admits a unique fixed point, the offered price distribution $G$. Theorem (ref) implies that we can nonparametrically identify the potential outcome distributions $G$ from any selected outcome distribution $\tilde{G}$, given the selection function $f$. Notably, the theorem imposes no assumptions on the functional form of the potential outcome distributions, allowing the outcome equation to be fully nonparametric and nonseparable in the error terms, and it applies to any form of the selection function $f$ satisfying the smoothness and positivity condition.

Moreover, the result in Theorem (ref) provides a constructive method for solving for $G$. Take any $\Psi\in \prod_j \Delta([\underline p_j, \overline p_j])$, by Theorem (ref), $$D(T^n \Psi,G)=D(T^n \Psi,TG)\leq \rho D(T^{n-1} \Psi,G)\leq \rho^{n-1} D(T \Psi,G),$$ where $D(T \Psi,G)$ is finite. This implies $$\lim_{n\to \infty}D(T^n \Psi,G)=0, $$ $$\lim_{n\to \infty}T^n \Psi=G. $$ Thus, we can simply take an initial guess for the potential outcome distributions and iteratively apply the operator. As the number of iterations approaches infinity, this process converges to the potential outcome distributions associated with the selection function.

The crux and the bulk of the proof for Theorem (ref) is to provide a bound on the ratio $$\sup_{\Psi,\Phi\in \prod_j \Delta([\underline p_j, \overline p_j])}\frac{D(T\Psi,T\Phi)}{D(\Psi,\Phi)}.$$ This is difficult for two reasons. First, the domain of the supremum, $\prod_j \Delta([\underline p_j, \overline p_j])$, is a large space. For instance, if $J=10$, the supremum is over 20 functions. Second, the selection function $f$ is very general as we have not imposed much structure. In the proof of this theorem in Appendix (ref), we employ a change-of-measure technique, also known as the tilted measure, and combine it with insights from transportation problem.\footnote{In Appendix (ref), we establish a connection between our contraction result and quantal response equilibria mckelvey1995quantal.}

Note that the condition of Theorem (ref) is a joint constraint on the selection function and the price range. The bound on the modulus, $\rho$, consists of the product between the number of alternatives, the price range $\overline p_j-\underline p_j$, and the maximum semi-elasticity difference.\footnote{Note that by definition $\rho$ is unitless. Changing the unit of price does not affect $\rho$. } Our condition requires this product to be small. We emphasize that this is a sufficient condition; the mapping may remain a contraction even when the product is above 1.

To build intuition for the role of each component in the bound on the modulus, we now examine these factors in turn. First, if we expand the support from $[\underline p_j, \overline p_j]$ to $[\underline p'_j, \overline p'_j]$, where $$\underline p'_j< \underline p_j< \overline p_j<\overline p'_j,$$ while keeping $\tilde G$ unchanged, $\rho$ weakly increases, making it more difficult for the operator $T$ to contract. This result is intuitive. The larger domain $ \prod_{k\neq j} \Delta([\underline p_k, \overline p_k])\times \Delta([\underline p'_j, \overline p_j'])$ nests more collections of probability measures, making it more challenging to control $\frac{D(T\Psi,T\Phi)}{D(\Psi,\Phi)}$ for all $\Psi$ and $\Phi$ in this domain.

Second, the maximum semi-elasticity difference $M_j$ is small when the own-price responsiveness of product $j$ stays stable even as competitors’ prices vary. This arises, for example, when product $j$ is highly differentiated and exhibits weak substitutability with other products. In the extreme case where the demand for product $j$ is completely independent of competitors’ prices, we have $M_j=0$. In this situation, the selection probability for product $j$ does not depend on the distribution of competing prices $\boldsymbol{p}_{-j}$, so the integral over $\boldsymbol{p}_{-j}$ in Equation (ref) reduces to a constant. Consequently, the offered price distribution for product $j$ can be obtained directly from its selected price distribution given the selection function $f$.

Extending this intuition to more general settings, when $M_j$ is small, the influence of competing offer distributions on the demand for product $j$ is limited. As a result, if we plug a conjecture $\Psi$ into the selection probability, even if this conjecture is not perfectly accurate, the resulting selection probability will remain close to the truth. Using this approximate selection probability to recover the offered price distribution therefore yields an estimate that is also close to the true distribution. Thus, smaller values of $M_j$ make it easier for the operator $T$ to be a contraction.

Finally, the effect of $J$ on $\rho$ is more subtle because $J$ not only directly affects the dimensionality of the unknown offered price distributions but also influences the maximum semi-elasticity difference $M_j$. For example, consider the multinomial logit model, arguably the most popular model for discrete choices due to its analytical form and ease of estimation:

equation[equation omitted — 128 chars of source]

where $\gamma$ represents the consumer's price sensitivity. We derive the semi-elasticity for the logit model, $$\frac{\partial\ln f_j(p_j,\boldsymbol{p}_{-j})}{\partial p_j}=\gamma (1-f_j(\boldsymbol{p})) .$$ When $J$ is large, the choice probability for each alternative may be small, so that the log derivative is approximately equal to $\gamma$. As a result, the maximum semi-elasticity difference is close to $0$.\footnote{Empirically, in markets with a large number of alternatives, a small number of goods may capture large market shares, while the rest have only negligible shares gandhi2023estimating. In such settings, the maximum semi-elasticity difference can still be small, because the alternatives with large market shares tend to exhibit weak substitutability with those that are rarely selected. } Hence, the modulus can remain small even when the number of alternatives $J$ is large.

Special Cases

Thus far, we have not imposed any structure on the selection function. For a general selection function, we have to take the supremum over $(\boldsymbol{p}_{-j},\boldsymbol{p}_{-j}')$ to compute the maximum semi-elasticity difference. Now we impose an assumption on the selection function to determine where the supremum is attained.

asmp[Log Supermodularity] For all $j\in\mathcal{J}$ and $p_j\in[\underline p_j, \overline p_j]$, $\frac{\partial\ln f_j(p_j,\boldsymbol{p}_{-j})}{\partial p_j}$ is weakly increasing in each $p_k$ with $k\neq j$.

Given log supermodularity, the maximum semi-elasticity difference is attained at the boundary, $$M_{j}=\sup_{p_j} \bigg|\frac{\partial\ln f_j(p_j,\overline{\boldsymbol{p}}_{-j})}{\partial p_j}-\frac{\partial\ln f_j(p_j,\underline{\boldsymbol{p}}_{-j})}{\partial p_j} \bigg|.$$ What is left in the definition of maximum semi-elasticity difference is the supremum over $p_j$. It turns out that we can use $\overline p_j-\underline p_j$ in the definition of $\rho$ to eliminate the supremum over $p_j$ and give a tighter bound. The result is as follows. Let $$\rho^*=\frac{J-1}{4}\max_{j\in\mathcal{J}} [\ln f_j(\overline{\boldsymbol p})-\ln f_j(\underline p_j,\overline{\boldsymbol p}_{-j})-\ln f_j(\overline p_j,\underline{\boldsymbol p}_{-j})+\ln f_j(\underline{\boldsymbol p})].$$

thmSuppose that Assumption (ref) holds. If $\rho^*<1$, the operator $T$ is a contraction with modulus less than $\rho^*$.
proofSee Appendix (ref).

Under Assumption (ref), the modulus $\rho^\ast$ takes a much simpler form and is straightforward to compute. The log-supermodularity assumption holds in models widely adopted by empirical researchers. For example, the multinomial logit model satisfies Assumption (ref). The binary probit model described in Section (ref) also satisfies the log-supermodularity condition in Assumption (ref), so Theorem (ref) applies.\footnote{To see this, we compute the log derivative for the binary probit model: $$\frac{\partial\ln f_1(p_1,p_2)}{\partial p_1}=\frac{\gamma\phi_{\mathcal{N}}(\Delta)}{1-\Phi_{\mathcal{N}}(\Delta)},$$ $$\frac{\partial^2\ln f_1(p_1,p_2)}{\partial p_1 \partial p_2}=\gamma^2 \frac{d}{d\Delta}\bigg[\frac{\phi_{\mathcal{N}}(\Delta)}{1-\Phi_{\mathcal{N}}(\Delta)}\bigg],$$ where $\Delta=\gamma(p_2-p_1)$ and the term in the square bracket is known as the hazard rate or inverse Mills ratio. As Gaussian satisfies increasing hazard rate baricz2008mills, the log-supermodularity condition in Assumption (ref) holds.} However, Assumption (ref) may not hold for probit models with three or more alternatives; in such cases, the more general results in Theorem (ref) can be applied.

For illustration, we compute $\rho^\ast$ for the simple multinomial logit model in Equation (ref) with two alternatives. We fix $\gamma=1$, normalize $\xi_1=0$, and set the lower bounds of both price distributions to zero. In Figure (ref), the x- and y-axes represent the upper bounds of the price distributions for alternatives 1 and 2, respectively. For various values of $\xi_2$, we plot the contour of the region where $\rho^\ast>1$, and we highlight this region with shaded areas.

Figure (ref) shows that as the price range widens, the bound on the modulus is more likely to exceed 1, so the shaded areas are concentrated in the upper‐right corner of the figure. Moreover, as alternatives become more differentiated (i.e., as $\xi_2$ increases), the shaded regions shrink. This pattern is consistent with our theoretical results: the maximum semi-elasticity difference $M_j$ is small when alternatives are highly differentiated, which allows the operator $T$ to be a contraction even over generally wider price ranges. Finally, we reiterate that $\rho^\ast<1$ is only a sufficient condition. Even if this condition is violated, the operator may still be a contraction.

figure[figure omitted — 573 chars of source]
comment

To summarize, our contraction results provide a novel method for identifying the potential outcome distribution from the selected outcome distribution, given any selection function $f$---whether parametric or nonparametric, and regardless of whether it is microfounded in a utility maximization problem. Moreover, the identification is constructive: starting with an initial guess, iterative application of the operator converges to the potential outcome distributions associated with the selection function. This powerful identification result exhausts all the information contained in the selected outcome distributions. Then the estimation of the selection model essentially reduces to recovering the selection function from observed choice patterns. We discuss the estimation strategy in the next section.

Estimation

We now turn to the estimation of the model's primitives, which include (1) the unobserved offered price distributions $G$ and (2) the parameters in the selection function $f$. We propose a two-step estimation procedure. In the first step, we estimate the selected outcome distribution conditional on both observable and unobservable covariates using instruments. Once the selected outcome distribution has been recovered, for any given selection function $f$, the potential outcome distribution can be recovered iteratively using the contraction mapping results in Section (ref). The second step nests this fixed-point problem within an estimation routine that recovers the parameters of the selection function from agents' observed choice patterns. Using the resulting parameter estimates, we then re-run the fixed-point algorithm to recover the offered price distribution $G$.

In the data, for each individual $i$, we observe their choice $y_i \in \mathcal{J}$ and the price of the selected product $p_i$. Let $x_{ij}$ denote a vector of observable characteristics, and define $x_i=(x_{i1}', \cdots, x_{iJ}')'\in X$. We let $x_i^\ast \in X^\ast$ denote an unobservable characteristic which may affect both the selection decision and the distribution of potential outcomes.

We assume that the selection function $f$ is derived from a standard multinomial choice model with an indirect utility given by $$u_{ij}=v_j(p_{ij}, x_{ij}, x_i^\ast, \varepsilon_{ij}; \theta), $$ where $v_j$ is a known function indexed by a finite-dimensional parameter vector $\theta$. Here, $p_{ij}$ is the offered price of alternative $j$ for individual $i$, and the vector of unobserved shocks $\varepsilon_i=(\varepsilon_{i1}, \cdots, \varepsilon_{iJ})$ follows a known joint distribution, such as Type 1 extreme value. Each individual chooses the alternative that maximizes utility, and the selection function $f$ is captured by the parameter $\theta$. Throughout the paper, we use $\theta_0$ to denote the true parameter.

For example, a widely used specification takes the following form:

align[align omitted — 152 chars of source]

where $\xi_j$ represents a scalar-valued unobserved characteristic of alternative $j$, such as product quality or brand loyalty. The term $x_i^\ast \kappa_j$ allows preferences for product $j$ to vary with the unobservable characteristic $x_i^\ast$. In this example, $\theta=(\gamma, \beta, \boldsymbol{\xi},\boldsymbol{\kappa})$, where $\boldsymbol{\xi}=(\xi_1, \cdots, \xi_J)$ and $\boldsymbol{\kappa}=(\kappa_1, \cdots, \kappa_J)$.

Two-Step Estimation Strategy

\paragraph{Step 1: Estimating Selected Outcome Distribution}

The key inputs for our contraction-mapping results are the selected outcome distributions $\tilde{G}$ conditional on $(x, x^\ast)$. When all relevant covariates are observed, so that no unobserved component $x^\ast$ is present, $\tilde{G}$ conditional on $x$ can be easily estimated nonparametrically from the data, for example using kernel methods. We therefore do not elaborate on this case. The more challenging setting arises when an unobserved covariate $x^\ast$ is present. In this case, $\tilde{G}$ conditional on $(x, x^\ast)$ cannot be directly estimated from the observed data, and additional information about the unobserved covariates is required in order to recover this distribution.

We follow the instrumental variable approach of hu2008identification to estimate the selected outcome distribution conditional on the unobservable $x^\ast$ in the first step. We assume that the variables $\omega_i=\{x_i,y_i,p_i,z_{1i},z_{2i}\}$ are observed in an i.i.d. sample and take values in a finite support.\footnote{The finite support assumption is not essential for identifying and estimating the selected outcome distributions in the first step. hu2008instrumental extends the results in hu2008identification to settings with continuously distributed variables, so similar identification argument and estimation procedure remain valid without discreteness. This assumption is adopted here primarily to simplify the asymptotic normality results, which we discuss further in Section (ref).} The variables $z_1$ and $z_2$ serve as instrumental variables and are required to satisfy the following condition:

align[align omitted — 197 chars of source]

where $h(\cdot)$ represents probability mass functions. Equation ((ref)) shows that the joint distribution of $(p,z_1)$ conditional on $(z_2,x,y)$ can be expressed as a mixture over the latent variable $x^\ast$. This condition requires, first, that the two instrumental variables are informative about the latent variable $x^\ast$, and second, that once we condition on $x^\ast$, the price and the instruments are independent.\footnote{Alternatively, suppose we have three instruments $(z_1, z_2, z_3)$ such that,

align*[align* omitted — 191 chars of source]

Under this condition, the instruments are allowed to depend arbitrarily on the price, while only requiring independence across instruments conditional on the latent variable. } In practice, finding such instruments is often feasible. In insurance pricing, for example, the latent variable $x^\ast$ may represent a consumer’s unobserved risk type, which influences premiums. Realized claims can serve as proxy variables for this latent type. In labor applications, the latent variable might correspond to a worker’s unobserved productivity, which affects wages. Measures such as work-performance evaluations or test scores can provide useful proxies in these settings.

Theorem 1 in hu2008identification shows that, under additional rank and ordering assumptions, the unknown probability mass functions on the right-hand side of Equation ((ref)), $h=(h_{p|x^*,x,y}, h_{z_1|x^*,x,y}, h_{x^*|z_2,x,y})\in H$, are nonparametrically identified. We do not restate these additional assumptions here and instead refer readers to hu2008identification for the technical details.

Given Equation ((ref)), a maximum likelihood estimator of $h$ can be obtained in a straightforward manner. We denote this estimator by $\hat{h}=(\hat h_{p|x^*,x,y},\hat{h}_{z_1|x^*,x,y}, \hat{h}_{x^*|z_2,x,y})$. The term $\hat h_{p|x^*,x,y}$ represents the estimate of the selected price distribution conditional on $(x,x^\ast)$, which corresponds to $\tilde G(x,x^*)$. The only distinction is that $h_{p|x^*,x,y}$ is a probability mass function, whereas $\tilde G(x,x^*)$ is its associated cumulative mass function. We do not distinguish between these two objects in what follows. Finally, by taking the expectation of $\hat{h}_{x^*|z_2, x, y}$ with respect to the distribution of $z_2$, we obtain an estimate of the distribution of the latent variable $x^\ast$ conditional on $(x,y)$, which we denote by $\hat{h}_{x^*|x,y}$. From this object, we can easily derive the choice probability for each alternative $j \in \mathcal{J}$ conditional on $(x, x^\ast)$, which are essential for identifying parameters of the selection function, as discussed in Section (ref).

\paragraph{Step 2: Estimating Selection Function Parameters and Offered Price Distributions}

Given the first-step estimates $\hat h_{p|x^*,x,y}$ and $\hat h_{x^*|x,y}$, we propose a semiparametric maximum likelihood estimator for parameter $\theta$ in the selection function:

align[align omitted — 99 chars of source]

where

align[align omitted — 360 chars of source]

Equation ((ref)) derives the probability that alternative $j$ is chosen conditional on $(x, x^\ast)$ for any utility parameters $\theta$ and any selected outcome distributions $\tilde{G}$. This probability is obtained by integrating the selection function $f_j(\boldsymbol{p}; x, x^\ast, \theta)$ with respect to the distribution of offered prices for all competing alternatives. Recall that $F$ denotes the mapping from the offered price distribution $G$ to the selected outcome distribution $\tilde{G}$ as defined in Equations (ref) and (ref). The inverse mapping $F^{-1}$ in Equation ((ref)) maps $\tilde{G}$ back to $G$.

To recover the offered price distribution, we rely on the contraction-mapping results in Theorem (ref), which guarantees that we can replicate $F^{-1}$ by iterating the operator $T$ until convergence. Note that the operator $T$ depends on two components: (1) the parameters of the selection function, $ \theta$, and (2) the selected outcome distributions $\tilde{G}$, for which a first-step estimate $\hat h_{p|x^*, x, y}$ is obtained. We use $\hat T_{\theta,\hat{h}}$ to denote the operator constructed given $\theta$ and $\hat{h}$. Let $\hat T^{\infty}_{\theta,\hat{h}} \Psi$ denote the limit of the iterates of $\hat T_{\theta,\hat{h}}$ starting from an initial distribution $\Psi$.\footnote{In practice, the algorithm used to solve the fixed point is terminated after a finite number of iterations. We show that the resulting approximation error is asymptotically negligible, provided that the number of iterations grows fast enough compared to the logarithm of the sample size. Further details are provided at the end of Section (ref).}

Our second-step estimation follows a nested fixed-point algorithm. In the inner loop, for any candidate value of the parameter $\theta$ in the selection function, we obtain the fixed point of the operator $T$ as $\hat T^{\infty}_{\theta,\hat{h}} \Psi$. Given the resulting offered price distribution, we compute agents’ choice probabilities using Equation (ref) and then construct the sample analogue of the likelihood function in Equation ((ref)). In the outer loop, we then search over $\theta$ to maximize this likelihood.

Once $\hat{\theta}$ is obtained, a plug-in estimator of the offered price distribution $G$ can be constructed by $$\hat{G}=\hat T^{\infty}_{\hat{\theta},\hat{h}} \Psi.$$ This step essentially repeats the inner-loop procedure, except that we replace $\theta$ with its estimate $\hat{\theta}$.

Consistency and Asymptotic Normality

We now discuss the asymptotic properties of our proposed estimators $\hat{\theta}$ and $\hat{G}$. When constructing the model-implied choice probabilities in Equation (ref), the inverse mapping $F^{-1}$ appears, which maps the selected price distribution $\tilde G$ back to the offered price distribution $G$. We therefore begin by analyzing the properties of this inverse mapping $F^{-1}$.

propSuppose $\rho<1$. The mapping $F$ is a homeomorphism. Moreover, both $F$ and $F^{-1}$ are Lipschitz continuous, with Lipschitz constants $1+\rho$ and $\frac{1}{1-\rho}$, respectively.
proofSee Appendix (ref).

Proposition (ref) has three important implications. First, because $F$ is a homeomorphism, its inverse $F^{-1}$ is well-defined, and we have $G=F^{-1}(\tilde G)$. Second, the continuity of $F^{-1}$ implies that if a consistent estimator $\tilde G_n$ of the selected outcome distribution is used in place of $\tilde{G}$, then $$F^{-1}(\tilde G_n)\overset{p}{\to} F^{-1}(\tilde G) = G \quad\text{as}\quad \tilde G_n\overset{p}{\to} \tilde G.$$ Finally, since $F^{-1}$ is Lipschitz continuous, $F^{-1}(\tilde G_n)$ converges to $G$ at the same rate as $\tilde G_n$ converges to $\tilde G$.

We now turn to the consistency and asymptotic normality of our estimators. To establish consistency, we rely on the fundamental consistency theorem for extremum estimators (Theorem 2.1 in newey1994large). We construct the true population objective function as follows: $$Q_0(\theta)=\mathbb E_{x,x^*}\sum_{j=1}^J\bigg( \int_{\boldsymbol{p}} f_j(\boldsymbol{p};x,x^*, \theta_0)dG(x,x^*)(\boldsymbol{p})\bigg) \ln\big(Prob_j( x,x^*, \theta,\tilde G)\big),$$ where $\int_{\boldsymbol{p}} f_j(\boldsymbol{p};x,x^*, \theta_0)dG(x,x^*)(\boldsymbol{p})$ represents the true probability of selecting alternative $j$ conditional on $x$ and $x^*$.

We maintain the previous assumptions on the selection function, namely that $f_j\in \mathscr{C}^1\colon \prod_j [\underline p_j, \overline p_j]\to (0,1]$. The following additional technical conditions are required to establish the consistency of $\hat{\theta}$.

asmp(i) The space $\Theta$ of parameter $\theta$ is compact; (ii) for each $x,x^*$, the selection function $f(\boldsymbol{p};x,x^*, \theta)$ is jointly continuous in $\theta$ and $\boldsymbol{p}$; (iii) the condition in Theorem (ref) holds for all $\theta\in \Theta$, that is, $\sup_{\theta\in\Theta} \rho(\theta)\leq \bar \rho<1$ for some $\bar \rho$.
asmp[Identification] There does not exist $\theta'\in \Theta$, $\theta'\neq \theta_0$, offered price distributions $G,\,G'\in \big(\prod_j \Delta([\underline p_j, \overline p_j])\big)^{X\times X^*}$ such that for all $j\in \mathcal J$ and $x,x^*$, $$F(G(x,x^*);\theta_0 ,x,x^*)=F(G'(x,x^*);\theta',x,x^* ),$$ $$\int_{\boldsymbol{p}} f_j( \boldsymbol{p}; x,x^*, \theta_0)dG(x,x^*)( \boldsymbol{p} )=\int_{\boldsymbol{p}} f_j( \boldsymbol{p}; x,x^*, \theta')dG'(x,x^*)( \boldsymbol{p} ).$$

Assumption (ref) (i) and (ii) are standard regularity conditions. Assumption (ref) (iii) ensures that for all $\theta \in \Theta$, the operator $T$ is a contraction. Assumption (ref) imposes the identification condition, which requires that there does not exist another parameter that can yield the same selected price distribution and choice probabilities.

The identification condition merits additional discussion. The unknown objects in our model are the parameter vector $\theta$ in the selection function $f$ and the offered price distribution $G$. A key insight from our contraction mapping result (Theorem (ref)) is that, for any given selection function $f$, the operator $T$ admits a unique fixed point, and this fixed point corresponds to the offered price distribution associated with $f$. In other words, given $f$, the offered price distribution $G$ is fully nonparametrically identified from the accepted price distribution. This is not an assumption; it is an implication of the model’s structure. As a result, identification of the full model reduces to identification of the parameter vector $\theta$ in the selection function.

Assumption (ref) essentially requires that variation in the observed choice probabilities conditional on $(x,x^\ast)$—which we have already identified in the first-step estimation—is sufficient to uniquely pin down the selection-function parameters. For example, under the commonly used demand specification in Equation (ref), the unknown parameters include the price sensitivity parameter $\gamma$, the coefficients $(\beta, \kappa_j)$ on the covariates $(x, x^\ast)$, and the unobserved product characteristics $\xi_j$ (with one of them normalized to zero without loss). The number of unknowns is $dim(x_{ij})+2J$, whereas the number of moments (i.e., conditional choice probabilities) available for identification is $|X||X^\ast|(J-1)$. As the dimensions of $x$ and $x^\ast$ increase, the variation in choice probabilities expands, generating an overidentified system for the utility parameters.

Moreover, if additional instrumental variables are available, such as exogenous cost shocks that shift the offered price distribution, these provide extra moment conditions for identifying the price sensitivity parameter as in the classical demand estimation literature. In the paper, we provide a high-level version of the identification condition for simplicity, but our framework can readily incorporate any additional instrumental variables when available. These extra moments can be included in the outer loop of the Step 2 estimation procedure described in Section (ref). We summarize the consistency result in the following theorem.

thm[Consistency] Under Assumptions (ref) and (ref), $\hat\theta\overset{p}{\to }\theta_0$, $\hat T^{\infty}_{\hat{\theta},\hat{h}} \Psi \overset{p}{\to } G$.
proofSee Appendix (ref).

Next, we show that the estimator defined in Equation (ref) is asymptotically normal. Let $$\mathfrak g(\omega ;\theta, h)=\nabla_\theta \bigg(\sum_{x^*} h_{x^*|x,y}(x^*|x, y )\ln Prob_{y }( x ,x^* , \theta, h_{p|x^*, x, y})\bigg ),$$ where $\nabla_\theta$ denotes the gradient operator with respect to $\theta$. The estimator $\hat\theta$ solves the first-order condition $$\frac{1}{n}\sum_{i=1}^n \mathfrak g(\omega_i;\theta, \hat h)=0.$$ Moreover, we define $$\mathfrak m(\omega_i, h)=\nabla_{h}\ln\bigg( \sum_{x^*} h_{p|x^*, x, y}(p_i|x^*, x_i, y_i) h_{z_1|x^*, x, y}(z_{1i}|x^*, x_i, y_i) h_{x^*|z_2, x, y}(x^*|z_{2i}, x_i, y_i) \bigg). $$

We stack $\mathfrak g$ and $\mathfrak m$ to form $$\tilde{\mathfrak g}(\omega, \theta, h)=[\mathfrak g(\omega, \theta, h)', \mathfrak m(\omega, {h})']',$$ then the estimators in the first two steps can be viewed as a GMM estimator. We impose the following standard regularity conditions.

asmp(i) $(\theta_0,h_0)$ is in the interior of $\Theta\times H$. (ii) $f$ is twice continuously differentiable in $\theta$. (iii) $\mathbb E \nabla_{\theta,h} \tilde {\mathfrak g}(\omega;\theta_0,h_0)$ is nonsingular.
thm[Asymptotic Normality] Suppose that Assumption (ref), (ref), and (ref) hold. Then $\hat\theta$, $\hat h$, $\hat T^{\infty}_{\hat{\theta},\hat{h}} \Psi$ are $\sqrt{n}$-asymptotically normal and $\sqrt{n} (\hat\theta-\theta_0)\overset{d}{\to} \mathcal N(0,V)$.\footnote{See the analytical form of $V$ in the proof of Theorem (ref).}
proofSee Appendix (ref).

So far, the asymptotic results have been stated under the assumption that the operator is iterated infinitely many times. In practice, however, the iteration used to obtain the offered price distribution is stopped after a finite number of steps. The resulting approximation error is asymptotically negligible as long as the number of iterations grows fast enough relative to the logarithm of the sample size. Formally, let $m(n)$ denote the number of iterations given the sample size $n$. Consistency of our estimator (Theorem (ref)) can be achieved as long as $\lim_{n\to+\infty} m(n)\to\infty$. Asymptotic normality (Theorem (ref)) continues to hold if in addition, $\liminf_{n\to+\infty}\frac{m(n)}{\ln n}>\frac{1}{2}(\ln(1/\bar\rho))^{-1}$.

Finally, we discuss the finite support assumption imposed on the outcome $p_i$. This assumption is not essential for the consistency result in Theorem (ref). As long as the estimator of $h$ is consistent, our proposed estimator remains consistent even when $p_i$ is continuous. Assuming that $p_i$ has finite support mainly keeps the proof of asymptotic normality in Theorem (ref) tractable. If $p_i$ is instead continuous, establishing asymptotic normality for a semiparametric two-step estimator typically requires a first-order expansion around the nonparametric estimator (see Theorem 8.1 in newey1994large). In our setting, this would require expanding the function $\mathfrak g$ around $\hat h_{p|x^*,x,y}$. A standard argument would apply if $\hat h_{p|x^*, x, y}$ entered Equation (ref) directly. However, in our case it enters only through $F^{-1}$, for which no analytic expression is available. As a result, working with the infinite-dimensional distribution $\tilde G$ is extremely challenging.

In practice, when the selected outcome distribution is estimated nonparametrically, even if $p_i$ is conceptually continuous, the estimator necessarily evaluates its CDF on a finite grid of points. For this reason, this assumption does not impose a substantive restriction in applied work.

Monte Carlo Simulations

To examine how our estimators for $\theta$ and the offered price distribution $G$ perform in finite samples, we conduct a Monte Carlo simulation experiment with $J=2$. The utility of individual $i$ from the two alternatives are specified as follows:

align*[align* omitted — 148 chars of source]

where $p_{ij}$ and $\xi_j$ are, respectively, the offered price and unobserved heterogeneity for alternative $j$; $x_{i1} \in \{0, 1\}$ is a binary observable with $Pr(x_{i1}=1)=0.5$ that shifts individual $i$'s choice probabilities; $x_i^\ast \in \{-1, 1\}$ is a binary unobservable with $Pr(x_{i}^\ast=1)=0.5$; and $\varepsilon_i \sim N(0,1)$ is the error term. Throughout the main simulation exercises, we set the utility parameters as follows: $\gamma=1$, $\xi_1=0$, $\xi_2=0.5$, $\beta=0.5$ and $\kappa=0.1$.\footnote{Note that $x_{i1}$ is included in the utility specification to facilitate the implementation of the two-step method for sample selection. Our method does not require this type of excluded variable. We therefore also consider a simulation exercise in which $x_{i1}$ is omitted, that is, $\beta=0$. The results are reported in Tables (ref) and (ref) in Appendix (ref).} Let $y_i \in \{1, 2\}$ denote the choice of individual $i$.

We consider four data generating processes for the offered prices. Let $x_{i2}$ denote the observable characteristic of individual $i$ that enters the pricing equation. We assume that $x_{i2}$ takes values in $\{0, 0.25, 0.5, 0.75, 1\}$ with equal probability.

enumerate[DGP 1:] • $\log(p_{ij}) = \delta_{0j}+\delta_{1j}x_{i2}+\delta_{2j}x_i^\ast+\eta_{ij}$, where $\eta_{ij} \sim N(0, \sigma_j^2)$. For alternative 1, we set $\delta_{01}=0.2, \delta_{11}=0.5, \delta_{21}=0.1, \sigma_1=0.1$. For alternative 2, we set $\delta_{02}=0.1, \delta_{12}=1,\delta_{22}=0.1, \sigma_2=0.2$. • $\log(p_{ij}) = \delta_{0j}+\delta_{1j}x_{i2}^2+\delta_{2j}x_i^\ast+\eta_{ij}$, where $\eta_{ij} \sim N(0, \sigma_j^2)$. For alternative 1, we set $\delta_{01}=0.2, \delta_{11}=0.5, \delta_{21}=0.1, \sigma_1=0.1$. For alternative 2, we set $\delta_{02}=0.1, \delta_{12}=1,\delta_{22}=0.1, \sigma_2=0.2$. • $\log(p_{ij}) = \exp\left( (\delta_{0j}+\delta_{1j}x_{i2})(\delta_{2j}x_i^\ast+\eta_{ij})\right)$, where $\eta_{ij} \sim N(1, \sigma_j^2)$. For alternative 1, we set $\delta_{01}=0.2, \delta_{11}=0.3, \delta_{21}=0.1, \sigma_1=0.1$. For alternative 2, we set $\delta_{02}=0.1, \delta_{12}=0.5,\delta_{22}=0.1, \sigma_2=0.2$. • $\log(p_{ij}) = (\delta_{0j}+\delta_{1j}x_{i2}^2)(\delta_{2j}x_i^\ast+\eta_{ij})^{-1}$, where $\eta_{ij} \sim N(-2, \sigma_j^2)$. For alternative 1, we set $\delta_{01}=0.2, \delta_{11}=0.1, \delta_{21}=0.1, \sigma_1=0.1$. For alternative 2, we set $\delta_{02}=0.1, \delta_{12}=0.3,\delta_{22}=0.1, \sigma_2=0.2$.

Across all data generating processes, the unobserved characteristic $x_i^\ast$ enters the pricing equations for both alternatives, which induces correlation in prices conditional on observables. In addition, $x_i^\ast$ also enters the utility specification, allowing the unobserved type to jointly affect the prices individuals face and their preferences over alternatives. DGP 1 specifies an additively separable linear pricing equation, which is commonly assumed in empirical applications. DGP 2 introduces a nonlinear term. DGPs 3 and 4 consider scenarios where the pricing function takes a nonseparable form.\footnote{Although all the offer price distributions admit unbounded support, in simulation we shall assume that the realized price range coincides with the true price range. Given a large sample size, the realized price range supports almost all the probability mass of the offered price distribution. Later we show that the estimation of the offered price distribution performs well.}

For each DGP, we simulate offered prices and individual choices. To implement our estimator, we require an instrument $z_i$ to recover the selected price distribution conditional on $(x_{i1},x_{i2},x_i^\ast)$ in the first step, since $x_i^\ast$ is unobserved. We construct such an instrument by assuming $z_i \sim Poisson(x_i^\ast)$ when $x_{i}^\ast=1$, and $z_i=0$ otherwise. This choice is motivated by settings where $x_i^\ast$ can be interpreted as an individual’s unobserved risk type, and such risk types may be reflected in the ex post realization of accidents, which are often modeled using a Poisson distribution. Because we impose a parametric relationship between the instrument and the unobservable, only one instrument is needed.

We assume that the econometrician observes $(y_i, x_{i1}, x_{i2}, p_i, z_i)$, where $p_i$ denotes the price of the chosen alternative. Using these data, we apply the procedure described in Section (ref) to estimate the parameters of the selection function, $\theta=(\gamma, \xi_2, \beta, \kappa)$ with $\xi_1$ normalized to 0, along with the offered price distribution for each alternative.\footnote{We estimate the cumulative distribution function of prices at 300 grid points.} For comparison, we first implement the classic Heckman parametric two-step method, assuming that the pricing equations are linearly separable and that the error terms in the selection and pricing equations follow a bivariate normal distribution. We also compare our estimator with the quantile selection model of arellano2017quantile. To implement their approach, we follow the standard practice of assuming that the quantile functions are linear in $x_{i2}$ and that the dependence structure is governed by a Gaussian copula.\footnote{Although arellano2017quantile discuss identification under more general settings, their empirical implementation focuses on cases in which the copula depends on a low-dimensional vector of parameters, which is the specification we adopt here.} For each design, we run 500 simulations with sample sizes of 2,000 and 5,000 observations.

Table (ref) reports the Monte Carlo biases, standard deviations, and root mean squared errors for the estimates of $\theta$ obtained using our method. Overall, the estimator performs well in finite samples across all DGPs, including those with nonseparable pricing equations. The biases are small, and the root mean squared errors remain modest for all parameters in the selection function. The standard deviation decreases as the sample size increases in all simulation designs.

table[table omitted — 2,085 chars of source]

For the cumulative distribution functions of $log(price)$, Tables (ref) and (ref) report the integrated squared biases and integrated mean squared errors for our proposed estimator, the Heckman two-step estimator, and the copula-based sample-selection correction estimator for quantile regression, separately for the two alternatives. Each row of the tables corresponds to the price distribution conditional on a specific value of $x_{i2}$. We also plot the true CDFs for alternatives 1 and 2 alongside the estimates produced by these models conditional on $x_{i2}=0.25$ and $x_{i2}=0.75$ in Figures (ref) and (ref), respectively. To save space, we report the CDF results only for the sample size of 2,000 observations, and we omit the figures for other values of the observable covariates.

Our method allows for nonparametric estimation of the offered price distributions, whereas the alternative approaches impose parametric restrictions on either the conditional mean or quantiles of the pricing distributions, or on the dependence structure through the copula. Tables (ref) and (ref) show that our estimator achieves very low integrated squared bias and integrated mean squared error for the CDFs of $log(price)$ across all simulation designs and for all values of $x_{i2}$. In contrast, while the classic Heckman two-step method and the quantile selection model perform well in DGP 1, their biases and mean squared errors increase substantially as the pricing equation becomes more complex in DGPs 2–4. These results are expected, since the parametric assumptions underlying these methods, such as linear conditional mean or quantile functions and a Gaussian copula, are severely violated in these designs.

Figures (ref) and (ref) provide a visual illustration of these results. We can see that across all simulation designs, the estimated CDFs of $log(price)$ for both alternatives produced by our functional contraction approach closely track the true CDFs, as indicated by the black curves with “+” markers and the red solid curve in Figures (ref) and (ref). By comparison, the biases of the Heckman two-step method (blue dashed curves) and the quantile selection model (purple dash–dotted curves) can be substantial, particularly in DGPs 3 and 4. The direction and magnitude of these biases also vary with the values of the observable covariates.

table[table omitted — 3,478 chars of source]
table[table omitted — 3,478 chars of source]
figure[figure omitted — 557 chars of source]
figure[figure omitted — 557 chars of source]

Another key advantage of our approach is that it does not require an instrument to exogenously shift the selection probability. It is well known in the literature that the two-step method is nearly unidentified when the same regressors are used in both the selection function and the outcome equation even under strong parametric restrictions. This occurs because the inverse Mills ratio is approximately linear over a wide range of its argument. In practice, it is also difficult to find variables that affect selection but can be excluded from the outcome equation.

In contrast, our approach does not require such an excluded variable. To illustrate this, we conduct a set of Monte Carlo simulations where the excluded variable $x_{i1}$ is removed from the indirect utility, using the same four DGPs for $log(price)$. The results for this specification are reported in Tables (ref)--(ref) in Appendix (ref). As shown, our estimator performs well in finite samples, even without an additional excluded variable to exogenously shift the selection probability. Our estimator consistently shows low biases across different DGPs and exhibits a decreasing standard deviation as the sample size increases.

Our method requires the econometrician to correctly specify the functional form of the selection function. To evaluate how the estimator performs under misspecification, we conduct a series of Monte Carlo simulations in which the econometrician assumes that $\varepsilon$ follows a logistic distribution, while in truth it is generated from a normal distribution. In Tables (ref)--(ref) in Appendix (ref), we report the estimation results for the utility parameters and CDFs of $log(price)$ under this misspecification. For the utility parameters, we rescale the estimates by the scale parameter of the logit model to make them comparable to those in the original probit specification. After this adjustment, the biases are small. The estimator for the offered price distributions also performs well: the integrated squared biases and mean squared errors of the CDFs remain close to those in Tables (ref) and (ref). These results suggest that the estimator of the offered price distributions is robust to misspecification of the selection function, an appealing feature in practice, particularly when the econometrician has limited prior information about the correct functional form.

Finally, we briefly discuss how the functional contraction performs computationally in practice. We compute $\rho^\ast$ for all four simulation designs and find that it is below 1 in every case. The average numbers of iterations needed to reach convergence (with a tolerance of $10^{-5}$) are 3.8, 3.8, 3.4, and 2.1 for DGPs 1 through 4, respectively (averaged over 500 replications). These results indicate that the proposed estimator is computationally efficient, converges rapidly, and remains stable across a range of data generating processes, making it well suited for applied work.

Discussion of Empirical Applications

Our estimator introduced in Section (ref) is broadly applicable to a wide range of empirical settings. It effectively addresses the challenge of selection bias that arises when only the outcomes of chosen alternatives are observed. The method has three features that are particularly important for empirical applications. First, it imposes no parametric or separability restrictions on the potential outcome distributions and allows them to vary flexibly across alternatives. Second, the framework accommodates unobservable characteristics in both the outcome distributions and the selection model, capturing selection on unobservables. Third, the selection function can incorporate alternative-specific unobserved heterogeneity and does not require an excluded variable, which is desirable in many empirical settings.

An important empirical application that illustrates these advantages is consumer demand estimation in markets where only transaction prices are observed. In classic differentiated product demand estimation pioneered by berry1994estimating and berry1995automobile, the price of a product is often assumed to be uniform across all consumers (e.g., the list price of a vehicle). But this assumption does not hold in contexts involving price discrimination or personalized pricing d2019automobile, sagl2023dispersion, buchholz2020value, dube2023personalized, discount negotiation goldberg1996dealer, allen2014price, or risk-based pricing crawford2018asymmetric, cosconati2024competing. In these contexts, researchers can relatively easily gather data on the transaction prices consumers pay, but it is challenging to gain access to competing prices offered to consumers.

In a companion paper with coauthors cosconati2024competing, we apply our method to estimate demand and insurance companies' information technology in the auto insurance market, where only the transaction prices of selected insurance plans are observed. In this market, insurance companies employ risk-based pricing. For each consumer, an insurance company generates a noisy estimate of their risk type and prices accordingly. Our goal is to quantify the heterogeneity in insurers' information technology, as measured by the dispersion of their risk estimates. Since the shape of the offered price distribution reflects the distribution of risk estimates, allowing for flexible estimation of the offered price distribution is crucial.

In this application, we assume that the offered prices across different firms are independent conditional on observable characteristics and the consumer’s true unobserved risk type. At the same time, the consumer’s risk type may also influence their preferences over insurance products. For example, higher-risk consumers may prefer insurers with higher service quality. We therefore allow the true risk type to affect both the pricing distributions and the utility parameters. Our data include realized claim records for each consumer over multiple years, and we use these records as instruments for the latent risk type in the first-step estimation.

We nonparametrically estimate each insurance company’s offered price distribution using our functional contraction approach. In Figure (ref), we plot the CDFs of offered price for several firms based on estimates in cosconati2024competing. The distributions differ substantially across firms, indicating significant heterogeneity in their pricing strategies. Building on this result, we estimate each firm's information precision parameter using supply-side model restrictions. These estimates provide important insights for analyzing competition under heterogeneous information structures in this market.

figure[figure omitted — 340 chars of source]

From a practical point of view, our iterative procedure to numerically solve for the offered price distributions given demand parameters is easy to implement and performs well in practice. In our empirical application using data from 11 insurers, the iterative algorithm converges very quickly, typically requiring only 6--7 iterations.

The usefulness of our method is not limited to consumer demand. It can also be applied to auction models and Roy models, where similar selection issues arise. For example, in multi-attribute auctions, our approach can be used to nonparametrically recover the full bid distribution and the auctioneer’s scoring weights when only the winning bids and the winner’s identity are observed, even in the presence of bidder asymmetry.\footnote{ Flexibly accommodating bidder asymmetries is a well-known challenge in auction models athey2007nonparametric. Bidder asymmetries may arise from factors such as distance to the contract location flambard2006asymmetry, information advantages hendricks1988empirical,de2009effect, varying risk attitudes campo2012risk, or strategic sophistication hortaccsu2019does. } Auctions in many settings have used the scoring rule that departs from the pure price-based criterion by accounting for quality differences,\footnote{See, for example, asker2008properties,lewis2011procurement,nakabayashi2013small,yoganarasimhan2016estimation,takahashi2018strategic,krasnokutskaya2020role,allen2024resolving.} and our framework can flexibly accommodate these multi-attribute scoring mechanisms with both observable and unobservable components. A similar application arises in Roy models, where our method can recover the distribution of potential wages when only realized wages in the chosen sector are observed. Our framework enables researchers to recover these distributions flexibly and without relying on excluded variables in the selection equation, which are frequently difficult to justify in applied settings. This capability provides a valuable tool for studying key questions in labor economics, such as occupational choice and wage inequality.

Conclusion

We introduce a novel method for estimating nonseparable selection models. We show that for a given selection function, potential outcome distributions are nonparametrically identified from the distribution of selected outcomes and can be recovered using a simple iterative algorithm. We achieve this by constructing an operator whose fixed point is the potential outcome distributions and proving that this operator is a functional contraction. Building on this theoretical result, we propose a two-step estimation strategy for both the selection function and potential outcome distributions. The consistency and asymptotic normality of the proposed estimators are established.

Our method has several important features. First, we allow the outcome equation to be fully nonparametric and nonseparable in error terms, and we recover the entire distribution of potential outcomes rather than focusing on specific moments or quantiles. In essence, we correct for sample selection bias by examining how the bias is systematically generated by the selection model. Second, our approach allows for fully heterogeneous effects of covariates on outcomes, which is a crucial feature for empirical analysis, as discussed in chernozhukov2025distribution. Another key advantage of our approach is that it does not rely on instruments to exogenously shift selection probabilities, which are often challenging to find in empirical settings, or on identification-at-infinity arguments. Finally, our approach also accommodates asymmetry in outcome distributions across alternatives and flexibly incorporates unobserved alternative-specific heterogeneity in the selection model.

We find that the proposed estimation strategy performs well in both simulations and real-world data applications; see our demand estimation using insurance market data in the companion paper cosconati2024competing. The approach is straightforward to implement and computationally efficient, making it highly appealing to empirical researchers. More broadly, the estimator can be applied in a wide range of settings in which only selected outcomes are observed, including consumer demand models with only transaction prices, auctions with incomplete bid data, and various selection models in labor economics. Our method is particularly valuable in applications where the entire distribution of outcomes is of interest.

singlespace