EconBase
← Back to paper

Identification and Estimation of Demand Models with Endogenous Product Entry and Exit

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

146,167 characters · 0 sections · 137 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and Estimation of Demand Models with Endogenous Product Entry and Exit

\thispagestyle{empty}

abstractFirms are more likely to introduce products in markets where they anticipate stronger demand. They also possess information that is unobserved to researchers. This creates endogenous selection bias in the estimation of demand parameters. With differentiated products, the entry decision violates the monotonicity conditions required for standard selection-correction methods to yield consistent demand estimates. Existing studies address this issue either by imposing strong assumptions about firms’ information on demand at the time of entry or by jointly estimating a full equilibrium model of demand, pricing, and entry. Both strategies make the estimation of demand heavily reliant on supply-side assumptions. We propose a new semiparametric estimation method that addresses these limitations. Our approach exploits the correlation across products in their market-entry decisions to identify entry probabilities conditional not only on observable characteristics but also on latent variables that capture unobserved interdependencies among firms’ entry choices. We refer to these probabilities as latent propensity scores. We show that the selection bias term in the demand equation is a convolution of these latent propensity scores and is therefore identifiable. Building on this result, we develop a two-step semiparametric estimator in the spirit of standard sample-selection correction methods. Applying our method to data from the airline industry, we find that conventional approaches to correcting for selection bias substantially underestimate price elasticities of demand. Keywords: Demand for differentiated product; Product entry; Selection bias; Airline markets. JEL codes: C14, C34, C35, C57, D22, L13, L93.

\setcounter{page}{1}

doublespacing\section{Introduction} Estimating demand systems for differentiated products typically relies on data spanning multiple geographic markets and time periods. In these settings, it is common for some products to be unavailable in certain markets or at particular times. Firms tend to introduce products in markets where they anticipate stronger demand, drawing on information about market conditions that is unobservable to researchers. As a result, the observed pattern of product availability is not random but reflects firms’ private expectations about demand. This endogenous selection into markets can generate substantial bias in the estimation of demand parameters in regression-based models. This issue is prevalent across various industries, including airlines (berry2006airline, berry2006airline; berry2010tracing, berry2010tracing; aguirregabiria2012dynamic, aguirregabiria2012dynamic), supermarket chains (smith2004supermarket, smith2004supermarket), radio stations (sweeting2013dynamic, sweeting2013dynamic), personal computers (eizenberg2014upstream, eizenberg2014upstream), and ice cream (draganska_2009, draganska_2009). The selection problem in this structural model of demand and product entry exhibits a distinctive feature that sets it apart from more conventional cases. Specifically, the demand unobservables are multi-dimensional and have a non-additive effect on firms' expected profits. This breaks a key monotonicity condition typically required for the selection equation. Without this condition, the selection propensity score---the probability of product entry given exogenous observables---cannot serve as a sufficient statistic to control for selection bias in the estimation of demand parameters (angrist_1997, angrist_1997). Furthermore, the model involves multiple equilibria in both the entry and pricing games. The possibility that different equilibria are selected across markets introduces additional non-monotonicity in the selection equation. As a result, standard identification results and two-step estimation methods that rely on the propensity score are not applicable in this context (e.g., ahn_powell_1993, ahn_powell_1993; das2003, das2003; aradillas2007pairwise, aradillas2007pairwise; newey_2009, newey_2009).\footnote{Importantly, instrumental variable approaches cannot address this form of selection bias. Consistent estimation typically requires control-function methods that explicitly model the selection process. See vella_1998, heckman_navarro_2004, wooldridge_2015.} The growing interest in estimating models of oligopoly competition that endogenize firms’ product-entry decisions across geographic markets has made the associated selection problem increasingly salient. The standard approach in this literature begins with the estimation of a demand system. However, in the absence of instrumental-variable or control-function methods to address selection bias, these studies typically impose strong assumptions about firms’ information sets at the time of entry. Such assumptions effectively rule out endogenous product selection based on unobserved demand shocks. Examples include aguirregabiria2012dynamic, fan_2013, sweeting2013dynamic, eizenberg2014upstream, and fan_yang_2020. Motivated by the importance of this issue, ciliberto2021market and li2022repositioning develop methods that jointly estimate the full structural model of demand, price competition, and product entry. Although these approaches fully account for selection bias in demand estimation, they make demand identification heavily dependent on supply-side assumptions---such as the nature of competition, the functional form of cost functions, and the distributional assumptions on unobservables. The main contribution of this paper is to establish new, more general conditions for the sequential (two-step) identification of demand parameters when product entry is endogenous. Our approach leverages the cross-product correlation in firms’ market-entry decisions to recover entry probabilities that are conditioned not only on observable characteristics but also on latent variables capturing unobserved interdependencies among firms’ choices. We refer to these as latent propensity scores. These probabilities are constructed by integrating over the distribution of unobservables that satisfy a monotonicity condition, while conditioning on those that violate it. Our identification result proceeds in two steps. First, we establish the nonparametric identification of the latent propensity scores. This step exploits a key feature of the model: the unobservables that violate monotonicity in a product’s entry decision are precisely the demand shocks of other products that could potentially enter the market. These unobservables generate the interdependence among firms’ entry decisions. Consequently, the joint distribution of entry decisions follows a mixture model structure, where the unobservables driving this interdependence act as the mixing variables. Second, we show that the selection-bias term in the demand equation can be expressed as a convolution of these latent propensity scores, and is therefore identifiable. Building on our constructive proof of identification, we propose a transparent and computationally simple two-step estimator that jointly corrects for endogenous product selection and price endogeneity in demand estimation. In the first step, we estimate each product’s latent propensity score using a semiparametric mixture model that captures unobserved interdependencies in firms’ entry decisions. In the second step, we recover the demand parameters through a control-function Generalized Method of Moments (GMM) procedure that accounts for both endogenous product availability and price endogeneity. This approach yields consistent estimates under minimal assumptions about firms’ information, the structure of competition, and the functional forms on the supply side. We illustrate the proposed method using data from the airline industry. The results demonstrate the importance of accounting for endogenous product entry when estimating demand parameters and highlight the limitations of conventional selection-correction approaches. Specifically, standard methods that impose strong informational or structural restrictions substantially underestimate price elasticities of demand. We also uncover significant selection bias in the estimation of marginal costs derived from Bertrand pricing equations. Moreover, our reduced-form estimation of entry probabilities---capturing rich correlations in firms’ entry decisions---provides economically meaningful insights. In particular, we find that models that ignore or restrict correlated unobservables in market-entry decisions tend to overstate the degree of market contestability, predicting a higher likelihood of new entry following mergers than what is supported by the data. Our paper contributes to the literature on sample selection bias in demand estimation when zeros arise from firms’ market entry decisions, including the seminal works of draganska_2009, conlon_mortimer_2013, ciliberto2021market, and li2022repositioning. These studies develop methods for estimating structural models that integrate differentiated-product demand systems à la berry1995automobile with market or product entry games following bresnahan1990entry, bresnahan1991entry and berry1992. Their approach involves joint estimation of demand, marginal cost, and entry-cost parameters using nested fixed-point algorithms. While powerful, these methods rely on strong parametric assumptions about functional forms and the distribution of unobservables. In contrast, our paper proposes a sequential estimation strategy that identifies the demand parameters without imposing specific assumptions about the supply side. This approach ensures robustness to a wide range of supply-side structures and greatly simplifies computation by avoiding the need to solve for equilibrium outcomes. Moreover, the same economic logic points to extensions to richer environments, including dynamic games of market entry and exit, though such extensions require adapting the first-step identification argument to the corresponding state variables.\footnote{Given the estimated demand parameters and unobservables from our method, one can subsequently recover marginal and entry costs under weaker parametric assumptions than those required in joint structural estimation. As in ciliberto2021market and li2022repositioning, our estimates can be used to conduct a variety of counterfactual experiments that account for the endogeneity of product entry and exit---an essential feature when simulating merger effects, as demonstrated by li2022repositioning. Section (ref) provides details on the implementation of these counterfactuals.} Our approach contributes to the growing literature on structural models of oligopoly competition that endogenize firms’ product entry decisions while explicitly incorporating demand systems for differentiated products. Contributions in this line of research include aguirregabiria2012dynamic, fan_2013, sweeting2013dynamic, eizenberg2014upstream, schaumans_verboven_2015, fan_yang_2020, bontemps2023price, caoui_steck_2026, and liu_luo_2025demand. These studies estimate structural parameters through a sequential approach that begins with the estimation of the demand system. To address potential selection bias from endogenous product entry, they impose restrictive assumptions about firms’ information sets---specifically, that firms lack information about unobserved components of demand when making entry decisions. These assumptions effectively rule out selection on unobservables and simplify identification, but at the cost of misspecification biases. In contrast, we relax this restriction, allowing firms to possess information about demand shocks at the time of entry. This not only addresses selection bias in demand estimation but also corrects the misspecification it induces in the entry game, where firms’ entry choices are endogenously correlated through shared information about demand fundamentals. Our estimation method contributes to the literature on semiparametric estimation of sample selection models (see, e.g., das2003,newey_2009,powell_2001,aradillas2007pairwise). We extend two-step propensity-score control function approaches to settings where the unobservables in the selection equation violate the standard monotonicity condition. Specifically, when the selection decision arises within a system of simultaneous selection equations, and the non-monotonic unobservables are those generating dependence across selection decisions, we show that it is still possible to identify a control function that corrects for selection bias. As far as we know, this is a novel result in this literature. Our approach can be applied to other sample selection problems that share this structural feature, such as labor market models with two-sided matching (choo_siow_2006, galichon_salanie_2022), models of joint household decisions (browning_2014), or peer effects models with endogenous network formation (graham_2017, depaula_2018). The remainder of this paper is structured as follows. Section (ref) introduces the class of models and underlying assumptions. Section (ref) describes the structure of the selection problem within this class of models. Section (ref) outlines our identification results and estimation method. In Section (ref), we illustrate our method with an empirical application to the US airline industry. Finally, Section (ref) provides a summary and concluding remarks. \section{Model } The framework follows the canonical model of demand and oligopoly competition in differentiated product markets in industrial organization. We outline it here to define notation and highlight the main assumptions. Proposition (ref) derives a simple property of this model that is fundamental to understanding the form of the selection bias examined in the paper. The demand system follows the Berry-Levinsohn-Pakes (BLP) framework (berry1995automobile, berry1995automobile). For the sake of notational simplicity, we focus on single-product firms. In Section (ref), we discuss the intuition for adapting the model and the selection-correction argument to the case of multi-product firms. There are $J$ firms indexed by $j\in\mathcal{J}=\{1,2,..., J\}$ and $T$ markets indexed by $t \in \{1,2,..., T\}$, where a market can be a geographic location, a period, or a combination of both. Consumers living in a market $t$ can buy only the products available in that market. Firms' market entry decisions, prices, and quantities are determined as an equilibrium of a two-stage game. In the first stage, firms maximize their expected profit by choosing whether or not to be active in the market. In the second stage, prices and quantities of the active firms are determined as a Nash-Bertrand equilibrium of a pricing game. This two-stage game is played separately across markets.\footnote{While this assumption is standard in the literature on empirical industrial organization, there are important exceptions, such as structural models of entry that allow potential entrants to internalize network externalities across markets, e.g., bontemps2023price,jia2008happens,aguirregabiria2012dynamic. However, these structural models of network formation do not consider the endogenous sample selection problem we study in this paper.} Demand and price competition are static. Appendix (ref) discusses how the same logic can be extended to dynamic games of firms' product entry and exit. \subsection{Demand} The indirect utility of household $h$ in market $t$ from buying product $j$ is: \begin{equation} U_{hjt} \equiv \delta(p_{jt},\boldsymbol{x}_{jt}) + v(p_{jt},\boldsymbol{x}_{jt},\boldsymbol{\upsilon}_{ht}) + \varepsilon_{hjt}, \end{equation} where $p_{jt}$ and $\boldsymbol{x}_{jt}$ are the price and other characteristics, respectively, of product $j$ in market $t$; $\delta_{jt}\equiv\delta(p_{jt},\boldsymbol{x}_{jt})$ is the average (indirect) utility of product $j$ in market $t$; and $v(p_{jt},\boldsymbol{x}_{jt},\boldsymbol{\upsilon}_{ht})+\varepsilon_{hjt}$ represents a household-specific deviation from the average utility. The term $v(p_{jt},\boldsymbol{x}_{jt},\boldsymbol{\upsilon}_{ht})$ depends on the vector of random coefficients $\boldsymbol{\upsilon}_{ht}$ with distribution $F_{\upsilon}(\cdot|\boldsymbol{\sigma})$, where $\boldsymbol{\sigma}$ is a vector of parameters. The term $\varepsilon_{hjt}$ is unobserved to the researcher and is i.i.d. over $(h,j,t)$ with type I extreme value distribution. Following the standard specification, the average utility of product $j$ is: \begin{equation} \delta_{jt} \equiv \alpha \; p_{jt} + \boldsymbol{x}_{jt}^{\prime} \; \boldsymbol{\beta} + \xi_{jt}, \end{equation} where $\alpha$ and $\boldsymbol{\beta}$ are parameters. Variable $\xi_{jt}$ captures the characteristics of product $j$ in market $t$ unobserved to the researcher. Throughout the paper, we normalize $\mathbb{E}(\xi_{jt}) = 0$. Similarly, the component of utility that depends on consumer-level random coefficients takes the multiplicative form: \begin{equation} v(p_{jt},\boldsymbol{x}_{jt}, \boldsymbol{\upsilon}_{ht}) \; = \; \left( p_{jt}, \, \boldsymbol{x}_{jt} \right)^{\prime} \; \boldsymbol{\Omega}_{\boldsymbol{\sigma}} \; \boldsymbol{\upsilon}_{ht} \end{equation} where $\boldsymbol{\Omega}_{\boldsymbol{\sigma}}$ is a $(K+1) \times (K+1)$ matrix that is a known, continuously differentiable function of the parameter vector $\boldsymbol{\sigma}$, and $\boldsymbol{\upsilon}_{ht}$ is a vector of random variables with a known distribution. The outside option is represented by $j=0$ and its indirect utility is normalized to $U_{h0t} = \varepsilon_{h0t}$. We denote by $\boldsymbol{\theta} \equiv (\alpha, \, \boldsymbol{\beta}^{\prime}, \, \boldsymbol{\sigma}^{\prime})^{\prime}$ the column vector of demand parameters. Let $a_{jt} \in \{0,1\}$ denote the indicator that product $j$ is offered in market $t$, and define $\boldsymbol{a}_{t} \equiv (a_{jt}: j \in \mathcal{J})$ as the vector collecting the offer indicators for all products in market $t$. The outside option $j=0$ is always offered in every market. Every household chooses the product that maximizes its utility. Let $s_{jt}$ be the market share of product $j$ in market $t$, i.e., the proportion of households choosing product $j$: \begin{equation} s_{jt} = d_{jt}(\boldsymbol{\delta}_{t}, \, \boldsymbol{a}_{t}, \, \boldsymbol{\sigma}) \equiv \int \frac{a_{jt} \; e^{ \delta_{jt} + \left[ p_{jt}, \, \boldsymbol{x}_{jt} \right]^{\prime} \; \boldsymbol{\Omega}_{\boldsymbol{\sigma}} \; \boldsymbol{\upsilon} }} {1+{ \sum \nolimits_{i=1}^{J}} a_{it} \; e^{ \delta_{it} + \left[ p_{it}, \, \boldsymbol{x}_{it} \right]^{\prime} \; \boldsymbol{\Omega}_{\boldsymbol{\sigma}} \; \boldsymbol{\upsilon} }} \; dF_{\upsilon}(\boldsymbol{\upsilon}). \end{equation} This system of $J$ equations represents the demand system in market $t$. We can represent this system in a vector form as $\boldsymbol{s}_{t} = \boldsymbol{d}_{t}(\boldsymbol{\delta}_{t}, \, \boldsymbol{a}_{t}, \, \boldsymbol{\sigma})$. For our analysis, it is convenient to define the subsystem of demand equations that includes the market shares, average utilities, and characteristics of the products that are offered. Define $\mathcal{J}_{t}^{\boldsymbol{a}} \equiv \{ j \in \mathcal{J} : a_{jt}=1 \}$, $\boldsymbol{s}_{t}^{\boldsymbol{a}} = (s_{jt} : j \in \mathcal{J}_{t}^{\boldsymbol{a}})$, and $\boldsymbol{\delta}_{t}^{\boldsymbol{a}} = (\delta_{jt} : j \in \mathcal{J}_{t}^{\boldsymbol{a}})$. We represent this system as: \begin{equation} \boldsymbol{s}_{t}^{\boldsymbol{a}} = \boldsymbol{d}^{\boldsymbol{a}}_{t} (\boldsymbol{\delta}_{t}^{\boldsymbol{a}}, \, \boldsymbol{\sigma}), \end{equation} Proposition (ref) establishes that, for any configuration of $\boldsymbol{a}_{t}$, the demand system (ref) satisfies the invertibility property with respect to $\boldsymbol{\delta}^{\boldsymbol{a}}_{t}$ (berry1994estimating, berry1994estimating). \begin{prop} Suppose that the outside option $j=0$ is always offered. Fix any value of the vector $\boldsymbol{a}_{t}\in\{0,1\}^{J}$ and define the set of feasible interior market shares for the offered products as \[ \mathcal{S}^{\boldsymbol{a}} \equiv \left\{ \boldsymbol{s}^{\boldsymbol{a}} \in (0,1)^{|\mathcal{J}_{t}^{\boldsymbol{a}}|} : \sum_{j\in\mathcal{J}_{t}^{\boldsymbol{a}}} s_{j} < 1 \right\}. \] Then the demand system in equation (ref) defines a one-to-one mapping from $\mathbb{R}^{|\mathcal{J}_{t}^{\boldsymbol{a}}|}$ into $\mathcal{S}^{\boldsymbol{a}}$. Therefore, for every $\boldsymbol{s}_{t}^{\boldsymbol{a}} \in \mathcal{S}^{\boldsymbol{a}}$, the inverse function $\boldsymbol{\delta}_{t}^{\boldsymbol{a}} = \left(\boldsymbol{d}_{t}^{\boldsymbol{a}}\right)^{-1} \left( \boldsymbol{s}_{t}^{\boldsymbol{a}}, \, \boldsymbol{\sigma} \right)$ exists and is unique. \qquad $\blacksquare$ \end{prop} \noindentProof of Proposition (ref): See Appendix (ref). For a product offered in market $t$, we have: \begin{equation} d_{jt}^{-1}\left( \boldsymbol{s}_{t}^{\boldsymbol{a}}, \, \boldsymbol{\sigma} \right) \; = \; \alpha \; p_{jt} + \boldsymbol{x}_{jt}^{\prime} \; \boldsymbol{\beta} + \xi_{jt} \ \ if and only if \ \ a_{jt}=1. \end{equation} Importantly, after applying Berry’s inversion, the selection condition for the existence of the regression equation for a product–market observation $(j,t)$ depends only on product $j$ (and the outside option $0$) being offered in market $t$, and not on which other products are offered in that market. Consequently, the selection bias in estimating the demand for product $j$ can be expressed in terms of the following conditional expectation: \begin{equation} \mathbb{E}\left( \xi_{jt} \mid a_{jt}=1 \right). \end{equation} Whenever $a_{jt}=1$, we write the inverse-demand term explicitly as $d_{jt}^{-1}\left(\boldsymbol{s}_{t}^{\boldsymbol{a}_{t}},\boldsymbol{\sigma}\right)$, where $\boldsymbol{s}_{t}^{\boldsymbol{a}_{t}}$ is the vector of observed market shares of the products offered in market $t$. This characterization of the selection term follows from working directly with the inverse demand system, as represented by equation ((ref)).\footnote{To appreciate the value of this property, consider instead the case of the Almost Ideal Demand System (AIDS) (deaton1980almost, deaton1980almost). In the AIDS, each value of the vector $\boldsymbol{a}_{t}$ implies a different set of regressors and slope parameters in the regression equation that relates the demand of product $j$ to the log-prices of the offered products. Therefore, in the AIDS model, the selection bias within the demand equation for product $j$ does not depend solely on the availability of that particular product but rather on the availability profile of all products within the system. In other words, the selection term cannot be represented in terms of $\mathbb{E}\left( \xi_{jt}\;|\;a_{jt}=1\right)$ but must instead be expressed in terms of $\mathbb{E}\left( \xi_{jt} \; | \; a_{jt}=1, \; \boldsymbol{a}_{-jt}=\boldsymbol{a}_{-j} \right) $. Consequently, in the AIDS model, we have a different selection term for each value of the vector $\boldsymbol{a}_{-j}$ representing the availability of products other than $j$. This structure makes the selection problem multi-dimensional and significantly complicates identification and estimation when the number of products $J$ is large.} As discussed in Appendix (ref), Proposition (ref) is unaffected in the case of multi-product firms, and so is the structure of the resulting selection term, which can still be represented as $\mathbb{E}\left( \xi_{jt}\;|\;a_{jt}=1\right)$ even if the firm owns other products. The following Example illustrates Proposition (ref) in the case of a nested logit model. \begin{example}[Nested logit model] The $J$ products are partitioned into $R+1$ mutually exclusive groups indexed by $r \in \{0, 1, ..., R\}$. We denote by $r_{j}$ the group to which product $j$ belongs. The outside good is the single element of group $r=0$. The indirect utility function is $U_{hjt} \equiv \delta_{jt} + v_{ht,r_{j}} + (1-\sigma) \; \varepsilon_{hjt}$, where variables $v$ and $\varepsilon$ are independently distributed, $\varepsilon$ and $v + (1-\sigma) \varepsilon$ are i.i.d. type I extreme value, $\sigma \in [0,1]$ is a parameter, and $v$ has a $C(\sigma)$ distribution as defined in cardell_1997. This model implies: \begin{equation} s_{jt} \, = \, d_{j}\left( \boldsymbol{\delta}_{t} ^{\boldsymbol{a}}, \sigma \right) \, = \, d_{j|r_{j}} \left( \boldsymbol{\delta}_{t} ^{\boldsymbol{a}}, \sigma \right) \, \cdot \, d_{r_{j}} \left( \boldsymbol{\delta}_{t} ^{\boldsymbol{a}}, \sigma \right) \end{equation} with: \begin{equation} d_{j|r_{j}} \left( \boldsymbol{\delta}_{t} ^{\boldsymbol{a}}, \sigma \right) \; = \; \frac{a_{jt} \; e^{\delta_{jt}/(1-\sigma)}} {{\textstyle\sum\nolimits_{i\in r_{j}}} a_{it} \; e^{\delta_{it}/(1-\sigma)}} \quad and \quad d_{r_{j}} \left( \boldsymbol{\delta}_{t} ^{\boldsymbol{a}}, \sigma \right) \; = \; \frac{\left[ {\sum\nolimits_{i\in r_{j}}} a_{it} \; e^{\delta_{it}/(1-\sigma)}\right] ^{1-\sigma}} {\sum\nolimits_{r=0}^{R} \left[ {\textstyle \sum \nolimits_{i\in r}} a_{it}\;e^{\delta_{it}/(1-\sigma)}\right] ^{1-\sigma}} \end{equation} If $a_{0t}=1$ and $a_{jt}=1$, this model implies that $s_{0t} > 0$ and $s_{jt}>0$, and the inverse function $d_{jt}^{-1}\left( \boldsymbol{s}_{t}^{\boldsymbol{a}}, \, \sigma \right)$ exists regardless of the value of $a_{it}$ for any product $i$ different from $j$. It is straightforward to show that this inverse function has the following form: \begin{equation} d_{jt}^{-1}\left( \boldsymbol{s}_{t}^{\boldsymbol{a}}, \, \sigma \right) \; = \; \ln\left( \frac{s_{jt}}{s_{0t}}\right) - \sigma\;\ln\left( \frac{s_{jt}}{ {\textstyle\sum\nolimits_{i\in r_{j}}} s_{it}}\right), \end{equation} and it implies the regression equation: \begin{equation} \ln\left( \frac{s_{jt}}{s_{0t}}\right) = \sigma\; \ln \left( \frac{s_{jt}}{{\textstyle\sum\nolimits_{i\in r_{j}}} s_{it}}\right) + \alpha \; p_{jt} + \boldsymbol{x}_{jt}^{\prime} \; \boldsymbol{\beta}+\xi_{jt}. \end{equation} Given $s_{0t}>0$, this regression equation holds whenever $a_{jt}=1$. \qquad $\blacksquare$ \end{example} \subsection{Price competition} Let $\Pi_{jt}$ be the profit of firm $j$ if active in market $t$. This profit equals revenues minus costs: \begin{equation} \Pi_{jt} \; = \; p_{jt} \; q_{jt} \; - \; c(q_{jt}, \, \boldsymbol{x}_{jt}, \, \omega_{jt}) \; - \; fc_{jt}, \end{equation} where $q_{jt}$ is the quantity sold (i.e., market share $s_{jt}$ times market size $H_{t}$), $c(q_{jt}, \, \boldsymbol{x}_{jt}, \, \omega_{jt})$ is the variable cost function, and $fc_{jt}$ is the fixed entry cost. Variable $\omega_{jt}$ is unobserved to the researcher. Given firms' entry decisions, the best response function in the Bertrand pricing game implies the following system of pricing equations: \begin{equation} p_{jt} \; = \; mc_{jt} \; - \; d_{jt}\left( \boldsymbol{\delta}_{t} ^{\boldsymbol{a}}, \, \boldsymbol{\sigma} \right) \; \left[ \dfrac{\partial d_{jt}\left( \boldsymbol{\delta}_{t} ^{\boldsymbol{a}}, \, \boldsymbol{\sigma} \right) } {\partial p_{jt}} \right]^{-1} for every j \in \mathcal{J}_{t}^{\boldsymbol{a}}, \end{equation} where $mc_{jt}$ is the marginal cost $\partial c_{jt}/\partial q_{jt}$. A solution to this system of equations is a Nash-Bertrand equilibrium. Let $\boldsymbol{x}_{t} \equiv (\boldsymbol{x}_{jt}:j\in\mathcal{J})$ denote the vector of exogenous variables observed by the researcher that affect demand or costs, with support $\mathcal{X}$ (each element of which may be continuous or discrete). The vectors $\boldsymbol{\xi}_{t}$ and $\boldsymbol{\omega}_{t}$ are defined analogously. Let $\boldsymbol{a}_{-jt}$ denote the vector of entry decisions for all firms other than $j$. We define the function \begin{equation} VP_{jt} \; = \; VP_{j} \left( \boldsymbol{a}_{-jt}, \, \boldsymbol{x}_{t}, \, \boldsymbol{\xi}_{t}, \, \boldsymbol{\omega}_{t} \right) \end{equation} as firm $j$’s \textit{indirect variable profit}, obtained by substituting into the expression $p_{jt}$ $q_{jt} -c(q_{jt};\boldsymbol{x}_{jt},\omega_{jt})$ the equilibrium values of prices and quantities from the Nash-Bertrand equilibrium given $(a_{jt}=1, \, \boldsymbol{a}_{-jt}, \, \boldsymbol{x}_{t}, \, \boldsymbol{\xi}_{t}, \, \boldsymbol{\omega}_{t})$.\footnote{The pricing game may admit multiple equilibria. We do not impose any restriction on equilibrium selection and allow each market to select its own equilibrium. For notational simplicity, we do not explicitly include an unobservable variable---say $\tau_{t}$---to index the equilibrium selected in the Bertrand game, although it can be interpreted as part of the broader vector of unobservables.} \subsection{Market entry game} This section introduces a model of product entry that encompasses a broad class of games studied in the literature. It nests complete-information frameworks such as those in ciliberto2009market and ciliberto2021market, as well as incomplete-information settings with common knowledge unobservables, as in grieco2014discrete and aguirregabiria2019identification. The model also allows for flexible information structures regarding firms’ knowledge of demand shocks at the time of entry, ranging from cases with full information to those with complete uncertainty, and including intermediate scenarios with imperfect signals. This general formulation ensures that the identification results developed in this paper apply to a wide spectrum of market entry environments. Firms’ entry decisions arise as the equilibrium outcome of this game. The payoff from remaining inactive is normalized to zero. Prior to making their entry decisions, firms may face uncertainty about their potential profits if active in the market. Their information about demand and cost fundamentals is therefore central to the entry process, as it shapes both individual incentives and the joint distribution of firms’ equilibrium entry decisions. Assumption (ref) summarizes our conditions on the information structure and the unobservables to the researcher.\footnote{The entry game has multiple equilibria. We adopt the same general approach as for the pricing game. We do not impose any restrictions on equilibrium selection, but for notational simplicity, we do not explicitly include an unobservable variable to index the selected equilibrium. It can be interpreted that vector $\boldsymbol{\kappa}_t$ includes the equilibrium selection index.} \begin{assumption} At the time firm $j$ makes its entry decision in market $t$, its information set consists of $(\boldsymbol{x}_{t}, \boldsymbol{\kappa}_{t}, \boldsymbol{\eta}_{jt})$. \begin{itemize} • The vector $\boldsymbol{x}_{t}$ of variables observable to the researcher is common knowledge among all firms. • The vector $\boldsymbol{\kappa}_{t}$ represents all information about demand and cost fundamentals $(\boldsymbol{\xi}_{t}, \boldsymbol{\omega}_{t})$ and fixed costs that is common knowledge among firms but unobserved by the researcher. In one possible scenario, $\boldsymbol{\kappa}_{t}$ may include the entire vector $(\boldsymbol{\xi}_{t}, \boldsymbol{\omega}_{t})$, implying that firms face no uncertainty about demand or variable costs at the time of entry. • The vector $\boldsymbol{\eta}_{jt}$ represents firm $j$'s private information about its entry cost. The vectors $\boldsymbol{\eta}_{jt}$ are assumed to be independently distributed across firms and independent of $(\boldsymbol{\xi}_{t}, \boldsymbol{\kappa}_{t}, \boldsymbol{x}_{t})$. As a special case, variable $\boldsymbol{\eta}_{jt}$ may have a degenerate distribution, in which case the entry game reduces to one of complete information. • All the unobservables for the researcher, $(\boldsymbol{\xi}_{t}, \boldsymbol{\omega}_{t}, \boldsymbol{\kappa}_{t}, \boldsymbol{\eta}_{jt})$, are assumed independent of the exogenous observables in $\boldsymbol{x}_{t}$. \qquad $\blacksquare$ \end{itemize} \end{assumption} Assumption (ref) provides a flexible specification of firms’ information sets at the time of entry. By allowing the common-knowledge component $\boldsymbol{\kappa}_{t}$ to range from a minimal set of market-level signals to the full vector of demand and cost fundamentals $(\boldsymbol{\xi}_{t}, \boldsymbol{\omega}_{t})$, the assumption nests environments with substantial uncertainty as well as those in which firms face effectively complete information about market conditions. Likewise, by introducing firm-specific private information $\boldsymbol{\eta}_{jt}$, the assumption accommodates a broad class of incomplete-information entry games, while also allowing the special case of complete information when $\boldsymbol{\eta}_{jt}$ is degenerate. Finally, part (d), which assumes independence between the unobserved shocks and the observable covariates $\boldsymbol{x}_{t}$, is entirely standard in empirical IO. This exogeneity condition underpins identification in both demand estimation and market entry models and aligns with conventional econometric practice in the literature. To simplify notation, in expressions for expected profits we sometimes write $\boldsymbol{\xi}_{t}$ as shorthand for the full vector of demand and cost fundamentals $(\boldsymbol{\xi}_{t}, \boldsymbol{\omega}_{t})$. When the distinction matters, we keep the two components separate. Let $\pi_{j}(\boldsymbol{a}_{-j},\boldsymbol{x}_{t},\boldsymbol{\kappa}_{t},\boldsymbol{\eta}_{jt})$ be firm $j$'s expected profit given its information about demand and costs and conditional on the hypothetical entry profile $\boldsymbol{a}_{-j} \in \{0,1\}^{J-1}$. Under Assumption (ref): \begin{equation} \pi_{j}(\boldsymbol{a}_{-j},\boldsymbol{x}_{t},\boldsymbol{\kappa}_{t},\boldsymbol{\eta}_{jt}) \; = \; {\int} VP_{j}(\boldsymbol{a}_{-j},\boldsymbol{x}_{t},\boldsymbol{\xi}_{t}) \; dF_{j,\xi} \left( \boldsymbol{\xi}_{t} \; | \; \boldsymbol{\kappa}_{t} \right) - fc(\boldsymbol{x}_{jt}, \boldsymbol{\kappa}_{t}, \boldsymbol{\eta}_{jt}), \end{equation} where $F_{j,\xi}\left( \boldsymbol{\xi}_{t} \; | \; \boldsymbol{\kappa}_{t} \right)$ is a CDF and represents firm $j$'s beliefs about the distribution of $\boldsymbol{\xi}_{t}$ conditional on $\boldsymbol{\kappa}_{t}$. Function $fc(\boldsymbol{x}_{jt}, \boldsymbol{\kappa}_{t}, \boldsymbol{\eta}_{jt})$ represents the fixed cost and entry cost of operating in the market. In the special case where firms face no uncertainty about market conditions at the time of entry, the common-knowledge vector is $\boldsymbol{\kappa}_{t} = \boldsymbol{\xi}_{t}$ and firm $j$’s profit function simplifies to $\pi_{j}(\boldsymbol{a}_{-j},\boldsymbol{x}_{t},\boldsymbol{\xi}_{t},\boldsymbol{\eta}_{jt}) \; = \; VP_{j}(\boldsymbol{a}_{-j},\boldsymbol{x}_{t},\boldsymbol{\xi}_{t}) - fc(\boldsymbol{x}_{jt}, \boldsymbol{\eta}_{jt})$. Assumption (ref) states that this entry game can accommodate complete information if the distribution of each $\boldsymbol{\eta}_{jt}$ is degenerate; otherwise, it is a game of incomplete information. Below, we describe an equilibrium of the game as a Bayesian Nash Equilibrium (BNE). However, this solution concept encompasses a complete information Nash Equilibrium (NE) when each $\boldsymbol{\eta}_{jt}$ has a degenerate probability distribution. Given $(\boldsymbol{x}_{t},\boldsymbol{\kappa}_{t})$, a Bayesian Nash Equilibrium (BNE) of this game can be represented as a $J$-tuple of entry probabilities, one for each firm, $(P_{jt}:j\in\mathcal{J})$. To describe this BNE, we first define a firm's expected profit function that accounts for its uncertainty about other firms' entry decisions. \begin{equation} \pi_{j}^{P}(\boldsymbol{x}_{t},\boldsymbol{\kappa}_{t},\boldsymbol{\eta}_{jt}) =\sum_{\boldsymbol{a}_{-j}\in\{0,1\}^{J-1}}\left( {\prod\limits_{i\neq j}} \left[ P_{it}\right] ^{a_{i}} \left[ 1-P_{it}\right] ^{1-a_{i}}\right) \pi_{j}(\boldsymbol{a}_{-j},\boldsymbol{x}_{t}, \boldsymbol{\kappa}_{t},\boldsymbol{\eta}_{jt}). \end{equation} Firm $j$'s best response is to enter the market if and only if this expected profit exceeds zero. Considering this, we can define a BNE in this game as follows. \begin{define} \textbf{Bayesian Nash Equilibrium.} Under Assumption (ref) and given $(\boldsymbol{x}_{t},\boldsymbol{\kappa}_{t})$, a Bayesian Nash Equilibrium (BNE) can be represented as a $J$-tuple of probabilities $\{ P_{jt} \equiv P_{j}(\boldsymbol{x}_{t},\boldsymbol{\kappa}_{t}): j \in \mathcal{J} \}$ that solves the following system of $J$ best response equations in the space of probabilities: \begin{equation} P_{jt} \; = \; \int \mathbb{1} \left\{ \pi_{j}^{P}(\boldsymbol{x}_{t},\boldsymbol{\kappa}_{t},\boldsymbol{\eta}_{jt}) \geq 0 \right\} \; dF_{\eta} \left( \boldsymbol{\eta}_{jt} \right). \qquad \blacksquare \end{equation} \end{define} \subsection{Special cases in the literature} This framework offers a general formulation that encompasses, as special cases, existing models of endogenous product entry with a structural demand system. To highlight this generality, Table (ref) summarizes several recent influential contributions, focusing on features relevant to this paper: the presence and nature of endogenous selection on unobservables, and in particular whether selection occurs on demand-side unobservables. As shown in Table (ref), most empirical applications of endogenous product entry assume that firms do not know demand or marginal-cost unobservables at the time they make entry decisions. In other words, these models rule out---by assumption---the possibility of selection on demand (variable-profit) unobservables. Two important exceptions are the models in ciliberto2021market and li2022repositioning. \begin{table} \caption{Models of endogenous product entry with structural differentiated-product demand} \begin{tabular}{p{2.3cm} | p{2.8cm} | p{2.8cm} | p{3.0cm} | p{2.8cm} | p{1.9cm} } \hline \hline \multicolumn{1}{c|}{Paper} & {Industry - \; \; \; \; \; Product Entry} & {Selection on \; \; \; demand (or MC) \; \; \; \; \; \; unobservables} & {What firms know about $\boldsymbol{\xi}_{t}$ at entry} & {Unobservables in entry cost} & {Entry game} \\ \hline \qquad \qquad \qquad \qquad aguirregabiria2012dynamic & \qquad \qquad \qquad \qquad US airlines - \; \; \; Route (city-pair) & \qquad \qquad \qquad \qquad NO, once airline & route FEs are accounted for. & \qquad \qquad \qquad \qquad \textcolor{blue}{$\boldsymbol{\xi}_{t}$ is unknown}. \qquad Firms know airline & route FEs but not demand shocks $\boldsymbol{\xi}_{t}$. & \qquad \qquad \qquad \qquad \textcolor{blue}{Only $\boldsymbol{\eta}_{jt}$}. Private information shocks assumed independent of $\boldsymbol{\xi}_{t}$. & \qquad \qquad \qquad \qquad Incomplete information dynamic game \\ \hline \qquad \qquad \qquad \qquad sweeting2013dynamic & \qquad \qquad \qquad \qquad US radio - \; \; \; \; \; \; Station genre & \qquad \qquad \qquad \qquad NO, once observable lagged variables are accounted for. & \qquad \qquad \qquad \qquad \textcolor{blue}{$\boldsymbol{\xi}_{t}$ is unknown}. Follows AR(1). Firms know lagged but not current $\xi$. & \qquad \qquad \qquad \qquad \textcolor{blue}{Only $\boldsymbol{\eta}_{jt}$}. Private information shocks assumed independent of $\boldsymbol{\xi}_{t}$. & \qquad \qquad \qquad \qquad Incomplete information dynamic game \\ \hline \qquad \qquad \qquad \qquad eizenberg2014upstream & \qquad \qquad \qquad \qquad US home PC - \; \; \; \; \; \; PC models & \qquad \qquad \qquad \qquad NO. Selection on entry cost unobservables, but are assumed independent of $\boldsymbol{\xi}_{t}$. & \qquad \qquad \qquad \qquad \textcolor{blue}{$\boldsymbol{\xi}_{t}$ is unknown}. \textit{"Firms only observe the realizations of $\boldsymbol{\xi}_{t}$ after committing to product choices".} & \qquad \qquad \qquad \qquad \textcolor{blue}{Includes common knowledge $\boldsymbol{\kappa}_{t}$} but are assumed independent of $\boldsymbol{\xi}_{t}$. & \qquad \qquad \qquad \qquad Complete information static game \\ \hline \qquad \qquad \qquad \qquad fan_yang_2020 & \qquad \qquad \qquad \qquad US smartphones - \; \; \; \; \; \; Phone models & \qquad \qquad \qquad \qquad NO. Selection on entry cost unobservables, but are assumed independent of $\boldsymbol{\xi}_{t}$. & \qquad \qquad \qquad \qquad \textcolor{blue}{$\boldsymbol{\xi}_{t}$ is unknown}. \qquad Firms know brand FEs but not demand shocks $\boldsymbol{\xi}_{t}$. & \qquad \qquad \qquad \qquad \textcolor{blue}{Includes common knowledge $\boldsymbol{\kappa}_{t}$} but are assumed independent of $\boldsymbol{\xi}_{t}$. & \qquad \qquad \qquad \qquad Complete information static game \\ \hline \qquad \qquad \qquad \qquad ciliberto2021market & \qquad \qquad \qquad \qquad US airlines - \; \; \; \; \; \; Route (city-pair) & \qquad \qquad \qquad \qquad YES & \qquad \qquad \qquad \qquad \textcolor{blue}{$\boldsymbol{\xi}_{t}$ is known}. \qquad Demand unobservables are known to firms when making product choices. & \qquad \qquad \qquad \qquad \textcolor{blue}{Includes common knowledge $\boldsymbol{\kappa}_{t}$} that can be correlated with $\boldsymbol{\xi}_{t}$. & \qquad \qquad \qquad \qquad Complete information static game \\ \hline \qquad \qquad \qquad \qquad li2022repositioning & \qquad \qquad \qquad \qquad US airlines - \; \; \; \; \; \; Route (city-pair) & \qquad \qquad \qquad \qquad YES & \qquad \qquad \qquad \qquad \textcolor{blue}{$\boldsymbol{\xi}_{t}$ is known}. \qquad Demand unobservables are known to firms when making product choices. & \qquad \qquad \qquad \qquad \textcolor{blue}{Includes common knowledge $\boldsymbol{\kappa}_{t}$} that can be correlated with $\boldsymbol{\xi}_{t}$. & \qquad \qquad \qquad \qquad Complete information static game \\ \hline \qquad \qquad \qquad \qquad bontemps2023price & \qquad \qquad \qquad \qquad US airlines - \; \; \; \; \; \; Airline's network of non-stop routes & \qquad \qquad \qquad \qquad NO. Selection on entry cost unobservables, but are assumed independent of $\boldsymbol{\xi}_{t}$. & \qquad \qquad \qquad \qquad \textcolor{blue}{$\boldsymbol{\xi}_{t}$ is unknown}. \qquad Firms don't know demand unobservables when making network choices. & \qquad \qquad \qquad \qquad \textcolor{blue}{Includes common knowledge $\boldsymbol{\kappa}_{t}$} but are assumed independent of $\boldsymbol{\xi}_{t}$. & \qquad \qquad \qquad \qquad Complete information static game \\ & & & & & \\ \hline \hline \end{tabular} \end{table} \section{Structure of the selection problem } \subsection{Selection-bias function} For the selection problem, it is convenient to work directly with the inverse-demand outcome from Proposition (ref). Recall from equation (ref) that, for any observation with $a_{jt}=1$, the demand inversion yields $d_{jt}^{-1}(\boldsymbol{s}_{t}^{\boldsymbol{a}_{t}}, \boldsymbol{\sigma}) = \alpha \, p_{jt} + \boldsymbol{x}_{jt}^{\prime} \boldsymbol{\beta} + \xi_{jt}$. Whether product $j$ is offered is determined by firm $j$'s equilibrium entry decision: \begin{equation} a_{jt} \; = \; \mathbb{1}\left\{ \pi_{j}^{P}(\boldsymbol{x}_{t},\boldsymbol{\kappa}_{t},\boldsymbol{\eta}_{jt}) \geq 0 \right\}. \end{equation} Since the sample of observations for which the demand equation holds is selected, we decompose the unobservable $\xi_{jt}$ into its conditional mean and a residual: $\xi_{jt} = \lambda_{j}(\boldsymbol{x}_{t}) + \widetilde{\xi}_{jt}$, where $\lambda_{j}(\boldsymbol{x}_{t}) \equiv \mathbb{E}(\xi_{jt} \mid \boldsymbol{x}_{t}, a_{jt}=1)$ is the \textit{selection-bias function} and $\widetilde{\xi}_{jt}$ is mean independent of $(\boldsymbol{x}_{t}, a_{jt}=1)$ by construction. This gives the following regression equation for any observation with $a_{jt}=1$: \begin{equation} d_{jt}^{-1}\left(\boldsymbol{s}_{t}^{\boldsymbol{a}_{t}},\boldsymbol{\sigma}\right) \; = \; \alpha \; p_{jt} + \boldsymbol{x}_{jt}^{\prime} \; \boldsymbol{\beta} + \lambda_{j}(\boldsymbol{x}_{t}) + \widetilde{\xi}_{jt}. \end{equation} The structure of the selection-bias function plays a key role in the identification of demand parameters $(\alpha, \, \boldsymbol{\beta}, \, \boldsymbol{\sigma})$. To characterize this structure, define the \textit{propensity score}---the probability of product $j$ being offered, conditional on observables---as: \begin{equation} \overline{P}_{j}\left(\boldsymbol{x}_{t}\right) \;\equiv\; \Pr\left(a_{jt} =1\;|\;\boldsymbol{x}_{t}\right) \;=\; {\int} \mathbb{1}\left\{ \pi_{j}^{P}(\boldsymbol{x}_{t}, \boldsymbol{\kappa}_{t},\boldsymbol{\eta}_{jt}) \geq 0 \right\} \; f_{\kappa}(\boldsymbol{\kappa}_{t}) \, f_{\eta}(\boldsymbol{\eta}_{jt}) \, d \boldsymbol{\kappa}_{t} \, d \boldsymbol{\eta}_{jt}. \end{equation} For values of $\boldsymbol{x}_{t}$ such that $\overline{P}_{j}(\boldsymbol{x}_{t}) > 0$, the selection-bias function can be written as: \begin{equation} \lambda_{j}(\boldsymbol{x}_{t}) \; = \; \frac{ \mathbb{E}\left( \xi_{jt} \, a_{jt} \;\middle|\; \boldsymbol{x}_{t} \right) }{ \overline{P}_{j}\left(\boldsymbol{x}_{t}\right) }. \end{equation} In the econometrics literature on sample selection, it is well known that estimating equation (ref) by instrumental variables---treating $\lambda_{j}(\boldsymbol{x}_{t}) + \widetilde{\xi}_{jt}$ as the composite error---is generally infeasible. The reason is that $\lambda_{j}(\boldsymbol{x}_{t})$ is an unknown function of all exogenous variables in the model, so any candidate instrument would also enter the error term through the selection function and therefore violate the exclusion restriction (wooldridge_pd_book_2010, wooldridge_pd_book_2010). A natural alternative is a control-function approach that explicitly accounts for the selection component $\lambda_{j}(\boldsymbol{x}_{t})$. However, without additional structure, this term remains an unrestricted function of the same exogenous variables that appear in demand. Consequently, the direct effect of $\boldsymbol{x}_{jt}$ on consumer demand (captured by $\boldsymbol{\beta}$) cannot be separated from its indirect effect operating through selection. In other words, absent further restrictions, the demand parameters are not identified. In this setting, the standard strategy in the literature is to impose conditions under which the selection term depends only on the propensity score $\overline{P}_{j}(\boldsymbol{x}_{t})$, that is, $\lambda_{j}(\boldsymbol{x}_{t}) = \rho_{j}\left(\overline{P}_{j}(\boldsymbol{x}_{t})\right)$. This restriction is powerful because it reduces the infinite-dimensional nuisance function $\lambda_{j}(\boldsymbol{x}_{t})$ to a single-index object. Once this dimensionality reduction holds, the effect of observables on demand can be separated from their effect through selection, restoring identification. Estimation then follows a standard two-step procedure. In the first step, the propensity score $\overline{P}_{j}(\boldsymbol{x}_{t})$ is estimated nonparametrically using data on $(a_{jt}, \boldsymbol{x}_{t})$. In the second step, one recovers the structural parameters using semiparametric methods, such as the series estimators in das2003 (das2003) and newey_2009 (newey_2009), or the pairwise differencing approaches in powell_2001 (powell_2001) and aradillas2012pairwise (aradillas2012pairwise). The remaining endogeneity of price arising from the standard simultaneity problem in demand estimation can be handled in the usual way: valid instruments are characteristics of other products, $\boldsymbol{x}_{-jt}$---the familiar BLP-type instruments. \subsection{Failure of the ordinary propensity score} Unfortunately, the conditions under which the selection-bias correction depends solely on the propensity score can fail in models of endogenous product entry and oligopoly competition, even under simple specifications. Intuitively, entry decisions in these environments depend on equilibrium profitability, which is jointly determined by unobserved demand and cost variables of all the competing products. As a result, the selection rule need not satisfy the single-index structure required for the propensity score to eliminate selection bias. Proposition (ref) presents this benchmark result for the standard propensity-score logic. Under the exogeneity conditions maintained in the paper, LATE-style monotonicity of the entry indicator is equivalent to a scalar threshold representation of that indicator, as in Theorem 1 of vytlacil_2002. These structural conditions are then sufficient for conditioning on the ordinary propensity score to eliminate selection bias, following Propositions 2 and 3 in angrist_1997. \begin{prop} Under independence between the unobservables $(\boldsymbol{\xi}_{t}, \boldsymbol{\kappa}_{t}, \boldsymbol{\eta}_{jt})$ and the observables $\boldsymbol{x}_{t}$, as in Assumption (ref)[d], consider the following conditions: \begin{itemize} • Monotonicity: For any $\boldsymbol{x}$ and $\boldsymbol{x}^{\prime}$, either $\mathbb{1}\{ \pi_{j}^{P}(\boldsymbol{x}, \boldsymbol{\kappa}_{t},\boldsymbol{\eta}_{jt}) \ge 0 \} \geq \mathbb{1}\{ \pi_{j}^{P}(\boldsymbol{x}^{\prime}, \boldsymbol{\kappa}_{t},\boldsymbol{\eta}_{jt}) \ge 0 \}$ for all $(\boldsymbol{\kappa}_{t},\boldsymbol{\eta}_{jt})$, or the inequality is reversed for all $(\boldsymbol{\kappa}_{t},\boldsymbol{\eta}_{jt})$. • Single-index representation: There are real-valued functions $\gamma_{1j}$ and $\gamma_{2j}$ such that: $\pi_{j}^{P}(\boldsymbol{x}_{t}, \boldsymbol{\kappa}_{t},\boldsymbol{\eta}_{jt}) \geq 0 \Longleftrightarrow \gamma_{1j}(\boldsymbol{x}_{t}) \; \geq \; \gamma_{2j}(\boldsymbol{\kappa}_{t},\boldsymbol{\eta}_{jt})$. • Conditional independence: $\Pr(\xi_{jt}, a_{jt} \mid \boldsymbol{x}_{t}, \overline{P}_{j}(\boldsymbol{x}_{t})) \; = \; \Pr(\xi_{jt}, a_{jt} \mid \overline{P}_{j}(\boldsymbol{x}_{t}))$. \end{itemize} Under the regularity conditions in Theorem 1 of vytlacil_2002 (vytlacil_2002), conditions (a) and (b) are equivalent. Furthermore, if either condition (a) or (b) holds, then (c) is satisfied. Finally, condition (c) is necessary and sufficient for the ordinary propensity score $\overline{P}_{j}(\boldsymbol{x}_{t})$ to control for selection bias. \qquad $\blacksquare$ \end{prop} Although the conditions in Proposition (ref) involve endogenous objects and are not straightforward to verify, the proposition is useful as a benchmark. The following example uses a minimal two-product environment satisfying Assumption (ref) to show that condition (c) easily fails in oligopoly models with endogenous entry, and that, as a consequence, the ordinary propensity score often cannot control for selection bias in demand estimation. \begin{example}[\textbf{\textit{Failure of propensity-score sufficiency under oligopoly entry}}] Consider a simple setting satisfying Assumption (ref) and the following features: (i) a standard logit demand system (no random coefficients); (ii) constant marginal costs $c_{i}(\boldsymbol{x}_{it})$ and exogenous price--cost margins $PCM_{i}(\boldsymbol{x}_{t})$, for $i \in \mathcal{J}$; (iii) the $J-1$ products $i \neq j$ are always offered, while product $j$’s entry is endogenous; (iv) the demand shocks $\boldsymbol{\xi}_{t} = (\xi_{it}: i \in \mathcal{J})$ are common knowledge among firms at the entry stage, and there is no private information: in the notation of Assumption (ref), we have $\boldsymbol{\kappa}_{t} = \boldsymbol{\xi}_{t}$ , and $\boldsymbol{\eta}_{jt}$ is degenerate; (v) entry costs depend only on observables; and (vi) $PCM_{j}(\boldsymbol{x}_t) \, H_t > EC_{j}(\boldsymbol{x}_{jt})$.\footnote{Without this condition, the probability of entry for firm $j$, conditional on observables $\boldsymbol{x}$, would be zero: even if the firm’s market share were arbitrarily close to one, its profit would still be negative. In that case, the selection problem would be irrelevant for those values of $\boldsymbol{x}$.} For the remainder of this example, we omit the subscript $t$ to simplify the notation. \textbf{Entry condition.} Firm $j$'s expected profit, upon entry, is $\pi_{j}^{P} = PCM_{j}(\boldsymbol{x}) \, H \, s_j - EC_{j}(\boldsymbol{x}_{j})$. Therefore, the entry condition $\pi_{j}^{P} \geq 0$ can be represented as: \begin{equation} \frac{e^{v_{j}(\boldsymbol{x}) + \xi_j}} {1 + e^{v_{j}(\boldsymbol{x}_{j}) + \xi_j} + \sum_{i \neq j} e^{v_{i}(\boldsymbol{x}) + \xi_i}} \; \geq \; \frac{EC_{j}(\boldsymbol{x}_{j})} {PCM_{j}(\boldsymbol{x}) \; H} \end{equation} where $v_{i}(\boldsymbol{x}) \equiv \boldsymbol{x}_{i}^{\prime}\boldsymbol{\beta} - \alpha [PCM_{i}(\boldsymbol{x}) + c_{i}(\boldsymbol{x}_{i})]$. After simple operations, we can represent this entry condition as a threshold rule for the demand unobservable $\xi_j$:\footnote{Multiplying equation (ref) by the denominators, we get: \[ PCM_j(\boldsymbol{x}) \; H \; e^{v_j(\boldsymbol{x}) + \xi_j} \geq EC_j(\boldsymbol{x}_j) \; \Bigl( 1 + e^{v_j(\boldsymbol{x}) + \xi_j} + \sum_{i \neq j} e^{v_{i}(\boldsymbol{x}) + \xi_i} \Bigr). \] Moving the term involving $e^{v_j(\boldsymbol{x}) + \xi_j}$ to the left: \[ \left( PCM_j(\boldsymbol{x}) \; H - EC_j(\boldsymbol{x}_j) \right) \; e^{v_j(\boldsymbol{x}) + \xi_j} \geq EC_j(\boldsymbol{x}_j) \; \Bigl( 1 + \sum_{i \neq j} e^{v_{i}(\boldsymbol{x}) + \xi_i} \Bigr). \] Taking logarithms and rearranging terms, we get equation (ref)} \begin{equation} a_{j} \; = \; \mathbb{1} \Big\{ \xi_j \; \geq \; \ln \Bigl( 1 + \sum_{i \neq j} e^{v_{i}(\boldsymbol{x}) + \xi_i} \Bigr) - B_{j}(\boldsymbol{x}) \Big\}, \end{equation} with $B_{j}(\boldsymbol{x}) \; \equiv \; v_j(\boldsymbol{x}) + \ln \left[ PCM_j(\boldsymbol{x}) \; H - EC_j(\boldsymbol{x}_j) \right] - \ln \left[ EC_j(\boldsymbol{x}_j) \right].$ \textbf{Failure of monotonicity and single-index (conditions (a)–(b), Proposition (ref)).} In this example, failure of monotonicity can be verified directly. Consider $\boldsymbol{x}$ and $\boldsymbol{x}'$ such that $[v_{k}(\boldsymbol{x}') - v_{k}(\boldsymbol{x})] > [B_{j}(\boldsymbol{x}') - B_{j}(\boldsymbol{x})] > 0$ for some $k \neq j$, and $v_i(\boldsymbol{x}') - v_i(\boldsymbol{x}) = 0$ for all $i \neq j,k$. Define the function: \begin{equation} \Delta_{j}(\xi_{k}) \equiv \ln \Bigl( 1 + \sum_{i \neq j} e^{v_{i}(\boldsymbol{x}^{\prime}) + \xi_i} \Bigr) - \ln \Bigl( 1 + \sum_{i \neq j} e^{v_{i}(\boldsymbol{x}) + \xi_i} \Bigr) - [B_{j}(\boldsymbol{x}^{\prime}) - B_{j}(\boldsymbol{x})] \end{equation} This function is continuous and strictly increasing, with $\Delta_{j}(\xi_{k}) \to -[B_{j}(\boldsymbol{x}') - B_{j}(\boldsymbol{x})] < 0$ as $\xi_{k} \to -\infty$, and $\Delta_{j}(\xi_{k}) \to [v_{k}(\boldsymbol{x}') - v_{k}(\boldsymbol{x})] - [B_{j}(\boldsymbol{x}') - B_{j}(\boldsymbol{x})] > 0$ as $\xi_{k} \to +\infty$. Hence, a crossing point exists: the ordering of the entry threshold reverses, so monotonicity fails. By equivalence, the single-index representation also fails. \textbf{Failure of sufficiency of the propensity score (condition (c), Proposition (ref)).} Again, consider $\boldsymbol{x}$ and $\boldsymbol{x}'$ such that $[v_{k}(\boldsymbol{x}') - v_{k}(\boldsymbol{x})] > [B_{j}(\boldsymbol{x}') - B_{j}(\boldsymbol{x})] > 0$ for some $k \neq j$, and $v_i(\boldsymbol{x}') - v_i(\boldsymbol{x}) = 0$ for all $i \neq j,k$. By continuity of $\overline{P}_{j}(\boldsymbol{x})$ and $\Pr(\xi_j, a_j \mid \boldsymbol{x})$ with respect to $v_{k}(\boldsymbol{x})$ and $B_{j}(\boldsymbol{x})$, and given the structure of the entry condition in this example, one can adjust these values so that $\overline{P}_{j}(\boldsymbol{x}') = \overline{P}_{j}(\boldsymbol{x})$, while still having $\Pr(\xi_j, a_j \mid \boldsymbol{x}') \neq \Pr(\xi_j, a_j \mid \boldsymbol{x})$. Thus, the propensity score is not sufficient. \qquad $\blacksquare$ \end{example} \begin{comment} \begin{example}[\textbf{\textit{Failure of propensity-score sufficiency under oligopoly entry}}] Consider a simple setting satisfying Assumption (ref) with two single-product firms, $j$ and $k$, with the following features: (i) a standard logit demand system (no random coefficients); (ii) constant marginal costs $c_{i}(\boldsymbol{x}_{it})$ and exogenous price--cost margins $PCM_{i}(\boldsymbol{x}_{t})$, for $i \in \{j,k\}$; (iii) product $k$ is always offered, while product $j$’s entry is endogenous; (iv) the demand shocks $(\xi_{jt}, \xi_{kt})$ are common knowledge among firms at the entry stage, and there is no private information; (v) entry costs depend only on observables, $EC_{j}(\boldsymbol{x}_{jt})$, with $PCM_{j}(\boldsymbol{x}) > EC_{j}(\boldsymbol{x}_{j})$; (vi) $\xi_{jt}$ and $\xi_{kt}$ are i.i.d.\ standard normal and independent of $\boldsymbol{x}_{t}$. In the notation of Assumption (ref), set $\boldsymbol{\kappa}_{t} = (\xi_{jt}, \xi_{kt})$ , and $\boldsymbol{\eta}_{jt}$ is degenerate. For the remainder of this example, we omit the subscript $t$ to simplify the notation. \underline{Entry condition.} Define $v_{i}(\boldsymbol{x}) \equiv \boldsymbol{x}_{i}^{\prime}\boldsymbol{\beta} - \alpha [PCM_{i}(\boldsymbol{x}) + c_{i}(\boldsymbol{x}_{i})]$ for $i \in \{j,k\}$, and $B_{j}(\boldsymbol{x}) \equiv v_{j}(\boldsymbol{x}) + \ln\frac{PCM_{j}(\boldsymbol{x}) - EC_{j}(\boldsymbol{x}_{j})}{EC_{j}(\boldsymbol{x}_{j})}$. Then $\pi_{j}^{P} \geq 0$ reduces to the threshold rule \begin{equation} a_{j} = \mathbb{1}\left\{ \kappa_{j} \geq t_{\boldsymbol{x}}(\kappa_{k}) \right\}, \qquad t_{\boldsymbol{x}}(\kappa_{k}) \equiv \ln\!\left(1 + e^{v_{k}(\boldsymbol{x}) + \kappa_{k}}\right) - B_{j}(\boldsymbol{x}). \end{equation} The threshold depends on $\boldsymbol{x}$ through two independent channels: $B_{j}(\boldsymbol{x})$, which shifts the threshold uniformly (reflecting firm $j$’s own profitability), and $v_{k}(\boldsymbol{x})$, which controls the threshold’s sensitivity to the rival’s component $\kappa_{k}$ of $\boldsymbol{\kappa}$ (reflecting the intensity of competition). \underline{Condition (c) fails.} Since $\kappa_{j} \sim N(0,1)$, the identity $\mathbb{E}(\kappa_{j} \cdot \mathbb{1}\{\kappa_{j} \geq c\}) = \phi(c)$ gives \[ \overline{P}_{j}(\boldsymbol{x}) = \mathbb{E}_{\kappa_{k}}\!\left[ \Phi(-t_{\boldsymbol{x}}(\kappa_{k})) \right], \qquad \lambda_{j}(\boldsymbol{x}) = \frac{ \mathbb{E}_{\kappa_{k}}\!\left[ \phi(t_{\boldsymbol{x}}(\kappa_{k})) \right] }{ \overline{P}_{j}(\boldsymbol{x}) }, \] where $\phi$ and $\Phi$ denote the standard normal density and CDF, respectively, and the factorization uses the independence of $\kappa_{j}$ and $\kappa_{k}$. Writing $p(\kappa_{k}) \equiv \Phi(-t_{\boldsymbol{x}}(\kappa_{k}))$ for the entry probability conditional on the rival’s component of $\boldsymbol{\kappa}$, and defining $G(p) \equiv \phi(\Phi^{-1}(p))$, we have $\overline{P}_{j} = \mathbb{E}[p(\kappa_{k})]$ and $\lambda_{j} = \mathbb{E}[G(p(\kappa_{k}))] / \overline{P}_{j}$. A direct calculation shows $G^{\prime}(p) = -\Phi^{-1}(p)$ and $G^{\prime\prime}(p) = -1/\phi(\Phi^{-1}(p)) < 0$, so $G$ is strictly concave. Since the threshold $t_{\boldsymbol{x}}(\kappa_{k})$ is strictly increasing in $\kappa_{k}$ (for any finite $v_{k}$), the random variable $p(\kappa_{k})$ is non-degenerate. By Jensen’s inequality, \[ \mathbb{E}\!\left[G(p(\kappa_{k}))\right] \;<\; G\!\left(\mathbb{E}[p(\kappa_{k})]\right) = G(\overline{P}_{j}), \] and therefore \[ \lambda_{j}(\boldsymbol{x}) \;<\; \frac{\phi\!\left(\Phi^{-1}(\overline{P}_{j}(\boldsymbol{x}))\right)} {\overline{P}_{j}(\boldsymbol{x})}. \] The right-hand side is the inverse Mills ratio---the classical Heckman selection correction that would obtain if the entry rule were a single-threshold model (i.e., if the rival were irrelevant). Oligopoly competition drives $\lambda_{j}$ strictly below this benchmark. Moreover, $\lambda_{j}$ is not a function of $\overline{P}_{j}$ alone. Fix any $\overline{P}_{j} = p_{0} \in (0,1)$. For each value of $v_{k}$, there exists a unique $B_{j}(v_{k})$ that maintains $\overline{P}_{j} = p_{0}$ by the intermediate value theorem, since $\overline{P}_{j}$ is continuous and strictly increasing in $B_{j}$ with limits $0$ and $1$. Along this curve, the mapping $v_{k} \mapsto \lambda_{j}(v_{k}, B_{j}(v_{k}))$ is continuous (by dominated convergence), strictly below $G(p_{0})/p_{0}$ for every finite $v_{k}$ (by Jensen’s inequality, since $p(\kappa_{k})$ is non-degenerate), and approaches $G(p_{0})/p_{0}$ as $v_{k} \to -\infty$ (the degenerate limit where the rival is irrelevant and $p(\kappa_{k}) \to p_{0}$ almost surely). If this mapping were constant at some value $c$ for all finite $v_{k}$, then its limit would also equal $c$; but $c < G(p_{0})/p_{0}$ by the strict Jensen bound, contradicting the fact that the limit equals $G(p_{0})/p_{0}$. Hence the mapping is non-constant: there exist covariate pairs with the same $\overline{P}_{j}$ but different $\lambda_{j}$, and condition (c) in Proposition (ref) fails. \underline{Conditions (a) and (b) also fail.} By Proposition (ref), failure of (c) rules out both (a) monotonicity and (b) the single-index representation. In this simple example, failure of monotonicity can also be verified directly. Define \[ h(\kappa_{k}) \;\equiv\; \ln\!\left(1 + e^{v_{k}^{\prime} + \kappa_{k}}\right) - \ln\!\left(1 + e^{v_{k} + \kappa_{k}}\right), \] where $v_{k}^{\prime} > v_{k}$. Then $h$ is continuous, strictly increasing, with $h(\kappa_{k}) \to 0$ as $\kappa_{k} \to -\infty$ and $h(\kappa_{k}) \to v_{k}^{\prime} - v_{k}$ as $\kappa_{k} \to +\infty$. The entry threshold difference between $\boldsymbol{x}^{\prime}$ and $\boldsymbol{x}$ is $t_{\boldsymbol{x}^{\prime}}(\kappa_{k}) - t_{\boldsymbol{x}}(\kappa_{k}) = h(\kappa_{k}) - \Delta B_{j}$, where $\Delta B_{j} \equiv B_{j}(\boldsymbol{x}^{\prime}) - B_{j}(\boldsymbol{x})$. For covariate vectors satisfying $0 < \Delta B_{j} < v_{k}^{\prime} - v_{k}$---a market condition change that improves firm $j$’s own profitability but strengthens the rival by more---this difference is negative as $\kappa_{k} \to -\infty$ and positive as $\kappa_{k} \to +\infty$. By continuity, there exists a crossing point: the entry indicator ordering reverses, and monotonicity fails. \qquad $\blacksquare$ \end{example} \end{comment} This example shows that, even in simple environments, the ordinary propensity score need not be sufficient to control for selection bias. The reason is that the rival's component of the latent state, $\kappa_{kt}$, affects firm $j$'s entry threshold through oligopoly competition. As a result, selection depends on the full latent state $\boldsymbol{\kappa}_{t}$, whereas the ordinary propensity score averages over $\boldsymbol{\kappa}_{t}$ and therefore loses relevant information. This motivates characterizing the selection-bias function in terms of entry probabilities conditional on $(\boldsymbol{x}_{t},\boldsymbol{\kappa}_{t})$. \subsection{Latent propensity scores and the selection-bias function} In this section, we derive a representation of the selection-bias function $\lambda_{j}(\boldsymbol{x}_{t})$ in terms of equilibrium entry probabilities conditional on $(\boldsymbol{x}_{t}, \boldsymbol{\kappa}_{t})$. Define the equilibrium entry probability of product $j$, conditional on both observables and common-knowledge unobservables, as \[ P_{j}(\boldsymbol{x}_{t}, \boldsymbol{\kappa}_{t}) \;\equiv\; \Pr\left( a_{jt} = 1 \;\middle|\; \boldsymbol{x}_{t}, \boldsymbol{\kappa}_{t} \right). \] We refer to $P_{j}(\boldsymbol{x}_{t}, \boldsymbol{\kappa}_{t})$ as the \textit{latent propensity score}, to highlight both its connection to the ordinary propensity score and its dependence on the latent state $\boldsymbol{\kappa}_{t}$. The ordinary propensity score is simply the average of the latent propensity scores over the distribution of $\boldsymbol{\kappa}$: \begin{equation} \overline{P}_j\left( \boldsymbol{x}_{t} \right) \; = \; \int P_{j}\left( \boldsymbol{x}_{t}, \boldsymbol{\kappa} \right) \, f_{\kappa} \left( \boldsymbol{\kappa} \right) \, d \boldsymbol{\kappa}. \end{equation} Proposition (ref) shows that, under only the information-structure conditions in Assumption (ref), the selection-bias function admits a mixture representation in terms of these latent propensity scores. Unlike the ordinary propensity-score approach, this representation does not require any monotonicity or single-index assumption. \begin{prop} Under Assumption (ref), and for values of $\boldsymbol{x}_{t}$ such that $\overline{P}_j\left( \boldsymbol{x}_{t} \right) > 0$, the selection-bias function $\lambda_j(\boldsymbol{x}_{t})$ admits the following mixture representation: \begin{equation} \lambda_{j}(\boldsymbol{x}_{t}) \; = \; \int \left[ \frac{P_{j}\left( \boldsymbol{x}_{t}, \boldsymbol{\kappa} \right)}{\overline{P}_j\left( \boldsymbol{x}_{t} \right)} \right] \, \mu_{j}(\boldsymbol{\kappa}) \, f_{\kappa}(\boldsymbol{\kappa}) \, d \boldsymbol{\kappa}. \end{equation} with $\mu_{j}(\boldsymbol{\kappa}) \equiv \mathbb{E}( \xi_{jt} \mid \boldsymbol{\kappa}_{t} = \boldsymbol{\kappa})$. $\qquad \blacksquare$ \end{prop} \textit{Proof:} In Appendix (ref). Proposition (ref) provides the first building block for our identification strategy and estimation procedure to control for endogenous product selection in the estimation of demand systems. In Section (ref), we establish semiparametric identification of the latent propensity score function using the empirical joint distribution of entry decisions across all products, $\Pr(a_{1t}, a_{2t}, \dots, a_{Jt} \mid \boldsymbol{x}_{t})$. The key intuition is that cross-product dependence in entry decisions reveals how these decisions are jointly affected by the common unobserved component $\boldsymbol{\kappa}_{t}$. Variation in this dependence structure allows us to recover the latent propensity score without imposing restrictive parametric assumptions. Combined with Proposition (ref), this implies that once the selection-bias function is approximated by sieve, it can be represented as a linear index in constructed regressors. As a consequence, the demand parameters can be identified. \section{Identification and estimation } \subsection{Setting} Suppose that each of the $J$ firms is a potential entrant in every local market. The researcher observes these firms in a random sample of $T$ markets. For every market $t$, the researcher observes the vector of exogenous variables $\boldsymbol{x}_{t} \in \mathcal{X}$ and the vectors of firms' entry decisions $\boldsymbol{a}_{t} \in \{0,1\}^{J}$. The space $\mathcal{X}$ can be discrete or continuous. For those firms active in market $t$, the researcher observes prices $\boldsymbol{p}_{t}$ and market shares $\boldsymbol{s}_{t}$. Recall the vector of demand parameters $\boldsymbol{\theta} \equiv (\alpha, \boldsymbol{\beta}^{\prime}, \boldsymbol{\sigma}^{\prime})^{\prime}$. Let $\boldsymbol{P} \equiv \{ P_{j}(\boldsymbol{x}, \boldsymbol{\kappa}) : \forall (j, \boldsymbol{x}, \boldsymbol{\kappa}) \}$ be the collection of equilibrium latent propensity scores, and let $\boldsymbol{f}_{\kappa} \equiv \{ f_{\kappa}(\boldsymbol{\kappa}) : \forall \boldsymbol{\kappa} \}$ denote the distribution of the unobserved heterogeneity $\boldsymbol{\kappa}$. Finally, let $\boldsymbol{\mu} \equiv \{ \mu_{j}(\boldsymbol{\kappa}) : \forall (j,\boldsymbol{\kappa}) \}$ denote the collection of conditional expectations $\mu_{j}(\boldsymbol{\kappa}) = \mathbb{E}(\xi_{jt} \mid \boldsymbol{\kappa}_t = \boldsymbol{\kappa})$. We adopt a two-step sequential identification strategy for $\boldsymbol{\theta}$. In the first step, we use the empirical distribution of firms’ entry decisions to identify the equilibrium probabilities $\boldsymbol{P}$ and the distribution $f_{\kappa}$. In the second step, we exploit the structure of the selection-bias function in equation (ref) to identify the demand parameters $\boldsymbol{\theta}$ and the incidental parameters $\boldsymbol{\mu}$. The econometric model consists of three sets of equations: (i) the demand regression equations, \begin{equation} d_{jt}^{-1}\left(\boldsymbol{s}_{t}^{\boldsymbol{a}_{t}},\boldsymbol{\sigma}\right) = \alpha \; p_{jt} + \boldsymbol{x}_{jt}^{\prime} \; \boldsymbol{\beta} + \lambda_{j}(\boldsymbol{x}_{t}) + \widetilde{\xi}_{jt}, \end{equation} (ii) the equation for the selection-bias function, \begin{equation} \lambda_{j}(\boldsymbol{x}_{t}) \; = \; \int \left[ \frac{P_{j}\left( \boldsymbol{x}_{t}, \boldsymbol{\kappa} \right)}{\overline{P}_j\left( \boldsymbol{x}_{t} \right)} \right] \, \mu_{j}(\boldsymbol{\kappa}) \, f_{\kappa}(\boldsymbol{\kappa}) \, d \boldsymbol{\kappa}, \end{equation} and (iii) the joint distribution of firms’ entry decisions conditional on the observed state variables implied by the equilibrium of the entry game. Under Assumption (ref), this distribution has a nonparametric mixture representation: \begin{equation} \Pr \left(a_{1t}, a_{2t}, \dots, a_{Jt} \mid \boldsymbol{x}_{t} \right) \; = \; \int \Bigg( \prod_{j=1}^{J} P_{j}(\boldsymbol{x}_t, \boldsymbol{\kappa}) ^{a_{jt}} \; \left[1-P_{j}(\boldsymbol{x}_t, \boldsymbol{\kappa}) \right]^{1-a_{jt}} \Bigg) \, f_{\kappa}( \boldsymbol{\kappa} ) \, d\boldsymbol{\kappa}. \end{equation} Conditional on $(\boldsymbol{x}_t, \boldsymbol{\kappa}_t)$, firms’ entry decisions depend only on their private shocks, which are assumed to be independent across firms. As a result, entry decisions are conditionally independent given $(\boldsymbol{x}_t, \boldsymbol{\kappa}_t)$. All residual dependence across firms’ entry decisions---beyond what is explained by observables---is therefore driven by the unobservables in $\boldsymbol{\kappa}_{t}$. This structure provides the key source of identification of $(\boldsymbol{P}, f_{\kappa})$. The joint distribution of $(a_{1t}, a_{2t}, \dots,$ $a_{Jt})$ conditional on $\boldsymbol{x}_t$, and in particular the cross-sectional correlation in entry decisions, reveals how the common unobservable $\boldsymbol{\kappa}_{t}$ shifts firms’ entry probabilities. In other words, the dependence across firms’ decisions encodes information about both the distribution $f_{\kappa}$ and the probability functions $P_j(\boldsymbol{x}_t, \boldsymbol{\kappa}_t)$. We establish identification of $(\boldsymbol{\theta}, \boldsymbol{\mu}, \boldsymbol{P}, f_{\kappa})$ using a two-step sieve-based approach. Specifically, we construct a sequence of sieve spaces over the support of the continuous unobserved heterogeneity $\boldsymbol{\kappa}_t$ and use these approximating spaces to represent the unknown objects $\boldsymbol{\mu}$, $\boldsymbol{P}$, and $f_{\kappa}$ as functions of $\boldsymbol{\kappa}$. \subsection{Two-Step Sieve-Based Approach} Equation (ref) shows that the selection-bias function depends on the latent heterogeneity only through three primitive objects: the latent propensity score $P_j(\boldsymbol{x}, \boldsymbol{\kappa})$, the density $f_{\kappa}(\boldsymbol{\kappa})$, and the conditional mean function $\mu_j(\boldsymbol{\kappa})$. This decomposition is key, as it reduces the infinite-dimensional dependence on $\boldsymbol{\kappa}$ to a structured combination of these objects and provides the foundation for the sieve-approximation approach that follows. The latent heterogeneity vector $\boldsymbol{\kappa}_t$ may be continuously distributed with support $\mathcal{K}$ on $\mathbb{R}^{d_{\kappa}}$. For example, in our differentiated-products setting, $\boldsymbol{\kappa}_t$ may collect the product-specific demand and cost unobservables, $\boldsymbol{\kappa}_t = (\xi_{jt}, \omega_{jt} : j \in \mathcal{J}) \in \mathbb{R}^{2J}$. Accordingly, $P_j(\boldsymbol{x}, \boldsymbol{\kappa})$, $f_{\kappa}(\boldsymbol{\kappa})$, and $\mu_j(\boldsymbol{\kappa})$ are all functions defined on the space $\mathbb{R}^{d_{\kappa}}$. To approximate these objects, we employ a discrete sieve. For each integer $L \geq 1$, let $\{ \mathcal{K}_1, \mathcal{K}_2, \dots, \mathcal{K}_L \}$ denote a partition of the support $\mathcal{K} \subseteq \mathbb{R}^{d_{\kappa}}$ of $\boldsymbol{\kappa}_t$. Each element of the partition can be interpreted as a \textit{latent market type} (or latent sieve class), thereby providing a finite-dimensional approximation to the underlying continuous heterogeneity. For each latent market type $\ell = 1, 2, \dots, L$, define its probability \begin{equation} \widetilde{f}_{\ell} \; \equiv \; \Pr \left( \boldsymbol{\kappa}_t \in \mathcal{K}_{\ell} \right) \; = \; \int_{\boldsymbol{\kappa} \in \mathcal{K}_{\ell}} f_{\kappa}(\boldsymbol{\kappa}) \, d\boldsymbol{\kappa}. \end{equation} Define the type-specific average propensity score \begin{equation} \widetilde{P}_{j,\ell}(\boldsymbol{x}) \; \equiv \; \frac{1}{\widetilde{f}_{\ell}} \int_{\boldsymbol{\kappa} \in \mathcal{K}_{\ell}} P_j(\boldsymbol{x}, \boldsymbol{\kappa}) \, f_{\kappa}(\boldsymbol{\kappa}) \, d \boldsymbol{\kappa}, \end{equation} as well as the type-specific average of the function $\mu_j(\boldsymbol{\kappa})$: \begin{equation} \widetilde{\mu}_{j, \ell} \; \equiv \; \frac{1}{\widetilde{f}_{\ell}} \int_{\boldsymbol{\kappa} \in \mathcal{K}_{\ell}} \mu_j(\boldsymbol{\kappa}) \, f_{\kappa}(\boldsymbol{\kappa}) \, d\boldsymbol{\kappa}. \end{equation} These definitions induce a finite-mixture approximation to the distribution of entry decisions: \begin{equation} \Pr \left(a_{1t}, a_{2t}, \dots, a_{Jt} \mid \boldsymbol{x}_{t} \right) \; \approx \; \widetilde{\Pr}^{L}( \boldsymbol{a}_{t} \mid \boldsymbol{x}) \; \equiv \; \sum_{\ell=1}^{L} \widetilde{f}_{\ell} \, \Bigg( \prod_{j=1}^J \widetilde{P}_{j,\ell}(\boldsymbol{x})^{a_{jt}} \left[1-\widetilde{P}_{j,\ell}(\boldsymbol{x})\right]^{1-a_{jt}} \Bigg). \end{equation} Likewise, the finite partition induces the sieve approximation \begin{equation} \lambda_{j}(\boldsymbol{x}) \; \approx \; \widetilde{\lambda}^{L}_{j}(\boldsymbol{x}) \; \equiv \; \sum_{\ell=1}^{L} \left[ \frac{\widetilde{P}_{j,\ell}(\boldsymbol{x})} {\overline{P}_{j}(\boldsymbol{x})} \right] \widetilde{\mu}_{j,\ell} \, \widetilde{f}_{\ell} \end{equation} Expression (ref) is the discrete-support analog of the mixture representation in Proposition (ref). Once the latent space has been partitioned into $L$ classes, the selection-bias function is summarized by the finite-dimensional objects $\{ \widetilde{f}_{\ell} \}_{\ell=1}^{L}$, $\{ \widetilde{P}_{j,\ell}(\boldsymbol{x})\}_{\ell=1}^{L}$, and $\{\widetilde{\mu}_{j,\ell}\}_{\ell=1}^{L}$. Therefore, if the first step delivers sufficiently accurate approximations to the class probabilities and the class-specific propensity scores, and if the second step provides a sufficiently accurate approximation to the function $\mu_{j}(\boldsymbol{\kappa})$, then the induced approximation to $\lambda_j(\boldsymbol{x})$ will also be accurate. The finite-mixture model should therefore be interpreted as a sieve, not as a literal restriction that the true latent heterogeneity has finite support. As $L$ increases and the partition becomes finer, the type probabilities $\widetilde{f}_{\ell}$ approximate the true distribution $f_{\kappa}$, the type-specific choice probabilities $\widetilde{P}_{j, \ell}(\boldsymbol{x})$ approximate the true latent propensity score function $P_j(\boldsymbol{x},\boldsymbol{\kappa})$, and the type-specific averages $\widetilde{\mu}_{j, \ell}$ approximate the true conditional expectation function $\mu_j(\boldsymbol{\kappa})$. Under standard regularity conditions---specifically, boundedness and continuity of the functions in $\boldsymbol{\kappa}$, uniformly in $\boldsymbol{x}$---there exists a sequence of partitions indexed by $L$ such that the approximation $\widetilde{\Pr}^{L}( \boldsymbol{a} \mid \boldsymbol{x})$ converges to the true $\Pr(\boldsymbol{a}\mid \boldsymbol{x})$ for every $\boldsymbol{a}\in\{0,1\}^J$ and every $\boldsymbol{x}$, and the approximation $\widetilde{\lambda}^{L}_{j}(\boldsymbol{x})$ converges to the true $\lambda_j(\boldsymbol{x})$ for every $j$ and every $\boldsymbol{x}$. Hence, the finite-support model can be viewed as a sieve approximation to the continuous-support structural model. Note that the parameters $\widetilde{\mu}_{j\ell}$ have a clear structural interpretation as $\mathbb{E}( \xi_{jt} \mid \boldsymbol{\kappa}_t \in \mathcal{K}_\ell )$. Thus, although the mixture model cannot be linked a priori with a specific partition of the continuous unobserved space, the estimation delivers an economically meaningful partition ex post. In particular, the estimates $\widetilde{\mu}_{j\ell}$ map each latent market type to the average unobserved demand for each product. For example, one type may feature high unobserved demand for products 1 and 2 but low for others---and this type may be associated with a low entry probability for product 1---while another type may exhibit high demand for products 1 and 3 and a high entry probability for product 1. This provides direct insight into how latent market types relate product selection to underlying demand heterogeneity. \subsection{First Step Identification} For a given number of support points $L$, equation (ref) defines a nonparametric finite mixture model. This representation is closely related to discrete choice games with incomplete information and finite-support unobserved heterogeneity, as studied in aguirregabiria2019identification and xiao2018identification. These papers draw on results from the nonparametric finite mixture literature---such as hall2003nonparametric (hall2003nonparametric), allman2009identifiability (allman2009identifiability), and kasahara2014non (kasahara2014non)---to establish identification of the mixture components $\{ \widetilde{f}_{\ell} \}_{\ell=1}^{L}$, the component-specific choice probabilities $\{ \widetilde{P}_{j,\ell}(\boldsymbol{x}) \}_{\ell=1}^{L}$, and the number of latent types $L$. In particular, Theorem 4 and Corollary 5 in allman2009identifiability (allman2009identifiability) provide primitive conditions under which these objects are nonparametrically identified. In nonparametric finite-mixture models, identification is obtained only up to a permutation of the latent types (\textit{label swapping}). A key implication of our framework is that the latent heterogeneity $\boldsymbol{\kappa}$ is independent of the observed covariates $\boldsymbol{x}$. This independence ensures that any admissible relabelling of the latent types must be global, i.e., it applies uniformly across all values of $\boldsymbol{x}$. As a result, label swapping has no substantive consequences for subsequent analysis. In particular, it does not affect the construction of the control function in the second step of our identification and estimation procedure, since all objects entering that step are invariant to a common permutation of the latent types. \subsection{Second Step Identification} The second-step regression equation is: \begin{equation} d_{jt}^{-1}\left(\boldsymbol{s}_{t}^{\boldsymbol{a}_{t}},\boldsymbol{\sigma}\right) \; = \; \alpha \, p_{jt} + \boldsymbol{x}_{jt}^{\prime} \boldsymbol{\beta} + \sum_{\ell=1}^{L} \widetilde{\mu}_{j \ell} \left[ \frac{\widetilde{P}_{j,\ell}(\boldsymbol{x}_t)} {\overline{P}_{j}(\boldsymbol{x}_t)} \, \widetilde{f}_{\ell} \right] + \widetilde{\xi}_{jt}, \end{equation} Hence, given the first-step objects, the selection-bias function is linear in the unknown coefficients $\{ \widetilde{\mu}_{j \ell} \}_{\ell=1}^{L}$. This is an important implication of Proposition (ref): the second-step unknown function $\mathbb{E}(\xi_{jt}\mid \boldsymbol{\kappa}_t=\boldsymbol{\kappa})$ enters the regression only through its type-averages, and those type-averages appear linearly once the first-step latent propensity scores and type probabilities are known. Before establishing identification of the parameters in this regression equation, note a potential perfect-collinearity issue and the linear restriction on the $\widetilde{\mu}_{j \ell}$ that resolves it. First, by definition, the constructed regressors for the control function sum to one: \begin{equation} \sum_{\ell=1}^{L} \frac{\widetilde{P}_{j,\ell}(\boldsymbol{x}_t) \, \widetilde{f}_{\ell} } {\overline{P}_{j}(\boldsymbol{x}_t)} \; = \; \frac{\overline{P}_{j}(\boldsymbol{x}_t)} {\overline{P}_{j}(\boldsymbol{x}_t)} \; = \; 1 \end{equation} Second, the definition of the $\widetilde{\mu}_{j \ell}$ parameters implies that \begin{equation} \sum_{\ell=1}^{L} \widetilde{\mu}_{j\ell} \, \widetilde{f}_\ell \; = \; \sum_{\ell=1}^{L} \mathbb{E}( \xi_{jt} \mid \boldsymbol{\kappa} \in \mathcal{K}_{\ell}) \, \Pr(\boldsymbol{\kappa} \in \mathcal{K}_{\ell}) \; = \; \mathbb{E}( \xi_{jt} ) \; = \; 0 \end{equation} Therefore, only $L-1$ of the $L$ parameters $\widetilde{\mu}_{j \ell}$ are free. Taking type $L$ as the reference type: \begin{equation} \widetilde{\mu}_{jL} \; = \; -\sum_{\ell=1}^{L-1} \frac{\widetilde f_\ell}{\widetilde f_L} \; \widetilde{\mu}_{j\ell}. \end{equation} Substituting (ref) into the selection-bias term, \begin{equation} \begin{array}[c]{rcl} \sum_{\ell=1}^{L} \widetilde{\mu}_{j \ell} \left[ \frac{\widetilde{P}_{j,\ell}(\boldsymbol{x}_t)} {\overline{P}_{j}(\boldsymbol{x}_t)} \, \widetilde{f}_{\ell} \right] & = & \sum_{\ell=1}^{L-1} \widetilde{\mu}_{j \ell} \left[ \frac{\widetilde{P}_{j,\ell}(\boldsymbol{x}_t)} {\overline{P}_{j}(\boldsymbol{x}_t)} \, \widetilde{f}_{\ell} \right] - \left( \sum_{\ell=1}^{L-1} \frac{\widetilde f_\ell}{\widetilde f_L} \; \widetilde{\mu}_{j\ell} \right) \left[ \frac{\widetilde{P}_{j,L}(\boldsymbol{x}_t)} {\overline{P}_{j}(\boldsymbol{x}_t)} \, \widetilde{f}_{L} \right] \\ \\ & = & \sum_{\ell=1}^{L-1} \widetilde{\mu}_{j \ell} \left[ \frac{\widetilde{P}_{j,\ell}(\boldsymbol{x}_t) - \widetilde{P}_{j,L}(\boldsymbol{x}_t)} {\overline{P}_{j}(\boldsymbol{x}_t)} \, \widetilde{f}_{\ell} \right] \end{array} \end{equation} Then the second-step regression can be written as \begin{equation} d_{jt}^{-1}\left(\boldsymbol{s}_{t}^{\boldsymbol{a}_{t}},\boldsymbol{\sigma}\right) \; = \; \alpha \, p_{jt} + \boldsymbol{x}_{jt}^{\prime} \boldsymbol{\beta} + \sum_{\ell=1}^{L-1} \widetilde{\mu}_{j\ell} \, r_{j\ell t} + \widetilde{\xi}_{jt}, \end{equation} where the coefficient vector $\widetilde{\boldsymbol\mu}_j \equiv (\widetilde\mu_{j1},\dots,\widetilde\mu_{j, L-1})^{\prime}$ contains the $L-1$ free parameters, and the constructed regressors $(r_{j 1 t}, r_{j 2 t}, ..., r_{j, L-1, t})$ have the following definition: \begin{equation} r_{j \ell t} \; = \; \frac{\widetilde{P}_{j,\ell}(\boldsymbol{x}_t) - \widetilde{P}_{j,L}(\boldsymbol{x}_t)} {\overline{P}_{j}(\boldsymbol{x}_t)} \, \widetilde{f}_{\ell} \end{equation} Conditional on $\boldsymbol{\sigma}$, equation (ref) is linear in the remaining parameters $(\alpha, \boldsymbol{\beta}, \widetilde{\boldsymbol{\mu}}_j)$. However, $\boldsymbol{\sigma}$ enters nonlinearly through the demand inverse $d_{jt}^{-1}(\boldsymbol{s}_{t}^{\boldsymbol{a}_{t}}, \boldsymbol{\sigma})$, so identification of the full parameter vector requires additional conditions. Proposition (ref) provides sufficient conditions for local identification of all parameters through a standard rank condition. \begin{prop} Let $\widetilde{\boldsymbol\mu}_{j} \equiv (\widetilde\mu_{j1},\dots,\widetilde\mu_{j,L-1})^{\prime}$ and define the parameter vector $\boldsymbol\vartheta_{j} \equiv (\alpha,\boldsymbol\beta^{\prime},\boldsymbol\sigma^{\prime},\widetilde{\boldsymbol\mu}_{j}^{\prime})^{\prime}$. Let $\boldsymbol{z}_{jt}$ denote a $[\dim(\boldsymbol{\sigma})+1] \times 1$ vector of instruments consisting of functions of the characteristics of products other than $j$, and define $\boldsymbol{r}_{jt} \equiv \big( r_{j1t}, r_{j2t}, \dots, r_{j,L-1,t}\big)^{\prime}$ and $\boldsymbol{w}_{jt} \equiv \big( \boldsymbol{z}_{jt}^{\prime}, \boldsymbol{x}_{jt}^{\prime}, \boldsymbol{r}_{jt}^{\prime} \big)^{\prime}$. Consider the vector-valued moment function evaluated over the selected sample: \begin{equation*} \boldsymbol{m}_{j}(\boldsymbol\vartheta_{j}) \;\equiv\; \mathbb{E} \left( \boldsymbol{w}_{jt} \, \Big[ d_{jt}^{-1}\left(\boldsymbol{s}_{t}^{\boldsymbol{a}_{t}},\boldsymbol{\sigma}\right) - \alpha \, p_{jt} - \boldsymbol{x}_{jt}^{\prime}\boldsymbol{\beta} - \boldsymbol{r}_{jt}^{\prime}\widetilde{\boldsymbol\mu}_{j} \Big] \;\middle|\; a_{jt} = 1 \right). \end{equation*} Suppose that the following conditions hold. \begin{itemize} • The objects $\Big\{ \widetilde f_\ell, \widetilde P_{j,\ell}(\boldsymbol{x}_{t}) \Big\}_{\ell=1}^{L}$ are identified from the first step. • The instruments are valid, so that $\mathbb{E}\left( \boldsymbol{w}_{jt} \, \widetilde{\xi}_{jt} \mid a_{jt} = 1 \right) = \boldsymbol{0}$. • The function $d_{jt}^{-1}\left(\boldsymbol{s}_{t}^{\boldsymbol{a}_{t}},\boldsymbol{\sigma}\right)$ is continuously differentiable in $\boldsymbol{\sigma}$, and the Jacobian matrix $\displaystyle \boldsymbol{M}_{j}(\boldsymbol\vartheta_{j0}) \equiv \left. \frac{\partial \boldsymbol{m}_{j}(\boldsymbol\vartheta_{j})} {\partial \boldsymbol\vartheta_{j}^{\prime}} \right|_{\boldsymbol\vartheta_{j}=\boldsymbol\vartheta_{j0}}$ has full column rank at the true parameter value $\boldsymbol\vartheta_{j0}$. \end{itemize} Then, $\boldsymbol\vartheta_{j0}$ is locally identified. \qquad $\blacksquare$ \end{prop} \subsubsection{Identification of marginal costs and fixed entry costs} As is standard in models of competition, once demand parameters are identified and a form of competition is specified (e.g., Bertrand competition), equilibrium price-cost margins $PCM_{jt}$ can be recovered from firms’ profit maximization conditions. Combining these with observed prices yields realized marginal costs $mc_{jt}$. Estimating a marginal cost function---from the regression of $mc_{jt}$ on observable product characteristics $\boldsymbol{x}_{jt}$ with an additive unobservable $\omega_{jt}$---raises the same selection bias issues as in demand estimation. In particular, the regression involves the selection-bias function $\lambda^{mc}_{j}(\boldsymbol{x}_{t}) = \mathbb{E}( \omega_{jt} \mid \boldsymbol{x}_{t}, a_{jt} = 1 )$. Fortunately, this selection-bias function has the same structure as in Proposition (ref), though with different parameters: \begin{equation} \lambda^{mc}_{j}(\boldsymbol{x}_{t}) \; = \; \int \left[ \frac{P_{j}\left( \boldsymbol{x}_{t}, \boldsymbol{\kappa} \right)}{\overline{P}_j\left( \boldsymbol{x}_{t} \right)} \right] \, \mu^{mc}_{j}(\boldsymbol{\kappa}) \, f_{\kappa}(\boldsymbol{\kappa}) \, d \boldsymbol{\kappa}. \end{equation} with $\mu^{mc}_{j}(\boldsymbol{\kappa}) \equiv \mathbb{E}( \omega_{jt} \mid \boldsymbol{\kappa}_{t} = \boldsymbol{\kappa} )$. Therefore, the same identification strategy can be applied to recover both the parameters in the marginal cost function and in the selection-bias function. Given estimates of the demand and marginal cost parameters, together with the selection-bias parameters $\widetilde{\boldsymbol{\mu}}$ and $\widetilde{\boldsymbol{\mu}}^{mc}$, we can compute equilibrium variable profits for any given $\boldsymbol{x}_t$, latent market type $\ell = 1, \ldots, L$, and counterfactual entry configuration $\boldsymbol{a} \in \{0,1\}^{J}$. A key feature of our approach is that we estimate the expected unobserved demand and marginal costs, $\widetilde{\mu}_{j\ell}$ and $\widetilde{\mu}^{mc}_{j\ell}$, for every product $j$ and latent type $\ell$, regardless of whether the product is observed in the market. This allows us to compute counterfactual equilibrium outcomes for all products. In particular, for a given $\boldsymbol{x}_t$ and latent type $\ell$, we can obtain equilibrium price-cost margins, market shares, and variable profits by setting $\xi_j = \widetilde{\mu}_{j\ell}$ and $\omega_j = \widetilde{\mu}^{mc}_{j \ell}$ for all $j \in \mathcal{J}$. Let $VP_{j \ell}(\boldsymbol{a}, \boldsymbol{x}_t)$ denote the equilibrium variable profit of firm $j$ evaluated at latent market type $\ell$ under entry configuration $\boldsymbol{a} \in \{0,1\}^{J}$. Using firms’ entry probabilities, expected variable profits at the time of entry decisions can be written as: \begin{equation} VP^{P}_{j \ell}(\boldsymbol{x}_t) \; = \; \sum_{\boldsymbol{a}_{-j} \in \{0,1\}^{J-1}} VP_{j \ell}(a_j=1, \boldsymbol{a}_{-j}, \boldsymbol{x}_t) \, \prod_{i \neq j} \widetilde{P}_{i,\ell}(\boldsymbol{x}_t)^{a_{i}} \, \left[ 1 - \widetilde{P}_{i,\ell}(\boldsymbol{x}_t) \right]^{1-a_{i}}. \end{equation} Up to this point, the private information entering fixed costs has been allowed to be the general vector $\boldsymbol{\eta}_{jt}$. To recover fixed costs from equilibrium entry probabilities, we now impose the additional restriction that this private information can be summarized by a scalar shock $\eta_{jt}$ entering additively in fixed costs. Specifically, suppose fixed costs take the form $fc_{j \ell}(\boldsymbol{x}_{jt}) + \sigma_{\eta_j}\eta_{jt}$, where $\eta_{jt}$ has mean zero and known strictly increasing CDF $F_{\eta}$. Then, the equilibrium entry probabilities satisfy $\widetilde{P}_{j,\ell}(\boldsymbol{x}_t) = F_{\eta}\left( \left[ VP^{P}_{j \ell}(\boldsymbol{x}_t) - fc_{j \ell} (\boldsymbol{x}_{jt}) \right] / \sigma_{\eta_j} \right)$, and can therefore be inverted to obtain the regression-like equation: \begin{equation} F_{\eta}^{-1} \left( \widetilde{P}_{j,\ell}(\boldsymbol{x}_t) \right) \; = \; \frac{1}{\sigma_{\eta_j}} VP^{P}_{j \ell}(\boldsymbol{x}_t) \; - \; \frac{1}{\sigma_{\eta_j}} fc_{j \ell} (\boldsymbol{x}_{jt}) \end{equation} Under the exclusion restriction that firm $j$’s fixed cost depends only on its own characteristics $\boldsymbol{x}_{jt}$ (and not on those of other firms), this equation identifies both the scale parameter $\sigma_{\eta_j}$ and the fixed cost function $fc_{j \ell}(\boldsymbol{x}_{jt})$ for each product and latent market type. \subsection{Estimation } In this section, we present a two-step estimation method that mimics our two-step sieve identification approach. In the first step, for a given number of market types $L$, we use a nonparametric sieve maximum likelihood method to estimate the distribution of unobserved market types, and the vector of entry probabilities for each unobserved type. In the second step, we construct the control variables $r$ and apply the Generalized Method of Moments (GMM) to estimate demand parameters and control function parameters. Estimating the number of latent market types $L$ warrants special attention, and we discuss it in Section (ref). \subsubsection{First step: Estimation of conditional choice probabilities (CCPs) and distribution of latent types} For a given number of latent types $L$, we approximate the nonparametric functions $\{ \widetilde{P}_{j \ell}(\boldsymbol{x}_{t}): \forall (j, \ell) \}$ using sieves as functions of $\boldsymbol{x}_{t}$ (hirano2003efficient, hirano2003efficient, chen_2007, chen_2007). For estimation, we specialize the general CDF $F_{\eta}$ above to the Logistic case, so that $F_{\eta} = \Lambda$. Let $\boldsymbol{b}_{t} \equiv$ $( b_{1}(\boldsymbol{x}_{t}), \dots,$ $b_{N_{X}}(\boldsymbol{x}_{t}) )^{\prime}$ be a vector with a finite number $N_{X}$ of basis functions. For any product $j$ and any latent type $\ell$, the entry probability function $\widetilde{P}_{j \ell}(\boldsymbol{x}_{t})$ has the following sieves binary logit structure: \begin{equation} \widetilde{P}_{j \ell}(\boldsymbol{x}_{t}) \; = \; \Lambda \left( \boldsymbol{b}^{\prime}_{t} \; \boldsymbol{\gamma}_{j \ell} \right), \end{equation} where $\Lambda(\cdot)$ is the logistic function, and $\boldsymbol{\gamma}_{j \ell}$ is a vector of parameters of dimension $N_{X} \times 1$. Let $\boldsymbol{\gamma} \equiv \{ \boldsymbol{\gamma}_{j \ell}: \forall (j, \ell)\}$ be the vector of $J \, L \, N_{X}$ parameters in the sieve approximation to the entry probabilities. And let $\widetilde{\boldsymbol{f}} \equiv \{ \widetilde{f}_{\ell}: \forall \ell\}$ be the vector of probabilities for the latent types. The log-likelihood function of the finite mixture model is: \begin{equation} \ln \mathcal{L}_{1st}(\widetilde{\boldsymbol{f}}, \boldsymbol{\gamma}) \; = \; \sum_{t=1}^{T} \ln \left( \sum_{\ell=1}^{L} \widetilde{f}_{\ell} \; \prod_{j=1}^{J} \Lambda \left( \boldsymbol{b}^{\prime}_{t} \; \boldsymbol{\gamma}_{j \ell} \right) ^{a_{jt}} \left[ 1 - \Lambda \left( \boldsymbol{b}^{\prime}_{t} \; \boldsymbol{\gamma}_{j \ell} \right) \right]^{1-a_{jt}} \right). \end{equation} We estimate $(\widetilde{\boldsymbol{f}}, \boldsymbol{\gamma})$ by maximum likelihood estimation (MLE) using the Expectation-Maximization (EM) algorithm (pilla_lindsay_2001, pilla_lindsay_2001).\footnote{Recent applications of MLE--EM methods to nonparametric mixture models in discrete choice settings include bunting_2022 (bunting_2022), bunting_diegert_2022 (bunting_diegert_2022), hu_xin_2022 (hu_xin_2022), and williams_2020 (williams_2020).} The sieve is constructed using polynomial basis functions in $\boldsymbol{x}_{t}$, $\{ b_{n}(\boldsymbol{x}_{t}): n=1, 2, \dots N_X\}$. We select the dimension of this basis using the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC), thereby balancing approximation flexibility against overfitting. Section (ref) describes in detail the separate selection of the number of latent types $L$. When $\boldsymbol{x}_{t}$ has discrete support, the nonparametric MLE achieves $\sqrt{T}$-consistency and asymptotic normality. With continuous covariates, the first-step nonparametric estimator converges at a slower rate. Crucially, however, this slower convergence does not compromise inference on the demand parameters in the second step. Under standard regularity conditions, the second-step estimator remains $\sqrt{JT}$-consistent and asymptotically normal. This result follows from the general semiparametric efficiency arguments in hirano2003efficient (hirano2003efficient) and das2003 (das2003), among others. \subsubsection{Second step: Estimation of demand parameters} Given estimates from the first step, $(\widehat{\widetilde{\boldsymbol{f}}}, \widehat{\boldsymbol{\gamma}})$, the estimation of demand parameters in the second step is based on applying standard GMM to the regression demand equations augmented by the linear in parameters control function, as presented in equation (ref). By definition, the constructed regressors for the control function are: \begin{equation} \widehat{r}_{j \ell t} \; = \; \frac{ \widehat{\widetilde{P}}_{j,\ell}(\boldsymbol{x}_t) - \widehat{\widetilde{P}}_{j,L}(\boldsymbol{x}_t) } {\widehat{\overline{P}}_{j}(\boldsymbol{x}_t)} \, \widehat{\widetilde{f}}_{\ell} \end{equation} Following das2003 (das2003), this two-step estimator of demand parameters is $\sqrt{JT}$-consistent and asymptotically normal. However, given the sequential nature of the estimator, the standard errors of the estimates in the second step need to be corrected. The correct standard errors can be computed using either the asymptotic approximations and formulas in newey_2009 (newey_2009) or the linearized bootstrap procedure we detail in Appendix (ref). Conditional on the selected first-step sieve specification, a key computational advantage of this bootstrap procedure is that it does not require repeated estimation of the first step. \subsubsection{Estimating the number of latent types } $L$ is a tuning parameter for approximating the selection-bias term well enough to identify and estimate demand parameters consistently. For this reason, the main criterion to select $L$ should be based on the demand equation and not so much on the goodness-of-fit of the joint entry distribution in the first step of the method. A value of $L$ that is nearly optimal for approximating the entry distribution need not be optimal for approximating the control function. Specifically, a criterion that gives substantial weight to the goodness-of-fit of the entry distribution might select a value of $L$ that is too small for the optimal approximation of the selection-bias function. This can easily happen when increasing $L$ produces only a modest improvement in the likelihood of the entry model, but generates new directions of variation in the control variables $r$ that matter a lot for the second-step regression and for the estimates of demand parameters. Therefore, selecting $L$ using only the first-step fit may be misaligned with our ultimate objective. That said, the goodness-of-fit in the first step should not be completely ignored when selecting $L$. The first step imposes a feasibility constraint: $L$ must be small enough that the finite-mixture model is identified and can be estimated reliably. For these reasons, we combine two BIC criteria to select $L$. The first-step criterion is a likelihood-based BIC from the estimation of the entry distribution mixture model: \begin{equation} \mathrm{BIC}_{\text{1st}}(L) \; = \; - 2 \, \ln \mathcal{L}_{\text{1st}}(\widehat{\widetilde{\boldsymbol{f}}}, \widehat{\boldsymbol{\gamma}}, L) + (J \, L \, N_{X} + L -1) \, \ln(T) \end{equation} where $\mathcal{L}_{\text{1st}}(\widehat{\widetilde{\boldsymbol{f}}}, \widehat{\boldsymbol{\gamma}}, L)$ is the likelihood function in the first-step estimation, and $(J \, L \, N_{X} + L -1)$ is the number of parameters. The second-step criterion is a BIC based on the demand equation residuals: \begin{equation} \mathrm{BIC}_{\text{2nd}}(L) \; = \; (T J) \, \ln \widehat{\sigma}^2_{\xi}(L) + (L-1)J \, \ln(TJ) \end{equation} where $\widehat{\sigma}^2_{\xi}(L)$ is the sample variance of the residuals in the estimation of the system of demand equations. Note that $(L-1)J$ represents the number of parameters in the control function. We use the minimization of $\mathrm{BIC}_{\text{2nd}}(L)$ as our selection criterion, and use the first-step criterion $\mathrm{BIC}_{\text{1st}}(L)$ only as a constraint and diagnostic. \section{Empirical application } \subsection{Data and descriptive statistics} We apply our method to estimate demand in the US airline industry. The challenge of endogenous product entry in demand estimation in this industry has recently been explored by ciliberto2021market (ciliberto2021market) and li2022repositioning (li2022repositioning). \textit{\textbf{Data sources.}} We use publicly available data from the US Department of Transportation for our analysis. Our working sample consists of the DB1B and T100 datasets. Specifically, we use quarterly data spanning 2012-Q1 to 2013-Q4 for routes between the airports at the 100 largest Metropolitan Statistical Areas (MSA) in the United States. These account for 108 airports, as there are a few MSAs with more than one airport. \textit{\textbf{Airlines.}} The airlines included in our analysis are American (AA), Delta (DL), United (UA), US Airways (US), Southwest (WN), a combined group of Low-Cost Carriers (LCC), and a combined group of the remaining carriers (Others).\footnote{Following ciliberto2021market (ciliberto2021market), the list of airlines included in the group LCC is: Alaska, JetBlue, Frontier, Allegiant, Spirit, Sun Country, and Virgin. The carriers in the group Others are small regional carriers, charters, and private jets.} Given the large number of carriers included in Others, we do not consider this combined group as a player in the entry game. In the notation of our model, $j$ always indexes an airline. \textit{\textbf{Markets in the demand model.}} In the demand model, a market $t$ is defined as a \textit{directional airport pair} in a given quarter. For example, LGA$\to$ORD in Q1 2012 and ORD$\to$LGA in Q1 2012 are two distinct demand markets. In each demand market $t$, consumers choose among the airlines offering non-stop flights on that directional route (up to seven airlines, plus the outside option). Thus, $s_{jt}$ denotes the market share of airline $j$ on directional route $t$. \textit{\textbf{Markets in the entry model.}} In the entry model, a market $t$ is defined as a \textit{non-directional airport pair} in a given quarter, where, for example, Chicago O'Hare (ORD) to New York La Guardia (LGA) is the same market as LGA to ORD. Each non-directional entry market thus corresponds to two directional demand markets. There are potentially 5,778 non-directional markets between the 108 airports, i.e., $108\times107/2$. However, many of these markets have not had an incumbent airline with non-stop flights for several decades. These are typically airport pairs that are geographically too close or in smaller MSAs. In our sample, we only consider non-directional markets which were served in at least 50 quarters between 1994 and 2018. This results in 2,230 non-directional markets and 17,155 market-quarter observations.\footnote{Given 2,230 non-directional markets and eight quarters, the total number of market-quarter observations in our sample is $2,230 \times 8 = 17,840$. We however discard from the analysis 685 market-quarter observations for which we either do not observe some of the regressors or none of the six airlines included in the entry model is a potential entrant.} The estimated control-function variables $\widehat{\boldsymbol{r}}_{jt}$ entering the demand equation for a given directional market are constructed from the entry probabilities estimated at the corresponding non-directional market level. \textit{\textbf{Potential entrants.}} We consider an airline a potential entrant in a non-directional airport pair in a given quarter if it operates non-stop flights from either origin or destination airport (toward or from any airport), while an airline is an \textit{entrant} in a non-directional airport pair in a given quarter if it operates non-stop flights between the origin and destination airports. \textit{\textbf{Market size.}} Following the empirical literature on the airline industry, we define market size as the geometric mean of the populations in the metropolitan areas (MSAs) of the two airports and market distance as the geodesic distance between the two airports. \textit{\textbf{Observable variables in $\boldsymbol{x}_{jt}$.}} The vector of exogenous observable variables at the airline-market level includes: market size, market distance, squared market distance, the airline's own hub-size in the origin airport, its hub-size in the destination airport, airline indicators, and quarter indicators. We define the hub-size of an airline in an airport as the number of non-stop routes that the airline operates from that airport. Table (ref) presents the distribution of the number of entrants and averages of the market characteristics. Notably, in a significant portion of these markets (almost $30\%$), there are no airlines providing non-stop flights, and they are exclusively served with stop flights. Among the markets with non-stop flights, more than $90\%$ are monopolies or duopolies. Furthermore, there is a strong positive correlation between the number of incumbents, market size, and distance. \begin{table}[ht] \caption{Distribution of Markets by Number of Entrants } \begin{tabular}{r|ccc} \hline \hline & Frequency & Avg. market size & Avg. market distance \\ Number of airlines & \# markets-quarters (%) & in millions of people & in miles \\ \hline & & & \\ 0 airlines & 5,117 (29.83%) & 7.09 & 734 \\ 1 airline & 8,217 (47.90%) & 8.82 & 913 \\ 2 airlines & 2,637 (15.37%) & 10.95 & 960 \\ 3 airlines & 869 (5.07%) & 13.00 & 1,117 \\ 4 airlines & 233 (1.36%) & 12.60 & 1,140 \\ 5 airlines & 72 (0.42%) & 20.16 & 1,255 \\ $\geq$ 6 airlines & 10 (0.06%) & 17.54 & 320 \\ & & & \\ \hline Total & 17,155 (100.00%) & 8.95 & 882 \\ \hline \hline \end{tabular} \end{table} Table (ref) presents entry frequencies for each airline and the average market size and distance associated with their entry. We observe significant variation in airlines' entry probabilities, with WN and AA having the highest (27.5%) and the lowest (10.6%) entry probabilities, respectively. Furthermore, there is substantial heterogeneity in the correlations between entry, market size, and distance among airlines. For example, while WN enters markets that are not significantly different in size from the markets it does not enter (8.7 million people versus 9 million people), AA tends to enter markets with much larger average population (13.3 million people versus 8.4 million people). Different entry strategies are also evident on the basis of market distance. DL and US typically enter markets with an average distance of around 875--950 miles, whereas the markets served by LCC have an average distance of 1,171 miles. \begin{table}[ht] \caption{Entry Frequency by Airline } \begin{tabular}{r|ccc} \hline \hline & Frequency & Avg. market size & Avg. market distance \\ Airline & \# markets-quarters (%) & in millions of people & in miles \\ \hline & & & \\ WN & 4,714 (27.48%) & 8.71 & 989 \\ DL & 3,285 (19.15%) & 10.68 & 875 \\ UA & 3,244 (18.91%) & 11.56 & 968 \\ LCC & 2,386 (13.91%) & 11.42 & 1,171 \\ US & 2,001 (11.66%) & 9.52 & 894 \\ AA & 1,820 (10.61%) & 13.28 & 965 \\ & & & \\ \hline \hline \end{tabular} \end{table} \subsection{First step: Estimation of the model of market entry} For the entry decisions, we consider the nonparametric sieve finite mixture Logit described in equation (ref). We explore various specifications of the mixture Logit model based on the polynomial order in $\boldsymbol{x}_{t}$ used to construct the basis $\boldsymbol{b}_{t}$ and the number of latent market types $L$. As our estimates of the demand parameters are robust to the selection of the basis $\boldsymbol{b}_{t}$ in the entry model, we only present results here for the specification with $\boldsymbol{b}_{t} = \boldsymbol{x}_{t}$. Regarding the number of latent types $L$, Table (ref) presents the goodness-of-fit statistics obtained from estimating four nested specifications of the mixture Logit model. The goodness-of-fit is guided by the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC), along with the convergence performance of the EM algorithm, the accuracy of the parameter estimates of the entry model, and the robustness of the estimates of the demand model. The introduction of latent market types improves the entry model's goodness-of-fit. Comparing the specification without $\boldsymbol{\kappa}_{t}$ and the one with two unobserved market types in Table (ref), we see a substantial increase in the log-likelihood and a decrease in both AIC and BIC. This form of unobserved market heterogeneity captures a strong correlation among airline entry decisions, a correlation not captured by the observable market and airline characteristics in $\boldsymbol{x}_{t}$. The inclusion of additional unobserved market types continues to positively impact the goodness-of-fit. However, this improvement has diminishing returns and it is very small when moving from three to four unobserved market types. While the EM algorithm converges rapidly to the MLE in the specifications with two and three unobserved market types, we experience convergence issues in the specification with four unobserved market types. In this case, we obtain imprecise estimates for some of the parameters of the entry model. These considerations, combined with the marginal improvement observed in the AIC and BIC criteria, lead us to favor the specification with $L = 3$. Moreover, this choice is also motivated by the implied estimates of the demand model. As we illustrate below, the estimated own-price elasticities of demand with $L = 3$ and $L = 4$ are practically indistinguishable. In contrast, with $L \leq 2$ we obtain estimates of the own-price elasticities which are substantially smaller. \begin{table}[ht] \caption{Estimation of Market Entry Model---Goodness-of-Fit Statistics} \begin{tabular}{lcccc} \hline \hline & {Logit} & {Mixture Logit} & {Mixture Logit} & {Mixture Logit} \\ Statistics & {\# types = 1} & {\# types = 2} & {\# types = 3} & {\# types = 4} \\ \hline & & & & \\ Observations & $17,155$ & $17,155$ & $17,155$ & $17,155$ \\ Parameters & $72$ & $145$ & $218$ & $287$ \\ Log-likelihood & $-20,378$ & $-18,985$ & $-18,022$ & $-17,621$ \\ AIC & $40,900$ & $38,261$ & $36,481$ & $35,817$ \\ BIC & $41,458$ & $39,385$ & $38,170$ & $38,041$ \\ & & & & \\ \hline \hline \end{tabular} \begin{tablenotes} \scriptsize\end{tablenotes} \end{table} \subsection{Estimation of demand parameters} For the demand system, we follow ciliberto2021market (ciliberto2021market) and estimate a nested logit model with two nests: a nest for all the airlines and another nest for the outside option. \begin{equation} \ln\left(\frac{s_{jt}}{s_{0t}}\right) \; = \; \alpha\; p_{jt} + \boldsymbol{x}_{jt}^{\prime}\;\boldsymbol{\beta} + \sigma \; \ln \left( \frac{s_{jt}}{1-s_{0t}} \right) + \widehat{\boldsymbol{r}}_{jt}^{\prime} \; \widetilde{\boldsymbol{\mu}}_{j} + \widetilde{\xi}_{jt}. \end{equation} We compute each directional route-specific market share in a given quarter $s_{jt}$ as the total number of passengers who traveled that directional route with a non-stop flight of a specific airline in that given quarter (times 10, as the data are a survey of 10% of total traffic) divided by market size. The vector of characteristics $\boldsymbol{x}_{jt}$ includes market distance and market distance squared, airline $j$'s hub-size in the origin airport, airline $j$'s hub-size in the destination airport, and airline $\times$ quarter fixed effects (indicators). The expression for the selection term, $\widehat{\boldsymbol{r}}_{jt}^{\prime} \; \widetilde{\boldsymbol{\mu}}_{j}$, varies with the specification of the market entry model, from the more restrictive parametric Logit model to the more general semiparametric finite mixture Logit model. \begin{itemize} • \textit{Parametric Logit without latent types.} We consider the entry model $a_{jt} = \mathbb{1}\{ \eta_{jt} \leq \boldsymbol{x}_{jt}^{\prime} \boldsymbol{\gamma}_{j}^{P} \}$, with $\eta_{jt} \sim Logistic$, and $\xi_{jt} = \widetilde{\mu}_{j} \; \eta_{jt} + v_{jt}$, with $v_{jt}$ independent of $\eta_{jt}$ and $\boldsymbol{x}_{t}$. In this parametric specification, the selection term is given by the expected value of a truncated Logistic variable, which can be interpreted as the Logit analogue of the Heckman selection correction term: \begin{equation} \mathbb{E}\left( \xi_{jt} \mid a_{jt}=1, \boldsymbol{x}_{jt} \right) \; = \; \widetilde{\mu}_{j} \; \mathbb{E}\left( \eta_{jt} \mid \eta_{jt} \leq \boldsymbol{x}_{jt}^{\prime} \boldsymbol{\gamma}_{j}^{P} \right) \; = \; \widetilde{\mu}_{j} \; m_{j}(\boldsymbol{x}_{jt}), \end{equation} where $m_{j}(\boldsymbol{x}_{jt})$ is the expectation of a truncated Logistic in terms of the truncation probability: \begin{equation} m_{j}(\boldsymbol{x}_{jt}) \equiv \ln \overline{P}_{j}(\boldsymbol{x}_{jt}) + \frac{ 1-\overline{P}_{j}(\boldsymbol{x}_{jt}) }{ \overline{P}_{j}(\boldsymbol{x}_{jt}) } \ln \left( 1-\overline{P}_{j}(\boldsymbol{x}_{jt}) \right), \end{equation} and $\overline{P}_{j}(\boldsymbol{x}_{jt}) \equiv \Lambda \left( \boldsymbol{x}_{jt}^{\prime} \boldsymbol{\gamma}_{j}^{P} \right)$ is the propensity score implied by the Logit entry model. • \textit{Semiparametric without latent types.} The entry model is still the Logit $a_{jt} = \mathbb{1}\{ \eta_{jt} \leq \boldsymbol{x}_{jt}^{\prime} \boldsymbol{\gamma}_{j}^{P} \}$, with $\eta_{jt} \sim Logistic$, but now $\mathbb{E}\left( \xi_{jt} \; \vert \; a_{jt}=1, \boldsymbol{x}_{jt} \right)$ is approximated by a third order polynomial in the scalar control variable $m_{j}(\boldsymbol{x}_{jt})$. This is a one-dimensional control variable because, under the single-index restriction without latent types, the selection term depends on observables only through the propensity score.\footnote{Both in this and in the case of the semiparametric mixture Logit, estimates are very similar by approximating the selection function with polynomials of higher orders.} This semiparametric approach to control for selection follows the standard series-based control-function strategy in newey_2009 (newey_2009). • \textit{Finite mixture Logit with latent types.} This is our entry model described above, with $\widehat{\boldsymbol{r}}^{\prime}_{jt} = (\widehat{r}_{j 1 t}, \widehat{r}_{j 2 t}, \dots, \widehat{r}_{j, L-1, t})$ constructed from the first-step estimates. \end{itemize} For all the two-stage least squares (2SLS) estimators, we use as instruments the number of competitors in the market and the average hub-size of the rest of the airlines, separately for origin and destination. We compute standard errors using the linearized bootstrap procedure detailed in Appendix (ref). Table (ref) presents the estimates of the demand parameters, while Table (ref) reports the average demand elasticities and Lerner indexes derived from these estimates. Comparing the estimates obtained using ordinary least squares (OLS) with those from the standard 2SLS method---not accounting for potential selection bias---we observe a significant change in all parameter estimates when addressing the endogeneity of prices and within-nest market shares. Controlling for endogeneity meaningfully affects the average estimated own-price elasticity, which decreases from $-1.60$ to $-5.55$, and the corresponding average Lerner index, which decreases from $68.8\%$ to $19.9\%$. Turning to the consequences of controlling for endogenous market entry, we note the important role played by finite mixture unobserved heterogeneity. The estimates of parameters $\alpha$ and $\sigma$ of a finite mixture model with $L=3$ are, compared to those of “Semiparametric” (assuming $L=1$), $15.9\%$ and $28.8\%$ higher (in absolute terms). These changes translate into an increase in the average estimated own-price elasticities of around $30\%$. Consequently, the corresponding average estimated Lerner index decreases from $18.9\%$ to $15.1\%$. These effects are of substantial importance and lead to meaningful economic implications. Parameter estimates and implied own-price elasticities of the standard 2SLS (not controlling for selection) and those of “Heckman” or “Semiparametric” (assuming $L=1$) are relatively similar. In contrast, parameter estimates and corresponding own-price elasticities remarkably change when we allow $L>1$. Although the estimated own-price elasticities of a model with $L=2$ are still meaningfully different from those of a model with $L=3$, the estimates implied by models with $L=3$ and $L=4$ are essentially indistinguishable. Collectively, these results stress the importance of allowing for “some” unobserved market heterogeneity to effectively control for endogenous selection, but also that as few unobserved market types as three may already be sufficient. \begin{table}[ht] \caption{Estimation of Demand Parameters } \resizebox{0.9\textwidth}{!}{ \begin{tabular}{l|cc|ccccc} \hline \hline & & & & & & & \\ & \multicolumn{2}{c|}{\textit{Not control. for sel.}} & \multicolumn{5}{c}{\textit{Controlling for endogenous selection}} \\ & {OLS} & {2SLS} & {2SLS} & {2SLS} & {2SLS} & {2SLS} & {2SLS} \\ & & & {Heckman} & {Semipar.} & {Fin.-Mix.} & {Fin.-Mix.} & {Fin.-Mix.} \\ & & & $L=1$ &$L=1$ & $L=2$ & $L=3$ & $L=4$ \\ \hline & & & & & & & \\ Price (100\$) ($\alpha$) & $-0.643$ & $-2.180$ & $-2.193$ & $-2.261$ & $-2.392$ & $-2.621$ & $-2.697$ \\ & $(0.0105)$ & $(0.1378)$ & $(0.2065)$ & $(0.2077)$ & $(0.2201)$ & $(0.2448)$ & $(0.2716)$\\ & & & & & & & \\ Within Share ($\sigma$) & $0.371$ & $0.409$ & $0.413$ & $0.431$ & $0.494$ & $0.555$ & $0.546$ \\ & $(0.0058)$ & $(0.0351)$ & $(0.0529)$ & $(0.0559)$ & $(0.0622)$ & $(0.0717)$ & $(0.0821)$ \\ & & & & & & & \\ Distance (1000mi) & $0.729$ & $2.130$ & $2.196$ & $2.264$ & $2.387$ & $2.503$ & $2.624$ \\ & $(0.0306)$ & $(0.1372)$ & $(0.2074)$ & $(0.2055)$ & $(0.2133)$ & $(0.2390)$ & $(0.2648)$ \\ & & & & & & & \\ Distance$^2$ & $-0.216$ & $-0.424$ & $-0.453$ & $-0.462$ & $-0.493$ & $-0.525$ & $-0.502$ \\ & $(0.0112)$ & $(0.0244)$ & $(0.0398)$ & $(0.0392)$ & $(0.0401)$ & $(0.0440)$ & $(0.0483)$ \\ & & & & & & & \\ hub-size orig. (100s) & $1.637$ & $2.272$ & $1.999$ & $1.320$ & $1.709$ & $1.677$ & $1.444$\\ & $(0.0263)$ & $(0.0382)$ & $(0.0767)$ & $(0.0919)$ & $(0.1085)$ & $(0.1206)$ & $(0.1244)$ \\ & & & & & & & \\ hub-size dest. (100s) & $1.613$ & $2.242$ & $1.995$ & $1.310$ & $1.703$ & $1.674$ & $1.436$ \\ & $(0.0267)$ & $(0.0385)$ & $(0.0784)$ & $(0.0933)$ & $(0.1106)$ & $(0.1228)$ & $(0.1266)$ \\ & & & & & & & \\ \hline & & & & & & & \\ Airline$\times$Quarter FE & Y & Y & Y & Y & Y & Y & Y \\ \# control var. entry & 0 & 0 & 6 & 18 & 36 & 54 & 72 \\ Observations & $35,763$ & $35,763$ & $35,763$ & $35,763$ & $35,763$ & $35,763$ & $35,763$ \\ & & & & & & & \\ \hline \hline \end{tabular}} \begin{tablenotes} • \scriptsize{Bootstrap standard errors are based on the linearized bootstrap procedure in Appendix (ref) and account for first-step estimation error conditional on the selected first-step specification.} \end{tablenotes} \end{table} \begin{table}[ht] \caption{Average Own-Price Elasticities and Lerner Indexes } \resizebox{0.9\textwidth}{!}{ \begin{tabular}{l|cc|ccccc} \hline \hline & & & & & & & \\ & \multicolumn{2}{c|}{\textit{Not control. for sel.}} & \multicolumn{5}{c}{\textit{Controlling for endogenous selection}} \\ & {OLS} & {2SLS} & {2SLS} & {2SLS} & {2SLS} & {2SLS} & {2SLS} \\ & & & {Heckman} & {Semipar.} & {Fin.-Mix.} & {Fin.-Mix.} & {Fin.-Mix.} \\ & & & $L=1$ & $L=1$ & $L=2$ & $L=3$ & $L=4$ \\ \hline & & & & & & & \\ \multicolumn{1}{r|}{\textit{Own-Price Elasticity}} & $-1.596$ & $-5.549$ & $-5.601$ & $-5.849$ & $-6.524$ & $-7.605$ & $-7.746$ \\ & & & & & & & \\ \multicolumn{1}{r|}{\textit{AA}} & $-1.722$ & $-6.013$ & $-6.071$ & $-6.363$ & $-7.143$ & $-8.399$ & $-8.543$ \\ \multicolumn{1}{r|}{\textit{DL}} & $-1.761$ & $-6.082$ & $-6.133$ & $-6.382$ & $-7.024$ & $-8.067$ & $-8.236$ \\ \multicolumn{1}{r|}{\textit{UA}} & $-1.887$ & $-6.573$ & $-6.636$ & $-6.936$ & $-7.766$ & $-9.090$ & $-9.253$ \\ \multicolumn{1}{r|}{\textit{US}} & $-1.665$ & $-5.801$ & $-5.856$ & $-6.122$ & $-6.854$ & $-8.023$ & $-8.167$ \\ \multicolumn{1}{r|}{\textit{WN}} & $-1.354$ & $-4.680$ & $-4.719$ & $-4.913$ & $-5.411$ & $-6.220$ & $-6.350$ \\ \multicolumn{1}{r|}{\textit{LCC}} & $-1.370$ & $-4.808$ & $-4.857$ & $-5.095$ & $-5.784$ & $-6.870$ & $-6.977$ \\ \multicolumn{1}{r|}{\textit{Others}} & $-1.332$ & $-4.705$ & $-4.757$ & $-5.006$ & $-5.750$ & $-6.915$ & $-7.009$ \\ & & & & & & & \\ \multicolumn{1}{r|}{\textit{Lerner Index}} & $68.8\%$ & $19.9\%$ & $19.7\%$ & $18.9\%$ & $17.2\%$ & $15.1\%$ & $14.7\%$ \\ & & & & & & & \\ \multicolumn{1}{r|}{\textit{AA}} & $62.7\%$ & $18.0\%$ & $17.9\%$ & $17.1\%$ & $15.5\%$ & $13.5\%$ & $13.2\%$ \\ \multicolumn{1}{r|}{\textit{DL}} & $60.4\%$ & $17.5\%$ & $17.3\%$ & $16.7\%$ & $15.3\%$ & $13.4\%$ & $13.1\%$ \\ \multicolumn{1}{r|}{\textit{UA}} & $56.9\%$ & $16.4\%$ & $16.2\%$ & $15.6\%$ & $14.1\%$ & $12.3\%$ & $12.1\%$ \\ \multicolumn{1}{r|}{\textit{US}} & $65.9\%$ & $19.0\%$ & $18.9\%$ & $18.1\%$ & $16.5\%$ & $14.5\%$ & $14.2\%$ \\ \multicolumn{1}{r|}{\textit{WN}} & $78.4\%$ & $22.8\%$ & $22.6\%$ & $21.8\%$ & $20.1\%$ & $17.8\%$ & $17.4\%$ \\ \multicolumn{1}{r|}{\textit{LCC}} & $82.1\%$ & $23.5\%$ & $23.3\%$ & $22.2\%$ & $19.9\%$ & $17.1\%$ & $16.8\%$ \\ \multicolumn{1}{r|}{\textit{Others}} & $79.2\%$ & $22.5\%$ & $22.3\%$ & $21.3\%$ & $18.9\%$ & $16.0\%$ & $15.8\%$ \\ & & & & & & & \\ \hline Observations & $35,763$ & $35,763$ & $35,763$ & $35,763$ & $35,763$ & $35,763$ & $35,763$ \\ \hline \hline \end{tabular}} \begin{tablenotes} • \scriptsize \end{tablenotes} \end{table} Figure (ref) plots the empirical distributions of the estimated own-price elasticities. Each row corresponds to an airline, while each column to a different 2SLS estimator: the first column plots results for the estimator that does not control for selection, the second column plots results for the estimator that controls for selection using a sieve method but no mixture, and the third column plots results for the estimator with three unobserved market types. The histograms in this figure are constructed based on estimates of own-price elasticities at the airline-market-quarter level. The equation describing each own-price elasticity only depends on data on price $p_{jt}$, market shares $s_{jt}$ and $s_{0t}$, and parameter estimates $\widehat{\alpha}$ and $\widehat{\sigma}$. It is important to note that the data regarding prices and market shares remain constant across the various columns in the figure. Therefore, any change in empirical distributions can only be attributed to changes in the values of the estimates $\widehat{\alpha}$ and $\widehat{\sigma}$ across the different estimators. \begin{figure}[ht] \caption{Distribution of Estimated Own-Price Elasticities (Airline-Market-Quarter level) } \end{figure} The empirical distributions in the first two columns of Figure (ref) are very similar. In contrast, the empirical distributions based on the finite mixture estimates show substantially different locations and dispersions. Across all airlines, the larger estimates of $\widehat{\alpha}$ and $\widehat{\sigma}$ using the mixture method lead to a leftward shift and an amplification in the spread of the empirical distributions. These changes in the empirical distributions' location and dispersion may have important economic implications in any application that requires demand estimates as input for further analyses---irrespective of whether endogenous product entry and/or exit is in itself of any economic interest. \subsection{Estimation of costs and counterfactual experiments } In this paper, we focus on the consistent estimation of demand parameters in the presence of endogenous product entry. However, relying on the structure of our model, it is straightforward for researchers to estimate marginal costs, entry costs, and the joint distribution of unobservable variables. Given these estimated primitives, a variety of counterfactual experiments can be performed. In this subsection, we discuss these additional estimation procedures in the context of our empirical application. \subsubsection{Marginal costs} Based on an assumption about the nature of competition, such as Bertrand-Nash competition, the researcher would be able to estimate marginal costs at the airline-market-quarter level as the residuals from the pricing equation. It is important to note that these marginal costs can be computed only for those airlines that are observed to be active in the market. For some empirical questions, given the marginal costs, the researcher may need to further estimate a marginal cost function: that is, a function that represents the effect of product characteristics and output on marginal costs. For this purpose, the researcher needs to estimate the parameters of a regression in which the dependent variable is the marginal cost estimate and the explanatory variables are the exogenous characteristics $\boldsymbol{x}_{jt}$ and, in the case of non-linear returns to scale, the output $q_{jt}$. As in the case of demand, this regression is subject to selection bias due to endogenous product entry. Remarkably, the structure of the selection term in this equation mirrors that in the demand equation. We can then control for selection bias in the estimation of the marginal cost function using exactly the same control variables that we have used for the estimation of the demand parameters. We now illustrate these points in the context of our application. Following ciliberto2021market, we assume that the airlines engage in Bertrand-Nash competition and that each airline has marginal cost function that does not depend on output. Then, given demand equation (ref), the marginal cost function of airline $j$ in market-quarter $t$ can be estimated from the following pricing equation: \begin{equation} p_{jt} + \frac{1-\sigma}{\alpha(1-\sigma s_{jt|g}-(1-\sigma)s_{jt})} = mc_{jt}, \end{equation} where $g$ denotes the nest that contains all the airlines, $s_{jt|g} \equiv s_{jt}/(1-s_{0t})$ is the within-nest market share, and the marginal cost $mc_{jt}$ is specified as: \begin{equation} mc_{jt} \; = \; \boldsymbol{x}_{jt}^{\prime}\;\boldsymbol{\varphi} + \widehat{\boldsymbol{r}}_{jt}^{\prime} \; \widetilde{\boldsymbol{\mu}}^{\text{mc}}_{j} + \widetilde{\omega}_{jt}, \end{equation} with both $\boldsymbol{x}_{jt}$ and $\widehat{\boldsymbol{r}}_{jt}$ defined as in the case of demand equation (ref), while $\widetilde{\boldsymbol{\mu}}^{\text{mc}}_{j}$ is a vector of $L-1$ parameters $\widetilde{\mu}^{\text{mc}}_{j \ell}$ with $\widetilde{\mu}^{\text{mc}}_{j \ell} = \mathbb{E}( \omega_{jt} \mid \boldsymbol{\kappa}_t \in \mathcal{K}_\ell)$. \begin{table}[ht] \caption{Average Marginal Costs } \resizebox{0.9\textwidth}{!}{ \begin{tabular}{l|cc|ccccc} \hline \hline & & & & & & & \\ & \multicolumn{2}{c|}{\textit{Not control. for sel.}} & \multicolumn{5}{c}{\textit{Controlling for endogenous selection}} \\ & {OLS} & {2SLS} & {2SLS} & {2SLS} & {2SLS} & {2SLS} & {2SLS} \\ & & & {Heckman} & {Semipar.} & {Fin.-Mix.} & {Fin.-Mix.} & {Fin.-Mix.} \\ & & & $L=1$ & $L=1$ & $L=2$ & $L=3$ & $L=4$ \\ \hline & & & & & & & \\ \multicolumn{1}{r|}{\textit{Marginal Cost (100\$)}} & $0.766$ & $1.718$ & $1.721$ & $1.736$ & $1.769$ & $1.810$ & $1.817$ \\ & & & & & & & \\ \multicolumn{1}{r|}{\textit{AA}} & $0.901$ & $1.829$ & $1.832$ & $1.847$ & $1.881$ & $1.924$ & $1.930$ \\ \multicolumn{1}{r|}{\textit{DL}} & $1.049$ & $2.032$ & $2.036$ & $2.050$ & $2.082$ & $2.123$ & $2.130$ \\ \multicolumn{1}{r|}{\textit{UA}} & $1.134$ & $2.072$ & $2.075$ & $2.090$ & $2.123$ & $2.165$ & $2.171$ \\ \multicolumn{1}{r|}{\textit{US}} & $0.830$ & $1.779$ & $1.782$ & $1.797$ & $1.830$ & $1.871$ & $1.878$ \\ \multicolumn{1}{r|}{\textit{WN}} & $0.464$ & $1.461$ & $1.464$ & $1.478$ & $1.510$ & $1.549$ & $1.557$ \\ \multicolumn{1}{r|}{\textit{LCC}} & $0.434$ & $1.330$ & $1.333$ & $1.349$ & $1.384$ & $1.427$ & $1.432$ \\ \multicolumn{1}{r|}{\textit{Others}} & $0.362$ & $1.220$ & $1.224$ & $1.239$ & $1.276$ & $1.319$ & $1.323$ \\ & & & & & & & \\ \hline Observations & $35,763$ & $35,763$ & $35,763$ & $35,763$ & $35,763$ & $35,763$ & $35,763$ \\ \hline \hline \end{tabular}} \begin{tablenotes} • \scriptsize \end{tablenotes} \end{table} Table (ref) reports the average marginal costs obtained from equation (ref) and the demand estimates in Table (ref) (see Appendix Figure (ref) for the corresponding empirical distributions), while Table (ref) presents our estimates of $\boldsymbol{\varphi}$ from equation (ref). The estimates of $\boldsymbol{\varphi}$ in each column of Table (ref) rely on the corresponding demand estimates of Table (ref), so that, for example, the first column of Table (ref) reports estimates of $\boldsymbol{\varphi}$ obtained by using the estimates of $\alpha$ and $\sigma$ (i.e., plugging them in the left-hand side of (ref)) from the first column of Table (ref). Collectively, these results illustrate that although endogeneity of prices and of within-nest market shares play an important role in the implied marginal cost estimates from equation (ref), endogenous selection seems to have less of an impact. Moreover, the parameter estimates of equation (ref) (which uses the estimated $mc_{jt}$ as a dependent variable) look remarkably similar across \textit{all} columns of Table (ref), including in the case of the OLS. From these findings, we can conclude that---at least in our sample---the unobserved component of entry $\eta_{jt}$ appears to be strongly correlated with the unobserved component of demand $\xi_{jt}$ but not with that of marginal cost $\omega_{jt}$. In other words, heterogeneity in airlines' entry decisions appears to be primarily explained by demand-side rather than by marginal cost-side unobserved heterogeneity. \begin{table}[ht] \caption{Estimation of Marginal Cost Parameters } \resizebox{0.9\textwidth}{!}{ \begin{tabular}{l|cc|ccccc} \hline \hline & & & & & & & \\ & \multicolumn{2}{c|}{\textit{Not control. for sel.}} & \multicolumn{5}{c}{\textit{Controlling for endogenous selection}} \\ & {OLS} & {2SLS} & {2SLS} & {2SLS} & {2SLS} & {2SLS} & {2SLS} \\ & & & {Heckman} & {Semipar.} & {Fin.-Mix.} & {Fin.-Mix.} & {Fin.-Mix.} \\ & & & $L=1$ &$L=1$ & $L=2$ & $L=3$ & $L=4$ \\ \hline & & & & & & & \\ Distance (1000mi) & $0.971$ & $0.927$ & $0.938$ & $0.935$ & $0.934$ & $0.937$ & $0.965$ \\ & $(0.014)$ & $(0.014)$ & $(0.023)$ & $(0.022)$ & $(0.024)$ & $(0.023)$ & $(0.025)$ \\ & & & & & & & \\ Distance$^2$ & $-0.150$ & $-0.139$ & $-0.146$ & $-0.144$ & $-0.147$ & $-0.149$ & $-0.149$ \\ & $(0.006)$ & $(0.006)$ & $(0.008)$ & $(0.008)$ & $(0.009)$ & $(0.009)$ & $(0.010)$ \\ & & & & & & & \\ hub-size orig. (100s) & $0.247$ & $0.382$ & $0.237$ & $0.103$ & $0.326$ & $0.348$ & $0.288$\\ & $(0.013)$ & $(0.013)$ & $(0.024)$ & $(0.031)$ & $(0.034)$ & $(0.034)$ & $(0.034)$ \\ & & & & & & & \\ hub-size dest. (100s) & $0.243$ & $0.377$ & $0.241$ & $0.105$ & $0.330$ & $0.353$ & $0.290$ \\ & $(0.013)$ & $(0.013)$ & $(0.024)$ & $(0.031)$ & $(0.035)$ & $(0.034)$ & $(0.034)$ \\ & & & & & & & \\ \hline & & & & & & & \\ Airline$\times$Quarter FE & Y & Y & Y & Y & Y & Y & Y \\ \# control var. entry & 0 & 0 & 6 & 18 & 36 & 54 & 72 \\ Observations & $35,763$ & $35,763$ & $35,763$ & $35,763$ & $35,763$ & $35,763$ & $35,763$ \\ & & & & & & & \\ \hline \hline \end{tabular}} \begin{tablenotes} • \scriptsize{Bootstrap standard errors are based on the linearized bootstrap procedure in Appendix (ref) and account for first-step estimation error conditional on the selected first-step specification.} \end{tablenotes} \end{table} \begin{figure} \caption{Distribution of Estimated Marginal Costs (Airline-Market-Quarter level) } \end{figure} \section{Conclusions } In local geographic markets, we typically find only a subset of all the differentiated products in an industry. Firms strategically select specific products that better match the preferences of local consumers. When making market entry decisions, firms possess information about the demand for their products, particularly regarding unobservable demand components. Firms tend to enter markets with higher expected demand. Neglecting this selection process can introduce significant biases in the estimation of demand parameters. This issue is common across various demand applications and industries. Existing methods to address this issue typically rely on strong parametric assumptions about demand unobservables and firms' information. In this paper, we investigate the identification of demand parameters within a structural model that encompasses demand, price competition, and market entry (static or dynamic), while specifying the distribution of demand unobservables in a nonparametric finite mixture manner. The paper makes three main contributions. First, it establishes sequential identification of the demand parameters in this model. We demonstrate that the selection term in the demand equation results from a convolution of the probabilities of product entry for each discrete unobserved market type and the densities associated with these market types. We show that data on firms' product entry decisions nonparametrically identify the probabilities of product entry conditional on the market type and the density of unobserved market types. Under mild conditions on the observable variables, demand parameters are identified after controlling for the nonparametric entry probabilities and densities for each market type. Second, we propose a simple two-step estimator to address endogenous selection. In the first step, we estimate a nonparametric finite mixture model to determine the choice probabilities of product entry. In the second step, demand parameters are estimated using a Generalized Method of Moments (GMM) approach that accounts for both endogenous product availability and price endogeneity. Third, we illustrate the proposed method by applying it to data from the airline industry. The findings highlight the importance of allowing for a finite mixture of unobserved market types when controlling for endogenous product entry, as failure to do so can lead to significant biases.