Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
78,681 characters · 20 sections · 50 citation commands
Uniformly Consistent Semi-nonparametric Demand Estimation with Micro-Data
\noindentKeywords: differentiated-products demand; micro data; sieve minimum distance; nonparametric IV; incidental parameters; uniform consistency.
\noindentJEL Codes: C14, C25, C36, L13.
Micro-level choice data provide a strong source of identifying variation for differentiated-product demand systems. With market-level data, researchers observe aggregate shares, prices, product characteristics, and instruments. Identification then generally relies on cross-market variation in observed market outcomes -- such as prices and market shares -- and on instruments that shift endogenous prices without shifting demand. With micro data, we can observe different consumers within markets that face the same prices, product attributes, and market-level demand shocks, but differ in their observed characteristics, thereby granting us another source of variation. berry_nonparametric_2024 show that this within-market variation can identify nonparametric demand systems using only price instruments. Their result is an identification theorem. This paper studies the subsequent estimation problem.
The object of interest is a semi-nonparametric demand system of the form: \[ \mathbb{E}[Q_{it}\mid C_{it}] = \sigma\Big(g_0(Z_{it},X_t)+\psi_0(P_t,X_t)+h_{0t}\Big) \] where $Q_{it} \in \{0,1\}^J$ is the vector of inside-good choice indicators for consumer $i$ in market $t$, $Z_{it}$ denotes consumer covariates, $P_t$ denotes the vector of prices, $X_t$ denotes observed product and market characteristics, and $h_{0t} \in \mathbb{R}^J$ is a vector of product-level demand shocks common to all consumers in market $t$. Here $C_{it}$ denotes the information observed by consumer $i$ in market $t$. The link function $\sigma$ is multinomial logit, but the functions comprising the index, $g_0$ and $\psi_0$, are unknown and flexible. Thus, $g_0$ captures observed consumer heterogeneity, while $\psi_0$ captures the systematic role of prices and observed product or market characteristics. The logit link is used for tractability\footnote{The goal for future work would be to relax this restriction.}; the model remains semi-nonparametric because the index functions are not parametrically specified.
The main difficulty in estimation is that the within-market data do not separately identify $\psi_0(P_t,X_t)$ and $h_{0t}$. Holding market $t$ fixed, both terms are constant across consumers and enter individual choice probabilities additively. Therefore, I define the composite market intercept: \[ a_{0t} := \psi_0(P_t,X_t) + h_{0t}\] similar to hahn_time-invariant_2005. This transformation will clarify the source of identification, and it will also help us handle the incidental parameters problem often generated when $T \to \infty$. Moreover, the micro data identify how choices vary with $Z_{it}$ within a market and therefore provide information about $g_0$ and $a_{0t}$. They do not, by themselves, decompose $a_{0t}$ into the price-side function $\psi_0(P_t,X_t)$ and the demand shock, $h_{0t}$. This decomposition requires excluded price instruments, $W_t$. Therefore, the estimator must combine two sources of variation: within-market variation in $Z_{it}$, which identifies the consumer-heterogeneity component and the composite intercept, and cross-market variation in $(P_t,W_t,X_t)$, which separates the composite intercept into $\psi_0(P_t,X_t)$ and $h_{0t}$.
I propose a profiled sieve minimum-distance estimator that implements this logic. The estimator has three main steps. First, for each candidate heterogeneity function $g$, I profile the market-specific composite intercepts $a_t$ by matching observed and predicted within-market shares. Second, I estimate $g_0$ using micro moments formed from the profiled individual choice residuals. Third, I estimate $\psi_0$ using the price-side conditional moment: \[ \mathbb{E}[a_{0t}-\psi_0(P_t,X_t)\mid W_t,X_t]=0.\] This moment will follow from the definition of $a_{0t}$, the normalization $\mathbb{E}[h_{0t} \mid X_t] = 0$, and the excluded-instrument restriction $\mathbb{E}[h_{0t} \mid W_t,X_t] = \mathbb{E}[h_{0t} \mid X_t]$ -- the latter two of which come directly from berry_nonparametric_2024. I implement the above moment condition by projecting the implied shock, $a_t-\psi(P_t,X_t)$, onto a growing basis of functions with arguments $(W_t,X_t)$. This is a sieve minimum-distance (SMD), or conditional-moment, step in the sense of ai_efficient_2003. The price-side moment is also closely related to the nonparametric IV (NPIV) logic in newey_instrumental_2003: completeness of the conditional distribution of prices given instruments is what converts the conditional moment into identification of the unknown function $\psi_0$.
The main result of this paper establishes uniform consistency of the profiled SMD estimator. The asymptotic framework has both the number of markets and the within-market sample sizes growing: $T \to \infty$ and $\min_{t\le T} n_t \to \infty$\footnote{This differs from large-product asymptotic frameworks for differentiated-products demand, where the number of products grows; see, for example, freyberger_asymptotic_2015 and lu_semi-nonparametric_2023.}. This sampling scheme is relevant because it would ordinarily imply the aforementioned incidental parameters problem: the number of market-specific intercepts grows to infinity with the number of markets. However, each intercept is to be estimated from a growing number of consumers. In conjunction with my profiling approach, the market intercepts will not cause a problem for consistency as they will not be estimated with fixed information. Under regularity, identification, sieve approximation and uniform convergence conditions, \[ \left\lVert \hat g - g_0 \right\rVert_\infty + \left\lVert \hat\psi - \psi_0 \right\rVert_\infty = o_p(1).\] The logic comes in three steps: first, the inner share-matching problem consistently profiles $a_t$ uniformly over candidate functions $g$. Second, the feasible profiled criterion converges uniformly to its population counterpart. And lastly, the population criterion is uniquely minimized at the true parameter values, subject to identification normalizations.
My work contributes to several literatures. First, it contributes to the econometric and industrial organization literature on differentiated-products demand. berry_estimating_1994 and berry_automobile_1995 provide the canonical framework for estimating demand with endogenous prices and market-level data. Subsequent work developed richer sources of variation, micro moments, and improved computational methods; examples include berry_differentiated_2004, nevo_measuring_2001, petrin_quantifying_2002, dube_improving_2012, reynaert_improving_2014, and gandhi_measuring_2025. Rather than specifying a parametric random-coefficients utility model and deriving an estimating equation through a market-share inversion, I begin from the Berry-Haile micro-data identification framework and impose a tractable semi-nonparametric link for estimation.
Second, my work is closely related to Berry and Haile's work on nonparametric identification of differentiated-products demand. berry_connected_2013 study connected substitutes and demand inversion, while berry_identification_2014 develop nonparametric identification results using market-level data and instruments. Most relevant to my work, berry_nonparametric_2024 show that micro data can identify demand using only price instruments. I use this identification result as the foundation for my estimator. In this sense, my paper is an estimation counterpart to the Berry-Haile micro-data identification result.
Third, the paper contributes to the literature on flexible and nonparametric demand estimation. compiani_market_2022, wang_sieve_2023, and tebaldi_nonparametric_2023 study flexible demand estimation in differentiated-products settings. monardo_measuring_2025 develops a scalable method that exploits restrictions such as permutation invariance to reduce the dimensionality of demand estimation. lu_semi-nonparametric_2023 studies semi-nonparametric estimation of random-coefficients logit demand using aggregate market-share data and a large-(J) asymptotic framework. I take a complimentary approach by using micro-level choice data and exploiting within-market variation across consumers instead of using market-level data to flexibly estimate aggregate demand or the distribution of random coefficients. This difference changes both the source of identifying variation and the relevant asymptotic framework: the number of markets and the number of consumers per market grow, while the number of products is treated as fixed.
Fourth, the paper contributes to the literature on nonparametric IV and sieve minimum distance. newey_instrumental_2003 show how completeness identifies nonparametric IV models, while hall_nonparametric_2005, horowitz_applied_2011, and related work study the ill-posedness and regularization issues that arise in nonparametric IV problems. ai_efficient_2003 develop efficient estimation for conditional moment models with unknown functions, and chen_sieve_2015 develops inference methods for semi/nonparametric conditional moment models. The price-side of my estimator follows this logic: the residual, \(a_t-\psi(P_t,X_t)\), must have zero conditional mean given \((W_t,X_t)\), and the conditional mean is estimated by projection onto an expanding basis. This use of growing unconditional moments to approximate a conditional moment is also related to donald_empirical_2003. The distinctive feature of my work is that the residual contains a profiled object, \(a_t\), estimated from within-market micro data.
Finally, this paper relates to work on fixed effects and incidental parameters in nonlinear models. The inner profile step estimates a vector of market-specific intercepts whose dimension grows with $T$. In classical nonlinear panel settings, such nuisance parameters create incidental-parameters bias when each fixed effect is estimated with a fixed number of observations hahn_jackknife_2004. In my setting, each market intercept is estimated using $n_t$ within-market observations, and $\min_{t\le T} n_t \to \infty$. This is adjacent to the logic in hahn_time-invariant_2005 where the nuisance parameters do not contaminate consistency of the common structural functions when the information per nuisance parameter diverges.
The rest of the paper proceeds as follows. Section (ref) provides a brief motivation for the estimator. Section (ref) presents the model and defines the composite market intercept. Section (ref) states the identification argument and the normalizations used to separate the additive components. Section (ref) defines the profiled sieve minimum-distance estimator. Section (ref) states the main uniform consistency theorem. Section (ref) presents Monte Carlo evidence, and Section (ref) concludes. The appendix contains the Berry-Haile primitive assumptions, all of the formal proofs, simulation details, and supplemental discussions regarding the curse of dimensionality.
Demand estimation is often used to answer counterfactual questions regarding markets. How would prices change after a cost shock? How much would demand would divert from one product to another after a price increase? What are the implied markups for different firms? How would consumer surplus change after a new product, tax, or merger? In differentiated-product markets, these answers depend on the shape of the demand curve. A demand model that fits market shares but imposes the wrong substitution patterns or the wrong price curvature can give misleading counterfactual conclusions.
In applications, researchers would use the empirical distribution of observed consumers, $n_t$, in market $t$ to compute the aggregate inside-good share: \[ \hat S_t(p) := \frac1{n_t}\sum_{i=1}^{n_t} \sigma\!\left( \hat g(Z_{it},X_t)+\hat\psi(p,X_t)+\hat h_t \right) \] where $p$ is a counterfactual price vector. From this object, one can compute own- and cross-price elasticities, diversion ratios, markup estimates, and cost shock pass-through. These objects depend on the slope or curvature of demand; hence, flexibly recovering the demand surface should improve the accuracy of the aforementioned estimated counterfactual objects.
This section presents the structural model for which I develop an estimator. This is done in two steps. First, I describe the high-level model studied by berry_nonparametric_2024 that motivates this paper. In particular, demand is generated by an index that separates within-market consumer heterogeneity from market-level demand shocks. Second, I impose additional structure used for estimation: a multinomial-logit link, a flexible price-side function, and a composite market intercept that can be profiled from within-market micro data. The location normalizations needed to separate the additive components are imposed in Section (ref), not in the model statement itself.
Markets are indexed by $t=1,\dots,T$, and consumers in market $t$ are indexed by $i=1,\dots,n_t$. Each consumer chooses among $J$ inside goods and an outside option. Let $\mathcal{J}=\{1,\dots,J\}$ be the set of alternatives and let $Q_{it}=(Q_{1it},\dots,Q_{Jit})'\in\{0,1\}^J$ denote the vector of inside-good choice indicators. Thus $Q_{jit}=1$ if consumer $i$ in market $t$ chooses product $j$, while the outside option is chosen when all components of $Q_{it}$ are equal to zero.
Consumer covariates are denoted by $Z_{it}\in\mathcal{Z}\subseteq\mathbb{R}^{d_z}$. Market-level prices, observed product/market characteristics, and excluded price instruments are denoted by $P_t\in\mathcal{P}\subseteq\mathbb{R}^J$, $X_t\in\mathcal{X}\subseteq\mathbb{R}^{d_x}$, and $W_t\in\mathcal{W} \subseteq \mathbb{R}^{d_w}$, respectively. Market-level unobserved demand shocks are collected in $\Xi_t\in\mathbb{R}^J$. Here $d_z$ denotes the dimension of the consumer covariates, $d_x$ denotes the dimension of the observed product/market characteristics, and $d_w$ denotes the dimension of the excluded price instruments. Throughout the paper, a prime denotes transpose. Following berry_nonparametric_2024, the full choice environment faced by consumer $i$ in market $t$ is \[ C_{it}:=(Z_{it},P_t,X_t,\Xi_t). \] The structural choice probability vector is
with outside-good probability $s_0(C_{it})=1-\sum_{j=1}^J s_j(C_{it})$. From here, we get to a model of demand through the Berry--Haile representation -- i.e., the existence of a finite-dimensional index that enters an otherwise nonparametric demand link.
Assumption (ref) is the economic foundation for using micro data. Within a market, the objects $(P_t,X_t,\Xi_t)$ are fixed across consumers, while $Z_{it}$ varies. Thus, within-market variation in $Z_{it}$ shifts the index without changing the market-level environment. This is the source of variation that Berry and Haile exploit to identify conditional demand without instruments for quantities.
For the estimator, it is convenient to rewrite the Berry--Haile index as a sum of a consumer-heterogeneity function and a market-level shock. I write this as
where $g_0:\mathcal{Z}\times\mathcal{X}\to\mathbb{R}^J$ captures the component that varies with consumer covariates, and $h_{0t}\in\mathbb{R}^J$ captures the market-level demand shock. Specifically, $g_0(Z_{it},X_t)\equiv\Gamma_0(Z_{it},X_t)+\mathbb{E}[\Xi_t\mid X_t]$ and $h_{0t}\equiv\Xi_t-\mathbb{E}[\Xi_t\mid X_t]$, such that their sum recovers equation (ref). At this point, the decomposition between $g_0$ and $h_{0t}$ is only defined up to location normalizations. Those normalizations are stated in Section (ref).
This paper estimates a specialization of (ref). The fully nonparametric Berry--Haile link $\sigma^{BH}$ is replaced by a multinomial-logit link, but the systematic role of prices and observed product/market characteristics is left flexible through an unknown function, $\psi_0$.
Under Assumption (ref), the maintained estimable demand model is
This is the main modeling restriction imposed by this paper beyond the Berry--Haile identification framework. The logit link makes the estimator tractable and gives a simple share-matching problem within each market. At the same time, the model remains semi-nonparametric because both $g_0$ and $\psi_0$ are unknown functions. The function $g_0$ captures how consumer covariates shift relative tastes for inside goods, while $\psi_0$ captures the systematic role of prices and observed product/market characteristics. The vector $h_{0t}$ captures product-level demand shocks that are common to all consumers within market $t$.
The object directly learned from the within-market micro data is not $h_{0t}$ by itself. Holding market $t$ fixed, both $\psi_0(P_t,X_t)$ and $h_{0t}$ are constant across consumers and enter (ref) only through their sum. Following the fixed-effect profiling logic in berry_automobile_1995 and hahn_time-invariant_2005, I define the composite market intercept
Then the model can be written as
This representation is central for both computation and asymptotics. For a candidate $g$, the inner step of the estimator will profile the market-specific vector $a_t$ from the within-market choice data. The price-side step then decomposes the profiled intercept into the systematic component $\psi_0(P_t,X_t)$ and the structural shock $h_{0t}$ using excluded price instruments.
The market-specific objects $a_{0t}$ and $h_{0t}$ are nuisance objects. They are nevertheless economically important because, after estimating $(g_0,\psi_0)$, the structural market shock is recovered as $h_{0t}=a_{0t}-\psi_0(P_t,X_t)$.
The target structural parameter is
Here $\mathcal{G}$ and $\boldsymbol{\Psi}$ are function spaces. Let $\mathcal D_g:=\mathop{\mathrm{supp}}(Z_{it},X_t)$ and $\mathcal D_\psi:=\mathop{\mathrm{supp}}(P_t,X_t).$ The space $\mathcal{G}$ contains normalized, bounded, smooth functions $g:\mathcal D_g\to\mathbb{R}^J$, and the space $\boldsymbol{\Psi}$ contains bounded, smooth functions $\psi:\mathcal D_\psi\to\mathbb{R}^J$. Heuristically,
where $C^s_B(D)$ denotes a bounded H\"older ball of smoothness $s$ with radius $B$ on the compact domain $D$. The smoothness and boundedness restrictions are regularity conditions used for compactness and sieve approximation. The normalization in $\mathcal{G}$ removes the part of $g$ that is constant across consumers within a market, while the normalization $\mathbb{E}[h_{0t}\mid X_t]=0$ fixes the location of $\psi_0$ relative to the structural shock. Section (ref) presents and discusses these normalizations formally. Similarly, Appendix (ref) defines these function spaces more explicitly and explains why these restrictions imply compactness under the sup norm.
The consistency analysis uses a double-asymptotic sampling scheme. Markets are independent across $t$, and conditional on the market-level environment, the observations within each market are independent.
The number of market-specific intercepts grows with $T$, but each intercept is estimated using a growing number of within-market observations. This is why the market effects are nuisance parameters without creating the usual fixed-$T$ incidental-parameters problem, as discussed in hahn_time-invariant_2005. The estimator combines many markets with large within-market micro samples.
This section states the identification logic needed for the main text. In general, berry_nonparametric_2024 shows that the high-level demand system in equation ((ref)) can be nonparametrically identified under the assumptions stated in Appendix (ref). This paper fundamentally relies on their identification result; however, since I impose additional restrictions in equation ((ref)), I will show that they too are identified.
Towards this, I state the main assumptions my estimator relies on and explain how they correspond to the Berry--Haile primitives. There are two objects to identify. The first is the within-market component of demand: the consumer heterogeneity function, $g_0$, and the composite market intercept, $a_{0t}$. The second is the price-side decomposition of the composite intercept $a_{0t} = \psi_0(P_t, X_t) + h_{0t}$.
The first step follows from the Berry--Haile index, support, invertibility, and normalization assumptions. The second step follows from their price-instrument and completeness assumptions.
The additive index contains multiple components, so location normalizations are required before identification can be stated. These normalizations do not impose economic substitution patterns; they simply assign the level of the index to particular components.
The first normalization follows from the construction of $h_{0t}$. The third is the standard outside-option normalization in multinomial logit models. The second normalization, coming directly from berry_nonparametric_2024, fixes the location of $g_0$ relative to the market-level intercept. Without (ref) or (ref), a component of $g_0$ that depends only on $X_t$ would be indistinguishable from the market intercept. See Appendix (ref) for a full equivalence-class argument.
The logit specification of $\sigma(\cdot)$ gives a simple way to see what the within-market data identify. Let $s_j(C_{it})$ and $s_0(C_{it})$ denote the inside-good and outside-good probabilities generated by (ref). Then, for each $j\in\mathcal{J}$,
Within a market, $a_{0jt}$ is constant across consumers, while $g_{0j}(Z_{it},X_t)$ varies with $Z_{it}$. Hence, after the normalization on $g_0$, within-market variation identifies the heterogeneity function and the composite market intercepts. The estimator will implement this logic by profiling $a_t$ market-by-market for each candidate $g$.
The within-market step identifies the composite object $a_{0t}$, but it does not by itself separate $\psi_0(P_t,X_t)$ from $h_{0t}$. This separation is obtained from excluded price instruments. The relevant price-side restriction is inherited from the Berry--Haile instrument assumption in Appendix (ref).
Assumption (ref) is a restatement of the Berry--Haile price-instrument and completeness conditions. Equation (ref) says that, after conditioning on $X_t$, the excluded instruments do not predict the structural demand shock. Combining (ref) with the normalization $\mathbb{E}[h_{0t}\mid X_t]=0$ gives \[ \mathbb{E}[h_{0t}\mid W_t,X_t]=0. \] Since $h_{0t}=a_{0t}-\psi_0(P_t,X_t)$, the identified composite intercept satisfies the price-side conditional moment
This NPIV-style moment condition is what we use to identify and eventually estimate $\psi_0$.
This section defines the estimator that will then be analyzed in Section (ref). Its construction mirrors the identification argument in that the within-market data identify, for each candidate heterogeneity function, the composite market intercept that shifts all consumers in the same market. The cross-market price-side moment then separates this composite intercept into the systematic price component and the structural demand shock.
Before introducing the sample criterion, it is useful to isolate the two population restrictions that make the estimator valid. I define the structural choice residual at the truth by
The first restriction is the structural residual condition,
which follows directly from the maintained choice model. The second restriction is the price-side conditional moment,
which follows from $a_{0t}=\psi_0(P_t,X_t)+h_{0t}$, the normalization $\mathbb{E}[h_{0t}\mid X_t]=0$, and the price-instrument condition $\mathbb{E}[h_{0t}\mid W_t,X_t]=\mathbb{E}[h_{0t}\mid X_t]$.
In words, the estimator works as follows. First, choose finite-dimensional sieve spaces for $g_0$ and $\psi_0$. Second, for any candidate $g$, choose the market intercepts $a_t$ that make predicted market shares match observed market shares. Third, use the resulting profiled choice residuals to estimate $g_0$ from within-market variation in $Z_{it}$. Fourth, choose $\psi$ so that the recovered shocks $a_t-\psi(P_t,X_t)$ have zero conditional mean given $(W_t,X_t)$. Fifth, after estimating $(g_0,\psi_0)$, recover the market-specific composite intercepts and structural demand shocks.
Let $\mathcal D_g:=\mathop{\mathrm{supp}}(Z_{it},X_t)$ and $\mathcal D_\psi:=\mathop{\mathrm{supp}}(P_t,X_t)$. For two vector-valued functions $g,\tilde g:\mathcal D_g\to\mathbb{R}^J$ and $\psi,\tilde\psi:\mathcal D_\psi\to\mathbb{R}^J$, define
The product norm on the structural parameter space is
The term “product norm” means the distance used on the Cartesian product $\Theta=\mathcal{G}\times\boldsymbol{\Psi}$: it measures the distance between two structural parameters by adding the supremum-norm distance between their $g$ components and the supremum-norm distance between their $\psi$ components. Thus convergence in $\left\lVert \cdot \right\rVert_{\Theta}$ is equivalent to the joint statement $\left\lVert g-\tilde g \right\rVert_\infty\to0$ and $\left\lVert \psi-\tilde\psi \right\rVert_\infty\to0$. The subscript $\infty$ denotes the usual uniform, or $\ell_\infty$, norm: for vector-valued functions, I take the largest absolute error over products and over the relevant support. Because the objects of interest are functions, uniform consistency is the natural notion of consistency as it requires the entire estimated demand system to converge over its domain, rather than only at fixed points.
Let $b^{K_g}(z,x)=(b_1(z,x),\dots,b_{K_g}(z,x))'\in\mathbb{R}^{K_g}$ be a basis for the consumer-heterogeneity function, and let $r^{K_\psi}(p,x)=(r_1(p,x),\dots,r_{K_\psi}(p,x))'\in\mathbb{R}^{K_\psi}$ be a basis for the price-side function. For each product $j\in\mathcal{J}$, write \[ g_{\beta,j}(z,x)=b^{K_g}(z,x)'\beta_j, \qquad \psi_{\vartheta,j}(p,x)=r^{K_\psi}(p,x)'\vartheta_j, \] where $\beta_j\in\mathbb{R}^{K_g}$, $\vartheta_j\in\mathbb{R}^{K_\psi}$, $\beta=(\beta_1',\dots,\beta_J')'\in\mathbb{R}^{JK_g}$, and $\vartheta=(\vartheta_1',\dots,\vartheta_J')'\in\mathbb{R}^{JK_\psi}$. The finite-dimensional structural sieve is
where $\Theta_K$ depends on the structural dimensions $(K_g,K_\psi)$, while the full sample criterion -- defined below -- also depends on the conditional-moment dimension, $K_q$. Thus $K_g$ controls both the sieve for $g$ and the dimension of the micro moment, $K_\psi$ controls only the sieve for $\psi$, and $K_q$ controls only the series approximation to the price-side conditional mean. The component sieves are \[
\] for compact coefficient sets $B_{K_g}$ and $V_{K_\psi}$, respectively.\footnote{In practice, this can be implemented either by imposing the resulting linear restrictions on $\beta$ or by using a centered basis, for example $b^{K_g}(z,x)-\mathbb{E}[b^{K_g}(Z_{it},x)\mid X_t=x]$ under the conditional-mean normalization, or $b^{K_g}(z,x)-b^{K_g}(z^0(x),x)$ under the baseline normalization. The point is that the estimator never treats an $X$-only term in $g$ as separately estimable from the market intercept.}
Fix a candidate $g\in\mathcal{G}_{K_g}$. In market $t$, the object that is constant across consumers is not $h_{0t}$ alone, but the composite intercept $a_{0t}=\psi_0(P_t,X_t)+h_{0t}$. The estimator therefore profiles $a_t$ market-by-market. Let $\hat s_t=n_t^{-1}\sum_{i=1}^{n_t}Q_{it}$. For $a\in\mathcal{A}\subset\mathbb{R}^J$, define
The profiled composite intercept is
The log-sum-exp criterion is used only as a convex objective whose gradient equals the negative of the within-market share-matching moment. When the minimizer is interior, its first-order condition is exactly the within-market share-matching equation
Thus the inner step is not introducing a separate likelihood estimator for the structural functions; it is just a convenient way to solve the just-identified within-market moment for the composite intercept.
For later reference, let $a_t^*(g)$ denote the population counterpart of (ref). At the truth, $a_t^*(g_0)=a_{0t}$; Appendix (ref) proves this claim and gives primitive conditions under which the sample profile converges uniformly to $a_t^*(g)$.
The micro block estimates $g_0$ using within-market variation in $Z_{it}$. Define the sample-profiled residual
and define its population-profile analogue by
For each $K_g$, the feasible micro moment is
Its population counterpart is
The subscript $K_g$ is substantive as the dimension of the moment changes with the number of basis functions. Here $\otimes$ denotes the Kronecker product: if $u\in\mathbb{R}^{K_g}$ and $v\in\mathbb{R}^J$, then $u\otimes v\in\mathbb{R}^{K_gJ}$ stacks all products $u_kv_j$. This moment is centered at the truth only after the profile identity $a_t^*(g_0)=a_{0t}$ has been established.
With a positive definite weighting matrix $W_{g,K_g}\in\mathbb{R}^{JK_g\times JK_g}$, define the micro-block criterion
The same basis $b^{K_g}$ appears in the sieve for $g$ and in the micro moments. This is intentional as the moment projects the residual onto the same directions used to perturb the unknown function. In a finite-dimensional logit model these would be score equations for the coefficients of $g$; here they are used as GMM moments for a growing sieve.
For a candidate pair $(g,\psi)\in\Theta_K$, define the implied structural demand shock
The population analogue is $h_t^*(g,\psi):=a_t^*(g)-\psi(P_t,X_t)$. At the truth, this implied shock equals $h_{0t}$ once $a_t^*(g_0)=a_{0t}$. This connects the profiled market intercept and the price-side conditional moment: the object projected onto functions of $(W_t,X_t)$ is centered at the truth.
Let $R_t:=(W_t,X_t) \in \mathcal{R}$ where $ \mathcal R:=\mathop{\mathrm{supp}}(R_t)\subseteq\mathcal{W}\times\mathcal{X}$ so that $\mathcal R$ is the support of the conditioning variables in the price-side conditional moment. Let $q^{K_q}(r)=(q_1(r),\dots,q_{K_q}(r))'\in\mathbb{R}^{K_q}$ be a series basis for functions of $r\in\mathcal R$. Following ai_efficient_2003, the population conditional moment for the price-side block is
At the truth, the price-side conditional moment is zero.
Let $\bm Q_T\in\mathbb{R}^{T\times K_q}$ be the matrix with $t$-th row $q^{K_q}(R_t)'$, and assume $\bm Q_T'\bm Q_T$ is nonsingular on the event used to compute the criterion. Following ai_efficient_2003, the series estimate of the conditional mean of the implied shock is
In this expression, $\hat h_t(g,\psi)q^{K_q}(R_t)'$ is a $J\times K_q$ matrix, so the bracketed sum is also $J\times K_q$. The corresponding SMD criterion is
where $\hat\Sigma(R_t)\in\mathbb{R}^{J\times J}$ is uniformly positive definite. Equivalently, when $\hat\Sigma(R_t)=I_J$, this is GMM with the increasing set of moments generated by $q^{K_q}(R_t)\otimes\{\hat a_t(g)-\psi(P_t,X_t)\}$; as $K_q$ grows, these unconditional moments approximate the conditional moment in (ref), as in ai_efficient_2003 and donald_empirical_2003.
The full sample criterion is
The estimator is any approximate minimizer\footnote{The use of an approximate minimizer is standard in extremum-estimation consistency arguments. It avoids requiring the numerical optimizer to attain the exact global minimum of a nonconvex finite-sample criterion; the proof only needs the optimizer to achieve the infimum up to an $o_p(1)$ error. See newey_chapter_1994.} over the structural sieve:
The final profiled composite intercepts and recovered structural shocks are
The estimated demand system can then be used for counterfactual analysis. Holding the recovered shock fixed, predicted demand in market $t$ at price vector $p$ is
Differentiating (ref) with respect to prices gives local elasticities and diversion ratios. Combining those derivatives with an ownership matrix gives Bertrand markups and pass-through. Appendix (ref) details these formulas. The present paper establishes consistency of the structural objects entering (ref); inference for the resulting IO functionals is left for future work.
This section states the main consistency argument. I keep the main text at the level needed to understand why the estimator works and defer the primitive profiling and empirical-process details to Appendix (ref)--(ref). The proof has two moving parts. First, each market-specific intercept is consistently profiled because $n_{\min}:=\min_{t\le T}n_t\to\infty$. Second, after replacing the population profile by this estimated profile, the outer criterion is a moving-sieve minimum-distance criterion. The word “moving” is important as the micro moment contains the basis vector $b^{K_g}$, and therefore the population micro criterion changes with $K_g$.
For each sieve dimension $K=(K_g,K_\psi,K_q)$, I define the population micro criterion
The population price-side criterion is the full conditional-mean criterion
Thus the population criterion associated with the finite sample criterion is
The price-side sample criterion depends on $K_q$ through the series regression used to estimate the conditional mean, but its population target is the full conditional moment in (ref). The micro population criterion, by contrast, is indexed by $K_g$ because the moment vector itself has dimension $JK_g$.
At the truth, both terms are zero for every $K$. The micro term is zero by Proposition (ref), and the price-side term is zero by Proposition (ref). Therefore \[ \mathcal Q_K(g_0,\psi_0)=0 \qquad\text{for every }K. \]
This assumption plays the usual role in sieve extremum estimation, with one additional requirement needed because the criterion moves with $K$. Compactness prevents the minimization problem from escaping to irregular functions with no convergent subsequence, while sieve approximation says that the finite-dimensional spaces eventually approach the true structural functions. The criterion-preservation condition says that these approximants are not merely close in sup norm; they are also close in the metric actually used by the moving population objective. This is automatic under suitable basis-envelope and weighting restrictions, but it is not implied by ordinary pointwise continuity of a fixed criterion because $\mathcal Q_K$ changes with $K$.
This assumption is the formal statement that the inner step is well posed. Uniqueness and separation ensure that the population intercept $a_t^*(g)$ is not ambiguous. The uniform law of large numbers is what lets each empirical market problem approximate its population counterpart, uniformly across candidate $g$ and markets. The Lipschitz condition says that a small error in $g$ cannot generate a large error in the population profile. For the logit link, these restrictions follow from bounded indices, compactness of $\mathcal{A}$, and the strong convexity of the log-sum-exp objective in $a$.
The next result is the first connection between the infeasible population problem and the feasible estimator: uniformly over candidate heterogeneity functions and markets, the empirical profile behaves like the population profile.
This assumption is the argmin identification condition appropriate for a moving-sieve criterion. Here $K\to\infty$ is shorthand for taking the sieve and series dimensions $(K_g,K_\psi,K_q)$ along the same admissible sequence used by the estimator as the sample grows. Since the sample grows through both $T\to\infty$ and $n_{\min}\to\infty$, the dimension sequence may depend on both $T$ and $n_{\min}$; the growth restrictions in Assumption (ref) require it to grow slowly enough relative to both sources of sampling information. The liminf is needed because $\mathcal Q_K$ is not a single fixed population objective; rather, its micro component changes when the dimension of $b^{K_g}$ changes. The condition says that, eventually along this admissible sequence, the population criterion remains uniformly bounded away from zero outside every sup-norm neighborhood of the truth.
This is the high-level stochastic equicontinuity condition for the full estimator; Appendix (ref) details the primitive sufficient conditions for each part. The micro block is a sieve GMM moment based on individual observations. Because it uses the generated profile $\hat a_t(g)$, Appendix (ref) first compares it with an oracle micro moment using $a_t^*(g)$. The price-side block is an Ai--Chen conditional-moment estimator implemented by a growing basis in $(W_t,X_t)$. The relevant high-level conditions are standard in sieve conditional-moment estimation; see ai_efficient_2003, chen_large_2007, and chen_sieve_2015. Its generated-profile error is controlled in the empirical $L_2$ norm induced by the realized market values $R_1,\ldots,R_T$.
The next result is the second connection between the infeasible population problem and the feasible estimator. It says that, after profiling the market intercepts, the feasible moving criterion is uniformly close to the moving population criterion that identifies the truth.
The preceding assumptions separate the consistency proof into three claims: the profile step is uniformly well behaved, the feasible moving criterion converges uniformly to its population analogue, and the moving population criterion separates the truth from all other normalized sieve parameters. Combining these claims gives the main result.
This section reports two Monte Carlo exercises. The first studies the consistency logic developed in Theorem (ref): as the number of markets and the number of consumers per market increase, the feasible profiled estimator should recover the structural functions and the generated market-level objects. The second compares the proposed flexible price-side estimator with standard differentiated-products demand benchmarks in counterfactual exercises of the kind used in empirical Industrial Organization. The point of the second exercise is to isolate a setting in which the price side of demand is nonlinear and to ask whether imposing a standard linear-in-price random-coefficients logit structure distorts counterfactual predictions. The comparison deliberately favors existing methods by using a heterogeneity structure that a random-coefficients logit model can represent well. The main difference across estimators is therefore the flexibility of the systematic price-side function.
The main text focuses on demonstrating uniform consistency of the estimator and comparing it to existing alternatives. Appendix (ref) details the data-generating process, the numerical implementation, the exact norm definitions, and the additional component diagnostics.
The baseline design has two inside goods, a scalar consumer covariate $Z_{it}$, a scalar observed market characteristic $X_t$, and two excluded instruments $W_t=(W_{1t},W_{2t})'$. Markets are independent. In each market, the structural composite intercept is \[ a_{0t}=\psi_0(P_t,X_t)+h_{0t}, \] and individual choices are drawn from the multinomial-logit demand system \[ \Pr(Q_{itj}=1\mid Z_{it},P_t,X_t,h_{0t}) = \sigma_j\{g_0(Z_{it},X_t)+a_{0t}\}, \qquad j=1,2. \] The price equation is constructed so that prices are endogenous: prices depend directly on the unobserved demand shock $h_{0t}$ as well as on the excluded instruments. The estimator observes $(Q_{it},Z_{it})$ within each market and observes $(P_t,X_t,W_t)$ at the market level. It does not observe $a_{0t}$, $h_{0t}$, $g_0$, or $\psi_0$.
The full-estimator simulation uses the correctly specified finite-sieve design. The heterogeneity function, $g_0$, lies in the same finite basis used by the estimator, and the price-side function, $\psi_0$, lies in the same finite basis used by the price block. Thus, any remaining error is sampling and optimization error rather than approximation error. The full estimator is run on the grid \[ T\in\{50,100,200\}, \qquad n_t\in\{100,250,500\}, \] with 50 Monte Carlo replications at each design point. The reported estimator is the feasible estimator: for every candidate $g$, the market intercepts are profiled from the simulated micro data, the micro moments are formed from the profiled residuals, and the price-side moment is evaluated using the generated shocks $\hat a_t(g)-\psi(P_t,X_t)$.
Figure (ref) presents the main simulation. Each panel reports the median error across 50 replications at a given pair $(T,n_t)$. The main pattern is consistent with the theorem: the sup norm errors for $g_0$ and $\psi_0$, and the average errors for $a_{0t}$ and $h_{0t}$, all decline as the amount of market-level and within-market information increases.
The sup-norm results are deliberately stringent. Unlike an integrated or root-mean-squared functional error, the sup norm records the largest absolute error over products and over the evaluation grid. It is, therefore, sensitive to local estimation error at tail grid points and should not be expected to be as smooth as an average error metric in a finite Monte Carlo design. Nevertheless, the same qualitative pattern emerges. The common structural functions are recovered more accurately in the high-information region of the design, especially as the number of markets increases. This is consistent with the role of cross-market variation in the outer problem: more markets improve the micro moment for $g_0$ and the price-side moment for $\psi_0$.
The market-level objects behave in the way predicted by the profiling argument. The profiled composite intercept $\hat a_t$ improves as the within-market sample size increases, because each $a_{0t}$ is learned from the consumers inside market $t$. Moving from $(T,n_t)=(50,100)$ to $(200,500)$ reduces the median $a_t$ error from 0.707 to 0.293. The recovered shock $\hat h_t=\hat a_t-\hat\psi(P_t,X_t)$ also improves over the grid, falling from 0.532 at $(50,100)$ to 0.255 at $(200,500)$. This decline reflects both components of the estimator: larger $n_t$ improves the profiled intercepts, while larger $T$ improves the price-side decomposition of $a_{0t}$ into $\psi_0(P_t,X_t)$ and $h_{0t}$.
The finite-sample pattern is not perfectly monotone cell-by-cell. This is especially true for the sup-norm errors for $g_0$ and $\psi_0$, because a single grid point can determine the reported value. This should not be a cause for major concern. The purpose of the simulation is not to estimate a rate or to claim monotonicity in every finite-sample cell; rather, it is to verify the qualitative consistency logic of the estimator. The relevant diagnostic is that the largest-errors-over-the-grid for the common functions, and the average errors for the generated market-level objects, are smaller in the high-information designs than in the low-information designs.
Taken together, the simulation supports the main consistency theorem. The inner profile behaves as a well-posed market-by-market share-matching step, the price-side SMD/GMM block recovers the systematic price component using excluded instruments, and the feasible estimator jointly recovers $g_0$, $\psi_0$, $a_{0t}$, and $h_{0t}$ as the amount of within-market and cross-market information increases.
The second exercise studies whether the flexibility of the price-side function matters for economically relevant counterfactuals. I compare the proposed estimator with three standard benchmarks: aggregate Logit-IV, BLP random-coefficients logit, and Micro-BLP. The counterfactual DGP has three inside goods and an outside option. Individual utility is \[ U_{ijt} = g z_i+\psi_{0j}(P_t,X_t)+h_{jt}+\varepsilon_{ijt}, \qquad j=1,2,3, \qquad U_{i0t}=\varepsilon_{i0t}, \] where $\varepsilon_{ijt}$ is Type-I extreme value, $z_i\sim N(0,1)$, and the heterogeneity loading is $g=0.3$. Thus the heterogeneity component is deliberately simple: the scalar term $g z_i$ enters each inside good symmetrically. Equivalently, \[ g_0(z_i,X_t)=g z_i\cdot \mathbf 1_J . \] This specification creates observed heterogeneity in consumers' inside-good tastes relative to the outside option, but it does not create product-specific heterogeneity across inside goods.
This design is intentionally favorable to the BLP benchmarks. A conventional random-coefficients logit model can represent this source of heterogeneity as a random coefficient on the inside-good constant. Moreover, Micro-BLP is given additional micro information, described below, that is directly informative about this heterogeneity. Hence, the simulation does not maliciously disadvantage BLP by giving the proposed estimator a rich nonparametric heterogeneity function that BLP cannot approximate. Instead, the main source of misspecification for the benchmarks is isolated on the price side: the true function $\psi_0(P_t,X_t)$ is nonlinear in prices, while aggregate Logit-IV, BLP, and Micro-BLP impose the standard linear price term in mean utility. The exercise therefore asks whether flexible recovery of the price-side function matters for counterfactual predictions, even in a setting where the heterogeneity structure is chosen to be relatively favorable to random-coefficients logit.
Prices are endogenous in the DGP. They depend both on excluded cost shifters and on the unobserved demand shocks. The excluded cost shifters are used as instruments. The simulation also draws first and second choices for every consumer. First choices generate the market shares used by all estimators. Second choices are used to construct a micro moment for Micro-BLP: the share of inside-good buyers whose second choice is the outside option. This moment is informative about the inside-versus-outside margin and therefore about the random-coefficient heterogeneity in $g z_i$. In this sense, Micro-BLP is a strong benchmark as it is given micro information that is well aligned with the particular dimension of heterogeneity present in the DGP.
The exercise considers two regimes. In the oracle regime, all estimators are given the true counterfactual demand shocks. This isolates the consequences of functional-form misspecification. In the full regime, each estimator must recover its own demand shocks from the simulated data. The main text reports the full regime because it is the relevant empirical case. Each simulation uses $T=300$ markets, $n_t=500$ consumers per market, and 40 Monte Carlo replications. The sieve order for the proposed estimator is selected by cross-validation and is not chosen using knowledge of the truth.
Figure (ref) reports three counterfactual objects in the full regime. Panel (a) reports the error in predicted counterfactual share changes after a price shock to the first inside good. Panel (b) reports the error in diversion ratios. Panel (c) reports the error in welfare changes, measured as the change in expected inclusive value. In all three panels, displayed lines are medians across replications and shaded regions are 10--90 percent bands.
The flexible estimator performs best on all three reported counterfactual objects. For share changes, the error of the sieve estimator increases much more slowly as the counterfactual price shock becomes larger. This is the object most directly tied to recovery of the demand surface away from observed prices. If the nonlinear price function is approximated by a linear price index, share predictions can deteriorate quickly as the counterfactual moves farther from the observed price distribution. The proposed estimator is designed to avoid this restriction by estimating $\psi_0(P_t,X_t)$ flexibly.
The same pattern appears for diversion ratios. Diversion is a local substitution object, so errors in the slope and curvature of the demand surface are amplified. The BLP and Micro-BLP benchmarks improve on aggregate Logit-IV by allowing consumer heterogeneity on the inside-versus-outside margin. However, that is not the main source of counterfactual difficulty in this DGP. The difficulty is that the price-side component of utility is nonlinear. Because the BLP benchmarks retain a linear price term, their random-coefficients structure cannot fully correct the misspecification in the price surface. The proposed estimator remains below the Logit-IV, BLP, and Micro-BLP benchmarks throughout the plotted range.
The welfare results are smoother, as expected. Welfare changes aggregate over the demand system and are therefore less sensitive than diversion ratios to local errors at a particular price point. Nevertheless, the same ranking appears: the flexible estimator delivers the lowest median welfare error across the counterfactual grid. This suggests that the gains from flexible price-side estimation are not limited to pointwise share predictions. They also matter for welfare calculations, even though welfare is a more averaged object.
These results should be interpreted narrowly. The simulation is not designed to show that BLP performs poorly in general. On the contrary, the heterogeneity structure is chosen to be relatively favorable to BLP and Micro-BLP. The point is instead more specific: even when the random-coefficients component is easy for BLP-style estimators to represent, counterfactual performance can still suffer if the systematic price-side function is misspecified. The proposed estimator performs well in this exercise because it targets precisely this object. The simulation therefore illustrates the economic value of the estimator developed in this paper; that is, flexible recovery of $\psi_0(P_t,X_t)$ can matter substantially for the counterfactual objects that motivate IO demand estimation, including share changes, diversion ratios, and welfare.
This paper develops a profiled sieve minimum-distance estimator for a semi-nonparametric differentiated products demand model with micro data. It is motivated by the observation that micro-level choice data contain within-market variation that is absent from market-level demand systems. Consumers in the same market face the same prices, product characteristics, and market-level demand shocks, but differ in observed characteristics. This variation identifies the consumer-heterogeneity component of demand and the composite market intercept. Excluded price instruments then separate the composite intercept into a systematic price-side function and a structural demand shock. The estimator formalizes this logic by combining a within-market profiling step with a cross-market conditional-moment step.
The main result establishes uniform consistency of the common structural functions. The profiled composite intercepts and recovered structural shocks are also consistently recovered. The Monte Carlo evidence supports this logic: the common functions and the generated market-level objects are recovered more accurately as both the number of markets and the number of consumers per market increase.
The paper is intentionally limited in scope. Its purpose is to show that the Berry--Haile micro-data identification argument can be turned into a feasible estimator and that the resulting profiled estimator is uniformly consistent. Several important issues are left open. First, the maintained model uses the multinomial-logit map as a tractable link function. This makes the profile step convex and transparent, but it also imposes the familiar substitution restrictions associated with logit demand. A natural next step is to relax this restriction and move closer to the fully nonparametric demand link in the Berry--Haile framework. Second, the paper studies consistency rather than rates, limiting distributions, or confidence sets. The assumptions are therefore designed to control the profile step and the moving SMD criterion at the level needed for consistency, not to drive an asymptotic distribution result. Third, the price-side block inherits the usual difficulties of nonparametric IV and conditional-moment estimation. Completeness is a strong identification condition, and finite-sample performance will depend on the strength of the instruments, the choice of basis, the amount of regularization, and the growth rates of the sieve dimensions.
Future work should therefore develop the asymptotic distribution theory for the estimator. This requires understanding the joint contribution of the within-market profile error, the growing micro moment, and the price-side SMD projection. In the current estimator, the profiled intercepts are treated in the way needed for consistency, but they are not bias-corrected. This is likely not enough for inference. The profile step generates an additional estimation effect, which may contribute to first-order bias or otherwise complicate the limiting distribution of smooth functionals even when it vanishes asymptotically. A promising direction is to construct a Hahn--Newey-style debiased version of the profiled estimator, following the logic of bias correction for nonlinear models with many nuisance parameters hahn_jackknife_2004. Such a correction would aim to remove the leading effect of estimating the market intercepts and could provide the asymptotic linearity needed for rates of convergence, asymptotic normality, and feasible inference. This would connect the present consistency result to the broader sieve-inference literature for semi/nonparametric conditional moment models, including chen_sieve_2015.
One major limitation of the method, even if the above advancements are made, is the curse of dimensionality; see Appendix (ref). Sieve-based nonparametric estimators require extremely large amounts of data once the covariate space or the number of products begins to grow -- even to relatively modest degrees. Methods that impose shape restrictions or truncate the sieve-approximation are viable, but either run at odds with the spirit of flexible demand estimation or may not be guaranteed to approximate the unknown function. Hence, future work might consider targeting lower dimensional counterfactual functionals directly rather than estimating a globally flexible demand system.
More broadly, the estimator developed here should be viewed as a first step toward a flexible micro-data approach to demand estimation. The paper shows that within-market consumer heterogeneity and cross-market price instruments can be combined in a single profiled SMD framework, and that this framework consistently recovers the structural functions of interest. The next step would be to turn this consistency result into an asymptotic distribution result and to study the counterfactual objects that motivate flexible demand estimation in industrial organization: elasticities, diversion ratios, pass-through, markups, and welfare. Those objects depend on the same structural functions estimated here. Establishing valid inference for them makes the estimator practical.