Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
107,460 characters · 26 sections · 42 citation commands
From Unstructured Data to Demand Counterfactuals: Theory and Practice
\onehalfspacing
\thispagestyle{empty}
\setcounter{page}{1}
\pagenumbering{arabic}
Many questions in economics and other social sciences require researchers to estimate demand for differentiated products. A common strategy is to estimate discrete-choice models which specify the utility of a product as a function of the product's price and other observable attributes, often allowing for rich types of consumer heterogeneity (BLP_1995,BLP_micro_2004, henceforth BLP). Among other applications, this approach has been used to study the impact of horizontal mergers nevo2000mergers, new product launches hausman1994valuation, petrin2002quantifying, trade policy goldberg1995product, school choice bayer2007unified, neilson2017targeted, two-sided markets fan2013ownership, lee2013vertical, and the evolution of markups over time grieco2024evolution.
The success of demand models in predicting (counterfactual) quantities of interest (counterfactuals hereafter) hinges on their ability to capture substitution patterns. Doing so requires using product attributes as model inputs that correctly reflect the underlying dimensions of differentiation. Choosing the correct attributes to use as model inputs poses a fundamental measurement challenge berry2021foundations. First, consumer choices are often driven by hard-to-quantify characteristics, such as visual design, user friendliness, or style. In these cases, a growing literature shows that product images, descriptions, and reviews contain valuable information to capture substitution patterns compiani2025demand, han2025copyright, lee2025generative. Consumer surveys may also provide measures of product differentiation Magnolfi_et_al_2022. To use these high-dimensional, unstructured data in demand models, researchers often transform them into lower-dimensional numerical variables, or embeddings, using machine learning (ML) methods. Second, even numeric attributes could be mismeasured nevo2001measuring, allcott2014gasoline, or could be high-dimensional and collinear, requiring dimension-reduction Backus_et_al_2021. In all of these cases, the variables used as inputs in the demand model are proxies for the true attributes that drive consumer choices. It is essential that these proxies adequately capture the true dimensions of differentiation: poor proxies can lead to biased estimates of demand model parameters and, in turn, biased counterfactuals.
In this paper, we propose a simple, post-estimation bias correction for counterfactuals. We take the naive estimator that treats the proxies as if they were the true dimensions of differentiation (as is implicitly done and reported in practice) and add a correction term designed to achieve two goals. First, it is chosen to mitigate the bias in counterfactuals arising from mismeasurement of the true dimensions of differentiation. Second, it is chosen so that the bias-corrected estimator is efficient, meaning that it has the lowest possible asymptotic variance among a broad class of estimators. We also provide simple formulas for standard errors, making valid inference easy. In addition, we show how two simple diagnostics can be used to help assess the adequacy of different proxies in capturing substitution. Together, these diagnostics help guide the choice of how many and which proxies/attributes should be included in the model---practical questions that researchers need to answer in any instance. Answering these questions is further complicated by the fact that, unlike standard prediction problems, the counterfactual is not observed in the data.
We develop the bias corrections and diagnostics for two widely-used empirical frameworks. The first follows BLP_1995,BLP_micro_2004: prices vary across markets, instrumental variables are used to address price endogeneity, and market-level data may be supplemented with individual choice data for a subset of markets. Many papers in industrial organization (IO) and fields using IO tools fit in this category. The second framework consists of models estimated on individual choice data with product-level fixed effects, which are very common in marketing applications (see dube2019handbook for a review). While we illustrate our approach for workhorse specifications (e.g., mixed logit with normal random coefficients), the method does not rely on specific parametric functional forms.
The bias corrections and diagnostics are computationally light and integrate easily into the standard demand estimation workflow. The bias corrections take as inputs the naive parameter estimates, which treat the proxies as the true dimensions of differentiation, and require neither bootstrapping nor any optimization. Similarly, the diagnostics are Lagrange Multiplier (LM) statistics evaluated at the naive estimates. All bias corrections, standard errors, and diagnostics admit closed-form expressions. These involve first derivatives of choice probabilities and counterfactuals, which are easily computed using automatic differentiation.
The key insight underlying our approach is that mismeasurement of the true dimensions of differentiation using proxies induces a form of model misspecification, as distinct from a measurement error problem.\footnote{In our setting, the relevant unit of observation is the market and/or individual level, whereas mismeasurement occurs at the product level. By contrast, mismeasurement is at the observation level in a standard measurement error problem. Misspecification and measurement error both cause bias, but do so for different reasons and require different corrections.} We address this misspecification by reparameterizing the model with a composite parameter that captures how proxies interact with structural parameters to affect utilities. Different proxies correspond to different values of the composite parameter. Framing the model in this way allows us to target counterfactuals using standard two-step estimation methods, where the first-step estimator corresponds to the composite parameter value pinned down by the proxies and naive parameter estimates.
An advantage of this approach is that it allows us to correct bias while remaining agnostic about the form of mismeasurement. This is particularly valuable because proxies are often obtained via black-box ML models, making it difficult to justify specific assumptions on the nature of mismeasurement. Importantly, the approach accommodates proxies that depend on the choice data, such as when they are obtained by fine tuning ML models on the same data that is used to estimate the demand model. For instance, researchers may fine tune neural networks or LLMs to obtain embeddings of product descriptions and images that better fit the observed substitution patterns than those produced by off-the-shelf algorithms. Moreover, because we correct how proxies and structural parameters jointly affect utility rather than the proxies themselves, our approach does not require practitioners to take a stand on the units of the proxies and/or true dimensions of differentiation. This is especially important for proxies for hard-to-quantify characteristics like visual design or user friendliness that lack natural units of measurement.
Simulations confirm that the bias correction improves performance for a range of levels of mismeasurement of the true dimensions of differentiation. Specifically, the corrected estimator has lower bias and lower variance than the naive estimator across all levels of mismeasurement. The bias correction leads to slightly higher variance in the knife-edge case in which the proxies perfectly capture differentiation, but the efficiency loss is small.\footnote{The fact that there is an efficiency loss in this knife-edge case is to be expected: the naive approach maintains the assumption that the proxies are measured without error, whereas the corrected estimator does not. What is surprising is that the efficiency loss is relatively small.} Simulations also confirm that our diagnostics convey useful information for selecting which proxies to use when estimating counterfactuals.
Finally, using the experimental data from compiani2025demand, we show that the bias correction materially improves the model's ability to predict counterfactual choices following product removals. To this end, we leverage the fact that the data features both consumers' first and second choices. We estimate model parameters on the first choice data alone, then compare how well the naive and bias-corrected estimators predict second choices. The second-choice data provide a ground truth to assess the effectiveness of our approach. The bias correction meaningfully improves the model's ability to predict a product's closest substitute, improving the hit rate from 40% to 70% for our preferred specification. Further, our diagnostics correctly identify the set of proxies that perform best at the counterfactual prediction task, indicating that they can be valuable tools for practitioners.
We emphasize that our approach is also helpful for practitioners using standard numeric attributes. As noted above, mismeasurement may be a concern even in this case, particularly when dimension-reduction methods are used to shrink the attribute set. Our bias corrections provide a practical remedy. Further, the choice of which attributes to include is generally ad hoc even with numeric attributes. Our diagnostics help guide practitioners in making these decisions. Beyond mismeasurement concerns, our approach yields easy-to-compute, efficient estimators of counterfactuals (even when model parameters are estimated inefficiently) across many empirical settings, including combined market-level and microdata petrin2002quantifying, BLP_micro_2004. Our standard-error formulas also allow easy inference without bootstrapping. To the best of our knowledge, these contributions are new.\footnote{grieco2025optimal study efficient estimation of model parameters in mixed logit models with combined market-level and microdata. Our focus is instead on efficient estimation of counterfactuals. For counterfactuals that depend on data moments in addition to model parameters (e.g., average welfare and average price elasticity), efficient estimators of model parameters do not necessarily lead to efficient estimators of counterfactuals brown1998efficient,ai2012semiparametric.}
Our approach is related to double/debiased ML (DML), which has recently been used in single-equation demand estimation with unstructured data bach2024adventures. Both aim to estimate a target parameter in the presence of nuisance parameters. In our setting, the target is the counterfactual and the nuisance are both the latent dimensions of differentiation, which are “estimated” using proxies, and the demand model parameters.\footnote{In contrast, bach2024adventures treats the embeddings as perfect proxies for product attributes. Correspondingly, it uses DML to correct the estimation of nuisance functions as for partially linear regression, not to correct for mismeasurement of product attributes.} Standard DML methods typically require models for the nuisance parameters and access to the data used to estimate them. In contrast, our approach accommodates proxies that are the outputs of black-box ML models trained on data to which the researcher might have limited to no access. To do so, we reparameterize the model via a composite parameter and rely on standard two-step estimation methods, including an orthogonalization step which shares similarities with DML.
A recent literature recognizes that naively treating ML-generated variables as data leads to measurement-error bias and develops corrections for it. But as noted above, the problem we study is one of model misspecification rather than measurement error, since the unit of observation (markets and/or individuals) is different from the unit of mismeasurement (products). Moreover, almost all strategies in this literature rely on validation data linking ML-generated variables and their ground-truth values.\footnote{See, e.g., fong2021machine,allon2023machine,angelopoulos2023prediction,egami2023using,zhang2023debiasing,carlson2025unifying and references therein. These works build on an earlier literature on auxiliary data chen2005measurement,chen2008semiparametric.} In our setting, however, the true dimensions of differentiation are latent and can at best be only imperfectly proxied via survey data, rendering these methods inapplicable. One exception is battaglia2024inference who develop analytical bias corrections without validation data, but their approach is specific to linear regression.
The remainder of the paper is structured as follows. Section (ref) presents our bias corrections and diagnostics for BLP-type models, while Section (ref) does the same for models with individual-level choice data and product fixed effects. Each section first presents the model, develops the bias corrections and diagnostics, and concludes with a practitioner's guide detailing the steps involved and giving practical recommendations. Simulations and the empirical application are presented in Sections (ref) and (ref), respectively, with additional empirical results deferred to Appendix (ref). Section (ref) presents all theoretical results while all proofs are presented in Appendix (ref).
We first consider a setting where prices vary at the market level and identification is achieved through instruments.
Following an established literature berry2014identification, freyberger2015asymptotic, we assume that the researcher has data from a large number $T$ of markets in which (subsets of) $J$ goods are sold.\footnote{For simplicity, we assume that $J$ is fixed. It is straightforward to extend our approach to asymptotic thought experiments where $J$ grows slowly with the number of markets $T$.} In addition to the outside option (denoted by 0), each market $t$ features products $\mathcal{J}_t \subseteq \left\{1, \ldots, J \right\} $, for which the researcher has access to data on prices $ p_t = (p_{jt})_{j \in \mathcal{J}_t}$, exogenous product attributes $ x_t = (x_{jt})_{j \in \mathcal{J}_t}$, and market shares $s_t = (s_{jt})_{j \in \mathcal{J}_t}$. Consumer choices are also driven by unobserved quality levels $\xi_t = (\xi_{jt})_{j \in \mathcal{J}_t}$. The model predicts market shares as a function of $p_t$, $x_t$, $\xi_t$, and a parameter vector $\theta$:
Prices $p_t$ are endogenous and may be correlated with the unobservables $\xi_t$. To address this, we rely on a vector of instrumental variables $w_t = (w_{jt})_{j \in \mathcal J_t}$ that satisfy
where $z_{jt} \equiv (x_{jt}, w_{jt})$. We also accommodate “micro BLP” settings BLP_micro_2004,berry2024nonparametric where, in addition to the above, individual-level data on choices and demographics are available in a subset of markets. The microdata consists of choice indicators $d_{it} = (d_{ijt})_{j \in \mathcal J_t}$ taking the value $1$ if $i$ chose $j$ in market $t$ and $0$ otherwise, and demographic variables $y_{it}$ that vary at the consumer level, such as income, and/or $\bar y_{ijt}$ that vary at the product-consumer level, such as distance between a household's home and a school or a hospital. We assume the microdata is available for a fixed set of markets which, without loss, we label $t=1, \ldots, \tau$.\footnote{The case where microdata is present for all $T$ markets is simpler and the associated derivations are available upon request.} We treat the microdata within each market $t$ as a repeated cross section of size $N_t$.
This model subsumes many empirical specifications used in the literature. We provide two simple examples below to fix ideas.
So far the model is standard. Our point of departure is to partition \[ x_{jt} \equiv (\bar x_{jt},e_j), \] where $\bar x_{jt}$ is a vector of conventional observed product attributes, such as product size, and $e_j$ is an $r$-vector of product attributes, such as visual design or user friendliness, that are difficult to capture using standard numeric data. Accordingly, we treat $e = (e_j')_{j=1}^J$ as known to the consumer but latent to the researcher. What is available to the researcher are proxies $\tilde e = (\tilde e_j')_{j=1}^J$ for the true underlying $e$. Note that $e$ does not vary across markets, consistent with the fact that difficult-to-quantify characteristics such as visual design or user friendliness are often fixed product characteristics.
In the leading case we study, the econometrician observes unstructured data $U_j$ and computes a low-dimensional representation $\tilde e_j$ of $U_j$, often referred to as an embedding, via ML methods. In this scenario, the embeddings $\tilde e_j$ act as proxies for the true latent $e_j$. To maximize generality, we stay agnostic on the form that $U_j$ takes. It could be text (product descriptions and reviews), images, audio/video components, or a combination thereof compiani2025demand,han2025copyright, or consumer preferences inferred from surveys Magnolfi_et_al_2022. Similarly, we are agnostic on the ML method used to compute $\tilde e_j$.
Unlike prior work, we wish to account for the fact that proxies $\tilde e_j$ are not the ground truth but rather are approximations to the true latent $e_j$. Different ML methods correspond to different approximations and produce different biases in downstream estimates of counterfactuals. Our first main goal is to develop estimators of model parameters and counterfactuals that are immune to this bias. Another goal is to shed light on what a “good” proxy might look like from the perspective of estimation and inference on counterfactuals. This objective is fundamentally different from the standard problem of choosing proxies for a prediction problem, since in our case the counterfactual is not observed in the data.
We consider a broad class of counterfactuals that can be written as
where the expectation is over the distribution of $(p_t, \xi_t, \bar x_t)$ across markets $t$. For instance, $\kappa$ might represent an average price elasticity, average equilibrium price or consumer welfare measure (possibly after a counterfactual change on the supply side). Expression (ref) also subsumes counterfactuals for a specific market, such as the price elasticity at a given $(p, \xi, \bar x)$. In this case, $k$ is a deterministic function of $\theta$ and the expectation becomes redundant. As we discuss below, $\kappa$ could also represent certain elements of $\theta$, such as the average price coefficient.
Given the observed aggregate data and a candidate set of proxies $\tilde e$, estimation typically proceeds using GMM based on the moment\footnote{With slight abuse of notation, we now write $\sigma_j$ as functions of $\bar x_t$ and $e$, with the understanding that $\sigma_j$ depends on $e$ only through $(e_j)_{j \in \mathcal J_t}$, and similarly for $\hat \xi_t$.} \[ \frac 1T \sum_{t=1}^T Z_t \hat \xi_{t}(s_t,p_t, \bar x_t, \tilde e; \theta), \] where $\hat \xi_{t}(s_t,p_t, \bar x_t, e; \theta) = (\hat \xi_{j}(s_t,p_t, \bar x_t, e; \theta))_{j \in \mathcal J_t}$ is defined implicitly via
and $Z_t$ is a $\dim(z) \times |\mathcal J_t|$ matrix whose columns contain $z_{jt}$ for $j \in \mathcal J_t$. Here $z_{jt}$ may be the original instruments in ((ref)) as well as transformations thereof, such as when sieves are used to approximate optimal instruments. When the researcher also has microdata, the GMM criterion can be combined with a minimum-distance criterion based on the micro moments conlon2025incorporating. Let $\bar m_t = \frac{1}{N_t} \sum_{i=1}^{N_t} m_{it}$, where $m_{it} = m(y_{it},\bar y_{it}, d_{it})$ is a known function of demographic and choice data for individual $i$ in market $t$. Let $m(p_t, \xi_t, \bar x_t, e; \theta)$ denote the model-implied expectation of $m_{it}$ conditional on the market-level data: \[ m(p_t, \xi_t, \bar x_t, e; \theta) = \mathbb{E}[m_{it} | p_t, \xi_t, \bar x_t, e]. \] Thus \[ \bar m_t - m(p_t, \hat \xi_{t}(s_t, p_t, \bar x_t, \tilde e; \theta), \bar x_t, \tilde e; \theta), \quad t = 1,\ldots,\tau, \] give an additional set of micro-moments to match when estimating $\theta$.
Given an estimate $\hat \theta$ of $\theta$, the counterfactual $\kappa$ is usually estimated as
We call this the naive estimator of $\kappa$ since it does not account for the fact that the proxies $\tilde e$ might differ from the true latent attributes $e$. This mismeasurement has the potential to affect the estimator via two channels: (i) directly, since $\tilde e$ is an argument of $k$, and (ii) indirectly through both $\hat \theta$ and $\hat \xi_t$.
We now introduce a bias-correction procedure that is designed to mitigate the bias from using $\tilde e$ in place of the true $e$. We begin by restricting attention to models in which the attributes $e$ and model parameters $\theta$ enter choice probabilities ((ref)) via a lower-dimensional composite parameter \[ \gamma \equiv \gamma(\theta, e). \] Many models feature this property. We illustrate it in the BLP example.
As can be seen from this example, parameters that do not interact with $e$ are left unchanged, as is the case for the average price coefficient $\bar \alpha$, for instance. For the remaining components that interact with $e$, we expand the parameter space to capture the effect of joint shifts in $e$ (as, for instance, when $\tilde e$ is used in place of $e$) and/or $\theta$. A similar reparameterization for Example (ref) is provided in Appendix (ref).
This reparameterization allows us to simplify notation as follows. First, we note that the right-hand side of ((ref)) depends on $(\theta, e)$ only via $\gamma(\theta,e)$. Thus we write $\hat \xi_{jt}(\gamma(\theta, e)) = \hat \xi_{j}(s_t,p_t, \bar x_t, e; \theta)$ for $j \in \mathcal J_t$ (suppressing dependence on $s_t,p_t, \bar x_t$) and let $\hat \xi_t(\gamma(\theta, e)) = (\hat \xi_{jt}(\gamma(\theta, e)))_{j \in \mathcal J_t}$. We similarly restrict attention to counterfactuals that depend on $(\theta, e)$ only via $\gamma(\theta, e)$ and write $k_t(\gamma(\theta, e)) = k(p_t, \hat \xi_t(\gamma(\theta, e)), \bar x_t, e; \theta)$. This includes many counterfactuals of interest, such as elasticities with respect to prices or $\bar x$, equilibrium prices, and welfare changes associated with changes in prices or $\bar x$. It precludes quantifying the effect of changes in the latent attributes $e$ or measuring heterogeneity in preferences for $e$, but these are of little meaning when $e$ has no natural scale or interpretation. When microdata are also available, we note that the choice probabilities in ((ref)) depend on $(\theta, e)$ only via $\gamma(\theta, e)$ and write $m_t(\gamma(\theta, e)) = m(p_t, \hat \xi_t(\gamma(\theta, e)), \bar x_t, e; \theta)$. Finally, we let $\hat \gamma = \gamma(\hat \theta, \tilde e)$ denote the value of $\gamma$ at the estimated structural parameters $\hat \theta$ using the candidate proxies $\tilde e$.
With this notation, we can define the bias-corrected estimator
where $\hat c$ is a $\dim(z) \times 1$ vector of weights for the aggregate moments and $\hat d_1,\ldots,\hat d_{\tau}$ are $\dim(m) \times 1$ vectors of weights for the micro moments. Without microdata, the bias-corrected estimator is simply
We give closed-form expressions for $\hat c$ and $\hat d_1,\ldots,\hat d_\tau$ below. The bias corrections are easy to implement: they simply take the naive estimator $\frac 1T \sum_{t=1}^T k_t(\hat \gamma)$ and add a weighted average of the estimation moments. As such, they require minimal computation beyond what is needed to estimate model parameters $\theta$ in the first place.
The idea behind ((ref)) is to choose the weights $\hat c$ and $\hat d_1,\ldots,\hat d_\tau$ so that $\hat \kappa_{bc}$ does not depend on $\hat \gamma$ to first order.\footnote{There is a long tradition of using corrections such as these in two-step estimation. See, e.g., andrews1994asymptotics and newey1994asymptotic. Of course, similar debiasing ideas underlie the DML literature.} This means that, to first order, $\hat \kappa_{bc}$ behaves like the right-hand side of (ref) with $\hat \gamma$ replaced by the true value $\gamma_0 = \gamma(\theta_0, e_0)$, where $\theta_0$ are the true structural parameters and $e_0$ are the true latent attributes. In doing so, this purges the first-order effect of proxying $e$ with $\tilde e$. In Section (ref), we show that, as a result, the asymptotic distribution of $\hat \kappa_{bc}$ is centered around the true counterfactual $\kappa_{0}$ and does not depend on $\hat \gamma$. This has two important implications: first, $\kappa_{bc}$ is immune to any bias arising from proxying $e$ with $\tilde e$, and second, standard errors do not need to be corrected when $\tilde e$ is chosen in a data-dependent way, e.g., by fine tuning an ML model on choice data.
For the intuition, consider the case without microdata. A Taylor expansion of the naive estimator yields \[ \hat \kappa \approx \frac{1}{T} \sum_{t=1}^T \left( k_t(\gamma_0) + \frac{\partial k_t(\hat \gamma)}{\partial \gamma'}(\hat \gamma - \gamma_0)\right). \] We wish to eliminate the second problematic term depending on $\hat \gamma - \gamma_0$. To do so, we replicate the dependence of $\hat \kappa$ on $\hat \gamma - \gamma_0$ using the estimation moments. This means that the vector of weights $\hat c$ will be chosen so that \[ \frac{1}{T} \sum_{t=1}^T \frac{\partial k_t(\hat \gamma)}{\partial \gamma'} = \hat c' \left( \frac{1}{T} \sum_{t=1}^T Z_t \frac{\partial \hat \xi_t(\hat \gamma)}{\partial \gamma'} \right). \] There are many different weights $\hat c$ with this property; the weights introduced below are designed to minimize the asymptotic variance of $\hat \kappa_{bc}$. Substituting in the previous display and “undoing” the Taylor expansion, we get \[
\] This suggests that to correct bias we want to adjust the naive estimator by subtracting $\frac{1}{T}\sum_{t=1}^T ( \hat c' Z_t \hat \xi_t(\hat \gamma) - \hat c' Z_t \hat \xi_t(\gamma_0) )$. The final (infeasible) term depending on $\gamma_0$ has mean zero by virtue of (ref), so we drop it, leading to the corrected estimator (ref). This correction therefore ensures that, to first order, $\hat \kappa_{bc}$ depends only on $\gamma_0$.
What assumptions are needed for this result? Besides standard regularity conditions, we require that the discrepancy between $\hat \gamma$ and the true value $\gamma_0$ not be too large relative to sampling error (see Section (ref) for a discussion). Figure (ref) shows that, in simulations, $\hat \kappa_{bc}$ has negligible bias up to moderate amounts of mismeasurement (and thus moderate deviations of $\hat \gamma$ from $\gamma_0$), while for high amounts of mismeasurement (and thus large deviations of $\hat \gamma$ from $\gamma_0$), the bias of $\hat \kappa_{bc}$ is still well below that of the naive estimator. We also implicitly require that the $e_j$ and $\tilde e_j$ have the same dimension $r$. In the next subsection, we provide two diagnostics to help researchers choose among proxies so that both these conditions are plausibly satisfied.
To introduce the expressions for the weights $\hat c$ and $\hat d_1,\ldots,\hat d_\tau$, let \[
\] denote the sample variance of the estimation moments, where $\bar g = \frac 1T \sum_{t=1}^T Z_t \hat \xi_t(\hat \gamma)$. Define
where $\hat k = \frac 1T \sum_{t=1}^T \dot k_t(\hat \gamma)$ is $\dim(\gamma) \times 1$, $\hat K = \frac 1T \sum_{t=1}^T k_t(\hat \gamma) Z_t \hat \xi_t(\hat \gamma)$ is $\dim(z) \times 1$, $\hat G = \frac 1T \sum_{t=1}^T Z_t \dot \xi_t(\hat \gamma)$ is $\dim(z) \times \dim(\gamma)$, $\hat M_t = \dot m_t(\hat \gamma)$ is $\dim(m) \times \dim(\gamma)$, and $\dot k_t(\gamma) = \frac{\partial k_t(\gamma)}{\partial \gamma}$, $\dot \xi_{t}(\gamma)' = \frac{\partial \hat \xi_{t}(\gamma)'}{\partial \gamma}$, and $\dot m_{t}(\gamma)' = \frac{\partial m_{t}(\gamma)'}{\partial \gamma}$ are $\dim(\gamma) \times 1$, $\dim(\gamma) \times J_t$, and $\dim(\gamma) \times \dim(m)$, respectively. The weights to plug into (ref) are
and
Without microdata, $\hat d_1 = \ldots = \hat d_\tau = 0$ and $\hat H = \hat G' \hat V^{-1} \hat G$.
We defer formal statements of the results sketched out above to Section (ref), and instead highlight a few key properties of the corrected estimator.
Next, we propose two diagnostics that practitioners can use to assess the suitability of a candidate set of proxies $\tilde e$. The first speaks to whether $\hat \gamma = \gamma(\hat \theta, \tilde e)$ is sufficiently close to the truth $\gamma_0 = \gamma(\theta_0, e_0)$. The second addresses the question of whether the dimension of $\tilde e$ matches that of $e_0$. Both diagnostics are based on LM statistics evaluated at $\hat \gamma$, so they require minimal additional computation.
The bias correction is based on linearization and thus requires the discrepancy between $\hat \gamma$ and $\gamma_0$ to not be too large relative to sampling error, as discussed above. Here we show that a simple LM statistic can be used to validate this condition. This diagnostic is also helpful to guide the choice among sets of embeddings (e.g., embeddings obtained from various data source and/or ML model combinations).
The first diagnostic is
where $\hat S = \hat G' \hat V^{-1} ( \frac 1T \sum_{t=1}^T Z_t \hat \xi_t(\hat \gamma)) + \sum_{t=1}^\tau \hat M_t' \hat V_t^{-1} (m_t(\hat \gamma) - \bar m_t)$ represents the “score” at $\hat \gamma$ and $\hat H$ is given in ((ref)). This diagnostic can be interpreted as the LM statistic in a test of the null hypothesis that $\gamma_0$ is in the set of composite parameters spanned by the candidate proxies $\tilde e$, ignoring the fact that $\tilde e$ is possibly stochastic.\footnote{Formally, a test of $\mathbb{H}_0: \gamma_0 \in \Gamma(\tilde e) := \{\gamma(\theta,\tilde e) : \theta \in \Theta\}$ against the alternative $\mathbb{H}_1: \gamma_0 \in \Gamma \setminus \Gamma(\tilde e)$, where $\Theta$ and $\Gamma$ are the parameter spaces for $\theta$ and $\gamma$, respectively.} Proposition (ref) below shows that $LM_1$ behaves like $\|\sqrt T (\hat \gamma - \gamma_0)\|^2$ as the sample size grows large. In other words, researchers can validate the assumption that the discrepancy between $\hat \gamma$ and $\gamma_0$ is sufficiently small, as needed in Proposition (ref), by checking whether $LM_1$ is below a threshold. In particular, for any sequence $C_T = o(T^{1/4})$, we have that $LM_1 \leq C_T^2$ implies $\|\hat \gamma - \gamma_0\| \leq \mathrm{constant} \times C_T/\sqrt T = o(T^{-1/4})$ with probability approaching one. It can also be shown under a slight strengthening of the conditions of Proposition (ref) that $LM_1$ can be used to bound $\|\tilde e - e_0\|$, providing a measure of how well the proxies $\tilde e$ capture the true latent attributes $e$ driving consumer choices. This is especially useful as the true $e_0$ can never be observed. As a result, $LM_1$ also serves as a model-based criterion to target when fine tuning.
When estimating this type of model, practitioners also have to choose how many attributes to include. In our notation, this corresponds to choosing the dimension of $\tilde e$. This is particularly delicate in the cases where the candidate proxies don't have a natural economic interpretation, as is typically the case when they are obtained from black-box ML algorithm or by applying principal component analysis (PCA) to a rich set of numeric attributes. For example, in the application of Section (ref), we use PCA to reduce the dimensionality of proxies obtained from pre-trained algorithms; the relevant question is then how many principal components to include in the model.
We provide guidance on this by again considering a diagnostic based on an LM statistic. The idea is to augment $\tilde e$ with a vector $\eta \in \mathbb R^J$ representing some excluded but potentially important product attributes and augment $\theta$ with an additional component $\psi \in \Psi$ representing coefficients on $\eta$. The second diagnostic is based on an LM statistic for the null that $\psi = 0$. Like $LM_1$, this diagnostic depends on $\hat \theta$ only and therefore requires minimal additional computation.
To introduce the diagnostic, we extend $\gamma$ to $\zeta = (\gamma,\psi)$. With slight abuse of notation, we now write $\hat \xi_{jt}$ and $m_t$ on this extended space as functions $\hat \xi_{jt}(\zeta;\eta)$ and $m_t(\zeta; \eta)$ of $\zeta$ and $\eta$, with the understanding that $\hat \xi_{jt}(\gamma) = \hat \xi_{jt}((\gamma, 0);\eta)$ and $m_{t}(\gamma) = m_{t}((\gamma, 0);\eta)$. Let
where $w_t(\gamma, \eta) = (w_{jt}(\gamma,\eta)')_{j \in \mathcal J_t}$ is $|\mathcal J_t| \times \dim(\psi)$ and $u_t(\gamma, \eta)$ are $\dim(\psi) \times \dim(m)$, with \[
\] We take limits to deal with parameters that are at the boundary when $\psi = 0$, such as the variance of the random coefficients on $\eta$. For the intuition, the left-most terms in ((ref)) and ((ref)) are the Jacobian of the moments with respect to $\psi$. These can depend to first order on $\hat \gamma$. The second parts of these expressions eliminate this dependence with a similar correction to $\hat \kappa_{bc}$. We estimate the variance of $\hat \lambda(\eta)$ and $\hat \lambda_t(\eta)$ using
Finally, define \[ \hat W(\eta) = \left(\hat \lambda(\eta) + \sum_{t=1}^\tau \hat \lambda_t(\eta)\right)' \left( \hat \Lambda(\eta) + \sum_{t=1}^\tau \hat \Lambda_t(\eta) \right)^{-1} \left(\hat \lambda(\eta) + \sum_{t=1}^\tau \hat \lambda_t(\eta)\right). \] Without microdata, $\hat W(\eta)$ simplifies to $\hat W(\eta) = \hat \lambda(\eta)' \hat \Lambda(\eta)^{-1} \hat\lambda(\eta)$. The statistic $\hat W(\eta)$ can be shown to behave like a $\chi^2_{\dim(\psi)}$ random variable under the null (i.e., when the true $\psi = 0$), but it depends on the nuisance parameter $\eta$. Thus, we define our second diagnostic as
where we take the supremum over $\eta$ in the unit sphere (since the scale of $\eta$ is not important) that are orthogonal to the column span of $\tilde e$.
Following, e.g., hansen1996inference, critical values can be computed by simulation. Draw $(\varpi_t^*)_{t=1}^T, (\varpi_{i1}^*)_{i=1}^{N_1},\ldots, (\varpi_{i\tau}^*)_{i=1}^{N_\tau}$ iid from a $N(0,1)$ distribution, and set \[ \hat W^*(\eta) = \left(\hat \lambda^*(\eta) + \sum_{t=1}^\tau \hat \lambda_t^*(\eta)\right)' \left( \hat \Lambda(\eta) + \sum_{t=1}^\tau \hat \Lambda_t(\eta) \right)^{-1} \left(\hat \lambda^*(\eta) + \sum_{t=1}^\tau \hat \lambda_t^*(\eta)\right), \] where \[
\] For each collection of $N(0,1)$ draws, compute \[ LM_2^* = \sup_{\eta \in S^J : \eta \perp C(\tilde e)} \hat W^*(\eta)^2. \] Let $\hat \xi_{0.95}^*$ denote the 95th percentile of $LM_2^*$ across a large number of independent draws. This quantity is easy to compute, as only the right-most terms in the expressions for $\hat \lambda^*$ and $\hat \lambda^*_t$ need to be recomputed for different draws and these terms do not depend on $\eta$. We reject the null that the dimension of $\tilde e$ is adequate if $LM_2 > \hat \xi_{0.95}^*$.
Given a counterfactual of interest $\kappa$ and a candidate set of proxies $\tilde e$:
Our approach also provides a data-driven way to robustly estimate counterfactuals and validate some of the assumptions implicitly made in the demand estimation literature in contexts where only standard numeric attributes are available. The typical workflow assumes that product attributes are measured without error and are of adequate dimension. Our bias correction allows practitioners to relax the assumption of correct measurement. To do so, the bias-corrected estimator can be implemented as above, where now $\tilde e$ simply represents the numeric attributes that may be mismeasured. The resulting bias-corrected estimator is robust to such mismeasurement. Similarly, our diagnostics may be used to choose among attributes and assess whether the dimension of a candidate set of attributes is adequate.
Even when mismeasurement is not a concern, our bias-corrected estimator and standard error formulas can be used to perform efficient inference on counterfactuals (see Proposition (ref) below for a formal statement). In this case, there is no need to reparameterize the model to account for mismeasurement and both are implemented as described above with $\gamma \equiv \theta$. This approach offers a few advantages: (i) it yields efficient estimates of counterfactuals even when $\hat \theta$ is inefficient, (ii) standard errors are available in closed form, avoiding the need to bootstrap; and (iii) it allows for combined market-level and microdata.
Next, we consider settings with individual-level price variation and choice data. Following an established literature dube2019handbook, we assume the researcher includes product fixed effects to account for systematic differences across products and is willing to rule out any remaining price endogeneity.
The researcher has data on a large number $n$ of consumers in a single market in which $J$ goods are sold,\footnote{As before, we assume that $J$ is fixed but it is straightforward to extend our analysis to asymptotic thought experiments where $J$ grows slowly with $n$. For ease of exposition, we present results for a single market, though our approach extends easily to settings with multiple markets.} and identifies the outside option with $j = 0$. For each consumer $i$, the researcher observes individual choices $d_i = (d_{ij})_{j=1}^J$ where $d_{ij} = 1$ if $i$ chooses good $j$ and $0$ otherwise, $p_i = (p_{ij})_{j=1}^J$ which collects prices and other variables (e.g., rankings on the results page) that vary across consumers, and a vector of demographic variables $y_i$. For each product $j$, the data also may contain attributes $x_j$ that are common across all consumers. The model predicts choice probabilities as a function of $p_i$, $y_i$, $x = (x_j)_{j=1}^J$, and a parameter vector $\theta$:
This model subsumes many empirical examples. Here we give just one standard workhorse model.
As before, we partition $x_j \equiv (\bar x_j,e_j)$, where $\bar x_j$ is a vector of standard observed product attributes, such as product size, and $e_j$ is an $r$-vector representing product attributes that are harder to capture using standard numeric data and which we treat as latent to the econometrician. We again assume the researcher has proxies $\tilde e = (\tilde e_j')_{j=1}^J$ for the true underlying $e = (e_j')_{j=1}^J$ and stay agnostic on the form that $\tilde e$ takes. A leading case is again the scenario where $\tilde e_j$ are embeddings computed to represent unstructured data $U_j$, though our approach may equally be used in scenarios where $\tilde e_j$ represents some potentially mismeasured product attributes.
We are interested in estimating a counterfactual of the form \[ \kappa = \mathbb{E}[ k(p_i, y_i, \bar x, e; \theta )], \] where the expectation is over the distribution of $(p_i,y_i)$. Here $\kappa$ might represent an average price-elasticity of consumers, average equilibrium price, or average welfare measure. It could also represent a quantity that doesn't depend on the distribution of $(p_i,y_i)$, such as the price-elasticity or welfare measure for an individual with given $y$ facing given prices $p$, in which case the expectation is redundant.
In the usual workflow, model parameters are estimated using the observed data and a candidate set of proxies $\tilde e$. For instance, one could use maximum likelihood. Given an estimate $\hat \theta$ of $\theta$, the counterfactual $\kappa$ is usually estimated as \[ \hat \kappa = \frac 1n \sum_{i=1}^n k(p_i, y_i, \bar x, \tilde e; \hat \theta ). \] As before, we refer to this as the naive estimator of $\kappa$ since it does not account for the fact that the proxies $\tilde e$ might differ from the true latent attributes. Mismeasurement of $e$ affects the naive estimator of $\kappa$ both directly, since $\tilde e$ is an argument of $k$, and indirectly through bias in the first-stage estimate $\hat \theta$, since $\tilde e$ enters the likelihood.
We now introduce a bias-corrected estimator of $\kappa$ that is designed to mitigate the effects of using $\tilde e$ in place of the true $e$. As before, we consider models in which $e$ and $\theta$ enter choice probabilities ((ref)) via a composite parameter \[ \gamma \equiv \gamma(\theta, e). \] As before, many common specifications have this property. We illustrate it in our leading example.
We shall implicitly assume in what follows that the right-hand side of ((ref)) depends on $(\theta, e)$ only via $\gamma(\theta, e)$. We similarly restrict attention to counterfactuals that depend on $(\theta, e)$ only via $\gamma(\theta, e)$ and write \[
\] We then define the bias-corrected estimator
where $\hat \gamma = \gamma(\hat \theta, \tilde e)$, and $\hat c_i$, $d_i = (d_{ij})_{j=1}^J$, and $\sigma_i(\gamma) = (\sigma_{ij}(\gamma))_{j=1}^J$ are $J \times 1$, vectors, with
where $V_i(\gamma) = \mathrm{diag} (\sigma_i(\gamma)) - \sigma_i(\gamma) \sigma_i(\gamma)'$ and $\hat H= \frac 1n \sum_{i=1}^n \dot \sigma_i(\hat \gamma)' V_i (\hat \gamma)^{-1} \dot \sigma_i(\hat \gamma)$ are of dimension $J \times J$, $\hat k = \frac 1n \sum_{i=1}^n \dot k_i(\hat \gamma)$ and $\dot k_i(\gamma) = \frac{\partial k_i(\gamma)}{\partial \gamma}$ are $\dim(\gamma) \times 1$, and $\dot \sigma_i(\gamma)' = \frac{\partial \sigma_i(\gamma)'}{\partial \gamma}$ is $\dim(\gamma) \times J$. Importantly, $\hat \kappa_{bc}$ involves closed-form expressions of objects that can be easily computed given the estimate $\hat \theta$. As a result, it requires minimal computation beyond what is needed to estimate the model.
As before, $\hat \kappa_{bc}$ takes the naive estimator and adds an adjustment term that purges the effect of $\hat \gamma$ on $\hat \kappa_{bc}$ to first order. Proposition (ref) in Section (ref) shows that the asymptotic distribution of $\hat \kappa_{bc}$ is centered around the true counterfactual $\kappa_0$ and does not depend on $\hat \gamma$. This means $\hat \kappa_{bc}$ is immune to any bias arising from proxying $e$ with $\tilde e$, and its variance is not impacted by data-dependent $\tilde e$. We defer the formal statement of these results to Section (ref) and instead highlight a few key properties.
Here we propose two diagnostics that can be used to assess the suitability of a candidate set of proxies $\tilde e$. The first can be used to assess whether $\hat \gamma = \gamma(\hat \theta, \tilde e)$ is sufficiently close to the truth $\gamma_0 = \gamma(\theta_0, e_0)$ as required by our theory in Section (ref). The second can be used to determine whether the true $e$ are higher dimensional than the proxies $\tilde e$. Both diagnostics are again based on LM statistics so that they require minimal additional computation. Later, we will show that these two diagnostics perform well in finite samples via both simulations and in our empirical application.
The standard demand estimation workflow implicitly assumes that product attributes are measured without error and are of the correct dimension. Our diagnostics may be used to assess the validity of these implicit assumptions even in standard contexts with only quantifiable attributes.
As in Section (ref), a simple LM statistic can be used to check whether $\hat \gamma$ is sufficiently close to $\gamma_0$. This diagnostic is also helpful to guide the choice among sets of embeddings obtained from various data source and/or ML model combinations. Let
where $\hat S = \frac 1n \sum_{i=1}^n \sum_{j=0}^J \frac{d_{ij}}{\sigma_{ij}(\hat \gamma)} \dot \sigma_{ij}(\hat \gamma)$ is the score and $\hat H$ is the (expected) Hessian, defined below ((ref)). As before, this diagnostic can be interpreted as an LM statistic for a test of the null hypothesis that $\gamma_0$ is in the set of composite parameters spanned by $\tilde e$. Proposition (ref) below shows that $LM_1$ behaves like $\|\sqrt n (\hat \gamma - \gamma_0)\|^2$ as $n$ grows large. This allows researchers to validate the assumption that the discrepancy between $\hat \gamma$ and $\gamma_0$ is sufficiently small, as needed in Proposition (ref), by checking whether $LM_1$ is below a threshold. In particular, for any sequence $C_n = o(n^{1/4})$, we have that $LM_1 \leq C_n^2$ implies $\|\hat \gamma - \gamma_0\| \leq \mathrm{constant} \times C_n/\sqrt n = o(n^{-1/4})$ with probability approaching one. Further, Proposition (ref) shows that $LM_1$ can be used to bound the discrepancy between $\tilde e$ and $e_0$.
The construction follows similar ideas to Section (ref). We augment $\tilde e$ with a vector $\eta$ representing additional attributes not included in $\tilde e$ and augment $\theta$ with an additional component $\psi \in \Psi$ representing coefficients on $\eta$. Correspondingly, we extend $\gamma$ to $\zeta = (\gamma, \psi)$. With slight abuse of notation, we now write choice probabilities on this extended space as $\sigma_{ij}(\zeta; \eta)$, with the understanding that $\sigma_{ij}(\gamma) = \sigma_{ij}((\gamma, 0);\eta)$.
Consider an LM test of the null hypothesis that $\psi = 0$. Such a test could be based on the score
where \[ w_{ij}(\gamma, \eta) = \lim_{\psi \to 0} \frac{\partial \sigma_{ij}((\gamma, \psi);\eta)}{\partial \psi}. \]
The statistic ((ref)) can still depend to first order on $\hat \gamma$. To eliminate this dependence, we perform a similar correction to $\hat \kappa_{bc}$. Let $g_i(\gamma;\eta) = w_i(\gamma, \eta)' V_i(\gamma)^{-1} (d_i - \sigma_i(\gamma))$, where $w_i(\gamma; \eta) = (w_{ij}(\gamma; \eta)')_{j=1}^J$ is $J \times \dim(\psi)$. Then define the $\dim(\psi) \times J$ matrix \[ \hat c_i(\eta)' = \left( \frac 1n \sum_{l=1}^n w_l(\hat \gamma;\eta)' V_l(\hat \gamma)^{-1} \dot \sigma_l(\hat \gamma) + \dot g_l(\hat \gamma;\eta) \right) \hat H^{-1} \dot \sigma_i(\hat \gamma)' V_i(\hat \gamma)^{-1}, \] where $\dot g_i(\gamma;\eta)' = \frac{\partial g_i(\gamma; \eta)}{\partial \gamma}$ is $\dim(\gamma) \times \dim(\psi)$, and let \[
\] Finally, let \[ \hat W(\eta) = \hat \lambda(\eta)'\hat \Lambda(\eta)^{-1} \hat \lambda(\eta). \] It can be shown that the statistic $\hat W(\eta)$ behaves like a $\chi^2_{\dim(\psi)}$ random variable under the null (i.e., when the true $\psi = 0$), but it depends on the nuisance parameter $\eta$. Thus, we again define our second diagnostic as \[ LM_2 = \sup_{\eta \in S^J : \eta \perp C(\tilde e)} \hat W(\eta). \] Critical values can be computed by simulation: for iid $N(0,1)$ random variables $(\varpi_i^*)_{i=1}^n$, compute \[ LM_2^* = \sup_{\eta \in S^J : \eta \perp C(\tilde e)} \hat W^*(\eta)^2, \] where $\hat W^*(\eta) = \hat \lambda^*(\eta)'\hat \Lambda(\eta)^{-1} \hat \lambda^*(\eta)$, with $\hat \lambda^*(\eta) = \frac{1}{\sqrt n } \sum_{i=1}^n \varpi_i \hat c_i(\eta)'(d_i - \sigma_i(\hat \gamma))$. Note $\lambda^*(\eta)$ factors into the product of terms involving $\eta$, which only need to be computed once, and terms involving $(\varpi_i^*)_{i=1}^n$, which are trivial to compute. Let $\hat \xi_{0.95}^*$ denote the 95th percentile of $LM_2^*$ across a large number of independent sequences $(\varpi_i^*)_{i=1}^n$. The null that the dimension of $\tilde e$ is adequate can be rejected if $LM_2 > \hat \xi^*_{0.95}$.
Given a counterfactual of interest $\kappa$ and a candidate set of proxies $\tilde e$:
We first illustrate our approach in simulations. We consider the model in Section (ref) with 10 products and 10,000 consumers. These figures are in line with the data used in the empirical application in Section (ref). We model utility as a function of price and two-dimensional latent attributes $e$ (in addition to idiosyncratic shocks). The latent attributes $e_j$ are drawn iid $N(0,1)$ across products. Individual-level prices are drawn iid from a $N(5,1)$ distribution and vary across simulations. The random coefficient on price is $N(-1,0.3^2)$ and the coefficients on the latent attributes are iid $N(0,0.75^2)$. We keep the latent attributes $e$ fixed in the data generating process and vary the amount of mismeasurement in the proxies $\tilde e$ that are used in estimation. Specifically, for every $j$, we let:
where $\eta_j$ is a two-dimensional standard normal random vector drawn iid across simulations and $\rho$ determines the amount of mismeasurement in $\tilde e$. When $\rho=0$, the proxies exactly match the latent attributes, whereas as $\rho$ increases towards 1 the proxies are increasingly mismeasured. We note that this simulation design imposes very few restrictions on the form of mismeasurement. In particular, depending on the draw of $\eta_j$, each element of $\tilde e_j$ could be smaller or larger than the corresponding element of $e_j$ and this can freely vary across goods $j$.\footnote{Note that (ref) is such that $\tilde e_j$ has roughly the same amount of variation across goods $j$ as $e_j$ does. This allows us to isolate the effect of mismeasurement in a way that is not confounded by changes in the scale of the proxies $\tilde e_j$ used in estimation.}
We focus on estimation of the fraction of consumers that switch from one product to another one when the former is removed from the choice set. Figure (ref) plots histograms of the naive estimator that takes the proxies $\tilde e $ as true and the distribution of our bias corrected estimator across simulations. As the level of mismeasurement increases, the distribution of the naive estimator moves away from the true value of the counterfactual (roughly 0.05), whereas the corrected estimator remains centered around the true value. Interestingly, for larger levels of mismeasurement, the distribution of the naive estimator ends up being centered around the counterfactual prediction of the logit model with no random coefficients. This is intuitive: as the proxies $\tilde e$ become increasingly noisy, they capture less of the substitution patterns in the data, and the estimated variance of their random coefficients shrinks towards zero. This finding also serves as a warning that mismeasurement in the proxies can defeat the purpose of estimating a random coefficients model in the first place: if the mismeasurement bias is not properly accounted for, the model may revert to the restrictive substitution patterns that the model was specifically intended to relax.
To better assess the trade-offs involved in our bias correction, Figure (ref) shows how the bias and RMSE of the two estimators vary with the amount of mismeasurement $\rho$. When $\tilde e$ is measured with no error, both the naive and the bias-corrected estimator have very low bias. Our estimator has a marginally higher RMSE, indicating that its variance is slightly higher than that of the naive estimator. This is intuitive: since the naive estimator leverages the assumption that the proxies $\tilde e$ are correct and ours does not, we obtain slightly less precise estimates when that assumption happens to be correct. However, this is a knife-edge case. When mismeasurement is present, our estimator consistently achieves lower bias and RMSE than the naive estimator. The comparison is especially striking for small to moderate mismeasurement ($\rho \in [0,0.3]$), where the bias correction is able to remove essentially all of the bias. As expected, when the mismeasurement becomes very large ($\rho>0.5$), the bias correction starts to also perform worse. This is because the bias correction requires that $\hat \gamma$ be within a vicinity of $\gamma_0$ that is roughly double the order of sampling error.
Finally, Figure (ref) shows that the $LM_1$ diagnostic discussed in Section (ref) is able to correctly rank proxies. In particular, the average $LM_1$ statistic increases monotonically with the average distance between the $\hat \gamma$ induced by the proxies and $\gamma_0$. This confirms that the diagnostic can be valuable in guiding researchers towards proxies that are relatively close to the true latent attributes.
We now apply our method to the experimental data from compiani2025demand. The data records the choices made by 9,265 participants when faced with a choice of ten e-books. In a first task, participants were asked to choose their preferred e-book based on information displayed to them, including (randomized) prices, standard attributes (author, year of publication, genre and number of pages), and unstructured information (cover images, titles, plot descriptions and reviews). In a second task, each participant's first choice was removed and they were asked to choose again from the remaining nine books. compiani2025demand estimate a range of models on the first choice data and compare their performance in predicting second choices. This gives a direct measure of how well different models capture counterfactual substitution patterns. Specifically, the paper compares mixed logit models based on standard attributes with mixed logit models that leverage proxies extracted from unstructured data. The key findings are that (i) unstructured data is predictive of substitution patterns, and (ii) book descriptions and reviews, when processed with transformer-based text models, perform particularly well at predicting substitution.
The results in compiani2025demand treat the proxies as if they were correctly specified. However, there are good reasons to believe that mismeasurement might play an important role. First, the unstructured data are processed using pre-trained ML models that are not targeted towards predicting substitution patterns.\footnote{Specifically, images are processed via classification models trained to assign each image to one of many classes; texts are processed using bag-of-words and transformer models: the former simply capture word frequency, whereas the latter are trained to predict the next word in a text.} While the resulting proxies are found to be predictive of substitution patterns, they may not perfectly capture the underlying attributes that drive consumers' choices. Second, the dimension of the proxies is reduced via PCA before inputting them into the demand model, which is likely to introduce further mismeasurement. This also raises the question of how many principal components should be included in the model.
Here we investigate whether applying our bias correction method and diagnostics helps better capture substitution patterns. We use the approach from Section (ref) since we have individual-level data and prices are randomized, so endogeneity is not a concern. We focus on the ability of different models to correctly predict the closest substitute for any given book. The second choice data give us a direct measure of this: for a given book A, its closest substitute is the book that most people switch to when A is removed from the choice set.
We note that this exercise sets a high bar for our approach. Unlike in a simulation, the demand model might be misspecified even if the proxies are correctly specified. For instance, the model assumes a normal distribution for the random coefficients but the true distribution of preference heterogeneity might be different. As a result, this exercise tests whether our approach works well even in cases where all assumptions needed for the theoretical results might not hold exactly. Further, by looking at substitution patterns in response to product removals that are not part of the estimation data, this provides a direct test of the model's ability to correctly predict counterfactuals.
Figure (ref) shows the results. For each specification---defined as a combination of unstructured data source and ML model used to extract proxies from it---we report the fraction of the ten books for which the model correctly identifies the closest substitute (as measured by the second choice data).\footnote{The model prediction of the closest substitute is the product with the highest average (across consumers) second-choice probability once product $A$ is removed. Here, $\kappa$ corresponds to the probability that $B$ is a consumer's second choice conditional on $A$ being their first choice, averaged across consumers, which we compute across all $(A,B)$ pairs.} The hashed bars show the performance of the naive approach that uses the estimates from compiani2025demand, whereas the green solid bars show the performance of the bias-corrected estimator. Three specifications are ruled out by the diagnostic $LM_2$, indicating that they don't feature a sufficient number of random coefficients. For 11 out of the 13 remaining specifications (around 85%), the bias correction weakly improves performance and the magnitude of the improvement is large in several cases. In particular, for specifications using reviews data, the fraction of correctly predicted substitutes goes from 40% to 60-70%. For comparison, a coin flip would achieve a hit rate of 11%. Further, the specification fitting the data best as measured by the $LM_1$ diagnostic is among those achieving the best counterfactual performance (70% with bias correction). These results confirm that our approach is able to meaningfully improve counterfactual predictions and guide researchers towards the best-performing specifications.
Let $\Gamma$ denote the set of all values of $\gamma(\theta, e)$ as $\theta$ varies over the parameter space $\Theta$ and $e$ varies over all $J \times r$ matrices with linearly independent rows. We shall implicitly assume in what follows that the true latent attributes $e_0$ and the proxies $\tilde e$ are $J \times r$ with linearly independent rows. We shall also implicitly assume that $\Gamma$ is convex and open. This is true for the $\gamma(\theta, e)$ given in Examples (ref) and (ref), for which $\Gamma$ is the product of copies of $\mathbb R$ and $(0,\infty)$ and is therefore convex and open.\footnote{ For instance, the operation $l$ stacking the lower-triangular entries of the Cholesky factor maps the manifold of symmetric positive definite matrices into the product of copies of $\mathbb R$ and $(0,\infty)$. A similar result holds for the reduced-rank Cholesky decomposition $l_r$ neuman2023restricted.} In Section (ref), the counterfactual function $k_t$, $\hat \xi_t$ from ((ref)), and micro-moments $m_t$ depended on $(\theta, e)$ only through the value of $\gamma(\theta, e)$. We can therefore view $k_t$, $\hat \xi_t$, and $m_t$ as random functions defined on $\Gamma$.
We first give results for the case of combined market-level data and microdata. We let $\chi_t = (p_t, \xi_t, \bar x_t, e)$ and let $\mathcal M$ denote the $\sigma$-algebra generated by $\chi_1,\ldots,\chi_\tau$. Let $V = \mathrm{Var} ( Z_t \hat \xi_t(\gamma_0) )$ be the $\dim(z) \times \dim(z)$ covariance matrix of the aggregate moments at the true parameters, and let $V_t = \mathrm{Var} \left( m_{it} | \chi_t \right)$ be the $\dim(m) \times \dim(m)$ covariance matrix of the micro moments for market $t$ conditional on $\chi_t$. Let
where $r_1,\ldots,r_\tau$ are defined in Assumption (ref) below, and let $h = \mathbb{E}[ \dot k_t(\gamma_0) ] - G' V^{-1} K$, where $K = \mathbb{E} [ k_t(\gamma_0) Z_t \hat \xi_t(\gamma_0) ) ]$ is $\dim(z) \times 1$, $G = \mathbb{E}[Z_t \dot \xi_t(\gamma_0)]$ is $\dim(z) \times \dim(\gamma)$, and $M_t = \dot m_t(\gamma_0)$ is $\dim(m) \times \dim(\gamma)$. Here $V$ and $h$ are deterministic whereas $V_1, \ldots, V_\tau$ and $H$ are $\mathcal M$-measurable random matrices. Let $N$ be a neighborhood of $\gamma_0$.
Assumption (ref)(ref)-(ref) are standard smoothness, moment, and rank conditions, respectively. Assumption (ref)(ref) treats the sample size $T$ of the aggregate data and the sample sizes of the microdata as comparable. This is designed to give a meaningful approximation to common empirical scenarios where $T$ is in the high tens or hundreds and $N_t$ is in the hundreds or thousands for each market (e.g., petrin2002quantifying and grieco2024evolution).
The next result shows that the bias-corrected estimator $\hat \kappa_{bc}$ from (ref) is asymptotically centered at the true counterfactual $\kappa_0 = \mathbb{E}[k_t(\gamma_0)]$ and its asymptotic variance is independent of $\hat \gamma$, $\hat \theta$, and $\tilde e$. Because the effect of market-level variables in markets for which there is microdata persists in the limit, we use the notion of stable convergence. We say a sequence of random variables $Z_T$ converges in distribution to $Z$ ($\mathcal M$-stably) if $\lim_{T \to \infty} \Pr( Z_T \leq z, A) = \Pr(Z \leq z, A)$ for all continuity points $z$ of the distribution of $Z$ and all $\mathcal M$-measurable events $A$. Convergence in distribution is a special case corresponding to replacing $\mathcal M$ with the trivial $\sigma$-algebra $\{\emptyset, \Omega\}$.
The (random) asymptotic variance $V_{bc}$ can easily be estimated using $\hat V_{bc}$ in equation ((ref)). Standard errors $\mathrm{s.e.}(\hat \kappa_{bc}) = \sqrt{\hat V_{bc}/T}$ are consistent under Assumption (ref) and valid inference can be performed based on $t$-statistics $(\hat \kappa_{bc} - \kappa_0)/\mathrm{s.e.}(\hat \kappa_{bc})$ using the standard $N(0,1)$ critical values.
We next provide a sense in which $\hat \kappa_{bc}$ is efficient. The proof of Proposition (ref) shows that $\hat \kappa_{bc}$ belongs to the class $\mathcal K$ of estimators $\hat \kappa$ of $\kappa_0$ that satisfy
where $c$ and $d_1,\ldots,d_\tau$ are any $\mathcal M$-measurable random vectors that satisfy
(almost surely). For instance, similar arguments as in the proof of Proposition (ref) show that any $\hat \kappa$ obtained by plugging-in any $\hat \gamma = \gamma_0 + o_p(T^{-1/4})$ into ((ref)) for some arbitrary weights $\hat c$ and $\hat d_1,\ldots,\hat d_\tau$ converging to $c$ and $d_1,\ldots,d_\tau$ belongs to this class. Condition (ref) typically ensures that such an estimator satisfies (ref) uniformly for $\gamma$ local to $\gamma_0$. In general, there are many different weights $\hat c$ and $\hat d_1,\ldots,\hat d_\tau$ whose probability limits $c$ and $d_1,\ldots,d_\tau$ will correspond to different (random) asymptotic variances. The following result shows that $\hat \kappa_{bc}$ has the smallest asymptotic variance among this class of estimators of $\kappa_0$.
An important practical take-away from Proposition (ref) is that microdata should be used, when available, to improve the efficiency of estimators of counterfactuals. Any estimator that discards the microdata by implicitly setting $d_1,\ldots,d_\tau = 0$ will have an unnecessarily large variance (and hence standard errors).
Of course, in many scenarios microdata may not be available. Here we state a simpler version of Proposition (ref) tailored to this case. Recall that the bias-corrected estimator in this case is given in (ref), where $\hat c$ is given in (ref) with $\hat H = \hat G' \hat V^{-1} \hat G$. To introduce the assumptions, let $V$ and $G$ be as above, and let $H = G' V^{-1} G$.
The next result is a special case of Proposition (ref) and is stated without proof.
The asymptotic variance can be easily estimated using the formula $\hat V_{bc}$ in (ref). Standard errors are then computed as $\sqrt{\hat V_{bc}/T}$. These are consistent under the conditions of Proposition (ref).
We now present a result that provides a formal sense in which the diagnostic $LM_1$ in (ref) behaves like $\|\sqrt T(\hat \gamma - \gamma_0)\|^2$ as the sample size grows large. We first state the result then discuss its implications. In what follows, we abbreviate “with probability approaching one” to “wpa1.” Recall $H$ from (ref) and let $\lambda_{\min}(H)$ denote its smallest singular value, which is uniformly bounded away from zero by Assumption (ref)(ref).
Proposition (ref) shows $LM_1$ behaves like $\|\sqrt T (\hat \gamma - \gamma_0)\|$. With $C_T = o(T^{1/4})$, wpa1 we have that $LM_1 \leq C_T^2$ implies $\|\hat \gamma - \gamma_0\| \leq \mathrm{constant} \times C_T/\sqrt T = o(T^{-1/4})$.
The proof of Proposition (ref) shows that the “wpa1” qualifier depends on whether a $\chi^2_{\dim(\gamma)}$ random variable is less than $\epsilon^2 C_T^2$. With $\epsilon = 1$, say, this suggests taking $C_T^2$ to be at least as large as the 95th or 99th percentile of the $\chi^2_{\dim(\gamma)}$ distribution. To check a convergence rate of $\sqrt{(\log T)/T}$, for instance, one could use $C_T^2 = \chi^2_{\dim(\gamma), 0.95} \log T$.
We again let $\Gamma$ denote the set of all values of $\gamma(\theta, e)$ as $\theta$ varies over the parameter space $\Theta$ and $e$ varies over all $J \times r$ matrices with linearly independent rows, assume $e_0$ and $\tilde e$ are $J \times r$ with linearly independent rows, and that $\Gamma$ is convex and open. This is true for the $\gamma(\theta, e)$ given in Example (ref), for which $\Gamma$ is the product of copies of $\mathbb R$ and $(0,\infty)$ and is therefore convex and open (see footnote (ref)). In Section (ref), the counterfactual function $k_i$ and choice probabilities $\sigma_i$ depended on $(\theta, e)$ only through the value of $\gamma(\theta, e)$. We therefore treat $k_i$ and $\sigma_i$ as random functions on $\Gamma$.
We now derive the theoretical properties of the bias-corrected estimator $\hat \kappa_{bc}$ from ((ref)). We first outline some standard smoothness, moment, and rank assumptions. In what follows, we use a dot (as above) and double dot to denote first and second derivatives with respect to $\gamma$. Let $H = \mathbb{E} [ \dot \sigma_i(\gamma_0)' V_i(\gamma_0)^{-1} \dot \sigma_i(\gamma_0) ]$ and $N$ be a neighborhood of $\gamma_0$.
Let $\kappa_0 = \mathbb{E}[k_i(\gamma_0)]$ denote the true value of the counterfactual. The following result shows that the asymptotic distribution of $\hat \kappa_{bc}$ is centered at $\kappa_0$ and its variance is independent of $\hat \gamma$, $\hat \theta$, and $\tilde e$.
The asymptotic variance $V_{bc}$ can be easily estimated using $\hat V_{bc}$ in (ref). Standard errors are then computed as $\sqrt{\hat V_{bc}/n}$. These are consistent under the conditions of Proposition (ref). We note that, as before, Proposition (ref) requires that $\hat \gamma$ be in a vicinity of $\gamma_0$, and refer the reader to Remark (ref) for a discussion.
The following result is analogous to Proposition (ref) and shows $LM_1$ behaves like $\|\sqrt n (\hat \gamma - \gamma_0)\|$.
The implications of Proposition (ref) are similar to before. In particular, for $C_n = o(n^{1/4})$, we have wpa1 that $LM_1 \leq C_n^2$ implies $\|\hat \gamma - \gamma_0\| \leq \mathrm{constant} \times C_n/\sqrt n = o(n^{-1/4})$. The proof of Proposition (ref) shows that the “wpa1” qualifier depends on whether a $\chi^2_{\dim(\gamma)}$ random variable is less than $\epsilon^2 C_n^2$. To check a convergence rate of $\sqrt{(\log n)/n}$, for instance, one could use something like $C_n^2 = \chi^2_{\dim(\gamma), 0.95} \log n$.
With some additional structure, we can also use $LM_1$ to deduce a similar bound on the proxies $\tilde e$. To introduce the assumptions, let $\theta_*$ and $e_*$ be such that $\gamma(\theta_*,e_*) = \gamma_0$. We do not require that $\theta_*$ and $e_*$ are the true structural parameters and attributes, only that they induce $\gamma_0$. Let $\hat G_\theta = \frac{\partial \gamma(\hat \theta, \tilde e)'}{\partial \theta}$, $\hat G_e = \frac{\partial \gamma(\hat \theta, \tilde e)'}{\partial{\mathrm{vec}(e)}}$, $G_\theta = \frac{\partial \gamma(\theta_*, e_*)'}{\partial \theta}$, and $G_e = \frac{\partial \gamma(\theta_*, e_*)'}{\partial{\mathrm{vec}(e)}}$ (these are well defined under Assumption (ref) below). Also let $C(G_\theta)$ denote the column span of $G_\theta$ and $M = I - H^{1/2} G_\theta'(G_\theta H G_\theta')^{-1} G_\theta H^{1/2}$ denote the projection onto $C(G_\theta)^\perp$.
Let $\sigma_{\min}(MH^{1/2} G_e')$ denote the smallest singular value of the matrix $MH^{1/2} G_e'$. Note this is positive by Assumption (ref)(ref).
The proof of Proposition (ref) shows that the “wpa1” qualifier depends on whether a $\chi^2_{\mathrm{rank}(M)}$ random variable is less than $\epsilon^2 C_n^2$. With $\epsilon = 1$, say, this suggests taking $C_n^2$ to be at least as large as the 95th or 99th percentile of the $\chi^2_{\mathrm{rank}(M)}$ distribution.
In this paper, we develop a toolkit to correct bias and perform valid inference on counterfactuals when the product attributes used in demand estimation may only imperfectly capture the latent attributes that drive substitution. A leading case is when consumer choice is driven by difficult-to-quantify characteristics and unstructured data, such as product images, descriptions, review text, or consumer surveys, are converted into numerical variables using ML methods. As e-commerce continues to expand and such data play an increasingly central role in driving consumer choices, the need to incorporate these sources into demand estimation will only grow. In addition, our methods may be applied as simple post-estimation robustness checks even with standard numeric attributes when mismeasurement is a concern. All our methods require minimal additional computation once model parameters are estimated and can be easily integrated in the canonical demand estimation workflow.
{\onehalfspacing }