Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
61,033 characters · 9 sections · 35 citation commands
Maximum Likelihood Estimation of Differentiated Products Demand Systems
\onehalfspacing
Demand estimation is a fundamental tool in applied economics. One widely used demand model for differentiated goods markets is berry1995automobile (BLP), which builds on the earlier work of mcfadden86 concerning the mixed logit. The BLP model is traditionally estimated by generalized method of moments (GMM), but the estimates are often imprecise. It has long been known that the estimation performance is dramatically improved by adding a model of the supply side, typically one in which firms engage in Bertrand-Nash pricing. A small but influential literature has documented other ways of improving estimation performance: replacing a step that inverts from market shares to mean utilities with a set of constraints on the difference between actual and predicted market shares (the MPEC approach of dube2012improving), using empirical likelihood instead of GMM conlon2013empirical), choosing the right instruments (reynaert2014improving,gandhi2019measuring) and using high-quality software (conlon2020best).
This paper adds to this literature by revisiting an old idea: estimating BLP by maximum likelihood. The theoretical advantages of doing so are clear. First, maximum likelihood is statistically efficient, which translates into more precise estimates. Second, maximum likelihood does not require the researcher to choose their instruments. In fact, provided the model is correctly specified, one can think of MLE as a magical procedure that automatically performs GMM with the optimal instruments. This may be important in practice, for two reasons. First, choosing good instruments can be hard, as attested to by the papers on this topic cited above. Second, BLP is often used for demand analysis in merger cases, and in such cases any degrees of freedom, such as instrument choice, can be exploited by competing expert witnesses in their analysis of the case.
There are, however, some drawbacks. Both price and quantity are endogenous, and so a well-specified likelihood requires modeling how both price and quantity arise, necessitating models of both demand and supply. This requires committing to some model of pricing. Moreover, the researcher needs to specify a distribution for the demand and cost shocks. \ We investigate these trade offs here.
The paper proceeds in three parts. First, we outline the BLP model and derive the maximum likelihood estimator (MLE) under Bertrand-Nash pricing. Somewhat surprisingly, this appears to be new, as the prior literature we review below has assumed that the demand error enters the pricing equation linearly, which is inconsistent with Bertrand-Nash.
The second part of the paper investigates the performance of the MLE relative to the best practice in implementing GMM (for best practice, we follow conlon2020best). We start with a variety of well-specified scenarios, in which we would expect MLE to outperform GMM. We find that it does, and quite handily - bias and mean squared error are smaller, standard errors are tighter, and coverage is more precise. We then consider three scenarios in which the model is wrong in some way. The first tests the robustness of MLE to a different error structure, namely Laplace errors. We find that the MLE continues to outperform the GMM benchmark. The second considers a misspecified supply side, where estimation assumes costs are log-linear rather than linear in characteristics. Here we find that both GMM and MLE perform poorly in estimating the demand parameters, as expected, but MLE performs worse. Finally, we consider a case where the ownership matrix is mis-specified, which we intend as reduced form shorthand for a mis-specification of the game being played on the supply side. The results of GMM and MLE are comparable in this case.
In the final part of the paper, we replicate BLP, using their original data, and building on the replication exercise done in andrews2017measuring (AGS). We compare the MLE estimates to a series of benchmarks: the original estimates in BLP, the replication results of AGS, and the best practice estimates of conlon2020best (CG). The MLE estimates are very similar to the GMM estimates both in terms of the parameters and implied own price elasticities. Notably, the price coefficient estimate is close to the best practice of CG, unlike BLP and AGS. In addition the standard errors are substantially smaller; often less than 25% on a typical parameter estimate. This performance is “out of the box”; unlike the GMM estimators, we didn't have to choose our set of instruments.
Much of the work in this paper lies in the details. The likelihood requires computation of a Jacobian term, a large matrix which we analytically derive in the appendix using the implicit function theorem. The Jacobian of the likelihood in turn includes terms from the Hessian of the share function. Because numerical errors can easily accumulate we take extra care to implement state of the art computational methods. We do this by utilizing the PyTorch framework developed in paszke2019pytorch. One particular feature of PyTorch that helps with optimizing is Automatic Differentiation (AD). For optimizing the objective, finite difference methods can lack speed and reliability compared to analytic derivatives. AD exploits the chain rule and the fact that computer code is made up of elementary mathematical operations to compute gradients that are as precise as analytically derived gradients.
Taken together, these simulation and replication exercises suggest that MLE is a useful addition to the set of tools available to estimate BLP, allowing researchers to avoid choosing moment conditions, and having a substantial advantage in statistical precision, at least where the researcher is willing to commit to a model of the supply side.
\paragraph{Related literature.} The idea of estimating differentiated products demand systems by maximum likelihood is by no means new. With individual choice data, it is in fact common (see e.g. honka2017advertising and abaluck2021consumers for more recent examples). The marketing literature has also considered using maximum likelihood with aggregate data. jiang2009bayesian suggests a Bayesian Markov Chain Monte Carlo (MCMC) approach for demand estimation in the absence of endogenous prices. They find in simulations that their estimates have lower mean squared error. In an extension of their main model, they add a linear supply side pricing equation and suggest how instrumental variables may be used in the presence of endogeneity. park2009simulated perform a similar analysis, using maximum likelihood rather than a Bayesian approach, and again insisting on a linear pricing equation and bivariate normal errors. Finally, in an appendix to their paper, dube2012improving once again employ a linear pricing equation when considering the performance of a maximum likelihood estimator of BLP. Our paper differs from these prior contributions in that we set up the supply-side pricing equation to be consistent with Bertrand-Nash pricing, leading to an equation in which the demand error enters non-linearly.
Our model closely follows berry1995automobile. Our exposition is deliberately concise, and we refer the reader to berry1995automobile for additional details. We use $i$ to index individuals, $j$ for products, and $t$ for markets. We present our full notation in Table (ref).
\paragraph{Utility.} An individual $i$ buying product $j$ in market $t$ obtains utility equal to:
where $\ensuremath{\boldsymbol{\mathbf{x}}}_{jt}$ represents the observable characteristics of a product, $p_{jt}$ is the associated price of a product, $\xi_{jt}$ is a scalar unobservable product characteristic, $\delta_{jt}$ is the product's mean utility, and $\epsilon_{ijt}$ is an idiosyncratic shock that is i.i.d. and follows a type-1 extreme value distribution. Product characteristics are allowed to have heterogeneous effects on individuals' utilities through $\ensuremath{\boldsymbol{\mathbf{\beta}}}_i$. Each observable characteristic coefficient $k$ follows $\beta_{ik} = \beta_k + \tilde{\beta}_{ik}$, where $\tilde{\beta}_{ik}$ is an i.i.d. random variable drawn from $N(0, \sigma_{\beta})$. Mean (across consumers) utility takes the form:
and thus individual utility can be re-written as:
\paragraph{Demand.} Individual consumers make product choices that maximize their utility. The observed market-level shares of each product are the result of aggregating these individual-level decisions. The assumption on the idiosyncratic shock $\epsilon_{ijt}$ and the existence of an outside option where $u_{0t}=0$ yields market shares:
\paragraph{Supply.} For the supply side of the market, we assume a marginal cost structure of:
where $c_{jt}$ is the marginal cost of producing a product, $\ensuremath{\boldsymbol{\mathbf{w}}}_{jt}$ is a vector of characteristics and $u_{jt}$ is a latent scalar supply side cost shock. Firm profits are given by:
As in BLP, we assume that firms have the objective of maximizing their profits in a given market across all of their products. This leads to the following first order conditions (FOCs):
where $\pi_{ft}$ is the profit of firm $f$ in market $t$. This can be written in matrix form as:
where $J^{s,p}_{t}[j, j'] = \frac{\partial s_{j'}}{\partial p_j}$ and $O_t$ is the ownership matrix:
Given $\{\ensuremath{\boldsymbol{\mathbf{p}}}_{t}, \ensuremath{\boldsymbol{\mathbf{s}}}_{jt}\}$ and a set of parameters, Eq. (ref) can be used to estimate marginal costs. These marginal costs can be used to estimate the parameters of Eq. (ref):
This section presents the two main estimation procedures for this model: first, we provide a brief discussion of estimation using Generalized Method of Moments (GMM), and then we provide a more elaborate description of the MLE estimator. We will use notation similar to prior literature: $\Theta = \{\theta, \beta, \gamma\}$ represents all linear and non-linear parameters in the model. The parameters $(\beta,\gamma)$ enter the mean utility and cost equations linearly, and we refer to them as the linear parameters. We refer to the remaining parameters $\theta = \{\alpha, \sigma_{\beta}\}$ as the non-linear parameters.
For GMM estimation, the supply and demand residuals are used to form the moment conditions using valid instruments. Identification relies on finding such valid instruments, which typically arise from exclusion restrictions, i.e., observables that enter only one of the demand and supply equations. For a given set of parameters $\Theta$, Equations (ref) and (ref) imply the following residuals:
As berry1994estimating shows, the share equation (ref) is invertible in the sense that given the data $\{\ensuremath{\boldsymbol{\mathbf{x}}}_{t}, \ensuremath{\boldsymbol{\mathbf{p}}}_{t}, \ensuremath{\boldsymbol{\mathbf{s}}}_{t}\}$ and non-linear parameters $\theta$, there is a unique value of the mean utilities $\delta_{jt}$ that rationalize the observed market shares. We denote this by $\ensuremath{\boldsymbol{\mathbf{\delta}}}_t(\theta)$. Then as seen in the first equality of Equation (ref), the implied costs can be recovered from the prices, ownership matrix and the Jacobian of the shares with respect to prices. Notice that these also depend only on the non-linear parameters, and we write the implied costs as a function $\ensuremath{\boldsymbol{\mathbf{c}}}_t(\theta)$.
The GMM estimator requires demand-side instruments $Z_{jt}^d$ and supply-side instruments $Z_{jt}^s$ that satisfy the moment conditions $E[\xi_{jt}Z_{jt}^d] = 0$ and $E[u_{jt}Z_{jt}^s] = 0$. Prior work suggests various ways to construct these instruments berry1995automobile,gandhi2019measuring,conlon2020best. Given instruments, one can form the empirical analogs of the moment conditions:
and proceed to minimize the GMM objective:
where $\Psi$ is a positive definite weighting matrix. In most cases, GMM estimation proceeds in two steps. In the first step, $\Psi$ is typically set to either the identity matrix, or two-stage least squares (2SLS) weighting matrix.\footnote{In our simulations and replication exercise, we initialize $\Psi$ using the 2SLS weighting matrix as in conlon2020best.} In the second step $\Psi$ is replaced with a heteroscedasticity robust weighting matrix that uses the first-step estimates. More details on this estimation routine can be found in BLP, AGS, and CG.\footnote{CG succinctly outlines the estimation procedure in their Algorithm 1.}
\paragraph{Optimization} We follow common practice and structure the optimization problem as an outer loop that searches over non-linear parameters, and an inner loop that recovers the optimal linear parameters given the non-linear parameters from the outer loop. As noted above, for any guess of the non-linear parameters $\theta$ we can generate the implied mean utilities $\ensuremath{\boldsymbol{\mathbf{\delta}}}_t(\theta)$ and marginal costs $\ensuremath{\boldsymbol{\mathbf{c}}}_t(\theta)$. We may then concentrate out the linear parameters $(\beta,\gamma)$ by setting them to the values that minimize the GMM objective for that $\theta$, a process referred to as the “inner loop”. The inner loop has an analytic solution, using an IV-GMM estimator, as described in CG among other references.
\paragraph{Standard Errors} We use the standard formula for heteroskedasticity robust standard errors in GMM. In practice the estimated standard errors depend on how one estimates the gradient of the GMM objective: either by finite differences, automatic differentiation, or using an analytic expression for the gradient. We primarily rely on automatic differentiation.
In this section, we outline our maximum likelihood estimation (MLE) procedure. In this context MLE requires a parametric specification of the joint distribution of the error terms, $\xi$ and $u$. If this assumption holds, MLE is the most statistically efficient estimation procedure available. The identification requirements for MLE are otherwise the same as for GMM: the researcher should have access to a cost shifter to identify the mean price coefficient, and choice set variation to identify the distribution of random coefficients.\footnote{Instead of a cost shifter, one may instead impose that the cost and demand errors are independent, in which case the cost residual itself is a valid instrument for price mackeymiller21.} Thus MLE requires demand and supply-side instruments just as GMM does, but doesn't require the explicit construction of moment conditions. Since MLE can be thought of as GMM with optimal instruments (the first order conditions of the likelihood in the parameters), one can think of MLE as automatically picking the best instruments. This is distinct from the optimal GMM instruments of chamberlain1987asymptotic, which are only semi-parametrically optimal, since that procedure makes no distributional assumptions on the distribution of the residuals.\footnote{In simulations, CG find that using the chamberlain1987asymptotic instruments generally performs better than other approaches for picking instruments, but the performance gains they document from doing so are small compared to those we find from the use of MLE.}
The likelihood will depend on the joint density of the observed and endogenous variables in the model: shares ($\ensuremath{\boldsymbol{\mathbf{s}}}_t$) and prices ($\ensuremath{\boldsymbol{\mathbf{p}}}_t$). As mentioned previously, to estimate the likelihood, we make distributional assumptions. Particularly, we assume that the residuals of the demand and supply unobserved characteristics are i.i.d. (homoskedasticity of the residuals) and follow a joint normal distribution:
The conditional density of interest is $f(\ensuremath{\boldsymbol{\mathbf{s}}}_{t}, \ensuremath{\boldsymbol{\mathbf{p}}}_{t}| \ensuremath{\boldsymbol{\mathbf{x}}}_{t}, \ensuremath{\boldsymbol{\mathbf{w}}}_{t} ;\ensuremath{\boldsymbol{\mathbf{\Theta}}})$, where $f(\cdot)$ is probability density function. We can re-write this density by using multi-variate change of variables and the Inverse Function Theorem:
where $|\cdot|$ is the determinant and $\ensuremath{\boldsymbol{\mathbf{J}}}(\ensuremath{\boldsymbol{\mathbf{s}}}_{t}, \ensuremath{\boldsymbol{\mathbf{p}}}_{t}| \ensuremath{\boldsymbol{\mathbf{x}}}_{t}, \ensuremath{\boldsymbol{\mathbf{w}}}_{t};\theta)$ are the full derivatives of shares and prices with respect to mean utility and costs, a $2N_t \times 2N_t $ matrix:
We derive the Jacobian in Appendix (ref). Notice that it doesn't depend on the parameters $\{\ensuremath{\boldsymbol{\mathbf{\beta}}}, \ensuremath{\boldsymbol{\mathbf{\gamma}}}, \ensuremath{\boldsymbol{\mathbf{\Sigma}}}\}$. The density of $f(\ensuremath{\boldsymbol{\mathbf{\delta}}}_{t}, \ensuremath{\boldsymbol{\mathbf{c}}}_{t} | \ensuremath{\boldsymbol{\mathbf{x}}}_{t}, \ensuremath{\boldsymbol{\mathbf{w}}}_t;\ensuremath{\boldsymbol{\mathbf{\Theta}}})$ can be written as a result of the structure of the model (particularly equations (ref) and (ref)) and the distributional assumptions in equation (ref):
where $\mu$ is a vector-valued function (returning a two-vector), and $\Sigma$ is a $2 \times 2$ matrix. For ease of notation, we denote the data in market $t$ as $\ensuremath{\boldsymbol{\mathbf{\mathcal{D}}}}_t$ and the data across all markets as $\ensuremath{\boldsymbol{\mathbf{\mathcal{D}}}}$. The log-likelihood can be written by expanding out the joint normal distribution for a market $t$ and taking the natural logarithm of all the terms in Equation (ref):
Write $
=
- \mu(\ensuremath{\boldsymbol{\mathbf{x}}}_{jt}, \ensuremath{\boldsymbol{\mathbf{w}}}_{jt};\ensuremath{\boldsymbol{\mathbf{\Theta}}})$, substitute into the likelihood and sum across markets to get the aggregate log-likelihood:
We will refer to the log-likelihood given above as the unconcentrated version of the log-likelihood.
\paragraph{Optimization} To simplify optimization, we will concentrate out a number of parameters. First consider the variance-covariance matrix $\Sigma$. Fixing $\Theta$ and therefore the residuals $\{\xi_{jt}(\Theta),u_{jt}(\Theta)\}$, the likelihood maximizing choice of $\Sigma$ is the implied variance-covariance matrix of the residuals. This is a function of the parameters, which we denote by $\Sigma^*(\Theta)$. As shown in Appendix (ref), making this substitution simplifies the middle term of the log-likelihood in eq (ref) to a constant, resulting in the following likelihood:
Notice that the linear parameters do not enter the Jacobian, only the implied variance-covariance matrix. Thus similar to the GMM procedure, we can concentrate out the linear parameters by solving an inner loop problem of the form:
Now recall that the implied variance covariance-matrix $\Sigma^*(\theta, \beta, \gamma)$ is a $2 \times 2$ matrix, with a determinant equal to $\sigma^2_{\xi}(\theta, \beta, \gamma)\sigma^2_{u}(\theta, \beta, \gamma) - (\sigma_{\xi,u}(\theta, \beta, \gamma))^2$. In the absence of a covariance term, minimizing this objective would be equivalent to minimizing the demand and cost residuals separately. This could be done by separate OLS regressions. But the presence of a covariance term complicates things, leading to a non-linear objective with no closed form solution.\footnote{In the GMM estimation procedure, OLS also fails unless there is no covariance in the errors. But there IV-GMM delivers a closed-form solution for $(\beta,\gamma)$.} In Appendix (ref) we show that $\beta$ and $\gamma$ can be solved in closed form holding the other parameter fixed and so we estimate them jointly by alternating least squares until convergence.
This yields a concentrated likelihood function that depends only on the non-linear parameters, to be maximized in an outer loop:
where we will call part I of the likelihood the covariance term and part II the Jacobian term.
\paragraph{Standard Errors} To compute standard errors, we use the inverse of the Fisher information matrix, which is given in Appendix (ref). For accurate estimation of the standard errors, we use the Fisher information matrix of the unconcentrated likelihood, to allow the uncertainty in the inner loop estimation procedures to be accounted for.
\paragraph{Computational Details and Algorithm} The optimization of the likelihood initially follows the GMM procedure in inverting from market shares to mean utilities and marginal cost vectors $\{\ensuremath{\boldsymbol{\mathbf{\delta}}},\ensuremath{\boldsymbol{\mathbf{c}}}\}$. In order to invert from market shares to mean utilities, we use the SQUAREM method proposed by varadhan2008simple and recommended by CG. The marginal cost is calculated from the Bertrand-Nash equations detailed in the Model section.
The rest of the optimization is best understood when thinking about the covariance term and the Jacobian term of the log-likelihood in equation (ref) separately. The covariance term is evaluated by concentrating out the linear terms using the objective in equation (ref).
The terms entering the Jacobian terms depend on partial hessians of the shares with respect to price. They are derived in detail in Appendix (ref). Note that these Hessians are the computational bottleneck of the estimation routine, having dimensions of size $(T, N_i, max(N_t), max(N_t), max(F_t))$\footnote{The $max(N_t)$ is a result of how we programmed markets. Essentially, in order to have all markets concatenated in a matrix or tensor, the max number of products is used. Additionally, the $max(F_t)$ dimension is a result of the particular hessian objects having a dimension that has to only do with products that are within the same firm.}.
Given data, Algorithm (ref) outlines the procedure of the estimation objective. The code developed for this application uses the PyTorch package. PyTorch provides a high-performance library for optimizing both the GMM and MLE objectives.\footnote{Two reasons for using PyTorch are: (1) utilizing the efficient tensor and matrix operations, (2) utilizing automatic differentiation for optimization.}
In this section, we evaluate the performance of GMM and MLE on simulated data, where the ground truth is known. Our simulation procedure is based on conlon2020best and armstrong2016large.
For each simulation, we set the number of markets to $T = 20$. The number of firms $F_t$ in each market is randomly chosen from the set $\{2, 5, 10\}$. The number of products for each firm $\mathcal{J}_{ft}$ is also chosen randomly from $\{3, 4, 5\}$. The structural error terms are given by: \[
\sim N(0, \Sigma) where \Sigma =
. \] The linear demand characteristics are $[1, x_{jt}, p_{jt}]$ and the linear supply characteristics are $[1, x_{jt}, w_{jt}]$. Both exogenous characteristics ($x_{jt}, w_{jt}$) are drawn from a standard uniform distribution. We allow for two random coefficients on the demand side. The first is on the demand characteristic $x_{jt}$, where $\tilde{\beta}_{i} \sim N(0, \sigma_{\beta}^2)$ and $\sigma_{\beta} = 3$. The second one is on price, where $\tilde{\alpha}_{i} \sim N(0, \sigma_{\alpha}^2)$ and $\sigma_{\alpha}=0.2$. To compute the values of the endogenous variables ($p_{jt}, s_{jt}$) we use the morrow2011fixed method as described in conlon2020best. We use Gauss-Hermite quadrature with a product rule to integrate over the 2-dimensional distribution of random coefficients. We provide details in Appendix (ref). We set the ground-truth values for the demand-side parameters to $[\beta_0, \beta_x, \alpha] = [−7, 6, −1]$, and the supply-side parameters to $[\gamma_0, \gamma_x, \gamma_w] = [2, 1, 0.2]$. To ensure reproducibility, we fix a random seed for each simulation.
We refer to the setup above as “No covariance”, and also consider the following variations:
For GMM estimation, we construct local differentiation instruments following gandhi2019measuring. We also create a predicted price instrument, by regressing price on valid instruments reynaert2014improving.
We compare MLE and GMM for each of the above scenarios on three metrics: parameter estimates, standard errors, and own-price elasticities. For each scenario, we run 1000 simulations, with a different random seed for each simulation\footnote{ Each random seed is also randomly chosen. The random seed is assigned for reproducibility.}. For each simulation, optimization starts at a random initial guess which is drawn uniformly within a 50% range of the true values. Optimization is repeated 3 times with different starting values, after which the parameter estimates are chosen based on the smallest objective value.
To validate that we have correctly coded up the scenarios in CG, as well as the our particular implementation of GMM, we compare our estimates to those produced by PyBLP. The parameter estimates are in most cases indistinguishable. In addition, the mean bias and mean absolute bias we report in our GMM results closely match those in Table 5 of CG.
The scenarios that we run simulations for can be viewed in three different categories. First, the correctly specified scenarios, where both the MLE and the GMM estimators are correctly specified. Second, the MLE misspecified scenarios, which are the Laplace scenarios (GMM is correctly specified in this case). Last, there are the fully misspecified scenarios, where both MLE and GMM are misspecified, these are the Supply Misspecification and Ownership Misspecification scenarios.
\paragraph{Mean Bias and RMSE.} We begin by summarizing the bias and root mean squared error (RMSE) of the estimators, in two Tables: (a) the RMSE of the estimators can be found in Table (ref) and (b) the mean bias can be found in Table (ref). The results show that in general the MLE estimates perform somewhat better when it comes to mean bias, and clearly outperform the GMM estimates when looking at the RMSE values. Another way of putting this is for a particular realization of the data, MLE estimates are generally closer to the true values. A visualization of this can be seen in Figure (ref), where we see the histogram of the parameter estimates of each of the three parameters for the no covariance and low covariance scenarios.
For RMSE, MLE clearly outperforms the GMM for all the correctly specified and MLE misspecified scenarios. In the fully misspecified scenarios, GMM performs better when there is a supply side misspecification and MLE performs better when the ownership matrix is misspecified.
\paragraph{Standard Errors and Coverage.} Standard errors and confidence intervals are fundamental for inference. These results are summarized in two tables: (a) the mean standard error for a scenario in Table (ref) and (b) the coverage percentile of the true values for the 95% confidence intervals in Table (ref).\footnote{For the standard errors of the $\sigma_{price}$ parameter in the GMM procedure, we find that about $10\%$ of the simulations have numerically infeasible values (standard error values ranging from 1000 to 1e13). To validate that this is not an error in our standard error computations, we take a handful of simulations with such issues and run the pyBLP package provided by CG. The outputs were $NaN$ for the corresponding standard error values. It therefore appears that this problem is not specific to our code, but to the GMM procedure implemented here and in pyBLP itself. We drop these cases in the relevant columns of Tables (ref) and (ref); if they were included the performance of GMM would be much worse.} The coverage percentile represents the percent of simulations where:
The mean standard errors for MLE are tighter under all correctly specified scenarios as well as the Laplace scenarios (the one exception to this is the mean standard error of $\alpha$ in the high covariance scenario). Table (ref) shows the coverage of each estimator for each parameter, for a nominal 95%. In general GMM tends to badly under-cover, with true coverage of between 80-90% for the mean price parameter, 90-95% for the standard deviation of the random coefficient on $x$ and 70-80% for the standard deviation of the random coefficient on price. By contrast for MLE those coverages are respectively 88-95%, 90-95% and 90-95%, so that the undercoverage problem is less severe. The sole exception to this general pattern is the case where the cost structure is misspecified, where MLE performs terribly.
\paragraph{Own-price elasticities} Demand systems are often used to generate own-price elasticities. Given the interest in these statistics, we examine what the mean bias and absolute mean bias of product-level own-price elasticities are for a given set of estimated parameters. Figure (ref) shows this for the no covariance and low covariance scenarios. The results support the conclusion from the previous sections, the MLE produces tighter results in most of our scenarios.
\paragraph{Jacobian.} One notable feature of MLE is the presence of a Jacobian in the likelihood. This accounts for the non-linear mapping from marginal costs and mean utilities to the (endogenous) prices and quantities. One might think that minimizing the covariance term in the likelihood would be enough to ensure good performance, since doing so in some sense minimizes the variance of the demand and cost residuals. We investigate the importance of the Jacobian term in the likelihood here.
Figure (ref) plots the likelihood decomposition into the covariance and Jacobian terms separately, where we vary the $\sigma_x$ parameter at the true $\alpha=1$ value (for a random simulation). The covariance term is minimized below the truth, and the Jacobian term is necessary to correct the bias in the covariance term and arrive at the correct solution. The Jacobian thus plays an important role, in some sense regularizing the estimator by penalizing parameters under which the derivatives of the endogenous variables in the exogenous variables are big (i.e. parameters at which the predictions are unstable).
In this section, we describe our replication of BLP'95 using MLE. We use the same data as berry1995automobile, provided by andrews2017measuring. The data contains US automobile data from 1971 to 1990. The dataset also includes 5 product characteristics, horse power to weight ratio (hpwt), if a car has air conditioning (air), miles per dollar of gasoline (mpd), width times length (space), and miles per gallon (mpg).
We start by replicating the GMM estimation procedure. We follow the procedure outlined in AGS, which is based off of BLP. This includes dropping certain instruments due to collinearity as well as starting the estimation procedure from the estimated parameters in BLP. As in BLP, we have 2217 model/years with 997 distinct models. The instruments that we use are BLP sums of characteristic instruments. With this setup, the local GMM estimation produces estimates that are almost identical to the AGS parameters.
Next, we estimate the parameters using MLE. For the MLE setup, we run the procedure on the same dataset. Our integration approach follows AGS by using the same random draws and importance sampling. One change we make is to double the number of draws to ensure symmetry. More specifically, we take the set of draws, multiply the vector associated with air conditioning with negative one, and concatenate these to the original draws. This leaves us with two times the total draws.\footnote{When estimating the model by MLE using the AGS draws, the parameter estimates were mostly similar, but $\sigma_{air}$ value was estimated to be 0. This is a result of an artificial lack of symmetry around zero of the objective function in $\sigma_{air}$ due to sampling variance in the draws; doubling the number of draws fixes this problem.}
To get a sense of how reasonable our results are, we compare our estimates to existing benchmarks in the literature that use this dataset. These benchmarks are the estimates reported in BLP, AGS, and CG. These benchmarks differ in two main ways, the instruments selected and the choice of draws (and weights) used for integration. BLP use the sum of rival characteristics (BLP) instruments and integrate using importance sampling. AGS follows BLP closely in order to replicate their estimates. CG uses the chamberlain1987asymptotic optimal instruments and quasi-Monte Carlo integration (10,000 scrambled Halton draws in each market).
The estimation results can be found in Table (ref) with standard error comparisons in parentheses below the parameter estimates. The parameter estimates of MLE are closely aligned with the GMM estimates. In particular, the $\alpha$ parameter is closest to the CG best practice estimates. The sigma estimates tend to align with the BLP results, and generally indicate larger heterogeneity in the data than the CG results. The similarity of the parameter estimates is surprising since the effort put into calibrating the MLE estimator to the automobile data was minimal i.e. these results are “out-of-the-box”. No calibration or instrument construction was necessary. Standard error estimates of the MLE parameters are significantly tighter for MLE, regularly one quarter of the smallest GMM standard error. The standard error on $\alpha$ is 0.769, compared to a range of 5.6-11.7 for the GMM. This is nearly 10-fold reduction compared to GMM. Since we don't know the ground truth we cannot comment on whether the estimated standard errors imply confidence intervals with correct coverage, but the simulations above give us confidence that if the model is correctly specified the coverage will be approximately correct.
Finally, we look at the model level comparison of the own-price elasticities of the GMM and MLE estimators in Figure (ref). We can see that there is a fairly large mass around the diagonal, indicating generally similar estimates. The most significant trend off of the diagonal is a section of the AGS estimates. This trend can be fully attributed to the large difference in the $\sigma_{air}$ estimates, where AGS estimates this parameter to be 4.2, compared to BLP, CG, and MLE estimating it to between 1.4-2. Since the air conditioning variable is binary, it leads to two separate trend lines, where the one above the diagonal representing models that have air conditioning.
Overall, we were able to use MLE on the BLP automobile data with relative ease and produce estimates that are similar to prior work but with tighter standard errors and without having to the construct valid instruments.
We develop and evaluate a maximum likelihood estimator for differentiated products demand. Unlike the traditionally used GMM estimator, the ML estimator requires stricter distributional assumptions, while offering the advantages of statistical efficiency and limiting choices required by the researcher to fit the model. While we are not the first to propose an ML estimator, prior work has relied on reduced-form supply specifications, whereas our supply-side pricing is consistent with Bertrand-Nash pricing.
In the simulation section, we show that under correct specification, our method outperforms GMM as well as showing its robustness and sensitivity to different types of misspecification. Finally, we replicate BLP using the new estimation method, finding that the estimates are very similar by comparing the estimated values to other benchmarks. These results demonstrate that MLE can be a useful addition to the tools available to researchers estimating BLP.