Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
103,970 characters · 21 sections · 63 citation commands
Bundle Choice Model with Endogenous Regressors: An Application to Soda Tax
The estimation of complementary choices addresses important economic and policy questions. Individuals and firms make complementary choices, such as the demand for milk and cereal lee2013direct, firm-level computer purchases hendel1999estimating, and library subscriptions nevo2005academic. Understanding these choices not only sheds light on consumption behavior but also informs broader economic policies and strategies. liu2010complementarities find that complementarities drive firms' bundling of cable TV, local phone, and broadband. ershov2024estimating show that merger analyses ignoring complementarity may overstate consumer harm. In the context of a soda tax, sugar-sweetened beverages and other unhealthy foods may be complements. If this is the case, empirical studies that ignore these complementarities may fail to capture the tax’s broader effects on unhealthy complementary goods and underestimate its full health benefits.
Standard discrete choice models typically assume that all choices are mutually exclusive, thereby ruling out joint consumption and complementarity. Relaxing this assumption induces the curse of dimensionality, as it requires redefining the choice set to include all possible bundles of choices. Consequently, most empirical studies gentzkow2007valuing,deza2015there focus on applications involving only two or three products and impose restrictive assumptions. The need to account for endogenous regressors, a common issue in demand estimation, further exacerbates these dimensionality challenges. The few recent empirical studies iaria2020identification,wang2023testing,ershov2024estimating are limited to addressing market-level endogeneity, such as market-level prices, or solely accounting for individual fixed effects.
This paper proposes a Bayesian factor-augmented bundle choice model to estimate joint consumption, and the substitutability and complementarity of multiple goods, in the presence of endogenous regressors. The model addresses dimensionality challenges by employing a sparse factor approach to account for taste correlation and endogeneity arising from unobserved individual heterogeneities in preferences and other unobserved variables. Time-varying factor loadings allow for more general individual-level and time-varying heterogeneity and endogeneity, while the sparsity induced by the shrinkage prior on loadings balances flexibility and parsimony. A computationally efficient posterior inference algorithm is developed with a Markov chain Monte Carlo (MCMC) sampling scheme, which also enables posterior density predictions for estimating choice probabilities, price elasticities, and conducting counterfactual policy analysis.
To address endogeneity, the proposed model extends the model with complementarity gentzkow2007valuing and incorporates a Bayesian instrumental variables approach rossi2012bayesian,lopes2014bayesian. It generalizes the two primary treatments of endogeneity in existing bundle choice models: (1) market-level endogenous regressors (e.g., prices) driven by time-varying unobserved variables (e.g., unobserved choice/product attributes and demand shocks) and addressed using market-level instruments iaria2020identification,ershov2024estimating, and (2) individual-level endogenous regressors induced by time-invariant unobserved individual heterogeneity (e.g., taste for quality) and addressed using a fixed-effects approach wang2023testing. The proposed model addresses individual-level endogenous regressors resulting from time-varying unobserved variables using individual-level instruments. This generality allows for broader applicability, and the richer individual-level information aids in identification and empirical implementation.
I employ a factor approach to model the high-dimensional error correlations that lead to taste correlation and endogeneity. Accounting for taste correlation is a central challenge when estimating complementarity gentzkow2007valuing, as the observed joint consumption of two goods may result from either genuine complementarity or positive taste correlation. The factor approach provides a flexible and parsimonious method for modeling complex correlation structures—between choices (taste correlation), between choices and endogenous regressors (endogeneity), and across time periods. The time-varying factor loadings allow unobserved individual heterogeneity in preferences and demand shocks, which cause taste correlation and endogeneity, to vary over time. In contrast, existing approaches for handling taste correlation (e.g., random effects) are not readily scalable to applications involving endogeneity and individual-level, time-varying settings due to dimensionality constraints.
Standard factor models require researchers to choose the number of latent factors to balance flexibility and parsimony, leading to model selection challenges. Inspired by advances in sparse Bayesian factor analysis, I adopt a cumulative shrinkage prior bhattacharya2011sparse,durante2017note on the factor loadings. This prior introduces sparsity into the factor loading matrix by shrinking unimportant columns and elements to zero, allowing the model to automatically determine the number of effective factors and loadings. This approach addresses model selection issues in existing Bayesian choice models with a factor structure loaiza2022scalable,jacobi2024working by jointly inferring both model parameters and complexity. Additionally, the asymmetric nature of the prior resolves the label-switching identification problem, ensuring a unique factor structure.
The proposed model integrates bundle choice and factor structures within a Bayesian multinomial probit framework mcculloch1994exact, facilitating the use of data augmentation techniques to bypass the need for evaluating high-dimensional integrals when calculating choice probabilities. I develop a computationally efficient posterior inference algorithm using a Markov chain Monte Carlo (MCMC) sampling scheme. To improve algorithmic and computational efficiency, the model is fully vectorized, which mitigates dimensionality challenges and enhances the tractability of applications with modern choice data. This algorithm also enables posterior density predictions for estimating choice probabilities, price elasticities, and conducting counterfactual policy analysis.
Monte Carlo simulation studies confirm that the proposed model accurately recovers the true parameters and provides reliable estimates of price elasticity. Comparisons between the proposed model and simpler benchmark models (time-invariant or exogenous) highlight the importance of controlling for price endogeneity when estimating complementarity, as well as the necessity of accounting for time-varying unobserved individual heterogeneities and demand shocks. Furthermore, the simulation studies demonstrate that as the number of individuals ($N$) or time periods ($T$) increases, the estimated own- and cross-price elasticities converge to their true values.
I apply the proposed model to evaluate the implications of implementing a sugary drink tax (also known as a soda tax) in the context of complementarities. A primary rationale for introducing a soda tax is to reduce sugar consumption. However, concerns have emerged allcott2019should. First, such a tax, if applied only to sugary soft drinks, may lead to increased consumption of alternative sugary products. Ignoring the potential for substitution could result in an overestimation of the tax's impact on overall sugar intake. Second, sugar-sweetened beverages and other unhealthy food items may be complements. If this is the case, existing empirical studies might underestimate the health benefits derived from a soda tax. Despite extensive research on substitution effects (see allcott2019should for an extensive review), the potential complementarity between sugary beverages and other unhealthy foods remains understudied. This gap exists because standard demand models do not allow for the simultaneous estimation of substitutability and complementarity within a unified framework.
The proposed bundle demand model overcomes the modeling challenge, allowing substitutability and complementarity between the goods. I estimate the substitution patterns between sugary soft drinks, diet soft drinks, carbonated water, milk drinks and salty snacks, using the IRI Marketing Data Set bronnenberg2008database. The empirical results confirm that soda taxes decrease the intake of sugar from sugar-sweetened beverages, consequently leading to the intended health benefits. The results also suggest that the relationship between sugary soft drinks and milk drinks demonstrates independence. This finding implies that a soda tax would not lead consumers to substitute milk drinks for sugary soft drinks as an alternative sugar source. Furthermore, sugary soft drinks and salty snacks are identified as complements. This relationship suggests that the implementation of a soda tax could yield additional health benefits by marginally decreasing the consumption of salty snacks along with sugary drinks, thus extending the health benefits of the tax beyond merely reducing sugar consumption.
There is a growing empirical literature on the estimation of demand for bundles and complementarity gentzkow2007valuing,crawford2012welfare,liu2010complementarities,ho2012use,gentzkow2014competition,grzybowski2016substitution,thomassen2017multi,crawford2018welfare,ershov2024estimating,ouyang2023semiparametric. This paper is along the line of bundle choice models accounting for endogeneity iaria2020identification,wang2023testing,ershov2024estimating. The sparse factor approach addresses the dimensionality challenges and accommodate more general individual-level and time-varying heterogeneity and endogeneity, and individual-level instruments. The factor approach provides an alternative to the demand inversion in models with a BLP structure berry1995automobile.
This paper introduces the first Bayesian bundle demand model for panel data, extending the Bayesian framework for cross-sectional data as developed by sovinsky2024working for bundle demand and jacobi2024working for bundle demand with endogenous choice sets. Panel data leads to gains in identification from repeated observations.
This paper employs the Bayesian factor approach to model taste correlation and endogeneity. Known for its flexibility and parsimony in modeling error correlations, as in lopes2004bayesian, philipov2006factor, chib2006analysis, han2006asset, lopes2007factor, nakajima2012dynamic, zhou2014bayesian, ishihara2017portfolio, fruehwirthschnatter2023sparse, kastner2019sparse, the factor approach has been utilized in various microeconometric contexts. These include estimating the causal effect of an endogenous binary treatment on correlated panel outcomes heckman2014treatment, jacobi2016bayesian, wagner2023factor, and in modeling error (taste) correlations in the multinomial probit model loaiza2022scalable. jacobi2024working were the first to integrate this approach into a cross-sectional bundle demand model with endogenous choice sets. My model builds upon their work by incorporating more flexible structures and allowing for time-varying factor loadings. I employ the sparse factor approach bhattacharya2011sparse,durante2017note,kastner2019sparse to balance flexibility and parsimony and addresses model selection issue.
The application to sugary drink tax contributes to the burgeoning literature aiming to simulate the effects of soda taxes using behavior estimates from markets and periods without an existing soda tax harding2017effects, bonnet2013assessment, wang2015soda, allcott2019design, chernozhukov2019quantifying, dubois2020well. To my knowledge, this study is the first to investigate the complementarity between sugar-sweetened beverages, that are subject to taxation, and other unhealthy foods. Furthermore, the proposed model uses individual-level prices, and accounts for price endogeneity, thereby providing more reliable elasticity estimates.
The rest of the paper is organized as follows. Section (ref) presents the Bayesian factor-augmented bundle demand model that accommodates endogenous regressors. Section (ref) discusses the choice of prior distributions and develops posterior inference and counterfactual density predictions for choice probability and price elasticity. Section (ref) presents the simulation studies. Section (ref) evaluates the implications of implementing a sugary drink tax in the context of complementarities. The paper concludes in Section (ref).
This section presents the Bayesian factor-augmented bundle demand model that accommodates endogenous regressors. The model is an extension of Gentzkow's model with complementarity gentzkow2007valuing, with three distinct features. First, as introduced in Subsection (ref), the base model assumes standard normal idiosyncratic errors to facilitate Bayesian posterior inference. Second, Subsection (ref) extends the base model to include endogenous regressors, using price as an example. Third, instead of modeling taste correlations with time-invariant random effects, Subsection (ref) introduces a factor approach to model error correlations induced by unobserved taste correlation and unobserved utility-price confounders.
Subsection (ref) extends the two-good illustration to the general model incorporating $J$ goods and $J_p$ endogenous regressors. It also presents the model vectorization to further illustrate (intertemporal) error correlations and facilitate posterior inference in the subsequent section. Subsection (ref) compares three alternative modeling approaches for error correlations, and Subsection (ref) discusses identification issues.
Consider a scenario with two goods, denoted as $j=1,2$. Each individual $i=1,\ldots,N$ at time $t=1,\ldots,T$ chooses to consume at most one unit of each good. The individual's indirect utility from consuming each good, $j=1,2$, is
where $p_{ijt}$ denotes the price of good $j$ faced by individual $i$ at time $t$. The price interacts with the price coefficient $\alpha$. The vector $x_{ijt}$ includes covariates associated with the individual $i$, good $j$ and time $t$. These covariates may contain good-specific intercepts, individual and product characteristics, market information and fixed effects. The parameter $\beta_j$ is a vector of good-specific preference parameters. The structural error, $\nu_{ijt}$, accounts for individual $i$'s good-specific unobserved tastes, which may vary over time. Importantly, these tastes may be correlated across the two goods.
Choosing among 2 goods defines 4 choices, denoted as $\mathcal{R}=\{0,1,2,(1,2)\}$. $r=0$ represents the outside option of consuming none of the goods, and $r=(1,2)$ the choice of consuming both goods jointly. The utility an individual obtains from consuming bundle $r\in\mathcal{R}$ is given by
where $\epsilon_{irt}$ denote the idiosyncratic errors or i.i.d. utility shocks. The utility for the outside option is normalized to contain only the error term, i.e., $u_{i0r}=\epsilon_{i0t}$, as the absolute levels of utilities are not identifiable. Let the idiosyncratic utility shocks follow the standard normal distribution, denoted as $\epsilon_{irt}\sim\mathcal{N}(0,1)$ for $r\in\mathcal{R}$. This induces a multinomial probit (MNP) model.
The bundle-effects parameter $\Gamma_{(1,2)}$ represents the additional utility (or disutility) from consuming good 1 and 2 together. It captures the substitutability and complementarity between the goods. gentzkow2007valuing shows that two goods are complements if $\Gamma>0$, substitutes if $\Gamma<0$, or independent if $\Gamma=0$, in terms of Hicksian cross-price elasticity of demand. Furthermore, $\Gamma$ can also reflect shopping costs and preference for variety that lead to individuals purchasing a variety of goods in a time period iaria2020identification.
As standard random utility models, the observed choice, $y_{it}$, is then determined by the highest latent utility, as $y_{it}=\arg\max_r(u_{irt})$. The error distribution leads to the choice probability:
Here, $u_{i,-r,t}$, with $-r\in\mathcal{R}\backslash r$, represents the latent utilities for the unchosen bundles, and $F(\epsilon)$ is the cumulative distribution function of the idiosyncratic errors. The integral’s range can be expressed in terms of $\epsilon_{irt}$ following equation (ref). For example, when good 1 is chosen, i.e. $y_{it}=1$, the integral range can be written as a region within the support of the joint distribution $F(\epsilon_{i0t},\epsilon_{i1t},\epsilon_{i2t},\epsilon_{i(1,2)t})$, as shown below:
In practice, I employ the Bayesian data augmentation technique to avoid the need for evaluating the high-dimensional integrals above for calculating choice probabilities.
The correlation between the unobserved tastes (structural errors) $\nu_{i1t}$ and $\nu_{i2t}$ plays a crucial conceptual role in identifying bundle effects $\Gamma_{(1,2)}$ and, in turn, assessing complementarity. The observed simultaneous use of two goods can arise from either genuine complementarity between them ($\Gamma_{(1,2)}>0$) or positive taste correlation ($\text{corr}(\nu_{i1t},\nu_{i2t})>0$). Distinguishing bundle effects from this (unobserved) taste correlation constitutes the central modeling and identification challenge in bunde demand models. As noted by gentzkow2007valuing, failing to account for possible correlations in the indirect utilities of bundles containing overlapping products may lead to the identification of spurious bundle effects and complementarity.
Endogeneity typically refers to situations where observed regressors are correlated with error terms, due to omitted variables or simultaneity. A classic example in choice models is the endogeneity of price, which arises from unobserved product characteristics berry1995automobile,nevo2000practitioner. This subsection also uses endogenous price as an example to extend the model to accommodate endogenous regressors.
Consider the price $p_{ijt}$ faced by individual $i$ at time $t$. This model allows prices to differ across individuals, which is a common occurrence in individual-level data due to variations in product baskets, membership discounts, volume discounts, bargaining, and other reasons.
The endogenous price $p_{ijt}$ for each good $j=1,2$ is modeled by a “first-stage" equation:
where $z_{ijt}^p$ contains a set of instrumental variables, which also includes the exogenous covariates in the utility equations (structural equations). $\theta_j^p$ is a vector of parameters to be estimated. The endogeneity arises from the correlation between the structural errors $\nu_{ijt}^p$ and $\nu_{ijt}$ from the reduced form and structural equations, respectively. This link induces a correlation between the endogenous price $p_{ijt}$ and unobserved characteristics $\nu_{ijt}$. Equivalently, we could consider that unobserved confounders, which determine both prices and utilities, are captured by the structural errors $\nu_{ijt}^p$ and $\nu_{ijt}$. This approach aligns with the Bayesian treatment of instrumental variables in linear models rossi2012bayesian. Modeling and identifying these error correlations is crucial in addressing endogeneity, an issue closely linked to unobserved tastes in the base model.
Error correlations play an important role in identifying complementarity in the presence of endogenous regressors. To consolidate this point, I define joint errors $\varepsilon$ as the summations of structural errors $\nu$ and idiosyncratic errors $\epsilon$ in the bundle utility equations (ref) (substituting the good utility equation (ref)), and the first-stage equations (ref). The idiosyncratic error only outside option $u_{i0t}=\epsilon_{i0t}$ is omitted.
where
and $\mathbf{I}_5$ denotes an identity matrix with dimension 5.
To estimate complementarity (bundle effects $\Gamma_{(1,2)}$), it is necessary to address the error correlation $\text{corr}(\nu_{i1t},\nu_{i2t})$. gentzkow2007valuing and sovinsky2024working tackle this issue using random effects. Their models assume that the vector of unobserved tastes is time-invariant and follows a multivariate normal distribution $(\nu_{i1},\nu_{i2},\ldots)'\sim\mathcal{N}(\mathbf{0},\Sigma)$ with the full covariance matrix $\Sigma$ capturing the correlation. deza2015there models unobserved tastes using a latent class model, where the taste correlations are captured by class-specific intercepts. Similarly, addressing endogenous regressors involves modeling error correlations, $\text{corr}(\nu_{i1t},\nu_{i1t}^p)$ and $\text{corr}(\nu_{i2t},\nu_{i2t}^p)$. Standard Bayesian approaches employ a bivariate normal distribution rossi2012bayesian or a mixture of bivariate normal distributions conley2008semi to model errors in the first-stage and structural equations.
The existing approaches largely rely on multivariate (mostly Gaussian) distributional assumptions for modeling errors. However, these established methods face challenges in the current context. First, the number of parameters in the full covariance matrix increases exponentially with the numbers of goods and endogenous regressors. While the models previously mentioned address error correlations stemming either from taste correlations or endogeneity, they do not tackle both simultaneously. In contrast, the model presented in this paper tackles error correlations due to both sources, resulting in higher-dimensional correlations. For instance, with $J=6$ goods and an equal number of endogenous regressors (prices), $J_p=6$, a full covariance matrix $\Sigma$ for time-invariant random effects would have 78 parameters. Identifying a large covariance matrix poses practical challenges. Moreover, most existing models are cross-sectional, e.g. sovinsky2024working, rossi2012bayesian and conley2008semi. For panel models, such as gentzkow2007valuing and deza2015there, unobserved tastes remain time-invariant, i.e., $\nu_{ij}$. In contrast, the model presented here, formulated for panel data, allows for time-varying unobserved tastes, $\nu_{ijt}$. This added flexibility leads to a larger parameter space. Consider the same example with $J=J_p=6$, but now over $T=12$ time periods. A time-varying random-effects model without any specific structure would comprise 10,440 variance and covariance parameters.
To address the challenges within a unified framework, I employ a factor approach to model the structural errors $\nu_{it}$, as follows:
where $\Lambda_t$ is a $4$-by-$L$ factor loading matrix, and $f_i\sim\mathcal{N}(\mathbf{0},\mathbf{I}_L)$ is an $L$-by-$1$ vector of independent factors with zero means and an identity covariance matrix of dimension $L$. The joint errors, as defined in Equation (ref), now take the form $\varepsilon_{it}=I_\nu\Lambda_t f_i+\epsilon_{it}\sim\mathcal{N}(\mathbf{0},\Omega_{it})$ with the covariance decomposition
To illustrate how the factor structure captures key error correlations, consider a tri-factor ($L=3$) example, as follows:
In this specification, the three latent factors account for taste correlation and unobserved confounders of price and utility for goods 1 and 2, respectively. The covariance matrix $\Omega_{it}$ of the joint errors $\varepsilon_{it}$, defined in Equation (ref), takes the following form: {\tiny
}The term $\lambda_{11t}\lambda_{21t}$ captures the correlation between unobserved tastes for good 1 and 2. The terms $\lambda_{12}\lambda_{p12}$ and $\lambda_{23}\lambda_{p23}$ capture the error correlations resulting from price endogeneity.
The factor approach is a flexible and parsimonious method for modeling correlations resulting from endogeneity, between choices and across time periods. In the above example, a time-invariant random-effects model with a full covariance matrix for the structural errors $\nu_{it}$ would require 10 parameters. While, the factor structure models these correlations more parsimoniously with 6 parameters. This method finds application in various models, such as estimating the causal effect of an endogenous binary treatment on correlated panel outcomes jacobi2016bayesian,wagner2023factor, and in modeling error (taste) correlations in multinomial probit model loaiza2022scalable. jacobi2024working introduce the factor approach into a cross-sectional bundle demand model with choice set endogeneity. This paper extends their approach to a panel data setting with more general and time-varying factor loading matrices, and applies it to models with endogenous regressors.
This subsection discusses the identification of the exogenous bundle demand model, the extended components of endogenous regressor, and the augmented factor loadings in sequence, in the context of 2 goods.
Identification of the bundle demand model involves the mean utility parameters $\theta_1$ and $\theta_2$, bundle effects $\Gamma_{(1,2)}$, and the error distribution $F(\nu_{i1t},\nu_{i2t})$. The mean utility parameters $\theta_1$ and $\theta_2$ are identified by the observed choice probabilities $\Pr(y_{it}=1)$ and $\Pr(y_{it}=2)$. The key challenge lies in separately identifying the bundle effects $\Gamma_{(1,2)}$ and the error correlation $\mathrm{corr}(\nu_{i1t}, \nu_{i2t})$. If goods 1 and 2 are frequently consumed together (high $\Pr(y_{it}=(1,2))$), this can be attributed to either a high value of $\Gamma_{(1,2)}$ (indicating complementarity) or a high value of $\mathrm{corr}(\nu_{i1t}, \nu_{i2t})$ (indicating positively correlated tastes for the goods).
One source of identification is exclusion restrictions. That is, there exists a variable, such as price, that enters the utility of one product, $\bar{u}_{ijt}$, but not the other, $\bar{u}_{i,-j,t}$. As argued by gentzkow2007valuing, if the goods are complements ($\Gamma_{(1,2)}>0$), shifting up the utility of good 1 (e.g., by decreasing its price $p_1$) should increase the probability of consuming good 2. In contrast, if $\Gamma_{(1,2)}=0$, the probability of consuming good 2 should remain unchanged. Formalizing Gentzkow's identification argument, fox2017note prove that an excluded variable for each product utility $\bar{u}_{ijt}$ with large support (i.e., spanning the real line $\mathbb{R}$) provides the separate identification of the bundle effects $\Gamma{(1,2)}$ and the joint error distribution $F(\nu_{i1t},\nu_{i2t})$. Additionally, keane1992note provides evidence on the role of exclusion restrictions in identifying the covariance parameters in multinomial probit models through Monte Carlo tests and an application to real-world data.
Identification of the bundle demand model with endogeneity builds upon the argument presented above. The identification relies on excluded variables, or instruments, as in instrumental variables models. lewbel2000semiparametric develops identification and estimation of binary and multinomial choice models with endogeneity, leveraging the special regressor assumption. Under this assumption, each utility and first-stage equation requires a special regressor with large support (real line $\mathbb{R}$) that is excluded from other equations. As a side note, fox2017note employs a similar assumption.
The identification of the factor-augmented model with endogeneity involves two steps: (1) identifying model parameters, $\theta$ and $\Gamma$, and error distribution, $F(\nu)$, and (2) recovering the factor loadings $\Lambda$ from $F(\nu)$. Meanwhile, in most cases, since the estimation of the actual factor loadings is not the primary concern but rather than a means to estimate and predict the covariance structure of $F(\nu)$, a unique identification of the factor loadings $\Lambda$ is not necessary. This allows leaving the factor loading matrix completely unrestricted.
This subsection presents the general model incorporating $J$ goods and $J_p$ endogenous regressors. The model is then vectorized to further illustrate (intertemporal) error correlations and facilitate posterior inference in the subsequent section.
Index goods by $j\in\mathcal{J}=\{1,\ldots,J\}$. Each individual $i\in\mathcal{N}=\{1,\ldots,N\}$ at time $t\in\mathcal{T}_i=\{1,\ldots,T_i\}$ chooses to consume a bundle (subset) of the goods, i.e., $r\subseteq\mathcal{J}$. Denote all time periods by $\mathcal{T}=\bigcup_{i\in\mathcal{N}}\mathcal{T}_i$ and the complete choice set by $\mathcal{R}=\{r\subseteq\mathcal{J}\}$. Note that the outside option (consuming none of $\mathcal{J}$) is in the choice set, i.e., $r=\varnothing\in\mathcal{R}$. As a matter of convention, I relabel the outside option as $r=0$.
Restate the indirect utilities for individual $i$ from consuming a single good $j$ at time $t$ (structural equation) and the first-stage equations for endogenous regressors $p_{ijt}$ in $z_{ijt}$, as defined in Equations (ref) and (ref),
where vector $z_{ijt}$ collects the endogenous and exogenous regressors that enter good utility equation, $\theta_j$ is a vector of good-specific preference parameters, and $\mathcal{J}_p$ denotes the set of endogenous regressors.
Similar to Equation (ref), define the utility of individual $i$ from consuming bundle $r$ at time $t$ as
for inside options $r\neq0$, and $u_{i0t}=\epsilon_{i0t}$ for the outside option. The first term on the right-hand side sums the utilities derived from each good $j$ in the bundle $r$. The second term, bundle effects or demand synergies, $\Gamma_{irt}=\sum_{j_1,j_2\in r}w_{i(j_1,j_2)t}'\gamma_{(j_1,j_2)}$ captures the extra utility or disutility from consuming each pair of goods $j_1$ and $j_2$ jointly in bundle $r$.\footnote{gentzkow2007valuing finds that allowing tri-good bundle effects barely changes the results. Besides, multi-good bundle effects lead to the curse of dimensionality in the number of goods.} The bundle effects are modeled as a function of observable characteristics (and an intercept) $w_{i(j_1,j_2)t}$. The idiosyncratic utility shocks follow the standard normal distribution, denoted as $\epsilon_{irt}\sim\mathcal{N}(0,1)$ for $r\in\mathcal{R}$, inducing a multinomial probit (MNP) model. The observed choice, $y_{it}$, is determined by the highest latent utility, $y_{it}=\arg\max_{r\in\mathcal{R}}(u_{irt})$.
I highlight model features that extend the illustrative model presented in Subsection (ref). First, the general model incorporates $J$ goods with $J_p=|\mathcal{J}_p|$ endogenous regressors. As the number of (inside) options $R=2^J-1$ increases exponentially with $J$, it is recommended in practice to limit $J\leq7$. To accommodate a larger number of goods, one may consider limiting the choice set to single-good and two-good bundles, as in iaria2020identification. Additionally, $J_p$ does not necessarily equal to $J$; each good's utility, $\bar{u}_{ijt}$, can include a different number of endogenous regressors. Second, bundle effects, $\Gamma_{irt}$, are now time-varying and heterogeneous across individuals, as a function of observable characteristics $w_{i(j_1,j_2)t}$. Third, the model allows for common coefficients across regressors in different equations. For example, the same price coefficient $\alpha$ interacts with prices in all utility equations $\alpha p_{ijt}$ for all $j\in\mathcal{J}$. This homogeneous-effects restriction can be similarly applied to other regressors in $z_{ijt}$, $z^p_{ijt}$ and $w_{i(j_1,j_2)t}$. Lastly, the model allows sparsity in the factor loading matrix $\Lambda_t$ implied by prior knowledge or economic theory, as in example (ref).
Next, I define the vectorized model. Substitute good utility (ref) into bundle utility (ref):
Stack bundle utilities over inside options $r\in\mathcal{R}\backslash0$, and first-stage equations over $j\in\mathcal{J}_p$:
where parameter vectors are defined as $\theta=(\theta_1',\ldots,\theta_J')'$, $\gamma=(\gamma_{(1,2)}',\ldots,\gamma_{(J-1,J)}')'$, $\theta^p=({\theta^p_1}',\ldots,{\theta^p_{J_p}}')'$. Similarly, structural error vectors are $\nu^u_{it}=(\nu_{i1t},\ldots,\nu_{iJt})'$ and $\nu^p_{it}=(\nu^p_{i1t},\ldots,\nu^p_{iJ_pt})'$. Further stack the latent utilities and endogenous regressors $y_{it}^*=(u_{it}',p_{it}')'$:
where $\Theta=(\theta',\gamma',{\theta^p}')'$ and $\nu_{it}=({\nu_{it}^u}',{\nu_{it}^p}')'=\Lambda_t f_i$ augmented with the factor structure. Lastly, stack over $t\in\mathcal{T}_i$ for each individual $i$, and then over $i\in\mathcal{N}$:
Details on model vectorization, along with an example with $J=3$ goods and $J_p=3$ endogenous regressors, are provided in Appendix (ref). The vectorized model promotes the forthcoming comparison of alternative modeling approaches for error correlations, defines more compact notations, and facilitates an efficient Markov-Chain Monte Carlo (MCMC) sampling scheme for posterior inference.
This subsection compares three alternative modeling approaches for error correlations. They differ in the distributional assumptions imposed on the unobserved tastes and utility-price confounders (structural errors), $\nu_{it}=(\nu_{i1t},\ldots,\nu_{iJt},\nu^p_{i1t},\ldots,\nu^p_{iJ_pt})$.
Both the RE and FA approaches restrict unobserved tastes and utility-price confounders, denoted as $\nu_i$, to be time-invariant. Consequently, they are inadequate for capturing temporal demand shocks that impact both utilities and market equilibrium prices. Similarly, these approaches fall short in accounting for temporal utility shocks that influence preferences for multiple goods.
Furthermore, the random-effects approach cannot be readily extended to account for time-varying $\nu_{it}$ due to dimensionality challenges. A time-varying random effects would be defined on the vector $\nu_{i}=(\nu_{i1},\ldots,\nu_{iT})'\sim\mathcal{N}(\mathbf{0},\Sigma)$. Consider an application with $J=6$ goods and an equal number of endogenous regressors $J_p=6$, across $T=12$ periods. The full covariance matrix $\Sigma$ would then have $(J+J_p)T\times((J+J_p)T+1)/2=10,440$ free parameters.
To illustrate how the three approaches capture error correlations induced by unobserved tastes and utility-price confounders, one can examine the covariance matrix $\Omega_i$ of the joint error vector $\varepsilon_i$. The process of model vectorization defines joint error vectors at different levels of vectorization. In Equations (ref), (ref), for $t\in\mathcal{T}_i$ and $i\in\mathcal{N}$,
The covariance parameters in $\Omega_{it}$ capture error correlations between latent utilities (taste correlations) and between utility and endogenous regressors (endogeneity) within each time period $t\in\mathcal{T}$.
In Equation (ref) and its RE and time-invariant FA counterparts, for $i\in\mathcal{N}$,
Covariance matrix $\Omega_{i}$ further captures intertemporal error correlations.
Table (ref) summarizes the key variances and covariances in the joint error covariance matrix $\Omega_i$ across models. It highlights the differences in how the RE (Random Effects), FA (Factor-Augmented), and TV-FA (Time-Varying Factor-Augmented) approaches capture error correlations induced by unobserved tastes and utility-price confounders. Notably, the intertemporal correlations are fixed for both the RE and FA models.
This section discusses the selection of prior distributions and the sampling scheme for the three groups of model parameters: augmented latent utilities $u$, equation parameters $\Theta$, and factors and loadings $\lambda$ and $f$.
The joint posterior distribution of the model parameters, conditional on observed data (bundle choices $y$ and endogenous regressors $p$) can be written as a product of the likelihood and the prior, $p(u,\Theta,\lambda,f)\propto p(y,u,p|\Theta,\lambda,f)\times p(\Theta)p(\lambda)p(f)$. The likelihood component, augmented with latent utilities, of the joint posterior is as follows:
Details on the Bayesian probabilistic model are provided in Subsection (ref).
Bayesian model specification is completed by assigning prior distributions to model parameters. For the equation parameters $\Theta=\left(\theta',\gamma',{\theta^p}'\right)'$, which consist of preference parameters $\theta=(\theta_1',\ldots,\theta_J')'$, bundle-effects parameters $\gamma=(\gamma_{(1,2)}',\ldots,\gamma_{(J-1,J)}')'$ and first-stage parameters $\theta^p=({\theta^p_1}',\ldots,{\theta^p_{J_p}}')'$, I apply a standard multivariate normal distribution priors to equation parameters, given by:
where $m_\Theta$ and $V_\Theta$ are hyperparameters.
For the non-zero factor loadings in the factor loading matrix $\Lambda_t$, I assume {i.i.d.} normal distribution conjugate priors
with a common hyperparameter $\sigma^2_\lambda$.
Note that the augmented latent utilities $u_{irt}$ do not require a prior, and the distribution of (augmented) latent factors $f_i\sim\mathcal{N}(\mathbf{0},\mathbf{I}_L)$ is specified as part of the model.
The model specification also involves choosing the number of factors $L$. Table (ref) provides guidelines for choosing the number of latent factors $L$ in Factor-Augmented (FA) and Time-Varying Factor-Augmented (TV-FA) models. As a general guideline, I match the number of free parameters modeling the error correlations $\Omega_{it}$ in the time-invariant factor-augmented model with that of the random-effects to ensure flexibility and that no restrictions are applied by the factor structure. When allowing the factor loading matrix $\Lambda_t$ to be time-varying, I add two more factors to the specification to capture the allowed time variations.
The vectorized model facilitates an efficient Markov-Chain Monte Carlo (MCMC) sampling scheme. The scheme for sampling from the posterior distribution involves the following steps:
Step 1 is a standard Gibbs sampling step for the multinomial probit model (MNP), see mcculloch1994exact. Since the error correlations are fully captured by the latent factors $f$ and factor loadings $\lambda$, the latent utilities $u_{irt}$ are conditionally independent. This simplifies the standard procedures.
The model vectorization, introduced in Subsection (ref), enables the joint sampling of all equation parameters $\Theta=\left(\theta',\gamma',{\theta^p}'\right)'$ in Step 2. Similarly, it facilitates sampling all latent factors $f$ and factor loadings $\lambda$ in one block in Step 3 and 6, respectively.
I develop the sampling Steps 3 to 6 for the factor model based on jacobi2016bayesian and fruehwirthschnatter2023sparse. Full details on this MCMC scheme are provided in Appendix (ref).
The key quantities of interest in bundle demand models are the price responses and complementarity, both of which are measured by the own- and cross-price elasticity of demand,
and the counterfactual choice probability implied by policy changes,
This subsection develops a Bayesian density predictions to estimates these quantities.
The evaluation of both price elasticity and counterfactual choice probability hinges on the prediction of bundle choice $y_{it}$ given adjusted covariates $z_{ijt}^{cf}$ and $w_{i(j_1,j_2)t}^{cf}$. The full likelihood function (ref) specifies how the bundle choice $y_{it}$ is generated:
where $z_{it}=\{z_{ijt}^{cf}\}_{j\in\mathcal{J}}$ contains price and other observed covariates $z_{ijt}^{cf}=(p_{ijt}^{cf},{x_{ijt}^{cf}}')'$, and $w_{it}^{cf}=\{w_{i(j_1,j_2)t}^{cf}\}_{j_1,j_2\in\mathcal{J}}$. The densities $p(y_{it}|u_{it})$ and $p(u_{it}|\cdot)$ are defined by $y_{it}=\arg\max_{r\in\mathcal{R}}(u_{irt})$ and Equation (ref), respectively.
Combine with the posterior density $p(\Theta,\Lambda,f|y)$, the predictive density for $y_{it}$ is given by
The choice probabilities of consuming good $j$ and bundle $r$ are formally defined as
where the integrals are over the empirical distribution of the individuals $i\in\mathcal{N}$ and time periods $t\in\mathcal{T}$. It specifically calculated by
as functions of the predicted bundle choice $y_{it}$. Their predictive densities are readily computed.
The evaluation of bundle and good price elasticity uses a two-sided numerical differentiation approach based on the above Bayesian predictive approach. Specifically, I simulate the bundle and good choice probability at three points, (1) decrease the observed price $p_{ijt}$ by 5%, $p_{ijt}^\text{backward}=p_{ijt}\times(1-0.05)$, (2) original price $p_{ijt}$, and (3) increase by 5%, $p_{ijt}^\text{forward}=p_{ijt}\times(1+0.05)$. Then I calculate the elasticities by
where bundle elasticity $\mathcal{E}_{jr}$ is the percentage change in the probability use of bundle $r$ with respect to 1% increase in the price of good $j$. Similarly, substance elasticity $\mathcal{E}_{jk}$ is the percentage change in the probability use of good $k$ with respect to 1% increase in the price of good $j$. For both equations, the right hand side is the multiplication of the differentiation respect to a 10% change (backward plus forward) in price and normalization of the 10% change.
I explore the performance of the proposed model, the time-varying factor-augmented bundle demand model with endogenous regressors, comparing it to the benchmark time-invariant random-effects and factor-augmented models with or without endogenous regressors. Furthermore, I explore the asymptotic behavior of the proposed model with various sample sizes.
For each data set, there are $J=3$ goods and $J_p=3$ good prices $p_{ijt}$ are endogenous. The covariates (and parameters) are simulated (and set to) mimicking the data application.
For the utility equations, the price coefficient is set to $\alpha=-1$. It is common to all three good utilities. The remaining preference parameters are $\beta_1=(1,0.2,0.1)'$, $\beta_2=(2,0.2,0.1)'$ and $\beta_3=(2,0.1,0.05)'$ for three goods, respectively. Similarly, there is a common parameter for all bundle effects, $\tilde{\gamma}=0.05$. The remaining bundle-effects parameters are bundle-specific intercepts, $\tilde{\gamma}_{(1,2)}=2$, $\tilde{\gamma}_{(1,3)}=0$ and $\tilde{\gamma}_{(2,3)}=-1$. Intuitively, good 1 and 2 are complements, 1 and 3 are independent, and 2 and 3 are substitutes, according to the extra utility or disutility from the bundle effects. For the first-stage equations, the parameters are set to $\theta^p_1=(0,1,0.5,0,0.01)'$, $\theta^p_2=(0,1,0.5,0,0)'$ and $\theta^p_3=(0,1,0.5,0,-0.01)'$ for $j=1,2,3$, respectively.
The error correlations are induced by two latent factors $f_i\sim\mathcal{N}(\mathbf{0},\mathbf{I}_2)$ and time-varying factor loadings $\Lambda_t=(\lambda_{1t},\lambda_{2t})$ with $\lambda_{1t}=(1,0,-1,1,0,-1)'$ and $\lambda_{2t}\sim\mathcal{N}(\mathbf{0},\mathbf{I}_6)$. The six rows of $\Lambda_t$ correspond to the three good utility and three first-stage equations, respectively. Note that the $\Lambda_t$ cannot be recovered due to rotational invariance. It is a means of introducing error correlations, with the same spirit of the model where the loading is a means of capturing error correlations.
The observed covariates in the utility equations are $z_{ijt}=(p_{ijt},1,x_{1i},x_{2i})'$, for $j\in\mathcal{J}$, where $p_{ijt}$ is the (endogenous) price, $1$ corresponds to the intercept, $x_{1i}\sim\mathcal{N}(10,0.7225)$, and $x_{2i}\sim\mathcal{TN}_{(0,\infty)}(10,36)$. The bundle-effects covariates are $w_{i(j_1,j_2)t}=(\tilde{w}_i,1)'$ for $j_1,j_2\in\mathcal{J}$, interacting with $\gamma=(\tilde{\gamma},\tilde{\gamma}_{(j_1,j_2)})'$, where $\tilde{w}_i=1,2,\ldots,6$ mimicking the family size. The instruments and exogenous regressors are $z^p_{ijt}=(p_{jt},1,z_{ij},x_{1i},x_{2i})'$ where $p_{jt}$ mimics the market-level price $p_{1t}\sim\mathcal{N}(7,0.04)$, $p_{2t}\sim\mathcal{N}(6,0.01)$ and $p_{3t}\sim\mathcal{N}(5,0.01)$. The instruments are $z_{ijt}\sim\mathcal{N}(0,1)$, for individual $i$, time $t$, and good $j\in\mathcal{J}_p$.
This subsection compare the proposed time-varying factor-augmented bundle demand model with endogenous regressors with the benchmark time-invariant random-effects and factor-augmented models with or without endogenous regressors. The true data-generating process (DGP) is outlined in Subsection (ref). The $J_p=3$ prices are endogenous and the unobserved tastes and the unobserved utility-price confounders are time-varying. So the time-varying factor-augmented (TV-FA) model corresponds to the true DGP. I simulate 50 Monte Carlo trials. For each trial, the sample contains $N=1000$ individuals and $T=12$ time periods, which is close to the dimension of the real data application. For each sample and model, the MCMC simulates 10,000 draws for use after a 10,000-draw burn-in period.
The key quantities of interest in bundle demand models are price responses and complementarity, both measured by own- and cross-price elasticity of demand. Therefore, the comparison focuses on the estimated price elasticity, specifically the distance between the true elasticities evaluated with the true model parameters and the estimated elasticities evaluated with the estimated parameters.
Table (ref) reports the true price elasticity and the differences between the true and estimated elasticities from the five models. The first column shows the true values of elasticities from the true DGP, bundle demand model with endogenous regressors (Endo) and time-varying (TV) unobserved tastes and the unobserved utility-price confounders. The remaining five columns correspond to the time-invariant (TI) random-effects (RE) exogenous (Exo) model, TI factor-augmented (FA) exogenous model, TI-FA model with endogenous regressors (Endo), time-varying (TV) FA exogenous (Exo) model, and lastly the correctly specified time-varying factor-augmented (TV-FA) model with endogenous regressors (Endo). The root mean squared errors (RMSE) are calculated from the squared difference between posterior mean and true value of price elasticity for each Monte Carlo trail.
This simulation study suggests that the correctly specified and most flexible TV-FA Endo model outperforms the four benchmark models. The comparison between the exogenous TV-FA Exo and endogenous TV-FA Endo models highlights the importance of controling for price endogeneity when estimating complementarity. The comparison between the time-invariant FA Endo (TI) and the time-varying FA Endo models highlights that accounting for time-varying unobserved tastes and the unobserved utility-price confounders is also important in the bundle demand models.
This subsection investigate the asymptotic behavior of the proposed model. The true data-generating process (DGP) is outlined in Subsection (ref). The $J_p=3$ prices are endogenous and the unobserved tastes and the unobserved utility-price confounders are time-varying. I consider five sample sizes with $N=100,1000,1000$ individuals and $T=6,12,24$ time periods. For the sample size, 50 Monte Carlo trials are simulated and estimated by the correctly specified time-varying factor-augmented (TV-FA) endogenous model.
Table (ref) reports the asymptotic results. The first column shows the true values of elasticities from the true DGP, bundle demand model with endogenous regressors (Endo) and time-varying (TV) unobserved tastes and the unobserved utility-price confounders. The remaining five columns correspond to five sample sizes. The results show clear pictures of convergence. As number of individuals $N$ or number of time periods $T$ increases, the distances between the true price elasticities and the estimated elasticities shrink.
This section applies the proposed model to examine the implications of implementing a sugary drink tax (also known as a soda tax) within the context of complementarities.
I combine household-level and store-level information on carbonated beverages, milk, and salty snacks from the IRI Marketing Data Set bronnenberg2008database. The household-level data, collected from two BehaviorScan markets in Pittsfield, Massachusetts, and Eau Claire, Wisconsin, include records of shopping trips, purchases, and demographic information. Store-level data encompass transactions for each product, identified by Universal Product Code (UPC), across each store and week. Additionally, this data set provides product attributes across five levels (category, small category, parent company, vendor, brand) and the equivalent volume of the package.
Table (ref) defines the $J=5$ goods. The first good, identified as a collection of sugary soft drinks, is targeted by the soda tax. This category encompasses beverages sweetened with cane, corn, or other sugars. The second good includes diet soft drinks, which are low-calorie beverages that are either sugar-free or sweetened with alternatives such as Aspartame, Sucralose, or Stevia. Regions such as France, the Philippines, the United Arab Emirates, and Philadelphia (Pennsylvania, United States) impose the tax on artificially sweetened beverages at the same rate as sugar-sweetened beverages. Conversely, these beverages are excluded from the tax in other areas, such as Canada and Berkeley (California, United States).
The third good, carbonated water, comprises seltzer, tonic water, and club soda, which are typically exempt from the tax. The fourth good includes milk drinks, such as flavored milk, eggnog, buttermilk, milkshakes, and non-dairy drinks, also generally not subject to the tax. These are included in the analysis as alternative sugar sources. The fifth good, consisting of salty snacks (e.g., cheese snacks, corn snacks, potato chips, ready-to-eat popcorn, tortilla chips), is considered a potential complement to soft drinks. This section explores the substitution and complementarity among these five goods.
The sample consists of 4,143 households that purchased any of the five goods at least once over the 52 weeks of 2011. Observed choices are defined as combinations of household and week. A household is considered to have made a choice decision if it visited a store and purchased any products, including but not limited to the five goods, within a given week. This results in a total of 163,434 choice decisions. The fourth column of Table (ref) presents the probability of purchasing each of the five goods in any week, while the fifth column shows the proportion of the 4,143 households that purchased each good at least once throughout the year.
Table (ref) defines all 32 possible bundles of goods, ranging from the empty bundle (no purchase) to combinations of one to five goods. The last column reports the weekly choice probability for each bundle based on 163,434 choices. Among these choices, 43.78% involve no purchase, 32.96% chose a single good, 16.82% chose two goods, 5.61% chose three goods, 0.81% chose four goods, 0.02% chose all goods.
Table (ref) includes the average prices of the five goods in the last column. I construct $\textit{price}_{ijt}$ for good $j$ faced by household $i$ in week $t$, using $i$'s complete purchase history, $i$'s shopping trips in week $t$, and the average product prices at the stores visited by $i$ during week $t$. The steps for constructing $\textit{price}{ijt}$ are as follows:
First, an aggregate product basket for each good is defined, based on 5-level product categories and the package's equivalent volume. The aggregate baskets for the five goods consist of 292, 228, 32, 50, and 642 products, respectively. These numbers are fewer than the total number of unique products identified by UPC codes, as a product can have multiple UPCs across different stores and cities. Table (ref) outlines the definitions of the aggregate product baskets for the five goods. Each household $i$'s specific product basket for good $j$ is a subset of this aggregate basket. The purchase record of household $i$ in 2011 then determines the weight of each product in their basket, proportional to the total volume of good $j$ that the household purchased, relative to all volumes purchased of that good."
Second, for each shopping trip of household $i$ in week $t$, the effective price of good $j$ is equal to the weighted average prices of the products in $i$'s basket, with weights derived from $i$'s basket. If household $i$ makes multiple shopping trips in week $t$, then $\textit{price}_{ijt}$ equals the lowest of these weighted average prices across all trips.
Prices $\text{price}_{ijt}$ might be endogenous due to three factors. First, the product baskets chosen by household $i$ are likely correlated with $i$'s preferences through unobserved characteristics. Second, the shopping trips of household $i$ may also be linked to their preferences. Lastly, demand shocks can influence households' preferences for products as well as the prices of these products. To address endogeneity, I consider two sets of instrumental variables.
The first set includes Hausman-type instruments, following the approach of nevo2000practitioner. For Pittsfield, the instruments comprise the unit prices of the same aggregate product basket in three surrounding cities in the market-level data: Boston, Hartford, and New York. Similarly, the instruments for Eau Claire are derived from surrounding cities, including Chicago, Green Bay, Indianapolis, Milwaukee, Minneapolis/St. Paul, and Omaha. These instrumental variables are correlated with the product prices but are unlikely to be influenced by demand shocks occurring in Eau Claire or Pittsfield, given the geographical and market separation.
Furthermore, another instrument is the number of different stores that household $i$ visited in week $t$. Visiting more stores is assumed to expose the household to lower effective prices. Meanwhile, the number of different stores visited, distinct from the number of shopping trips, is presumed to be uncorrelated with the household's preferences.
I specify the good utility as
where $\textit{price}_{ijt}$ represents the price of good $j$ faced by individual $i$ in week $t$. As the goods differ in nature and thus exhibit varying price responses, I allow the price coefficient $\alpha_j$ to vary across goods. $\textit{\# of trips}_{it}$ denotes the number of shopping trips by household $i$ in week $t$. A visit is recorded if the household makes any purchase, regardless of whether it includes any of the goods of interest. $\textit{income}_i$ is the self-reported annual household income, expressed in log thousands of U.S. dollars. $\textit{family size}_i$ represents the family size, which ranges from 1 to 8. $\textit{child 0-5}_i$, $\textit{child 6-11}_i$, and $\textit{child 12-17}_i$ are dummy variables indicating the presence of children within those age ranges in the household. It is noted that more than one of these dummy variables can be equal to one, indicating the presence of multiple children in different age categories. $\textit{Pittsfield}_i$ is a dummy variable identifying households from Pittsfield; otherwise, the household resides in Eau Claire. Additional household characteristics, such as the age and gender of the household head, level of education, and employment status, were considered but did not yield statistically significant coefficients or alter the outcomes.
The bundle effects are specified as follows:
where $\gamma_{(j_1,j_2)0}$ captures the baseline additional utility or disutility from the joint consumption of goods $j_1$ and $j_2$. The other coefficients $\gamma_{(j_1,j_2)k}$ account for the heterogeneities in preferences.
Table (ref) presents the posterior means and standard deviations (in parentheses) of the utility parameters for various goods, as derived from the time-varying factor-augmented (TV-FA) bundle demand model with endogenous regressors. The results are in line with expectations, with price coefficients being negative and precisely estimated, evidenced by small posterior standard deviations. Additionally, an observation is the impact of household income on consumption preferences. Wealthier households demonstrate a shift away from sugary soft drinks towards diet soft drinks and carbonated water, as shown by the negative income effects for sugary soft drinks and positive effects for diet soft drinks and carbonated water. Furthermore, households with children aged 0-5 do not exhibit significant differences in consumption preferences among the five goods. However, the presence of older children (ages 6-11 and 12-17) markedly increases the utility derived from milk drinks and salty snacks.
Table (ref) presents the posterior means and standard deviations (in parentheses) of the bundle-effects parameters, from the time-varying factor-augmented (TV-FA) bundle demand model with endogenous regressors. Each column in the top, middle, and bottom panels corresponds to a pair of goods, $j_1$ and $j_2$, illustrating the interaction between different product combinations. With multiple goods involved, the bundle-effects parameters directly are not readily interpretable. Therefore, the following subsection presents and discusses price elasticity, which determines the complementarity and substitutability among the goods.
Table (ref) reports the estimated own- and cross-price elasticities of demand from the full model, specifically the time-varying factor-augmented (TV-FA) bundle demand model with endogenous regressors. Own-price elasticities, displayed along the diagonal, show the percentage change in the choice probability of good $k$ due to a one percent change in its price, holding other prices constant. Aligned with expectation, all estimated own-price elasticities are negative. The statistically significant own-price elasticity of sugary soft drinks confirms that soda taxes decreases the intake of sugar from sugar-sweetened beverages, and consequently leads to the intended health benefits.
The off-diagonal entries report cross-price elasticities, showing the percentage change in the choice probability of good $k$ (column) when the price of good $j$ (row) increases by one percent. Notably, the relationship between sugary soft drinks and milk drinks exhibits independence, as indicated by the insignificant cross-price elasticity $(j,k)=(1,4)$. This suggests that a soda tax does not drive consumers to milk drinks as an alternative source of sugar. Furthermore, sugary soft drinks and salty snacks are identified as weak complements, as shown by the significant but modest cross-price elasticity $(j,k)=(1,5)$. This suggests that implementing a soda tax could have ancillary health benefits by slightly reducing the consumption of salty snacks alongside sugary drinks, thereby contributing to a broader health benefit beyond the primary target of reducing sugar intake.
This paper introduces a Bayesian factor-augmented bundle demand model for panel data. Its primary objective is to unravel the complexities of demand estimation, specifically focusing on the substitutability and complementarity among a variety of goods in the context of endogenous regressors. To tackle the challenges of taste correlation and endogeneity, the model employs a factor structure. Notably, this structure is enhanced with time-varying factor loadings, allowing it to adeptly capture the dynamic nature of unobserved tastes and utility-price confounders over time.
The empirical results from these applications indicate significant price endogeneity and fluctuating unobserved tastes and utility-price confounders over time. An insight from this study is the tendency of beer products to demonstrate independent demand, as suggested by the statistically insignificant cross-price elasticities. This finding provides an understanding of consumer behavior in the beverage market. The research also highlights a consumer penchant for variety, evidenced by the escalating bundle effects correlating with family size.
Extending upon Gentzkow’s model with complementarity, this research incorporates endogenous regressors into the framework. This is particularly pivotal given the often overlooked, yet crucial role of endogeneity in demand estimation. The factor approach introduced in this model is both flexible and parsimonious, effectively addressing the complexities brought about by endogeneity across both choices and time periods.
The paper employs a multinomial probit (MNP) model with standard normal idiosyncratic errors to facilitate Bayesian posterior inference. The posterior inference is developed through a Markov-Chain Monte Carlo sampling scheme. The factor structure's comprehensive capture of error correlations streamlines the estimation process. Furthermore, vectorization of the model boosts the efficiency of sampling algorithms.
Through extensive Monte Carlo simulation studies, the performance of the proposed model is scrutinized in comparison to benchmark models, including both time-invariant random-effects and factor-augmented models, with and without endogenous regressors. These studies confirm the model's ability to accurately recover true parameters and reliably estimate price elasticity. Importantly, they bring to light the crucial need to account for price endogeneity in complementarity estimation and the significance of incorporating time-varying unobserved tastes and utility-price confounders.
The application to the sugary drink tax investigates concerns that such a tax might lead to an increased consumption of alternative sugary products and examines the potential health benefits from reducing the consumption of unhealthy snacks. The estimation of substitution patterns between sugary soft drinks, diet soft drinks, carbonated water, milk drinks, and salty snacks confirms that soda taxes decrease the intake of sugar from sugar-sweetened beverages, consequently leading to the intended health benefits. The results also suggest an independence between sugary soft drinks and milk drinks, implying that a soda tax does not cause consumers to substitute sugary soft drinks with milk drinks as an alternative source of sugar. Additionally, sugary soft drinks and salty snacks are identified as weak complements. This suggests that a soda tax could yield secondary health benefits by marginally decreasing the consumption of salty snacks alongside sugary drinks, thus extending the health benefits of the tax beyond merely reducing sugar intake alone.