Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
40,005 characters · 15 sections · 33 citation commands
The Mixed Aggregate Preference Logit Model: A Machine Learning Approach to Modeling Unobserved Heterogeneity in Discrete Choice Analysis
Keywords: mixed logit, machine learning, heterogeneity, discrete choice modeling, choice modeling, neural networks, random utility theory
Discrete choice models are leveraged in a wide range of research communities, including Transportation, Economics, Healthcare, Business, and others Train2009,Haghani2021. The mixed logit model---alternatively known as the mixed multinomial, random parameters logit, or random coefficients logit model---is particularly important due to its capability, as proven by McFadden2000, to approximate choice probabilities from “any discrete choice model derived from random utility maximization.” However, a limitation of this theoretical work is that “it provides no practical indication of how to choose parsimonious mixing families” McFadden2000. This means modelers must make simplifying assumptions about how they believe heterogeneity should be distributed across a population for each feature in the model, including assuming independent mixing distributions, convenient mixing distributions (e.g., normal or log-normal), or preference homogeneity for a subset of product features. Furthermore, available software packages for estimating mixed logit models only support a limited number of heterogeneity distributions to choose from, e.g. Helveston2023, Arteaga2022. Consequently, mixed logit's theoretical potential as a flexible approach for modeling unobserved preference heterogeneity often remains unrealized in practice, and many papers leveraging random utility-based mixed logit models make limiting assumptions Forsythe2023,Helveston2015, Philip2021, Guo2021,Kavalec1999, Revelt1998.
This paper introduces the Mixed Aggregate Preference Logit (MAPL, pronounced “maple”) model, a novel class of discrete choice models that aims to relax many of the assumptions required in traditional mixed logit. MAPL models enable the ability to flexibly model various decision-making paradigms (e.g., utility maximization or regret minimization), heterogeneous consumer preferences, and governing functional forms via a new conception of unobserved heterogeneity. Whereas traditional discrete choice modeling frameworks like mixed logit parameterize preference heterogeneity through assumptions about feature-specific heterogeneity distributions, MAPL models directly relate model inputs to parameters of alternative-specific distributions of aggregate preference heterogeneity. In doing so, MAPL models also eliminate the need to specify a functional form for the latent decision model (e.g. a linear-in-parameter utility model for mixed logit models), including the need to make feature-specific assumptions about how heterogeneity is distributed.
By increasing the flexibility in how unobserved heterogeneity can be modeled and eliminating the need for restricting functional form assumptions, MAPL models significantly narrow the gap between the theoretical possibilities and practical applications of incorporating unobserved heterogeneity into discrete choice models. MAPL models also facilitate the incorporation of machine learning methods into discrete choice modeling, building on recent literature connecting these fields VanCranenburgh2022. By incorporating unobserved heterogeneity into machine learning modeling, MAPL models can achieve superior predictive performance compared to state-of-the-art neural network models and achieve at least as good performance as correctly-specified mixed logit models.
The structure of this paper is as follows: (ref) reviews traditional discrete choice modeling paradigms and their connections to machine learning. (ref) presents a formal description of the MAPL model. (ref) presents a simulation experiment where the predictive performance of a MAPL model is compared against that of different mixed logit models and neural network models. Lastly, (ref) summarizes the paper's findings and directions for future work.
Discrete choice modeling has a rich history lasting over one hundred years McFadden2000Nobel and is leveraged in a wide variety of fields Train2009,Haghani2021. Common discrete choice models include the multinomial logit (MNL), nested logit (NL), and mixed logit (MXL) models Train2009. The mixed logit model is particularly notable due to its flexibility and theoretical capabilities in terms of modeling unobserved heterogeneity McFadden2000. In practice, however, modelers often make several simplifying assumptions when specifying mixed logit models, including assuming independent mixing distributions, convenient mixing distributions (e.g., normal or log-normal), or preference homogeneity for a subset of product features Forsythe2023,Helveston2015,Philip2021,Revelt1998, Guo2021,Kavalec1999. Depending on the software used, these assumptions are not necessarily made by choice as not all software capable of estimating mixed logit models support alternative assumptions.
Recent work has sought to improve some of these restrictions by proposing more flexible mixing distributions Fosgerau2013,Train2016, Krueger2020, but more flexible mixing distributions can lead to longer estimation times Bansal2018. Researchers have developed more efficient estimation packages Arteaga2022,Helveston2023,Molloy2021, but the simulation requirements for equally-precise mixed logit estimation can still scale considerably with the number of features with random taste heterogeneity, larger data sizes Czajkowski2019, and in contexts where a model must be estimated multiple times (e.g., multi-start algorithms to search for global optima).
In more recent years, researchers have begun integrating discrete choice modeling with machine learning techniques VanCranenburgh2022. For instance, Nobel laureate Daniel McFadden notes in his Nobel lecture that the latent class logit can be represented as an artificial neural network McFadden2000Nobel. Machine learning models in the context of discrete choice models generally offer superior predictive capability Salas2022 and require less time to estimate Wang2021,Garcia-Garcia2022, but are often not immediately interpretable VanCranenburgh2022. Several recent works combine traditional discrete choice modeling techniques with machine learning to yield flexible and interpretable models. Studies such as Wong2021,Arkoudi2023,Han2022,Sifringer2020 integrate the multinomial logit with neural networks in a variety of ways to relax functional form assumptions while still providing interpretable parameters. Sifringer2020 additionally focuses on combining nested logit models with neural networks. Even latent class logit models have been combined with neural networks to improve model flexibility Lahoz2023. Currently, the integration of machine learning techniques with mixed logit modeling techniques is missing from the literature.
The MAPL model makes two notable contributions to existing literature. First, it provides a novel framework for modeling unobserved heterogeneity that allows for model inputs to directly predict parameters of alternative-specific distributions of unobserved preference heterogeneity. This structure also removes the need to make assumptions about the functional form for the latent decision model, further relaxing the assumptions required by the modeler. Second, the model allows for a novel extension of machine learning models that explicitly models continuous, unobserved consumer heterogeneity akin to the mixed logit. In making this integration, MAPL models have the capacity to continue improving by integrating future innovations that emerge from the field of machine learning. These contributions have the opportunity to greatly impact discrete choice modeling through alternative estimation schemes, expanded flexibility in modeling heterogeneous preference, and improved prediction performance.
The MAPL model is distinguished by its conception of unobserved heterogeneity. Traditional mixed logit models require the modeler specify feature-specific heterogeneity distributions, which may or may not be accurate assumptions. In contrast, MAPL models directly relate model inputs to distributional parameters of aggregate observables-related preference (see Figure (ref)). This frees the modeler from having to make feature-specific assumptions about heterogeneity and provides greater flexibility in fitting models with heterogeneity in preferences.
Specifying a MAPL model centers around three modeling decisions: the choice preference framework, the aggregate preference distribution specification, and the preference distribution parameter estimator.
The choice preference framework dictates the modeled decision-making paradigm, which is chosen based on the modeler's belief regarding individuals' decision-making processes. MAPL models assume there is some continuously distributed latent quantity associated with each choice alternative that can be transformed to estimate choice probabilities (e.g., via the logit function). As this is a specific functional form assumption not required by MAPL, we prefer to use the generic term “preference" over the commonly-used “utility.” While a utility maximization paramdigm is commonly assumed in many domains, others prefer a regret minimization framework Chorus2012. Moreover, future modelers might consider alternative conceptions of latent preferences that a MAPL model could incorporate. Still, it is incumbent upon the modeler to leverage institutional knowledge or empirical evidence to decide upon the relationship between preference quantities and choice probability estimates.
The aggregate preference distribution specification is the explicitly assumed distribution for aggregate preference heterogeneity. For instance, assuming a normal distribution is assuming that the aggregate preference for a given product is distributed unimodally and symmetrically across the population of interest. Alternatively, if one chose a mixture of normals, one would be assuming that there is a finite set of subpopulations that have distinct, normally distributed preferences. Aggregate preference distributions are generally assumed to be univariate,\footnote{MAPL models do not preclude assuming multi-dimensional distributions; for instance, one could theoretically model a multi-dimensional distribution of utility and regret if applicable.} but can be any parameterizable distribution with a finite integral, such as a Normal, Lognormal, or Triangular. Moreover, the model can incorporate more flexible non- or semi-parametric distributional forms like those described in Fosgerau2013,Train2016,Krueger2020. We demonstrate the use of the semi-parametric distribution from Fosgerau2013 in our simulation experiment in (ref). Modelers should make distributional assumptions based on institutional knowledge or empirical evidence.
The choice of a parameter estimator relates model inputs to the parameters governing consumer preference. This decision reflects the modeler's belief regarding the complexity associated with transforming model inputs to distributional parameters. This could be as simple as a linear model or as complex as a neural network, which is used in the simulation experiment in (ref). Any estimator that can predict scalar values can be incorporated into a MAPL model. Again, this decision should be made based on institutional knowledge or empirical evidence as well as the goals of the modeling exercise (e.g. predictive accuracy).
Assuming $n$ total choice alternatives, individual $i$ has chosen an alternative $j$ within choice set $t$ according to some decision-making paradigm. Choice decisions are represented by $y_{ijt}\in \left[ 0,1 \right]$ (y when observations are stacked), and $k$ explanatory variables (e.g., product features, consumer demographics, etc.) are contained within the vector $\text{\textbf{x}}_{ijt}\in \mathbb{R}^k$ ($\text{\textbf{X}}$ when observations are stacked). These model inputs as well as a vector of $m$ estimable parameters ($\bm{\beta}\in \mathbb{R}^m$) serve as arguments to some function $\text{\textbf{G}}\left( \bm{\beta}, \text{\textbf{X}} \right)\in \mathbb{R}^{s\cdot n}$ that maps $k$ model inputs to $s$ parameters for each observation. The full set of $s\cdot n$ parameters define a vector of preference distributions $\bm{\nu} \sim \bm{\Psi}\left( \text{\textbf{G}}\left( \bm{\beta}, \text{\textbf{X}} \right) \right)right)}$. Next, some linking function (e.g., the logit function) $\text{\textbf{h}}\left( \bm{\nu} \right)$ estimates choice probabilities ($\hat{\text{\textbf{p}}}$) based on a given vector of preference values. Some measure of the estimated choice probabilities that accounts for heterogeneity ($\hat{\text{\textbf{p}}}^*$) is used to compare model predictions to outcomes. A common choice for $\hat{\text{\textbf{p}}}^*$ will likely be the expected value of choice probabilities, which can be found by integrating said linking function across preference distributions. Similar to the integration over parameters in the mixed logit framework Train2009, (ref) shows the relevant integral required to identify the expected choice probabilities for a set of alternatives where $f$ is the appropriate probability density function for the $\bm{\Psi}\left( \text{\textbf{G}}\left( \bm{\beta}, \text{\textbf{X}} \right) \right)right)}$ vector of distributions. Parameters $\bm{\beta}$ are estimated by solving (ref) where $\bm{\beta}$ minimizes some measure of fit $d$ (e.g., log-likelihood, KL-divergence).
To examine the performance and flexibility of MAPL, we conduct a simulation experiment where we simulate sets of choice data according to multiple different data generating processes (DGPs) that follow different variations of mixed logit utility models with different utility structures and different heterogeneity distributions. Our experiment is motivated by the fact that modelers do not know a priori the underlying utility functional form nor the distributional form of any unobserved heterogeneity in the true DGP. When using mixed logit, inaccurate assumptions about the functional form or heterogeneity distributions will lead to poorer performance. In contrast, the flexibility of the MAPL model should produce good performance for each DGP without making any assumptions about the utility functional form or feature-level assumptions about preference heterogeneity.
To test this, we employ four different mixed logit DGPs to simulate different sets of choice data. Each DGP has a fixed parameter, $\beta_0$, and two random parameters, $\beta_1$ and $\beta_2$. In DGP 1, $\beta_1$ and $\beta_2$ follow independent normal distributions. DGP 2 is the same as 1 except that the two distributions are correlated. For DGPs 3 and 4, we use independent normal distributions for $\beta_1$ and $\beta_2$, except that DGP 3 includes an additonal interaction parameter between $x_0$ and $x_1$, and DGP 4 includes a nonlinear parameter for $x_1^2$. (ref) below summarizes each of the DGPs. All models follow a random utility theory decision-making paradigm where the utility is formed by a summation of a portion relating to $x$ features, $\nu_{ijt}$, and an error term, $\epsilon_{ijt}$, representing a Gumbel-distributed error term such that $u_{ijt} = \nu_{ijt} + \epsilon_{ijt}$.
Product features (the various $x$ values) in each respective DGP are uniformally distributed on a domain of $[-1, 1]$. With each DGP, we simulate 10,000 individuals making 10 successive choices each from among three choice alternatives. Choice probabilities for each alternative are derived according to traditional random utility theory Train2009 using simulation to obtain the mean $P_{jt}$ across 1,000 draws of the logit function, $P_{jt} = \exp({\nu_{jt}}) / \sum_{k=1}^{J} \exp({\nu_{kt}})$. The coefficients used in each DGP are as follows: $\beta_{0} = -1$, $\mu_{1} = 1$, $\mu_{2} = 2$, $\sigma_{1} = 1$, $\sigma_{2} = 1.5$, $\sigma_{12} = 0.7$, $\beta_{3} = 2$
For each DGP, we estimate a variety of choice-models specifications that are common among pracitioners: a simple multinomial logit (MNL), a Mixed Logit (MXL) with independent mixing distributions, a Simple Neural Network (NN), a Deep NN, and two forms of MAPL models: one using normal distributions for the aggregate preference heterogeneity distributions and one using the semi-parametric distribution posed by Fosgerau2013 with 12 distributional parameters. The MAPL models were trained using neural networks with two hidden layers composed of 64 neurons each as well as a normalization and dropout layer after each hidden layer. Importantly, the neural network transforms model inputs for each choice alternative independently and have identical weights. We train each model repeatedly across twenty instantiations of the DGPs using negative log-likelihood as the loss function for training with an 80/20 train/test split and 2,000 epochs per training session. Full model specifications can be found in this paper's replication code available at \url{https://github.com/crforsythe/MAPL-Paper}.
(ref) presents the differences between the estimated log-likelihood values of various model specifications and the true log-likelihood value across all DGPs, where the true log-likelihood is taken as the estimated log-likelihood of a mixed logit model that exactly matches the true underlying DGP. Results are presented as box plots since we conducted the experiment 20 times for each combination of DGP and model. Several observations are of note.
First, we see that models that misspecify unobserved heterogeneity have relatively poor performance. The MNL models (which cannot account for unobserved heterogeneity) have errors of approximately 9% even with correct utility functional forms and errors above 20% with incorrect utility functional forms. Likewise, heterogeneity misspecification is all but guaranteed for the traditional machine learning models (Simple NN and Deep NN) as these models are unable to account for unobserved heterogeneity. However, it is notable that these models perform equally well across different utility functional forms, such as the DGPs with interactions and nonlinearities.
Second, the mixed logit model with independent normals performs perfectly (0% error) in the case where the true DGP is exactly the same, as would be expected. This model also performs very well even when heterogeneity is correlated, suggesting that the common assumption of uncorrelated heterogeneity may be reasonable. However, misspecification in utility form leads to large discrepancies in estimate and true log-likelihood values: 12% when omitting the non-linear term and 14% when omitting the interaction term.
Finally, the results confirm that the MAPL models were robust to all DGP specifications, even though they require no a priori specification of utilty. A MAPL model with a simple assumption of aggregate preference heterogeneity being normally-distributed never exeeded more than 4% error, and this gap shrunk to $<$1% for all iterations when using the more flexible Fosgerau-Mabit distribution, suggesting that flexible heterogeneity distributions are able to capture the underlying true heterogeneity. This illustrates the significance of incorporating unobserved heterogeneity into machine learning-based modeling approaches.
While the results in (ref) are promising, it is important to note that this experiment was conducted using a relatively large sample size (10,000 simulated sets of 10 choices, for a total of 100,000 choice observations). While such sizes may be available in some domains, in others sample sizes may be constrained, such as in stated-choice experiments where sizes are often limited by research budgets. To examine the potential impact of sample size on the performance MAPL models, we repeat one of the simulation scenarios using different sample sizes ranging from 500 to 10,000 sets of 10 observations (5,000 to 100,000 choice observations). We use DGP1 where $\beta_1$ and $\beta_2$ follow independent normal distributions for the data simulation and fit a MAPL model with the Fosgerau-Mabit aggregate preference heterogeneity distribution.
(ref) shows the results of this sensitivity analysis, revealing that the performance of MAPL models is constrained by sample size, which is expected for models using neural networks van2021artificial. Errors are higher with sample sizes under 1,000 simulated people and then quickly fall to below 2% and stabilize at under 1% at 4,000 people and more. However, even at only 500 simulated people, errors never cross above 6%, which is still significantly below the errors of MNL and NN models with 10,000 simulated people. We also observe that the variation in the 20 independent iterations of the simulation shrinks with increased sample size.
This paper introduces the Mixed Aggregate Preference Logit (MAPL) model, a novel class of discrete choice models that significantly advances how researchers can incorporate unobserved heterogeneity into choice analysis. By directly characterizing alternative-specific distributions of preference heterogeneity as functions of model inputs, MAPL models offer several advantages over traditional approaches. First, they eliminate the need to specify feature-specific heterogeneity distributions, freeing modelers from potentially restrictive assumptions. Second, they remove the requirement to define a specific functional form for the underlying decision-making process. Instead, MAPL models integrate machine learning techniques to flexibly capture complex relationships between inputs and preference distributions.
Our simulation experiment demonstrates that the flexibility of MAPL models yields superior predictive performance compared to traditional neural networks that cannot account for unobserved heterogeneity. Importantly, a single MAPL model specification performs well across multiple data-generating processes with different utility structures and heterogeneity distributions, without requiring the modeler to correctly identify the true underlying model. This robustness is particularly valuable in real-world applications where the true decision-making process and heterogeneity structure are unknown.
The benefits of MAPL models extend beyond prediction performance. For policy analysts, market researchers, and product designers, MAPL models offer more reliable predictions of choice behavior under varying scenarios, potentially leading to better-informed decisions. Furthermore, MAPL models can serve as diagnostic tools for traditional choice modelers, helping identify potential misspecification in utility functions or heterogeneity distributions. By comparing the performance of a MAPL model using flexible distributions (such as the semi-parametric distribution proposed by Fosgerau2013) with that of traditional mixed logit specifications, researchers can detect areas where traditional models may be missing important preference structures.
Despite these advantages, MAPL models do face certain limitations. As shown in our sample size sensitivity analysis, their performance depends on having sufficient data. However, MAPL models with smaller sample sizes still out-perform other models even when those models are estimated with substantially larger sample sizes. Additionally, the current formulation focuses on alternative-specific distributions, which may not capture all forms of preference heterogeneity, particularly those that manifest as correlations across alternatives.
Several promising directions for future research emerge from this work. First, developing robust techniques for interpreting MAPL models represents a crucial area for further work. While MAPL models offer excellent predictive capabilities, extracting meaningful behavioral insights remains challenging. We aim to adapt interpretability methods from machine learning literature Linardatos2021,van2021artificial to help researchers understand the patterns of heterogeneity captured by these models. Of particular interest will be recovering economically meaningful quantities such as willingness-to-pay distributions, elasticities, and marginal effects from trained MAPL models. Second, future research should explore theoretical connections between MAPL models and traditional choice models, particularly establishing conditions under which MAPL models can exactly recover traditional specifications. This work would help position MAPL models within the broader discrete choice literature and clarify when they offer meaningful advantages over existing approaches. Third, computational efficiency remains an important consideration. While MAPL models reduce the dimensionality of integration compared to traditional mixed logit, the training process for complex neural network components can be computationally intensive. Research into more efficient training algorithms and model architectures could further enhance the practical utility of MAPL models. Finally, extensions of the MAPL framework to incorporate dynamic choice behavior, reference dependence, or other behavioral economic phenomena represent promising directions for expanding the approach. The flexibility of the MAPL modeling framework makes it well-suited to capturing these more complex choice processes.
In conclusion, MAPL models offer a powerful new approach to modeling unobserved heterogeneity in discrete choice analysis by combining the strengths of traditional econometric models with the flexibility of machine learning methods. They address key limitations in existing approaches and open new possibilities for understanding and predicting choice behavior across a wide range of applications.
\noindentDGP Data Generating Process \newline
\noindentMAPL Mixed Aggregate Preference Logit \newline
\noindentMNL Multinomial Logit \newline
\noindentMXL Mixed Logit \newline
\noindentNL Nested Logit \newline
\noindentNN Neural Network
Connor R. Forsythe: Conceptualization, Formal analysis, Investigation, Methodology, Software, Supervision, Validation, Writing – original draft, Writing – review & editing. Cristian Arteaga: Data curation, Formal analysis, Investigation, Software, Supervision, Validation, Writing – review & editing. John Paul Helveston: Data curation, Formal analysis, Investigation, Software, Visualization, Writing – review & editing.
We would like to acknowledge Dr. Jeremy Michalek and Dr. Kate Whitefoot for their feedback that helped guide the direction of inquiry for this project.
{0pt}