Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
97,832 characters · 30 sections · 128 citation commands
We present a mixed multinomial logit (MNL) model, which leverages the truncated stick-breaking process representation of the Dirichlet process as a flexible nonparametric mixing distribution. The proposed model is a Dirichlet process mixture model and accommodates discrete representations of heterogeneity, like a latent class MNL model. Yet, unlike a latent class MNL model, the proposed discrete choice model does not require the analyst to fix the number of mixture components prior to estimation, as the complexity of the discrete mixing distribution is inferred from the evidence. For posterior inference in the proposed Dirichlet process mixture model of discrete choice, we derive an expectation maximisation algorithm. In a simulation study, we demonstrate that the proposed model framework can flexibly capture differently-shaped taste parameter distributions. Furthermore, we empirically validate the model framework in a case study on motorists' route choice preferences and find that the proposed Dirichlet process mixture model of discrete choice outperforms a latent class MNL model and mixed MNL models with common parametric mixing distributions in terms of both in-sample fit and out-of-sample predictive ability. Compared to extant modelling approaches, the proposed discrete choice model substantially abbreviates specification searches, as it relies on less restrictive parametric assumptions and does not require the analyst to specify the complexity of the discrete mixing distribution prior to estimation.
The representation of inter-individual taste heterogeneity is a key concern of discrete choice analysis, as information on the distribution of tastes is critical for demand forecasting, welfare analysis and market segmentation. In many empirical settings, the analyst cannot perfectly explain taste heterogeneity in terms of observed individual characteristics, and taste heterogeneity remains to a substantial extent random from the analyst's point-of-view bhat_accommodating_1998. If decision-makers are assumed to employ a decision strategy that is consistent with random utility maximisation, a mixed random utility model such as the mixed multinomial logit (M-MNL) model or the mixed multinomial probit (M-MNP) model can accommodate any empirical random heterogeneity distribution by marginalising the discrete choice kernel over some mixing distribution, which describes the unobserved distribution of tastes in the sample mcfadden_mixed_2000,train_discrete_2009. However, the ability of mixed random utility models to recover any true heterogeneity distribution is only predicated on an existence proof mcfadden_mixed_2000 and therefore, the analyst is required to select an appropriate mixing distribution to capture an unobserved heterogeneity distribution in a given empirical setting.
There are two principal approaches to account for unobserved taste heterogeneity in mixed random utility models wedel_discrete_1999: Parametric mixed random utility models are based on the assumption that individual taste parameters are drawn from a sample-level continuum of tastes with a specific shape. In nonparametric mixed random utility models, individuals are probabilistically assigned to a countable, typically finite number of segments with homogeneous tastes. If the stochastic error terms of the utilities are assumed to be independent and identically distributed Gumbel random variates, the former approach can be referred to as parametric mixed multinomial logit (PM-MNL) model mcfadden_mixed_2000,train_discrete_2009, while a standard implementation of the latter approach is known as latent class multinomial logit (LC-MNL) model bhat_endogenous_1997,greene_latent_2003,kamakura_probabilistic_1989.\footnote{The remainder of this contribution focusses on heterogeneity representations within the M-MNL model, but for completeness, we point out that flexible representations of unobserved taste heterogeneity can also be accommodated by the M-MNP model bhat_new_2017,bhat_new_2012.}
Both the PM-MNL and LC-MNL models are widely used in disciplines studying individual choice behaviour but are subject to limitations: The distributional assumptions of parametric mixture models may be overly rigid and may not yield convincing representations of inter-individual random taste heterogeneity, as the shape of the estimated taste parameter distribution is constrained to be equal to the functional form of the imposed parametric random distribution vij_random_2017. For valid inferences in PM-MNL models, it is therefore imperative to correctly specify the mixing distribution of the randomised taste parameters hess_estimation_2005. In practice however, the analyst is unlikely to be able to exhaust the hypothesis space of theoretically-feasible parametric distribution functions keane_comparing_2013. Nonparametric mixture approaches such as the LC-MNL model free the analyst from rigid distributional assumptions greene_latent_2003 but are cumbersome to estimate, as the number of mixture components needs to be determined exogenously.
In response to the limitations of the PM-MNL and the LC-MNL models, further semi-nonparametric and nonparametric variations of the M-MNL models have been proposed vij_random_2017,yuan_guide_2015.
Several semi-nonparametric approaches combine discrete and continuous representations of unobserved heterogeneity by employing mixing distributions that are finite mixtures of continuous distributions. For example, bujosa_combining_2010, fosgerau_comparison_2009 and keane_comparing_2013 use mixing distributions that are finite mixtures of independent normal distributions. Similarly, greene_revealing_2013 employ a finite mixture of independent triangular distributions as a random taste parameter distribution. train_em_2008 allows for correlation between random taste parameters within mixture components by leveraging a finite mixture of multivariate Gaussians as a mixing distribution. Semi-nonparametric approaches relying on mixing distributions that are finite mixtures of continuous distributions are conceptually appealing, as unbounded continuous distributions can be closely approximated by a finite mixture of multivariate Gaussians keane_comparing_2013. However, such semi-nonparametric approaches are computationally demanding, as the analyst must perform post-hoc model selection to determine the appropriate number of mixture components.
Another class of semi-nonparametric approaches leverages flexible functionals such as Legendre polynomials fosgerau_practical_2007 and B-splines bastin_estimating_2010 to accommodate unobserved taste heterogeneity in M-MNL models. These approaches allow for flexible representations of unobserved heterogeneity but require the analyst to configure the complexity of the employed functional prior to estimation. Relatedly, train_mixed_2016 proposes a semi-nonparametric M-MNL model where an additional discrete mixing distribution is imposed on the parameters of a variety of flexible functionals such as step, spline or polynomial functions. The framework can flexibly recover differently-shaped taste parameter distributions but requires the analyst to select both the complexity of the discrete mixing distribution for the parameters of the functional as well as the functional itself prior to estimation.
The LC-MNL model is the simplest implementation of a M-MNL model with a nonparametric discrete mixing distribution with a finite number of support points, whose locations and associated probability mass need to be estimated. In practice, high-dimensional LC-MNL models are often plagued by identification issues due to the multi-modality of the log-likelihood function. Therefore, a stream of literature proposes nonparametric mixing distributions with structured support points dong_comparison_2014,train_em_2008,vij_random_2017. Effectively, these approaches implement multi-variate histogram estimators by defining a multi-dimensional grid on the coefficients space such that the heterogeneity distribution in question can be closely approximated by estimating the amount of probability mass positioned on the vertices of the grid. M-MNL models with gridded mixing distributions can capture complex heterogeneity distributions vij_random_2017 but require the analyst to specify the complexity of the grid of mass points a priori.
Extant M-MNL models with nonparametric discrete mixing distributions including the LC-MNL model are finite-dimensional nonparametric models, which rely on an a priori specification of a finite, comparatively large parameter space to flexibly cover the desired hypothesis space of heterogeneity distributions. train_em_2008 therefore suggests to refer to finite-dimensional nonparametric models as ”super-parametric“ models due to their large number of parameters. However, to be precise, finite-dimensional and infinite-dimensional nonparametric models can be distinguished gelman_bayesian_2013. The defining characteristic of infinite-dimensional nonparametric models is that the complexity of a model, i.e. the size of the parameter space, is endogenised gelman_bayesian_2013. Such infinite-dimensional models are known as Bayesian nonparametric models, because model complexity is incorporated into the posterior density via stochastic process priors and estimated conditional on the observed data gershman_tutorial_2012,orbanz_bayesian_2011.
Dirichlet process mixture models antoniak_mixtures_1974 are a flexible class of Bayesian nonparametric models, which preclude the a priori specification of the number mixture components in a discrete mixture model by exploiting the properties of the Dirichlet process ferguson_bayesian_1973. The Dirichlet process is a stochastic process, whose realisations are probability distributions. The Dirichlet process exhibits two useful properties: First, realisations from a Dirichlet process are discrete and second, repeated samples from a realisation from a Dirichlet process are clustered with non-zero probability, while the expected number of clusters grows only logarithmically in the sample size gershman_tutorial_2012,teh_dirichlet_2011. When used as a nonparametric prior in discrete mixture models, the Dirichlet process induces a random partition where each individual cluster of the partition is characterised by its own parameter vector for the probability distribution of the response variable. In such Dirichlet process mixture models, the number of mixture components need not be fixed a priori, as it is inferred from the evidence. Because the complexity of the discrete mixing distribution is not fixed prior to estimation, Dirichlet process mixture models can be conceived as infinite mixture models rasmussen_infinite_1999. The truncated stick-breaking construction ishwaran_gibbs_2001 is a finite-dimensional, precise approximation of the Dirichlet process and permits tractable inference in Dirichlet process mixture models.
In this paper, we present a M-MNL model, which leverages the truncated stick-breaking process representation of the Dirichlet process as a flexible nonparametric mixing distribution. The proposed model is a Dirichlet process mixture multinomial logit (DPM-MNL) model and accommodates discrete representations of heterogeneity, like a LC-MNL model. However, unlike a LC-MNL model, the proposed DPM-MNL model does not require the analyst to fix the number of mixture components prior to estimation, as the complexity of the discrete mixing distribution is inferred from the evidence. For posterior inference in the proposed DPM-MNL model, we derive an expectation maximisation (EM) algorithm dempster_maximum_1977. In a simulation study, we demonstrate that the proposed DPM-MNL model can flexibly capture differently-shaped taste parameter distributions. Furthermore, we empirically validate the model framework in a case study on motorists' route choice preferences and find that the proposed model outperforms an LC-MNL model and PM-MNL models with common continuous mixing distributions in terms of in-sample fit and out-of-sample predictive ability. Compared to extant modelling approaches, the proposed DPM-MNL model substantially abbreviates specification searches, as it relies on less restrictive parametric assumptions and does not require the analyst to specify the complexity of its discrete mixing distribution a priori.
Dirichlet process mixture models with kernels for continuous dependent data have been presented in the domains of statistics and machine learning teh_dirichlet_2011. In the context of discrete choice analysis, Dirichlet process mixture modelling methods have been leveraged to construct discrete choice models with discrete ansari_semiparametric_2006,kim_assessing_2004 and continuous burda_bayesian_2008,li_bayesian_2013 representations of taste heterogeneity. kim_assessing_2004 present Dirichlet process mixture models with multinomial logit and multinomial probit kernels. Similarly, ansari_semiparametric_2006 allow for a discrete representation of heterogeneity by defining taste parameters as a Dirichlet process mixture in a variety of Thurstonian models including the multinomial probit model. burda_bayesian_2008 accommodate continuous representations of taste heterogeneity in a Dirichlet process mixture models, where the kernels are mixed logit models with multivariate normal heterogeneity distributions. li_bayesian_2013 present a mixed multinomial probit model, where the mixing distribution of some coefficients is a Dirichlet process mixture of normal distributions.
All of the four applications of Dirichlet process mixture modelling methods in discrete choice analysis employ Markov Chain Monte Carlo methods for model inference and present case studies in the realms of consumer behaviour. Relative to this literature, the DPM-MNL model presented in the current paper affords two innovations: First, we demonstrate that an EM algorithm can be used for posterior inference in a Dirichlet process mixture model of discrete choice. Second, we empirically validate the proposed model framework in a case study on travel demand and comprehensively benchmark the proposed DPM-MNL model against established M-MNL models with both continuous and discrete mixing distributions in terms of both in-sample fit and out-of-sample predictive ability.
Moreover, our paper contributes to two other strands of literature: First, we introduce a Dirichlet process mixture model with a multinomial logit kernel into the domain of behavioural travel demand analysis and demonstrate the value of the proposed model framework in a case study on motorists' route choice preferences. Second, we contribute to a growing body of literature concerned with the development and application of EM algorithms for the estimation of complex discrete choice models bhat_endogenous_1997,sohn_expectation-maximization_2016,train_em_2008,train_discrete_2009,vij_random_2017. While it is well known that the EM algorithm can facilitate inference in finite-dimensional discrete mixture M-MNL models bhat_endogenous_1997,train_em_2008,train_discrete_2009,vij_random_2017, we demonstrate that the computational benefits of the EM algorithm generalise to the infinite-dimensional setting.
The remainder of this paper assumes the following structure: The subsequent section provides conceptual and technical prerequisites to the proposed Dirichlet process mixture model of discrete choice. The formulation of the proposed model framework is presented in Section (ref), and Section (ref) explicates the inference approach. The simulation and case studies are presented in Sections (ref) and (ref). Section (ref) concludes by summarising the proposed modelling approach, by acknowledging limitations and by pointing at directions for future research.
In this section, we provide the conceptual and technical prerequisites to the proposed Dirichlet process mixture model of discrete choice. First, we motivate Dirichlet process mixture models by the desideratum to obviate the a priori specification of the number of mixture components in a finite mixture model (Section (ref)). Next, we introduce the Dirichlet process, which is foundational to Dirichlet process mixture models (Section (ref)), and describe alternative representations of the Dirichlet process (Section (ref)). These alternative representations illustrate the clustering and discreteness properties of the Dirichlet process and are necessary for applications of the Dirichlet process in statistical models. Ultimately, we present the generative process of an exemplative Dirichlet process mixture model (Section (ref)).
In finite mixture models, analytical units are probabilistically assigned to one and only one of a finite number of mixture components, each of which is characterised by its own parameter vector for the component-specific probability distribution of the dependent variable. For example, in the LC-MNL model, decision-makers are distributed over a finite number of taste segments, each of which has its own taste vector parameterising the component-specific multinomial logit (MNL) kernel, from which the observed choices are drawn. The LC-MNL model assumes the following data generating process (DGP):
where $n \in \{1,\dots,N\}$ indexes individuals and $t \in \{1,\dots,T_{n}\}$ indexes choice occasions for individual $n$. $q_{n} \in \{1,\ldots,K\}$ denotes an individual's component assignment, which is drawn from a categorical distribution with parameter $\boldsymbol{\pi}$. The random variable $y_{n,t}$ denotes the chosen alternative and takes values in the set of available alternatives $C_{n,t}$. $y_{n,t}$ is assumed to be drawn from an MNL kernel. Thence, the probability that individual $n$ chooses alternative $j \in C_{n,t}$ on choice occasion $t$ is
where $\boldsymbol{\beta}_{k}$ is the component-specific taste vector. $\boldsymbol{X}_{n,t}$ is a matrix of covariates and $\boldsymbol{X}_{n,t,j}$ is a row of $\boldsymbol{X}_{n,t}$. $V(\boldsymbol{X}_{n,t,j}, \boldsymbol{\beta}_{k})$ is a function giving the deterministic component of utility.
In finite mixture models, the number of mixture components $K$ represents an exogenous model parameter. For this reason, the final specification of a finite mixture model must be determined based on a consideration of post-hoc model selection criteria, which are applied to a candidate set of models with varying numbers of mixture components. As such specification searches often require considerable time and effort, we would like to obviate the a priori specification of the number of mixture components in a discrete mixture model without compromising on the distributional flexibility afforded by a nonparametric mixing distribution. A discrete mixture model where the number of mixture components need not be fixed a priori, as it is inferred from the evidence, is an infinite mixture model rasmussen_infinite_1999. Such a model can be realised with the help of the Dirichlet process ferguson_bayesian_1973.
To allow a mixture model to endogenously adapt the complexity of its discrete mixing distribution to the evidence, we can leverage the properties of the Dirichlet process ferguson_bayesian_1973, which is a stochastic process, whose draws are probability distribution over some measurable space $\Theta$ gelman_bayesian_2013,gershman_tutorial_2012,teh_dirichlet_2011. A Dirichlet process is parameterised by a positive, real-valued concentration parameter $\alpha$ and a base measure $\mbox{G}_{0}$, which itself is a probability distribution on $\Theta$. In that vein, we let $\mbox{G} \sim \mbox{DP}(\alpha,\mbox{G}_{0})$ denote a sample from a Dirichlet process.
The Dirichlet process is an infinite-dimensional generalisation of the Dirichlet distribution (which gives a probability simplex associated with the outcomes of a multinomial event); ferguson_bayesian_1973 shows that if $\mbox{G} \sim \mbox{DP}(\alpha,\mbox{G}_{0})$, any finite, mutually exclusive partition $A_{1}, \ldots, A_{k}$ of a space $\Theta$ exhibits a Dirichlet distribution, i.e.
$\Theta$ may be any measurable space such as the $n$-dimensional real space and $\mbox{G}$ defines how the probability mass is distributed over the partitioned space, i.e. $\mbox{G}(A_{l})$ is the amount of probability mass in region $A_{l}$. $\mbox{G}_{0}$ is an initial guess about $\mbox{G}$ and $\alpha$ controls the proximity of $\mbox{G}_{0}$ and $\mbox{G}$.
The mean of a Dirichlet process is its baseline distribution teh_dirichlet_2011, i.e. $\mathbf{E} \left [ \mbox{G}(A) \right ] = \mbox{G}_{0}(A)$. The concentration parameter $\alpha$ controls the variance of the Dirichlet process teh_dirichlet_2011, i.e. $\mbox{Var} \left [ \mbox{G}_{0}(A) \right ] = \frac{\mbox{G}_{0}(A) \left( 1 - \mbox{G}_{0}(A) \right )}{\alpha + 1}$. For $\alpha \rightarrow 0$, draws from the Dirichlet process agglomerate around the mean. For $\alpha \rightarrow \infty$, $G \rightarrow \mbox{G}_{0}$, i.e. the Dirichlet process approaches the baseline distribution. Figure (ref) illustrates the behaviour of a Dirichlet process with a standard normal base measure for different values of $\alpha$.\footnote{Note that the draws in Figure (ref) were generated with the help of the stick-breaking process construction ishwaran_gibbs_2001,sethuraman_constructive_1994 of the Dirichlet process. We will introduce this alternative representation of the Dirichlet process in Section (ref), once the theory has been further developed.}
The Dirichlet process exhibits two important properties, which qualify it as a nonparametric prior in clustering and segmentation models: First, realisations from the Dirichlet process are discrete and second, repeated samples from a realisation $\mbox{G}$ from the Dirichlet process are clustered with non-zero probability. ferguson_bayesian_1973 provides formal proofs for both properties. However, the properties of the Dirichlet process can also be illustrated in more intuitive ways through alternative representations of the Dirichlet process.
\FloatBarrier
The formal definition of the Dirichlet process ((ref)) is not immediately useful for practical applications. For this reason, three alternative representations of the Dirichlet process have been proposed, namely the Blackwell-MacQueen urn scheme blackwell_ferguson_1973, the Chinese Restaurant process aldous_exchangeability_1985 and the stick-breaking process construction sethuraman_constructive_1994. The Blackwell-MacQueen urn scheme and the Chinese Restaurant process are closely related to one another and are helpful in building intuition about the clustering property of the Dirichlet process. The stick-breaking construction illustrates discreteness property of the Dirichlet process and is leveraged in the proposed Dirichlet process mixture model of discrete choice.
\FloatBarrier
The Blackwell-MacQueen urn scheme blackwell_ferguson_1973 elucidates the clustering property of the Dirichlet process. Consider the following DGP:
i.e. repeated draws $\theta_{1:N}$ are taken from a realisation $\mbox{G}$ from a Dirichlet process. blackwell_ferguson_1973 show that $\theta_{1:N}$ constitute a P\`olya sequence: Conditional on previous draws $\theta_{1:n-1}$, the probability that a new draw $\theta_{n} \sim \mbox{G}$ assumes a new value $\theta^{*} \sim \mbox{G}_{0}$ is $P(\theta_{n} = \theta^{*} \vert \theta_{1:n-1}) = \frac{\alpha}{\alpha + n - 1}$, and the probability that $\theta_{n}$ assumes an existing value $\theta_{k}$ is $P(\theta_{n} = \theta_{k} \vert \theta_{1:n-1}) = \frac{n_{k}}{\alpha + n - 1}$, where $n_{k}$ denotes the number of times $\theta_{k}$ appears in $\theta_{1:n-1}$. The latter probability is non-zero so that $\theta_{1:N}$ are clustered with non-zero probability. We also observe that the probability that $\theta_{n}$ assumes an existing value $\theta_{k}$ is proportional to $n_{k}$. Hence, clustering under the Dirichlet process is subject to preferential attachment and clusters that are large are relatively more likely to grow in size, as new draws are taken. On the other hand, the probability that $\theta_{n}$ assumes a new value is proportional to the concentration parameter $\alpha$: So the larger $\alpha$ is, the more likely are draws distributed over different clusters (see Figure (ref)). Furthermore, it can be shown that the expected number of distinct clusters $\mathbf{E}(K)$ under a Dirichlet process only grows logarithmically in the sample size $N$ and is proportional to $\alpha$, i.e. $\mathbf{E}(K) = \alpha \ln N$ for $\alpha < \frac{N}{\ln N}$ gershman_tutorial_2012,teh_dirichlet_2011.
Ergo, the Dirichlet process induces a random partition where the individual clusters of the partition are each characterised by a specific realisation $\theta_{k}$ from $\mbox{G}_{0}$. The Chinese Restaurant process representation of the Dirichlet process aldous_exchangeability_1985 induces a random partition like the Blackwell-MacQueen urn scheme, but does not assign draws $\theta_{k} \sim \mbox{G}_{0}$ to the clusters. The Chinese Restaurant process representation receives its name from a metaphor describing the process of sequentially seating customers in a restaurant and provides a more vivid illustration of the clustering property of the Dirichlet process: Consider a restaurant with an infinite number of tables, each of which provides seating for an infinite number of customers. The first customer entering the restaurant sits at any table. The second customer takes a seat at the same table as the first customer with probability $\frac{1}{1 + \alpha}$ and at another table with probability $\frac{\alpha}{1 + \alpha}$. The $n$th customer sits at an occupied table with probability proportional to the number of customers already seated at the table and at an unoccupied table with probability proportional to $\alpha$. More formally, the Chinese Restaurant process can be represented as follows: The probability of assignment $q_{n}$ of analytical unit $n \in \{1,\ldots,N \}$ to cluster $k \in \{1,\ldots,K \}$ is $P(q_{n} = k \vert q_{1:n-1}) = \frac{n_{k}}{\alpha + n - 1}$, if $n$ joins existing cluster $k$ with size $n_{k}$, and $P(q_{n} = k \vert q_{1:n-1}) = \frac{\alpha}{\alpha + n - 1}$, if $n$ initiates a new cluster. Figure (ref) illustrates a partition induced by the Chinese Restaurant process.
\FloatBarrier
ferguson_bayesian_1973 shows that the realisations from the Dirichlet process are probability-weighted point masses, i.e.
where $\pi_{k} \in [0,1]$ with $\sum_{k=1}^{\infty} \pi_{k} = 1$ is a probability weight and $\delta_{\theta_{k}}$ is a probability point mass centred at $\theta_{k} \sim \mbox{G}_{0}$. This atomic representation of a realisation from the Dirichlet process demonstrates the discreteness property of the Dirichlet process and is exploited by the stick-breaking process construction sethuraman_constructive_1994 of the Dirichlet process.
The stick-breaking process construction of the Dirichlet process defines a realisation from a Dirichlet process $\mbox{G} \sim \mbox{DP}(\alpha, \mbox{G}_{0})$ as a discrete mixture of point masses, whereby the component weights are factorisations of Beta-distributed random variables, i.e.
with
where $\pi_{k} \in [0,1]$ with $\sum_{k=1}^{\infty} \pi_{k} = 1$ is a probability weight and $\delta_{\theta_{k}}$ is the associated point mass centred at $\theta_{k}$, which is a unique realisation from $\mbox{G}_{0}$. The stick-breaking process receives its name from the metaphor describing the process of breaking a stick of unit length into an infinite number of pieces gelman_bayesian_2013,teh_dirichlet_2011. Figure (ref) illustrates the stick-breaking process: Beginning with a stick of unit length, we break the stick at $\eta_{1} \sim \text{Beta}(1,\alpha)$ and assign $\pi_{1}$ to the piece we broke off. We draw $\eta_{2} \sim \text{Beta}(1,\alpha)$ and break a piece of size $\pi_{2} = \eta_{2} (1 - \eta_{1})$ from the remaining $1 - \eta_{1}$ stick. Subsequently, we continue to break off sticks of sizes $\pi_{3}, \ldots , \pi_{k}$ of the remainder of the stick.
The truncated stick-breaking process ishwaran_gibbs_2001 is a finite-dimensional approximation of the infinite-dimensional stick-breaking process. Under the truncated stick-breaking process representation, $\mbox{G}$ is given by
with
where the truncation level $K$ is chosen by the analyst. To assure that $\sum_{k=1}^{K} \pi_{k} = 1$, the final random variable is degenerate, i.e. $\eta_{K} = 1$ so that $\pi_{K} = 1 - \sum_{k=1}^{K - 1} \pi_{k}$. This stick-breaking construction of a probability vector can be referred to as Griffiths-Engen-McCloskey (GEM) distribution pitman_sequential_2006. We write $\boldsymbol{\pi} \sim \mbox{GEM}_{K}(\alpha)$ to denote that the probability vector $\boldsymbol{\pi}$ is realisation from a truncated stick-breaking process with concentration parameter $\alpha$ and truncation level $K$.
At first glance, the use of a truncated stick-breaking process prior appears to defeat the purpose of the Bayesian nonparametric modelling paradigm, as we are essentially defining a finite mixture model of dimension $K$. However, the truncated stick-breaking process prior induces a shrinkage on the number of effectively populated mixture components, while maintaining the computational advantages of a finite mixture model gelman_bayesian_2013. In fact, the residual probability $\pi_{K}$ is negligibly small for reasonably large $K$ and most $\alpha$ values that are encountered in practice ishwaran_markov_2000,ohlssen_flexible_2007. In Section (ref), we provide a detailed discussion about the choice of $K$.
For inference in the proposed Dirichlet process mixture model of discrete choice, we exploit the fact that the truncated stick-breaking representation of the Dirichlet process is a generalised Dirichlet distribution connor_concepts_1969, which is the joint distribution of $K-1$ independent Beta-distributed random variables: Let
and $\eta_{K} = 1$. Furthermore, define $\pi_{k} = \eta_{k} \prod_{l=1}^{k-1}(1 - \eta_{l})$. Then, the joint density of $\pi_{1:K}$ is given by
$B(a,b)$ denotes the Beta function evaluated at $\{a,b\}$. The joint density $\pi_{1:K}$ under the truncated stick-breaking process is obtained by letting $a_{k} = 1$ and $b_{k} = \alpha$ $\forall$ $k = 1,\ldots,K - 1$. The generalised Dirichlet distribution is the conjugate prior of the multinomial distribution. The compound distribution of this conjugate pair is known as generalised-Dirichlet-multinomial distribution and assumes the following DGP:
where $\boldsymbol{x}$ is a $K$-dimensional vector of category counts. The marginal density of $\boldsymbol{x}$ is given by zhou_mm_2010:
where $N = \sum_{k=1}^{K} x_{k}$ and $m_{k} = \sum_{k'=k}^{K} x_{k'}$. $\Gamma(\cdot)$ denotes the Gamma function.
The Dirichlet process can be used as a nonparametric prior in mixture models to construct Dirichlet process mixture models antoniak_mixtures_1974. In Dirichlet process mixture models, the number of mixture components is not fixed, but is endogenously determined based on the evidence blei_variational_2006,gelman_bayesian_2013,mcauliffe_nonparametric_2006. Staying with the example of a mixture model with MNL kernels ((ref)--(ref)), we can write out the generative process of an exemplative Dirichlet process mixture model:
where
In practical terms, the Dirichlet process prior in the generative process ((ref)--(ref)) defines an infinite mixture model by precluding the a priori specification of the number of mixture components. Due to the clustering property of the Dirichlet process, $\boldsymbol{\beta}_{1:N}$ are clustered with non-zero probability and the sample can be partitioned ex post into countable segments based on the distinct component-specific parameter values blei_variational_2006.
We now present the formulation of a Dirichlet process mixture model, where the component-specific probability distribution functions are MNL kernels.\footnote{For completeness, we point out that it is straightforward to generalise our proposed model framework including the corresponding inference approach presented in Section (ref) to other kernels that are commonly considered in discrete choice analysis.} Our model formulation involves the truncated stick-breaking process construction of the Dirichlet process as a computationally efficient means to allow the proposed discrete choice model to adapt the complexity of its discrete mixing distribution to the evidence.
The generative process of the proposed Dirichlet process mixture multinomial logit (DPM-MNL) model is visualised in Figure (ref) and can be described as follows: Decision-makers are indexed by $n \in \{1, \ldots, N \}$ and are distributed over $K$ mixture components indexed by $k \in \{1, \ldots, K \}$. Each mixture component is characterised by a taste vector $\boldsymbol{\beta}_{k}$ with prior $\mbox{G}_{0}$. The latent variable $q_{n}$ indicates a decision-maker's component allocation such that $q_{n} = k$, if decision-maker $n$ is assigned to mixture component $k$. $q_{n}$ controls an individual's tastes $\boldsymbol{\beta}_{k}$ such that an observed choice $\boldsymbol{y}_{n,t}$ is a function of tastes $\boldsymbol{\beta}_{k}$ and covariates $\boldsymbol{X}_{n,t}$. $q_{n}$ is a realisation from a categorical distribution with parameter $\boldsymbol{\pi}$, which in turn is obtained via a $\mbox{GEM}_{K}$ distribution with truncation level $K$ and concentration parameter $\alpha$. $\alpha$ is drawn from some distribution f.
Stated succinctly, the generative process of the DPM-MNL model is:
Theoretically, the generative process of the DPM-MNL allows for a large number of populated mixture components. In practice however, the effective number of mixture components, i.e. the number of non-empty components, is considerably smaller than $K$ due to the clustering property of the Dirichlet process. As a reference, Figure (ref) shows the generative process of an LC-MNL model: The generative processes of the DPM-MNL and LC-MNL models resemble each other closely. Yet, in the case of the DPM-MNL model, the truncated stick-breaking process prior on $\boldsymbol{\pi}$ induces shrinkage on the number of effectively populated mixture components, and a prior $\mbox{G}_{0}$ is placed on the component-specific taste parameters $\beta_{1:K}$.
\FloatBarrier
Next, our goal is to define the posterior probability of the DPM-MNL model parameters $\alpha$ and $\boldsymbol{\beta}_{1:K}$ conditional on the observed choices $\boldsymbol{y} = \left \{ y_{n,1:T_{n}} \right \}_{n=1}^{N}$ and the covariates $\boldsymbol{X} = \{ \boldsymbol{X}_{n,1:T_{n}} \}_{n=1}^{N}$. First, we re-iterate that the component-specific probability distribution functions are MNL models. Hence, the probability of individual $n$ choosing alternative $j$ on choice occasion $t$ conditional on assignment to component $k$ is given by
where $y_{n,t} \in C_{n,t}$ denotes the observed choice for individual $n$ on occasion $t$. $\boldsymbol{X}_{n,t}$ is a matrix of covariates and $\boldsymbol{X}_{n,t,j}$ is a row of $\boldsymbol{X}_{n,t}$. $\boldsymbol{\beta}_{k}$ is a vector of component-specific taste parameters, and $V(\boldsymbol{X}_{n,t}, \boldsymbol{\beta}_{k})$ is a function giving the deterministic component of utility. $C_{n,t}$ denotes the choice set. ((ref)) is iterated over alternatives $j \in C_{n,t}$ and choice occasions $t = 1,\ldots,T_{n}$ to obtain the probability of observing choice vector $\boldsymbol{y}_{n} = y_{n,1:T_{n}}$ for individual $n$ conditional on component allocation $q_{n}$:
where $\boldsymbol{X}_{n} = \boldsymbol{X}_{n,1:T_{n}}$ and $[A]$ is the Kronecker delta, which equals one, if $A$ is true, and zero otherwise. Note that ((ref)) encapsulates the assumption that an individual's repeated choices are independent from one another conditional on the individual's component assignment. In discrete choice analysis, it is standard practice to adopt this exact conditional independence assumption, when longitudinal choice data are analysed with the help of mixed random utility models revelt_mixed_1998.
Since $q_{n}$ is unobserved, we marginalise over possible values of $q_{n}$ and condition the probability of observing choice vector $\boldsymbol{y}_{n}$ on the prior distribution over components $\boldsymbol{\pi}$ instead of on component assignment $q_{n}$:
The probability of observing the choice data $\boldsymbol{y} = \boldsymbol{y}_{1:N}$ for the sample is obtained by iterating ((ref)) over individuals:
where $\boldsymbol{X} = \{ \boldsymbol{X}_{n,1:T_{n}} \}_{n=1}^{N}$. Due to the stick-breaking construction of the component weights, $\boldsymbol{\pi}$ is a function of $K-1$ Beta random variables denoted by $\boldsymbol{\eta} = \eta_{1:K-1}$. The joint density of $\boldsymbol{\eta}$ is
Since $\boldsymbol{\eta}$ is also unobserved, we marginalise ((ref)) over ((ref)) to obtain the likelihood of the choice data $\boldsymbol{y}$ conditional on the DPM-MNL model parameters $\alpha$ and $\boldsymbol{\beta}$:
where
By Bayes' rule, the posterior probability is equal to the likelihood times the prior divided by the evidence. Hence,
where $P(\boldsymbol{y} \vert \alpha, \boldsymbol{\beta}_{1:K} ,\boldsymbol{X})$ is defined in ((ref)). $f(\alpha)$ denotes the prior density of $\alpha$. $\boldsymbol{\beta}_{k}$, $k = 1, \ldots, K$ are sampled from the base measure $\mbox{G}_{0}$; $g(\boldsymbol{\beta})$ denotes the prior density of $\boldsymbol{\beta}_{1:K}$. The denominator in ((ref)) represents the model evidence, which does not depend on the DPM-MNL model parameters $\alpha$ and $\boldsymbol{\beta}$.
Having defined the posterior probability of the DPM-MNL model parameters, we wish to devise a method for posterior inference. Two factors complicate inference in the DPM-MNL model: First, exact inference is not possible, as both the numerator and the denominator of the posterior probability ((ref)) involve integrations that are not analytically tractable. Second, the likelihood function ((ref)) involves a summation of the $K$ component-specific MNL kernels. Hence, the likelihood function is likely to exhibit multiple modes and a closed-form expression for the gradient of the log-likelihood function does not exist.
In general, inference in Dirichlet process mixture models has been carried out using Markov Chain Monte Carlo (MCMC) methods neal_markov_2000,ishwaran_gibbs_2001 and variational inference methods blei_variational_2006. MCMC methods treat the latent component assignments as random model parameters and approximate a posterior distribution through samples from a Markov Chain, whose stationary distribution is the posterior distribution of interest blei_variational_2006. Variational inference methods hinge on finding a variational distribution over the latent variables to approximate the difficult-to-compute posterior distribution wainwright_graphical_2008. The variational distribution is characterised by its own variational parameters, which are chosen such that the variational distribution and the posterior of interest are close to one another. Subsequently, inference on the model parameters proceeds with the variational distribution as a surrogate posterior distribution.
Both inference approaches have been successfully employed in numerous empirical applications carvalho_particle_2010,wang_fast_2011 but are subject to limitations: MCMC methods are computationally intensive and do not scale well to larger datasets (for an argumentation in the context of Dirichlet process mixture models, see wang_fast_2011, wang_fast_2011; for an argumentation in the context of discrete choice models, see braun_variational_2010, braun_variational_2010). In particular, when models depend on discrete latent quantities---such as labels in a mixture model---, MCMC samplers may exhibit poor mixing, and label-switching issues need to be addressed gelman_bayesian_2013. Variational inference methods allow for fast inference but give estimates that may not be asymptotically efficient. Thus, variational inference methods are most suitable for repeated inference on very large datasets and for applications where precise parameter estimates are not a primary concern blei_variational_2017.
In this paper, we conceive inference in a Dirichlet process mixture model as a missing data problem and leverage the expectation maximisation (EM) algorithm dempster_maximum_1977,mclachlan_em_2008 for maximum a posteriori (MAP) estimation of the DPM-MNL model parameters. MAP estimation is a computationally efficient Bayesian inference approach, whose objective is to identify the mode of the log-posterior density. From a practical point-of-view, MAP estimation extends maximum-likelihood estimation by accounting for prior information about the distribution of the model parameters. The implementation of MAP estimation approaches is generally possible via gradient-based optimisation routines. However, in the case of the proposed DPM-MNL model, gradient-based optimisation approaches are not feasible due to the complications outlined in the first paragraph of this subsection.
In principle, an EM algorithm for MAP estimation problems alternates between an expectation step (E-step) and a maximisation step (M-step) until a convergence criterion is satisfied. In the E-step, the expectation of the posterior density is computed. The objective of the subsequent M-step is to identify the mode of the expected posterior density by maximising over the set of unknown model parameters. Our main rationale for leveraging the EM algorithm for posterior inference in the DPM-MNL model is that the M-step results in a set of standard optimisation problems that are much simpler than the direct optimisation of the posterior density. In addition, label-switching is not a concern due to the deterministic nature of the EM algorithm. For general discussions of the properties of the EM algorithm, we refer to the literature dempster_maximum_1977,mclachlan_em_2008,train_discrete_2009.
The EM algorithm is a fast, scalable and precise inference approach, which is suitable for the estimation of a variety of latent variable models mclachlan_em_2008. In the domain of discrete choice analysis, bhat_endogenous_1997 presents an EM algorithm for the estimation of an LC-MNL model. train_em_2008,train_discrete_2009 devises EM algorithms for the estimation of M-MNL models with a variety of nonparametric and semi-nonparametric mixing distributions. Moreover, vij_random_2017 introduce an EM algorithm for inference in a M-MNL model with a gridded mixing distribution. sohn_expectation-maximization_2016 derives an EM algorithm for the integrated choice and latent variable model walker_generalized_2002.
We begin the derivation of the EM algorithm for the DPM-MNL model by writing out the joint distribution of the DPM-MNL model parameters $\alpha$, $\boldsymbol{\beta}_{1:K}$ and the latent variables $q_{1:N}$:
where $v_{k} = \sum_{k'=k}^{K} \sum_{n=1}^{N} [q_{n} = k']$. The joint probability ((ref)) consists of four factors (one in each line): The first factor follows from the fact that the distribution of the latent variables $q_{1:N}$ is the generalised-Dirichlet-multinomial distribution. This is because $N$ samples from a categorical distribution define a multinomial distribution with $N$ trials, and the multinomial distribution is conjugate to the GEM distribution, a special case of the generalised Dirichlet distribution. The second factor factorises the component-specific MNL kernels for each component, each choice occasion and each decision-maker. The third factor represents the prior density of $\alpha$, and the fourth factor factorises the prior densities of the taste vectors $\boldsymbol{\beta}_{1:K}$.
By Bayes' rule the joint probability ((ref)) is proportional to the posterior density of interest, i.e. $P(\alpha, \boldsymbol{\beta}_{1:K} \vert \boldsymbol{y}, \boldsymbol{X}) \propto P(\alpha, \boldsymbol{\beta}_{1:K}, q_{1:N}, \boldsymbol{y}, \boldsymbol{X})$, where the normalising constant, i.e. the marginal likelihood in the denominator, can be disregarded as it does not depend on the DPM-MNL model parameters. Finding the mode of the posterior density is therefore equivalent to finding the mode of the joint probability. The EM algorithm facilitates the optimisation of the joint probability by imputing the missing data in the E-step. In the sequel, we present the E- and M-steps of the algorithm and assume that the algorithm is initialised with a set of starting values $\{ \alpha^{(\tau)}, \boldsymbol{\beta}^{(\tau)} \}$.
\paragraph{E-step} In the E-step, we compute the expectation $Q (\alpha, \boldsymbol{\beta} \vert \alpha^{(\tau)}, \boldsymbol{\beta}^{(\tau)})$ of the complete-data posterior density with respect to the latent variables conditional on the current parameter estimates. In the case of ((ref)), the probability mass functions $[q_{n} = k]$ for $n = 1,\ldots,N$, $k = 1,\ldots,K$ are the sufficient statistics required for the estimation of the unknown model parameters. Hence, we can determine the expectation of the complete-data posterior density by computing the expectation of each sufficient statistic $[q_{n} = k]$ conditional on the current parameter estimates $\alpha^{(\tau)}$, $\boldsymbol{\beta}^{(\tau)}$. From Bayes' rule, we have
where
Consequently, the expectation of the complete-data posterior density is
where $v_{k} = \sum_{k'=k}^{K} \sum_{n=1}^{N} \omega_{n,k'}$.
\paragraph{M-step} In the M-step, we identify the mode of the surrogate function ((ref)) by maximising over the set of unknown model parameters. We update the model parameters by solving
which can be separated into into two comparatively easy optimisation problems. First, the concentration parameter is updated by numerically solving:
where $w_{k} = \sum_{k'=k}^{K} \sum_{n=1}^{N} \omega_{n,k'}$. Note that the last summand of the objective function represents the log prior density of $\alpha$. Second, the component-specific parameters are updated:
where the first summand of the objective function is essentially the log-likelihood of a weighted multinomial logit model. The second summand of the objective function represents the prior density over the taste vector $\boldsymbol{\beta}_{k}$. If the logarithm of the prior density and the gradient of the logarithm of the prior density exist in closed form, ((ref)) can be solved with the help of standard gradient-based optimisation routines.
Parameter estimates for the DPM-MNL model are obtained by cycling through the E- and M-steps presented above, until a convergence criterion is satisfied. As each iteration of the EM algorithm results in an improvement of the expected value of the log-posterior density, convergence can be assessed by considering the improvement of the expected log-posterior density relative to the previous iteration. In the subsequent applications of the inference approach, the execution of the EM algorithm is terminated, if the improvement of the expected log-posterior density between successive iterations is less than 0.01% of the current expected log-posterior density.
The EM algorithm is a deterministic algorithm, which, given a certain set of starting values will always converge to the same local optimum. However, the algorithm does not guarantee convergence to a global optimum. In practice, it is therefore critical to choose good starting values to avoid that the algorithm terminates in a local optimum. The analyst may use her intuition or may employ a systematic approach to choose starting values. In the subsequent applications of the EM algorithm, we adopt train_em_2008's (train_em_2008) procedure, i.e. we randomly partition the sample into $K$ groups and estimate separate MNL models for each of the groups. The MNL estimates are then used as starting values for the component-specific taste parameters and mixture components are assigned equal weights.
Moreover, the truncation level $K$ of the truncated stick-breaking process prior must be set by the analyst. Generally, the choice of $K$ affects the quality of the approximation of the Dirichlet process via the truncated stick-breaking process, but also impacts the computational tractability of the inference approach ohlssen_flexible_2007. In that vein, the literature has considered different truncation levels: For example, ishwaran_gibbs_2001 as well as li_bayesian_2013 use $K = 150$, while blei_variational_2006 employ a truncation level of $K = 20$; gelman_bayesian_2013 suggest that $K \in [20;50]$ suffices for most practical applications. In general, $K$ is naturally truncated by the sample size. In addition, the analyst may have a strong prior belief about the number of heterogeneity components that are required to explain the observed data and may set $K$ based on her prior belief. Either way, $K$ should be set such that components with a large indices have low probabilities of being occupied gelman_bayesian_2013. To this end, it is advisable to place an informative prior on the concentration parameter $\alpha$, since $\alpha$ has an immediate effect on the number of occupied mixture components ohlssen_flexible_2007. In the subsequent applications of the inference approach, we follow li_bayesian_2013 and set $K = 150$ to assure a close approximation of the Dirichlet process. In addition, we follow ishwaran_approximate_2002 and let $\alpha \sim \mbox{Gamma}(2,2)$, i.e. the prior density of $\alpha$ is Gamma with shape 2 and scale 2.
Thus far, we have not explicitly specified the base measure $\mbox{G}_{0}$. To facilitate model inference, it is generally useful to let $\mbox{G}_{0}$ be conjugate to the kernel gelman_bayesian_2013. However, in the case of the proposed DPM-MNL model, conjugacy is not a desideratum, as the MNL kernel does not have a conjugate prior. Instead, we would like to employ a base measure where the log-density and the gradient of the log-density can be expressed in closed form to allow for the application of standard gradient-based optimisation routines. gelman_bayesian_2013 advise against the use of diffuse base measures, as a high variance of the base measure may penalise the addition of new mixture components and may thus limit the flexibility of the mixing distribution. An obvious choice for $\mbox{G}_{0}$ is the normal distribution. In the subsequent applications of the DPM-MNL model, we let $\mbox{G}_{0} = \mbox{N}(0,5^{2})$. For taste parameters that are constrained to be strictly positive or negative---such as the taste parameter capturing sensitivity to cost---, we employ half-normal priors with scale 5. For priors to be meaningful, the scale of the corresponding parameters may have to be adjusted. In the subsequent applications of the DPM-MNL model, we scale the covariates such that the absolute values of the coefficient estimates of a standard MNL model applied to the same dataset as the DPM-MNL model are between 0.1 and 1.
In this section, we present a simulation study consisting of four Monte Carlo experiments to demonstrate the behaviour of the proposed DPM-MNL model for different taste parameter distributions.
For the simulation study, we generate multiple synthetic samples, each comprising 2,000 individuals, who are pseudo-observed to each complete eight choice tasks. The choice scenarios include three unlabelled alternatives, which are characterised by three attributes, namely in-vehicle travel time, out-of-vehicle travel time and travel cost. Table (ref) in the appendix details the data generating process of the simulated attribute levels.
The individuals are assumed to be utility maximisers and to evaluate alternatives based on the following utility specification:
where $n$ indexes individuals, $t$ indexes choice occasions and $j$ indexes alternatives. $\beta_{n,\mbox{ivtt}}$, $\beta_{n,\mbox{ovtt}}$ and $\beta_{n,\mbox{cost}}$ represent taste parameters, which respectively pertain to in-vehicle travel time (ivtt), out-of-vehicle travel time (ovtt) and travel cost. $\epsilon_{n,t,j}$ is a disturbance assumed to be i.i.d. $\mbox{Gumbel} \left( 0,\frac{\pi^2}{6} \right)$. $\beta_{n,\mbox{ivtt}}$, $\beta_{n,\mbox{ovtt}}$ and $\beta_{n,\mbox{cost}}$ vary randomly across individuals. The utility function is specified in willingness-to-pay space so that $\beta_{n,\mbox{ivtt}}$ and $\beta_{n,\mbox{ovtt}}$ are identical to the implicit values of the corresponding attributes.
In each of the four Monte Carlo experiments, the distributions of $\beta_{n,\mbox{ivtt}}$, $\beta_{n,\mbox{ovtt}}$ and $\beta_{n,\mbox{cost}}$ are manipulated: In the first experiment, $\beta_{n,\mbox{ivtt}}$ and $\beta_{n,\mbox{ovtt}}$ are sampled from a bivariate normal distribution so that the joint distribution of the two parameters is uni-modal with light tails. In the second experiment, $\beta_{n,\mbox{ovtt}}$ is log-normally distributed, while $\beta_{n,\mbox{ivtt}}$ is normally distributed; hence, the joint distribution of the two parameters is uni-modal and exhibits a heavy tail for one of the marginals. In the third experiment, $\beta_{n,\mbox{ivtt}}$ and $\beta_{n,\mbox{ovtt}}$ are sampled from a two-component-mixture of normals so that the joint distribution of the two parameters is bi-modal. In the fourth experiment, the joint distribution of $\beta_{n,\mbox{ivtt}}$ and $\beta_{n,\mbox{ovtt}}$ is tri-modal, as the two parameters are drawn from a three-component mixture of normals. In each of the four experiments, the negative of $\beta_{n,\mbox{cost}}$ is log-normally distributed to assure strict negativity of $\beta_{n,\mbox{cost}}$. In each of the four experiments, the location parameter of the distribution of $\beta_{n,\mbox{cost}}$ is set such that the error rate is roughly 7%, i.e. in 7% of the cases, decision-makers deviate from the deterministically best alternative due to the stochastic component. Table (ref) in the appendix gives the data generating processes of the taste parameters for the four Monte Carlo experiments.
For each of the four Monte Carlo experiments, we estimate the proposed DPM-MNL model, using the inference approach presented in Section (ref). We implement the inference method by writing our own MATLAB code and use Train's (train_em_2008) procedure to obtain starting values.
Figure (ref) visualises the true and the estimated taste parameter distributions for the simulation study. Each subfigure (see Figures (ref)--(ref)) corresponds to one of the four Monte Carlo experiments. In each subfigure, the first row shows histograms of the true joint density of $\beta_{n,\mbox{ivtt}}$ and $\beta_{n,\mbox{ovtt}}$; the second figure in the first row of each subfigure additionally displays the estimated point masses of the DPM-MNL model. By design, the DPM-MNL model yields discrete, non-smooth heterogeneity representations. To allow for a comparison of the true and the estimated distribution function, we estimate kernel density functions with normal kernel and bandwidth 2.5 for both the true and the estimated taste parameter distributions.\footnote{In the case of the DPM-MNL model, the kernel density functions are estimated based on 2,000 random draws from the estimated taste parameter distributions.} Plots of these kernel density functions are given in the second row of each subfigure.
Overall, we observe that the DPM-MNL performs well at recovering differently-shaped taste parameter distributions. The DPM-MNL model captures unobserved taste heterogeneity by parsimoniously placing few mass points in the hypothesis space. The estimated locations of these mass points may not necessarily coincide with the modes of the true heterogeneity distributions. However, the estimated kernel density functions show that the DPM-MNL model is able to correctly identify the effective support and the modes of the true taste parameters distributions in all four Monte Carlo experiments.
The estimated concentration parameters of the $\mbox{GEM}_{K}$ distribution are 6.9, 7.0, 7.5 and, respectively, 7.5 for each of the four Monte Carlo experiments. Given these estimates of $\alpha$, the expected numbers of mixture components with an occupancy of at least one decision-maker are 41, 42, 46 and 44, respectively, in each of the four Monte Carlo experiments.
In this section, we empirically validate the proposed model framework in a case study on motorists' route choice preferences.
Data for our analysis are sourced from the German Value of Time and Reliability Study, which was commissioned by the German Federal Ministry of Transport and Digital Infrastructure to obtain estimates of travellers' valuation of travel time and reliability for the Federal Transport Investment Plan 2030, a strategic programme for the appraisal of federal transport infrastructure projects in Germany. The German Value of Time and Reliability Study involved a series of stated choice experiments to elicit tastes with respect to strategic level of service attributes in the context of mode choice, route choice, travel itinerary choice, workplace choice and residential location choice. More information about the scope of the study and the data collection is provided by axhausen_ermittlung_2015 and ehreke_experiences_2014. For the case study, we consider a stated choice experiment on motorists' route choice preferences. The choice tasks required respondents to choose the best of two route alternatives, each of which was characterised by five attributes, namely free-flow travel time, access time, time spent in congested traffic conditions, probability of a significant delay and travel cost. The sample considered for the case study comprises 3,579 cases and 455 individuals.
The stated choice data are used to estimate the proposed DPM-MNL model. In addition, we estimate MNL, PM-MNL and LC-MNL models to benchmark the performance of the proposed DPM-MNL model against established modelling approaches. The DPM-MNL model is estimated, using the inference approach presented in Section (ref); we write our own MATLAB code to implement the inference method and employ Train's (train_em_2008) method to obtain starting values. The PM-MNL models are estimated, using maximum-simulated likelihood methods train_discrete_2009 in conjunction with PythonBiogeme bierlaire_pythonbiogeme:_2016. For each individual, 2,000 simulation draws generated via the Modified Latin Hypercube Sampling method hess_use_2006 are used. For the estimation of the LC-MNL models, we employ the EM algorithm as outlined in train_em_2008 and write our own MATLAB code to implement the inference approach. Again, Train's (train_em_2008) procedure is employed to obtain starting values.
Two different PM-MNL models are considered: The first model assumes that the implicit attribute values are normally distributed, while the second model assumes log-normally distributed implicit attributes values. In both models, the taste parameter capturing sensitivity to cost is assumed to be log-normally distributed. In the first model, utility is specified in willingness-to-pay space to assure that the implicit attribute values follow the desired random distribution. In the second model, it is not imperative to define the utility function in willingness-to-pay space, as the ratio of two log-normally distributed random variables is again log-normally distributed. Therefore as well as for numerical reasons, utility in the second model is specified in preference space. In both model specifications, the random taste parameters are assumed to be independent from one another.
In the case of the LC-MNL model, we are required to estimate multiple model specifications with varying numbers of mixture components, since model complexity represents an exogenous model parameter in a LC-MNL model. A final model specification can be selected based on a consideration of statistical information criteria such as the Akaike Information Criterion akaike_new_1974 and the Bayesian Information Criterion schwarz_estimating_1978. In that vein, we initiate the specification search by first estimating a two-component LC-MNL model and increment the number of mixture components in subsequent estimation runs. A comparison of the estimated LC-MNL model specifications is given in Table (ref). In general, AIC and BIC can attain their minima at different model specifications, because the BIC penalises model complexity more strictly than the AIC. In the present application, the BIC suggests that a model specification with six latent classes is optimal, while the AIC suggests that a model specification with 14 latent classes should be preferred.
\FloatBarrier
The performance of the DPM-MNL, MNL, PM-MNL and LC-MNL models is assessed in terms of both in-sample fit and out-of-sample predictive ability. To evaluate the out-of-sample predictive ability of each of the considered models, we employ a ten-fold cross-validation approach. To this end, the sample is randomly divided into ten subsamples of equal size. In each rotation of the validation procedure, each of the considered models is trained on nine of the subsets, while the held-out subsample is used to compute the predictive log-likelihood for the fold. In each of the ten rotations of the cross-validation procedure, a different subsample is held-out for validation so that in the end, each of the ten subsamples is used once for validation. The average of the predictive log-likelihood values for each of the folds gives the ten-fold cross-validated log-likelihood.
Table (ref) gives log-likelihood values indicating the in-sample fit and the out-of-sample predictive ability for each of the estimated models. The hypothesis of taste homogeneity can be soundly rejected, as the MNL model is substantially outperformed by the competing models. Moreover, the PM-MNL models with discrete heterogeneity representations outperform the two M-MNL with continuous parametric mixing distribution in terms of both in-sample fit and out-of-sample predictive ability. The DPM-MNL provides the best in-sample fit and out-of-sample predictive ability of all considered models. We re-iterate that inference in the LC-MNL model required an extensive specification search (see Table (ref)), whereas inference in the DPM-MNL model was instantaneous.
Figure (ref) shows the estimated cumulative distribution functions of the implicit attribute values for the DPM-MNL, the two PM-MNL models and the LC-MNL model with 14 mixture components. In addition, Figure (ref) visualises the estimated cumulative distribution function of the taste parameter capturing sensitivity to cost for the four models. It can be seen that the four models yield distinct representations of taste heterogeneity: By definition, the normal distribution is symmetric and exhibits comparatively light tails. The log-normal only has support on the strictly positive real line and exhibits heavy tails. Both the DPM-MNL and LC-MNL models represent heterogeneity in a discrete fashion so that the estimated cumulative distribution functions are not smooth.
A closer inspection of the cumulative distribution functions of the implicit attribute values for the the DPM-MNL model reveals several interesting features of the discrete heterogeneity representation produced by the DPM-MNL and LC-MNL model (see Figure (ref)). First, the heterogeneity distributions under the DPM-MNL and LC-MNL models are not confined to follow any specific parametric function. Therefore, the two models are able to exhibit comparatively light tails in the second orthant without compromising on the flexibility of the heterogeneity representation in the first orthant. By comparison, the normal distribution can assign non-zero amounts of probability mass to the first orthant, but symmetry constraints confine its ability to flexibly represent heterogeneity in the first orthant. Likewise, the strictly positive support of the log-normal distribution comes at the expense of heavy tails in the first orthant. Second, the DPM-MNL and LC-MNL models are capable of endogenously recovering patterns of attribute non-attendance hess_its_2013. For example, in the case of the attributes access time and delay probability, we observe that a small, yet noticeable section of the cumulative distribution functions are almost perfectly aligned with the ordinate, which suggests that a noteworthy proportion of the subjects are insensitive to these attributes.
\FloatBarrier
Furthermore, we observe that the estimated cumulative distribution functions of the random taste parameter capturing sensitivity to cost differ across the considered models (see Figure (ref)). In the case of the PM-MNL model with normally distributed implicit attribute values, the estimated cumulative distribution function exhibits a rather heavy tail. However, the cumulative distribution functions of the other models do not suggest that such extreme sensitivities are present. We also note that an upper bound of $-0.001$ was imposed on the estimate on cost sensitivity in the DPM-MNL and LC-MNL models to assure finite moments for the distributions of the implicit attribute values daly_assuring_2012. However, this bound is only relevant for a negligible proportion of the mixture components, and the vast majority of the subjects attend to the cost attribute.
The differences in in-sample fit and in the shapes of the estimated heterogeneity distributions across the considered models are also reflected in quantitative differences in willingness-to-pay measures. Table (ref) gives summary statistics describing the estimated distributions of the implicit attribute values. In accordance with Figure (ref), we observe that the estimated heterogeneity distributions differ noticeably, yet not starkly, in terms of their median values. However, differences in the dispersions of the estimated heterogeneity distributions are more pronounced, as the interquartile ranges (IQRs) and the interdecile ranges (IDRs) differ substantially across the estimated models. Due to the light tails of the normal distributions, the estimated heterogeneity distributions of the PM-MNL model with normal heterogeneity are the least dispersed. By contrast, the estimated heterogeneity distributions of the LC-MNL model with 14 components and the PM-MNL model with log-normal heterogeneity exhibit considerably larger IQRs and IDRs. Due to the influence of the base measure, the marginal heterogeneity distributions of the DPM-MNL model exhibit lighter tails than the LC-MNL model with 14 components and the PM-MNL model with log-normal heterogeneity.
\FloatBarrier
Since the DPM-MNL model also identifies dependencies between random taste parameters, we can meaningfully examine the joint distribution functions of selected pairings of sensitivities and implicit attribute values. Figure (ref) displays the estimated joint densities of two pairings of implicit attribute values. For both density functions, we observe that the densities are most highly concentrated in regions with low attribute valuations. However, both joint density functions also exhibit outlying peaks, where the valuations of one or both attributes assume comparatively large values.
For completeness, we report that the estimate of the concentration parameter $\alpha$ is 11.7. Hence, the expected number of mixture components with at least one occupant is 44.
The representation of unobserved taste heterogeneity is a key concern of discrete choice analysis. The literature distinguishes two principal approaches to accommodate unobserved taste heterogeneity in M-MNL models, namely PM-MNL and LC-MNL models. Both approaches are widely used in disciplines studying individual choice behaviour but are subject to limitations: PM-MNL models rely on potentially restrictive distributional assumptions, as the shape of the estimated taste parameter distribution is constrained to be equal to the functional form of the imposed parametric random distribution vij_random_2017. LC-MNL models free the analyst from parametric assumptions greene_latent_2003 but require the analyst to exogenously determine the number of mixture components.
In this paper, we leverage Bayesian nonparametric methods in conjunction with the Dirichlet process to capture unobserved taste heterogeneity in an infinite-dimensional generalisation of the LC-MNL model. Specifically, we present a M-MNL model, which employs the truncated stick-breaking construction of the Dirichlet process as a flexible nonparametric discrete mixing distribution. The resulting model is a Dirichlet process mixture multinomial logit (DPM-MNL) model, which fuses the benefits of a normative mixed random utility model with a flexible statistical method for the representation of unobserved heterogeneity. The DPM-MNL model relies on less restrictive parametric assumptions than a PM-MNL model and in contrast to a LC-MNL model, the analyst is not required to specify the complexity of the discrete mixing distribution a priori, as the DPM-MNL model infers the effective number of non-empty mixture components from the evidence. For posterior inference in the proposed DPM-MNL model, we derive an EM algorithm. In a simulation study, we demonstrate that the proposed model framework can flexibly capture differently-shaped taste parameter distributions. Furthermore, the proposed model framework is empirically validated in a case study on motorists' route choice preferences. We find that in terms of in-sample fit and out-of-sample predictive ability, the proposed DPM-MNL model outperforms a LC-MNL model as well as PM-MNL models with common mixing distributions.
Compared to extant modelling approaches, our proposed DPM-MNL model offers two benefits: First, only limited parametric assumptions about the nature of the base measure are required for the DPM-MNL model to be able to flexibly recover differently-shaped taste parameter distributions. Second, the DPM-MNL model captures heterogeneity distributions nonparametrically, but does not require the analyst to specify the size of the parameter space a priori, as the complexity of the discrete mixing distribution is inferred from the evidence. As a result, the DPM-MNL model can flexibly recover differently-shaped taste parameter distributions and can identify patterns of attribute non-attendance. Overall, the flexible and adaptive nature of the DPM-MNL model framework substantially reduces the need for parametric assumptions and allows for the identification of patterns of heterogeneity, which could otherwise only be revealed through extensive specification searches.
Our proposed model framework is not devoid of limitations: First, the DPM-MNL model yields rather unstructured representations of heterogeneity. Even though it is straightforward to compute summary statistics for the estimated taste parameter distributions (see Table (ref)), certain empirical applications may require smooth representations of taste heterogeneity. Future work could thus explore ways to incorporate semi-nonparametric heterogeneity representations into the proposed model framework by allowing for parametric heterogeneity within the mixture components burda_bayesian_2008,li_bayesian_2013. A second limitation of the DPM-MNL model concerns the nature of the Dirichlet process. In any Dirichlet process mixture model, the component weights are constructed from independent and identically distributed Beta random variables. Other stochastic processes such as the Pitman-Yor process pitman_two-parameter_1997 are not subject to this restriction but also exhibit desirable clustering and discreteness properties. Future research may thus explore the use of alternate stochastic processes, which induce random partitions, to accommodate unobserved heterogeneity in discrete choice models. A third limitation of the DPM-MNL model concerns its dependency on the specification of a parametric base measure. In general, the choice of the base measure in Dirichlet process mixture models is opportunistic in that it is driven by computational considerations rather than by behavioural principles gorur_dirichlet_2010,mcauliffe_nonparametric_2006. Future research may therefore investigate methods that preclude the a priori specification of a parametric base measure in a DPM-MNL model. For example, mcauliffe_nonparametric_2006 present an approach to infer the base measure of the Dirichlet process nonparametrically in an empirical Bayes fashion.
In addition, there are several broader directions for future research to build on the work presented in the current paper: First, it may be worthwhile to apply the proposed model framework to other datasets to assess whether the findings presented in the current paper generalise to other empirical contexts. Second, future research could investigate methods to combine systematic and random representations of heterogeneity bhat_endogenous_1997,bhat_incorporating_2000 within a DPM-MNL model. To that end, covariate-dependent stick-breaking processes ren_logistic_2011,rodriguez_nonparametric_2011 could be leveraged. Ultimately, future research may develop comprehensive simulation-based and empirical comparisons of different methods to accommodate unobserved taste heterogeneity in discrete choice models fosgerau_comparison_2009,keane_comparing_2013. In particular, it may be worthwhile to contrast the performance of the proposed DPM-MNL model with other nonparametric vij_random_2017, semi-nonparametric bastin_estimating_2010,fosgerau_practical_2007,train_em_2008,train_mixed_2016 and parametric bhat_new_2012,bhat_new_2017,train_mixed_2005 approaches to account for unobserved heterogeneity in discrete choice models.
Our proposed DPM-MNL model enriches the methodological toolbox of discrete choice analysts by allowing for a flexible, adaptive and parsimonious nonparametric representation of heterogeneity distributions in conjunction with an elegant inference approach. However, there is no one way of catering for taste heterogeneity in discrete choice models, and in any given empirical application, the preferred way of representing taste heterogeneity may depend on a variety of factors such as considerations of parsimony and tractability, emphasis on data fit or predictive performance as well as the analyst's prior belief about the nature of the latent taste parameter distribution.