EconBase
← Back to paper

A Multinomial Probit Model for Asymmetric Choice Responses

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

51,687 characters

A Multinomial Probit Model for Asymmetric Choice Responses




\def\spacingset#1{\renewcommand{\baselinestretch}
{#1}\small\normalsize} \spacingset{1}




\if11
{
  \title{\bf A Multinomial Probit Model\\ for\\ Asymmetric Choice Responses}
  \author{Cash Looi\thanks{
    Correspondence to: Department of Econometrics \& Business Statistics, Monash University, Clayton VIC 3800, Australia, e-mail: \textsf{[email removed]}}\hspace{.2cm}\\
    Rub\'en Loaiza-Maya \\
    Didier Nibbering\\
    Department of Econometrics and Business Statistics, Monash University}
  \maketitle
} \fi

\if01
{
  \bigskip
  \bigskip
  \bigskip
  \begin{center}
    {\LARGE\bf Title}
\end{center}
  \medskip
} \fi

\bigskip
\begin{abstract}
Standard multinomial probit (MNP) models specify symmetric latent utility distributions, implying that choice probabilities respond symmetrically to positive and negative covariate shifts of the same magnitude. This restriction is often implausible in empirical choice settings and can lead to misleading elasticity and substitution predictions.
We propose a skewed multinomial probit (SMNP) model that captures asymmetric choice responses by specifying a multivariate skew-normal distribution for the latent utilities. The model preserves the flexible substitution patterns of the MNP framework, introduces alternative-specific skewness parameters, and nests the standard MNP model when skewness is zero.
Introducing skewness creates identification and computational challenges because the skewness parameters interact with the MNP scale normalization and disrupt the conditional Gaussian updating structure used in Bayesian MNP estimation. We address these challenges through a covariance reparameterization that enforces identification and positive definiteness by construction, interpretable priors on the identified parameter space, and a double data-augmentation scheme that yields a Metropolis-Hastings within Gibbs sampler. Numerical experiments and applications to consumer choice data show that SMNP recovers asymmetric choice responses, improves probabilistic prediction, and produces economically meaningful differences in price elasticities and substitution patterns.
\end{abstract}

\noindent
{\it Keywords:}
Skew-normal distribution; discrete choice; data augmentation; price elasticities; substitution patterns
\vfill

\newpage
\spacingset{1.8}


\section{Introduction}
Discrete choice models often impose symmetric choice responses: they imply that a positive change in an attribute of a choice alternative affects choice probabilities by the same magnitude as an equivalent negative change. This restriction arises when a linear latent utility function is combined with a symmetric error distribution. However, in many decision-making settings, this symmetry is difficult to justify. People may react differently to increases and decreases in the same attribute, so the probability of choosing an alternative may change more strongly in one direction than in the other.

The idea of asymmetric choice responses has a long foundation in behavioral economics. Prospect Theory argues that agents respond more strongly to perceived losses than to perceived gains \citep{kahnemann1979prospect}. This has motivated a large empirical literature on asymmetric responses to price changes, especially in economics and marketing. Recent studies document loss aversion in settings ranging from child health care \citep{iizuka2023asymmetric} and soft drinks \citep{biondi2020between} to ketchup and milk \citep{caputo2018choice}. Capturing such asymmetric choice responses matters for governments that evaluate welfare policies and for firms that design pricing strategies.

This paper proposes a skewed multinomial probit (SMNP) model that allows for asymmetric choice responses. The model builds on the multinomial probit (MNP) framework, which is well suited to choice applications because it allows the latent utilities of different alternatives to be correlated. These correlations capture general substitution patterns among alternatives and avoid the independence of irrelevant alternatives restriction imposed by other choice models \citep{hausman1978conditional}. However, the standard MNP model still imposes symmetric choice responses by specifying a multivariate normal distribution for the latent utilities conditional on observed covariates. We relax this restriction by replacing the multivariate normal distribution with a multivariate skew-normal distribution. The resulting SMNP model introduces alternative-specific skewness parameters, allowing both the degree and direction of asymmetry to vary across alternatives in the choice set.

Introducing skewness into the MNP framework creates new identification challenges, which we address with a novel reparameterization of the implied covariance matrix. MNP models already require a scale normalization in the covariance matrix because the observed choices are invariant to rescalings of the latent utilities. In the SMNP model, this normalization interacts with the skewness parameters because the covariance structure depends on both the scale and the skewness parameters in the multivariate skew-normal distribution. As a result, one cannot freely estimate the skewness parameters and covariance matrix. We solve this problem by developing a reparameterization that fixes the scale and ensures a positive definite covariance matrix.

We address the computational challenges in the estimation of the SMNP model by developing Bayesian inference based on data augmentation and Markov chain Monte Carlo (MCMC) sampling. In the MNP model, likelihood evaluation is costly because choice probabilities involve multivariate integrals. Bayesian estimation reduces this burden by augmenting the model with latent utilities and sampling them inside a MCMC scheme \citep{albert1993bayesian}. This strategy relies on the conditional normality of the latent utilities and therefore does not directly transfer to SMNP. We address this problem by introducing a second layer of augmentation based on the conditionally Gaussian representation of the multivariate skew-normal distribution \citep{fruhwirth2010bayesian}. The double data augmentation yields closed-form updates for all parameters but one, which we update with a single Metropolis-Hastings step.

This paper contributes to a large literature that captures asymmetric choice responses by modifying the observable component of utility in discrete choice models. Typically, researchers split a focal attribute, such as price, into gains and losses relative to a reference point \citep{kalwani1990price, krishnamurthi1992asymmetric, hardie1993modeling, briesch1997comparative,mazumdar2005reference, caputo2018choice, biondi2020between}. The model then detects asymmetry when the coefficient on losses differs from the coefficient on gains. \citet{bhat2012new} introduce asymmetry through random coefficients with a skew-normal distribution. Alternatively, asymmetric responses can be generated by using more flexible predictor functions, such as nonlinear transformations, polynomial terms, or splines. Instead of imposing asymmetry through observed attributes, we allow asymmetry to arise from the conditional distribution of the latent utilities. This parsimonious way of capturing asymmetric choice responses does not require reference points, random coefficients, or predictor transformations.

A second strand of literature introduces asymmetry through the error distribution or link function, primarily for binary response data: \citet{chen1999new, bazan2010framework,kim2001bayesian,kim2002binary,kim2008flexible,zhang2023tractable}. However, choice applications often involve more than two alternatives, and extending to multinomial outcomes creates additional identification and computational challenges. Recent work has derived Bayesian conjugacy properties for a class of skewed MNP models assuming a fixed scale matrix \citep{fasano2022class, anceschi2023bayesian,karling2024conjugacy}. Conditioning on a fixed scale matrix is restrictive in applications, where the dependence across latent utilities is unobserved and often central for capturing substitution patterns. Moreover, these results do not deliver joint inference over all model parameters, nor do they address the identification problems that arise when skewness is introduced. We fill this gap by developing a fully identified SMNP model and an empirically implementable Bayesian inference method.

We illustrate the practical relevance of our method with two numerical experiments and two empirical applications. The numerical experiments compare the SMNP and MNP models under data generating processes with asymmetric and symmetric choice responses. We find that SMNP recovers asymmetric choice responses and improves predictive accuracy relative to the MNP. When choice responses are symmetric, SMNP performs comparably to the MNP model. The empirical applications to laundry detergent and ketchup purchases show that asymmetries matter in practice. The SMNP detects asymmetry in both applications, these asymmetries change the implied own- and cross-price elasticities, and improves prediction. The MNP understates price sensitivity for detergent brands, and misallocates substitution across ketchup brands. These results indicate that the SMNP model can help practitioners make better pricing and substitution decisions.

The remainder of this article is structured as follows. Section \ref{skewed} introduces the SMNP model specification, while Section \ref{bayes} covers aspects of Bayesian estimation pertaining to data augmentation, prior specification, and the proposed MCMC sampler. Section \ref{sim} evaluates the proposed model through the numerical simulation studies. Section \ref{empirical} then applies the proposed method to real consumer datasets, followed by a discussion in Section \ref{discuss}.

\section{Skewed multinomial probit model}\label{skewed}
\subsection{Model specification}\label{model}
Let $Y_i\in \{1,2,\dots,J+1\}$ be an observed discrete choice with $J+1$ the number of choice alternatives, and $i=1,2,\dots,N$ with $N$ the number of individuals. It is assumed that each $Y_i$ is associated with a $J\times1$ vector of latent utilities $Z_{i}=(z_{i1}, z_{i2},\dots, z_{iJ})^\top$:
\begin{equation}\label{eq:Y}
Y_i=
\begin{cases}
J+1,\quad  &\text{if }\ \max(Z_i)<0 ,\\
j,\quad  &\text{if }\ z_{ij}=\max(Z_i)>0,\\
\end{cases}
\end{equation}
where $\max(Z_i)$ denotes the largest value among the elements in the vector $Z_i$. Category $J+1$ is referred to as the base category. The latent utilities are modeled as
\begin{equation}\label{eq:Zmsn}
   Z_i=X_i\beta+\varepsilon_i, \quad \quad \varepsilon_i \sim SN_J(0,{\Sigma},{\alpha}),
\end{equation}
where $X_i$ represents the $J \times k$ regressor matrix and $\beta$ is a $k \times 1$ vector of coefficients. The $J\times 1$ disturbance vector $\varepsilon_i$ follows a $J-$variate skew normal distribution defined in \cite{azzalini1996multivariate} and \cite{azzalini1999statistical}, with $\Sigma$ a $J\times J$ positive definite scale matrix and ${\alpha} = (\alpha_1, \dots, \alpha_J)^\top \in \mathbb{R}^J$ a vector of shape parameters governing asymmetry.

The model specified in \eqref{eq:Y} and \eqref{eq:Zmsn} considers only $J$ utilities for each choice $Y_i$, while there are $J+1$ choice alternatives. This reduction to a $J$-dimensional space is to address the location identification problem that follows if a unique utility is specified for each choice alternative \citep{bunch1991estimability}. In addition, we have to address the scale identification problem: $Y_i(cZ_i) = Y_i(Z_i)$ for any $c \in \mathbb{R}^+$. We follow the common approach by fixing the first diagonal element of $\Sigma$ to one.

The regressor matrix typically includes $J$ alternative-specific intercepts, a $k_d$-dimensional vector $x_{i,d}$ of individual-specific covariates, and a $(J+1)\times k_a$ matrix $x_{i,a}$ of $k_a$ alternative-specific covariates. Therefore, $k=J + Jk_d+k_a$ and $X_i$ can be written as
\begin{equation}\label{eq:X}
   X_i = \left[I_J \quad x_{i,d}^\top \otimes I_J \quad Tx_{i,a}    \right],
\end{equation}
where $I_J$ denotes the $J\times J$ identity matrix, $ T = \bigl[\, I_J \;\; -\mathbf{1}_J \,\bigr]$,  and $\mathbf{1}_J$ is a $J \times 1$ vector of ones. Specifically, the transformation matrix $T$ constructs the differences in the alternative-specific covariates between each alternative and the $(J+1)$-th alternative.


\subsection{Asymmetric choice responses}
Since the multivariate skew normal distribution is closed under affine transformations \citep{Azzalini_2013}, the latent utilities satisfy $
Z_i \sim SN_J(X_i\beta,{\Sigma},{\alpha})$. When the shape parameter vector is a $J-$dimensional vector of zeros, such that $\alpha=0_J$, the distribution reduces to a $J$-variate normal distribution with mean $X_i\beta$ and covariance matrix $\Sigma$. Thus, the SMNP model encompasses the multinomial probit model (MNP) as a special case.

The shape vector $\alpha$ governs asymmetry in the multivariate distribution of the latent utilities, but its components may be difficult to interpret. We therefore use the skewness-parameter vector $\delta = (1+\alpha^\top\Sigma\alpha)^{-1/2}\Sigma\alpha$, whose elements directly determine the direction and magnitude of skewness in the marginal latent utilities. Throughout the paper, we refer to $\delta$ as the skewness parameter vector. Since $\Sigma$ is positive definite, $\alpha=0_J$ if and only if $\delta=0_J$.

Skewness in the conditional distribution of the latent utilities changes how choice probabilities respond to covariate shifts. Under the symmetric MNP model, positive and negative changes in a covariate of the same magnitude generate symmetric probability responses around a given reference point. The SMNP model relaxes this restriction. It allows the probability of choosing an alternative to react more strongly to a change in one direction than to an equally sized change in the opposite direction.

\noindent\textbf{Example:} Figure~\ref{fig:cpplot} illustrates this mechanism in a three-alternative choice setting. We compare three specifications that have the same conditional mean and covariance for the latent utilities but differ in skewness: a negatively skewed SMNP (dashed purple line), a symmetric MNP (solid black line), and a positively skewed SMNP (dashed orange) specification. Panel (a) shows that the skewness parameter changes the shape of the choice probability curve. Panel (b) makes the asymmetry explicit by plotting the absolute deviation of the choice probability from 0.5 around the corresponding reference value of the covariate.

\begin{figure}[tb!]
\caption{Asymmetric latent utilities and choice probability responses}
\centering
\includegraphics[width=\textwidth]{fig/motivationplots.eps}
\captionsetup{font=footnotesize}\caption*{This figure shows the impact of asymmetric latent utilities on choice probabilities. Panel (a) shows the choice probability for category $2$ as a function of its alternative-specific covariate $x_2$ for negatively skewed ($\delta=(0,-1.437)^\top$, dashed purple), symmetric ($\delta=(0,0)^\top$, solid black), and positively skewed ($\delta=(0,1.437)^\top$, dashed orange) latent utilities. Panel (b) shows the absolute deviation of the choice probabilities from the reference probability $0.5$ as a function of the deviations in $x_2$. Probabilities are simulated using \eqref{eq:Y} and \eqref{eq:Zmsn} with $\beta = (0,0,0.7)^\top$, $\Sigma=[1\,0.5;0.5\,1]$ in the MNP, and $x_2$ a standard normal distribution.}
\label{fig:cpplot}
\end{figure}

Panel (b) shows that, under the symmetric MNP model, equal positive and negative deviations in the covariate produce equal absolute changes in choice probability. This may be unrealistic in many empirical choice settings. For instance, positive and negative price changes may have different effects on purchase probabilities. In the SMNP model, these deviations produce different changes, generating asymmetric choice responses. For instance, with negatively (positively) skewed latent utilities, negative deviations in the covariate result in smaller (larger) changes in choice probability relative to changes as a result of positive deviations in the covariate. These asymmetries may reflect loss aversion, brand loyalty, or switching costs in consumer choice settings.


\section{Bayesian inference for the identified SMNP model} \label{bayes}
Estimating the SMNP model requires addressing two challenges. The first is computational. To avoid evaluation of the multivariate integrals in the likelihood function, Bayesian estimation augments the MNP model with latent utilities. In the SMNP model, however, augmenting only with the latent utilities does not deliver the same conjugate Gaussian updating structure. The second challenge is identification. As discussed in Section~\ref{model}, MNP models require a scale normalization. In the SMNP model, this normalization interacts with the skewness parameters. Hence, the skewness vector and the scale matrix cannot be sampled independently.

We address both challenges jointly. We first introduce a second layer of augmentation that gives the skew-normal distribution a conditionally Gaussian representation. We then impose the MNP scale normalization in this augmented parameterization. This yields an identified model and a tractable Metropolis-Hastings within Gibbs sampler.

\subsection{Double data augmentation}
Let $Z=(Z_1^\top,\dots,Z_N^\top)^\top$ and $X=(X_1^\top,\dots,X_N^\top)^\top$, and let $\theta$ denote the parameter vector. If we augment the observed choices only with the latent utilities, the posterior is
\begin{equation}\label{eq:post1}
p(\theta, Z\mid Y,X) \propto  \prod^N_{i=1} p(Y_i\mid Z_i)p(Z_i\mid X_i, \theta)p(\theta),
\end{equation}
where $ p(Y_i\mid Z_i)$ is implied by \eqref{eq:Y} and $p(\theta)$ denotes the prior distribution \citep{albert1993bayesian}. For the standard MNP model, this representation leads to conditionally normal latent utilities, and Gibbs sampling is directly applicable to the posterior in \eqref{eq:post1}. For the SMNP model, $p(Z_i\mid X_i,\theta)$ is skew-normal, so this augmentation alone does not admit conjugate conditional posteriors for the latent utilities $Z_i$.

We obtain a conditional normal representation in the skew-normal utilities by using the stochastic representation of the multivariate skew-normal distribution, see \citet{fruhwirth2010bayesian}. In particular, we write
 \begin{equation}\label{eq:equil}
\begin{split}
p(Z_i\mid X_i, \theta)
& = 2\phi_J(Z_i;X_i\beta; {\Sigma})\Phi({\alpha}^\top({Z_i}-X_i\beta))\\
&= \int_{0}^\infty \phi_J(Z_i;X_i\beta+\delta w_i,\Sigma_u)p(w_i) dw_i,
\end{split}
\end{equation}
where $\phi_{J}\left(Z;\mu,\Sigma\right)$ denotes the $J$-variate normal density with mean $\mu$ and covariance matrix $\Sigma$, $\Phi(\cdot)$ the univariate standard normal cumulative distribution function, $p(w_i) = 2\phi_1(w_i;0,1)\mathbb{I}(w_i>0)$ the univariate truncated standard normal density, and $\Sigma_u = \Sigma-\delta\delta^\top$. From here onward, we parameterize the model in terms of $\delta$ and $\Sigma_u$,  noting that $\alpha= (1-\delta^\top\Sigma^{-1}\delta)^{-1/2}\Sigma^{-1}\delta$ and $\Sigma = \Sigma_u+\delta\delta^\top$.

The representation in \eqref{eq:equil} allows for a second layer of data augmentation using $w_i$. The resulting doubly augmented posterior becomes
\begin{equation}\label{eq:augpost}
p(\theta, Z, w \mid Y,X) \propto \prod^N_{i=1}p(Y_i\mid Z_i)\phi_J(Z_i;X_i\beta+\delta w_i,\Sigma_u)p(w_i)p(\theta),
\end{equation}
where $w=(w_1,\dots,w_N)^\top$. Conditional on $w$, the latent utilities are Gaussian and a conjugate conditional posterior for $Z_i$ is available. Hence, this second augmentation step turns the skew-normal latent utility model into a conditionally Gaussian regression model, which yields closed-form conditional updates for most model components.

\subsection{Scale identification in the doubly augmented parameterization} \label{identification}
The double augmentation introduces the covariance matrix $\Sigma_u$, but the scale normalization $\Sigma_{11}=1$ discussed in Section~\ref{model} applies to the skew-normal scale matrix $\Sigma$. Since $\Sigma_u= \Sigma-\delta\delta^\top$, the scale restriction imposes that the first element of $\Sigma_u$ is fixed at $1-\delta^2_1$. This restriction shows why the covariance matrix and skewness vector cannot be treated as independent parameters. Sampling an unrestricted positive definite matrix for $\Sigma_u$ would generally violate the scale normalization of the SMNP model.

We impose the restriction directly by parameterizing $\Sigma_u$ as
\begin{equation}\label{eq:sig_u}
\Sigma_u=
\begin{bmatrix}
1-\delta_1^2 & \gamma^\top\\
\gamma &  \psi+ \dfrac{\gamma\gamma^\top}{1-\delta_1^2}
\end{bmatrix},
\end{equation}
where $\gamma$ is a $(J-1)$-dimensional vector and $\psi$ is a $(J-1)$-dimensional positive definite square matrix. If there is no skewness and $\delta_1=0$, this representation coincides with the parametrization of the covariance matrix in the MNP model of \cite{mcculloch2000bayesian}.

The parameterization in \eqref{eq:sig_u} is key for identification in the SMNP model which allows for $\delta_1\neq 0$. If $|\delta_1|<1$ and $\psi$ is positive definite, the Schur complement of $(1-\delta_1^2)$ in $\Sigma_u$ is exactly $\psi$, so $\Sigma_u$ is positive definite. Moreover, the implied scale matrix satisfies $\Sigma_{11}=(1-\delta_1^2)+\delta_1^2=1$. Thus, every posterior draw generated under this parameterization automatically satisfies the scale normalization and remains inside the positive definite parameter space.
The identified parameter vector is now $\theta = (\beta^\top,\delta^\top,\gamma^\top,\text{vec}(\psi)^\top)^\top$. Given draws of $\delta_1$, $\gamma$, and $\psi$, we construct $\Sigma_u$ using \eqref{eq:sig_u} and recover $\Sigma=\Sigma_u+\delta\delta^\top$. This implies that $\delta_1$ enters all expressions of the augmented likelihood that depend directly on $\Sigma_u$. Moreover, $\delta_1$ and $\psi$ require priors that ensure $|\delta_1|<1$ and the strictly positive definiteness of $\psi$.

\subsection{Prior specification under the identified parameterization}\label{prior}
We place priors directly on the identified parameter vector $\theta = (\beta^\top,\delta^\top,\gamma^\top,\text{vec}(\psi)^\top)^\top$. We use a multivariate normal prior on the model coefficients, $\beta \sim \mathcal{N}_k(0_k,10I_k)$. For the asymmetry parameter $\delta$, we use a multivariate normal prior with a truncation on $\delta_1$: $\delta \sim \mathcal{N}_J(0_J,\tau_\delta I_J)\,\mathbb{I}(|\delta_1|<c)$, where $c<1$ is a numerical bound that enforces the scale-identification constraint. We set $c=0.995$. Centering this prior at zero is important because $\delta=0_J$ corresponds to the standard MNP model. Hence, the prior does not force asymmetry into the latent utilities. The hyperparameter $\tau_\delta$ controls how much prior mass is assigned to asymmetric latent utility distributions.

For the covariance parameters, we specify $\gamma \sim \mathcal{N}_{J-1}(B_\gamma,\tau_\gamma I_{J-1})$ and $\psi \sim \mathcal{IW}_{J-1}(J+3,V)$. Because $\Sigma=\Sigma_u+\delta\delta^\top$, the prior on $\delta$ affects the implied prior distribution of the scale matrix $\Sigma$. We therefore choose the hyperparameters of the priors for $\gamma$ and $\psi$ jointly with the prior for $\delta$, so that the implied prior mean of $\Sigma$ is the equicorrelated matrix $\mathbb{E}(\Sigma) = \frac{1}{2}\,1_J1_J^\top + \frac{1}{2}\,I_J$ proposed in \cite{geweke1994alternative}. This construction centers the model on a familiar MNP covariance structure while allowing departures from symmetry through $\delta$.

The prior specification balances two goals. First, it preserves the standard MNP model as a meaningful special case by placing prior mass around $\delta=0_J$. Second, it leaves enough prior variation in $\delta$ to allow the data to reveal skewness when asymmetric choice responses are present. Values of $\tau_\delta$ that are too small shrink the model toward the symmetric MNP specification, while values that are too large can force dependence in the latent utilities to be captured through skewness rather than through the covariance structure. Appendix~\ref{prioreliccit} derives the hyperparameter restrictions that ensure the implied prior for $\Sigma_u$ remains positive definite and gives the resulting choices of $\tau_\delta$, $B_\gamma$, $\tau_\gamma$, and $V$.

\subsection{Posterior simulation}
The double augmentation and identified parameterization yield a tractable Metropolis-Hastings within Gibbs sampler. Conditional on $Z$ and $w$, the model is Gaussian, so all parameters except $\delta_1$ have standard full conditional distributions. This parameter enters both the skewness term  and the normalized covariance matrix $\Sigma_u$, because $(\Sigma_u)_{11}=1-\delta_1^2$. As a result, its full conditional distribution is not available in closed form.

We update $\delta_1$ using a univariate Metropolis-Hastings step. To respect the constraint $|\delta_1|<c$, we transform $\delta_1$ to the real line and construct a proposal distribution from a Laplace approximation to the transformed conditional posterior. All other parameters are updated using Gibbs steps. The sampler therefore isolates the nonconjugacy created by scale identification under skewness in a single one-dimensional Metropolis-Hastings update.

Starting from initial values $\{\delta_{1:J}^{(0)},w^{(0)},\gamma^{(0)},\psi^{(0)},Z^{(0)},\beta^{(0)}\}$, we iterate the following steps for $r=1,2,\dots,R$:
\begin{description}
    \item[\textbf{Step 1:}] Generate $\delta_1^{(r)} \sim p(\delta_1 \mid\delta_{1:J}^{(r-1)}, w^{(r-1)},\gamma^{(r-1)},\psi^{(r-1)}, Z^{(r-1)}, \beta^{(r-1)},Y,X)$.
    \item[\textbf{Step 2:}] Generate $\delta_{2:J}^{(r)} \sim p(\delta_{2:J} \mid\delta_{1}^{(r)}, w^{(r-1)},\gamma^{(r-1)},\psi^{(r-1)}, Z^{(r-1)}, \beta^{(r-1)},Y,X)$.
    \item[\textbf{Step 3:}] Generate $w^{(r)} \sim p(w \mid \delta^{(r)},\gamma^{(r-1)},\psi^{(r-1)}, Z^{(r-1)},\beta^{(r-1)},Y,X)$.
    \item[\textbf{Step 4:}] Generate $\gamma^{(r)} \sim p(\gamma \mid \delta^{(r)}, w^{(r)},\psi^{(r-1)}, Z^{(r-1)},\beta^{(r-1)},Y,X)$.
    \item[\textbf{Step 5:}] Generate $\psi^{(r)} \sim p(\psi \mid \delta^{(r)}, w^{(r)},\gamma^{(r)}, Z^{(r-1)},\beta^{(r-1)},Y,X)$.
    \item[\textbf{Step 6:}] Generate $Z^{(r)} \sim p(Z \mid \delta^{(r)}, w^{(r)},\gamma^{(r)},\psi^{(r)},\beta^{(r-1)},Y,X)$.
    \item[\textbf{Step 7:}] Generate $\beta^{(r)} \sim p(\beta \mid \delta^{(r)}, w^{(r)},\gamma^{(r)},\psi^{(r)},Z^{(r)},Y,X)$.
\end{description}

At each iteration, $\Sigma_u^{(r)}$ and $\Sigma^{(r)}$ are deterministic functions of $\delta^{(r)}$, $\gamma^{(r)}$, and $\psi^{(r)}$. The sampler therefore produces posterior draws that satisfy both positive definiteness and the scale normalization. Appendix~\ref{MCMC} provides the full conditional distributions and further implementation details.


\subsection{Predictive distribution and scoring rules}
In discrete choice models, the model parameters are generally not directly interpretable, and interest usually focuses on the predictive distribution of $Y_i$:
\begin{equation}
Pr(Y_i =j\mid X_i,Y,X)= \int \int p(Y_i =j\mid Z_i)p(Z_i\mid X_i, \theta)p(\theta \mid Y,X) \ dZ_i \ d\theta.
\end{equation}
Note that the parameter vector $\theta$ for SMNP and MNP represent different sets of model parameters. As this integral is not available in closed form, we evaluate it using MCMC draws. For each retained posterior draw $\theta^{(m)}$ with $m=1,\dots, M$, we generate $Z_i^{(m)} \sim p(Z_i\mid X_i, \theta^{(m)})$, and subsequently  $Y_i^{(m)} \sim  p(Y_i\mid Z_i^{(m)})$. The predictive probability is evaluated as
\begin{equation}
\hat{Pr}(Y_i =j\mid X_i,Y,X)=  \frac{1}{M} \sum^M_{m=1} \mathbb{I}(Y_i^{(m)}=j),
\end{equation}
and the point prediction $\hat Y_i$ equals the mode of the predictive distribution.

We evaluate predictive performance using the hit rate, and the proper scoring rules such as the log score and the censored likelihood score. The hit-rate evaluates point predictions:
\begin{equation}
HR=  \frac{1}{N} \sum^N_{i=1} \mathbb{I} \left( \hat Y_i =Y_i \right).
\end{equation}
The log score evaluates the full predictive distribution:
\begin{equation}
LS=  \frac{1}{N} \sum^N_{i=1} \sum^{J+1}_{j=1} \ln(\hat{Pr}(Y_i = j \mid X_i,Y,X)) \mathbb{I}(Y_i = j).
\end{equation}
Finally, the censored likelihood score of \citet{diks2011likelihood} measures predictive accuracy for a specific category $j$: $CLS_j = $
\begin{equation}
\frac{1}{N} \sum_{i=1}^N \left[ \ln(\hat{Pr}(Y_i = j\mid X_i,Y,X)) \mathbb{I}(Y_i = j) + \ln(\hat{Pr}(Y_i \neq j \mid X_i,Y,X)) \mathbb{I}(Y_i \neq j) \right].
\end{equation}
For all three accuracy metrics, larger values indicate better predictive performance.

\section{Numerical experiments} \label{sim}
We conduct two numerical experiments to evaluate the inferential and predictive accuracy of the SMNP relative to the MNP. The first experiment generates data from a skewed latent utility distribution, and the second experiment from a normal latent utility distribution. This allows us to assess whether the SMNP can recover asymmetric choice responses and whether the SMNP collapses to the MNP  when the skewness parameters are not needed.

\subsection{Design}\label{sec:design}

We consider a three alternative choice model with $N=20,000$ observations. We generate two datasets. The first dataset is generated from an SMNP model with skewed latent utilities, using $(\delta_0=(0.8,1.3)^\top)$. The second dataset is generated from the corresponding MNP model by setting $(\delta_0=(0,0)^\top)$. All other components of the data-generating process are kept fixed across the two experiments.

The model includes two alternative-specific intercepts and one alternative-specific regressor. We generate $x_{i,a}=(x_{i,1},x_{i,2},x_{i,3})^\top$, where $x_{i,j}\sim \text{Log Normal}(0,1)$ independently across observations and alternatives. We then use $\log x_{i,a}$ as the covariate, so that the regressor can be interpreted as log price. The true coefficient vector is $\beta_0=(-0.5,-0.8,-0.3)^\top$, where the first two elements are alternative-specific intercepts, and the third element is the price coefficient. The diagonal elements of the true scale matrix $\Sigma_0$ are $1$ and $1.79$, and the off-diagonal equals $1.04$. Given these parameters, we generate data from \eqref{eq:Y} and \eqref{eq:Zmsn}.

For each dataset, we randomly assign half of the observations to the training sample and half to the test sample. We estimate both the SMNP and MNP models on the training sample. For the SMNP model, we use the prior specification in Section~\ref{prior}. For the MNP model, we use the same prior for $\beta$ and the covariance prior of \citet{mcculloch2000bayesian}, centered on an equicorrelated matrix with off-diagonal elements equal to 0.5. We run the SMNP and the MNP samplers for 1,000,000 iterations, discarding the first half of each chain as burn-in. For the construction of predictive probabilities and scoring rules, we thin the retained draws to 10,000 posterior draws.

\subsection{Results with asymmetric choice responses}
We first examine the experiment in which the data are generated from the SMNP model. Figure~\ref{fig:idendeltaMSN} shows the posterior densities of the two skewness parameters in the SMNP model. The true parameter values, indicated by the vertical lines, are covered with high posterior probability mass, while the posteriors do not allocate any density mass at zero. This shows that the SMNP model captures the asymmetry in the latent utility distribution.

\begin{figure}[tb!]
\centering
\includegraphics[width=\textwidth]{fig/deltapost_MSNDGP.eps}
\caption{Posterior densities $\delta_1$ and $\delta_2$ in the SMNP model for the experiment with skewed latent utilities. The vertical lines represent true parameter values.}
\label{fig:idendeltaMSN}
\end{figure}

Next, we compare the posterior choice probabilities from the SMNP and MNP models with the true choice probabilities implied by the data-generating process. Figure~\ref{fig:choiceprob} shows choice probabilities for an increasing price level for a specific alternative, with the prices of the other alternatives held constant at their mean. The choice probabilities from SMNP, MNP, and data generating process (DGP) are represented by the solid yellow, dotted black, and dashed blue lines, respectively. The SMNP model closely tracks the true probability curves. The MNP model, in contrast, displays systematic deviations, especially in the tails of the price range. For instance, Panel (a) demonstrates that the MNP model underestimates the true choice probabilities for prices below $0.2$. In Panel (b), the MNP model exhibits bias at both ends of the price level. These differences show that the symmetric latent utility specification distorts the implied substitution patterns when the true latent utilities are skewed.

\begin{figure}[tb!]
\centering
\includegraphics[width=\textwidth]{fig/MSNsim_cp.eps}
\caption{Posterior choice probabilities for the experiment with skewed latent utilities.
The probability of choice $2$ and choice $3$ are a function of their price, with the prices of the other choices fixed at their mean. The solid yellow and dotted black lines represent the estimated choice probabilities using the SMNP and MNP models, respectively. The dashed blue line indicates the true choice probability.}
\label{fig:choiceprob}
\end{figure}

We examine how these discrepancies in choice probabilities translate into differences in predictive accuracy across the models. Panel A of Table~\ref{tab:MSNpredictive} reports the predictive performance under the asymmetric DGP, comparing the MNP and SMNP models against predictions from the true DGP, hereafter referred to as the oracle. In-sample, the SMNP model attains a higher hit rate, and performs better on the log score and all censored likelihood scores. The log-score and $CLS_3$ of SMNP are close to these metrics for the oracle, whereas the MNP model produces statistically significantly lower values according to the Giacomini--White test. Out-of-sample, the SMNP model also improves both point and probabilistic prediction. The SMNP log score and censored likelihood scores are also close to the oracle values, whereas the MNP model differs significantly from the oracle for the log score and $CLS_3$. These results show that the SMNP model improves probabilistic prediction when asymmetric choice responses are present.

\begin{table}[tb!]
\centering
\caption{Predictive performance numerical experiments} \label{tab:MSNpredictive}
\begin{tabular}{lrrrrrr}
\toprule
& \multicolumn{3}{c}{In-sample} & \multicolumn{3}{c}{Out-of-sample} \\
\cmidrule(lr){2-4} \cmidrule(lr){5-7}
Metric & MNP & SMNP & Oracle & MNP & SMNP & Oracle \\
\midrule
\multicolumn{7}{l}{\textit{Panel A: Asymmetric choice responses}} \\
\midrule
$HR$ & $0.5470$ & $\mathbf{0.5483}$ & $0.5469$ & $0.5480$ & $\mathbf{0.5499}$ & $0.5492$ \\
$LS$ & $-0.9201\rlap{$^{*}$}$ & $\mathbf{-0.9166}$ & $-0.9166$ & $-0.9216\rlap{$^{*}$}$ & $\mathbf{-0.9192}$ & $-0.9192$ \\
$CLS_{1}$ & $-0.5707$ & $\mathbf{-0.5700}$ & $-0.5703$ & $-0.5704$ & $\mathbf{-0.5700}$ & $-0.5701$ \\
$CLS_{2}$ & $-0.5984$ & $\mathbf{-0.5978}$ & $-0.5977$ & $-0.5981$ & $\mathbf{-0.5973}$ & $-0.5974$ \\
$CLS_{3}$ & $-0.4725\rlap{$^{*}$}$ & $\mathbf{-0.4693}$ & $-0.4693$ & $-0.4750\rlap{$^{*}$}$ & $\mathbf{-0.4727}$ & $-0.4726$ \\
\midrule \multicolumn{7}{l}{\textit{Panel B: Symmetric choice responses}} \\
\midrule
$HR$ & $\mathbf{0.6095}$ & $0.6078$ & $0.6083$ & $\mathbf{0.6071}$ & $0.6059$ & $0.6072$ \\
$LS$ & $-0.8824$ & $\mathbf{-0.8821}$ & $-0.8824$ & $\mathbf{-0.8837}$ & $-0.8838$ & $-0.8835$ \\
$CLS_{1}$ & $-0.4696$ & $\mathbf{-0.4696}$ & $-0.4696$ & $\mathbf{-0.4608}$ & $-0.4608$ & $-0.4607$ \\
$CLS_{2}$ & $-0.4621$ & $\mathbf{-0.4619}$ & $-0.4622$ & $\mathbf{-0.4724}$ & $-0.4724$ & $-0.4722$ \\
$CLS_{3}$ & $-0.6425$ & $\mathbf{-0.6423}$ & $-0.6425$ & $\mathbf{-0.6420}$ & $-0.6421$ & $-0.6421$ \\
\bottomrule
\end{tabular}\\
\vspace{1ex}
\raggedright
\footnotesize{\textit{Notes:} Bolded values indicate the superior predictive performance between the SMNP and MNP specifications within each DGP and sample split. The symbol $^{*}$ indicates that the difference between the model and the oracle is statistically significant at the $5\%$ significance level according to a \cite{giacomini2006tests} test on the differences in proper scoring rules.}
\end{table}


\subsection{Results with symmetric choice responses}
Second, we examine the experiment in which the data are generated from the MNP model. Figure~\ref{fig:idendeltaMVN} shows the posterior densities of the skewness parameters in the SMNP model. The posterior mass is centered around zero, indicating that the model recovers the symmetric latent utility structure. This also follows from the posterior choice probabilities reported in Appendix~\ref{A:numerical}, which are nearly identical between the MNP and SMNP models. In terms of predictive performance, Panel B of Table~\ref{tab:MSNpredictive} shows that the MNP model performs slightly better. However, the log scores and censored likelihood scores of both models are close to the oracle values, and the Giacomini--White tests do not reject equal predictive accuracy relative to the oracle.

\begin{figure}[tb!]
\centering
\includegraphics[width=\textwidth]{fig/deltapost_MVNDGP.eps}
\caption{Posterior densities $\delta_1$ and $\delta_2$ in the SMNP model for the experiment with symmetric latent utilities. The vertical lines represent true parameter values.}
\label{fig:idendeltaMVN}
\end{figure}

Taken together, the two experiments show that the flexibility of the SMNP model is useful when asymmetry is present, but does not impose a meaningful predictive cost when the standard symmetry assumption is valid. When the latent utilities are skewed, it recovers the asymmetric choice responses and improves predictive accuracy relative to the MNP model. When the latent utilities are symmetric, it shrinks toward the MNP special case and does not sacrifice predictive performance. In sum, the SMNP model provides a compelling alternative to the MNP model: it helps avoid potential misspecification from symmetric latent utilities while allowing researchers to capture more realistic asymmetric choice responses.

\section{Empirical applications}\label{empirical}
This section demonstrates the practical relevance of our proposed SMNP model by applying it to two consumer choice datasets. Section~\ref{detergent} examines a laundry detergent purchase dataset and Section~\ref{ketchup} a ketchup purchase dataset. The implementation and prior settings are discussed in Section~\ref{sec:design}. We discuss posterior distributions estimated on the full data sets.
For the evaluation of predictive performance, we repeatedly randomly  split the data into $80\%$ training and $20\%$ test sets. For each repetition $t=1,\dots,300$, the model is estimated on the training sample and we compute both in-sample and out-of-sample predictive metrics. We report average metrics across the repeated sampling iterations, and apply Giacomini--White tests to $\{HR^{(t)},LS^{(t)},CLS^{(t)}\}^{300}_{t=1}$ of SMNP and MNP.


\subsection{Laundry detergent purchases}\label{detergent}
The laundry detergent dataset includes $2,657$ observations of consumer purchases across six laundry detergent brands, together with the log price per ounce of each brand. The brands in the dataset are Tide $(26.38\%)$, EraPlus $(19.08\%)$, Surf $(15.28\%)$, Solo $(9.52\%)$, All $(3.27\%)$, and Wisk $(26.46\%)$. \citet{chintagunta1998empirical} discuss the data, and the data are available in \cite{imai2005mnp}. We follow previous analyses of this dataset and include brand intercepts and log prices in the latent utility specification \citep{imai2005bayesian,burgette2021symmetric,loaiza2022scalable,loaiza2023fast}.

The SMNP model picks up asymmetry in the latent utility distribution. Figure \ref{fig:detergent_delta} shows the posterior densities of the skewness parameters in the SMNP model. We find that the posterior density for $\delta_2$ allocates all probability mass at positive values, and  for $\delta_5$ most probability mass is also allocated to the right of zero. The posterior densities for $\delta_1, \delta_3$, and $\delta_4$ allocate high posterior density at zero. These results suggest asymmetry in the marginal latent utility distribution of EraPlus relative to the base category Wisk.

\begin{figure}[tb!]
\centering
\includegraphics[width=\textwidth]{fig/deltas_detergent_small.eps}
\caption{Posterior densities for $\delta_1$ to $\delta_5$, representing the skewness parameters for detergent brands Tide, EraPlus, Surf, Solo, and All, respectively, with Wisk the base category. }
\label{fig:detergent_delta}
\end{figure}

To investigate how this asymmetry impacts choice responses, we examine how the price of EraPlus affects choice probabilities. Figure~\ref{fig:elasticity_detergent} shows the price elasticities implied by the SMNP and MNP models as the price of EraPlus changes. Panel (a) reports the own-price elasticity for EraPlus and Panel (b) the cross-price elasticity for Solo.
The SMNP implies more nonlinear elasticity functions than the MNP model. For prices between 0.5 and 1, the SMNP model predicts a sharper decline in demand for EraPlus and a stronger shift toward Solo, indicating that the MNP model may understate both demand losses and competitive substitution. For lower or higher prices, the SMNP and MNP show similar price elasticities.

\begin{figure}[tb!]
\centering
\includegraphics[width=\textwidth]{fig/detergent_small_elasticity.eps}
\caption{Price elasticity of EraPlus and Solo given the price of EraPlus, with the prices of the other brands fixed at their mean. The solid yellow and dotted black lines represent the estimated elasticities using the SMNP and MNP models, respectively. The price elasticities are constructed from the choice probability functions, which are provided in Appendix~\ref{A:application}.}
\label{fig:elasticity_detergent}
\end{figure}

Pricing decisions are usually made around a specific current price. At the sample mean price of EraPlus ($1.06$), the SMNP implies stronger price sensitivity than the MNP model. The own-price elasticity of EraPlus is larger in magnitude, and the cross-price elasticity for Solo is also higher. Hence, a manager using the MNP around this price would understate both the demand loss from a price increase and the substitution toward Solo. The same logic applies in the opposite direction: for a price reduction, the MNP would understate the potential gain in demand for EraPlus. The SMNP model therefore provides a more informative basis for weighing the risks of price increases against the gains from discounts.

The predictive results in Table~\ref{tab:emppredictive} support this interpretation. For the laundry detergent dataset, the SMNP model improves the hit-rate, log score, and most censored likelihood scores both in-sample and out-of-sample. These improvements are mostly statistically significant. Hence, the differences between SMNP and MNP are not only economically meaningful, but they are also linked to better predictions of SMNP. In addition, we also apply the Giacomini--White test directly to the



\begin{table}[tb!]
\centering
\caption{Predictive performance empirical applications}
\label{tab:emppredictive}
\setlength{\tabcolsep}{3.5pt}
\begin{tabular}{lrrrrrrrr}
\toprule
& \multicolumn{4}{c}{Laundry detergent} & \multicolumn{4}{c}{Ketchup} \\
\cmidrule(lr){2-5} \cmidrule(lr){6-9}
& \multicolumn{2}{c}{In-sample} & \multicolumn{2}{c}{Out-of-sample}
& \multicolumn{2}{c}{In-sample} & \multicolumn{2}{c}{Out-of-sample} \\
\cmidrule(lr){2-3} \cmidrule(lr){4-5} \cmidrule(lr){6-7} \cmidrule(lr){8-9}
Metric & MNP & SMNP & MNP & SMNP & MNP & SMNP & MNP & SMNP \\
\midrule
$HR$      & $0.4963$ & $\mathbf{0.4974}$ & $0.4958$ & $\mathbf{0.4967}$ & $0.5805$ & $\mathbf{0.5806}$ & $0.5802$ & $\mathbf{0.5804}$ \\
$LS$      & $-1.3183$ & $\mathbf{-1.3153}\rlap{$^{*}$}$ & $-1.3231$ & $\mathbf{-1.3202}\rlap{$^{*}$}$ &  $-0.9813$ & $\mathbf{-0.9803}\rlap{$^{*}$}$ & $-0.9832$ & $\mathbf{-0.9826}\rlap{$^{*}$}$ \\
$CLS_{1}$ & $-0.4678$ & $\mathbf{-0.4677}\rlap{$^{*}$}$ & $-0.4689$ & $\mathbf{-0.4689}$ & $-0.4323$ & $\mathbf{-0.4321}\rlap{$^{*}$}$ & $\mathbf{-0.4328}$ & $-0.4328$ \\
$CLS_{2}$ & $-0.3966$ & $\mathbf{-0.3940}\rlap{$^{*}$}$ & $-0.3967$ & $\mathbf{-0.3942}\rlap{$^{*}$}$ & $-0.1729$ &  $\mathbf{-0.1719}\rlap{$^{*}$}$ & $-0.1741$  & $\mathbf{-0.1732}\rlap{$^{*}$}$ \\
$CLS_{3}$ & $-0.3185$ & $\mathbf{-0.3174}\rlap{$^{*}$}$ & $-0.3191$ & $\mathbf{-0.3180}\rlap{$^{*}$}$ & $\mathbf{-0.4778}\rlap{$^{*}$}$ & $-0.4778$ & $\mathbf{-0.4782}\rlap{$^{*}$}$ & $-0.4783$  \\
$CLS_{4}$ & $-0.2823$ & $\mathbf{-0.2815}\rlap{$^{*}$}$ & $-0.2832$ & $\mathbf{-0.2825}\rlap{$^{*}$}$ & $-0.5853$ & $\mathbf{-0.5852}\rlap{$^{*}$}$ & $-0.5858$ & $\mathbf{-0.5858}\rlap{$^{*}$}$ \\
$CLS_{5}$ & $-0.1273$ & $\mathbf{-0.1271}\rlap{$^{*}$}$ & $-0.1291$ & $\mathbf{-0.1289}\rlap{$^{*}$}$ & -- & -- & -- & -- \\
$CLS_{6}$ & $-0.4733$ & $\mathbf{-0.4728}\rlap{$^{*}$}$ & $-0.4744$ & $\mathbf{-0.4738}\rlap{$^{*}$}$ & -- & -- & -- & -- \\
\bottomrule
\end{tabular}\\ \vspace{1ex} \raggedright
\footnotesize{ \textit{Notes:} The predictive performance measures are averaged across $T=300$ repeated random sample splitting iterations. For laundry detergent, $CLS_1$ to $CLS_6$ refer to Tide, EraPlus, Surf, Solo, All, and Wisk. For ketchup, $CLS_1$ to $CLS_4$ refer to Hunts, Del Monte, STB, and Heinz. Bolded values indicate superior predictive performance between the MNP and SMNP specifications within each application and sample split. The symbol $^{*}$ denotes a statistically significant difference in predictive performance.} \end{table}

\subsection{Ketchup purchases}\label{ketchup}
The ketchup dataset consists of $4,956$ observations of consumer purchases across four ketchup brands. The ketchup brands are Hunts $(20.56\%)$, Del Monte $(5.17\%)$, a store brand referred to as STB $(23.31\%)$, and Heinz $(50.97\%)$. \citet{kim1995modeling} discuss the data, and the data is available in \cite{ecdat}. We maintain the same utility specification with brand intercepts and log prices.

This application provides a second test of whether asymmetric latent utilities matter in consumer choice data. Figure~\ref{fig:ketchup_delta} shows that the posterior density of $\delta_2$, corresponding to Del Monte relative to the base category Heinz, places most of its mass above zero. The posteriors for Hunts and the store brand are located closer to zero. The strongest evidence of asymmetry therefore again appears for a specific alternative rather than uniformly across all brands.

\begin{figure}[tb!]
\centering
\includegraphics{fig/deltas_ketchup.eps}
\caption{Posterior densities for $\delta_1$ to $\delta_3$, representing the skewness parameters for ketchup brands Hunts, Del Monte, and STB respectively, with Heinz the base category.}
\label{fig:ketchup_delta}
\end{figure}

Figure~\ref{fig:elasticity_ketchup} shows how the elasticities differ between the SMNP and MNP models. Panel (a) shows the cross-price elasticity for Del Monte and Panel (b) for STB, as a function of the price for Hunts. The SMNP changes which competing brands are predicted to benefit from a price change for Hunts. Around the sample mean price of Hunts ($1.34$), it implies a higher cross-price response for Del Monte and a lower cross-price response for STB relative to the MNP model. Hence, the MNP may not only misstate the magnitude of price sensitivity, but it may also misallocate the predicted substitution pattern across competing brands.

\begin{figure}[tb!]
\centering
\includegraphics[width=\textwidth]{fig/ketchup_elasticity.eps}
\caption{Price elasticity given the price of Hunts indicated on the x-axis.  The solid yellow and dotted black lines represent the estimated elasticity using the SMNP and MNP models, respectively.}
\label{fig:elasticity_ketchup}
\end{figure}

The predictive results in Table~\ref{tab:emppredictive} also favor the SMNP model. The SMNP improves the  hit rates and log scores both in-sample and out-of-sample.  The largest predictive gain occurs for $CLS_2$, corresponding to Del Monte, which aligns with the posterior evidence of nonzero skewness for $\delta_2$. The MNP model has a higher censored likelihood score for the third category corresponding to the store brand, which has a posterior skewness parameter density centered close to zero. This suggests that the model is most useful for alternatives where the data contain evidence of asymmetric choice responses.

Taken together, the two applications suggest that the SMNP model can help practitioners avoid misleading price-sensitivity estimates and make better pricing and substitution decisions than those based on the MNP model alone. The detergent application shows how asymmetric latent utilities can change the estimated shape and magnitude of own- and cross-price elasticities for a focal brand. The ketchup application shows that the same modeling flexibility can alter the implied allocation of substitution across competitors. In both cases, the SMNP model improves prediction.


\section{Discussion} \label{discuss}
This paper extends the MNP by allowing the conditional distribution of the latent utilities to be multivariate skew-normal, in which the degree and direction of asymmetry can differ across alternatives. This preserves the main advantage of the MNP: Latent utilities can be correlated and it therefore captures substitution patterns across alternatives. At the same time, it relaxes the symmetry restriction imposed by the multivariate normal distribution.

The main methodological contribution is to make this extension identified and estimation feasible. We propose a reparameterization that enforces the scale normalization and positive definiteness of the covariance matrix. We also develop a Bayesian estimation strategy based on double data augmentation. The resulting MCMC sampler only includes a single one-dimensional Metropolis-Hastings step, and updates all other parameters using Gibbs.

Numerical experiments and empirical applications demonstrate that the SMNP model is a useful tool for multinomial choice analysis. It nests the standard MNP model, captures asymmetric choice responses, and provides economically meaningful differences in elasticities and substitution patterns. The model offers a parsimonious way to reduce misspecification risk and obtain more reliable predictions for pricing and policy decisions.

The SMNP model captures asymmetric choice responses through the distribution of latent utilities, but it does not by itself identify the behavioral mechanism generating the asymmetry. The empirical evidence that asymmetric choice responses matter creates opportunities for future research. For instance, approximate Bayesian methods may be useful for larger choice sets or richer SMNP models with random coefficients. In these settings, the proposed exact Bayesian sampler may suffer from slow mixing due to the double data augmentation and the Metropolis-Hastings update for the scale-restricted skewness parameter.

\bibliography{cash.bib}
\clearpage