Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
99,463 characters · 21 sections · 121 citation commands
A recursive logit model with choice aversion and its application to transportation networks (Forthcoming at Transportation Research Part B: Methodological)
\pagestyle{myheadings} \thispagestyle{plain} \markboth
{\bf Keywords:} choice aversion, recursive logit, IIA, directed networks, transportation networks. JEL classification: D001, C00, C51, C61
\thispagestyle{empty}
Discrete choice models have been used extensively to understand the behavior of participants (who we will refer to as users) in transportation networks (mcffadens,mcf1 and BenAkiva_Lerman1985). In this context, users choose a path that maximizes total utility (or, equivalently, minimizes their total cost) of traveling. A prominent model that has arisen from this literature is the Multinomial Logit (MNL) Model, whose main advantage is its tractability and closed-form choice probabilities.
Despite its popularity, the MNL presents some drawbacks when applied to the case of transportation networks. In particular, the MNL can predict unrealistic choice probabilities for paths sharing common edges in the network. The root of this problem traces back to the Independence of Irrelevant Alternatives (IIA) axiom that is required to derive the MNL.
To explain the severity of the overlapping paths problem, consider the simple transportation network displayed in Figure (ref).\footnote{This example follows the discussion in BenAkivaBierlarie1999.}
In this network we have three paths: $(a_1,a_3)$, $(a_1,a_4)$, and $(a_2)$. Assume the utility of each path is 1. At the edge level, assume that the utility of edge $a_1$ is $1-\theta$, the utility of edges $a_3$ and $a_4$ is $\theta$, and, finally, the utility of edge $a_2$ is 1. The value of $\theta$ is assumed to satisfy $0<\theta<1$.
For the setting described above, the MNL will predict that each path is chosen with probability $1/3$. However, these choice probabilities are unrealistic when paths lack distinctiveness or independence from another. In particular, an assignment in which path $(a_2)$ is chosen with probability $1/2$ and paths $(a_1,a_3)$ and $(a_1,a_4)$ are chosen with probability 1/4 is more sensible. More explicitly, this latter solution takes into account the fact that paths $(a_1,a_3)$ and $(a_1,a_4)$ share the common edge $a_1$ and in terms of total cost they are equivalent to each other as well as the cost of path $(a_2)$. The reason why the MNL model cannot accommodate situations where the path costs are correlated is that the MNL relies on the property of Independence of Irrelevant Alternatives (IIA).\footnote{We remark that whether the IIA property holds or not in the network discussed above does not depend on the choice of utilities associated to each edge. In fact, a similar conclusion holds when we allow for general link utilities. The violation of IIA is a result that paths $(a_1,a_3)$ and $(a_1,a_4)$ share the common link $a_1$. This latter fact is what generates the failure of IIA. } As the discussion above shows, the IIA property is hardly satisfied even in simple transportation networks.\footnote{An in-depth discussion on path correlation and IIA can be found in mcfadden1974frontiers, which discusses the canonical red bus/blue bus problem.}
Recognizing this pitfall, the transportation literature has proposed several corrections to the MNL. This class of extensions is known as Path Size Logit (PSL). The idea of this class of models is to correct the problem of overlapping paths by adding an extra correlation-penalizing term to path costs. Thus, when the choice probabilities are generated through a standard MNL, the correction will account for the degree of overlapping between different paths.
While its usefulness in correction is clear, the PSL class has two problems. First, the type of correction employed by the different models do not have a theoretical justification in terms of users' behavior. In particular, the parameters describing the corrections do not have a direct interpretation from an economic viewpoint. Second, it is not clear how PSL models can be used to carry out welfare analysis in transportation networks. For instance, it is hard to interpret and predict the changes on welfare when edges are added to, or severed from, the network.
In this paper, we propose the use of a recursive logit (RL) model which incorporates the idea of choice aversion (choice overload) in users' behavior. In doing so we adapt the approach of FudenbergStra2015 to the context of directed graphs. Simply put, the choice aversion hypothesis states that an increase in the number of alternatives to choose from may lead to adverse consequences, such as lesser motivation to actually choose or lower satisfaction ex post (cf. SheenaLepper2000 and Scheibehenneetal2010).
Formally, we consider a transportation network with source node $s$ and designated sink node $t$. In this setting, we model users' behavior as a sequential choice process: when assessing an edge $a$ at some node $i\neq t$, users evaluate both the flow utility and the appropriate continuation value associated to such an edge. Following FudenbergStra2015, we introduce a term that penalizes the size of each choice set that stems subsequently from every current edge under scrutiny. In particular, when considering an edge $a$ at node $i$, users will penalize the number of outgoing edges at $i$. In other words, when facing a set of alternatives in order to depart from a specific node, users incorporate the size of the ensuing choice set when they appraise the continuation value of each outgoing edge. Formally, at each node $i\neq t$, we consider the penalty $\kappa_i\log|A_i^+|$, where $|A_i^+|$ is the cardinality of the set of outgoing edges at node $i$ and $\kappa_i\geq 0$ is a parameter that captures users' choice aversion degree at node $i$.
We make three contributions. First, we show that RL with choice aversion can overcome the problem of overlapping in transportation networks. In particular, the parameters $\kappa_i$ plays a critical role in the form of the correction. We show that our model performs as well as the recent Adaptive PSL model introduced by DUNCAN20201. However, our correction has two main advantages. First, it is a simple correction based on users' optimal behavior. Second, the parameters $\kappa_i$ have a clear interpretation in terms of users' attitude with respect to size of choice sets.
In our second contribution, we show how our model captures violations of regularity (LuceSuppes1965). Formally, we show that removing an edge in a particular node can decrease the choice probabilities of some paths in the network. To grasp how this result works, we note that removing an edge $a$ at node $i$ is equivalent to removing the set of all paths in which $a$ is a member. When removing an edge in the traditional MNL (i.e., the choice aversion model with $\kappa_i=0$ for all $i\neq s,t$), the choice probability of the remaining paths increase proportionally. However, in our model, removing the edge $a$ not only reduces the set of available paths (passing through node $i$) but also decreases the choice aversion costs associated to these paths. This latter reduction makes the set of paths passing through node $i$ comparatively more valuable than those not using node $i$. As a consequence, the path choice probabilities of paths passing through node $i$ may increase due to the reduction in the set of available paths and the choice aversion cost reduction.
We formalize the failure of regularity in terms of a precise relationship between the parameters $\kappa_i$ and the choice probabilities of paths using the node where the edge is removed. As far as we know, this result is new to the literature on recursive models in transportation networks. From a practical standpoint, accounting for the failure of regularity provides useful information to planners who must predict flows in transportation networks.
In our final contribution, we study how our model can capture a Braess's-like paradox. In particular, we show how adding edges to the network can decrease users' welfare. Similarly to regularity failure, we also characterize this result in terms of the degree of choice aversion $\kappa$. Unlike existing models and extensions, the behavioral foundation of the choice aversion model and its formulation allows for a tradeoff between instantaneous route utilities and choice aversion penalization such that decreases in users' welfare can be observed even in the absence of congestion. To our knowledge, this is also a novel result that can shed light on the design of transportation networks and its effects on welfare.
To ground these contributions, we conclude with an empirical analysis of the choice aversion model and measure its performance relative to other RL models. The results of these maximum likelihood estimation routines provide empirical support that real-world network users experience choice overload and display choice averse behavior. Additionally, we show that the simplicity of the choice aversion model allows for the model to be estimated much more quickly than other RL models, all while providing similar corrections to path predictions as well as maintaining a clear economic and behavioral interpretation.
RL models have been studied by bc, Fosgerauetal2013, MAI2015100, and Mai2016.\footnote{For an up to date survey on recursive models in traffic networks we refer the reader to ZIMMERMANN2020100004.} Our paper differs from theirs in at least two dimensions. First, we extend their recursive approach to incorporate the notion of choice aversion. Second, we show how our model can handle the problem of path overlapping, violations of regularity, and Braess's-like paradoxes. It is worth pointing out that the papers by MAI2015100 and Mai2016 correct the overlapping problem by using a Nested Recursive Logit (NRL) and a Generalized Recursive Extreme Value model, which allow for correlation among the stochastic terms. Our approach preserves the independence of the random terms, where the choice aversion penalty allows us to capture and correct the degree of overlapping between paths.
With respect to PSL models, the existing literature is extensive (BenAkivaBierlarie1999 and Frejinger_Bierlarie2003), and DUNCAN20201 present an up-to-date discussion of the PSL approach.\footnote{We must also mention that the papers by Chu1989, KoppelmanChieh2000, Vosha1997, Vosha1998, and WenKoppelman2001 propose variants of the MNL model in an attempt to generate flexible discrete choice models that allow for correlation between different paths.} In addition, they propose an alternative correction denominated as the Adaptive Path Size Logit (APSL) model. Our results differ from theirs in the type of the behavioral foundation we use.
From a behavioral standpoint, a similar approach to this paper is found in FOSGERAU20191 and JIANG2020188, who incorporate a rational inattention model into the context of transportation networks. While choice aversion and rational inattention are interconnected in terms of information processing, we model a particular variant of costly decision-making in the form of aversion to increasing choices, rather than mutual information through observing a signal. This modeling choice allows us to generate clear path choice predictions, violations of regularity, and Braess's-like paradox phenomena.
Finally, we mention that our paper is related to Acemogluetal2017. They study a congestion game introducing the notion of information-constrained Wardrop Equilibrium, which allows users to possess different information sets about the congestion level in the network. They show that providing more information to users can make them worse off. Our work differs from theirs in at least three aspects. First, we focus on the notion of choice aversion while they focus on the concept of information sets. Second, we show our Braess's-like result in the context of a RL model, while they study a deterministic and non-recursive model. Finally, we show that Braess's paradox may occur even in uncongested networks while they focus in the case of congested networks.
The rest of the paper is organized as follows. Section (ref) introduces the RL model with choice aversion. Section (ref) explores the use of choice aversion in the path choice model and in comparison to existing Path Size Logit (PSL) models. Section (ref) discusses the failure of regularity. Section (ref) discusses a type of Braess's paradox observed as a consequence of choice aversion. Section (ref) conducts an empirical analysis to estimate the role of choice aversion in real-world contexts. Section (ref) concludes. Technical details, proofs, and further welfare analyses are gathered in Appendixes (ref), (ref), and (ref), respectively.
In this section we propose a recursive discrete choice model in directed networks. Formally, we model a set of users as solving a dynamic programming problem over a directed, acyclic graph. In a noticeable departure from previous literature, we adapt the choice aversion formulation of FudenbergStra2015 into the context of directed graphs, and then we analyze the consequences on equilibria and welfare.\footnote{Throughout the text we use the term choice overload to refer to choice aversion.}
Consider a directed (not assumed acyclic) graph $G = (N,A)$ where $N$ is the set of nodes and $A$ the set of edges, respectively. We denote the set of ingoing edges to node $i$ by $A_i^-$, and the set of outgoing edges from node $i$ by $A_i^+$. We refer accordingly to the out-degree of node $i$ as $|A_i^+|$.
Without loss of generality, we assume that $G$ has a single source-sink pair, where $s$ and $t$ stand for the source (origin) and sink (destination) nodes, respectively. Let $j_a$ be the node $j$ that has been reached through edge $a$. We therefore define a path as a sequence of edges $(a_1,\ldots a_K)$ with $a_{k+1} \in A_{j_{a_{k}}}^+$ for all $k< K$.
The set of paths connecting nodes $s$ and $t$ is denoted by $\mathcal{R}$. The set of paths connecting nodes $s$ and $i\neq t$ is denoted by $\mathcal{R}_{si}$. Similarly, the set of all paths connecting nodes $i\neq s$ and $t$ is denoted as $\mathcal{R}_{it}$. Let ${\mathcal R}_i$ denote the set of paths passing through $i$. Finally, let ${\mathcal R}_i^c$ denote the set of paths not passing through node $i$.\footnote{We note that ${\mathcal R}_i, {\mathcal R}_{si},{\mathcal R}_{it}$, and ${\mathcal R}_i^c$ are subsets of ${\mathcal R}$. }
A deterministic utility component $u_a>0$ is associated with each edge $a\in A_i^+$ for all $i\neq t$. Path utilities are assumed to be edge additive; that is, for a path $r=(a_1,\ldots a_K)\in {\mathcal R}$, its associated utilities are given by $\sum_{k=1}^Ku_{a_k}$.
We assume that at node $s$ there is a unitary mass of network users who must choose a path from the set ${\mathcal R}$. For the sake of exposition, the mass of users is summarized by the canonical vector $e_s$, which has a 1 in the position of node $s$ and zero elsewhere. The dimension of $e_s$ is $|N|-1$.
We now develop a RL choice model over $G$ that incorporates choice overload by means of an specific kind of penalty on ensuing choice sets stemming from each edge appraisal. In particular, we adapt the choice aversion approach from FudenbergStra2015 into the environment described by $G$ as follows: for each $a\in A^+_i$ we associate a collection of i.i.d. random variables $\{\epsilon_{a}\}_{a\in A^+_{i}}$ such that the recursive utility associated to edge $a$ is defined as:
where $u_a$ denotes the instantaneous utility associated to edge $a$ and the term $$\mathbb{E}\left(\max_{a^\prime \in A^+_{j_a}}\{V_{a^\prime}+\epsilon_{a^\prime}-\kappa_{j_a}\log |A^+_{j_a}| \}\right)=\mathbb{E}\left(\max_{a^\prime \in A^+_{j_a}}\{V_{a^\prime}+\epsilon_{a^\prime}\}\right)-\kappa_{j_a}\log |A^+_{j_a}| $$ is the adjusted continuation value associated to the selection of $a$. Notice that the latter term includes the factor $\kappa_{j_a}\log |A^+_{j_a}|$, which is a penalty term that captures the size of the set $A^+_{j_a}$, where $\kappa_{j_a}\geq 0$.\footnote{FudenbergStra2015 study a RL model in the context of intertemporal choice. In doing so, they consider a discount factor $\delta\in (0,1)$. We focus on a digraph $G$ without discounting.}
Following FudenbergStra2015, we impose the following assumption on the random variables $\{\epsilon_{a}\}_{a\in A^+_{i}}$.
Under this assumption, Eq. ((ref)) can be expressed as:
where $\log\left(\sum_{a^\prime \in A^+_{j_a}}e^{V_{a^\prime}}\right)-\kappa_{j_a}\log |A^+_{j_a}|$ provides a closed-form expression for the adjusted continuation value.\footnote{See Train1, Chapter 3.}
Let us define $\varphi_{j_a}(V)\triangleq \log\left(\sum_{a^\prime \in A^+_{j_a}}e^{V_{a^\prime}}\right)$ and $\hat{\varphi}_{j_a}(V)\triangleq\varphi_{j_a}(V)-\kappa_{j_a}\log|A_{j_a}^+|$ for all $j_a\neq t$. Accordingly, Eq. ((ref)) can be rewritten as:
The previous expression deserves some remarks. First, the adjusted continuation value in Eq. ((ref)) captures the complexity of the choice sets $A_{j_a}^+$, as measured by $\kappa_{j_a}\log |A^+_{j_a}|$, with $\kappa_{j_a}\geq 0$. Intuitively, the expression $\kappa_{j_a}\log |A^+_{j_a}|$ penalizes the size of the choice sets at different nodes, where the parameter $\kappa_{j_a}\geq 0$ measures users' attitudes towards the size of $A_{j_a}^+$. In particular, $V_a$ is a decreasing function of $\kappa_{j_a}$.
Second, when $\kappa_{j_a}=0$ for $j_a \notin \{s,t\}$, the expression in Eq. ((ref)) boils down to a traditional RL model where users are choice-loving: they always prefer to add additional items into the menu, as in the “preference for flexibility” of Kreps1979. To see this, note when $\kappa_{j_a}=0$, the function $\varphi_{j_a}(V)$ is increasing in $|A_{j_a}^+|$. As a consequence, the recursive utility $V_a$ increases as the size of $A_{j_a}^+$ increases. This latter feature implies that traditional RL models in transportation networks (e.g., bc, Fosgerauetal2013, MAI2015100, and Mai2016) can be associated with an intrinsic taste for plentiful options.
On the other hand, the case of $\kappa_{j_a}\in(0,1)$ may be interpreted as a situation where the users prefer to include additional edges (alternatives) to the menu, provided the new options are not too much worse than the current average. The case of $\kappa_{j_a}=1$ captures a situation where the users want to remove choices (links) that are worse than the average: they worry about choosing such additional alternatives by accident given appraisal costs---such as Ortoleva13's thinking aversion---that may offset the benefits of the corresponding random draw associated to $\epsilon$. Finally, the case of $\kappa_{j_a}> 1$ should be interpreted as a situation in which the users only wish to add alternatives that are perceived to be sufficiently better.
We remark that the three cases of possible values that $\kappa_{j_a}$ may take provide us the flexibility to capture different users' attitudes with respect to the size of the choice set at each node. In sum, the collection of parameters $\{\kappa_{j_a}\}$ encapsulates the scale of penalties on the set of ensuing actions arising from each nonterminal node, which unlocks keen consequences on users' attitudes towards marginally increasing the set of edges.\footnote{In our analysis we focus on a digraph with a single origin-destination pair. However, our analysis can be adapted to the case where we have multiple origin-destination pairs. In particular, we can specify a parameter $\kappa_{j_a}^{s_l t_l}$ where $s_l t_l$ denotes a particular origin-destination pair for $l=1,\ldots, L.$}
Following FudenbergStra2015, we identify $\{\kappa_{j_a}\}_{a\in A^+_{i}}$ as the users' choice aversion parameters.
Each user is looking for an optimal path connecting $s$ and $t$. Now, when they reach node $i\neq t$, they observe the realization of the random utilities $V_a+\epsilon_{a}$ for all $a\in A_i^+$, and consequently choose the alternative $a\in A^+_{i}$ with the highest utility.
This process is repeated at each subsequent node giving rise to a recursive discrete choice model, where the expected flow entering node $i\neq t$ splits among the alternatives $a\in A_i^+$ according to the choice probability:
Due to Assumption (ref), Eq. ((ref)) can be rewritten as:
As $\kappa_{j_a}$ increases, the edge choice probability $\mathbb{P}(a|A^+_i)$ is increasingly penalized by the size of the choice set $|A_{j_a}^+|$, reflecting the cost of choice overload onto a user's edge utility from nodes with large choice sets.
This is a fundamental difference with the traditional RL model, which assumes $\kappa_{j_a}=0$ as we mentioned before.
Mathematically, the recursive process just described induces a Markov chain over the graph $G$, where the transition probabilities are given by Eq. ((ref)).
Let $x_i$ be the expected flow entering at node $i$ towards sink node $t$. Then the flow received by edge $a$ is given by:
with $f=(f_a)_{a\in A}$ denoting the expected flow vector.
In addition, let $\hat{\mathbb{P}}=(\mathbb{P}_{ij})_{i,j\neq t}$ denote the restriction to the set of nodes $N\setminus\{t\}$. Then the expected demand vector $x=(x_i)_{i\neq t}$ may be expressed as $x=e_s+\hat{\mathbb{P}}^Tx$ which generates the following stochastic conservation flow equations
A flow vector $f$ satisfying ((ref)) is called feasible. It is worth mentioning that there exists a unique flow vector $x^*$ satisfying the flow constraints ((ref)). In fact, using bc, it is possible to show that $[I-\mathbb{P}^\top]^{-1}$ is well-defined. Then $x^*$ is the unique vector that satisfies $x^*=[I-\mathbb{P}^\top]^{-1}e_s$ and $f_a=x_i^*\mathbb{P}(a|A_i^+)$ for all $a\in A_i^+$, $i\neq t.$
In this section we note that the solution of our recursive choice model can be equivalently written in terms of path choice probabilities. In doing so, assume that for each path $r\in{\mathcal R}$ the utility associated to it is a random variable defined as
where $U_r=\sum_{a\in r}(u_a-\kappa_{j_a}\log|A_{j_a}^+|)=\sum_{a\in r}u_a-\sum_{a\in r}\kappa_{j_a}\log|A_{j_a}^+|$ and $\{\epsilon_r\}_{r\in\mathcal{R}}$ is a collection of absolutely continuous random variables satisfying Assumption (ref).
Under these conditions, the probability of choosing path $r$ is defined as:
Equations ((ref)) and ((ref)) jointly define a path choice model over ${\mathcal R}$, where we again refer to the Gumbel assumption to obtain:
However, it is well known that the path choice probability $\mathbb{P}_r$ can be decomposed in terms of the edge probabilities (e.g., Fosgerauetal2013 and Gilbertetal2014). Formally, we have that for each path $r=(a_1,\ldots,a_K)\in {\mathcal R}$ with $K\geq 2$, the following equality holds
The previous characterization will play a key role in later sections. Intuitively, Eq. ((ref)) establishes that the $\mathbb{P}_r$ can be decomposed in terms of the recursive choice probabilities. This equivalence allows us to highlight the role and effect of the terms $\kappa_{j_a}\log|A_{j_a}^+|$ in the path choice probabilities $\mathbb{P}_r$. In particular, Eq. ((ref)) takes the form
where $u_r=\sum_{a\in r}u_a$ and $\rho_r\triangleq\sum_{a\in r}\kappa_{j_a}\log |A_{j_a}^+|$.
Fosgerauetal2013 discuss the RL when the graph $G$ may contain cycles, a situation where paths may contain loops and be arbitrarily long. They show that, in general, there will be infinitely many potential paths that connect an origin to a destination, but the probability of choosing a particular path is given by the MNL model, where the choice set is discrete but infinite. In particular, under the existence of cycles, the set of paths ${\mathcal R}$ is replaced by the infinite set $\Omega$, which contains all possible paths. Accordingly, the probability of choosing a particular path $r$ is given by: $$\mathbb{P}_r=\frac{e^{u_r-\rho_r}}{\sum_{r^\prime\in\Omega}e^{u_{r^\prime}-\rho_{r^\prime}}}\quad\quad\mbox{for all $r\in \Omega$}.$$ From the previous formula, it is easy to see that we can incorporate cycles in the same way as Fosgerauetal2013. However, an important difference with Fosgerauetal2013 is the fact that $\mathbb{P}_r$ incorporates the terms $\rho_r$. In particular, the ratio of choice probabilities between $r$ and $r^\prime$ depends on the difference $(u_r-\rho_r)-(u_{r^\prime}-\rho_{r^\prime})$.
In the context of path choice models, it is well known that the traditional MNL model is restricted by the IIA property, which does not hold in the context of route choice due to the overlapping paths problem. The main implication of overlapping paths is that the traditional MNL produces unrealistic path choice probabilities (BenAkiva_Ramming1998 and BenAkivaBierlarie1999).
In order to solve this problem, the route choice literature has proposed the Path Size Logit (PSL) approach. In simple terms, PSL models extend the MNL by adding a correction term to path utilities which account for the degree of overlapping among paths. For instance, BenAkiva_Lerman1985, BenAkivaBierlarie1999, Frejinger_Bierlarie2003, and, recently, DUNCAN20201 propose different corrections to the MNL in order to solve the overlapping problem. From a behavioral point of view, PSL models try to correct the fact that the IIA property should be relaxed in contexts where paths are not distinct or independent.
Following Luce1959, the IIA property can stated as follows: Given $r,r^\prime\in{\mathcal R}$,
From ((ref)) it follows that the IIA property establishes that the ratio between the probabilities of choosing $r$ and $r^\prime$ is independent of the choice set containing $r$ and $r^\prime$. In other words, the comparison between paths $r$ and $r^\prime$ should not be affected by expanding (shrinking) the set of paths, e.g., by adding (severing) links in some nodes.
PSL models described above allow for violations of Eq. ((ref)). In this section we show how our recursive model also allows for violations of condition ((ref)).\footnote{Our model allows for violations of the IIA property at the route level. However, we remark that at the node level, edge choice probabilities satisfy the IIA condition.} Thus, the notion of choice aversion can be seen as a natural mechanism that overcomes the problem of overlapping paths in transportation networks.
To see how our model works, for ease of exposition assume $\kappa_i=\kappa$ for all $i\notin \{s,t\}$, let us rewrite Eq. ((ref)) in \S (ref) as follows:
where $u_r=\sum_{a\in r}u_a$, $\gamma_r\triangleq\sum_{a\in r}\log |A_{j_a}^+|$, and $\kappa\geq 0$ is the (homogeneous) choice aversion parameter.\footnote{Note that equation ((ref)) simplifies to ((ref)) when the choice aversion parameters are fixed across nodes. In this case, $\rho_r = \kappa \gamma_{r}$.}
Expression ((ref)) deserves some comments. First, for $\kappa>0$, the term $\kappa\gamma_r$ can be seen as a penalty term that accounts for the size of choice sets at each of the nodes accessed along path $r$.\footnote{Note that paths passing through a common node $i$ would share the penalty $\kappa\log |A_i^+|$.} However, different from traditional corrections in the PSL literature, $\kappa\gamma_r$ is link additive. Second, the penalty $\kappa\gamma_r$ accounts for the degree of overlapping among different paths. In particular, adding or severing links to some of the nodes crossed by path $r$ will modify the value of $\kappa\gamma_r$.
To gain some intuition , consider two paths $r$ and $r^\prime$ with associated probabilities $\mathbb{P}_r$ and $\mathbb{P}_{r^\prime}$ respectively. Computing the probability ratio between $r$ and $r^\prime$ we get:
Expression ((ref)) shows that the ratio ${\mathbb{P}_r\over \mathbb{P}_{r^\prime}}$ depends on the ratio between the utilities associated to paths $r$ and $r^\prime$ times the ratio between $\kappa\gamma_r$ and $\kappa\gamma_{r^\prime}$. This latter term incorporates information about $r$ and $r^\prime$ regarding choice sets at each node crossed by these paths. This information is captured by the terms $\kappa\gamma_r$ and $\kappa\gamma_{r^\prime}$ respectively.
Concretely, in Eq. ((ref)), it is easy to see that adding or deleting links in a particular node crossed by $r$ (or $r^\prime$) will affect the ratio ${e^{-\kappa\gamma_r}\over e^{-\kappa\gamma_{r^\prime}}}$, and as a consequence ${\mathbb{P}_r\over \mathbb{P}_{r^\prime}}$ will be modified. In other words, ${\mathbb{P}_r\over \mathbb{P}_{r^\prime}}$ is not independent of the set ${\mathcal R}$. Furthermore, we note that when $\kappa=0$, then ${\mathbb{P}_r\over \mathbb{P}_{r^\prime}}= {e^{u_{r}}\over {e^{u_{r^\prime}}}}$ and we recover the IIA property in the MNL model.
It is worth remarking that if paths $r$ and $r^\prime$ cross through the same nodes, we get $\kappa\gamma_r=\kappa\gamma_{r^ \prime}$, so that ${\mathbb{P}_r\over \mathbb{P}_{r^\prime}}={e^{u_r}\over e^{u_{r^\prime}}}.$ Thus, the factor ${e^{-\kappa\gamma_r}\over e^{-\kappa\gamma_{r^\prime}}}$ captures the degree of overlapping between different paths.
In order to see how our model overcomes the overlapping problem, we study a concrete case. Let us reintroduce the network in Figure (ref) where the set of paths is given by $\mathcal{R}=\{r_1,r_2,r_3\}$ with $r_1=(a_1,a_3)$, $r_2=(a_1,a_4)$, and $r_3=(a_2)$. For this example, let $u_a = -c_a$ and assume $c_{a_1}=1.9$, $c_{a_3}=c_{a_4}=0.1$, and $c_{a_2}=2$.
As we have previously discussed for this network structure, paths $r_1$ and $r_2$ overlap, sharing the common edge $a_1$. Under this parameterization of instantaneous utility, coupled with $\kappa=0$, it follows that $u_{r_1}=u_{r_2}=u_{r_3}=-2$, and, consequently, the logit choice rule ((ref)) assigns one third of flow to each path. In other words, with $\kappa=0$, we get $\mathbb{P}_{r_1}=\mathbb{P}_{r_2}=\mathbb{P}_{r_3}={1\over 3}$. However, since paths $r_1$ and $r_2$ are identical in utility, the assignment $\mathbb{P}_{r_1}=\mathbb{P}_{r_2}={1\over 4}$ and $\mathbb{P}_{r_3}={1\over 2}$ is a more appropriate allocation.
However, as $\kappa \rightarrow 1$, the logit model with choice aversion predicts a flow allocation approaching $\left({1\over 4},{1\over 4}, {1\over 2}\right)$. To understand this, we look at the probability ratio between paths $r_1$, $r_2$, and $r_3$:
From the previous expression, it follows that the value of $\kappa$ will affect the ratios ${\mathbb{P}_{r_1}\over \mathbb{P}_{r_3}}$ and ${\mathbb{P}_{r_2}\over \mathbb{P}_{r_3}}$, but not ${\mathbb{P}_{r_1}\over \mathbb{P}_{r_2}}$. In particular, Eq. ((ref)) shows that the ratios ${\mathbb{P}_{r_1}\over \mathbb{P}_{r_3}}$ and ${\mathbb{P}_{r_2}\over \mathbb{P}_{r_3}}$ are decreasing in $\kappa$. In other words, as the degree of choice aversion increases, the probabilities $\mathbb{P}_{r_1}$ and $\mathbb{P}_{r_2}$ decrease while the probability associated to $r_3$ increases.
Figure (ref) shows how the route choice probabilities in Figure (ref) respond to $\kappa \in [0,2.5]$ under choice aversion.
Now assume that a new edge $\hat{a}$ is added at node $i_1$ which connects to destination node $t$. This implies that the new set of paths is $\hat{\mathcal{R}}=\mathcal{R}\cup\{\hat{a}\}$. In terms of Eq. ((ref)), adding $\hat{a}$ implies that: $${\mathbb{P}_{r_1}\over \mathbb{P}_{r_2}}=1\quad\text{and}\quad {\mathbb{P}_{r_1}\over \mathbb{P}_{r_3}}={\mathbb{P}_{r_2}\over \mathbb{P}_{r_3}}=e^{-\kappa\log 3}.$$ This latter expression makes explicit the fact that changes in $\mathcal{R}$ will change the probability ratio between different paths, i.e., IIA does not hold.
What the analysis just laid out shows---which applies to the general case of directed networks---is that from the vantage point of path selection, choice aversion is a robust way to derive path choice probabilities, even in the case of overlapping of different routes. This robustness feature makes our approach similar to the class of PSL models, which is widely used in applied work (e.g. DUNCAN20201).
In this section we compare our approach with the APSL model, a state-of-the-art extension of the class of PSL models introduced by DUNCAN20201.\footnote{For ease of exposition, we provide the details of the APSL and other PSL models in Appendix (ref).} Figure (ref) displays a network topology similar to one featured in Fosgerauetal2013. We test the performance of the choice aversion model on this network in calculating path choice probabilities. We assume that this is a directed acyclical graph, where the set of paths is given by $\mathcal{R}=\{r_1,r_2,r_3,r_4\}$ with $r_1=(12,23,35)$, $r_2=(12,23,34,45)$, $r_3=(12,24,45)$, and $r_4=(15)$. Similar to the previous section, let $u_a = -c_a$ for each edge $a$.
This graph represents a more complex uncongested network topology where the cost of all routes $r_i \in \mathcal{R}$ are equal. Thus, the only difference in routes 1 through 4 are the choice sets at each node along the path. The MNL model (equivalent to $\kappa=0$ in the choice aversion model) predicts equal path choice probabilities, i.e., $\mathbb{P}_{r_i}={1\over 4}$ for $i=1,2,3, 4$.
Figure (ref) shows choice probabilities as $\kappa$ increases from 0 to 10. As the choice aversion penalization grows larger, $\mathbb{P}_{r_4}$ approaches 1, since it is the only route with no alternatives for the user to decide between once the user travels beyond node 1. On the other hand, while $r_1$ and $r_2$ have equivalent choice aversion terms, $r_3$ has the advantage of lacking an additional downstream choice in comparison, allowing $\mathbb{P}_{r_3} > \mathbb{P}_{r_1} = \mathbb{P}_{r_2}$ for $\kappa > 0$ until the choice aversion parameter grows so large that the respective route choice probabilities converge and approach zero.
The prediction in Figure (ref) differs from route choice probabilities generated by the RL in Fosgerauetal2013 and the APSL model proposed by DUNCAN20201. This latter model is shown in Figure (ref). In the Adaptive PSL model, the overlapping path correction is based on correlation of routes through link-path incidence rather than a penalization for size of choice sets along the path. As a result, we see that for the APSL model, $\mathbb{P}_{r_1} = \mathbb{P}_{r_3}$ for all values of $\beta$, not $\mathbb{P}_{r_1} = \mathbb{P}_{r_2}$ as in the choice aversion model.
In other words, the APSL model assigns equal probability to paths 1 and 3 because they share equal edge costs and degrees of overlap with other paths, even as the sequence of costs differs. On the other hand, the choice aversion model assigns equal flow to paths 1 and 2 because they are equal in cost and number of outgoing links at path nodes, where we note that one outgoing link results in a choice aversion penalization of zero.\footnote{We thank an anonymous referee by suggesting this paragraph.}
We remark that the behavioral nature of the choice aversion model allows for a different type of correction than APSL model. The choice aversion and the APSL model work in a similar way in the sense of overcoming the overlapping path problem. However, the choice aversion model has a simple and clear behavioral interpretation. This feature sets the choice aversion model apart from other APSL and other PSL models both in terms of interpretation and performance.\footnote{We note that for paths crossing the same nodes, the IIA property holds. To see this, consider paths $r$ and $r^\prime$, crossing exactly the same nodes. Thus in this case $\kappa\gamma_r=\kappa\gamma{r^\prime}$ and ${\mathbb{P}_r\over \mathbb{P}_{r^\prime}}=e^{u_r-u_{r^\prime}}$. }
The paper by Fosgerauetal2013 proposes a link additive correction attribute called Link Size (LS). The LS correction is similar in spirit to PSL models. However, instead of correcting for the number of paths sharing a given link, the LS correction considers the expected link flow as a proxy for the amount of overlap between paths. Our choice aversion correction $\kappa \gamma_r$ is similar to the LS approach in the sense that is link additive, but the mechanism behind the correction of the overlapping paths problem is quite different. Our correction penalizes choice sets while LS penalizes common links. In other words, we focus at the node level while Fosgerauetal2013 focus at the link level. Thus, both approaches can be seen as alternatives to each other.\\
In addition, we mention that our approach differs from Fosgerauetal2013 in two important aspects. First, we show that our approach can generate violations of regularity. Second, we show how our model can be useful to characterize changes on welfare when the network is modified. We discuss these two topics in \S (ref) and (ref), respectively, and we provide an empirical comparison of these models in \S (ref).
The paper by MAI2015100 introduces a Nested Recursive Logit (NRL) model to overcome the overlapping paths problem. Using an approach similar to the Nested Logit (McFadden1978), their model allows for path utilities to be correlated where the links can have different scale parameters.
Formally, MAI2015100 extend the RL model in Fosgerauetal2013, by allowing the scale parameter $\mu_a$ of the Gumbel-distributed random variables $\{\epsilon_{a}\}_{a\in A} $ to be link-specific (in contrast with our Assumption (ref)). To see how their approach works, consider a link $k$ and a node $i=j_k$ where $k\in A_i^-$. Let $\mu_{j_k}$ be an link-specific parameter. Accordingly, the continuation value associated to node $j_k$ is defined as:
From ((ref)) it follows that the link-specific scale parameter $\mu_{j_k}$ affects the continuation values. Based on expression ((ref)), MAI2015100 show that the IIA property can be relaxed using different scale parameters. More importantly, they show how by allowing correlation between path utilities, the NRL generates more realistic path choice probabilities.
Our approach shares the recursive structure of the NRL. However, both models differ in a few key aspects. First, the choice aversion term $\kappa_{j_k}\log |A_{j_k}^+|$ modifies the level of the associated continuation value, while the link-specific parameter $\mu_{j_k}$ modifies the shape of it. Second, as we discussed above, the inclusion of the choice aversion term allows us to relax the IIA condition. MAI2015100 relax the IIA assumption by allowing correlation across stochastic terms $\{\epsilon_a\}$ at each node and through the scale parameters $\{\mu_{j_k}\}$. Thus, choice aversion exploits the topology of the network to relax IIA, while the NRL exploits the correlation between the random variables $\{\epsilon_a\}$ at each node. Third, as we shall see in \S (ref), our approach allows for the failure of the regularity property in the path choice probabilities. This pattern is not captured by the NRL model. A final difference is the fact that our approach allows us to characterize how modifications to the network can be welfare-improving or not. In \S(ref), we discuss this type of exercise.
It is worth remarking that MAI2015100 model the parameter $\mu_{j_k}$ as a function of $|A_{j_k}^+|$. Thus, they are able to capture the effect of $|A_{j_k}^+|$ in the edge and choice probabilities. This latter fact can be seen as an alternative way of capturing the notion of choice aversion. However, there are some key differences between their approach and our choice aversion model.
In order to visualize the differences between the NRL and choice aversion models, we compare path predictions using the example network shown in Figure (ref), which was first introduced in MAI2015100.
Figure (ref) has five nodes \{A, B, C, D, E\} and 10 links between the source node A and sink node D. There are six paths from A to D: $(a,a_1)$, $(a,a_2)$, $(a,a_3,e)$, $(b,b_1,e)$, $(b,b_2)$, and $(b,b_3)$. We denote these paths by $r_1$, $r_2$, $r_3$, $r_4$, $r_5$, and $r_6$, respectively. Following MAI2015100, link $f$ is not considered in the example routes, but its existence affects the continuation values for overlapping paths in the NRL model. Naturally, in the choice aversion model, link $f$ is included as an option in the choice set at node $E$.
Following MAI2015100, we use the parameterization $\mu_a = 0.5$ and $\mu_b = 0.8$. In addition, we model $\mu_a$ and $\mu_b$ as follows:
where $OL_{j_k}\triangleq|A_{j_k}^+|$ refers to the number of outgoing links. Intuitively, $\omega_{j_k}$ captures how the number of outgoing links affects the magnitude of $\mu_{j_k}$. Given that $\mu_{j_k}\in(0,1]$, we can see from ((ref)) that $\omega_{j_k}=(\log\mu_{j_k}) /OL_{j_k}\leq 0$. MAI2015100 also provide further empirical evidence that $\omega_{j_k}<0$. Thus, in order to compare our approach with theirs, we set $\kappa_{j_k} = -\omega_{j_k}$.
Table (ref) displays the differences in choice probabilities predicted by NRL and choice aversion models for the network in Figure (ref). For the choice aversion model, we also include a baseline parameterization where $\kappa=1$ for all nodes $i \neq t$. Here, we see the difference in predictions between the choice aversion model, which penalizes large choice sets directly, and the NRL model, where more options indirectly adjust the scaling of error terms. In other words, we incorporate the effect of a large choice set in a linear way, while MAI2015100's approach is nonlinear.
In terms of path choices, the NRL predicts a higher choice probability for routes $r_1$ and $r_6$ relative to the choice aversion model. The reason for this comes from the fact that from a behavioral standpoint, when $\omega<0$, the NRL model predicts that the role of the random shocks in user utilities is smaller. In other words, as the magnitude of $OL_i$ increases, users' choice of paths becomes more precise. This effect is particularly important for paths with a large degree of overlapping. In contrast, the choice aversion model predicts lower likelihoods of selecting the least costly routes as choice sets grow larger, which is aligned with the idea that in larger sets of options, the cost of thinking associated to choice aversion increases (Ortoleva13 and FudenbergStra2015).
Finally, we close this section noticing that, given the recursive nature of the choice aversion problem, we can combine our approach with the NRL. To see this note that Eq. ((ref)) can be rewritten as:
Studying the properties of combining these two models is left for future research.
Mai2016 proposes a way to estimate a recursive route choice model, which generalizes other existing recursive models in the literature. In particular, this approach can incorporate Multivariate Extreme Value models at each node, which allows him to accommodate general patterns of substitution between different paths. However, Mai2016 does not consider the problem treated in this paper. In particular, his model can generate neither a failure of regularity nor Braess's-like phenomena. Furthermore, and similar to Fosgerauetal2013 and MAI2015100, the model in Mai2016 predicts that adding links (or paths) to the network always increases users' welfare.
In the standard MNL (i.e., $\kappa_i=0$ for all $i\neq s,t$), adding an additional alternative to the choice set cannot increase the probability that an existing action is selected (and vice versa). This is known as the regularity property in Random Utility models (LuceSuppes1965).
Formally, in the context of a route choice, the regularity property states that, given the set of routes ${\mathcal R}$ and ${\mathcal R}^\prime$ with ${\mathcal R}\subseteq {\mathcal R}^\prime$, we will have:
Intuitively, condition ((ref)) establishes that removing a path from a set ${\mathcal R}^\prime$ should (weakly) increase the choice probability of the remaining paths. Alternatively, adding new paths to the choice set ${\mathcal R}$ should (weakly) decrease the choice probability of the existing paths in ${\mathcal R}$.
In this section, we show that there exists a critical value of $\kappa_i$, which allows us to understand how varying the network $G$ can generate violations of regularity. In order to gain some intuition, we discuss a simple network that allows us to show how regularity may break down. In particular, we study another nested network structure considered in MAI2015100. Figure (ref) is a simplified version of Figure (ref), such that links $e$ and $f$ are no longer present. There are still six possible paths from A to D: $(a,a_1)$, $(a,a_2)$, $(a,a_3)$, $(b,b_1)$, $(b,b_2)$, and $(b,b_3)$. Again, we denote these paths by $r_1$, $r_2$, $r_3$, $r_4$, $r_5$, and $r_6$, respectively.
Tables (ref) and (ref) display route choice probabilities when links $a_1$, $a_2$, $b_1$, and $b_2$ are removed from the nested network in Figure (ref). Table (ref) shows the results when choice parameters are homogeneous (i.e., $\kappa_{B}=\kappa_{C}=\kappa=1$), and Table (ref) shows the results when $\kappa_C=2$, all else held equal.
From Table (ref), we note that when link $a_1$ is removed, the choice probabilities of remaining paths \{$r_2$, \dots, $r_6$\} increase. Note that the probabilities of paths $r_2$ and $r_3$ increase in the same proportion (126%). Similarly, the probabilities of paths $r_4$, $r_5$, and $r_6$ also increase in the same proportion (51%). Each increase is even more pronounced in Table (ref) where $\kappa_{C}=2$. However, in both tables, the increase is not proportional across paths crossing different nodes. This feature comes from the IIA property, which holds within nodes (nests) but not across them.
Now, consider the case of removing edge $a_2$. In Table (ref), the probabilities of paths $r_1$ and $r_3$ increase proportionally (38%). However, the probability of paths $r_4$, $r_5$, and $r_6$ actually decrease. This counterintuitive result is a consequence of the fact that removing edge $a_2$ not only reduces the number of available paths but also decreases the cost associated to paths $r_1$ and $r_3$. In other words, this effect can be decomposed into two parts. First, the IIA property implies that the probabilities of $r_{1}$, $r_3$, $r_4$, $r_5$, and $r_6$ will increase upon removing edge $a_2$. The second force behind this counterintuitive result is that removing edge $a_2$ reduces the choice set when taking paths $r_1$ and $r_3$. For a choice averse user, this latter effect implies that $\kappa\log 3$ reduces to $\kappa\log 2$, which makes paths $r_1$ and $r_3$ relatively more attractive than $r_4$, $r_5$, and $r_6$, such that the probabilities of $r_4$, $r_5$, and $r_6$ decrease. When this second effect dominates, we will observe the failure of regularity associated with removing edge $a_2$.
It may be tempting to think that the effect in path choices described above may be driven by the assumption that $\kappa$ is homogeneous. However, Table (ref) shows that a similar pattern occurs when we consider $\kappa_B=1$ and $\kappa_C=2$. In this case, the failure of regularity is even more pronounced through relatively larger changes to remaining path probabilities.
We formalize the previous intuition in Proposition (ref) below. In doing so, recall that ${\mathcal R}_i$ is the set of paths passing through node $i$. Similarly, the set of paths not passing through node $i$ is defined as ${\mathcal R}_i^c$. In addition, define ${\mathcal R}_{ia}\subseteq {\mathcal R}$ as the set of paths passing through node $i$ after removing edge $a$ at node $i$. Finally, let $\mathbb{P}({\mathcal R}_i)=\sum_{r\in{\mathcal R}_i}\mathbb{P}_r$ and $\mathbb{P}({\mathcal R}_{ia})=\sum_{r\in {\mathcal R}_{ia}}\mathbb{P}_r$.
Some remarks are in order. First, Proposition (ref) provides a simple condition to know when removing an edge at node $i$ will decrease the probability of paths not crossing node $i$. Condition ((ref)) captures the fact that removing an edge at node $i$ will not only modify the set of available paths but also the choice aversion cost. Concretely, condition ((ref)) establishes a lower bound on the parameter $\kappa_i$ in terms of the path choice probabilities $\mathbb{P}({\mathcal R}_i)$ and $\mathbb{P}({\mathcal R}_{ia})$ and the magnitude of $|A_i^+|$ and $|A_i^+-1|$. To the best of our knowledge, this result is new to the literature on recursive discrete choice models in transportation networks.
Second, we note that, from a practical point of view, Proposition (ref) is useful because it allows us to understand how modifying the structure of the network $G$ can have different effects in terms of flow prediction (choice probability). In particular, ((ref)) establishes a simple condition to understand the choice behavior as a function of the parameter $\kappa_i$, which is node-specific. Thus, when designing interventions in transportation networks, Proposition (ref) helps us to understand how users' behavior will react to such interventions.
Third, Proposition (ref) predicts that removing edge $a\in A_i^+$ can decrease the probability of paths passing through nodes different from $i$. We have identified this phenomenon as a failure of regularity in the sense of LuceSuppes1965. However, behind this result is the factor that reducing the cardinality of $A_i^+$ reduces the penalization associated with choice aversion. This reduction effect can dominate, making paths crossing node $i$ more attractive than before. As a consequence, the choice probabilities of paths crossing node $i$ should increase while the choice probability of paths not crossing node $i$ should decrease.
In the context of rational inattention, MatejkaMckay2015 have shown that regularity may fail in the logit model. Our result is different in two aspects. First, we study a RL model with choice aversion in the context of transportation networks. Second, our result highlights the role of choice aversion by providing a specific condition on $\kappa_i$. MatejkaMckay2015 use the idea of information acquisition in order to derive their result.\footnote{We note that MatejkaMckay2015's analysis has been extended to the general class of additive random utility models by Fosgerauetal2020.}
Finally, we mention that a particular case of Proposition (ref) is when $ \kappa_i=\kappa$ for all $i\in A$.
In order to see how Proposition (ref) applies in the concrete case of the network in Figure (ref), Table (ref) summarizes the information after removing edges $a_1$, $a_2$, $b_1$, and $b_2$, respectively. The main message from this table is the simplicity in checking Proposition (ref).
In previous sections we have shown how the recursive choice aversion model corrects the problem of predicting routing behavior when there are overlapping paths. Similarly, we have shown how this model may generate violations of regularity when some edge of the network is removed.
In this section, we show how choice aversion can capture changes to users' welfare when the network topology is modified. Formally, we show that under the presence of choice aversion, adding edges to the network can decrease users' welfare. In particular, we show how a type of Braess's paradox (Braess1 and Braess2) can emerge even in the case of uncongested networks.
Following mcf1, we define the users' welfare as follows:
where the last equality follows from Assumption (ref) and $\boldsymbol{\kappa}=(\kappa_i)_{i\neq s, t}$. Notice that this definition exploits the equivalence in Eq. ((ref)) and makes explicit the dependence of users' welfare on the choice aversion parameters $\boldsymbol{\kappa}$. We say that $\boldsymbol{\kappa}^\prime>\boldsymbol{\kappa}$ iff $\kappa^\prime_i>\kappa_i$ for all $i\neq s,t.$
Following the literature on discrete choice models, expression ((ref)) can be interpreted as the inclusive value of paths in ${\mathcal R}$, which is equivalent to say that $\mathcal{W}(\boldsymbol{\kappa})$ measures the inclusive value of the source node $s$. In particular, $\mathcal{W}(\boldsymbol{\kappa})$ represents the expected utility faced by the network users. We now establish some properties of $\mathcal{W}(\boldsymbol{\kappa})$.
Part a) of Prop. (ref) shows an intuitive result: users' welfare decreases when the choice aversion parameter vector $\boldsymbol{\kappa}$ increases. Part b) establishes that an increase in the utility of an edge $a$, $u_a$, subsequently increases users' welfare $\mathcal{W}(\boldsymbol{\kappa})$ by the amount of flow received by edge $a$. Finally, part c) of Prop. (ref) allows us to predict precisely how much users' welfare $\mathcal{W}(\boldsymbol{\kappa})$ decreases when the choice aversion parameter $\kappa_i$ increases. Quantitatively, $\mathcal{W}(\boldsymbol{\kappa})$ decreases by the sum of the path choice probabilities of all paths passing through node $i$, multiplied by the natural log of the size of the choice set at node $i$, $A_i^+$.
Next, we show how $\mathcal{W}(\boldsymbol{\kappa})$ changes in response to adding or deleting edges to the network $G.$ To that end, we connect changes on $\boldsymbol{\kappa}$ with changes on $\mathcal{W}(\boldsymbol{\kappa})$ when the network $G$ is modified. Formally, we have the following:
It is worth stressing three implications of this result. First, Proposition (ref) underscores the way in which local effects\textemdash namely, the addition of alternatives at the node level\textemdash propagate throughout the transportation network and unlock aggregate welfare effects. Second, it connects the axiomatic characterization of stochastic choice in dynamic settings from FudenbergStra2015 with transportation networks. In particular, our Proposition (ref) extends their Proposition 3 to the case of transportation networks and graph-related settings. Finally, and most notably, it shows that Braess's network paradox may equivalently stem from choice aversion with no allusion to congestion whatsoever.
From an economic standpoint, Proposition (ref) characterizes how adding an edge to the network $G$ is welfare-improving as a function of the value of $\kappa_i$. In particular, after the new edge $a^\prime$ has been added, Eq. ((ref)) can be easily checked to know whether $\mathcal{W}(\boldsymbol{\kappa})$ increases or decreases. Subsequently, the choice aversion model predicts that there exists a range of values for the choice aversion parameter $\kappa_i$ where the addition of new edges can lead to a decrease in welfare, even if the utility of newly created routes is higher than existing routes. Thus, we can predict the emergence of Braess's paradox-like phenomena in the case of uncongested transportation networks through Proposition (ref). \S (ref) discusses this result in detail.
We note that from an empirical point of view, Eq. ((ref)) can be estimated providing a simple test to understand when modifications to the network are welfare-improving or not.
Finally, we point out that, in Appendix (ref), we carry out a welfare analysis comparing our approach with several PSL models, including the APSL.
Proposition (ref) provides a simple condition that characterizes the way in which an improvement at the node level may increase consumer surplus. This result can be connected to Braess's paradox, which states that adding free routes into a transportation network makes every decision maker worse off (Braess1,Braess2). The connection we establish is that we may generate this phenomenon as a consequence of choice aversion in a general directed network.
In order to see how our result is related to Braess's paradox, consider the special directed networks given in Figures (ref) and (ref). Figure (ref) shows a parallel serial network with paths $r_1=(a_1,a_2)$ and $r_2=(a_3,a_4)$. The set of paths therefore is given by ${\mathcal R}=\{r_1,r_2\}$, and welfare described by $$\mathcal{W}(\boldsymbol{\kappa})=\log(e^{U_{r_1}}+e^{U_{r_2}}).$$
Now assume that the network in Figure (ref) is modified by means of adding a new edge starting at node $i_1$ and ending at node $i_2$, as Figure (ref) shows. Now the new set of paths is given by $\tilde{\mathcal{R}}=\{r_1,r_2,r_3\}$, where $r_1$ and $r_2$ are defined as before and $r_3=(a_1,a_5,a_4)$. Users' surplus for this modified network is given by: $$\tilde{\mathcal{W}}(\boldsymbol{\kappa})=\log(e^{U_{r_1}}+e^{U_{r_2}}+e^{U_{r_3}}).$$ Proposition (ref) allows us to characterize the effects of adding edge $a_{5}$ into the original network. In particular, $\tilde{\mathcal{W}}(\boldsymbol{\kappa})>\mathcal{W}(\boldsymbol{\kappa})$ iff $$\mathbb{P}(a_{5}|\{a_2,a_{5}\})>1-\left(\frac{1}{2}\right)^{\kappa_{i_1}}.$$ In other words, moving from network (ref) to (ref) improves users' welfare if and only if the probability of choosing $a_5$ is strictly greater than $1-\left(\frac{1}{2}\right)^{\kappa_{i_1}}$, which is automatically satisfied when $\kappa_{i_1}=0$. To state this condition in terms of $\kappa_{i_1}$, we assume $u=u_{a_1}=u_{a_2}=u_{a_3}=u_{a_4}>0$ and $u_{a_5}=\epsilon$ with $\epsilon>0$. Given this parameterization, it is easy to see that $\mathcal{W}(\boldsymbol{\kappa})= \log 2e^{2 u}=\log 2+2u$ and $\tilde{\mathcal{W}}(\boldsymbol{\kappa})=\log \left(e^{2u}+e^{2u-\kappa_{i_1}\log 2}+e^{2u+\epsilon-\kappa_{i_1}\log 2}\right)=2u+\log\left(1+e^{-\kappa_{i_1}\log 2}(1+e^\epsilon)\right)$. Then $\tilde{\mathcal{W}}(\boldsymbol{\kappa})>\mathcal{W}(\boldsymbol{\kappa})$ if and only if $$ {e^{u+\epsilon}\over e^{u}+e^{u+\epsilon}}={e^{\epsilon}\over 1+e^{\epsilon}}>1-\left(\frac{1}{2}\right)^{\kappa_{i_1}}$$ which can be rewritten as $$\kappa_{i_1}<{\log (1+e^\epsilon)\over \log (2)}\triangleq\alpha.$$ It is easy to see that for $\epsilon>0$ we have $\tilde{\mathcal{W}}(\boldsymbol{\kappa})>\mathcal{W}(\boldsymbol{\kappa})$ if and only if $\kappa_{i_1}<\alpha$.\footnote{Note that for $\epsilon>0$ we have $\alpha>1$$ $.}
As we mentioned in \S(ref), the case of $\kappa_{i_1}\geq 1$ corresponds to a situation where users only want to add alternatives that are sufficiently better that the existing ones. In this case, network users want to add edge $a_5$ whenever $\kappa_{i_1}\in (0,\alpha)$. In particular, for values of $\kappa_{i_1}\geq \alpha$, improving the network can make all users worse off. Thereby, under choice aversion, we can obtain the very same type of paradox described by Braess1 and Braess2. However, while this paradox is usually obtained as a consequence of network users' selfish behavior in a congestion game, we derive it as a consequence of choice aversion, which provides a different perspective to the problem of network design.
In this section, we move beyond calibration of the choice aversion model and econometrically estimate the choice aversion parameter $\kappa$. Using observed data on paths taken across a given real-world network structure, we can employ maximum likelihood estimation (MLE) procedures to estimate $\kappa$. Previous literature on RL models in transportation networks use MLE routines to analyze the performance and strength of existing models (Fosgerauetal2013, MAI2015100, ZIMMERMANN2017183, ZIMMERMANN2020100004). Following this precedence, we exploit a similar approach, incorporating the choice aversion penalization in the utility function of the network user to estimate the value of $\kappa$ alongside other network attributes.\footnote{The MATLAB code used for the empirical exercises in this section are available online or upon request, along with all other code used for this paper.}
Through the following empirical analysis, we demonstrate a set of key strengths of the choice aversion model. First, we show that the obtained estimates for the choice aversion parameter are similar to calibrations given for the simulations on example networks, thus validating the legitimacy of the model and previous inferences from such exercises. Second, we show that the choice aversion model exhibits a similar performance to other RL models in terms of mechanical path prediction on transportation flows. Finally, because of the parsimonious nature of the choice aversion penalty, we show that the model is less computationally demanding, performing at much faster speeds than other corrective additions to the RL model.
We first test the choice aversion model on an illustrative example network used as a tutorial for estimation of the RL model (ZIMMERMANN2017slides) and depicted in Figure (ref). This network contains 17 nodes and 29 edges, with 2 origin-destination source-sink pairs (1 to 17 and 2 to 17). There are 100 synthetic path observations over this network generated by a linear utility specification with three attributes: link length, existence of a traffic signal, and a link constant.
For this network, we estimate the traditional RL model, a RL that incorporates the choice aversion attribute, and a RL model with the Link Size attribute introduced by Fosgerauetal2013 and discussed in Section (ref). For the straightforward RL specification, we follow the original authors and estimate coefficients for link length, traffic signal, and a link constant. In addition to this specification, we incorporate the choice aversion term to estimate $\kappa$, the choice aversion parameter.\footnote{For simplicity, we assume that $\kappa$ is homogeneous across nodes in the network. We anticipate future empirical work to relax and test this condition.} For the sake of comparison, we also incorporate the Link Size attribute, which allows users to account for expected flows based on link attributes. We test these specifications both separately and together, as well as allowing for an interaction term between choice aversion and Link Size. These specifications are detailed formally in Eq. ((ref)).\footnote{Note that $u_a + \kappa_a \log |A_{j_a}^+|$ for $\kappa_a < 0$ is equivalent to $u_a - \kappa_a \log |A_{j_a}^+|$ for $\kappa_a > 0$. We have made this modification for the empirical section of the paper to be consistent with testing our hypothesis alongside other attributes in the RL framework.}
Table (ref) displays the estimates obtained through the MLE routine conducted across the network for each specification previously mentioned. As may be expected for this tutorial data, the coefficients for link length and traffic signal are the only attributes where significant estimates are consistently observed across specifications.
In contrast, it does not appear that either choice aversion or Link Size are important attributes to network users under this construction, although they appear to cut into the explanatory power of the preceding variables in some specifications. Given this, we note that the estimated choice aversion parameter $\hat{\kappa}_{CA}$ is consistently negative, while the Link Size attribute ranges between positive and negative values. Additionally, we note that incorporating the Link Size attribute in the specification dramatically decreases the speed of the MLE routine.
Next, we test our model on GPS vehicle data observed in the city of Borl{\"a}nge, Sweden. This traffic data set is well known and has been previously used to test RL models in Fosgerauetal2013 and MAI2015100. Over a span of two years, a sample of 1,832 trips (using a minimum of five links) were observed across a network with 3,077 nodes and 7,459 links (21,452 link pairs). This network includes 466 unique destinations and 1,420 different origin-destination pairs. Most importantly, network users have over 37,000 link choices, making this real-world network a suitable application for testing the choice aversion model.
Using this data, we proceed with the standard RL model with and without choice aversion and the Link Size attribute as in Section (ref). The specifications used for the MLE routines in this section are identical to those shown by Eq. ((ref)), except there is no traffic signal attribute to this data.\footnote{Fosgerauetal2013 and MAI2015100 include additional network attributes regarding turn angles in their empirical analysis. We omit these from our analysis for the sake of simplicity and clarity.}
Table (ref) displays the results for each specification on the Borl{\"a}nge network. For this data, the estimated choice aversion parameter $\hat{\kappa}_{CA}$ is significant and negative, ranging between $-0.199$ and $-0.628$. This result provides evidence that network users are indeed choice averse and to a degree that is consistent with calibrations presented in Sections (ref)-(ref). We note that $\hat{\kappa}_{CA}$ and $\hat{\beta}_{LS}$ are similar in range of values and are both significant across specifications, demonstrating that choice aversion has the capability of providing predictive power on transportation networks like Link Size and other additions to the RL model.
We note that for a complex network such as Borl{\"a}nge, choice aversion does not even double the speed of the MLE routine, while incorporating the Link Size attribute may cause the time required for convergence to increase exponentially. While we also observed a decrease in speed of Link Size for the tutorial data, the decline observed with this data is much sharper moving from the RL to the RL-LS (by over a factor of 75). From an applied perspective, the combination of speed, tractability, and economic interpretation of the choice aversion model should prove attractive when compared to other RL models.
The last column in Table (ref) displays a coefficient for Link Size that is not significant, while also displaying a negative and significant estimate for the interaction term $\hat{\kappa}_{\text{CA} \times \text{LS}}$. It is intuitive that expected flows and choice aversion are in tandem with one another in terms of affecting individual routing decisions. We interpret this result as an indicator that there is more to understand about the role of choice aversion among users in complex, congested networks. We leave this discussion to be explored more deeply in future empirical and experimental work.
Finally, we note that using the RL-LS rather than the RL-CA results in a larger improvement of the log-likelihood compared to the RL, suggesting that the RL-LS may outperform the RL-CA in this regard. In Appendix (ref), we conduct an out-of-sample validation analysis to compare performance across RL models. The main conclusion of this exercise is that the RL-LS and RL-CA exhibit similar out-of-sample performance.
The recursive choice model with choice aversion is a highly tractable extension of the standard RL model that can be used to predict reasonable route choice probabilities and provide welfare interpretations in transportation networks. Upon testing our approach against existing PSL models, our model exhibits the power to provide reasonable corrections and predictions with the benefit of a microfoundation in choice overload.
In addition, we explore how the choice aversion model allows for a break in the regularity condition typically preserved by path choice probabilities in logit models and extensions. In doing so, we show that, conditional on the degree of choice aversion, removing edges in the network can lead to a decrease in choice probability of certain existing paths.
We also simulate the welfare implications of the choice aversion model and find a novel prediction: even in uncongested networks, a decrease in welfare akin to Braess's paradox can arise when costless edges are added. Here, we also provide a simple characterization for welfare changes conditional on choice aversion which is testable in empirical settings.
Furthermore, we estimate our model providing evidence that users in transportation networks suffer from choice aversion. An important extension of this work is to consider the role of choice aversion in the context of congested traffic networks and the respective modeling approach. One way to study this question is to extend the model in bc by introducing choice aversion in their recursive approach.
Finally, we remark that there is much work to be done regarding the empirical support of choice aversion in transportation networks. We anticipate there is more to understand regarding the role of node-specific (or nest-specific) choice aversion within complex real-world environments. Similarly, we point out that our empirical analysis assumes a homogeneous choice aversion parameter. However, in real-life transportation networks, it is reasonable to assume that $\kappa$ depends on users' characteristics. We leave for future research the econometric analysis of this type of model.
Future work regarding experiments on behavior of participants in transportation networks would help establish a better understanding of the significance of choice overload in making routing choices.