Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
77,490 characters · 17 sections · 68 citation commands
Behavioral Foundations of Nested Stochastic Choice and Nested Logit
{{\bf Keywords:} Nested Logit; Nested Stochastic Choice; Luce Model; IIA; Similarity Effect; Regularity; Revealed Similarity; Cross-Nested Logit; Nest Identification.}
{\bf JEL:} D01, D81, D9.
Nested logit (Ben-Akiva1973nestedlogit, mcfadden1978modeling) is the most widely applied generalization of multinomial logit (or Luce's (1959) model) due to its ability to capture various substitution patterns.\footnote{Nested logit has been used to study transportation demand (anderson1992multiproduct, forinash1993application), airline competition (lurkin2018), automobile demand (brownstone1989efficient, goldberg1995product), telephone use (train1987demand, lee1999calling), and much more. See Chapter 4 of train2009discrete for an excellent discussion.} In nested logit, each alternative belongs to a nest (or subset) of “similar" alternatives, and choice may be decomposed into two Luce procedures: the probability that $a$ is chosen from menu $A$ is the probability that $a$'s nest is chosen from among nests available in $A$ multiplied by the conditional probability that $a$ is chosen from that nest. Despite its immense importance, nested logit has escaped behavioral characterization. In this paper, we provide the first behavioral characterization of nested logit through the introduction of a fully non-parametric version that we term Nested Stochastic Choice (NSC). This axiomatic characterization sheds light on the implicit assumptions behind nested logit and related models and leads to a tractable method to identify (unobserved) nests from data.
Nested logit was developed to address the limitations of multinomial logit when dealing with “similar alternatives.” In the Luce model, choice probabilities are proportional to a utility index and hence satisfy Independence of Irrelevant Alternatives (IIA): probability ratios are menu-independent. The similarity effect (debreu1960review, tversky1972eba) is a violation of IIA in which the introduction of an alternative to a menu has a much larger effect on the choice probabilities of alternatives of a “similar type” than on those of a “different type.” Nested logit allows the similarity effect by assuming nested (similar) alternatives are more substitutable (e.g., because they receive a correlated utility shock). Our main result shows that nesting of similar alternatives, a key behavioral feature of nested logit (and also NSC), is captured by a weakening of IIA that allows for the similarity effect. This finding reveals a deep connection between nested logit and the similarity effect.
To illustrate the similarity effect, consider the red bus/blue bus example from debreu1960review. A commuter making a choice between a red bus and a train may choose either with probability $0.5$. If the option to take a blue bus is introduced, it is plausible that it will have no effect on the commuter's likelihood of taking the train; the blue bus only affects the probability of selecting the red bus (e.g., by reducing it to $0.25$). The intuition behind this is that the buses are similar to each other in a way that neither one is to the train. Nested logit handles this by placing the two buses into a bus nest. While this example provides an extreme case of the similarity effect (the buses are perfect substitutes, or “duplicates”), the principle that “similar alternatives affect each other” readily extends to many situations of interest to economists, firms, and policy-makers, such as a consumer's choice of vehicle or apartment.\footnote{In these cases, the correct nest specification is not easily observed by the analyst. For instance, apartments in a city might be nested based on subjective neighborhoods, which may depend on a variety of factors. In turn, some of these factors may be observable while others may be subjective or difficult to observe.}
Our analysis of nested logit relies on the introduction of NSC, a non-parametric version of nested logit. Formally, a stochastic choice function $p$ is an NSC if there exist nests $X_1, \ldots, X_K$ that partition the set of all alternatives $X$ and functions $v$ and $u$ such that, for each choice set $A$, the probability of choosing $a \in A \cap X_i$ is given by
The NSC is defined by two Luce rules where the attractiveness of the nest $A\cap X_i$ is measured by $v(A\cap X_i)$, and the attractiveness of the alternative $a$ is measured by $u(a)$ (or simply Luce's utility of $a$). In terms of the red bus/blue bus example, $v$ governs the choice between general modes, “a bus” or “a train,” while $u$ governs the choice between specific alternatives, the red or blue bus.
Notice that nested logit is the special case of NSC in which, for each $i\le K$,
Hence, NSC is nested logit without any assumptions on the relationship between $v$ and $u$. It is commonly assumed in the applied literature that $\eta_i\le 1$, as this ensures that nested logit is a Generalized Extreme Value (GEV) model mcfadden1978modeling, and therefore this restriction is sufficient for consistency with the random utility model (RUM). However, this restriction ($\eta_i\le 1$) is not necessary for (ref) to be a RUM and nested logit has been estimated without this restriction.\footnote{This parameter restriction is sometimes referred to as the Daly-Zachary-McFadden condition (Daly1978, mcfadden1978modeling), as they showed that this is sufficient for consistency with RUM for arbitrary values of the other variables (e.g., utilities). However, this restriction is not always imposed. For instance, train1987demand provide estimates of a model for which $\eta_i > 1$, remarking that it represents greater substitutability between nests than within nests (see also train1989, lee1999calling, Foubert2007). Further, nested logit with $\eta_i>1$ may still be consistent with RUM (see borsch1990 and herriges1996).} Therefore, for simplicity of exposition we will refer to (ref) as a nested logit for any parameter value. To provide further clarity, we may sometimes refer to nested logit satisfying the restriction ($\eta_i\le 1$) as the random utility (RU) nested logit. We provide behavioral foundations for NSC and nested logit, along with a characterization of random utility nested logit.
We utilize a revealed preference approach to identify the subjective/endogenous nest structure of the NSC. This is achieved by introducing a notion of revealed similarity. In nested logit, if two alternatives are in the same nest, then their probability ratio will always be independent of other alternatives. Consistent with nested logit, we use this insight to define a notion of similarity: alternatives $a$ and $b$ are revealed categorically similar, denoted $a \sim_p b$, if IIA holds between $a$ and $b$ at any menu. Otherwise, they are revealed categorically dissimilar. Therefore, the core notion of similarity underpinning nested logit is binary: alternatives are similar or not.
Equipped with this notion of reveled similarity, we can weaken IIA to allow for the similarity effect. To do so, we decompose IIA into two axioms. The first axiom, \nameref{ISA} (ISA), imposes an IIA condition between $a$ and $b$ in the presence of a third alternative $x$, when $x$ is symmetrically related to $a$ and $b$ in terms of revealed similarity (i.e., both are revealed categorically similar or dissimilar to $x$). In terms of the red bus/blue bus example, IIA should hold between the buses in the presence of the train. The second axiom, \nameref{IAA} (IAA), “completes” ISA by imposing an IIA condition when $x$ is asymmetrically similar to $a$ and $b$ (i.e., $x$ is similar to one and not the other). For example, IAA implies that the introduction of the blue bus impacts the red bus and the train equally, which directly rules out the similarity effect.
Our main result is that ISA characterizes NSC ((ref)). Further, since ISA is the minimal departure from IIA that allows for the similarity effect (i.e., violations of IAA), our analysis reveals that NSC is the model obtained when IIA is relaxed to allow for the similarity effect.
To see how our axiomatization provides a clearer picture of nested logit, recall its standard textbook description. For instance, Chapter 4 of train2009discrete states that for nested logit “IIA holds over alternatives in each nest and independence of irrelevant nests\footnote{If $a$ and $b$ are from distinct nests, then the addition of an alternative $c$ from a third nest will not affect the relative probabilities of $a$ and $b$.} (IIN) holds over alternatives in different nests.” NSC also satisfies these properties. Indeed, ISA ensures the existence of endogenous nests and imposes exactly these properties on them. But this finding shows that “IIA within a nest $+$ IIN” are not sufficient for a nested logit representation, as ISA is equivalent to NSC. Since there are missing behavioral assumptions behind nested logit, the textbook description is incomplete.
To provide a complete picture of nested logit, we establish two characterizations of nested logit as well as a characterization of random utility nested logit. Our first characterization shows that an NSC is a nested logit if and only if it satisfies \nameref{LRI}. This axiom requires that the natural logarithm of certain probability ratios featuring collections of similar alternatives is menu-independent. The explicit use of a functional form in the axiom allows us to establish necessary and sufficient conditions for the functional form assumed in nested logit with finite data.
Our second characterization is based upon a novel monotonicity condition, \nameref{RLI}, that is necessary for nested logit and becomes sufficient under a mild richness assumption. To understand this axiom, note that nested logit ((ref)) requires that the attractiveness of a nest is increasing in the sum of utilities. \nameref{RLI} implies that this feature must hold in a relative sense; the attractiveness of a nest relative to another nest is increasing in the sum of utilities, holding the alternatives in the other nest fixed. Finally, random utility nested logit is characterized by one additional axiom, Regularity, a well-known monotonicity property that all random utility models must satisfy (under the same richness assumption). We summarize all of our characterization results in (ref).
\nameref{ISA} implies that the revealed similarity relation $\sim_p$ is transitive, which ensures that the nests form a partition. Hence, our axiomatic characterizations show that the notion of (categorical) similarity in nested logit, as well as in NSC, is quite structured. In some applications, an analyst may want to allow for more flexible forms of substitutability. For this reason, cross-nested logit, a generalization of nested logit, has been proposed and widely applied in empirical work (see vovsha1997application, ben1999discrete). The main difference between nested logit and cross-nested logit is that alternatives may belong to several nests in cross-nested logit. Since we are focused on the problem of recovering the nesting structure, we consider a generalization of cross-nested logit that relaxes the typical parameter restriction, and refer to this as the unrestricted cross-nested logit. We show that the unrestricted cross-nested logit does not have testable implications. Therefore, our results reveal the trade-off between nested logit and cross-nested models: relaxing the partition structure of nested logit results in an overly permissive model. In other words, the behavioral content of cross-nested logit is essentially driven by the analyst's assumption of the nest structure and parameter restrictions.
In practice, $p$ is estimated from observed choices and “IIA like" conditions never hold exactly. However, we show that the true nest structure can still be identified, for any NSC, by solving a minimization problem. In particular, our axiomatic characterization allows us to derive a “distance" function $D$ that measures, for a given nest structure, the degree of violations of IIA within and across nests for a given set of observations. When the data are close to the true (or theoretical) $p$, the true nest structure will be the unique minimizer of $D$. In applied settings where the researcher has several nest structures in mind, $D$ may also be useful as a selection criteria.
Because the number of possible nests grows exponentially as the number of alternatives increases, the full minimization problem may become intractable quickly. However, this issue can be managed due to insights from our similarity relation; one only needs to check nest structures that are consistent with an empirical approximation of $\sim_p.$ In fact, we show that there are at most $|X|$ potential nests that we need to check, where $|X|$ is the number of alternatives. We illustrate our theoretical finding and our data-driven algorithm to reduce the number number of candidate nests with a simulation exercise.
The rest of the paper is organized as follows. In (ref), we discuss setup and notation as well as define NSC and nested logit. In (ref), we define revealed similarity and the similarity effect ((ref)) before characterizing NSC ((ref)) and nested logit ((ref)). We discuss ways of extending our notion of similarity and the testable implications of unrestricted cross-nested logit in (ref). The identification of nest structure from choice data is presented in (ref). We conclude with a discussion of related literature in (ref). We also discuss the relationship between the similarity effect and regularity in Appendix (ref).
All of our models are developed in the standard stochastic choice setup. Accordingly, let $X$ be a finite set of alternatives and $\mathscr{A}$ be the collection of all nonempty subsets of $X$ (menus). Let $\mathds{R}_{+}$ ($\mathds{R}_{++}$) denote the non-negative (positive) real numbers.
Throughout this paper, we assume that $p$ is positive; i.e., $p(a, A)>0$ for all $A\in\mathscr{A}$ and $a\in A$. For notational simplicity, we write $A\cup x$ instead of $A\cup \{x\}$.
The Luce model is the most widely-known and influential stochastic choice model. In this model, choice probabilities are proportional to a utility index: $p(a, A)=u(a)/\sum_{b\in A} u(b)$. In his seminal paper, luce1959individual proves that a stochastic choice function can be represented by the Luce model if and only if it satisfies IIA for every pair of alternatives.
It is well known that IIA may fail when similar alternatives are added to the menu, as was illustrated by Debreu's (debreu1960review) famous “red bus/blue bus” example. The nested logit is the most commonly applied generalization of Luce's model and was developed to accommodate violations of IIA like the similarity effect. We now formally define nested logit and the NSC, which is a novel, non-parametric version of nested logit.
The NSC is defined by two Luce procedures, where $v$ governs the choice over nests (e.g., transportation modes or neighborhoods) and $u$ governs the choice over the particular alternatives in the selected nest (e.g., the red/blue bus or a specific apartment). Note that the nest value function $v$ is not necessarily related to alternative utilities $u$, which enables the NSC to capture rich behavior (see (ref)). Despite this generality, the NSC may be falsified with relatively few observations. This is because behavior is disciplined by $u$ and the partition structure of the nests, both of which are menu-independent.\footnote{It is straightforward to derive from the representation that for any $A\subseteq X$ with $|A|=3$, there is a distinct pair $a, b\in A$ such that $\frac{p(a, \{a, b\})}{p(b, \{a, b\})}=\frac{p(a, A)}{p(b, A)}$. This is because for any three alternatives, either (i) at least two belong to the same nest or (ii) all three belong to distinct nests. Hence there must exist some pair for which IIA holds and therefore NSC may be rejected with only three alternatives.}
The nested logit imposes a specific parametric relationship between $v$ and $u$. Notice that the NSC, and consequently nested logit, reduces to the Luce model when there is a single nest. Additionally, it is simple to see from (ref) that any NSC satisfies “IIA within a nest $+$ IIN." These properties are often taken as the hallmark of nested logit, yet they apply to all NSC (with endogenous nests). Since NSC permits behavior that nested logit does not (three examples are discussed in section (ref)), this means that there are additional behavioral assumptions underpinning nested logit. We elucidate these assumptions in section (ref).
In nested logit, $1-\eta_i$ is usually considered a measure of correlation or substitutability between alternatives in nest $i$. When $\eta_i < 1$ alternatives within the same nest are substitutes. Further, it is well known that when $\eta_i < 1$, the nested logit is always a RUM for any profile of utilities.
When $\eta_i >1$, choice frequencies may (but do not always) violate regularity, a necessary property of every random utility model (RUM) which states that the probability of choosing some alternative must never increase as the menu expands.\footnote{There is some experimental evidence that violations of regularity occur when similar alternatives are introduced, in-line with convex aggregation. This has been observed in humans \citep*{rieskamp2006extending} and animals \citep*{shafir2002}. Recently, Batley2016 estimated nested logit parameters to check for consistency with regularity (and various forms of stochastic transitivity) and found that parameter values consistent with violations of regularity provided the best fit.} Behaviorally, we can interpret $\eta_i >1$ as indicating complementarities among alternatives.\footnote{Relatedly, ortoleva2019deliberate suggests that regularity may be violated due to deliberate randomization between complementary lotteries.} In certain contexts, we may even anticipate $\eta_i >1$. For instance, Foubert2007 study the effects of “product bundling” and find a parameter greater than one, consistent with the effectiveness of bundling.\footnote{Indeed, in regard to whether $\eta_i$ should be less than or greater than one, train1987demand state that “...the value of [$\eta_i$] indicates relative substitutability within and among nests, and neither possibility can be rule out a priori.”} As NSC allows for violations of regularity, formally defined below, the NSC is not nested by RUM.
Finally, note that when $\eta_i=1$ the nest value is exactly proportional to the sum of Luce utilities. If this proportionality happens for every nest $i$, the model reduces to Luce. In fact, in this case the Luce model has multiple NSC (and nested logit) representations with different partitions and identification of a unique nest structure is not possible. To rule this out, we say that an NSC $p$ with $(v, u, \{X_i\}^K_{i=1})$ is nondegenerate if there is at most one nest where this proportionality occurs: there is at most one $i\le K$ such that for some $a\in X_i$, \[\frac{\sum_{x\in A_i}u(x)}{v(A_i)}= \frac{u(a)}{v(a)}\text{ for any }A_i\subseteq X_i\text{ with }a\in A_i.\]This restriction rules out cases when $v$ is always proportional to the sum of Luce utilities. Further, the Luce model has a unique nondegenerate NSC representation in which there is a single nest, $X_1=X$.\footnote{Indeed, if there are $i, j$ such that $\frac{\sum_{x\in A_i}u(x)}{v(A_i)}=\frac{u(a)}{v(a)}$ and $\frac{\sum_{y\in A_j}u(y)}{v(A_j)}=\frac{u(b)}{v(b)}$ for any $A_i\subseteq X_i, A_j\subseteq X_j$, $a\in A_i$, and $b\in A_j$, then $X_i\cup X_j$ should be treated as one nest. It is also not difficult to show that the set of degenerate NSC is measure zero with respect to the set of all NSC.} This nondegeneracy condition will be crucial for the unique identification of nests, but it is not required for the sufficiency part of our characterization (Theorem 1).
Following the intuition behind the similarity effect, we introduce a notion of revealed similarity that will be essential to our analysis. Consider the effect of adding an alternative $x$ on the probabilities of choosing $a$ and $b$ from some menu $A$. Adding $x$ might decrease these probabilities as it competes with $a$ and $b$. If $x$ disproportionately affects one of them, say $a$ relative to $b$, this reveals that $a$ and $b$ are dissimilar. Conversely, if $x$ takes away from $a$ and $b$ proportionally, then this reveals that $a$ and $b$ are similar (symmetric) in menu $A$. We take a conservative approach and call two alternatives similar only if this is true for any menu $A$ (they are symmetric to all other alternatives).
The similarity effect is often defined using an exogenously given similarity relation. With our formal notion of revealed similarity, we may establish a fully behavioral definition of the similarity effect given $\sim_p$.\footnote{There are other ways to define similarity and other properties one might demand of a similarity relation. For instance, rubinstein1988similarity studies similarity and choice under risk. In his paper, the similarity relation is reflexive and symmetric, like ours, but also must satisfy a form of betweenness with respect to objective attributes and violates transitivity, unlike ours. natenzon2018random introduces a notion of comparative similarity based on absolute rather than relative choice frequencies. This similarity notion is not related to IIA and will not induce a partition structure on the set of alternatives.} Recalling the red bus/blue bus example, adding the blue bus had a larger effect on the red bus than on the train. Hence the blue bus “takes more away” from similar alternatives than from dissimilar alternatives.
Intuitively, $x$ hurts the revealed categorically similar alternative $a$ more than a revealed categorically dissimilar alternative $b$. Since $a$ and $x$ are closer substitutes, $x$ competes more with $a$ than it does with $b$.
In order to introduce our axiom, we consider the general effect of introducing an alternative $x$ on the choice probabilities of two alternatives $a$ and $b$. IIA requires that the relative probability between $a$ and $b$ is always independent of $x$. However, as the similarity effect suggests, (asymmetric) similarity between $x$ and $a, b$ might affect the relative probabilities. We therefore divide IIA into two logically independent axioms based on the revealed similarity between $x$ and $a, b$.
The first axiom requires that the relative probability between $a$ and $b$ is independent of $x$ when $a$ and $b$ are revealed categorically (dis)similar to $x$. Intuitively, if $a$ and $b$ are symmetric from the perspective of $x$, then $x$ should symmetrically influence $a$ and $b$; it does not affect the relative probability between $a$ and $b$. The similarity effect directly contradicts the second axiom yet is unrelated to the first axiom.
We show in (ref) that NSC is characterized by \nameref{ISA}, and thus NSC is precisely the generalization of Luce's model that accommodates the similarity effect.
(ref) characterizes NSC when there are at least three nests; $a\not\sim_p b$, $b\not\sim_p c$, and $a\not\sim_p c$ for distinct alternatives.\footnote{When the assumption is violated (i.e., there are only two nests), we can still obtain the characterization result by modifying \nameref*{ISA}. In particular, we can impose a modification of Luce's (1959) Product Rule instead of the second part of \nameref*{ISA}. It is well known that IIA is equivalent to the Product Rule for menus with two alternatives (see Luce (1959)).} While the proof is in the appendix, we discuss briefly how our axiom characterizes NSC. It should be apparent from the definition that $\sim_p$ is reflexive and symmetric. It turns out that the first part of \nameref{ISA} ($a\sim_p x$ and $b\sim_p x$) implies that $\sim_p$ is transitive.\footnote{More general notions of similarity may be intransitive (e.g., due to context dependence). Since we take a conservative definition of similarity, we find transitivity quite reasonable in our setting. That is, by requiring $\frac{p(a, A)}{p(b, A)}=\frac{p(a, \{a, b\})}{p(b, \{a, b\})}$ for any menu, we eliminate much of the context dependence. Further, transitivity of this revealed similarity relation is implicitly assumed in nested logit. See (ref).} Hence, transitivity of $\sim_p$ immediately generates a partition $X_1, \ldots, X_k$ of $X$ (or disjoint nests) such that any two alternatives in $X_i$ are revealed categorically similar.\footnote{Transitivity of $\sim_p$ is imposed in li2016associationistic, which will be carefully discussed in (ref). A weak version of transitivity of $\sim_p$ is also used in echenique2018perception.} However, by itself it imposes no particular structure on choice, nor does it establish a relationship between the partition and choices (except that IIA is satisfied within each nest). The essential structure of NSC is captured by the second part of \nameref{ISA} ($a\nsim_p x$ and $b\nsim_p x$). Therefore, almost all of the proof is devoted to showing that the second part of \nameref{ISA} implies a nested choice structure consistent with this partition.
Lastly, we state the uniqueness properties of the NSC representation. The following proposition shows that the nest structure is unique, the nest utility $v$ is unique up to a positive scalar, and Luce's utility $u$ is unique up to a positive scalar at each nest.
The most well-known special case of NSC is nested logit, which was specifically created to handle the similarity effect. The difference between NSC and nested logit is that the latter imposes a special structure on the nest values, $v(A\cap X_i)=\big(\sum_{a\in A\cap X_i} u(a)\big)^{\eta_i}$ with $\eta_i>0$, which has non-trivial behavioral consequences.
Despite its widespread use, nested logit has not been subject to careful axiomatic analysis in the way that other choice models have been. We provide two characterizations of nested logit that clarify the behavioral assumptions embedded in this model. The first characterization uses one additional axiom that imposes a menu independence condition on certain probability ratios.
Log Ratio Invariance requires that the ratio $\log\!\big(\frac{p(A,\, A\,\cup\, x)}{p(x,\, A\,\cup\, x)}\big{/}\frac{p(a,\, \{a,\, x\})}{p(x,\, \{a,\, x\})}\big)$ and $\log\!\big(\frac{p(A, \,A\,\cup\, a)}{p(a, \,A\,\cup\, a)}\big)$ are proportional.
The explicit use of a functional form in Log Ratio Invariance allows us to establish testable implications for the functional form assumed in nested logit even with finite data.
In the rest of this section, we discuss an alternative axiom that captures the essential features of nested logit without an explicit functional form and shows that is characterizes nested logit under a richer domain assumption. To state this axiom, first notice that an important behavioral property of nested logit, beyond its treatment of similarity (ISA), is that the probability of choosing a nest depends on the total attractiveness of the nest: $v(A\cap X_i)$ is increasing in $\sum_{a\in A\cap X_i} u(a)$.
This behavior is characterized by a simple monotonicity property imposed among similar alternatives. Suppose that $A, A'\in\mathscr{A}$ are menus such that all alternatives in $A\cup A'$ are revealed similar. When $A$ is more attractive than $A'$, then alternatives in $A$ are always chosen more frequently than alternatives in $A'$ when they are compared with any other alternative $x$. More formally, $p(A, A\cup A')\ge p(A', A\cup A')$ implies $p(A, A\cup x) \ge p(A', A'\cup x)$ for any $x\not\in A\cup A'$. This can be viewed as an additional form of context independence, as it requires that there is no interaction between alternatives in $A\cup A'$ and $x$ which might create a choice frequency reversal.
Because nested logit involves a power function, it satisfies a stronger version of the monotonicity property above. In particular, the monotonicity property holds even in relative terms: if $A$ is relatively more attractive than $A'$ when they are compared to any other menus, $B$ and $B'$, then alternatives in $A$ will be chosen relatively more frequently than $A'$ when they are chosen against $x$.
In our next result, we prove that Relative Likelihood Independence is a necessary condition for nested logit. Moreover, it implies that $v(A\cap X_i)$ is increasing in $\sum_{a\in A\cap X_i} u(a)$.
While \nameref{RLI} is not sufficient for nested logit, this is essentially due to the limitations of finite data. Indeed, we show that Relative Likelihood Independence characterizes nested logit when the following richness condition is satisfied.
On its own, \nameref{richness} is relatively mild and simply ensures that there are alternatives for each utility value. However, under this condition \nameref{RLI} yields the well-known functional form of nested logit. Consequently, \nameref{RLI} captures all remaining behavioral features of nested logit.
In applied settings, nested logit is often restricted to $\eta_i \in (0,1]$, as this is sufficient for it to be a RUM. Since any RUM satisfies \nameref{reg}, a random utility nested logit must as well. It is well known that \nameref{reg} is necessary but not sufficient for a model to be a RUM in general. However, we show that \nameref{reg} is sufficient for the nested logit to be a RUM under \nameref{richness}.
If $\eta_i > 1$, nested logit will violate regularity for some specifications of $u$. Therefore \nameref{richness} strengthens the bite of \nameref{reg} and $\eta_i \in (0,1]$ is ensured.
It is a matter of folk knowledge that the aforementioned “IIA within a nest and IIN” properties serve as the behavioral underpinnings of nested logit. However, our results show that “IIA within a nest and IIN” (with an endogenous nest structure) in fact characterize NSC, and nested logit requires an additional property (\nameref{RLI}). In this subsection, we present three special cases of NSC that are distinct from nested logit. These examples illustrate some natural choice behaviors that \nameref{RLI} rules out, further clarifying the behavioral assumptions behind nested logit.
One interesting example of NSC that is distinct from nested logit is the Linear NSC. In this example, the nest value is linear in total nest utility, in contrast to the power function used in nested logit. For each nest $i$, there exist parameters $\lambda_{i}\ge 0$ and $\nu_{i},$ so that
In the Linear NSC, the attractiveness of a nest depends on both its instrumental utility, through $\lambda_i$, and an intrinsic “category” utility, through $\nu_i$. It turns out that the Linear NSC is a special case of both Elimination-by-Aspects (EBA) of tversky1972eba and the Attribute Rule (AR) of gul2014random. Consequently, the Linear NSC is also a RUM.
In nested logit, the nest parameter $\eta_i$ captures substitutability of alternatives. While the standard nested logit only allows for a single substitution parameter for each nest, NSC can accommodate menu-dependent substitutability. For instance, consider the following example where substitutability depends on the size of the menu, capturing the idea that consumers may find it more difficult to distinguish between alternatives in larger option sets.
For each nest $i$, there exists a threshold $\tau_i \in \{1,\ldots,|X_i|\}$, and nest parameters $\eta_{i}, \tilde{\eta}_{i} > 0,$ so that
If $1-\eta_{i} > 1-\tilde{\eta}_{i}$, this means the agent perceives fewer differences among alternatives as the nest becomes “more represented.” That is, when $ |A\cap X_i|$ exceeds $\tau_i$, alternatives are “more substitutable.” In this example, $\tau_i$ has a natural interpretation as the consumer's “distinction capacity.” Further, if $1-\eta_{i} > 0 > 1-\tilde{\eta}_{i}$, then whether the alternatives are complements or substitutes depends on the size of the nest. Lastly, when $\eta_i$ and $\tilde{\eta}_{i}$ are both less than one and are sufficiently close, this example is also consistent with RUM.
The NSC also allows for cases where the nest value is not directly tied to alternative utility. We consider a particular example in which $v$ is determined by the “salience” of alternatives. For some function $S:X \rightarrow \mathbb{R}_{++}$,
In this specification, the value of a nest is determined by the “attractiveness” or “noticeability” of its most salient alternative. To illustrate its behavioral implications, suppose there are three alternatives, $X=\{x,y,z\}$, with nests $X_1=\{x,z\}$ and $X_2=\{y\}$. When $z$ is highly salient but low utility, $S(z) u(x) > S(x)[u(x)+u(z)]$, then $\frac{p(x,\{x,y,z\})}{p(y,\{x,y,z\})} > \frac{p(x,\{x,y\})}{p(y,\{x,y\})}$. Examples of such $z$ include brands offering a high-end good with a high price to attract attention, expecting all consumers to purchase their “moderate” offering $x$. When the value of $S(z)$ is large enough relative to the value of $u(z)$, this may induce a violation of regularity. Similar examples can generate “spillover” effects. For example, one successful or attractive product may funnel attention to others, causing demand spillover. This is the traditional rationale behind the use of “loss-leaders" (lal1994) or “attention-grabbers” (eliaz2011attention).
We say that two alternatives are revealed categorically similar if IIA is satisfied between them at all menus. Requiring this eliminates the menu dependence of similarity, and so our notion of revealed similarity captures a form of absolute or fundamental similarity. Consequently, similarity is symmetric and, under \nameref{ISA}, transitive. One drawback is that this notion does not allow statements about comparative similarity; two alternatives are similar or not. Additionally, in some cases impressions of similarity may be context-dependent.\footnote{There is a sense in which our notion is somewhat moderate. Consider Debreu's red bus/blue bus example. In this case, the similar alternatives (buses) are in fact identical, often called duplicates (or in some cases replicas). Duplicates are not merely similar alternatives; they are similar and provide the exact same utility value. For example, in gul2014random the use of duplicates is essential to their characterization of the Attribute Rule. Formally, $x$ and $y$ are duplicates if $p(a, A\cup x)=p(a, A\cup y)$ for any $A$ and $a\in A$. However, our notion of similarity is not tied to utility. A commuter may regard all buses as similar (i.e., they belong to the same nest), yet nothing in our model restricts an agent from exhibiting a preference over different buses (e.g., because some bus routes may be faster or cheaper than others).} Because of these apparent limitations, we consider two ways in which to extend NSC to accommodate more complex notions of similarity.
The first extension of NSC relaxes the requirement that an alternative must belong to a single nest. In section (ref), we consider the (unrestricted) cross-nested logit vovsha1997application and the (more general) generalized nested logit (wen2001gnl), which allow for each alternative to be “allocated” across several nests. While overlapping nests allows for the most flexible notion of similarity, these models have no testable implications if the nests are not known a priori and parameter values are unrestricted. Thus we demonstrate an important trade-off between nested and cross-nested models.
The second extension of NSC allows for “intermediate nests." These intermediate nests are often visually represented through a multi-level decision tree. Within this structure, we can allow for a more nuanced notion of similarity through the introduction of a second similarity relation that is conceptually related to our core similarity notion. This secondary relation captures “context-dependent” similarity and allows for comparative statements. We provide an axiomatic characterization of this model ((ref)) in appendix (ref).\footnote{Just as our similarity relation identifies endogenous nests, this secondary relation identifies endogenous, intermediate nests. Thus, (ref) shows that we may identify an endogenous tree structure.}
In NSC, each alternative belongs to one, and only one, nest. This feature of NSC places restrictions on the similarity relation. Because of these restrictions, in some settings, applied researchers have proposed allowing alternatives to exist in multiple nests. This leads to a class of models known as “cross-nested logits” (see vovsha1997application, ben1999discrete, wen2001gnl, papola2004cnl, and bierlaire2006cnl). In the cross-nested logit and the generalized nested logit, each alternative is allocated among the various nests. This allocation is specified with a vector of weights, one for each alternative, which describes to what extent an alternative belongs to each nest.
We show that any stochastic choice rule $p$ can be rationalized by some unrestricted cross-nested logit. That is, for any $p$, there exist some nesting structure, $X_1, \ldots, X_K$, allocations to these nests, $(\alpha^k_x)^K_{k=1}$, and utilities so that the resulting unrestricted cross-nested logit generates identical choice frequencies. Hence, unlike the nested logit and Luce models, there can be no behavioral characterization of the unrestricted cross-nested logit. The only testable implications of the model are due to the analysts' assumptions about alternative categorization and parameter restrictions.
Our result relies on the key insight that the cross-nested logit is behaviorally equivalent to a form of menu-dependent utility. We first prove this equivalence as (ref) and show how we can go from menu-dependent utility to weighted allocations and back. This equivalence between the cross-nested logit and menu-dependent utility allows us to reduce the problem of finding weights to the problem of finding menu-dependent utility values for each menu that satisfy the cross-nested logit equation. The bulk of the proof is dedicated to showing that the existence of these menu-dependent utilities is equivalent to the existence of a fixed point for some self-map. The proof is completed by applying Brouwer's fixed point theorem.
This result precisely shows the trade-offs between using nested logit and related models: either accept a restrictive form of similarity or impose assumptions on nest structure and model parameters. As we mentioned previously, further assumptions on parameters or nest structure may lead to testable restrictions. In the literature, similar to nested logit, it is commonly assumed that $\lambda\le 1$, since this is sufficient for cross-nested logit to be RUM. As with our handling of nested logit, we refer to such specifications as the random utility cross-nested logit. Note that our result shows that this restriction is not necessary for consistency with RUM; By (ref), every RUM has an unrestricted cross-nested logit representation with $\lambda >1$.
In any case, a random utility cross-nested logit must have, at least, the same testable restrictions as RUM. However, our result suggests that random utility cross-nested logit may not have any testable restrictions beyond RUM. In fact, although it does not prove our hypothesis, fosgerau2013choice prove that the set of random utility cross-nested logit models is dense in the set of RUMs. Note that our (ref) is quite different from the result of fosgerau2013choice for the following reasons: (i) we prove an exact result while they prove an approximation result, (ii) they focus on random utility cross-nested logit, and (iii) our proof techniques are completely different because their proof relies on the properties of the CDF for GEV distributions while we use Brouwer's fixed point theorem.
In most applications of nested logit to market data, researchers assume nests based on knowledge of alternative attributes. This is potentially problematic, as in many environments there are many plausible structures. When studying vehicle choice, the researcher might construct nests based on brand, body type (e.g., sedan vs. truck), or country of origin.\footnote{A common approach to this type of problem is to utilize multiple levels of nesting (which we characterize in appendix (ref)). Even under this approach, the order of the levels matters.} In other environments, classification may be subjective. When studying choice over apartments, nests might depend on both observable attributes and a myriad of unobservables: subjective impressions of neighborhoods, proximity to landmarks, or a host of other features. If the nest structure is misspecified, this may lead to biased conclusions regarding substitutability of goods and systematically inaccurate predictions.\footnote{greene2003econometric provides an excellent summary of this issue: “To specify the nested logit model, it is necessary to partition the choice set into branches [nests]. Sometimes there will be a natural partition ... In other instances, however, the partitioning of the choice set is ad hoc and leads to the troubling possibility that the results might be dependent on the branches so defined. ... There is no well-defined testing procedure for discriminating among tree structures, which is a problematic aspect of the model."}
We show in (ref) that the true (unobserved) nest structure can be identified from the data by solving a minimization problem. Any potential nest structure has implications for when IIA may and may not be violated between alternatives. For a hypothesized nest structure, $\mathcal{Y}$, we propose a measure of the total magnitude of IIA violations within and across the proposed nests, $D(\mathcal{Y})$. We show that the true nest is a minimizer of $D$ and that it will be the unique minimizer of $D$ under a mild identification assumption ((ref)). In cases where the researcher has several potential nest structures in mind, such as in vehicle choice, our procedure for nest identification could be useful for nest selection. The researcher can calculate $D$ for the particular nests in mind and select the one that best fits.
Because the number of possible nests grows exponentially as the number of alternatives increases, the full minimization problem becomes intractable. However, this issue can be managed due to insights from our similarity relation; by (ref), one only needs to check nest structures that are consistent with an empirical approximation of $\sim_p.$ Note that in finite data sets it is unlikely that IIA will hold between any alternatives (e.g., since we observe a finite sample from the true distribution). However, one can measure the magnitude of the the IIA violation between $a, b$ across various menus in the data. If this magnitude is below some threshold $\epsilon$, then we conclude that $a$ and $b$ are approximately similar: $a \sim_{\epsilon} b$. When $\sim_{\epsilon}$ is transitive, then there are at most $|X|$ potential nests that we need to check, as stated in (ref).
To analyze the problem of nest identification, we consider a data set denoted $\mathcal{O}=\{A, \{p_{t}(\cdot, A)\}^{N_A}_{t=1}\}_{A\in\mathscr{A}}$, where $N_A$ is the number of observations of menu $A$ and $p_{t}(a, A)=1$ means that $a$ was chosen from $A$ at observation $t\le N_A$. We also require $\sum_{x\in A}p_{t}(x, A)=1$ for each $A$, so that $p_{t}(a, A)=1$ implies $p_{t}(b, A)=0$ for any $b\in A\setminus\{a\}$. Note that we may always write \[p_{t}(a, A)=\overline{p}(a, A)+\epsilon_{t, a, A},\] where $\overline{p}(a, A)$ is the probability that $a$ is chosen from $A$ according to the NSC with $(v, u, \{X_i\}^K_{i=1})$. Then, the observed choice frequency of $a$ from $A$ in $\mathcal{O}$ is \[p(a, A)\equiv\frac{\sum^{N_A}_{t=1}p_t(a, A)}{N_A}=\overline{p}(a, A)+\epsilon_{a, A}\text{ where }\epsilon_{a, A}\equiv\frac{\sum^{N_A}_{t=1}\epsilon_{t, a, A}}{N_A}.\] We assume that $\{p_t(\cdot, A)\}^{N_A}_{t=1}$ are independently drawn according to $\overline{p}(\cdot, A)$. Then, by the classical Glivenko-Cantelli theorem, $\epsilon_{a, A}\xrightarrow{a.s.} 0$.\footnote{All convergence statements in this paper are with respect to $N^*\to\infty$ where $N^*=\min_{A\in\mathscr{A}} N_A$.} For notational simplicity, we write \[r_A(A', B')\equiv\frac{p(A', A)}{p(B', A)}\text{ and }\bar{r}_A(A', B')\equiv \frac{\overline{p}(A', A)}{\overline{p}(B', A)}\text{ for any }A, A', B'\in\mathscr{A}.\] Finally, let $\mathscr{X}$ denote the set of all partitions of $X$ and $\mathcal{X}^*$ denote the true partition $\{X_i\}^K_{i=1}$.
Consider the following minimization problem.
Intuitively, $D_1(\mathcal{Y})$ measures the degree to which the data violates IIA among elements in the same nest in $\mathcal{Y}$, while $D_2(\mathcal{Y})$ measures the degree to which the data violates IIA among different nests in $\mathcal{Y}$. These measures are motivated by our axiom \nameref{ISA}: $D_1$ measures the extent to which the first part of \nameref{ISA} holds, and $D_2$ measures the extent to which the second part of \nameref{ISA} holds.
Similarly, let us define loss functions $D^*, D^*_1, D^*_2$ when there is no observational noise; these are defined by replacing $p$ with $\bar{p}$ in Equations (ref)-(ref). Moreover, let $\hat{\mathcal{X}}=\arg\min_{\mathcal{Y}\in\mathscr{X}}D(\mathcal{Y})$. Note that $\hat{\mathcal{X}}$ is an M-estimator (takeshi1985advanced). Hence, by standard results (newey1994large), $\hat{\mathcal{X}}$ is a strongly consistent estimator of $\mathcal{X}^*$ if $\mathcal{X}^*$ is the unique minimizer of $D^*$. Indeed, $\mathcal{X}^*$ is a minimizer of $D^*$ since $D^*(\mathcal{X}^*)=0$. It turns out that it is the unique minimizer under the following identification assumption.
We now can state our strong consistency result.
(ref) shows that the true nest structure can be found by solving (ref). The intuition behind the result is as follows. As we see in our axiomatization, $a\sim_p b$ if and only if $a, b\in X_i$ for some $i$. Hence, IIA is satisfied between $a$ and $b$ when $a, b\in X_i$ and IIA is violated at least once between $a$ and $b$ when $a\in X_i$ and $b\in X_j$. Hence, the distance $\sum_{A, B\in\mathscr{A}, a, b\in A\cap B\cap Y}\Big(\log\big(r_A(a, b)\big)-\log\big(r_B(a, b)\big)\Big)^2$ between $a$ and $b$ is smaller whenever $a, b\in X_i$. Hence, minimizing $D_1(\mathcal{Y})$ helps us to correctly identify that elements from different nests are in fact from different nests.
However, it is important to notice that $D_1(\mathcal{Y})$ alone is not sufficient to identify $\mathcal{X}^*$. For example, suppose $X=\{a_1, \ldots, a_5\}$ and $\mathcal{X}^*$ is given by $X_1=\{a_1, a_2, a_3\}$ and $X_2=\{a_4, a_5\}$. Since the data provide a noisy measure of $\bar{p}$, it is possible that $D_1$ is minimized at some finer partition, say $Y_1=\{a_1\}$, $Y_2=\{a_2, a_3\}$, and $Y_3=\{a_4, a_5\}$. Note that $\mathcal{Y}$ splits $X_1$, and since $D_1$ measures IIA violations within nests, $D_1(\mathcal{Y})\le D_1(\mathcal{X}^*)$ because $\mathcal{Y}$ never combines two elements from different nests into the same nest (i.e., it is finer than $\mathcal{X}^*$).
This example illustrates a potential problem. $D_1$ by itself tends to select finer partitions (it wants to create “too many nests”). The second term, $D_2$, corrects this problem. If $\mathcal{Y}$ were the true nest structure, our axiomatization (i.e., the second half of \nameref{ISA}) requires that the relative likelihoods between alternatives in $Y_2$ (for instance, $a_2$) and alternatives in $Y_3$ (for instance, $a_4$) are unaffected by the presence of $a_1$. Accordingly, $\mathcal{Y}$ is penalized by $D_2$ if introducing $a_1$ changes the relative likelihoods between alternatives in $Y_2$ and $Y_3$. Importantly, since the true nest structure, $\mathcal{X}^*$, groups $a_1$ with $a_2$ and $a_3$, $\mathcal{X}^*$ will not be penalized, and so $D_2(\mathcal{Y})> D_2(\mathcal{X}^*)$ almost surely.\footnote{We say $Z_n>Z'_n$ almost surely when there is $N$ such that Pr$(Z_n>Z'_n)=1$ for any $n>N$.} Thus $D_2(\cdot)$ enables us to correctly conclude that $a_1$ and $a_2$ do in fact belong to the same nest.
Notice that solving (ref) is quite different from the typical exercise of selecting a nest structure in the literature. In a typical nested logit estimation, a researcher assumes a nest structure and then runs a maximum likelihood (ML) estimation to identify model parameters. To compare different nest structures, the researcher has to run a full ML estimation for each nest structure. Hence, it is computationally expensive to compare many different nest structures. However, our (ref) provides a data-driven way to compare different nest structures without estimating the full parametric model. Moreover, (ref) does not rely on the functional form of nested logit, since it applies to any NSC.
Finally, note that $\mathcal{X}^*$ is a minimizer of $D$ without any further assumptions; our identifying (ref) is only required to ensure that $\mathcal{X}^*$ is the unique minimizer. Consequently, when (ref) is violated, $\mathcal{X}^*$ will always be contained in the set of minimizers of $D$. This suggests that $D$ may still be used for nest selection and that solving (ref) can facilitate identification of the true nest structure.
There is a practical concern with directly applying (ref) to identify the nest structure because $|\mathscr{X}|$ grows exponentially as $|X|$ increases.\footnote{Unlike the standard method of estimating nested logit, it is not computationally expensive to solve (ref) by going through all possible partitions of $X$ when $|X|$ is small. For instance, there are 877 different partitions when $|X|=7$. Indeed, many papers in the literature study situations with relatively few alternatives (e.g., transportation modes or cellphone providers), and (ref) can be applied to these situations directly.} Therefore, we further refine our result by showing that we only need to compare at most $|X|$ different partitions, rather than $|\mathscr{X}|$. This dramatically reduces the number of calculations; comparing $|X|$ different partitions is computationally inexpensive even when $X$ contains hundreds of alternatives. To establish this result, we introduce the following measure of IIA violations. For any $a, b\in X$, let
The value of $d(a, b)$ captures the total “distance” between $a$ and $b$, in terms of IIA violations in the data $\mathcal{O}$. Consistent with our axiomatization, and the intuition behind $D_1$, the value of $d(a,b)$ is smaller when $a$ and $b$ are from the same nest than when they are from different nests. While conceptually similar to $D_1$, note that it is defined over the alternatives, not on nest structures. This crucial distinction allows us to use $d$ to narrow our candidate nests purely based on the data.
(ref) shows that for large enough $N^*$, there exists a “separating" threshold that correctly identifies whether two alternatives belong, or do not belong, to the same nest. If the researcher knows $\epsilon^*$, then identifying the nest structure is a straightforward task due to this result. But when $\epsilon^*$ is unknown, (ref) is not sufficient to identify the nest structure. However, the insights provided by (ref) allow us to show that in order to identify the correct nest structure for any NSC, only $|X|$ different partitions are worth considering. In fact, we will explicitly construct the set of partitions that need to be considered using $d$ and prove that this set contains the true nest $\mathcal{X}^*$.
In order to construct the set of relevant partitions, we introduce the following “approximately similar" relation: for any $\epsilon\ge 0$ and $a, b\in X$, let $a\sim_\epsilon b$ if $d(a, b)<\epsilon$. When $\sim_\epsilon$ is transitive, let $\mathcal{X}_\epsilon\equiv X/\sim_{\epsilon}$, which is the partition of $X$ such that for any $A\in\mathcal{X}_\epsilon$, $a\in A$, and $b\in X$, $a\sim_{\epsilon} b$ if and only if $b\in A$.
Since we have finite data, if $\epsilon$ and $\epsilon'$ are close enough, they will result in the same relation ($\sim_{\epsilon}=\sim_{\epsilon'}$), except for certain knife-edge cases (which happens at most $|X|$ times). Notice that for smaller $\epsilon$, we are “more discriminating” in declaring similarity and this results in a finer partition. For larger $\epsilon$, we are “less discriminating” in declaring similarity and this results in a coarser partition. Let $\overline{\epsilon}\equiv \max_{a, b\in X} d(a, b)$, the maximal distance calculated in the data. Then for any $\epsilon > \overline{\epsilon}$, the resulting relation $\sim_{\epsilon}$ is complete, which reduces to the Luce model ($\epsilon=0$ also gives the Luce model). Consequently, we never need to use an $\epsilon$ above $\overline{\epsilon}$. Because of these two key features of $\sim_{\epsilon}$, it turns out that the set $\mathscr{X}^*\equiv\{\mathcal{X}_\epsilon\}_{\epsilon\in [0, \overline{\epsilon}]}$ is the desired collection of partitions and contains at most $|X|$ different elements.
Combining Propositions (ref)-(ref), we can immediately show that $\mathscr{X}^*$ contains the true partition structure, and it can be found by solving (ref), as formalized below. Let $\hat{\mathcal{X}}^*=\arg\min_{\mathcal{Y}\in\mathscr{X}^*}D(\mathcal{Y})$.
The minimization problem (ref) is not computationally demanding since $|\mathscr{X}^*|\le |X|$. Hence, we can find the true nest structure even if there are many products. In practice, computing $\mathscr{X}^*$ from choice frequencies is quite simple. First note that any partition of $X$ can be represented by an $|X|\times |X|$ matrix $M$ such that $M_{x, y}=1$ when $x$ and $y$ are from the same nest and $0$ otherwise. Hence, to compute $\mathscr{X}^*$, we follow the following steps:
The first step determines the collection of relevant thresholds from the data to construct candidate relations $\sim_{\epsilon}$. The second step generates $|X|(|X|-1)/2$ matrices, which represent the similarity thresholds found in the previous step. The third step reduces the number to $|X|$, as we prove in (ref), since $\sim_{\epsilon}$ must be transitive.
To illustrate our algorithm and our theoretical result on identification, we ran the following simulation with six alternatives. We assumed that the true nest structure is $X_1=\{x_1, x_2, x_3\}$ and $X_2=\{x_4, x_5, x_6\}$, with $X=X_1\cup X_2$, and calculated the fraction of trials in which our procedure correctly identified the nest structure. To do so, we randomly generated values for $u$ and $v$ and calculated $\overline{p}$, which is the NSC given by $(v, u, \{X_1, X_2\})$. To introduce sampling error, we drew independent errors from a uniform distribution $U[0, \delta]$ and perturbed $\overline{p}$.\footnote{Specifically, for each menu $A$ and each simulation trial $t$, we independently draw errors $\{\zeta_{a, A,t}\}_{a\in A}$ from $U[0, \delta]$ and construct $\overline{p}^t(\cdot, A)$ as follows: $\overline{p}^t(a, A)=\frac{\overline{p}^t(a, A)+\zeta_{a, A, t}}{\sum_{b\in A}\overline{p}^t(b, A)+\zeta_{b, A, t}}$ for each $a\in A$.} As shown by (ref), when $\delta$ is small enough, the true nest structure will be identified correctly. This was confirmed by our simulation.
We considered six different values for $\delta$ ($\{0, 0.01, 0.025, 0.035, 0.05, 0.075\}$) and ran a total of 2400 trials, the results of which are summarized in (ref). For $\delta \in \{0, 0.01, 0.025, 0.035\}$, the nest structure was correctly identified in all trials. For $\delta=0.05$ ($0.075$), the nest structure was correctly identified 395 (391) times out of 400 trials. In other words, in line with our theoretical result ((ref)), when error is relatively small (e.g., $\delta \le 0.035$) the true nest is recovered $100\%$ of the time. Even for relatively large errors (e.g., $\delta \ge 0.05$), we recover the true nests over $97\%$ of the time.
The main contributions of our paper are the characterizations of nested logit and NSC. While nested logit is the most commonly applied model that deals with the similarity effect, many other models have been proposed. Two prominent such models are Elimination-By-Aspects (EBA) of tversky1972eba and the Attribute Rule (AR) of gul2014random. Both EBA and AR are RUM and generalize the Luce model. While each of these models has an intersection with the NSC, neither one nests nor is nested by NSC.
In Tversky's EBA, each alternative is a collection of aspects. The decision maker randomly selects one of these aspects from the aspects available in the menu, via a Luce rule, and eliminates alternatives that do not have the selected aspect. The decision maker repeats this procedure until a single alternative remains. EBA is conceptually similar to an $N$-step nested logit, where $N$ is the total number of alternatives. The Linear NSC is a special case of EBA, but EBA is disjoint from nested logit.
In the AR of gul2014random, each alternative has many attributes. A decision maker randomly selects one attribute from the attributes available in the menu via a Luce rule. When the selected attribute is $\omega$, alternative $x$ will be chosen with a probability that is proportional to $\gamma^\omega_x$, where $\gamma^\omega_x$ is the intensity of $\omega$ in $x$. The AR is conceptually similar to cross-nested logit. In fact, one can show that the AR is a special case of a non-parametric version of cross-nested logit in which weights assigned to nests are menu-independent (i.e., $\gamma^\omega_x$ is menu-independent). Because of the menu independence of $\gamma^\omega_x$, the AR is more restrictive than unrestricted cross-nested logit. The Linear NSC is a special case of the AR, but nested logit is not.
Other recent papers dealing with the similarity effect are farolucereplicates and li2016associationistic. Both are special cases of NSC, but are generally distinct from nested logit.
farolucereplicates introduces the Luce Model with Replicas (LR), which is a special case of Linear NSC in which nest values and Luce utilities are constant: $v(A)=v_i$ for each $A\subseteq X_i$ and $u(a)=u(b)$ for any $a, b\in X_i$. In terms of behavior, Faro's model only allows for restrictive forms of the similarity effect in which similar alternatives are replicas.
li2016associationistic present the Associationistic Luce Model (AL), which is a special case of NSC with $v(A)=\sum_{a\in A} \gamma(a)$ for some function $\gamma$. Because of the additive structure of $v$, AL is significantly more restrictive than NSC. In fact, the Luce model is the only intersection between nested logit and AL. We note that the AL also allows for violations of regularity (e.g., the attraction effect). However, since $v$ is increasing, the AL cannot simultaneously allow for violations of regularity and the similarity effect (see Appendix (ref)). In terms of axiomatic foundations, they also use the revealed similarity relation $\sim_p$, and impose transitivity of $\sim_p$ as one of their axioms.
NSC has a large overlap with RUM, which goes back to Block1960rum, falmagne1978representation, and barbera1986falmagne. For example, both random utility nested logit and Linear NSC are RUM. In addition to EBA, AR, and random utility nested logit, many special cases of RUM have been proposed, including: gul2006reu, in which each preference has an expected utility representation; apesteguia2017scrum, in which the collection of preferences satisfy the single-crossing property; and manzini2014stochastic, in which randomness occurs due to stochastic consideration. Our characterization of random utility nested logit contributes to this area of the stochastic choice literature.
NSC has an interpretation as a sequential choice model, in which a nest is chosen and then an alternative. manzini2012categorize study a deterministic choice model in which a decision maker categorizes alternatives before choosing. The decision maker first selects the “best” category according to some ordering, then selects their most preferred alternative according to another. Categories however do not need to form a partition, unlike in NSC. ravid2018focus introduce the following stochastic choice model that involves a sequence of binary comparisons: \[p(x, A)=\frac{\prod_{y\in A\setminus\{x\}}\pi(x, y)}{\sum_{z\in A} \prod_{t\in A\setminus\{z\}}\pi(z, t)}.\] This model, which is a special case of marley1991, is disjoint from nested logit but has an interesting connection to NSC. In particular, when $\pi(x, y)=\frac{1}{u(y)}$ and $\pi(x, z)=\frac{1}{w(z)}$ for any $x, y\in X_i$ and $z\in X_j$, we obtain an NSC with $v(A\cap X_i)=\big(\sum_{y\in A\cap X_i} u(y)\big)\frac{\prod_{y\in A\cap X_i} w(y)}{\prod_{y\in A\cap X_i} u(y)}$.