Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
96,447 characters · 16 sections · 60 citation commands
Peer Effects in Random Consideration Sets
In the last few years, the basic rational choice model of decision-making has been revised in different ways, partly recognizing that cognitive factors might play an important role in determining the choices of people. These revisions have produced, among others, the consideration set models. In these models, people do not consider all the available options at the moment of choosing, but a subset of them. It is still an open question in the literature how the consideration sets are formed and what their main determinants are. We think that peer effects might be of first-order importance in the formation process of the consideration sets.\footnote{This possibility has been (explicitly or implicitly) discussed by other researchers in specific contexts ---e.g., the choices of peers may help us discover a new television show godes2004using, a new welfare program caeyers2014peer, a new retirement plan duflo2003role, a new restaurant qiu2018learning, or an opportunity to protest enikolopov2020social.} To study this possibility, this paper builds a dynamic model where the choices of friends affect the subset of options that a person ends up considering. We show the empirical content of the model and apply our main results to an experimental dataset that has been designed to evaluate the visual focus of attention of people. In our application, peer effects have a large impact on the attention mechanism.
In the dynamic model we build, people are linked through a social network. At a randomly given time, a person gets the opportunity to select a new option out of a finite set of alternatives. The person sticks to her new option until the revision opportunity arises again. People do not consider all the available options at the moment of revising their selection. Instead, each person first forms a consideration set and then picks her most preferred option from it. The distinctive feature of the model is that the probability that a given alternative enters the consideration set depends on the number of friends that are currently adopting that option. This theoretical model leads to a sequence of choices that evolves through time according to a Markov random process. The dynamic system induced by our model has a unique equilibrium that specifies the fraction of time that each vector of joint choices prevails in the long run.
In the consideration set models, the relative frequency by which a person chooses each alternative does not necessarily respect her ranking of preference. Formally, let us say that a person makes a mistake when, at the moment of choosing, she selects an alternative that is dominated by another option according to her preference order. In our model, a person could seldom choose her most preferred option at equilibrium. The set of connections, or the network structure, shapes the nature and the strength of the mistakes that people make. Thus, our model can be thought of as a model of peer effects in mistakes. We show via example that having friends with similar preferences ---a phenomenon known as homophily--- reduces the frequency by which people make mistakes. In addition, the example shows that having more friends also results in fewer mistakes.
Having established equilibrium existence and characterized equilibrium behavior, we consider a researcher who observes a long sequence of choices made by the members of the network. We show that all primitives of the model can be uniquely recovered. These primitives include the ranking of preferences, the attention mechanism (or consideration probabilities), and the network structure. There are two aspects of our nonparametric identification results that deserve special attention. First, we allow unrestricted heterogeneity across people with respect to all parts of the model. Second, in contrast to most other works on consideration sets, we do not rely on the variation of either covariates or the set of available options (or menu) to recover these components.
In our dynamic model, the observed choices of network members are generated by a system of conditional choice probabilities: each of these conditional choice probabilities specifies the distribution of choices of a given person conditional on the choices of others (at the moment of revising her selection). The identification strategy we offer is a two-step procedure. First, we show how to identify the primitives of the model using these conditional choice probabilities. Second, we study identification of the conditional choice probabilities from observed data.
To recover the primitives of the model from the conditional choice probabilities we exploit that, in our framework, changes in the choices of friends induce stochastic variation in the consideration set probabilities. We use this variation to recover the set of connections between the people in the network and their ranking of preferences. We then use this information to recover the attention mechanism of each person, i.e., the probability of including a specific option in the consideration set as a function of the number of friends who are currently choosing it. As we explained earlier, this identification result is nonparametric and allows for full heterogeneity across people regarding preferences and attention probabilities.
To identify the conditional choice probabilities, we consider two datasets: continuous-time data and discrete-time data with arbitrary time intervals. These two datasets coincide in that they provide a long sequence of choices from people in the network. They differ in the timing at which the researcher observes these choices. In continuous-time datasets, the researcher observes people's choices in real time. We can think of this dataset as the “ideal dataset.” With the proliferation of online platforms and scanners, this sort of data might be available in some applications. In this dataset, the researcher directly recovers the conditional choice probabilities. In discrete-time datasets, the researcher observes the joint configuration of choices at fixed time intervals (e.g., the choice configuration is observed every Monday). In this case, the conditional choice probabilities are not directly observed or recovered, they need to be inferred from the data. Adding an extra mild condition, we show that the conditional choices are also uniquely identified. For this last result we invoke insights from blevins2017identifying,blevins2018identification. Briefly, the reason for which the second dataset suffices to recover the primitives of the model lies on the fact that the transition rate matrix of a continuous-time process with independent revision times across people is rather parsimonious. In particular, the probability that two or more people revise their selected options at the same time is zero. As a consequence, the transition rate matrix has zeros in many known locations. The non-zero elements can then be nicely recovered and they constitute a one-to-one mapping with the conditional choice probabilities.
Our initial results rely on a few simplifying restrictions. In particular, the preferences of each person are deterministic; we assume that there exists one alternative (the default) that is picked if and only if nothing else is considered; and we let the distribution of consideration sets be multiplicatively separable across alternatives and the probability of including each option depends on the number (but not the identity) of the friends that selected that option. We extend our model in various directions to relax these assumptions. First, we consider random preferences (on top of random consideration sets). The network structure and the attention mechanism are identified without any new assumptions. The random preferences are identified if and only if each person has “enough” friends ---the actual number of friends needed depends on the number of options. Second, we analyze a model with no default. All the primitives of this model are identified if there are more than three options in the set of available alternatives. Finally, we consider a set of extensions where consideration sets are formed arbitrarily (e.g., no multiplicative separability assumption). In these extensions, different friends may have different effects on the attention mechanism of a given person. In this model, we still can uniquely recover preferences and the network structure. The consideration probabilities, in general, are partially identified. However, we show that different forms of symmetry between consideration set probabilities are sufficient for their identification.
After presenting the main identification results, we propose a consistent maximum-likelihood estimator of all parameters of the model that uses discrete-time data. Our Monte Carlo simulations imply that our estimator performs well even for relatively small datasets.
Last, but not least, we apply our model to the experimental dataset used by bai2019predicting, which has been designed to evaluate the visual focus of attention of people. In the experiment, a group of people seated around a circle were asked to play the Resistance game. This is a party game in which people communicate with each other to find the deceptive players. A tablet was placed in front of each player and it recorded the direction of the sight of the player every third of a second (i.e., we consider each player looking to her left, right, or center). The structural model we offer allows us to disentangle the effects of two key forces in the visual focus of attention of people: directional sight preferences and peer effects in gazing. Regarding directional sight preferences, our estimates provide very strong support to the left-to-right bias. That is, the observation that people tend to look to the left more often than to the right.\footnote{See, for example, the experimental findings of spalek2004supporting,spalek2005left and reutskaja2011search. The left-to-right bias has also been documented in marketing research \burl{https://www.nngroup.com/articles/horizontal-attention-original-research/}. See also maass2003directional and reference therein.} (This observation has guided, among others, the design of many online platforms!) Regarding peer effects in gazing, our results support the observation that people redirect their visual attention by following the gaze orientation of others, a phenomenon called “gaze following.”\footnote{See, for example, gallup2012visual.}
We finally relate our results with the existing literature. From a modeling perspective, our setup combines the dynamic model of social interactions of blume1993statistical,blume1995statistical with the (single-agent) model of random consideration sets of manski1977structure and manzini2014stochastic. By adding peer effects in the consideration sets we can use variation in the choices of others as the main tool to recover preferences. The literature on identification of single-agent consideration set models has mainly relied on variation of the set of available options or menus. The latter includes aguiar2017random, ABDsatisficing16, brady2016menu, caplin2018rational, cattaneo2017random, horan2019random, lleras2017more, manzini2014stochastic, and masatliogluRA. (See aguiar2019doesrandom for a comparison of several consideration set models in an experimental setting.) Other papers have relied on the existence of exogenous covariates that shift preferences or consideration sets. The latter include barseghyan2019discrete, barseghyan2019heterogeneous, crawford2021survey, conlon2013demand, draganska2011choice, gaynor2016free, goeree2008limited, mehta2003price, and roberts1991development. Variation of exogenous covariates has also been used by abaluck2017consumers via an approach that exploits symmetry breaks with respect to the full consideration set model. aguiar2020identification, crawford2021survey, and dardanoni2020inferring use repeated choices but do not allow for peer effects (i.e., they work with panel data).
There is a vast econometric literature on identification of models of social interactions where choices of peers affect preferences but not the choice sets (see blume2011identification, bramoulle2020peer, de2017econometrics, and graham2015methods for comprehensive reviews of this literature). We depart from this literature in that in our framework the direct interdependence between choices (endogenous effects) is captured by consideration sets (choice sets), not preferences.\footnote{Implicitly, preferences in our model capture the so-called contextual effects -- effects of predetermined factors that are not affected by peers.} As we mentioned earlier, we can recover from the data the set of connections between the people in the network. In the context of linear models, a few recent papers have made progress in the same direction. Among them, blume2015linear, bonaldi2015empirical, de2018recovering, and manresa2013estimating. In the context of discrete-choice, chambers2019behavioral also identifies the network structure but in their model peers do not affect consideration sets but directly change preferences (among other differences).
In the paper, we connect the equilibrium behavior of our model with the Gibbs equilibrium. This connection is similar to the one in blume2003equilibrium. We also distinguish our approach with the more standard models of peer effects in preferences. To this end we embed the models of brock2001discrete, brock2002multinomial into our dynamic revision process.
Let us finally mention two other papers that incorporate peer effects in the formation of consideration sets. borah2018 do so in a static framework and rely on variation of menus for identification. lazzati2018diffusion considers a dynamic model but the time is discrete and she focuses on two binary options that can be acquired together.
The rest of the paper is organized as follows. Section (ref) presents the model, the main assumptions, and some key insights of our approach (including how it differs from the more standard approach of peer effects in preferences). Section (ref) describes the equilibrium behavior. Section (ref) studies the empirical content of the model. Section (ref) extends the initial idea to contemplate random preferences (besides random consideration sets), the case of no-default option, and more general formation process for the consideration sets. Section (ref) applies our model to experimental data on visual focus of attention. Section (ref) concludes, and all the proofs are collected in Appendix (ref). Appendices (ref) and (ref) cover the Gibbs random field model and provide some simulation results, respectively. Appendix (ref) contains additional results for our empirical application.
This section describes the model and the main assumptions we invoke in the paper. It also compares our approach to the more standard model of peer effects in preferences.
Network and Choice Configuration There is a finite set of people connected through a social network. The network is described by a simple graph $\Gamma =\left( \mathcal{A},e\right)$, where $\mathcal{A}=\left\{ 1,2,\dots,A\right\}$ is the finite set of nodes (or people) and $e$ is the set of edges. Each edge identifies two connected people and the direction of the connection. For each Person $a\in \mathcal{A}$ her set of friends (or reference group) is defined as follows:
There is a set of alternatives $\overline{\mathcal{Y}}=\mathcal{Y}\cup\left\{ o\right\}$, where $\mathcal{Y}=\left\{1,2,\dots,Y\right\}$ is a finite set of options and $o$ is a default option. Each Person $a$ has a strict preference order $\succ_{a}$ over the set of options $\mathcal{Y}$. All people agree in that the default option is the least preferred.\footnote{Although we assume the same default option for all people, all our results go through if people have different default options as far as the researcher knows the identity of the worst option for each person.} We extend the analysis to random preferences in Section (ref) and fully relax the specification of the default option in Section (ref). We refer to a vector $\mathbf{y}=\left( y_{a}\right)_{a\in \mathcal{A}}\in \overline{\mathcal{Y}}^{A}$ as a choice configuration.
Choice Revision We model the revision of choices as a standard continuous-time Markov process. In particular, we assume that people are endowed with independent Poisson alarm clocks with rates $\mathbf{\lambda }=\left( \lambda_{a}\right)_{a\in \mathcal{A}}$.\footnote{See blume1993statistical, blume1995statistical for theoretical models that rely on Poisson alarm clocks and blevins2018identification for a nice discussion of the advantages of this type of revision process from an applied perspective.} At randomly given moments (exponentially distributed with mean $1/\lambda_{a}$) the alarm of Person $a$ goes off.\footnote{That is, each Person $a$ is endowed with a collection of random variables $\left\{\tau_{n}^{a}\right\}_{n=1}^{\infty}$ such that each difference $\tau_{n}^{a}-\tau_{n-1}^{a}$ is exponentially distributed with mean $1/\lambda_{a}$. These differences are independent across people and time.} When this happens, the person selects the most preferred alternative among the ones she is actually considering. Formally, if $\mathcal{C}\subseteq\mathcal{Y}$ is her consideration set, then the choice of Person $a$ can be represented by an indicator function
that takes value 1 if $v$ is the most preferred option in $\mathcal{C}$ according to $\succ_{a}$. If, at the moment of choosing, the consideration set of Person $a$ does not include any alternative in $\mathcal{Y}$, then the person selects the default option.
Peer Effects in the Formation of Consideration Sets In our model, whether Person $a$ pays attention to a particular alternative depends on her own choice and the configuration of choices of her friends at the moment of revising her selection. We indicate by $\operatorname{Q}_{a}\left( v \mid \mathbf{y}\right)$ the probability that Person $a$ pays attention to alternative $v$ given a choice configuration $\mathbf{y}$. It follows that the probability of facing consideration set $\mathcal{C}$ has the form of
By combining preferences and stochastic consideration sets, the probability that Person $a$ selects (at the moment of choosing) alternative $v\in\mathcal{Y}$ is given by
The default option is assumed to be always considered. Since it is also assumed to be the worst for every person, the default option is chosen by the person if, and only if, nothing else is considered. Thus, the probability of selecting $o$ is just $\prod\nolimits_{v\in\mathcal{Y}}\left( 1-\operatorname{Q}_{a}\left( v \mid \mathbf{y}\right) \right)$. Leaving aside peer effects, this process of formation of consideration sets is analogous to the one studied by manski1977structure and manzini2014stochastic. In Section (ref) we extend our analysis to a more general setting.
Altogether, the three elements just described characterize our initial model of peer effects in random consideration sets. (As mentioned above, we consider later several extensions.) This model leads to a sequence of joint choices that evolves through time according to a Markov random process. An equilibrium in this model is an invariant distribution on the space of joint choices of the dynamic process, that describes the frequency by which each member of the group selects each option. We next present three key assumptions we invoke in the paper.
Our results build on three assumptions. Let $\operatorname{N}_{a}^{v}\left( \mathbf{y}\right)$ be the number of friends of Person $a$ who select option $v$ in choice configuration $\mathbf{y}$. Formally,
We indicate by $\left\vert \mathcal{N}_{a}\right\vert$ the cardinality of $\mathcal{N}_{a}$. The three assumptions are as follows.
(A1) For each $a\in\mathcal{A}$, $v\in \mathcal{Y}$, and $\mathbf{y}\in \overline{\mathcal{Y}}^{A}$, $1>\operatorname{Q}_{a}\left( v \mid \mathbf{y}\right) >0$.
(A2) For each $a\in\mathcal{A}$, $\left\vert\mathcal{N}_{a}\right\vert>0$.
(A3) For each $a\in\mathcal{A}$, $v\in\mathcal{Y}$, and $\mathbf{y}\in \overline{\mathcal{Y}}^{A}$,
Assumption A1 states that, for any choice configuration, the probability of considering each option is strictly positive and lower than one, independently on how many friends have selected that option. This assumption captures the idea that people can eventually pay attention to an alternative for various reasons that are outside the control of our model (e.g., watching an ad on television or receiving a coupon). It also allows people to eventually disregard any further consideration of a given option, including the one that she is currently adopting. Assumption A1 guarantees that each subset of options is (ex-ante) considered with nonzero probability. Assumption A2 requires each person to have at least one friend. Assumption A3 states that the probability that a given person pays attention to a specific option depends on the current choice of the person and the number (but not the identity) of friends that currently selected it. Assumption A3 also states that each person pays more attention to a particular option if more of her friends are adopting it.
In our model, the probability that Person $a$ selects option $v$ (at the moment of choosing) given network configuration $\mathbf{y}$ takes the form of \[ \operatorname{P}_{a}\left( v \mid \mathbf{y}\right) = \operatorname{Q}_{a}\left( v \mid y_{a},\operatorname{N}_{a}^{v}\left( \mathbf{y}\right)\right) \prod\nolimits_{v'\in \mathcal{Y}, v' \succ_{a} v}\left( 1-\operatorname{Q}_{a}\left( v' \mid y_{a},\operatorname{N}_{a}^{v'}\left( \mathbf{y}\right)\right)\right). \] Let us formally say a person makes a mistake whenever she selects a dominated option. Note that the preference order of Person $a$ over the available alternatives is fixed. The choices of friends simply redirect the attention of the person from one alternative to another one. As a consequence, at any given time, the probability that Person $a$ selects a low ranked alternative that is popular among her friends can be much higher than the probability of selecting her most preferred option. In summary, in our model, the network structure does not shape preferences (or payoffs across alternatives) but the strength of the mistakes that people make. We show later, via example, that having similar friends -- a phenomenon known as homophily -- can sharply reduce the frequency of mistakes that the person makes. Having more friends acts in the same direction.
Our model markedly differs from the more standard model of peer effects in preferences. To illustrate some of the main differences, note that in the same dynamic revision process we can easily embed a model of peer effects in preferences by modifying the conditional choice probabilities. Let us assume that individual choices are path independent, in the sense that choosing from a large set of options can be broken up into choosing from smaller subsets (Luce's Choice Axiom). If we add peer effects in preferences into the resulting choice model, then the probability that option $v$ is chosen given network configuration $\mathbf{y}$ can be expressed as follows \[ \operatorname{P}_{a}\left( v \mid \mathbf{y}\right) =\ \mathrm{U}_{a}^{v}\left(y_{a},\operatorname{N}_{a}^{v}\left(\mathbf{y}\right)\right)/(\sum\nolimits_{v'\in\overline{\mathcal{Y}}}\mathrm{U}_{a}^{v'}\left(y_{a},\operatorname{N}_{a}^{v'}\left(\mathbf{y}\right)\right)), \] where $\mathrm{U}_{a}^{v}\left(y_{a},\operatorname{N}_{a}^{v}\left(\mathbf{y}\right)\right)$ is the utility Person $a$ gets from alternative $v$ given her previous choice $y_a$ and the number of friends who also picked it, $\operatorname{N}_{a}^{v}\left(\mathbf{y}\right)$.\footnote{If we further assume, as in brock2001discrete, brock2002multinomial, that the payoff Person $a$ gets from alternative $v$ is the sum of a deterministic private utility the person receives from the choice, $\delta_{a,v}$, a social component that depends on the number of friends that select that option, $\mathrm{S}_{a,v}\left(\mathbf{y}\right)\equiv \mathrm{S}_{a,v}\left(y_{a},\operatorname{N}_{a}^{v}\left( \mathbf{y}\right)\right)$, and an idiosyncratic term that is doubly exponentially distributed with index parameter $1$, then $\mathrm{U}_{a}^{v}\left(y_{a},\operatorname{N}_{a}^{v}\left(\mathbf{y}\right)\right)$ takes the familiar exponential form $\exp\left(\delta_{a,v}+\mathrm{S}_{a,v}\left(y_{a},\operatorname{N}_{a}^{v}\left(\mathbf{y}\right)\right)\right)$.} In this model, the choices of friends affect the value that Person $a$ assigns to each option. Thus, at each given time, the relative frequency by which Person $a$ selects option $v$ over option $v'$ reflects her relative expected utility of the first option with respect to the second one. It follows that the alternative with the highest expected utility is always selected more often. This is in sharp contrast with what happens in our peer effects model in consideration sets, as we explained above.
Let us finally remark that, expressed as we did, these two models also have different empirical implications. In particular, in the random consideration set model, the probability that Person $a$ selects alternative $v$ does not vary if some of her friends switch from the default option to any option that is dominated by alternative $v$ according to $\succ_a$. In contrast, in the model of peer effects in preferences, the probability of selecting alternative $v$ strictly increases under the same scenario. It follows that if we can recover the conditional choice probabilities from the data (as we do in Section (ref)), then the two models can be set apart.
The independent and identically distributed Poisson alarm clocks, which lead the selection revision process, guarantee that at each moment of time at most one person revises her selection almost surely. Thus, the transition rates between choice configurations that differ in more than one person changing the current selection are zero. The advantage of this fact for model identification is that there are fewer terms to recover. blevins2017identifying,blevins2018identification offers a nice discuss of this feature and its advantage over discrete time models. Formally, the transition rate from choice configuration $\mathbf{y}$ to any different one $\mathbf{y}'$ is as follows
In the statistical literature on continuous-time Markov processes these transition rates are the out of diagonal terms of the transition rate matrix (also known as the infinitesimal generator matrix). The diagonal terms are simply build from these other values as follows
We will indicate by $\mathcal{M}$ the transition rate matrix. In our model, the number of choice configurations is $\left(Y+1\right)^{A}$. Thus, $\mathcal{M}$ is a $\left(Y+1\right)^{A}\times \left(Y+1\right)^{A}$ matrix. There are many different ways of ordering the choice configurations and thereby writing the transition rate matrix. To avoid any sort of ambiguity in the exposition, we will let the choice configurations be ordered according to the lexicographic order with $o$ treated as zero. Constructed in this way the first element of $\mathcal{M}$ is $\mathcal{M}_{11}=\operatorname{m}\left(\left(o,o,\dots,o\right)^{\prime}\mid \left(o,o,\dots,o\right)^{\prime}\right)$. Formally, let $\iota\left( \mathbf{y}\right) \in\left\{1,2,\dots,\left(Y+1\right)^{A}\right\} $ be the position of $\mathbf{y}$ according to the lexicographic order. Then,
An equilibrium in this model is an invariant distribution $\mu: \overline{\mathcal{Y}}^{A} \rightarrow\left[0,1\right]$, with $\sum\nolimits_{\mathbf{y}\in \overline{\mathcal{Y}}^{A}}\mu \left( \mathbf{y}\right) =1$, of the dynamic process with transition rate matrix $\mathcal{M}$. It indicates the likelihood of each choice configuration $\mathbf{y}$ in the long run. This equilibrium behavior relates to the transition rate matrix in a linear fashion
The next proposition establishes equilibrium existence and uniqueness for our model.
Remark. In Appendix (ref) we relate the equilibrium notion we use with the so-called Gibbs equilibrium. This alternative solution concept has been extensively used in Economics to study social interactions. (See allen1982some, blume1993statistical, blume1995statistical, and blume2003equilibrium, among many others.) The existence of Gibbs equilibrium requires extra symmetry restrictions among people. While these additional restrictions make the model more tractable, they are hard to justify in empirical applications. This is why we selected a more general approach.
The first example describes the equilibrium behavior of a very simple specification of our model.
Example 1. The network consists of two identical people that select between two alternatives, option $1$ and the default option $o$. The rates for their Poisson alarm clocks are $1$. Let us also assume, to simplify the setup, that the probability of paying attention to a particular option only depends on the current choice of the other person (i.e., $\operatorname{Q}_{a}\left( v \mid y_a, \operatorname{N}_{a}^{v}\left( \mathbf{y}\right) \right)=\operatorname{Q}_{a}\left( v \mid \operatorname{N}_{a}^{v}\left( \mathbf{y}\right) \right)$) and is the same for both people (i.e., $\operatorname{Q}(\cdot)=\operatorname{Q}_1(\cdot)=\operatorname{Q}_2(\cdot)$). Thus, for $a=1,2$, we get that
The transition rate matrix $\mathcal{M}$ is as follows.
The transition rate matrix $\mathcal{M}$ is naturally more complex when there are more alternatives or more people. However, the structure of the zeros in $\mathcal{M}$ is rather similar. In particular, there are many zeros in known locations. As we mentioned earlier, this feature of the model is particularly attractive for identification and estimation.
The invariant distribution of choices, or steady-state equilibrium, satisfies $\mu\mathcal{M}=\mathbf{0}$. Solving this system of equations, we get that the steady-state equilibrium is
The steady-state equilibrium is a joint distribution on the pair of choice configurations. It states the fraction of time that each pair of choices $(y_{1},y_{2})$ prevails in the long run.$\blacksquare $
The second example shows how the strength of the mistakes that people make relates to the structure of the network. We will use the simulated data in the example later to show how the main parts of the model can, indeed, be estimated from a sequence of choices.
Example 2. There are five people in the network. Their reference groups are as follows \[ \mathcal{N}_{1}=\left\{ 2\right\},\quad\mathcal{N}_{2}=\left\{ 1\right\},\quad\mathcal{N}_{3}=\left\{ 1,2\right\},\quad\mathcal{N}_{4}=\left\{ 5\right\} ,\quad\mathcal{N}_{5}=\left\{ 4\right\}. \] Note that each person has at least one friend, so Assumption A2 is satisfied. Moreover, the network is directed since $3\notin \mathcal{N}_{1},\mathcal{N}_{2}$ and $1,2\in\mathcal{N}_{3}$. There are two possible alternatives $\mathcal{Y} = \left\{ 1,2\right\}$ in addition to the default option $o$. The preferences of these people are as follows \[ 2\succ_{1}1,\quad 1\succ_{2}2,\quad 2\succ_{3}1,\quad 1\succ_{4}2,\quad 1\succ_{5}2. \] That is, Persons 2, 4, and 5 prefer option 1, and Persons 1 and 3 prefer option 2. We will assume the attention mechanism is invariant across people and alternatives. To keep the exercise simple, as in Example 1, we will also assume the probability of paying attention to a given option only depends on the choices of friends. In this case, we indicate by $\operatorname{Q}\left( v \mid \operatorname{N}_{a}^{v}\left(\mathbf{y}\right) \right)$ the probability that Person $a$ pays attention to option $v\in \mathcal{Y}$ if the number of friends that are currently selecting that option is $\operatorname{N}_{a}^{v}\left( \mathbf{y}\right)$. We initially let \[ \operatorname{Q}\left( v \mid 0\right) =\frac{1}{4},\quad\operatorname{Q}\left( v \mid 1\right) =\frac{3}{4},\quad\operatorname{Q}\left( v \mid 2\right) =\frac{7}{8}. \] The rates for their Poisson alarm clocks are 1.
The equilibrium behavior of this choice model is a joint distribution $\mu $ with support on 243 choice configurations $\left(3^{5}\right)$. We simulated a long sequence of choices and calculated the equilibrium behavior. (See Appendix (ref) for more details.) From the equilibrium distribution we can obtain the marginal equilibrium probabilities for each person. Each of these marginals specifies the fraction of time that a given person selects each alternative in the long run. These marginal probabilities are shown in Table (ref).
It is interesting to observe that, for some people in the network, the relative frequency of choosing each alternative does not respect their order of preferences. In particular, notice that Persons 4 and 5 select the default more often than option 2, which dominates the former according to their preferences.
As before, let us say that a person makes a mistake when does not select her most preferred alternative. From the equilibrium marginals, we can calculate the probabilities of making mistakes. These probabilities are displayed in Table (ref).
Note that Persons 2 and 4 are identical in all respects except in the type of friend they have. In particular, Person 4 shares with her friend the same preferences. The opposite is true for Person 2. This difference leads Person 4 to make fewer mistakes. Hence, having friends with similar preferences (homophily) seems to help each person to consider more often her best alternative. Also, note that Person 3 has more friends and makes fewer mistakes. Thus, having more friends is also helpful in this application.
To show a bit more how the network structure shapes people's mistakes, we add two more connections in the model. In particular, let us assume that Person 3 is a friend of Persons 1 and 2. That is, \[ \mathcal{N}_{1}=\left\{ 2,3\right\},\quad \mathcal{N}_{2}=\left\{1,3\right\},\quad \mathcal{N}_{3}=\left\{ 1,2\right\},\quad \mathcal{N}_{4}=\left\{ 5\right\},\quad \mathcal{N}_{5}=\left\{ 4\right\}. \] In this case, the network becomes undirected. Repeating the previous exercise, the new network generates the marginal equilibrium distributions depicted in Table (ref).
From these marginals, the probabilities of mistakes are presented on Table (ref).
The probabilities of making mistakes decrease for Persons 1, 2, and 3. But notice that the change is larger for Persons 1 and 3 as they share the same preferences over the alternatives.$\blacksquare $\
This section provides conditions under which the researcher can uniquely recover (from a long sequence of choices) the set of connections $\Gamma =\left( \mathcal{A},e\right)$, the profile of strict preferences $\succ=\left( \succ_{a}\right)_{a\in \mathcal{A}}$, the attention mechanism $\operatorname{Q}=\left( \operatorname{Q}_{a}\right)_{a\in \mathcal{A}}$, and the rates of the Poisson alarm clocks $\lambda=\left( \lambda_{a}\right)_{a\in \mathcal{A}}$.
We separate the identification analysis in two parts. Let $\operatorname{P}=\left( \operatorname{P}_{a}\right)_{a\in \mathcal{A}}$ be the profile of choice probabilities of people in the network. Each $\operatorname{P}_{a}\left( v \mid \mathbf{y}\right) :\overline{\mathcal{Y}}\times\overline{\mathcal{Y}}^{A}\rightarrow\left( 0,1\right)$ specifies the (ex-ante) probability that Person $a$ selects option $v$ when the choice configuration is $\mathbf{y}$. Recall that, in our model,
and the probability of selecting the default option $o$ is $\prod\nolimits_{v\in\mathcal{Y}}\left( 1-\operatorname{Q}_{a}\left( v \mid \mathbf{y}\right) \right)$. First, we show that each set of conditional choice probabilities $\operatorname{P}$ maps into a different set of connections, profile of strict preferences, and attention mechanism. Thus, knowledge of the set of the conditional choice probabilities allows us to uniquely recover all the elements of the model. Second, we build identification of the conditional choice probabilities $\operatorname{P}$ from a long sequence of choices.
Under Assumptions A1 and A2, changes in the choices of friends induce stochastic variation of the consideration sets. Assumption A3 guarantees this variation is monotonic in the sense that the probability of considering one option increases with the number of friends that are currently adopting it. This stochastic variation in choices allows us to recover the set of connections between the people in the network and the ranking of preferences of each of them. We then sequentially identify the attention mechanism of each person moving from the most preferred alternative to the least preferred one. Proposition (ref) presents our first identification result.
The next example sheds some light on the identification strategy in Proposition (ref).
Example 3. Suppose there are three people $\mathcal{A}=\left\{ 1,2,3\right\}$ that select between two alternatives $\mathcal{Y}=\left\{ 1,2\right\} $ and the default option $o$. The researcher knows $\operatorname{P}_{1}$, $\operatorname{P}_{2}$, and $\operatorname{P}_{3}$. Let us consider Person 1. Let $\mathbf{y}$ be such that $y_1=o$. The probability that Person 1 selects the default option $o$ (given a profile of choices $\mathbf{y}$ with $y_1=o$) is
Under A3, we get that $2\in \mathcal{N}_{1}$ if and only if
Indeed, if $2\in \mathcal{N}_{1}$, then the probability of choosing the default option by Person 1 should decrease if Person $2$ picks something else. Also, if $2\not\in \mathcal{N}_{1}$, then the probability of choosing the default option by Person 1 should be invariant to choices of Person 2. Similarly, $3\in \mathcal{N}_{1}$ if and only if $\operatorname{P}_{1}\left( o \mid o,o,o\right) >\operatorname{P}_{1}\left( o \mid o,o,1\right)$. Thus, we can learn the set of friends of Person 1 from observed $\operatorname{P}_{1}$. Let us assume that $\mathcal{N}_{1}=\left\{ 2\right\}$. To recover the preferences of Person 1 note that
Thus, $2\succ_{1}1$ if and only if
Suppose that, indeed, we get that $2\succ_{1}1$. We can finally recover the attention mechanism (for $y_{_{1}}=o$) via the next four probabilities in the data
By considering two other choice profiles $\mathbf{y}$ with $y_1=1$ and $y_1=2$ (instead of $y_1=o$), respectively, we can fully recover the attention mechanism of Person 1. By a similar exercise we can recover the sets of friends, preferences, and the attention mechanisms for Persons 2 and 3. $\blacksquare $
This section studies identification of the conditional choice probabilities, $\operatorname{P}$, and the rates of the Poisson alarm clocks from two different datasets. These two datasets coincide in that they contain long sequences of choices from people in the network. They differ in the timing at which the researcher observes these choices. In Dataset 1 people's choices are observed in real time. This allows the researcher to record the precise moment at which a person revises her strategy and the configuration of choices at that time. In Dataset 2 the researcher simply observes the joint configuration of choices at fixed time intervals.
Let us assume the researcher observes people's choices at time intervals of length $\Delta$ and can consistently estimate $\Pr\left(\mathbf{y}^{t+\Delta }=\mathbf{y}'\mid\mathbf{y}^{t}=\mathbf{y}\right)$ for each pair $\mathbf{y}',\mathbf{y}\in\overline{\mathcal{Y}}^{A}$. We will capture these transition probabilities by a matrix $\mathcal{P}\left( \Delta \right)$. (Here again, we will assume that the choice configurations are ordered according to the lexicographic order when we construct $\mathcal{P}\left(\Delta \right)$.) The connection between $\mathcal{P}\left( \Delta \right)$ and the transition rate matrix $\mathcal{M}$ described in Equation ((ref)) is given by
where $e^{\left( \Delta \mathcal{M}\right)}$ is the matrix exponential of $\Delta \mathcal{M}$. The two datasets we consider differ regarding $\Delta$: in Dataset 1 we let the time interval be very small. This is an ideal dataset that registers people's choices at the exact time in which any given person revises her choice. As we mentioned earlier, with the proliferation of online platforms and scanner this sort of data might indeed be available for some applications. In Dataset 2 we allow the time interval to be of arbitrary size. The next table formally describes Datasets 1 and 2
In both cases, the identification question is whether (or under what extra restrictions) it is possible to uniquely recover $\mathcal{M}$ from $\mathcal{P}\left( \Delta \right)$. The first result in this section is as follows.
The proof of Proposition (ref) relies on the fact that when the time interval between the observations goes to zero, then we can recover $\mathcal{M}$. There are at least two well-known cases that produce the same outcome without assuming $\Delta \rightarrow 0$. One of them requires the length of the interval $\Delta$ to be below a threshold $\overline{\Delta }$. The main difficulty of this identification approach is that the value of the threshold depends on the details of the model that are unknown to the researcher. The second case requires the researcher to observe the dynamic system at two different intervals $\Delta_{1}$ and $\Delta_{2}$ that are not multiples of each other. (See, e.g., blevins2017identifying, and the literature therein.)
The next proposition states that, by adding an extra restriction, the transition rate matrix can be identified from people's choices even if these choices are observed at the endpoints of discrete time intervals. In this case, the researcher needs to know the rates of the Poisson alarm clocks or normalize them in empirical work.
The key element in proving Proposition (ref) is that the transition rate matrix in our model is rather parsimonious. To see why, recall that, at any given time, only one person revises her selection with nonzero probability. This feature of the model translates into a transition rate matrix $\mathcal{M}$ that has many zeros in known locations (see Example 1). It is important to remark that this aspect of the model is not shared by models of discrete-time revision processes, where people in the network revise choices simultaneously at fixed time intervals.
This section extends the initial model to allow randomness in preferences and in consideration sets. In this case, the choice rule $\operatorname{R}_{a}\left( \cdot \mid \mathcal{C}\right)$ from Section (ref) is not an indicator function but a distribution on $\mathcal{Y}$. We naturally let $\operatorname{R}_{a}\left( v \mid \mathcal{C}\right) =0$ if $v\notin \mathcal{C}$. Note that the relation between preferences and consideration sets is completely unrestricted.
Keeping unchanged the other parts of the model, the probability that Person $a$ selects (at the moment of choosing) alternative $v\in \mathcal{Y}$ is given by
The probability of selecting the default option $o$ is (as before) $\prod\nolimits_{v\in \mathcal{Y}}\left( 1-\operatorname{Q}_{a}\left( v \mid y_{a}, \operatorname{N}_{a}^{v}\left( \mathbf{y}\right) \right) \right)$. The next example illustrates the random choice rule with the well-known logit model.
Example 4. If we use the logit model to represent the random preferences of Person $a$, then the probability that the person selects alternative 1 when alternative 2 is also part of her consideration set would be given by
In this expression, $\mathrm{U}_{a}^{1}$ and $\mathrm{U}_{a}^{2}$ are the mean expected utilities that Person $a$ gets from alternatives 1 and 2, respectively. $\blacksquare $
Under this variant of the initial model, the identification of $\operatorname{P}$ follows directly from Propositions (ref) and (ref). We will thereby focus on recovering the set of connections, the choice rule, and the attention mechanism from $\operatorname{P}$. The main result is as follows.
Remark. The last result extends to the case in which the random preferences include the default option $o$ with only one caveat. In this case, the attention mechanism can be recovered up to ratios of the form $\operatorname{Q}_{a}\left( v \mid y_{a},\operatorname{N}_{a}^{v}\left( \mathbf{y}\right)\right) /\operatorname{Q}_{a}\left( v \mid y_{a},0\right)$. That is, we can only recover how much extra attention a person pays to each option as more of her friends select that option.
As in our previous results, under Assumptions A1-A3, observed variation in the choices of friends induce stochastic variation of the consideration sets and this variation suffices to recover the connections between the people in the network and the attention mechanism. The only difference with respect to the case of deterministic preferences is that with random preferences we need a larger number of friends for each person, i.e., $\left\vert \mathcal{N}_{a}\right\vert \geq Y-1$ for each $a\in \mathcal{A}$. The extra condition guarantees the matrix of coefficients for the $\operatorname{R}_{a}'$s in expression ((ref)) is full column rank. Indeed, we show that $\left\vert\mathcal{N}_{a}\right\vert \geq Y-1$ is not only sufficient, but necessary, to this end. We illustrate the last result by a simple extension of Example 3 above.
Example 3 (continued). Let us keep all the structure of Example 3 except for people's preferences, which we now assume are random. The identification of the set of connections and the attention mechanism follows from similar ideas. Thus, we will only focus on recovering $\operatorname{R}_{1}$, $\operatorname{R}_{2}$, and $\operatorname{R}_{3}$. Consider the following system of equations for Person 1 that connects the observed $\operatorname{P}$ with $\operatorname{Q}$ and $\operatorname{R}$
The fact that $\operatorname{R}_{1}\left( 1 \mid \left\{ 1\right\} \right)$ and $\operatorname{R}_{1}\left( 1 \mid \left\{ 1,2\right\} \right)$ can be recovered follows because, by Assumption A3, we have that the determinant of the matrix of the coefficients in the above system of equations is different from zero:
The extra condition, $\left\vert \mathcal{N}_{a}\right\vert \geq Y-1$, and Assumption A3 guarantee that the matrix of coefficients for the $\operatorname{R}_{a}$'s is always full column rank. $\blacksquare $
In the initial model, the default option plays a special role: it is chosen if, and only if, nothing else is considered. In some settings, such default option may not exist.\footnote{Also, see horan2019random.} This section offers a variant of the model that accommodates to this possibility.\footnote{For alternative ways to close the model see, for example, barseghyan2019discrete.}
Let us assume there is no default option $o$, so that $\overline{\mathcal{Y}}=\mathcal{Y}$. The formation process of the consideration set is as before except that, since there is no default option, we need to specify what people do when the consideration set is empty. Given the dynamic nature of our model, we will simply assume that each person sticks to her previous choice if no alternative receives further consideration. Formally, the probability that Person $a$ selects (at the moment of choosing) alternative $v\in\mathcal{Y}$ is given by
In our initial model, the probability that Person $a$ faces consideration set $\mathcal{C}$ given a choice configuration $\mathbf{y}$ takes the form of
This approach entails a multiplicative separable specification of alternatives in the consideration sets. It also assumes that the probability of considering an option depends only on the total number of friends who pick that option, but not on the identity of these friends (i.e., the effect of friends' choices is symmetric across friends). We next show that neither of these assumptions is essential for our approach. Indeed, all previous insights can be used here to extend the initial results.
For each Person $a$ and configuration $\mathbf{y}$, let $\eta_{a}(\cdot \mid \mathbf{y})$ be an attention index function from $2^{\mathcal{Y}}$ to the positive reals. The value $\eta_a(\mathcal{C} \mid \mathbf{y})$ captures the attention that Person $a$ pays to the set of alternatives $\mathcal{C}\in2^{\mathcal{Y}}$ given the choice configuration $\mathbf{y}$. The attention-index measures how enticing a consideration set is (see aguiar2019doesrandom for further details). We define the probability of facing consideration set $\mathcal{C}$ as
In this model, the consideration set probabilities are as in brady2016menu. Since $\eta_a$ can be identified only up to scale, we normalize the attention-index for the empty set to be 1 (i.e., $\eta(\emptyset \mid \mathbf{y})=1$ for all $\mathbf{y}\in \overline{\mathcal{Y}}^{A}$).\footnote{If one instead normalizes $\sum\nolimits_{\mathcal{D}\subseteq\mathcal{Y}}\eta_a(\mathcal{D} \mid \mathbf{y})=1$, then we get the consideration rule of aguiar2017random.} This more general setting covers our initial model, which is based on manzini2014stochastic, as a special case. By combining preferences and this specification of stochastic consideration sets, the probability that Person $a$ selects (at the moment of choosing) alternative $v\in \mathcal{Y}$ is given by
The probability of selecting the default option $o$ is just $\eta_{a}\left( \emptyset \mid \mathbf{y}\right)/\sum\nolimits_{\mathcal{D}\subseteq\mathcal{Y}}\eta_a(\mathcal{D} \mid \mathbf{y})$. To better understand the connection between our initial model and this extension note that, for any $v\in\mathcal{Y}$ and $\mathcal{C}\subseteq\mathcal{Y}$ such that $v\not\in\mathcal{C}$, we have that \[ \dfrac{\eta_a(\mathcal{C}\cup\{v\} \mid \mathbf{y})}{\eta_a(\mathcal{C}\mid \mathbf{y})}=\dfrac{\operatorname{Q}_a(v\mid \mathbf{y})}{1-\operatorname{Q}_a(v\mid \mathbf{y})}. \] Thus, \[ \operatorname{Q}_a(v\mid \mathbf{y})=\dfrac{\eta_a(\mathcal{C}\cup\{v\} \mid \mathbf{y})}{\eta_a(\mathcal{C}\cup\{v\} \mid \mathbf{y})+\eta_a(\mathcal{C}\mid \mathbf{y})}. \] Based on this alternative specification, we accommodate Assumptions A1 and A3 as follows.
(A1') For each $a\in \mathcal{A}$, $v\in\mathcal{Y}$, and $\mathbf{y}\in \overline{\mathcal{Y}}^{A}$, there exists $\mathcal{C}\in 2^\mathcal{Y}$ such that $v\succ_a v'$ for all $v'\in\mathcal{C}$ and $\eta_{a}\left(\mathcal{C}\cup\{v\} \mid \mathbf{y}\right)>0$.
(A3') For each $a\in \mathcal{A}$, $\mathcal{C}\in 2^\mathcal{Y}$, and $\mathbf{y},\mathbf{y}^*\in \overline{\mathcal{Y}}^{A}$, such that $\mathbf{y}$ is different from $\mathbf{y}^*$ just in one component $a^*$,
Assumption A3'(i) states that the attention a person pays to a given set is invariant to the choices of those who are not connected with the person and to alternatives that do not enter the set. Assumption A3'(ii) means that switches of friends to a new alternative boost the attention for all sets that contain this new alternative. Note that Assumption A3' does not assume that different friends affect consideration probabilities symmetrically. That is, the model allows the possibility that some friends have a bigger effect than others. Since the probability of facing consideration set $\mathcal{C}$ is proportional to the inverse of the total attention, $\sum\nolimits_{\mathcal{D}\subseteq\mathcal{Y}}\eta_a(\mathcal{D} \mid \mathbf{y})$, even if peers are switching between alternatives that are not elements of $\mathcal{C}$, the probability of facing $\mathcal{C}$ may change.
The next proposition states that, with this general alternative specification, the network structure and the profile of preferences can be uniquely recovered from the conditional choice probabilities. Though the attention mechanism is just partially-identified, it is point identified under additional restrictions that reduce the dimensionality of the problem; we describe some of these restrictions in the next result.
Condition (i) in Proposition (ref) restates our main result using the consideration set formation model of manzini2014stochastic ---see also manski1977structure. Condition (ii) shows that the model analyzed in dardanoni2020inferring ---where the sets of equal cardinality are considered with equal probability--- also imposes enough restrictions to recover consideration probabilities. Conditions (iii) and (iv) are new. Condition (iii) is similar to condition (i), but defines “aggregate” attention as a sum of singleton-attentions. Condition (iv) postulates that consideration probability of a set only depends on the best alternative in that set.
This section provides an estimator of the model parameters that uses discrete-time data (Dataset 2) and assesses its quality in Monte Carlo simulations. Let $\theta =\left( \Gamma ,\succ,\operatorname{Q}\right)$ be an element of the space of possible parameters we want to estimate. (We normalize the intensity parameter $\lambda_{a}$ to $1$ for all $a\in\mathcal{A}$.) Note that since the number of people in the network and the number of choices is finite, the parameter space for $\Gamma$ and $\succ$ is a finite set. For each $\theta$ we can construct the transition rate matrix $\mathcal{M}\left( \theta \right)$ using Equations ((ref)) and ((ref)). In turn, this information allows us to calculate the transition matrix
Given a sample of network configurations $\{\mathbf{y}_{t}\}_{t=0}^T$, we can use the latter to build the log-likelihood function $\mathrm{L}_{T}\left(\theta \right) =\Sigma_{t=0}^{T-1}\ln \mathcal{P}_{\iota \left( \mathbf{y}_{t}\right),\iota \left(\mathbf{y}_{t+1}\right) }\left( \theta ,\Delta\right)$, where $\iota \left(\mathbf{y}\right) \in \left\{ 1,2,\dots,\overline{\mathcal{Y}}^{A}\right\}$ is the position of $\mathbf{y}$ according the lexicographic order, and $\mathcal{P}_{k,m}\left( \theta,\Delta \right)$ is the $(k,m)$-th element of the matrix $\mathcal{P}\left(\theta ,\Delta \right)$. Finally, the maximum likelihood estimator of the true parameter value can be defined as
To evaluate the finite-sample properties of our estimator we conducted several Monte Carlo experiments. In all experiments we simulated discrete-time data from the specification presented in Example 2 in Section (ref). In particular, we assumed that \[ 2\succ_{1}1,\quad 1\succ_{2}2,\quad 2\succ_{3}1,\quad 1\succ_{4}2,\quad 1\succ_{5}2, \] \[ \operatorname{Q}\left( v \mid 0\right) =\frac{1}{4},\quad\operatorname{Q}\left( v \mid 1\right) =\frac{3}{4},\quad\operatorname{Q}\left( v \mid 2\right) =\frac{7}{8}, \] and \[ \mathcal{N}_{1}=\left\{ 2,3\right\},\quad \mathcal{N}_{2}=\left\{1,3\right\},\quad \mathcal{N}_{3}=\left\{ 1,2\right\},\quad \mathcal{N}_{4}=\left\{ 5\right\},\quad \mathcal{N}_{5}=\left\{ 4\right\}. \] First, we estimate $\operatorname{Q}$ under the assumption that the network structure and preferences are known. The experiment was replicated 1000 times for 6 different sample sizes. The results of these simulations are presented in Table (ref). The estimator of $\operatorname{Q}$ performs well in terms of the mean bias and the root mean squared error. As expected, the bias and the root mean squared error decrease with the sample size. (See Appendix (ref) for more details on how the data was generated.)
Next we estimate the whole parameter vector since our method allows to consistently estimate the network structure and the preference orders of individuals as well. Since in our example, without any restrictions, there are $2^{A(A-1)}=1,048,576$ possible networks and $(Y!)^{A}=32$ strict preference orders, to make the problem computationally tractable we restricted the parameter space for $\Gamma$ by making the following assumptions: (i) each person has at most two friends; (ii) the attention mechanism is invariant across people and alternatives; and (iii) the network is undirected. As a result, the number of possible networks becomes $112$ (the number of possible preference orders is still $32$). The experiment was conducted $500$ times for different sample sizes. Table (ref) presents the results of these simulations. With just $50$ observations the network structure is correctly estimated $94.4$ percent times. For the sample size of $500$ the network structure and the preferences are correctly estimated in all simulations.
In this section, we apply our framework to an experimental dataset used by bai2019predicting to test models that aim to predict the visual focus of attention of people. In each experiment, a group of people were asked to play the Resistance game. This is a party game in which people communicate with each other in order to find the deceptive players.\footnote{See \url{https://en.wikipedia.org/wiki/The_Resistance_(game)} for a more detailed description of the game.} The game was implemented in several rounds and lasted about 30 minutes. For the purpose of our analysis, it is important to remark that all participants were seated around a circle. A tablet was placed in front of each player and it recorded the direction of the player's sight every 1/3 of a second. The structural approach we offer allows us to disentangle and estimate two key forces in explaining the data: directional sight preferences of players and peer effects in gazing.
The two determinants of visual focus of attention that we want to recover for each person are motivated by two well-known observations: First, visual designers create online platforms under the premise that people scan certain areas of the screen before others. In particular, it is believed that people spend more time looking to the left side of the screen as compared to the right side.\footnote{See, for example, \burl{https://www.nngroup.com/articles/horizontal-attention-original-research/}.} The so-called left-to-right bias has been also documented by the experimental studies of spalek2004supporting,spalek2005left, and reutskaja2011search.\footnote{See also maass2003directional and reference therein.} Second, in social environments, it has been observed that people automatically redirect their visual attention by following others' gaze orientation, a phenomenon called gaze following.\footnote{See, for example, gallup2012visual and \burl{https://www.nationalgeographic.com/science/article/what-are-you-looking-at-people-follow-each-others-gazes-but-without-a-tipping-point\#close.21}.} We believe our model is particularly well-suited to separate these two forces from the observed repeated choices.
Before presenting the main results, we describe the dataset we use and the main assumptions we invoke in the application.
\noindentData The dataset contains five experiments of up to eight people. For our study, we use the data coming from the two experiments that have five players. We refer to them as Experiments 1 and 2. Experiments 1 and 2 contain 4248 and 7875 observations, respectively. Although the frequency of sight measure seems to be rather high (3 measurements per second), 19 and 26 percent of the observations in Experiments 1 and 2, respectively, have more than one player changing the direction of sight between consecutive measurements. Thus, we treat the datasets generated by the two experiments as discrete time data (i.e., Dataset 2 in Section (ref)). \noindentChoice sets To capture possible preferences for direction in visual focus of attention, we aggregate the data to three options: left, right, and tablet.\footnote{This aggregation has several other advantages. For instance, the induced model has more players than alternatives. It also reduces the cardinality of the outcome space and the possible number of preference orders ($3^5=243$ vs. $5^5=3125$).} (The original dataset has five options as each person can look at one of her four opponents or her tablet.) As a result, every measurement indicates the direction of sight of every player at the moment of recording (e.g., Player 1 looks to the left). Formally, from the initial dataset we construct a new dataset on the joint configuration of choices $\mathbf{y}=(y_a)_{a\in\mathcal{A}}\in {\mathcal{Y}}^A=\times_{a\in\mathcal{A}}\{t_a,l_a,r_a\}$, where: $y_a=t_a$ if, at the moment of recording, Player $a$ looks at her own tablet; $y_a=l_a$ if Player $a$ looks to her left; and $y_a=r_a$ if Player $a$ looks to her right. Players were seated clockwise. Thus, the left side for Players $a=1,2,3$, is defined as $t_a=\{a+1,a+2\}$ and for Players 4 and 5 it is $t_4=\{5,1\}$ and $t_5=\{1,2\}$, respectively. Similarly, the right side is defined as $r_a=\{a-1,a-2\}$ for Players $a=3,4,5$, and as $r_2=\{1,5\}$ and $r_1=\{5,4\}$ for Players 2 and 1, respectively. Note that the choice sets differ across players.\footnote{For instance, if Players 1 and 4 look to the left, they look at different sets of players: Player 1 looks at either Player 2 or Player 3, Player 4, however, looks at either Player 5 or Player 1.} This feature can be incorporated in our model, as none of our results crucially depend on it. \noindentConsideration probabilities and choices of friends We assume that consideration probabilities on the direction of sight (i.e., left, right, or tablet) vary with the number of peers that look in the same direction. That is, each player is more likely to consider looking at a particular area if other players are doing so. For each Player $a$ we say Player $a'$ is looking in the direction $y_a$ if $y_a\cap y_{a'}\neq \emptyset$. For example, suppose that (at the moment of deciding where to look at) Player 1 observes that $y_2=t_2$, $y_3=r_3$, $y_4=l_4$, and $y_5=l_5$. That is, $3$ players are looking in the direction of Player 1 (Players $3$, $4$, and $5$), $3$ players are looking in the direction of Player 2 (Players $2$, $3$, and $5$), only Player 4 is looking in the direction of Player $5$, and no one is looking in the direction of Players 3 and 4. As a result, we state that 3 people are looking to the “left” side of Player 1 ($l_1\cap y_{a'}\neq\emptyset$ only for $a'=2,3,5$), 1 person is looking to the “right” side of Player 1 ($r_1\cap y_{a'}\neq\emptyset$ only for $a'=4$), and 3 people are looking at Player 1 ($t_1\cap y_{a'}\neq\emptyset$ for $a'=3,4,5$). In this situation, the probability that Player 1 considers left is $\operatorname{Q}_1\left( l_1 \mid 3\right)$, right is $\operatorname{Q}_1\left( r_1 \mid 1\right)$, and tablet is $\operatorname{Q}_1\left( t_1 \mid 3\right)$. \noindentNetwork structure Given the small number of players and the fact that participants are playing a party game with monetary prizes, we think it is natural to assume the network is complete. That is, we assume each player is connected to all other players in the group. \noindentPreferences and the default option Since, in this setting, there is no reason for us to treat one of the directions as the default option (i.e., always considered and the least preferred), we use the specification of the model with no default option presented in Section (ref) for the estimates. Also, we do not impose any restrictions on preferences neither for each person nor across different people. Thus, in total, we have 3 possible preference orders for each player and $3^5=243$ possible combinations of preference orders for the five players. Before presenting the main estimates, let us display some raw data description. Using the data from the experiments we first construct marginal shares of each option for every player in Experiments 1 and 2. This information is displayed in Tables (ref) and (ref) below.
Looking at Table (ref) we see that, in Experiment 1, all the players spend very little time looking at the tablet. In addition, all players but Player 1 look to the left more often than to the right. Table (ref) displays rather different results for Experiment 2. In particular, the second experiment has more balanced shares. Also, looking at own tablet is not longer the least selected option for all the players. For example, Player 2 looks to her tablet more often than to the left or to the right. In addition, three out of the five players look to the right more often than to the left. Based on these initial results, one may expect considerable differences in the two experiments regarding preferences for directional sight. Interestingly, we get that this is not the case.
We estimate a few specifications of the model. We present two of them next and display the other ones in Appendix (ref). In Model I(a), we assume there is no heterogeneity in consideration probabilities across options and players, i.e., $\operatorname{Q}(\cdot)=\operatorname{Q}_{a}\left( v \mid \cdot\right)$ for all $a$ and $v$. Model I(b) adds heterogeneity in consideration probabilities across players. All the models we estimate allow for unrestricted heterogeneity in preferences regarding directional visual sight. Figure (ref) shows the estimates for the consideration probabilities of Model I(a) for the two experiments. Note that although monotonicity of $\operatorname{Q}$ in the number of friends was not imposed in the estimation, the estimated probabilities are indeed monotone in Experiment 2 and are very close to be monotone in Experiment 1. The estimated preferences for directional sight coincide for all players in the two experiments. Specifically, all players prefer looking to the left, then to tablet, and then to the right. That is, we obtain that, for each $a\in\mathcal{A}$,
Thus, in the first specification of our model, all the players (in Experiments 1 and 2) show the left-to-right bias.
Figure (ref) shows the estimates for the consideration probabilities for Model I(b) in Experiments 1 and 2. Similar to the case with homogeneity, many of the consideration probabilities for each direction of sight are indeed increasing in the number of players looking at that direction. The estimated preference orders coincide in Models I(a) and (b), for the two experiments. That is, here again, all the players prefer looking to the left, then to tablet, and then to the right.
To sum up, despite considerable differences in raw marginal shares (see Tables (ref) and (ref)), the preferences for directional sight of all players in the two experiments coincide. Specifically, all players prefer to look from left to right after controlling for unobserved consideration sets that are determined by visual focus of attention of others. As a robustness check, in order to address possible concerns with the dynamic of players' interactions in different stages of the game, we estimated the previous two models using the last half of observations and the middle half of observations (i.e., we did not use the first and the last quarter of observations). In both cases the estimated consideration probabilities are similar to the ones estimated using the whole sample, and the estimated preferences of all players are again $l_a\succ_a t_a\succ_a r_a$ (see Appendix (ref) for further details). As we mentioned earlier, this type of preferences are consistent with the left-to-right bias invoked by online platform designers that has been also documented by experimental work.
This paper adds peer effects to the consideration set models. It does so by combining the dynamic model of social interactions of blume1993statistical, blume1995statistical with the (single-agent) model of random consideration sets of manski1977structure and manzini2014stochastic. The model we build differs from most of the social interaction models in that the choices of friends do not affect preferences but the subset of options that people end up considering. In our model, the network structure affects the nature and strength of the mistakes that people make. Thus, we can think of our model as a model of peer effects in mistakes.
From an applied perspective, changes in the choices of friends induce stochastic variation of the considerations sets. We show that this variation can be used to recover the main parts of the model. On top of nonparametrically recovering the preference ranking of each person and the attention mechanism, we identify the set of connections or edges between the people in the network. The identification strategy allows unrestricted heterogeneity across people.
We propose a consistent estimator of model parameters and apply it to an experimental dataset. The structural approach we offer allows us to identify and estimate two main determinants on people visual focus of attention: directional sight preferences and peer effects in gazing. Our results are consistent with the existing literature. We believe our approach to peer effects in consideration sets could be incorporated in various empirical studies of technology adoption and diffusion. For instance, aral2009distinguishing develop a dynamic matched sample estimation framework to distinguish influence and homophily effects in the day-by-day adoption of a mobile service application (i.e., Yahoo! Go). A similar dynamic matched sample estimation framework could be used to recover preferences over adoption and the influence of neighbor adoption rates on likelihood of considering the possibility of incorporating the mobile service application.
\phantomsection\addcontentsline{toc}{section}{\refname}