Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
88,218 characters · 14 sections · 89 citation commands
Revealed Social Networks
People rarely make decisions in isolation and are often influenced by their neighbors and peers.\footnote{Social influence has been studied in an array of contexts including test scores sacerdote2001peer, worker productivity mas2009peers, alcohol use kremer2008peer, risky behavior card2013peer, and tax compliance fortin2007tax.} This influence can take the form of a social norm, a common convention, or conformism. The linear-in-means model serves as a foundational framework for estimating peer effects in economics and social science research. The model supposes that an agent's choice is a convex combination of their own ideal point and the weighted average of each other agent's choice. The objective of this paper is to investigate the behavioral implications of the linear-in-means model leveraging revealed preference techniques \`a la Afriat. Moreover, we establish sufficient conditions to jointly recover each agent’s preferences and the underlying social network, including the strength of every connection.
As manski1993identification emphatically observed, the identification of social influence parameters has long been recognized as a challenging task. The crux of the reflection problem lies in disentangling and separately identifying exogenous group effects from endogenous peer effects. In this paper, we tackle the identification of network effects through the lens of stochastic discrete choice theory—an approach that yields several novel, intuitive, and surprisingly powerful insights.
We consider a choice procedure based on the model of manski1993identification wherein each agent chooses an action which is a convex combination of their ideal point and a weighted average of the other agents' actions. Formally,
where the vectors of $v_i$ and $p_i^N$ correspond the agent $i$'s ideal point and the action taken by agent $i$ in group $N$, respectively. Further, $\pi_i^N(j)\geq 0$ denotes the impact of agent $j$'s action on agent $i$ in the context of group $N$. While $p_i^N$s are observable, $v_i$ and $\pi_i^N$s are not observable and need to be identified from observed choices.
Consider three college friends—Ann, Ben, and Can—and their sports choices. While Can was studying abroad, Ann chose to play tennis 50% of time, volleyball 10% of the time, and walk 40% of the time, while Ben's frequencies were 70%, 10%, and 20%, respectively. This data alone does not solve the reflection problem. An observer cannot discern whether behavioral similarities arise from peer influence or similar preferences/backgrounds that would produce identical choices even without interaction. These explanations remain observationally equivalent without further data. Our approach resolves this by leveraging natural group variation. For instance, suppose that Can returns but Ben leaves to study abroad the following semester. Now, we observe that Ann's choices are tennis (10%), volleyball (50%), and walking (40%), while Can's frequencies were 10%, 70%, and 20%, respectively. By comparing choice behavior across the two periods (the first semester choices of Ann and Ben without Can and the second semester choices of Ann and Can without Ben) we can uniquely isolate peer effects (point identification), provided the observed choices satisfy our model's assumptions (formalized later).\footnote{In this example, our unique identification reveals that Ann's favorite exercise is walking, even though it is not the most frequently chosen by Ann. Indeed, her ideal point is $v_{Ann}=(0.1,0.1,0.8)$. On the other hand, tennis and volleyball are Ben's and Can's favorite sports, respectively. The data further show that Ann weighs her friends' choices twice as heavily as her own ideal point (a 2:1 ratio).}
Our identification strategy relies on (1) group variation $\{$Ann, Ben$\}$ vs $\{$Ann, Can$\}$ and (2) multi-dimensional choice objects (tennis, volleyball, walking).\footnote{In this paper, our main form of variation is group/participation variation. However, our main results can be applied in a setting where there is no group variation but there is variation of some observable variable that is a perfect instrument for variation in the underlying network structure of a fixed group.} A key insight of our analysis is that a necessary condition for recovery of the underlying network structure is that the number of agents in each of our groups must be no more than the dimension of our outcome variable. Notice that we have made no assumptions (1) on the underlying social network structure across groups, (2) on the observable characteristics of these activities, and (3) on the preferences of individuals. Hence, our identification results complement existing work relying on these variations/assumptions.\footnote{Much of the work on identifying parameters in the linear-in-means model assumes (partial) observability of the underlying network structure bramoulle2009identification,blume2015linear. Recently, lewbel2023social and de2024identifying provide sufficient conditions under which the underlying network structure can be recovered without data on the network itself. Unlike ours, these papers require enough variations on observable characteristics. Since our approach is complementary to these papers, one might allow a more general model where point identification is possible under weaker conditions.}
Given that our identification results rely on the particular choice procedure we have adopted, it is natural to consider the falsifiability of this model. We show that our model is fully characterized by a version of classical no-money-pump conditions in the spirit of afriat1967construction. This result renders our model behaviorally testable without any restrictions across groups or about the weights assigned to each agent. Our test characterizes datasets of the previously described form which are consistent with the linear-in-means model via an easily solvable linear program. We interpret this linear program as a no money-pump condition on an outside observer who is making bets on the choices of an agent. Unlike standard no-money-pump conditions, which are typically given by two conditions, feasibility of a bet and (expected) profitability of the bet, our condition has a third part which imposes incentive compatibility of a bet. Incentive compatibility captures the idea that if the outside observer is betting on one decision maker across each group, there is no group where they would prefer to bet on a different decision maker.\footnote{Testing models of peer effects is difficult as many people choose peers who are similar to them in observable characteristics. As a result of this difficulty, many studies that aim to test models of peer effects, including the linear-in-means model, do so through natural/quasi/pure experimental methods sacerdote2014experimental,basse2024randomization. In addition, experimental methods are also often used when quantifying the impact of peer effects. See agranov2021importance as an example.}
Thus far, we have made no assumptions about the underlying social network structure across groups. While this is a strong point of our testing and identification results, this poses a problem for the predictive power of our model. To solve this problem, we consider a refinement of the linear-in-means model, which assumes a common social network structure across groups. The strength of network connections from one person to another only varies across groups due to renormalization. Under this assumption, we provide sufficient conditions that allow us to predict choices in any possible group. To test the validity of this assumption, we develop an extension of our testing procedure for the general case. This version of the linear-in-means model is characterized by a strengthening of our previous no incentive compatible money pump condition. The updated version of incentive compatibility corresponds to the idea that if the outside observer is betting on one decision maker across each group, there is no other agent that the outside observer would rather bet on across all of the same groups.
Finally, we consider a version of the linear-in-means model more in line with the original reflection problem posed by manski1993identification. In this version of the model, each person influences each other person uniformly. This corresponds to the unweighted average choice of the group being a common social norm within the group. In this case, we provide a test in terms of a finite set of linear inequalities. In the context of linear social influence models, such as the linear-in-means model, the influence one person exerts on another is proportional to the difference in their choices. Following this logic, we call the difference between agent $i$'s choice and agent $j$'s choice the peer effect of agent $j$ on agent $i$. This version of the model is characterized by three restrictions on the peer effects between agents. Our first axiom asks that the peer effect of agent $j$ on agent $i$ is group invariant. Our second axiom is a condition about the symmetry of agent $i$'s peer effect across groups. The last axiom asks that the total peer effect on agent $i$ in group $N$ is bounded above by agent $i$'s choice in group $N$.
The rest of this paper is organized as follows. In Section (ref), we formally introduce the linear-in-means model and our notation. In Section (ref), we introduce and discuss our testing and identification results for the three different specifications of the linear-in-means model. In Section (ref), we discuss how our results extend to alternative types of data. Finally, we conclude with a discussion of the related literature in Section (ref).
Our interest is in studying the linear-in-means model of social interaction. We build on the base model in two meaningful ways. First, instead of restricting an agent's choice to be from an interval, we allow agents to choose a distribution over a finite set of goods. Second, we consider stochastic choice data that arises when the set of agents present varies.
Denote by $\mathcal{A}$ the grand set of agents. Assume that $|\mathcal{A}| \geq 2$. A typical group of agents will be denoted $N$, where $\varnothing\neq N \subseteq \mathcal{A}$. We let $\mathcal{N}\subseteq 2^{\mathcal{A}}$ be any set of groups. For any agent $i\in \mathcal{A}$, let $\mathcal{N}_i=\{N\in\mathcal{N}:i\in N\}$ denote the set of groups to which $i$ belongs. We sometimes abuse notation and use $N \setminus i$ to denote the group of agents formed by removing agent $i$ from group $N$. Let $X$ be some finite set of alternatives. Enumerate these alternatives from $1$ to $|X|$. We use $(0,\dots,0,1,0,\dots,0)$, where $1$ is in the $n$th dimension, to denote the $n$th alternative and use $x,y \in X$ to denote arbitrary alternatives. The data in our model consists of a probability distribution over $X$, for each $N\in\mathcal{N}$ and each $i\in N$. Formally, for $i\in N\in\mathcal{N}$, this is denoted $p_i^N\in\Delta(X)$. We use $p_i^N(x)$ to denote the choice probability of good $x$ by agent $i$ in group $N$ and $p^N$ to denote the matrix where each row corresponds to $p_i^N$ for a different $i \in N$. We will sometimes use $p_{-i}^N$ to denote the matrix which is formed by taking $p^N$ and removing the row corresponding to agent $i$.
In the linear-in-means model, each agent's action is a mixture of their own ideal point and the actions of others. We use $v_i \in \Delta(X)$ to denote agent $i$'s ideal point. This corresponds to the action agent $i$ would take in isolation. We introduced $p_i^N$ as our data, but it also corresponds to the action taken by agent $i$ in group $N$. The amount that agent $j$'s action impacts agent $i$'s action may differ from the amount that agent $k$'s action impacts agent $i$'s action. In fact, the impact of agent $j$'s action on agent $i$ may depend on the context or the group of agents currently present. We use $\pi_i^N(j)\geq 0$ to denote the impact of agent $j$'s action on agent $i$ in the context of group $N$. Similarly, we use $\pi_i^N(i)>0$ to denote the impact of agent $i$'s ideal point on agent $i$ in group $N$. We assume that $\sum_{j \in N}\pi_{i}^N(j)=1$. These impact or influence weights enter into agent $i$'s action in a linear manner.
Equation (ref) tells us that agent $i$'s action in group $N$ is given as a convex combination of their ideal point and a weighted average of the other agents' actions. We now discuss three key modeling assumptions which differ from the most general version of the linear-in-means model.
Our first assumption is that an agent's outcome is restricted to the simplex. This means that the total value of agent $i$'s outcome and the total value of agent $j$'s outcome are restricted to be the same. This rules out situations where the total value of each agent's outcome can vary, either across groups or across agents.\footnote{One such example of a situation is the case of exam scores. Suppose each dimension corresponds to an agent's score on an exam in a specific subject. If we want to allow for the sum of scores to vary across agents or groups, this simplex assumption is a meaningful restriction.} In Section (ref) we discuss how our main results on testing and identification extend to the case when outcomes are allowed to be anywhere in a convex subset of Euclidean space.
Our second key assumption is that $\pi_i^N \geq 0$ and $\sum_{j \in N} \pi_i^N(j) =1$. In other words, we assume that the total influence an agent faces, including their impact on themselves, is non-negative and constant across groups.\footnote{This assumption on the linear-in-means model is partially due to our data assumptions. The fact that each $p_i^N$ lies in the simplex necessitates $\sum_{j \in N} \pi_i^N(j)=1$. If this were not the case, our data would lie outside the simplex.} Under this assumption, the linear-in-means model is best interpreted as a model of conformism. Our agent is influenced by the weighted average of the other agents' choices and has some incentive to match this average. In our main results, we cannot dispense with this assumption.
Our last key assumption is about the value of $\pi_i$ across the dimensions of the choice problem. Since we are working in a multidimensional setup, we make assumptions on our network which have no content in the one-dimensional case. In our setup, since we are considering multiple dimensions, the importance of agent $j$ to agent $i$ could in theory depend on the dimension or good in consideration. Our specification of the linear-in-means model imposes that these influence parameters, $\pi_i^N(j)$, are constant across each dimension. Our main results also rely on this assumption and it cannot be relaxed.
While this assumption is not applicable to every environment (see belhaj2014competing and zenou2024games), we argue that this assumption is reasonable in certain cases. As a first example, we suggest there may be certain sets of choices that share the same network. Consider the choices of professors in a department. When considering how to split their time between research, teaching, and departmental service, it is reasonable to expect that one professor's influence on another is constant across these dimensions. As a second example, we suggest that this assumption may be more reasonable for agents in certain age ranges (or more generally, specific groups). Our prime example is the choice of children or teenagers who have yet to build distinct social networks (i.e. work colleagues, research networks, in-laws, etc.) and are mostly impacted by their school friends.
A key part of our analysis is observing data across various groups. As there are no restrictions on who is present in each of these groups, the underlying network structure can vary from group to group. This group and network variation can be modeled through different assumptions on $\pi_i^N$. Intuitively, $\pi^N$ is a matrix that captures weighted directed influence. By assuming that $\sum_{j \in N} \pi_i^N(j)=1$, we are assuming that $\pi^N$ is a stochastic matrix. When we observe group variation, each group $N$ is subject to its own $\pi^N$. To demonstrate the richness of our framework and motivate the analysis to follow, we discuss seven examples of linear-in-means models.
Each model in the list above is interesting in its own right. However, due to space constraints, a comprehensive study of all of them is infeasible. Instead, we focus on the extreme cases--namely, the ULM, LLM, and GLM. Our analysis of the most general form of the GLM will provide insights into identification across these models. That said, each model benefits from additional identification power due to its specific structural assumptions. We therefore encourage future research to explore these models in greater depth.
Our goal in this paper is to study data that arises from the ULM, LLM, and GLM versions of the linear-in-means model. Our focus is on testing these models and identifying $v_i$ and $\pi_N$. With this in mind, we introduce our definition of consistency.
Note that these three models are nested: ULM $\subset$ LLM $\subset$ GLM.
In this section, we discuss two foundations for the linear-in-means model as well as potential data-generating procedures under these foundations. In the standard setup of the linear-in-means model, when $|X|=2$, which we call the one-dimensional case as one dimension is a sufficient statistic for the other, $p_i^N$ is often interpreted as an effort level or a test score. Notably, our focus will be on situations where our data can be multidimensional. Our model allows us to capture higher granularity in choice. As we will see in Section (ref), the multi-dimensional aspect of our model allows us to recover generic joint identification of preference and network parameters which generically fails in the one dimension case.\footnote{As we will see later, the reflection problem of manski1993identification is partially a result of the one dimension case being the standard case.}
Our first interpretation of the linear-in-means model is as arising as the equilibrium best response of a complete information game. Consider two agents who have preferences over how they split their time between labor, leisure, and volunteering. In addition, these two agents have a desire to spend time with each other. Upon making their preferences known, the agents commit to how they spend their time. In the context of this story, an analyst has access to data on how these two agents have used their time across these three dimensions. Formally, each agent has preferences over the elements of the simplex but faces perturbations of their preferences due to the choices of their peers.\footnote{In this case, an agent can be thought of as having a preference over their choice frequencies rather than repeatedly maximizing (potentially different) static utility functions. This corresponds to an agent who is deliberately stochastic cerreia2019deliberately or faces a perturbed utility function fudenberg2015stochastic subject to social influence.} Specifically, the choices described in Equation (ref) arise from the group $N$ playing a complete information game where each agent's utility is given by the following.\footnote{In Equation (ref), note that if we choose to optimize in each dimension, ignoring across dimension constraints, we are left with $p_i^N(x)$ satisfying Equation (ref). Since each dimension is subject to the same $\pi_i^N$, the proposed $p_i^N(x)$ satisfy non-negativity and $\sum_{x \in X} p_i^N(x) =1$. Notably, if we relax the assumption of $\pi_i^N(x)=\pi_i^N(y)$, allowing for non-trivial multiplexing across dimensions, our non-negativity constraints may potentially bind and our formulation would not result from optimization of a quadratic loss utility function. See belhaj2014competing for a discussion of this in the two dimension case.}
Under this interpretation, each agent's connection weights $\pi_i^N(j)$ depend on the relative costliness of deviating from agent $j$'s choice or their own bliss point. This is in line with the observation made in blume2015linear, boucher2016some, kline2020econometric, and ushchev2020social that the one-dimensional linear-in-means model arises from agents maximizing quadratic loss functions. Under this interpretation of the model, we can think of our data as arising from each agent's choice of time usage. We also note that golub2020expectations offers an interpretation of the linear-in-means model as arising from an incomplete information game where, in Equation (ref), $v_i$ corresponds to agent $i$'s expectation of some underlying random variable and $p_i^N(j)$ corresponds to agent $i$'s expectation of agent $j$'s action.
Our second interpretation of the linear-in-means model is as the time average of a discrete choice procedure. Consider two agents who repeatedly go out on dinner dates with each other. Each agent has a base random utility function over the food on the menu, but this utility function is perturbed by a preference to conform with their date's expected choice. In this story, an analyst has access to the repeated choices or the long run average of these repeated choices of each agent. Formally, in each period, each agent chooses the alternative which maximizes
where $c_j$ corresponds to the choice of agent $j$ and $\mathbb{E}$ corresponds to the expected value. We make two further assumption. First, $\epsilon_x$ is distributed according to the Gumbel distribution, leading to logit choice frequencies mcfadden1974conditional.\footnote{Note that the assumption of Gumbel errors can be relaxed given that we jointly change the specification of the utility function.} Second, each agent's beliefs are stationary and correct. This means that $\mathbb{E}\left[\mathbf{1}\{c_j=x\}\right]$ is equal to $p_j^N(x)$.\footnote{While we make no claim about convergence of such a process, this correct expectations assumption can be thought of as arising from a stationary distribution of the process where, in each period, agents choose to maximize $\ln\left(\pi_i^N(i) v_i + \sum_{j \in N \setminus i}\pi_{i}^N(j) \hat{p}_j^N(x)\right) + \epsilon_x$ where $\hat{p}_j^N(x)$ is the empirical frequency thus far of the choice of $x$ by agent $j$ in group N.} Under these two assumptions, the long run choice frequencies are given by Equation (ref).
In this section, we study the behavioral implications of the general, the Luce, and the uniform linear-in-mean models. We begin with the general case and provide two characterizations of datasets consistent with the general case as well as partial and point identification results for $v_i$ and $\pi^N$. Our first characterization is via the non-empty intersection of a collection of convex sets. These convex sets correspond to the set of $v_i$ which feasibly induce the observed choice of agent $i$ the context of a given group. Our partial identification of $v_i$ builds on this result. Our second characterization provides an existential linear program which fails to hold if and only if the dataset is consistent with the general linear-in-means model. This linear program can be interpreted as a no money pump condition with an added incentive compatibility condition which allows for heterogeneous peer effects across groups. We then proceed to the Luce and uniform linear-in-mean models and provide refinements of these results.
We now begin our analysis of the general linear-in-means model. Recall that in GLM, for each group of agents $N$, each agent's choice $p_i^N$ is written as a convex combination of $v_i$ and $p_j^N$ across all $j \in N \setminus i$. Further, there are no restrictions across groups about the weights assigned to each agent other than each agent $i$ puts some weight on $v_i$. This means that testing GLM amounts to finding if there is some $v_i$ which can induce, for all $N \in \mathcal{N}_i$, $p_i^N$ as a convex combination of $v_i$ and $p_j^N$. This observation tells us two things. First, testing GLM can be done agent by agent. Whether or not agent $i$ has a feasible rationalizing $v_i$ can be tested independently of agent $j$. Second, a key part of testing GLM is finding the set of feasible $v_i$ for agent $i$ given $p^N$ for each $N \in \mathcal{N}_i$.
With this in mind, suppose we observe data $p^N$ and we see that agent $i$'s choices, $p_i^N$, lie on the interior of the convex hull of the other agents' choices, which we write $int \Delta(p_{-i}^N)$. In this case, no matter what $v_i \in \Delta(X)$ we consider, $p_i^N$ can be written as a convex combination of the other agents' choices. Thus we can simply ask that agent $i$ puts a vanishingly small amount of weight on $v_i$ in their convex combination and rationalize this $v_i$. Now suppose we observe data such that $p_i^N$ does not lie in $int \Delta(p^N_{-i})$. In this case, a $v_i$ is feasible if and only if $p_i^N$ can be written as a convex combination of $v_i$ and some convex combination of $p_j^N$ for $j \in N \setminus i$. Thus the set of points on the “opposite side" of $p_i^N$ from $\Delta(p^N_{-i})$ correspond to the set of feasible $v_i$. Formally, this is given by the following equation.
Figure (ref) visualizes the set of feasible $v_i$ for agent $1$ when there are three goods and three agents. These two observations tell us that the testable content of GLM amounts to checking, for each $i$, if there is some feasible $v_i$ which is in each $co^{-1}(\Delta(p^N),p_i^N)$.
Our previous observations tell us that testing GLM amounts to checking if, for each $i$, a collection of convex sets has a point of mutual intersection. Keeping in line with samet1998common and morris1994trade, we can transform this condition into a no money-pump condition. With this in mind, we consider a setting where an outside observer is able to make bets on an agent and their choices in each group $N$.
We restrict attention to bets which we call feasible. Effectively, feasibility is a statement about initial investment and says that, once we aggregate across each group $N$ in $\mathcal{N}_i$, placing a bet on agent $i$ choosing alternative $x$ should be ex-ante costly.
However, for an outside observer to make a bet, it should be ex-post profitable to them. We ask that this individual rationality condition holds for each group $N \in \mathcal{N}_i$.
We impose one last condition on these bets. Notably, we have defined these bets as bets on a specific agent $i$. To this end, we restrict to bets which are incentive compatible. Here, incentive compatibility means that the outside observer cannot gain by placing the bet on agent $j$ instead of agent $i$ at any $N \in \mathcal{N}_i$.
Finally, we say that there is no incentive compatible money-pump if there is no bet that satisfies these three conditions. This effectively asks that there is no way for an outside observer to guarantee that they make an expected profit using an incentive compatible bet.
We are now ready to state our characterization of GLM.
We leave all proofs to the appendix. The equivalence between (1) and (2) follows from our discussion at the start of Section (ref). If $co^{-1}(\Delta(p^N),p_i^N)$ corresponds to the set of feasible $v_i$ for agent $i$ in group $N$, then there needs to be some $v_i$ common to this set across all groups. The equivalence between (1) and (3) is partially a result of (2). As mentioned previously, since testing for GLM amounts to testing for a point of mutual intersection, we can transform this via linear programming duality to get our no money pump condition. However, we note that our no money pump condition does not follow immediately from (2) and an application of the result of samet1998common and billot as, in our proof, we work with polyhedral sets rather than compact sets. This variation allows us to recover the exact form of our no money pump condition. In Appendix (ref) we discuss the relation and application of samet1998common to our Theorem (ref). In this case, we get a type of no trade condition which characterizes GLM.
Before moving on, we note that condition (2) from Theorem (ref) reduces to an easily checkable condition in the one-dimensional case. In the one-dimensional case, the choice frequency of one alternative is a sufficient statistic for the choice frequency of the other alternative. To simplify notation in the one dimension case, we use $p_i^N$ to denote the probability that agent $i$ chooses alternative $x$ in group $N$. Now let $\mathcal{N}_i^- \subseteq \mathcal{N}_i$ denote the set of groups $N$ satisfying $p_i^N \leq p_j^N$ for each $j \in N \setminus i$. Similarly, let $\mathcal{N}_i^+ \subseteq \mathcal{N}_i$ denote the set of groups $N$ satisfying $p_i^N \geq p_j^N$ for each $j \in N \setminus i$.
We now turn our attention to the identification properties of GLM. Specifically, our focus is on the joint recovery of an agents' bliss point as well as the network structure of each group. Going back to manski1993identification, it is known that recovering influence parameters is generally a hard problem. In this section we highlight that, conditional on observing group variation, many of these identification problems are due to the one-dimensional formulation of the original reflection problem. Our identification argument for $(v_i,\pi_i^N)$ proceeds in two steps. First, we show that, generically and with specific group variation, we are able to recover $v_i$ so long as $|X| \geq 3$ (i.e. in any case other than the one dimension case). Further, this identification generically fails in the one dimension case. Second, conditional on recovering $v_i$, we show that $\pi_i^N(j)$ is generically identified from $v_i$ and $p_{-i}^N$ so long as $|X| \geq |N|$. These two results highlight that multiple dimensions are useful in recovering each agent's bliss point, $v_i$, as well as their peer influence parameters, $\pi_i^N$. We begin by discussing the partial identification of $v_i$ that arises without sufficient group variation.
Our first goal is to characterize the sharp identified set of $v_i$. Recall our discussion of $co^{-1}(\Delta(p^N),p_i^N)$ prior to Theorem (ref); it corresponds to the set of feasible $v_i$ given observed choices $p^N$. Theorem (ref) tells us that our dataset is consistent with GLM if and only if, for each $i \in \mathcal{A}$, $\{co^{-1}(\Delta(p^N),p_i^N)\}_{N \in \mathcal{N}_i}$ have a point of mutual intersection. This amounts to testing if there is some ideal point $v_i$ which is feasible in each group $N \in \mathcal{N}_i$. It follows from similar logic that any point in the mutual intersection of $\{co^{-1}(\Delta(p^N),p_i^N)\}_{N \in \mathcal{N}_i}$ is a feasible value for $v_i$.
Proposition (ref) formally states the observation from the previous paragraph. As mentioned earlier, one of our main goals is to give conditions under which $v_i$ and $\pi_i^N$ are jointly point identified. Here, point identified means that the sharp identified set is a singleton. As a first step in our two step procedure, our next result gives a sufficient condition for point identification of $v_i$. Let $\mathcal{N}_i^{ext} \subseteq \mathcal{N}_i$ denote the set of groups $N$ with $p_i^N \not \in \Delta(p_{-i}^N)$.
Figure (ref) offers a visualization of Corollary (ref). While we do not formally show it here, Corollary (ref) naturally extends. Consider $co^{-1}(\Delta(p^N),p_i^N)$ for a group of agents $N$. Note that, in the case of $N \in \mathcal{N}_i^{ext}$, if we drop the non-negativity restriction on $v$ in our definition of $co^{-1}(\Delta(p^N),p_i^N)$, then $co^{-1}(\Delta(p^N),p_i^N)$ defines a polyhedral cone. The extremal rays of this cone take the form $p_i^N-p_j^N$ plus some location translation. In out setting, linear independence of the extremal rays of two of these $n$ dimensional cones will make the intersection of these two sets $(n-1)$ dimensions. Thus, when we observe $n$ groups of $n$ agents, with each group being contained in $\mathcal{N}_i^{ext}$, and the set of extremal rays of $co^{-1}(\Delta(p^N),p_i^N)$ across all $n$ groups of $n$ agents being linearly independent, $\bigcap_{N \in \mathcal{N}_i^{ext}}\{co^{-1}(\Delta(p^N),p_i^N)\}$ is a $(n-n)=0$ dimensional set. By Theorem (ref) this set is non-empty and thus we get point identification. We now return to our example with Ann, Ben, and Can to highlight Corollary (ref).
Before moving on, we point out that point identification of $v_i$ generically fails in the one-dimensional case. Recall the three conditions of Corollary (ref). The last two conditions arise as a result of non-generic situations. Specifically, if we observe data where $p_i^N \neq p_i^M$ for all $N,M \in \mathcal{N}_i$, then it is sufficient to just check the first condition. However, in this case, $v_i$ is necessarily not point identified. In the one-dimensional case, the sharp identified set corresponds to the interval $[\max_{\mathcal{N}_i^+} p_i^N, \min_{N \in \mathcal{N}_i^-} p_i^N]$ and this set is singleton only in a non-generic case. We visualize this failure of point identification in Figure (ref) and show it in Example (ref).
We now turn our attention to the recovery of $\pi_i^N$. This recovery utilizes knowledge of $v_i$ and thus is dependent on our previous arguments in Corollary (ref).
The identification of $\pi_i^N$ follows from basic properties of simplices. Specifically, any point in the convex hull of a simplex can be written as a convex combination, with unique weights, of the extreme points of the simplex. The affine independence condition ensures that $v_i$ and $\{p_j^N\}_{j \in N\setminus i}$ form the extreme points of a simplex. We point out that affine independence of $v_i$ and $\{p_j^N\}_{j \in N\setminus i}$ puts joint restrictions on the number of agents and alternatives. Notably, when $|N| \leq |X|$, the set of points given by $v_i$ and $\{p_j^N\}_{j \in N\setminus i}$ is generically affinely independent. On the other hand, if $|N| > |X|$, the set of points $v_i$ and $\{p_j^N\}_{j \in N\setminus i}$ is never affinely independent. In terms of dimension counting, in group $N$, there are $|N|$ agents each with $|N|-1$ unknowns giving a total of $|N|(|N|-1)$ unknowns in $\pi^N$. Each agent has $|X|-1$ observables per group, and so the total number of observables in group $N$ is given by $|N|(|X|-1)$. This tells us that we have more observables than unknowns if $|X| \geq |N|$. In order to highlight this result, we return to Ann, Ben, and Can.
We conclude our discussion of identification by putting our results in the context of the reflection problem of manski1993identification. Our focus is on a related but slightly different version of the linear-in-means model proposed in manski1993identification. First, our version of the model is primarily concerned with conformism ($\pi_i^N \geq 0$ and $\sum_{j \in N}\pi_i^N(j)=1$). Second, we allow for weighted averages (from unknown networks) while the version of manski1993identification assumes an unweighted average across peers. We make weaker assumptions in one direction and stronger assumptions in other directions. manski1993identification shows that identification fails in his setup. This continues to be true even if our conformism assumptions are applied to the setup of manski1993identification, primarily due to the fact that the reflection problem considers one-dimensional outcome variables. Our results in this section show how utilizing multidimensional outcome variables, along with group variation, allows the analyst to jointly recover agents' bliss points and the underlying network structure. We take this to be prescriptive. While outcome variables can sometimes be sufficiently summarized in a one-dimensional statistic, such as labor time, by gathering more granular data and considering a higher dimensional outcome variable, we can actually recover and quantify heterogeneous peer effects (subject to the conditions of our results).
As we have seen, there are conditions in GLM under which we are able to recover both an agent's ideal point $v_i$ and their social interaction parameters $\pi_i^N$ for groups that we observe. However, these identified parameters have little predictive power for groups we do not observe in GLM. Since GLM allows arbitrary variation in $\pi_i^N$ across groups, the most we can predict for agent $i$'s choice in an unobserved group $N$ is that it lies within the convex hull of $\{v_j\}_{j \in N}$. Our goal in this section is to consider the Luce linear-in-means model which allows us to connect our social interaction parameters across groups. LLM allows analysts to predict choices in unobserved groups subject to the conditions for identification from Section (ref) holding.
Recall that LLM supposes that each agent $i$ has a weighting function, $w_i(j)$, which corresponds to the absolute importance of agent $j$ to agent $i$. In a group $N$, the relative importance of agent $j$ to agent $i$ is given by the renormalization of this weighting function, $\pi_i^N(j)=\frac{w_i(j)}{\sum_{k \in N} w_i(k)}$. The interpretation here is that there is a single underlying network fixed across each group. We first discuss the identification and predictive properties of LLM. We focus on the case when $v_i$ is known. For a specific group $N$, if the conditions of Proposition (ref) hold, we can pin down $\pi_i^N$. This allows us to pin down the relative weights of each agent $j \in N\setminus i$ as $\frac{w_i(j)}{w_i(k)}= \frac{\pi_i^N(j)}{\pi_i^N(k)}$. This is the exact condition for identification that is used in the Luce model of luce1959individual. This means that if $N = \mathcal{A}$, we can predict choice (for agent $i$) on every possible group of agents. However, we may not always observe choice on $N = \mathcal{A}$ or, if we do, the conditions of Proposition (ref) may not be satisfied at $N=\mathcal{A}$. To regain full predictive power in LLM, we need the conditions of Proposition (ref) to hold for a collection of groups $\{N_l\}_{l=1}^L$ such that every two agents $k$ and $j$ can be compared through a string of groups.
Proposition (ref) is an immediate result of our Proposition (ref), Corollary 4 of alos2024characterization, and the fact that $i$ is in both $N_1$ and $N_2$. Further, Proposition (ref) gives conditions under which we can predict the choice of agent $i$ in any group $N$, conditional on knowing $v_j$ for each $j \in N$. We show how to operationalize this predictive power by revisiting Ann, Ben, and Can one last time.
Figure (ref) compares GLM and LLM in terms of their prediction power on the Marschak-Machina triangle in a domain of three alternatives, $X = \{x, y, z\}$ and three agents $\{1,2,3\}$. This is the simplest environment in which we can illustrate the differences between these models. In the figure, stochastic choices for three binary groups ($\{1,2\}, \{1,3\}$ and $\{2,3\}$) are fixed and denoted by different colored dots on the edges of the dotted triangle. The blue dots indicate agent $3$'s choices, $p_3^{\{1,3\}}$ and $p_3^{\{2,3\}}$. We then illustrate the predictions for $p_3^{\{1,2,3\}}$ for GLM and LLM. The blue-shaded triangle identifies all possibilities for $p_3^{\{1,2,3\}}$ for GLM given these binary group choices. This figure illustrates both the predictive and explanatory power of GLM: If $p_3^{\{1,2,3\}}$ lies outside this triangle, the data cannot be explained by GLM.
LLM restricts choices for $p_3^{\{1,2,3\}}$ even further. Indeed, LLM predicts that there is only a single possibility for $p_3^{\{1,2,3\}}$ indicated by a blue dot inside GLM's prediction. If $p_3^{\{1,2,3\}}$ is not equal to this point, then the data cannot be explained by the Luce linear-in-means model. Note that the prediction of LLM belongs to the shaded triangle indicating that LLM is a special case of GLM.
While we focus on LLM to recover predictive power in the linear-in-means model, one could instead consider any known mapping from $\pi_i^N$ to $\pi_i^M$ for two different groups $N$ and $M$ to recover predictive power. We focus on LLM for two reasons. First, LLM is an intuitive criterion that captures the case of there being a fixed underlying network. Second, as we now discuss, we can test LLM using a natural extension of our no incentive compatible money pump condition. Recall our story about an outside observer making bets on an agent and their choices. In the case of GLM, we asked that these bets were strictly feasible, individually rational for each group containing $i$, and incentive compatible for each group containing $i$. When moving from GLM to LLM, we are moving to a model where the social influence parameters $\pi_i^N$ and $\pi_i^M$ are actually connected across groups. Our condition for testing LLM extends the no incentive compatible money pump condition taking this across group connection into account. Specifically, we weaken incentive compatibility so that the outside observer cannot gain by placing a bet on $j$ instead of $i$ across every $N \in N_i \cap N_j$.
Observe that if our dataset satisfies no weakly incentive compatible money pump, then it satisfies no incentive compatible money pump. It then follows from Theorem (ref) that the collection of sets $\{co^{-1}(\Delta(p^N),p_i^N)\}_{N \in \mathcal{N}_i}$ have a point of mutual intersection. In the case of GLM, this is the set of feasible $v_i$. The strengthening of incentive compatibility to weak incentive compatibility is exactly what guarantees us that the social influence parameters $\pi_i^N$ follow a Luce rule across groups.
In the previous sections, we considered the linear-in-means model allowing for heterogeneous social interaction terms. The interpretation of the heterogeneity is that the importance of agent $j$ and agent $k$ may differ to agent $i$. In this section, we consider the hypothesis that agent $i$ belonging to a group impacts their behavior but no one agent in that group is any more important than the other. We model this hypothesis through the uniform linear-in-means model. This is a special case of LLM when $w_i(j)=w_i(k)$ for each $j,k \in \mathcal{A}$. In this section, we maintain the assumption that each observed group has the same size, for all $N,M \in \mathcal{N}$, $|N|=|M|$. This is done so we can focus on the empirical content of variation of group make-up rather than group size. We relax this assumption in Appendix (ref) where we give all the axioms and results from this section allowing for group size variation.
In the context of a linear influence model, such as ours, $p_i^N - p_j^N$ corresponds to a (rescaling) of the influence agent $j$ has on agent $i$'s choice. Consequently, we call $p_i^N - p_j^N$ the peer effect of agent $j$ on agent $i$ in group $N$. All of our axioms for ULM are stated in terms of peer effects. Before stating our axioms, we need one definition.
A cycle captures cycles of influence. Agent $i_1$ influences agent $j_1$ in group $N_1$ and then agent $j_1=i_2$ influences agent $j_2$ in group $N_2$. This is repeated until we return to agent $i_1$. Our first axiom puts restrictions on the sum of peer effects across cycles.
To best understand cyclic constancy, we first introduce a second axiom implied by cyclic constancy.
Observe that constant peer effects is implied by cyclic constancy when we consider cycles of length two. Constant peer effects tells us that the peer effect of agent $j$ on agent $i$ is group invariant. Returning to cyclic constancy, it says that the peer effect of agent $i$ on agent $j$ plus the peer effect of agent $j$ on agent $k$ should be equal to the peer effect of agent $i$ on agent $k$ (and so on for longer cycles). Further, there is no way to break this equality by going to different groups during a cycle. In this manner, cyclic constancy tells us that agent $i$ to agent $j$ peer effects are group invariant and that a long chain of peer effects corresponding to the indirect peer effect of agent $i$ on agent $j$ equals the direct peer effect of agent $i$ on agent $j$. Our next axiom puts restrictions on the peer effect of agent $i$ across different groups.
Symmetric peer effects tell us that, once we take agent $i$'s actual choice into account, the total peer effect of agent $i$ on agents in group $N$ which are not in $M$ is equal to the total peer effect of agent $i$ on agents in group $M$ which are not in group $N$. The combination of cyclic constancy and symmetric peer effects tells us that the total peer effect of agent $i$ is the same in groups $N$ and $M$, once we take into account their actual choice in each group. Our last axiom restricts the total amount of peer effect that agent $i$ has in a group.
Bounded total peer effects tell us that the total amount of peer effect agent $i$ has in group $N$ can be no more than their actual choice $p_i^N$. With this in mind, we are now ready to give our characterization of ULM. Recall that this result assumes that $|N|=|M|$ and that this assumption is relaxed in the appendix.
We first note that, in Theorem (ref), we can replace cyclic constancy with constant peer effects and the equivalence holds. Our focus on cyclic constancy is due to the following discussion. Consider the following the equation.
In Equation (ref), $\hat{v}_i$ is not restricted to have non-negative elements (but still satisfies $\sum_{x \in X}v_i(x)=1$) and we ask $\sum_{x \in X} O^N(x)=0$. Choice induced by Equation (ref) differs from ULM in two important ways. First, an agent's ideal point $\hat{v}_i$ no longer needs to lie within the simplex. Second, every agent $i$ in group $N$ is subject to some group specific shock to tastes given by $O^N$. Under one additional assumption on $\mathcal{N}$, we show in Appendix (ref) that choice according to Equation (ref) is characterized by cyclic constancy. The addition of symmetric peer effects is exactly what rules out group specific shocks, reducing $ O^N$ to be the zero vector across all $N \in \mathcal{N}$. Finally, bounded total peer effects is what induces each agent's ideal point to be within the simplex.
We conclude our analysis with the observation that ULM is trivially point identified. Since each agent's social interaction parameters are pinned down to be $\frac{1}{|N|}$, the only parameter left to identify is each agent's ideal point. This can be easily recovered from the following equation.
Equation (ref) also offers an alternative characterization of ULM. A dataset is consistent with ULM if and only if the value of the right hand side of Equation (ref) is group invariant and lies in the simplex.
In this section, we discuss how our results from Section (ref) extend under alternative assumptions on the data. Our first and main extension is to consider relaxing the assumption that our data satisfies $\sum_{x \in X} p_i^N(x)=1$ (as well as $p_i^N \geq 0$). In this case, all of our results extend (mostly) unchanged. We then consider three additional forms of variation and show how they can provide even stronger results than what we have discussed in Section (ref).
Thus far, we have taken the perspective of stochastic choice. Agents in our model choose a probability distribution. However, nothing precludes us from investigating more sophisticated choice environments. The classic one-dimensional framework has as its choice space $[0,1]$. The unit interval can be interpreted in many ways---any cardinal unidimensional choice can be modeled this way. We have interpreted the unit interval as a probability distribution over two objects, leading to the simplex as the natural generalization. Similarly, if we were to imagine an individual deciding how to spend their day, $[0,1]$ might measure the proportion of hours devoted to sleep vs. waking, whereas a model with three alternatives might further refine waking hours into work and leisure.
However, suppose that $[0,1]$ represents the score on a test. We might also be interested in scores on multiple exams: a member of $[0,1]^2$ might reflect the scores on an English language exam and on a Turkish language exam. In fact, Equation (ref) is meaningful whenever there is a convex subset of a vector space $Y$, and each $p_i^N$ lies in $Y$. The particular case of $Y=\Delta(X)$ is only one example. All three of our particular models (general, Luce, and uniform) also have ready interpretation in terms of more general vector spaces.
Let us illustrate this point via example, whereby our choice space is $Y=[0,1]^T$. We can think of $T$ as indexing a collection of tests. Any $p_i^N\in Y$ is a vector of test scores. Our equation $p_i^N = \pi_i^N(i)v_i+\sum_{J\in j\in N\setminus i}\pi_i^N(j)p_j^N$ asserts that an individual's test scores are an average of some “baseline scores” ($v_i$) and the scores of other individuals present.
In this more general setup, the definition of inverse cone remains unchanged and still captures the set of feasible $v_i$ given the data.
Theorem (ref) parts 1 and 2 remain equivalent. More generally, the duality conditions (such as incentive compatible money pump) will change from domain to domain, but can be characterized for polyhedral domains. The downside is that, outside of the simplex domain, these dual conditions become more difficult to interpret. In regards to identification, since $co^{-1}(Y,p_i^N)$ still captures the set of feasible $v_i$, all of the results from Section (ref) go through unchanged.
One remaining caveat is that, in more general vector spaces, we may allow for the possibility that individual weights need not sum to one. As mentioned in Section (ref), this causes problems for many of our results. For example, if $Y= \mathbb{R}^M_+$, then it is mathematically meaningful to allow any weights $\pi_i^N$ for which $\pi_i^N(j)\geq 0$ for all $j$ and $\pi_i^N(i)>0$. Which sets of weights might be mathematically meaningful depends, in general, on the domain under consideration. The economic interpretation will likely dictate which set of weights to choose.
Throughout this paper, we have focused on group variation being the main source of variation in the data. We now consider three alternative form of variation; network variation, characteristic variation, and product-attribute variation. We discuss each of these forms of variation in a series of three remarks.
In this paper we study the testing and identification properties of the linear-in-means model when an analyst observes group variation. As part of our analysis, we show that the reflection problem of manski1993identification is a generic problem when each agent's decision space is one dimension but stops being generic when we move to higher dimensions. This highlights the identifying power of more granular data in the linear-in-means model and we emphasize this as a key takeaway of our analysis. We now conclude with a discussion of this paper's place in the related literature.
Our paper is related to several strands of literature. We begin by discussing the strand which studies the linear-in-means model of social interactions. Predating the modern literature on the linear-in-means model, keynes1937general considers a model of financial markets via a story of beauty contests. In this setting, an agent wishes to take the action that coincides with the average action of the rest of the population. In our setup, this corresponds to $\pi_i^N(i)=0$ and $\pi_i^N(j)=\frac{1}{|N|-1}$. More recently, ushchev2020social studies the microfoundations and comparative statics of the linear-in-means model allowing for arbitrary network structure. As mentioned earlier, ushchev2020social along with blume2015linear, boucher2016some, and kline2020econometric show that the linear-in-means choice rule can be achieved as the best response to a quadratic loss utility function in a complete information game where each agent knows each $v_i$ and $\pi_i^N$. golub2020expectations consider an extension of this setup where agents have incomplete information and relate the linear-in-means model to higher-order expectations as well as conventions in networks. We build on this literature by studying the empirical content of the linear-in-means model.
More closely related to our paper is the strand of literature which focuses studying the identification properties of the linear-in-means model as well as other related models of peer effects. A seminal contribution in this literature is manski1993identification and his discussion of the reflection problem. Important to our analysis is the following takeaway from the reflection problem of manski1993identification. It is in general difficult to identify the social influence of a group on an agent due to the endogenous nature of outcomes. In our setting, this corresponds to identifying both the underlying network structure and the corresponding weights on directed edges in this network. Much of the literature following manski1993identification aims to identify social interaction parameters when the underlying network structure is (partially) known. This literature is extensive, so we list the following, all of which provide various conditions in order to recover identification of social interaction parameters in linear-in-means style models; lee2007identification, graham2008identifying, bramoulle2009identification, de2010identification, blume2011identification, boucher2014peers, blume2015linear, de2017econometrics. More recently, boucher2024toward extends the linear-in-means to a CES in means model of social influence and provides identifications in their setting.
Perhaps most closely related to our analysis within this literature is the work of lewbel2023social and de2024identifying. Both of these papers study the joint identification of social interaction parameters as well as the identification of the underlying network structure.\footnote{breza2020using and Griffith2023 also considers joint identification of peer effects and network structure. However, breza2020using asks for a stronger form of data where the analyst knows the distribution over how many connection each agent has to other agents with specific characteristics. Similarly, Griffith2023 makes assumptions on the underlying network density in order to recover the unobserved network structure. battaglini2022endogenous also estimates network connections in a related model of endogenous network formation.} While lewbel2023social utilizes group variation and de2024identifying utilizes long run panel data, a common tool for identification in each of these studies is the use of characteristic variation in order to recover identification. The models of these two paper also differs in the fact that the characteristics of each agent show up in the choice procedure of other agents. Adapting to our context, the core components of these models can be summarized by
where $z_i$ is the characteristic vector of agent $i$, $\alpha$ is the average impact of characteristics on each agent's bliss point, $\beta$ corresponds to the average impact of other agents' choice on each agent's choice, and $\gamma$ corresponds to the average impact of each other agents' characteristics on each agent's choice. de2024identifying interprets $\beta$ as the endogenous peer effects and $\gamma$ as the exogenous peer effects from the reflection problem of manski1993identification. Our analysis differs from this literature in that we require no knowledge of characteristics (other than identity) to jointly recover each agent's bliss point as well as the underlying network structure. The key to our identification result is that our outcome vector $p_i^N$ is multidimensional. In fact, generically, our results show that identification fails in our setup when the outcome variable is one dimension.
Our paper also contributes to the choice theoretic literature studying social interactions. To our knowledge, this literature begins with cuhadaroglu2017choosing who studies a two period model with social influence taking effect in the second period. borah2018choice also considers a two period model of social influence. In the first period, social influence is used to form an agent's consideration set and the second period is used for choice. kashaev2023peer consider a model of social influence where the choices of an agent's peers directly form their consideration set. One of the main goals of this work is the actual identification of underlying social parameters, which they are able to achieve through dynamic choice variation. bhushan2023beliefs considers a model of choice where agents influence each other through their beliefs. An agent's beliefs are determined through a process similar to ULM. Choices then correspond to subjective expected utility given these beliefs. Most closely related to our work in this literature is the work of chambers2023behavioral who consider a version of the linear-in-means model. They focus on a setting with menu variation (i.e. variation of $X$) with a fixed group and network structure. This differs from our analysis in that our focus is on group variation and we can accommodate network variation.
More generally, our paper is related to the literature studying the empirical content of strategic settings. Part of this literature focuses on testing the empirical content of specific solution concepts. sprumont2000testable studies the testable content of Nash equilibrium. The work of haile2008empirical finds that quantal response equilibrium has no empirical content. Similarly, bossert2013every and rehbeck2014every find that backwards induction has no empirical content when we only observe the induced choice function. Another portion of this literature focuses on characterizing the empirical content of specific (types of) games. To begin, lee2012testable characterizes the testable content of zero-sum games. carvajal2013revealed provides a revealed preference style characterization of the Cournot model of competition. Finally, lazzati2023ordinal characterizes the empirical content of Nash equilibrium in monotone games. As mentioned earlier, the linear-in-means model can be thought of as arising from Nash equilibrium play where each agent's utility is given by Equation (ref). As a result of this equivalence, we characterize Nash equilibrium play in the corresponding game when we observe player variation.
Finally, our paper is also related to the literature in stochastic choice studying agents who have preferences for non-deterministic bundles. The idea of deliberately stochastic preferences goes back to at least machina1985stochastic. Recently, cerreia2019deliberately axiomatizes data that arises from agents choosing with deterministic preferences over lotteries. Similarly, fudenberg2015stochastic characterizes stochastic choice data that arises from agents who have cardinal preferences over each alternative but face a perturbation to their utility function within the simplex. allen2019identification studies the identification properties of a similar class of perturbed utility functions. There has been little work on these types of perturbed utility functions with social influence components. hashidate2023social is an exception and considers a perturbed utility function where the perturbation corresponds to a norm. Recall that the choices in the linear-in-means model can be induced by a perturbed utility model where an agent's base utility is given by a quadratic loss function with reference to their ideal point. The perturbation then corresponds to a sum of quadratic loss functions with each one referencing another agent's choice.