EconBase
← Back to paper

Emergence of Cooperation in the thermodynamic limit

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

27,479 characters · 8 sections · 39 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

The emergence of Cooperation in the thermodynamic limit

abstractPredicting how cooperative behavior arises in the thermodynamic limit is one of the outstanding problems in evolutionary game theory. For two player games, cooperation is seldom the Nash equilibrium. However, in the thermodynamic limit cooperation is the natural recourse regardless of whether we are dealing with humans or animals. In this work, we use the analogy with the Ising model to predict how cooperation arises in the thermodynamic limit.

{\bf Keywords:} Nash equilibrium; Prisoner's dilemma; game of Chicken; Ising model\\ {\bf PACS numbers:} 02.50.Le; 87.23.Kg; 01.80.+b; 05.50.+q

Introduction

The solution to any game theoretic problem involves finding an equilibrium strategy known as the Nash equilibrium, whence deviating from this strategy brings in more loss to the player. An introduction to Nash equilibrium for two player two strategy games can be found in Ref. 4. Finding the Nash equilibrium analytically for a two player game isn't difficult, however, many a situation arises wherein one needs to go beyond two players. To analyze such situations in a game theoretic setting one needs to go beyond two players to a game with infinite number of players, i.e., the thermodynamic limit. A particularly interesting problem arises in the context of evolution, where we see that cooperation arises even when defection is the preferred choice of individuals, see Ref. 5. Game theory is inextricably linked with artificial intelligence. The problem of many or infinite number of agents collaborating is an outstanding problem not just in game theory but in artificial intelligence tooai. Cooperation in short term might not seem beneficial however, in the long run the population which cooperates survives. This scenario is generally tackled numerically and dynamically using replicator equations as has been shown in Refs. 12,nowak. Further, in Ref. 13 the authors use the replicator equations to numerically calculate how the fraction of cooperators and defectors evolve in time when playing the Prisoner's dilemma game. The evolution of cooperators and defectors is not just restricted to Prisoner's dilemma, this has also been observed in context of social dilemmas like the Vaccination game, in Refs. 21,22. The Vaccination game, which is a tool to monitor public health, is essentially a variation of the Prisoner's dilemma wherein instead of coooperators and defectors one has immunizers and anti-immunizers. In this paper, on the other hand, we use statistical mechanics tools to analytically predict whether and how cooperation emerges in the thermodynamic limit and this is one of the main attractions of our work. An approach has also been suggested by Szabo and Borsos by constructing potentials corresponding to particular games and then comparing with models borrowed from statistical mechanics. However, our approach is analytical and mathematically much more simpler than the method of Szabo and Borsos, see Ref. 19. This is an interesting topic for evolutionary game theory and useful application of artificial intelligence, as any ecosystem tends to evolve towards an equilibrium with a large number of living beings. We show that 1D statistical models can correctly be applied to such social dilemmas.

Using the analogy with the 1D Ising model, we try to understand the equilibrium strategy in a population and predict how cooperative behavior emerges in the thermodynamic limit. We model a situation similar to 1D Ising model, where the sites are replaced by players and spin up or spin down correspond to the strategies $s_1$ or $s_2$ adopted by the players. Magnetization in Ising model is defined as the difference in the number of spin up and spin down particles. Similarly, Magnetization in game theory can be defined as the difference in the number of players choosing strategy $s_1$ or $s_2$. We first relate the Ising model to the payoff's in game theory and then apply this method first to Prisoner's dilemma and then to game of Chicken (variant of Hawk-Dove game). There have been earlier attempts to use the 1D Ising model to find the equilibrium strategy in the thermodynamic limit, see Ref. 5. We find that there are some unphysical implications of the results of Ref. 5. This paper is organized as follows-we next discuss how to connect the 1D Ising model to the payoffs of game theory by extending the analogy of Ref. 1 to the thermodynamic limit, then we calculate the game magnetization which gives the Nash equilibrium strategy for Prisoner's dilemma in the thermodynamic limit. We observe how cooperators arise in Prisoner's dilemma in the thermodynamic limit even when defection is the Nash equilibrium. Further, we deal with the problems associated with the model of Ref. 5 in brief. Later we do a similar analysis for the game of Chicken which has no unique pure strategy Nash equilibrium in the two player case. We find how in the thermodynamic limit majority of cooperators can emerge. We end with the conclusions.

1D Ising model and game theory

There have been several works which have found it useful to explain the behavior in social dilemmas by looking at 1D models like the 1D Ising model, see Ref. 14 wherein the equivalence between the voter model and 1D Ising model has been brought out. Further, the 1D Sznajd model has been used as a template to explain spreading of opinion in a society, see Ref. 15. These are just two examples which have used 1D Ising models to model social dilemmas, many others also do exist, see Ref. 14 for more examples. The most well known among such models is the 1D Ising model6 which consists of spins that can be in either of the two states $+1$ ($\uparrow$) or $-1$ ($\downarrow$). The spins are arranged in a line, and can only interact with their nearest neighbors. The Hamiltonian of such a system can be written as-

equation[equation omitted — 87 chars of source]

where $J$ is the coupling between the spins, $h$ is the external magnetic field and $\sigma$'s denote the spin. The partition function corresponding to the Hamiltonian ((ref)) is

equation[equation omitted — 146 chars of source]

$\beta=\frac{1}{k_{B} T}$, with $k_B$ being Boltzmann's constant. $\sigma_i$ denotes either the spin up $(+1)$ or spin down $(-1)$. In order to carry out the spin sum, we define a matrix $T$ with elements as follows,

eqnarray[eqnarray omitted — 93 chars of source]

Using the transfer matrix $T$ and carrying out the spin sum via the completeness relation, the partition function from Eq. ((ref)) in the large $N$ limit can be written as-

equation[equation omitted — 90 chars of source]

Since the Free energy $F=-k_{B}T \ln Z$, the magnetization is

equation[equation omitted — 105 chars of source]

In Fig. (ref), we plot the magnetization versus external magnetic field $h$ for different values of inverse temperature($\beta$).

figure[figure omitted — 202 chars of source]

In Ref. [1], it has been shown that a one-to-one correspondence can be made between 1D Ising model Hamiltonian and the payoff matrix for a particular game. We first look at a general payoff matrix for two player game and try to understand the methodology,

equation[equation omitted — 133 chars of source]

where $U(s_i,s_j)$ is the payoff function with $a, b, c, d$ as the payoffs for row player and $a', b', c', d'$ are the payoffs for column player, $s_1$ and $s_2$ denote the strategies adopted by the two players. The Nash equilibrium as defined before is the strategy, deviating from which brings loss to the players. For the payoff's defined in Eq. ((ref)), the Nash equilibrium, in case one takes the condition of a symmetric game, i.e., $b'=c, c=b'$ and $a=a', b=b'$ with $c<d<a<b$, is the strategy $(s_{2},s_{2})$. To make a one-to-one correspondence of the game payoff's with Ising model we need a set of transformations to the payoff's which will achieve that. The transformations of payoffs of the players are as follows:

equation[equation omitted — 187 chars of source]

We choose the transformations as $\lambda=-\frac{a+c}{2},\lambda'=-\frac{a'+b'}{2}$ and $\mu=-\frac{b+d}{2},\mu'=-\frac{c'+d'}{2}$. Under such a transformation the Nash equilibrium doesn't change (see Supplementary material section 1.2 for more details where we prove this result using fixed point analysis, see also Ref. 11). Since Ising model Hamiltonian considered above (ref), assumes that the coupling $J$ is symmetric, thus to model the game (ref) using Ising model we must consider only symmetric games, i.e., $a=a',\ b=c',\ c=b'$ and $d=d'$. After imposing the above conditions the transformed payoff matrix from (ref) becomes-

equation[equation omitted — 227 chars of source]

To calculate the Nash equilibrium of a generalised two player game in the thermodynamic limit, we have to relate the transformed payoff matrix of the classical game as in Eq. ((ref)) to the Ising model Hamiltonian with two spins. When $N=2$, the Hamiltonian Eq. (ref) can be written as-

equation[equation omitted — 93 chars of source]

So the individual energies of the spins 1 and 2 can be written as:

equation[equation omitted — 101 chars of source]

It is to be noted that equilibrium in Ising model corresponds to minimizing the energies of spins. Now for symmetric coupling as in Eq.'s ((ref),(ref)) minimizing Hamiltonian $H$ with respect to spins $\sigma_1, \sigma_2$ is same as maximizing $-H$ with respect to $\sigma_1, \sigma_2$. In game theory, on the other hand players search for the Nash equilibrium with maximum payoffs. This implies maximizing the payoff function $U$ in Eqs ((ref)-(ref)) with respect to strategies $s_i,s_j$ which for the two player Ising model is equivalent to maximizing $-E_i$, see Eq. (ref) with respect to spins $\sigma_i,\sigma_j$. Thus, the Ising game matrix can be written as-

equation[equation omitted — 185 chars of source]

Comparing the matrix elements of the transformed payoff matrix- Eq. ((ref)) to the Ising game matrix Eq. (ref), we get the relation between parameters of Ising model ($J$ and $h$) and the payoffs of two player game as-

equation[equation omitted — 58 chars of source]

Substituting $J$ and $h$ in terms of payoff's in the equation for magnetization(ref), gives us the game magnetization ($m_{g}$), defined as the fraction of player's choosing strategy $s_1$ over $s_2$ in the thermodynamic limit-

equation[equation omitted — 219 chars of source]

This completes the connection of the payoffs from a two player game to Ising model relating spins in the thermodynamic limit. $\beta$ in Ising model is the inverse temperature. Decreasing $\beta$ or increasing temperature leads to randomness in spin orientation. Thus, as $\beta\rightarrow 0$ then game magnetization vanishes, see (ref), due to increase in randomness in the strategic choices of the players. In the following sections we will apply this to some famous two player games so as to analyze them in the thermodynamic limit. Note that this approach can't be compared to Nowak's approach as in contrast to replicator equations, the above suggested approach is not dynamical but completely analytical. However, in the next section we qualitatively compare our results with the conclusions presented in Refs. 12,13.

Prisoner's dilemma

In this game, the police are questioning two suspects in separate cells. Each has two choices: to cooperate with each other and not confess the crime (C), or defect to the police and betray each other(D). We construct the Prisoner's dilemma payoff matrix by taking the matrix elements from Eq. (ref) as $a=r$, $d=p$, $b=s$ and $c=t$, with $t>r>p>s$ where $r$ is the reward, $t$ is the temptation, $s$ is the sucker's payoff and $p$ is the punishment. Thus, the payoff matrix is-

equation[equation omitted — 123 chars of source]

The values in the payoff matrix can be explained as follows- reward $r$ means 1 year in jail while punishment $p$ means 10 years in jail, sucker's payoff $s$ represents a life sentence while temptation $t$ implies no jail time. Independent of the other suspects choice, one can improve his own position by defecting. Therefore, the Nash equilibrium in this case is to defect. Following on from the calculations for the general two player game as in Eq. (ref) and Eq. (ref) as applied to Prisoners dilemma game matrix Eq. (ref), we get- $J=\frac{r-t+p-s}{4}$ and $h=\frac{r+s-t-p}{4}$. From Ising model, the game magnetization($m_g$) in the thermodynamic limit Eq. ((ref)) is-

equation[equation omitted — 134 chars of source]
figure[figure omitted — 339 chars of source]

Plotting game magnetization as in Eq. (ref) as function of punishment, with the condition: $s<p<r$, we see in the thermodynamic limit the Nash equilibrium is always the defect strategy. A phase transition would occur only if $p<r+s-t$ which is not possible as the Prisoner's dilemma has the condition: $p>s$ and $t>r$. When $\beta$ decreases, game magnetization decreases, which implies that number of cooperators increases. As $\beta\rightarrow 0$, $m_{g}\rightarrow$ 0, implying equal number of cooperators and defectors. At finite and large $\beta$ as seen from Fig. (ref), in the thermodynamic limit for $p>1.5$ almost all are defectors. However, in the range $0<p<1.5$ there is a decrease in the number of defectors so much so that around $p=.5$, $25~\%$ of the population tend to cooperate for $\beta=1.0$. From the numerical calculations presented in Ref. 13, we also see that the number of defectors always remains a majority and the number of cooperators in any particular generation increases if the reward increases which is similar to the results presented above. Further, in Ref. 12 it is shown that in the iterative Prisoner's dilemma with finite number of players cooperation can become Nash equilibrium if tit for tat scheme is allowed for some fraction of population but not the entire population. In Ref. 12, the thermodynamic limit is not diretly dealt with however, they infer via natural selection that for the iterative Prisoner's Dilemma the Nash equilibrium would be everyone defecting. Contrary to this we show that even if defection is the Nash equilibrium in the thermodynamic limit there still exist a finite minority of cooperators which too increase as the reward increases. Thus, our results show that even in the thermodynamic limit of Prisoner's dilemma cooperation emerges. In the next section we approach this problem via the method proposed in Ref. 5 and unravel some deficiencies in the method of Ref. 5.

Problems with the approach of Ref. 5

The connection between Ising model and game theory as shown above is not the only approach available. In Ref. 5 too, it has been shown that in the thermodynamic limit games can be modeled using 1D Ising model. However, when one analyses the Prisoner's dilemma game using the approach of Ref. 5, the results are not compatible with the basic tenets of the game for some cases, as shown below.

When reward $r$ approaches temptation $b$:

We start with payoff matrix(ref) used in Ref. 5, to describe Prisoner's dilemma-

equation[equation omitted — 125 chars of source]

where $r=b-c$, with $b>r>0$ and $b>c>0$. Eq. ((ref)) is the payoff matrix used in Ref. 5. This is similar to Eq. ((ref)) with payoffs for reward as $r$, temptation as $b$, sucker's payoff as $-c$ and punishment as $0$ with the condition $r=b-c$. The game magnetization as derived in Ref. 5 is given as-

equation[equation omitted — 75 chars of source]

This game magnetization is independent of temptation $b$ unlike that derived in Eq. ((ref)) using our approach. Although in Ref. [5] it has been shown that for all values of reward $r$, the dominant choice is to defect but this is not true in the limiting case when $r$ approaches $b$. We analyze the same situation using the payoff matrix of the Prisoner's dilemma-

equation[equation omitted — 129 chars of source]

In Eq. (ref), we see an inconsistency, when reward $r$ equals the temptation $b$, there is no unique Nash equilibrium, i.e., both strategies cooperation and defection are equiprobable. The players can equally choose between cooperation and defection and hence game Magnetization $m_{g}$ should be $0$. However, from Ref. [5] the game magnetization (ref) is negative (see Fig. (ref) inset) which means that defect is the Nash equilibrium which is not correct.

The reward $r$ approaches $0$:

Another situation where Ref. [5]'s results are negated is when $r=0$, $m_g$ tends to $0$ as in Eq. ((ref)) implying equal number of cooperators and defectors. However, when we look at the payoff matrix Eq. ((ref)) for $r=0$, we have-

equation[equation omitted — 113 chars of source]

one can see defect (D,D) is still the Nash equilibrium. Using our approach, see the calculations as done in Eqs. (ref)-(ref) and using payoff matrix (ref) we get the game magnetization as $m_{g}=\tanh(\beta \frac{r-b}{2})$ where we have substituted $t=b,s=-c,p=0$ with the condition $r=b-c$ in Eq. ((ref)). In Fig. (ref) one sees $m_{g} \rightarrow -1$ as $r\rightarrow 0$, which is the correct result using our approach. In inset of Fig. (ref), the game magnetization using the approach of Ref. [5] however tends to $0$ which is obviously incorrect.

figure[figure omitted — 706 chars of source]

From Fig. (ref) as reward $r$ approaches temptation $b$, game magnetization $m_{g}\rightarrow 0$. Further, when reward $r$ approaches $0$, game magnetization $m_{g}\rightarrow -1$. Our approach corrects the problems in Ref. 5 in the limiting cases when $r\rightarrow 0$ and $r \rightarrow b$. This is also elaborately dealt with in Ref. 18 along with the case of Public goods game with and without punishment. In the supplementary material accompanying this article we deal elaborately with the reasons behind the problems in approach of Ref. 5. Next, we extend this approach to the game of Chicken.

Game of Chicken

The name “Chicken" has its origins in a game in which two teenagers drive their vehicles towards each other at high speeds4. Each has two strategies: one is to swerve and the other is going straight. If one teenager swerves and the other drives straight, then the one who swerved will be called a "Chicken" or coward. “Hawk$-$Dove" game, on the other hand refers to a situation in which players compete for a shared resource and can choose either mediate (Dove strategy) or fight for the resource (Hawk strategy). The parameterized payoff matrix from Eq. ((ref)) by taking $a=-s,\ b=r,\ c=-r$ and $d=0$ for the game of Chicken is given by-

equation[equation omitted — 151 chars of source]

where $``r"$ denotes the reputation and $``s"$ denotes the cost of injury and $s>r>0$. If one teen swerves before the other, then the one who drives straight gains in reputation while the other loses reputation. However, if both drive straight, there is a crash, and both are injured. There are two pure strategy Nash equilibriums (straight, swerve) and (swerve, straight). Each gives a payoff of $r$ to one player and -$r$ to the other. There is another mixed strategy Nash equilibrium given by $(\sigma,\sigma)$, where [$\sigma$ = $p$.straight+$(1-p)$.swerve] where $p=\frac{r}{s}$ ($p$ is the probability to choose straight). In “Hawk-dove" the reputation from game of “Chicken" is replaced by the value of resource and the cost of injury doesn't change. Similar to game of “Chicken", the Hawk-Dove game has two pure strategy Nash equilibrium: (Hawk, Dove) and (Dove, Hawk) and a mixed strategy Nash equilibrium ($\sigma,\sigma$): [$\sigma$ = $p$.Hawk +$(1-p)$.Dove]. Thus, from a game-theoretic point of view, “Chicken" and “Hawk$-$Dove" are identical.

figure[figure omitted — 550 chars of source]

We analyze the game of Chicken in the thermodynamic limit. Following on from the calculations for the general two player game as in Eq. (4) and Eq. (10) as applied to the game of Chicken payoff's in Eq. ((ref)) we get $J=-\frac{s}{4}$ and $h=\frac{2r-s}{4}$. Thus, in the thermodynamic limit of the game of Chicken the game magnetization is-

equation[equation omitted — 180 chars of source]

From Eq. ((ref)), the condition for change of sign in “$m_{g}$" is given by-

equation[equation omitted — 117 chars of source]

Plotting game magnetization $m_g$ as in Eq. ((ref)), we see from Fig. (ref) that as the reputation $r$ increases more than $s/2$, more players choose straight as choosing swerve would bring in more loss to one's reputation. Similarly, when $r<s/2$ then players would rather choose to swerve and not get injured. Further, it should be noted from Eq. ((ref)) that as the cost of injury increases (see inset of Fig. (ref)), the game magnetization becomes more positive implying that more players choose to swerve or cooperate.

Discussion and Conclusions

Since game of Chicken and Hawk-Dove game are equivalent in game theory, it can be inferred from the above results that for Hawk-Dove game as the value of resource increases keeping the cost of injury constant, then more fraction of players choose the Hawk strategy (defect) or fight for the resource. Further, when the cost of injury increases then the players are reluctant to fight for the resource as getting injured is more expensive. Thus, larger fraction of players end up choosing Dove strategy(cooperate), i.e., sharing the resource when cost of injury is high. Contrary to the notion as in two player games that the players would always opt for the Nash equilibrium strategy, in the thermodynamic limit this is not true. In the thermodynamic limit our results show that a larger fraction of the players would choose the Nash equilibrium strategy but not every player. For example, when the temptation decreases in Prisoner's dilemma the fraction of cooperators increases even when the Nash equilibrium is to defect. A natural extension in the thermodynamic limit would be that every player would choose to defect as Nash equilibrium for the two player case is (D, D) however, there is a finite fraction of players who choose cooperation which increases as the temptation (t) decreases. Further, we see in game of Chicken that even if the reputation becomes high still there is a small fraction of players who choose to swerve(cooperate) and lose. This shows that in the thermodynamic limit cooperation does emerge even when defection would be the preferred choice of the individual players. Again, we see in Prisoner's dilemma that in thermodynamic limit slightly reducing punishment below $r/3$ where $r$ is the reward increases the fraction of cooperators by a large amount even when the Nash equilibrium is to defect. Even in game of “Chicken", when cost of injury is low the best choice for the players is to choose straight or defect. However, we find that still there exist a large fraction of players who choose to swerve or cooperate. In related works, see Refs. 10,sar-benj-physa, we extend our model of predicting cooperative behavior in the thermodynamic limit to the quantum regime.