Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
32,897 characters · 6 sections · 26 citation commands
Two-step Estimation of Network Formation Models with Unobserved Heterogeneities and Strategic Interactions
\abovedisplayskip=15pt \belowdisplayskip=15pt
\centerline{Abstract} In this paper, I characterize the network formation process as a static game of incomplete information, where the latent payoff of forming a link between two individuals depends on the structure of the network, as well as private information on agents' attributes. I allow agents' private unobserved attributes to be correlated with observed attributes through individual fixed effects. Using data from a single large network, I propose a two-step estimator for the model primitives. In the first step, I estimate agents' equilibrium beliefs of other people's choice probabilities. In the second step, I plug in the first-step estimator to the conditional choice probability expression and estimate the model parameters and the unobserved individual fixed effects together using Joint MLE. Assuming that the observed attributes are discrete, I showed that the first step estimator is uniformly consistent with rate $N^{-1/4}$, where $N$ is the total number of linking proposals. I also show that the second-step estimator converges asymptotically to a normal distribution at the same rate.
The social network is an important feature to take into account when studying many economic behaviors, from peer effects in education and crime to the dynamics of product adoption and financial contagions. However, most network studies of these behaviors are challenged by the endogeneity of the network. This highlights the importance of developing econometric models of network formation. Moreover, the network formation process is itself an interesting subject to study, since it sheds light on real-world behaviors such as how people engage with each other on social media platforms. \\ Two features are crucial in a network formation model. The first feature is strategic interactions. The incentives of forming a link in a network are not only affected by the two agents' characteristics but also the linking decisions of other agents, such as the "popularity effect"-- an agent is more likely to form a link with another agent who has many friends. The second feature involves unobserved agent-level heterogeneities, which are typically private information that is known only to the agent themselves, such as an individual's personality traits on a dating app. The agent-level unobserved heterogeneities are correlated with observed characteristics but are unobserved to other agents or researchers. Depicting these two features is essential for effectively modeling the network formation process and accurately inferring agents' preferences. Motivated by this, I study a directed network formation model with individual-specific unobserved heterogeneities and strategic interactions. In the incremental utility of a link from person $i$ to $j$, I include the linking choices of the person $j$ to capture the popularity effect, and include individual fixed effects to capture agent-level unobserved heterogeneities, while remaining agnostic about the conditional distribution of the agent-level unobservables, not requiring it to be known to the researchers.\\ There's growing literature on the estimation of network formation models. Among them, this paper is most related to leung2015two and ridder2022estimation. Both of them study the estimation of network formation games with incomplete information and strategic interactions and assume that the private information is independent of observed characteristics. leung2015two lets the payoff depend on network structure in a separable way, through the sum of incremental utilities from each link. Then the optimal link choices are myopic, in the sense that an agent chooses to form a link with another member if the expected utility of forming that link is greater than 0. To be specific, let $G_{ij}$ denote the linking proposal from individual $i$ to $j$, and let $X_i, X_j$ denote observed characteristics of the two individuals; let $\epsilon_{ij}$ denote unobserved link-specific characteristics that are independent with $X$. leung2015two's model yields the following optimal linking decision:
where $w$ is a known function capturing the homophily effect, and $\sigma$ denotes the equilibrium. ridder2022estimation considers a more general case in which the utility function depends on the choice of potential partners in a non-separable way, for example, allowing the utility to depend on links-in-common. Using the Legendre transform, they show that even under this general case, the optimal linking choice is still equivalent to a sequence of myopic link choices. For estimation, both of the two papers assume that the data observed comes from a symmetric equilibrium, whereby agents with the same observable characteristics have the same equilibrium linking probabilities, i.e. $P(G_{ij}|X_{ij}=x, X)=P(G_{kl}|X_{kl}=x, X)$. Then the conditional linking probabilities can be estimated in the first step, by taking the empirical frequency with which agents with the same observable characteristics link to each other. In terms of strategic interactions, this paper adopts the same framework as leung2015two, including only the popularity effect and keeping the dependence on the network structure to be separable, which is simpler than ridder2022estimation's framework. Different from the two papers, this paper studies the case when private information is correlated with observables by including individual fixed effects in the utility. For estimation, this paper also adopts a two-step procedure and estimates the realized equilibrium beliefs in the first step. This allows us to circumvent the difficulty to specify the equilibrium selection mechanism when there might be multiple equilibria.\\ This paper is also closely related to graham2017econometric, which studies a network formation model with dyadic link formation. In their model, the linking decision between individual $i$ and $j$ only depends on the characteristics of $i$ and $j$ and there are no strategic interactions. Let $A_i, A_j$ denote individual fixed effects unobserved to researchers. The linking decision in graham2017econometric is
Same as graham2017econometric, this paper also incorporates unobserved individual fixed effects. The difference is that my model contains strategic interactions, so the information structure matters. I assume that individual fixed effects $A_i$ are private information that is i.i.d. across individuals. The agents know the distribution of the individual fixed effects so that they can form beliefs of the expected "type"\footnote{I don't assume $A_i$ to have discrete distribution, though.} of other people. From the modeling point of view, this paper studies a model which is a combination of ((ref)) and ((ref)). Note that a special case of ((ref)) is when $\epsilon_{ij}$ can be written as the sum of an individual "random effect" $A_i$ and an idiosyncratic error $\nu_{ij}$. This is different from this paper's setting since $A_i$ is assumed to be independent with $X$ in leung2015two.\\ Another strand of literature on estimating strategic network formation models assumes complete information, such as miyauchi2016structural and sheng2020structural. These models are the hardest to deal with because they generally admit multiple equilibria and thus achieve set but not point identification of the model parameters. This paper shies away from these cases by assuming incomplete information.\\ The rest of the paper is organized as follows. In section 2, I develop the model and derive the optimal link choices. In section 3, I propose a two-step estimation procedure and show the consistency of the first-step estimator. In section 4, I showed the asymptotic distribution of the estimators. The last section concludes.
I consider the directed network formation model in this paper. The formation process is a static game of incomplete information. An agent's payoff of forming a link depends on idiosyncratic private information. Given the belief of other people's linking decisions, agents form their own links simultaneously. Formally, the network formation game is set up as follows:\\ There are $n$ agents indexed by $i\in \mathcal I=\{1,2,...,n\}$. Each agent chooses whether or not to link with the other $n-1$ agents. Player $i$'s action vector $G_i=(G_{i1}, G_{i2},...,G_{ij},..., G_{in})'$ where $j\neq i$ is chosen from the action profile ${A}$ which has $2^{n-1}$ components. The payoff function of individual $i$ is
The deterministic part of incremental utility from link $ij$ is specified as
where the first term captures the homophily effect. $w$ is a known function. $X=(X_1',..., X_n')'$ is public information for all agents and is observable to researchers. For simplicity, write $w(X_i,X_j)=W_{ij}$ from now on. The second term $A_i$ is individual-specific heterogeneity, which is unobserved both to other agents and researchers. Let $F_{A|X}$ be the distribution of $A_i$ conditional on observables, which is assumed to be independent and identical across $i$, and known to all agents, but not necessarily known to researchers. $A_i$ can be correlated with $X$. The third and the last term capture the popularity effect. The realization of $\varepsilon_i=(\varepsilon_{i1},...,\varepsilon_{in})'$ is agent $i$'s private information which is also unobserved to researchers. The model is therefore a static game with incomplete information, and the solution concept is Bayesian Nash Equilibrium. Different from leung2015two, my model allows private information to be correlated with common information while doesn't require the conditional distribution of private information to be known to researchers. Also, I allow "asymmetric" equilibrium which will be mentioned later in this part.\\ For the above model, I impose the following assumptions:
Let $\delta_j(X,A_j,\varepsilon_j)$ denote agent $j$'s (pure) strategy. Let $\sigma_j(a|X, A_j)= Pr\left(\delta_j(X,A_j,\varepsilon_j)=a|X,A_j\right)$ denote the agent $i$'s belief that agent $j$ of type $A_j$ chooses action $a$, given commonly known information $X$ and agent $i$'s private information. By Assumption (ref) (b) and (c), actions $G_i$ and $G_j$, $i\neq j$ are independent given commonly known attributes $X$. This fact simplifies the proof of consistency by weakening the correlation between links. Since agent $i$ actually doesn't known the realization of $A_j$, so agent $i$'s expected utility from choosing action $g_i\in S$ is $\sum_{g_{-i}} U_i(g_i,g_{-i}, X, A_i, \varepsilon_i)\mathbb E_{A_{-j}}\big[\sigma_{-i}(g_{-i}|X, A_{-j})\big]$. Therefore,
A (Bayesian) equilibrium $\sigma^*(X, A_i)$ is a belief function that solves the fixed point equation:
for all $X\in \mathbf{X}$, agents $i\in \mathcal I$ and actions $a\in S$.\\ I consider "symmetric" equilibria in which pairs of agents with the same observable attributes and the same type ($A_i$) have the same conditional linking probabilities. For any $(X,A_i,\varepsilon_{ij})$ and "symmetric" belief profile $\sigma$ in a neighborhood of an "symmetric" equilibrium $\sigma^*$ , player $i$'s optimal strategy $G_{i}(X, A_i, \varepsilon_i, \sigma)=\big(G_{ij}(X, A_i, \varepsilon_i, \sigma)\big)_{j\neq i}$ is given by:
Assuming a symmetric equilibrium exists, the model is incomplete because there could be multiple equilibria for any realization of $(X, A, \varepsilon)$. For completeness of the model, I specify the equilibrium selection mechanism in the following assumption. The mechanism, however, is not explicitly used in writing the likelihood function in part 3, because by using two-step estimation, I can avoid specifying the equilibrium theoretically. For the convenience of defining equilibrium selection mechanisms, I add subscript $n$ to $G$, $X$, $A$, and $\varepsilon$. The equilibrium selection mechanism is a measurable function $\lambda_n: (X_n, \nu_n,\beta_0) \mapsto \sigma_n\in \mathcal{G}(X_n, A_n, \beta_0)$, where $\mathcal{G}(X_n, A_n, \beta_0)$ is the set of symmetric equilibria.
Define $P_{ij}(X,A_{i},\sigma)$ to be the probability that individual $i$ proposes to form a link with $j$ conditional on $X$ $A_{i}$, and $\sigma$. According to ((ref)) and Assumption (ref) (c),
Define $p_{ij}(X, A_{i})$ to be the equilibrium probability that agent $i$ proposes a link to agent $j$. which is realized in the data. Equilibrium condition requires that
For notation simplicity, denote $q_{jk}(X, \sigma^*):=\mathbb E_{A_{j}}\big[Pr(G_{ jk}(X,A_{j}, \varepsilon_{j},\sigma)=1|X, A_{j}, \sigma^*)\big]$, which is the probability that agent $j$ proposes a link to $k$ conditional on $X$ and the realized equilibrium $\sigma^*$. Then ((ref)) can be rewritten as
Although $p_{ jk}(X, A_{j})$ is not identified from data, $q_{ij}(X)$ is identified. With abuse of notations, let $q_{ st}(X)=\mathbb E_{A_{j}}\big[p_{ jk}(X_{j}=x_s, X_{k}=x_t, X,A_{j})\big]$. \\ Consider the empirical frequency of pairs with the same observable characteristics proposing to form a link:
First, I want to show that $q_{st}(X, \sigma^*)$ can be consistently estimated by $\hat q_{n, st}$ under the payoff function specified in (ref). Formally, I want to prove the following lemma:
For the convenience of the following analysis, I introduce a change of notation:
and
then by Lemma (ref), $\sup_{s,t}|\hat Z_{s,t}- Z_{s,t}|=O_p\big(n^{-1/2}\big)$\\ With the estimates $\hat q_n=\{\hat q_{st}\}_{\forall s,t}$, I propose to estimate the parameter $\beta$ and individual fixed effects $\{A_i\}_{i=1}^n$ jointly by MLE.\\ By Assumption (ref) (c), the conditional likelihood of the network is
By ((ref)) and ((ref)),
Construct the log-likelihood function:
Let $\hat \beta$ and $\hat A$ be the maximizer of the log-likelihood with $q$ replaced by $\hat q_n$.
By first concentrating out $A$, the estimators are given by:
where
By rearranging the sample score of ((ref)), it can be shown that $\hat A(\beta)$, when it exists, is the unique solution to the fixed point problem:
where
In this part, I first show the consistency of $\hat\beta$ and $\hat A$. Because link proposals from the same individual are correlated, the first step estimator has a slow convergence rate $\sqrt{n}$, which is equivalent to the usual convergence rate of $N^{1/4}$, since the number of summands in the likelihood function $N=n(n-1)$. As is well discussed in the nonlinear panel literature, there could be an estimation bias of $\hat\beta$ caused by the incidental parameters problem (e.g. hahn2004jackknife, arellano2007understanding). However, as I will show in this part, the effect of second-step bias is dominated by the slow convergence rate of the first step, so a bias term won't show up in the asymptotic distribution.
Compactness of the support (Assumption (ref) (a)(b) and Assumption (ref)) implies that
for some $0<\kappa<1$ and for all $A_i\in\mathbb A$, $\beta\in \mathbb B$ and $\forall q\in (k, 1-k)$.
With a more involved argument, I can actually show the uniform convergence rate of $\hat A$
To state the form of the asymptotic distribution, define
In this section, I implement the proposed method in a simulation study. Assume the following utility specification:
where $X_i$ is a random variable taking values in $\{1,-1\}$ with equal probability, and $\epsilon_{ij}$ follows the Logistic distribution. The distribution of $A_i$ is generated according to
with $\alpha_L< \alpha_H$ and $a_i\sim N(0, 0.1), V_i\sim N(0, \sqrt{0.1})$, and they are independent. In the simulation exercise, I consider three scenarios. In the first two scenarios, $A_i$ is correlated with $X$. In Scenario 1, I let $\alpha_L=-2/3, \alpha_H=-1/6$, and $\gamma=0$, so that the correlation between $A_i$ and $X_i$ is only through the value of $X_i$. In Scenario 2, I let $\alpha_L=-2/3, \alpha_H=-1/6$, and $\gamma=1$, so that the correlation between $A_i$ and $X_i$ is determined not only by the value of $X_i$ but also by the identity of $i$ (captured by the random variable $a_i$). In Scenario 3, I let $\alpha_L=-1/2, \alpha_H=-1/2$ and $\gamma=0$, so that $A_i$ is independent with $X$. The true values of the parameters are $(\beta_1, \beta_2, \beta_3)=(-2,1,1)$. The network is generated according to the $n-$player incomplete information game described in Section 2, with $n$ taking values of $50, 100, 250$, and $500$. For each value of $n$, I generate a single network and use the method proposed in this paper and leung2015two to estimate the parameters. When using leung2015two's estimator, the private information $\eta_{ij}$ is the sum of $A_i$ and $\epsilon_{ij}$ with $A_i\perp\epsilon_{ij}$. Each experiment is repeated $1000$ times. I report the means and standard errors of the estimated parameters in the tables below.
As can be seen in Table (ref) and (ref), when the private information is correlated with observed individual characteristics $X$, this paper's approach yields good estimates for the parameters, while leung2015two's estimator doesn't perform well, both in terms of the mean and variance of the estimators. This is not surprising since leung2015two assumes that private information and observable individual characteristics are independent. Under the correlated scenario, leung2015two's estimator will not be consistent. Table (ref) shows the simulation results when the individual private information $A$ is independent with observed characteristics $X$. Not surprisingly, both this paper's estimator and leung2015two's estimator perform reasonably well, except that leung2015two's estimator has larger variances.
In this paper, I characterize the network formation process as a static game of incomplete information, where the latent payoff of forming a link between two individuals depends on the structure of the network, as well as private information on agents’ attributes. I allow agents’ private unobserved attributes to be correlated with observables (i.e. existence of individual fixed effects). Using data from a single large network, I propose a two-step estimator for the model primitives. In the first step, I estimate agents’ equilibrium beliefs of other people’s choice probabilities. In the second step, I plug in the first-step estimator to the conditional choice probability expression and estimate the model parameters and the unobserved individual fixed effects together using Joint MLE. Assuming that the observed attributes are discrete, I showed that the first step estimator is uniformly consistent with the rate $n^{-1/2}$, where $n$ is the number of individuals in the network. This rate corresponds to the usual $N^{-1/4}$ rate where $N$ stands for the total number of linking proposals and is the effective sample size. The slow convergence rate is translated to the second step so that the usual asymptotic bias of order $N^{-1/2}$ caused by the "incidental parameter problem" won't show up in the asymptotic distribution. The second-step estimator $\hat\beta$ subtracted by its mean converges asymptotically to a normal distribution at the rate $N^{-1/4}$. Monte Carlo Simulation shows that the estimator proposed in this paper performs well in finite samples.
\setcounter{section}{1}