Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
59,133 characters · 14 sections · 43 citation commands
Identification of Peer Effects using Panel Data
This paper provides new identification results for panel data models of peer effects, through which outcomes depend on peers' unobservable (to the researcher) heterogeneity. Our framework also allows for correlated effects, modelled as unobserved group heterogeneity which is permitted to be correlated with individual heterogeneity in an unrestricted manner. We extend existing identification results to apply both to a general network structure (e.g., a social network) and to allow for correlated effects, and apply our results to study innovation take-up among physicians.
Identification depends on a conditional mean restriction requiring exogenous mobility of individuals between groups over time, as first formalised in abowd99. That is, though our model of correlated effects allows, for example, that high outcome individuals be systematically located in high outcome groups, their mobility between groups ought not be determined by transitory outcome shocks. Not all patterns of mobility suffice for identification, and we provide identifying and non-identifying examples. We also provide an extension of our identification results to allow for endogenous peer effects, through which outcomes depend directly on peers' outcomes. With endogenous effects, identification can also be attained using the conditional mean restriction, however, for certain network structures, additional conditional variance restrictions are necessary. We conduct a Monte-Carlo with many individuals and few time periods and demonstrate that the NLS estimator first proposed by arci12 works well in practice. Increasing the number of time periods, the rate of mobility and the richness of the network (e.g., social network data) improves the performance of the estimator.
Our empirical work considers innovation take-up in cancer treatment. The innovation we consider is keyhole surgery for colorectal cancer, and our data are from the English National Health Service. Colorectal cancer is the third most common cancer worldwide arnold17. In England, it accounts for 10% of cancer deaths and is the most expensive cancer to treat laudicella16. Keyhole surgery for colorectal cancer is an important innovation. It has been shown to lead to better patient outcomes than the alternative open procedure, particularly in the short term. Moreover, in the National Health Service, keyhole surgery is less costly, primarily due to shorter post-surgery hospital stays by patients lacy02,nelson04,laudicella16. Despite this, its take-up was slow, increasing from 1% of eligible surgeries when it was first introduced in 2000 to 49% by 2014.
We use matched patient-surgeon-hospital-year data from 2000 to 2014 to estimate peer effects in take-up of keyhole surgery, measured by the fraction of colorectal cancer surgeries performed by keyhole. We find positive and statistically significant peer effects. Our results suggest that a standard deviation increase in the average latent take-up of other surgeons in the same hospital leads to a 5 percentage point increase in take-up. We decompose this effect by additionally estimating the effect of peer experience, which we find accounts for some, but not all, of the peer effect.
Research has largely focussed on settings in which peer effects operate through observable characteristics. That is, in addition to correlated effects, outcomes depend on peers' observables manski93,moffitt01, lee07,bramoulle09,calvo09,davezies09,degiorgi10,goldsmith13,blume15,depaula17,cohen18,bramoulle19. These papers consider a sample comprising a cross-section of groups, for which identification requires within-group variation in peers. Our panel data allow us to relax these requirements. First, the researcher need not observe individual characteristics, and if observed, need not know which ones are exogenous nor which ones are appropriate to include. Second, variation induced by mobility of individuals between groups over time means that there is no need for within-group variation in peers. This implies that identification is possible under the linear-in-means network structure, in which peer effects operate through the group average, and which precludes identification when only a cross-section of groups is available manski93,bramoulle09.
Another strand of the literature considers peer effects operating through individual unobservables using a cross-section of groups graham08,rose17. Relative to these papers, we allow for unobserved group level heterogeneity to be arbitrarily correlated with individual heterogeneity, and, though our results can be applied to any network, we show that panel data can be used to identify peer effects for the linear-in-means network, which are not identifiable with only a cross-section of networks when there are correlated effects rose17.
The most closely related work considers panel data models of peer effects. Key contributions are mas09, arci12 and cornel17, which, like us, study models of peer effects operating through the unobserved heterogeneity of peers. We build on their work by providing identification conditions which are straightforward to verify in practice, and which can be applied to a general network structure (e.g., social networks) in the presence of correlated effects. We also extend our results to the canonical model of endogenous peer effects, in which outcomes are simultaneously determined.\footnote{arci12 consider a variant of endogenous effects through which outcomes depend on the expected (as opposed to realized) outcomes of others, and study identification in an example with 2 individuals. This model does not allow simulataneity in outcomes.}
Beyond the peer effects literature our work can be viewed as extending the worker-firm fixed effects framework for wage decomposition used in labor economics (e.g., abowd99) to allow for within-firm interactions of workers. That is, in addition to worker and firm heterogeneity, wages may depend on the composition of other workers in the firm as well as their wages. Such spillovers would be expected to operate in firms in which workers work in teams mas09,cornel17.
Our empirical work constributes to the literature on innovation take-up, particularly in the healthcare context. Related papers are agha18 and barrenho20. agha18 study how the take-up of new cancer drugs depends on the presence of local opinion leaders. To do this, they compare diffusion patterns across regions, separating correlated regional demand for new technology from information spillovers. They find that take-up is fastest in the region in which the lead author on the clinical trial practices. However, their work does not directly study peer effects in innovation. barrenho20 study the take-up of keyhole surgery, and estimate peer effects through surgeon observables. Our empirical contribution is to additionally allow peer effects to operate through surgeons' latent propensity to innovate, which we show matters above and beyond the effect of peer experience.
We proceed as follows. In Section (ref), we present our baseline model. In Section (ref) we provide identification results, apply them to two examples and discuss estimation. In Sections (ref) and (ref) we conduct a Monte Carlo experiment and apply our method to surgeons' take-up of keyhole surgery for colorectal cancer. In Section (ref) we conclude. Proofs and extensions are in the Appendix.
If $M$ is a strictly positive integer, we denote $[M]=\{1,2,...,M\}$. If $\mathbf{A}$ and $\mathbf{B}$ are $M\times P$ and $M\times Q$ matrices, we denote the $M\times (P+Q)$ matrix obtained by concatenating $\mathbf{A}$ and $\mathbf{B}$ by $[\mathbf{A},\mathbf{B}]$. If element $(i,j)$ of $\mathbf{A}$ is $\mathbf{A}_{ij}$, we write $\mathbf{A}=(\mathbf{A}_{ij})_{i\in [M],j\in[P]}$, and if $\mathbf{a}$ is a vector, $\mathbf{a}_i$ denotes entry $i$. We use $\mathbf{I}_M$ for the $M$ dimensional identity, $\iota_M$ for the $M\times 1$ vector of ones and $\mathbf{0}$ to denote a matrix of zeros. If its dimensions are ambiguous we write $\mathbf{0}_{M,P}$ to denote a $M\times P$ matrix of zeros. We use $\mathbf{1}(\cdot)$ to denote the indicator function.
We consider a pattern of mobility of $N$ workers between $M$ groups over $T$ periods. Our interest lies in the typical setting in which $T$ is small and $N$ can be large. A mobility pattern is characterised by the $N\times M$ matrices of group membership indicators $\mathbf{C}_1,\mathbf{C}_2,...,\mathbf{C}_T$ and the $N\times N$ interaction matrices $\mathbf{G}_1,\mathbf{G}_2,...,\mathbf{G}_T$, the elements of which encode the peer effect exerted by one individual on another. The groups are such that in each period, every individual is in exactly one group and every group has at least one individual in at least one period. We use $g(i,t)\in[M]$ to denote the group of individual $i\in[N]$ in year $t\in[T]$ and $N_{g(i,t)}$ to denote its size.
The $N\times 1$ vector of continuous outcomes in period $t\in[T]$ is $\bold{y}_t$, which is determined by
where $\alpha$ is an $N\times 1$ vector of time-invariant individual unobserved heterogeneity, $\gamma$ an $M\times 1$ vector of unobserved group heterogeneity, and $\epsilon_{t}$ an $N\times 1$ disturbance. In Section (ref), we consider an extension in which $\gamma$ is permitted to vary by period. The parameter of interest is $\rho$, which captures the peer effect.
Unless otherwise stated $\mathbf{G}_t$ is unrestricted and can be interpreted as the adjacency matrix of a weighted, directed network linking $N$ individuals. In particular, denoting entry $(i,j)$ of $\mathbf{G}_t$ by $\mathbf{G}_{ijt}$, it need not be the case that $\mathbf{G}_{ijt}=0$ when $g(i,t)\neq g(j,t)$, nor when $i=j$, nor when $\mathbf{G}_{jit}=0$. From this point onwards, we refer to $\mathbf{G}_t$ as the network. A typical example of $\mathbf{G}_t$ is the linear-in-means network,
This implies that the peer effect operates through the group average of $\alpha_i$, and is a natural choice when only group membership indicators are available. If more detailed data on the network structure are available (e.g., social network data), this information can be incorporated into $\mathbf{G}_t$. Clearly, $\mathbf{G}_t$ can evolve over time as individuals move between groups.
In our empirical application, $\mathbf{y}_t$ is surgeons' take-up of keyhole surgery for colorectal cancer, measured by the fraction of eligible surgeries performed by keyhole in year $t$, $\mathbf{C}_t$ comprises indicators for the hospital in which a surgeon practices, and we take $\mathbf{G}_t$ to be linear-in-means, linear-in-others'-means (in which the focal surgeon is excluded from the group average, see (ref)) and a persistent version of linear-in-others'-means, in which links persist when a surgeon moves to a new hospital and are weighted by the number of years worked at the same hospital. The latter captures cumulative peer exposure, which may be better suited to innovation take-up. The vector $\alpha$ captures surgeons' latent propensity to take up keyhole surgery. It can account for heterogeneity in education/training, ability (e.g., dexterity) and taste for innovation. The vector $\gamma$ captures hospital level heterogeneity including resources (e.g., equipment) and patient composition.
Stacking (ref) by period yields $\mathbf{y}=\left(\mathbf{J}+\rho\mathbf{G}\right)\alpha+\mathbf{C}\gamma+\epsilon$, where $\mathbf{y}$ and $\epsilon$ are $NT\times 1$, $\mathbf{C}=(\mathbf{C}_1',\mathbf{C}_2',\hdots, \mathbf{C}_{T}')'$, $\mathbf{J}=(\mathbf{I}_N,\mathbf{I}_N,\hdots, \mathbf{I}_{N})'$ and $\mathbf{G}=(\mathbf{G}_1',\mathbf{G}_2',\hdots, \mathbf{G}_{T}')'$. Since $\sum_{i=1^N}\mathbf{J}_{ki}=\sum_{f=1}^M\mathbf{C}_{kf}=1$ for all $k\in[NT]$, we use the normalization $\gamma_M=0$ to obtain
where $\mathbf{D}$ comprises the first $M-1$ columns of $\mathbf{C}$ and from this point forwards $\gamma=(\gamma_1,\gamma_2,...,\gamma_{M-1})'$. This is without consequence for identification of $\rho$.
Our identification results depend on variation in the network (both over individuals and over time) and mobility of individuals between groups over time. It is well known that mobility serves to separate the individual and correlated effects (e.g., abowd99). As we show below, mobility also serves to separate individual and correlated effects from the peer effect. This is because, in the typical setting in which there are no between group links,\footnote{i.e., if $g(i,t)\neq g(j,t)$ then $\mathbf{G}_{ijt}=0$, though our results do not require this.} if an individual moves from one group to another she ceases to interact with others in her previous group and begins to interact with others in her new group.
Our identification results treat the joint distribution of $\mathbf{y},\mathbf{G},\mathbf{D}$ as observable. This is consistent with the researcher accessing a sample from this distribution. For small $T$, our approach is identical in spirit to that used throughout the peer effects literature (i.e., with $T=1$), in which the researcher is assumed to observe a cross-section of groups (see bramoulle19 for a review).
A more challenging but sometimes more realistic alternative is that the researcher observes a sample of individuals from a single group, so that, depending on the network structure, all individuals are potentially linked to one another, at least indirectly. It is more challenging because additional structure is required to deal with the dependence between individuals, though this is primarily a concern for inference rather than identification. goldsmith13 describe how asymptotic analysis could be implemented based on a random variable measuring `distance' between individuals. Their argument is as follows. If distant individuals have low probability of link formation (e.g., due to homophily), and distance has large support, it may be possible to construct blocks of individuals such that each pair of blocks could be treated as close to independent. This is in the same spirit as certain types of asymptotic analysis for time-series, and allows the researcher to view a sample of individuals from a single group similarly to a sample of many groups.
goldsmith13 use the above arguments to justify conducting identification analysis as if the joint distribuion of $\mathbf{y},\mathbf{G},\mathbf{X}$ were observable, where $\mathbf{X}$ is a matrix of exogenous characteristics through which peer effects may operate.\footnote{Their model does not include correlated effects, hence the absence of $\mathbf{D}$.} Their arguments can be directly applied in our context to justify analysis of a sample from the joint distribution of $\mathbf{y},\mathbf{G},\mathbf{D}$ if $T$ is small. The only caveat is that $\mathbf{y},\mathbf{G},\mathbf{D}$ must include all observed time periods for the individuals and groups, so that temporal dependence is not an issue in (hypothetically) constructing blocks of observations. In our application, such blocks could be thought of as corresponding to regions of England because mobility is primarily between hospitals in the same region goldacre13,barrenho20. In other applications such as wage decomposition in labor economics, the relevant partition could be by industry and/or by region.
We study identification of $\rho$ based on the conditional mean restriction
which implies exogeneity of the network and mobility of individuals between groups over time with respect to the outcome shock $\epsilon$. The former is typical in the peer effects literature and the latter is typical in the wage decomposition literature. We do not restrict $\mathbb{E}[\alpha|\mathbf{G},\mathbf{D}]$ nor $\mathbb{E}[\gamma|\mathbf{G},\mathbf{D}]$, which allows for a limited form of network endogeneity with respect to unobserved individual and group heterogeneity.
We say that $\rho$ is identified when it can be uniquely recovered from the right-hand side of (ref). Our results are thus asymptotic in nature (see manski95), and hence charaterize whether peer effects can be distentangled from individual and group heterogeneity if there is no limit to the number of mobility patterns observed. To simplify the exposition, following bramoulle09 and abowd99, the remainder of the paper presents the case in which $\mathbf{G}$ and $\mathbf{D}$ are treated as fixed. To allow for the random case, we simply replace unconditional expectations with expectations conditional on $\mathbf{G},\mathbf{D}$, and the identification results below hold if there exists a realization in the support of $\mathbf{G},\mathbf{D}$ which satisfies the relevant condition. In the same way, it is straightforward to allow $N,M$ and $T$ to vary by mobility pattern. Identical arguments are used throughout the peer effects literature (see, e.g., bramoulle09). Returning to the fixed case, (ref) becomes
where $\mu^\alpha=\mathbb{E}[\alpha]$ and $\mu^\gamma=\mathbb{E}[\gamma]$. Since (ref) is non-linear in the parameter $\rho$ and the (unknown) $\mu^\alpha$, establishing identification is non-trivial, depending both on the properties of the $NT\times 2N+M-1$ matrix $[\mathbf{J},\mathbf{G},\mathbf{D}]$ and on the value of $\mu^\alpha$. For example, it is clear that $\rho$ is not identified when $\mu^\alpha=\mathbf{0}$.
Mobility of individuals between groups over time is necessary for identification. In the absence of mobility, $[\mathbf{J},\mathbf{G},\mathbf{D}]$ has rank at most $N$, so (ref) yields $N$ equations in $N+M$ unknowns. For the same reason, we also require $T\geq 2$. This is because the peer effect operates through time-invariant individual unobserved heterogeneity, rather than through individual observables. However, mobility alone is not sufficient for identification, as made clear in Example 2 below.
We now present our first identification result, which makes use of the within-group annihilator for the correlated effects, given by $\mathbf{W}=\mathbf{I}_{NT}-\mathbf{D}(\mathbf{D}'\mathbf{D})^{-1}\mathbf{D}$, and a decomposition of vectors $\mathbf{v}=(\mathbf{v}_1',\mathbf{v}_2')'$ which lie in the null-space of $[\mathbf{WJ},\mathbf{WG}]$ such that $\mathbf{v}_1$ and $\mathbf{v}_2$ are both $N\times 1$.
Notice that full column rank of $[\mathbf{J},\mathbf{G},\mathbf{D}]$ is not necessary because $\mathbf{v}=\mathbf{0}$ need not be the only vector in the null-space of $[\mathbf{WJ},\mathbf{WG}]$. This is because $[\mathbf{J},\mathbf{G},\mathbf{D}]$ has $2N+M-1$ columns but there are only $N+M$ unknowns. Requiring $[\mathbf{J},\mathbf{G},\mathbf{D}]$ to have full column rank is too strong because it rules out $T=2$.\footnote{This is because full column rank requires $NT\geq 2N+M-1$, hence $T\geq2+(M-1)/N$. $T=2$ is not immediately ruled out when $M=1$, but $\rho$ is not identifiable in this case because there can be no mobility if there is only 1 group.} A common empirical setting is when the rows of $\mathbf{G}$ sum to one (e.g., linear-in-means), such that peer effects operate through a weighted average. If there are no other collinearities among the columns of $[\mathbf{J},\mathbf{G},\mathbf{D}]$ then one can apply the following.
Corollary (ref) can be shown using the decomposition $[\mathbf{WJ},\mathbf{WG}]=\mathbf{S}\mathbf{R}$ where $\mathbf{S}$ is the $NT\times 2N-1$ full rank matrix formed by concatenating $\mathbf{WJ}$ and the first $N-1$ columns of $\mathbf{WG}$ and
From the structure of $\mathbf{R}$, it is immediate that vectors $\mathbf{v}$ in the null-space of $[\mathbf{WJ},\mathbf{WG}]$ are of the form $\mathbf{v}_1=c\iota_N$, $\mathbf{v}_2=-\mathbf{v}_1$ for $c\in\mathbb{R}$. Applying Proposition (ref) yields identification of $\rho$ if there exists $(i,j)\in[N]^2$ such that $\mu^\alpha_i\neq\mu^\alpha_j$. If this condition is violated then $\mu^\alpha=a\iota_N$ for some $a\in\mathbb{R}$, and $(\mathbf{J}+\rho\mathbf{G})\mu^\alpha=(1+\rho)a\iota_{NT},$ so only $(1+\rho)\mu^\alpha$ is identifiable. Intuitively, we require $\mu^\alpha_i\neq\mu^\alpha_j$ because if individuals are homogeneous then no amount of mobility can lead to changes in the average of $\alpha_i$ over peers, hence outcomes do not vary in response to changes in peer composition over time.
The identification conditions above depend on $\mu^\alpha$, which is not observed. We now ask how `large' is the set of values of $\mu^\alpha$ for which $\rho$ is identified. This is the notion of generic identification (see lewbel19). If $\rho$ is generically identified, then it is identified for all values of $\mu^\alpha$ with the exception of a few pathological cases, which are `unlikely' to arise in practice.
Corollary (ref) means that $\rho$ is generically identified if ${\rm rank}([\mathbf{WJ},\mathbf{WG}])\geq N+1$. This is because $\mathbf{v}$ lies in a subspace of $\mathbb{R}^{2N}$ of dimension at most $N-1$, hence $\delta_1\mathbf{v}_1+\delta_2\mathbf{v}_2$ lies in a subspace of $\mathbb{R}^N$ of dimension at most $N-1$.
As with the well known rank condition in linear models (e.g., no perfect multicollinearity among regressors for linear regression), the researcher can check whether the rank requirements of Corollaries (ref) and (ref) hold in the observed data. To do this, one interprets $N$ as the total number of observed individuals, $M$ as the total number of observed firms and $T$ as the total number of observed time periods, and constructs the observed values of $\mathbf{J},\mathbf{G}$ and $\mathbf{D}$ accordingly.\footnote{In our identification analysis, these quantities are the number of individuals, firms and time periods in a single mobility pattern (i.e., a draw from the joint distribution of $\mathbf{y},\mathbf{G},\mathbf{D}$). The observed data comprise the realizations of many such patterns.} If the rank requirement of either Corollary holds then the researcher can be confident of identification.
Proposition (ref) can be viewed as panel data analogue of the identification results of bramoulle09, which imply that, for $T=1$ and exogenous observed characteristic(s) $\mathbf{X}$ of dimension $N\times K$, the peer effect $\dot\rho$ in the model
is identified if and only if $\mathbf{WJ}$ and $\mathbf{WG}$ are linearly independent (i.e., if there does not exist nonzero $\lambda\in\mathbb{R}^2$ such that $\lambda_1\mathbf{WJ}+\lambda_2\mathbf{WG}=\mathbf{0}$).\footnote{Since $\alpha$ does not appear in (ref), there is no need to impose $\gamma_M=0$, which is no longer a normalization. In this case $\mathbf{D}$ ought to be replaced by $\mathbf{C}$ and $\mathbf{W}$ by the annihilator for $\mathbf{C}$ when referring to the model in (ref). We continue to use the notation $\mathbf{D}$ and $\mathbf{W}$ to facilitate the comparison with our results.} This is a weaker rank requirement than that in Proposition (ref), and can be satisfied when $T=1$. This is because the peer effect is assumed to operate through the observed $\mathbf{X}$ rather than through the unobserved $\alpha$. In practice the researcher may not know what to include in $\mathbf{X}$ and/or $\mathbf{X}$ may be of large dimension and/or endogenous. Proposition (ref) shows that identification can be attained with $T\geq 2$ without requiring any knowledge on $\mathbf{X}$.\footnote{Of course this requires that $\mathbf{X}$ be time-invariant. However, common choices of the components of $\mathbf{X}$ such as gender and education are also time-invariant.} An additional advantage of panel data is that it facilitates identification for some network structures for which peer effects are not otherwise identifiable. For example, if $T=1$ and the network is linear-in-means, then $\mathbf{WG}=\mathbf{0}$ and (ref) does not identify the peer effect. In contrast, as we show in Example 1 below, peer effects are identifiable when $T=2$ provided that there is mobility of individuals between groups over time.
It is straightforward to extend our model to include exogenous characteristics, yielding
where $\mathbf{X}$ is $NT\times K$, $\mathbf{F}$ is a $NT\times NT$ block diagonal matrix with blocks $\mathbf{G}_1,\mathbf{G}_2,...,\mathbf{G}_T$, $\mu^\alpha(\mathbf{X})=\mathbb{E}[\alpha|\mathbf{X}]$ and $\mu^\gamma(\mathbf{X})=\mathbb{E}[\gamma|\mathbf{X}]$. Provided that the entries of $\mathbf{X}$ vary over time, the parameters $\rho_1$ and $\rho_2$ are separately identifiable under similar conditions to Proposition (ref). For brevity, we do not pursue this formally, though we do estimate such a specification in our empirical application.
In the Appendix we provide analagous idenification results which additionally allow for endogenous peer effects, through which outcomes are simultaneously determined. We also provide conditional variance restrictions in the spirit of graham08 and rose17, which can be used when the conditional mean restriction does not suffice for identification.
We now consider two examples of our baseline identification results with $N=M=T=2$.\\
Example 1: An identifying mobility pattern. Consider the following mobility pattern in which individuals one and two are respectively in groups 1 and 2 in the first period. In the second period, individual one remains in group 1 and individual two moves from group 2 to group 1. Under the linear-in-means network, this mobility pattern yields
which respectively have rank $4$ and $3$. Since $[\mathbf{J},\mathbf{G},\mathbf{D}]$ has rank $2N+M-2$, by Corollary (ref), $\rho$ is identified if $\mu^\alpha_1\neq\mu^\alpha_2$. Since ${\rm rank}[\mathbf{WJ},\mathbf{W}\mathbf{G}]=N+1$, Corollary (ref) states that $\rho$ is generically identified. This is because the subset of $\mathbb{R}^2$ such that $\mu^\alpha_1=\mu^\alpha_2$ has measure zero.
For the intuition, consider the underlying system of equations for the outcomes $y_{it}$ of individual $i$ in period $t$,
When individual two moves groups, individual one obtains a new peer, hence the peer effect on individual one changes from $\rho\mu^\alpha_1$ in the first period to $\rho(\mu^\alpha_1+\mu^\alpha_2)/2$ in the second period, whilst the individual and correlated effect are unchanged. The change due to the peer effect is given by the expected change in the outcome of individual one between the first and second period $\mathbb{E}[y_{11}]-\mathbb{E}[y_{12}]=\rho(\mu^\alpha_1-\mu^\alpha_2)/2$. To identify $\rho$ we now need to identify $\mu^\alpha_1-\mu^\alpha_2$. We can use the second period, in which both individuals are in the same group, hence have the same correlated and peer effects. This means that any difference in their expected outcomes is due to differences in their individual effects, so $\mathbb{E}[y_{12}]-\mathbb{E}[y_{22}]=\mu^\alpha_1-\mu^\alpha_2$. If $\mu^\alpha_1\neq\mu^\alpha_2$ we obtain
If we were to observe this mobility pattern repeatedly, we could estimate the expectations using sample means, yielding an estimator of $\rho$. Of course, in practice we do not repeatedly observe the same mobility pattern, but a variety of patterns, the information from which we combine through an estimator based on the conditional mean restriction (ref). Nevertheless, for the purposes of identification we require only that there exists a single identifying mobility pattern realized with non-zero probability. Finally, notice that the above arguments are unchanged when the normalization $\gamma_M=0$ is not used, in which case one has $\mathbb{E}[y_{21}]=\mu^\alpha_2+\rho\mu^\alpha_2+\mu^\gamma_2$ in (ref).\\
Example 2: A non-identifying mobility pattern. Now modify Example 1 such that both individuals move in the second period. This means that the individuals are in different groups in the first period and swap groups in the second period, implying that $\mathbf{G}=\mathbf{J}$. The null-space of $[\mathbf{WJ},\mathbf{W}\mathbf{G}]$ comprises $\mathbf{v}=(\mathbf{v}_1',\mathbf{v}_2')'=(\mathbf{u}',-\mathbf{u}')'$ for any $\mathbf{u}\in\mathbb{R}^2$. For any value of $\mu^\alpha\in\mathbb{R}^2$, there clearly exists $\mathbf{u}\in\mathbb{R}^2$ and scalars $\delta_1$ and $\delta_2\neq 0$ such that $(\delta_1-\delta_2)\mathbf{u}=\mu^\alpha$, so by Proposition (ref), $\rho$ is not identified. Note also that the identification condition in Corollary (ref) is violated since $[\mathbf{J},\mathbf{G},\mathbf{D}]$ has rank $3<2N+M-2=4$ and the generic identification condition in Corollary (ref) is violated since $[\mathbf{WJ},\mathbf{W}\mathbf{G}]$ has rank $2<N+1=3$.
For the intuition, consider again the underlying system of equations
The correlated effect $\mu^\gamma_1$ is identified by individual one moving groups ($\mu^\gamma_1=\mathbb{E}[y_{11}]-\mathbb{E}[y_{12}]$) but the remaining equations are only sufficient to identify $(1+\rho)\mu^\alpha$. The reason for this is that there is no variation in the peers of either individual because the individuals are never in the same group in the same period. Mobility of individuals between groups over time is insufficient for identification of $\rho$ because it does not induce changes in peer groups. This contrasts with the canonical wage decomposition model imposing $\rho=0$, for which group-swapping would be sufficient to identify $\mu^\alpha_1,\mu^\alpha_2,\mu^\gamma_1$.\footnote{Note that identification of $\mu^\alpha$ and $\mu^\gamma$ is only up to the normalization $\gamma_M=0$.} Since it does not identify $\rho$, no matter how many times we observe this mobility pattern in our data, it cannot be used to construct an estimator of $\rho$.
If correlated effects are time-varying, then $\gamma_{g(i,t)}$ is replaced by $\gamma_{g(i,t)t}$, in which case $\mathbf{C}$ is a $NT\times MT$ block diagonal matrix with blocks $\mathbf{C}_1,\mathbf{C}_2,...,\mathbf{C}_T$, $\mathbf{D}$ comprises the first $MT-1$ columns of $\mathbf{C}$ and $\gamma$ is $(MT-1)\times 1$. All Propositions then apply as stated, provided that $\mathbf{W}$ is modified accordingly. The parameters are not identifiable under the linear-in-means network because the peer effect varies only at the group-period level, so cannot be separated from the correlated effect.
arci12 propose NLS estimation of (ref), treating $\alpha,\rho,\gamma$ as parameters to be estimated (i.e., a fixed effects approach). The authors study its properties under the linear-in-others'-means network,
and without correlated effects. Assuming that $\mathbb{E}[\epsilon_{it}\epsilon_{js}]=0$ for all $i\neq j, t\neq s$, $\mathbb{E}[\epsilon_{it}^2|g(i,t)]=\mathbb{E}[\epsilon_{jt}^2|g(j,t)]$ for all $i,j,t$ such that $g(i,t)=g(j,t)$, $\mathbb{E}[\epsilon_{it}\alpha_j]=0$ for all $i,j,t$ and $\rho<\min_{i,t}N_{g(i,t)}$, the authors show that the estimator of $\rho$ is consistent and asymptotically normal in the number of individuals provided that there are at least two time periods. The authors discuss a variant of correlated effects appropriate to their application to peer effects in education, which is allowed to vary over time but is restricted to be the same across multiple groups, and argue that they expect similar behavior of the estimator in this case. The proposed NLS estimator is based on the conditional mean restriction (ref), hence, subject to identification, could equally be applied to other network structures and specifications of the correlated effect. We do not formally establish its properties because our focus is on identification and our empirical application. However we do explore this in our Monte-Carlo experiment.
The design is tailored to our empirical application, with 700 individuals, 140 groups, 15 time periods, mobility rate $p=0.03$ (this is the probability that an individual moves groups from one period to the next, see summary statistics in Table (ref)) and $\rho=0.5$. We also consider designs with 2 time periods and mobility rate $p=0.1$.
The data generating process is as follows. In the first period, all individuals are randomly assigned to groups of size five. In each subsequent period, each individual moves group with probability $p$, in which case she draws a new group with uniform probability over all groups. The expected group size is 5 for all groups in all periods. The network structures we consider are linear-in-means, linear-in-others'-means, and a variant of linear-in-others'-means in which links persist when individuals move groups and are weighted by the number of periods spent in the same group, and a social network, in which there is within-group variation in peers.
For the persistent network, we let $\mathbf{A}_s$ be the $N\times N$ binary adjacency matrix in period $s$, with element $(i,j)$ equal to 1 if $g(i,s)=g(j,s)$ and $i\neq j$. Then we define $\mathbf{G}_t$ by taking $\sum_{s=1}^t\mathbf{A}_s$ and rescaling its rows to sum to 1. This means that $\mathbf{G}\alpha$ captures both contemporaneous and cumulative exposure to others.
The social network is constructed as follows. In the first period, each individual draws two links uniformly over other individuals in the same group. In each subsequent period, links persist whilst individuals remain in the same group. If an individual loses a link(s) due to mobility, a replacement link(s) is drawn uniformly over the other individuals in the group with whom there is not already a link. Links need not be reciprocal.\footnote{If an individual is in a group of size 1 they have no link.} If there exists a link between $i$ and $j$ in period $t$ then $\mathbf{G}_{ijt}$ is equal to the inverse of the number of links that $i$ has in period $t$. Otherwise $\mathbf{G}_{ijt}=0$. If there are three or fewer individuals in the group, each individual is linked to all other individuals.
We take $\alpha_i=1+W_i+\eta_i$, where $W_i\in[0,1]$ is the number of moves made by individual $i$ divided by $(T-1)$ and $\eta_i\sim\mathcal{N}(0,1)$. We set $\epsilon_{it}\sim\mathcal{N}(0,1/2)$ and consider cases in which $\gamma=\mathbf{0}$ (which is also imposed on the estimator) and in which $\gamma_m$ is the mean of $\alpha_i$ over all members of group $m$ and all time periods. This design means that high $\alpha_i$ individuals are more mobile and tend to be located in high $\gamma_m$ groups. We choose the variances of $\alpha_i$ and $\epsilon_{it}$ as above so that, on average, a standard deviation increase in the peer effect on individual $i$ in period $t$ (i.e., in $\sum_{j=1}^N\mathbf{G}_{ijt}\alpha_j$) leads to a 0.15-0.3 (depending on the network structure) standard deviation increase in $y_{it}$. This matches the effect sizes we find in our empirical application. We simulate 500 datasets for each experiment. Every dataset verifies the generic identification condition in Corollary (ref).
Table (ref) reports the results. The top panel reports the NLS estimator of $\rho$ described in Section (ref). To provide a benchmark for comparison, the bottom panel reports the infeasible OLS estimator that would be used if $\alpha$ were observable (i.e., for the model in (ref) taking $\mathbf{X}=\mathbf{J}\alpha$). Columns 1-2 of Table (ref) show designs with 15 periods and mobility rate 0.03, which match our empirical application. The NLS estimator of $\rho$ is centered on the true value over all networks and with and without correlated effects. The variance of the NLS estimator is of a similar order of magnitude to the infeasible OLS estimator, though of course it is larger. Columns 3-4 reduce the number of periods to 2, which causes an increase in the variance of the estimator. The final 4 columns increase the rate of mobility to 0.1, which reduces the variance of the estimator. Comparing the rows of Table (ref), we find that peer effects are most precisely estimated for the social network. This is likely because the social network exhibits the most within-group variation in peers.
We study surgeons' take-up of keyhole surgery for colorectal cancer in the English National Health Service (henceforth, NHS). As discussed in the introduction, colorectal cancer is prevalent, costly to treat, and accounts for a high proportion of cancer deaths worldwide. An important innovation in its treatment is keyhole surgery, which reduces costs and improves patient outcomes relative to the older open procedure. Despite this, take-up of keyhole surgery in England was slow, increasing from 1% of eligible surgeries in 2000, the year it was introduced in England, to 49% by 2014 (see Table (ref)). Our goal is to study the extent to which its diffusion was driven by peer effects, hence by mobility of surgeons between hospitals over time.
The NHS is an ideal setting for our empirical work because it treats almost all cancer patients in England,\footnote{There is a small private sector in England that mainly provides care for planned procedures for which there are long waiting lists. Private sector provision for (any) cancers during the period we examine was very limited and primarily focused on treatment of overseas patients.} it is a public system in which surgeons are salaried employees of one hospital at any point in time (hence their renumeration does not depend on the treatment provided), all hospitals operate under the same financial rules set by central government and have the technology necessary for keyhole surgery, and a two week waiting time guarantee for cancer referral implies that allocation of patients to surgeons is close to random, based on surgeon availability in a local hospital barrenho20.
Our data are from barrenho20, comprising a panel of NHS surgeons and their take-up of keyhole surgery from its introduction in 2000 through to 2014. The data are obtained by merging Hospital Episode Statistics, which provides treatment information for all patients in the English NHS with NHS Workforce Statistics and the General Medical Council register. This provides matched patient-surgeon-hospital-year data, which is collapsed into a surgeon-hospital-year panel. To be included in the estimation sample, surgeons must be observed at least twice, must perform more than 5 colorectal cancer surgeries,\footnote{This is to avoid including surgeons who do not routinely perform the procedure.} and hospital-year pairs must have at least two surgeons. The resulting unbalanced panel comprises 11,923 observations of 1,363 surgeons over 15 years across 194 hospitals.
The dependent variable is surgeon take-up of keyhole surgery for colorectal cancer, measured by the fraction of eligible colorectal cancer sugeries performed by keyhole in the focal year.\footnote{Keyhole surgery is suitable for some, but not all, patients. The sample of patients considered is restricted to those for which the surgeon has a choice between keyhole and the alternative open surgery. A detailed description of patient eligibility can be found in barrenho20.} The surgeons we observe are senior physicians, known as Consultants. They are entirely autonomous, and, for the patients we consider, have discretion to choose either keyhole or open surgery. We also observe surgeon demographics (gender and age, though we omit gender from our models because it is time-invariant), and compute an experience measure which is the cumulative number of eligible colorectal cancer surgeries performed (both keyhole and open) by the beginning of the focal year. To account for increasing average take-up over time, we also include year fixed effects in all specifications.
We apply our model for surgeon $i$ in hospital $g(i,t)$ in year $t$,\footnote{We set a surgeon's hospital to be that at which they practiced for the majority of days of the year.} and consider the linear-in-means and linear-in-others'-means networks, as well as the persistent linear-in-others'-means network used in our Monte-Carlo experiment. Recall that in the persistent network, a surgeon is a peer if they have ever worked concurrently in the same hospital, and peers are weighted by number of years worked together, with weights summing to one. This network captures both contemporaneous and cumulative exposure to others, hence measures a `stock' of peer influence, rather than the `flow' measured by the other networks.
Table (ref) summarises the data. In a typical hospital in a typical year there are 5 surgeons performing colorectal cancer surgery. The largest has 16 surgeons. A typical surgeon is in their early 40s and has performed around 140 colorectal cancer surgeries by the beginning of a typical year. We only observe surgeons beginning in 2000, hence experience is equal to zero in this year for all surgeons. The minimum value of experience can be zero in any year due to entry of newly qualified surgeons.\footnote{For this reason, in addition to taste and ability, $\alpha$ may be thought of as including experience in colorectal cancer surgery prior to 2000.}
Our identification results are based on a sample from the joint distribution of $\mathbf{y},\mathbf{G},\mathbf{D}$. As argued in Section (ref), we also expect them to be applicable if the sample can be partitioned into blocks based on individuals' `distance' to one another, such that the dependence between each pair of blocks is limited. Mobility of surgeons between hospitals is largely within the local region goldacre13,barrenho20, hence such blocks could (hypothetically) be constructed by partitioning regions. Due to this, we expect that our theory provides a reasonable approximation.
Identification is based on exogenous mobility of surgeons between hospitals over time. We expect this to be the case because all hospitals have the required technology for keyhole surgery, the NHS is a public system with similar working conditions and renumeration nationwide, and colorectal cancer forms only a small part of surgeons' workloads. The leading reason for mobility is to relocate closer to the pre-medical school family home goldacre13. barrenho20 provide empirical evidence of exogenous mobility, showing that past take-up does not predict mobility, and that, conditional on moving, there is no association between take-up of the surgeon and take-up of the hospital moved to. These arguments support the conditional mean restriction in (ref). Since $\mathbb{E}[\alpha|\mathbf{G},\mathbf{D}]$ and $\mathbb{E}[\gamma|\mathbf{G},\mathbf{D}]$ are unrestricted, we allow for dependence between a surgeon's latent propensity to take-up keyhole surgery and the nature of their network. Moreover, we do not rule out high take-up surgeons being systematically located in high take-up hospitals, though we do rule out their mobility between hospitals being driven by transitory take-up shocks.
We now discuss the extent of identifying variation in the data. Though only 3% of surgeons move hospital in a typical year (see Table (ref)), our panel is relatively long, comprising 15 years in total. Moreover, each move changes the peer groups of all those in the hospital left and all those in the hospital joined. The median number of surgeons in a hospital-year pair is 5, hence a surgeon moving from one median sized hospital to another changes the peer groups of 10 surgeons. For this reason, 60% of surgeons' peers change in a typical year. These changes in peer groups can generate large fluctuations in average peer take-up because peer groups are small. Our Monte-Carlo experiment also demonstrates that the observed mobility rate is sufficient to accurately estimate $\rho$.
All estimated models verify generic identification of $\rho$. In the baseline model without surgeon demographics, the rank of $[\mathbf{WJ},\mathbf{WG}]$ is 2361 for the linear-in-means network, 2678 for the linear-in-others'-means network and 2683 for the persistent network. Since there are $N=1363$ surgeons, the rank requirement of Corollary (ref) is satisfied.\footnote{For these calculations, year dummies are included to construct the annihilator $\mathbf{W}$, hence the partialling out is with respect to both hospital and year dummies. For models with surgeon demographics, these are also included to construct $\mathbf{W}$.}
Table (ref) reports the results. The first row reports estimates of $\rho$. We find positive peer effects in all specifications. The peer effect is precisely estimated, and statistically distinguishable from zero in all specifications. This is in line with our Monte-Carlo results, and, though their application to peer effects in education is different to ours, with the results of arci12.
The vector $\alpha$ combines both surgeons' ability and taste for innovation, but we cannot decompose the two with the available data. Clearly, the size of the peer effect is interpretable only relative to the variance of $\alpha$. To provide an interpretable benchmark, we treat $\mathbf{G}$ as fixed and $\alpha_i$ as i.i.d., with variance $\sigma^2_\alpha$ estimated using the estimated $\alpha$, as proposed by arci12. For surgeon $i$ in year $t$, the effect on take-up of a standard deviation increase in average peer latent propensity to innovate is $\rho\sigma_\alpha\left(\sum_{i=1}^N\mathbf{G}_{ijt}^2\right)^{1/2}$. For the linear-in-means network this is $\rho\sigma_\alpha N_{g(i,t)}^{-1/2}$, and for linear-in-others'-means we replace $N_{g(i,t)}$ by $N_{g(i,t)}-1$. Since the effect of a standard deviation increase in own latent propensity to innovate is $\sigma_\alpha$, the magnitude of the peer effect relative to the own effect is $\rho\left(\sum_{i=1}^N\mathbf{G}_{ijt}^2\right)^{1/2}$. Table (ref) reports its average over all observations. With hospital fixed effects, it is around 0.1 for linear-in-means, 0.15-0.2 for linear-in-others'-means and 0.35 for the persistent network. Our estimates imply that a standard deviation increase in average peer latent propensity to innovate increases take-up by 3 percentage points for linear-in-means, 5 percentage points for linear-in-others'-means and 10 percentage points for the persistent network. Our finding that the persistent network corresponds to the largest effect size suggests that both contemporaneous and cumulative peer exposure play a role.
For each network, we also estimate a specifications which includes the age and experience of a surgeon and their peers (i.e., the model in (ref)). We find an inverted-U profile in age, with take-up estimated to peak between 45 and 49. The coefficient on experience is positive, suggesting that an additional hundred career colorectal cancer surgeries leads to around a 5 percentage point increase in take-up. The effect of peer age is close to zero, whilst peer experience plays a similar role to own experience. The estimate of $\rho$ is smaller than in models in which peer effects operate through $\alpha$ alone, though it remains positive and statistically significant at conventional levels. This suggests that some, but not all, of the peer effect operates through experience.
This paper provides and implements new identification results for panel data models with peer effects. Our results suggest that these can typically be separated from correlated effects provided either that there is sufficient mobility in the data or that the network data are sufficiently detailed. We find positive peer effects in surgeons' take-up of keyhole surgery for colorectal cancer, which operate through peer experience and latent propensity to take-up the innovation. Our results imply that exposure to others with high take-up and experience may be useful to increase innovation take-up. Based on this, policymakers might consider programmes which expose low take-up surgeons to high take-up and/or experienced surgeons. For example, one might conceive a targeted secondment programme.
{\singlespacing }