Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
69,191 characters · 18 sections · 83 citation commands
Peer Effects with Miss-specified Peer Groups
\onehalfspacing
We address two challenges faced by researchers studying peer effects, each of which can lead to miss-specification of peer groups. The first challenge is that standard methods require that the researcher has access to group data: a sample of groups which includes the outcome and characteristics of all members (see bramoulle20 for a recent review). Such data are often proprietary or restricted-access breza20, and widely available individual level survey data cannot be used. Without group data, empirical practice is either to drop individuals with missing peer data, leading to sample selection and loss of information, or to use only non-missing peers, leading to measurement error. Measurement error also arises if the researcher is unaware that peers are missing. We refer to this as the missing data problem, and propose a solution which corrects for measurement error and makes full use of the available information.
The second challenge is that the researcher has to choose the relevant peer group, often from a set of candidate group structures. For example, studies based on the Dartmouth room-mate data, in which college freshman were randomly allocated to dorm rooms, have noted that it isn't clear whether peer effects operate at the room or the floor levels sacerdote01,glaeser03,angrist14, and the relevant group may be different for different outcomes, such as academic attainment and fraternity membership sacerdote01. Empirical practice is to conduct a robustness test by re-estimating the peer effects for each candidate group. This implicitly assumes that the relevant group is the same for all individuals (i.e., it is deterministic), but there is no reason for this to be the case. In the Dartmouth context, some dorms may simply be more sociable than others, and the researcher is unlikely to know which ones. Moreover, the researcher does not know which (if any) of the estimates are valid, which can be problematic when there are large differences in the estimates. We refer to this as the group uncertainty problem, and propose a solution based on random peer group structure.
We first show that missing data and group uncertainty are examples of a wider class of miss-specification in which the members of the specified peer group are a subset of the members of the (true) group. This is clear for missing data if the specified group comprises the observed group members. For group uncertainty, it is applicable in the typical empirical setting in which the candidate groups are nested and the specified group is the smallest group. For example, in the Dartmouth context rooms are nested within floors and groups can be specified based on rooms.
Under subset-miss-specification, we show that peer effects can be identified using assumptions which allow the distribution of the group size to be inferred. With missing data, we show that widely used assumptions on the missingness mechanism allow the group size to be inferred by restricting the distribution of the group size conditional on the specified group size (i.e., on the number of observed individuals from that group). For group uncertainty, we suppose that there are two nested candidate groups (e.g., dorm rooms and floors), from which the relevant group is exogenously determined at random.\footnote{The extension to three or more candidate groups is straightforward.} This also restricts the distribution of the group size conditional on the sizes of the two candidate groups.
We apply our approach to settings with both small and medium-large group sizes. Our Monte-Carlo experiment is designed to mimic the Dartmouth room-mate data used in sacerdote01,glaeser03 and angrist14, in which students were assigned to dorm rooms of size 2, 3 or 4. Our results demonstrate that the biases arising from ignoring missing data and group uncertainty can be large, but are corrected by the GMM and NLS estimators we propose.
Our empirical application studies how peer ability impacts lawyers' decision to exit the local legal market using unique employer-employee matched data comprising all lawyers practicing in Shanghai, China. Peer groups are specified based on the number of junior lawyers practicing in the same firm, with a mean of 11.2 and a maximum of 482. We find that the likelihood of a lawyer's quit-to-exit increases in the proportion of high-ability peers, which is consistent with the invidious comparison model hoxby2005taking,antecol2016peer. We apply our results on missing data to show that similar estimates and qualitative conclusions could have been reached had we only had access to an individual level sample of lawyers, rather than the group data we use. Finally, we apply our results on group uncertainty to study whether peer effects operate at the firm or firm-cohort levels, and find evidence of considerable heterogeneity in the relevant peer group across firms.
The literature on sampled networks, mismeasured networks and missing data has recently attracted much interest. lewbel191 and hardy19 consider the implications of measurement error in the network through which peer effects operate. lewbel191 also show that peer effects can be identified and consistently estimated when there is no information on the network beyond group identifiers. boucher20 study peer effects when the network is unknown but that its distribution is known or can be consistently estimated.
chand11 studies the implications of using sampled network data in applied work, provide analytical corrections in some examples of interest and develop a more general graphical reconstruction approach. In the context of peer effects, the analytical correction can be applied if the researcher has data on the network, outcomes and exogenous characteristics for a subsample of individuals and knows the identities, outcomes and exogenous characteristics of all of their peers. This means that data are only missing for peers-of-peers. To apply graphical reconstruction, for every individual, the researcher must observe the outcomes, exogenous characteristics and variables which can be used to predict the propensity of individuals to form links in a network formation model.
breza20 use a parametric network formation model to show that network structure can be identified and used to construct network statistics (e.g., centrality measures) if the researcher does not collect network data for all individuals, but instead collects relational data for a sub-sample of individuals and basic characteristics of all individuals.\footnote{Relational data are collected by asking questions of the form `Consider all the individuals in your group with whom you do activity X. How many of these have trait Y?' breza20} sojourner13, wang13 and liu17 consider the case in which the network and all individuals are observed without error but there are missing data in the outcomes and/or exogenous characteristics.
All of the above papers suppose that the researcher observes at least some data on every individual. In contrast, we allow for some individuals to be entirely missing and do not require knowledge on whether there even exist such individuals. These gains come from exploiting the structure of the group interactions model we consider, which, though widely used in practice (e.g., moffitt01,sacerdote01,glaeser03,angrist14,boucher14,cornel17) and well understood theoretically (e.g., lee07,bramoulle09,davezies09,bramoulle20), is less general than some.\footnote{We discuss this issue further in Section (ref).}
Our work is also related to the literature on multiple peer effects. goldsmith13, arduini20 and reza21 study a setting in which individuals are exposed to multiple peer effects, each operating through a different network. In contrast, we suppose that individuals are exposed to peer effects operating through one network, but there is uncertainty as to which is the relevant network. Our model is thus not a model of multiple peer effects, but one of peer effects with uncertain peer groups. The approaches are not observationally equivalent. With multiple peer effects, the endogenous peer effects propogate through a network formed by taking the union of the links over multiple networks. In contrast, with uncertain peer groups, the endogenous effects propogate through only one network, the identity of which is unknown.
Our empirical application contributes to the literature on peer effects in competitive environments, particularly with respect to peers' ability impact on lower- and medium-achieving individuals. One strand of this literature focuses on sports competitions and tournaments and mostly find a negative impact. brown2011quitters, smith2013peers and emerson2018peer find a negative peer ability effect, in contrast to guryan2009peer who find no effect. yamane2015peer find positive effects when subjects are ahead and negative effects when they are behind. Another strand of this literature focuses on educational outcomes and reports mixed findings. Some studies find significant peer effects (for example, hoxby00 and sacerdote01 find a positive effect and antecol2016peer find negative effect), and others find small to no effects (for example, foster2006s and angrist2004does). brady2017bad find negative effects at broader level but positive effects at smaller level, suggesting that peer effects operate differently for different groups. The overall perception is that effects are not identical across different ability groups. Low-achieving groups are adversely affected by peer ability and may benefit from tracking compared to ability-mixing groups duflo2011pee,carrell2013natural,booij2017ability. Our work builds on these papers by examining career decisions in the highly competitive and incentivized labor market for junior lawyers. Our results are more in line with the literature finding peer ability hurts subjects' persistence in education and in their chosen career thiemann2022persistent, howell2021learning.
We consider a group interactions model in which the data comprise an i.i.d. sample of $G_0$ groups. The asymptotic is $G_0\to\infty$, which is analagous to the large-$N$ fixed-$T$ asymptotic used for microeconometric panel data bramoulle09. Individual $i\in\{1,2,...,N_0\}$ is in exactly one group denoted $g_0(i)\in\{1,2,...,G_0\}$. The size of $i$'s group is $n_{g_0(i)}\geq 2$. When it is not required, we omit the dependence on $i$ and simply refer to $g_0$ and $n_{g_0}$. The continuous outcome $y_{i}$ of individual $i$ is determined by
where $\alpha_{g_0}$ is group unobserved heterogeneity, $x_{i}$ is an exogenous characteristic and $\epsilon_{i}$ is an error term satisfying
where $\mathbf{x}_{g_0}=(x_{j})_{j:g_0(j)=g_0(i)}$ is $n_{g_0}\times 1$.\footnote{Since we only consider the within-transformation of the fixed effects model below, (ref) can be relaxed to $\mathbb{E}[\epsilon_{i}|n_{g_0},\mathbf{x}_{g_0}]=0$. We maintain the form of (ref) to maintain comparability with the literature.} The moment condition supposes that group size and the characteristics of the group members are exogenous with respect to $\epsilon_{i}$. However, both can be arbitrarily dependent on $\alpha_{g_0}$, which allows for selection of individuals into groups. The model includes correlated effects (captured by $\alpha_{g_0}$), endogenous peer effects ($\beta$) and contextual peer effects ($\delta$).
For simplicity of exposition, following bramoulle09 we present results for the case where there is a single characteristic $x_{i}$. All results continue to apply in the more general case in which $x_{i}$, $\gamma$ and $\delta$ are vectors. As explained in Section (ref), this model is well understood from a theoretical perspective and is pervasive in applied work. Moreover, it can be microfounded based on a game in which utility depends on one's own action and the average action of the others in the group (see blume15,bramoulle20).
To ensure that the reduced form of (ref) exists, as is standard in the literature, we suppose that $|\beta|<1$ (i.e., that the endogenous peer effect is not explosive), yielding
where $u_i$ is the reduced form error. Applying the within-group transformation, $$\widetilde{y}_{i}\triangleq y_{i}-\frac{1}{n_{g_0}}\sum_{j:g_0(j)=g_0(i)}y_{j},$$ to the reduced form yields,
Though other transformations can be used, we focus on the within-group transformation since it delivers the least information loss (see Section 3.3 of bramoulle09).
The researcher specifies the model in Section (ref) with group $g(i)\in\{1,...,G\}$ for individual $i$. The size of group $g(i)$ is $n_{g(i)}$. As with $g_0(i)$, we omit the dependence of $g(i)$ on $i$ unless it is required. The specified groups may cover only a subset $\mathcal{S}\subseteq\{1,...,N_0\}$ of individuals of cardinality $N\leq N_0$. This could arise, for example, if the available data comprise a random sample of the $N_0$ individuals. However, we maintain that every individual in $\mathcal{S}$ is in exactly one specified group. We also make use of the within-specified-group transformation, $$\overline{y}_{i}\triangleq y_{i}-\frac{1}{n_{g}}\sum_{j:g(j)=g(i)}y_{j},$$ which would be applied to eliminate specified group fixed effects. To build intuition on the implications of miss-specification for identification, we now consider two examples and derive their implications for the reduced form, which will later be used for our identification analysis.
Example 1 - Superset miss-specification. Suppose that two groups, each of size $n_0/2$, are combined to form a specified group of size $n_0$. Solving for the reduced form and applying the within-specified-group transformation yields, for individual $i$ in group 1
where,
There are three sources of miss-specification, each corresponding to a term in equation (ref). First, the researcher miss-specifies the correlated effect, yielding the term $a$. Second, by specifying only one group the researcher erroneously includes the exogenous characteristics and outcomes of those in group 2 in the structural equation determining the outcomes for group 1. This yields the term $b$. Finally, the researcher miss-specifies the group size to be $n_0$ instead of $n_0/2$. This leads the reduced form parameter to be incorrectly specified as $\pi(n_0)$ rather than $\pi(n_0/2)$, yielding the term $c$. Due to these terms, the error $\overline{v}_i$ depends on $\overline{x}_i$ and $n_0$.
Example 2 - Subset miss-specification. Now switch the roles of the groups and the specified groups in Example 1, such that the researcher divides one group into two specified groups, each of size $n_0/2$. Then we obtain, for individual $i$ in group 1,
where this time,
Terms such as $a$ and $b$ in Example 1 do not arise. This is because all members of specified group 1 have the same correlated effect as those in specified group 2, and, the sum of the outcomes and exogenous characteristics over all members of specified group 2 (respectively, specified group 1) is common to all members of specified group 1 (respectively, specified group 2). The within-specified-group transformation thus eliminates both sources of miss-specification. Miss-specification only manifests through the group size (i.e., through $c$).
In Example 1 the specified group members are a superset of the group members, and there are three sources of miss-specification. In Example 2 they are a subset, and there is one source of miss-specification.\footnote{Though we do not present such an example, the case in which we have neither a subset nor a superset is similar to the superset case.} Examples 1-2 suggest that subset miss-specification is more readily addressed because one only needs to be concerned with specification of the group size (i.e., the term $c$). This is important, because there is no clear solution to terms such as $a$ and $b$ without making strong restrictions on the distributions of $\mathbf{x}_{g_0}$ and $\alpha_{g_0}$. As we show below, the intuition that subset miss-specification is more easily addressed is not specific to Examples 1-2. We also argue below that subset miss-specification is more relevant in practice. We thus proceed under the following assumption,
and the remainder of the paper focuses on miss-specification arising due to incorrect specification of the group size.
Assumption (ref) covers the case of missing data by allowing the researcher to observe $N<N_0$ individuals with outcomes, characteristics and group identifiers. For example, in the context of education, the researcher might access a sample of students' test-scores, characteristics and classroom/school identifiers (e.g., davezies09) or have access to a large educational survey such as the student level PISA survey, which contains school and grade identifiers but includes only a subset of students.\footnote{Other examples include risky behaviours and neighbourhood effects. For the former, survey data on smoking, drinking and illicit drug use among students can be incomplete lundborg06. For the latter, the neighbourhood average outcome may be measured using survey data which includes geographic identifiers (e.g., census tract). For example, bertrand00 use the 5% public use microsample of the 1990 Census to study welfare take-up. glaeser03 use the same data for wages.}
Assumption (ref) also covers the case in which there is group uncertainty but the candidate groups are nested. It is satisfied provided that the researcher uses the smallest group structure to specify $g$. For example, glaeser03 and angrist14 consider peer groups based on dorm rooms and floors,\footnote{Similarly, peer effects in education may operate at the classroom, grade or school levels. Neighbourhood effects may operate at at the two, three or four digit postcode level.} so Assumption (ref) holds if the specified groups are rooms. In our empirical application, we postulate that peer effects among lawyers may operate either at the firm or firm-cohort levels, in which case Assumption (ref) holds if firm-cohorts are specified.
We now study identification under Assumption (ref). Our identification analysis follows bramoulle09, hence we say that the structural parameters are (point) identified if and only if they can be uniquely recovered from the reduced form parameters presented below (i.e., there is an injective relationship). Our results are thus asymptotic in nature (see manski95), and characterize whether endogenous, contextual and correlated effects can be distentangled if there is no limit to the number of groups $G_0$.
Under Assumption (ref), the reduced form for individual $i\in\mathcal{S}$ is
and taking conditional expectations yields
where $\varphi(n,\mathbf{x})\triangleq \mathbb{E}[\pi(n_{g_0})|n_{g}=n,\mathbf{x}_{g}=\mathbf{x}]$. To ease the notational burden we do not make explicit that the expectations are also conditional on $i\in\mathcal{S}$. Below, we consider only cases in which this omission is innocuous and omit the condition $i\in\mathcal{S}$ for the remainder of the paper.
Equation (ref) shows that identification prospects depend on whether there is within-specified-group variation in the exogenous characteristic, on whether $\mathbb{E}[\overline{u}_i|n_{g},\mathbf{x}_{g}]=0$, and if both of these conditions are satisfied, on whether the structural parameters can be uniquely recovered from the reduced form parameters $\varphi(n,\mathbf{x})$, which are identified for all $(\mathbf{x},n\geq 2)$ in the support of $(\mathbf{x}_{g},n_g)$.\footnote{We require $n\geq 2$ because $\varphi(1,\mathbf{x})$ is not identifiable since $n_g=1$ implies $\overline{y}_i=\overline{x}_i=0$.}
In general, even under Assumption (ref), the moment condition in (ref) does not imply $\mathbb{E}[\overline{u}_i|n_{g},\mathbf{x}_{g}]=0$. Suppose however that we can make a restriction such that it is satisfied, and also that there is within-specified-group variation in the exogenous covariate. Then the structural parameters are point identified if they can be uniquely recovered from the reduced form parameters $\varphi(n,\mathbf{x})$ and the equations
for all $(\mathbf{x},n\geq 2)$ in the support of $(\mathbf{x}_{g},n_g)$. The distribution $n_{g_0}|n_{g},\mathbf{x}_{g}$ is not observed, hence the structural parameters are not point identified unless it can be restricted.\footnote{Partial identification is possible because the probabilities $\mathbb{P}[n_{g_0}=m|n_{g},\mathbf{x}_g]$ for $m=2,3....$ are non-negative and sum to 1. We do not pursue partial identification, since, as shown below, typical empirical settings faced by researchers can lead to point identification.}
We now consider three assumptions, each of which guarantees $\mathbb{E}[\overline{u}_i|n_{g},\mathbf{x}_{g}]=0$ and restricts the distribution of $n_{g_0}|n_{g},\mathbf{x}_{g}$. Each assumption includes the standard model of peer effects with known peer groups as a special case, respectively when $n_g=n_{g_0}$ (Assumption (ref)), when $\rho=1$ (Assumption (ref)) and when $\psi\in\{0,1\}$ (Assumption (ref)).
To rule out pathological cases and ease the notational burden, we present our our results for the case in which there is within-specified-group variation in the exogenous characteristic, and that the support of the distribution of $n_{g_0}|\mathbf{x}_{g_0}$ does not depend on $\mathbf{x}_{g_0}$. The results in Propositions (ref) and (ref) continue to apply if the latter is relaxed, provided that there exists an element in the support of the exogenous characteristic such that the support condition holds for the conditional distribution of $n_{g_0}$.
We first consider identification when there are missing data. Our analysis makes uses of the indicator $s_i\in\{0,1\}$ for the condition $i\in\mathcal{S}$.
Assumption (ref) covers the simplest case in which the group size is known but individuals are sampled from the group. In practice, it means that the researcher has access to a sample of individuals with outcomes, characteristics, group membership indicators and group size. The distribution of $s_i$ (i.e., the inclusion probability) can be heterogeneous across individuals, but ought not to depend on the structural error (i.e., there should be no sample selection on unobservables). Identification under Assumption (ref) was first considered by davezies09. We include it for comparison with our approach, which allows the group size to be unknown.
Assumption (ref) relaxes the requirement that the researcher knows the group size. In practice, it allows for the case in which the researcher observes an individual level sample of outcomes, exogenous characteristics and group membership indicators. In such a sample, the researcher knows the number of individuals sampled from each group ($n_{g}$) but does not know the group size ($n_{g_0}$), nor whether there are any missing individuals in each group. Relative to Assumption (ref), we additionally assume that the sample of observed individuals is unbiased and uncorrelated. These are mild assumptions likely to be satisfied in well designed individual level samples under common sampling schemes sampling.
Assumptions (ref) and (ref) both maintain that the sampling of individuals does not induce sample selection, whilst Assumption (ref) additionally maintains that the sampling does not depend on exogenous covariates nor on the group size. Viewed from the missing data perspective, Assumption (ref) is essentially a type of Missing Completely at Random assumption, which is widely used (often implicitly) in empirical work. Together, $\mathbb{E}[s_i|\mathbf{x}_{g_0},n_{g_0}]=\rho\in(0,1]$ and $\mathbb{COV}[s_i,s_j|\mathbf{x}_{g_0},n_{g_0}]=0$ imply
which leads to identification of the distribution $n_{g_0}|n_g,\mathbf{x}_g$ from the observed distribution $n_g|n_g\geq 1,\mathbf{x}_g$.\footnote{Conditioning on $n_{g}\geq 1$ is important, because for some groups the researcher may not observe any individuals. Hence it is the distribution $n_{g}|n_{g}\geq 1,\mathbf{x}_g$ which is observed rather than $n_{g}|\mathbf{x}_g$.}
We now briefly discuss possible generalizations of Assumption (ref), and the extent to which they facilitate identification. First, one might consider $\rho=\rho_{g_0}$ to allow for group heterogeneity. Such an assumption does not provide information to identify $n_{g_0}|n_g,\mathbf{x}_g$ because $\mathbb{P}[n_{g_0}=m|\mathbf{x}_{g}]$ and $\rho_{g_0}$ only appear multiplicatively in the expression for $\mathbb{P}[n_{g}=n|n_{g}\geq 1, \mathbf{x}_{g}]$ for all $m$ and $n$. For the same reason, using $\rho=\rho(n_{g_0})$ does not provide identifying information.
Alternatively one might consider $\rho=\rho(z_i)$ for some exogenous observable $z_i$ (e.g., $z_i=x_i$). Viewed from the missing data perspective, this corresponds to a Missing At Random assumption. In this case $n_{g}|n_{g_0},\mathbf{x}_{g_0},\mathbf{z}_{g_0}$ follows a Poisson Binomial distribution, which could, in principle, be used for identification. However, if $\mathbf{z}_{g_0}$ is observed then $n_{g_0}$ is known and identification is attained instead under the weaker Assumption (ref). In contrast, the distribution $n_{g}|n_{g_0},\mathbf{x}_{g},\mathbf{z}_{g}$ does not take a known form, hence is of limited use for the identification in the typical setting in which we have no information on the missing individuals (i.e., there are some individuals for whom $z_i$ is unobserved).
We do not pursue the above generalisations formally because the focus of our analysis is not on sample selection at the individual level, but instead on replacing a group level sample with an individual level sample. Hence we consider Assumption (ref) as the benchmark case in which the researcher has access to a well designed sample of individuals rather than a well designed sample of groups. We now present our identification result.
Proposition (ref) shows that the structural parameters are identifiable if $\gamma\beta+\delta\neq 0$. This is a well known necessary identification condition in models with correctly specified groups and both endogenous and contextual effects (see for example bramoulle09 or rose17), which is generically satisfied over the parameter space. If it is violated then the endogenous and contextual effects exactly offset one another such that the reduced form effect is $\pi(n)=\gamma$, which does not vary with $n$, and hence no amount of variation in group sizes can identify $\beta$ and $\delta$. Both results in Proposition (ref) also require that the support of the distribution of $n_{g_0}$ has at least three elements. This is also necessary condition for identification when the groups are correctly specified and there are endogenous, contextual and correlated effects lee07,davezies09,bramoulle09, hence it is also necessary if the groups are (possibly) miss-specified.
Result 1 was first established by davezies09 for missing data with known group sizes whereas result 2 applies when the group sizes are unknown. Point identification of $\rho$ comes in part from information on the lower bound of the support of $n_{g_0}$. The result uses that $n_{g_0}\geq 2$, which is maintained throughout the theoretical and empirical literature (e.g., lee07,bramoulle09,davezies09,bramoulle20,moffitt01,sacerdote01,glaeser03,angrist14,boucher14,cornel17). Indeed, groups of size one cannot provide information on peer effects and are not consistent with our model. However, if individuals are missing, then we may observe groups with only one member even though all groups have at least two members (i.e., $\mathbb{P}[n_g=1|\mathbf{x}_g,n_{g}\geq 1]>0$ if $\rho<1$). Hence the distribution of $n_g$ provides information on $\rho$, which combined with (ref), leads to identification of $\rho$ and the distribution $n_{g_0}|n_g,\mathbf{x}_g$. Equation (ref) can then be used to identify the peer effects.
In some applications a stronger support restriction may be used if it is known that $n_{g_0}\geq \underline{n}$ for some known $\underline{n}>2$. For example, this is likely to be satisfied in studies of peer effects in education, where classrooms might reasonably be assumed to contain at least a handful of students. Though not required for identification, this can improve the precision of the GMM estimator we propose below. In other applications $n_{g_0}=1$ may be feasible, in which case one can never rule out $\rho=1$ and $n_{g_0}=n_g$ so that peer effects are not identifiable. For this reason, a restriction on the lower bound of the support of $n_{g_0}$ is necessary for identification of the peer effects.
Result 2 also uses the mild assumption that the distribution of $n_{g_0}$ is bounded, which is needed to identify $\rho$ and the distribution of $n_{g_0}|\mathbf{x}_g$ from the distribution of $n_g|n_g\geq 1,\mathbf{x}_g$ as a solution of a non-linear system comprising a finite number of equations. The upper bound need not be known since it is identifiable from the distribution of $n_g|n_g\geq 1$. This assumption is also used, for example, by lee07, davezies09, lewbel19 and boucher20.
We now consider subset miss-specification resulting from group uncertainty.
Assumption (ref) covers the case in which there is group uncertainty but the candidate groups are naturally nested. Though an extension to three or more is straightforward, for simplicity of exposition we do not pursue it formally. We also require a minor strengthening of the moment condition, replacing (ref) with $\mathbb{E}[\epsilon_{i}|n_{g_1},n_{g_2},\mathbf{x}_{g_2}]=0$, and hence $n_{g}$ with $n_{g_1},n_{g_2}$ in equations (ref) and (ref). The stronger moment condition requires that both potential group sizes and the exogenous characteristic of all individuals in the two potential groups be strictly exogenous with respect to the structural error.
Under Assumption (ref), the reduced form is
where $\varphi'(n_1,n_2,\mathbf{x})\triangleq \mathbb{E}[\pi(n_{g_0})|n_{g_1}=n_1,n_{g_2}=n_2,\mathbf{x}_{g_2}=\mathbf{x}]=\psi\pi(n_{1})+(1-\psi)\pi(n_{2})$, hence identification hinges on the properties of the joint distribution of $n_{g_1}$ and $n_{g_2}$, specifically, through the bi-partite graph of the support of $(n_{g_1},n_{g_2})$, which we denote by $\mathcal{G}_{n_{g_1},n_{g_2}}$. Figure (ref) depicts an example.
Our approach can be generalized to allow $\psi=\psi(\mathbf{z}_{g_2})$ (e.g., $\mathbf{z}_{g_2}=\mathbf{x}_{g_2}$), provided that the conditional expectations and probabilities in Assumption (ref) continue to hold conditional also on $\mathbf{z}_{g_2}$. We discuss the corresponding modification of Proposition (ref) at the end of this section. We now present our identification result.
The conditions $\gamma\beta+\delta\neq 0$ and that the supports of the distributions of $n_{g_1}$ and $n_{g_2}$ each have least three elements are necessary even when there is no miss-specification (i.e, when it is known that either $\psi=0$ or $\psi=1$) lee07,davezies09,bramoulle09. In addition to these, we require a mild condition on the joint support of the candidate group sizes $(n_{g_1},n_{g_2})$, expressed through the bi-partite graph $\mathcal{G}_{n_{g_1},n_{g_2}}$. When this condition is satisfied, using standard arguments for bi-partite fixed effects models (e.g., abowd99), we are able to identify $\psi\pi(n_1)+(1-\psi)\pi(n_2)$ for all $(n_1,n_2)$ in a subset of the support of $(n_{g_1},n_{g_2})$. This subset is given by all pairs which appear in the same connected component of $\mathcal{G}_{n_{g_1},n_{g_2}}$. Importantly, there is no requirement that the supports of $n_{g_1}$ and $n_{g_2}$ overlap. Figure (ref) provides such an example.
The intuition for identification of $\psi\pi(n_1)+(1-\psi)\pi(n_2)$ comes from a regression of $\overline{y}_i$ on the interaction of $\overline{x}_i$ with dummies for each value of $n_{g_1}$ and the interaction of $\overline{x}_i$ with dummies for each value of $n_{g_2}$. Together, both sets of dummies sum to two for every individual in the sample (because each individual is in exactly two candidate groups, a small one and a larger one), hence one dummy must be omitted. For this reason, we identify the pairwise sums of $\psi\pi(n_1)$ and $(1-\psi)\pi(n_2)$ for all pairs in the same connected component. If we observe a connected component containing at least three vertices corresponding to $n_{g_1}$ we can contrast $\psi\pi(n_{g_1})$ over three different values of $n_{g_1}$ holding $n_{g_2}$ constant. This yields three equations in the four unknowns $\psi,\gamma,\delta,\beta$. We obtain an additional equation by repeating the exercise for a different value of $n_{g_2}$, which leads to identification.
Under the generalization of Assumption (ref) to $\psi=\psi(\mathbf{z}_{g_2})$, Proposition 2 holds provided that the support conditions on $n_{g_1}$ and $n_{g_2}$ hold conditionally on $\mathbf{z}_{g_2}=\mathbf{z}$ for some element $\mathbf{z}$ of the support of $\mathbf{z}_{g_2}$. Similarly, one replaces $\mathcal{G}_{n_{g_1},n_{g_2}}$ by $\mathcal{G}_{n_{g_1},n_{g_2}|\mathbf{z}_{g_2}=\mathbf{z}}$. This rules out pathological cases such as $z_i=n_{g_2(i)}$.
For missing data with known group sizes, we estimate $\gamma,\delta,\beta$ by non-linear least squares based on
We focus here on the non-linear least squares estimator since it is more easily adapted to the case of unknown group sizes. We present an alternative instrumental variables estimator in the appendix. If the group sizes are unknown, we use
where $\rho\in(0,1]$, $0\leq\mathbf{q}_m\leq 1$ for $m=1,2,...$, $\mathbf{q}_1=0$, $\sum_{m=1}^\infty\mathbf{q}_m=1$, and
For the parameters $\mathbf{q}$ and $\rho$ we can use maximum likelihood based on
In practice we do not sequentially estimate the parameters by maximum likelihood and then by non-linear least squares. Instead, we estimate the parameters jointly by GMM based on the non-linear least squares moment conditions and equality to zero of the expectation of the score of the log-likelihood function based on (ref). This framework allows the researcher to use GMM standard errors (as opposed to having to adjust for a two-stage estimator or use a boostrap) and to include additional information when it is available. For example, if the researcher knows which observed groups have missing members (i.e., they observe an indicator for $n_g=n_{g_0}$), they can also use the moment $\mathbb{E}[\mathbf{1}(n_g=n_{g_0})|n_g]=\rho^{n_g}$. Similarly, one can also make a parametric assumption on the distribution of $n_{g_0}$ (i.e., use $\mathbf{q}=\mathbf{q}(\boldsymbol{\lambda})$ for a parameter $\boldsymbol{\lambda}$).
Identically to the empirical likelihood estimator of the probability mass function, the score equations yield $\widehat{\mathbf{q}}_m=0$ for all $m>\widehat{\overline{n}}$, where $\widehat{\overline{n}}$ is the sample maximum of $n_g$, hence the concentrated model is
For group uncertainty, we apply non-linear least squares based on
We tailor the design to the well known data on roommates at Dartmouth college studied by sacerdote01, glaeser03 and angrist14. The original data comprise 1589 freshman at Dartmouth college who were randomly assigned to dorm-rooms. Fifty three percent of dorm-rooms were doubles, 44 percent were triples and the remaining 3 percent were quads. The first part of our Monte-Carlo experiment supposes that only a random sample of these students is available. The second part considers the issue of the definition of the relevant peer group, which could be the room or the floor sacerdote01,glaeser03,angrist14.
We consider a design satisfying Assumption (ref). The original data comprise a complete sample of freshman (i.e., $\rho=1$). Our design varies $\rho\in\{0.3,0.5,0.7,0.9,1\}$. For $\rho=1$ we set $s_i=1$ for all $i\in\{1,...,N_0\}$ and for $\rho<1$ we set $p_i=\rho+q_i$, where $q_i\overset{i.i.d.}{\sim}\text{Uniform}(-0.1,0.1)$ and $s_i=1$ with probability $p_i$. We set $G_0$ to be the nearest integer to $M/(\rho \mathbb{E}[n_{g_0}])$, where $M$ is discussed below. For each dataset, we draw $G_0$ dorm-rooms from the distribution of dorm-room sizes, which we take to be $n_{g_0}-2\sim {\rm Binomial}(2,0.25)$ so as to match the sample mean and support of the Dartmouth data. We use a parametric distribution so as to evaluate the performance of the GMM estimator both with and without a parametric restriction on the group size distribution.
Since we draw a fixed number of dorm rooms, each with a random size, the number of students $N_0$ varies from one dataset to the next. We then draw an independent random sample of $N$ students, which includes student $i$ with probability $p_i$. Hence $N$ also varies form one dataset to the next. Allowing $G_0$ to depend on $\rho$ as above implies that the size of the observed sample is $N\approx M$ no matter the value of $\rho$. Setting $M=1600$, our design matches the sample size of the Dartmouth data. We thus consider the performance of our method had the Dartmouth data been obtained from a random sample of freshman, rather than all freshman, holding the number of observations constant as we vary $\rho$. We also consider $M=8000$ to study the performance of the method under different sample sizes.
sacerdote01 exploits random assignment of freshmen to dorms in the Dartmouth data to identify peer effects in educational attainment, measured by freshman GPA. It is argued that endogenous effects are difficult to identify due to the reflection problem (see manski93). sacerdote01 focuses instead on credibly identifying the reduced form effect of room-mate high school attainment on freshman GPA (see column 5 of Table 3 in sacerdote01). The estimated peer effect is positive and statistically significant, though smaller in magnitude than own high-school attainment. This reduced form regression has $R^2=0.19$.
To match this setting as well as possible, in our design we set the own effect to be $\gamma=1$ and the endogenous effect to be $\beta=0$, hence the reduced form effect of room-mate's high-school attainment is equal to the contextual effect, which we set to be $\delta=0.5$. We set $x_i\overset{i.i.d.}{\sim}\mathcal{N}(0,1)$, $\epsilon_i\overset{i.i.d.}{\sim}\mathcal{N}(0,\sigma^2)$ and $\alpha_{g_0}\overset{i.i.d.}{\sim}\mathcal{N}(1,\sigma^2)$ and choose $\sigma^2=2(\gamma^2+\delta^2/\mathbb{E}[n_{g_0}-1])$. For the expected dorm-room size, this choice of $\sigma^2$ corresponds to fraction 0.8 of $\mathbb{V}[y_i]$ being due to $\alpha_{g_0(i)}+\epsilon_i$ and 0.2 being due to $x_i\gamma+(1-n_{g_0})^{-1}\sum_{j:g_0(j)=g_0(i),j\neq i}x_j$, hence a population $R^2$ of 0.2 in a reduced form regression without fixed effects, as estimated by sacerdote01.
Table (ref) presents the results. We compare four models; one in which the observed groups are treated as if they are the groups (`MS', estimated by NLS), one in which the group size is known (`K', estimated by NLS), one in which the group size is unknown (`U', estimated by GMM) and a modification of U in which the parametric assumption $n_{g_0}-2\sim {\rm Binomial}(2,\omega)$ is additionally imposed (`U-P', estimated by GMM). Model MS is miss-specified when $\rho<1$. The other models are correctly specified but use information and assumptions in different ways. Model (K) is the full information benchmark with which we compare models (U) and (U-P). Models (U) and (U-P) take the upper bound on the support of the group sizes to be known as $\overline{n}=4$, which, in the context of the Dartmouth data implies that the researcher knows that there are no rooms of more than four students.
Beginning with the contextual effects only specification ($\beta=0$ is imposed) and $\rho=1$, we see that the results are very similar (identical to three decimal places) for all four models. As $\rho$ decreases, the distributions of the NLS estimator of $\delta$ and $\gamma$ of model (MS) shift closer to zero. For $\rho\leq0.5$, the bias becomes sizeable. In contrast, the estimators for the other models remain approximately unbiased for all values of $\rho$, though their root mean squared error (RMSE) increases as $\rho$ decreases. This is at least partly because, as $\rho$ decreases the number of groups with only one observed member increases, and these groups provide no information on the parameters due to the within-specified-group transformation (i.e., because $\overline{y}_i=\overline{x}_i=0$ when $n_g=1$).
The NLS estimator for model (K) performs at least as well as the GMM estimator for model (U) in terms of RMSE. This is unsurprising since it requires known group sizes in order to be implemented. For $\rho=0.9$ the difference in RMSE is small (0.086 vs 0.099 for $\delta$ when $M=8000$), though the gap widens as $\rho$ decreases. This is likely because, as $\rho$ decreases, there is less variation in the specified group size $n_g$, and hence, for a given specified group size, more uncertainty as to the group size $n_{g_0}$. The additional parametric assumption on the group size distribution used in model (U-P) makes little difference when $\rho\geq 0.7$, but for smaller values the difference in RMSE between the GMM estimators of models (U) and (U-P) becomes non-neglible.
Moving on to the specification with contextual and endogenous effects, it is clear that, though identified due to there being at least three group sizes, there is insufficient group size variation in this design to reliably separate the two, as suggested by sacerdote01. This is likely because only around 6% of rooms were quads. In particular, the RMSE on $\beta$ is large and all estimators of $\beta$ are biased upwards, whereas estimators of $\delta$ tend to be biased downwards. Nevertheless, the qualiative conclusions regarding the relative performance of the estimators of the four models are the same as for the specifications with contextual effects only.
sacerdote01, glaeser03 and angrist14 all discuss the appropriate definition of the peer group, which may either be the room, the floor or the entire dorm. We focus here on the distinction between the room and the floor, and consider a design in which Assumption (ref) holds. To match the Dartmouth data, we first draw the rooms as described above. We then draw $f_1\in\{1,2,...,5\}$ uniformly, and take rooms $\{1,2,...,f_1\}$ to comprise the first floor. We then draw another integer $f_2\in\{1,2,...,5\}$ and take rooms $\{f_1+1,...,f_1+f_2\}$ to comprise the second floor. We proceed in this way until every room has been allocated to a floor. In the Dartmouth data, the sample mean number of students per floor is close to 8. This data generating process yields an expected number of students per floor of 7.5. All other aspects of the data generating process remain unchanged and we vary $\psi\in\{0.2,0.4,0.6,0.8\}$.
Table (ref) presents the results. We compare four models; one in which the specified group is the room (`R'), one in which the specified group is the floor (`F'), one in which the researcher knows whether the group is the room or the floor (`K') and one in which the group is unknown (`U'). All are estimated by NLS. Models (R) and (F) are miss-specified because have $0<\psi<1$ in all designs. The other models are correctly specified but use information and assumptions in different ways. Model (K) is the full information benchmark with which we compare model (U).
For brevity, the discussion focuses on the contextual effects only specifications. Specifying the group to be the room (model (R)) leads to downwards bias of $\delta$. The bias and RMSE grow as $\psi$ decreases. This is because the proportion of incorrectly specified groups grows as $\psi$ decreases. For the same reason the RMSE from specifying the group to be the floor (model (F)) increases as $\psi$ increases.
The results suggest that the estimator of model (F) is unbiased even when $\psi>0$. This is coincidental. It just so happens that in this design the three sources of bias (the terms $a,b$ and $c$ discussed in Example 1 in Section (ref)) offset one another. To demonstrate this, Table (ref) modifies the design by instead setting $\alpha_{g_0(i)}\overset{i.i.d.}{\sim}\mathcal{N}(n_{g_0}^-1\sum_{j:g_0(j)=g_0(i)}x_j,\sigma^2)$. This has no impact on the estimators of models (R), (K) and (U) because all are based on room fixed effects.\footnote{For these estimators the results are numerically identical to results reported in Table (ref) because identical seeds were used to draw the data for each replication.} However, it introduces substantial bias for the estimator of model (F), which uses floor fixed effects.
Comparing the results for models (K) and (U), as expected the RMSE of the latter is larger since it does not require knowledge of the groups. Nevertheless, depending on the sample size and value of $\psi$, the NLS estimator of model (U) has small bias and sufficiently small RMSE so as to be distinguishable from zero. Considering now the specifications with endogenous and contextual effects, though weakly identified, the main qualiative conclusions regarding the relative performance of the estimators of the four models are the same as for the specifications with contextual effects only. Notice also that once endogenous effects are introduced the estimator of model (F) often has larger bias and RMSE than that of model (U). This is particularly true for $\delta$.\\
In this section, we use unique employer-employee matched data to study how lawyers' quit-to-exit responses depend on their own ability and that of their peers. The dataset is collected from the Shanghai Bar Association and includes a full history of law firm composition and lawyers' career changes from 2009-2016. Exploiting lawyers' annual registration records, we observe when they cancel legal practice in Shanghai, which we refer to as their quit-to-exit. On quitting, lawyers may change occupations, practice law in other regions, or return to education. Since legal practice is among the highest-income occupations and Shanghai is one of China's highest-paying cities, we argue that quit-to-exit in this context is likely a result of lesser persistence in a highly competitive, incentivized environment, similar to exits in the education and labor literature thiemann2022persistent, wasserman2018gender.
Our analysis focuses on lawyers' quit-to-exit in 2016 and how it is affected by their own ability and that of their peers. Throughout our analysis, we focus on associate lawyers, excluding partners and directors. This is because associates and partners/directors are unlikely to perceive one another as direct competitors. The behavior, ability and characteristics of partners and directors in the firm is accounted for by firm fixed effects. From this point onwards, we refer to associate lawyers simply as lawyers.
To measure lawyers' ability, we merge the registration records with Shanghai court judgements. In this dataset, lawyers' performance can be more reliably measured for civil litigation, so we focus on law firms and lawyers whose main business is in civil law. Our analysis is thus restricted to a sample of law firms in which most lawyers complete at least one civil case annually and lawyers who do not declare criminal law as one of the three specialized fields.\footnote{The performance in criminal law is not well measured because: 1) We only observe a small share of criminal judgements and many are restricted for confidentiality reason; 2) For criminal cases, the case fees are identical within the same case category (e.g. theft or homicide), so we cannot use case fees to measure case size as what we do for civil cases. Nevertheless, our results are robust to use number of civil and criminal cases to create the ability measure and run the analysis for all (civil and criminal) lawyers, and are robust to including all law firms in the analysis (see Table (ref)).} To measure performance in civil litigation, we use lawyers' fee-weighted caseload in the previous three years (2013-2015). Though we observe case outcomes (win or lose), we do not use this to measure ability because it is likely that high-ability lawyers take on more difficult cases.
To compute our ability measure, we perform the following steps: 1) We extract the case fee (measured in thousand yuan) for each civil case, which is the amount paid to file the case to the court and is calculated based on the disputed amount of the case; 2) We divide the case fee by the number of lawyers representing the case, according to the lawyers being listed in the court judgement; and 3) We repeat the previous two steps for each civil case and sum all fee-weighted cases undertaken from 2013-2015, and 4) We compute the annual average caseload by dividing by the number of years of practice between 2013 and 2015.\footnote{For lawyers who started practice after 2013, the annual caseload in 2013 (and 2014 if started after 2014) would be zero. To include these lawyers and make them comparable, we use the annual average caseload, rather than the three-year sum, by excluding the year(s) before the lawyer's legal practice date.}
The fee-weighted annual average caseload is a proxy for ability because it reflects past performance and income generated, hence can be used to predict one's career prospects. Generally, lawyers in China charge by case, rather than by work hours liu2006client. For civil litigation, lawyers' fees are typically charged based on the disputed amount michelson2006practice, thus high-ability lawyers are more likely to work on high fee cases, regardless of whether cases are assigned by partners or obtained by the lawyer themselves. Nevertheless, we are aware that non-litigation business is not measured (e.g., Initial Public Offerings and Mergers/Acquisitions). To this end, we replicate our analysis in small- to medium-size firms because major non-litigation business is concentrated in large firms.
The effect of interest is that of peer ability on quit-to-exit. To measure peer ability we use both the peer average fee-weighted caseload and the proportion of high-ability peers, where high-ability is defined as having a fee-weighted caseload in the top quartile (i.e., $>24.44$ in our data). We believe that the latter better captures the competitive environment through which lawyers compete for promotion to a small number of senior positions (team leaders, partners and directors).
We consider two group structures through which peer effects may operate. Group F supposes that peer effects operate among all lawyers working at the same law firm. Alternatively, we might expect stronger peer effects among similar-age lawyers, who are likely to be at a similar career stage. To account for this possibility, we split lawyers into two cohorts with a cut-off age of 35 (inclusive) for the younger cohort. Group F-C supposes that peer effects only operate through lawyers in the same cohort at the same law firm.
Our sample includes 8,448 lawyers working at 755 law firms during 2016. Table (ref) summarizes the data. As shown in the left panel, 3.6% of lawyers quit-to-exit, 45% are female, 38% have a graduate degree (Masters or PhD), the average age is around 36, average experience is around 8 years, and the average fee-weighted annual caseload is 20 thousand yuan. In the right panel of Table (ref), we present summary statistics for the 732 small- to medium-size firms, which are largely similar to the sample of all firms.
We estimate (ref) taking the dependent variable to be an indicator for quit-to-exit in 2016. Even though the dependent variable is binary, the linear model is well defined from both an econometric and economic perspective, and has desirable computational and statistical properties, especially when there are group fixed effects (see bou20).
Individual covariates comprise an indicator for being female, an indicator for having a graduate degree, age, years of experience, quadratic terms in age and experience, and fee-weighted caseload. We consider contextual effects only (i.e., $\beta=0$ is imposed) because we do not expect strategic interactions in quit-to-exit behavior. To account for within-firm dependence, we cluster standard errors by firm. All soecifications include either firm or firm-cohort fixed effects.
Our baseline results are reported in Table (ref). Models (1)-(2) respectively use group structures F and F-C. We find a positive yet statistically insignificant peer effect in fee-weighted caseload. In group F, the magnitude of the peer effect is around one third of the effect of a lawyer's own fee-weighted caseload, which is negative and statistically significant at the 0.01 level. For group F-C, the magnitude of the peer effect is similar to that of own fee-weighted caseload. Models (3)-(4) consider instead the proportion of high-ability peers in groups F and F-C respectively, for which we find larger peer effects. The coefficient of 0.0544 in model (3) is statistically significant at the 0.05 level and indicates that a one standard deviation (17.2 pp) increase in the proportion of high-ability peers leads to 26.0 percent (0.9 pp) increase in lawyers' quit-to-exit probability. We find a similar effect in model (4). In both models we find a notable gender gap, in that female lawyers are around 35.8-41.1 percent (1.3-1.5 pp) more likely to quit than male lawyers. This is consistent with previous findings that women are more likely to discontinue in competitive environments hunt2016women, wasserman2018gender. We also find evidence that quits are negatively but concavely correlated with own age and negatively correlated with own fee-weighted caseload, which is consistent with jovanovic1979job.
Next, in models (5)-(6) we estimate the impact of peer ability separately for those above and below 35. We find that the peer effect is stronger in the younger cohort. The coefficient of 0.0767 in model (5) is statistically significant at the 0.01 level and indicates that a one standard deviation (19.3 pp) increase in the proportion of high-ability peers leads to 35.9 percent (1.48 pp) increase in lawyers' quit-to-exit likelihood. In contrast, we find a smaller effect among those older than 35, which is not statistically distinguishable from zero. We also find a larger gender gap among those below 35. Overall, we find stronger quit-to-exit effects among those under 35, which is consistent with jovanovic1979job.
In Table (ref), we restrict our analysis to small- to medium- size firms to alleviate the concern that our ability measure cannot capture non-litigation performance. We limit the sample of firms to those have 50 or fewer lawyers (i.e. 732 out of 755 firms). Our results are qualitatively identical the baseline specification. Another concern is on lawyers close to the retirement age, who may intentionally reduce caseload before the retirement. To that end, we replicate our preferred models (models (3)-(4) in Table (ref)) and we exclude lawyers beyond age 55, the age at which most employees in China are eligible to claim their pension. As shown in models (1)-(2) of Table (ref), we still find similar adverse peer ability effects, indicating that retiring lawyers are not driving the estimated effects. Next, in models (3)-(4), we conduct another robustness check by including all law firms in Shanghai. The similar effects we find in these two columns suggest that caseload broadly serves as a reliable ability proxy in the legal profession context. Overall, the similar effects found in alternative samples in models (1)-(4) and in Table (ref) imply that the estimated peer ability effects generally exist in our context. Lastly, in models (5)-(6), we consider an alternative definition using the number of high-ability peers in the specific groups F and F-C. The results are also in line with our baseline findings.
In Table (ref), we conduct a heterogeneity analysis. In models (1)-(2), we examine the impact of high-ability peers on lawyers whose fee-weighted caseload is below median. Consistent with antecol2016peer and booij2017ability, these low- to medium-ability individuals drive the negative peer ability effects. In models (3)-(4), we study whether women are more sensitive to peer ability effects. Despite the significant gender gap of persistence, we find no evidence suggesting that women are more sensitive to peer ability effects. In models (5)-(6), we check whether the magnitude of peer ability effects vary by firm size by interacting the proportion of high-ability peers with an indicator for firms below the median size. The coefficient on the interaction term suggests that lawyers in small firms are not additively more prone to peer ability effects.
We now ask whether we could have obtained similar findings had only a random sample of lawyers with firm identifiers been available. To do this, we draw random 1000 random sub-samples of lawyers using the sampling process from our Monte-Carlo experiment (see Section (ref)). We focus on the baseline specification for group F (model (3) of Table (ref)), consider 50% and 70% subsamples (i.e., $\rho=0.5$ and $\rho=0.7$). The minimum value of $n_{g_0}$ in our sample is 2, which we use as a lower bound for the estimator with unknown group sizes. We make a parametric restriction on the distribution of $n_{g_0}-2$, supposing that it follows a negative binomial distribution. Figure (ref) depicts the empirical distribution for the original sample and a fitted negative binomial distribution, which provides a good approximation.
The results are reported in Figures (ref) ($\rho=0.5$) and (ref) ($\rho=0.7$). As in our Monte-Carlo experiment, we consider the estimator based on miss-specified peer group sizes (i.e., incorrectly supposing that $n_g=n_{g_0}$), known sizes and unknown sizes. Looking at the left hand columns we see that miss-specification has little consequence for estimates of the effect of a lawyer's own characteristics. However, looking at the right hand columns we see that miss-specification biases estimates of peer effects towards zero. Comparing Figure (ref) with Figure (ref), we see that the bias in in the peer effects is larger for $\rho=0.5$ than for $\rho=0.7$. Unsurprisingly, the variability of the estimators is also larger for $\rho=0.5$ than for $\rho=0.7$. These findings are qualitatively identical to those of our Monte-Carlo experiment, but use real data and larger group sizes.
Model (7) of Table (ref) allows for uncertainty as to whether F or F-C is the relevant peer group. The point estimate of $\psi$ (i.e., the probability of F-C) is 0.562, which suggests considerable heterogeneity, though it is not precisely estimated. Despite this, we find similar peer effects to the models based on groups F and F-C. We obtain similar results in the sample of small-medium firms (see Table (ref)).
Our identification results and empirical work demonstrate that it is possible to conduct empirical analysis of peer effects despite missing data and group uncertainty. Regarding missing data, we require only that the researcher has access to a sample of individuals with outcomes, exogenous characteristics and group identifiers. We propose a method which does not require information on group size, nor on individuals which are not sampled, nor on whether such individuals even exist. In principle, and subject to the limitations discussed below, this opens up the possibility of peer effects studies based on widely available individual level survey data (provided that it contains group identifiers). We also show that peer effects can be identified under group uncertainty. Future work may extend our results to incorporate both missing individuals and group uncertainty simultaneously.
A limitation of our work is that our results are specific to the group interactions model we consider. Under more general network structures (e.g., social networks), it is not the case that the specified group fixed effects reduce the problem of miss-specification to the problem of inferring the group size. Future work may seek to bridge this divide.
\singlespacing \onehalfspacing