EconBase
← Back to paper

Efficient Peer Effects Estimators with Group Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

107,696 characters · 6 sections · 81 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Efficient Peer Effects Estimators with Group Effects

abstractWe study linear peer effects models where peers interact in groups and individual's outcomes are linear in the group mean outcome and characteristics. We allow for unobserved random group effects as well as observed fixed group effects. The specification is in part motivated by the moment conditions imposed in graham_identifying_2008. We show that these moment conditions can be cast in terms of a linear random group effects model and that they lead to a class of GMM estimators with parameters generally identified as long as there is sufficient variation in group size or group types. We also show that our class of GMM estimators contains a Quasi Maximum Likelihood estimator (QMLE) for the random group effects model, as well as the Wald estimator of graham_identifying_2008 and the within estimator of lee_identification_2007 as special cases. Our identification results extend insights in graham_identifying_2008 that show how assumptions about random group effects, variation in group size and certain forms of heteroscedasticity can be used to overcome the reflection problem in identifying peer effects. Our QMLE and GMM estimators accommodate additional covariates and are valid in situations with a large but finite number of different group sizes or types. Because our estimators are general moment based procedures, using instruments other than binary group indicators in estimation is straight forward. Our QMLE estimator accommodates group level covariates in the spirit of Mundlak and Chamberlain and offers an alternative to fixed effects specifications. This model feature significantly extends the applicability of Graham's identification strategy to situations where group assignment may not be random but correlation of group level effects with peer effects can be controlled for with observable group level characteristics. Monte-Carlo simulations show that the bias of the QMLE estimator decreases with the number of groups and the variation in group size, and increases with group size. We also prove the consistency and asymptotic normality of the estimator under reasonable assumptions.

Introduction

\global\long\global\long\global\long\global\long\global\long\global\longPeer effects are of great interest to empirical researchers and policy makers. The idea that individuals are affected by their peers motivates policies that try to manipulate peer composition for better outcomes. Peer effects are often confounded by group level effects. An example are teacher effects in a class room setting. Identifying peer effects is notoriously challenging due to the reflection problem manski_identification_1993,angrist_perils_2014 as well as due to spurious peer effects originating from group level effects. Random group allocation may be one way to overcome these identification problems. With groups formed at random, a random effects specification for group level characteristics can be adopted. An alternative approach consists in postulating that, conditional on observed group level characteristics, group level effects can be viewed as randomly assigned. Regression control techniques based on observed group characteristics then lead to a similar random effects specification, but without the need to appeal to random group assignment. We propose estimators that can accommodate both scenarios.

Random group assignment plays a prominent role in the empirical peer effects literature in a number of fields including education, labor, firm, finance and development studies. Recent examples from this literature include \textcolor{black}{sacerdote_peer_2001,duflo_role_2003,zimmerman_peer_2003,stinebrickner_what_2006,kang_classroom_2007,graham_identifying_2008,guryan_peer_2009,carrell_does_2009,carrell_natural_2013,duflo_peer_2011,sojourner_identification_2013,booij_ability_2017,garlick_academic_2018,fafchamps_networks_2018,cai_interfirm_2018,frijters_heterogeneity_2019.} Assuming group effects to be independent of observed individual and group characteristics is plausible when groups are formed at random. Ignoring group effects or assuming fixed group effects lee_identification_2007 leads to consistent but less efficient estimators. Random group effects themselves have important empirical interpretations. For example, researchers in education policy often treat random class effects as unobserved teacher effects (e.g., nye_how_2004,rivkin_teachers_2005,chetty_how_2011). Absent random group assignment, the estimators we propose can accommodate observed group level effects that can come from information about group characteristics such as the training and experience of teachers, or averages of individual group member characteristics. Group level characteristics can be interpreted as parametrizations of group effects in the spirit of mundlak_pooling_1978 and chamberlain_analysis_1980. The choice between a random effects or fixed effects estimator then depends less on random group assignment but more on whether group specific effects are believed to be observable or not. In some cases there may be independent interest in the effects of group specific covariates. An example is the effect of teacher training on student performance. In such cases a random effects estimator is the preferred choice because fixed effects estimators are often unable to identify these types of group level effects.

Our analysis extends insights in graham_identifying_2008 that show how assumptions about random group effects, variation in group size and certain forms of heteroscedasticity can be used to overcome the reflection problem in identifying peer effects. We give an interpretation of the conditional variance estimator (CVE) of graham_identifying_2008 in terms of a GMM estimator based on moment conditions for the within-group variance and between-group variance. We show that the moment conditions underlying graham_identifying_2008 are the score function of a quasi maximum likelihood estimator (QMLE) for a random group effects model. The QMLE can be shown to be the best GMM estimator in the class of estimators using moment conditions for the within and between variances of outcomes individually, rather than combining them into a single moment condition as is the case for the CV estimator, or focusing only on the within variation as is the case of the CMLE of lee_identification_2007.

One limitation of the conditional variance estimator proposed by Graham is the fact that it amounts to a difference in difference identification strategy for the variances that requires groups to fall into two size categories. As shown in graham_identifying_2008 the resulting procedure takes the form of a Wald estimator for a set of binary instruments. This setting is restrictive in applications where groups may not be easily separated into two categories or where a more general set of instruments needs to be considered. The estimators that we propose are general GMM based procedures that accommodate additional covariates as well as offer flexibility in terms of the instruments and the number of moment conditions that are being used. We illustrate these points by explicitly considering moment based estimators that exploit exogenous variation in group size as well as general group level heteroscedasticity as instruments. In contrast to the CMLE of lee_identification_2007, which is a member of the class of GMM estimators we consider, our QMLE uses both the within and between variance. This leads to efficiency gains under correct specification but comes at the cost of potential miss-specification bias if the assumption of observed fixed and unobserved random group effects is incorrect. The trade-offs are similar to related results for fixed and random effects in the panel literature.

Our work is also related to the literature in spatial econometrics started by the work of cliff_spatial_1973,cliff_spatial_1981 and anselin_spatial_1988.\footnote{anselin_thirty_2010 offers a brief review of the development of spatial econometrics literature over the past thirty years.} Recently, there is a growing number of studies using spatial methods to model social network effects, e.g., lee_identification_2007, bramoulle_identification_2009, and kuersteiner_dynamic_2020. The strength of social links can be characterized by proximity in the social network space. We extend kelejian_estimation_2006 and lee_identification_2007 by considering a random group effects specification. Spatial models were traditionally estimated with maximum likelihood (ML), e.g., ord_estimation_1975. kelejian_generalized_1998,kelejian_generalized_1999 develop generalized method of moments (GMM) estimators based on linear and quadratic moments. While this paper utilizes a quasi-maximum likelihood estimation method, the score function depends on linear quadratic forms of the error terms. Properties of quadratic moment conditions were introduced by kelejian_generalized_1998,kelejian_generalized_1999 in the cross section case, and kapoor_panel_2007 and kuersteiner_dynamic_2020 in a panel setting. Moreover, kelejian_asymptotic_2001 and kelejian_specification_2010 develop a central limit theorem for linear quadratic forms, which is the basis for the asymptotic analysis in this paper.

The linear-in-means peer effect model in manski_identification_1993 is a special case of a spatial model with group-wise equal dependence, see kelejian_2sls_2002 and kelejian_estimation_2006. kelejian_2sls_2002 were the first to study the group-wise equal dependence spatial model. They show that if there is one group in a single cross section and the model has equal spatial weights, two-stage least squares (2SLS), GMM and QMLE methods all yield inconsistent estimators, although consistent estimation with 2SLS and GMM is possible for panel data. However, kelejian_estimation_2006 point out that if group fixed effects are incorporated and the panel is balanced, the estimators are inconsistent. The results in kelejian_estimation_2006 show the importance of variation in group size in identification of spatial models with blocks of equal weights. The QMLE developed in this paper and the conditional maximum likelihood estimator in lee_identification_2007 both rely on group size variation for identification although we show that identification exploiting heteroscedastic errors is also possible. Extensions include lee_specification_2010 who allow for specific social structure within each group and liu_gmm_2010 and liu_endogenous_2014 who allow for non-row normalized weight matrices. The linear spatial model has also been applied to the empirical evaluation of peer effects by lin_identifying_2010 and boucher_peers_2014. bramoulle_identification_2009 study a broader range of social interaction models and give conditions for identification.

The paper is organized as follows. In Section (ref) we consider identification of endogenous peer effects in a simple setting without covariates for the CV, CML and QML estimators. Section (ref) presents the full model that allows for covariates and general variation in group size. Section (ref) summarizes the technical conditions we impose and presents theoretical results for the QMLE. Section (ref) contains a small Monte Carlo experiment. Proofs are collected in an appendix.

Peer Effects with Random Group Effects

We start the discussion by presenting a simple model without covariates, to introduce and discuss basic features of our new quasi-maximum likelihood estimator (QMLE), and connect it to the conditional variance (CV) estimator in graham_identifying_2008 and the conditional maximum likelihood (CMLE) estimator in lee_identification_2007. The model decomposes variation in outcomes of a cross-section of individuals into idiosyncratic noise, group level random effects and correlation that is due to group level interaction. Quadratic moment conditions implied by this random effects specification lead to efficient GMM, quasi maximum likelihood, and under additional distributional assumptions, maximum likelihood estimators. Estimators based on these moment conditions include the CV estimator of graham_identifying_2008, the QMLE as well as the CMLE of lee_identification_2007 as special cases.

Let $y_{ir}$ be an observed outcome of individual $i$ in group $r$ which has $m_{r}$ members, let $\alpha_{r}$ be an unobserved group level effect and let $\epsilon_{ir}$ be unobserved individual specific characteristics. We observe data for $R$ groups as well as a categorical variable $D_{r}$ which determines group type. An example is when there are three group sizes such that $D_{r}\in\left\{ 'small','medium','large'\right\} .$ However, $D_{r}$ could be a characteristic that is not necessarily related to group size. An example is when groups are defined by classrooms of schools in urban, suburban or rural districts and $D_{r}$ is used to denote urbanicity. Classes could also be categorized by sociodemographic composition such as whether English or other languages are the native language spoken by students in the class. We allow for type-dependent heteroscedasticity. Types add flexibility to the specification by relaxing the constraints the model imposes on the relationship between group variance and group size. In some cases type specific heteroscedasticity provides identifying variation that is separate from group size variation.

The peer effects model is stated in terms of a structural equation

equation[equation omitted — 109 chars of source]

where $\bar{y}_{\left(-i\right)r}=\frac{1}{m_{r}-1}\sum_{j\neq i}^{m_{r}}y_{jr}$ is the leave-out-mean of the outcome variable. The parameter $\lambda$ captures the endogenous peer effects, see manski_identification_1993. The structural form emphasizes the decomposition of $y_{ir}$ into a social interaction term $\lambda\bar{y}_{(-i)r}$, a group level effect $\alpha_{r}$ and an idiosyncratic error term $\epsilon_{ir}.$ For example, when $y_{ir}$ is a measure of student performance and $r$ is a class-room index then $\alpha_{r}$ can be interpreted as a class-room or teacher effect while $\epsilon_{ir}$ are unobserved student characteristics for student $i$ in classroom $r.$ Cross-sectional independence of $\epsilon_{ir}$ can be justified by random group assignment such as in the application of Graham (2008). The assumptions we impose on $\epsilon_{ir}$ and $\alpha_{r}$ are in line with the random effects panel literature where group level dependence of unobservables is modeled with the common factor $\alpha_{r}.$ We leave possible generalizations of this framework to cases where $\epsilon_{ir}$ is allowed to be dependent for future work.

Following Graham (2008) who emphasizes random assignments of individuals to groups, we assume that $\alpha_{r}$ is a random effect independent of $\epsilon_{ir}.$ As shown by Graham (2008) for a slightly different model based on full rather than leave-out means, the random effects nature of the model leads to a set of quadratic moment conditions that can be exploited for identification. We expand on these ideas by showing that the implied moment conditions are related to the moment conditions of a random effects pseudo likelihood estimator. Transformations of these moments turn out to coincide with moments used by graham_identifying_2008 as well as lee_identification_2007 who considers a fixed effects version of the model. lee_identification_2007 focuses on identification of $\lambda$ based on group size variation. Here we emphasize a random effects specification where identification is driven by heterogeneity at the group level that could result from sources including but not limited to class size variation. A literature on linear instrumental variables methods gives conditions under which $\lambda$ can be identified in models that have additional exogenous covariates $Z_{r}$, e.g., angrist_perils_2014 or bramoulle_identification_2009.\footnote{The leave-out-mean $\bar{y}_{\left(-i\right)r}$ can be viewed as a special case of a Cliff-Ord-type (cliff_spatial_1973,cliff_spatial_1981) spatial lag. kelejian_generalized_1998 give an early basic condition for identification by IV.} Besides the conventional instrumental variables strategies, alternative strategies are also available, see lee_identification_2007, graham_identifying_2008 for a modified model or kuersteiner_dynamic_2020. Letting $Y_{r}=\left(y_{1r},...,y_{m_{r}r}\right)^{'}$, $\epsilon_{r}=\left(\epsilon_{1r},...,\epsilon_{m_{r}r}\right)^{'}$, $\iota_{m_{r}}=\left(1,...,1\right)^{\prime}$ and $W_{m_{r}}=\frac{1}{m_{r}-1}(\iota_{m_{r}}\iota_{m_{r}}^{\prime}-I_{m_{r}})$, the model can be written in matrix notation as

equation[equation omitted — 114 chars of source]

To isolate or identify the social interaction effect, we impose the following restrictions on unobservables.

assumptionFor $r=1,\ldots,R$ the $r$-th group is associated with a categorical variable $D_{r}\in\{1,2,...,J\}$ with $J\geqslant1$ being fixed and finite, and for each category $j\in\{1,2,...,J\}$ there is at least one group $r$ with $D_{r}=j$. For $r=1,...,R$ and $i=1,...,m_{r}$ the disturbance terms $\epsilon_{ir}$ are independently distributed across all $i$ and $r$, with $E\left[\epsilon_{ir}|D_{r},m_{r}\right]=0$ and $E\left[\epsilon_{ir}^{2}|D_{r},m_{r}\right]=\sigma_{\epsilon0,D_{r}}^{2}$ , $0<\underbar{\ensuremath{a}}_{\epsilon}\leqslant\sigma_{\epsilon0,D_{r}}^{2}\leqslant\overline{\text{\ensuremath{a}}}_{\epsilon}<\infty$ and where $\sigma_{\epsilon0,D_{r}}^{2}$ is a function only of $D_{r}$. There exists some $\eta_{\epsilon}>0$ such that $E[|\epsilon_{ir}|^{4+\eta_{\epsilon}}]<\infty$.

Note that the variance $E\left[\epsilon_{ir}^{2}|D_{r},m_{r}\right]=\sigma_{\epsilon0,D_{r}}^{2}$ has the representation $\sigma_{\epsilon0,D_{r}}^{2}=\sigma_{\epsilon0,1}^{2}1\left\{ D_{r}=1\right\} +...+\sigma_{\epsilon0,J}^{2}1\left\{ D_{r}=J\right\} $ where $\sigma_{\epsilon0,1}^{2},....,\sigma_{\epsilon0,J}^{2}$ are fixed parameters to be estimated.

assumptionFor $r=1,...,R$, the group effects $\alpha_{r}$ are independently and identically distributed, with $E\left[\alpha_{r}|D_{r},m_{r}\right]=0$ and $E\left[\alpha_{r}^{2}|D_{r},m_{r}\right]=\sigma_{\alpha0}^{2}$, where $\text{\ensuremath{0\leq}}\sigma_{\alpha0}^{2}\leq\overline{\text{\ensuremath{a}}}_{\alpha}<\infty$. There exists some $\eta_{\alpha}>0$ such that $E\left[|\alpha_{r}|^{4+\eta_{\alpha}}\right]<\infty$. Also, $\{\alpha_{r}:r=1,...,R\}$ are independent of $\{\epsilon_{ir}:i=1,...,m_{r};r=1,...,R\}$.

Assumption (ref) implies in particular that individuals do not self select into groups based on unobserved characteristics and Assumption (ref) suggests that there is no matching between group characteristics and individual characteristics. This no sorting or matching assumption can sometimes be motivated by specific empirical designs. For example, in the Project STAR experiment that graham_identifying_2008 considers, kindergarten students and teachers are randomly assigned to classrooms. This random assignment mechanism justifies interpreting $\alpha_{r}$ as the classroom or teacher effect. It also justifies assuming that $\alpha_{r}$ and $\epsilon_{ir}$ are mutually independent random variables, see Graham (2008) Assumption 1.1. Assumption (ref) allows $\epsilon_{ir}$ to be homoscedastic across all groups when $J=1$ or heteroscedastic across different categories of $D_{r}$ when $J\geqslant2$. This formulation contains the case considered by graham_identifying_2008 where $J=2$ as a special case.

Assumptions (ref) and (ref) above imply moment conditions. These moment conditions take the form of restrictions on the within and between group variance. As discussed in more detail below, these moment conditions are fundamental to the ML estimator. In particular, we show that the score of the ML estimator is a weighted average of those fundamental moment conditions.

To derive the moment conditions, define the composite error term $U_{r}=\alpha_{r}\iota_{m_{r}}+\epsilon_{r}$ where $U_{r}$ is an $m_{r}\times1$ vector with elements $u_{ir}=\alpha_{r}+\epsilon_{ir}.$ Let $\bar{u}_{r}$ and $\bar{\epsilon}_{r}$ be the mean of $u_{ir}$ and $\epsilon_{ir}$ in group $r$. Let $\ddot{U}_{r}=U_{r}-\bar{u}_{r}\iota_{m_{r}}$ be the vector of within-group deviations from the mean of $U_{r}$ and let $\ddot{Y}_{r}$ and $\ddot{\epsilon}_{r}$ be defined in a similar manner. It can be shown that $\bar{y}_{r}=\bar{u}_{r}/(1-\lambda)=(\alpha_{r}+\bar{\epsilon}_{r})/(1-\lambda)$ with $\bar{u}_{r}=\alpha_{r}+\bar{\epsilon}_{r}$, and $\ddot{Y}_{r}=\frac{m_{r}-1}{m_{r}-1+\lambda}\ddot{U}_{r}=\frac{m_{r}-1}{m_{r}-1+\lambda}\ddot{\epsilon}_{r}$. Two conditional moment conditions, one for the within-group variance, the other for the between group variance, arise from the model in ((ref)) under Assumptions (ref) and (ref). The expected value of the within-group and between-group squares of group $r$ are

equation[equation omitted — 311 chars of source]
equation[equation omitted — 251 chars of source]

where $\sigma_{\epsilon,D_{r}}^{2}=\sigma_{\epsilon,1}^{2}1\left\{ D_{r}=1\right\} +...+\sigma_{\epsilon,J}^{2}1\left\{ D_{r}=J\right\} .$

To see how these moment conditions can achieve the identification of $\lambda$ consider the case where $\sigma_{\epsilon,D_{r}}^{2}=\sigma_{\epsilon,D_{s}}^{2}$ but $m_{r}\neq m_{s}$. Then, Equation (ref) implies that

equation[equation omitted — 247 chars of source]

Alternatively consider the case where $m_{r}=m_{s}=m$ and $\sigma_{\epsilon,D_{r}}^{2}\neq\sigma_{\epsilon,D_{s}}^{2}$, then combining (ref) and (ref) gives

equation[equation omitted — 301 chars of source]

Expressions on the left hand side of both (ref) and (ref) in principle can be solved for $\lambda$ if we restrict $\lambda\in(-1,1)$ and $m_{r}\geqslant2$, as both expressions are monotonic functions of $\lambda$. Equation (ref) is a modified version of Equation (9) in graham_identifying_2008 that accounts for the leave-out-mean specification we consider. The numerator differences out the variance of $\alpha_{r}$ which is assumed constant across types. This restriction is also imposed by graham_identifying_2008 in his Assumption 1.2. In Lemma (ref) below we outline the exact conditions under which identification is possible.

The discussion above shows that under additional assumptions on $\lambda$ and group size, identification of $\lambda$ is possible through moment conditions related to within and between variance when there is variation in either group size $m_{r}$ or idiosyncratic error variance $\sigma_{\epsilon,D_{r}}^{2}$. We now formalize the discussion into Lemma (ref) below. Let the parameter vector be $\theta=\left(\lambda,\sigma_{\alpha}^{2},\sigma_{\epsilon,1}^{2},...,\sigma_{\epsilon,J}^{2}\right)^{\prime}$ and, for clarity, let the true parameter vector be denoted by $\theta_{0}=\left(\lambda_{0},\sigma_{\alpha0}^{2},\sigma_{\epsilon0,1}^{2},...,\sigma_{\epsilon0,J}^{2}\right)^{\prime}$. For identification, we further assume that group size $m_{r}\geqslant2$ and impose the following assumption on $\lambda$.

assumptionThe parameter of the endogenous peer effects $\lambda_{0}\in\Lambda$, where $\Lambda$ is a compact subset of $(-1,1)$. Assume that $\theta_{0}\in\Theta$ with $\Theta=\Lambda\times\text{\ensuremath{[0,\overline{\text{\ensuremath{a}}}_{\alpha}]\times[\underbar{\ensuremath{a}}_{\epsilon},\overline{\text{\ensuremath{a}}}_{\epsilon}]\times\ldots\times[\underbar{\ensuremath{a}}_{\epsilon},\overline{\text{\ensuremath{a}}}_{\epsilon}]}}$ compact.

The estimation procedures we propose in this paper can be implemented with the availability of a general set of valid instruments and are valid for cases where $J\geqslant1$ as long as $J$ is fixed and finite. In the simple model without covariates the available instruments are group size $m_{r}$ and categorical variable $D_{r}$. These instruments are valid if assignment to groups is random in a way that generates random variation in group size or category. Utilizing Equation (ref) and (ref), and using group size $m_{r}$ and the categorical variable $D_{r}$ as instruments yields the following conditional moment restriction $E[\chi_{r}(\theta_{0})|m_{r},D_{r}]=0$ with

equation[equation omitted — 385 chars of source]

Identification of the parameter $\theta$ is possible with variation in group size for a given category or variation in the idiosyncratic variance over categories for the same group size. This is summarized in the following lemma. The proof of the lemma is given in Appendix (ref).

lemSuppose Assumptions (ref)-(ref) hold. Then the parameter $\theta_{0}$ is identified under the following two scenarios: (i) There are two groups $r$ and $s$ such that $m_{r}\neq m_{s}$ and $D_{r}=D_{s}$, and therefore $\sigma_{\epsilon0,D_{r}}^{2}=\sigma_{\epsilon0,D_{s}}^{2}$. Then the parameter $\theta_{0}$ is identified in $\Theta.$ In particular, the moment conditions $E[\chi_{q}^{w}(\theta)|m_{q},D_{q}]=0$ and $E[\chi_{q}^{b}(\theta)|m_{q},D_{q}]=0$ for $q=r,s$ with $\chi_{r}^{w}(\theta)$ and $\chi_{r}^{b}(\theta)$ defined in ((ref)) identify $\lambda_{0}$, $\sigma_{\epsilon0,D_{r}}^{2}$,and $\sigma_{\alpha0}^{2}$. The remaining parameters $\sigma_{\epsilon0,j}^{2}$ are identified by $E(\chi_{q}^{w}(\theta)|m_{q},D_{q})=0$ for $q\neq r$ or $s.$ (ii) There are two groups $r$ and $s$, such that $m_{r}=m_{s}$ and $\sigma_{\epsilon0,D_{r}}^{2}\neq\sigma_{\epsilon0,D_{s}}^{2}$. Then the parameter $\theta_{0}$ is identified in $\Theta.$ In particular, the moment condition $E[\nu_{q}(\theta)|m_{q},D_{q}]=0$, $q=r,s$ uniquely identifies $\lambda_{0}$ and $\sigma_{\alpha0}^{2}$, where \begin{align} \nu_{q}(\theta) & =\chi_{q}^{b}(\theta)-\frac{\chi_{q}^{w}(\theta)}{m_{q}(m_{q}-1)}=(1-\lambda)^{2}\bar{y}_{q}^{2}-\sigma_{\alpha}^{2}-\frac{(m_{q}-1+\lambda)^{2}\ddot{Y}_{q}^{\prime}\ddot{Y}_{q}}{m_{q}(m_{q}-1)^{3}} \end{align} with $\chi_{q}^{w}(\theta)$ and $\chi_{q}^{b}(\theta)$ defined in ((ref)). The remaining parameters $\sigma_{\epsilon0,j}^{2}$ are identified by $E(\chi_{q}^{w}(\theta)|m_{q},D_{q})=0$.

Full identification is achieved in Scenario (i) with group size variation in at least one category. As an example, consider types that describe urbanicity such that $D_{r}=D_{s}=1$ denotes two classrooms $r$ and $s$ that are both located in an urban school but where $m_{r}\neq m_{s}$ such that the classrooms differ in size, while the remaining categories $d=2,...,J$ may have the same group sizes. In this setting $\theta_{0}$ is identified without any further constraints on the variances $\sigma_{\epsilon,j}^{2}$. If the number of distinct group sizes exceeds the number of categories $J$ then it automatically must be the case that there exist some category that is associated with at least two distinct group sizes. Note that the result holds irrespective of whether the constraint of homoscedastic errors $\sigma_{\epsilon0,D_{r}}^{2}=\sigma_{\epsilon0,D_{s}}^{2}$ is imposed on the model or not. From Scenario (i) we see that variation in group size alone can provide variation that is sufficient for identification. Furthermore, in the homoscedastic case where only a common variance parameter $\sigma_{\epsilon}^{2}$ is specified, two distinct group sizes are sufficient for identification by the result in Scenario (i). This corresponds to the identification result of the conditional maximum likelihood estimator (CMLE) in lee_identification_2007, the score function of which can be written as $\varphi(m_{r})\chi_{r}^{w}(\theta)$, where $\varphi(m_{r})$ is a function of $m_{r}$.

While variation in group size serves as the source of identification in Scenario (i), identification based on the moment condition $E[\chi_{r}(\theta)|m_{r},D_{r}]=0$ is also possible without group size variation as long as there is some other form of group heterogeneity. As is shown in the proof for Scenario (ii) of Lemma (ref), utilizing $m_{q}=m$ and $E\left[\nu_{q}(\theta)|m_{q},D_{q}\right]=0$ for $q=r,s$ yields (ref). From (ref) we see that the endogenous peer effect parameter $\lambda$ is identified if there is heteroscedasticity across groups of the same size for at least one size, and that $\lambda$ can be estimated from the sample analog of (ref). The intuition of identification in Scenario (ii) echoes that of the conditional variance (CV) estimator of graham_identifying_2008. Similar to Graham (2008), (ref) is based on the relationship between the within-group and between-group variance as captured by $\nu_{r}(\theta)$ , and can be used to construct a Wald type moment condition like in (ref) using the categorical variable as the instrument.

The above discussion focused on identification based on the moment vector $\chi_{r}(\theta)$. We next discuss the importance of these moment conditions for efficient estimation, and their relationship to the score of the Gaussian ML estimator. The optimal moment function corresponding to $\chi_{r}(\theta)$ is given by $\chi_{r}^{\ast}(\theta)=\varphi^{*}(m_{r},D_{r})\chi_{r}(\theta)$ where, focusing on the case with $J=2$ for exposition, \footnote{See our Online Appendix for details. The derivation uses Lemma (ref) and the special properties of matrices $\Omega(\theta)$, $I-\lambda W$ and $W$ described in Appendix (ref). In the Online Appendix we also give an explicit expression for the variance covariance matrix of $\chi_{r}(\delta).$}

align[align omitted — 853 chars of source]

Clearly, it follows that $E[\chi_{r}^{\ast}(\theta_{0})]=0$ by iterated expectations. We note that the moment condition in ((ref)) underlying the CV estimator is based on a linear transformation of $\varphi^{*}\left(m_{r},D_{r}\right).$ Furthermore, as we shall see in the next section, under the additional assumption that $\alpha$ and $\epsilon$ follow a Gaussian distribution, the score function of the log likelihood for group $r$ is exactly the negative of $\chi_{r}^{\ast}(\theta)$, that is \[ \partial lnL_{r}(\theta_{0})/\partial\theta=-\chi_{r}^{\ast}(\theta_{0}), \] where $\ln L_{r}(\theta)$ denotes the log likelihood function for group $r$ conditional on $\left(m_{1},...,m_{R},D_{1},...,D_{R}\right)$. From these observations we see that the matrices $\varphi^{\ast}\left(m_{r},D_{r}\right)$ can be viewed to provide the optimal weighting for the basic moment functions $\chi_{r}(\delta)$; compare also the corresponding discussion for the general model for more details.

The result that $\partial\ln L_{r}(\theta_{0})/\partial\theta=-\chi_{r}^{\ast}(\theta_{0})$ for the score function under Gaussianity establishes the asymptotic efficiency of the GMM estimator based on $E\left[\chi_{r}(\theta)|m_{r},D_{r}\right]=0$ under the assumption of Gaussian distributions for the unobservables. When the unobservables are not Gaussian then the GMM estimator has the interpretation of a quasi maximum likelihood estimator (QMLE). Similarly, in lee_identification_2007 the score function of the conditional maximum likelihood estimator (CMLE) for group $r$ is the optimal moment function corresponding to $E\left[\chi_{r}^{w}(\theta)|m_{r}\right]=0$ under the assumption of homoscedastic and normally distributed errors $\epsilon_{ir}$. While the CMLE of lee_identification_2007 is not efficient under the assumptions we postulate in this paper, it shares robustness properties of within group panel estimators in cases where the group effects are possibly correlated with covariates in the model. Under those circumstances, random effects quasi maximum likelihood estimators are generally not expected to be consistent.

Our discussion so far highlights variance as the source of identification, with variation in either size $m_{r}$ or variance of the idiosyncratic error terms $\sigma_{\epsilon,D_{r}}^{2}$ across groups as conditions. We show that variation in group size and error term variance is a source of identification in the QMLE, CVE and CMLE. In all, the CMLE utilizes how the within-group variance changes with $\lambda$ and size when error terms are homoscedastic, while the CVE exploits the relationship between the within-group variance and between-group variance in relation to $\lambda$ and size when there is either variation in group size or heteroscedasticity across groups. Our QMLE uses both pieces of information. All three estimators remain valid without covariates, and may achieve identification as long as there are at least two different group sizes in the limit in the case of homoscedasticity. This complements other results in the literature. For example, Proposition 4 in bramoulle_identification_2009 states that in the setting of lee_identification_2007, $\lambda$ is identified by instrumenting $(I-W)WY$ with $(I-W)W^{2}Z$, $(I-W)W^{3}Z$, etc., in line with the spatial literature on the estimation of Cliff-Ord type models. Their result is due to the fact that they only exploit restrictions for the conditional mean of $\epsilon$. In graham_identifying_2008 as well as in this paper additional constraints on the distribution of $\alpha$ and $\epsilon$ are imposed and shown to be useful in the identification of peer effects. Under these conditions including $Z$ offers additional sources of variation, but identification is possible with or without it.

Adding covariates is critically important in empirical applications. Consider adding the covariate matrix $Z$. This leads to two additional moment conditions $E\left[\ddot{Z}_{r}^{\prime}\ddot{U}_{r}|m_{r},D_{r}\right]=0$ and $E\left[\bar{z}_{r}^{\prime}\bar{u}_{r}|m_{r},D_{r}\right]=0$, where $\bar{z}_{r}=\iota_{m_{r}}^{\prime}Z_{r}/m_{r}$ is the group mean of $Z_{r}$ and $\ddot{Z}_{r}=Z_{r}-\iota_{m_{r}}\bar{z}_{r}$ is the deviation from group mean. Moreover, $\ddot{Y}_{r}$ and $\bar{y}_{r}$ now need to be replaced by $\ddot{Y}_{r}-\frac{m_{r}-1}{m_{r}-1+\lambda}\ddot{Z}_{r}\beta$ and $\bar{y}_{r}-\frac{\bar{z}_{r}\beta}{1-\lambda}$ respectively. The score function of the QMLE then is the same as the moment conditions of the best GMM corresponding to these two moment functions in addition to the moments $E\left[\chi_{r}(\theta)|m_{r},D_{r}\right]=0$. In the same way, in the presence of covariates and assuming homoscedasticity of $\epsilon$, Lee's CMLE estimator is based on $E\left[\ddot{Z}_{r}^{\prime}\ddot{U}_{r}\right]=0$ in addition to $E\left[\chi_{r}^{w}(\theta)|m_{r}\right]=0$ and the relative efficiency considerations discussed in this section continue to apply to the situation with covariates.

General Model

In this section we generalize the model to allow for individual characteristics, average individual characteristics of peers and group level covariates. We assume that we have access to observations on $R$ groups belonging to $J$ categories, where $1\leqslant J<\infty$ is fixed. We consider asymptotics where the number of groups $R$ tends to infinity and where the number of group sizes is finite. For the asymptotic identification of $\lambda_{0}$ and $\sigma_{\alpha0}^{2}$ this setup assumes that in the limit we observe infinitely many groups for at least two group sizes or two categories, echoing the requirement of variation in either group sizes or categories for identification in Section (ref). In designs that allow for heteroscedasticity, we also need infinitely many groups for each category $j\in\left\{ 1,...,J\right\} $ to identify the remaining variance parameters $\sigma_{\epsilon0,j}^{2}$. Let $r=1,...,R$ denote the group index, let $D_{r}$ denote the category of group $r$, and let $m_{r}$ denote the size of group $r$. The total sample size is then given by $N=\sum_{r=1}^{R}m_{r}$. Suppose further that interactions occur within each group, but not across groups, and that peer effects work through the mean outcome and mean characteristics of peers in the same group. The linear-in-means peer effects model that includes endogenous as well as exogenous peer effects then is given by

equation[equation omitted — 163 chars of source]

where $y_{ir}$ is the outcome variable of individual $i$ in group $r$, $\bar{y}_{(-i)r}=\frac{1}{m_{r}-1}\sum_{j\neq i}^{m_{r}}y_{jr}$ is the average outcome of $i$'s peers, $x_{1,ir}$ and $x_{2,ir}$ are both row vectors of predetermined characteristics of individual $i$ in group $r$, $\bar{x}_{2,(-i)r}=\frac{1}{m_{r}-1}\sum_{j\neq i}^{m_{r}}x_{2,jr}$ is a vector of average characteristics of $i$'s peers, $x_{3,r}$ is a vector of observed group characteristics. The variables in $x_{1,ir}$ and $x_{2,ir}$ can be non-overlapping, partially overlapping or totally overlapping. The error term consists of two components, the group effect $\alpha_{r}$ and the disturbance term $\epsilon_{ir}$. We treat $x_{1,ir}$, $x_{2,ir}$, $x_{3,r}$, $D_{r}$ and $m_{r}$ as non-stochastic, while noting that at the expense of more complex notation we could also think of the analysis as being conditional on these variables. In this model, peer effects work through the mean peer outcome $\bar{y}_{(-i)r}$ and mean peer characteristics $\bar{x}_{2,(-i)r}$. The two terms are also known as the leave-out-mean of $y$ and $x_{2}$, as they are means of the group leaving out oneself. In Manski's terminology, $\lambda\bar{y}_{(-i)r}$ in ((ref)) reflects endogenous peer effects, and $\bar{x}_{2,(-i)r}\beta_{3}$ is the exogenous peer effect, also referred to as contextual peer effects. The covariates $\bar{x}_{2,(-i)r}$ and $x_{3,r}$ contain group level information and can be interpreted as parametrizations of group level fixed effects in the spirit of Mundlak (1978) and Chamberlain (1980). For example, $x_{3,r}$ can contain full group averages of individual characteristics or be composed of other characteristics that only vary at the group level. The CMLE, as in the conventional panel case, cannot account for this group level information. This can be a limitation in cases where the effects of group level characteristics are of independent interest in the analysis. An example is the effects of teacher education and training on class test scores.

Let $z_{ir}=(1,x_{1,ir},\bar{x}_{2,(-i)r},x_{3,r})$ be the row vector of all exogenous variables, let $\beta=(\beta_{1},\beta_{2}^{\prime},\beta_{3}^{\prime},\beta_{4}^{\prime})^{\prime}$ be the corresponding coefficients vector, and let $k_{Z}$ denote the number of columns in $z_{ir}$. A compact form of model ((ref)) is

equation[equation omitted — 98 chars of source]

The model can be further written as a Cliff-Ord type spatial model. To see this let $I_{m}$ denote the $m-$dimensional identity matrix, let $\iota_{m}$ denote the $m-$dimensional column vector of ones, and define the weight matrix $W_{m_{r}}$ for group $r$ as $W_{m_{r}}=\frac{1}{m_{r}-1}(\iota_{m_{r}}\iota_{m_{r}}^{\prime}-I_{m_{r}})$. The off-diagonal elements of this matrix are all equal to $\frac{1}{m_{r}-1}$ and diagonal elements are 0. Let $Y_{r}=(y_{1r},...,y_{m_{r}r})^{\prime}$, $Z_{r}=(z_{1r}^{\prime},...,z_{m_{r}r}^{\prime})^{\prime}$, $\epsilon_{r}=(\epsilon_{1r},...,\epsilon_{m_{r}r})^{\prime}$, then the model for group $r$ can be expressed in matrix form as

equation[equation omitted — 84 chars of source]

where $U_{r}=\alpha_{r}\iota_{m_{r}}+\epsilon_{r}.$ Let $Y=[Y_{1}^{\prime},Y_{2}^{\prime},...,Y_{R}^{\prime}]^{\prime}$, $Z=[Z_{1}^{\prime},Z_{2}^{\prime},...,Z_{R}^{\prime}]^{\prime}$, $U=[U_{1}^{\prime},U_{2}^{\prime},...,U_{R}^{\prime}]^{\prime}$, and $W=diag_{r=1}^{R}\{W_{m_{r}}\}$ such that the model for the whole sample is given by

equation[equation omitted — 62 chars of source]

In the spatial literature $W$ is referred to as a spatial weight matrix and $WY$ as a spatial lag. In analyzing the model in ((ref)) we maintain the random effects specification detailed in Assumptions 1 and 2 of Section (ref), which imply that $\alpha_{r}\sim(0,\sigma_{\alpha}^{2})$ and $\epsilon_{ir}\sim(0,\sigma_{\epsilon,D_{r}}^{2})$, where $D_{r}\in\{1,...,J\}$ with $J\geqslant1$ fixed and finite. The specification allows for heteroscedasticity at the group level as long as there are only a finite number of different parameters. For example, we could allow for $\sigma_{\epsilon,D_{r}}^{2}$ to be different for small and large groups, or more generally for all groups of a certain size $m_{r}$. On the other hand we do not cover the case where $\sigma_{\epsilon,D_{r}}^{2}$ differs for each individual group $r,$ as this would lead to an infinite dimensional parameter space.

The parameters of interest are $\lambda,\sigma_{\alpha}^{2},\sigma_{\epsilon,1}^{2},...,\sigma_{\epsilon,J}^{2}$ and $\beta$. Their respective true values are $\lambda_{0},\sigma_{\alpha0}^{2},\sigma_{\epsilon0,1}^{2},...,\sigma_{\epsilon0,J}^{2}$ and $\beta_{0}$. In analyzing the model it will be convenient to concentrate the log-likelihood function with respect to $\beta$ for given values of $\theta=(\lambda,\sigma_{\alpha}^{2},\sigma_{\epsilon1}^{2},...,\sigma_{\epsilon J}^{2})^{\prime}$. Let $\Theta$ denote the parameter space for $\theta$ , and let $\delta=(\theta^{\prime},\beta^{\prime})^{\prime}$ denote the vector of all parameters. Under Assumptions 1 and 2 the expression for the variance covariance matrix $\Omega_{0}$ of $U$ is \[ \Omega_{0}=\Omega(\theta_{0})=\operatorname{diag}_{r=1}^{R}\Omega_{r}(\theta_{0})=\operatorname{diag}_{r=1}^{R}\{\sigma_{\epsilon0,D_{r}}^{2}I_{m_{r}}+\sigma_{\alpha0}^{2}\iota_{m_{r}}\iota_{m_{r}}^{\prime}\}. \]

To define the quasi-maximum likelihood estimator (QMLE) for the peer effects model in ((ref)) note that solving $Y$ from ((ref)) yields the reduced from:

equation[equation omitted — 81 chars of source]

If $\alpha_{r}$ and $\epsilon_{ir}$ follow normal distributions,

equation[equation omitted — 126 chars of source]

The corresponding log likelihood function is

align[align omitted — 228 chars of source]

and the corresponding QMLE is given by

equation[equation omitted — 168 chars of source]

It is convenient to concentrate out $\beta$ and to obtain the QMLE for $\theta$ first. The first order condition for $\beta$ is

equation[equation omitted — 142 chars of source]

which leads to

equation[equation omitted — 140 chars of source]

Plugging $\hat{\beta}_{N}(\theta)$ back into ((ref)) yields the following concentrated log likelihood function,

align[align omitted — 259 chars of source]

where

equation[equation omitted — 150 chars of source]

Then the QMLE for $\theta,$ $\hat{\theta}_{N}=(\hat{\lambda}_{N},\hat{\sigma}_{\alpha,N}^{2},\hat{\sigma}_{\epsilon1,N}^{2},...,\hat{\sigma}_{\epsilon J,N}^{2})^{\prime}$ is given by

equation[equation omitted — 93 chars of source]

Plugging $\hat{\theta}_{N}$ back into (ref), the QMLE for $\beta$ is

equation[equation omitted — 199 chars of source]

A formal result regarding the asymptotic identification of the model parameters is given in the next section. We next provide some intuition for that result, by extending our earlier discussion of identification for the canonical model without covariates to our model ((ref)) with covariates. Let $\left\Vert .\right\Vert $ be the Euclidean norm on $\mathbb{R^{\text{\ensuremath{k}}}}.$ Using the relationships $\ddot{U}_{r}=\frac{\left(m_{r}-1+\lambda_{0}\right)}{m_{r}-1}\ddot{Y}_{r}-\ddot{Z}_{r}\beta_{0}$ and $\bar{u}_{r}=(1-\lambda_{0})\bar{y}_{r}-\bar{z}_{r}\beta_{0}$, the moment functions related to the full model can be written as follows \[ \chi_{r}(\theta)=\left[

array[array omitted — 129 chars of source]

\right]=\left[

array[array omitted — 426 chars of source]

\right] \] where $\chi_{r}^{w}(\delta)$ and $\chi_{r}^{b}(\delta)$ summarize the restrictions on the unobservables, and are natural extensions of the moment conditions considered before in (ref) for the model without covariates. The additional moment restrictions $\chi_{r}^{zw}\left(\delta\right)$ and $\chi_{r}^{zb}\left(\delta\right)$ relate to the exogeneity of $Z_{r}$ relative to $\epsilon_{r}$ and $\alpha_{r}$. A formal asymptotic identification result will be given in the next section. Intuitively, for given $\lambda$ the last two moment conditions identify $\beta$, while the first two identify $\lambda,\sigma_{\alpha}^{2},\sigma_{\varepsilon,j}^{2},j=1,...,J$ in an analogous manner as described in the discussion of Lemma (ref) for the model without covariates.

As for the model without covariates there is a representation of the score of the log-likelihood in terms of the fundamental moment conditions. To describe the relationship between moments and the score we define the matrix

align[align omitted — 989 chars of source]

Furthermore observe that the log-likelihood function can be written as $\ln L_{N}(\delta)=-\frac{N}{2}\ln(2\pi)+\sum_{r=1}^{R}\ln L_{r}(\delta)$ where

align*[align* omitted — 235 chars of source]

is the log-likelihood function for group $r$. Then it can be shown that\footnote{See our Online Appendix for details. The derivation uses Lemma (ref) and the special properties of matrices $\Omega(\theta)$, $I-\lambda W$ and $W$ described in Appendix (ref). In the Online Appendix we also give an explicit expression for the variance covariance matrix of $\chi_{r}(\delta).$}

eqnarray*[eqnarray* omitted — 158 chars of source]

As is well known, the score of the log-likelihood function, $S(\delta)=-\sum_{r=1}^{R}\frac{\partial\ln L_{r}(\delta)}{\partial\delta}$ can be interpreted as a moment function corresponding to the moments $E\left[S(\delta_{0})\right]=-\sum_{r=1}^{R}E\left[\frac{\partial\ln L_{r}(\delta_{0})}{\partial\delta}\right]=0$. Furthermore, under a Gaussian assumption the score is an optimal moment function.\footnote{Observe that

align*[align* omitted — 432 chars of source]

in light of the information matrix equality.} From this we see that the matrices $\varphi(m_{r},D_{r})$ can be viewed to provide the optimal weighting for the basic moment functions $\chi_{r}(\delta)$. Under Gaussian assumptions the optimal GMM estimator coincides with the maximum likelihood estimator and is asymptotically efficient under the stated assumptions.

Theoretical Results

We next state our assumptions for the general model. We maintain Assumptions (ref)-(ref) on $\epsilon$, $\alpha$ and $\lambda$. In the following we add assumptions regarding the exogenous variables, and the sizes and relative magnitudes of groups in the sample. Let $\mathcal{I}_{m,j}\subset\{1,...,R\}$ be the index set of all groups in category $j$ with size equal to $m$. Thus if $r\in\mathcal{I}_{m,j}$, then $D_{r}=j$ and $m_{r}=m$. Let $R_{m,j}$ be the cardinality of $\mathcal{I}_{m,j}$, in other words $R_{m,j}$ is the number of groups in category $j$ with size equal to $m$, and let $R_{j}$ be the number of groups in category $j$, that is $R_{j}=\sum_{r=1}^{R}1(D_{r}=j)=\sum_{m=2}^{\bar{M}}R_{m,j}$, where the upper bound\textcolor{red}{ }\textcolor{black}{$\bar{M}$} on the group size is specified in the next assumption below. Furthermore let $\omega_{m,j}=R_{m,j}/R$ denote the share of groups in category $j$ with size equal to $m$, and let $\omega_{j}=R_{j}/R=\sum_{m=2}^{\bar{M}}\omega_{m,j}$ be the share of groups in category $j$. Below we maintain the following assumption regarding the group sizes and their relative magnitudes.

assumption\textcolor{black}{(a) The sample size $N$ goes to infinity; (b) The group size is bounded in the sense that there exists some positive constant $\bar{M}$ such that $2\leqslant m_{r}\leqslant\bar{M}<\infty$ }for\textcolor{black}{ $r=1,2,...,R$; (c) The limit $\omega_{m,j}^{\ast}=\lim_{N\rightarrow\infty}\omega_{m,j}$ exists and $\omega_{m,j}^{\ast}<1$ for all $2\leqslant m\leqslant\bar{M}$ and $j$, and $\omega_{j}^{\ast}=\lim_{N\rightarrow\infty}\omega_{j}=\sum_{m=2}^{\bar{M}}\omega_{m,j}^{*}>0$ for all $j$.}

The restriction that the minimal group size is 2 rules out singleton groups. A member of such a group has no peers. Assumption (ref)(b) imposes a fixed upper bound on group size. In many applications this is not a serious constraint. The assumption is more restrictive than Lee (2007) who allows for group size to grow with sample size. It is worth pointing out that increasing group sizes generally reduce the convergence rates for estimators of peer effects parameters, and as demonstrated by kelejian_2sls_2002 in some cases lead to inconsistency of these estimators.

Assumption (ref)(c) states that asymptotically, no single type-group size combination can dominate the sample by requiring that \textcolor{black}{$\omega_{m,j}^{\ast}<1$ for all $2\leqslant m\leqslant\bar{M}$ and $j$}. In addition, all types $j$ occur in the sample in an asymptotically non-negligible way because \textcolor{black}{$\omega_{j}^{\ast}>0$ for all $j$. On the other hand, we do allow that for certain combinations of $j$ and $m$ the limit }$\omega_{m,j}^{\ast}$ is zero, allowing for some group sizes of type $j$ to occur infrequently or not at all in the sample.

Observe that $N=\sum_{r=1}^{R}m_{r}=\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}mR_{m,j}$. Since group size is bounded, the number of groups $R$ goes to infinity as $N$ goes to infinity. Since $\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}R_{m,j}=R$, we have $\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}\omega_{m,j}=\sum_{j=1}^{J}\omega_{j}=1$ and thus $\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}\omega_{m,j}^{\ast}=\sum_{j=1}^{J}\omega_{j}^{\ast}=1$. Since $\omega_{j}^{\ast}>0$ \textcolor{black}{by As}sumption (ref)(c) it follows that also $R_{j}$ goes to infinity, which is needed to facilitate the consistent estimation of \textcolor{black}{$\sigma_{\epsilon,j}^{2}$}. Assumption (ref)(c) implies that the limit of the average group size is given by

align[align omitted — 213 chars of source]

Clearly $2\leqslant m^{\ast}\leqslant\bar{M}$, since $2\leqslant m_{r}\leqslant\bar{M}$.

Observe that in light of Assumptions (ref), (ref), and (ref) the parameter space $\Theta$ for $\theta=\left(\lambda,\sigma_{\alpha}^{2}\sigma_{\epsilon,1}^{2},...,\sigma_{\epsilon,J}^{2}\right)^{\prime}$ is a compact subset of the Euclidean space $\mathbb{R}^{2+J}$. Observe further that

align[align omitted — 116 chars of source]

where $I_{m_{r}}^{\ast}=I_{m_{r}}-\iota_{m_{r}}\iota_{m_{r}}^{\prime}/m_{r}$ and $J_{m_{r}}^{\ast}=\iota_{m_{r}}\iota_{m_{r}}^{\prime}/m_{r}$ are symmetric, idempotent, orthogonal, and sum to the identity matrix. Furthermore from the results in Appendix (ref) we have $|I_{m_{r}}-\lambda W_{m_{r}}|=[1+\lambda/(m_{r}-1)]^{m_{r}-1}(1-\lambda)$.\footnote{In Appendix (ref) we review additional properties of matrices of the form $pI_{m}^{\ast}+sJ_{m}^{\ast}$ , which will be used repeatedly in this paper. In particular, their multiplication is commutative. The products of such matrices are also of the form of $pI_{m}^{\ast}+sJ_{m}^{\ast}$, and $|pI_{m}^{\ast}+sJ_{m}^{\ast}|=p^{m-1}s$, $(pI_{m}^{\ast}+sJ_{m}^{\ast})^{-1}=\frac{1}{p}I_{m}^{\ast}+\frac{1}{s}J_{m}^{\ast}$.} Thus the matrix $I_{m_{r}}-\lambda W_{m_{r}}$ is nonsingular if $1+\lambda/(m_{r}-1)\neq0$ and $1-\lambda\neq0$. Assumption (ref) ensures the non-singularity of $I_{m_{r}}-\lambda W_{m_{r}}$, and hence the non-singularity of $I-\lambda W=\operatorname{diag}_{r=1}^{R}\{I_{m_{r}}-\lambda W_{m_{r}}\}$, since for $m_{r}\geqslant2$ and $\lambda<1$ we have $1+\lambda/(m_{r}-1)>0$ and $1-\lambda>0$.

Let $\bar{z}_{r}=\frac{1}{m_{r}}\iota_{m_{r}}^{\prime}Z_{r}$ be the row vector of column means of $Z_{r}$, and let $\ddot{Z}_{r}=Z_{r}-\iota_{m_{r}}\bar{z}_{r}$ be the deviations from the column means. Then $Z_{r}^{\prime}I_{m_{r}}^{\ast}Z_{r}=\ddot{Z}_{r}^{\prime}\ddot{Z}_{r}$, $Z_{r}^{\prime}J_{m_{r}}^{\ast}Z_{r}=m_{r}\bar{z}_{r}^{\prime}\bar{z}_{r}$.

assumption(a) The $N\times k_{Z}$ matrix $Z$ is non-stochastic, with $rank(Z)=k_{Z}>0$ for $N$ sufficiently large. The elements of $Z$ are uniformly bounded in absolute value. (b)For $2\leqslant m\leqslant\bar{M}$, and $1\leqslant j\leqslant J$ the following limits exist: \[ \lim_{N\rightarrow\infty}N^{-1}\sum_{r\in\mathcal{I}_{m,j}}\ddot{Z}_{r}^{\prime}\ddot{Z}_{r}=\ddot{\varkappa}_{m,j}, \] \[ \lim_{N\rightarrow\infty}N^{-1}\sum_{r\in\mathcal{I}_{m,j}}m\bar{z}_{r}^{\prime}\bar{z}_{r}=\bar{\varkappa}_{m,j}, \] \[ \lim_{N\rightarrow\infty}N^{-1}\sum_{r\in\mathcal{I}_{m,j}}\bar{z}_{r}=\bar{z}_{m,j}. \] (c) For at least one pair of $(m,j)$ such that $\omega_{m,j}^{\ast}>0$, and N sufficiently large, the smallest eigenvalues of $N^{-1}\sum_{r\in\mathcal{I}_{m,j}}Z_{r}^{\prime}Z_{r}=N^{-1}\sum_{r\in\mathcal{I}_{m,j}}\ddot{Z}_{r}^{\prime}\ddot{Z}_{r}+N^{-1}\sum_{r\in\mathcal{I}_{m,j}}m\bar{z}_{r}^{\prime}\bar{z}_{r}$ are bounded away from zero, uniformly in N, by some finite constant $\underline{\xi}_{Z}>0$.

Suppose we have some $N\times N$ matrix $A_{N}(\theta)=diag_{r=1}^{R}\{p(m_{r},D_{r},\theta)I_{m_{r}}^{\ast}+s(m_{r},D_{r},\theta)J_{m_{r}}^{\ast}\}$, where $p(m_{r},D_{r},\theta)$ and $s(m_{r},D_{r},\theta)$ are positive, uniformly continuous and bounded on $\Theta$. An example of an expression of this form is $\Omega(\theta)^{-1}$ which is obtained in closed form in Equation (ref) in Appendix (ref). Then under Assumption (ref)(b), the limiting matrix of $N^{-1}Z^{\prime}A_{N}(\theta)Z$ always exists, is continuous in $\theta$ and takes the form \[ \lim_{N\rightarrow\infty}\frac{1}{N}Z^{\prime}A_{N}(\theta)Z=\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}[p(m,j,\theta)\ddot{\varkappa}_{m,j}+s(m,j,\theta)\bar{\varkappa}_{m,j}]. \] Furthermore, $N^{-1}Z^{\prime}A_{N}(\theta)Z$ converges to its limiting matrix uniformly on $\Theta$. With $p(m_{r},D_{r},\theta)>0$ and $s(m_{r},D_{r},\theta)>0$, Assumption (ref)(a) ensures that $N^{-1}Z^{\prime}A_{N}(\theta)Z$ and its limiting matrix are invertible, with the elements of the inverse matrix uniformly bounded in absolute value. In the special case when $A_{N}(\theta)$ is the identity matrix, $lim_{N\rightarrow\infty}\frac{1}{N}Z^{\prime}Z=\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}[\ddot{\varkappa}_{m,j}+\bar{\varkappa}_{m,j}]$, which has the smallest eigenvalue bounded above zero by some finite constant $\text{\ensuremath{\underbar{\ensuremath{\xi}}}}_{Z}>0$. See Lemma (ref) for details and a proof.

As shown by Lemma (ref) in Section (ref), identification of $\lambda$ and $\sigma_{\alpha}^{2}$ requires variation in the group size or variance of the error terms. The following assumption ensures this so that in the limit we have non-negligible samples for at least two different group sizes or two different categories with different variances of the idiosyncratic errors $\epsilon_{ir}$.

assumptionFor some sizes $m$ and $m^{\prime}$, and some categories $j$ and $j^{\prime}$ we have $\omega_{m,j}^{\ast}>0$ and $\omega_{m^{\prime},j^{\prime}}^{\ast}>0$, and either of the following two scenarios hold, (a) $m\neq m^{\prime}$, and $\sigma_{\epsilon0,j}^{2}=\sigma_{\epsilon0,j^{\prime}}^{2}$ for some $j,j^{\prime}\in\left\{ 1,...,J\right\} $. (b) $m=m^{\prime}$, and $\sigma_{\epsilon0,j}^{2}\neq\sigma_{\epsilon0,j^{\prime}}^{2}$ for some $j,j^{\prime}\in\left\{ 1,...,J\right\} $ with $j\neq j^{\prime}.$

The conditions in Assumption (ref) are the asymptotic analogs of identification conditions imposed in Lemma (ref). Assumption (ref) by itself is not sufficient for identification because it only implies that no single pair $\left(m,j\right)$ asymptotically dominates the sample. Assumption (ref) alone does not guarantee that there is enough variation in the underlying group sizes $m$ or the variances $\sigma_{\epsilon0,j}^{2}$. For example, it is possible under Assumption (ref) that all groups are of the same size and that all variances $\sigma_{\epsilon0,j}^{2}$ are the same. Assumption (ref) rules out such cases. Assumption (ref)(a) is related to Assumption 6.1 and Footnote 9 of lee_identification_2007 which requires group size variation to achieve identification for the case where group sizes are bounded, the only case we consider. Assumption (ref)(b) has no analog in Lee (2007) because of his Assumption 1 which imposes homoscedasticity on the errors $\epsilon_{ir}.$ We show that identification is possible purely based on group level heteroscedasticity even if all group sizes are the same. This insight also extends the analysis of graham_identifying_2008 where types and class sizes are linked.

Below we give results on the consistency and asymptotic normality of the QMLE $\hat{\delta}_{N}=(\hat{\theta}_{N}^{\prime},\hat{\beta}_{N}^{\prime})^{\prime}$ defined in ((ref)).

thmSuppose Assumptions 1-6 hold, then (a) The parameter $\delta_{0}$ is asymptotically identified in the sense that it is the unique maximizer of the criterion $\bar{R}(\theta,\beta)=\lim_{N\rightarrow\infty}E\left[\frac{1}{N}\textrm{ln}L(\theta,\beta)\right]$. (b) The QMLE $\widehat{\delta}_{N}$ is consistent, i.e., $\widehat{\delta}_{N}\overset{p}{\rightarrow}\delta_{0}$ as $N\rightarrow\infty$.

A detailed proof of the theorem is given in Appendices (ref) and (ref). As can be seen from the proof, the argumentation that ensures part (a) of the theorem is analogous to the argumentation used in establishing Lemma 2.1. Here is a sketch of the proof to provide some intuition. The limiting expected value of the concentrated log likelihood function $Q_{N}(\theta)$ is \[ \bar{Q}^{\ast}(\theta)=C^{\ast}+\frac{1}{2m^{\ast}}\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}\omega_{m,j}^{\ast}g(m,j,\theta)+Q^{(2)\ast}(\theta), \] where $C^{\ast}$ is a constant term, $g(m,j,\theta)=\ln|G(m,j,\theta)|-\operatorname{tr} G(m,j,\theta)$ with \[ G(m,j,\theta)=\frac{\sigma_{\epsilon0,j}^{2}}{\sigma_{\epsilon,j}^{2}}\left(\frac{m-1+\lambda}{m-1+\lambda_{0}}\right)^{2}I_{m}^{\ast}+\frac{(\sigma_{\epsilon0,j}^{2}+m\sigma_{\alpha0}^{2})}{(\sigma_{\epsilon,j}^{2}+m\sigma_{\alpha}^{2})}\left(\frac{1-\lambda}{1-\lambda_{0}}\right)^{2}J_{m}^{\ast}, \] and $Q^{(2)\ast}(\theta)=\lim_{N\rightarrow\infty}\bar{Q}_{N}^{(2)}(\theta)$ where $\bar{Q}_{N}^{(2)}(\theta)=-\frac{1}{2N}\tilde{\eta}_{Z}(\theta)^{\prime}\tilde{M}_{Z}(\theta)\tilde{\eta}_{Z}(\theta)$ with $\tilde{M}_{Z}(\theta)=I-\Omega(\theta)^{-1/2}Z(Z^{\prime}\Omega(\theta)^{-1}Z)^{-1}Z^{\prime}\Omega(\theta)^{-1/2}$ and \[ \tilde{\eta}_{Z}(\theta)=\Omega(\theta)^{-1/2}(I-\lambda W)(I-\lambda_{0}W)^{-1}Z\beta_{0}. \] It is easy to see that $\theta_{0}$ is a global maximizer of $Q^{(2)\ast}(\theta)$, given that $-\bar{Q}_{N}^{(2)}\left(\theta\right)$ is the quadratic form of an idempotent and thus positive semi-definite matrix, and $Q^{(2)\ast}(\theta_{0})=0$. However, this does not ensure that $\theta_{0}$ is a unique global maximizer. Identification thus comes from $\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}\omega_{m,j}^{\ast}g(m,j,\theta)$. Note that for any symmetric positive definite $m\times m$ matrix $A$, $\ln|A|-\operatorname{tr}(A)\leqslant-m$ with equality if and only if $A$ is an identity matrix.\footnote{To see this, note that under the maintained assumptions the eigenvalues of $A$, say, $\lambda_{i}$, are positive and $\ln\left\vert A\right\vert -\operatorname{tr}\left[A\right]=\sum_{i=1}^{m}\left[\ln(\lambda_{i})-\lambda_{i}\right]$. The claim is seen to hold by observing that the function $f(x)=\ln(x)-x\leq-1$ for $x\in(0,\infty)$ with a unique maximum at $x=1$, and observing that $A=I_{m}$ if and only if $\lambda_{i}=1$ for $i=1,\ldots,m$.} For any $m$ and $j$, $g(m,j,\theta)$ is maximized if and only if $G(m,j,\theta)=I_{m}$, which is equivalent to $E\left[\chi_{r}(\theta)|m_{r}=m,D_{r}=j\right]=0$ with $\chi_{r}(\theta)=(\chi_{r}^{w}(\theta),\chi_{r}^{b}(\theta))$ defined in (ref). It now follows from an asymptotic analogue of Lemma (ref) that in either case (i) or (ii) of Assumption (ref), $\theta_{0}$ is the only solution to $E\left[\chi_{r}(\theta)|m_{r}=m,D_{r}=j\right]=0$ and $E\left[\chi_{r}(\theta)|m_{r}=m^{\prime},D_{r}=j^{\prime}\right]=0$. Thus for any $\theta\neq\theta_{0}$, \[ \min\left(g(m,j,\theta_{0})-g(m,j,\theta),g(m^{\prime},j^{\prime},\theta_{0})-g(m^{\prime},j^{\prime},\theta)\right)>0. \] As a result, $\theta_{0}$ is the unique global maximizer of $\bar{Q}^{\ast}(\theta)$ when one of the two scenarios holds true for some $\omega_{m,j}^{\ast}>0$ and $\omega_{m^{\prime},j^{\prime}}^{\ast}>0$.

To study the asymptotic distribution of the estimator, first note that under Assumptions (ref) and (ref), the third and fourth moments of $\epsilon_{ir}$ and $\alpha_{r}$ exist. Let $E\left[\epsilon_{ir}^{3}|D_{r}=j\right]=\mu_{\epsilon0,j}^{(3)}$, $E\left[\epsilon_{ir}^{4}|D_{r}=j\right]=\mu_{\epsilon0,j}^{(4)}$, $E\left[\alpha_{r}^{3}\right]=\mu_{\alpha0}^{(3)}$ and $E\left[\alpha_{r}^{4}\right]=\mu_{\alpha0}^{(4)}$. Also, define $\Gamma_{0}$ and $\Upsilon_{0}$ as

eqnarray*[eqnarray* omitted — 337 chars of source]

As shown in Appendix (ref), the two limiting matrices exist. Specific expressions are given in Appendix (ref). When $\epsilon_{ir}$ and $\alpha_{r}$ both follow normal distributions, $\Upsilon_{0}=\Gamma_{0}$.

The next lemma shows that $\Gamma_{0}$ is p.d. under the maintained assumptions. The lemma also provides a sufficient condition on the moments of $\varepsilon$ under which $\Upsilon_{0}$ is p.d..

lemSuppose Assumptions (ref)-(ref) hold, then $\Gamma_{0}$ is positive definite. Under the additional assumption that $\mu_{\varepsilon0,j}^{(4)}-\sigma_{\varepsilon0,j}^{4}>(\mu_{\epsilon0,j}^{(3)})^{2}/\sigma_{\varepsilon0,j}^{2}$ for all $j\in\{1,...,J\}$, $\Upsilon_{0}$ is also positive definite.

The proof of the lemma is in Appendix (ref). Note that from Holder's inequality we have $\mu_{\varepsilon0,j}^{(4)}-\sigma_{\varepsilon0,j}^{4}\geqslant(\mu_{\epsilon0,j}^{(3)})^{2}/\sigma_{\varepsilon0,j}^{2}$. The sufficient condition is mild in that it only postulates that the inequality holds strongly. Of course, the condition holds, e.g., for the Gaussian distribution.

With both $\Upsilon_{0}$ and $\Gamma_{0}$ ensured to be positive definite, we have the following theorem.

thmUnder Assumptions (ref)-(ref), and assuming that $\delta_{0}$ is in the interior of the parameter space $\Theta$ defined in Assumption (ref) and that $\mu_{\varepsilon0,j}^{(4)}-\sigma_{\varepsilon0,j}^{4}>(\mu_{\epsilon0,j}^{(3)})^{2}/\sigma_{\varepsilon0,j}^{2}$ for $j\in\{1,...,J\}$, we have $\sqrt{N}(\hat{\delta}_{N}-\delta_{0})\xrightarrow{d}N(0,\Gamma_{0}^{-1}\Upsilon_{0}\Gamma_{0}^{-1})$ as $N\rightarrow\infty$.

The proof of the theorem is given in Appendix (ref). We next discuss consistent estimators for the matrices $\Gamma_{0}$ and $\Upsilon_{0}$ composing the asymptotic variance covariance matrix. An inspection shows that $\Gamma_{0}=\Gamma(\delta_{0},s_{0})$ and $\Upsilon_{0}=\Upsilon(\delta_{0},\mu_{\alpha0}^{(3)},\mu_{\alpha0}^{(4)},\mu_{\epsilon0,1}^{(3)},...,\mu_{\epsilon0,J}^{(3)},\mu_{\epsilon0}^{(4)}...,\mu_{\epsilon0,J}^{(4)},s_{0})$, with $s_{0}=(s_{0,1},...,s_{0,J},m^{\ast})$ and \[ s_{0,j}=[\ddot{\varkappa}_{2,j},...\ddot{\varkappa}_{\bar{M},j},\bar{\varkappa}_{2,j},...,\bar{\varkappa}_{\bar{M},j},\bar{z}_{2,j},...,\bar{z}_{\bar{M},j},\omega_{2,j}^{\ast},...,\omega_{\bar{M},j}^{\ast}], \] and where the functions $\Gamma(.)$ and $\Upsilon(.)$ are continuous. Since the functions $\Gamma(.)$ and $\Upsilon(.)$ are continuous, consistent estimators for $\Gamma_{0}$ and $\Upsilon_{0}$ can be readily obtained by replacing the arguments of those functions by consistent estimators thereof. Let $\hat{s}_{N}$ be the sample analogue of $s_{0}$, then clearly $\hat{s}_{N}\overset{p}{\rightarrow}s_{0}$ in light of Assumptions (ref) and (ref). Recall further that by Theorem (ref) the QMLE estimator $\hat{\delta}_{N}$ is consistent for $\delta_{0}$, and suppose we have consistent estimators for $\mu_{\alpha0}^{(3)}$, $\mu_{\alpha0}^{(4)}$, $\mu_{\epsilon0,1}^{(3)},...,\mu_{\epsilon0,J}^{(3)}$, and $\mu_{\epsilon0,1}^{(4)}...,\mu_{\epsilon0,J}^{(4)}$, denoted as $\hat{\mu}_{\alpha}^{(3)},\hat{\mu}_{\alpha}^{(4)},\hat{\mu}_{\epsilon,1}^{(3)},...\hat{\mu}_{\epsilon,J}^{(3)},\hat{\mu}_{\epsilon,1}^{(4)},...\hat{\mu}_{\epsilon,J}^{(4)}$. Now define $\hat{\Gamma}_{N}$ and $\hat{\Upsilon}_{N}$ as

align[align omitted — 348 chars of source]

then it follows from Slutsky's theorem that $\hat{\Gamma}_{N}$ and $\hat{\Upsilon}_{N}$ are consistent estimators for $\Gamma_{0}$ and $\Upsilon_{0}$. A consistent estimator for the variance covariance matrix of the limiting distribution is given by $\hat{\Gamma}_{N}^{-1}\hat{\Upsilon}_{N}\hat{\Gamma}_{N}^{-1}$.

The above discussion assumed the availability of consistent estimators for the third and fourth moment of the error components. In the following we now define consistent estimators for $\mu_{\alpha0}^{(3)}$, $\mu_{\alpha0}^{(4)}$ and $\mu_{\epsilon0,j}^{(3)}$, $\mu_{\epsilon0,j}^{(4)}$, $j=1,...,J$. To motivate the estimators consider the composite error term for individual $i$ in group $r$, $u_{ir}=\alpha_{r}+\epsilon_{ir}$, and let $\bar{u}_{r}=\frac{1}{m_{r}}\sum_{i=1}^{m_{r}}u_{ir}$, and $\ddot{u}_{ir}=u_{ir}-\bar{u}_{r}$ . Then $\bar{u}_{r}=\alpha_{r}+\bar{\epsilon}_{r}$ and $\ddot{u}_{ir}=\epsilon_{ir}-\bar{\epsilon}_{r}$, where $\bar{\epsilon}_{r}$ is the group mean of $\epsilon_{ir}$. It is readily verified that under Assumptions (ref) and (ref), we have \[ E\left[\ddot{u}_{ir}^{3}\right]=(1-\frac{3}{m_{r}}+\frac{2}{m_{r}^{2}})\mu_{\epsilon0,D_{r}}^{(3)}, \]

\[ E\left[\ddot{u}_{ir}^{2}\bar{u}_{r}\right]=\frac{(m_{r}-1)}{m_{r}^{2}}\mu_{\epsilon0,D_{r}}^{(3)}, \]

\[ E\left[\bar{u}_{r}^{3}\right]=\mu_{\alpha0}^{(3)}+\frac{\mu_{\epsilon0,D_{r}}^{(3)}}{m_{r}^{2}}, \]

\[ E\left[\ddot{u}_{ir}^{4}\right]=\frac{m_{r}^{3}-4m_{r}^{2}+6m_{r}-3}{m_{r}^{3}}\mu_{\epsilon0,D_{r}}^{(4)}+\frac{3(m_{r}-1)(2m_{r}-3)}{m_{r}^{3}}\sigma_{\epsilon0,D_{r}}^{4}, \]

\[ E\left[\bar{u}_{r}^{4}\right]=\mu_{\alpha0}^{(4)}+\frac{1}{m_{r}^{3}}\mu_{\epsilon0,D_{r}}^{(4)}+3\frac{m_{r}-1}{m_{r}^{3}}\sigma_{\epsilon0,D_{r}}^{4}+\frac{6}{m_{r}}\sigma_{\alpha0}^{2}\sigma_{\epsilon0,D_{r}}^{2}. \] Next define for group $r$ ,

\[ f_{\epsilon,r}^{(3)}=

cases\frac{1}{m_{r}}\sum_{i=1}^{m_{r}}\ddot{u}_{ir}^{3}/(1-\frac{3}{m_{r}}+\frac{2}{m_{r}^{2}}) & m_{r}\geqslant3\\ \frac{1}{m_{r}}\sum_{i=1}^{m_{r}}\ddot{u}_{ir}^{2}\bar{u}_{r}/(\frac{1}{m_{r}}-\frac{1}{m_{r}^{2}}) & m_{r}=2

, \]

\[ f_{\alpha,r}^{(3)}=\bar{u}_{r}^{3}-f_{\epsilon,r}^{(3)}/m_{r}^{2}\text{,} \] \[ f_{\epsilon,r}^{(4)}=\frac{m_{r}^{3}}{m_{r}^{3}-4m_{r}^{2}+6m_{r}-3}\text{\text{[(\ensuremath{\frac{1}{m_{r}}\sum_{i=1}^{m_{r}}\ddot{u}_{ir}^{4}})-\ensuremath{\frac{3(m_{r}-1)(2m_{r}-3)}{m_{r}^{3}}\sigma_{\epsilon0,D_{r}}^{4}}]}, } \]

\[ f_{\alpha,r}^{(4)}=\bar{u}_{r}^{4}-f_{\epsilon,r}^{(4)}/m_{r}^{3}-\frac{3(m_{r}-1)}{m_{r}^{3}}\sigma_{\epsilon0,D_{r}}^{4}-\frac{6}{m_{r}}\sigma_{\alpha0}^{2}\sigma_{\epsilon0,D_{r}}^{2}. \] Then $E\left[f_{\epsilon,r}^{(3)}\right]=\mu_{\epsilon0,D_{r}}^{(3)}$, $E\left[f_{\alpha,r}^{(3)}\right]=\mu_{\alpha0}^{(3)}$, $E\left[f_{\epsilon,r}^{(4)}\right]=\mu_{\epsilon0,D_{r}}^{(4)}$, $E\left[f_{\alpha,r}^{(4)}\right]=\mu_{\alpha0}^{(4)}$. By Lemma (ref)(a), $\frac{1}{R_{j}}\sum_{r=1}^{R}1(D_{r}=j)f_{\epsilon,r}^{(l)}\overset{p}{\rightarrow}\mu_{\epsilon0,j}^{(l)}$ and $\frac{1}{R}\sum_{r=1}^{R}f_{\alpha,r}^{(l)}\overset{p}{\rightarrow}\mu_{\alpha0}^{(l)}$ for $l=3,4$ and $j=1,...,J$ as $R$ goes to infinity.

To construct feasible counterparts of these estimates, consider the estimated disturbances $\hat{u}_{ir}=y_{ir}-\hat{\lambda}\bar{y}_{(-i)r}-z_{ir}\hat{\beta}$, where $\hat{\lambda}$ and $\hat{\beta}$ denote the QML estimators, and let $\hat{\bar{u}}_{r}=\frac{1}{m_{r}}\sum_{i=1}^{m_{r}}\hat{u}_{ir}$ and $\hat{\ddot{u}}_{ir}=\hat{u}_{ir}-\hat{\bar{u}}_{ir}$. Feasible counterparts, say, $\hat{f}_{\epsilon,r}^{(3)}$, $\hat{f}_{\alpha,r}^{(3)}$, $\hat{f}_{\epsilon,r}^{(4)}$, $\hat{f}_{\alpha,r}^{(4)}$ of $f_{\epsilon,r}^{(3)}$, $f_{\alpha,r}^{(3)}$, $f_{\epsilon,r}^{(4)}$, $f_{\alpha,r}^{(4)}$ can now be defined by replacing $\bar{u}_{r}$ and $\ddot{u}_{ir}$ with $\hat{\bar{u}}_{r}$ and $\hat{\ddot{u}}_{ir}$, and $\sigma_{\alpha0}^{2}$ and $\sigma_{\epsilon0,j}^{2}$ with their QML estimators. Now consider the following estimators for the third and fourth moments of the error components: $\hat{\mu}_{\alpha}^{(3)}=\sum_{r=1}^{R}\hat{f}_{\alpha,r}^{(3)}/R$, $\hat{\mu}_{\alpha}^{(4)}=\sum_{r=1}^{R}\hat{f}_{\alpha,r}^{(4)}/R$, $\hat{\mu}_{\epsilon,j}^{(3)}=\sum_{r=1}^{R}1(D_{r}=j)\hat{f}_{\epsilon,r}^{(3)}/R_{j}$, $\hat{\mu}_{\epsilon,j}^{(4)}=\sum_{r=1}^{R}1(D_{r}=j)\hat{f}_{\epsilon,r}^{(4)}/R_{j}$, $j=1,...,J$.

The next theorem establishes that valid inference based on standardized statistics is possible. At the core of this result is the fact that $\hat{\Gamma}_{N}\overset{p}{\rightarrow}\Gamma_{0}$ and $\hat{\Upsilon}_{N}\overset{p}{\rightarrow}\Upsilon_{0}$ as shown in Appendix (ref).

thmUnder Assumptions (ref)-(ref), and assuming that $\mu_{\varepsilon0,j}^{(4)}-\sigma_{\varepsilon0,j}^{4}>(\mu_{\epsilon0,j}^{(3)})^{2}/\sigma_{\varepsilon0,j}^{2}$ for $j\in\{1,...,J\}$, and $\hat{\Gamma}_{N}$, $\hat{\Upsilon}_{N}$ defined in ((ref)) and ((ref)) we have $\sqrt{N}\left(\hat{\Gamma}_{N}^{-1}\hat{\Upsilon}_{N}\hat{\Gamma}_{N}^{-1}\right)^{-1/2}(\hat{\delta}_{N}-\delta_{0})\xrightarrow{d}N(0,I)$ as $N\rightarrow\infty$.

The proof of the theorem is in Appendix (ref).

Monte Carlo Results

We conduct Monte-Carlo (MC) experiments to assess the finite sample properties of the quasi-maximum likelihood (QML) estimator $\hat{\delta}_{N}$. The data generating mechanism is determined by the main model in (ref). For simplicity, $x_{1,ir}$, $x_{2,ir}$ and $x_{3,ir}$ each only includes a scalar variable. We set the true value of the parameters to $\lambda_{0}=0.5$, $\sigma_{\alpha0}^{2}=0.25$, $\beta_{10}=1$, $\beta_{20}=1$,$\beta_{30}=1$, and $\beta_{40}=1$, while $\sigma_{\epsilon0}^{2}=1$ in the case of homoscedasticity. The model for the data generating process (DGP) is thus

equation[equation omitted — 126 chars of source]

The inputs $x_{j,ir}$ $\alpha_{r}$ and $\epsilon_{ir}$ are generated as follows. In the case when $x_{1}=x_{2}$, $x_{1,ir}=x_{2,ir}\sim\textrm{i.i.d.}\,N(0,1)$. In the case when $x_{1}\neq x_{2}$, $x_{1,ir}$ and $x_{2,ir}$ are generated mutually independently, each drawn from an $\textrm{i.i.d.}\,N(0,1)$. We then calculate the leave-out-mean $\bar{x}_{2,(-i)r}=\frac{1}{m_{r}}\sum_{j\neq i}x_{2j,r}$. Group characteristics are drawn as $x_{3,r}\sim\textrm{i.i.d.}\,N(0,1)$. In the case of homoscedastic normal errors in Tables (ref) to (ref), the idiosyncratic error terms $\epsilon_{ir}$ are i.i.d $N(0,1)$ and group effects $\alpha_{r}$ are i.i.d $N(0,0.25)$. Both $\epsilon_{ir}$ and $\alpha_{r}$ are drawn independently of $x_{1,ir}$, $x_{2,ir}$, $x_{3,r}$, and of each other. The dependent variable $y_{ir}$ is calculated using Equation (ref). In Table (ref), we use homoscedastic but nonnormal errors. In the case of the Skew normal distribution, we set the location parameter to 0, scale to 1 and shape to $0.9/\sqrt{1-0.9^{2}}$. Therefore, Skewness is 0.472 and Kurtosis is 3.321. In the case of the student distribution, degrees of freedom are set to 6. Therefore, Skewness is 0 and Kurtosis is 6. In both cases, $\alpha_{r}$ and $\epsilon_{ir}$ are independently drawn from identical distributions and then standardized to have mean 0 and variance 0.25 and 1 respectively. In Table (ref), group effects $\alpha_{r}$ are still i.i.d $N(0,0.25)$, $\epsilon_{ir}$ follow normal distributions but are allowed to be heteroscedastic. In the first case (Columns 1-2), we randomly select half of the groups into category 1, with $\epsilon_{ir}$ i.i.d $N(0,0.5)$. The other half of the groups have $\epsilon_{ir}$ i.i.d $N(0,1.5)$. In the second case (Columns 3-4), $\epsilon_{ir}$ are i.i.d $N(0,1)$. But we randomly divide the groups into two categories and allow for heteroscedasticity of $\epsilon_{ir}$ between categories in estimation. In the third case (Columns 5-6), groups are randomly divided into two categories, with $\sigma_{\epsilon r}^{2}\in\{0.5,1.5\}$ and $\epsilon_{ir}$ i.i.d $N(0,\sigma_{\epsilon r}^{2})$. In the fourth case, groups are randomly divided into four categories with $\sigma_{\epsilon r}^{2}\in\{0.4,0.8,1.2,1.6\}$ and $\epsilon_{ir}$ i.i.d $N(0,\sigma_{\epsilon r}^{2})$.

The number of groups $R$ is selected from the set $\{50,100,200,400,800,1600\}$. In Tables (ref), (ref) and (ref), group size $m_{r}$ is drawn from a discrete uniform distribution $\mathcal{U}\{2,6\}$ so that the average group size is 4. Small group sizes are motivated by applications to college room mates, friendship networks in the Add Health data set or golf tournaments, see sacerdote_peer_2001, goldsmith-pinkham_social_2013 and guryan_peer_2009. In Tables (ref) and (ref), group size is drawn from $\mathcal{U}\{13,25\}$. The distribution is motivated by Project STAR where class size ranges from 13 to 25. We also consider the case when $m_{r}$ is drawn from $\mathcal{U}\{3,5\}$, $\mathcal{U}\{4,8\}$, $\mathcal{U}\{8,30\}$ and $\mathcal{U}\{10,22\}$ in Table (ref) to examine how the distribution of group size affects the performance of the estimator. Note that $\mathcal{U}\{3,5\}$ has the same mean as $\mathcal{U}\{2,6\}$ but smaller variance, $\mathcal{U}\{4,8\}$ has the same variance as $\mathcal{U}\{2,6\}$ but larger mean. Meanwhile $\mathcal{U}\{8,30\}$ has the same mean as $\mathcal{U}\{13,25\}$ but larger variance, $\mathcal{U}\{10,22\}$ has the same variance as $\mathcal{U}\{13,25\}$ but smaller mean.

In Tables 1-(ref), we compare our QML estimator with the conditional maximum likelihood (CML) estimator of lee_identification_2007. Table (ref) does not present CMLE estimates as it does not allow for heteroscedasticity. lee_identification_2007 assumes normality of the error terms. Our discussion suggests that the CMLE is in fact consistent under nonnormal errors, as it can be viewed as a GMM estimator based on the moment conditions from the within equation. When group effects are in fact independent of the observed characteristics, the CML estimator is still consistent but less efficient than our QML estimator. The comparison thus helps to evaluate the efficiency gain of our estimator over the CML estimator in finite samples. The CML estimator is based on the within-group variation hence $\sigma_{\alpha}^{2}$, $\beta_{1}$ and $\beta_{4}$ are not identified.

We generate 5000 repetitions for each of the experiments. Tables (ref)-(ref) summarize the results of the Monte Carlo (MC) experiments. Each panel displays the MC median, MC robust standard errors (Rob.Std.Dev), MC sample standard deviation (Std.Dev.), MC median of the estimated standard deviation (est.Std.Dev), and the mean rejection rate of the Wald test with significance level 0.05 of our QMLE and Lee's CMLE across 5000 repetitions. The robust standard errors are defined as IQ/1.35, where IQ denotes the inter-quantile range, that is $IQ=C_{0.75}-C_{0.25}$ with $C_{0.75}$ and $C_{0.25}$ being the 75th and 25th percentile respectively. If the distribution of the estimate is normal, IQ/1.35 is (apart from rounding errors) equal to the standard deviation. The null hypothesis for the Wald test is that the estimate equals its true value. Critical values for the test are obtained at 5% significance level and are based on the asymptotic approximation in Theorem (ref).

Identification of our models is more challenging, the larger group sizes are, all else equal. This follows from work of kelejian_2sls_2002. Identification is also more difficult when there is less variation in group sizes, or less variation in type specific variances or both. Finally, identification is more difficult in designs where $x_{1,ir}=x_{2,ir}$ because the implied correlation between $x_{1,ir}$ and $\bar{x}_{2,(-i)r}$ reduces the overall variation in the covariates. Standard finite sample theory for the Gaussian regression model shows that maximum likelihood estimators for the variance parameters are biased in finite samples. In fixed effects panel regressions this finite sample bias can lead to inconsistent estimates of the variance parameter due to incidental parameter bias, as demonstrated by neyman_consistent_1948. In the current context, we expect the CML estimator to suffer from such incidental parameter bias because the moment conditions that identify $\lambda$ depend on the estimated variances. We also expect Wald type statistics, such as the t-ratio, to perform poorly in designs where identification is problematic, in line with insights from dufour_impossibility_1997.

Tables (ref) and (ref) contain results for small groups and homoscedastic Gaussian errors. In Table (ref) where $x_{1}\neq x_{2},$ both the QMLE and CMLE perform well, with the CMLE being more biased for the parameter $\lambda$ in sample sizes where $R$ is below 200. The QMLE is generally less biased and significantly more precise than CMLE, demonstrating the expected efficiency gains of QMLE. Size is better controlled for CMLE but the size distortions for the parameters $\lambda$ and $\beta$ do not exceed 7% in the smallest sample sizes even for the QMLE. Size distortions for the t-ratios of the two estimated variance parameters are somewhat larger, reaching 11.6% for the t-ratio for $\sigma_{\alpha}^{2}$ when $R=50.$ The size distortion seems to be due both to some estimator bias as well as standard errors that are a bit too small. Size distortions for all parameters disappear in the larger samples. In Table (ref) where $x_{1}=x_{2}$ the CMLE for $\lambda$ is even more biased in small samples, and considerably more volatile than in the design in Table (ref). The performance of the QMLE is not very different from the case with $x_{1}\neq x_{2}.$ The standard deviation measured by IQ/1.35 is somewhat larger than when $x_{1}\neq x_{2}$, as are size distortions, confirming the intuition that this design is more difficult to identify.

Tables (ref) and (ref) differ from Tables (ref) and (ref) in that they consider the same designs but with larger group sizes, now drawn from the uniform distribution on the interval $\left[13,25\right].$ In Table (ref) we consider the case with $x_{1}\neq x_{2}.$ The QMLE remains roughly unbiased across all sample sizes. The robust standard deviation roughly doubles relative to the small group size case and the size properties for t-ratios of the parameters $\lambda$ and $\beta$ deteriorate in samples where $R\leq100$ with size reaching around 10% in some cases. Size remains well controlled in larger samples with $R\geq200$. The size distortions for the variance parameters are not much affected by the larger class sizes. The CMLE is even more biased when $R=50$ but less biased for larger sample sizes compared to Tables (ref) and (ref). This is consistent with incidental parameter bias which is expected to decrease with increasing group size. In addition the CMLE now is significantly less precise. This is in line with results by Lee (2007). Table (ref) contains results for the case $x_{1}=x_{2}$ and large group sizes. The QMLE remains largely unbiased across all sample sizes but there is notable loss in estimator precision as measured by IQ/1.35, indicating the more challenging estimation environment. In line with theoretical predictions, estimator precision increases monotonically with sample size. Size distortions are now pronounced with empirical size reaching more than 20% in the smaller samples. The CMLE controls size well across all four designs. This comes at the cost of much less precisely and sometimes more biased estimated parameters.

Table (ref) explores the effects that variation in group size has on both estimators. The case with $\mathcal{U}\{3,5\}$ maintains the same mean group size as in Table (ref) but reduces the group size variance. We only report results for $\lambda.$ The bias of the QMLE is not affected while the CMLE is somewhat less biased. The variance of both estimators increases. For the QMLE size distortions are somewhat larger than in Table (ref). The design with $\mathcal{U}\{4,8\}$ increases the mean while leaving the variance of class sizes unchanged relative to Table (ref). Overall, the results for this case are quite similar to the scenario with $\mathcal{U}\{3,5\}$. The designs with $\mathcal{U}\{8,30\}$ and $\mathcal{U}\{10,22\}$ both improve identification relative to the design in Table (ref). For the QMLE this results in unchanged good bias properties except when $R=50$ where we now see a small amount of bias, somewhat lower variance and slightly improved size properties. For the CMLE bias increases while variance somewhat improves relative to the results in Table (ref) and the size properties remain similar.The larger bias for the CMLE may be related to a larger fraction of smaller classes in both designs. Smaller group sizes tend to amplify incidental parameter bias.

Table (ref) explores the effects that non-Gaussian error distributions have on the estimators. For the Skew Normal distribution we see little difference to the results in Table (ref) both for the QMLE and the CMLE estimator. The QMLE is also robust to the second design which uses a t-distribution with 6 degrees of freedom. The CMLE is more sensitive to this fat-tailed distribution. It is somewhat more biased and has higher variance compared to the Gaussian case. In addition, we now observe size distortions for the t-ratio related to the parameter $\lambda.$ These size distortions don't disappear in larger samples and seem to be due to the fact that the standard errors show a significant downward bias. This is most likely due to the fact that Lee (2007) bases standard errors on Gaussian error distributions. The final set of results we discuss are in Table (ref) where we examine the effects of heteroscedasticity on the QMLE. We do not report results for the CMLE since this estimator was designed for the homoscedastic case only. The first set of results are based on a design where class size varies according to a $\mathcal{U}\{2,6\}$ distribution and where we maintain $x_{1}=x_{2}$. Compared to a homoscedastic design the QMLE is somewhat less variable with no change in bias. The size properties of the t-ratio are overall comparable between the two cases, with slightly smaller size distortions in the heteroscedastic case when $R=50.$ We also consider a scenario where group size is fixed at $m=4$ while the type specific variances vary. While the QMLE continues to be nearly unbiased it has a higher variance. The size properties of t-ratios are slightly worse than in the homoscedastic case. For larger sample sizes both standard errors and t-ratios are well behaved.

Conclusion

In this paper, we show that moment conditions underlying the conditional variance method of Graham (2008) can be related to and motivated from a general class of linear peer effects models with random group effects. When augmented with group specific covariates our specification of the peer effects model is appropriate for settings where people are randomly assigned to groups or where group level heterogeneity is credibly controlled for with observed group level characteristics. We show that the quasi maximum likelihood estimator (QMLE) related to a linear Gaussian specification, as well as Graham's estimator and the fixed effects estimator of Lee (2007) are contained in the class of GMM estimators we consider. Under Gaussian error assumptions the QMLE is the most efficient estimator in this class. We study conditions of identification, extending results in Graham (2008) and Lee (2007) for a simple model without covariates and a general model with covariates estimated by QML. We also establish that our QMLE is asymptotically normal and we construct consistent standard error formulas. Monte Carlo results show that our QML estimator has good small sample properties.

\addcontentsline{toc}{section}{\refname}

thebibliography{52} \expandafter\ifx\csname natexlab\endcsname\relax\def\natexlab#1{#1}\fi \ifx\xfnm\relax \def\xfnm[#1]{\unskip,\space#1}\fi \bibitem[{Angrist(2014)}]{angrist_perils_2014} Angrist, J.D., 2014. \newblock The perils of peer effects. \newblock Labour Economics 30, 98--108. \bibitem[{Anselin(1988)}]{anselin_spatial_1988} Anselin, L., 1988. \newblock Spatial {{Econometrics}}: {{Methods}} and {{Models}}. volume 4. \newblock {Springer Science & Business Media}. \bibitem[{Anselin(2010)}]{anselin_thirty_2010} Anselin, L., 2010. \newblock Thirty years of spatial econometrics. \newblock Papers in Regional Science 89, 3--25. \bibitem[{Booij et al.(2017)Booij, Leuven and Oosterbeek}]{booij_ability_2017} Booij, A.S., Leuven, E., Oosterbeek, H., 2017. \newblock Ability {{Peer Effects}} in {{University}}: {{Evidence}} from a {{Randomized Experiment}}. \newblock The Review of Economic Studies 84, 547--578. \bibitem[{Boucher et al.(2014)Boucher, Bramoull{\'e}, Djebbari and Fortin}]{boucher_peers_2014} Boucher, V., Bramoull{\'e}, Y., Djebbari, H., Fortin, B., 2014. \newblock Do {{Peers Affect Student Achievement}}? {{Evidence}} from {{Canada Using Group Size Variation}}. \newblock Journal of Applied Econometrics 29, 91--109. \bibitem[{Bramoull{\'e} et al.(2009)Bramoull{\'e}, Djebbari and Fortin}]{bramoulle_identification_2009} Bramoull{\'e}, Y., Djebbari, H., Fortin, B., 2009. \newblock Identification of peer effects through social networks. \newblock Journal of Econometrics 150, 41--55. \bibitem[{Cai and Szeidl(2018)}]{cai_interfirm_2018} Cai, J., Szeidl, A., 2018. \newblock Interfirm {{Relationships}} and {{Business Performance}}*. \newblock The Quarterly Journal of Economics 133, 1229--1282. \bibitem[{Carrell et al.(2009)Carrell, Fullerton and West}]{carrell_does_2009} Carrell, S.E., Fullerton, R.L., West, J.E., 2009. \newblock Does {{Your Cohort Matter}}? {{Measuring Peer Effects}} in {{College Achievement}}. \newblock Journal of Labor Economics 27, 439--464. \bibitem[{Carrell et al.(2013)Carrell, Sacerdote and West}]{carrell_natural_2013} Carrell, S.E., Sacerdote, B.I., West, J.E., 2013. \newblock From {{Natural Variation}} to {{Optimal Policy}}? {{The Importance}} of {{Endogenous Peer Group Formation}}. \newblock Econometrica 81, 855--882. \bibitem[{Chamberlain(1980)}]{chamberlain_analysis_1980} Chamberlain, G., 1980. \newblock Analysis of {{Covariance}} with {{Qualitative Data}}. \newblock The Review of Economic Studies 47, 225--238. \bibitem[{Chetty et al.(2011)Chetty, Friedman, Hilger, Saez, Schanzenbach and Yagan}]{chetty_how_2011} Chetty, R., Friedman, J.N., Hilger, N., Saez, E., Schanzenbach, D.W., Yagan, D., 2011. \newblock How {{Does Your Kindergarten Classroom Affect Your Earnings}}? {{Evidence}} from {{Project Star}}. \newblock The Quarterly Journal of Economics 126, 1593--1660. \bibitem[{Chung(2001)}]{chung_course_2001} Chung, K.L., 2001. \newblock A Course in Probability Theory. \newblock {Academic press}. \bibitem[{Cliff and Ord(1973)}]{cliff_spatial_1973} Cliff, A.D., Ord, J.K., 1973. \newblock Spatial Autocorrelation. volume 5. \newblock {Pion London}. \bibitem[{Cliff and Ord(1981)}]{cliff_spatial_1981} Cliff, A.D., Ord, J.K., 1981. \newblock Spatial Processes: Models & Applications. volume 44. \newblock {Pion London}. \bibitem[{Dhrymes(1978)}]{dhrymes_mathematics_1978} Dhrymes, P.J., 1978. \newblock Mathematics for Econometrics. \newblock Technical Report. {Springer}. \bibitem[{Duflo et al.(2011)Duflo, Dupas and Kremer}]{duflo_peer_2011} Duflo, E., Dupas, P., Kremer, M., 2011. \newblock Peer {{Effects}}, {{Teacher Incentives}}, and the {{Impact}} of {{Tracking}}: {{Evidence}} from a {{Randomized Evaluation}} in {{Kenya}}. \newblock The American Economic Review 101, 1739--1774. \bibitem[{Duflo and Saez(2003)}]{duflo_role_2003} Duflo, E., Saez, E., 2003. \newblock The {{Role}} of {{Information}} and {{Social Interactions}} in {{Retirement Plan Decisions}}: {{Evidence}} from a {{Randomized Experiment}}*. \newblock The Quarterly Journal of Economics 118, 815--842. \bibitem[{Dufour(1997)}]{dufour_impossibility_1997} Dufour, J.M., 1997. \newblock Some {{Impossibility Theorems}} in {{Econometrics With Applications}} to {{Structural}} and {{Dynamic Models}}. \newblock Econometrica 65, 1365--1387. \bibitem[{Fafchamps and Quinn(2018)}]{fafchamps_networks_2018} Fafchamps, M., Quinn, S., 2018. \newblock Networks and {{Manufacturing Firms}} in {{Africa}}: {{Results}} from a {{Randomized Field Experiment}}. \newblock The World Bank Economic Review 32, 656--675. \bibitem[{Frijters et al.(2019)Frijters, Islam and Pakrashi}]{frijters_heterogeneity_2019} Frijters, P., Islam, A., Pakrashi, D., 2019. \newblock Heterogeneity in peer effects in random dormitory assignment in a developing country. \newblock Journal of Economic Behavior & Organization 163, 117--134. \bibitem[{Garlick(2018)}]{garlick_academic_2018} Garlick, R., 2018. \newblock Academic {{Peer Effects}} with {{Different Group Assignment Policies}}: {{Residential Tracking}} versus {{Random Assignment}}. \newblock American Economic Journal: Applied Economics 10, 345--369. \bibitem[{{Goldsmith-Pinkham} and Imbens(2013)}]{goldsmith-pinkham_social_2013} {Goldsmith-Pinkham}, P., Imbens, G.W., 2013. \newblock Social {{Networks}} and the {{Identification}} of {{Peer Effects}}. \newblock Journal of Business & Economic Statistics 31, 253--264. \bibitem[{Graham(2008)}]{graham_identifying_2008} Graham, B.S., 2008. \newblock Identifying {{Social Interactions Through Conditional Variance Restrictions}}. \newblock Econometrica 76, 643--660. \bibitem[{Guryan et al.(2009)Guryan, Kroft and Notowidigdo}]{guryan_peer_2009} Guryan, J., Kroft, K., Notowidigdo, M.J., 2009. \newblock Peer {{Effects}} in the {{Workplace}}: {{Evidence}} from {{Random Groupings}} in {{Professional Golf Tournaments}}. \newblock American Economic Journal: Applied Economics 1, 34--68. \bibitem[{Johnson and Horn(1985)}]{johnson_matrix_1985} Johnson, C.R., Horn, R.A., 1985. \newblock Matrix Analysis. \newblock {Cambridge university press Cambridge}. \bibitem[{Kang(2007)}]{kang_classroom_2007} Kang, C., 2007. \newblock Classroom peer effects and academic achievement: {{Quasi-randomization}} evidence from {{South Korea}}. \newblock Journal of Urban Economics 61, 458--495. \bibitem[{Kapoor et al.(2007)Kapoor, Kelejian and Prucha}]{kapoor_panel_2007} Kapoor, M., Kelejian, H.H., Prucha, I.R., 2007. \newblock Panel data models with spatially correlated error components. \newblock Journal of Econometrics 140, 97--130. \bibitem[{Kelejian and Prucha(1998)}]{kelejian_generalized_1998} Kelejian, H.H., Prucha, I.R., 1998. \newblock A {{Generalized Spatial Two-Stage Least Squares Procedure}} for {{Estimating}} a {{Spatial Autoregressive Model}} with {{Autoregressive Disturbances}}. \newblock The Journal of Real Estate Finance and Economics 17, 99--121. \bibitem[{Kelejian and Prucha(1999)}]{kelejian_generalized_1999} Kelejian, H.H., Prucha, I.R., 1999. \newblock A {{Generalized Moments Estimator}} for the {{Autoregressive Parameter}} in a {{Spatial Model}}. \newblock International Economic Review 40, 509--533. \bibitem[{Kelejian and Prucha(2001)}]{kelejian_asymptotic_2001} Kelejian, H.H., Prucha, I.R., 2001. \newblock On the asymptotic distribution of the {{Moran I}} test statistic with applications. \newblock Journal of Econometrics 104, 219--257. \bibitem[{Kelejian and Prucha(2002)}]{kelejian_2sls_2002} Kelejian, H.H., Prucha, I.R., 2002. \newblock {{2SLS}} and {{OLS}} in a spatial autoregressive model with equal spatial weights. \newblock Regional Science and Urban Economics 32, 691--707. \bibitem[{Kelejian and Prucha(2010)}]{kelejian_specification_2010} Kelejian, H.H., Prucha, I.R., 2010. \newblock Specification and estimation of spatial autoregressive models with autoregressive and heteroskedastic disturbances. \newblock Journal of Econometrics 157, 53--67. \bibitem[{Kelejian et al.(2006)Kelejian, Prucha and Yuzefovich}]{kelejian_estimation_2006} Kelejian, H.H., Prucha, I.R., Yuzefovich, Y., 2006. \newblock Estimation {{Problems}} in {{Models}} with {{Spatial Weighting Matrices Which Have Blocks}} of {{Equal Elements}}*. \newblock Journal of Regional Science 46, 507--515. \bibitem[{Kuersteiner and Prucha(2020)}]{kuersteiner_dynamic_2020} Kuersteiner, G.M., Prucha, I.R., 2020. \newblock Dynamic {{Spatial Panel Models}}: {{Networks}}, {{Common Shocks}}, and {{Sequential Exogeneity}}. \newblock Econometrica 88, 2109--2146. \bibitem[{Lee(2007)}]{lee_identification_2007} Lee, L.F., 2007. \newblock Identification and estimation of econometric models with group interactions, contextual factors and fixed effects. \newblock Journal of Econometrics 140, 333--374. \bibitem[{Lee et al.(2010)Lee, Liu and Lin}]{lee_specification_2010} Lee, L.F., Liu, X., Lin, X., 2010. \newblock Specification and estimation of social interaction models with network structures. \newblock The Econometrics Journal 13, 145--176. \bibitem[{Lin(2010)}]{lin_identifying_2010} Lin, X., 2010. \newblock Identifying {{Peer Effects}} in {{Student Academic Achievement}} by {{Spatial Autoregressive Models}} with {{Group Unobservables}}. \newblock Journal of Labor Economics 28, 825--860. \bibitem[{Liu and Lee(2010)}]{liu_gmm_2010} Liu, X., Lee, L.f., 2010. \newblock {{GMM}} estimation of social interaction models with centrality. \newblock Journal of Econometrics 159, 99--115. \bibitem[{Liu et al.(2014)Liu, Patacchini and Zenou}]{liu_endogenous_2014} Liu, X., Patacchini, E., Zenou, Y., 2014. \newblock Endogenous peer effects: Local aggregate or local average? \newblock Journal of Economic Behavior & Organization 103, 39--59. \bibitem[{Manski(1993)}]{manski_identification_1993} Manski, C.F., 1993. \newblock Identification of {{Endogenous Social Effects}}: {{The Reflection Problem}}. \newblock The Review of Economic Studies 60, 531--542. \bibitem[{Mundlak(1978)}]{mundlak_pooling_1978} Mundlak, Y., 1978. \newblock On the {{Pooling}} of {{Time Series}} and {{Cross Section Data}}. \newblock Econometrica 46, 69--85. \bibitem[{Newey(1991)}]{newey_uniform_1991} Newey, W.K., 1991. \newblock Uniform {{Convergence}} in {{Probability}} and {{Stochastic Equicontinuity}}. \newblock Econometrica 59, 1161--1167. \bibitem[{Neyman and Scott(1948)}]{neyman_consistent_1948} Neyman, J., Scott, E.L., 1948. \newblock Consistent {{Estimates Based}} on {{Partially Consistent Observations}}. \newblock Econometrica 16, 1--32. \bibitem[{Nye et al.(2004)Nye, Konstantopoulos and Hedges}]{nye_how_2004} Nye, B., Konstantopoulos, S., Hedges, L.V., 2004. \newblock How {{Large Are Teacher Effects}}? \newblock Educational Evaluation and Policy Analysis 26, 237--257. \bibitem[{Ord(1975)}]{ord_estimation_1975} Ord, K., 1975. \newblock Estimation {{Methods}} for {{Models}} of {{Spatial Interaction}}. \newblock Journal of the American Statistical Association 70, 120--126. \bibitem[{P{\"o}tscher and Prucha(1991)}]{potscher_basic_1991} P{\"o}tscher, B.M., Prucha, I.R., 1991. \newblock Basic structure of the asymptotic theory in dynamic nonlineaerco nometric models, part i: Consistency and approximation concepts. \newblock Econometric Reviews 10, 125--216. \bibitem[{P{\"o}tscher and Prucha(1994)}]{potscher_generic_1994} P{\"o}tscher, B.M., Prucha, I.R., 1994. \newblock Generic uniform convergence and equicontinuity concepts for random functions. \newblock Journal of Econometrics 60, 23--63. \bibitem[{Rivkin et al.(2005)Rivkin, Hanushek and Kain}]{rivkin_teachers_2005} Rivkin, S.G., Hanushek, E.A., Kain, J.F., 2005. \newblock Teachers, {{Schools}}, and {{Academic Achievement}}. \newblock Econometrica 73, 417--458. \bibitem[{Sacerdote(2001)}]{sacerdote_peer_2001} Sacerdote, B., 2001. \newblock Peer {{Effects}} with {{Random Assignment}}: {{Results}} for {{Dartmouth Roommates}}. \newblock The Quarterly Journal of Economics 116, 681--704. \bibitem[{Sojourner(2013)}]{sojourner_identification_2013} Sojourner, A., 2013. \newblock Identification of {{Peer Effects}} with {{Missing Peer Data}}: {{Evidence}} from {{Project STAR}}. \newblock The Economic Journal 123, 574--605. \bibitem[{Stinebrickner and Stinebrickner(2006)}]{stinebrickner_what_2006} Stinebrickner, R., Stinebrickner, T.R., 2006. \newblock What can be learned about peer effects using college roommates? {{Evidence}} from new survey data and students from disadvantaged backgrounds. \newblock Journal of Public Economics 90, 1435--1454. \bibitem[{Zimmerman(2003)}]{zimmerman_peer_2003} Zimmerman, D.J., 2003. \newblock Peer {{Effects}} in {{Academic Outcomes}}: {{Evidence}} from a {{Natural Experiment}}. \newblock Review of Economics and Statistics 85, 9--23.