The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
107,690 characters
Efficient Peer Effects Estimators with Group Effects
\title{Efficient Peer Effects Estimators with Group Effects}
\author{Guido M. Kuersteiner, Ingmar R. Prucha, and Ying Zeng\thanks{Kuersteiner: Department of Economics, University of Maryland, College
Park, MD 20742, United States ([email removed]); Prucha: Department
of Economics, University of Maryland, College Park, MD 20742, United
States ([email removed]); Zeng: Department of Public Finance, School
of Economics, Xiamen University, Xiamen 361005, China ([email removed]).}}
\maketitle
\begin{abstract}
We study linear peer effects models where peers interact in groups
and individual's outcomes are linear in the group mean outcome and
characteristics. We allow for unobserved random group effects as well
as observed fixed group effects. The specification is in part motivated
by the moment conditions imposed in \citet{graham_identifying_2008}.
We show that these moment conditions can be cast in terms of a linear
random group effects model and that they lead to a class of GMM estimators
with parameters generally identified as long as there is sufficient
variation in group size or group types. We also show that our class
of GMM estimators contains a Quasi Maximum Likelihood estimator (QMLE)
for the random group effects model, as well as the Wald estimator
of \citet{graham_identifying_2008} and the within estimator of \citet{lee_identification_2007}
as special cases. Our identification results extend insights in \citet{graham_identifying_2008}
that show how assumptions about random group effects, variation in
group size and certain forms of heteroscedasticity can be used to
overcome the reflection problem in identifying peer effects. Our QMLE
and GMM estimators accommodate additional covariates and are valid
in situations with a large but finite number of different group sizes
or types. Because our estimators are general moment based procedures,
using instruments other than binary group indicators in estimation
is straight forward. Our QMLE estimator accommodates group level covariates
in the spirit of Mundlak and Chamberlain and offers an alternative
to fixed effects specifications. This model feature significantly
extends the applicability of Graham's identification strategy to situations
where group assignment may not be random but correlation of group
level effects with peer effects can be controlled for with observable
group level characteristics. Monte-Carlo simulations show that the
bias of the QMLE estimator decreases with the number of groups and
the variation in group size, and increases with group size. We also
prove the consistency and asymptotic normality of the estimator under
reasonable assumptions.
\newpage{}
\end{abstract}
\section{Introduction}
\global\long\global\long\global\long\global\long\global\long\global\longPeer effects are of great interest to empirical researchers and policy
makers. The idea that individuals are affected by their peers motivates
policies that try to manipulate peer composition for better outcomes.
Peer effects are often confounded by group level effects. An example
are teacher effects in a class room setting. Identifying peer effects
is notoriously challenging due to the reflection problem \citep{manski_identification_1993,angrist_perils_2014}
as well as due to spurious peer effects originating from group level
effects. Random group allocation may be one way to overcome these
identification problems. With groups formed at random, a random effects
specification for group level characteristics can be adopted. An alternative
approach consists in postulating that, conditional on observed group
level characteristics, group level effects can be viewed as randomly
assigned. Regression control techniques based on observed group characteristics
then lead to a similar random effects specification, but without the
need to appeal to random group assignment. We propose estimators that
can accommodate both scenarios.
Random group assignment plays a prominent role in the empirical peer
effects literature in a number of fields including education, labor,
firm, finance and development studies. Recent examples from this literature
include \textcolor{black}{\citet{sacerdote_peer_2001,duflo_role_2003,zimmerman_peer_2003,stinebrickner_what_2006,kang_classroom_2007,graham_identifying_2008,guryan_peer_2009,carrell_does_2009,carrell_natural_2013,duflo_peer_2011,sojourner_identification_2013,booij_ability_2017,garlick_academic_2018,fafchamps_networks_2018,cai_interfirm_2018,frijters_heterogeneity_2019}.}
Assuming group effects to be independent of observed individual and
group characteristics is plausible when groups are formed at random.
Ignoring group effects or assuming fixed group effects \citep{lee_identification_2007}
leads to consistent but less efficient estimators. Random group effects
themselves have important empirical interpretations. For example,
researchers in education policy often treat random class effects as
unobserved teacher effects (e.g., \citealt{nye_how_2004,rivkin_teachers_2005,chetty_how_2011}).
Absent random group assignment, the estimators we propose can accommodate
observed group level effects that can come from information about
group characteristics such as the training and experience of teachers,
or averages of individual group member characteristics. Group level
characteristics can be interpreted as parametrizations of group effects
in the spirit of \citet{mundlak_pooling_1978} and \citet{chamberlain_analysis_1980}.
The choice between a random effects or fixed effects estimator then
depends less on random group assignment but more on whether group
specific effects are believed to be observable or not. In some cases
there may be independent interest in the effects of group specific
covariates. An example is the effect of teacher training on student
performance. In such cases a random effects estimator is the preferred
choice because fixed effects estimators are often unable to identify
these types of group level effects.
Our analysis extends insights in \citet{graham_identifying_2008}
that show how assumptions about random group effects, variation in
group size and certain forms of heteroscedasticity can be used to
overcome the reflection problem in identifying peer effects. We give
an interpretation of the conditional variance estimator (CVE) of \citet{graham_identifying_2008}
in terms of a GMM estimator based on moment conditions for the within-group
variance and between-group variance. We show that the moment conditions
underlying \citet{graham_identifying_2008} are the score function
of a quasi maximum likelihood estimator (QMLE) for a random group
effects model. The QMLE can be shown to be the best GMM estimator
in the class of estimators using moment conditions for the within
and between variances of outcomes individually, rather than combining
them into a single moment condition as is the case for the CV estimator,
or focusing only on the within variation as is the case of the CMLE
of \citet{lee_identification_2007}.
One limitation of the conditional variance estimator proposed by Graham
is the fact that it amounts to a difference in difference identification
strategy for the variances that requires groups to fall into two size
categories. As shown in \citet{graham_identifying_2008} the resulting
procedure takes the form of a Wald estimator for a set of binary instruments.
This setting is restrictive in applications where groups may not be
easily separated into two categories or where a more general set of
instruments needs to be considered. The estimators that we propose
are general GMM based procedures that accommodate additional covariates
as well as offer flexibility in terms of the instruments and the number
of moment conditions that are being used. We illustrate these points
by explicitly considering moment based estimators that exploit exogenous
variation in group size as well as general group level heteroscedasticity
as instruments. In contrast to the CMLE of \citet{lee_identification_2007},
which is a member of the class of GMM estimators we consider, our
QMLE uses both the within and between variance. This leads to efficiency
gains under correct specification but comes at the cost of potential
miss-specification bias if the assumption of observed fixed and unobserved
random group effects is incorrect. The trade-offs are similar to related
results for fixed and random effects in the panel literature.
Our work is also related to the literature in spatial econometrics
started by the work of \citet{cliff_spatial_1973,cliff_spatial_1981}
and \citet{anselin_spatial_1988}.\footnote{\citet{anselin_thirty_2010} offers a brief review of the development
of spatial econometrics literature over the past thirty years.} Recently, there is a growing number of studies using spatial methods
to model social network effects, e.g., \citet{lee_identification_2007},
\citet{bramoulle_identification_2009}, and \citet{kuersteiner_dynamic_2020}.
The strength of social links can be characterized by proximity in
the social network space. We extend \citet{kelejian_estimation_2006}
and \citet{lee_identification_2007} by considering a random group
effects specification. Spatial models were traditionally estimated
with maximum likelihood (ML), e.g., \citet{ord_estimation_1975}.
\citet{kelejian_generalized_1998,kelejian_generalized_1999} develop
generalized method of moments (GMM) estimators based on linear and
quadratic moments. While this paper utilizes a quasi-maximum likelihood
estimation method, the score function depends on linear quadratic
forms of the error terms. Properties of quadratic moment conditions
were introduced by \citet{kelejian_generalized_1998,kelejian_generalized_1999}
in the cross section case, and \citet{kapoor_panel_2007} and \citet{kuersteiner_dynamic_2020}
in a panel setting. Moreover, \citet{kelejian_asymptotic_2001} and
\citet{kelejian_specification_2010} develop a central limit theorem
for linear quadratic forms, which is the basis for the asymptotic
analysis in this paper.
The linear-in-means peer effect model in \citet{manski_identification_1993}
is a special case of a spatial model with group-wise equal dependence,
see \citet{kelejian_2sls_2002} and \citet{kelejian_estimation_2006}.
\citet{kelejian_2sls_2002} were the first to study the group-wise
equal dependence spatial model. They show that if there is one group
in a single cross section and the model has equal spatial weights,
two-stage least squares (2SLS), GMM and QMLE methods all yield inconsistent
estimators, although consistent estimation with 2SLS and GMM is possible
for panel data. However, \citet{kelejian_estimation_2006} point out
that if group fixed effects are incorporated and the panel is balanced,
the estimators are inconsistent. The results in \citet{kelejian_estimation_2006}
show the importance of variation in group size in identification of
spatial models with blocks of equal weights. The QMLE developed in
this paper and the conditional maximum likelihood estimator in \citet{lee_identification_2007}
both rely on group size variation for identification although we show
that identification exploiting heteroscedastic errors is also possible.
Extensions include \citet{lee_specification_2010} who allow for specific
social structure within each group and \citet{liu_gmm_2010} and \citet{liu_endogenous_2014}
who allow for non-row normalized weight matrices. The linear spatial
model has also been applied to the empirical evaluation of peer effects
by \citet{lin_identifying_2010} and \citet{boucher_peers_2014}.
\citet{bramoulle_identification_2009} study a broader range of social
interaction models and give conditions for identification.
The paper is organized as follows. In Section \ref{sec:graham} we
consider identification of endogenous peer effects in a simple setting
without covariates for the CV, CML and QML estimators. Section \ref{sec:General-Model}
presents the full model that allows for covariates and general variation
in group size. Section \ref{sec:Theoretical-Results} summarizes the
technical conditions we impose and presents theoretical results for
the QMLE. Section \ref{sec:Monte_Carlo} contains a small Monte Carlo
experiment. Proofs are collected in an appendix.
\section{Peer Effects with Random Group Effects\label{sec:graham}}
We start the discussion by presenting a simple model without covariates,
to introduce and discuss basic features of our new quasi-maximum likelihood
estimator (QMLE), and connect it to the conditional variance (CV)
estimator in \citet{graham_identifying_2008} and the conditional
maximum likelihood (CMLE) estimator in \citet{lee_identification_2007}.
The model decomposes variation in outcomes of a cross-section of individuals
into idiosyncratic noise, group level random effects and correlation
that is due to group level interaction. Quadratic moment conditions
implied by this random effects specification lead to efficient GMM,
quasi maximum likelihood, and under additional distributional assumptions,
maximum likelihood estimators. Estimators based on these moment conditions
include the CV estimator of \citet{graham_identifying_2008}, the
QMLE as well as the CMLE of \citet{lee_identification_2007} as special
cases.
Let $y_{ir}$ be an observed outcome of individual $i$ in group $r$
which has $m_{r}$ members, let $\alpha_{r}$ be an unobserved group
level effect and let $\epsilon_{ir}$ be unobserved individual specific
characteristics. We observe data for $R$ groups as well as a categorical
variable $D_{r}$ which determines group type. An example is when
there are three group sizes such that $D_{r}\in\left\{ 'small','medium','large'\right\} .$
However, $D_{r}$ could be a characteristic that is not necessarily
related to group size. An example is when groups are defined by classrooms
of schools in urban, suburban or rural districts and $D_{r}$ is used
to denote urbanicity. Classes could also be categorized by sociodemographic
composition such as whether English or other languages are the native
language spoken by students in the class. We allow for type-dependent
heteroscedasticity. Types add flexibility to the specification by
relaxing the constraints the model imposes on the relationship between
group variance and group size. In some cases type specific heteroscedasticity
provides identifying variation that is separate from group size variation.
The peer effects model is stated in terms of a structural equation
\begin{equation}
y_{ir}=\lambda\bar{y}_{\left(-i\right)r}+\alpha_{r}+\epsilon_{ir},\label{eq:Graham_Mod_L-M-O}
\end{equation}
where $\bar{y}_{\left(-i\right)r}=\frac{1}{m_{r}-1}\sum_{j\neq i}^{m_{r}}y_{jr}$
is the leave-out-mean of the outcome variable. The parameter $\lambda$
captures the endogenous peer effects, see \citet{manski_identification_1993}.
The structural form emphasizes the decomposition of $y_{ir}$ into
a social interaction term $\lambda\bar{y}_{(-i)r}$, a group level
effect $\alpha_{r}$ and an idiosyncratic error term $\epsilon_{ir}.$
For example, when $y_{ir}$ is a measure of student performance and
$r$ is a class-room index then $\alpha_{r}$ can be interpreted as
a class-room or teacher effect while $\epsilon_{ir}$ are unobserved
student characteristics for student $i$ in classroom $r.$ Cross-sectional
independence of $\epsilon_{ir}$ can be justified by random group
assignment such as in the application of Graham (2008). The assumptions
we impose on $\epsilon_{ir}$ and $\alpha_{r}$ are in line with the
random effects panel literature where group level dependence of unobservables
is modeled with the common factor $\alpha_{r}.$ We leave possible
generalizations of this framework to cases where $\epsilon_{ir}$
is allowed to be dependent for future work.
Following Graham (2008) who emphasizes random assignments of individuals
to groups, we assume that $\alpha_{r}$ is a random effect independent
of $\epsilon_{ir}.$ As shown by Graham (2008) for a slightly different
model based on full rather than leave-out means, the random effects
nature of the model leads to a set of quadratic moment conditions
that can be exploited for identification. We expand on these ideas
by showing that the implied moment conditions are related to the moment
conditions of a random effects pseudo likelihood estimator. Transformations
of these moments turn out to coincide with moments used by \citet{graham_identifying_2008}
as well as \citet{lee_identification_2007} who considers a fixed
effects version of the model. \citet{lee_identification_2007} focuses
on identification of $\lambda$ based on group size variation. Here
we emphasize a random effects specification where identification is
driven by heterogeneity at the group level that could result from
sources including but not limited to class size variation. A literature
on linear instrumental variables methods gives conditions under which
$\lambda$ can be identified in models that have additional exogenous
covariates $Z_{r}$, e.g., \citet{angrist_perils_2014} or \citet{bramoulle_identification_2009}.\footnote{The leave-out-mean $\bar{y}_{\left(-i\right)r}$ can be viewed as
a special case of a Cliff-Ord-type (\citealt{cliff_spatial_1973,cliff_spatial_1981})
spatial lag. \citet{kelejian_generalized_1998} give an early basic
condition for identification by IV.} Besides the conventional instrumental variables strategies, alternative
strategies are also available, see \citet{lee_identification_2007},
\citet{graham_identifying_2008} for a modified model or \citet{kuersteiner_dynamic_2020}.
Letting $Y_{r}=\left(y_{1r},...,y_{m_{r}r}\right)^{'}$, $\epsilon_{r}=\left(\epsilon_{1r},...,\epsilon_{m_{r}r}\right)^{'}$,
$\iota_{m_{r}}=\left(1,...,1\right)^{\prime}$ and $W_{m_{r}}=\frac{1}{m_{r}-1}(\iota_{m_{r}}\iota_{m_{r}}^{\prime}-I_{m_{r}})$,
the model can be written in matrix notation as
\begin{equation}
Y_{r}=\lambda W_{m_{r}}Y_{r}+\alpha_{r}\iota_{m_{r}}+\epsilon_{r}.\label{eq:Model_Graham_Modified}
\end{equation}
To isolate or identify the social interaction effect, we impose the
following restrictions on unobservables.
\begin{assumption}
\label{assume:epsilon} For $r=1,\ldots,R$ the $r$-th group is associated
with a categorical variable $D_{r}\in\{1,2,...,J\}$ with $J\geqslant1$
being fixed and finite, and for each category $j\in\{1,2,...,J\}$
there is at least one group $r$ with $D_{r}=j$. For $r=1,...,R$
and $i=1,...,m_{r}$ the disturbance terms $\epsilon_{ir}$ are independently
distributed across all $i$ and $r$, with $E\left[\epsilon_{ir}|D_{r},m_{r}\right]=0$
and $E\left[\epsilon_{ir}^{2}|D_{r},m_{r}\right]=\sigma_{\epsilon0,D_{r}}^{2}$
, $0<\underbar{\ensuremath{a}}_{\epsilon}\leqslant\sigma_{\epsilon0,D_{r}}^{2}\leqslant\overline{\text{\ensuremath{a}}}_{\epsilon}<\infty$
and where $\sigma_{\epsilon0,D_{r}}^{2}$ is a function only of $D_{r}$.
There exists some $\eta_{\epsilon}>0$ such that $E[|\epsilon_{ir}|^{4+\eta_{\epsilon}}]<\infty$.
\end{assumption}
Note that the variance $E\left[\epsilon_{ir}^{2}|D_{r},m_{r}\right]=\sigma_{\epsilon0,D_{r}}^{2}$
has the representation $\sigma_{\epsilon0,D_{r}}^{2}=\sigma_{\epsilon0,1}^{2}1\left\{ D_{r}=1\right\} +...+\sigma_{\epsilon0,J}^{2}1\left\{ D_{r}=J\right\} $
where $\sigma_{\epsilon0,1}^{2},....,\sigma_{\epsilon0,J}^{2}$ are
fixed parameters to be estimated.
\begin{assumption}
\label{assume:alpha}For $r=1,...,R$, the group effects $\alpha_{r}$
are independently and identically distributed, with $E\left[\alpha_{r}|D_{r},m_{r}\right]=0$
and $E\left[\alpha_{r}^{2}|D_{r},m_{r}\right]=\sigma_{\alpha0}^{2}$,
where $\text{\ensuremath{0\leq}}\sigma_{\alpha0}^{2}\leq\overline{\text{\ensuremath{a}}}_{\alpha}<\infty$.
There exists some $\eta_{\alpha}>0$ such that $E\left[|\alpha_{r}|^{4+\eta_{\alpha}}\right]<\infty$.
Also, $\{\alpha_{r}:r=1,...,R\}$ are independent of $\{\epsilon_{ir}:i=1,...,m_{r};r=1,...,R\}$.
\end{assumption}
Assumption \ref{assume:epsilon} implies in particular that individuals
do not self select into groups based on unobserved characteristics
and Assumption \ref{assume:alpha} suggests that there is no matching
between group characteristics and individual characteristics. This
no sorting or matching assumption can sometimes be motivated by specific
empirical designs. For example, in the Project STAR experiment that
\citet{graham_identifying_2008} considers, kindergarten students
and teachers are randomly assigned to classrooms. This random assignment
mechanism justifies interpreting $\alpha_{r}$ as the classroom or
teacher effect. It also justifies assuming that $\alpha_{r}$ and
$\epsilon_{ir}$ are mutually independent random variables, see Graham
(2008) Assumption 1.1. Assumption \ref{assume:epsilon} allows $\epsilon_{ir}$
to be homoscedastic across all groups when $J=1$ or heteroscedastic
across different categories of $D_{r}$ when $J\geqslant2$. This
formulation contains the case considered by \citet{graham_identifying_2008}
where $J=2$ as a special case.
Assumptions \ref{assume:epsilon} and \ref{assume:alpha} above imply
moment conditions. These moment conditions take the form of restrictions
on the within and between group variance. As discussed in more detail
below, these moment conditions are fundamental to the ML estimator.
In particular, we show that the score of the ML estimator is a weighted
average of those fundamental moment conditions.
To derive the moment conditions, define the composite error term $U_{r}=\alpha_{r}\iota_{m_{r}}+\epsilon_{r}$
where $U_{r}$ is an $m_{r}\times1$ vector with elements $u_{ir}=\alpha_{r}+\epsilon_{ir}.$
Let $\bar{u}_{r}$ and $\bar{\epsilon}_{r}$ be the mean of $u_{ir}$
and $\epsilon_{ir}$ in group $r$. Let $\ddot{U}_{r}=U_{r}-\bar{u}_{r}\iota_{m_{r}}$
be the vector of within-group deviations from the mean of $U_{r}$
and let $\ddot{Y}_{r}$ and $\ddot{\epsilon}_{r}$ be defined in a
similar manner. It can be shown that $\bar{y}_{r}=\bar{u}_{r}/(1-\lambda)=(\alpha_{r}+\bar{\epsilon}_{r})/(1-\lambda)$
with $\bar{u}_{r}=\alpha_{r}+\bar{\epsilon}_{r}$, and $\ddot{Y}_{r}=\frac{m_{r}-1}{m_{r}-1+\lambda}\ddot{U}_{r}=\frac{m_{r}-1}{m_{r}-1+\lambda}\ddot{\epsilon}_{r}$.
Two conditional moment conditions, one for the within-group variance,
the other for the between group variance, arise from the model in
(\ref{eq:Model_Graham_Modified}) under Assumptions \ref{assume:epsilon}
and \ref{assume:alpha}. The expected value of the within-group and
between-group squares of group $r$ are
\begin{equation}
var_{r}^{w}=E\left[\frac{\ddot{Y}_{r}^{\prime}\ddot{Y}_{r}}{m_{r}-1}|m_{r},D_{r}\right]=E\left[\frac{(m_{r}-1)\ddot{U}_{r}^{\prime}\ddot{U}_{r}}{(m_{r}-1+\lambda)^{2}}|m_{r},D_{r}\right]=\frac{\left(m_{r}-1\right)^{2}}{\left(m_{r}-1+\lambda\right)^{2}}\sigma_{\epsilon,D_{r}}^{2},\label{eq:varw}
\end{equation}
\begin{equation}
var_{r}^{b}=E\left[\bar{y}_{r}^{2}|m_{r},D_{r}\right]=E\left[\frac{\bar{u}_{r}^{2}}{(1-\lambda)^{2}}|m_{r},D_{r}\right]=\frac{1}{(1-\lambda)^{2}}\left(\sigma_{\alpha}^{2}+\frac{\sigma_{\epsilon,D_{r}}^{2}}{m_{r}}\right),\label{eq:varb}
\end{equation}
where $\sigma_{\epsilon,D_{r}}^{2}=\sigma_{\epsilon,1}^{2}1\left\{ D_{r}=1\right\} +...+\sigma_{\epsilon,J}^{2}1\left\{ D_{r}=J\right\} .$
To see how these moment conditions can achieve the identification
of $\lambda$ consider the case where $\sigma_{\epsilon,D_{r}}^{2}=\sigma_{\epsilon,D_{s}}^{2}$
but $m_{r}\neq m_{s}$. Then, Equation \eqref{eq:varw} implies that
\begin{equation}
(\frac{m_{r}-1+\lambda}{m_{s}-1+\lambda})^{2}=\frac{(m_{r}-1)^{3}}{(m_{s}-1)^{3}}\frac{E\left[\ddot{Y}_{s}^{\prime}\ddot{Y}_{s}|m_{s},D_{s}\right]}{E\left[\ddot{Y}_{r}^{\prime}\ddot{Y}_{r}|m_{r},D_{r}\right]}.\label{eq:Wald_within}
\end{equation}
Alternatively consider the case where $m_{r}=m_{s}=m$ and $\sigma_{\epsilon,D_{r}}^{2}\neq\sigma_{\epsilon,D_{s}}^{2}$,
then combining \eqref{eq:varw} and \eqref{eq:varb} gives
\begin{equation}
\frac{(m-1+\lambda)^{2}}{(1-\lambda)^{2}}=\frac{E\left[\bar{y}_{r}^{2}|m,D_{r}\right]-E\left[\bar{y}_{s}^{2}|m,D_{s}\right]}{E\left[\ddot{Y}_{r}^{\prime}\ddot{Y}_{r}/[m(m-1)^{3}]|m,D_{r}\right]-E\left[\ddot{Y}_{s}^{\prime}\ddot{Y}_{s}/[m(m-1)^{3}]|m,D_{s}\right]}.\label{eq:Wald_L-O-M}
\end{equation}
Expressions on the left hand side of both \eqref{eq:Wald_within}
and \eqref{eq:Wald_L-O-M} in principle can be solved for $\lambda$
if we restrict $\lambda\in(-1,1)$ and $m_{r}\geqslant2$, as both
expressions are monotonic functions of $\lambda$. Equation \eqref{eq:Wald_L-O-M}
is a modified version of Equation (9) in \citet{graham_identifying_2008}
that accounts for the leave-out-mean specification we consider. The
numerator differences out the variance of $\alpha_{r}$ which is assumed
constant across types. This restriction is also imposed by \citet{graham_identifying_2008}
in his Assumption 1.2. In Lemma \ref{lem:ID_1} below we outline the
exact conditions under which identification is possible.
The discussion above shows that under additional assumptions on $\lambda$
and group size, identification of $\lambda$ is possible through moment
conditions related to within and between variance when there is variation
in either group size $m_{r}$ or idiosyncratic error variance $\sigma_{\epsilon,D_{r}}^{2}$.
We now formalize the discussion into Lemma \ref{lem:ID_1} below.
Let the parameter vector be $\theta=\left(\lambda,\sigma_{\alpha}^{2},\sigma_{\epsilon,1}^{2},...,\sigma_{\epsilon,J}^{2}\right)^{\prime}$
and, for clarity, let the true parameter vector be denoted by $\theta_{0}=\left(\lambda_{0},\sigma_{\alpha0}^{2},\sigma_{\epsilon0,1}^{2},...,\sigma_{\epsilon0,J}^{2}\right)^{\prime}$.
For identification, we further assume that group size $m_{r}\geqslant2$
and impose the following assumption on $\lambda$.
\begin{assumption}
\label{assume:lambda} The parameter of the endogenous peer effects
$\lambda_{0}\in\Lambda$, where $\Lambda$ is a compact subset of
$(-1,1)$. Assume that $\theta_{0}\in\Theta$ with $\Theta=\Lambda\times\text{\ensuremath{[0,\overline{\text{\ensuremath{a}}}_{\alpha}]\times[\underbar{\ensuremath{a}}_{\epsilon},\overline{\text{\ensuremath{a}}}_{\epsilon}]\times\ldots\times[\underbar{\ensuremath{a}}_{\epsilon},\overline{\text{\ensuremath{a}}}_{\epsilon}]}}$
compact.
\end{assumption}
The estimation procedures we propose in this paper can be implemented
with the availability of a general set of valid instruments and are
valid for cases where $J\geqslant1$ as long as $J$ is fixed and
finite. In the simple model without covariates the available instruments
are group size $m_{r}$ and categorical variable $D_{r}$. These instruments
are valid if assignment to groups is random in a way that generates
random variation in group size or category. Utilizing Equation \eqref{eq:varw}
and \eqref{eq:varb}, and using group size $m_{r}$ and the categorical
variable $D_{r}$ as instruments yields the following conditional
moment restriction $E[\chi_{r}(\theta_{0})|m_{r},D_{r}]=0$ with
\begin{equation}
\chi_{r}(\theta)=\left[\begin{array}{c}
\chi_{r}^{w}(\theta)\\
\chi_{r}^{b}(\theta)
\end{array}\right]=\left[\begin{array}{c}
\frac{(m_{r}-1+\lambda)^{2}\ddot{Y}_{r}^{\prime}\ddot{Y}_{r}}{(m_{r}-1)^{2}}-(m_{r}-1)\sigma_{\epsilon,D_{r}}^{2}\\
(1-\lambda)^{2}\bar{y}_{r}^{2}-\sigma_{\alpha}^{2}-\frac{\sigma_{\epsilon,D_{r}}^{2}}{m_{r}}.
\end{array}\right].\label{eq:mwb}
\end{equation}
Identification of the parameter $\theta$ is possible with variation
in group size for a given category or variation in the idiosyncratic
variance over categories for the same group size. This is summarized
in the following lemma. The proof of the lemma is given in Appendix
\ref{sec:prooflemma2}.
\begin{lem}
\label{lem:ID_1}Suppose Assumptions \ref{assume:epsilon}-\ref{assume:lambda}
hold. Then the parameter $\theta_{0}$ is identified under the following
two scenarios:
(i) There are two groups $r$ and $s$ such that $m_{r}\neq m_{s}$
and $D_{r}=D_{s}$, and therefore $\sigma_{\epsilon0,D_{r}}^{2}=\sigma_{\epsilon0,D_{s}}^{2}$.
Then the parameter $\theta_{0}$ is identified in $\Theta.$ In particular,
the moment conditions $E[\chi_{q}^{w}(\theta)|m_{q},D_{q}]=0$ and
$E[\chi_{q}^{b}(\theta)|m_{q},D_{q}]=0$ for $q=r,s$ with $\chi_{r}^{w}(\theta)$
and $\chi_{r}^{b}(\theta)$ defined in (\ref{eq:mwb}) identify $\lambda_{0}$,
$\sigma_{\epsilon0,D_{r}}^{2}$,and $\sigma_{\alpha0}^{2}$. The remaining
parameters $\sigma_{\epsilon0,j}^{2}$ are identified by $E(\chi_{q}^{w}(\theta)|m_{q},D_{q})=0$
for $q\neq r$ or $s.$
(ii) There are two groups $r$ and $s$, such that $m_{r}=m_{s}$
and $\sigma_{\epsilon0,D_{r}}^{2}\neq\sigma_{\epsilon0,D_{s}}^{2}$.
Then the parameter $\theta_{0}$ is identified in $\Theta.$ In particular,
the moment condition $E[\nu_{q}(\theta)|m_{q},D_{q}]=0$, $q=r,s$
uniquely identifies $\lambda_{0}$ and $\sigma_{\alpha0}^{2}$, where
\begin{align}
\nu_{q}(\theta) & =\chi_{q}^{b}(\theta)-\frac{\chi_{q}^{w}(\theta)}{m_{q}(m_{q}-1)}=(1-\lambda)^{2}\bar{y}_{q}^{2}-\sigma_{\alpha}^{2}-\frac{(m_{q}-1+\lambda)^{2}\ddot{Y}_{q}^{\prime}\ddot{Y}_{q}}{m_{q}(m_{q}-1)^{3}}\label{eq:nu}
\end{align}
with $\chi_{q}^{w}(\theta)$ and $\chi_{q}^{b}(\theta)$ defined in
(\ref{eq:mwb}). The remaining parameters $\sigma_{\epsilon0,j}^{2}$
are identified by $E(\chi_{q}^{w}(\theta)|m_{q},D_{q})=0$.
\end{lem}
Full identification is achieved in Scenario (i) with group size variation
in at least one category. As an example, consider types that describe
urbanicity such that $D_{r}=D_{s}=1$ denotes two classrooms $r$
and $s$ that are both located in an urban school but where $m_{r}\neq m_{s}$
such that the classrooms differ in size, while the remaining categories
$d=2,...,J$ may have the same group sizes. In this setting $\theta_{0}$
is identified without any further constraints on the variances $\sigma_{\epsilon,j}^{2}$.
If the number of distinct group sizes exceeds the number of categories
$J$ then it automatically must be the case that there exist some
category that is associated with at least two distinct group sizes.
Note that the result holds irrespective of whether the constraint
of homoscedastic errors $\sigma_{\epsilon0,D_{r}}^{2}=\sigma_{\epsilon0,D_{s}}^{2}$
is imposed on the model or not. From Scenario (i) we see that variation
in group size alone can provide variation that is sufficient for identification.
Furthermore, in the homoscedastic case where only a common variance
parameter $\sigma_{\epsilon}^{2}$ is specified, two distinct group
sizes are sufficient for identification by the result in Scenario
(i). This corresponds to the identification result of the conditional
maximum likelihood estimator (CMLE) in \citet{lee_identification_2007},
the score function of which can be written as $\varphi(m_{r})\chi_{r}^{w}(\theta)$,
where $\varphi(m_{r})$ is a function of $m_{r}$.
While variation in group size serves as the source of identification
in Scenario (i), identification based on the moment condition $E[\chi_{r}(\theta)|m_{r},D_{r}]=0$
is also possible without group size variation as long as there is
some other form of group heterogeneity. As is shown in the proof for
Scenario (ii) of Lemma \ref{lem:ID_1}, utilizing $m_{q}=m$ and $E\left[\nu_{q}(\theta)|m_{q},D_{q}\right]=0$
for $q=r,s$ yields \eqref{eq:Wald_L-O-M}. From \eqref{eq:Wald_L-O-M}
we see that the endogenous peer effect parameter $\lambda$ is identified
if there is heteroscedasticity across groups of the same size for
at least one size, and that $\lambda$ can be estimated from the sample
analog of \eqref{eq:Wald_L-O-M}. The intuition of identification
in Scenario (ii) echoes that of the conditional variance (CV) estimator
of \citet{graham_identifying_2008}. Similar to Graham (2008), \eqref{eq:nu}
is based on the relationship between the within-group and between-group
variance as captured by $\nu_{r}(\theta)$ , and can be used to construct
a Wald type moment condition like in \eqref{eq:Wald_L-O-M} using
the categorical variable as the instrument.
The above discussion focused on identification based on the moment
vector $\chi_{r}(\theta)$. We next discuss the importance of these
moment conditions for efficient estimation, and their relationship
to the score of the Gaussian ML estimator. The optimal moment function
corresponding to $\chi_{r}(\theta)$ is given by $\chi_{r}^{\ast}(\theta)=\varphi^{*}(m_{r},D_{r})\chi_{r}(\theta)$
where, focusing on the case with $J=2$ for exposition, \footnote{See our Online Appendix for details. The derivation uses Lemma \ref{lemma:moment}
and the special properties of matrices $\Omega(\theta)$, $I-\lambda W$
and $W$ described in Appendix \ref{subsec:matrix properties}. In
the Online Appendix we also give an explicit expression for the variance
covariance matrix of $\chi_{r}(\delta).$}
\begin{align}
\varphi^{*}\left(m_{r},D_{r}\right) & =E[\frac{\partial}{\partial\theta^{\prime}}\chi_{r}(\theta_{0})|m_{r},D_{r}]\left(E[\chi_{r}(\theta_{0})\chi_{r}(\theta_{0})^{\prime}|m_{r},D_{r}]\right)^{-1}\nonumber \\
& =\left(\begin{array}{cc}
\frac{1}{\left(m_{r}-1+\lambda\right)\sigma_{\epsilon0,D_{r}}^{2}} & -\frac{m_{r}}{(1-\lambda)(\sigma_{\epsilon0,D_{r}}^{2}+m_{r}\sigma_{\alpha0}^{2})}\\
0 & -\frac{m_{r}^{2}}{2(\sigma_{\epsilon0,D_{r}}^{2}+m_{r}\sigma_{\alpha0}^{2})^{2}}\\
-\frac{1\left\{ D_{r}=1\right\} }{2\sigma_{\epsilon0,1}^{4}} & -\frac{1\left\{ D_{r}=1\right\} m_{r}}{2(\sigma_{\epsilon0,1}^{2}+m_{r}\sigma_{\alpha0}^{2})^{2}}\\
-\frac{1\left\{ D_{r}=2\right\} }{2\sigma_{\epsilon0,2}^{4}} & -\frac{1\left\{ D_{r}=2\right\} m_{r}}{2(\sigma_{\epsilon0,2}^{2}+m_{r}\sigma_{\alpha0}^{2})^{2}}
\end{array}\right).\label{eq:varphi-1}
\end{align}
Clearly, it follows that $E[\chi_{r}^{\ast}(\theta_{0})]=0$ by iterated
expectations. We note that the moment condition in (\ref{eq:Wald_L-O-M})
underlying the CV estimator is based on a linear transformation of
$\varphi^{*}\left(m_{r},D_{r}\right).$ Furthermore, as we shall see
in the next section, under the additional assumption that $\alpha$
and $\epsilon$ follow a Gaussian distribution, the score function
of the log likelihood for group $r$ is exactly the negative of $\chi_{r}^{\ast}(\theta)$,
that is
\[
\partial lnL_{r}(\theta_{0})/\partial\theta=-\chi_{r}^{\ast}(\theta_{0}),
\]
where $\ln L_{r}(\theta)$ denotes the log likelihood function for
group $r$ conditional on $\left(m_{1},...,m_{R},D_{1},...,D_{R}\right)$.
From these observations we see that the matrices $\varphi^{\ast}\left(m_{r},D_{r}\right)$
can be viewed to provide the optimal weighting for the basic moment
functions $\chi_{r}(\delta)$; compare also the corresponding discussion
for the general model for more details.
The result that $\partial\ln L_{r}(\theta_{0})/\partial\theta=-\chi_{r}^{\ast}(\theta_{0})$
for the score function under Gaussianity establishes the asymptotic
efficiency of the GMM estimator based on $E\left[\chi_{r}(\theta)|m_{r},D_{r}\right]=0$
under the assumption of Gaussian distributions for the unobservables.
When the unobservables are not Gaussian then the GMM estimator has
the interpretation of a quasi maximum likelihood estimator (QMLE).
Similarly, in \citet{lee_identification_2007} the score function
of the conditional maximum likelihood estimator (CMLE) for group $r$
is the optimal moment function corresponding to $E\left[\chi_{r}^{w}(\theta)|m_{r}\right]=0$
under the assumption of homoscedastic and normally distributed errors
$\epsilon_{ir}$. While the CMLE of \citet{lee_identification_2007}
is not efficient under the assumptions we postulate in this paper,
it shares robustness properties of within group panel estimators in
cases where the group effects are possibly correlated with covariates
in the model. Under those circumstances, random effects quasi maximum
likelihood estimators are generally not expected to be consistent.
Our discussion so far highlights variance as the source of identification,
with variation in either size $m_{r}$ or variance of the idiosyncratic
error terms $\sigma_{\epsilon,D_{r}}^{2}$ across groups as conditions.
We show that variation in group size and error term variance is a
source of identification in the QMLE, CVE and CMLE. In all, the CMLE
utilizes how the within-group variance changes with $\lambda$ and
size when error terms are homoscedastic, while the CVE exploits the
relationship between the within-group variance and between-group variance
in relation to $\lambda$ and size when there is either variation
in group size or heteroscedasticity across groups. Our QMLE uses both
pieces of information. All three estimators remain valid without covariates,
and may achieve identification as long as there are at least two different
group sizes in the limit in the case of homoscedasticity. This complements
other results in the literature. For example, Proposition 4 in \citet{bramoulle_identification_2009}
states that in the setting of \citet{lee_identification_2007}, $\lambda$
is identified by instrumenting $(I-W)WY$ with $(I-W)W^{2}Z$, $(I-W)W^{3}Z$,
etc., in line with the spatial literature on the estimation of Cliff-Ord
type models. Their result is due to the fact that they only exploit
restrictions for the conditional mean of $\epsilon$. In \citet{graham_identifying_2008}
as well as in this paper additional constraints on the distribution
of $\alpha$ and $\epsilon$ are imposed and shown to be useful in
the identification of peer effects. Under these conditions including
$Z$ offers additional sources of variation, but identification is
possible with or without it.
Adding covariates is critically important in empirical applications.
Consider adding the covariate matrix $Z$. This leads to two additional
moment conditions $E\left[\ddot{Z}_{r}^{\prime}\ddot{U}_{r}|m_{r},D_{r}\right]=0$
and $E\left[\bar{z}_{r}^{\prime}\bar{u}_{r}|m_{r},D_{r}\right]=0$,
where $\bar{z}_{r}=\iota_{m_{r}}^{\prime}Z_{r}/m_{r}$ is the group
mean of $Z_{r}$ and $\ddot{Z}_{r}=Z_{r}-\iota_{m_{r}}\bar{z}_{r}$
is the deviation from group mean. Moreover, $\ddot{Y}_{r}$ and $\bar{y}_{r}$
now need to be replaced by $\ddot{Y}_{r}-\frac{m_{r}-1}{m_{r}-1+\lambda}\ddot{Z}_{r}\beta$
and $\bar{y}_{r}-\frac{\bar{z}_{r}\beta}{1-\lambda}$ respectively.
The score function of the QMLE then is the same as the moment conditions
of the best GMM corresponding to these two moment functions in addition
to the moments $E\left[\chi_{r}(\theta)|m_{r},D_{r}\right]=0$. In
the same way, in the presence of covariates and assuming homoscedasticity
of $\epsilon$, Lee's CMLE estimator is based on $E\left[\ddot{Z}_{r}^{\prime}\ddot{U}_{r}\right]=0$
in addition to $E\left[\chi_{r}^{w}(\theta)|m_{r}\right]=0$ and the
relative efficiency considerations discussed in this section continue
to apply to the situation with covariates.
\section{General Model\label{sec:General-Model}}
In this section we generalize the model to allow for individual characteristics,
average individual characteristics of peers and group level covariates.
We assume that we have access to observations on $R$ groups belonging
to $J$ categories, where $1\leqslant J<\infty$ is fixed. We consider
asymptotics where the number of groups $R$ tends to infinity and
where the number of group sizes is finite. For the asymptotic identification
of $\lambda_{0}$ and $\sigma_{\alpha0}^{2}$ this setup assumes that
in the limit we observe infinitely many groups for at least two group
sizes or two categories, echoing the requirement of variation in either
group sizes or categories for identification in Section \ref{sec:graham}.
In designs that allow for heteroscedasticity, we also need infinitely
many groups for each category $j\in\left\{ 1,...,J\right\} $ to identify
the remaining variance parameters $\sigma_{\epsilon0,j}^{2}$. Let
$r=1,...,R$ denote the group index, let $D_{r}$ denote the category
of group $r$, and let $m_{r}$ denote the size of group $r$. The
total sample size is then given by $N=\sum_{r=1}^{R}m_{r}$. Suppose
further that interactions occur within each group, but not across
groups, and that peer effects work through the mean outcome and mean
characteristics of peers in the same group. The linear-in-means peer
effects model that includes endogenous as well as exogenous peer effects
then is given by
\begin{equation}
y_{ir}=\beta_{1}+\lambda\bar{y}_{(-i)r}+x_{1,ir}\beta_{2}+\bar{x}_{2,(-i)r}\beta_{3}+x_{3,r}\beta_{4}+\alpha_{r}+\epsilon_{ir},\label{eq:main_orig}
\end{equation}
where $y_{ir}$ is the outcome variable of individual $i$ in group
$r$, $\bar{y}_{(-i)r}=\frac{1}{m_{r}-1}\sum_{j\neq i}^{m_{r}}y_{jr}$
is the average outcome of $i$'s peers, $x_{1,ir}$ and $x_{2,ir}$
are both row vectors of predetermined characteristics of individual
$i$ in group $r$, $\bar{x}_{2,(-i)r}=\frac{1}{m_{r}-1}\sum_{j\neq i}^{m_{r}}x_{2,jr}$
is a vector of average characteristics of $i$'s peers, $x_{3,r}$
is a vector of observed group characteristics. The variables in $x_{1,ir}$
and $x_{2,ir}$ can be non-overlapping, partially overlapping or totally
overlapping. The error term consists of two components, the group
effect $\alpha_{r}$ and the disturbance term $\epsilon_{ir}$. We
treat $x_{1,ir}$, $x_{2,ir}$, $x_{3,r}$, $D_{r}$ and $m_{r}$
as non-stochastic, while noting that at the expense of more complex
notation we could also think of the analysis as being conditional
on these variables. In this model, peer effects work through the mean
peer outcome $\bar{y}_{(-i)r}$ and mean peer characteristics $\bar{x}_{2,(-i)r}$.
The two terms are also known as the leave-out-mean of $y$ and $x_{2}$,
as they are means of the group leaving out oneself. In Manski's terminology,
$\lambda\bar{y}_{(-i)r}$ in (\ref{eq:main_orig}) reflects endogenous
peer effects, and $\bar{x}_{2,(-i)r}\beta_{3}$ is the exogenous peer
effect, also referred to as contextual peer effects. The covariates
$\bar{x}_{2,(-i)r}$ and $x_{3,r}$ contain group level information
and can be interpreted as parametrizations of group level fixed effects
in the spirit of Mundlak (1978) and Chamberlain (1980). For example,
$x_{3,r}$ can contain full group averages of individual characteristics
or be composed of other characteristics that only vary at the group
level. The CMLE, as in the conventional panel case, cannot account
for this group level information. This can be a limitation in cases
where the effects of group level characteristics are of independent
interest in the analysis. An example is the effects of teacher education
and training on class test scores.
Let $z_{ir}=(1,x_{1,ir},\bar{x}_{2,(-i)r},x_{3,r})$ be the row vector
of all exogenous variables, let $\beta=(\beta_{1},\beta_{2}^{\prime},\beta_{3}^{\prime},\beta_{4}^{\prime})^{\prime}$
be the corresponding coefficients vector, and let $k_{Z}$ denote
the number of columns in $z_{ir}$. A compact form of model (\ref{eq:main_orig})
is
\begin{equation}
y_{ir}=\lambda\bar{y}_{(-i)r}+z_{ir}\beta+\alpha_{r}+\epsilon_{ir}.\label{eq:main}
\end{equation}
The model can be further written as a Cliff-Ord type spatial model.
To see this let $I_{m}$ denote the $m-$dimensional identity matrix,
let $\iota_{m}$ denote the $m-$dimensional column vector of ones,
and define the weight matrix $W_{m_{r}}$ for group $r$ as $W_{m_{r}}=\frac{1}{m_{r}-1}(\iota_{m_{r}}\iota_{m_{r}}^{\prime}-I_{m_{r}})$.
The off-diagonal elements of this matrix are all equal to $\frac{1}{m_{r}-1}$
and diagonal elements are 0. Let $Y_{r}=(y_{1r},...,y_{m_{r}r})^{\prime}$,
$Z_{r}=(z_{1r}^{\prime},...,z_{m_{r}r}^{\prime})^{\prime}$, $\epsilon_{r}=(\epsilon_{1r},...,\epsilon_{m_{r}r})^{\prime}$,
then the model for group $r$ can be expressed in matrix form as
\begin{equation}
Y_{r}=\lambda W_{m_{r}}Y_{r}+Z_{r}\beta+U_{r},\label{eq:main_matrix}
\end{equation}
where $U_{r}=\alpha_{r}\iota_{m_{r}}+\epsilon_{r}.$ Let $Y=[Y_{1}^{\prime},Y_{2}^{\prime},...,Y_{R}^{\prime}]^{\prime}$,
$Z=[Z_{1}^{\prime},Z_{2}^{\prime},...,Z_{R}^{\prime}]^{\prime}$,
$U=[U_{1}^{\prime},U_{2}^{\prime},...,U_{R}^{\prime}]^{\prime}$,
and $W=diag_{r=1}^{R}\{W_{m_{r}}\}$ such that the model for the whole
sample is given by
\begin{equation}
Y=\lambda WY+Z\beta+U.\label{eq:main_matrix_2}
\end{equation}
In the spatial literature $W$ is referred to as a spatial weight
matrix and $WY$ as a spatial lag. In analyzing the model in (\ref{eq:main_matrix_2})
we maintain the random effects specification detailed in Assumptions
1 and 2 of Section \ref{sec:graham}, which imply that $\alpha_{r}\sim(0,\sigma_{\alpha}^{2})$
and $\epsilon_{ir}\sim(0,\sigma_{\epsilon,D_{r}}^{2})$, where $D_{r}\in\{1,...,J\}$
with $J\geqslant1$ fixed and finite. The specification allows for
heteroscedasticity at the group level as long as there are only a
finite number of different parameters. For example, we could allow
for $\sigma_{\epsilon,D_{r}}^{2}$ to be different for small and large
groups, or more generally for all groups of a certain size $m_{r}$.
On the other hand we do not cover the case where $\sigma_{\epsilon,D_{r}}^{2}$
differs for each individual group $r,$ as this would lead to an infinite
dimensional parameter space.
The parameters of interest are $\lambda,\sigma_{\alpha}^{2},\sigma_{\epsilon,1}^{2},...,\sigma_{\epsilon,J}^{2}$
and $\beta$. Their respective true values are $\lambda_{0},\sigma_{\alpha0}^{2},\sigma_{\epsilon0,1}^{2},...,\sigma_{\epsilon0,J}^{2}$
and $\beta_{0}$. In analyzing the model it will be convenient to
concentrate the log-likelihood function with respect to $\beta$ for
given values of $\theta=(\lambda,\sigma_{\alpha}^{2},\sigma_{\epsilon1}^{2},...,\sigma_{\epsilon J}^{2})^{\prime}$.
Let $\Theta$ denote the parameter space for $\theta$ , and let $\delta=(\theta^{\prime},\beta^{\prime})^{\prime}$
denote the vector of all parameters. Under Assumptions 1 and 2 the
expression for the variance covariance matrix $\Omega_{0}$ of $U$
is
\[
\Omega_{0}=\Omega(\theta_{0})=\operatorname{diag}_{r=1}^{R}\Omega_{r}(\theta_{0})=\operatorname{diag}_{r=1}^{R}\{\sigma_{\epsilon0,D_{r}}^{2}I_{m_{r}}+\sigma_{\alpha0}^{2}\iota_{m_{r}}\iota_{m_{r}}^{\prime}\}.
\]
To define the quasi-maximum likelihood estimator (QMLE) for the peer
effects model in (\ref{eq:main}) note that solving $Y$ from (\ref{eq:main_matrix_2})
yields the reduced from:
\begin{equation}
Y=(I-\lambda W)^{-1}Z\beta+(I-\lambda W)^{-1}U.\label{eq:y_sol_2}
\end{equation}
If $\alpha_{r}$ and $\epsilon_{ir}$ follow normal distributions,
\begin{equation}
Y\sim N((I-\lambda W)^{-1}Z\beta,(I-\lambda W)^{-1}\Omega(\theta)(I-\lambda W^{\prime})^{-1}).\label{eq:y_dis}
\end{equation}
The corresponding log likelihood function is
\begin{align}
\ln L_{N}(\theta,\beta) & =-\frac{N}{2}\ln(2\pi)+\frac{1}{2}\ln|(I-\lambda W)^{2}\Omega(\theta)^{-1}|\nonumber \\
& -\frac{1}{2}(Y-\lambda WY-Z\beta)^{\prime}\Omega(\theta)^{-1}(Y-\lambda WY-Z\beta).\label{eq:logL}
\end{align}
and the corresponding QMLE is given by
\begin{equation}
\hat{\delta}_{N}=(\hat{\theta}_{N}^{\prime},\hat{\beta}_{N}^{\prime})^{\prime}=\textrm{argmax}_{\theta,\beta}\ln L_{N}(\theta,\beta).\label{eq:Def_QMLE}
\end{equation}
It is convenient to concentrate out $\beta$ and to obtain the QMLE
for $\theta$ first. The first order condition for $\beta$ is
\begin{equation}
\frac{\partial\ln L_{N}(\theta,\beta)}{\partial\beta}=(Y-\lambda WY-Z\beta)^{\prime}\Omega(\theta)^{-1}Z=0,\label{eq:focdelta}
\end{equation}
which leads to
\begin{equation}
\hat{\beta}_{N}(\theta)=(Z^{\prime}\Omega(\theta)^{-1}Z)^{-1}Z^{\prime}\Omega(\theta)^{-1}(I-\lambda W)Y.\label{eq:gammahat}
\end{equation}
Plugging $\hat{\beta}_{N}(\theta)$ back into (\ref{eq:logL}) yields
the following concentrated log likelihood function,
\begin{align}
Q_{N}(\theta) & =\frac{1}{N}\ln L_{N}(\theta,\hat{\beta}_{N}(\theta))\nonumber \\
& =-\frac{\ln(2\pi)}{2}+\frac{1}{2N}\ln|(I-\lambda W)^{2}\Omega(\theta)^{-1}|-\frac{1}{2N}Y^{\prime}(I-\lambda W)^{\prime}M_{Z}(\theta)(I-\lambda W)Y,\label{eq:QN}
\end{align}
where
\begin{equation}
M_{Z}(\theta)=\Omega(\theta)^{-1}-\Omega(\theta)^{-1}Z(Z^{\prime}\Omega(\theta)^{-1}Z)^{-1}Z^{\prime}\Omega(\theta)^{-1}.\label{eq:Mz}
\end{equation}
Then the QMLE for $\theta,$ $\hat{\theta}_{N}=(\hat{\lambda}_{N},\hat{\sigma}_{\alpha,N}^{2},\hat{\sigma}_{\epsilon1,N}^{2},...,\hat{\sigma}_{\epsilon J,N}^{2})^{\prime}$
is given by
\begin{equation}
\hat{\theta}_{N}=\textrm{argmax}_{\theta}Q_{N}(\theta).\label{eq:varthetahat}
\end{equation}
Plugging $\hat{\theta}_{N}$ back into \eqref{eq:gammahat}, the QMLE
for $\beta$ is
\begin{equation}
\hat{\beta}_{N}=\hat{\beta}_{N}(\hat{\theta}_{N})=(Z^{\prime}\Omega(\hat{\theta}_{N})^{-1}Z)^{-1}Z^{\prime}\Omega(\hat{\theta}_{N})^{-1}(I-\hat{\lambda}_{N}W)Y.\label{eq:deltahathat1}
\end{equation}
A formal result regarding the asymptotic identification of the model
parameters is given in the next section. We next provide some intuition
for that result, by extending our earlier discussion of identification
for the canonical model without covariates to our model (\ref{eq:main_matrix_2})
with covariates. Let $\left\Vert .\right\Vert $ be the Euclidean
norm on $\mathbb{R^{\text{\ensuremath{k}}}}.$ Using the relationships
$\ddot{U}_{r}=\frac{\left(m_{r}-1+\lambda_{0}\right)}{m_{r}-1}\ddot{Y}_{r}-\ddot{Z}_{r}\beta_{0}$
and $\bar{u}_{r}=(1-\lambda_{0})\bar{y}_{r}-\bar{z}_{r}\beta_{0}$,
the moment functions related to the full model can be written as follows
\[
\chi_{r}(\theta)=\left[\begin{array}{c}
\chi_{r}^{w}(\theta)\\
\chi_{r}^{b}(\theta)\\
\chi_{r}^{zw}\left(\delta\right)\\
\chi_{r}^{zb}\left(\delta\right)
\end{array}\right]=\left[\begin{array}{c}
\left\Vert \frac{m_{r}-1+\lambda}{m_{r}-1}\ddot{Y}_{r}-\ddot{Z}_{r}\beta\right\Vert ^{2}-(m_{r}-1)\sigma_{\epsilon,D_{r}}^{2}\\{}
[(1-\lambda)\bar{y}_{r}-\bar{z}_{r}\beta]^{2}-\sigma_{\alpha}^{2}-\frac{\sigma_{\epsilon,D_{r}}^{2}}{m_{r}}\\
\ddot{Z}_{r}^{\prime}\left(\frac{m_{r}-1+\lambda}{m_{r}-1}\ddot{Y}_{r}-\ddot{Z}_{r}\beta\right)\\
\bar{z}_{r}^{\prime}\left((1-\lambda)\bar{y}_{r}-\bar{z}_{r}\beta\right)
\end{array}\right]
\]
where $\chi_{r}^{w}(\delta)$ and $\chi_{r}^{b}(\delta)$ summarize
the restrictions on the unobservables, and are natural extensions
of the moment conditions considered before in \eqref{eq:mwb} for
the model without covariates. The additional moment restrictions $\chi_{r}^{zw}\left(\delta\right)$
and $\chi_{r}^{zb}\left(\delta\right)$ relate to the exogeneity of
$Z_{r}$ relative to $\epsilon_{r}$ and $\alpha_{r}$. A formal asymptotic
identification result will be given in the next section. Intuitively,
for given $\lambda$ the last two moment conditions identify $\beta$,
while the first two identify $\lambda,\sigma_{\alpha}^{2},\sigma_{\varepsilon,j}^{2},j=1,...,J$
in an analogous manner as described in the discussion of Lemma \ref{lem:ID_1}
for the model without covariates.
As for the model without covariates there is a representation of the
score of the log-likelihood in terms of the fundamental moment conditions.
To describe the relationship between moments and the score we define
the matrix
\begin{align}
\varphi(m_{r},D_{r}) & =\left(\begin{array}{cccc}
\frac{1}{(m_{r}-1+\lambda_{0})\sigma_{\epsilon0,D_{r}}^{2}} & -\frac{m_{r}}{(1-\lambda_{0})(\sigma_{\epsilon0,D_{r}}^{2}+m_{r}\sigma_{\alpha0}^{2})} & \frac{1}{(m_{r}-1+\lambda_{0})\sigma_{\epsilon0,D_{r}}^{2}}\beta_{0}^{\prime} & -\frac{m_{r}}{(1-\lambda_{0})(\sigma_{\epsilon0,D_{r}}^{2}+m_{r}\sigma_{\alpha0}^{2})}\beta_{0}^{\prime}\\
0 & -\frac{m_{r}^{2}}{2(\sigma_{\epsilon0,D_{r}}^{2}+m_{r}\sigma_{\alpha0}^{2})^{2}} & 0 & 0\\
-\frac{1(D_{r}=1)}{2\sigma_{\epsilon0,1}^{4}} & -\frac{m_{r}1(D_{r}=1)}{2(\sigma_{\epsilon0,1}^{2}+m_{r}\sigma_{\alpha0}^{2})^{2}} & 0 & 0\\
\vdots & \vdots & \vdots & \vdots\\
-\frac{1(D_{r}=J)}{2\sigma_{\epsilon0,J}^{4}} & -\frac{m_{r}1(D_{r}=J)}{2(\sigma_{\epsilon0,J}^{2}+m_{r}\sigma_{\alpha0}^{2})^{2}} & 0 & 0\\
0 & 0 & -\frac{1}{\sigma_{\epsilon0,D_{r}}^{2}}I_{k_{Z}} & -\frac{m_{r}}{\sigma_{\epsilon0,D_{r}}^{2}+m_{r}\sigma_{\alpha0}^{2}}I_{k_{Z}}
\end{array}\right).\label{eq:varphim}
\end{align}
Furthermore observe that the log-likelihood function can be written
as $\ln L_{N}(\delta)=-\frac{N}{2}\ln(2\pi)+\sum_{r=1}^{R}\ln L_{r}(\delta)$
where
\begin{align*}
\ln L_{r}(\delta) & =\frac{1}{2}\ln|(I_{m_{r}}-\lambda W_{m_{r}})^{2}\Omega_{r}(\theta)^{-1}|\\
& -\frac{1}{2}(Y_{r}-\lambda W_{m_{r}}Y_{r}-Z_{r}\beta)^{\prime}\Omega_{r}(\theta)^{-1}(Y_{r}-\lambda W_{m_{r}}Y-Z_{r}\beta)
\end{align*}
is the log-likelihood function for group $r$. Then it can be shown
that\footnote{See our Online Appendix for details. The derivation uses Lemma \ref{lemma:moment}
and the special properties of matrices $\Omega(\theta)$, $I-\lambda W$
and $W$ described in Appendix \ref{subsec:matrix properties}. In
the Online Appendix we also give an explicit expression for the variance
covariance matrix of $\chi_{r}(\delta).$}
\begin{eqnarray*}
\frac{\partial\ln L_{r}(\delta_{0})}{\partial\delta} & = & -\chi_{r}^{\ast}(\delta_{0})=-\varphi(m_{r},D_{r})\chi_{r}\left(\delta_{0}\right).
\end{eqnarray*}
As is well known, the score of the log-likelihood function, $S(\delta)=-\sum_{r=1}^{R}\frac{\partial\ln L_{r}(\delta)}{\partial\delta}$
can be interpreted as a moment function corresponding to the moments
$E\left[S(\delta_{0})\right]=-\sum_{r=1}^{R}E\left[\frac{\partial\ln L_{r}(\delta_{0})}{\partial\delta}\right]=0$.
Furthermore, under a Gaussian assumption the score is an optimal moment
function.\footnote{Observe that
\begin{align*}
E\left[\frac{\partial S(\delta_{0})}{\partial\delta'}\right]\left[Var\left(S(\delta_{0})\right)\right]^{-1}S(\delta_{0}) & =\left(-\sum_{r=1}^{R}E\left[\frac{\partial^{2}lnL_{r}(\delta_{0})}{\partial\delta\partial\delta'}\right]\right)\left(\sum_{r=1}^{R}E\left[\frac{\partial lnL_{r}(\delta_{0})}{\partial\delta}\frac{\partial lnL_{r}(\delta_{0})}{\partial\delta'}\right]\right)^{-1}S(\delta_{0})\\
& =S(\delta_{0}).
\end{align*}
in light of the information matrix equality.} From this we see that the matrices $\varphi(m_{r},D_{r})$ can be
viewed to provide the optimal weighting for the basic moment functions
$\chi_{r}(\delta)$. Under Gaussian assumptions the optimal GMM estimator
coincides with the maximum likelihood estimator and is asymptotically
efficient under the stated assumptions.
\section{Theoretical Results\label{sec:Theoretical-Results}}
We next state our assumptions for the general model. We maintain Assumptions
\ref{assume:epsilon}-\ref{assume:lambda} on $\epsilon$, $\alpha$
and $\lambda$. In the following we add assumptions regarding the
exogenous variables, and the sizes and relative magnitudes of groups
in the sample. Let $\mathcal{I}_{m,j}\subset\{1,...,R\}$ be the index
set of all groups in category $j$ with size equal to $m$. Thus if
$r\in\mathcal{I}_{m,j}$, then $D_{r}=j$ and $m_{r}=m$. Let $R_{m,j}$
be the cardinality of $\mathcal{I}_{m,j}$, in other words $R_{m,j}$
is the number of groups in category $j$ with size equal to $m$,
and let $R_{j}$ be the number of groups in category $j$, that is
$R_{j}=\sum_{r=1}^{R}1(D_{r}=j)=\sum_{m=2}^{\bar{M}}R_{m,j}$, where
the upper bound\textcolor{red}{{} }\textcolor{black}{$\bar{M}$} on
the group size is specified in the next assumption below. Furthermore
let $\omega_{m,j}=R_{m,j}/R$ denote the share of groups in category
$j$ with size equal to $m$, and let $\omega_{j}=R_{j}/R=\sum_{m=2}^{\bar{M}}\omega_{m,j}$
be the share of groups in category $j$. Below we maintain the following
assumption regarding the group sizes and their relative magnitudes.
\begin{assumption}
\textcolor{black}{\label{assume:n}(a) The sample size $N$ goes to
infinity; (b) The group size is bounded in the sense that there exists
some positive constant $\bar{M}$ such that $2\leqslant m_{r}\leqslant\bar{M}<\infty$
}\textup{for}\textcolor{black}{{} $r=1,2,...,R$; (c) The limit $\omega_{m,j}^{\ast}=\lim_{N\rightarrow\infty}\omega_{m,j}$
exists and $\omega_{m,j}^{\ast}<1$ for all $2\leqslant m\leqslant\bar{M}$
and $j$, and $\omega_{j}^{\ast}=\lim_{N\rightarrow\infty}\omega_{j}=\sum_{m=2}^{\bar{M}}\omega_{m,j}^{*}>0$
for all $j$.}
\end{assumption}
The restriction that the minimal group size is 2 rules out singleton
groups. A member of such a group has no peers. Assumption \ref{assume:n}(b)
imposes a fixed upper bound on group size. In many applications this
is not a serious constraint. The assumption is more restrictive than
Lee (2007) who allows for group size to grow with sample size. It
is worth pointing out that increasing group sizes generally reduce
the convergence rates for estimators of peer effects parameters, and
as demonstrated by \citet{kelejian_2sls_2002} in some cases lead
to inconsistency of these estimators.
Assumption \ref{assume:n}(c) states that asymptotically, no single
type-group size combination can dominate the sample by requiring that
\textcolor{black}{$\omega_{m,j}^{\ast}<1$ for all $2\leqslant m\leqslant\bar{M}$
and $j$}. In addition, all types $j$ occur in the sample in an asymptotically
non-negligible way because \textcolor{black}{$\omega_{j}^{\ast}>0$
for all $j$. On the other hand, we do allow that for certain combinations
of $j$ and $m$ the limit }$\omega_{m,j}^{\ast}$ is zero, allowing
for some group sizes of type $j$ to occur infrequently or not at
all in the sample.
Observe that $N=\sum_{r=1}^{R}m_{r}=\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}mR_{m,j}$.
Since group size is bounded, the number of groups $R$ goes to infinity
as $N$ goes to infinity. Since $\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}R_{m,j}=R$,
we have $\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}\omega_{m,j}=\sum_{j=1}^{J}\omega_{j}=1$
and thus $\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}\omega_{m,j}^{\ast}=\sum_{j=1}^{J}\omega_{j}^{\ast}=1$.
Since $\omega_{j}^{\ast}>0$ \textcolor{black}{by As}sumption \ref{assume:n}(c)
it follows that also $R_{j}$ goes to infinity, which is needed to
facilitate the consistent estimation of \textcolor{black}{$\sigma_{\epsilon,j}^{2}$}.
Assumption \ref{assume:n}(c) implies that the limit of the average
group size is given by
\begin{align}
m^{\ast} & =\lim_{N\rightarrow\infty}\frac{N}{R}=\lim_{N\rightarrow\infty}\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}\frac{R_{m,j}}{R}m=\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}\omega_{m,j}^{\ast}m.\label{eq:nstar-1}
\end{align}
Clearly $2\leqslant m^{\ast}\leqslant\bar{M}$, since $2\leqslant m_{r}\leqslant\bar{M}$.
Observe that in light of Assumptions \ref{assume:epsilon}, \ref{assume:alpha},
and \ref{assume:lambda} the parameter space $\Theta$ for $\theta=\left(\lambda,\sigma_{\alpha}^{2}\sigma_{\epsilon,1}^{2},...,\sigma_{\epsilon,J}^{2}\right)^{\prime}$
is a compact subset of the Euclidean space $\mathbb{R}^{2+J}$. Observe
further that
\begin{align}
I_{m_{r}}-\lambda W_{m_{r}} & =(1+\frac{\lambda}{m_{r}-1})I_{m_{r}}^{\ast}+(1-\lambda)J_{m_{r}}^{\ast},
\end{align}
where $I_{m_{r}}^{\ast}=I_{m_{r}}-\iota_{m_{r}}\iota_{m_{r}}^{\prime}/m_{r}$
and $J_{m_{r}}^{\ast}=\iota_{m_{r}}\iota_{m_{r}}^{\prime}/m_{r}$
are symmetric, idempotent, orthogonal, and sum to the identity matrix.
Furthermore from the results in Appendix \ref{subsec:matrix properties}
we have $|I_{m_{r}}-\lambda W_{m_{r}}|=[1+\lambda/(m_{r}-1)]^{m_{r}-1}(1-\lambda)$.\footnote{In Appendix \ref{subsec:matrix properties} we review additional properties
of matrices of the form $pI_{m}^{\ast}+sJ_{m}^{\ast}$ , which will
be used repeatedly in this paper. In particular, their multiplication
is commutative. The products of such matrices are also of the form
of $pI_{m}^{\ast}+sJ_{m}^{\ast}$, and $|pI_{m}^{\ast}+sJ_{m}^{\ast}|=p^{m-1}s$,
$(pI_{m}^{\ast}+sJ_{m}^{\ast})^{-1}=\frac{1}{p}I_{m}^{\ast}+\frac{1}{s}J_{m}^{\ast}$.} Thus the matrix $I_{m_{r}}-\lambda W_{m_{r}}$ is nonsingular if
$1+\lambda/(m_{r}-1)\neq0$ and $1-\lambda\neq0$. Assumption \ref{assume:lambda}
ensures the non-singularity of $I_{m_{r}}-\lambda W_{m_{r}}$, and
hence the non-singularity of $I-\lambda W=\operatorname{diag}_{r=1}^{R}\{I_{m_{r}}-\lambda W_{m_{r}}\}$,
since for $m_{r}\geqslant2$ and $\lambda<1$ we have $1+\lambda/(m_{r}-1)>0$
and $1-\lambda>0$.
Let $\bar{z}_{r}=\frac{1}{m_{r}}\iota_{m_{r}}^{\prime}Z_{r}$ be the
row vector of column means of $Z_{r}$, and let $\ddot{Z}_{r}=Z_{r}-\iota_{m_{r}}\bar{z}_{r}$
be the deviations from the column means. Then $Z_{r}^{\prime}I_{m_{r}}^{\ast}Z_{r}=\ddot{Z}_{r}^{\prime}\ddot{Z}_{r}$,
$Z_{r}^{\prime}J_{m_{r}}^{\ast}Z_{r}=m_{r}\bar{z}_{r}^{\prime}\bar{z}_{r}$.
\begin{assumption}
\label{assume:z}(a) The $N\times k_{Z}$ matrix $Z$ is non-stochastic,
with $rank(Z)=k_{Z}>0$ for $N$ sufficiently large. The elements
of $Z$ are uniformly bounded in absolute value.
(b)For $2\leqslant m\leqslant\bar{M}$, and $1\leqslant j\leqslant J$
the following limits exist:
\[
\lim_{N\rightarrow\infty}N^{-1}\sum_{r\in\mathcal{I}_{m,j}}\ddot{Z}_{r}^{\prime}\ddot{Z}_{r}=\ddot{\varkappa}_{m,j},
\]
\[
\lim_{N\rightarrow\infty}N^{-1}\sum_{r\in\mathcal{I}_{m,j}}m\bar{z}_{r}^{\prime}\bar{z}_{r}=\bar{\varkappa}_{m,j},
\]
\[
\lim_{N\rightarrow\infty}N^{-1}\sum_{r\in\mathcal{I}_{m,j}}\bar{z}_{r}=\bar{z}_{m,j}.
\]
(c) For at least one pair of $(m,j)$ such that $\omega_{m,j}^{\ast}>0$,
and N sufficiently large, the smallest eigenvalues of $N^{-1}\sum_{r\in\mathcal{I}_{m,j}}Z_{r}^{\prime}Z_{r}=N^{-1}\sum_{r\in\mathcal{I}_{m,j}}\ddot{Z}_{r}^{\prime}\ddot{Z}_{r}+N^{-1}\sum_{r\in\mathcal{I}_{m,j}}m\bar{z}_{r}^{\prime}\bar{z}_{r}$
are bounded away from zero, uniformly in N, by some finite constant
$\underline{\xi}_{Z}>0$.
\end{assumption}
Suppose we have some $N\times N$ matrix $A_{N}(\theta)=diag_{r=1}^{R}\{p(m_{r},D_{r},\theta)I_{m_{r}}^{\ast}+s(m_{r},D_{r},\theta)J_{m_{r}}^{\ast}\}$,
where $p(m_{r},D_{r},\theta)$ and $s(m_{r},D_{r},\theta)$ are positive,
uniformly continuous and bounded on $\Theta$. An example of an expression
of this form is $\Omega(\theta)^{-1}$ which is obtained in closed
form in Equation \eqref{AppCon1} in Appendix \ref{subsec:matrix properties}.
Then under Assumption \ref{assume:z}(b), the limiting matrix of $N^{-1}Z^{\prime}A_{N}(\theta)Z$
always exists, is continuous in $\theta$ and takes the form
\[
\lim_{N\rightarrow\infty}\frac{1}{N}Z^{\prime}A_{N}(\theta)Z=\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}[p(m,j,\theta)\ddot{\varkappa}_{m,j}+s(m,j,\theta)\bar{\varkappa}_{m,j}].
\]
Furthermore, $N^{-1}Z^{\prime}A_{N}(\theta)Z$ converges to its limiting
matrix uniformly on $\Theta$. With $p(m_{r},D_{r},\theta)>0$ and
$s(m_{r},D_{r},\theta)>0$, Assumption \ref{assume:z}(a) ensures
that $N^{-1}Z^{\prime}A_{N}(\theta)Z$ and its limiting matrix are
invertible, with the elements of the inverse matrix uniformly bounded
in absolute value. In the special case when $A_{N}(\theta)$ is the
identity matrix, $lim_{N\rightarrow\infty}\frac{1}{N}Z^{\prime}Z=\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}[\ddot{\varkappa}_{m,j}+\bar{\varkappa}_{m,j}]$,
which has the smallest eigenvalue bounded above zero by some finite
constant $\text{\ensuremath{\underbar{\ensuremath{\xi}}}}_{Z}>0$.
See Lemma \ref{lem:Aux1} for details and a proof.
As shown by Lemma \ref{lem:ID_1} in Section \ref{sec:graham}, identification
of $\lambda$ and $\sigma_{\alpha}^{2}$ requires variation in the
group size or variance of the error terms. The following assumption
ensures this so that in the limit we have non-negligible samples for
at least two different group sizes or two different categories with
different variances of the idiosyncratic errors $\epsilon_{ir}$.
\begin{assumption}
\label{assume:id}For some sizes $m$ and $m^{\prime}$, and some
categories $j$ and $j^{\prime}$ we have $\omega_{m,j}^{\ast}>0$
and $\omega_{m^{\prime},j^{\prime}}^{\ast}>0$, and either of the
following two scenarios hold,
(a) $m\neq m^{\prime}$, and $\sigma_{\epsilon0,j}^{2}=\sigma_{\epsilon0,j^{\prime}}^{2}$
for some $j,j^{\prime}\in\left\{ 1,...,J\right\} $.
(b) $m=m^{\prime}$, and $\sigma_{\epsilon0,j}^{2}\neq\sigma_{\epsilon0,j^{\prime}}^{2}$
for some $j,j^{\prime}\in\left\{ 1,...,J\right\} $ with $j\neq j^{\prime}.$
\end{assumption}
The conditions in Assumption \ref{assume:id} are the asymptotic analogs
of identification conditions imposed in Lemma \ref{lem:ID_1}. Assumption
\ref{assume:n} by itself is not sufficient for identification because
it only implies that no single pair $\left(m,j\right)$ asymptotically
dominates the sample. Assumption \ref{assume:n} alone does not guarantee
that there is enough variation in the underlying group sizes $m$
or the variances $\sigma_{\epsilon0,j}^{2}$. For example, it is possible
under Assumption \ref{assume:n} that all groups are of the same size
and that all variances $\sigma_{\epsilon0,j}^{2}$ are the same. Assumption
\ref{assume:id} rules out such cases. Assumption \ref{assume:id}(a)
is related to Assumption 6.1 and Footnote 9 of \citet{lee_identification_2007}
which requires group size variation to achieve identification for
the case where group sizes are bounded, the only case we consider.
Assumption \ref{assume:id}(b) has no analog in Lee (2007) because
of his Assumption 1 which imposes homoscedasticity on the errors $\epsilon_{ir}.$
We show that identification is possible purely based on group level
heteroscedasticity even if all group sizes are the same. This insight
also extends the analysis of \citet{graham_identifying_2008} where
types and class sizes are linked.
Below we give results on the consistency and asymptotic normality
of the QMLE $\hat{\delta}_{N}=(\hat{\theta}_{N}^{\prime},\hat{\beta}_{N}^{\prime})^{\prime}$
defined in (\ref{eq:Def_QMLE}).
\begin{thm}
\label{theorem:Consistency}Suppose Assumptions 1-6 hold, then
(a) The parameter $\delta_{0}$ is asymptotically identified in the
sense that it is the unique maximizer of the criterion $\bar{R}(\theta,\beta)=\lim_{N\rightarrow\infty}E\left[\frac{1}{N}\textrm{ln}L(\theta,\beta)\right]$.
(b) The QMLE $\widehat{\delta}_{N}$ is consistent, i.e., $\widehat{\delta}_{N}\overset{p}{\rightarrow}\delta_{0}$
as $N\rightarrow\infty$.
\end{thm}
A detailed proof of the theorem is given in Appendices \ref{sec:Proof-of-cons}
and \ref{sec:Proof-of-cons-2}. As can be seen from the proof, the
argumentation that ensures part (a) of the theorem is analogous to
the argumentation used in establishing Lemma 2.1. Here is a sketch
of the proof to provide some intuition. The limiting expected value
of the concentrated log likelihood function $Q_{N}(\theta)$ is
\[
\bar{Q}^{\ast}(\theta)=C^{\ast}+\frac{1}{2m^{\ast}}\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}\omega_{m,j}^{\ast}g(m,j,\theta)+Q^{(2)\ast}(\theta),
\]
where $C^{\ast}$ is a constant term, $g(m,j,\theta)=\ln|G(m,j,\theta)|-\operatorname{tr} G(m,j,\theta)$
with
\[
G(m,j,\theta)=\frac{\sigma_{\epsilon0,j}^{2}}{\sigma_{\epsilon,j}^{2}}\left(\frac{m-1+\lambda}{m-1+\lambda_{0}}\right)^{2}I_{m}^{\ast}+\frac{(\sigma_{\epsilon0,j}^{2}+m\sigma_{\alpha0}^{2})}{(\sigma_{\epsilon,j}^{2}+m\sigma_{\alpha}^{2})}\left(\frac{1-\lambda}{1-\lambda_{0}}\right)^{2}J_{m}^{\ast},
\]
and $Q^{(2)\ast}(\theta)=\lim_{N\rightarrow\infty}\bar{Q}_{N}^{(2)}(\theta)$
where $\bar{Q}_{N}^{(2)}(\theta)=-\frac{1}{2N}\tilde{\eta}_{Z}(\theta)^{\prime}\tilde{M}_{Z}(\theta)\tilde{\eta}_{Z}(\theta)$
with $\tilde{M}_{Z}(\theta)=I-\Omega(\theta)^{-1/2}Z(Z^{\prime}\Omega(\theta)^{-1}Z)^{-1}Z^{\prime}\Omega(\theta)^{-1/2}$
and
\[
\tilde{\eta}_{Z}(\theta)=\Omega(\theta)^{-1/2}(I-\lambda W)(I-\lambda_{0}W)^{-1}Z\beta_{0}.
\]
It is easy to see that $\theta_{0}$ is a global maximizer of $Q^{(2)\ast}(\theta)$,
given that $-\bar{Q}_{N}^{(2)}\left(\theta\right)$ is the quadratic
form of an idempotent and thus positive semi-definite matrix, and
$Q^{(2)\ast}(\theta_{0})=0$. However, this does not ensure that $\theta_{0}$
is a unique global maximizer. Identification thus comes from $\sum_{j=1}^{J}\sum_{m=2}^{\bar{M}}\omega_{m,j}^{\ast}g(m,j,\theta)$.
Note that for any symmetric positive definite $m\times m$ matrix
$A$, $\ln|A|-\operatorname{tr}(A)\leqslant-m$ with equality if and only if $A$
is an identity matrix.\footnote{\label{footnote:max}To see this, note that under the maintained assumptions
the eigenvalues of $A$, say, $\lambda_{i}$, are positive and $\ln\left\vert A\right\vert -\operatorname{tr}\left[A\right]=\sum_{i=1}^{m}\left[\ln(\lambda_{i})-\lambda_{i}\right]$.
The claim is seen to hold by observing that the function $f(x)=\ln(x)-x\leq-1$
for $x\in(0,\infty)$ with a unique maximum at $x=1$, and observing
that $A=I_{m}$ if and only if $\lambda_{i}=1$ for $i=1,\ldots,m$.} For any $m$ and $j$, $g(m,j,\theta)$ is maximized if and only
if $G(m,j,\theta)=I_{m}$, which is equivalent to $E\left[\chi_{r}(\theta)|m_{r}=m,D_{r}=j\right]=0$
with $\chi_{r}(\theta)=(\chi_{r}^{w}(\theta),\chi_{r}^{b}(\theta))$
defined in \eqref{eq:mwb}. It now follows from an asymptotic analogue
of Lemma \ref{lem:ID_1} that in either case (i) or (ii) of Assumption
\ref{assume:id}, $\theta_{0}$ is the only solution to $E\left[\chi_{r}(\theta)|m_{r}=m,D_{r}=j\right]=0$
and $E\left[\chi_{r}(\theta)|m_{r}=m^{\prime},D_{r}=j^{\prime}\right]=0$.
Thus for any $\theta\neq\theta_{0}$,
\[
\min\left(g(m,j,\theta_{0})-g(m,j,\theta),g(m^{\prime},j^{\prime},\theta_{0})-g(m^{\prime},j^{\prime},\theta)\right)>0.
\]
As a result, $\theta_{0}$ is the unique global maximizer of $\bar{Q}^{\ast}(\theta)$
when one of the two scenarios holds true for some $\omega_{m,j}^{\ast}>0$
and $\omega_{m^{\prime},j^{\prime}}^{\ast}>0$.
To study the asymptotic distribution of the estimator, first note
that under Assumptions \ref{assume:epsilon} and \ref{assume:alpha},
the third and fourth moments of $\epsilon_{ir}$ and $\alpha_{r}$
exist. Let $E\left[\epsilon_{ir}^{3}|D_{r}=j\right]=\mu_{\epsilon0,j}^{(3)}$,
$E\left[\epsilon_{ir}^{4}|D_{r}=j\right]=\mu_{\epsilon0,j}^{(4)}$,
$E\left[\alpha_{r}^{3}\right]=\mu_{\alpha0}^{(3)}$ and $E\left[\alpha_{r}^{4}\right]=\mu_{\alpha0}^{(4)}$.
Also, define $\Gamma_{0}$ and $\Upsilon_{0}$ as
\begin{eqnarray*}
\Gamma_{0} & = & \lim_{N\rightarrow\infty}N^{-1}E\left[-\frac{\partial^{2}lnL_{N}(\delta_{0})}{\partial\delta\partial\delta^{\prime}}\right],\\
\Upsilon_{0} & = & \lim_{N\rightarrow\infty}N^{-1}E\left[\frac{\partial lnL_{N}(\delta_{0})}{\partial\delta}\frac{\partial lnL_{N}(\delta_{0})}{\partial\delta^{\prime}}\right].
\end{eqnarray*}
As shown in Appendix \ref{sec:Proof-of-asydis}, the two limiting
matrices exist. Specific expressions are given in Appendix \ref{sec:Variance-Covariance-Matrix}.
When $\epsilon_{ir}$ and $\alpha_{r}$ both follow normal distributions,
$\Upsilon_{0}=\Gamma_{0}$.
The next lemma shows that $\Gamma_{0}$ is p.d. under the maintained
assumptions. The lemma also provides a sufficient condition on the
moments of $\varepsilon$ under which $\Upsilon_{0}$ is p.d..
\begin{lem}
\label{lem:nd} Suppose Assumptions \ref{assume:epsilon}-\ref{assume:id}
hold, then $\Gamma_{0}$ is positive definite. Under the additional
assumption that $\mu_{\varepsilon0,j}^{(4)}-\sigma_{\varepsilon0,j}^{4}>(\mu_{\epsilon0,j}^{(3)})^{2}/\sigma_{\varepsilon0,j}^{2}$
for all $j\in\{1,...,J\}$, $\Upsilon_{0}$ is also positive definite.
\end{lem}
The proof of the lemma is in Appendix \ref{sec:Variance-Covariance-Matrix}.
Note that from Holder's inequality we have $\mu_{\varepsilon0,j}^{(4)}-\sigma_{\varepsilon0,j}^{4}\geqslant(\mu_{\epsilon0,j}^{(3)})^{2}/\sigma_{\varepsilon0,j}^{2}$.
The sufficient condition is mild in that it only postulates that the
inequality holds strongly. Of course, the condition holds, e.g., for
the Gaussian distribution.
With both $\Upsilon_{0}$ and $\Gamma_{0}$ ensured to be positive
definite, we have the following theorem.
\begin{thm}
\label{theorem:AsymptoticNormlity}Under Assumptions \ref{assume:epsilon}-\ref{assume:id},
and assuming that \textup{$\delta_{0}$ is in the interior of the
parameter space $\Theta$ defined in Assumption \ref{assume:lambda}
and that} $\mu_{\varepsilon0,j}^{(4)}-\sigma_{\varepsilon0,j}^{4}>(\mu_{\epsilon0,j}^{(3)})^{2}/\sigma_{\varepsilon0,j}^{2}$
for $j\in\{1,...,J\}$, we have $\sqrt{N}(\hat{\delta}_{N}-\delta_{0})\xrightarrow{d}N(0,\Gamma_{0}^{-1}\Upsilon_{0}\Gamma_{0}^{-1})$
as $N\rightarrow\infty$.
\end{thm}
The proof of the theorem is given in Appendix \ref{sec:Proof-of-asydis}.
We next discuss consistent estimators for the matrices $\Gamma_{0}$
and $\Upsilon_{0}$ composing the asymptotic variance covariance matrix.
An inspection shows that $\Gamma_{0}=\Gamma(\delta_{0},s_{0})$ and
$\Upsilon_{0}=\Upsilon(\delta_{0},\mu_{\alpha0}^{(3)},\mu_{\alpha0}^{(4)},\mu_{\epsilon0,1}^{(3)},...,\mu_{\epsilon0,J}^{(3)},\mu_{\epsilon0}^{(4)}...,\mu_{\epsilon0,J}^{(4)},s_{0})$,
with $s_{0}=(s_{0,1},...,s_{0,J},m^{\ast})$ and
\[
s_{0,j}=[\ddot{\varkappa}_{2,j},...\ddot{\varkappa}_{\bar{M},j},\bar{\varkappa}_{2,j},...,\bar{\varkappa}_{\bar{M},j},\bar{z}_{2,j},...,\bar{z}_{\bar{M},j},\omega_{2,j}^{\ast},...,\omega_{\bar{M},j}^{\ast}],
\]
and where the functions $\Gamma(.)$ and $\Upsilon(.)$ are continuous.
Since the functions $\Gamma(.)$ and $\Upsilon(.)$ are continuous,
consistent estimators for $\Gamma_{0}$ and $\Upsilon_{0}$ can be
readily obtained by replacing the arguments of those functions by
consistent estimators thereof. Let $\hat{s}_{N}$ be the sample analogue
of $s_{0}$, then clearly $\hat{s}_{N}\overset{p}{\rightarrow}s_{0}$
in light of Assumptions \ref{assume:n} and \ref{assume:z}. Recall
further that by Theorem \ref{theorem:Consistency} the QMLE estimator
$\hat{\delta}_{N}$ is consistent for $\delta_{0}$, and suppose we
have consistent estimators for $\mu_{\alpha0}^{(3)}$, $\mu_{\alpha0}^{(4)}$,
$\mu_{\epsilon0,1}^{(3)},...,\mu_{\epsilon0,J}^{(3)}$, and $\mu_{\epsilon0,1}^{(4)}...,\mu_{\epsilon0,J}^{(4)}$,
denoted as $\hat{\mu}_{\alpha}^{(3)},\hat{\mu}_{\alpha}^{(4)},\hat{\mu}_{\epsilon,1}^{(3)},...\hat{\mu}_{\epsilon,J}^{(3)},\hat{\mu}_{\epsilon,1}^{(4)},...\hat{\mu}_{\epsilon,J}^{(4)}$.
Now define $\hat{\Gamma}_{N}$ and $\hat{\Upsilon}_{N}$ as
\begin{align}
\hat{\Gamma}_{N} & =\Gamma(\hat{\delta}_{N},\hat{s}_{N}),\label{eq:Gamma_hat}\\
\hat{\Upsilon}_{N} & =\Upsilon(\hat{\delta}_{N},\hat{\mu}_{\alpha}^{(3)},\hat{\mu}_{\alpha}^{(4)},\hat{\mu}_{\epsilon,1}^{(3)},...\hat{\mu}_{\epsilon,J}^{(3)},\hat{\mu}_{\epsilon,1}^{(4)},...\hat{\mu}_{\epsilon,J}^{(4)},\hat{s}_{N}),\label{eq:Upsilon_hat}
\end{align}
then it follows from Slutsky's theorem that $\hat{\Gamma}_{N}$ and
$\hat{\Upsilon}_{N}$ are consistent estimators for $\Gamma_{0}$
and $\Upsilon_{0}$. A consistent estimator for the variance covariance
matrix of the limiting distribution is given by $\hat{\Gamma}_{N}^{-1}\hat{\Upsilon}_{N}\hat{\Gamma}_{N}^{-1}$.
The above discussion assumed the availability of consistent estimators
for the third and fourth moment of the error components. In the following
we now define consistent estimators for $\mu_{\alpha0}^{(3)}$, $\mu_{\alpha0}^{(4)}$
and $\mu_{\epsilon0,j}^{(3)}$, $\mu_{\epsilon0,j}^{(4)}$, $j=1,...,J$.
To motivate the estimators consider the composite error term for individual
$i$ in group $r$, $u_{ir}=\alpha_{r}+\epsilon_{ir}$, and let $\bar{u}_{r}=\frac{1}{m_{r}}\sum_{i=1}^{m_{r}}u_{ir}$,
and $\ddot{u}_{ir}=u_{ir}-\bar{u}_{r}$ . Then $\bar{u}_{r}=\alpha_{r}+\bar{\epsilon}_{r}$
and $\ddot{u}_{ir}=\epsilon_{ir}-\bar{\epsilon}_{r}$, where $\bar{\epsilon}_{r}$
is the group mean of $\epsilon_{ir}$. It is readily verified that
under Assumptions \ref{assume:epsilon} and \ref{assume:alpha}, we
have
\[
E\left[\ddot{u}_{ir}^{3}\right]=(1-\frac{3}{m_{r}}+\frac{2}{m_{r}^{2}})\mu_{\epsilon0,D_{r}}^{(3)},
\]
\[
E\left[\ddot{u}_{ir}^{2}\bar{u}_{r}\right]=\frac{(m_{r}-1)}{m_{r}^{2}}\mu_{\epsilon0,D_{r}}^{(3)},
\]
\[
E\left[\bar{u}_{r}^{3}\right]=\mu_{\alpha0}^{(3)}+\frac{\mu_{\epsilon0,D_{r}}^{(3)}}{m_{r}^{2}},
\]
\[
E\left[\ddot{u}_{ir}^{4}\right]=\frac{m_{r}^{3}-4m_{r}^{2}+6m_{r}-3}{m_{r}^{3}}\mu_{\epsilon0,D_{r}}^{(4)}+\frac{3(m_{r}-1)(2m_{r}-3)}{m_{r}^{3}}\sigma_{\epsilon0,D_{r}}^{4},
\]
\[
E\left[\bar{u}_{r}^{4}\right]=\mu_{\alpha0}^{(4)}+\frac{1}{m_{r}^{3}}\mu_{\epsilon0,D_{r}}^{(4)}+3\frac{m_{r}-1}{m_{r}^{3}}\sigma_{\epsilon0,D_{r}}^{4}+\frac{6}{m_{r}}\sigma_{\alpha0}^{2}\sigma_{\epsilon0,D_{r}}^{2}.
\]
Next define for group $r$ ,
\[
f_{\epsilon,r}^{(3)}=\begin{cases}
\frac{1}{m_{r}}\sum_{i=1}^{m_{r}}\ddot{u}_{ir}^{3}/(1-\frac{3}{m_{r}}+\frac{2}{m_{r}^{2}}) & m_{r}\geqslant3\\
\frac{1}{m_{r}}\sum_{i=1}^{m_{r}}\ddot{u}_{ir}^{2}\bar{u}_{r}/(\frac{1}{m_{r}}-\frac{1}{m_{r}^{2}}) & m_{r}=2
\end{cases},
\]
\[
f_{\alpha,r}^{(3)}=\bar{u}_{r}^{3}-f_{\epsilon,r}^{(3)}/m_{r}^{2}\text{,}
\]
\[
f_{\epsilon,r}^{(4)}=\frac{m_{r}^{3}}{m_{r}^{3}-4m_{r}^{2}+6m_{r}-3}\text{\text{[(\ensuremath{\frac{1}{m_{r}}\sum_{i=1}^{m_{r}}\ddot{u}_{ir}^{4}})-\ensuremath{\frac{3(m_{r}-1)(2m_{r}-3)}{m_{r}^{3}}\sigma_{\epsilon0,D_{r}}^{4}}]}, }
\]
\[
f_{\alpha,r}^{(4)}=\bar{u}_{r}^{4}-f_{\epsilon,r}^{(4)}/m_{r}^{3}-\frac{3(m_{r}-1)}{m_{r}^{3}}\sigma_{\epsilon0,D_{r}}^{4}-\frac{6}{m_{r}}\sigma_{\alpha0}^{2}\sigma_{\epsilon0,D_{r}}^{2}.
\]
Then $E\left[f_{\epsilon,r}^{(3)}\right]=\mu_{\epsilon0,D_{r}}^{(3)}$,
$E\left[f_{\alpha,r}^{(3)}\right]=\mu_{\alpha0}^{(3)}$, $E\left[f_{\epsilon,r}^{(4)}\right]=\mu_{\epsilon0,D_{r}}^{(4)}$,
$E\left[f_{\alpha,r}^{(4)}\right]=\mu_{\alpha0}^{(4)}$. By Lemma
\ref{lem:moment34}(a), $\frac{1}{R_{j}}\sum_{r=1}^{R}1(D_{r}=j)f_{\epsilon,r}^{(l)}\overset{p}{\rightarrow}\mu_{\epsilon0,j}^{(l)}$
and $\frac{1}{R}\sum_{r=1}^{R}f_{\alpha,r}^{(l)}\overset{p}{\rightarrow}\mu_{\alpha0}^{(l)}$
for $l=3,4$ and $j=1,...,J$ as $R$ goes to infinity.
To construct feasible counterparts of these estimates, consider the
estimated disturbances $\hat{u}_{ir}=y_{ir}-\hat{\lambda}\bar{y}_{(-i)r}-z_{ir}\hat{\beta}$,
where $\hat{\lambda}$ and $\hat{\beta}$ denote the QML estimators,
and let $\hat{\bar{u}}_{r}=\frac{1}{m_{r}}\sum_{i=1}^{m_{r}}\hat{u}_{ir}$
and $\hat{\ddot{u}}_{ir}=\hat{u}_{ir}-\hat{\bar{u}}_{ir}$. Feasible
counterparts, say, $\hat{f}_{\epsilon,r}^{(3)}$, $\hat{f}_{\alpha,r}^{(3)}$,
$\hat{f}_{\epsilon,r}^{(4)}$, $\hat{f}_{\alpha,r}^{(4)}$ of $f_{\epsilon,r}^{(3)}$,
$f_{\alpha,r}^{(3)}$, $f_{\epsilon,r}^{(4)}$, $f_{\alpha,r}^{(4)}$
can now be defined by replacing $\bar{u}_{r}$ and $\ddot{u}_{ir}$
with $\hat{\bar{u}}_{r}$ and $\hat{\ddot{u}}_{ir}$, and $\sigma_{\alpha0}^{2}$
and $\sigma_{\epsilon0,j}^{2}$ with their QML estimators. Now consider
the following estimators for the third and fourth moments of the error
components: $\hat{\mu}_{\alpha}^{(3)}=\sum_{r=1}^{R}\hat{f}_{\alpha,r}^{(3)}/R$,
$\hat{\mu}_{\alpha}^{(4)}=\sum_{r=1}^{R}\hat{f}_{\alpha,r}^{(4)}/R$,
$\hat{\mu}_{\epsilon,j}^{(3)}=\sum_{r=1}^{R}1(D_{r}=j)\hat{f}_{\epsilon,r}^{(3)}/R_{j}$,
$\hat{\mu}_{\epsilon,j}^{(4)}=\sum_{r=1}^{R}1(D_{r}=j)\hat{f}_{\epsilon,r}^{(4)}/R_{j}$,
$j=1,...,J$.
The next theorem establishes that valid inference based on standardized
statistics is possible. At the core of this result is the fact that
$\hat{\Gamma}_{N}\overset{p}{\rightarrow}\Gamma_{0}$ and $\hat{\Upsilon}_{N}\overset{p}{\rightarrow}\Upsilon_{0}$
as shown in Appendix \ref{app:Proofthm}.
\begin{thm}
\label{thm:ValidInference}Under Assumptions \ref{assume:epsilon}-\ref{assume:id},
and assuming that $\mu_{\varepsilon0,j}^{(4)}-\sigma_{\varepsilon0,j}^{4}>(\mu_{\epsilon0,j}^{(3)})^{2}/\sigma_{\varepsilon0,j}^{2}$
for $j\in\{1,...,J\}$, and $\hat{\Gamma}_{N}$, $\hat{\Upsilon}_{N}$
defined in (\ref{eq:Gamma_hat}) and (\ref{eq:Upsilon_hat}) we have
$\sqrt{N}\left(\hat{\Gamma}_{N}^{-1}\hat{\Upsilon}_{N}\hat{\Gamma}_{N}^{-1}\right)^{-1/2}(\hat{\delta}_{N}-\delta_{0})\xrightarrow{d}N(0,I)$
as $N\rightarrow\infty$.
\end{thm}
The proof of the theorem is in Appendix \ref{app:Proofthm}.
\section{Monte Carlo Results\label{sec:Monte_Carlo}}
We conduct Monte-Carlo (MC) experiments to assess the finite sample
properties of the quasi-maximum likelihood (QML) estimator $\hat{\delta}_{N}$.
The data generating mechanism is determined by the main model in \eqref{eq:main_orig}.
For simplicity, $x_{1,ir}$, $x_{2,ir}$ and $x_{3,ir}$ each only
includes a scalar variable. We set the true value of the parameters
to $\lambda_{0}=0.5$, $\sigma_{\alpha0}^{2}=0.25$, $\beta_{10}=1$,
$\beta_{20}=1$,$\beta_{30}=1$, and $\beta_{40}=1$, while $\sigma_{\epsilon0}^{2}=1$
in the case of homoscedasticity. The model for the data generating
process (DGP) is thus
\begin{equation}
y_{ir}=0.5\bar{y}_{(-i)r}+1+x_{1,ir}+\bar{x}_{2,(-i)r}+x_{3,r}+\alpha_{r}+\epsilon_{ir}.\label{eq:monte_model}
\end{equation}
The inputs $x_{j,ir}$ $\alpha_{r}$ and $\epsilon_{ir}$ are generated
as follows. In the case when $x_{1}=x_{2}$, $x_{1,ir}=x_{2,ir}\sim\textrm{i.i.d.}\,N(0,1)$.
In the case when $x_{1}\neq x_{2}$, $x_{1,ir}$ and $x_{2,ir}$ are
generated mutually independently, each drawn from an $\textrm{i.i.d.}\,N(0,1)$.
We then calculate the leave-out-mean $\bar{x}_{2,(-i)r}=\frac{1}{m_{r}}\sum_{j\neq i}x_{2j,r}$.
Group characteristics are drawn as $x_{3,r}\sim\textrm{i.i.d.}\,N(0,1)$.
In the case of homoscedastic normal errors in Tables \ref{table:simtab_1_0}
to \ref{table:simtab_3_1}, the idiosyncratic error terms $\epsilon_{ir}$
are i.i.d $N(0,1)$ and group effects $\alpha_{r}$ are i.i.d $N(0,0.25)$.
Both $\epsilon_{ir}$ and $\alpha_{r}$ are drawn independently of
$x_{1,ir}$, $x_{2,ir}$, $x_{3,r}$, and of each other. The dependent
variable $y_{ir}$ is calculated using Equation \eqref{eq:y_sol_2}.
In Table \ref{table:simtab_4_1}, we use homoscedastic but nonnormal
errors. In the case of the Skew normal distribution, we set the location
parameter to 0, scale to 1 and shape to $0.9/\sqrt{1-0.9^{2}}$. Therefore,
Skewness is 0.472 and Kurtosis is 3.321. In the case of the student
distribution, degrees of freedom are set to 6. Therefore, Skewness
is 0 and Kurtosis is 6. In both cases, $\alpha_{r}$ and $\epsilon_{ir}$
are independently drawn from identical distributions and then standardized
to have mean 0 and variance 0.25 and 1 respectively. In Table \ref{table:simtab_5_1},
group effects $\alpha_{r}$ are still i.i.d $N(0,0.25)$, $\epsilon_{ir}$
follow normal distributions but are allowed to be heteroscedastic.
In the first case (Columns 1-2), we randomly select half of the groups
into category 1, with $\epsilon_{ir}$ i.i.d $N(0,0.5)$. The other
half of the groups have $\epsilon_{ir}$ i.i.d $N(0,1.5)$. In the
second case (Columns 3-4), $\epsilon_{ir}$ are i.i.d $N(0,1)$. But
we randomly divide the groups into two categories and allow for heteroscedasticity
of $\epsilon_{ir}$ between categories in estimation. In the third
case (Columns 5-6), groups are randomly divided into two categories,
with $\sigma_{\epsilon r}^{2}\in\{0.5,1.5\}$ and $\epsilon_{ir}$
i.i.d $N(0,\sigma_{\epsilon r}^{2})$. In the fourth case, groups
are randomly divided into four categories with $\sigma_{\epsilon r}^{2}\in\{0.4,0.8,1.2,1.6\}$
and $\epsilon_{ir}$ i.i.d $N(0,\sigma_{\epsilon r}^{2})$.
The number of groups $R$ is selected from the set $\{50,100,200,400,800,1600\}$.
In Tables \ref{table:simtab_1_0}, \ref{table:simtab_1_1} and \ref{table:simtab_4_1},
group size $m_{r}$ is drawn from a discrete uniform distribution
$\mathcal{U}\{2,6\}$ so that the average group size is 4. Small group
sizes are motivated by applications to college room mates, friendship
networks in the Add Health data set or golf tournaments, see \citet{sacerdote_peer_2001},
\citet{goldsmith-pinkham_social_2013} and \citet{guryan_peer_2009}.
In Tables \ref{table:simtab_2_0} and \ref{table:simtab_2_1}, group
size is drawn from $\mathcal{U}\{13,25\}$. The distribution is motivated
by Project STAR where class size ranges from 13 to 25. We also consider
the case when $m_{r}$ is drawn from $\mathcal{U}\{3,5\}$, $\mathcal{U}\{4,8\}$,
$\mathcal{U}\{8,30\}$ and $\mathcal{U}\{10,22\}$ in Table \ref{table:simtab_3_1}
to examine how the distribution of group size affects the performance
of the estimator. Note that $\mathcal{U}\{3,5\}$ has the same mean
as $\mathcal{U}\{2,6\}$ but smaller variance, $\mathcal{U}\{4,8\}$
has the same variance as $\mathcal{U}\{2,6\}$ but larger mean. Meanwhile
$\mathcal{U}\{8,30\}$ has the same mean as $\mathcal{U}\{13,25\}$
but larger variance, $\mathcal{U}\{10,22\}$ has the same variance
as $\mathcal{U}\{13,25\}$ but smaller mean.
In Tables 1-\ref{table:simtab_4_1}, we compare our QML estimator
with the conditional maximum likelihood (CML) estimator of \citet{lee_identification_2007}.
Table \ref{table:simtab_5_1} does not present CMLE estimates as it
does not allow for heteroscedasticity. \citet{lee_identification_2007}
assumes normality of the error terms. Our discussion suggests that
the CMLE is in fact consistent under nonnormal errors, as it can be
viewed as a GMM estimator based on the moment conditions from the
within equation. When group effects are in fact independent of the
observed characteristics, the CML estimator is still consistent but
less efficient than our QML estimator. The comparison thus helps to
evaluate the efficiency gain of our estimator over the CML estimator
in finite samples. The CML estimator is based on the within-group
variation hence $\sigma_{\alpha}^{2}$, $\beta_{1}$ and $\beta_{4}$
are not identified.
We generate 5000 repetitions for each of the experiments. Tables \ref{table:simtab_1_0}-\ref{table:simtab_5_1}
summarize the results of the Monte Carlo (MC) experiments. Each panel
displays the MC median, MC robust standard errors (Rob.Std.Dev), MC
sample standard deviation (Std.Dev.), MC median of the estimated standard
deviation (est.Std.Dev), and the mean rejection rate of the Wald test
with significance level 0.05 of our QMLE and Lee's CMLE across 5000
repetitions. The robust standard errors are defined as IQ/1.35, where
IQ denotes the inter-quantile range, that is $IQ=C_{0.75}-C_{0.25}$
with $C_{0.75}$ and $C_{0.25}$ being the 75th and 25th percentile
respectively. If the distribution of the estimate is normal, IQ/1.35
is (apart from rounding errors) equal to the standard deviation. The
null hypothesis for the Wald test is that the estimate equals its
true value. Critical values for the test are obtained at 5\% significance
level and are based on the asymptotic approximation in Theorem \ref{thm:ValidInference}.
Identification of our models is more challenging, the larger group
sizes are, all else equal. This follows from work of \citet{kelejian_2sls_2002}.
Identification is also more difficult when there is less variation
in group sizes, or less variation in type specific variances or both.
Finally, identification is more difficult in designs where $x_{1,ir}=x_{2,ir}$
because the implied correlation between $x_{1,ir}$ and $\bar{x}_{2,(-i)r}$
reduces the overall variation in the covariates. Standard finite sample
theory for the Gaussian regression model shows that maximum likelihood
estimators for the variance parameters are biased in finite samples.
In fixed effects panel regressions this finite sample bias can lead
to inconsistent estimates of the variance parameter due to incidental
parameter bias, as demonstrated by \citet{neyman_consistent_1948}.
In the current context, we expect the CML estimator to suffer from
such incidental parameter bias because the moment conditions that
identify $\lambda$ depend on the estimated variances. We also expect
Wald type statistics, such as the t-ratio, to perform poorly in designs
where identification is problematic, in line with insights from \citet{dufour_impossibility_1997}.
Tables \ref{table:simtab_1_0} and \ref{table:simtab_1_1} contain
results for small groups and homoscedastic Gaussian errors. In Table
\ref{table:simtab_1_0} where $x_{1}\neq x_{2},$ both the QMLE and
CMLE perform well, with the CMLE being more biased for the parameter
$\lambda$ in sample sizes where $R$ is below 200. The QMLE is generally
less biased and significantly more precise than CMLE, demonstrating
the expected efficiency gains of QMLE. Size is better controlled for
CMLE but the size distortions for the parameters $\lambda$ and $\beta$
do not exceed 7\% in the smallest sample sizes even for the QMLE.
Size distortions for the t-ratios of the two estimated variance parameters
are somewhat larger, reaching 11.6\% for the t-ratio for $\sigma_{\alpha}^{2}$
when $R=50.$ The size distortion seems to be due both to some estimator
bias as well as standard errors that are a bit too small. Size distortions
for all parameters disappear in the larger samples. In Table \ref{table:simtab_1_1}
where $x_{1}=x_{2}$ the CMLE for $\lambda$ is even more biased in
small samples, and considerably more volatile than in the design in
Table \ref{table:simtab_1_0}. The performance of the QMLE is not
very different from the case with $x_{1}\neq x_{2}.$ The standard
deviation measured by IQ/1.35 is somewhat larger than when $x_{1}\neq x_{2}$,
as are size distortions, confirming the intuition that this design
is more difficult to identify.
Tables \ref{table:simtab_2_0} and \ref{table:simtab_2_1} differ
from Tables \ref{table:simtab_1_0} and \ref{table:simtab_1_1} in
that they consider the same designs but with larger group sizes, now
drawn from the uniform distribution on the interval $\left[13,25\right].$
In Table \ref{table:simtab_2_0} we consider the case with $x_{1}\neq x_{2}.$
The QMLE remains roughly unbiased across all sample sizes. The robust
standard deviation roughly doubles relative to the small group size
case and the size properties for t-ratios of the parameters $\lambda$
and $\beta$ deteriorate in samples where $R\leq100$ with size reaching
around 10\% in some cases. Size remains well controlled in larger
samples with $R\geq200$. The size distortions for the variance parameters
are not much affected by the larger class sizes. The CMLE is even
more biased when $R=50$ but less biased for larger sample sizes compared
to Tables \ref{table:simtab_1_0} and \ref{table:simtab_1_1}. This
is consistent with incidental parameter bias which is expected to
decrease with increasing group size. In addition the CMLE now is significantly
less precise. This is in line with results by Lee (2007). Table \ref{table:simtab_2_1}
contains results for the case $x_{1}=x_{2}$ and large group sizes.
The QMLE remains largely unbiased across all sample sizes but there
is notable loss in estimator precision as measured by IQ/1.35, indicating
the more challenging estimation environment. In line with theoretical
predictions, estimator precision increases monotonically with sample
size. Size distortions are now pronounced with empirical size reaching
more than 20\% in the smaller samples. The CMLE controls size well
across all four designs. This comes at the cost of much less precisely
and sometimes more biased estimated parameters.
Table \ref{table:simtab_3_1} explores the effects that variation
in group size has on both estimators. The case with $\mathcal{U}\{3,5\}$
maintains the same mean group size as in Table \ref{table:simtab_1_1}
but reduces the group size variance. We only report results for $\lambda.$
The bias of the QMLE is not affected while the CMLE is somewhat less
biased. The variance of both estimators increases. For the QMLE size
distortions are somewhat larger than in Table \ref{table:simtab_1_1}.
The design with $\mathcal{U}\{4,8\}$ increases the mean while leaving
the variance of class sizes unchanged relative to Table \ref{table:simtab_1_1}.
Overall, the results for this case are quite similar to the scenario
with $\mathcal{U}\{3,5\}$. The designs with $\mathcal{U}\{8,30\}$
and $\mathcal{U}\{10,22\}$ both improve identification relative to
the design in Table \ref{table:simtab_2_1}. For the QMLE this results
in unchanged good bias properties except when $R=50$ where we now
see a small amount of bias, somewhat lower variance and slightly improved
size properties. For the CMLE bias increases while variance somewhat
improves relative to the results in Table \ref{table:simtab_2_1}
and the size properties remain similar.The larger bias for the CMLE
may be related to a larger fraction of smaller classes in both designs.
Smaller group sizes tend to amplify incidental parameter bias.
Table \ref{table:simtab_4_1} explores the effects that non-Gaussian
error distributions have on the estimators. For the Skew Normal distribution
we see little difference to the results in Table \ref{table:simtab_1_1}
both for the QMLE and the CMLE estimator. The QMLE is also robust
to the second design which uses a t-distribution with 6 degrees of
freedom. The CMLE is more sensitive to this fat-tailed distribution.
It is somewhat more biased and has higher variance compared to the
Gaussian case. In addition, we now observe size distortions for the
t-ratio related to the parameter $\lambda.$ These size distortions
don't disappear in larger samples and seem to be due to the fact that
the standard errors show a significant downward bias. This is most
likely due to the fact that Lee (2007) bases standard errors on Gaussian
error distributions. The final set of results we discuss are in Table
\ref{table:simtab_5_1} where we examine the effects of heteroscedasticity
on the QMLE. We do not report results for the CMLE since this estimator
was designed for the homoscedastic case only. The first set of results
are based on a design where class size varies according to a $\mathcal{U}\{2,6\}$
distribution and where we maintain $x_{1}=x_{2}$. Compared to a homoscedastic
design the QMLE is somewhat less variable with no change in bias.
The size properties of the t-ratio are overall comparable between
the two cases, with slightly smaller size distortions in the heteroscedastic
case when $R=50.$ We also consider a scenario where group size is
fixed at $m=4$ while the type specific variances vary. While the
QMLE continues to be nearly unbiased it has a higher variance. The
size properties of t-ratios are slightly worse than in the homoscedastic
case. For larger sample sizes both standard errors and t-ratios are
well behaved.
\section{Conclusion}
In this paper, we show that moment conditions underlying the conditional
variance method of Graham (2008) can be related to and motivated from
a general class of linear peer effects models with random group effects.
When augmented with group specific covariates our specification of
the peer effects model is appropriate for settings where people are
randomly assigned to groups or where group level heterogeneity is
credibly controlled for with observed group level characteristics.
We show that the quasi maximum likelihood estimator (QMLE) related
to a linear Gaussian specification, as well as Graham's estimator
and the fixed effects estimator of Lee (2007) are contained in the
class of GMM estimators we consider. Under Gaussian error assumptions
the QMLE is the most efficient estimator in this class. We study conditions
of identification, extending results in Graham (2008) and Lee (2007)
for a simple model without covariates and a general model with covariates
estimated by QML. We also establish that our QMLE is asymptotically
normal and we construct consistent standard error formulas. Monte
Carlo results show that our QML estimator has good small sample properties.
\pagebreak{}
\bibliographystyle{elsarticle-harv}
\addcontentsline{toc}{section}{\refname}
\begin{thebibliography}{52}
\expandafter\ifx\csname natexlab\endcsname\relax\def\natexlab#1{#1}\fi
\ifx\xfnm\relax \def\xfnm[#1]{\unskip,\space#1}\fi
\bibitem[{Angrist(2014)}]{angrist_perils_2014}
Angrist, J.D., 2014.
\newblock The perils of peer effects.
\newblock Labour Economics 30,
98--108.
\bibitem[{Anselin(1988)}]{anselin_spatial_1988}
Anselin, L., 1988.
\newblock Spatial {{Econometrics}}: {{Methods}} and
{{Models}}. volume~4.
\newblock {Springer Science \& Business Media}.
\bibitem[{Anselin(2010)}]{anselin_thirty_2010}
Anselin, L., 2010.
\newblock Thirty years of spatial econometrics.
\newblock Papers in Regional Science 89,
3--25.
\bibitem[{Booij et~al.(2017)Booij, Leuven and Oosterbeek}]{booij_ability_2017}
Booij, A.S., Leuven, E.,
Oosterbeek, H., 2017.
\newblock Ability {{Peer Effects}} in {{University}}:
{{Evidence}} from a {{Randomized Experiment}}.
\newblock The Review of Economic Studies
84, 547--578.
\bibitem[{Boucher et~al.(2014)Boucher, Bramoull{\'e}, Djebbari and
Fortin}]{boucher_peers_2014}
Boucher, V., Bramoull{\'e}, Y.,
Djebbari, H., Fortin, B.,
2014.
\newblock Do {{Peers Affect Student Achievement}}? {{Evidence}}
from {{Canada Using Group Size Variation}}.
\newblock Journal of Applied Econometrics
29, 91--109.
\bibitem[{Bramoull{\'e} et~al.(2009)Bramoull{\'e}, Djebbari and
Fortin}]{bramoulle_identification_2009}
Bramoull{\'e}, Y., Djebbari, H.,
Fortin, B., 2009.
\newblock Identification of peer effects through social
networks.
\newblock Journal of Econometrics 150,
41--55.
\bibitem[{Cai and Szeidl(2018)}]{cai_interfirm_2018}
Cai, J., Szeidl, A., 2018.
\newblock Interfirm {{Relationships}} and {{Business
Performance}}*.
\newblock The Quarterly Journal of Economics
133, 1229--1282.
\bibitem[{Carrell et~al.(2009)Carrell, Fullerton and West}]{carrell_does_2009}
Carrell, S.E., Fullerton, R.L.,
West, J.E., 2009.
\newblock Does {{Your Cohort Matter}}? {{Measuring Peer
Effects}} in {{College Achievement}}.
\newblock Journal of Labor Economics 27,
439--464.
\bibitem[{Carrell et~al.(2013)Carrell, Sacerdote and
West}]{carrell_natural_2013}
Carrell, S.E., Sacerdote, B.I.,
West, J.E., 2013.
\newblock From {{Natural Variation}} to {{Optimal Policy}}?
{{The Importance}} of {{Endogenous Peer Group Formation}}.
\newblock Econometrica 81,
855--882.
\bibitem[{Chamberlain(1980)}]{chamberlain_analysis_1980}
Chamberlain, G., 1980.
\newblock Analysis of {{Covariance}} with {{Qualitative
Data}}.
\newblock The Review of Economic Studies
47, 225--238.
\bibitem[{Chetty et~al.(2011)Chetty, Friedman, Hilger, Saez, Schanzenbach and
Yagan}]{chetty_how_2011}
Chetty, R., Friedman, J.N.,
Hilger, N., Saez, E.,
Schanzenbach, D.W., Yagan, D.,
2011.
\newblock How {{Does Your Kindergarten Classroom Affect Your
Earnings}}? {{Evidence}} from {{Project Star}}.
\newblock The Quarterly Journal of Economics
126, 1593--1660.
\bibitem[{Chung(2001)}]{chung_course_2001}
Chung, K.L., 2001.
\newblock A Course in Probability Theory.
\newblock {Academic press}.
\bibitem[{Cliff and Ord(1973)}]{cliff_spatial_1973}
Cliff, A.D., Ord, J.K.,
1973.
\newblock Spatial Autocorrelation. volume~5.
\newblock {Pion London}.
\bibitem[{Cliff and Ord(1981)}]{cliff_spatial_1981}
Cliff, A.D., Ord, J.K.,
1981.
\newblock Spatial Processes: Models \& Applications.
volume~44.
\newblock {Pion London}.
\bibitem[{Dhrymes(1978)}]{dhrymes_mathematics_1978}
Dhrymes, P.J., 1978.
\newblock Mathematics for Econometrics.
\newblock Technical Report. {Springer}.
\bibitem[{Duflo et~al.(2011)Duflo, Dupas and Kremer}]{duflo_peer_2011}
Duflo, E., Dupas, P.,
Kremer, M., 2011.
\newblock Peer {{Effects}}, {{Teacher Incentives}}, and the
{{Impact}} of {{Tracking}}: {{Evidence}} from a {{Randomized Evaluation}} in
{{Kenya}}.
\newblock The American Economic Review
101, 1739--1774.
\bibitem[{Duflo and Saez(2003)}]{duflo_role_2003}
Duflo, E., Saez, E., 2003.
\newblock The {{Role}} of {{Information}} and {{Social
Interactions}} in {{Retirement Plan Decisions}}: {{Evidence}} from a
{{Randomized Experiment}}*.
\newblock The Quarterly Journal of Economics
118, 815--842.
\bibitem[{Dufour(1997)}]{dufour_impossibility_1997}
Dufour, J.M., 1997.
\newblock Some {{Impossibility Theorems}} in {{Econometrics
With Applications}} to {{Structural}} and {{Dynamic Models}}.
\newblock Econometrica 65,
1365--1387.
\bibitem[{Fafchamps and Quinn(2018)}]{fafchamps_networks_2018}
Fafchamps, M., Quinn, S.,
2018.
\newblock Networks and {{Manufacturing Firms}} in {{Africa}}:
{{Results}} from a {{Randomized Field Experiment}}.
\newblock The World Bank Economic Review
32, 656--675.
\bibitem[{Frijters et~al.(2019)Frijters, Islam and
Pakrashi}]{frijters_heterogeneity_2019}
Frijters, P., Islam, A.,
Pakrashi, D., 2019.
\newblock Heterogeneity in peer effects in random dormitory
assignment in a developing country.
\newblock Journal of Economic Behavior \& Organization
163, 117--134.
\bibitem[{Garlick(2018)}]{garlick_academic_2018}
Garlick, R., 2018.
\newblock Academic {{Peer Effects}} with {{Different Group
Assignment Policies}}: {{Residential Tracking}} versus {{Random
Assignment}}.
\newblock American Economic Journal: Applied Economics
10, 345--369.
\bibitem[{{Goldsmith-Pinkham} and Imbens(2013)}]{goldsmith-pinkham_social_2013}
{Goldsmith-Pinkham}, P., Imbens, G.W.,
2013.
\newblock Social {{Networks}} and the {{Identification}} of
{{Peer Effects}}.
\newblock Journal of Business \& Economic Statistics
31, 253--264.
\bibitem[{Graham(2008)}]{graham_identifying_2008}
Graham, B.S., 2008.
\newblock Identifying {{Social Interactions Through Conditional
Variance Restrictions}}.
\newblock Econometrica 76,
643--660.
\bibitem[{Guryan et~al.(2009)Guryan, Kroft and Notowidigdo}]{guryan_peer_2009}
Guryan, J., Kroft, K.,
Notowidigdo, M.J., 2009.
\newblock Peer {{Effects}} in the {{Workplace}}: {{Evidence}}
from {{Random Groupings}} in {{Professional Golf Tournaments}}.
\newblock American Economic Journal: Applied Economics
1, 34--68.
\bibitem[{Johnson and Horn(1985)}]{johnson_matrix_1985}
Johnson, C.R., Horn, R.A.,
1985.
\newblock Matrix Analysis.
\newblock {Cambridge university press Cambridge}.
\bibitem[{Kang(2007)}]{kang_classroom_2007}
Kang, C., 2007.
\newblock Classroom peer effects and academic achievement:
{{Quasi-randomization}} evidence from {{South Korea}}.
\newblock Journal of Urban Economics 61,
458--495.
\bibitem[{Kapoor et~al.(2007)Kapoor, Kelejian and Prucha}]{kapoor_panel_2007}
Kapoor, M., Kelejian, H.H.,
Prucha, I.R., 2007.
\newblock Panel data models with spatially correlated error
components.
\newblock Journal of Econometrics 140,
97--130.
\bibitem[{Kelejian and Prucha(1998)}]{kelejian_generalized_1998}
Kelejian, H.H., Prucha, I.R.,
1998.
\newblock A {{Generalized Spatial Two-Stage Least Squares
Procedure}} for {{Estimating}} a {{Spatial Autoregressive Model}} with
{{Autoregressive Disturbances}}.
\newblock The Journal of Real Estate Finance and Economics
17, 99--121.
\bibitem[{Kelejian and Prucha(1999)}]{kelejian_generalized_1999}
Kelejian, H.H., Prucha, I.R.,
1999.
\newblock A {{Generalized Moments Estimator}} for the
{{Autoregressive Parameter}} in a {{Spatial Model}}.
\newblock International Economic Review
40, 509--533.
\bibitem[{Kelejian and Prucha(2001)}]{kelejian_asymptotic_2001}
Kelejian, H.H., Prucha, I.R.,
2001.
\newblock On the asymptotic distribution of the {{Moran I}}
test statistic with applications.
\newblock Journal of Econometrics 104,
219--257.
\bibitem[{Kelejian and Prucha(2002)}]{kelejian_2sls_2002}
Kelejian, H.H., Prucha, I.R.,
2002.
\newblock {{2SLS}} and {{OLS}} in a spatial autoregressive
model with equal spatial weights.
\newblock Regional Science and Urban Economics
32, 691--707.
\bibitem[{Kelejian and Prucha(2010)}]{kelejian_specification_2010}
Kelejian, H.H., Prucha, I.R.,
2010.
\newblock Specification and estimation of spatial
autoregressive models with autoregressive and heteroskedastic disturbances.
\newblock Journal of Econometrics 157,
53--67.
\bibitem[{Kelejian et~al.(2006)Kelejian, Prucha and
Yuzefovich}]{kelejian_estimation_2006}
Kelejian, H.H., Prucha, I.R.,
Yuzefovich, Y., 2006.
\newblock Estimation {{Problems}} in {{Models}} with {{Spatial
Weighting Matrices Which Have Blocks}} of {{Equal Elements}}*.
\newblock Journal of Regional Science 46,
507--515.
\bibitem[{Kuersteiner and Prucha(2020)}]{kuersteiner_dynamic_2020}
Kuersteiner, G.M., Prucha, I.R.,
2020.
\newblock Dynamic {{Spatial Panel Models}}: {{Networks}},
{{Common Shocks}}, and {{Sequential Exogeneity}}.
\newblock Econometrica 88,
2109--2146.
\bibitem[{Lee(2007)}]{lee_identification_2007}
Lee, L.F., 2007.
\newblock Identification and estimation of econometric models
with group interactions, contextual factors and fixed effects.
\newblock Journal of Econometrics 140,
333--374.
\bibitem[{Lee et~al.(2010)Lee, Liu and Lin}]{lee_specification_2010}
Lee, L.F., Liu, X., Lin,
X., 2010.
\newblock Specification and estimation of social interaction
models with network structures.
\newblock The Econometrics Journal 13,
145--176.
\bibitem[{Lin(2010)}]{lin_identifying_2010}
Lin, X., 2010.
\newblock Identifying {{Peer Effects}} in {{Student Academic
Achievement}} by {{Spatial Autoregressive Models}} with {{Group
Unobservables}}.
\newblock Journal of Labor Economics 28,
825--860.
\bibitem[{Liu and Lee(2010)}]{liu_gmm_2010}
Liu, X., Lee, L.f., 2010.
\newblock {{GMM}} estimation of social interaction models with
centrality.
\newblock Journal of Econometrics 159,
99--115.
\bibitem[{Liu et~al.(2014)Liu, Patacchini and Zenou}]{liu_endogenous_2014}
Liu, X., Patacchini, E.,
Zenou, Y., 2014.
\newblock Endogenous peer effects: Local aggregate or local
average?
\newblock Journal of Economic Behavior \& Organization
103, 39--59.
\bibitem[{Manski(1993)}]{manski_identification_1993}
Manski, C.F., 1993.
\newblock Identification of {{Endogenous Social Effects}}:
{{The Reflection Problem}}.
\newblock The Review of Economic Studies
60, 531--542.
\bibitem[{Mundlak(1978)}]{mundlak_pooling_1978}
Mundlak, Y., 1978.
\newblock On the {{Pooling}} of {{Time Series}} and {{Cross
Section Data}}.
\newblock Econometrica 46,
69--85.
\bibitem[{Newey(1991)}]{newey_uniform_1991}
Newey, W.K., 1991.
\newblock Uniform {{Convergence}} in {{Probability}} and
{{Stochastic Equicontinuity}}.
\newblock Econometrica 59,
1161--1167.
\bibitem[{Neyman and Scott(1948)}]{neyman_consistent_1948}
Neyman, J., Scott, E.L.,
1948.
\newblock Consistent {{Estimates Based}} on {{Partially
Consistent Observations}}.
\newblock Econometrica 16,
1--32.
\bibitem[{Nye et~al.(2004)Nye, Konstantopoulos and Hedges}]{nye_how_2004}
Nye, B., Konstantopoulos, S.,
Hedges, L.V., 2004.
\newblock How {{Large Are Teacher Effects}}?
\newblock Educational Evaluation and Policy Analysis
26, 237--257.
\bibitem[{Ord(1975)}]{ord_estimation_1975}
Ord, K., 1975.
\newblock Estimation {{Methods}} for {{Models}} of {{Spatial
Interaction}}.
\newblock Journal of the American Statistical Association
70, 120--126.
\bibitem[{P{\"o}tscher and Prucha(1991)}]{potscher_basic_1991}
P{\"o}tscher, B.M., Prucha, I.R.,
1991.
\newblock Basic structure of the asymptotic theory in dynamic
nonlineaerco nometric models, part i: Consistency and approximation
concepts.
\newblock Econometric Reviews 10,
125--216.
\bibitem[{P{\"o}tscher and Prucha(1994)}]{potscher_generic_1994}
P{\"o}tscher, B.M., Prucha, I.R.,
1994.
\newblock Generic uniform convergence and equicontinuity
concepts for random functions.
\newblock Journal of Econometrics 60,
23--63.
\bibitem[{Rivkin et~al.(2005)Rivkin, Hanushek and Kain}]{rivkin_teachers_2005}
Rivkin, S.G., Hanushek, E.A.,
Kain, J.F., 2005.
\newblock Teachers, {{Schools}}, and {{Academic Achievement}}.
\newblock Econometrica 73,
417--458.
\bibitem[{Sacerdote(2001)}]{sacerdote_peer_2001}
Sacerdote, B., 2001.
\newblock Peer {{Effects}} with {{Random Assignment}}:
{{Results}} for {{Dartmouth Roommates}}.
\newblock The Quarterly Journal of Economics
116, 681--704.
\bibitem[{Sojourner(2013)}]{sojourner_identification_2013}
Sojourner, A., 2013.
\newblock Identification of {{Peer Effects}} with {{Missing
Peer Data}}: {{Evidence}} from {{Project STAR}}.
\newblock The Economic Journal 123,
574--605.
\bibitem[{Stinebrickner and Stinebrickner(2006)}]{stinebrickner_what_2006}
Stinebrickner, R., Stinebrickner, T.R.,
2006.
\newblock What can be learned about peer effects using college
roommates? {{Evidence}} from new survey data and students from disadvantaged
backgrounds.
\newblock Journal of Public Economics 90,
1435--1454.
\bibitem[{Zimmerman(2003)}]{zimmerman_peer_2003}
Zimmerman, D.J., 2003.
\newblock Peer {{Effects}} in {{Academic Outcomes}}:
{{Evidence}} from a {{Natural Experiment}}.
\newblock Review of Economics and Statistics
85, 9--23.
\end{thebibliography}
\pagebreak{}