EconBase
← Back to paper

A Nonparametric Test for Cross-Unit Spillovers

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

77,960 characters

A Nonparametric Test for Cross-Unit Spillovers





\begin{titlepage}

\title{A Nonparametric Test for Cross-Unit Spillovers\thanks{\textit{Acknowledgements}:  We are grateful to Marco Alfano, Aislinn Bohren, Arun Chandrasekhar, Ben Deaner, Guido Imbens, \'{A}ureo de Paula and Michael Vlassopoulos for helpful comments. We also thank audiences at UCL, Essex, Southampton, CFE London 2025, ISI Delhi ACEGD 2025, SEW/SEA Paris 2026 and the Celebrating James MacKinnon Conference Aarhus 2026. Comola acknowledges the financial support of the French National Research Agency (ANR-21-CE26-0002-01 and ANR-22-CE26-0001). Comunello and Gupta acknowledge the financial support of the Leverhulme Trust via grant RPG-2024-038. Gupta is supported in part by funding from the Social Sciences and Humanities Research Council of Canada.}}
\author{Margherita Comola\thanks{University Paris-Saclay (RITM) and Paris School of Economics. Email: [email removed]} \and Camila Comunello\thanks{Department of Economics, Universidad Torcuato Di Tella. Email:  [email removed]} \and Abhimanyu Gupta\thanks{Department of Economics, Queen's University, Dunning Hall, 94 University Avenue, Kingston K7L 3N6, Canada. Email:  [email removed]}}
\date{\today
}

\maketitle
\begin{abstract}
\noindent
Cross-unit dependence is pervasive in empirical applications and complicates econometric inference, especially when spillovers operate in nonlinear ways.
We propose a novel nonparametric test for cross-unit spillovers that may operate through peers’ attributes, peers’ outcomes, or both. The test is straightforward to implement, as it requires only estimation under the null hypothesis of no cross-unit spillovers, and is shown to have a convenient asymptotic standard normal distribution. It is also versatile, accommodating data generated by a wide range of interaction structures. We present four empirical illustrations showing that the proposed test can yield substantively different conclusions about the presence of cross-unit spillovers than existing approaches.
\\
\vspace{0in}\\
\noindent\textbf{Keywords:}  Nonparametric test; Cross-unit spillovers; Social interactions; Interference\\
\vspace{0in}\\
\noindent\textbf{JEL Codes:} C21, C14 \\
\bigskip
\end{abstract}
\setcounter{page}{0}
\thispagestyle{empty}
\end{titlepage}
\pagebreak \newpage


\section{Introduction}

Cross-unit dependence is pervasive in economic and social interactions and poses a fundamental challenge for econometric inference.\footnote{\cite{VivianoRudder2024} report that approximately 40\% of experimental papers published in top-five economics journals in 2020 discuss spillovers as a potential threat to identification.}
This
 paper deals with the two most common sources of cross-unit spillovers, which arise through the attributes channel and the outcome channel. To illustrate, consider the canonical  cross-sectional linear model, where the outcome of the unit (e.g. individual) is regressed on own attributes.
 First, the attributes of other units (`peers') may affect $i$'s outcome. We term this as \textbf{`covariate ($\boldsymbol{c}$) spillovers'}.  The second source of cross-unit dependence refers to the case where the outcomes of other units affect $i$'s outcome. We call this \textbf{`outcome ($\boldsymbol{y}$) spillovers'}. In this paper, we propose a test for unknown nonparametric spillovers operating through one or both channels, establish its asymptotic properties, and illustrate its applicability in four diverse settings. We term this the $\boldsymbol{s}$ \textbf{test}.

Cross-unit spillovers have received considerable attention from applied economists in a broad range of  contexts. These include, \emph{inter alia}, disease transmission \citep{MiguelKremer2004, Ozier2018}, educational outcomes \citep{ sacerdote2001peer, LaliveCattaneo2009, BoboniFinan2009,AvvisatiEtAl2013}, employment decisions \citep{DufloSaez2003, BrownLaschever2012} and technology adoption \citep{OsterThornton2012,BanerjeeEA2013,CaiEA2015}.

A strand of the literature on (broadly defined) `peer effects'  has explicitly modeled cross-unit dependence in observational data via the attribute and/or the outcome channel, depending on the setting. Oftentimes, cross-unit dependence is modeled solely through the attribute channel, even though outcome spillovers could also be incorporated due to economic considerations.\footnote{The exclusion restriction that peers’ attributes serve as a reduced-form sufficient statistic for their outcomes is frequently imposed. However, when the research design permits, spillovers operating through both peers’ covariates and peers’ realized outcomes can be jointly identified, offering a sharper understanding of the underlying economic mechanisms \citep{BDF2009, BursztynFiorin2017}.}
  A related line of work focuses on treatment-mediated spillovers, which are a first-order concern in the context of impact evaluation as they violate the Stable Unit Treatment Value Assumption (SUTVA), which asserts that an individual’s potential outcomes should be independent of peers’ treatment assignments.
  In response to this concern, it has become increasingly common to design cluster-randomized experiments generating exogenous variation in peers’ treatment status.\footnote{Cluster-randomized trials (also known as `two-stage randomization experiments', `randomized-saturation experiments', `partial-population experiments') randomly assign different treatment rates across different clusters.}


A practical challenge is that spillovers are typically modelled through linear functions of peers' attributes and outcomes. While convenient, such specifications may fail to detect economically meaningful forms of cross-unit dependence when agents respond nonlinearly to their social or economic environment. In many settings, behavior depends on the distribution of peers' characteristics rather than on simple averages. For example, a small fraction of adopters may induce complementary behavior, whereas widespread adoption may generate substitutability.\footnote{Nonlinear adoption dynamics have been documented both theoretically and empirically \citep{bandiera2006social, young2009innovation, acemoglu2016networks}.} As a result, spillover effects may vary across the distribution and partially offset one another, causing linear specifications to understate or fail to detect cross-unit dependence.



We develop a nonparametric test for cross-unit dependence arising through the attribute channel, the outcome channel, or both. The test is inspired by classical Lagrange Multiplier (LM) diagnostics, such as the RESET test. We construct a test statistic for the null of no spillovers against flexible nonparametric alternatives, which are approximated using a series expansion. Because the test follows an LM approach, estimation is required only under the null. As a result, the procedure accommodates rich forms of cross-unit dependence while requiring only the estimation of a familiar multiple linear regression model.
We establish that the test is asymptotically standard normal under the null and is consistent.
These results are derived under a cluster-robust framework, allowing for forms of within-cluster dependence commonly encountered in applied work.
Extensions to alternative error dependence structures—such as serial correlation or more general forms of spatial dependence—are conceptually straightforward.





Our test is far-reaching in that it is versatile in its data requirements, which is an important advantage. First, it accommodates data defined through a variety of interaction structures, including those based on blocks or links. Block-type data are partitioned into separate self-exclusive groups within which all units are assumed to interact.
Blocks may represent villages or schools or nuclear households for individuals,  geographical area and/or productive sectors for firms.
Alternatively, network-type data contain detailed links between units, which may or may not overlap  (e.g. $i$ is linked to $j$ and $j$ is linked to $k$, but $i$ is not linked to $k$). This is the case for self-declared link data in household surveys, or trade data among firms from administrative records.
 Our test accommodates both data structures. Second, it allows for heterogeneous spillovers via multiple interaction matrices, as we justify in Appendix \ref{sec:app_ext_mult}. Third, it allows the interaction structure to be incomplete or measured noisily. In applied work, interaction data are often measured poorly yet still convey useful information. By embedding our test within a latent-space framework, Appendix \ref{sec:app_ext_netform} states high level conditions under which the perturbations induced by measurement error are asymptotically negligible, ensuring that inference remains valid.




Most of the methodological literature on cross-unit dependence focuses on the estimation of causal parameters under interference. By and large, existing approaches rely on parametric assumptions to achieve this aim \citep{Rosenbaum2007,HudgensHalloran2008, Hirano2010, LiuHudgens2014, BairdEtAl2018, Arduini2020, McNealisMoodieDean2024_JRSSC, viviano_etal2024}. However, more recently nonparametric and semiparametric methods are being increasingly adopted.  \cite{VazquezBare2023} works in a nonparametric identification framework to propose estimators for spillover effects in experiments, while \cite{viviano2025} studies policy targeting with interference in a semiparametric framework. A semiparametric approach is also used in \cite{Hoshin2023}, which develops treatment effects models with strategic interaction in treatment decisions, while \cite{Hoshino2024} use nonparametric methods to conduct causal inference under unknown interference structure and in more recent work provide a specification test for the latter \citep{Hoshino2026}.

By proposing a nonparametric test for cross-unit spillovers, this paper complements the existing literature by providing practitioners with an easy-to-implement diagnostic tool to guide the design, validation, and refinement of estimation strategies. If no spillovers are detected, the test provides empirical support for ruling out interference concerns and proceeding with standard estimation strategies—such as linear intent-to-treat regressions. If instead the test detects nonlinear cross-unit dependence, the estimation strategy should be adapted to account for its potentially substantial effects.\footnote{While our test does not explicitly suggest the nonparametric functional form which best describes the data at hand, a variety of suitable econometric methods are available, see e.g. \cite{Jenish2012a, Jenish2016, Xu2015, Xu2018}.} Our test can also be useful prior to rolling out a large-scale survey, when researchers, based on the pilot, must decide whether and how to adjust the design to account for cross-unit dependence.\footnote{For instance, a researcher may apply our test to interaction data collected in a pilot study to assess whether detailed link information is required, or whether the treatment intensity should be exogenously varied across clusters in a subsequent scale-up. Because both design choices can entail substantial additional survey costs, they are warranted primarily when spillovers are expected to play an important role.}



We illustrate the implementation and economic relevance of our test through four empirical applications that revisit spillovers in distinct settings. In the first illustration, we examine peer effects in a high-skill environment, namely professional golf tournaments \citep{GuryanKroftNotowidigdo}.
The second studies the impact of interracial roommate assignment on stereotype attitudes and academic performance at the University of Cape Town \citep{CornoLaFerraraBurns}. The third examines a mixed-ability randomized deskmate intervention on student performance in Chinese elementary schools \citep{wuzhangwang2023}. The fourth revisits peer effects in academic research among French economists \citep{bosquetcombesetal2022}.
Across all four settings, we show that our test is able to detect cross-unit dependence through peers’ attributes and/or peers’ outcomes in many cases where linear functional forms fail to do so. To paraphrase a longstanding criticism of nonparametric specification tests, they can be akin to using a nuclear weapon against an ant, as few parametric specifications survive their scrutiny. Our applications, however, show that our test is sufficiently discerning to not reject the null when there is little evidence of spillovers.

Section \ref{sec:method} introduces the test statistic, while Section \ref{sec:asymptotics} characterizes its asymptotic behavior. In Section \ref{sec:implementation} we provide an implementation guide that includes advice on choosing the tuning parameters, and in Section \ref{illustrations} we illustrate the test by revisiting four existing studies. Section \ref{Conclusions} concludes. We show that our test controls size well in a Monte Carlo study of finite performance in Appendix \ref{app:sims}. In Appendix \ref{sec:app_ext} we extend our method to heterogeneous cross-unit dependence, and to embedded graphs with noisy measurement and parametric modeling of the underlying link formation structure.
Appendices \ref{appendix:proofs} and \ref{appendix:lemmas} contain theorem proofs and auxiliary lemmas respectively, while Appendix \ref{app_C} reproduces the original results for our empirical demonstrations.









\section{Method and test statistic}\label{sec:method}
We write the general model with both $ \boldsymbol {c}$ and $\boldsymbol {y}$ spillovers in scalar notation as
\begin{equation}\label{model_1}
	y_i=  f(w_i' y) + x_i^\prime \beta +\sum_{j=1}^l g_j(w_i'c_j)+ \epsilon_i, \ \ \ i=1,....,n,
\end{equation}
where $f(\cdot):\mathbb{R}\rightarrow\mathbb{R}$ is an unknown function that captures `outcome ($\boldsymbol{y}$) spillovers', and $g_j(\cdot):\mathbb{R}\rightarrow\mathbb{R}$ are also unknown functions that capture `attribute ($\boldsymbol{c}$) spillovers'. $y=\left(y_1,\ldots,y_n\right)'$ is an observed outcome vector and  $W=\left(w_1,\ldots,w_n\right)'$ is a social weight matrix representing cross-unit interactions that is either fixed or exogenous conditional on observables, and has zero diagonal.\footnote{Since our approach is cast within an instrumental-variables framework, it naturally accommodates endogeneity of the term $w_i' y$ whenever valid instruments are available. In the context of network-type data, it is compatible with instrumental-variable strategies designed to address outcome spillovers that are endogenous due to simultaneity \citep{BDF2009} or due to assortative link formation on unobservables \citep{jochmans2023}.} $x_i'$ is the $i$-th row of an $n\times k$ regressor matrix $X$ that can have endogenous elements as long as instruments are available, $c_j$, $j=1,\ldots,l$, are some $n\times 1$ vectors of exogenous regressors and $\epsilon_i$ is an unobserved disturbance.\footnote{In some applications, it is useful to write the model as \[
	y_i=  f\left(w_i' y^*\right) + x_i^\prime \beta +\sum_{j=1}^l g_j(w_i'c_j)+ \epsilon_i, \ \ \ i=1,....,n,
\] where $y^*$ measures the same outcome as $y$ but need not have the same entries. For instance, in a peer effects study, $y$ could be an academic outcome for a given subsample but $y^*$ could be the outcome for a different subsample.} The data are observed as $n_g$ observations in each of $g=1,\ldots,G$ clusters, so that $n=\sum_{g=1}^G n_G$. We take $n_g$ as fixed with $\sup_{g=1,\ldots,G}n_g<\infty$ and therefore $n\sim G$, i.e. our sample size grows like the number of clusters.  Our approach gives rise to three different testing options:
\begin{enumerate}
\item Jointly test for both $\boldsymbol {c}$ and $\boldsymbol {y}$ spillovers (the `$\boldsymbol{cy}$ test').

\item Omit $f(w_i'y)$ from the general model (\ref{model_1}) and test for only $\boldsymbol {c}$ spillovers (the `$\boldsymbol{c}$ test').

\item Omit $g_j(s)=0,\;j=1,\ldots,l,$ from the general model (\ref{model_1}) and test for only $\boldsymbol {y}$ spillovers (the `$\boldsymbol{y}$ test').
\end{enumerate}
\noindent The corresponding null hypotheses are:
\begin{eqnarray}
&&\boldsymbol{cy}\text{ test, }\mathcal{H}_0: f(s)=0 \text{ and } g_j(s)=0,\;j=1,\ldots,l,\label{nullxy}\\
&& \boldsymbol{c}\text{ test, }	\mathcal{H}_0: g_j(s)=0,\;j=1,\ldots,l, \label{nullx}\\
&& \boldsymbol{y}\text{ test, }		\mathcal{H}_0: f(s)=0,\label{nully}
\end{eqnarray}
for all $s\in support(s)$, which implies in all cases that the null model is $y_i=x_i'\beta+\epsilon_i$, i.e. a standard linear regression. Given these three possible testing choices, we combine them to define our recommended $\boldsymbol{s}$ test, which has the following rule-of-thumb decision rule (on which we elaborate in Section \ref{sec:implementation}):

\begin{definition}\label{def:s-test}(Rejection rule of the $\boldsymbol{s}$ test) Reject the null hypothesis of no spillovers if at least one of the $\boldsymbol{c}$, $\boldsymbol{y}$ or $\boldsymbol{cy}$ tests rejects the null hypothesis.
\end{definition}

The $\boldsymbol{cy}$ test jointly includes nonlinear spillovers in both the covariate/attribute and outcome channels, while the $\boldsymbol{c}$ test and $\boldsymbol{y}$ test examine the covariate and outcome channels individually, respectively.
Most applied work focuses on linear spillovers in covariates and extending that to nonlinearity via the $\boldsymbol{c}$ test seems natural, while augmenting for outcome spillovers through the $\boldsymbol{cy}$ test demonstrates the full power of our approach.
For ease of exposition we present asymptotic results and notation for the most general case covered by the $\boldsymbol{cy}$ test in (\ref{nullxy}).

Let $\psi_i(s)$, for $i=1,\ldots,p$, be a user-chosen set of basis functions (our applications use Hermite polynomials) such that
\begin{equation}\label{lambda_prime}
	f(s)=\sum_{i=1}^p \mu_{f,i} \psi_i(s)+r_f(s),\;\; g_j(s)=\sum_{i=1}^p \mu_{g_j,i} \psi_i(s)+r_{g_j}(s),\;j=1,\ldots,l,
\end{equation}
with $p=p_n$ a divergent deterministic sequence, i.e. $p\rightarrow\infty$ as $n\rightarrow\infty$, $\mu_f= (\mu_{f,1},\ldots, \mu_{f,p})^\prime$, $\mu_{g_j}= (\mu_{g_j,1},\ldots, \mu_{g_j,p})^\prime$  vectors of unknown series coefficients and $r_f(s)$, $r_{g_j}(s)$ approximation errors. We define our approximate null hypothesis as
\begin{equation}\label{approximate_null}
	\mathcal{H}_{0A}: \mu_f=0 \text{ and }\mu_{g_j}=0,\; j=1,\ldots,l, \text{ for some } \beta,
\end{equation}
which is a set of $q=p(l+1)$ restrictions, so that $q\rightarrow\infty$ as $n\rightarrow\infty$. This indicates that standard fixed-dimension asymptotics will not work for our test. Our test statistic is based on determining if the moment conditions for the instrumental variables (IV) estimate of $\beta$ under the null hypothesis are close enough to zero. OLS is of course a special case of this.

Now, for each $i=1,\ldots,p$, define the $n\times 1$ vector $\Upsilon_{f,i}(y)= \left(\psi_i(w_1'y),\ldots,\psi_i(w_n'y)\right)'$
and the $n\times 1$ vectors $\Upsilon_{g_j,i}(c_j)= \left(\psi_i(w_1'c_j),\ldots,\psi_i(w_n'c_j)\right)'$, and write
\[
\Upsilon_{f,g}=\begin{pmatrix} \Upsilon_{f,1}(y)& \ldots& \Upsilon_{f,p}(y)&\Upsilon_{g_1,1}(c_1)& \ldots& \Upsilon_{g_1,p}(c_1)&\ldots& \Upsilon_{g_l,1}(c_l)& \ldots& \Upsilon_{g_l,p}(c_l) \end{pmatrix},
\]
which is an $n\times q$ matrix. Denote $\mu=(\mu_f',\mu_{g_1}',\ldots,\mu_{g_l}')'$. Given  the expansion in (\ref{lambda_prime}), the series approximated IV objective function is
\begin{equation}\label{2slsobj}
	\mathcal{F}_p( \beta, \mu, y)= \frac{1}{n}\left( y-\Upsilon_{f,g}\mu-X\beta\right)^\prime \mathcal{P}_Z\left( y-\Upsilon_{f,g}\mu-X\beta\right),
\end{equation}
where $Z$ is an $n\times m$ matrix of valid instruments, with $m\geq q+k$, and $\mathcal{P}_Z= Z(Z^\prime Z)^{-1} Z^\prime$. Next, define the $n\times (q+k) $ matrix  $U= \begin{pmatrix} \Upsilon_{f,g}& X \end{pmatrix}$. OLS is a special case with $Z=U$.

Define the $(q+k)\times 1$ gradient vector $\tilde{d}( \beta, y)$ of (\ref{2slsobj}) under $\mathcal{H}_{0A}$ as
\begin{equation}\label{d}
	\tilde{d}(\beta, y)= \left.\frac{\partial \mathcal{F}( \mu, \beta, y)}{\partial (\mu, \beta)^\prime } \right\vert_{\mu=0}= -\frac{2}{n}U^\prime \mathcal{P}_Z  (y -  X\beta).
\end{equation}
Denoting by   $\hat{\beta}$ some consistent estimate of $\beta$, e.g. IV or OLS, under $\mathcal{H}_{0A}$,  the gradient evaluated at the corresponding residuals is
\begin{equation}\label{dhat}
	\hat{d}= \tilde{d}\left( \hat{\beta}, y\right)=  -\frac{2}{n}U^\prime \mathcal{P}_Z  (y - X\hat{\beta}).
\end{equation}
Let
$\hat{J}=n^{-1}Z^\prime U$, where $\hat{J}$ is $m \times (q+k) $. Next, define the $m\times m$ matrices $\hat{M}= n^{-1}{Z^\prime Z}$ and $\hat{\Phi}=n^{-1}{Z^\prime \hat{\Sigma} Z}$, with $\hat{\Sigma}=diag\left(\hat{\Sigma}_1,\ldots,\hat{\Sigma}_G\right)$ where $\hat{\Sigma}_g$ has typical $(i,j)$-th element $\hat{\epsilon}_i\hat{\epsilon}_j$ and $\hat{\epsilon}_i= y_i- x_i^\prime\hat{\beta}$, for $i,j=1,\ldots,n_g$ and $g=1,\ldots,G$. We write
\begin{equation}\label{Hhat}
	\hat{H}=4 \hat{J}^\prime \hat{M}^{-1}\hat{\Phi}\hat{M}^{-1} \hat{J},
\end{equation}
and define
our cluster robust test statistic as
\begin{equation}\label{statistic}
	\mathcal{S}= \frac{n\hat{d}^\prime \hat{H}^{-1} \hat{d} - q}{\sqrt{2q}}.
\end{equation}
This is a weighted measure of the distance of the gradient from zero, centred and rescaled to account for $q\rightarrow\infty$.  In practice, especially for small $q$, one can use $n\hat{d}^\prime \hat{H}^{-1} \hat{d}$ as the test statistic with $\chi^2_q$ critical values instead of standard normal ones. We study this in our simulations in Appendix \ref{app:sims}.


	\section{Asymptotic theory}\label{sec:asymptotics}
We commence this section by introducing some technical assumptions to establish the limiting behaviour of (\ref{statistic}) under $\mathcal{H}_{0A}$. Throughout we denote by $K$ a generic positive constant, arbitrarily large but independent of $p$ and $n$.
\begin{assumption}\label{ass:errors} $\epsilon _{i}$
	are random
	variables with zero mean and unknown variance $\sigma_{i} ^{2}\in[c,K]$, $c>0$, and,
	for some $\tau >0,$ $ \mathbb{E} \left\vert\epsilon _{i}\right\vert^{8+\tau }\leq K$ for $i=1,\ldots,n$. Furthermore $\epsilon_{ig}$ and $\epsilon_{jg'}$ are independent for $g\neq g'$, $g,g'=1,\ldots,G$, while $\mathbb{E}\left(\epsilon_{ig}\epsilon_{jg}\right)=\sigma_{ijg}<\infty$, $i\neq j$, and we accordingly write ${\Sigma}=diag\left({\Sigma}_1,\ldots,{\Sigma}_G\right)$.
\end{assumption}

\begin{assumption}\label{ass:regressors}  $ \mathbb{E}(x_{ir}^4 )\leq K$ and $\mathbb{E}(z_{is}^4 )\leq K$, for $i=1,\ldots,n$ and $r=1,\ldots,k$ and $s=1,\ldots,l$.
\end{assumption}
\noindent We also allow $cov(\epsilon_i, x_{ij}) \neq 0$, for some $j=1,\ldots, k$, i.e. $X$ might contain some endogenous columns. Let $X_1$ be the $n \times k_1$ matrix containing the subset of exogenous columns of $X$, while $X_2$ ($n\times k_2$, with $k_2=k-k_1$) contains the endogenous ones. Now, for a generic symmetric positive-definite matrix $A$, let $\overline{\textit{eig}}(A)$ and $\underline{\textit{eig}}(A)$ denote its largest and smallest eigenvalues, respectively. For a generic matrix $B$, denote by $\left\Vert B\right\Vert=\sqrt{\overline{\textit{eig}}(B'B)}$, i.e. the spectral norm of $B$, and by  $\left\Vert B\right\Vert_\infty$ its largest absolute row sum.



\begin{assumption}\label{ass:eigsandinsts}
The $n\times n$ matrix $\Sigma$ satisfies
 \begin{equation}\label{ass_6_Sigma}
		\limsup_{n\rightarrow\infty}\sup_{g=1,\ldots,G}\overline{\textit{eig}}(\Sigma_g) <\infty ,\;\;\; \liminf_{n\rightarrow\infty}\inf_{g=1,\ldots,G}\underline{\textit{eig}}\left(\Sigma_g \right) >0,
	\end{equation}
the $m\times m$ matrix $M=\mathbb{E}(\hat{M})$, with $m\geq q+k$ , satisfies
	\begin{equation}\label{ass_6_a}
		\limsup_{n\rightarrow\infty}\overline{\textit{eig}}(M) <\infty ,\;\;\; \liminf_{n\rightarrow\infty}\underline{\textit{eig}}\left(M \right) >0,
	\end{equation}
	and the $(q+k) \times (q+k)$ matrix
	$L=n^{-1}\mathbb{E}(U'U)$ satisfies
	\begin{equation}\label{ass_6_b}
		\limsup_{n\rightarrow\infty}\overline{\textit{eig}}(L) <\infty ,\;\;\; \liminf_{n\rightarrow\infty}\underline{\textit{eig}}\left(L \right) >0,
	\end{equation}
	for $n$ large enough.  For some $\nu>0$ satisfying $n/p^{(\nu + 1/2)}=o(1)$,
	\[
	\sup_z r_f(z)+\sup_{j=1,\ldots,l}\sup_z r_{g_j}(z)=O_p\left(p^{-\nu} \right),
	\] as $p\rightarrow\infty$.  $\mathbb{E}\left(u_{il_2}^4\right)\leq K$ for $i=1,\ldots,n, l_1=1,\ldots,m$ and $l_2=1,\ldots,q+k$, and $\epsilon_i$ and $z_j$ are uncorrelated for each $i,j=1,\ldots,n$.
\end{assumption}
\noindent Assumption \ref{ass:eigsandinsts} imposes regularity conditions and controls the approximation errors. Specifically, (\ref{ass_6_a})-(\ref{ass_6_b}) are asymptotic boundedness and no multicollinearity conditions for matrices of increasing dimension, while under Assumption \ref{ass:errors}, (\ref{ass_6_Sigma}) also ensures that $0<\sup_{i=1,\ldots,n}\Sigma_i< \infty$ in the special case of purely heteroskedasticity robust testing i.e. when $G=n$ and $\Sigma_i$ are scalars. For the instruments $Z$ we use at least $k_2$ columns of instruments for the endogenous covariates $X_2$, and also the columns of $X_1$, $WX_1$. We also use a set of instruments of the form $\psi_r\left(\sum_j w_{ij}x_{1, jl} \right)$, where $r=1,\ldots,p$, and $x_{1, jl}$ denotes the $(j,l)$th element of $X_1$, with $l=1,\ldots,k_1$.  For more discussion on approximation error decay rates see e.g. \cite{Chen2007}. Our next assumption sets a suitable bound on cross-sectional dependence, analogous to that in \cite{Lee2016}. Conditions such as linear process representations for the underlying random variables or the near-epoch dependence conditions of \cite{Jenish2012} imply that this assumption holds.
 \allowdisplaybreaks
\begin{assumption}\label{ass:Mhat}
	Let
	\[
		\xi = \underset{0\leq l,k \leq m} {\sup} \ \underset{j\neq i}{\underset{i=1}{\overset{n}\sum} {{\underset{j=1}{\overset{n}\sum}}}}\left\vert cov (z_{il}z_{ik},z_{jl}z_{jk}) \right\vert,
		\varkappa =\underset{0\leq l \leq m,\;\; 0\leq r \leq q+k} {\sup} \ \underset{j\neq i}{\underset{i=1}{\overset{n}\sum} {{\underset{j=1}{\overset{n}\sum}}}}\left\vert cov(z_{il}u_{ir},z_{jl}u_{jr})\right\vert,
	\]
	and assume
	\begin{equation}\label{delta_cond}
		\xi+\varkappa =O(n), \ \ \ \text{as} \ \ \ n\rightarrow \infty.
	\end{equation}
\end{assumption}
\noindent Our null asymptotic theory first approximates the test statistic $\mathcal{S}$ with a quadratic form in $\epsilon$, and then shows that this approximation is asymptotically standard normal. Write $J=E(\hat J)$ and define
\begin{align}
	d=d( \beta_0, y)=& -\frac{2}{n} J^\prime M^{-1/2}\left(I - M^{-1/2} N\left(N^\prime M^{-1} N \right)^{-1} N^\prime M^{-1/2} \right) M^{-1/2} Z^\prime \epsilon \notag \\=& -\frac{2}{n} J^\prime M^{-1/2}\mathcal{K}_{NM} M^{-1/2} Z^\prime \epsilon,
\end{align}
where $\mathcal{K}_{NM}= \left(I - M^{-1/2} N\left(N^\prime M^{-1} N \right)^{-1} N^\prime M^{-1/2} \right)$ is $m\times m$ and $N= \mathbb{E}(\hat{N})$, with $\hat{N}= n^{-1} Z' X$, the last being  an $m \times k$ matrix with full rank under (\ref{ass_6_b}) in Assumption \ref{ass:eigsandinsts}.
Set
\begin{align}\label{H}
	H= n\mathbb{E}(d d^\prime)  = 4 J^\prime M^{-1/2} \mathcal{K}_{NM} M^{-1/2}\Phi M^{-1/2} \mathcal{K}_{NM} M^{-1/2} J,
\end{align}
with $\Phi =n^{-1}\mathbb{E}(Z^\prime \Sigma Z)$. Under Assumptions \ref{ass:errors} and \ref{ass:eigsandinsts}, $H^{-1}$ exists and is non-singular for $n$ large enough, using Lemma \ref{lemma:Omegaeigs}. We now state the main result of this section.
\begin{theorem}\label{theorem:nulldist}
	Let Assumption \ref{ass:errors} hold with the representation $\epsilon_{i}=\sum_{r=1}^n b_{ir}\eta_r$, where $\eta_r$ are i.i.d. mean zero and unit variance random variables and the $b_{ir}<K$ are finite constants that are non-zero only for the group that contains observation $i$. Also let $\mathcal{H}_0$ and Assumptions \ref{ass:regressors}-\ref{ass:Mhat} hold together with $\nu>5/2$ and $p^3/n =o(1)$. Then
	\begin{equation}
		\mathcal{S}\overset{d}{\rightarrow} N(0,1), \text{ as } n\rightarrow \infty.
	\end{equation}
\end{theorem}

\noindent Theorem \ref{theorem:nulldist} provides asymptotic justification for using one-sided, standard normal critical values as observed also by \cite{Hong1995}. The `linear process' representation of the errors is convenient to establish a general result for clustered error structure and has been used in more general settings \citep{Kelejian2007, Conley2023}. Our next theorem relates to the power properties of our test. First consider the global alternative
\begin{equation}
	\mathcal{H}_{1A}: \ \  \mu_i \neq 0, \ \ \ \text{for some} \ i=1,\ldots,q, \text{ and any } \beta,
\end{equation}
where $\mu_i$ denotes the $i$-th element of $\mu$. We introduce the unrestricted quantities
\begin{equation}\label{epsilonU_def}
	\epsilon_{Ui}(\mu, \beta)=  y_i - \mu^\prime \upsilon_{f,g,i}- \beta^\prime x_i,i=1,\ldots,n,  \ \ \ \text{and} \ \ \ \  \tilde{\Phi}_{U}= \tilde{\Phi}_{U}(\mu,  \beta)=n^{-1}{Z^\prime \tilde{\Sigma}_U Z},
\end{equation}
where $\upsilon_{f,g,i}'=\left(\psi_1(w_i'y),\ldots,\psi_p(w_i'y),\psi_1(w_i'c_1),\ldots,\psi_p(w_i'c_1),\ldots,\psi_1(w_i'c_l),\ldots,\psi_p(w_i'c_l)\right)'$,
$\tilde{\Sigma}_{U}=diag\left(\tilde{\Sigma}_{U1},\ldots,\tilde{\Sigma}_{UG}\right)$, and $\tilde{\Sigma}_{Ug}$ has $(i,j)$-th element $\epsilon_{Ui}(\mu, \beta)\epsilon_{Uj}(\mu, \beta)$. Then $\hat{\Phi}= \tilde{\Phi}_U(0_{q\times 1}, \hat{\beta})$. Let $\gamma = (\mu, \beta)\in\Gamma= \Re^q \times  \Re^k$ and introduce:
\begin{assumption}\label{ass:power} \textit{For all sufficiently large $n$ and all $j=1,\ldots,q+k$,}
	\begin{equation}\label{cond_omegaU_1}
		\underset{\gamma \in \Gamma} {\sup} \;\;\overline{\textit{eig}}(\tilde{\Phi}_U)  +   \underset{\gamma \in \Gamma} {\sup}\;\; \overline{\textit{eig}}\left(\frac{\partial\tilde{\Phi}_U}{\partial \gamma_j}\right) =O_p(1)  ,
	\end{equation}
	and
	\begin{equation}\label{cond_omegaU_2}
		\left\{ \underset{\gamma \in \Gamma} {\inf}\;\;\underline{\textit{eig}}(\tilde{\Phi}_U)\right\}^{-1}+ \left\{ \ \underset{\gamma \in \Gamma} {\inf}\;\;\underline{\textit{eig}}\left(\frac{\partial\tilde{\Phi}_U(\gamma)}{\partial\gamma_j}\right)\right\}^{-1}=O_p(1).
	\end{equation}
\end{assumption}




\noindent This assumption imposes mild regularity under $\mathcal{H}_{1A}$, reminiscent of boundedness and invertibility conditions.  Our power result follows below.



\begin{theorem}\label{theorem:consistency}
	Under Assumptions \ref{ass:errors}-\ref{ass:power}, $\mathcal{H}_{1A}$, $\nu>5/2$, and $p^3/n =o(1)$, $\mathcal{S}$ provides a consistent test.
\end{theorem}









\section{Practical guidance for the $s$ test}\label{sec:implementation}

In this section, we discuss practical implementation issues and explain how we construct the recommended $\boldsymbol{s}$ test decision rule introduced in Definition \ref{def:s-test}.
\paragraph{Choice of tuning parameters.}
 Our testing procedure is nonparametric and as such requires the tuning parameter $p$ to be chosen. Given our rate condition $p^3/n\rightarrow 0$, a reasonable empirical choice might be $p=\max\left\{2,\left[n^{1/3}\right]\right\}$, where $[\cdot]$ denotes the closest integer, see e.g. \cite{Gupta2023}, and noting that we always want $p>1$ to distinguish our setting from the linear case when using polynomial bases. However, note that the rate $p^3/n\rightarrow 0$ is equivalent asymptotically to $p^3(l+1)/n\rightarrow 0$ because $l$ is fixed, but in finite samples the extra $l+1$ factor can play a role. Thus our recommendation is to use
\begin{eqnarray}
\text{For the }\boldsymbol{cy}\text{ test, with both } \boldsymbol {c} \text { and } \boldsymbol {y} \text{ spillovers}&:&  p_{cy}=\max\left\{2,\left[\frac{n^{1/3}}{l+1}\right]\right\}, \label{pxy_reco}\\
\text{For the }\boldsymbol{c}\text{ test, with only } \boldsymbol {c} \text{ spillovers}&:& p_c=\max\left\{2,\left[\frac{n^{1/3}}{l}\right]\right\},\label{px_reco}\\
\text{For the }\boldsymbol{y}\text{ test, with only } \boldsymbol {y} \text{ spillovers}&:& p_y=\max\left\{2,\left[n^{1/3}\right]\right\}. \label{py_reco}
\end{eqnarray}
\noindent In Appendix \ref{sec:app_ext_mult} we also show that we can extend our test to a setting with multiple $W$ matrices, say $\ell$, for which we recommend $p_{cy,m}=\max\left\{2,\left[{n^{1/3}}/{\ell (l+1)}\right]\right\}$ in (\ref{pxym_reco}) therein for the $\boldsymbol{cy}$ test.

\paragraph{Rank deficiency issues in $Z$.} In finite samples, the instrument matrix $Z$ may be rank deficient. In such cases, we recommend identifying a maximal set of linearly independent columns of $Z$, yielding an $n\times r$ submatrix with full column rank, where $r=rank(Z)$. The test statistics can then be constructed using this reduced instrument matrix. Standard numerical routines are available for extracting a linearly independent subset of columns, with restriction of course that $m\geq q+k$.

\paragraph{Critical values.} While our asymptotic theory justifies the use of one-sided standard normal critical values for the $\boldsymbol{c}$, $\boldsymbol{y}$ and $\boldsymbol{cy}$ tests, it is reasonable to suspect that these might not be completely reliable for small $(n,p)$. A recommended robustness check is to also compare the test statistic $\mathcal{S}$ to $(\chi^2_{q,\alpha}-q)/\sqrt{2q}$, where $\chi^2_{q,\alpha}$ is the critical value at the $\alpha\%$ significance level for a $\chi^2_q$ distribution.



\paragraph{Discussion of the rejection rule in Definition \ref{def:s-test}.} Our recommended $\boldsymbol{s}$-test rejection rule in Definition \ref{def:s-test} warrants further discussion. Our methodology yields three candidate tests for spillovers: through individual covariates ($\boldsymbol{c}$), through outcomes ($\boldsymbol{y}$), and through both channels jointly ($\boldsymbol{cy}$). One may therefore ask why we do not simply recommend one of these tests, rather than the composite decision rule underlying the $\boldsymbol{s}$ test in Definition \ref{def:s-test}. The reason is that reliance on any single channel can be misleading. To see this, consider the following cases:
\begin{enumerate}
    \item \textit{At least one of the $\boldsymbol{c}$ or $\boldsymbol{y}$ tests reject the null of no spillovers, but the $\boldsymbol{cy}$ test does not.}
    This case can arise when one out of the covariate or outcome channels exhibits strong spillovers but the other does not, or does so to a much weaker degree. If the strength of the channel through which spillovers occur is sufficiently weak relative to the no-spillover channel, the combined $\boldsymbol{cy}$ test statistic can become too small and mask the spillover. Degrees of freedom issues can also play a role; indeed, the $\boldsymbol{cy}$ test has $q=p(l+1)$ while the other tests have $q=p$ or $q=pl$. This can distort the test procedure in finite samples. Our $\boldsymbol{s}$ test rejection rule avoids these problems because it will detect spillovers via the $\boldsymbol{c}$ or $\boldsymbol{y}$ tests. \vspace{0.2cm}
    \item \textit{Neither the $\boldsymbol{c}$ or $\boldsymbol{y}$ tests reject the null of no spillovers, but the $\boldsymbol{cy}$ test does.} This can be caused by a high degree of collinearity between the series approximations of $f(s)$ and the $g_j(s)$, namely the columns of the matrix $\Upsilon_{f,g}$. Our theory rules out perfect multicollinearity, nevertheless a high degree of collinearity can make the $\boldsymbol{c}$ and $\boldsymbol{y}$ test statistics small and therefore unreliable on their own as a decision tool. Our $\boldsymbol{s}$ test rejection rule avoids this problem of high collinearity masking the spillover because it will detect the spillover via the $\boldsymbol{cy}$ test.
\end{enumerate}

\allowdisplaybreaks
\section{Four demonstrations}\label{illustrations}

In this section, we illustrate the scope of our nonparametric test by revisiting four existing studies that investigate peer effects across different contexts—high-skill professionals, interracial college roommates, elementary-school deskmates, and academic researchers—with mixed findings. We use Hermite polynomials as basis functions in all four examples. Our test detects spillovers in many cases where linear specifications fail to do so, thereby overturning several, but not all, of the original findings of no spillovers. These applications highlight the repercussions of ignoring nonlinearities in the spillover mechanism. Additional details on the original studies are reported in Appendix \ref{app_C}.






\subsection{Professional golf tournaments \citep{GuryanKroftNotowidigdo}} \label{golf}\label{applications}




Our first example builds on \cite{GuryanKroftNotowidigdo}, who study whether peer effects influence individual productivity in high-skill professional environments in the context of professional golf tournaments. They exploit a natural experiment within golf tournaments where playing partners are randomly assigned within predefined block–round categories.

Their original results are reproduced in Appendix \ref{app_C}, Table \ref{tab:GKNaugmented}, and include three specifications.\footnote{The large sample size and relatively few covariates imply large values for $p$ using (\ref{pxy_reco})-(\ref{py_reco}). This can cause rank deficiency issues in $Z$, practical guidance for which was provided in Section \ref{sec:implementation}.}  In specification (i), players' performance is modelled as a function of their own ability, measured by the corrected handicap score, and the average ability of their peers in the same block–round–tournament.\footnote{
For details about the way the corrected handicap score is calculated, see Appendix \ref{app_C}.} Specification (ii)  incorporates alternative measures of peer ability—average driving distance, number of putts, and number of greens hit—designed to distinguish motivation effects (e.g., higher effort induced by stronger partners) from learning effects (e.g. adapting to observed putting strategies). Specification (iii) introduces heterogeneity by interacting partners’ average ability with a player’s own baseline ability and years of professional experience.

While previous studies have found significant positive peer effects in low-skill labour markets \citep{BandieraBarankayRasul,MasMoretti}, \cite{GuryanKroftNotowidigdo} find limited evidence of peer effects in individual performance. In particular, they conclude against peer effects in specifications (i) and (ii), while specification (iii) provides some support for heterogeneous peer effects via the experience channel.






\subsubsection*{Test results}




Table \ref{tab:GKNjointtest} summarizes the authors' main results and the results from our tests.
The three rows correspond, in order of appearance, to the three original specifications reported in Appendix \ref{app_C}, Table \ref{tab:GKNaugmented}.
Columns 1--2 report the specification number and the construction of the attribute peer exposure variables $w'c$, respectively.
The `Original result' column reports the conclusions reached in the original paper regarding the null hypothesis of no linear peer effects, while the `$\boldsymbol{s}$ test result' column reports the conclusion of our test.
The `Rejection channel' column indicates the channels through which we reject the null: peers' attributes ($\boldsymbol{c}$), peers' outcomes ($\boldsymbol{y}$), or both ($\boldsymbol{cy}$).
Columns `$n$' and `$l$' report the sample size and the number of peer attribute terms used in the test, respectively.
Finally, the column `$p_c$, $p_y$, and $p_{cy}$' reports the number of basis functions used to test each spillover channel.


\begin{table}[p]\centering
\caption{Peer Exposure Design and Test Results: \citet{GuryanKroftNotowidigdo}}
\label{tab:GKNjointtest}
\small
\resizebox{\textwidth}{!}{
\begin{tabular}{
c
p{3cm}
c
c
c
c
c
c
}
\toprule
\multicolumn{8}{l}{\textbf{Dependent variable:} Score in a given tournament-round, $y_{i,tr}$ } \\
\multicolumn{8}{l}{\textbf{Peer exposure in outcomes:} $w_{i,tr}'y_{tr}$} \\
\addlinespace
\midrule


Spec. & Peer exposure $w'\boldsymbol{c}$
& Original result
& $\boldsymbol{s}$ test result
& Rejection channel
& $n$
& $l$
& $p_c$, $p_{y}$, $p_{cy}$ \\
\midrule

(i)
& $w_{i,tr}'Ability$
& Do not reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$, $\boldsymbol{cy}$
& 17,492
& 1
& 26, 26, 13\\ [2 em]

(ii)
& $\begin{aligned}[t]
& w_{i,tr}'DrivDist, \\
& w_{i,tr}'Greens, \\
& w_{i,tr}'Putts \\
& \end{aligned}$
& Do not reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$, $\boldsymbol{cy}$
& 17,182
& 3
& 9, 26, 6 \\ [5em]

(iii)
& $\begin{aligned}[t]
& w_{i,tr}'Ability, \\
&  Ability_i \times w_{i,tr}' Ability,\\
& Exp_i \times w_{i,tr}' Ability\\
& \end{aligned}$
& Reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$, $\boldsymbol{cy}$
& 17,492
& 3
& 9, 9, 4 \\
\bottomrule
\end{tabular}
}

\footnotesize
\justifying
\textit{Note:}
Each row corresponds to a specification in Table \ref{tab:GKNaugmented} in Appendix \ref{app:gkn}.
Columns 1-2 report the specification number and the construction of attribute peer exposure variables. The `Original result' column reports the conclusions reached in the original paper for the corresponding null hypothesis of no linear peer effects. The $\boldsymbol{s}$ test examines dependence operating through three rejection channels: the $\boldsymbol{c}$ channel through peers’ attributes, the $\boldsymbol{y}$ channel through peers’ outcomes, and the $\boldsymbol{cy}$ channel through both peers’ attributes and outcomes. Column $n$ reports the sample size and column $l$ the number of peer attribute terms. The column `$p_c$, $p_y$, $p_{cy}$' reports the selected number of basis functions used in each test. The outcome variable is the golf score for the round. $Ability_i$ is measured by the player's average handicap, $DrivDist_i$ is the average driving distance, $Greens_i$ is the average number of greens hit in regulation, and $Putts_i$ is the average number of putts per round, all averaged over the previous 2-3 years. $Exp_i$ is measured as years of experience. Peer exposure is measured using a social vector $w_{i,tr}$, where each player is exposed to the weighted average of the attributes/outcomes of peers within the same group-tournament-round. Sample weights are given by the inverse of the sample variance of the estimated ability of each player, in line with the original study. All specification controls are identical to those reported in Table \ref{tab:GKNaugmented} in Appendix \ref{app:gkn}. Tests are performed at the 95\% confidence level. Standard errors are clustered at the playing group level.
\end{table}


Our findings suggest the presence of peer effects in professional golf that may not be fully captured by the linear specifications in \cite{GuryanKroftNotowidigdo}. Indeed, our $\boldsymbol{s}$ tests detect spillovers in all three specifications, pointing to nonlinear effects that the original specifications fail to detect in cases (i) and (ii).


\begin{comment}

\subsection{Network expansion and firm performance \citep{CaiSzeidl}}\label{firm}


\subsubsection{Context and findings}

Firms in developing economies face not only financial and managerial constraints but also networking frictions—such as limited trust or information—that may prevent them from accessing knowledge, clients, and suppliers.\footnote{Such networking frictions are likely to be more binding in developing economies, where search and trust costs impede self-organization, whereas at higher levels of development similar business associations can often emerge without external coordination \citep{CaiSzeidl}.}  In our second illustration we revisit the study of \cite{CaiSzeidl}, who investigate whether an exogenous expansion in business networks can improve firm performance.

The intervention under study randomly assigned firms into small business-association groups of ten owner-managers. The treatment group managers met monthly for one year, while other firms' managers did not participate and served as a control group. The intervention outcomes were measured through detailed baseline, midline, and endline surveys covering sales, profits, employment, assets, inputs, management practices, and networks.
The authors first examine the direct impact of the intervention and find that treated firms experienced significant and sustained improvements in performance.\footnote{Sales increased by about 8\% at midline and 10\% at endline for treated firms relative to control firms, with similarly positive impacts on profits, employment, fixed assets, and input usage. Firms expanded both the number of clients and suppliers they interacted with, increased access to formal and informal borrowing, and improved management scores by roughly a fifth of a standard deviation.} To help uncover the mechanisms behind these gains, the authors then focus on peer composition within the networking groups and run a battery of peer-effect specifications that we focus on, reported in Table \ref{tab:caiszeidl_regress} of Appendix \ref{app:caiszeidl}.
By proxying peer quality by baseline employment size, \cite{CaiSzeidl} show that firms randomly assigned to groups with larger peers achieved faster growth across multiple dimensions, including sales, profits, management practices, and network expansion.



\subsubsection{Test results}
Table~\ref{tab:CaiSzeidltest} compares our nonparametric test results to the findings by \cite{CaiSzeidl}.
Each row (i)–(xiv) reports results from a separate regression corresponding  to the specifications
reported in Table \ref{tab:caiszeidl_regress}.
 As before, for each of the outcomes we test for dependence operating through peers’ attributes alone ($\boldsymbol{c}$ test), peers’ outcomes alone ($\boldsymbol{y}$ test) and through both peer attributes and outcomes ($\boldsymbol{cy}$ test).
 The columns $p_c$, $p_y$ and $p_{cy}$ report the number of basis functions used in each test, and $n$ reports the sample size.


Table~\ref{tab:CaiSzeidltest} highlights several differences between our nonparametric results and the original findings. Nevertheless, the $\boldsymbol{s}$ test reaches the same qualitative conclusions as the authors for three key outcomes: it rejects the null for sales and utility costs and does not reject it for the number of suppliers.

However, we also identify several outcomes for which our $\boldsymbol{s}$ test rejects the null, whereas the authors find no significant peer effects. In some cases, the results suggest that peers' observable characteristics help shape peer effects, as for No. of employees, Total Assets, Material Cost, and Productivity. For other outcomes, namely Bank loan, Innovation score, Book sales, and Tax/Sales, the $\boldsymbol{s}$ test indicates that peer effects operate predominantly through the $\boldsymbol{y}$ or $\boldsymbol{cy}$ channels.

Interestingly, we also observe  two outcomes (Profits, No. of clients) where the linear approach detects spillovers, while our $\boldsymbol{s}$ test fails to do so. \textcolor{red}{[discussion needed]}

Overall, this illustration suggests that peer effects in this setting are transmitted not only through peers' observable characteristics, but also through outcome-based spillovers and the interaction of these two channels. More generally, it highlights that the various components of the $\boldsymbol{s}$ test may yield different rejection patterns, helping to identify the mechanisms underlying cross-unit dependence.










\begin{table}[p]\centering
\caption{Peer Exposure Design and Test Results: \citet{CaiSzeidl} }
\label{tab:CaiSzeidltest}
\small
\resizebox{\textwidth}{!}{
\begin{tabular}{
p{.7cm}
c
c
c
c
c
c
c
c
c
}
\toprule

\multicolumn{10}{l}{\textbf{Peer exposure in attributes:} $  w_{i,t}' EmpSize$} \\
\multicolumn{10}{l}{\textbf{Peer exposure in outcomes:} $  w_{i,t}' y_t$} \\
\multicolumn{10}{l}{ $  l=1$} \\
\addlinespace
\midrule
& & \multicolumn{3}{c}{Peer exposure in attributes} & \multicolumn{2}{c}{\makecell{Peer exposure \\ in outcomes}}  & \multicolumn{2}{c}{\makecell{Peer exposure \\ in both}}\\
\cmidrule(lr){3-5}\cmidrule(lr){6-7}\cmidrule(lr){8-9}

Spec.
& $y$
& Original result
& $p_c$
& $\boldsymbol{c}$ test
& $p_{y}$
& $\boldsymbol{y}$ test
& $p_{cy}$
& $\boldsymbol{cy}$  test
& $n$\\
\midrule

(i)
& Sales
& Reject
& 16 & Do not reject
& 16  & Reject
& 8 & Reject
& 4,183 \\ [.5em]

(ii)
& Profits
& Reject
& 16 & Do not reject
& 16 & Do not reject
& 8 & Do not reject
& 4,076 \\ [.5em]

(iii)
& No. of employees
& Do not reject
& 16 & Reject
& 16  & Reject
& 8 & Do not reject
& 4,183 \\ [.5em]

(iv)
& Total Assets
& Do not reject
& 16 & Reject
& 16  & Reject
& 8 & Reject
& 4,183 \\ [.5em]

(v)
& Material Cost
& Do not reject
& 16 & Reject
& 16  & Do not reject
& 8  & Do not reject
& 4,148 \\ [.5em]

(vi)
& Utility Cost
& Reject
& 16 & Do not reject
& 16  & Reject
& 8 & Do not reject
& 4,086 \\ [.5em]

(vii)
& Productivity
& Do not reject
& 16 & Reject
& 16 & Reject
& 8 & Do not reject
& 4,183 \\ [.5em]

(viii)
& No. of clients
& Reject
& 16 & Do not reject
& 16  & Do not reject
& 8 & Do not reject
& 4,173 \\ [.5em]

(ix)
& No. of Suppliers
& Do not reject
& 16 & Do not reject
& 16 & Do not reject
& 8 & Do not reject
& 4,170 \\ [.5em]

(x)
& Bank loan
& Do not reject
& 16 & Do not reject
& 16  & Reject
& 8  & Do not reject
& 4,183 \\ [.5em]

(xi)
& Management score
& Reject
& 14 & Do not reject
& 14 & Reject
& 7 & Reject
& 2,774 \\ [.5em]

(xii)
& Innovation score
& Do not reject
& 11 & Do not reject
& 11  & Do not reject
& 6 & Reject
& 1,409 \\ [.5em]

(xiii)
& Reported - book sales
& Do not reject
& 16 & Do not reject
& 16  & Do not reject
& 8  & Reject
& 4,152 \\ [.5em]

(xiv)
& Tax/Sales
& Do not reject
& 16 & Do not reject
& 16 & Do not reject
& 8  & Reject
& 4,178 \\

\bottomrule
\end{tabular}
}
\justifying
\footnotesize
\textit{Note:} Each row (i)–(xiv) reports results from a separate regression corresponding to specifications (1)–(14) in Table 7 of \citet{CaiSzeidl}. The attribute variable $EmpSize_{i}$ is measured as the log baseline number of employees for each firm. The first column reports the specification number, $y$ reports the dependent variable. The `Original result' column reports the conclusions reached in the original paper for the corresponding null hypothesis of no peer effects. The $\boldsymbol{c}$ test examines dependence operating through peers’ attributes, the $\boldsymbol{y}$ test examines dependence operating through peers’ outcomes, and the $\boldsymbol{cy}$ test implements the test with both peers’ attributes and outcomes. The columns $p_c$, $p_y$ and $p_{cy}$ report the selected number of basis functions used in each test, $n$ reports the sample size. Peer exposure is measured using a simple average social vector, $w_{i,t}$, where each player is exposed to the average attribute/outcome of peers within the same meeting group at time $t$. All specification controls are identical to those reported in Table \ref{tab:caiszeidl_regress} in Appendix \ref{app:caiszeidl}. Tests are performed at the 95\% confidence level. Standard errors are clustered at the meeting group level.
\end{table}







\subsection{Student achievement \citep{BooijLeuvenOsterbeek}}



In our second demonstration we revisit \cite{BooijLeuvenOsterbeek}, who study a large-scale field experiment conducted at the University of Amsterdam in the Netherlands. The intervention randomly assigned
 first-year economics students to tutorial groups. By exogenously varying group composition, this design created substantial variation
to estimate how students' performance responds to changes in the academic environment.\footnote{Students attended weekly tutorials with the same group throughout the course, and all teaching assistants followed a common syllabus, minimizing the chance that differences in outcomes could be attributed to instruction quality.}


The authors study how students' performance (measured by the number of collected credits)  is affected by prior achievement of the peers in the same tutorial group (measured by their secondary-school GPA).
Their main results are reproduced in Appendix \ref{app_C}, Table \ref{tab:booij_regress}.  From columns (i) to (v), academic performance is regressed on various summary statistics of peers’ GPA (e.g., the mean, dispersion, and their interaction), using a sequence of increasingly rich specifications that allow for heterogeneous responses across the ability distribution. As shown in Table \ref{tab:booij_regress}, these linear regressions find limited evidence that peer achievement or peer heterogeneity affects student outcomes, except for some specifications involving higher-order interactions.\footnote{As clarified in Appendix \ref{app_C} `higher-order'
 interactions refers to interaction terms involving multiple variables (e.g.\ peer mean $\times$ peer dispersion $\times$ own GPA).
 }


\subsubsection*{Test results}


In Table \ref{tab:Booijtest} we apply our nonparametric test to re-examine the peer effects documented in \cite{BooijLeuvenOsterbeek}. Each row of Table \ref{tab:Booijtest} corresponds, in order of appearance, to a specification in Table \ref{tab:booij_regress}.
Since peer exposure variables are specification-specific, we report them as $w'c$ and $w'y$ respectively.
The remaining columns follow the conventions of Table \ref{tab:GKNjointtest}. The comparison with the results by \cite{BooijLeuvenOsterbeek}  reveals some differences between our nonparametric test and the authors’ original findings.
Overall, our $\boldsymbol{s}$ test rejects the null hypothesis in all specifications. This is contrary to the authors'  conclusions in specifications (i) to (iii), and provides evidence for nonlinear interactions. On the other hand, the results for specifications (iv) and (v) are consistent with the authors’ conclusion.






\begin{table}[p]\centering
\caption{Peer Exposure Design and Test Results: \citet{BooijLeuvenOsterbeek}}
\label{tab:Booijtest}
\small
\resizebox{\textwidth}{!}{
\begin{tabular}{
p{.8cm}
p{3.2cm}
p{.5cm}
p{3.2cm}
c
c
c
c
c
}
\toprule
\multicolumn{9}{l}{\textbf{Dependent variable:} Number of credits achieved} \\
\multicolumn{9}{l}{${n =} 1{,}876$} \\
\addlinespace
\midrule


Spec. & Peer exposure $w'\boldsymbol{c}$
&
& Peer exposure $w'\boldsymbol{y}$
& Original result
& $\boldsymbol{s}$ test result
& Rejection channel
& $l$
& $p_c$, $p_{y}$, $p_{cy}$
\\
\midrule

(i)
& $w_{i,\text{avg}}'GPA$
&
& $w_{i,\text{avg}}'y$
& Do not reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$, $\boldsymbol{cy}$
& 1
& 12, 12, 6
\\ [1.5em]

(ii)
& $w_{i,\text{avg}}'GPA$
&
& $w_{i,\text{avg}}'y$
& Do not reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$, $\boldsymbol{cy}$
& 1
& 12, 12, 6
 \\
\midrule

&
&
&
&
&
&
& $l,\ell$
& $p_{c}$, $p_{y,m}, p_{cy,m}$
\\
\midrule

(iii)
& $\begin{aligned}[t]
& w_{i,\text{avg}}'GPA, \\
& w_{i,\text{sd}}'GPA
\end{aligned}$
&
& $\begin{aligned}[t]
& w_{i,\text{avg}}'y, \\
& w_{i,\text{sd}}'y
\end{aligned}$
& Do not reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$, $\boldsymbol{cy}$
& 2, 2
& 6, 6, 3
 \\ [2.5em]

(iv)
& $\begin{aligned}[t]
& w_{i,\text{avg}}'GPA, \\
& w_{i,\text{sd}}'GPA, \\
& w_{i,\text{avg}}'GPA \times w_{i,\text{sd}}'GPA
\end{aligned}$
&
& $\begin{aligned}[t]
& w_{i,\text{avg}}'y, \\
& w_{i,\text{sd}}'y, \\
& w_{i,\text{avg}}'y \times w_{i,\text{sd}}'y
\end{aligned}$
& Reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$, $\boldsymbol{cy}$
& 3, 3
& 4, 4, 2
 \\ [4.7em]

(v)
& $\begin{aligned}[t]
& w_{i,\text{avg}}'GPA, \\
& w_{i,\text{sd}}'GPA, \\
& w_{i,\text{avg}}'GPA \times w_{i,\text{sd}}'GPA, \\
& GPA_i \times w_{i,\text{avg}}'GPA, \\
& GPA_i \times w_{i,\text{sd}}'GPA, \\
& GPA_i \times w_{i,\text{avg}}'GPA \times w_{i,\text{sd}}'GPA
\end{aligned}$
&
& $\begin{aligned}[t]
& w_{i,\text{avg}}'y, \\
& w_{i,\text{sd}}'y
\end{aligned}$
& Reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$, $\boldsymbol{cy}$
& 6, 2
& 2, 2, 2
 \\
\bottomrule
\end{tabular}
}

\justifying
\footnotesize
\textit{Note:} Each row corresponds to specifications (i)-(v) in Table 4 of \cite{BooijLeuvenOsterbeek}. Columns 1-3 report the specification number, the construction of attribute peer exposure variables, and the construction of outcome peer exposure variables. The `Original result' column reports the conclusions reached in the original paper for the corresponding null hypothesis of no linear peer effects. The $\boldsymbol{s}$ test examines dependence operating through three rejection channels: the $\boldsymbol{c}$ channel through peers’ attributes, the $\boldsymbol{y}$ channel through peers’ outcomes, and the $\boldsymbol{cy}$ channel through both peers’ attributes and outcomes. Column
$l$ reports the number of peer attribute terms. The column `$p_c$, $p_y$, $p_{cy}$' reports the selected number of basis functions used in each test. The outcome variable, student performance, is measured by the number of credits obtained in the first year of university. The ability attribute, $GPA$, refers to students' pre-university GPA. In specifications (i) and (ii), peer attribute exposure is measured using a weighted-average vector, $w_{i,\text{avg}}$, where each student is exposed to the average GPA of peers within the same tutorial group. In specifications (iii) and (iv), peer exposure additionally incorporates dispersion in peer characteristics through a standard-deviation vector, $w_{i,{\text{sd}}}$, which captures variation in peers’ GPA within the group, with specification (iv) adding their interaction. Accordingly, the second-to-last column now reports $l,\ell$ and the last column the recommended $p_{cy,m}$ from (\ref{pxym_reco}). In specification (v), the peer exposure terms in (iv) are further reweighted by individual GPA, interacting peer exposure with the student’s own GPA to allow the strength of spillovers to vary with academic ability. All specification controls are identical to those reported in Table \ref{tab:booij_regress} in Appendix \ref{app:booij}. Tests are performed at the 95\% confidence level. Standard errors are clustered at the tutorial group level.   \snaptodo[margin block/.append style=blue]{	\sloppy\textbf{Marghe}: here we dont say that controls are as in the original table}
    \snaptodo[margin block/.append style=green!50!black]{
	{\sloppy\textbf{Abhi}: tbd}}

\end{table}







\end{comment}

\subsection{Interracial contact, stereotypes, and academic performance \citep{CornoLaFerraraBurns}}\label{roommates}



In our second demonstration we revisit \cite{CornoLaFerraraBurns}, who exploit a policy implemented by the University of Cape Town to study whether interracial interaction affects stereotypes, attitudes, and academic performance in post-apartheid South Africa. The policy randomly allocates first-year students to roommates, providing exogenous variation in whether a student shares a double room with a roommate of a different race.


We focus on the analysis of whether mixed-race interaction affects academic performance.  The authors estimate the effect of a mixed room intervention on four academic outcomes --- GPA, the number of exams passed, eligibility to continue to the second year, and a composite index of academic performance --- separately for the White subsample, the Black subsample, and the full sample. Their original results are reproduced in Appendix \ref{app_C}, Table \ref{tab:CLBregress}. They find that Black students assigned to mixed-race rooms experience significant improvements in academic performance, whereas the estimated effects for White students are close to zero and statistically insignificant. In the full sample, the effects are positive and significant for all outcomes except GPA.





\subsubsection*{Test results}

Table \ref{tab:CLBtest} reports our $\boldsymbol{s}$ test results. Rows are ordered by panels A (Whites), B (Blacks) and C (full sample), and, within each panel, by specification (i)--(iv), matching the layout of Table \ref{tab:CLBregress} in Appendix \ref{app:clb}. Our nonparametric test reveals additional spillover patterns that complement the authors' original findings.
The $\boldsymbol{s}$ test agrees with the authors in nine of the twelve cases, including two cases in Panel A where it preserves the original conclusion of `Do not reject'.  On the other hand, in three cases, the original linear regressions detect no significant peer effects, whereas our $\boldsymbol{s}$ test does.



\begin{table}[p]\centering
\caption{Peer Exposure Design and Test Results: \citet{CornoLaFerraraBurns}}
\label{tab:CLBtest}
\small
\resizebox{\textwidth}{!}{
\begin{tabular}{
p{.9cm}
p{3.4cm}
c
c
c
c
c
c
}
\toprule
\multicolumn{8}{l}{\textbf{Peer exposure in attributes:} $w_{i}'Race$} \\
\multicolumn{8}{l}{\textbf{Peer exposure in outcomes:} $w_{i,\text{race}}'y$} \\
\multicolumn{8}{l}{Panels: A = Whites, B = Blacks, C = Full sample} \\
\multicolumn{8}{l}{$l=1$} \\
\addlinespace
\midrule


Spec.
& $y$
& Panel
& Original result
& $\boldsymbol{s}$ test result
& Rejection channel
& $n$
& $p_c$, $p_{y}$, $p_{cy}$
\\
\midrule

(i)
& GPA
& A
& Do not reject
& Do not reject
& --
& 117
& 5, 5, 2
 \\ [.5em]

(ii)
& No. of exams passed
& A
& Do not reject
& Reject
& $\boldsymbol{y}$
& 117
& 5, 5, 2 \\ [.5em]

(iii)
& Eligible to continue
& A
& Do not reject
& Reject
& $\boldsymbol{c}$
& 117
& 5, 5, 2 \\ [.5em]

(iv)
& Academic perf. index
& A
& Do not reject
& Do not reject
& --
& 117
& 5, 5, 2 \\
\addlinespace

(i)
& GPA
& B
& Reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{cy}$
& 332
& 7, 7, 3 \\ [.5em]

(ii)
& No. of exams passed
& B
& Reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$, $\boldsymbol{cy}$
& 332
& 7, 7, 3
 \\ [.5em]

(iii)
& Eligible to continue
& B
& Reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{cy}$
& 332
& 7, 7, 3\\ [.5em]

(iv)
& Academic perf. index
& B
& Reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{cy}$
& 332
&  7, 7, 3 \\
\addlinespace

(i)
& GPA
& C
& Do not reject
& Reject
& $\boldsymbol{cy}$
& 499
& 8, 8, 4
 \\ [.5em]

(ii)
& No. of exams passed
& C
& Reject
& Reject
& $\boldsymbol{y}$
& 499
& 8, 8, 4 \\ [.5em]

(iii)
& Eligible to continue
& C
& Reject
& Reject
& $\boldsymbol{y}$
& 498
& 8, 8, 4 \\ [.5em]

(iv)
& Academic perf. index
& C
& Reject
& Reject
& $\boldsymbol{y}$, $\boldsymbol{cy}$
& 498
& 8, 8, 4\\

\bottomrule
\end{tabular}
}

\justifying
\footnotesize


\textit{Note:} Each row corresponds to one panel $\times$ specification cell of Table \ref{tab:CLBregress} in Appendix \ref{app:clb}. Panels A--C are the White, Black, and full samples, respectively; specifications (i)--(iv) use GPA, number of exams passed, eligibility to continue, and the academic performance index as the dependent variable. The `Original result' column reports the conclusions reached in the original paper for the corresponding null hypothesis of no linear peer effects. The $\boldsymbol{s}$ test examines dependence operating through three rejection channels: the $\boldsymbol{c}$ channel through peers’ attributes, the $\boldsymbol{y}$ channel through peers’ outcomes, and the $\boldsymbol{cy}$ channel through both peers’ attributes and outcomes. Column $n$ reports the sample size.
The column `$p_c$, $p_y$, $p_{cy}$' reports the selected number of basis functions used in each test. The attribute variable $Race$ is the race indicator of each student, so that $w_i'Race$ is the race of $i$'s baseline roommate (see Footnote \ref{footnote:w_corno}).
All specification controls are identical to those reported in Table \ref{tab:CLBregress}. Tests are performed at the 95\% confidence level. Standard errors are clustered at the room level.


\end{table}






\subsection{Ability mix and student performance \citep{wuzhangwang2023}}\label{wzw}



In our third demonstration we revisit \cite{wuzhangwang2023}, who study a randomized deskmate intervention in elementary schools in China. Exploiting the fixed-seat system, under which students sit beside the same deskmate throughout the semester, the authors randomly pair previously high- and low-achieving students and examine effects on academic performance and `big five' personality traits.

The design features two treatment arms. In mixed-seating (MS) classes, students above and below the median of a prior examination are randomly paired as deskmates for 20 weeks. In mixed-seating with reward (MSR) classes, the same pairing is combined with a tournament-type incentive for high-achieving students: they receive a monetary award and a certificate of commendation if their deskmate's score improvement ranks in the top 10\% among lower-track students in the class. Control classes receive random seat assignment. The authors find that MS alone does not raise test scores relative to control, whereas MSR raises low-achieving students' mathematics scores by about 0.24 standard deviations; high-achieving students are largely unaffected academically. The MSR intervention also increases extraversion and agreeableness for both tracks.

We focus on their deskmate-level peer-effect specifications, which regress endline outcomes on the deskmate's baseline achievement. These results are reproduced in Appendix \ref{app:wzw}, Table \ref{tab:peer_effects_mixed_seating}. Statistically significant linear effects of deskmate baseline performance are largely absent in MS classes and remain limited under MSR---appearing mainly for selected personality traits among lower-track students---with little evidence that deskmate baseline scores raise endline academic performance measured by z-scores.


\subsubsection*{Test results}

In Table \ref{tab:WZWtest} we apply our $\boldsymbol{s}$ test to the deskmate peer structure. Following the authors' setup we define peer attribute exposure as the deskmate's baseline measurement of the outcome of interest, $w_i'\text{y}^{base}$, and peer outcome exposure as the deskmate's endline outcome, $w_i'y^{end}$. Rows are ordered by panels A (MS/low), B (MS/high), C (MSR/low) and D (MSR/high) and, within each panel, by specification (i)--(vi), matching Table \ref{tab:peer_effects_mixed_seating} in Appendix \ref{app:wzw}. The remaining columns follow previous conventions. Our $\boldsymbol{s}$ test rejects the null in most cases where the original linear specifications do not. In Panel~A, the $\boldsymbol{s}$ test rejects for every specification, whereas the original regressions do not reject at the 5\% level. In Panel~B, we fail to reject only for academic z-score and Agreeableness, matching the authors on those two cases but rejecting for the remaining traits. In Panel~C, we do not reject for academic z-score and Neuroticism but reject for Agreeableness, as in the original study; we additionally reject for Extraversion, Openness, and Conscientiousness. In Panel~D, we reject throughout, while the original regressions do not reject at 5\%. Overall, the results capture nonlinear deskmate dependence for both reward and no reward settings.


\begin{table}[p]\centering
\caption{Peer Exposure Design and Test Results: \citet{wuzhangwang2023}}
\label{tab:WZWtest}
\footnotesize
\resizebox{\textwidth}{!}{
\begin{tabular}{
c c c p{2.2cm} c c c
c
c
c
p{2.2cm}
c
c
}
\toprule
\multicolumn{7}{l}{\textbf{Peer exposure in attributes:} $w_{i}'y^{base}$ }\\
\multicolumn{7}{l}{\textbf{Peer exposure in outcomes:} $w_{i}'y^{end}$} \\
\multicolumn{7}{l}{Panels: A = MS/low, B = MS/high, C = MSR/low, D = MSR/high} \\
\multicolumn{7}{l}{$l=1$,  \hspace{0.5cm} $p_c$, $p_y$, $p_{cy} =$ 7, 7, 3} \\
\addlinespace
\midrule

Spec.
& $y$
& Panel
& Original result
& $\boldsymbol{s}$ test result
& Rejection channel
& $n$
 \\
\midrule

(i)
& Z-score
& A
& Do not reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{cy}$
& 317
\\

(ii)
& Extraversion
& A
&  Do not reject
& Reject
& $\boldsymbol{c}$
& 317
 \\

(iii)
& Agreeableness
& A
&  Do not reject
& Reject
& $\boldsymbol{cy}$
& 317
  \\

(iv)
& Openness
& A
& Do not reject
& Reject
& $\boldsymbol{c}$
& 317
  \\

(v)
& Neuroticism
& A
&  Do not reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$
& 317
  \\

(vi)
& Conscient.
& A
&  Do not reject
& Reject
&  $\boldsymbol{c}$
& 317
  \\
\addlinespace

(i)
& $\text{z-score}_{end}$
& B
& Do not reject
& Do not reject
& --
& 317
 \\

(ii)
& Extraversion
& B
& Do not reject
& Reject
& $\boldsymbol{c}$
& 317
  \\

(iii)
& Agreeableness
& B
&  Do not reject
&  Do not reject
& --
& 317
  \\

(iv)
& Openness
& B
&  Do not reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$
& 317
  \\

(v)
& Neuroticism
& B
&  Do not reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$, $\boldsymbol{cy}$
& 317
  \\

(vi)
& Conscient.
& B
&  Do not reject
& Reject
&  $\boldsymbol{c}$
& 317
  \\
\addlinespace

(i)
& $\text{z-score}_{end}$
& C
& Do not reject
& Do not reject
& --
& 297
 \\

(ii)
& Extraversion
& C
& Do not reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$, $\boldsymbol{cy}$
& 297
  \\

(iii)
& Agreeableness
& C
& Reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{cy}$
& 297
  \\

(iv)
& Openness
& C
& Do not reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{cy}$
& 297
  \\

(v)
& Neuroticism
& C
&  Do not reject
&  Do not reject
& --
& 297
  \\

(vi)
& Conscient.
& C
&  Do not reject
& Reject
&  $\boldsymbol{y}$, $\boldsymbol{cy}$
& 297
  \\
\addlinespace

(i)
& $\text{z-score}_{end}$
& D
& Do not reject
& Reject
& $\boldsymbol{c}$
& 297
  \\

(ii)
& Extraversion
& D
&  Do not reject
&  Reject
& $\boldsymbol{y}$, $\boldsymbol{cy}$
& 297
  \\

(iii)
& Agreeableness
& D
& Do not reject
& Reject
& $\boldsymbol{cy}$
& 297
  \\

(iv)
& Openness
& D
&  Do not reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$
& 297
  \\

(v)
& Neuroticism
& D
&  Do not reject
& Reject
&  $\boldsymbol{y}$
& 297
  \\

(vi)
& Conscient.
& D
&  Do not reject
& Reject
&  $\boldsymbol{y}$
& 297
  \\

\bottomrule
\end{tabular}
}

\justifying
\footnotesize
\textit{Note:} Each row corresponds to one panel $\times$ specification cell of Table \ref{tab:peer_effects_mixed_seating} in Appendix \ref{app:wzw}. Panels A--D are composed of lower- or upper-track students in the MS and MSR classes; specifications (i)--(vi) use academic z-score and the five personality traits at endline as the outcome of interest respectively. The `Original result' column reports the conclusions reached in the original paper for the corresponding null hypothesis of no linear peer effects. The $\boldsymbol{s}$ test examines dependence operating through three rejection channels: the $\boldsymbol{c}$ channel through peers’ attributes, the $\boldsymbol{y}$ channel through peers’ outcomes, and the $\boldsymbol{cy}$ channel through both peers’ attributes and outcomes. Column $n$ reports the sample size.
Tests are performed at the 95\% confidence level. Standard errors are clustered at the class level.
\end{table}



\subsection{Peer effects in academic research \citep{bosquetcombesetal2022}}\label{bch}



Our fourth demonstration focuses on \cite{bosquetcombesetal2022}, who study peer effects in academic research among economists in French universities. They exploit a national contest (\textit{concours d'agr\'egation}) that centrally allocates successful candidates across universities, generating quasi-exogenous variation in colleagues' field-specific productivity. Identification is sharpened by focusing on arrivals among the lowest-ranked successful candidates, whose remaining choice sets are tightly constrained, so that the field specialization of the arriving professor can be treated as approximately as good as random for the receiving university.

The authors find no peer effects when peers are defined as the entire university. Restricting attention to colleagues in the same JEL field
 and year, however, they find that one additional publication by field peers raises own productivity by a substantial range of about  0.6 publications on average. Accounting for heterogeneity in peer interactions, they further show that spillovers are weaker for women and older researchers. Their main results are reproduced in Appendix \ref{app_C}, Table \ref{tab:peer_effects_ols}.


\subsubsection*{Test results}

Table \ref{tab:BCHtest} reports our nonparametric test for the specifications in Table \ref{tab:peer_effects_ols}.
The attribute terms in $w'\boldsymbol{c}$ are the number of peers and the average peer outcome in a given university and year (restricted to the same JEL code from specification~(iii) onwards). The richer specifications~(iv)--(vi) include their interactions  with gender and/or age.

Our $\boldsymbol{s}$ test rejects the null in every specification. This aligns with the authors for the specifications with field-level peer effects (iii)--(vi),  but differs for  specifications (i)--(ii), where the original regressions do not reject.


\begin{table}[p]\centering
\caption{Peer Exposure Design and Test Results: \cite{bosquetcombesetal2022} }
\label{tab:BCHtest}
\small
\resizebox{\textwidth}{!}{
\begin{tabular}{
c
p{5.2cm}
c
c
p{1.6cm}
c
c
c
}
\toprule
\multicolumn{8}{l}{\textbf{Dependent variable:} Individual research output } \\


\addlinespace
\midrule
Spec.
& Peer exposure $w'\boldsymbol{c}$
& Original result
& $\boldsymbol{s}$ test result
& Rejection channel
& $n$
& $l$
& $p_c$, $p_y$, $p_{cy}$
\\

\midrule
(i)
& $\begin{aligned}[t]
& w_{i,ut}'Peer, \\
& w_{i,ut}'Output
\end{aligned}$
& Do not reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$
& 42{,}861
& 2
& 17, 17, 11 \\ [3em]

(ii)
& $\begin{aligned}[t]
& w_{i,ut}'Peer, \\
& w_{i,ut}'Output
\end{aligned}$
& Do not reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$,  $\boldsymbol{cy}$
& 771{,}498
& 2
& 46, 46, 31 \\ [3em]

(iii)
& $\begin{aligned}[t]
& w_{i,ut}'Peer, \\
& w_{i,uft}'Output
\end{aligned}$
& Reject
& Reject
& $\boldsymbol{cy}$
& 771{,}498
& 2
& 31, 31, 23 \\ [3em]

(iv)
& $\begin{aligned}[t]
& w_{i,ut}'Peer, \\
& w_{i,uft}'Output \\
& Woman_i \times w_{i,uft}'Output
\end{aligned}$
& Reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$
& 771{,}498
& 3
& 31, 31, 23 \\ [4.5em]

(v)
& $\begin{aligned}[t]
& w_{i,ut}'Peer, \\
& w_{i,uft}'Output \\
& Age_i \times w_{i,uft}'Output
\end{aligned}$
& Reject
& Reject
& $\boldsymbol{c}$
& 771{,}498
& 3
& 31, 31, 23 \\ [4.5em]

(vi)
& $\begin{aligned}[t]
& w_{i,ut}'Peer, \\
& w_{i,uft}'Output \\
& Woman_i \times w_{i,uft}'Output, \\
& Age_i \times w_{i,uft}'Output
\end{aligned}$
& Reject
& Reject
& $\boldsymbol{c}$, $\boldsymbol{y}$
& 771{,}498
& 4
& 23, 23, 19 \\

\bottomrule
\end{tabular}
}
\justifying
\footnotesize
\textit{Note:} Each row corresponds to a specification in Table \ref{tab:peer_effects_ols} in Appendix \ref{app_C}. In Specification (i) the individual output is aggregated over all JEL
codes
($n=42{,}861$), while in
specifications (ii)--(vi) individual output is defined at the JEL level ($n=771{,}498$). Columns 1--2 report the specification number and the construction of attribute peer exposure variables. The `Original result' column reports the conclusions reached in the original paper for the corresponding null hypothesis of no linear peer effects. The $\boldsymbol{s}$ test examines dependence operating through three rejection channels: the $\boldsymbol{c}$ channel through peers' attributes, the $\boldsymbol{y}$ channel through peers' outcomes, and the $\boldsymbol{cy}$ channel through both peers' attributes and outcomes. Column $n$ reports the sample size and column $l$ the number of peer attribute terms. The column `$p_c$, $p_y$, $p_{cy}$' reports the selected number of basis functions used in each test. Individual research output $y$ is the three-year moving average of articles divided by the number of co-authors, aggregated over fields in specification~(i) and measured at the JEL-code level thereafter. Peer output $w_{i,ut}'Output$ is the average of colleagues present in the department at date~$t$, where each colleague's productivity is their career-average publications per year;
 from specification~(iii), this average is restricted to the same JEL code, $w_{i,uft}'Output$.
Tests are performed at the 95\% confidence level. Standard errors are clustered at the department level.
\end{table}


\section{Conclusion}\label{Conclusions}
Cross-unit dependence has long been recognized as a central concern in  economics, reflecting both fundamental identification problems and first-order implications for econometric inference \citep{manski1993reflection,Conley1999}.
This paper proposes a novel nonparametric test for spillovers operating through peers’ attributes and/or outcomes, and provides a full asymptotic theory for it. The test has several appealing features. First, it can pick up nonlinear interactions of unknown form. Second, it only requires estimation under the null hypothesis of no spillovers, thereby avoiding nonparametric estimation altogether. Third, it is versatile, accommodating a wide range of data structures, including settings in which the interaction structure is incomplete or measured with error.

Our approach complements existing methods by offering a simple diagnostic to assess whether cross-unit dependence is present and whether linear approximations are likely to be informative. We illustrate its usefulness through four empirical applications, which suggest that the test can uncover forms of cross-unit dependence that are missed by standard specifications.
More broadly, our results reinforce the idea that the form and extent of spillovers should be informed by empirical evidence whenever possible.

\newpage

\bibliographystyle{chicago}
\bibliography{references}